Fabric defect detection method based on image reconstruction
By using a U-shaped self-coding model with attention mechanism to perform noise addition and denoising operations on defect-free fabric images, the problem of difficult to deal with complex fabric defect detection tasks in the prior art is solved, and efficient and robust fabric defect detection is achieved.
Patent Information
- Application Number
- CN202510259366.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-13
Smart Images

Figure CN120147282A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of defect detection and artificial intelligence, and provides a method for defect detection by only using defect-free positive samples for training and by repairing and comparing defect regions of a fabric. Background Art
[0002] Defects such as stains and shrinkage on the fabric surface will seriously affect the appearance, yield and value of the fabric, while structural defects such as broken threads and holes will also affect the strength and durability of the fabric, and may even pose a safety hazard in extreme cases. Therefore, defect detection is an important part of fabric production. It is necessary for manufacturers to detect the fabrics on the production line through effective defect detection means, both to improve the overall quality of the fabric and to prevent defective fabrics from flowing into the market from the production line. The most traditional fabric defect detection method mainly relies on human eye observation, and requires inspectors to directly find defects on the fabric with the naked eye. This method is inefficient and requires high labor costs. In addition, when the inspector is fatigued, the possibility of defect omission will increase. This makes the method of human eye observation gradually replaced by automated detection methods based on computer vision.
[0003] Traditional computer vision-based defect detection methods usually use algorithms designed by humans to extract feature information such as texture features, edge features, and frequency features in the fabric image, and distinguish defect regions based on these features. Such methods can achieve a high detection accuracy rate under the conditions of a small number of fabric colors, relatively fixed defect appearance features, and a relatively stable detection environment. However, in the face of complex detection tasks with a large number of color types, complex textures, strong uncertainty in defect appearance, and unstable detection environments, the robustness and generality of these detection methods cannot be guaranteed.
[0004] With the birth and development of artificial intelligence technology, a large number of defect detection methods based on artificial intelligence algorithms have been continuously proposed, overcoming the deficiencies of traditional methods. These artificial intelligence-based detection methods usually take a photo of the fabric as input, and output whether the fabric has defects and the specific positions of the defects. In the case where a large number of defect images can be obtained, a supervised training artificial intelligence model can be used to design a defect detection algorithm, allowing the model to directly contact defect images and learn defect features during the training stage, so as to directly locate defects through features during the detection stage. However, usually, the probability of defective fabrics produced on the production line is very low, it is difficult to collect a sufficient amount of defect images, and the collected defect images cannot guarantee that they contain all types of defect features. A supervised artificial intelligence model is not suitable for designing a defect detection algorithm. Therefore, designing a fabric defect detection method that does not require the use of defect-related data during the training stage, that is, using an unsupervised method for training, will effectively solve the problems of difficult acquisition of defect images and difficult training of supervised detection methods. Summary of the Invention
[0005] The object of the present invention is to provide a fabric defect detection method based on image reconstruction. This method uses a U-shaped autoencoder model with an attention mechanism module as the image reconstruction model, and uses defect-free normal fabric images for model training. In the defect detection stage, the noisy image to be tested is denoised multiple times to repair the defective areas in the image to be tested. The differences between the images before and after reconstruction are compared to determine the shape, position, and size of the defective areas. To address the problem of the long time consumption of iterative denoising operations, a method of resetting the step size is used to shorten the reconstruction time.
[0006] The specific technical solution for achieving the object of the present invention is as follows:
[0007] A fabric defect detection method based on image reconstruction, comprising the following steps:
[0008] Step a, collect images and organize the data set
[0009] Collect normal fabric images and make a training data set to enable the model to learn how to gradually remove the noise in the image, so as to reconstruct the noisy image into a noise-free image;
[0010] Step b, create an autoencoder model
[0011] Create a U-shaped autoencoder model with a self-attention mechanism. This model consists of an encoder, an intermediate layer, and a decoder. The encoder contains several code blocks composed of a convolutional layer, an activation function, a spatial self-attention module, a channel self-attention module, and a downsampling layer; the decoder contains several code blocks composed of a convolutional layer, an activation function, a spatial self-attention module, a channel self-attention module, and an upsampling layer; the intermediate layer contains two code blocks composed of a convolutional layer, an activation function, and a spatial self-attention module;
[0012] Step c, train the model
[0013] The training dataset constructed in step a is used to train the auto-encoding model created in step b. During training, there are two processes: a diffusion process of gradually adding noise and an inverse diffusion process of gradually removing noise, enabling the model to learn how to reconstruct images through denoising operations. Both the noise addition and denoising operations are carried out in units of time steps. After passing through t time steps in the diffusion or inverse diffusion process, it means that t noise addition or denoising operations have been performed. The total number of time steps for both the diffusion process and the inverse diffusion process is set to T. Each time the model parameters are adjusted, a batch of training images and a batch of random time step values are used. The images and time steps in the batch correspond one by one. Each training image is subjected to the corresponding number of noise addition steps according to the corresponding time step and then input into the model. Calculate the result of each image predicted by the model after one denoising operation, as well as the true result of each image after one denoising operation. Calculate the loss function between the output of the model and the true denoising result. After calculating the images and time steps in the batch, perform gradient descent operations to adjust the model parameters according to the value of the loss function. After each batch in the training dataset has been calculated by the model, return to the beginning of the training dataset to fetch a new batch and continue training. Stop training when the loss function converges or reaches the upper limit of the preset number of batches before training;
[0014] Step d, add noise to the image to be measured and reconstruct the image
[0015] Specify the number of noise addition steps T N , T N Less than the number of time steps T in the inverse diffusion process, perform T N steps of noise addition operations on the image to be measured and then pause, preparing to use the model obtained in step c for denoising operations to reconstruct the image to be measured. Before starting denoising, use the reset step size method to map the original T N steps of operations in the denoising process into a shorter denoising process with fewer steps to improve the reconstruction speed. After mapping, the model formally performs step-by-step denoising on the image to be measured according to the time step sequence of the shorter denoising process. During the denoising process, the model eliminates both the defective area and the noise, completing the reconstruction of the image to be measured;
[0016] Step e, process and compare the images to be measured before and after reconstruction, and locate the defective area
[0017] After reconstruction, perform Gaussian blur on the images to be measured before and after reconstruction, and perform pixel value standardization and normalization. Calculate the pixel-level structural similarity loss function in the images to be measured before and after reconstruction in a sliding window manner. The area composed of pixels whose loss function value exceeds the given threshold is the defective area.
[0018] Furthermore, step b specifically includes:
[0019] b-1: Create the model infrastructure. The model consists of an encoder, a middle layer, and a decoder. Both the encoder and the decoder contain N groups of code blocks. The last layer of the first N-1 code blocks in the encoder is a downsampling layer, which is used to halve the side length of the feature map. The last layer of the first N-1 code blocks in the decoder is an upsampling layer, which is used to double the side length of the feature map. The input image of the model is set as a square RGB image with side length w.
[0020] b-2: Create each code block in the encoder. Each code block contains multiple stacked convolutional blocks, and each convolutional block consists of a convolutional layer and an activation function. The input feature map resolution of the first code block is the same as the input image resolution of the model. A downsampling layer is added to the last of the first N-1 code blocks. Starting from the second code block, the input feature map resolution of each code block is equal to the output feature map resolution of the previous code block. When the side length of the input feature map of the code block satisfies a specified multiple relationship with the side length of the model input image, each convolutional block in the code block contains a parallel spatial self-attention module and channel self-attention module, which are added after the activation function and used to extract semantic information from the feature map.
[0021] b-3: Create two code blocks in the middle layer. Both code blocks contain a convolutional block, and each convolutional block consists of a convolutional layer and an activation function. The convolutional block in the first code block contains a spatial self-attention module, which is added after the activation function. The second convolutional block does not contain a self-attention module, and the resolutions of the input feature map and the output feature map are the same as those of the output feature map of the last code block in the encoder.
[0022] b-4: Create each code block in the decoder. Each code block contains multiple stacked convolutional blocks, and each convolutional block consists of a convolutional layer and an activation function. The input feature map resolution of the first code block is equal to the output feature map resolution of the middle layer. An upsampling layer is added to the last of the first N-1 code blocks. Starting from the second code block, the input feature map resolution of each code block is equal to the output feature map resolution of the previous code block. When the side length of the input feature map of the code block satisfies a specified multiple relationship with the side length of the model input image, each convolutional block in the code block contains a parallel spatial self-attention module and channel self-attention module, which are added after the activation function and used to extract semantic information from the feature map.
[0023] b-5: Create skip connections between the corresponding code blocks in the encoder and the decoder. For an integer i in the range [1, N], concatenate the output feature map of the i-th code block in the encoder and the input feature map of the (N + 1 - i)-th code block in the decoder in the channel dimension, which is used to fuse features at different levels and scales in the model.
[0024] b-6: Load the model into the video memory.
[0025] Further, step c specifically includes:
[0026] c-1: Specify an integer T greater than or equal to 100 as the number of time steps for both the diffusion process and the reverse diffusion process. The diffusion process is the process of gradually adding noise, and the reverse diffusion process is the process of gradually removing noise.
[0027] c-2: Set the noise weight β used when adding noise for each time step t t , and obtain the weight noise sequence:
[0028] {β 1 , β 2 , …, β T |0 < β 1 < β 2 < … < β T < 1}
[0029] where the sequence {β 1 , β 2 , …, β T} is a floating-point number sequence uniformly distributed starting from β 1 and ending with β T . After obtaining the noise weight sequence, calculate the image weight α used when adding noise for each time step t t , and obtain the image weight sequence:
[0030]
[0031] After obtaining the noise weight sequence {β 1 , β 2 , …, β T} and the image weight sequence {α 1 , α 2 , …, α T}, a single noise addition operation is represented by the following formula, that is, adding noise to the image that has undergone t - 1 noise addition operations once again to obtain the image with noise added t times:
[0032] x t = α t x t-1 + β t ε t
[0033] where ε t is a mixed noise signal obtained by weighted summation of a Gaussian noise signal and a simplex noise signal:
[0034] ε t = ρg t + γs t
[0035] where gt is a Gaussian noise signal and g t ~N(0, 1), s t is a simplex noise signal, ρ and γ are the weights of the noise signals. The Gaussian noise signal is used to let the model learn how to eliminate noise points, and the simplex noise signal is used to let the model learn how to eliminate large irregular color patches. To simplify the noise addition process, calculate the image merging weights for each time step to obtain the image merging weight sequence:
[0036]
[0037] and calculate the noise merging weights for each time step to obtain the noise merging weight sequence:
[0038]
[0039] and ensure to establish a single-step noise addition formula according to the merging weights:
[0040]
[0041] where x 0 represents the training image that has gone through all the preprocessing steps but has not been added with noise, and x t represents x 0 after t times of noise addition operations. The role of the single-step noise addition formula is to complete t times of noise addition operations in one step during the training and defect detection processes, reducing the processing time;
[0042] c-3: Take a batch containing M training images from the training dataset, and randomly generate a time step batch containing M time steps. Any time step value t in the time step batch satisfies that t is a positive integer and t ∈ [1, T]. Perform preprocessing on each image in the batch, including size adjustment operations and pixel value remapping operations. During the size adjustment operation, the image will be scaled and padded to become a square image with side length w. After completion, remap the pixel values using the following formula:
[0043]
[0044] where x represents the image after size adjustment, and x 0 represents the result of the image x after all preprocessing;
[0045] c-4: Generate a noise-added image x 0 corresponding to the time step t for each training image x t in the batch:
[0046]
[0047] and the image x with noise added t - 1 times t-1 :
[0048]
[0049] Take x t as the input data of the model, and x t-1 as the prediction target of the model;
[0050] c - 5: The auto - encoder model created in step b takes the time step t and the noisy image x t as inputs, predicts a denoising signal, and after weighted summation with x t can eliminate part of the noise in x t The model predicts the denoising signal and weighted summation with x t is a denoising operation. The image obtained after one denoising operation is represented by the following formula:
[0051]
[0052] where ∈ θ represents the auto - encoder model created in step b, and θ represents the learnable parameters in the model;
[0053] c - 6: Calculate the loss function L of the model. Here, the mean - square error loss between the image at time step t - 1 predicted by the model and the real image x at time step t - 1 t-1 is used as the loss function L, which is represented by the following formula:
[0054]
[0055] where w is the side length of the input and output images, and x t-1 (m, q) and respectively represent the color values of the pixel with abscissa m and ordinate q in x t-1 and . After obtaining the value of the loss function L, calculate the gradient of the model parameters and save it for subsequent model parameter adjustment;
[0056] c-7: For each image in the batch, perform the operations in steps c-4 to c-6. After all the images in the batch have been processed, perform gradient descent on the model parameters to adjust and optimize the model parameters. Then, take a new batch from the training dataset, generate the corresponding time-step batch, and perform the operations in steps c-4 to c-6 on the images and time steps in the new batch, enabling the model to learn how to gradually remove noise. After all batches in the training dataset have been computed by the model, return to the beginning of the training dataset, take a new batch and continue training, and stop training when the loss function converges or reaches the upper limit of the number of batches preset before training.
[0057] Further, the said step d specifically includes:
[0058] d-1: Specify the number of noise-adding steps T less than the number of time steps T of the diffusion process N ;
[0059] d-2: Use the reset step-size method for the denoising process, remap the originally used time-step sequence in the denoising process to a time-step sequence with a length less than T N so that the original denoising process with T N operations becomes a short denoising process with fewer than T N operations. Specify an integer S greater than or equal to 1 and less than T N as the number of denoising operations in the mapped short denoising process, and calculate the reference step size k:
[0060]
[0061] and the number of compensation time steps c:
[0062] c = mod(T N , S)
[0063] and calculate each time step τ i of the short denoising process to obtain the time-step sequence of the short denoising process:
[0064]
[0065] sequence in each time step τ i is calculated using the following formula:
[0066]
[0067] After resetting the step size, the original subsequence {1, 2,..., T N -1, T N} is mapped to the new subsequence {τ 1 , τ 2 ,..., τ S-1 , τ S}, TN The subsequent time steps are reserved for readjusting the image weights and the noise weights;
[0068] d-3: Calculate the noise weight sequence for the short denoising process. Let the auxiliary weight Traverse the time step sequence {1, 2, …, T} in ascending order. If the currently traversed time step t belongs to then calculate the noise weight β′ of time step t in the short denoising process t :
[0069]
[0070] Add β′ t to the noise weight sequence of the short denoising process, and set to be Continue to traverse the sequence {1, 2, …, T}. After the traversal is completed, obtain the noise weight sequence of the short denoising process and calculate the image weights i for each time step τ in the short denoising process to obtain the image weight sequence of the short denoising process
[0071]
[0072] d-4: Use the single-step noise addition formula to add noise to the image x to be measured 0 to obtain the image to be measured after adding noise T N times
[0073]
[0074] After readjusting the image weights and the noise weights corresponding to each time step, it is equivalent to adding noise in one step using time step τ S in the short denoising process to obtain x S ;
[0075] d-5: Let i = S, and use the model ∈ θ to repeatedly perform the denoising operation represented by the following formula:
[0076]
[0077] where τ i is the time step in the subsequence {τ 1 , τ 2 , …, τ S-1 , τ S}. Each time after the denoising operation, the value of i is decreased by 1 until the reconstructed image is obtained and then stop.
[0078] Further, step e specifically includes:
[0079] e-1: Perform Gaussian blur processing on the original image x of the image to be measured 0 and the reconstructed image to reduce the interference caused by the subtle differences in texture before and after reconstruction, and perform pixel value standardization and normalization;
[0080] e-2: Calculate the pixel-level structural similarity loss function L 0 between the original image x of the image to be measured and the reconstructed image in a sliding window manner. The image to be measured and the model input image have the same size, both being square images with side length w. Before calculation, specify the side length of the sliding window as 2b + 1, where b is a positive integer. Pad both the image to be measured and the reconstructed image with a width of b pixels on all four sides. For each pixel within the range of the original image, let the abscissa of the pixel in the padded image be m and the ordinate be q. The value ranges of both n and q are [b + 1, b + w]. Take a sliding window with side length 2b + 1 centered at the pixel (m, q) from the image to be measured and the reconstructed image respectively, and use the following formula to calculate the pixel-level structural similarity loss function at the pixel position with coordinates (m, q) for the two images: SSIM
[0081]
[0082] where I is the sliding window centered at the pixel (m, q) in the original image x of the image to be measured 0 and J is the sliding window centered at the pixel (m, q) in the reconstructed image , μ I and μ J are the mean pixel values of I and j respectively, σ I and σ J are the variances of the pixel values of I and j respectively, σ IJ is the covariance of the pixel values of I and j, C 1 and C 2 are constants. After calculating for all pixels (m, q), obtain a pixel-level structural similarity loss map with side length w;
[0083] e-3: Set a fixed threshold, find the pixels in the pixel-level structural similarity loss map whose loss function values are greater than the threshold, and determine these pixels as defect regions. Mark the corresponding pixels in white on a black background image of the same size as the image to be measured;
[0084] e-4: After checking all pixels, obtain a defect region mask image with a black background and white defect regions.
[0085] Beneficial effects
[0086] The method proposed by the present invention can train an image reconstruction model without using defect images. During detection, noise is first added to the image to be detected, and then the image reconstruction model is used to denoise. The noise and defect areas in the image to be detected are repaired and reconstructed into the texture of a normal fabric together, and then the image differences before and after reconstruction are compared to obtain the defect areas. The mixed noise of Gaussian noise and simplex noise used in the training and detection processes can add dense noise points and irregularly shaped color blocks to the image at the same time, enabling the model to learn how to repair defects with large color differences from the fabric without using defect images during the training process and without manually adding defect-related information. The method of resetting the step size during defect detection can significantly reduce the number of denoising operations of the model and improve the detection speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Figure 1 is a flowchart of the method of the present invention;
[0088] Figure 2 is a structural diagram of the U-shaped autoencoder model in an embodiment of the present invention;
[0089] Figure 3 is a schematic diagram of the convolutional block structure with a spatial self-attention module and a channel self-attention module; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0090] To make the advantages of the technical solution proposed by the present invention clearer, the method proposed by the present invention will be further described below in conjunction with the drawings and embodiments. The described embodiments are partial embodiments of the present invention, not all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments described in the present invention without creative efforts fall within the scope of protection of the present invention.
[0091] Embodiment
[0092] Refer to Figure 1 , a fabric defect detection method based on image reconstruction in this embodiment specifically includes:
[0093] Step a, collect images and organize the data set, including the following steps:
[0094] Collect normal fabric images without defects and make a training data set for the model to learn how to reconstruct the image with added noise;
[0095] Step b, refer to Figure 2 , Figure 3 , create an autoencoder model, including the following steps:
[0096] b-1: Create the model infrastructure, which consists of an encoder, a decoder, and an intermediate layer to form a U-shaped autoencoder model with self-attention mechanism. Both the encoder and the decoder contain 6 groups of code blocks. The last layer of the first 5 code blocks in the encoder is a downsampling layer, and the last layer of the first 5 code blocks in the decoder is an upsampling layer. The input image of the model is set as a square RGB image with side length w = 256.
[0097] b-2: Create each code block in the encoder. Each code block contains multiple stacked convolutional blocks. Each convolutional block consists of a convolutional layer and an activation function. The convolutional layer in the convolutional block is a three-dimensional convolutional layer with a channel dimension, and the activation function is the SiLU function. The input resolution of the first code block is the input image resolution of the model. The number of convolutional blocks in the first and second code blocks is 1 each, the number of convolutional blocks in the third and fourth code blocks is 2 each, and the number of convolutional blocks in the fifth and sixth code blocks is 4 each. The input resolution of the first code block is the input image resolution of the model. A downsampling layer is added at the end of the first 5 code blocks. The downsampling layer is a convolutional layer with learnable parameters. For the code blocks where the side length of the input feature map is exactly 1 / 8 and 1 / 16 of the side length of the model input image, that is, the fourth code block with an input feature map side length of 32 and the fifth code block with an input feature map side length of 16, a spatial self-attention module and a channel self-attention module are added in parallel after the activation function of each convolutional block.
[0098] b-3: Create the intermediate layer. The intermediate layer contains two code blocks. Each code block contains a convolutional block. For the convolutional block of the first code block, a spatial self-attention module is added after the activation function. The convolutional layer in the convolutional block is a three-dimensional convolutional layer with a channel dimension, and the activation function is the SiLU function. The resolution of the input and output feature maps of the intermediate layer is the same and is equal to the output feature map resolution of the sixth code block in the encoder.
[0099] b-4: Create each code block in the decoder. Each code block contains multiple stacked convolutional blocks. Each convolutional block consists of a convolutional layer and an activation function. The convolutional layer in the convolutional block is a three-dimensional convolutional layer with a channel dimension, and the activation function is the SiLU function. The number of convolutional blocks in the first and second code blocks is 4 each, the number of convolutional blocks in the third and fourth code blocks is 2 each, and the number of convolutional blocks in the fifth and sixth code blocks is 1 each. The input resolution of the first code block is the output feature map resolution of the intermediate layer. An upsampling layer is added at the end of the first 5 code blocks. The upsampling layer is a convolutional layer with learnable parameters. For the code blocks where the side length of the input feature map is exactly 1 / 8 and 1 / 16 of the side length of the model input image, that is, the second code block with an input feature map side length of 16 and the third code block with an input feature map side length of 32, a spatial self-attention module and a channel self-attention module are added in parallel after the activation function of each convolutional block.
[0100] b-5: Create skip connections between the corresponding code blocks of the encoder and decoder. For an integer i in the range [1, 6], concatenate the output feature map of the i-th code block in the encoder and the input feature map of the (7 - i)-th code block in the decoder along the channel dimension to fuse features of different levels and scales in the model;
[0101] b-6: Load the model into the video memory. The graphics card model used in this embodiment is NVIDIA GeForce RTX3090Ti.
[0102] Step c, train the model, including the following steps:
[0103] c-1: Take T = 1000 as the number of time steps for both the diffusion process and the reverse diffusion process;
[0104] c-2: Generate the noise weight β used for adding noise for each time step t t , to obtain the weight noise sequence:
[0105] {β 1 , β 2 , …, β 1000 |0 < β 1 < β 2 < … < β 1000 < 1}
[0106] where β 1 = 1×10 -5 , β 1000 = 0.02, and the sequence {β 1 , β 2 , …, β 1000} starts with β 1 . After obtaining the noise weight sequence, calculate the image weight α used for adding noise for each time step t t , to obtain the image weight sequence:
[0107]
[0108] The noise addition operation on the image is represented by the following formula:
[0109] x t = α t x t-1 + β t ε t
[0110] where ε t is a mixed noise signal obtained by weighted summation of a Gaussian noise signal and a simplex noise signal:
[0111] εt = ρg t + γs t
[0112] where g t is a Gaussian noise signal and g t ~ N(0, 1), s t is a simplex noise signal, ρ and γ are the weights of the noise signals, both set to 0.5. To simplify the noise addition process, calculate the image merging weights for each time step t:
[0113]
[0114] and the noise merging weights:
[0115]
[0116] and ensure establish the single-step noise addition formula:
[0117]
[0118] where x 0 represents the training image that has gone through all the preprocessing steps but has not been added noise, and x t represents the result of x 0 after t noise addition operations. The role of the single-step noise addition formula is to complete t noise addition operations in one step during the training and defect detection processes, reducing the processing time;
[0119] c-3: Set the batch size M to 4, take a batch of 4 training images from the training dataset, and randomly generate a time step batch containing 4 time steps. Any time step value t in the time step batch satisfies that t is a positive integer and t ∈ [1, 1000]. Perform preprocessing on each image in the batch, including size adjustment operations and pixel value remapping operations. In the size adjustment operation, the image will be scaled and padded to a square image with a side length of 256. After completion, remap the pixel values using the following formula:
[0120]
[0121] where x represents the image after size adjustment, and x 0 represents the result of the image x after all preprocessing;
[0122] c-4: Generate the noise-added image x 0 corresponding to the time step t for each training image x t in the batch:
[0123]
[0124] and the image x with noise added t - 1 times t-1 :
[0125]
[0126] Take x t as the input data of the model, and x t-1 as the prediction target of the model;
[0127] c - 5: The auto - encoder model created in step b takes the time step t and the noisy image x t as inputs, predicts a denoising signal, and after weighted summation with x t , can eliminate some noise in x t . The model predicts the denoising signal and weighted summation with x t is a denoising operation. The image obtained after one denoising operation is represented by the following formula:
[0128]
[0129] where ∈ θ represents the auto - encoder model created in step b, and θ represents the learnable parameters in the model;
[0130] c - 6: Calculate the loss function L of the model. Here, the mean - square error loss between the image at time step t - 1 predicted by the model and the real image x at time step t - 1 t-1 is used as the loss function L, which is represented by the following formula:
[0131]
[0132] where w is the side length of the input and output images, x t-1 (m, q) and respectively represent the color values of the pixel with abscissa m and ordinate q in x t-1 and . After obtaining the value of the loss function L, calculate the gradient of the model parameters and save it for subsequent adjustment of the model parameters;
[0133] After obtaining the value of the loss function L, calculate the gradient of the model parameters and save it for subsequent adjustment of the model parameters;
[0134] c-7: Set the upper limit of the number of training batches to 50000. For each image in the batch, perform the operations in steps c-4 to c-6. After all the images in the batch have been processed, perform gradient descent on the model parameters to adjust and optimize the model parameters. Then, take a new batch from the training dataset, generate the corresponding time-step batch, and perform the operations in steps c-4 to c-6 on the images and time steps in the new batch, enabling the model to learn how to gradually remove noise. After all batches in the training dataset have been calculated by the model, return to the beginning of the training dataset, take a new batch and continue training. Stop training when the loss function converges or reaches the pre-set upper limit of the number of batches before training.
[0135] Step d, add noise to the image to be measured and reconstruct the image, including the following steps:
[0136] d-1: Specify the number of noise-adding steps T N = 150;
[0137] d-2: Use the reset step size method for the denoising process. Set the number of times S that the model needs to perform inference in the short denoising process to 5, obtain the reference step size k = 30, the compensation time step number c = 0, and the time step sequence of the short denoising process:
[0138] {30, 60, 90, 120, 150, 151, 152, … 1000}
[0139] The time step lengths of the subsequence {1, 2, …, 149, 150} are reset and mapped to the new subsequence {30, 60, 90, 120, 150}, and the time steps after 150 are retained for readjusting the image weight and noise weight;
[0140] d-3: Calculate the noise weight sequence of the short denoising process. Let the auxiliary weight Traverse the time step sequence {1, 2, …, 1000} in ascending order. When the currently traversed time step t is in the time step sequence used in the short denoising process, calculate the new noise weight β′ of the time step t t :
[0141]
[0142] Add β′ t to the noise weight sequence of the short denoising process, and set to Continue to traverse the sequence {1, 2, …, 1000}. After traversal, obtain the noise weight sequence of the short denoising process And calculate the image weight i for each time step τ in the short denoising process Obtain the image weight sequence of the short denoising process
[0143]
[0144] d - 4: Use the single - step noise - adding formula to process the image x to be measured 0 Add noise to obtain the noisy image T N after adding noise to the image x to be measured 150 :
[0145]
[0146] d - 5: After readjusting the image weights and noise weights corresponding to each time step, let i = S, and use the model ∈ θ Repeatedly perform the following denoising operations:
[0147]
[0148] where τ i is a time step in the subsequence {τ 1 , τ 2 , …, τ S-1 , τ S}, and the value of i is subtracted by 1 after each denoising operation until the reconstructed image is obtained and then stop.
[0149] Step e, process and compare the images to be measured before and after reconstruction, and locate the defect area, including the following steps:
[0150] e - 1: Perform Gaussian blur processing on the original image x to be measured 0 and the reconstructed image to reduce the interference caused by the subtle differences in texture before and after reconstruction, and perform pixel value standardization and normalization;
[0151] e - 2: Calculate the pixel - level structural similarity loss function L 0 between the original image x to be measured and the reconstructed image in a sliding window manner. Specify the side length of the sliding window as 15, and perform padding with a width of 7 pixels on the upper, lower, left, and right sides of the image to be measured and the reconstructed image. For each pixel within the range of the original image, let the abscissa of the pixel in the padded image be m, and the ordinate be q. The value ranges of m and q are both [8, 263]. Take a sliding window with a side length of 2b + 1 centered on the pixel (m, q) from the image to be measured and the reconstructed image, and use the following formula to calculate the pixel - level structural similarity loss function of the two images at the pixel position with coordinates (m, q): SSIM where I is the original image x to be measured
[0152]
[0153] and the reconstructed image 0The sliding window centered on pixel (m, q), J is the reconstructed image The sliding window centered on pixel (m, q), μ I and μ J are the mean pixel values of I and J respectively, σ I and σ J are the variances of the pixel values of I and J respectively, σ IJ is the covariance of the pixel values of I and J, C 1 = 1×10 -4 , C 2 = 9×10 -4 , after calculating for all pixels (m, q), a pixel-level structural similarity loss map with side length w is obtained;
[0154] e-3: Set the fixed threshold to 0.4, find the positions in the pixel-level structural similarity loss map where the loss function value is greater than the threshold, and judge these positions as defect regions, marking the corresponding pixels with white pixels on a black background;
[0155] e-4: After checking all positions, a defect region mask image with a black background and white defect regions is obtained.
Claims
1. A fabric defect detection method based on image reconstruction, characterized in that: The method comprises the following steps: Step a: Collect images and organize data sets Collect normal fabric images and create a training dataset to let the model learn how to gradually remove noise from the image, so as to reconstruct the noisy image into a noise-free image; Step b: Create an autoencoder model Create a U-shaped autoencoder model with self-attention mechanism, which consists of an encoder, an intermediate layer and a decoder. The encoder contains several code blocks consisting of convolutional layers, activation functions, spatial self-attention modules, channel self-attention modules and downsampling layers; the decoder contains several code blocks consisting of convolutional layers, activation functions, spatial self-attention modules, channel self-attention modules and upsampling layers; the intermediate layer contains two code blocks consisting of convolutional layers, activation functions and spatial self-attention modules; Step c: training the model The training data set constructed in step a is used to train the autoencoder model created in step b. The training process goes through two processes: the diffusion process of gradually adding noise and the reverse diffusion process of gradually removing noise. The model learns how to reconstruct the image through denoising operations. Both denoising and denoising operations are performed in time steps. After t time steps in the diffusion or reverse diffusion process, it means that t denoising or denoising operations are performed. The total number of time steps in the diffusion process and the reverse diffusion process is set to T. Each time the model parameters are adjusted, a batch of training images and a batch of random time step values are used. The images in the batch correspond to the time steps one by one, and each The training images are denoised for the corresponding number of steps according to the corresponding time steps and input into the model. The model predicts the result of each image after a denoising operation and the actual result of each image after a denoising operation. The loss function between the output of the model and the actual denoising result is calculated. After the images and time steps in the batch are calculated, the gradient descent operation is performed according to the value of the loss function to adjust the parameters of the model. After all the batches in the training data set have been calculated by the model, the training data set is returned to the beginning to take out a new batch to continue training. The training is stopped when the loss function converges or reaches the upper limit of the number of batches preset before training. Step d: add noise to the image to be tested and reconstruct the image Specify the number of noise adding steps T N , T N Less than the number of time steps T of the reverse diffusion process, the image to be tested is T N After the step of denoising, the model obtained in step c is used for denoising to reconstruct the image to be tested. Before starting denoising, the T of the original denoising process is reset using the step size reset method. N The step operation is mapped into a short denoising process with fewer steps to improve the reconstruction speed. After the mapping is completed, the model formally performs step-by-step denoising on the image to be tested according to the time step sequence of the short denoising process. During the denoising process, the model eliminates the defect area and the noise together to complete the reconstruction of the image to be tested. Step e: Process and compare the images before and after reconstruction to locate the defect area. After the reconstruction is completed, the images to be tested before and after the reconstruction are Gaussian blurred, and the pixel values are standardized and normalized. The pixel-level structural similarity loss function in the images to be tested before and after the reconstruction is calculated in a sliding window manner. The area composed of pixels whose loss function values exceed a given threshold is the defect area.
2. The fabric defect detection method based on image reconstruction according to claim 1, characterized in that: The step b specifically comprises: b-1: Create the basic structure of the model. The model consists of an encoder, an intermediate layer, and a decoder. The encoder and decoder both contain N groups of code blocks. The last layer of the first N-1 code blocks of the encoder is a downsampling layer, which is used to halve the length of the feature map. The last layer of the first N-1 code blocks of the decoder is an upsampling layer, which is used to double the length of the feature map. The input image of the model is set to a square RGB image with a side length of w; b-2: Create each code block in the encoder. Each code block contains multiple stacked convolution blocks. Each convolution block consists of a convolution layer and an activation function. The input feature map resolution of the first code block is the input image resolution of the model. The first N-1 code blocks finally add a downsampling layer. From the second code block onwards, the input feature map resolution of each code block is equal to the output feature map resolution of the previous code block. When the side length of the input feature map of the code block meets the specified multiple relationship with the side length of the model input image, each convolution block in the code block contains a parallel spatial self-attention module and a channel self-attention module, which are added after the activation function to extract semantic information in the feature map. b-3: Create two code blocks in the middle layer. Both code blocks contain a convolution block. Each convolution block consists of a convolution layer and an activation function. The convolution block in the first code block contains a spatial self-attention module, which is added after the activation function. The second convolution block does not contain a self-attention module. The resolution of the input feature map and the output feature map are the same as the output feature map of the last code block of the encoder. b-4: Create each code block in the decoder. Each code block contains multiple stacked convolution blocks. Each convolution block consists of a convolution layer and an activation function. The input feature map resolution of the first code block is equal to the output feature map resolution of the intermediate layer. The upsampling layer is added to the first N-1 code blocks. From the second code block onwards, the input feature map resolution of each code block is equal to the output feature map resolution of the previous code block. When the side length of the input feature map of the code block meets the specified multiple relationship with the side length of the model input image, each convolution block in the code block contains a parallel spatial self-attention module and a channel self-attention module, which are added after the activation function to extract semantic information in the feature map. b-5: Create a skip connection between the corresponding code blocks in the encoder and decoder. For an integer i in the range [1, N], concatenate the output feature map of the i-th code block in the encoder and the input feature map of the N+1-i-th code block in the decoder in the channel dimension to fuse features of different levels and scales in the model. b-6: Load the model into video memory.
3. The fabric defect detection method based on image reconstruction according to claim 1, characterized in that: The step c specifically comprises: c-1: Specify an integer T greater than or equal to 100 as the number of time steps for the diffusion process and the reverse diffusion process. The diffusion process is the process of gradually adding noise, and the reverse diffusion process is the process of gradually removing noise. c-2: Set the noise weight β used when adding noise for each time step t t , and get the weight noise sequence: {β1,β2,…,β T |0<β1<β2<…<β T <1} The sequence {β1,β2,…,β T } starts with β1 and ends with β T After obtaining the noise weight sequence, the image weight αt used to add noise is calculated for each time step t to obtain the image weight sequence: After obtaining the noise weight sequence {β1, β2, …, β T } and the image weight sequence {α1, α2, …, α T }, the following formula is used to represent a denoising operation, that is, the image that has been denoised t-1 times is denoised once again to obtain an image that has been denoised t times: x t =a t x t-1 +b t e t Among them, ε t It is a mixed noise signal obtained by weighted summation of Gaussian noise signal and simplex noise signal: e t =ρg t +γs t where g t is a Gaussian noise signal and g t ~N(0,1),s t is the simplex noise signal, ρ and γ are the weights of the noise signal, Gaussian noise signal is used to let the model learn how to eliminate noise points, simplex noise signal is used to let the model learn how to eliminate large irregular color blocks, to simplify the noise adding process, calculate the image merging weights at each time step Get the image merging weight sequence: And calculate the noise merging weights for each time step Get the noise merging weight sequence: and ensure The single-step noise addition formula is established based on the merging weight: Where x0 represents the training image that has gone through all preprocessing steps but without adding noise, and x t It represents the result after x0 is subjected to t times of noise addition operation. The function of the single-step noise addition formula is to complete t times of noise addition operation in one step during the training and defect detection process, thus reducing the processing time. c-3: Take a batch of M training images from the training data set, and randomly generate a time step batch containing M time steps. Any time step value t in the time step batch satisfies t is a positive integer and t∈[1,T]. Preprocess each image in the batch, including resizing and pixel value remapping. In the resizing operation, the image will be scaled and padded to become a square image with a side length of w. After completion, the pixel values are remapped using the following formula: Where x represents the resized image, and x0 represents the result of image x after all preprocessing; c-4: Generate a noisy image x corresponding to time step t for each training image x0 in the batch t : And the image x that has been noisy t-1 times t-1 : x t As the input data of the model, x t-1 As the prediction target of the model; c-5: The autoencoder model created in step b combines the time step t and the noisy image x t As input, predict a denoised signal that is similar to x t After weighted summation, x can be eliminated t The model predicts the denoised signal and compares it with x t Weighted summation is a denoising operation. The image obtained after a denoising operation is It is expressed by the following formula: where ∈ θ represents the autoencoder model created in step b, and θ represents the learnable parameters in the model; c-6: Calculate the loss function L of the model, here we use the t-1 time step image predicted by the model and the real t-1 time step image x t-1 The mean square error loss between is used as the loss function L, which is expressed by the following formula: Where w is the side length of the input and output images, x t-1 (m, q) and Respectively represent x t-1 and The color value of the pixel with the horizontal coordinate m and the vertical coordinate q in the middle is used to obtain the value of the loss function L, and then the gradient of the model parameters is calculated and saved for subsequent model parameter adjustment; c-7: Perform steps c-4 to c-6 for each image in the batch. After all images in the batch have been processed, perform gradient descent on the model parameters to adjust and optimize the model parameters. Then take out a new batch from the training data set, generate the corresponding time step batch, and perform steps c-4 to c-6 on the images and time steps in the new batch to let the model learn how to gradually remove noise. After all batches in the training data set have been calculated by the model, return to the beginning of the training data set, take out a new batch and continue training. Stop training when the loss function converges or reaches the upper limit of the number of batches preset before training.
4. The fabric defect detection method based on image reconstruction according to claim 1, characterized in that: The step d specifically comprises: d-1: specifies the number of noise addition steps T which is less than the number of time steps T of the diffusion process N ; d-2: Use the re-step method for the denoising process to remap the time step sequence originally used in the denoising process to a length less than T N The time step sequence of denoising T N The original denoising process of times is changed to a denoising operation number less than T N Short denoising process, specified as greater than or equal to 1 and less than T N The integer S is used as the number of denoising operations in the short denoising process after mapping, and the benchmark step size k is calculated: And the number of compensation time steps c: c=mod(T N ,S) And calculate each time step τ of the short denoising process i , and get the time step sequence of the short denoising process: sequence At each time step τ i Calculated using the following formula: After resetting the step size, the original subsequence {1,2,…,T N -1, T N } is mapped into a new subsequence {τ1, τ2, …, τ S-1 , τ S }, T N The time steps after are retained and used to rescale the image weights and noise weights; d-3: Calculate the noise weight sequence of the short denoising process, and let the auxiliary weight Traverse the time step sequence {1, 2, ..., T} in ascending order. If the current traversed time step t belongs to Then calculate the noise weight β′ of time step t in the short denoising process t : β′ t Add the noise weight sequence of the short denoising process and Set as Continue to traverse the sequence {1, 2, ..., T}, and after the traversal is completed, the noise weight sequence of the short denoising process is obtained And for each time step τ in the short denoising process i Calculate the image weight Get the image weight sequence of the short denoising process d-4: Use the single-step noise addition formula to add noise to the image x0 to get the noise T N The image to be tested after After re-adjusting the image weights and noise weights corresponding to each time step, It is equivalent to using the time step τ in the short denoising process S One-step noise addition to get x S ; d-5: Let i = S, use model ∈ θ Repeatedly perform the denoising operation expressed by the following formula: where τ i is a subsequence {τ1,τ2,…,τ S-1 ,τ S }, the value of i is reduced by 1 after each denoising operation until the reconstructed image is obtained. Stop when.
5. The fabric defect detection method based on image reconstruction according to claim 1, characterized in that: The step e specifically comprises: e-1: the original image x0 and the reconstructed image to be tested Gaussian blur processing is performed to reduce the interference caused by subtle differences in texture before and after reconstruction, and pixel values are standardized and normalized; e-2: Calculate the original image x0 and the reconstructed image using a sliding window The pixel-level structural similarity loss function L between SSIM , the image to be tested has the same size as the model input image, both are square images with a side length of w. Before calculation, the sliding window side length is specified as 2b+1, where b is a positive integer. The four sides of the image to be tested and the reconstructed image are padded with a width of b pixels. For each pixel in the original image range, the horizontal coordinate of the pixel in the padded image is set to m, and the vertical coordinate is set to q. The value range of m and q is [b+1, b+w]. From the image to be tested and the reconstructed image, a sliding window with a side length of 2b+1 and a pixel (m, q) as the center is taken out respectively. The pixel-level structural similarity loss function of the two images at the pixel position with coordinates (m, q) is calculated using the following formula: Where I is the sliding window centered at pixel (m, q) in the original image x0 to be tested, and J is the reconstructed image The sliding window centered at pixel (m, q) in I and μ J are the mean pixel values of I and J, σ I and σ J are the pixel value variances of I and J, σ IJ is the pixel value covariance of I and J, C1 and C2 are constants, and after calculating all pixels (m, q), a pixel-level structural similarity loss map with a side length of w is obtained; e-3: Set a fixed threshold, find pixels whose loss function value is greater than the threshold in the pixel-level structural similarity loss map, judge these pixels as defect areas, and mark the corresponding pixels in white on a black background image of the same size as the image to be tested; e-4: After checking all pixels, a defect area mask image is obtained in which the background is black and the defect area is white.
Citation Information
Cited By
Lightweight multi-scale aluminum profile surface defect detection method
CN117911399A
Industrial product defect sample controllable generation method based on reference image guidance
CN120495301A
Defect sample generation method and device
CN120807410A
Woven fabric defect detection method based on AnDDPM unsupervised learning
CN121053436A