A fabric defect detection method with high-frequency feature constraint and edge entropy minimization
The fabric defect detection method based on high-frequency feature constraints and edge entropy minimization solves the problem of inaccurate detection of complex textured fabric images in existing technologies. It achieves higher precision in detecting minute defects and fine-grained defects, significantly improving the technical effect, detection accuracy and application precision, especially in the detection capabilities of high-frequency features and fine-grained features, thus meeting the needs of industrial production.
Patent Information
- Application Number
- CN202411051792.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-01
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-08-01
AI Technical Summary
Existing methods for detecting fabric defects and anomalies are insufficient when processing fabric images with complex textures and rich details, especially in the extraction and processing of high-frequency features and fine-grained features. This results in inaccurate detection of edge areas and minor defects, leading to low precision.
The detection method employs high-frequency feature constraints and edge entropy minimization. It constrains high-frequency information in the frequency domain through discrete Fourier transform, and optimizes the prediction confidence of the model for edge regions by combining the edge entropy minimization module. The model is trained using unsupervised learning, and automatic detection is only required for flawless fabric image data.
It significantly improves the accuracy of fabric defect and anomaly detection and the fineness of reconstructed images, enabling more accurate detection of minute defects and abnormal areas on the fabric surface, adapting to the needs of large-scale industrial production, reducing data preparation costs and avoiding manual annotation errors.
Smart Images

Figure CN119130905B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to cloth image defect anomaly detection technology, and particularly relates to a cloth defect detection method based on high-frequency feature constraint and edge entropy minimization. BACKGROUND
[0002] In industrial production, the control and improvement of product quality is always one of the core goals. Especially in the modern industrial environment, high-quality products not only determine the market competitiveness of enterprises, but also directly affect the economic benefits and brand reputation of enterprises.
[0003] In the textile industry, the quality of cloth directly affects the quality and market competitiveness of downstream products. Therefore, the detection of defects and anomalies on the surface of cloth has become a key link in the textile industry. However, in the process of cloth production, due to long-time operation or operation errors of machine equipment, faults or errors are inevitable, resulting in quality problems in the produced cloth. For example, holes, stains, fuzz and color difference in cloth, these defects and anomalies will have adverse effects on the appearance quality and use performance of cloth, and will bring direct economic losses to manufacturing enterprises. Therefore, real-time detection of defects and anomalies of cloth has become a crucial link in the production process. Traditionally, textile enterprises usually adopt manual detection to monitor the defects and anomalies of cloth. Although this method is direct, it has significant limitations. First, manual detection requires a large amount of human resources and has low detection efficiency, which is difficult to meet the needs of large-scale production. Second, the accuracy of manual detection is easily affected by the subjective factors of the detection personnel, and with the extension of working time, the accuracy and consistency of detection will gradually decrease. In addition, manual detection is difficult to accurately find small and complex defects, especially in high-precision cloth production, the effectiveness of this method is more insufficient. Therefore, developing efficient, accurate and automated cloth defect and anomaly detection technology has become a major issue to be solved in the current textile industry field.
[0004] In the context of modern industrial automation and intelligent manufacturing, utilizing computer vision and artificial intelligence technologies to achieve automated defect detection of cloth has become an important means and development trend. Computer vision technology acquires image data of cloth through camera equipment and analyzes it using image processing and pattern recognition algorithms, which can effectively detect defects on the surface of cloth. However, traditional computer vision methods mainly rely on manually designed feature extraction, which often struggles to handle different types of defect conditions in complex industrial environments. In recent years, with the rapid development of deep learning technology, image processing algorithms based on deep learning have made significant progress in computer vision. Deep learning can automatically learn high-level features in images through multi-layer neural network training, significantly improving image classification, object detection, and image segmentation performance. In the field of unsupervised learning, deep learning algorithms have shown great potential, enabling effective feature learning and defect detection without the need for large amounts of labeled data.
[0005] Existing unsupervised image defect detection methods mainly include algorithms based on feature embedding, simulation synthesis, and image reconstruction. Feature embedding-based methods use convolutional neural networks to extract high-level features from images, map these features to low-dimensional embedding spaces, and detect anomalies by measuring the deviation of features in the embedding space. Simulation synthesis methods train defect detection models by introducing anomalies into normal data or generating synthetic data. Image reconstruction-based methods are currently the most widely used, identifying abnormal regions by measuring the reconstruction ability of non-anomalous images, often using autoencoders, generative adversarial networks, and other models. Although image reconstruction-based defect detection methods perform well on general datasets, their performance significantly decreases in specific industrial scenarios due to the presence of more details, textures, and structures in cloth images. The detailed texture features in cloth images are often contained in high-frequency information, while neural networks tend to fit low-frequency information during training, resulting in the loss of some fine-grained information. In addition, the high-quality standards of cloth require defect detection models to accurately identify minor defects, which poses higher accuracy requirements for detection models. The uncertainty of existing methods in predicting the edges of defect regions further limits their effectiveness in cloth image defect detection.
[0006] To address the above problems, the present invention proposes a cloth defect detection method based on high-frequency feature constraint and edge entropy minimization.
[0007] Currently, there have been some studies on cloth image defect detection:
[0008] For example, Chinese invention patent application No. CN202010339971.4 discloses a cloth defect detection method based on One-Class deep support vector description. Under the condition of semi-supervision, the advantages of deep convolutional neural network are used to extract effective deep features of images. Through training, a hyper-sphere model that can accurately describe normal sample points is mapped in a high-dimensional space. End-to-end cloth defect detection is achieved. The relationship between the test sample and the hyper-sphere can be completely described by the network parameters and the hyper-sphere radius obtained by training, realizing the discrimination of defects. The problems of large memory occupation, slow detection speed, and limitation of defect types in previous models are solved.
[0009] For example, Chinese invention patent application No. CN202311189949.6 discloses a textile defect rapid detection method based on image features. By acquiring a textile gray image and performing block processing, all attribution pixel pairs and all original pixel pairs in each reference direction within each sliding window are obtained. The overall adjustment amount of the pixel pairs of the reset textile gray image in each reference direction is calculated according to these data, and then the error degree of the reset textile gray image in each reference direction is obtained. By analyzing the texture distribution regularity degree of the image block, the defects of the textile can be quickly detected. This method has the advantages of being fast, accurate, and efficient.
[0010] For another example, Chinese invention patent application No. CN202311052085.3 discloses a fabric defect detection method based on convolutional neural network and repetitive pattern analysis, which includes: collecting fabric defect images, performing pixel-level labeling and data enhancement processing on the fabric defect images, establishing a fabric defect image dataset, and dividing the dataset into a training set and a validation set; building a repetitive pattern detection model for repetitive pattern detection; inputting the fabric defect images into the repetitive pattern detection model to obtain the output feature map and adjusting the size to serve as the input label of the semantic segmentation network for defect detection, extracting the peak value and converting it into the target label of the semantic segmentation network for defect detection using one-hot encoding; building and training the semantic segmentation network, saving the necessary part of the semantic segmentation network for fabric defect detection after training, and obtaining a fabric defect detection model; using the saved fabric defect detection model and repetitive pattern detection model for fabric defect detection, which is suitable for periodic fabric defect detection and has high detection accuracy.
[0011] For another example, Chinese invention patent application No. CN202211231211.7 discloses a fabric flaw line defect identification method for textile production. The method obtains a fabric image after sewing is completed, and semantically segments a fabric region image to be detected. According to the region features on the fabric, a defect line is identified. According to the position of the defect line, the stretching degree of the corresponding line is adjusted. Based on industrial digitization, the present application uses image processing technology to realize identification and enhancement of the flaw line, so as to better facilitate defect detection and distinguish whether the fabric contains defects.
[0012] The defects of the above-mentioned invention patent applications are that these methods have limited processing capability when facing complex texture and detail-rich cloth images, especially in the extraction and processing of high-frequency features and fine-grained features, resulting in inaccurate detection of edge regions and subtle flaws and low precision. The problem addressed by the present patent is more in line with the actual needs of the current industry, and has greater application prospects. SUMMARY
[0013] The present application aims to provide a cloth flaw detection method with high-frequency feature constraint and edge entropy minimization.
[0014] The method comprises the following specific steps:
[0015] Step 1. Collect cloth training picture data sets without flaws and abnormalities, and initialize hyperparameters, initialize and fine-tune model parameters;
[0016] Step 2. Randomly sample an image sample x, obtain vector quantized image features through a feature encoder, generate flaw noise through an anomaly generator and add it to the original features of the image, input the features with anomalies into an image reconstruction decoder and an image restoration decoder respectively, and reconstruct the images without anomalies and with anomalies respectively;
[0017] Step 3. Calculate the high-frequency feature constraint loss Supervise the quality of the reconstructed anomaly-free image to ensure the reconstruction quality of the detail texture features;
[0018] Step 4. Calculate the difference image between the reconstructed original input image and the reconstructed anomaly-free image, and realize feature enhancement of the abnormal area;
[0019] Step 5. Input the reconstructed original input image and the reconstructed anomaly-free image and the difference image of the two into an anomaly detector, and calculate the anomaly detection loss L det ;
[0020] Step 6. Calculate the entropy loss Optimize the prediction of the edge region with low certainty, and improve the confidence of the model in predicting the edge region.
[0021] Step 7. Calculate joint loss and backpropagate to update parameters;
[0022] Step 8. Repeat steps 2 to 7 until the maximum number of iterations I is reached or the model parameters reach convergence;
[0023] Step 9. In the use phase of the model: input the test picture into the completed training fabric defect anomaly detection model, perform defect detection on the test picture, and select the result with confidence greater than a certain threshold Φ in the model output anomaly score result as the abnormal area.
[0024] Further, in step 1, data collection and model initialization, the specific steps are:
[0025] Step 101, in the actual fabric textile factory operation environment, use the camera module to capture and collect different types of flawless fabrics, so that each picture can clearly show the appearance of the fabric;
[0026] Step 102, set the balance parameters a and β and the maximum number of iterations I;
[0027] Step 103, use the fully pre-trained deep feature extractor, image restoration decoder and feature quantization codebook to initialize the model, and then fine-tune the model parameters using the data of the target data set, so as to minimize the image restoration loss and the difference between the feature space projection F calculated by the deep feature extractor and its codebook vectorized feature Q. The loss function can be represented as:
[0028]
[0029] In the above formula (1), F represents the feature space projection calculated by the deep feature extractor, Q represents the corresponding codebook vectorized feature, x rec represents the input image restored by the image restoration decoder, sg[·] represents the stop gradient calculation, L2 represents the L2 norm, which is used to measure the distance between the two points corresponding to the original image and the generated picture. After the first stage of training, the feature encoder, image restoration decoder and feature quantization codebook will be fixed.
[0030] Further, in step 2, encode and decode the input image and add abnormal features, the specific steps are:
[0031] Step 201, feature extraction: the input picture x is extracted to obtain the feature f(x) via the deep feature extractor. The deep feature extractor consists of 16 residual blocks, each of which contains two 3x3 convolution layers with a step of 1 and a 1x1 convolution layer. Each residual block contains a jump connection that directly adds the input of the block to the output;
[0032] Step 202, feature quantization: in order to generate abnormal region data closer to the normal distribution, the input picture is encoded by using vector quantization, which is a technology for mapping continuous latent space into discrete features. By introducing a discrete codebook in the latent space, the model can learn the high-level abstract representation of the data more effectively. By introducing a discrete codebook, the vector quantization divides the latent space into a limited number of discrete regions, each region is represented by a specific codebook vector. Specifically, the feature quantization process of the feature f(x) extracted by the deep feature extractor is represented as:
[0033]
[0034] In the above formula (2), f(x) represents the feature extracted by the deep feature extractor from the input picture x, f(x) ij represents the feature of the feature map (i,j) position, e k represents a vector in the codebook, where k is the index of the codebook vector, k∈(1…K), where K is the number of features in the codebook, argmin represents the value of e k that makes the function minimum, ||·|| represents the Euclidean distance, finally, each point in the feature space is replaced by the nearest neighbor vector in the codebook.
[0035] Step 203, use two-level vector quantization to improve the quality of the reconstruction result: use low-resolution and high-resolution codebooks to encode the input image in two levels, the quantized feature map of the low-resolution codebook is reduced by 4 times compared to the original input image, and the quantized feature map of the high-resolution codebook is reduced by 8 times. By this two-level encoding method, the multi-level features of the input image can be better captured, and the reconstruction result with higher quality can be generated.
[0036] Step 204, constructing cloth feature anomaly generation module: in the training stage, the model can only contact the input image without anomaly, and needs to generate anomaly for the input image. The current noise generation method is divided into image level and feature level. The image level noise generation can simulate limited anomalies, while the feature level noise generation is more robust. By introducing noise in specific features of the image, the feature level noise generation can simulate the change of abnormal features, thereby greatly increasing the diversity of abnormal information. In order to balance the randomness of abnormal area and at the same time make the model have better discrimination ability for more difficult to distinguish anomalies, two methods of random noise addition and quantized subspace noise addition are combined in anomaly generation: first, the mask Y of the abnormal area is randomly sampled, and the value of 1 represents the abnormal area. The random noise addition superimposes the generated random noise on the selected area features. The quantized subspace noise addition method is based on the quantized features, and randomly selects features from the feature space of the picture data to replace the selected area features, while limiting the similarity between the replaced features and the original normal features, and excluding the most similar 5% vectors;
[0037] Step 205, constructing cloth image decoder: in the image decoding stage, image reconstruction decoder and image restoration decoder are introduced respectively. The target of the image reconstruction decoder is to restore the abnormal area to the normal appearance of the sample observed during training, while the image restoration decoder directly restores the input features to the image.
[0038] Step 206, for the image reconstruction decoder: the decoder extracts features through two convolutional blocks. The first convolutional block contains a 3x3 convolutional layer to maintain the input channel number, followed by an instance normalization layer and a ReLU activation function, followed by another 3x3 convolutional layer to double the channel number, followed by another normalization layer and ReLU activation function. The second convolutional block operates in a similar manner to the first convolutional block. Each convolutional block is followed by a max pooling layer using a 2x2 window and a stride of 2 for down-sampling. Next, the feature map passes through a 1x1 convolutional layer to reduce the channel number to 64. Then the feature map passes through two up-sampling blocks, each using a 4x4 transpose convolution with a stride of 2 for up-sampling operation, and a ReLU activation function after each up-sampling block. After that, the feature map passes through a 3x3 convolutional layer and then enters two transpose convolutional layers to gradually reduce the channel number for reconstructing the RGB channels of the original image. After each transpose convolutional layer, a ReLU activation function is connected except for the last layer. Finally, the network outputs the reconstructed image.
[0039] Step 207, for image restoration decoder: the decoder uses a fixed pre-trained model to realize the direct reconstruction of the input feature, the decoder includes a convolutional layer for preliminary processing of the latent representation, deep feature extraction is performed through a residual stack, in which each residual block contains two convolutional layers with ReLU activation function, an upsampling is performed through a deconvolutional layer, a further feature extraction is performed through a convolutional layer with ReLU activation function, and the latent representation is mapped to the space of the original image through another deconvolutional layer, both the image reconstruction decoder and the image restoration decoder include operations on input features of two different resolutions, the low-resolution quantized encoded feature map is subjected to upsampling and convolution operation and then feature extraction through the residual module, the high-resolution feature map is also subjected to upsampling and convolution operation, the two feature maps are concatenated together, and the final reconstructed image is generated through transposed convolution.
[0040] Further, in step 3, the high-frequency feature constraint loss is calculated, and the specific steps are as follows:
[0041] Step 301, constructing a high-frequency feature constraint module: in the image reconstruction-based anomaly detection model, the higher the quality of the de-anomaly reconstructed image, the stronger the understanding and learning ability of the model to normal samples, which helps to improve the robustness and generalization ability of the model, making it more reliable in actual application, when the model is used for anomaly detection task, it can better distinguish normal and abnormal samples, a common way to enhance the reconstruction quality is to use Euclidean distance loss to directly constrain the consistency of the de-anomaly reconstructed image and the original image, but since the priority of neural network fitting frequency information is different in the whole training process, it is usually from low to high, which means that part of the high-frequency information may not be valued by the network, and the detail texture information is just contained in the high-frequency information;
[0042] In order to narrow the gap between the reconstructed de-anomaly picture and the original picture in high-frequency information, a high-frequency feature constraint module is set, and discrete Fourier transform (DFT) is used to convert the original input picture and the reconstructed de-anomaly picture into frequency domain:
[0043]
[0044] In the above formula (3), F(u,v) is the representation of the image in the frequency domain, (u,v) represents the coordinates in the frequency domain, the size of the input picture is MxN, x(m,n) represents the pixel value of the (m,n) point in the picture, e and i are natural constant and imaginary unit respectively, as shown in the following formula (4):
[0045]
[0046] In the above formula (4), cos represents cosine, sin represents sine, (u, v) represents coordinates in the frequency domain, the size of the input picture is MxN, x(m, n) represents the pixel value of the (m, n) point in the picture, e and i are natural constants and imaginary units respectively;
[0047] Step 302, after the discrete Fourier transform, the amplitude and phase of the image are two key elements describing the frequency domain information, the phase describes the spatial position and relative position of the frequency component in the image, determines the starting point of the sine wave in the image, thereby affecting the spatial layout of different frequencies in the frequency domain, the amplitude represents the response strength or energy of the image to a specific frequency, in the centralized amplitude graph, low frequency information is in the middle and high frequency information is in the periphery, in order to design a loss function that constrains the loss of high frequency information, for the real part and imaginary part of F(u, v) in formula (3), let the real part R(u, v) = a, the imaginary part I(u, v) = b, then F(u, v) is expressed as the following formula (5):
[0048] F(u, v) = R(u, v) + I(u, v)i = a + bi …… (5),
[0049] In the above formula (5), (u, v) represents coordinates in the frequency domain, R(u, v) and a represent the real part of the frequency domain information, I(u, v) and b represent the imaginary part of the frequency domain information, and i is an imaginary unit;
[0050] Step 303, calculate the amplitude of the image, as follows:
[0051]
[0052] In the above formula (6), (u, v) represents coordinates in the frequency domain, R(u, v) and a represent the real part of the frequency domain information, I(u, v) and b represent the imaginary part of the frequency domain information, and |·| represents the modulus of a complex number;
[0053] Step 304, calculate the phase of the image according to the following formula (7):
[0054]
[0055] In the above formula (7), (u, v) represents coordinates in the frequency domain, R(u, v) represents the real part of the frequency domain information, I(u, v) represents the imaginary part of the frequency domain information, and arctan is the tangent function, here the distance measurement is performed by mapping each frequency value to a Euclidean vector in two-dimensional space;
[0056] Step 305, calculate the spatial frequency value at the spectrum coordinate (u, v) of the input image according to the following formula (8):
[0057] F in (u, v) = ain +b in i……(8),
[0058] In the above formula (8), a in represents the real part of the input image spectrum, b in represents the imaginary part of the input image spectrum, and i is the imaginary unit;
[0059] Step 306, calculate the spatial frequency value at the spectral coordinates (u, v) of the de-anomaly reconstructed image according to the following formula (9):
[0060] F gen (u, v) = a gen +b gen i……(9),
[0061] In the above formula (9), a gen represents the real part of the de-anomaly reconstructed image spectrum, b gen represents the imaginary part of the de-anomaly reconstructed image spectrum, and i is the imaginary unit;
[0062] Step 307, calculate the global distance of the frequency using the square of the Euclidean distance, as follows:
[0063]
[0064] In the above formula (10), and represent two vectors mapped from F in (u, v) and F gen (u, v), based on the definition of amplitude and phase, the size of the vector and correspond to the amplitude, and the angles θ in and θ gen correspond to the phase, ||·|| represents the Euclidean distance, and |·| represents the modulus of a complex number;
[0065] Step 308, calculate the frequency distance between the original input image and the de-anomaly reconstructed image, represented by the average of the global distance, as follows:
[0066]
[0067] In the above formula (11), F in represents the input image spectrum, F gen represents the de-anomaly reconstructed image spectrum, (u, v) represents the coordinates in the frequency domain, the size of the input image is M x N, and |·| represents the modulus of a complex number;
[0068] Step 309, in order to make the model more focused on learning key fine-grained features, filter out low-frequency information to reduce the model's attention to overall structural information, capture subtle textures and detailed features in the image, here introduce a mask M r , assuming that the center point coordinates of the image are (0, 0), M r is M x N in size, the value of the central area rM x rN is 0, and the value of the remaining area is 1, where r ∈ (0, 1), the mask function is calculated as follows (12) :
[0069]
[0070] In the above formula (12), (u, v) represents the coordinates in the frequency domain, the size of the input picture is M x N, and r is the central area size ratio, is an indicator function used to determine whether the frequency domain coordinates (u, v) are outside the specified central area, if outside the area, the mask function M r (u, v) takes the value 1; if inside the area, it takes the value 0;
[0071] Step 310, calculate the high-frequency feature constraint loss as follows (13) :
[0072]
[0073] In the above formula (13), (u, v) represents the coordinates in the frequency domain, the size of the input picture is M x N, and r is the central area size ratio, F in represents the input image spectrum, F gen represents the abnormal reconstruction image spectrum.
[0074] Further, in step 4, a flaw feature enhancement module is constructed, the specific steps are as follows:
[0075] Step 401, construct a flaw feature enhancement module to accept two inputs of the reconstructed original input image and the reconstructed anomaly-removed image respectively;
[0076] Step 402, calculate the difference image of the reconstructed original input image and the reconstructed anomaly-removed image, the difference image is obtained by directly subtracting the former two to highlight the difference area between the images and achieve feature enhancement of the abnormal area;
[0077] Step 403, splice the reconstructed original input image, the reconstructed anomaly-removed image and the difference image of the two to input into the cloth flaw anomaly detection module.
[0078] Further, in step 5, a cloth flaw anomaly detection module is constructed, the specific steps are as follows:
[0079] Step 501, constructing a cloth defect anomaly detection module: the function of the anomaly detection module is to make anomaly prediction for each pixel point, and present it in the form of anomaly score, indicating the difference degree between the input image and the normal mode. The anomaly detection module accepts three input images, including the reconstructed original input image, the reconstructed anomaly-removed image, and the difference image of the above two images. An anomaly score is predicted through a deep defect detection network. The deep defect detection network includes an encoder for capturing hierarchical features and a decoder for reconstructing high-resolution anomaly prediction. The encoder part is composed of four convolutional blocks, each of which includes two 3x3 convolutional layers for feature extraction. The output channel number of the first convolutional layer is consistent with the input, and the output channel number of the second convolutional layer is twice that of the first convolutional layer. Each convolutional layer is followed by an instance normalization layer and a ReLU activation function. Each pair of convolutional layers is followed by a 2x2 max-pooling layer for down-sampling, gradually reducing the spatial size of the feature map while increasing the number of feature channels. The decoder part is composed of up-sampling and convolutional blocks, which are used to gradually restore the spatial resolution of the image. Each up-sampling step uses bilinear interpolation and is followed by a 3x3 convolutional layer to adjust the number of channels, and then a ReLU activation function. After each up-sampling step, a convolutional block is used for further feature refinement, which includes two 3x3 convolutional layers. In each up-sampling step, the decoder also performs a skip connection with the feature map of the corresponding layer of the encoder. Finally, the output of the decoder is mapped to the required output channel number through a 1x1 convolutional layer.
[0080] Step 502, calculating the defect detection loss: training by minimizing the anomaly detection loss function between the predicted score P and the anomaly region mask Y. The loss function is represented by the following formula (14):
[0081]
[0082] In the above formula (14), Y i is the true label of the i-th pixel point (0 represents normal or 1 represents anomaly), P i is the predicted value of the i-th pixel point, is a balance factor used to balance the proportion of positive and negative samples, and γ is an adjustment factor used to adjust the weight of difficult and easy samples.
[0083] Further, in step 6, an edge entropy minimization module is constructed, and the specific steps are as follows:
[0084] Step 601, when performing pixel-level anomaly region prediction, the model will produce low-entropy prediction in the central part close to the anomaly region, and high-entropy prediction in the edge part of the anomaly region. Specifically, anomaly features are usually more prominent in the central part of the anomaly region, and the model is more likely to learn common features in these central regions, resulting in confident predictions for the central regions. However, when the model faces the edge part of the anomaly region, due to the difference and complexity of the data distribution or lack of abnormal information, the model cannot accurately capture the detailed information of the edge part, resulting in high-entropy predictions in these areas. This shows that using only the above step 502 of the supervised mode is still insufficient to make the model produce high-deterministic predictions for the edge region of the anomaly, and one way to improve the detection performance of the edge part of the anomaly detection model is to encourage the model to produce high-confidence (low-entropy) predictions during target prediction. In this setting, an edge entropy minimization module is designed to directly punish low-confidence predictions using entropy loss.
[0085] Specifically, the edge entropy minimization module does not consider using the generated anomaly region mask Y to improve the confidence of the edge prediction, but rather considers a more direct constraint method, i.e., by minimizing the entropy of the prediction. According to the Shannon entropy principle, for an input image x, its entropy map is composed of independent pixel-level entropies, and the entropy value of each pixel point is as follows (15):
[0086]
[0087] In the above formula (15), represents the anomaly prediction value of the model for the image (m, n) position;
[0088] Step 602, to improve the confidence of the anomaly detection model in the edge uncertain region, a re-weighting mechanism is introduced. Since the prediction values of the edge uncertain region are all close to 0.5, the distance between the prediction absolute value and 0.5 is calculated, and a greater loss supervision is applied to the pixel points closer to 0.5. For the (m, n) position of the image, a distance weight is introduced as follows (16):
[0089]
[0090] In the above formula (16), λ is a hyperparameter that controls the decay rate of the weight, and its value range is between 0 and 1. The value of λ directly affects the decay rate of the distance weight. The larger the λ, the faster the decay rate, and the weight of the prediction value with a larger distance decreases faster. The smaller the λ, the slower the decay rate, and the weight of the prediction value with a larger distance decreases relatively slowly.
[0091] Step 603, the entropy loss is defined as the sum of all pixel-level entropies as follows (17):
[0092]
[0093] In the above formula (17), represents the entropy value at the pixel point (m, n), represents the distance weight at the pixel point (m, n).
[0094] Further, in step 7, the joint training module, the specific steps are:
[0095] Step 701, calculate the joint loss according to the following formula (18) to update the parameters of the image reconstruction decoder and the anomaly detector:
[0096]
[0097] In the above formula (18), is the anomaly detection loss, is the consistency loss of the de-anomaly reconstructed image and the original image, is the high-frequency feature constraint loss, is the entropy loss, and α and β are weight factors of and .
[0098] Further, in step 9, the actual use of the model, the specific steps are:
[0099] Step 901, for the input image with defects, it is first passed through the feature encoder to obtain the feature with the abnormal area;
[0100] Step 902, the feature with the abnormal area is sent into the image reconstruction decoder and the image restoration decoder respectively, and the de-anomaly and the image with the anomaly are reconstructed respectively;
[0101] Step 903, the generated anomaly and de-anomaly images and the difference image between the two are spliced and input into the anomaly detector;
[0102] Step 904, the anomaly possibility of each pixel point is predicted, and the score of the pixel point in the final prediction result of the defect area will be closer to 1, and the score of the non-defect area will be closer to 0;
[0103] Step 905, for the input image without defects, after the same inference process, the defect anomaly scores of the entire image area will be closer to 0.
[0104] The method described in the application has the following superior effects compared with the existing similar methods in the technical field:
[0105] 1. Using the method of the present application, the detection of fabric defect anomalies is more accurate. By introducing a high-frequency feature constraint module and an edge entropy minimization module, the accuracy of fabric defect anomaly detection is significantly improved. The high-frequency feature constraint module uses discrete Fourier transform to convert real images and reconstructed images into frequency domain representation. By reducing the difference between high-frequency information, the consistency of the reconstructed image and the original image in terms of detail texture is ensured. The edge entropy minimization module penalizes low-confidence predictions in the defect anomaly region by minimizing the entropy loss, enhancing the model's prediction certainty in the edge region of the defect anomaly. The combination of the high-frequency feature constraint module and the edge entropy minimization module enables the method of the present application to more accurately detect subtle defects and abnormal regions on the fabric surface, thereby improving the accuracy of detection.
[0106] 2. Using the method of the present application, the fabric defect area reconstruction is highly precise. Compared with existing fabric defect anomaly detection methods, the method of the present application supervises the quality of the reconstructed anomaly-free image through the high-frequency feature constraint module to ensure the reconstruction quality of detail texture features. The established module constrains high-frequency features in the frequency domain to ensure the consistency of the reconstructed image in terms of detail texture, thereby improving the precision of the reconstructed image. Whether it is an anomaly-free image or an image with anomalies, the reconstructed result can retain more detail information, significantly improving the precision of the reconstructed image.
[0107] 3. Using the method of the present application, the detection accuracy of the defect edge region is higher. Compared with existing fabric defect anomaly detection methods, the edge entropy minimization module of the method of the present application optimizes low-confidence predictions during the anomaly detection stage to improve the confidence of the model's prediction of the edge region. By minimizing the entropy loss, this module penalizes low-confidence predictions in the defect anomaly region, thereby enhancing the detection accuracy of the model in the edge region of the defect anomaly. In this way, the model not only detects defect regions but also accurately identifies the edges of defects, resulting in higher overall detection accuracy, especially when processing complex industrial fabric images.
[0108] 4. Using the method of the present application, the fabric defect anomaly detection training samples do not need to be labeled. Only unblemished fabric image data is required during model training, and no labeling of training samples is needed. This advantage significantly reduces the cost and time of data preparation, while avoiding subjective errors caused by manual labeling. Through unsupervised learning, the model can automatically learn and detect defect anomaly regions in fabric images, making the entire detection process more efficient and automated, and adapting to the needs of large-scale industrial production.
[0109] 5. The method can significantly enhance the fine-grained features of the reconstructed image and the defect anomaly detection effect of the edge region in the complex industrial environment, improve the fine degree of the reconstructed image and the precision of the defect anomaly detection, provide a new solution for the defect anomaly detection of the cloth image, and has wide application prospect and important practical significance. BRIEF DESCRIPTION OF DRAWINGS
[0110] Figure 1 The flowchart of the method.
[0111] Figure 2 The framework diagram of the training stage of the method.
[0112] Figure 3 is a framework diagram of the test reasoning stage of the method. Figure 3A The test reasoning stage framework diagram of the cloth image with defects, Figure 3B The test reasoning stage framework diagram of the cloth image without defects. DETAILED DESCRIPTION
[0113] In order to make the purpose, technical scheme and advantages of the method clearer, the method will be further described in detail below in combination with the drawings and specific examples.
[0114] The method comprises the following steps:
[0115] As shown in Figure 1 , Figure 2 Step 1, data acquisition and model initialization, the specific steps are:
[0116] Step 101, in the actual cloth textile factory operation environment, use the camera module to shoot and collect different types of defect-free cloth, so that each picture can clearly show the appearance of the cloth;
[0117] Step 102, set the balance parameters a and b and the maximum iteration number I;
[0118] Step 103, use the fully pre-trained deep feature extractor, image restoration decoder and feature quantization codebook to initialize the model, and then use the data of the target data set to fine-tune the model parameters, so as to minimize the image restoration loss and the difference between the feature space projection F calculated by the deep feature extractor and its codebook vectorization feature Q, and the loss function is represented as:
[0119]
[0120] In the above formula (1), F represents the feature space projection calculated by the deep feature extractor, Q represents the corresponding codebook vectorization feature, x recsg[img_recon(x, f(x), Q)] represents the input image restored by the image restoration decoder, sg[·] represents the stop gradient calculation, L2 represents the L2 norm used to measure the distance between the two points of the original image and the corresponding position of the generated picture, after the first stage training, the feature encoder, the image restoration decoder and the feature quantization codebook will be fixed;
[0121] Step 2, encoding and adding abnormal features to the input image, as shown in the figure, the specific steps are as follows: Figure 2
[0122] Step 201, feature extraction: the input picture x is extracted to obtain the feature f(x) through the deep feature extractor, the deep feature extractor is composed of 16 residual blocks, each residual block contains two 3x3 convolution layers with a step of 1 and a 1x1 convolution layer, and each residual block contains a jump connection which directly adds the input of the block to the output;
[0123] Step 202, feature quantization: in order to generate abnormal region data closer to the normal distribution, the input picture is encoded by using vector quantization, which is a technology that maps continuous latent space into discrete features. By introducing a discrete codebook in the latent space, the model can more effectively learn the high-level abstract representation of the data. By introducing a discrete codebook, the latent space is divided into a limited number of discrete regions, each of which is represented by a specific codebook vector. Specifically, the feature quantization process of the feature f(x) extracted by the deep feature extractor is represented as:
[0124]
[0125] In the above formula (2), f(x) represents the feature of the input picture x extracted by the deep feature extractor, f(x) ij represents the feature of the feature map (i,j) position, e k represents a vector in the codebook, where k is the index of the codebook vector, k∈(1…K), and K is the number of features in the codebook, argmin represents the value of e k that makes the function minimum, and ||·|| represents the Euclidean distance. Finally, each point in the feature space is replaced by the nearest neighbor vector in the codebook.
[0126] Step 203, using two-level vector quantization to improve the quality of the reconstruction result: using low-resolution and high-resolution codebooks to encode the input image in two levels, the quantized feature map of the low-resolution codebook is reduced by 4 times compared with the original input image, and the quantized feature map of the high-resolution codebook is reduced by 8 times. Through this two-level encoding method, the multi-level features of the input image can be better captured, and a higher quality reconstruction result can be generated.
[0127] Step 204, build cloth feature anomaly generation module: in the training stage, the model can only contact the input image without anomaly, and needs to generate anomaly for the input image. The current noise generation method is divided into image level and feature level. The image level noise generation can simulate limited anomalies, while the feature level noise generation is more robust. By introducing noise in specific features of the image, the feature level noise generation can simulate the change of abnormal features, thereby increasing the diversity of abnormal information. In order to balance the randomness of abnormal area and at the same time make the model have better discrimination ability for more difficult to distinguish anomalies, two methods of random noise addition and quantized subspace noise addition are combined in anomaly generation: first, the mask Y of the abnormal area is randomly sampled, and the value of 1 represents the abnormal area. The random noise addition superimposes the generated random noise on the selected area features. The quantized subspace noise addition method randomly selects features from the feature space of the picture data to replace the selected area features, while limiting the similarity between the replaced features and the original normal features, and excluding the most similar 5% vectors;
[0128] Step 205, build cloth image decoder: in the image decoding stage, image reconstruction decoder and image restoration decoder are introduced respectively. The goal of image reconstruction decoder is to restore the abnormal area to the normal appearance of the sample observed during training, while the image restoration decoder directly restores the input features to the image;
[0129] Step 206, for image reconstruction decoder: the decoder extracts features through two convolution blocks. The first convolution block contains a 3x3 convolution layer to maintain the input channel number, followed by an instance normalization layer and a ReLU activation function, followed by another 3x3 convolution layer to double the channel number, followed by another normalization layer and ReLU activation function. The second convolution block operates in a similar manner to the first convolution block. Each convolution block is followed by a max pooling layer, which uses a 2x2 window and a stride of 2 for down-sampling. Next, the feature map passes through a 1x1 convolution layer to reduce the channel number to 64. Then the feature map passes through two up-sampling blocks, each using a 4x4 transpose convolution with a stride of 2 for up-sampling operation, and a ReLU activation function after each up-sampling block. After that, the feature map passes through a 3x3 convolution layer and then enters two transpose convolution layers to gradually reduce the channel number for reconstructing the RGB channels of the original image. After each transpose convolution layer, a ReLU activation function is connected except for the last layer. Finally, the network outputs the reconstructed image;
[0130] Step 207, for the image restoration decoder: the decoder uses a fixed pre-trained model to realize the direct reconstruction of the input feature, the decoder includes a convolutional layer for preliminary processing of the latent representation, deep feature extraction is performed through a residual stack, in which each residual block contains two convolutional layers with ReLU activation function, an upsampling is performed through a deconvolutional layer, and a further feature extraction is performed through a convolutional layer with ReLU activation function, and then the latent representation is mapped to the space of the original image through another deconvolutional layer, both the image reconstruction decoder and the image restoration decoder include operations on input features of two different resolutions, the low-resolution quantized encoded feature map is subjected to upsampling and convolution operation and then feature extraction through the residual module, the high-resolution feature map is also subjected to upsampling and convolution operation, the two feature maps are concatenated together, and the final reconstructed image is generated through transposed convolution;
[0131] Step 3, calculate the high-frequency feature constraint loss, the specific steps are:
[0132] Step 301, construct a high-frequency feature constraint module: in the image reconstruction-based anomaly detection model, the higher the quality of the de-anomaly reconstructed image, the stronger the understanding and learning ability of the model to normal samples, which helps to improve the robustness and generalization ability of the model, making it more credible in actual application, when the model is used for anomaly detection task, it can distinguish normal and abnormal samples, and the way to enhance the reconstruction quality is to use Euclidean distance loss to directly constrain the consistency of the de-anomaly reconstructed image and the original image, since the priority of neural network fitting frequency information is different in the whole training process, it is usually from low to high, which means that part of the high-frequency information may not be valued by the network, and the detail texture information is just contained in the high-frequency information;
[0133] In order to narrow the gap between the reconstructed de-anomaly picture and the original picture in high-frequency information, a high-frequency feature constraint module is set, and discrete Fourier transform (DFT) is used to convert the original input picture and the reconstructed de-anomaly picture into frequency domain:
[0134]
[0135] In the above formula (3), F(u,v) is the representation of the image in the frequency domain, (u,v) represents the coordinates in the frequency domain, the size of the input picture is M*N, x(m,n) represents the pixel value of the (m,n) point in the picture, e and i are natural constant and imaginary unit respectively, as shown in the following formula (4):
[0136]
[0137] In the above formula (4), cos represents cosine, sin represents sine, (u, v) represents coordinates in the frequency domain, the size of the input picture is MxN, x(m, n) represents the pixel value of the (m, n) point in the picture, e and i are natural constants and imaginary units respectively;
[0138] Step 302, after the discrete Fourier transform, the amplitude and phase of the image are two key elements describing the frequency domain information, the phase describes the spatial position and relative position of the frequency component in the image, determines the starting point of the sine wave in the image, thereby affecting the spatial layout of different frequencies in the frequency domain, the amplitude represents the response strength or energy of the image to a specific frequency, in the centralized amplitude graph, the low frequency information is in the middle and the high frequency information is in the periphery, in order to design a loss function that constrains the loss of high frequency information, for the real part and imaginary part of F(u, v) in formula (3), let the real part R(u, v) = a, the imaginary part I(u, v) = b, then F(u, v) is expressed as the following formula (5):
[0139] F(u, v) = R(u, v) + I(u, v)i = a + bi …… (5),
[0140] In the above formula (5), (u, v) represents the coordinates in the frequency domain, R(u, v) and a represent the real part of the frequency domain information, I(u, v) and b represent the imaginary part of the frequency domain information, and i is the imaginary unit;
[0141] Step 303, calculate the amplitude of the image, as follows:
[0142]
[0143] In the above formula (6), (u, v) represents the coordinates in the frequency domain, R(u, v) and a represent the real part of the frequency domain information, I(u, v) and b represent the imaginary part of the frequency domain information, and |·| represents the modulus of a complex number;
[0144] Step 304, calculate the phase of the image according to the following formula (7):
[0145]
[0146] In the above formula (7), (u, v) represents the coordinates in the frequency domain, R(u, v) represents the real part of the frequency domain information, I(u, v) represents the imaginary part of the frequency domain information, and arctan is the tangent function, here the distance measurement is performed by mapping each frequency value to a Euclidean vector in two-dimensional space;
[0147] Step 305, calculate the spatial frequency value at the frequency spectrum coordinates (u, v) of the input image according to the following formula (8):
[0148] F in (u, v) = ain +b in i……(8),
[0149] In the above formula (8), a in represents the real part of the input image spectrum, b in represents the imaginary part of the input image spectrum, and i is the imaginary unit;
[0150] Step 306, calculate the spatial frequency value at the spectral coordinates (u, v) of the de-anomaly reconstructed image according to the following formula (9):
[0151] F gen (u, v) = a gen +b gen i……(9),
[0152] In the above formula (9), a gen represents the real part of the de-anomaly reconstructed image spectrum, b gen represents the imaginary part of the de-anomaly reconstructed image spectrum, and i is the imaginary unit;
[0153] Step 307, calculate the global distance of the frequency using the square of the Euclidean distance, as follows formula (10):
[0154]
[0155] In the above formula (10), and represent two vectors mapped from F in (u, v) and F gen (u, v), based on the definition of amplitude and phase, the size of the vector and correspond to the amplitude, and the angles θ in and θ gen correspond to the phase, ||·|| represents the Euclidean distance, and |·| represents the modulus of a complex number;
[0156] Step 308, calculate the frequency distance between the original input image and the de-anomaly reconstructed image, represented by the average value of the global distance, as follows formula (11):
[0157]
[0158] In the above formula (11), F in represents the input image spectrum, F gen represents the de-anomaly reconstructed image spectrum, (u, v) represents the coordinates in the frequency domain, the size of the input picture is M × N, and |·| represents the modulus of a complex number;
[0159] Step 309, in order to make the model more focused on learning key fine-grained features, filter out low-frequency information to reduce the model's attention to overall structural information, capture subtle textures and detailed features in the image, here introduce a mask M r , assuming that the center point coordinates of the image are (0, 0), then M r is of size M x N, the value of the central region rM x rN is 0, and the value of the remaining region is 1, where r ∈ (0, 1), the mask function is calculated as follows (12) :
[0160]
[0161] In the above formula (12), (u, v) represents the coordinates in the frequency domain, the size of the input picture is M x N, r is the area size ratio of the central region, is an indicator function, which is used to determine whether the frequency domain coordinates (u, v) are outside the specified central region, if outside the region, the mask function M r (u, v) takes the value 1; if inside the region, take the value 0;
[0162] Step 310, calculate the high-frequency feature constraint loss according to the following formula (13) :
[0163]
[0164] In the above formula (13), (u, v) represents the coordinates in the frequency domain, the size of the input picture is M x N, r is the area size ratio of the central region, F in represents the input image spectrum, F gen represents the abnormal reconstruction image spectrum;
[0165] Step 4, build a defect feature enhancement module, the specific steps are:
[0166] Step 401, build a defect feature enhancement module to respectively accept two inputs of the reconstructed original input image and the reconstructed abnormal image;
[0167] Step 402, calculate the difference image of the reconstructed original input image and the reconstructed abnormal image, the difference image is obtained by directly subtracting the first two, to highlight the difference area between the images, and realize the feature enhancement of the abnormal area;
[0168] Step 403, splice the reconstructed original input image, the reconstructed abnormal image and the difference image of the two, and input into the cloth defect anomaly detection module;
[0169] Step 5, build a cloth defect anomaly detection module, the specific steps are:
[0170] Step 501, constructing a cloth defect anomaly detection module: the function of the anomaly detection module is to make anomaly prediction for each pixel point, and present it in the form of anomaly score, indicating the difference degree between the input image and the normal mode. The anomaly detection module accepts three input images, including the reconstructed original input image, the reconstructed anomaly-removed image, and the difference image of the above two images. An anomaly score is predicted through a deep defect detection network. The deep defect detection network includes an encoder for capturing hierarchical features and a decoder for reconstructing high-resolution anomaly prediction. The encoder part is composed of four convolutional blocks, each of which includes two 3x3 convolutional layers for feature extraction. The output channel number of the first convolutional layer is consistent with the input, and the output channel number of the second convolutional layer is twice that of the first convolutional layer. Each convolutional layer is followed by an instance normalization layer and a ReLU activation function. Each pair of convolutional layers is followed by a 2x2 max-pooling layer for downsampling, gradually reducing the spatial size of the feature map while increasing the number of feature channels. The decoder part is composed of upsampling and convolutional blocks, which are used to gradually restore the spatial resolution of the image. Each upsampling step uses bilinear interpolation and is followed by a 3x3 convolutional layer to adjust the number of channels, and then a ReLU activation function. Each upsampling step is followed by a convolutional block, which includes two 3x3 convolutional layers for further feature refinement. In each upsampling step, the decoder also performs a skip connection with the feature map of the corresponding layer of the encoder. Finally, the output of the decoder is mapped to the required output channel number through a 1x1 convolutional layer.
[0171] Step 502, calculating defect detection loss: training by minimizing the anomaly detection loss function between the predicted score P and the anomaly region mask Y. The loss function is represented by the following formula (14):
[0172]
[0173] In the above formula (14), Y i is the true label of the i-th pixel point (0 represents normal or 1 represents anomaly), P i is the predicted value of the i-th pixel point, θ is a balance factor for balancing the proportion of positive and negative samples, and γ is an adjustment factor for adjusting the weight of difficult and easy samples.
[0174] Step 6, constructing an edge entropy minimization module, the specific steps are as follows:
[0175] Step 601: When performing pixel-level anomaly region prediction, the model will generate low-entropy predictions near the center of the anomaly region and high-entropy predictions at the edges of the anomaly region. Specifically, anomaly features are usually more prominent in the center of the anomaly region, and the model is more likely to learn the common features of these central regions, thus generating confident predictions about the central regions. However, when the model faces the edges of the anomaly region, due to the differences and complexity of the data distribution or insufficient anomaly information, the model cannot accurately capture the detailed information of the edges, resulting in low-confidence high-entropy predictions in these regions. Using only the supervision mode in step 502 above is still insufficient to enable the model to generate highly certain predictions for the anomaly edge regions. One way to improve the performance of the anomaly detection model in edge detection is to encourage the model to generate highly certain (low-entropy) predictions in the target prediction stage. Here, an edge entropy minimization module is set up to directly penalize low-confidence predictions using entropy loss.
[0176] The edge entropy minimization module does not consider using the generated abnormal region mask Y to improve the edge prediction confidence, but considers a more direct constraint method, namely, minimizing the prediction entropy. According to the Shannon entropy principle, for the input image x, its entropy map is composed of independent pixel-level entropy, and the entropy value of each pixel is as follows (15):
[0177]
[0178] In the above formula (15), This represents the model's outlier prediction for the location (m,n) in the image;
[0179] Step 602: To improve the confidence of the anomaly detection model in uncertain edge regions, a reweighting mechanism is introduced. Since the predicted values in uncertain edge regions are all close to 0.5, the distance between the absolute value of the prediction and 0.5 is calculated. Pixels closer to 0.5 are subject to greater loss supervision. For the (m,n) position of the image, a distance weight is introduced. The following formula (16):
[0180]
[0181] In the above formula (16), λ is a hyperparameter that controls the decay rate of the weights. The value range is between 0 and 1. The value of λ will directly affect the decay rate of the weights. The larger λ is, the faster the decay rate, and the weights of the predicted values that are farther away will decrease faster. The smaller λ is, the slower the decay rate, and the weights of the predicted values that are farther away will decrease relatively slowly.
[0182] Step 603: Entropy loss is defined as the sum of the entropies at all pixel levels, as shown in equation (17):
[0183]
[0184] In the above formula (17), represents the entropy value at the pixel point (m, n), represents the distance weight at the pixel point (m, n);
[0185] Step 7, the joint training module, the specific steps are:
[0186] Step 701, calculate the joint loss according to the following formula (18) to update the parameters of the image reconstruction decoder and the anomaly detector:
[0187]
[0188] In the above formula (18), is the anomaly detection loss, is the consistency loss of the de-anomaly reconstructed image and the original image, is the high-frequency feature constraint loss, is the entropy loss, and α and β are weight factors of and respectively;
[0189] Step 8, repeat steps 2 to 7 until the maximum iteration number I is reached or the model parameters reach convergence;
[0190] Step 9, actual use of the model, the specific steps are:
[0191] Step 901, for the input image with defects, as shown in Figure 3A , it is first passed through the feature encoder to obtain the feature with the abnormal area;
[0192] Step 902, the feature with the abnormal area is sent into the image reconstruction decoder and the image restoration decoder respectively, to reconstruct the de-anomaly and the image with the abnormal area respectively;
[0193] Step 903, the generated anomaly and de-anomaly images and the difference image between the two are spliced and input into the anomaly detector;
[0194] Step 904, the anomaly possibility of each pixel point is predicted, and the score of the pixel point in the final prediction result of the defect area will be closer to 1, and the score of the non-defect area will be closer to 0;
[0195] Step 905, for the input image without defects, as shown in Figure 3B , after the same inference process, the defect anomaly scores of the entire image area will be closer to 0.
[0196] The random generation method of the abnormal area of the method of the application is:
[0197] Different region shapes that can be generated on the image in the random generation mode of the abnormal region include rectangle, circle, triangle, ellipse, star, and in a reasonable region, the center point coordinates of the abnormal region are determined randomly, and the parameters required for generating the shape are determined randomly; taking the random generation of the elliptical region as an example, first, the width and height of the target image are obtained, then the center point coordinates (cx, cy) of the ellipse are randomly generated, wherein cx and cy are in the range of the width and height of the image respectively, then the length a and b of the major axis and the minor axis of the ellipse are randomly selected, and it is ensured that they are less than the corresponding size of the image, finally, the rotation angle of the ellipse is randomly selected, usually between 0 and 180 degrees, after all the parameters of the ellipse are determined, the corresponding elliptical region on the picture can be determined, in the actual training, the method described in the application randomly selects one of the two noise generation modes of random noise addition and quantization subspace noise addition with a probability of 0.5 to generate noise at the feature level.
[0198] The parameter setting of the model training of the method described in the application:
[0199] In the model initialization and model fine-tuning stage of step 1, the deep feature extractor, the image restoration decoder and the feature quantization codebook are trained for image reconstruction on the ImageNet dataset, a total of 200000 iterations are performed, the batch size is set to 32, and the learning rate is 2x10 -4 The codebook (Codebook) is composed of 4096 vectors with a length of 128 dimensions;
[0200] In the repeated training stage of step 8 in the summary section, the method described in the application performs a total of 5x10 4 iterations of training, the batch size is set to 8, and the learning rate is 2x10 -4 In addition, when the training is performed to 3.5x10 4 iterations, the learning rate is reduced to 2x10 -5 At the same time, r in formula (13) is 0.09, in formula (14) is 0.75, in formula (14) is 2, in formula (16) is 0.9, in formula (18) is 20, and in formula (18) is 10.
[0201] To demonstrate the performance of the method described in the present application on the task of fabric image defect detection, tests were conducted on actual production fabric image data, and the conventional index (Area Under the Receiver Operating Characteristic, AUROC) was used to test the performance of the model. The image-level AUROC is used to evaluate the abnormal detection ability of the model, and this index measures the ability of the model to distinguish between normal images and images containing defects at the image level, denoted as AUROC img ; the pixel-level AUROC is used to evaluate the abnormal positioning ability of the model, and this index measures the ability of the model to distinguish between normal pixels and defective pixels at the single pixel level, denoted as AUROC pix , to jointly evaluate the abnormal positioning performance of the model. Specifically:
[0202] True Positive (TP) refers to the number of samples correctly predicted as positive by the model. For example, in defect detection, if the model detects a defect and there is indeed a defect, this will be counted as TP.
[0203] False Negative (FN) refers to the number of positive samples that are incorrectly predicted as negative by the model. In the context of defect detection, if the model fails to detect an actual defect, this will be counted as FN.
[0204] False Positive (FP) refers to the number of negative samples that are incorrectly predicted as positive by the model. In defect detection, if the model incorrectly identifies a defect-free part as having a defect, this will be counted as FP.
[0205] True Negative (TN) refers to the number of samples correctly predicted as negative by the model, that is, when the model correctly identifies a defect-free part of the image, this will be counted as TN.
[0206] True Positive Rate (TPR) and False Positive Rate (FPR) are as follows:
[0207]
[0208] Precision and Recall are as follows:
[0209]
[0210] Recall = TPR,
[0211] Then the calculation method of AUROC img is:
[0212]
[0213] AUROC pix AUROC is calculated as follows:
[0214]
[0215] where N is the total number of pixels.
[0216] The fabric flaw detection result of the method in the actual production environment can be seen, the method can accurately identify the edge of the flaw when processing the industrial fabric image with complex texture and rich details, the overall detection precision is high, the consistency of the reconstructed image with the original image in the detail texture can be improved in the reconstruction process, the fineness of the reconstructed image is guaranteed, the accuracy of the fabric flaw anomaly detection, the reconstruction fineness and the edge area detection accuracy are comprehensively improved through the synergistic effect of each module, an efficient and reliable technical solution is provided for the industrial fabric quality control.
[0217] The present application is not limited to the above-mentioned embodiments, the above-mentioned embodiments and the description in the specification are only to illustrate the principles of the present application, various changes and improvements can be made without departing from the spirit and scope of the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims.
Claims
1. A fabric defect detection method based on high-frequency feature constraints and edge entropy minimization, comprising: Step 1. Collect a dataset of flawless and non-defective fabric training images, and initialize and fine-tune the hyperparameters and model parameters; Step 2. Randomly sample an image. The image features are obtained by vector quantization through the feature encoder, and defect noise is generated by the anomaly generator and added to the original features of the image. The features with anomalies are input into the image reconstruction decoder and the image restoration decoder respectively to reconstruct the images with and without anomalies. Step 3. Calculate the high-frequency feature constraint loss. Supervise the quality of the reconstructed anomaly-free image to ensure the quality of detailed texture feature reconstruction; Step 4. Calculate the difference image between the reconstructed original input image and the reconstructed anomaly-free image to enhance the features of the anomaly regions; Step 5. Input the reconstructed original input image, the reconstructed anomaly-free image, and the difference image between the two into the anomaly detector simultaneously, and calculate the anomaly detection loss. ; Step 6. Calculate the entropy loss using the edge entropy minimization module. Optimize the prediction of edge regions with low certainty to improve the confidence of the model in predicting edge regions; Step 7. Calculate the joint loss And propagate back to update the parameters; Step 8. Repeat steps 2 through 7 until the maximum number of iterations is reached. Or the model parameters have converged; Step 9. In the model usage phase, input the test image into the trained fabric defect anomaly detection model, perform defect detection on the test image, and select anomaly scores from the model output with a confidence level greater than a certain threshold. The result is considered an anomaly region.
2. The fabric defect detection method based on high-frequency feature constraints and edge entropy minimization according to claim 1, characterized in that, Step 1 involves collecting a dataset of flawless fabric training images, initializing hyperparameters, and fine-tuning model parameters. The specific steps are as follows: Step 101: In the actual working environment of a fabric textile factory, use a camera module to capture images of different types of flawless fabrics, so that each image can clearly show the appearance of the fabric. Step 102: Set the balance parameters and and maximum number of iterations ; Step 103: Initialize the model using a fully pre-trained deep feature extractor, image reconstruction decoder, and feature quantization codebook. Then, fine-tune the model parameters using data from the target dataset to minimize the image reconstruction loss and the feature space projection calculated by the deep feature extractor. Its codebook vectorization features The difference between them can be expressed by the loss function as: ……(1), In the above formula (1), This represents the feature space projection calculated by the deep feature extractor. This represents the corresponding codebook vectorization feature. This represents the input image that the image reconstruction decoder reconstructs. This indicates that gradient calculation has been stopped. express The norm is used to measure the distance between two corresponding points in the original image and the generated image. After the first stage of training, the feature encoder, image reconstruction decoder, and feature quantization codebook will be fixed.
3. The fabric defect detection method based on high-frequency feature constraints and edge entropy minimization according to claim 1, characterized in that, Step 2 involves randomly sampling an image sample. The image features are vector-quantized by the feature encoder, and defect noise is generated by the anomaly generator and added to the original image features. The features with anomalies are then input into the image reconstruction decoder and the image restoration decoder to reconstruct the images with and without anomalies, respectively. The specific steps are as follows: Step 201, Feature Extraction: Input Image Features are extracted by a deep feature extractor The deep feature extractor consists of 16 residual blocks. Each residual block contains two 3x3 convolutional layers with a stride of 1 and a 1x1 convolutional layer. Each residual block also contains a skip connection that directly adds the input of the block to the output. Step 202, Feature Quantization: To facilitate the generation of outlier region data that more closely approximates a normal distribution, vector quantization is used to encode the input image. Vector quantization is a technique that maps a continuous latent space to discrete features. By introducing a discrete codebook into the latent space, the model can more effectively learn high-level abstract representations of the data. By introducing a discrete codebook, vector quantization divides the latent space into a finite number of discrete regions, each represented by a specific codebook vector. Specifically, the features extracted by the deep feature extractor... The feature quantization process is represented as: ……(2), In the above formula (2), Indicates input image Features extracted by a deep feature extractor Representation of feature map Location characteristics, Represents a vector in the codebook, where It is the index of the codebook vector, and its value range is... ,in It is the number of features in the codebook. This represents the function reaching its minimum value. The value of , Representing the Euclidean distance, finally, each point in the feature space is replaced with the nearest neighbor vector in the codebook; Step 203: Improve the quality of reconstruction results by using two-level vector quantization: Encode the input image in two levels using low-resolution and high-resolution codebooks respectively. The quantized feature map of the low-resolution codebook is reduced by 4 times compared with the original input image, while the quantized feature map of the high-resolution codebook is reduced by 8 times. Through this two-level encoding method, the multi-level features of the input image can be captured better, resulting in higher quality reconstruction results. Step 204: Constructing a Fabric Feature Anomaly Generation Module: During the training phase, the model can only access input images without anomalies. Anomaly generation is required for the input images. Current noise generation methods are divided into image-level and feature-level. Image-level noise generation can simulate a limited number of anomalies, while feature-level noise generation is more robust. By introducing noise into specific features of the image, feature-level noise generation can simulate changes in abnormal features, increasing the diversity of anomaly information. While considering the randomness of the anomaly region, it allows the model to better distinguish anomalies that are difficult to differentiate. Anomaly generation combines two methods: random noise addition and quantized subspace noise addition. A mask for the anomaly region is generated by random sampling. In this method, regions with a value of 1 represent anomalies. Random noise addition adds generated random noise to the features of the selected region. The quantized subspace noise addition method is based on the quantized features. It randomly selects features from the feature space of the image data to replace the features of the selected region, while limiting the similarity between the replaced features and the original non-anomaly features, excluding the most similar 5% of vectors. Step 205: Constructing a fabric image decoder: In the image decoding stage, an image reconstruction decoder and an image restoration decoder are introduced respectively. The goal of the image reconstruction decoder is to restore the abnormal regions to the normal appearance of the samples observed during training, while the image restoration decoder directly restores the input features to the image. Step 206, For the image reconstruction decoder: The decoder extracts features through two convolutional blocks. The first convolutional block contains a 3x3 convolutional layer to maintain the number of input channels, followed by an instance normalization layer and a ReLU activation function. Then, another 3x3 convolutional layer doubles the number of channels, followed by another normalization layer and a ReLU activation function. The second convolutional block operates similarly to the first convolutional block. Each convolutional block is followed by a max-pooling layer, using a 2x2 window and a stride of 2 for downsampling. Next… The feature map is reduced to 64 channels by a 1x1 convolutional layer. Then, the feature map is passed through two upsampling blocks. Each upsampling block uses a 4x4 transposed convolution with a stride of 2 to perform the upsampling operation. After each upsampling block, a ReLU activation function is applied. After that, the feature map is passed through a 3x3 convolutional layer and then two transposed convolutional layers to gradually reduce the number of channels, which is used to reconstruct the RGB channels of the original image. After each transposed convolutional layer, except for the last layer, a ReLU activation function is applied. Finally, the network outputs the reconstructed image. Step 207: For the image reconstruction decoder: The decoder uses a fixed pre-trained model to directly reconstruct the input features. The decoder includes a convolutional layer for preliminary processing of the latent representation, and extracts deep features through residual stacking. In the residual stacking, each residual block contains two convolutional layers with ReLU activation functions. It is upsampled through a deconvolutional layer, and then further extracted through another convolutional layer with ReLU activation function. Finally, another deconvolutional layer maps the latent representation to the space of the original image. Both the image reconstruction decoder and the image reconstruction decoder include operations on two input features of different resolutions. The low-resolution quantized encoded feature map is upsampled and convolutionally processed before feature extraction through the residual module. The high-resolution feature map is also processed through upsampling and convolution. The two feature maps are concatenated together, and the final reconstructed image is generated through transposed convolution.
4. The fabric defect detection method based on high-frequency feature constraints and edge entropy minimization according to claim 1, characterized in that, Step 3 describes calculating the high-frequency feature constraint loss. The quality of the reconstructed anomaly-free image is supervised to ensure the quality of detailed texture feature reconstruction. The specific steps are as follows: Step 301: Construct a high-frequency feature constraint module: In an anomaly detection model based on image reconstruction, the higher the quality of the anomaly-reconstructed image, the stronger the model's understanding and learning ability of normal samples, which helps to improve the model's robustness and generalization ability, making it more reliable in practical applications. When the model is used for anomaly detection tasks, it can better distinguish between normal and abnormal samples. The way to enhance the reconstruction quality is to use Euclidean distance loss to directly constrain the consistency between the anomaly-reconstructed image and the original image. Since the neural network's fitting priority for frequency information is different throughout the training process, usually from low to high, this means that some high-frequency information is not valued by the network, while detailed texture information is included in the high-frequency information. To narrow the gap in high-frequency information between the reconstructed anomaly-free image and the original image, a high-frequency feature constraint module is set up, and the discrete Fourier transform is used to convert both the original input image and the reconstructed anomaly-free image to the frequency domain: ……(3), In the above formula (3), It is the frequency domain representation of the image. Represents the coordinates in the frequency domain, and the size of the input image is [value missing]. , Indicates the image pixel value of a point and Let be the natural constant and the imaginary unit, respectively, as shown in equation (4): ……(4), In equation (4) above, It represents the cosine. It represents the sine wave. Represents the coordinates in the frequency domain, and the size of the input image is [value missing]. , Indicates the image pixel value of a point and These are the natural constant and the imaginary unit, respectively. Step 302: After the discrete Fourier transform, the amplitude and phase of the image are two key elements describing the frequency domain information. The phase describes the spatial and relative positions of the frequency components in the image, determining the starting point of the sine wave in the image, thus affecting the spatial layout of different frequencies in the frequency domain. The amplitude represents the response intensity or energy of the image to a specific frequency. In the centered amplitude diagram, low-frequency information is in the middle while high-frequency information is around the edges. In order to design a loss function that constrains high-frequency information, the formula (3) is used. The real and imaginary parts, let the real part 6. Imaginary part ,but It can be expressed as the following formula (5): ……(5), In the above formula (5), Represents coordinates in the frequency domain. and The real part represents the frequency domain information. and Represents the imaginary part of the frequency domain information. The imaginary unit; Step 303: Calculate the amplitude of the image, as shown in equation (6): ……(6), In the above formula (6), Represents coordinates in the frequency domain. and The real part represents the frequency domain information. and Represents the imaginary part of the frequency domain information. This indicates finding the modulus of a complex number; Step 304: Calculate the phase of the image according to the following formula (7): ……(7), In the above formula (7), Represents coordinates in the frequency domain. The real part represents the frequency domain information. Represents the imaginary part of the frequency domain information. It is the tangent function, and the distance is measured here by mapping each frequency value to a Euclidean vector in two-dimensional space; Step 305: Calculate the spectral coordinates of the input image according to the following formula (8). Spatial frequency value at: ……(8), In the above formula (8), Represents the real part of the spectrum of the input image. Represents the imaginary part of the spectrum of the input image. The imaginary unit; Step 306: Calculate the spectral coordinates of the anomaly-reconstructed image according to the following formula (9). Spatial frequency value at: ……(9), In the above formula (9), This represents the real part of the spectrum of the anomaly-reconstructed image. This represents the imaginary part of the spectrum of the anomaly-reconstructed image. The imaginary unit; Step 307: Calculate the global distance of the frequency using the square of the Euclidean distance, as shown in equation (10): ……(10), In the above formula (10), and Indicates from and The two mapped vectors, based on the definitions of amplitude and phase, have different magnitudes. and Corresponding to amplitude, angle and Corresponding to phase, Represents Euclidean distance. This indicates finding the modulus of a complex number; Step 308: Calculate the frequency distance between the original input image and the anomaly-reconstructed image, expressed as the average of the global distances, as shown in equation (11): ……(11), In the above formula (11), Represents the spectrum of the input image. This represents the spectrum of the image after anomaly removal and reconstruction. Represents the coordinates in the frequency domain, and the size of the input image is [value missing]. , This indicates finding the modulus of a complex number; Step 309: To enable the model to focus more on learning key fine-grained features, low-frequency information is filtered out to reduce the model's attention to overall structural information and capture subtle textures and details in the image. A mask is introduced here. Assuming the center point coordinates of the image are ,but The size is its central area The value of is 0, and the value of the rest of the region is 1. The mask function is calculated as follows (12): ……(12), In the above formula (12), Represents the coordinates in the frequency domain, and the size of the input image is [value missing]. , The ratio of the area size of the central region. This is an indicator function used to determine frequency domain coordinates. Whether it is outside the specified center area; if it is outside the area, the mask function... The value is 1; if it is within the range, the value is 0. Step 310: Calculate the high-frequency feature constraint loss according to the following formula (13): ……(13), In the above formula (13), Represents the coordinates in the frequency domain, and the size of the input image is [value missing]. , The ratio of the area size of the central region. Represents the spectrum of the input image. This indicates the spectrum of the image after anomaly removal and reconstruction.
5. The fabric defect detection method based on high-frequency feature constraints and edge entropy minimization according to claim 1, characterized in that, Step 4 involves calculating the difference image between the reconstructed original input image and the reconstructed anomaly-free image to enhance the features of the anomaly regions. The specific steps are as follows: Step 401: Construct a defect feature enhancement module to accept two inputs: the reconstructed original input image and the reconstructed de-anomaly image. Step 402: Calculate the difference image between the reconstructed original input image and the reconstructed anomaly-removed image. The difference image is obtained by directly subtracting the two images to highlight the difference regions between the images and achieve feature enhancement of the anomaly regions. Step 403: The reconstructed original input image, the reconstructed anomaly-free image, and the difference image between the two are stitched together and input into the fabric defect anomaly detection module.
6. The fabric defect detection method based on high-frequency feature constraints and edge entropy minimization according to claim 1, characterized in that, Step 5 involves simultaneously inputting the reconstructed original input image, the reconstructed anomaly-free image, and the difference image between the two into the anomaly detector, and calculating the anomaly detection loss. The specific steps are as follows: Step 501: Construct a fabric defect anomaly detection module: The anomaly detection module's function is to generate anomaly predictions for each pixel, presenting them as anomaly scores representing the degree of difference between the input image and the normal pattern. The anomaly detection module accepts three input images: the reconstructed original input image, the reconstructed anomaly-free image, and the difference image between the two. Anomaly scores are predicted using a deep defect detection network. This deep defect detection network includes an encoder for capturing hierarchical features and a decoder for reconstructing high-resolution anomaly predictions. The encoder consists of four convolutional blocks, each containing two 3x3 convolutional layers for feature extraction. The first convolutional layer has the same number of output channels as the input, and the second convolutional layer has a different number of output channels than the first. The encoder consists of two convolutional layers, each followed by an instance normalization layer and a ReLU activation function. Each pair of convolutional layers is followed by a 2x2 max pooling layer for downsampling, progressively reducing the spatial size of the feature map while increasing the number of feature channels. The decoder consists of upsampling and convolutional blocks for progressively restoring the spatial resolution of the image. Each upsampling step uses bilinear interpolation followed by a 3x3 convolutional layer to adjust the number of channels, then a ReLU activation function. Each upsampling step is followed by a convolutional block containing two 3x3 convolutional layers for further feature refinement. In each upsampling step, the decoder also makes skip connections with the feature maps of the corresponding layers of the encoder. Finally, the decoder output is mapped to the desired number of output channels through a 1x1 convolutional layer. Step 502: Calculate the defect detection loss: by minimizing the prediction score. With anomaly region mask The anomaly detection loss function is used for training, and the loss function is expressed as follows (14): ……(14), In the above formula (14), It is the first The actual label of each pixel, where 0 indicates normal or 1 indicates abnormal. It is the first The predicted value of each pixel. It is a balancing factor used to balance the ratio of positive to negative samples. It is a moderating factor used to adjust the weights of easy and difficult samples.
7. The fabric defect detection method based on high-frequency feature constraints and edge entropy minimization according to claim 6, characterized in that, Step 6 describes calculating the entropy loss using the edge entropy minimization module. To optimize the prediction of edge regions with low certainty and improve the confidence of the model in predicting edge regions, the specific steps are as follows: Step 601: When performing pixel-level anomaly region prediction, the model will generate low-entropy predictions near the center of the anomaly region and high-entropy predictions at the edges of the anomaly region. Anomaly features are usually more prominent in the center of the anomaly region, and the model is more likely to learn the common features of these central regions, thus generating confident predictions for the central regions. However, when the model faces the edges of the anomaly region, due to the differences and complexity of the data distribution in this part or the lack of anomaly information, the model cannot accurately capture the detailed information of the edge part, resulting in low-confidence high-entropy predictions in these areas. This shows that using only the supervision mode of step 502 above is still not enough to make the model generate highly certain predictions for the edge regions of anomalies. One way to improve the performance of the anomaly detection model in edge detection is to encourage the model to generate highly certain predictions in the target prediction stage. Here, an edge entropy minimization module is set up, and the design uses entropy loss to directly penalize low-certainty predictions. Specifically, the edge entropy minimization module does not consider utilizing the generated anomalous region mask. To improve the confidence of edge prediction, a more direct constraint method is considered, namely, minimizing the prediction entropy. According to the Shannon entropy principle, for the input image... Its entropy map is composed of independent pixel-level entropy, and the entropy value of each pixel is as follows (15): ……(15), In the above formula (15), The model represents the image Anomaly prediction values for location; Step 602: To improve the confidence of the anomaly detection model in uncertain edge regions, a reweighting mechanism is introduced. Since the predicted values in uncertain edge regions are all close to 0.5, the distance between the absolute value of the prediction and 0.5 is calculated. Pixels closer to 0.5 are subject to greater loss supervision. Location, introducing distance weights : ……(16), In the above formula (16), It is a hyperparameter that controls the rate of weight decay, and its value ranges from 0 to 1. The value of directly affects the decay rate of the distance weight. The larger the value, the faster the decay rate; the weight of predictions from farther distances decreases even more rapidly. The smaller the value, the slower the decay rate; the weight of predictions that are farther away decreases relatively slowly. Step 603: Entropy loss is defined as the sum of the entropies at all pixel levels, as shown in equation (17): ……(17), In the above formula (17), Represents pixels The entropy value at that point, Represents pixels Distance weight at each location.
8. The fabric defect detection method based on high-frequency feature constraints and edge entropy minimization according to claim 1, characterized in that, Step 7 describes the calculation of joint loss. And backpropagate to update parameters, the specific steps are as follows: Step 701: Calculate the combined loss according to the following formula (18). This is used to update the parameters of the image reconstruction decoder and anomaly detector. ……(18), In the above formula (18), It is anomaly detection loss. It is the consistency loss between the anomaly-reconstructed image and the original image. It is a high-frequency feature constraint loss. It is entropy loss. and They are and Weighting factors.
9. The fabric defect detection method based on high-frequency feature constraints and edge entropy minimization according to claim 1, characterized in that, In step 9, during the model usage phase, the test image is input into the trained fabric defect anomaly detection model to detect defects in the test image. The model then selects anomaly scores from the outputs with a confidence level greater than a certain threshold. The specific steps for identifying the abnormal region are as follows: Step 901: For a flawed input image, it first passes through a feature encoder to obtain features with abnormal regions; Step 902: The features with abnormal regions are fed into the image reconstruction decoder and the image restoration decoder respectively to reconstruct images with and without abnormalities. Step 903: Stitch together the generated images with / without anomalies and the difference images between the two, and input them into the anomaly detector; Step 904: Predict the probability of anomalies for each pixel. The final prediction result will show that the score of the pixel in the defective area will be close to 1, and the score of the pixel in the defect-free area will be close to 0. Step 905: For an input image without defects, after the same reasoning process, the defect anomaly score of the entire image region will be close to 0.
Citation Information
Patent Citations
Cloth defect detection method based on One-Class depth support vector description
CN111709907A
Fabric flaw line defect identification method for textile production
CN115311264A
A rapid method for detecting textile defects based on image features
CN116934749B
Fabric defect detection method based on convolutional neural network and repetitive pattern analysis
CN117237946A
Machine vision identification method for warp and weft flaws of fabric
CN115311279A