Industrial flaw detection method based on frequency domain mask strategy

By adopting a self-supervised learning method based on frequency domain masking strategy in industrial defect detection, and using self-supervised pre-training and time-frequency consistency loss function, the problem of insufficient dependence and robustness of labeled data in the prior art is solved, and efficient and robust defect detection is achieved.

CN120088193AActive Publication Date: 2025-06-03NANJING UNIV OF SCI & TECH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510006288.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-06-03
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

Existing industrial defect detection methods rely on a large amount of labeled data and lack robustness for lighting and viewing angle changes, making it difficult to effectively detect minor defects on the product surface.

Method used

Using a self-supervised learning method based on frequency domain masking strategy, by performing occlusion processing in the spatial domain and frequency domain, a shared encoder and encoder decoder branch is constructed, and robust features are extracted using Vision Transformer, and pre-trained through the time-frequency consistency loss function, and finally connected to the detection network for defect detection.

Benefits of technology

It significantly reduces the labeling cost, improves the sensitivity to surface defects, enhances the robustness and generalization capabilities of the model, and can more accurately detect subtle defects on the product surface.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088193A_ABST
    Figure CN120088193A_ABST
Patent Text Reader

Abstract

The invention discloses a self-supervised industrial flaw detection method based on a frequency domain mask strategy, and the method comprises the steps: constructing a shared encoder based on a Vision Transform, carrying out the feature extraction of a covered industrial product surface image and a covered industrial product surface spectrogram, and obtaining a spatial domain shared coding feature and a frequency domain shared coding feature; two independent encoder and decoder branches are constructed, the two independent encoder and decoder branches comprise a spatial domain branch and a frequency domain branch, each branch takes a Transform block as an encoder and a convolutional layer as a decoder, and a spatial domain reconstructed image and a frequency domain reconstructed image are obtained; based on a pre-trained shared encoder and an encoder, connecting a detection network, constructing a self-supervised industrial flaw detection model, splicing the outputs of the encoders of a spatial domain branch and a frequency domain branch, inputting the spliced encoders into the detection network to identify a flaw area, and performing model fine tuning by using a labeled sample; and carrying out industrial flaw detection by using the adjusted self-supervised industrial flaw detection model. According to the method, the marking cost is greatly reduced while the industrial defect detection performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to industrial defect detection technology, and particularly to an industrial defect detection method based on a frequency domain masking strategy. Background Art

[0002] In industrial production, the detection and control of product surface quality are crucial. Defect detection is a key link to ensure product quality compliance, improve production efficiency, and reduce costs. Traditional defect detection mainly relies on manual visual inspection, which has problems such as low efficiency, high omission rate, and large interference from subjective factors. To overcome the above limitations, automated defect detection methods based on machine vision have received increasing attention.

[0003] In recent years, deep learning technologies represented by convolutional neural networks have made remarkable progress in the field of industrial vision detection. However, existing deep learning methods are mostly trained under the supervision of a large amount of labeled data, which requires high human and time costs. In industrial scenarios, large-scale defect annotation samples are often lacking, severely restricting the application of these methods. At the same time, factors such as the material, texture, and illumination of the industrial product surface are complex and variable, posing higher requirements for the robustness and generalization ability of the detection algorithm.

[0004] Facing these challenges, there is an urgent need for a new method that can make full use of unlabeled data and automatically learn robust feature representations from surface images. In recent years, self-supervised learning methods have provided new ideas for solving the problem of scarce labeled samples. Self-supervised learning designs prediction tasks, enabling the model to learn general visual representations from a large amount of unlabeled data. On this basis, fine-tuning with a small amount of labeled data can achieve good results in downstream tasks.

[0005] However, existing self-supervised learning methods mainly focus on constructing prediction tasks through various transformations in the spatial domain, such as image rotation, puzzle restoration, etc. The features learned by such methods lack robustness to changes in illumination, perspective, etc., and are difficult to directly apply to industrial detection scenarios. At the same time, existing methods do not pay enough attention to image details and textures, and the extracted features have limited sensitivity to surface micro-defects. Summary of the Invention

[0006] The object of the present invention is to propose an industrial defect detection method based on a frequency domain masking strategy.

[0007] The technical solution for realizing the object of the present invention is: A self-supervised industrial defect detection method based on a frequency domain masking strategy, comprising the following steps:

[0008] Step 1, obtain an industrial product surface image, perform masking processing on the industrial product surface image in both the spatial domain and the frequency domain to obtain a masked industrial product surface image and a masked industrial product surface frequency spectrum image;

[0009] Step 2: Build a shared encoder based on Vision Transformer to extract features from the masked industrial product surface image and the masked industrial product surface spectrogram, obtaining the spatial domain shared coding features and the frequency domain shared coding features;

[0010] Step 3: Build two independent encoder-decoder branches, including a spatial domain branch and a frequency domain branch. Each branch uses a Transformer block as the encoder and a convolutional layer as the decoder. Among them, the spatial domain branch reconstructs the spatial domain shared coding features to obtain a spatial domain reconstructed image, and the frequency domain branch reconstructs the frequency domain shared coding features to obtain a frequency domain reconstructed image;

[0011] Step 4: Build a time-frequency consistency loss function based on the cosine similarity loss, the frequency domain recovery loss, and the spatial domain recovery loss to complete the pre-training of the shared encoder and the encoder-decoder branches;

[0012] Step 5: Based on the pre-trained shared encoder and encoder, connect a detection network to build a self-supervised industrial defect detection model. Concatenate the encoder outputs of the spatial domain branch and the frequency domain branch, input them into the detection network to identify the defect area, and use the labeled samples to fine-tune the model;

[0013] Step 6: Collect the industrial product surface images in both the spatial domain and the frequency domain, and use the adjusted self-supervised industrial defect detection model to perform industrial defect detection.

[0014] Furthermore, in Step 1, obtain the industrial product surface image, and perform masking processing on the industrial product surface image in both the spatial domain and the frequency domain to obtain the masked industrial product surface image and the masked industrial product surface spectrogram. The specific method is as follows:

[0015] Step 1.1: Obtain the industrial product surface image X space ;

[0016] Step 1.2: Process the industrial product surface image X space by masking both the spatial domain and the frequency domain to more efficiently learn the feature representation, where:

[0017] In the spatial domain, simulate the situation of surface defects by occluding some pixel regions of the image to obtain the masked industrial product surface image;

[0018] In the frequency domain, first perform a discrete cosine transform on the industrial product surface image X space :

[0019] X freq = DCT(Xspace )

[0020] Among them, X freq is the frequency-domain representation of the surface image of the industrial product;

[0021] Then, perform a masking process on the frequency-domain representation of the surface image of the industrial product:

[0022] X masked = X freq ⊙M

[0023] Among them, X masked is the surface spectrum diagram of the industrial product after masking, M represents the dynamically generated masking mask, and ⊙ is the point-by-point multiplication operation.

[0024] Furthermore, the mask M is generated based on low-frequency and high-frequency separation masking, random frequency masking, or energy-based adaptive masking.

[0025] Furthermore, in step 2, construct a shared encoder based on Vision Transformer, and perform feature extraction on the masked surface image of the industrial product and the masked surface spectrum diagram of the industrial product to obtain the spatial-domain shared coding features and the frequency-domain shared coding features. The specific method is as follows:

[0026] Both the masked surface image of the industrial product and the masked surface spectrum diagram of the industrial product are sliced into N token image patches of a fixed size, and all these image patches are input into a linear projection layer (Linear Projection Layer). The corresponding output of each image patch is a tensor of length D. Then, the N token outputs of length D of these two pictures are concatenated respectively to generate the initial feature representation H 1 of the masked surface image of the industrial product and the initial feature representation H 2 of the masked surface spectrum diagram of the industrial product. Their specifications are both N token ×D. Then, further add these two feature representations H 1 and H 2 to the positional encoding (Positional Encoding) to respectively obtain the positional information of the masked surface image of the industrial product in the spatial domain and the positional information of the masked surface spectrum diagram of the industrial product in the frequency domain. Then, the feature representations H 1 and H 2 are respectively input into the shared encoder;

[0027] In the shared encoder, the feature representations H 1 and H 2 both sequentially pass through N shareA Transformer block, specifically, the feature representation H 1 is input to the first Transformer block of the shared encoder to obtain the output of this Transformer block, and then this output is input to the second Transformer block of the shared encoder to obtain the output of this Transformer block; the feature representations H 1 and H 2 at the output of the N share -th (i.e., the last) Transformer block are H space and H freq respectively, with the specification of N token ×D.

[0028] Furthermore, in step 3, two independent encoder-decoder branches are constructed, including a spatial domain branch and a frequency domain branch. Each branch uses a Transformer block as the encoder and a convolutional layer as the decoder. The specific method is as follows:

[0029] The spatial domain branch is used for image reconstruction of the output H space of the shared encoder. In this branch, after inputting H space into the encoder of N branch Transformer blocks and a one-layer convolutional layer decoder, the output is the reconstructed image of the industrial product surface image

[0030] The frequency domain branch is used for frequency domain reconstruction of the output H freq of the shared encoder. In this branch, after inputting H freq into the encoder of N branch Transformer blocks and the final decoder, the output is the reconstructed frequency spectrum of the industrial product surface frequency spectrum

[0031] Furthermore, in step 4, a time-frequency consistency loss function is constructed based on the cosine similarity loss, the frequency domain recovery loss, and the spatial domain recovery loss to complete the training of the encoder-decoder branches. The specific method is as follows:

[0032] The spatial domain recovery loss function is introduced with the goal of minimizing the pixel difference between the reconstructed image of the industrial product surface image and the industrial product surface image X space so as to reconstruct an image closer to the original surface in the spatial domain. The spatial domain recovery loss function is:

[0033]

[0034] where M spaceis a set of indexes of image patches whose space domain is occluded, which can be obtained when occluding the surface image of industrial products. is the i-th image patch corresponding tensor of the reconstructed image of the industrial product surface image, X space,i is the i-th image patch corresponding tensor of the industrial product surface image, N space is the number of occlusion patches;

[0035] Introduce a frequency domain recovery loss function to minimize the reconstructed spectrogram of the industrial product surface spectrogram and the industrial product surface spectrogram X freq The difference in DCT coefficients between them is the goal, so that an image closer to the original surface is reconstructed in the frequency domain. The frequency domain recovery loss function is:

[0036]

[0037] Introduce a cosine similarity loss between the output H space of the spatial domain branch and the output H freq of the frequency domain branch to enhance the consistency of the spatial domain and frequency domain feature representations. The cosine similarity loss is:

[0038]

[0039] where, N token is the total number of image patches, whose meaning is as described above, H space,i and H freq,i respectively represent the i-th image patch corresponding tensors of the output H space of the spatial domain branch and the output H freq of the frequency domain branch;

[0040] Construct the overall loss function, expressed as:

[0041] L total =λ 1 L space +λ 2 L freq +λ 3 L con

[0042] where λ 1 , λ 2 , λ 3 are weight coefficients. By minimizing the overall loss function, the pre-training process of the shared encoder and the spatial domain branch and the frequency domain branch is completed.

[0043] Further, in step 5, based on the pre-trained shared encoder and the encoder, and then connecting the detection network, a self-supervised industrial defect detection model is constructed. The encoder outputs of the spatial domain branch and the frequency domain branch are concatenated and input into the detection network to identify the defect area, and the labeled samples are used for model fine-tuning, where:

[0044] The detection network adopts Faster R-CNN or SSD.

[0045] A self-supervised industrial defect detection system based on a frequency domain masking strategy implements the self-supervised industrial defect detection method based on the frequency domain masking strategy to achieve self-supervised industrial defect detection based on the frequency domain masking strategy.

[0046] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the self-supervised industrial defect detection method based on the frequency domain masking strategy is implemented to achieve self-supervised industrial defect detection based on the frequency domain masking strategy.

[0047] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the self-supervised industrial defect detection method based on the frequency domain masking strategy is implemented to achieve self-supervised industrial defect detection based on the frequency domain masking strategy.

[0048] Compared with the prior art, the significant advantages of the present invention are:

[0049] (1) Effectively reduce the annotation cost. The present invention constructs a self-supervised pre-training task through spatial domain and frequency domain masking, makes full use of a large number of unlabeled product surface images to learn robust feature representations. On this basis, only a small amount of labeled data is required for fine-tuning to obtain an excellent performance defect detection model. Compared with the traditional method that requires large-scale manual annotation, the present invention can significantly save annotation time and labor costs.

[0050] (2) Improve the sensitivity to surface defects. The present invention randomly masks the image in the frequency domain, and the model pays more attention to fine-grained surface features such as texture and edges. At the same time, the time-frequency consistency constraint further enhances the complementarity of spatial domain and frequency domain features, which enables the model to more accurately capture the subtle defects and abnormalities on the product surface, and greatly improves the detection accuracy and recall rate.

[0051] (3) Enhance the robustness and generalization ability of the model. The present invention uses the Vision Transformer architecture to model the global context, enabling the model to better understand the association between defects and the surrounding areas. The pre-training based on frequency domain masking makes the model more robust to changes such as illumination and noise. In addition, the surface features learned through unsupervised pre-training have better transferability, enabling the model to better generalize to different types of industrial product detection tasks. Description of the Drawings

[0052] Figure 1 It is the overall flowchart of pre-training and fine-tuning in the embodiment of the present invention.

[0053] Figure 2 It is an example of low-frequency and high-frequency separation masking proposed in the embodiment of the present invention.

[0054] Figure 3 It is the schematic diagram of the model with a shared encoder and a dual-branch structure proposed in the embodiment of the present invention. Detailed Embodiment

[0055] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0056] An industrial defect detection method based on a frequency domain masking strategy, which randomly occludes the image spectrum to enable the model to learn finer-grained and more global surface features. At the same time, a time-frequency consistency constraint is introduced to enhance the complementarity of spatial domain and frequency domain features. On this basis, the Transformer architecture is used to model the global context to further improve the semantic expression ability of features.

[0057] As Figure 1 shown, the self-supervised industrial defect detection method based on the frequency domain masking strategy mainly includes the following steps: 1. Masking process; 2. Input to the shared encoder; 3. Dual-branch feature recovery; 4. Time-frequency consistency loss training; 5. Fine-tuning the model.

[0058] Step 1, Masking process:

[0059] Obtain the surface image of the industrial product, and perform masking processing on both the spatial domain and the frequency domain to obtain the masked surface image of the industrial product and the masked surface spectrum image of the industrial product; the specific process is as follows:

[0060] Step 1.1, Obtain the surface image X of the industrial product space , and these image data can be the surface image acquisition of various types of products on the production line, such as metal parts, circuit boards, fabrics, glass products, etc.

[0061] Step 1.2, process the industrial product surface image x by masking in both the spatial domain and the frequency domain to more efficiently learn the feature representation, where: space For more efficient learning of feature representations,

[0062] (1) Spatial domain

[0063] In the spatial domain, simulate the situation of surface defects by masking some pixel regions of the image to obtain the masked industrial product surface image;

[0064] (2) Frequency domain

[0065] In the frequency domain, first perform a discrete cosine transform on the industrial product surface image x space :

[0066] x freq = DCT(x space )

[0067] where x freq is the frequency domain representation of the industrial product surface image;

[0068] Then, perform a masking process on the frequency domain representation of the industrial product surface image:

[0069] X masked = X freq ⊙ M

[0070] where X masked is the masked industrial product surface spectrogram, M represents a dynamically generated masking mask, and ⊙ is the element-wise multiplication operation;

[0071] In the above description, the mask M can be designed based on low-frequency and high-frequency separation masking, random frequency masking, or energy-based adaptive masking, where:

[0072] A. Low-frequency and high-frequency separation masking: Implement different masking strategies for the low-frequency and high-frequency parts of the frequency domain representation of the industrial product surface image, and respectively focus on the macroscopic texture features and microscopic defect features of the surface; Figure 2 shows an example of applying low-frequency and high-frequency separation masking to the DCT spectrogram.

[0073] The left figure is the original spectrogram, and the right figure is the masked spectrogram. In the Figure 2 right figure, the upper left region of the spectrogram (i.e., the region enclosed by a square with a side length of 28 pixels) is defined as the low-frequency region, and the outside is defined as the high-frequency region, where 90% and 80% of the low-frequency region and the high-frequency region are randomly masked respectively, and masks of sizes 2×2 and 14×14 are applied to the former and the latter respectively.

[0074] B. Random frequency masking. On the frequency-domain representation X of the industrial product surface image freq set the masking ratio to 75%, and use a 14×14 masking block to randomly mask the frequency-domain representation X freq ;

[0075] C. Energy-based adaptive masking. Calculate the energy distribution map E of the frequency-domain representation X of the industrial product surface image, where E(i,j) represents the energy value at the frequency component (i,j); calculate the total energy E freq of the frequency-domain representation X freq : total

[0076]

[0077] where H and W represent the height and width of the frequency-domain representation X freq respectively; according to the energy distribution map E and the total energy E total , calculate the masking probability P mask (i,j) of each frequency component (i,j):

[0078]

[0079] Generate a mask for the frequency-domain representation X according to the masking probability P mask (i,j), and set the masking ratio to 30%. freq

[0080] Step 2. Input the shared encoder:

[0081] Construct a shared encoder based on Vision Transformer to extract features from the masked industrial product surface image and the masked industrial product surface spectrogram, and obtain the spatial-domain shared encoding features and the frequency-domain shared encoding features; the specific method is as follows:

[0082] The shared encoder consists of N share Transformer blocks, and its main function is to uniformly map the masked image data in the spatial domain and the frequency domain to a high-dimensional feature space, so as to provide a consistent input representation for subsequent two-branch processing.

[0083] Suppose both the masked industrial product surface image and the masked industrial product surface spectrogram are sliced into N token image blocks of a fixed size, and then all these image blocks are input into a linear projection layer (Linear ProjectionLayer). The corresponding output of each image block is a tensor of length D. Then, the N token outputs of length D of these two pictures are concatenated to generate the initial feature representation H of the masked industrial product surface image​​1 and the initial feature representation H of the spectrum diagram of the industrial product surface after masking 2 , both of which have the specification of N token ×D. Then, the two feature representations H 1 and H 2 are added to the positional encoding respectively to obtain the position information of the masked industrial product surface image in the spatial domain and the position information of the masked industrial product surface spectrum diagram in the frequency domain. Then, the feature representations H 1 and H 2 are respectively input into the shared encoder.

[0084] In the shared encoder, the feature representations H 1 and H 2 both sequentially pass through N share Transformer blocks. Specifically, the feature representation H 1 is input to the first Transformer block of the shared encoder to obtain the output of this Transformer block, and then this output is input to the second Transformer block of the shared encoder and the output of this Transformer block is obtained, and so on. The operation for the feature representation H 2 is the same. Let the outputs of the feature representations H 1 and H 2 in the N share th (i.e., the last) Transformer block be H space and H freq respectively, and their specifications are also N token ×D.

[0085] Step 3, dual-branch feature recovery:

[0086] After the shared encoder, the present invention designs two independent encoder-decoder branches, which are respectively used to model the features in the spatial domain and the frequency domain. Each branch contains N branch Transformer blocks as the encoder and one convolutional layer as the decoder.

[0087] The dual-branch structure is as shown in Figure 3 . In the dual-branch structure, the spatial domain branch is used to reconstruct the image of the output H space of the shared encoder. In this branch, after inputting H space into the encoder of N branch Transformer blocks and the decoder of one convolutional layer, the output is the reconstructed image of the industrial product surface image The present invention introduces a spatial domain recovery loss function in the spatial domain branch. This loss function aims to minimize the pixel difference between the reconstructed image of the industrial product surface image and the industrial product surface image X space so that the model can reconstruct an image closer to the original surface in the spatial domain. This loss function is defined as:

[0088]

[0089] where M space is the set of indices of the image patches masked in the spatial domain, which can be obtained when masking the industrial product surface image. is the i-th image patch corresponding tensor of the reconstructed image of the industrial product surface image, X space,i is the i-th image patch corresponding tensor of the industrial product surface image, and N space is the number of masking patches.

[0090] In the specific program implementation, let x represent the original unmasked image, xspace represent the spatial domain reconstructed image, and mask represent the spatial masking mask. First, calculate the interpolation of x and xspace pixel by pixel, then square the interpolation to obtain the pixel-level reconstruction error (x - xspace)^2. Next, multiply the pixel-level reconstruction error element-wise with mask, that is, mask*(x - xspace)^2, so as to set the error at the unmasked positions to zero and retain the error values at the masked positions. Finally, take the mean of all elements in the weighted reconstruction error tensor, that is, mean(mask*(x - xspace)^2), to obtain a scalar loss function value.

[0091] On the other hand, the frequency domain branch is used to perform frequency domain reconstruction on the output H freq of the shared encoder. In this branch, H freq is input into the encoder of N branch Transformer blocks and the final decoder, and the output is the reconstructed frequency spectrum of the industrial product surface frequency spectrum The present invention designs a frequency domain recovery loss function. This loss function aims to minimize the DCT coefficient difference between the reconstructed frequency spectrum of the industrial product surface frequency spectrum and the industrial product surface frequency spectrum X freq and is defined as:

[0092]

[0093] In the specific program implementation, the calculation of the frequency-domain reconstruction loss is similar to that in the spatial domain. First, the discrete cosine transform (DCT) is performed on the original image x and the frequency-domain reconstructed image x_freq, denoted as dct(x) and dct(x_freq) respectively, to map them into the frequency-domain representation, denoted as x_dct and x_freq_dct. Then, the difference between x_dct and x_freq_dct is calculated for each frequency component, i.e., x_dct - x_freq_dct, and then the square of the difference is taken and the mean is calculated, i.e., mean((x_dct - x_freq_dct)^2), to obtain the frequency-domain reconstruction loss.

[0094] Step 4: Training of the time-frequency consistency loss:

[0095] To enhance the consistency between the spatial-domain and frequency-domain feature representations, a cosine similarity loss is imposed between the output H space of the spatial-domain branch and the output H freq of the frequency-domain branch, which is defined as:

[0096]

[0097] where N token is the total number of image patches, whose meaning is as described above, and H space,i and H freq,i represent the tensors corresponding to the i-th image patch of the output H space of the spatial-domain branch and the output H freq of the frequency-domain branch, respectively.

[0098] In the specific program implementation, first, L2 normalization is performed on the spatial feature H_space and H_freq, i.e., the operations H_space / norm(H_space) and H_freq / norm(H_freq) are executed to obtain the results H_space_norm and H_freq_norm. Then, the inner product of the two normalized features is calculated, i.e., H_space_norm' * H_freq_norm, and the mean of all inner product results in the batch is taken and a negative sign is added, i.e., -mean(H_space_norm' * H_freq_norm), to obtain the final cosine similarity loss.

[0099] In summary, the overall loss function of this method is named the time-frequency consistency loss, which is defined as the weighted sum of three losses:

[0100] L total = λ 1 L space + λ 2 L freq + λ 3 L con

[0101] where λ1 , λ 2 , λ 3 is the weight coefficient. By minimizing the above formula, the pre-training process of the backbone model (i.e., the shared encoder and the spatial domain branch and frequency domain branch) is completed.

[0102] Step 5, fine-tune the model:

[0103] In the fine-tuning stage, the decoder part (i.e., the convolutional layer) of the double-branch structure is deleted from the program definition, and only the encoder of the double-branch is retained. The output H space of the spatial domain branch and the output H freq of the frequency domain branch are concatenated. The dimension of the concatenation result is 2×N token ×D. Next, according to the specific detection task objective, a neural network suitable for the detection task (such as Faster R-CNN, SSD, etc.) is selected, the above concatenation result is input into this neural network, and end-to-end supervised training is performed using a small number of labeled samples, so that the backbone model and the neural network suitable for the detection task are fine-tuned and the weights are updated according to the requirements of this defect detection data.

[0104] Embodiment

[0105] In order to verify the effectiveness of the method of the present invention in industrial defect detection tasks, experiments were carried out on the NEU-DET surface defect dataset. This dataset contains defect images of 6 types of material surfaces: metal surface, tile surface, glass surface, leather surface, fabric surface and wood surface. Each type of material surface has multiple common defect types in actual production. A 5-way 1-shot defect detection experiment was carried out on this dataset, that is, 5 defect categories were randomly selected, and 1 labeled sample was selected from each category for fine-tuning training.

[0106] The method of the present invention first performs spatial domain and frequency domain masking pre-training on 5000 unlabeled surface images of NEU-DET. The frequency domain masking includes three strategies: low-frequency and high-frequency separation masking, random frequency masking, and energy-based adaptive masking. The setting of hyperparameters is determined through multiple ablation experiments. Among them, the total number of layers of the encoder Transformer is 12, the first 8 layers are retained as the shared encoder in the pre-training stage, the frequency domain and spatial domain branches each consist of 4 Transformer layers, the total number of training epochs during pre-training is 300, and the weights of the loss function are respectively set as λ 1 = 0.25, λ 2 = 0.25, λ 3 = 0.50.

[0107] After the pre-training is completed, it enters the detector fine-tuning stage. SSD is selected as the backbone network of the multi-scale surface defect detector, and the model is fine-tuned for 10 rounds using 5-way 1-shot few-shot labeled data. It is compared with the fully supervised SSD, the SSD pre-trained directly on ImageNet, and the MAE pre-trained model without using the frequency domain masking strategy. The results are shown in Table 1. The frequency domain masking self-supervised pre-trained model of the present invention has achieved the best results in the defect detection MAP metrics of 6 materials, fully demonstrating the effectiveness and generalization ability of this method.

[0108] Table 1 Comparison of MAP performance of different surface defect detection methods in the NEU-DET 5-way 1-shot detection task

[0109] Method Metal surface Tile surface Glass surface Leather surface Fabric surface Wood surface Average MAP SSD 42.5 25.1 37.8 31.6 48.5 37.3 37.1 SSD-ImgNet 62.7 33.9 51.4 41.9 57.8 45.6 48.9 SSD-MAE 75.4 44.8 67.9 50.3 70.6 60.1 61.5 SSD-Ours 80.1 49.5 73.6 55.8 76.3 66.4 66.9

[0110] In summary, the self-supervised pre-training method based on frequency domain masking of the present invention provides a new idea for industrial product surface defect detection. Through the random masking operation of spatial domain and frequency domain images, the model is trained to learn finer-grained surface texture features. At the same time, the time-frequency consistency learning further improves the complementarity and consistency of spatial domain and frequency domain features. A large number of experiments show that the present invention has achieved excellent performance in various material surface defect detection tasks, greatly improving the detection accuracy, and at the same time significantly reducing the dependence on large-scale labeled data, and has high industrial application value. This method can also be extended and applied to other fine-grained surface analysis fields.

[0111] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0112] The above-described embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A self-supervised industrial defect detection method based on frequency domain masking strategy, characterized in that: The steps include: Step 1, obtaining an industrial product surface image, and performing masking processing on the industrial product surface image in two dimensions, namely, the spatial domain and the frequency domain, to obtain a masked industrial product surface image and a masked industrial product surface spectrum diagram; Step 2: construct a shared encoder based on Vision Transformer, extract features from the masked industrial product surface image and the masked industrial product surface spectrum, and obtain spatial domain shared coding features and frequency domain shared coding features; Step 3: construct two independent encoder-decoder branches, including a spatial domain branch and a frequency domain branch. Each branch consists of a Transformer block as an encoder and a convolutional layer as a decoder. The spatial domain branch reconstructs the shared coding features in the spatial domain to obtain a spatial domain reconstructed image, and the frequency domain branch reconstructs the shared coding features in the frequency domain to obtain a frequency domain reconstructed image. Step 4: construct a time-frequency consistency loss function based on cosine similarity loss, frequency domain recovery loss, and spatial domain recovery loss to complete the pre-training of the shared encoder and encoder-decoder branches; Step 5: Based on the pre-trained shared encoder and encoder, the detection network is connected to build a self-supervised industrial defect detection model, the encoder outputs of the spatial domain branch and the frequency domain branch are spliced, input into the detection network to identify the defect area, and the labeled samples are used to fine-tune the model; Step 6: Collect the surface images of the industrial products to be tested in the spatial domain and the frequency domain, and use the adjusted self-supervised industrial defect detection model to perform industrial defect detection.

2. The self-supervised industrial defect detection method based on frequency domain masking strategy according to claim 1 is characterized in that: Step 1: Obtain the surface image of the industrial product, and perform masking processing on the surface image of the industrial product in two dimensions, namely, the spatial domain and the frequency domain, to obtain the masked surface image of the industrial product and the masked surface spectrum of the industrial product. The specific method is as follows: Step 1.1, obtain the industrial product surface image X space ; Step 1.2: Mask the surface image X of the industrial product by masking the spatial domain and the frequency domain. space is processed to learn feature expressions more efficiently, where: In the spatial domain, by blocking part of the pixel area of ​​the image, the surface defects are simulated to obtain the masked surface image of the industrial product; In the frequency domain, we first analyze the surface image H of the industrial product space Perform a discrete cosine transform: X freq =DCT(X space ) Among them, X freq It is the frequency domain representation of the surface image of industrial products; Then, the frequency domain representation of the industrial product surface image is masked: X masked =X freq ⊙M Among them, X masked is the masked industrial product surface spectrum, M represents the dynamically generated mask, and ⊙ is the point-by-point multiplication operation.

3. The self-supervised industrial defect detection method based on frequency domain masking strategy according to claim 2 is characterized in that: The mask M is generated based on low-frequency and high-frequency separation masking, random frequency masking, or energy-based adaptive masking.

4. The self-supervised industrial defect detection method based on frequency domain masking strategy according to claim 1 is characterized in that: Step 2: Based on Vision Transformer, a shared encoder is constructed to extract features from the masked industrial product surface image and the masked industrial product surface spectrum to obtain spatial domain shared coding features and frequency domain shared coding features. The specific method is as follows: The masked industrial product surface image and the masked industrial product surface spectrum are divided into N fixed-size token image blocks, all of which are input into a linear projection layer. The corresponding output of each image block is a tensor of length D. Then, the two images are each tensor of length D. token The outputs are spliced ​​to generate the initial feature representation H1 of the masked industrial product surface image and the initial feature representation H2 of the masked industrial product surface spectrum map, and their specifications are both N token ×D, and then further add these two feature representations H1 and H2 with positional encoding to obtain the position information of the masked industrial product surface image in the spatial domain and the position information of the masked industrial product surface spectrum in the frequency domain, respectively. Then, the feature representations H1 and H2 are input into the shared encoder respectively; In the shared encoder, feature representations H1 and H2 are sequentially passed through N share Transformer blocks. Specifically, the feature representation H1 is input to the first Transformer block of the shared encoder to obtain the output of the Transformer block, and then this output is input to the second Transformer block of the shared encoder to obtain the output of the Transformer block; the feature representations H1 and H2 are in the Nth share The outputs of the last Transformer block are H space and H freq , specification is N token ×D.

5. The self-supervised industrial defect detection method based on frequency domain masking strategy according to claim 4 is characterized in that: Step 3: Construct two independent encoder-decoder branches, including a spatial domain branch and a frequency domain branch. Each branch consists of a Transformer block as an encoder and a convolutional layer as a decoder. The specific method is as follows: The spatial domain branch is used to share the output H of the encoder space Image reconstruction is performed. In this branch, H space Input to N branch After a Transformer block encoder and a convolutional layer decoder, the output is a reconstructed image of the surface image of the industrial product. The frequency domain branch is used to share the output H of the encoder freq To reconstruct the frequency domain, in this branch, H freq Input to N branch After the encoder of the Transformer block and the final decoder, the output is the reconstructed spectrum of the industrial product surface spectrum.

6. The self-supervised industrial defect detection method based on frequency domain masking strategy according to claim 4 is characterized in that: Step 4: construct a time-frequency consistency loss function based on cosine similarity loss, frequency domain recovery loss, and spatial domain recovery loss to complete the training of the encoder-decoder branch. The specific method is: The spatial domain restoration loss function is introduced to minimize the reconstructed image of the industrial product surface image. Industrial product surface image X space The pixel difference between them is taken as the goal, so that an image closer to the original surface can be reconstructed in the spatial domain, where the spatial domain restoration loss function is: Among them, M space is the index set of the image blocks masked in the spatial domain, which can be obtained when masking the surface image of the industrial product. is the tensor corresponding to the i-th image block of the reconstructed image of the industrial product surface image, X space,i is the tensor corresponding to the i-th image block of the industrial product surface image, N space is the number of masking blocks; A frequency domain restoration loss function is introduced to minimize the reconstructed spectrum of the surface spectrum of industrial products. Surface spectrum of industrial products X freq The goal is to reconstruct an image closer to the original surface in the frequency domain, where the frequency domain restoration loss function is: The output H of the spatial domain branch space The output of the frequency domain branch H freq The cosine similarity loss is introduced between them to enhance the consistency of spatial and frequency domain feature representations, where the cosine similarity loss is: Among them, N token is the total number of image blocks, its meaning is as mentioned above, H space,i and H freq,i They represent the spatial domain branch output H space The frequency domain branch output H freq The tensor corresponding to the i-th image block; Construct the overall loss function, expressed as: L total =λ1L space +λ2L freq +λ3L con Among them, λ1, λ2, λ3 are weight coefficients, and the pre-training process of the shared encoder and the spatial domain branch and the frequency domain branch is completed by minimizing the overall loss function.

7. The self-supervised industrial defect detection method based on frequency domain masking strategy according to claim 4 is characterized in that: Step 5: Based on the pre-trained shared encoder and encoder, the detection network is connected to build a self-supervised industrial defect detection model. The encoder outputs of the spatial domain branch and the frequency domain branch are spliced ​​and input into the detection network to identify the defect area. The labeled samples are used to fine-tune the model, where: The detection network uses Faster R-CNN or SSD.

8. A self-supervised industrial defect detection system based on a frequency domain masking strategy, implementing the self-supervised industrial defect detection method based on a frequency domain masking strategy described in any one of claims 1-7, to achieve self-supervised industrial defect detection based on a frequency domain masking strategy.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the self-supervised industrial defect detection method based on the frequency domain mask strategy according to any one of claims 1 to 7 is implemented to realize self-supervised industrial defect detection based on the frequency domain mask strategy.

10. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method for self-supervised industrial defect detection based on a frequency domain mask strategy according to any one of claims 1 to 7 is implemented to realize self-supervised industrial defect detection based on a frequency domain mask strategy.

Citation Information

Patent Citations

  • Industrial flaw detection method and device, storage medium and electronic equipment

    CN116245788A

  • Image reconstruction method and system based on self-encoding neural network

    CN117218149A

  • Weak supervision image tampering and forgery detection method based on mask consistency

    CN118155055A

  • Mask recovery enhancement-based low-resolution weak and small target detection method and system

    CN118864826A

  • Mask-guided frequency domain-spatial domain hybrid network target tracking method

    CN119006979A