Unsupervised defect detection method and system based on multi-scale characteristic distillation reconstruction fusion

An unsupervised defect detection method based on multi-scale feature distillation, reconstruction and fusion uses Perlin noise to generate abnormal images and combines teacher and student networks to solve the problem of inaccurate multi-scale feature processing in existing technologies and achieve efficient and accurate defect detection.

CN120612301APending Publication Date: 2025-09-09HANGZHOU DIANZI UNIV

Patent Information

Application Number
CN202510706580.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing industrial defect detection methods tend to ignore detailed information when processing multi-scale features and rely on high-quality labeled data, resulting in unsatisfactory detection results and poor generalization capabilities in different scenarios.

Method used

An unsupervised defect detection method based on multi-scale feature distillation, reconstruction and fusion is adopted. Perlin noise is used to generate abnormal images, and teacher and student networks are used for feature distillation and reconstruction. Transformer is combined to repair defect features, and defect area detection is achieved through a multi-scale feature segmentation network.

Benefits of technology

In the absence of labeled data, high-precision and robust defect detection is achieved, reducing dependence on manual labeling and improving detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612301A_ABST
    Figure CN120612301A_ABST
Patent Text Reader

Abstract

The invention provides an unsupervised defect detection method and system based on multi-scale feature distillation reconstruction fusion, and the method comprises the following steps: S1, employing an abnormal image generation network, and employing a normal sample to generate an abnormal image; s2, establishing a feature distillation reconstruction network, inputting a normal image into the teacher network, and supervising and training the student network through the pre-trained teacher network; the student network performs feature distillation and reconstruction on the input image; s3, adopting a defect feature repair network, taking the obtained defect features as input, and obtaining a plurality of repair features of different scales by using deconvolution; s4, keeping feature details in different scales through comparison of distillation and reconstruction features of different scales and repair features; s5, designing a multi-scale feature segmentation network, and detecting a defect region in the defect image by using features of different scales; and S6, inputting the to-be-predicted image into the teacher network and the student network, and obtaining a predicted abnormal mask through the multi-scale feature segmentation network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of the combination of industrial automation and artificial intelligence, and specifically relates to an unsupervised defect detection method and system based on multi-scale feature distillation, reconstruction and fusion. Background Art

[0002] At present, many industrial defect detection methods mainly rely on image processing technology and machine learning algorithms, but these methods often face some challenges, such as difficulty in obtaining high-quality data, high real-time requirements, and poor generalization ability in different scenarios. In order to improve the accuracy and efficiency of defect detection, researchers in this field have introduced deep learning technology, especially convolutional neural networks (CNNs), as mainstream feature extraction and classification tools. However, when processing multi-scale features, deep learning models tend to ignore some detailed information, resulting in unsatisfactory detection results. In addition, with the continuous development of unsupervised learning technology, defect detection methods based on unlabeled data for training have become an important research direction. Unsupervised learning can automatically discover potential patterns in data without manually labeled data, reducing the burden of manual labeling, and has strong adaptability and flexibility. Based on this, the present invention proposes an unsupervised defect detection method and system based on multi-scale feature distillation, reconstruction and fusion. Summary of the Invention

[0003] To address the aforementioned limitations of existing technologies, this paper proposes an unsupervised defect detection method and system based on multi-scale feature distillation, reconstruction, and fusion. By introducing multi-scale feature distillation, this method effectively integrates information at different scales, improving the accuracy and robustness of defect detection. Through unsupervised learning, the method automatically learns the feature distribution of data in the absence of labeled data, thereby achieving accurate defect detection.

[0004] The present invention adopts the following technical solutions:

[0005] The unsupervised defect detection method based on multi-scale feature distillation, reconstruction and fusion includes the following steps:

[0006] Step S1, using a Perlin noise-based abnormal image generation network to generate abnormal images using normal samples;

[0007] Step S2: Establish a feature distillation and reconstruction network including a teacher network and a student network. The teacher network inputs a normal image, and supervises the student network through the pre-trained teacher network; the student network performs feature distillation and reconstruction on the input image;

[0008] Step S3: Using a Transformer-based defect feature repair network, taking the defect features obtained in step S2 as input, and using deconvolution to obtain repair features of multiple scales;

[0009] Step S4: retaining feature details at different scales by distilling and comparing reconstructed features with repaired features at different scales;

[0010] Step S5: design a multi-scale feature segmentation network to detect defect areas in the defect image using features of different scales;

[0011] In step S6, the image to be predicted is input into the teacher network and the student network, and the predicted anomaly mask is obtained through the multi-scale feature segmentation network.

[0012] Preferably, step S1 is in the abnormal synthesis stage, using an abnormal image generation network based on Perlin noise, randomly selecting images from other external texture datasets (such as DTD dataset, DAGM dataset, BTAD dataset, etc.), and using Perlin noise to generate an abnormal mask to determine the location of the abnormal area. Subsequently, the corresponding area is extracted from the random image and superimposed on the normal sample, thereby synthesizing an image with an abnormal area while maintaining the normal sample structure, realizing the generation of abnormal samples, and generating an abnormal image I. a , generated as follows:

[0013] I a =(1-β)(M⊙I)+(1-M)⊙I+β(M⊙A)

[0014] Among them, I, I a They represent the original image without abnormalities and the image with abnormalities respectively; ⊙ represents element-by-element multiplication; β represents the mixing factor; M simulates abnormal areas of different shapes through Perlin noise and binarizes them into abnormality masks.

[0015] Preferably, step S2 is the first training phase. During this phase, the teacher network receives a normal image without synthetic anomalies as input. A feature distillation reconstruction network, consisting of the teacher and student networks, is established to generate distilled reconstructed features. The student network uses a U-Net architecture and incorporates specific 3×3 convolutional modules at key locations to extract features at different scales, ensuring that its output effectively matches the feature representation T of the teacher network. During training, the input image undergoes dimensionality increase and then dimensionality reduction to achieve effective feature reconstruction and alignment.

[0016] The teacher network uses a pre-trained network, which can more efficiently extract multiple scale features contained in the original image.

[0017] The student network designed in this step is based on the U-Net architecture and consists of an encoder and a decoder. It performs feature distillation and reconstruction on the original image and outputs the defect feature D. First, the 3-channel image is encoded into features at four scales through the encoder. The input is set to The output is The convolution kernel is The bias term is It is calculated as follows:

[0018]

[0019] in, Represents the value of the input feature map at the c1th channel and position (h·s+i,w·s+j), where s is the stride. Is the weight of the convolution kernel, which represents the weight value of the position on the output channel c2 and the input channel c1. is the bias of the c2th output channel.

[0020] Through the decoding step, the encoded features are gradually decoded into the required features, and the effect of reconstructing the features is achieved. The reconstructed features are compared with the features extracted by the teacher network, and the cosine similarity loss is used to calculate the feature Y extracted by the student network. i s As close as possible to the output T of the teacher network i , the loss is calculated as follows:

[0021]

[0022] Among them, T i 、Y i s are the outputs of the teacher network and the student network in the training phase, respectively. '·' represents the sum (dot product) after element-by-element multiplication, and ||·|| represents the L2 norm calculated by dimension. i and h i It is the size information of the corresponding channel.

[0023] Preferably, step S3 is in the first stage of training, and a defect feature repair network based on Transformer is used to generate repaired features. The network takes the defect feature D obtained in step S2 as input, and extracts high-dimensional feature representation D through the Transformer module. T , and uses a 3×3 convolution module to perform dimensionality reduction processing to obtain the restored feature representation. The process is as follows:

[0024] First, the defect feature D extracted in step S2 is divided into several patches, and each patch is mapped to a high-dimensional space to obtain the embedded patch P:

[0025] P=PatchEmbedding(D)

[0026] After patch segmentation, each patch is embedded into a vector, represented as:

[0027] P i =W patch ·x i

[0028] Among them, x i is the i-th patch in the defect signature, W patch is the weight matrix used for mapping, P i is the embedding of the i-th patch.

[0029] The obtained high-dimensional defect feature D T , use deconvolution to obtain multiple restoration features Y of different scales d , assuming the input of layer l is X (l) , the output is Y d , which is calculated as follows:

[0030]

[0031] Where Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key.

[0032] The repair feature after repairing the defect feature is obtained by deconvolution. The same loss as step S2 is used to make the repair feature Y d Also as close as possible to the outputs T and Y of the teacher and student networks s .

[0033] Preferably, step S4, at this time enters the second stage of training, through the distillation and reconstruction of different scales obtained in step S2 s Repair feature Y with step S3 d The comparison can preserve the feature details at different scales, such as local texture details, structure and contour information, semantics, etc.

[0034] At this stage, the images input to the student network and the auxiliary network are still images with synthetic anomalies I a , and the teacher network input is also an image with synthetic anomalies I a (i.e., the abnormal image of S1). Multi-scale comparison is performed using the multi-scale repair features output by the student network and auxiliary network and the abnormal defect features output by the teacher network. The results of the multi-scale comparison are then fused to obtain high-precision anomaly detection results.

[0035] Preferably, in step S5, a multi-scale feature segmentation network is designed to achieve high-precision detection of defect areas in defect images using features of different scales. The calculation method is as follows:

[0036] pre_mask = sigmoid(diff+d c )

[0037] Among them, diff is the output comparison between the three networks, d c It is the correction matrix in the network, pre_mask is the final predicted anomaly mask, and the predicted anomaly mask is compared with the real mask generated in the synthetic anomaly mask. By comparing and training the multi-scale feature segmentation network, a high-precision predicted anomaly mask is obtained.

[0038] Preferably, in step S6, the image actually to be predicted is input into the teacher network and the student network, and the predicted anomaly mask obtained by the multi-scale feature segmentation network is the predicted anomaly mask for the image to be predicted.

[0039] The present invention also discloses an unsupervised defect detection system based on multi-scale feature distillation, reconstruction and fusion, which is used to execute the above method and includes the following modules:

[0040] Abnormal image generation module: uses an abnormal image generation network based on Perlin noise and normal samples to generate abnormal images;

[0041] Feature distillation and reconstruction module: Establish a feature distillation and reconstruction network consisting of a teacher network and a student network. The teacher network inputs a normal image, and the pre-trained teacher network supervises the student network for training. The student network performs feature distillation and reconstruction on the input image.

[0042] Feature repair module: This module uses a Transformer-based defect feature repair network, takes the defect features obtained by the feature distillation and reconstruction modules as input, and uses deconvolution to obtain repair features at multiple scales.

[0043] Feature comparison module: preserves feature details at different scales by distilling and comparing reconstructed features with repaired features at different scales;

[0044] Defective area detection module: Design a multi-scale feature segmentation network and use features of different scales to detect defective areas in defective images;

[0045] Mask prediction module: The image to be predicted is input into the teacher network and the student network, and the predicted anomaly mask is obtained through the multi-scale feature segmentation network.

[0046] The present invention proposes an unsupervised defect detection method and system based on multi-scale feature distillation, reconstruction and fusion. The present invention achieves high-precision defect detection in the absence of labeled data by introducing multi-scale feature distillation technology. The present invention first uses Perlin noise to generate abnormal images, extracts and reconstructs image features of different scales through the feature distillation reconstruction process of the teacher network and the student network, and repairs the defect features through a Transformer-based approach. Then, the method fuses feature information of different scales through multi-scale comparison and feature segmentation networks, and finally achieves accurate prediction of defect areas. The present invention avoids dependence on labeled data, has strong robustness and efficiency, and can automatically learn and detect defects in images without labeled data. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0048] Figure 1 This is a flow chart of a high-precision and efficient unsupervised defect detection method based on multi-scale feature distillation, reconstruction and fusion according to a preferred embodiment of the present invention;

[0049] Figure 2 2 is a schematic diagram of the detection results using the preferred embodiment of the present invention;

[0050] Figure 3 This is a structural diagram of a multi-scale feature segmentation network according to a preferred embodiment of the present invention;

[0051] Figure 4 This is a block diagram of an unsupervised defect detection system based on multi-scale feature distillation, reconstruction and fusion in a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0052] To make the embodiments, technical solutions, and advantages of the present invention more apparent, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, rather than all of them. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0053] like Figure 1 As shown, this embodiment provides a high-precision and efficient unsupervised defect detection method based on multi-scale feature distillation, reconstruction and fusion, which includes the following steps:

[0054] Step S100: Generate an abnormal image using Perlin noise and generate an abnormal mask to provide synthetic abnormal data for the training phase. Specifically:

[0055] At this stage, we are in the abnormal synthesis stage. We use an abnormal image generation network based on Perlin noise. This network can use existing technology and normal samples to generate abnormal images. The generation method is as follows:

[0056] I a =(1-β)(M⊙I)+(1-M)⊙I+β(M⊙A)

[0057] Among them, I, I a They represent the original image without abnormalities and the image with abnormalities respectively; ⊙ represents element-by-element multiplication; β represents the mixing factor; M simulates abnormal areas of different shapes through Perlin noise and binarizes them into abnormality masks.

[0058] Step S200 marks the first stage of training. During this stage, the teacher network receives a normal image without synthetic anomalies as input. A feature distillation reconstruction network, consisting of the teacher and student networks, is established to generate distilled reconstructed features. The student network is supervised by the pre-trained teacher network. The student network uses a U-Net architecture as its foundation and incorporates a specific 3×3 convolutional module in each decoder layer to extract features at different scales, ensuring that its output effectively matches the feature representation T of the teacher network. During training, the input image undergoes dimensionality increase and then dimensionality reduction to achieve effective feature reconstruction and alignment.

[0059] The student network designed in this embodiment is based on the U-Net architecture and consists of an encoder and a decoder. It performs feature distillation and reconstruction on the original image and outputs the defect feature D. First, the 3-channel image is encoded into features at four scales through the encoder. The input is set to The output is The convolution kernel is The bias term is It is calculated as follows:

[0060]

[0061] in, Represents the value of the input feature map at the c1th channel and position (h·s+i,w·s+j), where s is the stride. Is the weight of the convolution kernel, which represents the weight value of the position on the output channel c2 and the input channel c1. is the bias of the c2th output channel.

[0062] Through the decoding step, the encoded features are gradually decoded into the required features, and the effect of reconstructing the features is achieved at the same time. The reconstructed features are compared with the features extracted by the teacher network, and the cosine similarity loss is used to calculate the features extracted by the student network so that the features are as close as possible to the output of the teacher network. i , the loss is calculated as follows:

[0063]

[0064] Among them, T i 、Y i s are the outputs of the teacher network and the student network in the training phase, respectively. '·' represents the sum (dot product) after element-by-element multiplication, and ||·|| represents the L2 norm calculated by dimension. i and h i It is the size information of the corresponding channel.

[0065] In this step, the teacher network extracts multi-scale features through pre-training, and the student network performs feature distillation and reconstruction through dimensionality increase and reduction, and optimizes the student network through feature comparison.

[0066] Step S300 is the first stage of training. A Transformer-based defect feature repair network is used. The network can use existing technology to generate repaired features. The network takes the defect feature D obtained in step S200 as input and extracts high-dimensional feature representation D through the Transformer module. T , and uses a 3×3 convolution module to perform dimensionality reduction processing to obtain the restored feature representation. The process is as follows:

[0067] First, the defect feature D extracted in step S200 is divided into several patches, and each patch is mapped to a high-dimensional space to obtain an embedded patch P:

[0068] P=PatchEmbedding(D)

[0069] After patch segmentation, each patch is embedded into a vector, represented as:

[0070] P i =W patch ·x i

[0071] Among them, x i is the i-th patch in the defect signature, W patch is the weight matrix used for mapping, P i is the embedding of the i-th patch.

[0072] The obtained high-dimensional defect feature D T, use deconvolution to obtain multiple restoration features of different scales. Assume that the input of the lth layer is X (l) , the output is Y d , which is calculated as follows:

[0073]

[0074] Where Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key.

[0075] The repair feature after repairing the defect feature is obtained by Transformer-based deconvolution. The same loss as S200 is used to make the repair feature Y d Also as close as possible to the outputs T and Y of the teacher and student networks s .

[0076] In this step, defect features are extracted from the student network, and after high-dimensional extraction and mapping to high-dimensional space through Transformer, dimensionality reduction processing is performed to obtain the repaired defect features.

[0077] Step S400, at this time, enters the second stage of training, through the distillation of different scales to reconstruct features and repair features with high contrast, retaining feature details in different scales, such as: local texture details, structure and contour information, semantics, etc.

[0078] At this stage, the images input to the student network and the auxiliary network are still images with synthetic anomalies I a , and the teacher network input is also an image with synthetic anomalies I a The multi-scale repair features output by the student network and the auxiliary network are used with the defect features with anomalies output by the teacher network for multi-scale comparison. The results of the multi-scale comparison are then fused to obtain high-precision anomaly detection results.

[0079] In this step, through multi-scale comparison, the repair features output by the student network and the auxiliary network are fused with the defect features of the teacher network to retain details and improve detection accuracy.

[0080] Step S500: design a multi-scale feature segmentation network, such as Figure 3 As shown in Figure 1, features of different scales are used to achieve high-precision detection of defect areas in defect images. The calculation method is as follows:

[0081] pre_mask = sigmoid(diff+d c )

[0082] Among them, diff is the output comparison between the three networks, d cIt is the correction matrix in the network, pre_mask is the final predicted anomaly mask, and the predicted anomaly mask is compared with the real mask generated in the synthetic anomaly mask. By comparing and training the multi-scale feature segmentation network, a high-precision predicted anomaly mask is obtained.

[0083] In step S600, the image that actually needs to be predicted is input into the teacher network, the student network, and the auxiliary network, and a prediction anomaly mask is generated through the multi-scale feature segmentation network to complete the defect prediction. This prediction anomaly mask is the prediction mask for the image that needs to be predicted, thereby achieving high-precision unsupervised defect detection.

[0084] refer to Figure 2 The data used in this embodiment are industrial product images with manually marked abnormal areas. The method of this embodiment is used to process the image to be detected, and the corresponding abnormality detection results can be output after processing. Figure 2 As shown in the figure, the detection heat map is the prediction result of the abnormal area in the implementation step S600, where the depth of the color represents the model's estimate of the abnormal probability of each area: the closer the color is to red, the higher the probability that the area is judged to be abnormal. On this basis, by setting a threshold (set to 0.5) for the detection heat map, an abnormality mask can be further generated: areas above the threshold are judged to be abnormal and marked as white, and areas below the threshold are considered normal and marked as black, thus obtaining the final abnormality mask map. Figure 2 It can be seen intuitively that the detection method proposed in the present invention is highly consistent with the manual labeling results in identifying abnormal areas, which verifies the effectiveness of the method of the present invention. It can be widely used in actual industrial scenarios, effectively reducing manual detection costs and improving production efficiency.

[0085] Refer to the attached Figure 4 This embodiment provides an unsupervised defect detection system based on multi-scale feature distillation, reconstruction and fusion, including the following modules:

[0086] Abnormal image generation module: Using an abnormal image generation network based on Perlin noise, we randomly select images from other external texture datasets and generate abnormal masks using Perlin noise to determine the location of abnormal areas. Subsequently, we extract the corresponding areas from the random images and superimpose them on the normal samples, thereby synthesizing images with abnormal areas while maintaining the normal sample structure, achieving the generation of abnormal samples and generating abnormal images. a , generated as follows:

[0087] I a =(1-β)(M⊙I)+(1-M)⊙I+β(M⊙A)

[0088] Among them, I, I aThey represent the original image without abnormalities and the image with abnormalities respectively; ⊙ represents element-by-element multiplication; β represents the mixing factor; M simulates abnormal areas of different shapes through Perlin noise and binarizes them into abnormality masks.

[0089] Feature Distillation and Reconstruction Module: The teacher network takes as input a normal image without synthetic anomalies. A feature distillation and reconstruction network is established, consisting of the teacher and student networks, to generate distilled reconstructed features. The student network uses a U-Net architecture and introduces specific 3×3 convolutional modules at key locations to extract features at different scales, ensuring that its output effectively matches the feature representation T of the teacher network. During training, the input image undergoes dimensionality increase and then dimensionality reduction to achieve effective feature reconstruction and alignment. The teacher network utilizes a pretrained network, which can more efficiently extract features at multiple scales from the original image.

[0090] The student network designed in this embodiment is based on the U-Net architecture and consists of an encoder and a decoder. It performs feature distillation and reconstruction on the original image and outputs the defect feature D. First, the 3-channel image is encoded into features at four scales through the encoder. The input is set to The output is The convolution kernel is The bias term is It is calculated as follows:

[0091]

[0092] in, Represents the value of the input feature map at the c1th channel and position (h·s+i,w·s+j), where s is the stride. Is the weight of the convolution kernel, which represents the weight value of the position on the output channel c2 and the input channel c1. is the bias of the c2th output channel.

[0093] Through the decoding step, the encoded features are gradually decoded into the required features, and the effect of reconstructing the features is achieved at the same time. The reconstructed features are compared with the features extracted by the teacher network, and the cosine similarity loss is used to calculate the features extracted by the student network so that the features are as close as possible to the output of the teacher network. i , the loss is calculated as follows:

[0094]

[0095] Among them, T i 、Y i sare the outputs of the teacher network and the student network in the training phase, respectively. '·' represents the sum (dot product) after element-by-element multiplication, and ||·|| represents the L2 norm calculated by dimension. i and h i It is the size information of the corresponding channel.

[0096] Feature repair module: A Transformer-based defect feature repair network is used to generate repaired features. The network takes the defect feature D obtained in step S2 as input and extracts high-dimensional feature representation D through the Transformer module. T , and uses a 3×3 convolution module to perform dimensionality reduction processing to obtain the restored feature representation. The process is as follows:

[0097] First, the defect feature D extracted in step S2 is divided into several patches, and each patch is mapped to a high-dimensional space to obtain the embedded patch P:

[0098] P=PatchEmbedding(D)

[0099] After patch segmentation, each patch is embedded into a vector, represented as:

[0100] P i =W patch ·x i

[0101] where x i is the i-th patch in the defect signature, W patch is the weight matrix used for mapping, P i is the embedding of the i-th patch.

[0102] The obtained high-dimensional defect feature D T , use deconvolution to obtain multiple restoration features Y of different scales d , assuming the input of layer l is X (l) , the output is Y d , which is calculated as follows:

[0103]

[0104] Where Q is the query matrix, K is the key matrix, V is the value matrix, d k is the dimension of the key.

[0105] The repair features after repairing the defect features are obtained through Transformer-based deconvolution. The same loss as S200 is used to make the repair features as close as possible to the output of the teacher and student networks.

[0106] Feature comparison module: Through distillation at different scales, it reconstructs features and repairs high contrast features, preserving feature details at different scales, such as local texture details, structure and contour information, and semantics.

[0107] At this stage, the student and auxiliary networks are still fed images with synthetic anomalies, while the teacher network is also fed images with synthetic anomalies. Multi-scale inpaint features output by the student and auxiliary networks are compared with defect features with anomalies output by the teacher network for multi-scale comparison. The results of these multi-scale comparisons are then fused to achieve highly accurate anomaly detection.

[0108] Defective area detection module: A multi-scale feature segmentation network is designed to use features at different scales to achieve high-precision detection of defective areas in defective images. The calculation method is as follows:

[0109] pre_mask = sigmoid(diff+d c )

[0110] Where diff is the output comparison between the three networks, d c It is the correction matrix in the network, pre_mask is the final predicted anomaly mask, and the predicted anomaly mask is compared with the real mask generated in the synthetic anomaly mask. By comparing and training the multi-scale feature segmentation network, a high-precision predicted anomaly mask is obtained.

[0111] Mask prediction module: The image that actually needs to be predicted is input into the teacher network, student network and auxiliary network. The prediction mask obtained by the multi-scale feature segmentation network is the predicted anomaly mask for the image that needs to be predicted.

[0112] From the above description and result graphs, it can be seen that the present invention proposes an unsupervised defect detection method and system based on multi-scale feature distillation, reconstruction and fusion. By introducing multi-scale feature distillation technology, high-precision defect detection is achieved in the absence of labeled data. The method first uses Perlin noise to generate abnormal images, extracts and reconstructs image features at different scales through the feature distillation reconstruction process of the teacher network and the student network, and repairs the defect features through patch embedding technology. Then, the method fuses feature information at different scales through multi-scale comparison and feature segmentation networks, and finally achieves accurate prediction of defect areas. This method avoids dependence on labeled data, has strong robustness and efficiency, and can automatically learn and detect defects in images without labeled data.

[0113] It should be noted that in the description of this invention, the use of terms such as "input," "extraction," "training," "local," "multi-scale," "distillation," "transfer," and "reconstruction" to refer to specific concepts is intended to clearly express and understand the technical features of this invention, and does not limit its implementation. In actual applications, relevant parameters and methods can be adjusted according to specific scenarios and requirements.

[0114] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "application", "combination", "connection", etc. should be understood in a broad sense. For example, it can be the direct use of relevant methods and technologies, or modification or expansion based on them; it can be the close combination of multiple methods and technologies to form an overall solution, or the selective adoption of some of them; it can be the integration of different data into a whole so that they can work together, or the interaction between different data can be achieved through interfaces or middleware. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0115] The term "comprise" or any other similar term is intended to cover non-exclusive inclusion, such that a process, article, or apparatus / device that comprises a list of elements includes not only those elements but also other elements not expressly listed or inherent to such process, article, or apparatus / device.

[0116] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.

Claims

1. An unsupervised defect detection method based on multi-scale feature distillation, reconstruction and fusion, characterized by: The method comprises the following steps: Step S1, using a Perlin noise-based abnormal image generation network to generate abnormal images using normal samples; Step S2: Establish a feature distillation and reconstruction network including a teacher network and a student network. The teacher network inputs a normal image, and supervises the student network through the pre-trained teacher network; the student network performs feature distillation and reconstruction on the input image; Step S3: Using a Transformer-based defect feature repair network, taking the defect features obtained in step S2 as input, and using deconvolution to obtain repair features of multiple scales; Step S4: retaining feature details at different scales by distilling and comparing reconstructed features with repaired features at different scales; Step S5: design a multi-scale feature segmentation network to detect defect areas in the defect image using features of different scales; In step S6, the image to be predicted is input into the teacher network and the student network, and the predicted anomaly mask is obtained through the multi-scale feature segmentation network.

2. The unsupervised defect detection method based on multi-scale feature distillation, reconstruction and fusion as claimed in claim 1, characterized in that: In step S1, the generation method is as follows: I a =(1-β)(M⊙I)+(1-M)⊙I+β(M⊙A) Among them, I, I a They represent the original image without abnormalities and the image with abnormalities respectively; ⊙ represents element-by-element multiplication; β represents the mixing factor; M simulates abnormal areas of different shapes through Perlin noise and binarizes them into abnormality masks.

3. The unsupervised defect detection method based on multi-scale feature distillation, reconstruction and fusion as claimed in claim 2, characterized in that: In step S2, the student network is based on the U-Net architecture and consists of an encoder and a decoder. It performs feature distillation and reconstruction on the original image and outputs defect features. First, the 3-channel image is encoded into features at four scales through the encoder, and the input is set to The output is The convolution kernel is The bias term is The calculation is as follows: in, represents the value of the input feature map at the c1th channel and position (h·s+i,w·s+j), where s is the stride; is the weight of the convolution kernel, which represents the weight value of the position on the output channel c2 and the input channel c1; is the bias of the c2th output channel; By decoding, the encoded features are gradually decoded into the required features, and the effect of reconstructing the features is achieved. The reconstructed features are compared with the features extracted by the teacher network and calculated using the cosine similarity loss. The loss calculation method is as follows: Among them, T i 、Y i s are the outputs of the teacher network and the student network in the training phase, '·' represents the sum of element-wise multiplication, ||·|| represents the L2 norm calculated by dimension, and w i and h i It is the size information of the corresponding channel.

4. The unsupervised defect detection method based on multi-scale feature distillation, reconstruction and fusion as claimed in claim 3, characterized in that: Step S3 is as follows: first, the defect feature D extracted in step S2 is divided into several patches, and each patch is mapped to a high-dimensional space to obtain an embedded patch P: P=PatchEmbedding(D) After patch segmentation, each patch is embedded into a vector, represented as: P i =W patch ·x i Among them, x i is the i-th patch in the defect signature, W patch is the weight matrix used for mapping, P i is the embedding of the i-th patch; The obtained high-dimensional defect feature D T , use deconvolution to obtain multiple restoration features Y of different scales d , let the input of the lth layer be X (l) , the output is Y d , the calculation method is as follows: Where Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key; The repair features after repairing the defect features are obtained through deconvolution.

5. The unsupervised defect detection method based on multi-scale feature distillation, reconstruction and fusion as claimed in claim 4, characterized in that: In step S4, the multi-scale repair features output by the student network and the auxiliary network are used to perform multi-scale comparison with the defect features with anomalies output by the teacher network, and then the results of the multi-scale comparison are fused to obtain the anomaly detection result.

6. The unsupervised defect detection method based on multi-scale feature distillation, reconstruction and fusion as claimed in claim 5, characterized in that: In step S5, the calculation method is as follows: pre_mask=sigmoid(diff+d c ) Among them, diff is the output comparison between the three networks, d c It is the correction matrix in the network, pre_mask is the final predicted anomaly mask, and the predicted anomaly mask is compared with the real mask generated in the synthetic anomaly mask. The multi-scale feature segmentation network is trained by comparison to obtain the predicted anomaly mask.

7. An unsupervised defect detection system based on multi-scale feature distillation, reconstruction and fusion, for executing the method according to any one of claims 1 to 6, characterized in that: Includes the following modules: Abnormal image generation module: uses an abnormal image generation network based on Perlin noise and normal samples to generate abnormal images; Feature distillation and reconstruction module: Establish a feature distillation and reconstruction network consisting of a teacher network and a student network. The teacher network inputs a normal image, and the pre-trained teacher network supervises the student network for training. The student network performs feature distillation and reconstruction on the input image. Feature repair module: This module uses a Transformer-based defect feature repair network, takes the defect features obtained in the feature distillation and reconstruction modules as input, and uses deconvolution to obtain repair features at multiple scales. Feature comparison module: preserves feature details at different scales by distilling and comparing reconstructed features with repaired features at different scales; Defective area detection module: Design a multi-scale feature segmentation network and use features of different scales to detect defective areas in defective images; Mask prediction module: The image to be predicted is input into the teacher network and the student network, and the predicted anomaly mask is obtained through the multi-scale feature segmentation network.

Citation Information

Patent Citations

  • Retina OCT image lesion classification method based on unsupervised heterogeneous distillation framework

    CN116091449A

  • Unsupervised defect detection method of tandem knowledge distillation added with Transform

    CN116468667A

  • Image anomaly detection method and system based on knowledge migration, terminal and medium

    CN116883773A

  • Single-modal image segmentation method based on cross-modal consistency

    CN118135222A

  • Heterogeneous knowledge distillation method and device based on Transform attention mechanism

    CN119337968A

Cited By

  • Defect detecting and positioning method based on regional anomaly generation and multi-level reverse distillation

    CN121527502A

  • Electronic component defect detection and segmentation method based on normal sample prompt

    CN121810590A