Method and System for Defect Detection and Location Based on Cross-Image Local Feature Alignment
By introducing a cross-image local feature alignment loss function in the distillation learning model, the problem of poor defect detection effect in the existing methods is solved, and more efficient object surface defect detection and positioning is achieved.
Patent Information
- Application Number
- CN202111502012.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-12-09
AI Technical Summary
The existing object surface defect detection and positioning methods based on distillation learning models fail to effectively utilize the local feature alignment information across images, resulting in poor detection results.
The distillation learning model is constructed, and the local feature alignment loss function across images is introduced during the training process, and the teacher model is used to guide the student model to learn feature extraction of normal image samples. The pixel-pixel local correspondence relationship of the training set is used to constrain the feature space across the image, so as to improve the model's sensitivity to fine-grained local pixel-level features.
It improves the accuracy and interpretability of object surface defect detection and positioning, enhances the model's sensitivity to local pixel-level features, and improves the detection effect.
Smart Images

Figure CN114170478B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object surface defect anomaly detection and localization, and more specifically, to a method and system for defect detection and localization based on cross-image local feature alignment. Background Art
[0002] With the development of computer vision research, object surface defect detection and localization technologies are widely used in fields such as industrial vision inspection and medical image lesion screening. The purpose of anomaly detection and localization is to screen out abnormal sample pictures and locate the abnormal areas in the samples.
[0003] Currently, methods for anomaly detection and localization can be divided into two types: reconstruction-based methods and representation similarity-based methods. Among them, reconstruction-based methods mainly train autoencoders, variational autoencoders, or generative adversarial networks to reconstruct normal samples. During testing, abnormal samples are identified because they cannot be well reconstructed. When determining whether an image is an abnormal sample, the reconstruction error of the entire image is used. When locating the abnormal area, the pixel-level reconstruction error is used. Reconstruction-based methods are very intuitive and interpretable, but their performance is often limited by the generative model. Because sometimes abnormal samples can also be well reconstructed, especially when abnormal samples and normal samples are highly similar, a reconstruction error failure phenomenon will occur. Representation similarity-based methods use deep neural networks to extract the representation of the entire image for anomaly detection and extract the representation of local image patches for anomaly localization. Although most representation similarity-based methods can achieve better results than reconstruction-based methods, they lack interpretability because in this method, the anomaly score comes from the distance between the representation of the test set image and the representation of the normal samples in the training set, and it is actually difficult to know which part of the abnormal image causes the high anomaly score. Moreover, the computational cost of anomaly localization based on image patches in this type of method is also relatively large.
[0004] In the method based on representation similarity, the method based on the distillation learning model is also one of them, which has good interpretability. For example, in the prior art, a positive sample industrial defect detection method based on knowledge distillation is disclosed. First, an industrial data set is constructed, and then preprocessing operations are performed on the industrial data set. The preprocessed industrial data set includes a positive sample set and an unlabeled defect sample set. Then, self-supervised contrast learning is used to pre-train the teacher network model on the formed industrial data set. On the basis of the formed positive sample set, the trained teacher network model is used to guide the training of the student network model. Finally, the trained teacher network model and the trained student network model are used to detect defects in the test pictures. Since the student network model only learns the ability to extract positive sample features, the features extracted from the defect area are quite different from those of the teacher network model, which can be used as the basis for defect judgment. In fact, currently, an industrial data set for object surface defect detection has a major feature, that is, most of the images are the same object that has been registered (the alignment of two or more images of the same target in spatial position). At this time, the cross-image local features are highly correlated. The cross-image local feature alignment information can ensure the sensitivity of the model to fine-grained local pixel-level features. However, the existing methods based on the distillation learning model do not apply this information. Therefore, the effect of defect detection and localization is not good. And how to apply this information to improve the effect of defect detection and localization has become a difficult problem to be solved. Summary of the Invention
[0005] To solve the problem that the detection effect of the traditional method based on the distillation learning model for defect detection and localization is not good, the present invention proposes a method and system for defect detection and localization based on cross-image local feature alignment. Based on the distillation learning model, the pixel-pixel local correspondence relationship of the training set across images is used to constrain the feature space, ensuring the sensitivity of the model to fine-grained local pixel-level features, and then improving the effect of object surface defect anomaly detection and localization.
[0006] To achieve the above technical effects, the technical solution of the present invention is as follows:
[0007] A method for defect detection and localization based on cross-image local feature alignment includes the following steps:
[0008] S1. Construct a distillation learning model, where the distillation learning model includes a teacher model T and a student model S;
[0009] S2. Determine the loss function of the distillation learning model;
[0010] S3. Construct a data set, divide the data set into a training set and a test set. The training set only contains normal image samples, and the test set includes normal image samples and abnormal image samples;
[0011] S4. Construct a cross-image local feature alignment loss function among several normal image samples in the same training batch;
[0012] S5. Integrate the local feature alignment loss function in S2 with the loss function of the distillation learning model in S4 to form the total loss function of the distillation learning model based on cross-image local feature alignment;
[0013] S6. Input the normal image samples in the training set into the teacher model T and the student model S simultaneously. Keep the parameters of the teacher model T fixed. Using the total loss function as the training guidance, use the teacher model T to guide the training of the student model S, thereby training the distillation learning model to obtain a trained distillation learning model;
[0014] S7. Use the image samples in the test set as the input samples of the trained distillation learning model. Starting from the gradient of the total loss function with respect to the input samples, use the trained distillation learning model to perform defect detection and localization on the image samples in the test set.
[0015] In this technical solution, first, a distillation learning model (including the teacher model T and the student model S) is built, and the loss function of the distillation learning model is determined. Then, a data set is constructed and divided into a training set and a test set. Incorporate the cross-image local feature alignment information among several normal image samples in the same training batch into the loss function of the distillation learning model to form the final loss function. Input the normal image samples into the teacher model T and the student model S simultaneously, keep the parameters of the teacher model T fixed, and use the teacher model T to guide the training of the student model S, so that the student model only obtains the ability to extract features of normal image samples, and the model can use the pixel-pixel local correspondence relationship across images in the training set to constrain the feature space, and use the cross-image feature alignment information to ensure the sensitivity of the model to fine-grained local pixel-level features, thereby ensuring the effect of object surface defect anomaly detection and localization.
[0016] Preferably, in the distillation learning model constructed in step S1, the teacher model T adopts the VGG16 network structure loaded with pre-trained weights on ImageNet, and the student model S adopts the same VGG16 network structure as the teacher model T and randomly initializes the weights.
[0017] Here, in order to shorten the training time of the initial distillation learning model, the teacher model T is loaded with pre-trained weights on ImageNet. The teacher model T has good feature extraction capabilities for both normal image samples and abnormal image samples. The student model S adopts the same VGG16 network structure as the teacher model T to facilitate the teacher model T to further guide the training of the student model S.
[0018] Preferably, the VGG16 network structure includes several modules.
[0019] In step S2, both the teacher model T and the student model S take the last layer of the last four modules in their respective network structures as their respective key layers, and the loss function of the distillation learning model is:
[0020] L1 = L val + λL dir
[0021] Where,
[0022]
[0023]
[0024] Where, L1 represents the loss function of the distillation learning model; CP i represents the output of the i-th key layer of the VGG16 network structure; CP0 represents the normal image features of the original input VGG16 network structure; represents the activation value of the i-th key layer of the teacher model T, and the activation value is the normal image features output by the key layer of the network structure; represents the activation value of the i-th key layer of the student model S; N i represents CP i the number of neurons in; represents the activation value of the j-th neuron in the i-th key layer of the teacher model T; represents the activation value of the j-th neuron in the i-th key layer of the student model S; N cp represents the total number of key layers; L val represents the sum of the Euclidean distances of the corresponding activation values in each key layer of the teacher model T and the student model S; vec() represents the vectorization function that converts a matrix with any dimension into a one-dimensional vector; L dir represents the cosine similarity of the vectors converted from each corresponding key layer of the teacher model T and the student model S; λ represents the hyperparameter set manually. The process of establishing this loss function can not only constrain the similarity of the activation values output by the key layers of the student model S and the teacher model T, but also constrain the similarity of the vector directions output by the key layers.
[0025] Preferably, assume that the cross-image local feature alignment loss function is constructed for K normal image samples in the same training batch in step S4. During the training process of the distillation learning model, the alignment losses of the activation value maps output by the first, second,..., K-th normal image samples and the other K-1 normal image samples in the corresponding VGG16 network structures of the teacher model T and the student model S are calculated pixel by pixel in turn, and then 1 / 2 is taken to eliminate the repeated calculations, and the cross-image local feature alignment loss function for K normal image samples in the same training batch is obtained. The expression is:
[0026]
[0027] Among them, L2 represents the cross-image local feature alignment loss function among K normal image samples in the same training batch; since the activation value maps output by the key layers of the VGG16 network structure are used to calculate the alignment loss pixel by pixel across images, the activation value maps are obtained by performing convolution calculations on the original input normal image samples using convolutional kernels, and a pixel position in the activation value map corresponds to the local features of at least 3×3 pixels in the original input normal image sample.
[0028] Preferably, the expression of the total loss function of the distilled learning model formed based on cross-image local feature alignment described in step S5 is:
[0029] L total = L1 + γL2
[0030] Among them, L total represents the total loss function of the distilled learning model based on cross-image local feature alignment; γ represents a hyperparameter set artificially during training.
[0031] Preferably, in step S6, the normal image samples in the training set are input into the teacher model T and the student model S simultaneously. Keeping the parameters of the teacher model T fixed, using the total loss function as the training guidance, when using the teacher model T to guide the training of the student model S, the training method is backpropagation and gradient descent. The output of the key layer of the student model S is the image features of the normal image samples extracted by the student model S. Using the output of the key layer of the student model S to fit the output of the key layer of the teacher model T, so that the student model S only obtains the ability to extract the features of normal image samples. During the training process, when the total loss function L total no longer decreases in 20 training rounds, the model training is completed, and the trained distilled learning model is obtained.
[0032] Here, when the value of the total loss function L total stabilizes in a state of no longer decreasing and converging, it also ensures the convergence of the cross-image local feature alignment loss function L2, ensuring the consistency of cross-image local features, that is, "alignment".
[0033] Preferably, in step S7, starting from the gradient of the total loss function with respect to the input sample, when using the trained distilled learning model to perform defect detection and localization on the image samples in the test set, assuming that the input sample in the test set is represented as x, the expression of the gradient map Λ of the total loss function with respect to the input sample is
[0034]
[0035] The pixel gradient value of the input sample x is directly obtained through a single backpropagation during the training process. Since the total loss function L of the distillation learning model based on cross-image local feature alignment total includes the loss function L1 of the distillation learning model and the local feature alignment loss function L2 across images. When using the test set for testing in step S7, when the input sample x is a normal image sample, both the loss function L1 of the distillation learning model and the local feature alignment loss function L2 across images are small, and the gradient is also small. When the input sample x is an abnormal image sample, the opposite is true. Let the gradient contribution threshold be ε. The gradient value generated at the pixel position of the input sample x with a gradient contribution greater than ε is large, and the pixel position corresponds to the abnormal defect area. That is, using the total gradient map of L
[0036] to perform abnormal localization and obtain the abnormal defect area that causes its value to increase. total Here, when using the trained distillation learning model to perform defect detection and localization on the image samples in the test set, it includes two levels. First, based on the value of the total loss function L total of the trained distillation learning model, during abnormal detection, it is determined whether the input sample is a normal image sample or an abnormal image sample. Second, the gradient of the total loss function L
[0037] with respect to the input sample x is of great significance. Among them, the pixels with a relatively large contribution to the gradient are the areas where the output differences between the key layers of the teacher model T and the student model T and the local feature alignment loss are large, and they are also the abnormal detection areas. total Preferably, when using the gradient map of L
[0038] to perform abnormal localization, it is implemented in combination with the SmoothGrad algorithm.
[0039] M = g σ (Λ)
[0040]
[0041] where M represents the result of Gaussian smoothing of the gradient map Λ; B represents a binary map with an elliptical or circular shape; and respectively represent the morphological erosion and dilation operations using the structural element B, also known as morphological opening; L map represents the abnormal localization map after Gaussian smoothing and morphological opening of the gradient map Λ. This processing process also reduces the noise of the gradient map Λ.
[0042] The present invention also proposes a defect detection and localization system based on cross-image local feature alignment, including:
[0043] A distillation learning model construction module for constructing a distillation learning model, where the distillation learning model includes a teacher model T and a student model S;
[0044] A loss function determination module for determining the loss function of the distillation learning model;
[0045] A dataset construction and partitioning module for constructing a dataset and partitioning it into a training set and a test set. The training set only contains normal image samples, and the test set includes normal image samples and abnormal image samples;
[0046] An alignment loss function construction module for constructing a cross-image local feature alignment loss function among several normal image samples in the same training batch;
[0047] A total loss function construction module that integrates the local feature alignment loss function and the loss function of the distillation learning model to form the total loss function of the distillation learning model based on cross-image local feature alignment;
[0048] A distillation learning model training module that inputs the normal image samples in the training set into both the teacher model T and the student model S at the same time, keeps the parameters of the teacher model T fixed, uses the total loss function as the training guidance, and uses the teacher model T to guide the training of the student model S, thereby training the distillation learning model to obtain a trained distillation learning model;
[0049] A testing module that uses the image samples in the test set as the input samples of the trained distillation learning model, starts from the gradient of the total loss function with respect to the input samples, and uses the trained distillation learning model to perform defect detection and localization on the image samples in the test set.
[0050] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0051] The present invention proposes a method and system for defect detection and localization based on cross-image local feature alignment. First, a distillation learning model is built and the loss function is determined. Then, the cross-image local feature alignment information is incorporated into the loss function of the distillation learning model to form the total loss function. Using the total loss function as the training guidance, the distillation learning model under the total loss function is trained with the training set in the dataset. Finally, starting from the gradient of the total loss function with respect to the input samples, the trained distillation learning model is used to perform defect detection and localization on the image samples in the test set. In the process of this method, the distillation learning model can use the pixel-to-pixel local correspondence relationship across images in the training set to constrain the feature space, thereby ensuring the sensitivity of the distillation learning model to fine-grained local pixel-level features, and further improving the effect of object surface defect anomaly detection and localization. Description of the Drawings
[0052] Figure 1 Schematic flow chart of the defect detection and localization method based on cross-image local feature alignment proposed in Embodiment 1 of the present invention;
[0053] Figure 2 Schematic diagram showing the VGG16 network structure adopted by both the student model S and the teacher model T proposed in Embodiment 1 of the present invention;
[0054] Figure 3 Schematic block diagram showing the overall process of defect detection and localization based on cross-image local feature alignment proposed in Embodiment 2 of the present invention;
[0055] Figure 4 Structure diagram of the defect detection and localization system based on cross-image local feature alignment proposed in Embodiment 3 of the present invention. Detailed implementation manners
[0056] The accompanying drawings are only for illustrative purposes and should not be construed as limitations on this patent;
[0057] For better illustration of this embodiment, some parts of the accompanying drawings are omitted, enlarged or reduced, and do not represent actual sizes;
[0058] For those skilled in the art, it is understandable that some well-known content descriptions in the accompanying drawings may be omitted.
[0059] The terms used to describe the positional relationships in the accompanying drawings are only for illustrative purposes and should not be construed as limitations on this patent;
[0060] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0061] Embodiment 1
[0062] The present invention proposes a defect detection and localization method based on cross-image local feature alignment in Embodiment 1. The flow chart of the method is as Figure 1 shown and specifically includes the following steps:
[0063] S1. Construct a distillation learning model, where the distillation learning model includes a teacher model T and a student model S;
[0064] In this embodiment, in the constructed distillation learning model, the teacher model T adopts the VGG16 network structure loaded with pre-trained weights on ImageNet, and the student model S adopts the same VGG16 network structure as the teacher model T and randomly initializes the weights. For the specific network structure adopted by the teacher model T or the student model S, it is not limited to the VGG16 network structure and can also be other network structures. The student model S can also adopt a network structure similar to but not exactly the same as the teacher model T, but more concise and compact. On the basis of having constructed a clear distillation learning model, further determine the loss function of the distillation learning model, that is, execute step S2:
[0065] S2. Determine the loss function of the distillation learning model;
[0066] In this embodiment, as Figure 2 shown, the VGG16 network structure adopted by the teacher model T or the student model S includes several modules, Figure 2 The block structure in Figure 2 represents a module. In step S2, both the teacher model T and the student model S take the last layer of the last four modules in their respective network structures as their respective key layers. For specific reference, see
[0067] L1 = L val + λL dir
[0068] where,
[0069]
[0070]
[0071] where, L1 represents the loss function of the distillation learning model; CP i represents the output of the i-th key layer of the VGG16 network structure; CP0 represents the normal image features of the original input VGG16 network structure; represents the activation value of the i-th key layer of the teacher model T, and the activation value is the normal image features output by the key layer of the network structure; represents the activation value of the i-th key layer of the student model S; N i represents CP i the number of neurons in; represents the activation value of the j-th neuron in the i-th key layer of the teacher model T; represents the activation value of the j-th neuron in the i-th key layer of the student model S; N cp represents the total number of key layers; L valdenotes the sum of the Euclidean distances of the corresponding activation values in each key layer of the teacher model T and the student model S; vec() represents the vectorization function that converts a matrix with any dimension into a one-dimensional vector; L dir denotes the cosine similarity of the vectors converted from each corresponding key layer of the teacher model T and the student model S; λ represents the hyperparameter set manually.
[0072] After determining the loss function of the distillation learning model, in order to introduce cross-image local feature alignment information, for the introduction of cross-image local feature alignment information, it is necessary to start from the data set. Therefore, the following steps S3 and S4 are executed in sequence:
[0073] S3. Construct a data set, divide the data set into a training set and a test set. The training set only contains normal image samples, and the test set includes normal image samples and abnormal image samples.
[0074] In this embodiment, the constructed data sets are MVTecAD and Head-CT respectively. MVTecAD is an industrial quality inspection data set, which contains 15 categories of industrial products. Each category of products is divided into a training set and a test set. The training set only contains normal image samples (about 300 for each category), and the test set contains normal image samples and different types of abnormal image samples (about 30 for each category) and binary maps calibrated with the abnormal regions of the abnormal image samples. Head-CT is a medical data set, which contains 100 normal brain CTs and 100 brain CTs with lesions in this example.
[0075] S4. Construct the cross-image local feature alignment loss function among several normal image samples in the same training batch;
[0076] Suppose to construct the cross-image local feature alignment loss function for K normal image samples in the same training batch in step S4. During the training process of the distillation learning model, calculate the alignment loss of the activation value maps output by the key layers of the corresponding VGG16 network structure of the teacher model T and the student model S for the first, second,..., Kth normal image samples and the other K-1 normal image samples pixel by pixel in sequence, and then take 1 / 2 to eliminate the duplicate calculations, and obtain the cross-image local feature alignment loss function among K normal image samples in the same training batch. The expression is:
[0077]
[0078] Among them, L2 represents the cross-image local feature alignment loss function among K normal image samples in the same training batch; since the activation value map output by the key layer of the VGG16 network structure is used to calculate the alignment loss pixel by pixel across images, the activation value map is obtained by performing convolution calculation on the original input normal image samples using convolution kernels, and a pixel position in the activation value map is equivalent to the local features of at least 3*3 pixels in the original input normal image samples.
[0079] S5. Integrate the local feature alignment loss function in S2 with the loss function of the distillation learning model in S4 to form the total loss function of the distillation learning model based on cross-image local feature alignment;
[0080] The expression of the total loss function of the distillation learning model based on cross-image local feature alignment formed is:
[0081] L total = L1 + γL2
[0082] Among them, L total represents the total loss function of the distillation learning model based on cross-image local feature alignment; γ represents a hyperparameter set artificially.
[0083] On the premise that the framework of the distillation learning model is fixed, after forming the total loss function of the distillation learning model based on cross-image local feature alignment, using the total loss function as the training guidance, at this time, training the distillation learning model can enable the model to utilize the cross-image feature alignment information to ensure the sensitivity of the model to fine-grained local pixel-level features. The specific training process executes step S6:
[0084] S6. Input the normal image samples in the training set into the teacher model T and the student model S simultaneously, keep the parameters of the teacher model T fixed, use the total loss function as the training guidance, and use the teacher model T to guide the training of the student model S, thereby training the distillation learning model to obtain a trained distillation learning model;
[0085] When inputting the normal image samples in the training set into the teacher model T and the student model S simultaneously, keeping the parameters of the teacher model T fixed, and using the total loss function as the training guidance to use the teacher model T to guide the training of the student model S, the training method is backpropagation and gradient descent. The output of the key layer of the student model S is the image features of the normal image samples extracted by the student model S. Use the output of the key layer of the student model S to fit the output of the key layer of the teacher model T, so that the student model S only obtains the ability to extract the features of normal image samples. During the training process, when the total loss function L total no longer decreases in 20 training epochs, the model training is completed, that is, when the total loss function L totalWhen the value stabilizes at a state where it no longer decreases and converges, it also ensures the convergence of the cross-image local feature alignment loss function L2, ensuring the cross-image local feature consistency, that is, "alignment", and obtaining a trained distillation learning model.
[0086] S7. Use the image samples in the test set as the input samples of the trained distillation learning model. Starting from the gradient of the total loss function with respect to the input samples, use the trained distillation learning model to perform defect detection and localization on the image samples in the test set.
[0087] Overall, first build a distillation learning model, which includes a teacher model T and a student model S. Determine the loss function of the distillation learning model. Then construct a dataset and divide it into a training set and a test set. Incorporate the cross-image local feature alignment information in a number of normal image samples in the same training batch into the loss function of the distillation learning model to form the final loss function. Input the normal image samples into both the teacher model T and the student model S at the same time, keep the parameters of the teacher model T fixed, and use the teacher model T to guide the training of the student model S so that the student model only obtains the ability to extract features of normal image samples, and the model can use the pixel-pixel local correspondence relationship across images in the training set to constrain the feature space, and rely on the cross-image feature alignment information to ensure the sensitivity of the model to fine-grained local pixel-level features, thereby ensuring the effect of object surface defect detection and localization.
[0088] In this embodiment, in the last step S7, when starting from the gradient of the total loss function with respect to the input samples and using the trained distillation learning model to perform defect detection and localization on the image samples in the test set, let the input sample summarized in the test set be represented as x, and the expression of the gradient map Λ of the total loss function with respect to the input sample is
[0089]
[0090] The pixel gradient value of the input sample x is directly obtained through one backpropagation during the training process. Since the total loss function L of the distillation learning model based on cross-image local feature alignment total includes the loss function L1 of the distillation learning model and the cross-image local feature alignment loss function L2. When using the test set for testing in step S7, when the input sample x is a normal image sample, both the loss function L1 of the distillation learning model and the cross-image local feature alignment loss function L2 are small, and the gradient is also small. When the input sample x is an abnormal image sample, the opposite is true. Let the gradient contribution threshold be ε. The gradient value generated at the pixel position of the input sample x with a gradient contribution greater than ε is large, and the pixel position corresponds to the abnormal defect area, that is, using L totalPerform anomaly localization on the gradient map to obtain the anomalous defect area that causes its value to increase. That is, when using the trained distillation learning model to perform defect detection and localization on the image samples in the test set, it includes two levels. First, based on the total loss function L of the trained distillation learning model total of the value, during anomaly detection, determine whether the input sample is a normal image sample or an anomalous image sample. Second, the gradient of the total loss function L total with respect to the input sample x is of great significance. Among them, the pixels with larger contributing gradients are the areas where the output differences between the key layers of the teacher model T and the student model T and the local feature alignment loss are larger, and they are also the anomaly detection areas.
[0091] When using the gradient map of L total for anomaly localization, it is implemented in combination with the SmoothGrad algorithm. This algorithm is a commonly used algorithm for localization and will not be elaborated here.
[0092] To improve the accuracy of defect detection and localization, use Gaussian smoothing and morphological opening operation to process the gradient map Λ. The process formula is as follows:
[0093] M = g σ (Λ)
[0094]
[0095] Among them, M represents the result of Gaussian smoothing of the gradient map Λ; B represents a binary map with the shape of an ellipse or a circle; and respectively represent the morphological erosion and dilation operations using the structural element B, also known as morphological opening operation; L map represents the anomaly localization map after Gaussian smoothing and morphological opening operation on the gradient map Λ. This processing process also reduces the noise of the gradient map Λ.
[0096] Embodiment 2
[0097] In this embodiment, overall, the implementation process of the defect detection and localization method described in Embodiment 1 is further elaborated in the form of a schematic block diagram. See the schematic block diagram in Figure 3 , and the input samples of the same batch are nut image samples. As can be seen from Figure 3 , the nut image samples include normal image samples and defective anomalous image samples (such as the last one shown). The nut image samples are used as inputs and enter the teacher model T and the student model S of the distillation learning model. The alignment loss and the original key layer output loss are fused and summed to obtain the total loss function L total In the total loss function L totalInclude the loss function L1 (output of the key layer) of the distillation learning model and the local feature alignment loss function L2 (alignment loss) across images. Using the total loss function as the training guidance, train the entire distillation learning model through gradient descent and backpropagation. After the gradient is backpropagated, obtain the anomaly localization gradient heat map and the binary map of the anomaly region, thereby completing the defect anomaly detection and localization.
[0098] Embodiment 3
[0099] As Figure 3 shown, to implement the methods in Embodiment 1 and Embodiment 2, this embodiment proposes a defect detection and localization system based on cross-image local feature alignment, including:
[0100] A distillation learning model construction module 101, used to construct a distillation learning model, and the distillation learning model includes a teacher model T and a student model S;
[0101] A loss function determination module 102, used to determine the loss function of the distillation learning model;
[0102] A dataset construction and division module 103, used to construct a dataset, divide the dataset into a training set and a test set, the training set only contains normal image samples, and the test set includes normal image samples and abnormal image samples;
[0103] An alignment loss function construction module 104, used to construct the local feature alignment loss function across images in several normal image samples of the same training batch;
[0104] A total loss function construction module 105, which integrates the local feature alignment loss function and the loss function of the distillation learning model to form the total loss function of the distillation learning model based on cross-image local feature alignment;
[0105] A distillation learning model training module 106, input the normal image samples in the training set into the teacher model T and the student model S at the same time, keep the parameters of the teacher model T fixed, use the total loss function as the training guidance, and use the teacher model T to guide the student model S to train, thereby training the distillation learning model to obtain a trained distillation learning model;
[0106] A test module 107, use the image samples in the test set as the input samples of the trained distillation learning model, starting from the gradient of the total loss function with respect to the input samples, use the trained distillation learning model to perform defect detection and localization on the image samples in the test set.
[0107] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.
Claims
1. A method for defect detection and localization based on cross-image local feature alignment, characterized in that Including the following steps: S1. Construct a distillation learning model, where the distillation learning model includes a teacher model T and a student model S; S2. Determine the loss function of the distillation learning model; S3. Construct a dataset, divide the dataset into a training set and a test set. The training set only contains normal image samples, and the test set includes normal image samples and abnormal image samples; S4. Construct a cross-image local feature alignment loss function for several normal image samples in the same training batch; Suppose to construct the cross-image local feature alignment loss function for K normal image samples in the same training batch in step S4. During the training process of the distillation learning model, calculate the alignment loss of the activation value maps output by the key layers of the corresponding VGG16 network structures of the first, second,..., Kth normal image samples and the other K - 1 normal image samples pixel by pixel in turn, and then take 1 / 2 to eliminate the duplicate calculations, obtaining the cross-image local feature alignment loss function for K normal image samples in the same training batch. The expression is: Among them, L2 represents the cross-image local feature alignment loss function among K normal image samples in the same training batch; N cp represents the total number of key layers in the student model S; because the activation value maps output by the key layers of the VGG16 network structure are used to calculate the alignment loss pixel by pixel across images, the activation value maps are obtained by performing convolution calculations on the original input normal image samples using convolutional kernels, and a pixel position in the activation value map is equivalent to the local features of at least 3*3 pixels in the original input normal image sample; S5. Integrate the local feature alignment loss function in S4 with the loss function of the distillation learning model in S2 to form the total loss function of the distillation learning model based on cross-image local feature alignment; the expression of the total loss function of the distillation learning model based on cross-image local feature alignment formed in step S5 is: L total = L1 + γL2 Among them, L1 represents the loss function of the distillation learning model; L total represents the total loss function of the distillation learning model based on cross-image local feature alignment; γ represents a hyperparameter set manually; S6. Input the normal image samples in the training set into the teacher model T and the student model S at the same time. Keep the parameters of the teacher model T fixed. Using the total loss function as the training guidance, use the teacher model T to guide the training of the student model S, thereby training the distillation learning model to obtain a trained distillation learning model; S7. Use the image samples in the test set as the input samples of the trained distillation learning model. Starting from the gradient of the total loss function with respect to the input samples, use the trained distillation learning model to perform defect detection and localization on the image samples in the test set.
2. The method for defect detection and localization based on cross-image local feature alignment according to claim 1, characterized in that, In the distillation learning model constructed in step S1, the teacher model T uses the VGG16 network structure loaded with pre-trained weights on ImageNet, and the student model S uses the same VGG16 network structure as the teacher model T and randomly initializes the weights.
3. The method for defect detection and localization based on cross-image local feature alignment according to claim 2, wherein The VGG16 network structure includes several modules. In step S2, both the teacher model T and the student model S take the last layer of the last four modules in their respective network structures as their key layers. The loss function of the distillation learning model is: L1 = L val + λL dir Where, Among them, L1 represents the loss function of the distillation learning model; CP i represents the output of the i-th key layer of the VGG16 network structure; represents the activation value of the i-th key layer of the teacher model T, and the activation value is the normal image feature output by the key layer of the network structure; represents the activation value of the i-th key layer of the student model S; N i represents CP i represents the number of neurons in represents the activation value of the j-th neuron in the i-th key layer of the teacher model T; represents the activation value of the j-th neuron in the i-th key layer of the student model S; N cp represents the total number of key layers; L val represents the sum of the Euclidean distances of the corresponding activation values in each key layer of the teacher model T and the student model S; vec() represents the vectorization function that converts a matrix with any dimension into a one-dimensional vector; L dir represents the cosine similarity of the vectors converted from each corresponding key layer of the teacher model T and the student model S; λ represents the hyperparameter set artificially.
4. The method for defect detection and localization based on cross-image local feature alignment according to claim 1, characterized in that In step S6, the normal image samples in the training set are input into the teacher model T and the student model S simultaneously. Keeping the parameters of the teacher model T fixed, and using the total loss function as the training guidance, when using the teacher model T to guide the training of the student model S, the training method is backpropagation and gradient descent. The output of the key layer of the student model S is the image features of the normal image samples extracted by the student model S. The output of the key layer of the student model S is used to fit the output of the key layer of the teacher model T, so that the student model S only obtains the ability to extract the features of normal image samples. During the training process, the total loss function L total When it no longer decreases in 20 training epochs, the model training is completed, and the trained distillation learning model is obtained.
5. The method for defect detection and localization based on cross-image local feature alignment according to claim 4, wherein In step S7, when starting from the gradient of the total loss function with respect to the input samples and using the trained distillation learning model to perform defect detection and localization on the image samples in the test set, assume that the input sample summarized in the test set is represented as x, and obtain the gradient map Λ of the total loss function with respect to the input sample. The expression is The pixel gradient value of the input sample x is directly obtained through a single backpropagation during the training process. Since the total loss function L of the distillation learning model based on cross-image local feature alignment total includes the loss function L1 of the distillation learning model and the local feature alignment loss function L2 across images. When using the test set for testing in step S7, when the input sample x is a normal image sample, both the loss function L1 of the distillation learning model and the local feature alignment loss function L2 across images are small, and the gradient is also small. When the input sample x is an abnormal image sample, the opposite is true. Let the gradient contribution threshold be ε. The gradient value generated at the pixel position of the input sample x with a gradient contribution greater than ε is large, and the pixel position corresponds to the abnormal defect area, that is, using L total 's gradient map for anomaly localization to obtain the abnormal defect area that causes its value to increase.
6. The method for defect detection and localization based on cross-image local feature alignment according to claim 5, characterized in that When performing anomaly localization using the gradient map of L total , it is implemented in combination with the SmoothGrad algorithm.
7. The method for defect detection and localization based on cross-image local feature alignment according to claim 5 or 6, characterized in that To improve the accuracy of defect detection and localization, use Gaussian smoothing and morphological opening operations to process the gradient map Λ. The process formula satisfied is: M = g σ (Λ) Among them, M represents the result of Gaussian smoothing of the gradient map Λ; B represents a binary map with an elliptical or circular shape; and respectively represent morphological erosion and dilation operations using the structure element B, also known as morphological opening operation; L map represents the anomaly localization map after Gaussian smoothing and morphological opening operation on the gradient map Λ.
8. A defect detection and positioning system based on cross-image local feature alignment, characterized in that Including: A distillation learning model construction module for constructing a distillation learning model, where the distillation learning model includes a teacher model T and a student model S; A loss function determination module for determining the loss function of the distillation learning model; The dataset construction and division module is used to construct a dataset and divide it into a training set and a test set. The training set only contains normal image samples, and the test set includes normal image samples and abnormal image samples; The alignment loss function construction module is used to construct a cross-image local feature alignment loss function for several normal image samples in the same training batch; Construct a cross-image local feature alignment loss function for K normal image samples in the same training batch. During the training process of the distillation learning model, calculate the alignment loss of the activation value maps output by the key layers of the corresponding VGG16 network structures of the teacher model T and the student model S for the first, second, …, Kth normal image samples and the other K - 1 normal image samples pixel by pixel in turn. Then take 1 / 2 to eliminate the duplicate calculations, and obtain the cross-image local feature alignment loss function for K normal image samples in the same training batch. The expression is: where, L2 represents the cross-image local feature alignment loss function for K normal image samples in the same training batch; because the activation value maps output by the key layers of the VGG16 network structure are used to calculate the alignment loss pixel by pixel across images, the activation value maps are obtained by performing convolution calculations on the original input normal image samples using convolution kernels. A pixel position in the activation value map is equivalent to the local feature of at least 3*3 pixels in the original input normal image sample; The total loss function construction module integrates the local feature alignment loss function and the loss function of the distillation learning model to form the total loss function of the distillation learning model based on cross-image local feature alignment; the expression of the total loss function of the distillation learning model based on cross-image local feature alignment formed is: L total = L1 + γL2 Among them, L total represents the total loss function of the distillation learning model based on cross-image local feature alignment; γ represents a hyperparameter set manually; The distillation learning model training module inputs the normal image samples in the training set into the teacher model T and the student model S at the same time, keeps the parameters of the teacher model T fixed, uses the total loss function as the training guidance, and uses the teacher model T to guide the training of the student model S, so as to train the distillation learning model and obtain a trained distillation learning model; The testing module uses the image samples in the test set as the input samples of the trained distillation learning model, starts from the gradient of the total loss function with respect to the input samples, and uses the trained distillation learning model to detect and locate defects in the image samples in the test set.
Citation Information
Patent Citations
Knowledge distillation-based positive sample industrial defect detection method
CN112991330A
Surface defect detection method and apparatus, and electronic device
WO2019233166A1