A lightweight industrial environment defect detection method based on a feature memory bank

By adopting semi-supervised feature knowledge distillation method and feature memory bank optimization technology in industrial environments, the problems of resource consumption and inference time in industrial scenarios are solved, and efficient detection of defects in lightweight industrial environments is achieved.

CN115690058BActive Publication Date: 2025-06-03FUZHOU UNIV

Patent Information

Application Number
CN202211380109.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-04
Publication Date
2025-06-03
Estimated Expiration
2042-11-04

AI Technical Summary

Technical Problem

The existing anomaly detection technology has problems such as model size expansion, long inference time and surge in resource consumption in industrial scenarios, which is difficult to meet the industrial production needs of shortage of computing resources.

Method used

The semi-supervised feature knowledge distillation method is used to pass the pre-trained large network knowledge to the lightweight designed student network, and the feature memory bank method is processed to establish a lightweight industrial environment defect detection method.

Benefits of technology

On the premise of not losing accuracy, the efficiency of abnormal detection is improved, the model resource occupation and inference time are reduced, and it is suitable for industrial scenarios with tight computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115690058B_ABST
    Figure CN115690058B_ABST
Patent Text Reader

Abstract

The present invention proposes a lightweight industrial environment defect detection method based on a feature memory bank. First, the top view / left view or right view of qualified products on the production line is photographed, preprocessed after being sorted into a data set. Secondly, the data set is input into a pre-trained standard teacher model and an initialized lightweight student model to perform feature knowledge distillation training on the student model. After that, the defect-free product data set is input into the lightweight student model trained by feature knowledge distillation. After extracting the corresponding feature set of the product image and performing multi-scale feature set fusion, maximum core set distance sampling is carried out to establish a feature memory bank of qualified products on the current production line. Finally, the production line images to be detected are input into the lightweight student model and multi-scale feature fusion is performed, and then compared with the feature memory bank by the nearest neighbor algorithm to detect whether the current product contains defects. The present invention can improve the efficiency of anomaly detection without sacrificing accuracy as much as possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a lightweight industrial environment defect detection method based on a feature memory bank. Background Art

[0002] Anomaly detection is also known as Anomaly Detection (AD). The main task of anomaly detection is to assume that a given image set is defined as a "normal set", and a classification model is trained using the "normal set" so that it can accurately classify "normal" images and "abnormal" images during the test process. Currently, the mainstream solutions to the anomaly detection (AD) problem are roughly divided into two types. One is to train a model using the "normal set" and directly output the classification probability, and the other solution is to use a feature memory bank to "remember" the features of the "normal set" and compare them with the features of the input image during the test to obtain the classification probability.

[0003] In recent years, although the model has had a relatively large improvement in detection accuracy, there are still problems such as the expansion of the model size, the increase in inference time, and the sharp increase in resource consumption. It is not suitable for industrial scenarios with a huge demand for anomaly detection technology. For industrial production scenarios with tight computing resources, improving the model running speed and reducing the model resource occupancy while minimizing the loss of accuracy are also important development directions for anomaly detection.

[0004] To solve the above problems, the present invention proposes a semi-supervised feature knowledge distillation method to transfer the knowledge of a pre-trained large network to a lightweight designed student network, and optimizes the process of the feature memory bank method so that it can improve the efficiency of anomaly detection while minimizing the loss of accuracy. Summary of the Invention

[0005] The present invention proposes a lightweight industrial environment defect detection method based on a feature memory bank, which can improve the efficiency of anomaly detection while minimizing the loss of accuracy.

[0006] The present invention adopts the following technical solutions.

[0007] A lightweight industrial environment defect detection method based on a feature memory bank includes the following steps:

[0008] Step S1: Use an industrial camera installed on the production line in cooperation with a fixed diffused light source to take a top view photo / left view photo or right view photo of the qualified products passing through the production line, and preprocess the photos after organizing them into a data set;

[0009] Step S2: Input the data set into a pre-trained standard teacher model and an initialized lightweight student model, and perform semi-supervised feature knowledge distillation training on the student model;

[0010] Step S3: Input the defect-free product dataset into the lightweight student model trained by feature knowledge distillation. After extracting the corresponding feature set of the product image and performing multi-scale feature set fusion, perform maximum core set distance sampling to establish a feature memory bank for qualified products of the current product line;

[0011] Step S4: Input the production line images to be detected into the lightweight student model and perform multi-scale feature fusion, and then compare them with the feature memory bank using the nearest neighbor algorithm to detect whether the current product contains defects.

[0012] The specific steps of the said Step S1 include the following steps:

[0013] Step S11: Install an industrial camera and a fixed diffused light source on the production line;

[0014] Step S12: Take the top view / left view or right view of normal products pre-screened manually on the production line in the production environment and aggregate them into the dataset of the current products;

[0015] Step S13: Perform data preprocessing on all the pictures in the dataset. The data preprocessing includes image size adjustment, image center cutting, and image normalization.

[0016] The specific steps of the said Step S2 include the following steps:

[0017] Step S21: Input the preprocessed dataset into the standard teacher model that has been pre-trained using other public databases, and extract the first, second, and third stage feature maps output by the standard teacher model, and merge them into the teacher feature map set of this dataset, called Feature teacher ;

[0018] Step S22: Input the preprocessed dataset into the initialized student model, and extract the first, second, and third stage feature maps output by the student model, and merge them into the student feature map set of this dataset, called Feature studen t ;

[0019] Step S23: Respectively expand the number of image channels in the first, second, and third stages of Feature student through the mapping layer to make it the same as the number of image channels in the first, second, and third stages of Feature teacher . The expanded student feature map set is called Feature expand_student ;

[0020] Step S24: Compare Feature teacher with Feature expand_studentThe input loss function calculates the multi-layer feature loss, and respectively associates the Feature teacher with the Feature student to the input layer correlation score calculation function to obtain the layer correlation score. After subtracting the two scores, the layer correlation loss is obtained. After adding the two losses, backpropagation is performed for knowledge distillation training from the teacher model to the student model.

[0021] The specific method for calculating the multi-layer feature correlation score is as follows: Expand Feature studdent through the mapping layer into Feature expand_student , and calculate the MSE loss for the feature maps of the first, second, and third stages in Feature expand_stydent with the feature maps of the first, second, and third stages of Fetature teacher respectively. Finally, the losses of the feature maps of the three stages are added up to obtain the total loss value. The calculation method is as follows:

[0022]

[0023] where Feature expand_student represents the set of expanded student feature maps, represents the feature map output at the i-th stage and expanded, Feature teacher represents the set of teacher feature maps, represents the feature map output at the i-th stage, Loss mse represents the calculated MSE loss value, and n represents the number of feature maps extracted from the teacher model and the student model respectively, with the value being 3 here.

[0024] The specific method for calculating the layer correlation loss is as follows: For each pair of feature maps of adjacent stages in Feature teacher , perform adaptive max pooling to the same length and width and then flatten them, and multiply the flattened matrices to obtain the correlation score between the feature maps. For Feature student , the layer correlation score calculation formula is the same. The calculation method is as follows:

[0025]

[0026]

[0027]

[0028]

[0029] where AdaptiveMaxPooling (H*W) represents the adaptive max pooling operation, which adaptively samples the feature map to height H and width W. Denote the feature map output at the i-th stage, Denote the feature map after pooling at the i-th stage. Flatten represents the operation of flattening the feature map, which converts the length and width dimensions of the feature map to the same dimension. Denote the feature map after pooling and flattening at the i-th stage. Transpose represents the transpose operation, and Matmul represents the matrix multiplication operation. Denote the FSP scores of the t-th group. s represents the upper limit of the number of groups for FSP scores, and its value should be 2 here. Score FSP Denote the total FSP score.

[0030] The step S3 includes the following steps;

[0031] Step S31: Input the defect-free product dataset into the lightweight student model trained by knowledge distillation and extract the first, second, and third stage feature map sets of its output, namely Feature student , and similarly input the defect-free product dataset into the teacher model and extract the first, second, and third stage feature map sets of its output, namely Feature teacher ;

[0032] Step S32: Perform efficient optimization multi-scale feature map fusion on the feature map set in Feature student , and the synthesized new feature map set is called Feature reshape ;

[0033] Step S33: For Feature reshape , use the maximum core set distance sampling algorithm for sampling, and the sampling result is the feature memory bank of the current dataset, which is called MemoryBank.

[0034] The specific method of the efficient optimization multi-scale feature map fusion is as follows:

[0035] Downsample the first stage feature map and upsample the third stage feature map, so that the length and width of the feature maps of the three stages are the same. Then perform average pooling operation on all feature maps, and then connect the feature maps of the three stages in the channel dimension to synthesize a new feature map. Finally, deform the new feature map to obtain the final feature memory bank. The specific algorithm is as follows:

[0036]

[0037] Feature pooling =AveragePooling (3*3) (Feature align ) Formula Seven;

[0038] Feature concat = Concat(Feature pooling) Formula VIII;

[0039] Feature reshape = Reshape (B*H*W,C) (Feature concat ) Formula IX;

[0040] Wherein represents the feature map of the i-th stage in Feature student , DownSample (H*W) represents downsampling the feature map to height H and width W, UpSample (H*W) represents upsampling the feature map to height H and width W, Feature align represents the set of feature maps after aligning the length and width, AveragePooling (3*3) represents the average pooling operation with a core length and width of 3*3, Feature pooling represents the set of feature maps after the average pooling operation, Concat represents the channel dimension connection operation of the set of feature maps, fusing multiple feature maps into a new feature map, that is, Feature concat , Reshape (B*H*W,C) represents reshaping the feature map with the shape of (B, H, W, C) into (B*H*W, C), where B represents the sample size, H represents the height, W represents the width, Feature reshape represents the reshaped feature map.

[0041] The step S4 includes the following steps;

[0042] Step S41: The product image dataset to be detected also goes through step S31 and step S32 to obtain Feature reshape ;

[0043] Step S42: Perform a weighted K-nearest neighbor search on each pixel block in Feature reshape with the feature memory bank obtained in step S33 to obtain the distance score Score reshape of each pixel block in Feature distance ;

[0044] Step S43: Take the maximum value in Score distance as the anomaly score of the input image; Take Score distanceUpsampled to the length and width of the input image, that is, the anomaly score of each pixel in the input image. The higher the anomaly score, the higher the possibility that the image or the pixel contains an anomaly.

[0045] The specific method of the weighted K-nearest neighbor search is:

[0046] First, Feature reshape Perform Euclidean distance calculation with MemoryBank, take the smallest K distance values, and perform weighted sum on these K values. The weighted value of the first distance value is 0.5, and the weighted value of the xth distance value is half of the weighted value of the x-1th distance value. The specific calculation is as follows:

[0047]

[0048] MinDistance=Min K (Distance) Formula XI;

[0049]

[0050] Among them, Distance represents MemoryBank and Feature reshape The Euclidean distance matrix, Min K It means listing the K minimum values ​​of each row in the matrix and outputting the result as MinDistance, MinDistance x Indicates the xth distance value from small to large, p x Represents weighted values, the first weighted value is 0.5, and the xth weighted value is half of the x-1th weighted value.

[0051] Compared with the prior art, the present invention has the following beneficial effects:

[0052] 1. Compared with the existing feature memory library method, the present invention updates the fast feature fusion operation and simplifies the feature memory library while maintaining the accuracy as much as possible, thereby improving the reasoning speed of the method and making it more suitable for industrial scenarios;

[0053] 2. The present invention innovatively uses a feature knowledge distillation method. Different from the traditional knowledge distillation method, the present invention only uses the "normal set" data, and by calculating the multi-layer feature loss and layer association loss of the output features of the teacher model and the output features of the student model, it can transfer the knowledge of the pre-trained teacher model to the student model that has only been initialized, while only requiring "thousands" of data;

[0054] 3. Inspired by the EfficientNet model, the present invention reconstructs and trains a lightweight backbone network, resulting in lower occupancy rates of system resources such as disks, video memories, and memories compared to existing methods. It is more suitable for industrial scenarios and reduces the cost of computing resources. Description of the Drawings

[0055] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:

[0056] Att Figure 1 is a flowchart of the present invention. Specific Embodiments

[0057] As Figure 1 shown, a lightweight industrial environment defect detection method based on a feature memory bank includes the following steps:

[0058] Step S1: Use an industrial camera installed on the production line in cooperation with a fixed diffused light source to take a top view photo / left view photo or right view photo of a qualified product passing through the production line, and after organizing the photos into a dataset, perform preprocessing;

[0059] Step S2: Input the dataset into a pre-trained standard teacher model and an initialized lightweight student model, and perform semi-supervised feature knowledge distillation training on the student model;

[0060] Step S3: Input the defect-free product dataset into the lightweight student model trained with feature knowledge distillation, extract the corresponding feature set of the product image, perform multi-scale feature set fusion, and then perform maximum core set distance sampling to establish a feature memory bank of qualified products on the current production line;

[0061] Step S4: Input the production line image to be detected into the lightweight student model and perform multi-scale feature fusion, and then compare it with the feature memory bank using the nearest neighbor algorithm to detect whether the current product contains defects.

[0062] The specific steps of step S1 include the following steps:

[0063] Step S11: Install an industrial camera and a fixed diffused light source on the production line;

[0064] Step S12: Take a top view / left view or right view of a normal product pre-screened by humans on the production line of the production environment, and aggregate them into the dataset of the current product;

[0065] Step S13: Perform data preprocessing on all the pictures in the dataset. The data preprocessing includes image size adjustment, image center cutting, and image standardization.

[0066] The specific steps of step S2 include the following steps:

[0067] Step S21: Input the preprocessed dataset into a standard teacher model that has been pre-trained using other public databases, and extract the first, second, and third stage feature maps output by the standard teacher model, and merge them into a set of teacher feature maps for this dataset, called Feature teacher ;

[0068] Step S22: Input the preprocessed dataset into the initialized student model, and extract the first, second, and third stage feature maps output by the student model, and merge them into a set of student feature maps for this dataset, called Feature student ;

[0069] Step S23: Respectively expand the number of image channels in the first, second, and third stages of Feature student through the mapping layer so that they are the same as the number of image channels in the first, second, and third stages of Feature teacher . The expanded set of student feature maps is called Feature expand_student ;

[0070] Step S24: Input Feature teacher and Featuree xpend_student into the loss function to calculate the multi-layer feature loss. Respectively input Feature teacher and Feature student into the layer association score calculation function to obtain the layer association scores. After subtracting the two scores, the layer association loss is obtained. After adding the two losses, backpropagation is performed for knowledge distillation training from the teacher model to the student model.

[0071] The specific method for calculating the multi-layer feature association score is as follows: Expand Feature student into Feature expand_student through the mapping layer, and calculate the MSE loss between the first, second, and third stage feature maps in Feature expand_student and the first, second, and third stage feature maps in Feature teacher respectively. Finally, the losses of the three stage feature maps are added up to obtain the total loss value. The calculation method is as follows:

[0072]

[0073] where Feature expand_student represents the expanded set of student feature maps, represents the feature map output at the i-th stage and expanded, Feature teacher represents the set of teacher feature maps, represents the feature map output at the i-th stage, Loss mseIt represents the calculated MSE loss value, and n represents the number of feature maps extracted from the teacher model and the student model respectively. Here, the value is 3.

[0074] The specific calculation method of the layer correlation loss is as follows: For each pair of feature maps in adjacent stages in Feature teacher perform adaptive max pooling to the same length and width, then flatten them, and multiply the flattened matrices to obtain the correlation scores between the feature maps. For Feature student , the calculation formula of the layer correlation score is the same, and the calculation method is as follows:

[0075]

[0076]

[0077]

[0078]

[0079] Among them, AdaptiveMaxPooling (H*W) represents the adaptive max pooling operation, which adaptively samples the feature map to height H and width W. represents the feature map output in the i-th stage. represents the feature map after pooling in the i-th stage. Flatten represents the operation of flattening the feature map, which converts the length and width dimensions of the feature map to the same dimension. represents the feature map after pooling and flattening in the i-th stage. Transpose represents the transpose operation, and Matmul represents the matrix multiplication operation. represents the FSP score of the t-th group, s represents the upper limit of the number of FSP score groups, and the value should be 2 here. Score FSP represents the total FSP score.

[0080] Step S3 includes the following steps;

[0081] Step S31: Input the defect-free product dataset into the lightweight student model trained by knowledge distillation and extract the first, second, and third stage feature map sets output by it, that is, Feature student , and similarly input the defect-free product dataset into the teacher model and extract the first, second, and third stage feature map sets output by it, that is, Feature teacher ;

[0082] Step S32: Perform efficient optimization multi-scale feature map fusion on the feature map sets in Feature student , and the synthesized new feature map set is called Feature reshape ;

[0083] Step S33: For Feature reshape , sample using the maximum core set distance sampling algorithm, and the sampling result is the feature memory bank of the current data set, which is called MemoryBank.

[0084] The specific efficiency-optimized multi-scale feature map fusion method is as follows:

[0085] Downsample the first-stage feature map and upsample the third-stage feature map so that the lengths and widths of the feature maps in the three stages are the same. Then perform average pooling on all feature maps. After that, concatenate the feature maps in the three stages in the channel dimension to synthesize a new feature map. Finally, deform the new feature map to obtain the final feature memory bank. The specific algorithm is as follows:

[0086]

[0087] Feature pooling = AveragePooling (3*3) (Feature align ) Formula Seven;

[0088] Feature cencat = Concat(Feature pooling ) Formula Eight;

[0089] Feature reshape = Reshape (B*H*W,C) (Feature concat ) Formula Nine;

[0090] Where represents the feature map of the i-th stage in Feature student , DownSample (H*W) represents downsampling the feature map to height H and width W, UpSample (H*W) represents upsampling the feature map to height H and width W, Feature align represents the set of feature maps with aligned lengths and widths, AveragePooling (3*3) represents the average pooling operation with a core length and width of 3*3, Feature pooling represents the set of feature maps after the average pooling operation, Concat represents the operation of concatenating the set of feature maps in the channel dimension to fuse multiple feature maps into a new feature map, that is, Feature concat , Reshape (B*H*W,C)Denote the feature map with shape (B, H, W, C) being reshaped into (B*H*W, C), where B represents the number of samples, H represents the height, W represents the width, and Feature reshape Denote the reshaped feature map.

[0091] The step S4 includes the following steps;

[0092] Step S41: The product image dataset to be detected is also processed through step S31 and step S32 to obtain Feature reshape ;

[0093] Step S42: Perform weighted K-nearest neighbor search on each pixel block in Feature reshape with the feature memory bank obtained in step S33 to obtain the distance score Score reshape for each pixel block in Feature distance ;

[0094] Step S43: Take the maximum value in Score distance as the anomaly score of the input image; Upsample Score distance to the length and width of the input image, which is the anomaly score of each pixel point in the input image. The higher the anomaly score, the higher the possibility that the image or the pixel point contains an anomaly.

[0095] The specific method of the weighted K-nearest neighbor search is as follows:

[0096] First, calculate the Euclidean distance between Feature reshape and MemoryBank, take the smallest K distance values, and perform a weighted sum on these K values. The weight value of the first distance value is 0.5, and the weight value of the x-th distance value is half of the weight value of the (x - 1)-th distance value. The specific calculation is as follows:

[0097]

[0098] MinDistance = Min K (Distance) Formula XI;

[0099]

[0100] where Distance represents the Euclidean distance matrix between MemoryBank and Feature reshape , Min K represents listing the K values of the minimum values in each row of the matrix and outputting the result as MinDistance, MinDistance x represents the x-th distance value from smallest to largest, and p xIndicates the weighted value. The first weighted value is 0.5, and the x-th weighted value is half of the (x - 1)-th weighted value.

[0101] The above are only the preferred embodiments of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by the present invention.

Claims

1. A lightweight industrial environment defect detection method based on a feature memory bank, characterized in that: It includes the following steps: Step S1: Use an industrial camera installed on the assembly line in cooperation with a fixed diffused light source to take a top view photo or a left view photo or a right view photo of a qualified product passing through the assembly line, and preprocess the photos after organizing them into a data set; Step S2: Input the data set into a pre-trained standard teacher model and an initialized lightweight student model, and perform semi-supervised feature knowledge distillation training on the student model; Step S3: Input the defect-free product data set into the lightweight student model trained by feature knowledge distillation, extract the corresponding feature set of the product image, perform multi-scale feature set fusion, and then perform maximum core set distance sampling to establish a feature memory bank of qualified products on the current production line; Step S4: Input the production line image to be detected into the lightweight student model and perform multi-scale feature fusion, and then compare it with the feature memory bank using the nearest neighbor algorithm to detect whether the current product contains defects; The specific steps of the above Step S2 include the following steps: Step S21: Input the preprocessed dataset into a standard teacher model that has been pre-trained using other publicly available databases, and extract the feature maps of the first, second, and third stages output by the standard teacher model, and merge them into a set of teacher feature maps for this dataset, called Feature teacher ; Step S22: Input the preprocessed dataset into the initialized student model, extract the feature maps of the first, second, and third stages output by the student model, and merge them into a set of student feature maps for this dataset, called Feature student ; Step S23: Take Feature student and separately expand the number of image channels in the first, second, and third stages through the mapping layer so that they are the same as the number of image channels in the first, second, and third stages of Feature teacher . Denote the expanded set of student feature maps as Feature expand_student ; Step S24: Take Feature teacher and Feature expand_student as inputs to the loss function to calculate the multi-layer feature loss. Separately, take Feature teacher and Feature student as inputs to the layer association score calculation function to obtain the layer association scores. Subtract the two scores to get the layer association loss. Add the two losses together and perform backpropagation to conduct knowledge distillation training from the teacher model to the student model; The specific method of the layer correlation score calculation function is as follows: Feature student is extended into Feature expand_student through the mapping layer, and the feature maps of the first, second, and third stages in Feature expand_student are respectively used to calculate the MSE loss with the feature maps of the first, second, and third stages of Feature teacher . Finally, the losses of the feature maps of the three stages are summed to obtain the total loss value. The calculation method is as follows: Among them, Feature expand_student represents the set of extended student feature maps, represents the feature map output at the i-th stage and extended, Feature teacher represents the set of teacher feature maps, represents the feature map output at the i-th stage, Loss mse represents the calculated MSE loss value, and n represents the number of feature maps extracted from the teacher model and the student model respectively, where the value is 3.

2. A lightweight industrial environment defect detection method based on a feature memory bank according to claim 1, characterized in that: The specific steps of the above Step S1 include the following steps: Step S11: Install an industrial camera and a fixed diffused light source on the assembly line; Step S12: Take a top view or a left view or a right view of a normal product pre-screened manually on the assembly line in the production environment, and aggregate them into a data set of the current product; Step S13: Perform data preprocessing on all the pictures in the data set. The data preprocessing includes image size adjustment, image center cutting, and image standardization.

3. A lightweight industrial environment defect detection method based on a feature memory bank according to claim 1, characterized in that: The specific calculation method of the layer association loss is as follows: For each pair of feature maps in adjacent stages in Feature teacher , perform adaptive max pooling to the same length and width, then flatten them, and multiply the flattened matrices to obtain the association scores between the feature maps. For Feature student , the calculation formula of the layer association score is the same, and the calculation method is as follows: Among them, AdaptiveMaxPooling (H*W) represents an adaptive max pooling operation that adaptively samples the feature map to a height H and a width W. represents the feature map output at the i-th stage. represents the feature map after pooling at the i-th stage. Flatten represents an operation to flatten the feature map, converting the two dimensions of length and width of the feature map to the same dimension. represents the feature map after pooling and flattening at the i-th stage. Transpose represents a transpose operation, and Matmul represents a matrix multiplication operation. represents the FSP score of the t-th group. s represents the upper limit of the number of groups for the FSP score, and the value should be 2 here. Score FSP represents the total FSP score.

4. A lightweight industrial environment defect detection method based on a feature memory bank according to claim 1, characterized in that: The above Step S3 includes the following steps; Step S31: Input the defect-free product dataset into the lightweight student model trained by knowledge distillation and extract the first, second, and third stage feature map sets output by it, namely Feature student , similarly, input the defect-free product dataset into the teacher model and extract the first, second, and third stage feature map sets output by it, namely Feature teacher ; Step S32: Perform efficiency-optimized multi-scale feature map fusion on the set of feature maps in Feature student , and the synthesized new set of feature maps is called Feature reshape ; Step S33: For Feature reshape , sample using the maximum core set distance sampling algorithm, and the sampling result is the feature memory bank of the current dataset, which is called MemoryBank.

5. A lightweight industrial environment defect detection method based on a feature memory bank according to claim 4, characterized in that: The specific method for optimizing the efficiency of multi-scale feature map fusion is as follows: Downsample the first-stage feature map and upsample the third-stage feature map, so that the length and width of the three-stage feature maps are the same, then perform average pooling operation on all the feature maps, and then connect the three-stage feature maps in the channel dimension to synthesize a new feature map. Finally, deform the new feature map to obtain the final feature memory bank; The specific algorithm is as follows: Feature pooling = AveragePooling (3*3) (Feature align ) Formula VII; Feature concat = Concat(Feature pooling ) Formula VIII; Feature reshape = Reshape (B*H*W,C) (Feature concat ) Formula Nine; Among them represents the feature map in the i-th stage of Feature student , DownSample (H*W) represents downsampling the feature map to height H and width W, UpSample (H*W) represents upsampling the feature map to height H and width W, Feature align represents the set of feature maps after aligning the length and width, AveragePooling (3*3) represents the average pooling operation with a core length and width of 3*3, Feature pooling represents the set of feature maps after the average pooling operation, Concat represents the channel dimension connection operation on the set of feature maps, fusing multiple feature maps into a new feature map, that is, Feature concat , Reshape (B*H*W,C) represents reshaping the feature map with the shape of (B, H, W, C) into (B*H*W, C), where B represents the sample size, H represents the height, W represents the width, Feature reshape represents the reshaped feature map.

6. A lightweight industrial environment defect detection method based on a feature memory bank according to claim 4, characterized in that: The above Step S4 includes the following steps; Step S41: The product image dataset to be detected also undergoes Step S31 and Step S32 to obtain Feature reshape ; Step S42: Weight each pixel block in Feature reshape with the feature memory bank obtained in step S33 to perform weighted K-nearest neighbor search, and obtain the distance score Score reshape for each pixel block in Feature distance ; Step S43: Score distance Take the maximum value among them as the anomaly score of the input image in Step S11; Upsample Score distance to the length and width of the input image, which is the anomaly score of each pixel in the input image. The higher the anomaly score, the higher the possibility that the image or the pixel contains anomalies.

7. A lightweight industrial environment defect detection method based on a feature memory bank according to claim 6, characterized in that: The specific method of the weighted K-nearest neighbor search is as follows: First, calculate the Euclidean distance between Feature reshape and the MemoryBank, and take the smallest K distance values. Then, perform a weighted sum on these K values. The weighting value of the first distance value is 0.5, and the weighting value of the x-th distance value is half of the weighting value of the (x - 1)-th distance value. The specific calculation is as follows: MinDistance = Min K (Distance) Formula XI; Among them, Distance represents the Euclidean distance matrix between MemoryBank and Feature reshape , Min K represents listing the K smallest values in each row of the matrix, and the result is output as MinDistance. MinDistance x represents the x-th distance value from the smallest to the largest, p x represents the weighting value. The first weighting value is 0.5, and the x-th weighting value is half of the (x - 1)-th weighting value.

Citation Information

Patent Citations

  • Knowledge distillation method based on semantic segmentation intra-class feature difference

    CN111062951A

  • Industrial image defect detection method based on knowledge distillation

    CN114663392A

Cited By

  • Unsupervised defect detection method based on memory bank reconstruction

    CN117274215A