Industrial defect classification detection method and device, medium and electronic equipment

Through cascading architecture design and knowledge distillation technology, lightweight neural network is built on edge devices, combined with server-side models, real-time high-precision screening of industrial defect detection is achieved, solving the problem of limited resources of edge devices and reducing hardware costs.

CN120374628AActive Publication Date: 2025-07-25JIANGXI NORMAL UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510874493.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-25
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to achieve real-time and high-precision requirements for industrial defect detection on edge devices with limited computing power and storage resources, and at the same time there is a problem of resource waste.

Method used

The cascade architecture design is adopted, and the edge-end defect classification model is constructed using lightweight neural networks and knowledge distillation technology. Combined with the server-end defect classification model, the preliminary screening of industrial defects and the judgment of complex samples is achieved through local image block matching and fine-grained feature matching.

Benefits of technology

Meet the real-time and high-precision needs of industrial defect detection, avoid resource waste, and reduce system hardware costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374628A_ABST
    Figure CN120374628A_ABST
Patent Text Reader

Abstract

The invention provides an industrial defect classification detection method and device, a medium and electronic equipment, and the method comprises the steps: obtaining a reference image, and collecting an industrial production product image as a to-be-detected image; extracting the original features of the to-be-detected image and the reference image, carrying out local image block matching, calculating a first abnormal score, and classifying the to-be-detected image as a complex image when the first abnormal score is greater than a preset first classification threshold value; and extracting shallow layer features of the complex image and the reference image, performing fine-grained matching of the features, calculating a second abnormal score of the complex image, and classifying the complex image as an abnormal image when the second abnormal score is greater than a preset second classification threshold. According to the method, models with different performances are designed to complete preliminary screening of images and complex sample judgment respectively, dual requirements of industrial defect detection on real-time performance and high precision can be met by utilizing a cooperative work mode, and resource waste can be avoided. And the hardware cost of the system can be effectively reduced at the equipment arrangement level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial image defect detection, and in particular to an industrial defect classification detection method, device, medium and electronic device. Background Art

[0002] In recent years, with the continuous improvement of the intelligent level of the manufacturing industry, on highly automated production lines, the number of products has increased sharply, and it has been unable to meet the actual needs to rely on manual means to complete the detection of product quality. Therefore, industrial defect detection based on deep learning technology has received more and more attention. The research in this field aims to apply advanced artificial intelligence technology based on deep neural networks to improve the detection accuracy of product defects, so as to ensure production efficiency and product reliability.

[0003] However, while deep neural networks achieve high-precision defect detection, they also bring more complex network structures, which put higher requirements on hardware performance, making it difficult to deploy on edge devices with limited computing power and storage resources, and the inference speed cannot meet the real-time requirements of the industrial environment, restricting its popularization in the industrial field.

[0004] In actual industrial production, different from the defect detection indicators in academic research (such as Pixel-AP, AUPRO, Image-AUC, etc.), defect detection usually pays more attention to the binary classification problem of images, that is, whether the image belongs to a defective sample or a non-defective sample.

[0005] In addition, most industrial sites generally use expensive dedicated defect detection equipment, which integrates high-performance hardware and customized software, and the cost is relatively high. However, in modern industrial production, most products are normal or have obvious defects, and only a small part have unobvious defects. Therefore, most samples do not need to be processed complexly. For those simple samples, if high-performance detection equipment is used for detection, it is also easy to cause waste of computing resources.

[0006] Therefore, it is necessary to provide an industrial defect classification detection method that can not only meet the real-time and high-precision requirements of industrial defect detection but also avoid resource waste. Summary of the Invention

[0007] The purpose of the present invention is to provide an industrial defect classification detection method, device, medium and electronic device, so as to achieve avoiding resource waste while meeting the real-time and high-precision requirements of industrial defect detection.

[0008] In a first aspect, the industrial defect classification and detection method provided by the present invention includes: obtaining a reference image, collecting an industrial production product image as an image to be detected; extracting the original features of the image to be detected and the original features of the reference image for local image patch matching and calculating a first anomaly score. When the first anomaly score is greater than a preset first classification threshold, classifying the image to be detected as a complex image; extracting the shallow features of the complex image and the shallow features of the reference image and performing fine-grained feature matching, and calculating a second anomaly score of the complex image according to the fine-grained feature matching result of the complex image and the reference image. When the second anomaly score is greater than a preset second classification threshold, classifying the complex image as an abnormal image.

[0009] The beneficial effects of the industrial defect classification and detection method provided by the present invention are as follows: making full use of cascade architecture to design models with different performances to respectively complete the preliminary screening of images and the judgment of complex samples. Through the way of collaborative work, it can not only meet the dual requirements of real-time performance and high precision for industrial defect detection, but also avoid waste of resources. At the equipment layout level, it can also effectively reduce the hardware cost of the system.

[0010] In a possible embodiment, extracting the original features of the image to be detected and the original features of the reference image for local image patch matching and calculating a first anomaly score includes: extracting the original features of the image to be detected and the original features of the reference image; performing local feature patch matching on the original features of the image to be detected and the original features of the reference image to match the nearest neighbor reference image feature patch for the feature patch of the image to be detected in the reference image; calculating the pixel-level image anomaly score of the image to be detected according to the nearest neighbor reference image feature patch that matches the reference image and the image to be detected; calculating the first anomaly score of the image to be detected according to the pixel-level image anomaly score of the image to be detected.

[0011] In a possible embodiment, the shallow features of the complex image and the shallow features of the reference image are extracted and fine-grained matching of the features is performed, and the second anomaly score of the complex image is calculated according to the fine-grained matching result of the features of the complex image and the reference image, including: extracting the shallow features of the complex image and the shallow features of the reference image; combining the shallow features of the complex image with the original features of the complex image and extracting the fine-grained features of the complex image, and combining the shallow features of the reference image with the original features of the reference image and extracting the fine-grained features of the reference image; flattening the fine-grained features of the complex image and the fine-grained features of the reference image to obtain a set of feature blocks of the complex image and a set of feature blocks of the reference image; matching the set of feature blocks of the complex image with the set of feature blocks of the reference image to retrieve the local neighboring feature blocks of the complex image; calculating an initial anomaly score of the complex image according to the set of feature blocks of the complex image and the local neighboring feature blocks of the complex image; and calculating the second anomaly score of the complex image according to the initial anomaly score of the complex image and the foreground information of the complex image.

[0012] In another possible embodiment, an edge-side defect classification model is constructed to determine whether the image to be detected is a complex image. The edge-side defect classification model includes: a backbone network for extracting the original features of the image; an image matching module for performing local feature block matching on the original features of the image to be detected and the original features of the reference image; and a calculation module for calculating the first anomaly score of the image to be detected.

[0013] In other possible embodiments, the backbone network is a lightweight neural network, and a knowledge distillation technique is designed to transfer the learning ability of the dense convolutional network to the lightweight neural network.

[0014] A server-side defect classification model is constructed to determine whether the complex image is an abnormal image. The server-side defect classification model includes: a shallow feature encoder for extracting the shallow features of the image; a zero convolutional layer for introducing the shallow features into the local retrieval branch; a local retrieval branch for extracting the fine-grained features of the image; and a calculation unit for calculating the second anomaly score of the complex image.

[0015] In a second aspect, the present invention also provides an industrial defect classification and detection device, including: An image acquisition unit for acquiring a reference image and collecting an industrial production product image as a to-be-detected image; a first classification unit for extracting the original features of the to-be-detected image and the original features of the reference image for local image patch matching and calculating a first anomaly score, and classifying the to-be-detected image as a complex image when the first anomaly score is greater than a preset first classification threshold; a second classification unit for extracting the shallow features of the complex image and the shallow features of the reference image and performing fine-grained feature matching, calculating a second anomaly score of the complex image according to the fine-grained feature matching result of the complex image and the reference image, and classifying the complex image as an abnormal image when the second anomaly score is greater than a preset second classification threshold.

[0016] In a third aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above industrial defect classification and detection method is implemented.

[0017] In a fourth aspect, the present invention also provides an electronic device, including: a processor and a memory; the memory is used for storing a computer program; the processor is used for executing the computer program stored in the memory so that the electronic device executes the above industrial defect classification and detection method.

[0018] For the beneficial effects of the above second aspect to the fourth aspect, reference may be made to the description of the first aspect above. Description of the Drawings

[0019] Figure 1 It is a schematic flowchart of an industrial defect classification and detection method provided by an embodiment of the present invention; Figure 2 It is a schematic flowchart of a distillation process provided by an embodiment of the present invention; Figure 3 It is a schematic flowchart of a bottleneck injection layer generalization training process provided by an embodiment of the present invention; Figure 4 It is a schematic flowchart of a server-side defect classification model detection process provided by an embodiment of the present invention; Figure 5 It is a schematic diagram of a shallow feature encoder structure provided by an embodiment of the present invention; Figure 6 It is a schematic flowchart of a server-side defect classification model training process provided by an embodiment of the present invention; Figure 7 It is a schematic diagram of an industrial defect classification and detection device provided by an embodiment of the present invention; Figure 8 It is a schematic diagram of an electronic device structure provided by an embodiment of the present invention. Detailed Embodiments

[0020] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meaning as understood by those of ordinary skill in the art to which the present invention pertains. The words such as "including" used herein mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items.

[0021] In view of the problems existing in the prior art, embodiments of the present invention provide an industrial defect classification detection method, device, medium, and electronic device.

[0022] See the attached Figure 1 In this embodiment, an industrial defect classification detection method is provided. The method includes: S101: Obtain a reference image, and collect an industrial production product image as the image to be detected.

[0023] In a possible embodiment, the reference image is a normal image of an industrial production product. The collected industrial production product image is an image of the industrial production product to be detected collected during the industrial production process.

[0024] In a specific embodiment, an image acquisition camera is built on the industrial production site by selecting appropriate industrial cameras, lenses, light sources, etc., and the image acquisition camera is installed on a three-axis motion platform to obtain industrial product images at the industrial production site. The image acquisition camera is connected to an edge device, and the collected images can be directly transmitted to the edge device for image detection.

[0025] S102: Extract the original features of the image to be detected and the original features of the reference image for local image block matching and calculate the first anomaly score. When the first anomaly score is greater than a preset first classification threshold, classify the image to be detected as a complex image.

[0026] In a possible embodiment, extracting the original features of the image to be detected and the original features of the reference image for local image block matching and calculating the first anomaly score includes: extracting the original features of the image to be detected and the original features of the reference image; performing local feature block matching on the original features of the image to be detected and the original features of the reference image to match the nearest neighbor reference image feature block for the feature block of the image to be detected in the reference image; calculating the pixel-level image anomaly score of the image to be detected according to the nearest neighbor reference image feature block that matches the reference image and the image to be detected; and calculating the first anomaly score of the image to be detected according to the pixel-level image anomaly score of the image to be detected.

[0027] In a specific embodiment, the preliminary detection of whether the image to be detected belongs to a complex image is performed at the edge device side. Exemplarily, the detection process specifically includes: extracting the original features of the image to be detected and the original features of the reference image, and performing local feature block matching on the original features of the image to be detected and the original features of the reference image. Specifically, the local feature block matching of the original features of the image to be detected and the original features of the reference image can be implemented by local retrieval, that is, extracting the finer features contained in the original features of the detection image and the original features of the reference image, flattening the finer features of the image to be detected and the finer features of the reference image respectively, and in the flattened finer features, matching the nearest neighbor reference image feature block for each feature block of the image to be detected. Calculating the cosine similarity between the two feature blocks according to the feature block of the image to be detected and the nearest neighbor reference image feature block to be matched, and the calculation result is the pixel-level image anomaly score of the image to be detected. For an image to be detected, its first anomaly score is the sum of the top T pixel-level image anomaly scores on the image to be detected.

[0028] In a possible embodiment, the determination of the preset first classification threshold is: sorting the first anomaly scores of all images to be detected to obtain , calculating the average value of the first anomaly scores of every two images to be detected as the candidate classification threshold, that is . For each candidate classification threshold , dividing the positive and negative samples according to this threshold and calculating the corresponding recall rate . Finally, according to the candidate classification threshold and the recall rate threshold of the edge classifier on the validation set selecting an optimal classification threshold as the first classification threshold, and the first classification threshold is the maximum value among all classification thresholds that can meet the recall rate requirements of the validation set: . When the first anomaly score is less than or equal to the first classification threshold, the corresponding image to be detected is determined to be a normal image, and when the first anomaly score is greater than the first classification threshold, the corresponding image to be detected is determined to be a complex image, and the complex image needs to be transmitted to the cloud server for further detection.

[0029] In a specific embodiment, a set of images to be detected is input , and the first anomaly score of the i-th sample calculated by the edge-side defect classification model is the sum of the top T pixel-level image anomaly scores on the image. The specific calculation satisfies the following formula: . Among them, , H represents the height of the picture, W represents the width of the picture, represents sorting all pixel values in the pixel-level anomaly score, and T is a positive integer.

[0030] In a possible embodiment, an edge-side defect classification model is constructed to determine whether the image to be detected is a complex image. The edge-side defect classification model includes: a backbone network for extracting the original features of the image; an image matching module for performing local feature block matching on the original features of the image to be detected and the original features of the reference image; and a calculation module for calculating the first anomaly score of the image to be detected.

[0031] In a possible embodiment, the backbone network is a lightweight neural network, and a knowledge distillation technique is designed to transfer the learning ability of the dense convolutional network to the lightweight neural network.

[0032] In a possible embodiment, due to the limited performance of edge devices, directly using a dense convolutional network as the backbone network is not ideal. Therefore, a lightweight neural network is used as the backbone network. At the same time, in order to improve the performance of the edge-side defect classification model, a knowledge distillation technique is adopted to transfer the learning ability of the dense convolutional network to the lightweight neural network to improve the model performance while maintaining a certain inference speed. In addition, during the distillation process, due to the differences in the network structures of the dense convolutional network and the lightweight neural network, especially the inconsistency in the dimensions of the output features of the target intermediate layer, directly transferring the features of the dense convolutional network to the lightweight neural network is not feasible. For this reason, the present invention proposes a bottleneck injection method, which reduces the dimension of the output features of the dense convolutional network by introducing a small convolutional layer with a small convolution kernel size and a small stride, so that the target intermediate layer features of the dense convolutional network and the lightweight neural network are consistent in the channel dimension.

[0033] In a specific embodiment, the dense convolutional network is DenseNet201, and the lightweight neural network is MobileNetV2. In this embodiment, taking DenseNet201 and MobileNetV2 as examples, it is illustrated how to reduce the dimension of the output features of the dense convolutional network so that the target intermediate layer features of the dense convolutional network and the lightweight neural network are consistent in the channel dimension. The specific selection of the dense convolutional network and the lightweight neural network can also be other networks known to those skilled in the art.

[0034] Exemplarily, a 1×1 bottleneck injection layer is added at the output ends of two dense blocks (features.block1, features.block2) of DenseNet201 to compress the channel dimension to align with the feature layers of MobileNetV2 (features.2, features.3). The process satisfies the following formula: ; ; . Among them, , , represents the input image, represents DenseNet201, represents MobileNetV2, represents the intermediate layer features extracted from DenseNet201, represents the intermediate layer features extracted from MobileNetV2, and the bottleneck injection layer adjusts the number of feature channels of to and adjusts it to .

[0035] See the attached drawings of the specification Figure 2 . The output features of DenseNet201 are extracted from its two dense blocks (features.block1, features.block2), denoted as D1 and D2 respectively, while MobileNetV2 is (features.2, features.3), denoted as M1 and M2 respectively. By distilling the features of the above network layers respectively, the consistency of the features extracted by both in terms of width and height is ensured, and the difference in the channel dimension is eliminated by the bottleneck injection layer, ensuring that the information transfer and matching during the distillation process are more precise, enabling the effective transfer of knowledge from DenseNet201 to MobileNetV2. The distillation process satisfies the following formula: ; ; ; . Among them represents the output features of the teacher model, represents the output features of the student model, is the divergence of the features. In order to comprehensively consider the contribution of features at different levels to the distillation effect, the weighted KL divergence is used as the distillation loss function, represents the weight coefficient of the th level, reflecting the contribution degree of different feature layers to the distillation loss.

[0036] After distillation, MobileNetV2 is used as the backbone network to extract the features of the image, and the obtained extraction result is the original feature of the image. After the backbone network, a module consistent with the CPR local retrieval branch is connected to perform local image patch matching. The module consistent with the CPR local retrieval branch specifically performs the operation of local feature patch matching on the original features of the image to be detected and the original features of the reference image.

[0037] See the attached drawings of the specification Figure 3 In a possible embodiment, in order to prevent overfitting problems during the distillation process, the bottleneck injection layer is generalized trained. Specifically, the MVTec 3D-AD dataset is selected and the bottleneck injection layer is trained with the standard CPR training process to obtain its weights. This dataset contains 10 subsets, and the training samples and test samples of all subsets are combined respectively, which can effectively help the bottleneck injection layer learn general feature representations and improve its generalization ability under different datasets and tasks. When performing the generalization training, the local retrieval branch is trained based on metric learning and an improved contrast loss function. During training, a training sample of one in is randomly selected as the reference sample for metric learning. Among the and fine-grained features extracted, three pairs of corresponding local feature patches at different positions are selected to establish sample pairs for metric learning, and the losses of the three sample pairs are calculated based on the improved contrast loss function for backpropagation. Among them, represents the set of the top K global nearest neighbor reference images of the training sample, represents any reference image a in the set of the top K global nearest neighbor reference images, and

[0038] S103: Extract the shallow features of the complex image and the shallow features of the reference image and perform fine-grained matching of the features. Calculate the second anomaly score of the complex image according to the fine-grained matching result of the features of the complex image and the reference image. When the second anomaly score is greater than the preset second classification threshold, classify the complex image as an abnormal image.

[0039] In a possible embodiment, the shallow features of the complex image and the shallow features of the reference image are extracted and fine-grained feature matching is performed. According to the fine-grained feature matching result of the complex image and the reference image, the second anomaly score of the complex image is calculated, including: extracting the shallow features of the complex image and the shallow features of the reference image; combining the shallow features of the complex image with the original features of the complex image and extracting the fine-grained features of the complex image, and combining the shallow features of the reference image with the original features of the reference image and extracting the fine-grained features of the reference image; flattening the fine-grained features of the complex image and the fine-grained features of the reference image to obtain the feature block set of the complex image and the feature block set of the reference image; matching the feature block set of the complex image with the feature block set of the reference image to retrieve the local neighbor feature blocks of the complex image; calculating the initial anomaly score of the complex image according to the feature block set of the complex image and the local neighbor feature blocks of the complex image; calculating the second anomaly score of the complex image according to the initial anomaly score of the complex image and the foreground information of the complex image.

[0040] In a specific embodiment, the detection of whether the complex image is an abnormal image is performed by the server-side defect classification model on the server. Exemplarily, see the accompanying Figure 4 drawings. The specific detection process includes: extracting the shallow features of the complex image and the shallow features of the reference image. The shallow features refer to the features containing low-level texture and edge information in the image, and the shallow features can make up for the deficiency of the high-level semantic features of the backbone network in the detail representation ability. The shallow features are new information for the model. In order to avoid destroying the original ability of the local retrieval branch and prevent the model training from being unstable, therefore, the shallow features are introduced into the local retrieval branch through the zero convolution layer. The processing process of the zero convolution layer for the shallow features can be expressed as the following function: . The shallow features corresponding to the global neighbor reference image of the complex image pass through processing to obtain , and the shallow features of the complex image pass through processing to obtain . After combining and with the corresponding original features and , the local retrieval branch extracts the fine-grained features therein. The feature extraction process of can be expressed by the following formula: Flatten to obtain the feature vector set: , where , denote the height of the extracted feature, and the subscript L denotes the identifier of the corresponding network. Similarly, denote the height of the extracted feature, denote the width of the extracted feature. The set of feature vectors is obtained in a similar way: . . Each image feature block in is subjected to region-restricted matching with the set of feature blocks in to retrieve the local neighboring feature blocks of , where denotes the cosine similarity between two feature vectors, denotes the size of the local retrieval range. Subsequently, the initial anomaly score of the complex image is calculated as follows: . The foreground information of is calculated by the foreground estimation branch : . The second anomaly score of the complex image is obtained by element-wise multiplication of the initial anomaly score of the complex image and the foreground information: .

[0041] In a specific embodiment, the foreground estimation predicts the foreground and background of the image, and the obtained foreground information (background pixel value is 0, foreground pixel value is 1) is multiplied element-wise with the initial anomaly score family, so that the prediction region is concentrated in the foreground of the image and the incorrect predictions in the background are ignored. What the foreground estimation branch performs is the operation of obtaining the above foreground information.

[0042] In a possible embodiment, a shallow feature encoder is constructed based on the Inception module for extracting shallow features of the image. Exemplarily, referring to the accompanying drawings of the specification Figure 5 , the shallow feature encoder includes three consecutive Inception modules, and each Inception module has the same structure as the Inception module in the local retrieval branch of the CPR, but the stride of the convolution and pooling operations is 2. The feature extraction process of the shallow feature encoder can be represented by the following function: .

[0043] In a possible embodiment, the global retrieval branch splits each feature in Feature block clustering center set , is also decomposed into feature blocks. During the retrieval of the global neighbors of the reference image set and its original features , through separately calculating and the histogram vectors of the feature blocks, and then obtaining according to the spatial distance between and the histogram vectors of.

[0044] In a possible embodiment, a server-side defect classification model is constructed to determine whether a complex image is an abnormal image. The server-side defect classification model includes: a shallow feature encoder for extracting shallow features of the image; a zero convolution layer for introducing the shallow features into a local retrieval branch; a local retrieval branch for extracting fine-grained features of the image; and a calculation unit for calculating a second abnormal score of the complex image.

[0045] In a specific embodiment, referring to the accompanying drawings of the specification Figure 6 , in order to obtain a better shallow feature encoder, zero convolution layer and local retrieval branch, the parameter learning process is optimized based on metric learning and contrast loss function. Specifically, the training samples of the server-side classifier are a subset of the training samples used by the edge-side classifier. Given a training image , in its global neighbor reference image set a is randomly selected as the reference image for metric learning. The original features of and are extracted by and . The shallow features processed by are and . After summing the shallow features and the corresponding original features of both, more fine-grained features are extracted through the local retrieval branch and , and further flattened to construct the corresponding feature vector sets and .

[0046] Exemplarily, for the training of metric learning, this study sets three different sample pairs, namely normal sample pairs, abnormal sample pairs and far sample pairs. The selection methods of the three sample pairs are specifically as follows: Normal sample pair: The normal pixel points of pixel points abnormal sample pair abnormal pixel points and pixel points at the same position pixel points distant sample pair normal pixel points of any pixel points of, and at the same time, the coordinate distance between the pixel points is greater than the set local retrieval range . Among them, the training set contains normal samples and abnormal samples, and the abnormal samples are marked with the abnormal areas in the images, so that it is possible to clearly know which pixel points are abnormal or normal.

[0047] The contrast loss function adopted by the server classifier is: , Training is completed when the loss function is minimized. Among them, The inner product of represents the cosine similarity of the sample pair, and are the boundary thresholds of the positive sample and the negative sample respectively, is the weight constant for balancing the positive and negative samples, is the weight of the distant sample pair, which is obtained by calculating the distance between and , represents the label of the sample, the positive sample is 1, the negative sample is 0, and the abnormal sample pair and the distant sample pair are both negative samples. represents the number of sample pairs.

[0048] The industrial defect classification and detection method provided by the present invention faces the challenges of industrial defect detection in binary image classification, cost control and real-time performance. Based on the cascade classifier and edge computing, a cascade classification scheme is designed. Through the collaborative work of edge devices and servers, it not only meets the dual requirements of industrial defect detection for real-time performance and high precision, but also effectively reduces the hardware cost of the system. In this scheme, the application of edge devices can replace a large number of traditional expensive defect detection machines. An edge-side defect classification model is designed on the edge device to complete the preliminary screening of most images that do not require complex processing. A server-side defect classification model with better performance is deployed on the server side to process complex samples that cannot be accurately judged by the edge device side. Through the technical solution of the present invention, the application of edge devices can reduce many traditional expensive defect detection machines, significantly reducing the hardware investment and maintenance costs.

[0049] Applying the industrial defect classification and detection method of the present invention to detect industrial production product images can make full use of models with different performance designed by the cascade architecture. Through the collaborative work mode, it can not only meet the dual requirements of real-time performance and high precision for industrial defect detection, but also avoid resource waste. At the equipment layout level, it can also effectively reduce the hardware cost of the system.

[0050] See the attached Figure 7 to the specification. In this embodiment, an industrial defect classification and detection device is further provided. This device is used to implement the above method embodiment. The device includes: An image acquisition unit 201, configured to acquire a reference image and collect an industrial production product image as an image to be detected.

[0051] A first classification unit 202, configured to extract the original features of the image to be detected and the original features of the reference image for local image block matching and calculate a first anomaly score. When the first anomaly score is greater than a preset first classification threshold, classify the image to be detected as a complex image.

[0052] A second classification unit 203, configured to extract the shallow features of the complex image and the shallow features of the reference image and perform fine-grained matching of the features, calculate a second anomaly score of the complex image according to the fine-grained matching result of the features of the complex image and the reference image. When the second anomaly score is greater than a preset second classification threshold, classify the complex image as an abnormal image.

[0053] All relevant contents of each step involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be elaborated here.

[0054] In some other embodiments of the present application, embodiments of the present application disclose an electronic device, such as Figure 8 shown. The electronic device 300 may include: one or more processors 301; a memory 302; a display 303; one or more applications (not shown); and one or more computer programs 304. The above components can be connected through one or more communication buses 305. Wherein the one or more computer programs 304 are stored in the above memory and are configured to be executed by the one or more processors 301. The one or more computer programs 304 include instructions, and the above instructions can be used to execute each step in Figure 1 and the corresponding embodiments.

[0055] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0056] In each of the embodiments of this application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0057] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk, or optical disc.

[0058] The above is only the specific implementation manner of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of this application should be covered by the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be subject to the protection scope of the claims.

Claims

1. An industrial defect classification and detection method, characterized in that Including: Obtain a reference image, and collect an image of an industrial production product as the image to be detected; Extract the original features of the image to be detected and the original features of the reference image for local image patch matching and calculate a first anomaly score. When the first anomaly score is greater than a preset first classification threshold, classify the image to be detected as a complex image; Extract the shallow features of the complex image and the shallow features of the reference image and perform fine-grained feature matching. Calculate a second anomaly score for the complex image according to the fine-grained feature matching result of the complex image and the reference image. When the second anomaly score is greater than a preset second classification threshold, classify the complex image as an abnormal image.

2. The method according to claim 1, characterized in that Extract the original features of the image to be detected and the original features of the reference image for local image patch matching and calculate a first anomaly score, including: Extract the original features of the image to be detected and the original features of the reference image; Perform local feature patch matching on the original features of the image to be detected and the original features of the reference image to match the nearest neighbor reference image feature patch for the feature patch of the image to be detected in the reference image; Calculate the pixel-level image anomaly score of the image to be detected according to the nearest neighbor reference image feature patch that matches the reference image and the image to be detected; Calculate the first anomaly score of the image to be detected according to the pixel-level image anomaly score of the image to be detected.

3. The method according to claim 1, wherein Extract the shallow features of the complex image and the shallow features of the reference image and perform fine-grained feature matching. Calculate a second anomaly score for the complex image according to the fine-grained feature matching result of the complex image and the reference image, including: Extract the shallow features of the complex image and the shallow features of the reference image; Combine the shallow features of the complex image with the original features of the complex image and extract the fine-grained features of the complex image. Combine the shallow features of the reference image with the original features of the reference image and extract the fine-grained features of the reference image; Flatten the fine-grained features of the complex image and the fine-grained features of the reference image to obtain a feature patch set of the complex image and a feature patch set of the reference image; Match the feature patch set of the complex image with the feature patch set of the reference image to retrieve the local nearest neighbor feature patch of the complex image; Calculate the initial anomaly score of the complex image according to the feature patch set of the complex image and the local nearest neighbor feature patch of the complex image; Calculate the second anomaly score of the complex image according to the initial anomaly score of the complex image and the foreground information of the complex image.

4. The method according to claim 1, characterized in that, Construct an edge-side defect classification model for determining whether the image to be detected is a complex image. The edge-side defect classification model includes: A backbone network for extracting the original features of an image; An image matching module for performing local feature patch matching on the original features of the image to be detected and the original features of the reference image; A calculation module for calculating the first anomaly score of the image to be detected.

5. The method according to claim 4, wherein The backbone network is a lightweight neural network, and a knowledge distillation technique is designed to transfer the learning ability of the dense convolutional network to the lightweight neural network.

6. The method according to claim 1, characterized in that, A server-side defect classification model is constructed to determine whether the complex image is an abnormal image. The server-side defect classification model includes: A shallow feature encoder for extracting the shallow features of the image; A zero convolution layer for introducing the shallow features into the local retrieval branch; A local retrieval branch for extracting the fine-grained features of the image; A calculation unit for calculating the second abnormal score of the complex image.

7. An industrial defect classification and detection device, characterized in that The device includes: An image acquisition unit for acquiring a reference image and collecting an industrial production product image as a to-be-detected image; A first classification unit for extracting the original features of the to-be-detected image and the original features of the reference image for local image patch matching and calculating a first abnormal score. When the first abnormal score is greater than a preset first classification threshold, classifying the to-be-detected image as a complex image; A second classification unit for extracting the shallow features of the complex image and the shallow features of the reference image and performing fine-grained feature matching, calculating the second abnormal score of the complex image according to the fine-grained feature matching result of the complex image and the reference image, and when the second abnormal score is greater than a preset second classification threshold, classifying the complex image as an abnormal image.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the industrial defect classification and detection method according to any one of claims 1 to 6.

9. An electronic device, characterized in that, It includes: A processor and a memory; The memory is used for storing a computer program; The processor is used for executing the computer program stored in the memory, so that the electronic device executes the industrial defect classification and detection method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image defect detection method and device, electronic equipment and storage medium

    CN113240642A

  • Positive sample-based foreign matter detection method and storage medium

    CN115294323A

  • Anomaly detection method and system based on image block feature cascade retrieval model

    CN118115822A

  • Industrial defect detection model construction method, industrial defect detection method, industrial defect detection device, industrial defect detection equipment and storage medium

    CN119090810A

  • Method and device for detecting printed matter

    CN120088185A