Training method, defect detection method, device, electronic device, storage medium and computer program product for unsupervised defect detection model based on continuous learning

By introducing a continuous learning mechanism into the unsupervised defect detection model, using feature perturbation and knowledge distillation technology, the problem of poor generalization ability of existing models is solved, and effective defect detection of continuous updating products on industrial production lines is achieved.

CN119672491BActive Publication Date: 2025-05-27INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411755378.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-05-27
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

When facing the dynamic product data that is continuously updated on the industrial production line, the existing unsupervised defect detection model has poor generalization capabilities and cannot effectively conduct defect detection.

Method used

The training method of the unsupervised defect detection model based on continuous learning is adopted. By obtaining the product sample image features of the current task, adding feature perturbations, and using the learnable query embedding vector for knowledge distillation, adjusting the parameters of the defect detection model to improve the generalization ability of the model.

Benefits of technology

By introducing prompt signals from historical tasks, ensure that the query embedding vectors of new and old tasks are as close as possible, helping the defect detection model to maintain memory of the old tasks while training new tasks, thereby improving anti-forgetfulness and generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672491B_ABST
    Figure CN119672491B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a training method, a defect detection method, a device, an electronic device, a storage medium, and a computer program product for an unsupervised defect detection model based on continuous learning, including: obtaining product sample image features corresponding to a current task; adding feature perturbations to the product sample image features corresponding to the current task; inputting the product sample image features corresponding to the current task after the feature perturbations are added and learnable query embedding vectors corresponding to the current task into a defect detection model; calculating a loss based on reconstructed image features and product sample image features; adjusting parameters based on the loss. In this way, by introducing a hint signal of a historical task, the query embedding vectors of new and old tasks can be made as close as possible, which can help the defect detection model maintain the memory of the old task while training the new task, thereby improving the anti-forgetting ability of the defect detection model and the generalization ability on the new task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and more particularly, to a training method, a defect detection method, an apparatus, an electronic device, a storage medium, and a computer program product for an unsupervised defect detection model based on continual learning. Background Art

[0002] In modern manufacturing, defect detection is generally performed in the last link of production and manufacturing, and its main purpose is to try to identify defects in products to ensure product quality. In the past, defect detection in industrial production was mainly completed by experienced workers. However, this method will consume a high labor cost, and if the workers are in a relatively fatigued state, it will also lead to the occurrence of false positives, where "false positives" refers to samples that are actually abnormal but are identified as normal samples. To solve the problem of high labor cost brought by the above-mentioned manual defect detection method, in related technologies, an unsupervised method is mainly used for defect detection of products.

[0003] However, the current unsupervised defect detection methods are usually effective for existing static product data. When facing the continuously updated dynamic product data on the industrial production line, the defect detection ability of the model will be greatly reduced, that is, the generalization ability of the current defect detection model is poor, resulting in the inability to effectively detect the continuously updated products on the industrial production line. Summary of the Invention

[0004] The present disclosure provides a training method, a defect detection method, an apparatus, an electronic device, a storage medium, and a computer program product for an unsupervised defect detection model based on continual learning, so as to at least solve the problem in the above-mentioned related technologies that the generalization ability of the defect detection model is poor, resulting in the inability to effectively detect the continuously updated products on the industrial production line.

[0005] According to the first aspect of the embodiments of the present disclosure, there is provided a training method for an unsupervised defect detection model based on continuous learning, including: obtaining product sample image features corresponding to a current task, where the current task is one of a plurality of tasks executed in sequence, each task in the plurality of tasks corresponds to multiple types of product sample images, and the product sample image features are image features extracted from the multiple types of product sample images; adding feature perturbations to the product sample image features corresponding to the current task; inputting the product sample image features corresponding to the current task after the feature perturbations are added and the learnable query embedding vector corresponding to the current task into the defect detection model to obtain reconstructed image features, where the learnable query embedding vector corresponding to the current task is a query embedding vector obtained by knowledge distillation of the trained query embedding vector corresponding to the previous task before the current task and the initialized query embedding vector of the current task; calculating a loss based on the reconstructed image features and the product sample image features; and training the defect detection model by adjusting the parameters of the defect detection model based on the loss.

[0006] Optionally, the obtaining of the product sample image features corresponding to the current task includes: performing feature extraction on the multiple types of product sample images corresponding to the current task to obtain the current product sample image features corresponding to the current task; randomly selecting some product sample image features from the multiple product sample image features corresponding to each task in at least one task before the current task; and determining the current product sample image features corresponding to the current task as the product sample image features obtained by combining the current product sample image features and the randomly selected partial product sample image features.

[0007] Optionally, the defect detection model includes a feature reconstruction network, and the feature reconstruction network includes an encoder and a decoder; the inputting of the product sample image features corresponding to the current task after the feature perturbations are added and the learnable query embedding vector corresponding to the current task into the defect detection model to obtain reconstructed image features includes: inputting the product sample image features corresponding to the current task after the feature perturbations are added into the encoder to obtain an encoding result; and inputting the encoding result and the learnable query embedding vector corresponding to the current task into the decoder to obtain the reconstructed image features.

[0008] Optionally, the encoder includes a plurality of cascaded encoders, and the decoder includes a plurality of cascaded decoders.

[0009] According to a second aspect of the embodiments of the present disclosure, a defect detection method is provided, including: obtaining target image features of an image including a product to be detected; inputting the target image features and a query embedding vector into a defect detection model trained according to the training method of the present disclosure to obtain target reconstructed image features, where the query embedding vector is a query embedding vector obtained by updating an initialized query embedding vector for a preset task during the training process of the defect detection model; and determining whether there is a defect in the product to be detected based on the target reconstructed image features and the target image features.

[0010] Optionally, the determining whether there is a defect in the product to be detected based on the target reconstructed image features and the target image features includes: calculating a target mean squared loss error between the target reconstructed image features and the target image features; determining that the product to be detected is normal when the target mean squared loss error is less than or equal to a preset threshold; otherwise, determining that the product to be detected is abnormal.

[0011] According to a third aspect of the embodiments of the present disclosure, a training device for a defect detection model is provided, including: an image feature acquisition module configured to acquire product sample image features corresponding to a current task, where the current task is one of a plurality of tasks executed in sequence, each task in the plurality of tasks corresponds to multiple types of product sample images, and the product sample image features are image features extracted from the multiple types of product sample images; a feature perturbation addition module configured to add feature perturbations to the product sample image features corresponding to the current task; a reconstruction module configured to input the product sample image features corresponding to the current task after the feature perturbations are added and the learnable query embedding vector corresponding to the current task into an unsupervised defect detection model based on continual learning to obtain reconstructed image features, where the learnable query embedding vector corresponding to the current task is a query embedding vector obtained by knowledge distillation of the trained query embedding vector corresponding to the previous task before the current task and the initialized query embedding vector of the current task; a loss calculation module configured to calculate a loss based on the reconstructed image features and the product sample image features; and a parameter adjustment module configured to train the defect detection model by adjusting parameters of the defect detection model based on the loss.

[0012] Optionally, the image feature acquisition module is configured to: extract features from multiple types of product sample images corresponding to the current task to obtain the current product sample image features corresponding to the current task; randomly select some product sample image features from the multiple product sample image features corresponding to each task in at least one task before the current task; and determine the product sample image features corresponding to the current task as the current product sample image features and the randomly selected partial product sample image features.

[0013] Optionally, the defect detection model includes a feature reconstruction network, and the feature reconstruction network includes an encoder and a decoder; the reconstruction module is configured to: input the product sample image features corresponding to the current task after adding feature perturbations into the encoder to obtain an encoding result; and input the encoding result and the learnable query embedding vector corresponding to the current task into the decoder to obtain the reconstructed image features.

[0014] Optionally, the encoder includes multiple cascaded encoders, and the decoder includes multiple cascaded decoders.

[0015] According to a fourth aspect of the embodiments of the present disclosure, there is provided a defect detection device, including: a to-be-detected image feature acquisition module configured to acquire target image features of an image including a to-be-detected product; a reconstructed feature acquisition module configured to input the target image features and a query embedding vector into a defect detection model trained according to the training method of the present disclosure to obtain target reconstructed image features, where the query embedding vector is a query embedding vector obtained by updating an initialized query embedding vector for a preset task during the training process of the defect detection model; and a defect determination module configured to determine whether there is a defect in the to-be-detected product based on the target reconstructed image features and the target image features.

[0016] Optionally, the defect determination module is configured to: calculate a target mean square loss error between the target reconstructed image features and the target image features; and determine that the to-be-detected product is normal when the target mean square loss error is less than or equal to a preset threshold; otherwise, determine that the to-be-detected product is abnormal.

[0017] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; where the processor is configured to execute the instructions to implement the training method of the unsupervised defect detection model based on continuous learning according to the present disclosure, or to implement the defect detection method according to the present disclosure.

[0018] According to a sixth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the training method of an unsupervised defect detection model based on continual learning according to the present disclosure, or execute the defect detection method according to the present disclosure.

[0019] According to a seventh aspect of the embodiments of the present disclosure, there is provided a computer program product including a computer program. When the computer program is executed by a processor, it implements the training method of an unsupervised defect detection model based on continual learning according to the present disclosure, or implements the defect detection method according to the present disclosure.

[0020] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0021] In the present disclosure, for each of multiple tasks, a learnable query embedding vector can be set. When a new task comes, the knowledge of the previous task, that is, the task whose execution order is before the current task, can be transmitted to the current task by means of knowledge distillation. In this way, by introducing the hint signal of the historical task, the query embedding vectors of the new and old tasks can be made as close as possible, that is, it can be ensured that the query in the decoding process does not deviate from the old task, and thus it can help the defect detection model to maintain the memory of the old task while training the new task, thereby improving the anti-forgetting ability of the defect detection model and the generalization ability on the new task.

[0022] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.

[0024] Figure 1 is a flowchart showing a training method of an unsupervised defect detection model based on continual learning according to an exemplary embodiment of the present disclosure;

[0025] Figure 2 is a schematic flowchart showing obtaining training sample images and dividing continual learning tasks according to an exemplary embodiment of the present disclosure;

[0026] Figure 3 is a schematic diagram showing extracting product sample image features from product sample images according to an exemplary embodiment of the present disclosure;

[0027] Figure 4It is a schematic diagram showing the addition of feature perturbations to product sample image features according to an exemplary embodiment of the present disclosure;

[0028] Figure 5 It is a schematic diagram showing the working process of a continuous distillation prompt module according to an exemplary embodiment of the present disclosure;

[0029] Figure 6 It is a schematic diagram showing the structure of a feature reconstruction network according to an exemplary embodiment of the present disclosure;

[0030] Figure 7 It is a flowchart showing a defect detection method according to an exemplary embodiment of the present disclosure;

[0031] Figure 8 It is a block diagram showing a training device for a defect detection model according to an exemplary embodiment of the present disclosure;

[0032] Figure 9 It is a block diagram showing a defect detection device according to an exemplary embodiment of the present disclosure;

[0033] Figure 10 It is a block diagram showing an electronic device according to an exemplary embodiment of the present disclosure. Detailed implementation manners

[0034] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0036] It should be noted here that "at least one of several items" in the present disclosure all represents the inclusion of three parallel situations: "any one of the several items", "any combination of multiple items of the several items", and "the whole of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example, "performing at least one of step one and step two" means the following three parallel situations: (1) performing step one; (2) performing step two; (3) performing step one and step two.

[0037] In modern manufacturing, defect detection is generally performed in the last stage of production, mainly aiming to identify defects in products to ensure product quality. In the past, defect detection in industrial production was mainly completed by experienced workers. However, this method consumes high labor costs, and if workers are in a relatively fatigued state, it will also lead to the occurrence of false positives, where "false positives" refer to samples that are actually abnormal but are identified as normal samples. To solve the problem of high labor costs brought about by the above-mentioned method of manual defect detection, in related technologies, an unsupervised method is mainly used for defect detection of products. It should be noted that the object of industrial defect detection and localization in the field of computer vision is an image containing the product to be detected, that is, by identifying and locating abnormal images to determine whether there are defects in the products shown in the images.

[0038] However, the current unsupervised defect detection methods are usually effective for existing static product data. When faced with continuously updated dynamic product data on the industrial production line, the defect detection ability of the model will be greatly reduced, that is, the generalization ability of the current defect detection model is poor, resulting in the inability to effectively detect defects in the continuously updated products on the industrial production line.

[0039] To solve the above problems existing in related technologies, the training method, defect detection method, device, electronic device, storage medium, and computer program product of an unsupervised defect detection model based on continual learning provided by the present disclosure can set a learnable query embedding vector for each task among multiple tasks. When a new task comes, the knowledge of the previous task, that is, the task before the current task in the execution order, can be transmitted to the current task through knowledge distillation. In this way, by introducing the hint signal of the historical task, the query embedding vectors of the new and old tasks can be made as close as possible, that is, it can be ensured that the decoding process will not deviate from the query of the old task, and thus it can help the defect detection model maintain the memory of the old task while training the new task, thereby improving the anti-forgetting ability of the defect detection model and its generalization ability on new tasks.

[0040] Figure 1 It is a flowchart showing the training method of an unsupervised defect detection model based on continual learning according to an exemplary embodiment of the present disclosure.

[0041] Refer to Figure 1, in step 101, the product sample image features corresponding to the current task can be obtained. Here, the current task can be one of multiple tasks executed in sequence, and each task among the multiple tasks can correspond to multiple types of product sample images. The above product sample image features can be image features extracted from multiple types of product sample images. Specifically, normal images containing various types of industrial items can be obtained as training sample images, and these training sample images can be randomly shuffled. Then, according to the number of tasks, these training sample images can be respectively assigned to different continual learning tasks to meet the requirements of continual learning.

[0042] Figure 2 is a schematic flowchart showing the process of obtaining training sample images and dividing continual learning tasks according to an exemplary embodiment of the present disclosure. Refer to Figure 2 , first, normal sample images containing various products can be obtained through an industrial camera or other imaging devices. Here, the "normal sample images" refer to sample images containing normal products, that is, sample images without abnormal products. Then, a large number of obtained normal sample images can be scattered according to product types and assigned to different continual learning tasks for continual learning. Exemplarily, assuming there are a total of N types of product types and the number of continual learning tasks is T, then these N types of product sample images can be respectively scattered and assigned to different continual learning tasks among the T continual learning tasks for continual learning.

[0043] It should be noted that in the present disclosure, the publicly available industrial defect detection dataset MVTecAD can be used to train the defect detection model. The MVTecAD dataset can contain normal training sample images of 15 different types of industrial products, that is, training sample images without abnormalities. Exemplarily, the 15 different types of industrial products may include industrial products of 10 object types and industrial products of 5 texture types. The "industrial products of object types" can include but are not limited to: camera lenses, bottles, pills, metal nuts, screws, capsules, etc.; the "industrial products of texture types" can include but are not limited to: industrial products made of wood, industrial products such as blankets with patterned images.

[0044] In addition, the continuous learning tasks can be set according to actual needs. Then, the normal training sample images of the above 15 different types of industrial products can be randomly shuffled and divided into different continuous learning tasks for continuous learning. Exemplarily, one way of task setting can be: a total of 5 continuous learning tasks are set, and each continuous learning task can correspond to the normal training sample images of 3 types of industrial products, that is, each continuous learning task can only learn 3 industrial product categories. At the same time, the continuous learning can also be carried out in the way of class incremental learning; another way of task setting can be: a total of 6 continuous learning tasks are set, and the first continuous learning task can learn 10 industrial product categories, and each of the subsequent 5 continuous learning tasks can correspond to incrementally learning one industrial product category. At the same time, the continuous learning can also be carried out in the way of class incremental learning.

[0045] In this way, by obtaining the normal training sample images of multiple different types of industrial products, it can be ensured that a large number of image samples of different types of industrial products can be effectively applied to the training of the defect detection model, the adaptability of the defect detection model to diverse image data can be guaranteed, and thus the defect detection model can have good robustness.

[0046] It should be noted that after obtaining product sample images of multiple types, a pre-trained convolutional neural network can also be used to extract image features from these product sample images, and then product sample image features can be obtained, that is, a feature map after multi-scale feature fusion can be obtained. Figure 3 is a schematic diagram showing the extraction of product sample image features from a product sample image according to an exemplary embodiment of the present disclosure. Refer to Figure 3 , the EfficientNet_b4 model pre-trained on the large-scale public dataset ImageNet can be used to extract features from the product sample images, and then the final product sample image features can be obtained through multi-scale feature fusion processing.

[0047] In this way, in the present disclosure, since the EfficientNet model has a powerful feature extraction ability, by leveraging the multi-scale feature extraction function of the EfficientNet model, the detailed information and global information in the product sample images can be captured, and the richness and accuracy of feature expression can be improved through multi-scale feature fusion processing.

[0048] According to an exemplary embodiment of the present disclosure, feature extraction can be performed on multiple types of product sample images corresponding to the current task to obtain the current product sample image features corresponding to the current task. Then, some product sample image features can be randomly selected from the multiple product sample image features corresponding to each task in at least one task before the current task. Next, the current product sample image features and the randomly selected part of the product sample image features from the multiple product sample image features corresponding to the previous task can be determined as the product sample image features corresponding to the current task, that is, the product sample image features corresponding to the current task itself and the randomly selected part of the multiple product sample image features corresponding to the previous task itself can be used together to train for the current task.

[0049] Specifically, a representative feature subset can be selected from the product sample image feature pool corresponding to each previous task of the current task as the coreset. Exemplarily, a representative feature subset can be selected from the product sample image feature pool corresponding to each previous task of the current task in a greedy selection manner as the coreset. Then, during the training for the current task, the above coreset can be replayed to retain the previous tasks of the current task, that is, the important representative feature information in the old tasks, thereby improving the recognition and localization capabilities of the defect detection model. That is, through this feature replay method, the defect detection model can retain the memory of the old tasks when training new tasks, and thus improve its comprehensive performance.

[0050] Exemplarily, assume that the product sample image feature pool corresponding to the first continuous learning task contains a total of 100 product sample image features, and assume that the product sample image feature pool corresponding to the second continuous learning task contains a total of 200 product sample image features. Then, when training for the second continuous learning task, several product sample image features can be randomly selected from the 100 product sample image features contained in the product sample image feature pool corresponding to the first continuous learning task. Exemplarily, 5 product sample image features can be randomly selected. Then, the 200 product sample image features corresponding to the second continuous learning task itself and the randomly selected several product sample image features from the 100 product sample image features under the first continuous learning task can be used together for the training of the second continuous learning task.

[0051] Similarly, assume that there are a total of 150 product sample image features in the product sample image feature pool corresponding to the third continuous learning task. Before training for the third continuous learning task, several product sample image features can be randomly selected from the 200 product sample image features included in the product sample image feature pool corresponding to the second continuous learning task. Then, several product sample image features randomly selected from the 100 product sample image features included in the product sample image feature pool corresponding to the first continuous learning task before, several product sample image features randomly selected from the 200 product sample image features corresponding to the second continuous learning task itself this time, and the 150 product sample image features corresponding to the third continuous learning task itself can be used together for the training of the third continuous learning task.

[0052] In the present disclosure, a class-incremental learning approach can be adopted, and the memory ability of the defect detection model for old tasks can be improved through feature replay. That is, the coreset downsampling method can be used for feature selection, and thus a representative feature subset can be selected from the feature pool corresponding to the previous task of the current task. That is, a more representative feature map can be obtained from the feature distribution corresponding to the previous task of the current task. In this way, during the process of training a new task, the more representative feature map corresponding to the previous task obtained by the coreset downsampling method can be used for replay, and thus the important feature information with representativeness in the old tasks can be retained, thereby improving the anti-forgetting ability and recognition and localization ability of the defect detection model.

[0053] In step 102, feature perturbation can be added to the product sample image features corresponding to the current task. Figure 4 is a schematic diagram showing the addition of feature perturbation to product sample image features according to an exemplary embodiment of the present disclosure. Referring to Figure 4 , perturbation can be sampled from a Gaussian distribution, and then the perturbation sampled from the Gaussian distribution can be used to add feature perturbation to the product sample image features corresponding to the current task to guide the defect detection model to learn the knowledge in normal product sample images through denoising.

[0054] Specifically, noise perturbation can be sampled and the sampled noise perturbation can be added to the product sample image feature f, that is, the feature map f, with a fixed jitter probability p. For a product sample image feature f with a dimension of C, the sampled perturbation can be:

[0055]

[0056] where α is a feature perturbation coefficient, which can be used to control the jitter scale of the noise level, and the noise perturbation D follows a Gaussian distribution N, where μ is the mean and σ2 is the variance.

[0057] In this disclosure, by adding feature perturbations, i.e., noise perturbations, to the product sample image features, i.e., the feature maps, the model can better learn and adapt to the noise in the images. That is, the addition of feature perturbations can enhance the reconstruction ability of the defect detection model for positive sample images, i.e., it can enhance the ability of the defect detection model to reconstruct the product sample images without anomalies, and thus can enhance the robustness and generalization ability of the defect detection model, enabling it to maintain a good ability to reconstruct normal images, i.e., defect-free images, in the face of various image noises.

[0058] It should be noted that in this disclosure, adding noise perturbations to the product sample image features is equivalent to artificially adding defects to the product sample images, i.e., artificially creating deficiencies. Moreover, the purpose of the feature reconstruction network is to reconstruct the normal sample images, i.e., to reconstruct the product sample images without defects, without anomalies, and in an ideal state before adding noise perturbations. Thus, by calculating the loss between the reconstructed image features and the original product sample image features and updating the parameters of the defect detection model based on the calculated loss, i.e., by calculating the loss between the reconstructed image features in the ideal state without anomalies and the normal product sample image features before adding noise perturbations and updating the parameters of the defect detection model based on the calculated loss, the reconstruction ability of the defect detection model can be gradually improved, i.e., the restoration ability of the defect detection model can be gradually improved, which also means that it can be ensured that the defect detection model can accurately restore the original state, i.e., the true state, i.e., the anomaly-free state, of the product sample image before adding noise perturbations.

[0059] After the defect detection model is trained, the image to be detected can be directly input into the defect detection model, and the normal image that the image to be detected should be in the ideal situation without defects, i.e., without abnormal deficiencies, can be reconstructed. Then, the original image to be detected can be compared with the normal image reconstructed by the defect detection model in the ideal situation. When the comparison result indicates that the gap between the two is relatively small, it means that the image to be detected is relatively close to the normal image in the ideal situation. The reason for the closeness of the two is likely that the image to be detected itself is almost a normal image without anomalies. Therefore, at this time, it can be determined that the image to be detected is a normal image without anomalies, i.e., it can be determined that the product included in the image to be detected is a normal product without anomalies.

[0060] Alternatively, when the comparison result indicates a relatively large gap between the two, it means that the image to be detected is quite different from the normal image under ideal conditions. The reason for the large gap between the two is likely that the image to be detected itself is an abnormal image with many defects. At this time, it can be determined that the image to be detected is an abnormal image, that is, it can be determined that the product included in the image to be detected is an abnormal product with defects. In this way, an effective identification of whether industrial products are abnormal is achieved.

[0061] In step 103, the product sample image features corresponding to the current task after adding feature perturbations and the learnable query embedding vector corresponding to the current task can be input into the defect detection model to obtain reconstructed image features. That is, the feature map after adding feature perturbations can be input into the encoder-decoder reconstruction network based on Transformer to reconstruct the feature map.

[0062] It should be noted that the "learnable query embedding vector corresponding to the current task" can be a query embedding vector obtained by performing knowledge distillation on the trained query embedding vector corresponding to the previous task before the current task and the initial query embedding vector of the current task. That is, in the process of continual learning, each continual learning task can learn a separate prompt embedding as the query of the decoder. And a continuous distillation prompt module can be introduced into the decoder. In this way, when a new task comes, the knowledge of the previous task can be passed to the current task by means of weight distillation, that is, the query of the previous task can be distilled in the subsequent task.

[0063] Figure 5 FIG. is a schematic diagram showing the working process of the continuous distillation prompt module according to an exemplary embodiment of the present disclosure. Referring to Figure 5 , the learnable query embedding vector of the current task can be optimized to distill the query embedding vector of the old task to obtain the knowledge of the old task. Then, the query embedding vector obtained after distillation can be input into the decoder included in the defect detection model. In addition, the multi-head self-attention module in Transformer can be modified to a multi-head neighbor mask self-attention module, so as to avoid the feature reconstruction network falling into the shortcut of directly copying the original image through the multi-head neighbor mask self-attention module, and thus the ability of the defect detection model to capture and reconstruct image features can be greatly improved.

[0064] In this way, in the present disclosure, when a new task comes, knowledge distillation can be performed on the trained query embedding vector corresponding to the old task before the new task, which can ensure that the query embedding vector of the new task is as close as possible to the query embedding vector of the old task, that is, it can be ensured that the query in the decoding process does not deviate from the old task, thereby improving the anti-forgetting ability of the defect detection model.

[0065] According to an exemplary embodiment of the present disclosure, the above defect detection model may include a feature reconstruction network, and the feature reconstruction network may include an encoder and a decoder. The feature map of the product sample image corresponding to the current task after adding feature perturbation can be input into the encoder to obtain an encoding result. Then, the encoding result and the learnable query prompt embedding corresponding to the current task can be input into the decoder to obtain a reconstruction token.

[0066] Figure 6 FIG. is a schematic structural diagram showing a feature reconstruction network according to an exemplary embodiment of the present disclosure. Referring to Figure 6 , the encoder can encode the feature map with added feature perturbation into the latent space, and can modify the multi-head self-attention module in each layer of the encoder into a multi-head neighbor masked self-attention module. Specifically, the encoder may include a multi-head neighbor masked self-attention module and a feed forward net. After adding feature perturbation to the feature tokens of the product sample image corresponding to the current task, the feature tokens of the product sample image corresponding to the current task with added feature perturbation can be input into the multi-head neighbor masked self-attention module included in the encoder. At this time, any one of K, V, and Q input into the multi-head neighbor masked self-attention module can be the above-mentioned product sample image feature with added feature perturbation. Then, the multi-head neighbor masked self-attention module can process K, V, and Q through the following formula:

[0067]

[0068] where D(K) is the channel dimension of K.

[0069] Next, residual processing can be performed on the output result of the multi-head neighbor masked self-attention module and the aforementioned product sample image feature with added feature perturbation to obtain a residual processing result. Then, the residual processing result can be input into the feed forward net to obtain an output result of the feed forward net. Next, residual processing can be performed again on the output result of the feed forward net and the aforementioned residual processing result to obtain the final output result of the encoder, that is, the encoding result.

[0070] According to an exemplary embodiment of the present disclosure, referring to Figure 6, the decoder can modify the multi-head attention module into a multi-head neighbor masked attention module. Exemplarily, the decoder can include two multi-head neighbor masked self-attention modules and a feedforward network. In addition, the key (K) and value (V) input to the first multi-head neighbor masked self-attention module can come from the output of the encoder; the query of the first multi-head neighbor masked self-attention module can use a learnable prompt embedding vector (query prompt embedding), and the learnable prompt embedding vector can be a query embedding vector obtained by performing knowledge distillation on the trained query embedding vector corresponding to the previous task before the current task and the initialized query embedding vector of the current task.

[0071] According to an exemplary embodiment of the present disclosure, the key (K) and value (V) of the second multi-headed neighbor mask self-attention module in the input decoder can come from the output of the previous decoder (output of lastlayer) cascaded with the current decoder, and the query of the second multi-headed neighbor mask attention module of the input decoder can come from the output of the first multi-headed neighbor mask attention module.

[0072] Specifically, the first multi-head neighbor masked attention module can process K, V, and Q to obtain the output result of the first multi-head neighbor masked attention module. Then, residual processing can be performed on the output result of the first multi-head neighbor masked attention module and the learnable query embedding vector input into the first multi-head neighbor masked attention module to obtain a first residual processing result.

[0073] Next, K, V contained in the output result of the previous decoder cascaded with the current decoder and the above-mentioned first residual processing result can also be used as the input of the second multi-head neighbor mask attention module, and then the output result of the second multi-head neighbor mask attention module can be obtained.

[0074] Then, the output result of the second multi-head neighbor mask attention module and the first residual processing result can be subjected to residual processing again to obtain a second residual processing result. Next, the second residual processing result can be input into the feedforward network to obtain the output result of the feedforward network. Then, the output result of the feedforward network and the second residual processing result can be subjected to residual processing again to obtain the final reconstructed image features.

[0075] According to an exemplary embodiment of the present disclosure, the encoder may include a plurality of cascaded encoders, and the decoder may include a plurality of cascaded decoders. Exemplarily, the number of cascades may be, but is not limited to, 4.

[0076] It should be noted that in the case of cascading multiple decoders, each decoder among the multiple decoders can include two multi-head neighbor mask self-attention modules and a feed-forward network. Moreover, the inputs of the first multi-head neighbor mask self-attention module included in each decoder, namely K and V, can both come from the encoding results of the last encoder among the cascaded multiple encoders; in addition, the inputs of the second multi-head neighbor mask self-attention module included in the first decoder, namely K and V, can also come from the encoding results of the last encoder among the cascaded multiple encoders; for the second multi-head neighbor mask self-attention module included in each decoder from the second decoder to the Nth decoder among the cascaded multiple decoders, the inputs of this second multi-head neighbor mask self-attention module, namely K and V, can come from the output results of the previous decoder cascaded with the current decoder.

[0077] In step 104, the loss can be calculated based on the reconstructed image features and the product sample image features. Exemplarily, the mean square loss error between the input feature map and the reconstructed feature map can be calculated.

[0078] In step 105, the defect detection model can be trained by adjusting the parameters of the defect detection model based on the loss. Specifically, the backpropagation algorithm can be used to update the parameters of the defect detection model to minimize the mean square loss error.

[0079] In this way, by continuously adjusting the parameters of the defect detection model based on the loss, it can help the defect detection model better reconstruct the image features, that is, it can improve the feature reconstruction ability and overall performance of the defect detection model.

[0080] It should be noted that after training the defect detection model, the performance of the defect detection model can also be tested. Specifically, the test sample image to be tested can be input into the already trained defect detection model, and then it can be predicted whether there are defects in the test sample image to be tested and the specific positions of the defects. Then, the test results of the defect detection model can be compared with the ground truth label corresponding to the test sample to be tested. Moreover, the more similar the test results are to the ground truth label, the better the performance of the defect detection model. In addition, the defect detection model can be further optimized based on the test results.

[0081] In the present disclosure, an efficient unsupervised industrial vision defect detection method based on continual learning is provided. By adding continuous prompt distillation in the feature reconstruction network, the performance and anti-forgetting ability of the defect detection model can be greatly improved; by coreset sampling for feature replay, the recognition and localization ability of the defect detection model can be further improved; the addition of feature perturbation enhances the reconstruction ability of the defect detection model for positive sample images; by using the neighbor mask attention mechanism in the feature reconstruction network of the Transformer-based encoder-decoder, the feature reconstruction network is effectively prevented from falling into the "copy shortcut", and the ability of the defect detection model to capture and reconstruct features is greatly improved. The defect detection method provided by the present disclosure can effectively improve the accuracy and efficiency of industrial vision defect detection, reduce the errors and costs of manual defect detection, and can improve the automation and intelligence levels of industrial product production lines.

[0082] Figure 7 is a flowchart showing a defect detection method according to an exemplary embodiment of the present disclosure.

[0083] Referring to Figure 7 , in step 701, the target image features of the image including the product to be detected can be obtained. Specifically, a pre-trained convolutional neural network can be used to perform image feature extraction on the image of the product to be detected, and then the target image features can be obtained, that is, the feature map after multi-scale feature fusion can be obtained. Exemplarily, the EfficientNet_b4 model pre-trained on the large-scale public dataset ImageNet can be used to perform feature extraction on the image of the product to be detected, and then the final target image features can be obtained through multi-scale feature fusion processing.

[0084] Thus, in the present disclosure, since the EfficientNet model has a strong feature extraction ability, by leveraging the multi-scale feature extraction function of the EfficientNet model, the detailed information and global information in the image of the product to be detected can be captured, and the richness and accuracy of feature expression can be improved through multi-scale feature fusion processing.

[0085] In step 702, the target image features and the query embedding vector can be input into the defect detection model trained according to the training method of the present disclosure to obtain the target reconstructed image features, where the query embedding vector can be the query embedding vector obtained by updating the initialized query embedding vector for a preset task during the training process of the defect detection model.

[0086] It should be noted that, as described in the previous embodiment, multiple continual learning tasks can be preset, and each continual learning task can correspond to a learnable query embedding vector. Additionally, when a new task comes, the knowledge of the previous task can be transferred to the current task through weight distillation, that is, the queries of the previous task can be distilled in subsequent tasks. Exemplarily, assuming a total of 5 continual learning tasks are set, for the 1st continual learning task, it can correspond to an initial query embedding vector, and this initial query embedding vector can contain multiple randomly generated numerical elements. During the training process of the 1st continual learning task, the initial query embedding vector corresponding to the 1st continual learning task can be continuously updated, and finally, the trained query embedding vector corresponding to the 1st continual learning task can be obtained.

[0087] Next, for the 2nd continual learning task, it can also correspond to an initial query embedding vector, and the trained query embedding vector corresponding to the 1st continual learning task before the 2nd continual learning task can be subjected to knowledge distillation with the initial query embedding vector of the 2nd continual learning task. Then, the query embedding vector obtained through knowledge distillation can be used as the input to the 1st multi-head neighbor masked self-attention module of the decoder included in the feature reconstruction network, so as to train the 2nd continual learning task. And during the training process of the 2nd continual learning task, the learnable query embedding vector corresponding to the 2nd continual learning task can be continuously updated, and finally, the trained query embedding vector corresponding to the 2nd continual learning task can be obtained.

[0088] Similarly, for the 3rd continual learning task, it can also correspond to an initial query embedding vector, and the trained query embedding vector corresponding to the 2nd continual learning task before the 3rd continual learning task can be subjected to knowledge distillation with the initial query embedding vector of the 3rd continual learning task. Then, the query embedding vector obtained through knowledge distillation can be used as the input to the 1st multi-head neighbor masked self-attention module in the decoder included in the feature reconstruction network, so as to train the 3rd continual learning task. And during the training process of the 3rd continual learning task, the learnable query embedding vector corresponding to the 3rd continual learning task can be continuously updated, and finally, the trained query embedding vector corresponding to the 3rd continual learning task can be obtained.

[0089] And so on. After completing the training for the above 5 continual learning tasks, the finally fixed query embedding vector can be obtained. This finally fixed query embedding vector can be used as the query input of the first multi-head neighbor mask self-attention module in the decoder included in the feature reconstruction network when reconstructing the target image features of the image of the product to be detected by the trained defect detection model.

[0090] In this way, during the training process of the defect detection model, when a new task arrives, knowledge distillation can be performed on the trained query embedding vectors corresponding to the old tasks before the new task, which can ensure that the query embedding vectors of the new task are as close as possible to those of the old tasks. That is, it can be ensured that the decoding process does not deviate from the queries of the old tasks, thereby improving the anti-forgetting ability of the defect detection model.

[0091] In step 703, based on the target reconstructed image features and the target image features, it can be determined whether the product to be detected has defects. That is, based on the normal image features of the image of the product to be detected reconstructed by the trained feature reconstruction network in the ideal state and the original image features of the image of the product to be detected, it can be determined whether the product to be detected has defects.

[0092] According to an exemplary embodiment of the present disclosure, the target mean square loss error between the target reconstructed image features and the target image features can be calculated.

[0093] In the case where the target mean square loss error is less than or equal to a preset threshold, it indicates that the image of the product to be detected is relatively close to the normal image in the ideal situation. The reason for the closeness of the two is likely that the image of the product to be detected is itself a normal image with almost no abnormalities. Therefore, at this time, it can be determined that the image of the product to be detected is a non-abnormal image, that is, it can be determined that the product to be detected included in the image of the product to be detected has no abnormalities, and that is, it can be determined that the product to be detected is a normal product.

[0094] Otherwise, that is, in the case where the target mean square loss error is greater than the preset threshold, it indicates that the image of the product to be detected is quite different from the normal image in the ideal situation. The reason for the large difference between the two is likely that the image of the product to be detected is itself an abnormal image with many defects. At this time, it can be determined that the image of the product to be detected is an abnormal image, that is, it can be determined that the product to be detected included in the image of the product to be detected has abnormalities, and that is, it can be determined that the product to be detected is a substandard and inferior product. In this way, the effective identification of whether industrial products have abnormalities is achieved.

[0095] In the present disclosure, an efficient unsupervised industrial vision defect detection method based on continual learning is provided. By adding continuous prompt distillation to the feature reconstruction network, the performance and anti-forgetting ability of the defect detection model can be greatly improved; by coreset sampling for feature replay, the recognition and localization ability of the defect detection model can be further improved; the addition of feature perturbation enhances the reconstruction ability of the defect detection model for positive sample images; by using the neighbor mask attention mechanism in the feature reconstruction network based on Transformer encoder-decoder, the feature reconstruction network is effectively prevented from falling into the "copy shortcut", and the ability of the defect detection model to capture and reconstruct features is greatly improved. The defect detection method provided by the present disclosure can effectively improve the accuracy and efficiency of industrial vision defect detection, reduce the errors and costs of manual defect detection, and can improve the automation level and intelligent level of the industrial product production line.

[0096] Figure 8 is a block diagram showing a training apparatus of a defect detection model according to an exemplary embodiment of the present disclosure.

[0097] Referring to Figure 8 , the training apparatus 800 of the defect detection model may include an image feature acquisition module 801, a feature perturbation addition module 802, a reconstruction module 803, a loss calculation module 804, and a parameter adjustment module 805.

[0098] The image feature acquisition module 801 can acquire product sample image features corresponding to the current task, where the current task can be one of a plurality of tasks executed in sequence, and each task in the plurality of tasks can correspond to multiple types of product sample images, and the above product sample image features can be image features extracted from multiple types of product sample images. Specifically, normal images containing various types of industrial articles can be acquired as training sample images, and these training sample images can be randomly shuffled. Then, according to the number of tasks, these training sample images can be respectively divided into different continual learning tasks to meet the requirements of continual learning.

[0099] According to an exemplary embodiment of the present disclosure, the image feature acquisition module 801 may extract features from multiple types of product sample images corresponding to the current task to obtain the current product sample image features corresponding to the current task. Then, it may randomly select some product sample image features from among the multiple product sample image features corresponding to each task in at least one task prior to the current task. Next, the current product sample image features and the randomly selected part of the product sample image features from the multiple product sample image features corresponding to the previous task may be determined as the product sample image features corresponding to the current task, that is, the product sample image features corresponding to the current task itself and the randomly selected part of the multiple product sample image features corresponding to the previous task itself may be used together to train for the current task.

[0100] Specifically, a representative feature subset may be selected from the product sample image feature pool corresponding to each previous task of the current task as the coreset. Exemplarily, a greedy selection method may be used to select a representative feature subset from the product sample image feature pool corresponding to each previous task of the current task as the coreset. Then, during the process of training for the current task, the above coreset may be replayed to retain the previous tasks of the current task, that is, the important representative feature information in the old tasks, thereby improving the recognition and localization capabilities of the defect detection model. That is, through this feature replay method, the defect detection model can retain the memory of the old tasks when training new tasks, and thus improve its comprehensive performance.

[0101] In the present disclosure, a class incremental learning method may be adopted, and feature replay may be used to improve the memory ability of the defect detection model for old tasks. That is, a coreset downsampling method may be used for feature selection, and thus a representative feature subset may be selected from the feature pool corresponding to the previous task of the current task, that is, a more representative feature map may be obtained from the feature distribution corresponding to the previous task of the current task. In this way, during the process of training a new task, the more representative feature map corresponding to the previous task obtained by coreset downsampling may be used for replay, and thus the important representative feature information in the old tasks can be retained, thereby improving the anti-forgetting ability and recognition and localization capabilities of the defect detection model.

[0102] The feature perturbation addition module 802 may add feature perturbations to the product sample image features corresponding to the current task.

[0103] Thus, in the present disclosure, by adding feature perturbations, i.e., noise perturbations, to the product sample image features, i.e., the feature maps, the model can better learn and adapt to the noise in the images. That is, the addition of feature perturbations can enhance the reconstruction ability of the defect detection model for positive sample images, i.e., it can enhance the ability of the defect detection model to reconstruct product sample images without anomalies, and thus can enhance the robustness and generalization ability of the defect detection model, enabling it to still maintain a good ability to reconstruct normal images, i.e., defect-free images, in the face of various image noises.

[0104] The reconstruction module 803 can input the product sample image features corresponding to the current task after the addition of feature perturbations and the learnable query embedding vector corresponding to the current task into the defect detection model to obtain reconstructed image features. That is, the feature map after the addition of feature perturbations can be input into the Transformer-based encoder-decoder reconstruction network to reconstruct the feature map.

[0105] It should be noted that the "learnable query embedding vector corresponding to the current task" can be a query embedding vector obtained by knowledge distillation of the trained query embedding vector corresponding to the previous task before the current task and the initialization query embedding vector of the current task. That is, in the continuous learning process, each continuous learning task can learn a separate prompt embedding as the query of the decoder. And a continuous distillation prompt module can be introduced into the decoder. In this way, when a new task comes, the knowledge of the previous task can be passed to the current task by means of weight distillation, that is, the query of the previous task can be distilled in the subsequent tasks.

[0106] Thus, in the present disclosure, when a new task comes, knowledge distillation can be performed on the trained query embedding vector corresponding to the old task before the new task, which can ensure that the query embedding vector of the new task is as close as possible to the query embedding vector of the old task. That is, it can be ensured that the query in the decoding process does not deviate from the query of the old task, thereby improving the anti-forgetting ability of the defect detection model.

[0107] According to an exemplary embodiment of the present disclosure, the above-mentioned defect detection model may include a feature reconstruction network, and the feature reconstruction network may include an encoder and a decoder. The reconstruction module 803 can input the product sample image features corresponding to the current task after the addition of feature perturbations into the encoder to obtain an encoding result. Then, the reconstruction module 803 can input the encoding result and the learnable query embedding vector corresponding to the current task into the decoder to obtain reconstructed image features.

[0108] According to an exemplary embodiment of the present disclosure, the above-mentioned encoder may include a plurality of cascaded encoders, and the above-mentioned decoder may include a plurality of cascaded decoders. Exemplarily, the number of cascades may be, but is not limited to: 4.

[0109] The loss calculation module 804 can calculate the loss based on the reconstructed image features and the product sample image features. Exemplarily, the mean square loss error between the input feature map and the reconstructed feature map can be calculated.

[0110] The parameter adjustment module 805 can train the defect detection model by adjusting the parameters of the defect detection model based on the loss. Specifically, the parameters of the defect detection model can be updated using the backpropagation algorithm to minimize the mean square loss error.

[0111] In this way, by continuously adjusting the parameters of the defect detection model based on the loss, it can help the defect detection model better reconstruct the image features, that is, it can improve the feature reconstruction ability and overall performance of the defect detection model.

[0112] In the present disclosure, an efficient unsupervised industrial vision defect detection method based on continual learning is provided. By adding continuous prompt distillation in the feature reconstruction network, the performance and anti-forgetting ability of the defect detection model can be greatly improved; by coreset sampling for feature replay, the recognition and localization ability of the defect detection model can be further improved; the addition of feature perturbation enhances the reconstruction ability of the defect detection model for positive sample images; by using the neighbor mask attention mechanism in the feature reconstruction network of the Transformer-based encoder-decoder, it effectively prevents the feature reconstruction network from falling into the "copy shortcut", greatly improving the feature capture and reconstruction ability of the defect detection model. The defect detection method provided by the present disclosure can effectively improve the accuracy and efficiency of industrial vision defect detection, reduce the errors and costs of manual defect detection, and can improve the automation and intelligence levels of industrial product production lines.

[0113] Figure 9 is a block diagram showing a defect detection device according to an exemplary embodiment of the present disclosure.

[0114] Referring to Figure 9 , the defect detection device 900 may include a target image feature acquisition module 901 for the image containing the product to be detected, a reconstructed feature acquisition module 902, and a defect determination module 903.

[0115] The target image feature acquisition module 901 for the image containing the product to be detected can acquire the target image features of the image containing the product to be detected. Specifically, a pre-trained convolutional neural network can be used to perform image feature extraction on the image of the product to be detected, and then the target image features can be obtained, that is, the feature map after multi-scale feature fusion can be obtained. Exemplarily, the EfficientNet_b4 model pre-trained on the large-scale public dataset ImageNet can be used to perform feature extraction on the image of the product to be detected, and then the final target image features can be obtained through multi-scale feature fusion processing.

[0116] Thus, in the present disclosure, since the EfficientNet model has a powerful feature extraction ability, by leveraging the multi-scale feature extraction function of the EfficientNet model, it is possible to capture the detailed information and global information in the image of the product to be detected, and the richness and accuracy of feature representation can be improved through multi-scale feature fusion processing.

[0117] The reconstructed feature acquisition module 902 can input the target image feature and the query embedding vector into an unsupervised defect detection model based on continual learning trained according to the training method of the present disclosure to obtain the target reconstructed image feature, where the query embedding vector can be a query embedding vector obtained by updating the initialized query embedding vector for a preset task during the training process of the defect detection model.

[0118] Thus, during the training process of the defect detection model, when a new task arrives, knowledge distillation can be performed on the trained query embedding vector corresponding to the old task before the new task, which can ensure that the query embedding vector of the new task is as close as possible to the query embedding vector of the old task, that is, it can be ensured that the decoding process does not deviate from the query of the old task, thereby improving the anti-forgetting ability of the defect detection model.

[0119] The defect determination module 903 can determine whether the product to be detected has a defect based on the target reconstructed image feature and the target image feature. That is, it can be determined whether the product to be detected has a defect based on the normal image feature of the image of the product to be detected reconstructed by the trained feature reconstruction network in the ideal state and the original image feature of the image of the product to be detected.

[0120] According to an exemplary embodiment of the present disclosure, the defect determination module 903 can calculate the target mean squared loss error between the target reconstructed image feature and the target image feature.

[0121] In the case where the target mean squared loss error is less than or equal to a preset threshold, it indicates that the image of the product to be detected is relatively close to the normal image in the ideal situation, and the reason for the closeness of the two is likely that the image of the product to be detected itself is almost a normal image without anomalies. Therefore, at this time, it can be determined that the image of the product to be detected is an image without anomalies, that is, it can be determined that the product to be detected included in the image of the product to be detected has no anomalies, and thus it can be determined that the product to be detected is a normal product.

[0122] Otherwise, that is, when the target mean squared loss error is greater than the preset threshold, it indicates that the image of the product to be detected is quite different from the normal image under ideal conditions. The reason for the large difference between the two is likely that the image of the product to be detected itself is an abnormal image with many defects. At this time, it can be determined that the image of the product to be detected is an abnormal image, that is, it can be determined that there are abnormalities in the product to be detected included in the image of the product to be detected, and it can also be determined that the product to be detected is a substandard and inferior product. In this way, the effective identification of whether industrial products are abnormal is achieved.

[0123] In the present disclosure, an efficient unsupervised industrial vision defect detection method based on continual learning is provided. By adding continuous prompt distillation in the feature reconstruction network, the performance and anti-forgetting ability of the defect detection model can be greatly improved; by coreset sampling for feature replay, the recognition and localization ability of the defect detection model can be further improved; the addition of feature perturbation enhances the reconstruction ability of the defect detection model for positive sample images; by using the neighbor mask attention mechanism in the feature reconstruction network based on the Transformer encoder-decoder, the feature reconstruction network is effectively prevented from falling into the "copy shortcut", greatly improving the ability of the defect detection model to capture and reconstruct features. The defect detection method provided in the present disclosure can effectively improve the accuracy and efficiency of industrial vision defect detection, reduce the errors and costs of manual defect detection, and can improve the automation and intelligent levels of the industrial product production line.

[0124] Figure 10 It is a block diagram showing an electronic device according to an exemplary embodiment of the present disclosure.

[0125] Referring to Figure 10 , the electronic device 1000 includes at least one memory 1001 and at least one processor 1002. Instructions are stored in the at least one memory 1001. When the instructions are executed by the at least one processor 1002, a training method or a defect detection method of an unsupervised defect detection model based on continual learning according to an exemplary embodiment of the present disclosure is executed.

[0126] As an example, the electronic device 1000 can be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instructions. Here, the electronic device 1000 does not have to be a single electronic device, and can also be any assembly of devices or circuits that can execute the above instructions (or instruction sets) alone or jointly. The electronic device 1000 can also be a part of an integrated control system or a system manager, or can be configured as a portable electronic device that is interconnected with a local or remote (for example, via wireless transmission) interface.

[0127] In the electronic device 1000, the processor 1002 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, and the like.

[0128] The processor 1002 may execute instructions or code stored in the memory 1001, where the memory 1001 may also store data. The instructions and data may also be sent and received over a network via the network interface device, where the network interface device may employ any known transmission protocol.

[0129] The memory 1001 may be integrated with the processor 1002, for example, by arranging RAM or flash memory within an integrated circuit microprocessor or the like. Additionally, the memory 1001 may include a separate device, such as an external disk drive, a storage array, or other storage devices usable by any database system. The memory 1001 and the processor 1002 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, etc., such that the processor 1002 can read files stored in the memory.

[0130] In addition, the electronic device 1000 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 1000 may be connected to each other via a bus and / or a network.

[0131] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium may also be provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the above-mentioned training method or defect detection method of the unsupervised defect detection model based on continuous learning. Examples of the computer-readable storage medium here include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, where the any other device is configured to store a computer program and any associated data, data files and data structures in a non-transitory manner and provide the computer program and any associated data, data files and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system such that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.

[0132] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, including a computer program, which when executed by a processor, implements the training method or defect detection method of the unsupervised defect detection model based on continuous learning according to the present disclosure.

[0133] According to the training method, defect detection method, device, electronic device, storage medium and computer program product of the unsupervised defect detection model based on continuous learning of the present disclosure, for each task among multiple tasks, a learnable query embedding vector can be set. When a new task comes, the knowledge of the previous task, that is, the task before the current task in the execution order, can be transferred to the current task through knowledge distillation. In this way, by introducing the hint signal of the historical task, the query embedding vectors of the new and old tasks can be made as close as possible, that is, it can be ensured that the decoding process will not deviate from the query of the old task. Furthermore, it can help the defect detection model maintain the memory of the old task while training the new task, thereby improving the anti-forgetting ability and generalization ability of the defect detection model on the new task.

[0134] According to the exemplary embodiment of the present disclosure, the way of class incremental learning can be adopted, and the memory ability of the defect detection model for the old task can be improved through feature replay. That is, the coreset downsampling method can be used for feature selection, and then a representative feature subset can be selected from the feature pool corresponding to the previous task of the current task, that is, a more representative feature map can be obtained from the feature distribution corresponding to the previous task of the current task. In this way, during the process of training the new task, the more representative feature map corresponding to the previous task obtained by the coreset downsampling method can be used for replay, and then the important feature information with representativeness in the old task can be retained, thereby improving the anti-forgetting ability and recognition and localization ability of the defect detection model.

[0135] According to the exemplary embodiment of the present disclosure, by adding feature perturbations, that is, noise perturbations, to the product sample image features, that is, the feature map, the model can better learn and adapt to the noise in the image. That is, the addition of feature perturbations can enhance the reconstruction ability of the defect detection model for the positive sample image, that is, the ability of the defect detection model to reconstruct the product sample image without abnormalities, and further enhance the robustness and generalization ability of the defect detection model, so that it can still maintain a good ability to reconstruct normal images, that is, defect-free images, in the face of various image noises.

[0136] According to the exemplary embodiment of the present disclosure, when a new task comes, knowledge distillation can be performed on the trained query embedding vectors corresponding to the old tasks before the new task, which can ensure that the query embedding vectors of the new task are as close as possible to those of the old task, that is, it can be ensured that the decoding process will not deviate from the query of the old task, thereby improving the anti-forgetting ability of the defect detection model.

[0137] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known or customary technical means in the art not disclosed herein. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0138] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A training method for an unsupervised defect detection model based on continuous learning, characterized in that: include: Acquire a product sample image feature corresponding to a current task, wherein the current task is one of a plurality of tasks to be performed in sequence, each of the plurality of tasks corresponds to a plurality of types of product sample images, and the product sample image feature is an image feature extracted from the plurality of types of product sample images; Adding feature disturbance to the product sample image features corresponding to the current task; Inputting the product sample image features corresponding to the current task after adding feature perturbations and the learnable query embedding vector corresponding to the current task into the defect detection model to obtain reconstructed image features, wherein the learnable query embedding vector corresponding to the current task is a query embedding vector obtained by performing knowledge distillation on the trained query embedding vector corresponding to the previous task before the current task and the initialization query embedding vector of the current task; Calculating a loss based on the reconstructed image features and the product sample image features; The defect detection model is trained by adjusting parameters of the defect detection model based on the loss.

2. The training method according to claim 1, characterized in that: The obtaining of product sample image features corresponding to the current task includes: Extracting features of multiple types of product sample images corresponding to the current task to obtain features of the current product sample images corresponding to the current task; Randomly select part of the product sample image features from a plurality of product sample image features corresponding to each task in at least one task before the current task; The current product sample image features and the randomly selected portion of product sample image features are determined as product sample image features corresponding to the current task.

3. The training method according to claim 1, characterized in that: The defect detection model includes a feature reconstruction network, and the feature reconstruction network includes an encoder and a decoder; The step of inputting the product sample image features corresponding to the current task after adding feature disturbance and the learnable query embedding vector corresponding to the current task into the defect detection model to obtain the reconstructed image features includes: Inputting the product sample image features corresponding to the current task after adding feature disturbance into the encoder to obtain an encoding result; The encoding result and the learnable query embedding vector corresponding to the current task are input into the decoder to obtain the reconstructed image features.

4. The training method according to claim 3, characterized in that: The encoder includes a plurality of cascaded encoders, and the decoder includes a plurality of cascaded decoders.

5. A defect detection method, characterized in that: include: Acquire target image features of an image containing a product to be inspected; Inputting the target image feature and the query embedding vector into a defect detection model trained according to any one of the training methods of claims 1 to 4 to obtain a target reconstructed image feature, wherein the query embedding vector is a query embedding vector obtained by updating an initialization query embedding vector for a preset task during the training process of the defect detection model; Based on the target reconstructed image features and the target image features, it is determined whether the product to be inspected has defects.

6. The defect detection method according to claim 5, characterized in that: The determining whether the product to be inspected has a defect based on the target reconstructed image feature and the target image feature includes: Calculating a target mean square loss error between the target reconstructed image features and the target image features; When the target mean square loss error is less than or equal to a preset threshold, determining that there is no abnormality in the product to be detected; Otherwise, it is determined that the product to be detected has an abnormality.

7. A training device for a defect detection model, characterized in that: include: An image feature acquisition module is configured to acquire a product sample image feature corresponding to a current task, wherein the current task is one of a plurality of tasks to be performed in sequence, each of the plurality of tasks corresponds to a plurality of types of product sample images, and the product sample image feature is an image feature extracted from the plurality of types of product sample images; A feature disturbance adding module, configured to add feature disturbance to the product sample image features corresponding to the current task; A reconstruction module is configured to input the product sample image features corresponding to the current task after adding feature perturbations and the learnable query embedding vector corresponding to the current task into an unsupervised defect detection model based on continuous learning to obtain reconstructed image features, wherein the learnable query embedding vector corresponding to the current task is a query embedding vector obtained by performing knowledge distillation on a trained query embedding vector corresponding to a previous task before the current task and an initialized query embedding vector of the current task; a loss calculation module, configured to calculate the loss based on the reconstructed image features and the product sample image features; A parameter adjustment module is configured to train the defect detection model by adjusting the parameters of the defect detection model based on the loss.

8. A defect detection device, characterized in that: include: A module for acquiring features of an image to be detected, configured to acquire features of a target image of an image containing a product to be detected; A reconstruction feature acquisition module, configured to input the target image feature and the query embedding vector into a defect detection model trained according to any one of the training methods of claims 1 to 4, to obtain a target reconstructed image feature, wherein the query embedding vector is a query embedding vector obtained by updating an initialization query embedding vector for a preset task during the training process of the defect detection model; The defect determination module is configured to determine whether the product to be inspected has defects based on the target reconstructed image features and the target image features.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the training method of an unsupervised defect detection model based on continuous learning as described in any one of claims 1 to 4, or to implement the defect detection method as described in any one of claims 5 to 6.

10. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the training method of an unsupervised defect detection model based on continuous learning as described in any one of claims 1 to 4, or to execute the defect detection method as described in any one of claims 5 to 6.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the training method of an unsupervised defect detection model based on continuous learning as described in any one of claims 1 to 4 is implemented, or the defect detection method as described in any one of claims 5 to 6 is implemented.

Citation Information

Patent Citations

  • System and method of machine learning using embedding networks

    CA3096145A1

  • Insulator defect detection method, device, equipment, medium and product

    CN117314881A