Abnormal Image Generation Method, Training Sample Generation Method, and Defective Product Detection Method

By optimizing the embedding vector and generating abnormal images, the problem of difficulty in obtaining negative samples in industrial product anomaly detection is solved, and the detection performance and semantic consistency of samples are improved.

CN118015399BActive Publication Date: 2025-06-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410158313.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-01
Publication Date
2025-06-10
Estimated Expiration
2044-02-01

AI Technical Summary

Technical Problem

In the prior art, it is difficult for industrial products to effectively utilize negative samples, especially the semantic gap between the abnormal images and the abnormal images in the real world, which affects the detection performance.

Method used

By inputting the actual abnormal images into the pretrained image to generate the model, and optimizing the embedding vector while keeping the model parameters unchanged, an optimized embedding vector representing the feature distribution of defective images is generated. Then, the target image and optimized embedded features of the good product are input into the image generation model to generate the corresponding generated abnormal images.

Benefits of technology

The detection performance of the abnormality detection algorithm is improved, and by generating negative samples closer to the real abnormality characteristics, the semantic gap between the generated abnormal image and the real abnormal image is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118015399B_ABST
    Figure CN118015399B_ABST
Patent Text Reader

Abstract

The present application provides an abnormal image generation method for defective products, a training sample generation method for a defective product detection model, a defective product detection method, related devices, an electronic device, and a medium, which are applied to the field of computer vision technology. The abnormal image generation method includes: the optimized embedding vector is obtained by optimizing the initialized embedding vector while keeping the parameters of the image generation model unchanged. Since only the embedding vector with fewer parameters is optimized, it is convenient to meet the sample requirements in the abnormal image generation process without a large number of actual abnormal images. Input the target image of a good product and the optimized embedding features into the image generation model. Since the optimized embedding vector contains the image feature distribution of defective products learned from actual abnormal images, the generated abnormal image corresponding to the target image can be obtained. The generated abnormal image can be used as a negative sample for training the abnormal detection algorithm, which is beneficial to improving the detection performance of the abnormal detection algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer vision technology, and in particular, to a method for generating abnormal images of defective products, a method for generating training samples of a defective product detection model, a defective product detection method, related devices, electronic devices, and computer-readable storage media. Background Art

[0002] With the upgrade of equipment manufacturing automation, the yield rate of good products (normal products) on industrial production lines can reach a relatively high level, for example, it can reach more than 90%. Although the defective product (abnormal product) rate can reach a relatively low level, high-accuracy detection is still required to improve the overall quality of products. Artificial intelligence algorithm models can be used to predict whether industrial products are good products or defective products to improve detection efficiency. However, it is difficult to obtain a large number of negative samples of abnormal products, which makes intelligent abnormal detection algorithms a practical and challenging task in industrial quality inspection.

[0003] In related technologies, abnormal features of industrial products are synthesized with noise or external data to obtain abnormal images, which are used as negative samples to solve the above problems. However, there is a large semantic gap between the abnormal images obtained by the above synthesis and the abnormal images existing in the real world (actual abnormal images), which affects the detection performance of abnormal detection algorithms. Summary of the Invention

[0004] The present application provides a method for generating abnormal images of defective products, a method for generating training samples of a defective product detection model, a defective product detection method, related devices, electronic devices, and computer-readable storage media, which can generate abnormal images of defective products according to the image feature distribution of defective products, and is beneficial to improving the detection performance of abnormal detection algorithms.

[0005] In a first aspect, the present application provides a method for generating abnormal images of defective products, the method includes: determining an optimized embedding vector, where the optimized embedding vector is obtained by inputting an actual abnormal image into a pre-trained image generation model and optimizing an initialized embedding vector while keeping the parameters of the image generation model unchanged, and the optimized embedding vector represents the image feature distribution of the defective product; and inputting the target image of the good product and the optimized embedding feature into the image generation model to obtain a generated abnormal image corresponding to the target image.

[0006] In one implementation, based on the foregoing solution, before the determining the optimized embedding vector, the method further includes: inputting the initialized embedding vector and the actual abnormal image into a pre-trained image generation model, and optimizing the embedding vector while keeping the parameters of the image generation model unchanged.

[0007] In one implementation, based on the foregoing solution, the above-mentioned image generation model includes a pixel space network and a latent space network;

[0008] The above-mentioned process of inputting the initialized embedding vector and the actual abnormal image into the pre-trained image generation model includes: inputting the above-mentioned actual abnormal image into the encoder of the above-mentioned pixel space network to convert the above-mentioned actual abnormal image into actual abnormal features, where the above-mentioned actual abnormal features will be processed by the noise addition network of the above-mentioned latent space network to obtain corresponding noise-added features; the denoising network in the above-mentioned latent space network performs denoising processing on the above-mentioned noise-added features based on the initialized embedding vector to obtain the noise predicted by the denoising network, where the above-mentioned initialized embedding vector is mapped to the above-mentioned denoising network in a cross-attention manner; the method further includes: determining a first objective function according to the noise predicted by the above-mentioned denoising network; and determining an optimized embedding vector according to the above-mentioned first objective function.

[0009] In one implementation, based on the foregoing solution, the above-mentioned denoising network includes T U-nets corresponding to T times of denoising processing, where T is a positive integer; the method further includes: determining a j-th key vector according to the above-mentioned initialized embedding vector and the first parameter of the j-th U-net, where j is a positive integer less than or equal to T; determining a j-th value vector according to the above-mentioned initialized embedding vector and the second parameter of the above-mentioned j-th U-net; and mapping the above-mentioned initialized embedding vector to the above-mentioned j-th U-net based on the above-mentioned j-th key vector and the above-mentioned j-th value vector.

[0010] In one implementation, based on the foregoing solution, the above-mentioned process of determining a first objective function according to the noise predicted by the denoising network includes: determining a first objective function according to the noise predicted by the denoising network and the noise sampled from a Gaussian distribution; the above-mentioned process of determining an optimized embedding vector according to the above-mentioned first objective function includes: determining an optimized embedding vector by determining the minimum value of the above-mentioned first objective function.

[0011] In one implementation, based on the foregoing solution, the above-mentioned process of inputting the initialized embedding vector and the actual abnormal image into the pre-trained image generation model to optimize the embedding vector while keeping the parameters of the above-mentioned image generation model unchanged includes: inputting the actual abnormal image and its corresponding mask image into the pre-trained image generation model; determining a second objective function based on the features of the above-mentioned actual abnormal image that are inside the boundary of the above-mentioned mask while keeping the parameters of the above-mentioned image generation model unchanged; and determining an optimized embedding vector according to the above-mentioned second objective function; where the inside of the boundary of the above-mentioned mask corresponds to the abnormal area in the above-mentioned actual abnormal image.

[0012] In one implementation, based on the foregoing solution, while keeping the parameters of the above image generation model unchanged, determining a second objective function based on the features inside the boundary of the mask in the above actual abnormal image includes: determining the difference between the noise sampled from a Gaussian distribution and the noise predicted by the above denoising network; and determining the second objective function according to the above difference and the features inside the boundary of the mask in the above actual abnormal image; determining the optimized embedding vector according to the above second objective function includes: determining the optimized embedding vector by determining the minimum value of the above second objective function.

[0013] In one implementation, based on the foregoing solution, inputting the target image of a good product and the above optimized embedding features into the above image generation model to obtain the generated abnormal image corresponding to the above target image includes: inputting the target image of a good product, the bounding box corresponding to the above target image, and the above optimized embedding features into the above image generation model to obtain the generated abnormal image corresponding to the above target image; wherein, the area inside the above bounding box in the above generated abnormal image is the abnormal area.

[0014] In one implementation, based on the foregoing solution, inputting the target image of a good product and the above optimized embedding features into the above image generation model to obtain the generated abnormal image corresponding to the above target image includes: inputting the target image of a good product into the above image generation model, and performing noise addition processing on the above target image for T steps through the noise addition network of the above image generation model to obtain the noise-added features of the above target image, the noise-added features including the noise-added features at t steps, T being a positive integer, and t being a positive integer not greater than T; performing denoising processing for T steps on the noise-added features of the above target image through the denoising network mapped with the above optimized embedding features; wherein, the generated abnormal features obtained by the denoising processing at t steps include: according to the first part of the features outside the above bounding box in the loaded features at t steps, and the second part of the features inside the above bounding box in the denoised features corresponding to the loaded features at t + 1 steps; and decoding the generated abnormal features after the denoising processing for T steps through the decoder of the above pixel space network to obtain the generated abnormal image corresponding to the above target image.

[0015] In a second aspect, there is provided an abnormal image generation device for defective products, the device including: a first determination module and a first generation module; wherein, the above first determination module is used to determine an optimized embedding vector, where the above optimized embedding vector is obtained by inputting an actual abnormal image into a pre-trained image generation model and optimizing the initialized embedding vector while keeping the parameters of the above image generation model unchanged, and the above optimized embedding vector represents the image feature distribution of the above defective products; and the above first generation module is used to input the target image of a good product and the above optimized embedding features into the above image generation model to obtain the generated abnormal image corresponding to the above target image.

[0016] In one implementation, based on the foregoing solution, the abnormal image generation device for defective products further includes: an embedding vector acquisition module; wherein, before the first determination module determines the optimized embedding vector, the embedding vector acquisition module is configured to: input the initialized embedding vector and the actual abnormal image into a pre-trained image generation model, and optimize the embedding vector while keeping the parameters of the image generation model unchanged.

[0017] In one implementation, based on the foregoing solution, the image generation model includes a pixel space network and a latent space network;

[0018] The embedding vector acquisition module includes: a first input unit, a denoising unit, a first determination unit, and a second determination unit;

[0019] Wherein, the first input unit is configured to: input the actual abnormal image into the encoder of the pixel space network to convert the actual abnormal image into actual abnormal features, and the actual abnormal features will be processed by the noise addition network of the latent space network to obtain corresponding noise-added features; the denoising unit is configured to: the denoising network in the latent space network performs denoising processing on the noise-added features based on the initialized embedding vector to obtain the noise predicted by the denoising network, wherein the initialized embedding vector is mapped to the denoising network in a cross-attention manner; the first determination unit is configured to: determine a first objective function according to the noise predicted by the denoising network; the second determination unit is configured to: determine the optimized embedding vector according to the first objective function.

[0020] In one implementation, based on the foregoing solution, the denoising network includes T U-nets corresponding to T denoising processes, where T is a positive integer; the embedding vector acquisition module further includes a mapping unit; wherein, the mapping unit is configured to: determine a j-th key vector according to the initialized embedding vector and the first parameter of the j-th U-net, where j is a positive integer less than or equal to T; determine a j-th value vector according to the initialized embedding vector and the second parameter of the j-th U-net; and map the initialized embedding vector to the j-th U-net based on the j-th key vector and the j-th value vector.

[0021] In one implementation, based on the foregoing solution, the first determination unit is specifically configured to: determine a first objective function according to the noise predicted by the denoising network and the noise sampled from a Gaussian distribution; the second determination unit is specifically configured to: determine the optimized embedding vector by determining the minimum value of the first objective function.

[0022] In one implementation, based on the foregoing solution, the above-mentioned embedded vector acquisition module includes: a second input unit, a third determination unit, and a fourth determination unit; wherein the second input unit is configured to: input an actual abnormal image and its corresponding mask image into a pre-trained image generation model; the third determination unit is configured to: based on the features inside the boundary of the mask in the actual abnormal image while keeping the parameters of the image generation model unchanged, determine a second objective function; the fourth determination unit is configured to: determine an optimized embedded vector according to the second objective function; wherein, the inside of the boundary of the mask corresponds to the abnormal area in the actual abnormal image.

[0023] In one implementation, based on the foregoing solution, the third determination unit is configured to: determine the difference between the noise sampled from a Gaussian distribution and the noise predicted by the denoising network; and, based on the difference and the features inside the boundary of the mask in the actual abnormal image, determine a second objective function; the fourth determination unit is specifically configured to: determine an optimized embedded vector by determining the minimum value of the second objective function.

[0024] In one implementation, based on the foregoing solution, the first generation module is specifically configured to: input a target image of a good product, the bounding box corresponding to the target image, and the optimized embedded feature into the image generation model to obtain a generated abnormal image corresponding to the target image; wherein, the area inside the bounding box in the generated abnormal image is the abnormal area.

[0025] In one implementation, based on the foregoing solution, the first generation module includes: a first generation unit and a second generation unit; wherein, the first generation unit is configured to: input a target image of a good product into the image generation model, perform T-step noise addition processing on the target image through the noise addition network of the image generation model to obtain a noise-added feature of the target image, the noise-added feature includes t-step noise-added features, T is a positive integer, and t + 1 is a positive integer not greater than T; perform T-step denoising processing on the noise-added feature of the target image through the denoising network mapped with the optimized embedded feature; wherein, the generated abnormal feature obtained by the t-step denoising processing includes: according to the first part of the features outside the bounding box in the t-step loaded feature, and, the second part of the features inside the bounding box in the denoised feature corresponding to the (t + 1)-step loaded feature; and, the second generation unit is configured to: after decoding the generated abnormal feature after T-step denoising processing through the decoder of the pixel space network, obtain a generated abnormal image corresponding to the target image.

[0026] In a third aspect, a method for generating training samples of a defective product detection model is provided. The method includes: determining positive sample images; and generating generated abnormal images corresponding to the positive sample images through the method for generating training samples of the defective product detection model to obtain negative sample images.

[0027] In a fourth aspect, a device for generating training samples of a defective product detection model is provided. The device includes: a second determination module and a second generation module; wherein, the second determination module is configured to determine positive sample images; and the second generation module is configured to generate generated abnormal images corresponding to the positive sample images through the method for generating abnormal images of defective products provided in the first aspect or any one of its implementations to obtain negative sample images.

[0028] In a fifth aspect, a method for detecting defective products is provided. The method includes: detecting a to-be-detected object through a trained defective product detection model to obtain a detection result; wherein, the detection result includes that the type of the to-be-detected object is a non-defective product or a defective product, and when the detection result is that the to-be-detected object is a defective product, the detection result further includes a positioning result of an abnormal area, and the defective product detection model is trained according to the samples provided by the method in the third aspect.

[0029] In a sixth aspect, a device for detecting defective products is provided. The device includes: a detection module; wherein, the detection module is configured to detect a to-be-detected object through a trained defective product detection model to obtain a detection result; wherein, the detection result includes that the type of the to-be-detected object is a non-defective product or a defective product, and when the detection result is that the to-be-detected object is a defective product, the detection result further includes a positioning result of an abnormal area, and the defective product detection model is trained according to the samples provided by the method in the third aspect.

[0030] In a seventh aspect, an electronic device is provided, including a processor and a memory; the memory is used for storing a computer program, and the processor is used for calling and running the computer program stored in the memory to execute the method for generating abnormal images of defective products in the first aspect and its various implementations, or to execute the method for generating training samples of the defective product detection model provided in the third aspect, or to execute the method for detecting defective products provided in the fifth aspect.

[0031] In an eighth aspect, a chip is provided for implementing the method in any one of the first aspects or its various implementations. Specifically, the chip includes: a processor, configured to call and run a computer program from a memory, so that a device installed with the chip executes the method for generating abnormal images of defective products in the first aspect and its various implementations, or to execute the method for generating training samples of the defective product detection model provided in the third aspect, or to execute the method for detecting defective products provided in the fifth aspect.

[0032] In a ninth aspect, a computer-readable storage medium is provided for storing a computer program, which causes a computer to execute the abnormal image generation method for defective products in the first aspect and its various implementation manners described above, or to execute the training sample generation method for the defective product detection model provided in the third aspect, or to execute the defective product detection method provided in the fifth aspect.

[0033] In a tenth aspect, a computer program product is provided, including computer program instructions, which cause a computer to execute the abnormal image generation method for defective products in the first aspect and its various implementation manners described above, or to execute the training sample generation method for the defective product detection model provided in the third aspect, or to execute the defective product detection method provided in the fifth aspect.

[0034] In an eleventh aspect, a computer program is provided, which when running on a computer, causes the computer to execute the abnormal image generation method for defective products in the first aspect and its various implementation manners described above, or to execute the training sample generation method for the defective product detection model provided in the third aspect, or to execute the defective product detection method provided in the fifth aspect.

[0035] In summary, in the solution provided in the embodiments of the present application, an actual abnormal image is input into a pre-trained image generation model, and while keeping the parameters of the image generation model unchanged, the initialized embedding vector is optimized, so as to obtain an optimized embedding vector representing the true image feature distribution of the defective product. Since the embodiments of the present application do not optimize the image generation model with more parameters, but optimize the embedding vector with fewer parameters, it is convenient to meet the sample requirements in the abnormal image generation process without a large number of actual abnormal images and is easy to implement. On the other hand, the target image of the non-defective product and the above optimized embedding features are input into the image generation model. Since the optimized embedding vector contains the image feature distribution of the defective product learned from the actual abnormal image, the generated abnormal image corresponding to the target image can be obtained. The above generated abnormal image can be used as a negative sample for training the anomaly detection algorithm, which is beneficial to improving the detection performance of the anomaly detection algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0037] Figure 1Schematic diagram of a system framework applicable to the embodiments of the present application;

[0038] Figure 2 Flow chart of the method for generating abnormal images of defective products provided by the embodiments of the present application;

[0039] Figure 3 Flow chart of the process for determining the optimized embedding vector provided by the embodiments of the present application;

[0040] Figure 4 Flow chart of the method for determining the optimized embedding vector provided by the embodiments of the present application;

[0041] Figure 5 Flow chart of the process for determining the optimized embedding vector provided by an embodiment of the present application;

[0042] Figure 6 Flow chart of the process for determining the optimized embedding vector provided by another embodiment of the present application;

[0043] Figure 7 Flow chart of the process for determining the optimized embedding vector provided by still another embodiment of the present application;

[0044] Figure 8 Flow chart of the process for determining the generation of abnormal images based on the optimized embedding vector provided by the embodiments of the present application;

[0045] Figure 9 Flow chart of the process for determining the generation of abnormal images based on the optimized embedding vector provided by an embodiment of the present application;

[0046] Figure 10 Flow chart of the process for determining the generation of abnormal images based on the optimized embedding vector provided by another embodiment of the present application;

[0047] Figure 11 Schematic diagram of generating negative samples of the defective product detection model according to a small number of actual abnormal images provided by an embodiment of the present application;

[0048] Figure 12 Schematic diagram of the structure of the abnormal image generation device for defective products provided by the embodiments of the present application;

[0049] Figure 13 Schematic diagram of the structure of the training sample generation device of the defective product detection model provided by the embodiments of the present application;

[0050] Figure 14 Schematic diagram of the structure of the defective product detection device provided by the embodiments of the present application;

[0051] Figure 15 Schematic diagram of the structure of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0052] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0053] It should be noted that the terms "first", "second", etc. in the description and claims of the present application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In the embodiments of the present invention and the present application, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices. In the description of the present application, unless otherwise specified, "a plurality of" means two or more than two.

[0054] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0055] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0056] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, and mechatronics. Among them, the pre-trained model, also known as the large model or the basic model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0057] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes to identify and measure targets, etc., for machine vision, and further perform graphics processing to make the computer process images that are more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. The large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the visual field such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0058] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The pre-trained model is the latest development result of deep learning, which integrates the above technologies.

[0059] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence-generated content (AIGC), conversational interaction, smart healthcare, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0060] The solution provided by the embodiments of this application involves technologies such as computer vision technology and machine learning in artificial intelligence, and will be specifically described through the following embodiments.

[0061] Exemplarily, Figure 1 Figure 1 is a schematic diagram of a system framework applicable to the embodiments of this application. As Figure 1 shown, the embodiments of this application will be described by taking hazelnut images as an example. Specifically, Image 11 represents a normal product, also known as a good product, that is, a product without defects; Image 12 represents an abnormal product, also known as a defective product, that is, a product with defects. In the embodiments of this application, the image of the defective product is denoted as an "abnormal image".

[0062] As described above, since it is difficult to obtain a large number of negative samples of defective products, unsupervised anomaly detection algorithms are used in related technologies to intelligently detect industrial products. Specifically, unsupervised anomaly detection algorithms usually only use the features of normal images to learn the normal distribution or reconstruct normal features for anomaly detection. However, due to the lack of discriminative representation, the performance is sub-optimal, and it is impossible to locate the abnormal areas in defective products. For example, it is impossible to achieve anomaly localization at the pixel level. To solve the problem of unable to locate abnormal areas, in related technologies, a synthetic method is used to obtain abnormal images as negative samples for supervised intelligent detection algorithms. For example, in the paper "Draem: a discriminatively trained reconstruction embedding for surface anomaly detection", an external texture dataset is used to construct abnormal images; in the paper "Cutpaste: Self-supervised learning for anomaly detection and localization", image patches are randomly cropped and pasted to other parts to create abnormal images; in the paper "Simplenet: A simple network for image anomaly detection and localization", the feature map is directly perturbed to simulate the features of abnormal images. However, there is a significant distribution gap between the synthetic abnormal images or features and the abnormal images or features in the real world. Obviously, the detection performance of the model trained with synthetic abnormal images as negative samples needs to be improved.

[0063] In view of the above technical problems existing in the related technologies, in the embodiments of the present application, the defective product image features distribution in the actual abnormal images is learned through the embedding vector, and the generated abnormal images are obtained based on the embedded features. Exemplarily, a pre-trained image generation model is loaded in the server 104, and the user can input the actual abnormal images into the above pre-trained image generation model through the terminal 102. Further, the server 104 is controlled to optimize the initialized embedding vector while keeping the parameters of the image generation model unchanged, so as to obtain an optimized embedding vector representing the defective product image features distribution. Since the embodiments of the present application do not optimize the image generation model with more parameters, but optimize the embedding vector with fewer parameters, it is convenient to meet the sample requirements in the abnormal image generation process without a large number of actual abnormal images, and it is easy to implement.

[0064] Exemplarily, after determining the optimized embedding vector as described above, the user can also input the target image of the qualified product and the above-mentioned optimized embedding features into the image generation model through the terminal 102. Since the optimized embedding vector contains the image feature distribution of the defective products learned from the actual abnormal images, the generated abnormal image corresponding to the above-mentioned target image can be obtained. The generated abnormal image can be used as a negative sample for training the anomaly detection algorithm, which is beneficial to improving the detection performance of the anomaly detection algorithm.

[0065] Through the above embodiments of the present application, negative samples of the anomaly detection algorithm can be generated, which can not only solve the problem that it is difficult to obtain a large number of negative samples of defective products, but also reduce the semantic gap between the generated abnormal images and the abnormal images existing in the real world (actual abnormal images) compared with the above-mentioned synthetic abnormal images, which is beneficial to improving the detection performance of the anomaly detection algorithm, thereby making the trained model more robust.

[0066] Exemplarily, after the server 104 is equipped with the trained anomaly detection algorithm, when the user inputs a test image through the terminal 102, the server 104 performs anomaly detection on the test image. Furthermore, the detection result can be displayed on the terminal 102 for the user to view. Among them, the detection result is as Figure 1 shown, and it can be a defective product (qualified product) or a defective product (defective product). Exemplarily, whether there is color filling in the circle is used to indicate the predicted category of the current test image.

[0067] Among them, the above-mentioned terminal 102 is a smart phone, a tablet, a smart audio-video interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a wearable smart device, a medical device, etc. As mentioned before, the above-mentioned terminal 102 is often configured with a display device, and the display device can also be a monitor, a display screen, a touch screen, etc. The touch screen can also be a touch panel, a touch panel, etc. But it is not limited thereto. The above-mentioned server 104 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but it is not limited thereto. The terminal 102 and the server 104 can be directly or indirectly connected through wired or wireless communication methods, and the embodiments of the present application do not limit this here.

[0068] The technical solutions of the embodiments of the present application will be described in detail below through some embodiments. These several embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0069] Figure 2 It is a schematic flowchart of an abnormal image generation method P200 for defective products provided by an embodiment of the present application. Among them, the execution subject of method P200 can be a server, such as Figure 1 the terminal 104 in Figure 1 ; it can also be the terminal 102 in Figure 15 ; or an electronic device storing relevant algorithms, such as Figure 2 the electronic device 1500 shown. Among them, method P200 is described by taking the execution subject as a server as an example. Referring to

[0070] In S210, an optimized embedding vector is determined, where the optimized embedding vector is obtained by inputting an actual abnormal image into a pre-trained image generation model and optimizing the initialized embedding vector while keeping the parameters of the image generation model unchanged, and the optimized embedding vector represents the image feature distribution of the defective product.

[0071] In an exemplary embodiment, Figure 3 It is a schematic flowchart of the method for determining the optimized embedding vector provided by an embodiment of the present application. Referring to Figure 3 In the embodiments of the present application, the actual abnormal image is the image of a defective product existing in the real world, and the actual abnormal image can reflect the real semantic features of the abnormality. The actual abnormal image and the initialized embedding vector are input into the pre-trained image generation model. It should be noted that when the number of actual abnormal images is limited, it is unrealistic to optimize an image generation model with millions of parameters; however, when using an embedding with fewer parameters (for example, only including hundreds of parameters), the optimization process becomes relatively simple, that is, even when the number of actual abnormal images is limited, the requirement for optimizing the embedding vector can be met; that is to say, the embodiments of the present application focus on learning an embedding that can effectively capture the real abnormal semantic features, rather than relying on fine-tuning a complex model.

[0072] Exemplarily, the above image generation model can be a diffusion model (DM), a latent diffusion model (LDM), and of course, it can also be other types of generation models, which are not limited in the embodiments of the present application.

[0073] In an exemplary embodiment, Figure 4 It is a schematic flowchart of the method P400 for determining the optimized embedding vector provided by an embodiment of the present application. Referring toFigure 4 , the illustrated embodiment includes S410 - S440.

[0074] In S410, the actual abnormal image is input into the encoder of the pixel space network of the image generation model to convert the actual abnormal image into actual abnormal features.

[0075] Exemplarily, in method P400, the above - mentioned image generation model takes the latent diffusion model as an example to describe an optimized embodiment of the embedding vector. Refer to Figure 5 , the image generation model 50 includes a pixel space network 51 and a latent space network 52. The pixel space network 51 includes a pre - trained auto - encoder model (AutoEncoder), and this auto - encoder model includes an encoder ε and a decoder D. Through the encoder ε, the input image can be compressed, then a diffusion operation is performed in the latent representation space, and finally the decoder D is used for image reconstruction to restore the image to the pixel space network. Transferring the image to the latent space network for denoising processing can retain only important and basic features while ignoring imperceptible high - frequency information in the image, thereby greatly reducing the computational complexity in the training and sampling stages, enabling the generation task to generate images on a consumer - level GPU in about 10 seconds, and thus greatly reducing the implementation threshold.

[0076] Exemplarily, refer to Figure 5 , input the actual abnormal image I T into the pixel space network 51, and after encoding by the encoder ε, obtain the actual abnormal feature z in the latent space network 52, which can also be denoted as ε(I T ).

[0077] In S420, the actual abnormal feature is processed by the noise - adding network of the latent space network to obtain the noise - added feature corresponding to the actual abnormal feature.

[0078] Exemplarily, refer to Figure 5 , process the actual abnormal feature z through the noise - adding network of the latent space network 52 to obtain the noise - added feature z T corresponding to the actual abnormal feature.

[0079] In an exemplary embodiment, the number of actual abnormal images may be one or more. For example, the number of actual abnormal images may be any one of 1 to 10. When the number of actual abnormal images is more than 1, multiple actual abnormal images input into the image generation model at the same time may belong to the same type of abnormality. For example, taking hazelnuts as an example, the abnormal type of hazelnuts is actually defective hazelnuts. More specifically, the types of defects may include notches or abrasions caused by extrusion, defects caused by insect infestation, and so on. In order to enable the embedding vector to finely learn the true semantic features, in the embodiments of the present application, actual abnormal images belonging to the same type of abnormality are input into the model. For example, different abnormal images with notches caused by extrusion are input into the pre-trained image generation model at the same time, so that the embedding vector can learn the semantic features of the notches caused by extrusion.

[0080] In S430, the denoising network in the latent space network performs denoising processing on the above-mentioned noise-added features based on the initialized embedding vector, and obtains the noise predicted by the denoising network, where the above-mentioned initialized embedding vector is mapped to the denoising network in a cross-attention manner.

[0081] Exemplarily, the initialized embedding vector is first mapped to the denoising network in the latent space network in a cross-attention manner.

[0082] In the embodiments of the present application, the denoising network in the latent space network is enhanced through a cross-attention mechanism, that is, the conditional temporal denoising autoencoder is implemented t = 1, …, T. The conditional temporal denoising autoencoder aims to optimize the embedding variable to learn the semantic features reflecting the true abnormality in the actual abnormal images while freezing the model parameters.

[0083] When the denoising network in the latent space network includes T U-Nets corresponding to T denoising processes, the embedding variable is mapped to the middle layer of the j-th (positive integer less than or equal to T) U-Net through the cross-attention layer. Exemplarily, according to the embedding vector and the first parameter of the j-th U-net, the j-th key vector K is determined; according to the embedding vector and the second parameter of the j-th U-net, the j-th value vector V is determined; based on the j-th key vector K and the j-th value vector V, the initialized embedding vector is mapped to the j-th U-net.

[0084] ​Specifically, mapping the initialized embedding vector to the j-th U-net can be achieved based on formula (1):

[0085]

[0086] Wherein, represents the initialized embedding vector, ε() represents the auto-encoder of the pixel space network, and I T represents the actual abnormal image, is the intermediate layer feature representation of the j-th U-net, d represents the matrix dimension of K and Q, and are matrices related to the parameters of the j-th U-Net.

[0087] It should be noted that in the embodiments of the present application and belong to matrices related to model parameters, and remain unchanged during the process of optimizing the embedding vector, which is beneficial to obtaining embedding features that can reflect real anomalies through a small number of actual abnormal images. It can be understood that referring to and , the denoising network may include T U-nets corresponding to T denoising processes. For each U-net, the same mapping process is performed. Figure 5 Exemplarily, the denoising network mapped with the embedding vector performs denoising processing on the noise-added feature corresponding to the actual abnormal feature to obtain the noise predicted by the denoising network.

[0088] Exemplarily, referring to

[0089] , after the denoising network mapped with the embedding vector performs denoising processing on the noise-added feature z Figure 5 corresponding to the actual abnormal feature, the noise z T predicted by the denoising network can be obtained, which can also be denoted as 0

[0090] In S440, a first objective function is determined according to the noise predicted by the denoising network, and an optimized embedding vector is determined according to the first objective function.

[0091] Exemplarily, a first objective function is determined according to the noise predicted by the denoising network and the noise sampled from the Gaussian distribution; specifically, the first objective function can be determined according to the following formula (2)

[0092]

[0093] Wherein, ∈ represents the noise sampled from the Gaussian distribution, represents the noise predicted by the denoising network.

[0094] Exemplarily, the optimized embedding vector is determined by determining the minimum value of the above first objective function. Specifically, the optimized embedding vector is determined according to the following formula (3)

[0095]

[0096] Through the solution provided by method P400, since the actual abnormal images required for optimizing the embedding vector are fewer, the optimized embedding vector can be obtained based on a relatively small number of actual abnormal images Specifically, in the embodiment provided by method P400, instead of learning the parameters of the entire network of the image generation model, the embedding is started and the parameters of the image generation model are frozen, so that the parameters of the embedding can be updated separately during the noise addition and denoising processes. In this way, the learned embedding captures the distribution of the provided real-world anomalies and can be used to guide the image generation in the subsequent stage. In a further embodiment, since the optimized embedding vector represents the semantic features reflecting the real anomalies in the actual abnormal images, it can act on the normal images of good products to obtain generated abnormal images (distinguished from "synthesized abnormal images" in the related art by "generated abnormal images").

[0097] In an exemplary embodiment, the above optimized embedding vector can be determined not only based on the image generation model 50 as shown in Figure 5 , but also based on the image generation model 60 as shown in Figure 6 . Specifically, Figure 6 in the embodiment shown in T , the input image is directly loaded and denoised in the pixel space network. Exemplarily, the actual abnormal image I

[0098] As mentioned above, the number of actual abnormal images can be one or more. For example, the number of actual abnormal images can be any one of 1 to 10. When the number of actual abnormal images is more than 1, the multiple actual abnormal images input into the image generation model at the same time can belong to the same type of anomaly. In order to enable the embedding vector to finely learn the real semantic features, in the embodiments of the present application, the actual abnormal images belonging to the same type of anomaly are input into the model. For example, different abnormal images of the notch caused by extrusion are input into the pre-trained image generation model at the same time, so that the embedding vector can learn the semantic features of the notch caused by extrusion.

[0099] The noise z is obtained by processing the actual abnormal image z through the noise addition network of the image generation model 60 T。In this embodiment, the embedding vectors to be initialized are still mapped to the denoising network in the way of cross-attention, that is, the conditional denoising autoencoder is implemented t = 1, …, T. The conditional temporal denoising autoencoder aims to optimize the embedding variables while freezing the model parameters to learn the semantic features reflecting real anomalies in the actual anomaly images. When the denoising network in the latent space network includes T U-Nets corresponding to T denoising processes, the embedding variables are mapped to the middle layer of the j-th U-Net through the cross-attention layer, where the cross-attention layer is implemented as shown in formula (4). where represents the initialized embedding vector, ε() represents the autoencoder of the pixel space network, I

[0100]

[0101] wherein, represents the initialized embedding vector, ε() represents the autoencoder of the pixel space network, I T represents the actual anomaly image, is the middle layer feature representation of the j-th U-net, d represents the matrix dimension of K and Q, and are matrices related to the parameters of the j-th U-Net.

[0102] It should be noted that in the embodiments of the present application and belong to the matrices related to the model parameters, and remain unchanged during the process of optimizing the embedding vectors, which is beneficial to obtaining the embedding features that can reflect real anomalies through a relatively small number of actual anomaly images. It can be understood that referring to and , the denoising network can include T U-nets corresponding to T denoising processes, and for each U-net, the same mapping process is performed. Figure 6

[0103] Figure 6 Referring to Figure 6 , the denoising network with the embedding vectors after mapping performs denoising processing on the noise z T corresponding to the actual anomaly image, and the predicted noise z 0 of the denoising network can be obtained, which can also be denoted as The third objective function is determined according to the noise predicted by the denoising network, and the optimized embedding vector is determined according to the third objective function.

[0104] The third objective function is determined according to the following formula (5)

[0105]

[0106] Among them, ∈ represents the noise sampled from the Gaussian distribution, and represents the noise predicted by the denoising network.

[0107] The optimized embedding vector is determined according to the following formula (6)

[0108]

[0109] In an exemplary embodiment, the above optimized embedding vector can be determined not only based on the manner as Figure 5 or Figure 6 shown. Since the abnormal region in the actual abnormal image is usually a small part of the entire object. Optimizing the above embedding vector on the entire actual abnormal image may cause the data distribution learned by the optimized embedding vector to deviate towards the features of the object itself rather than specifically focusing on the features of the abnormal region. For example, a hole in a carpet, where the abnormal region "hole" is a small part of the entire object "carpet" image; optimizing the embedding vector on the abnormal image of the carpet may cause the data distribution learned by the optimized embedding vector to deviate towards the entire carpet rather than focusing on the feature distribution at the abnormal region "hole". However, one of the main problems to be solved in the embodiments of the present application is to determine the generated abnormal image, focusing on the semantic features at the abnormal region. To solve this problem, the embodiments of the present application combine the segmentation mask M T of the actual abnormal image I T to optimize the initialized embedding vector

[0110] In an exemplary embodiment, referring to Figure 7 , the image generation model adopted in the embodiment shown in this figure is the same as the Figure 5 adopted. Exemplarily, the content of the pixel space network 51 input into the model includes the actual abnormal image I T and its corresponding mask M T . Exemplarily, Figure 7 shows a case of an actual abnormal image and its corresponding mask. In other embodiments, the input of the model can include multiple actual abnormal images and their respective corresponding masks, such as the i-th mask whose boundary corresponds to the abnormal region in the i-th actual abnormal image . For example, referring to Figure 7 , for the actual abnormal image I T of hazelnuts, the boundary of its corresponding mask M T is the broken edge of the hazelnut.

[0111] In this embodiment, after inputting the actual abnormal image and its corresponding mask image into the pre-trained image generation model, without changing the parameters of the image generation model, based on the features inside the boundary of the corresponding mask in the above actual abnormal image, the second objective function is determined. Exemplarily, the difference between the noise sampled from the Gaussian distribution and the noise predicted by the denoising network is determined; then, according to the difference and the features inside the boundary of the mask in the actual abnormal image, the second objective function is determined. Specifically, it is determined as the objective function of the model according to formula (7) (which can be denoted as the second objective function).

[0112] where ∈ represents the noise sampled from the Gaussian distribution, represents the noise predicted by the denoising network, I T represents the actual abnormal image, M T represents the mask of the actual abnormal image I T .

[0113] Furthermore, the optimized embedding vector is determined according to the above second objective function as the following formula (8):

[0114]

[0115] In the embodiment of determining the optimized embedding vector as shown in Figure 7 , by introducing the mask of the actual abnormal image, it is targeted to guide the capture of local details of the object. That is, by guiding the learning method for the abnormal area, it can ensure that the generated image exhibits semantic controllability while also controlling spatial controllability. Specifically, at the semantic level, through the actual abnormal image, it aims to generate an image consistent with the abnormality in the real world, maintaining the consistency of the object or texture (e.g., bottle) and the type of abnormality (e.g., breakage); at the spatial level, by providing the mask, the position, size, and quantity of the abnormal area can be precisely controlled. Based on semantic controllability and spatial controllability, it can prompt the optimized embedding vector to accurately learn the real abnormal semantic features, which is beneficial to generating high-quality abnormal images in the following embodiments.

[0116] It can be understood that the method of determining the embedding vector based on the actual abnormal image and its corresponding mask can be applied not only to the image generation model as shown in Figure 5 , but also to the image generation model as shown in Figure 6 , which will not be elaborated here.

[0117] Continue to refer to Figure 2, after determining the optimized embedded vector, execute S220: input the target image of the good product and the optimized embedded feature into the image generation model to obtain the generated abnormal image corresponding to the target image.

[0118] In an exemplary embodiment, since the above optimized embedded vector learns the real abnormal semantic features, the optimized embedded vector can be mapped to the denoising network of the model in a cross-attention manner to control the generation of abnormalities in the target image, and then output the generated abnormal image. Exemplarily, referring to Figure 8 , input the target image of the good product and the optimized embedded vector into the pre-trained image generation model, and the output of the model is the generated abnormal image corresponding to the above target image.

[0119] Exemplarily, referring to Figure 9 , in S220-1, input the target image I normal (i.e., the image without abnormal regions) into the encoder ε of the pixel space network of the image generation model 50, so that the input image I normal can be compressed through the encoder ε, and the target feature z in the latent space network 52 is obtained after encoding by the encoder ε, which can also be denoted as ε(I normal ). Then, perform a diffusion operation on the target feature z in the latent representation space, and finally use the decoder D for image reconstruction to restore the image to the pixel space network. Transfer the target image I normal to the latent space network 52 for denoising processing, which can retain only important and basic features and ignore the imperceptible high-frequency information in the image, thereby greatly reducing the computational complexity in the training and sampling stages.

[0120] Exemplarily, referring to Figure 9 , in S220-2, the target feature z is processed by the noise addition network of the latent space network 52 to obtain the noise-added feature z T corresponding to the target feature. In S220-3, the optimized embedded vector is mapped to the denoising network of the model 50 in a cross-attention manner . In the embodiment of generating the abnormal image in this application, the optimized embedded variable is frozen and injected into the image as a condition through the cross-attention module, so as to generate the expected abnormality. Specifically, the denoising network in the latent space network is enhanced through the cross-attention mechanism, that is, the conditional temporal denoising autoencoder conditional temporal denoising autoencoder aims to Under the influence of T Predict a corresponding denoised variable or predict noise, where z T is the input ε(I normal ) to generate abnormal images output by model 50 It contains semantic features that reflect the real anomaly.

[0121] In the case where the denoising network of model 50 is U-Net, the optimized embedding variables are transformed through the cross attention layer. Mapped to the middle layer of each U-Net (taking the h-th U-Net as an example, h is a positive integer less than or equal to T), the cross attention layer implements the following formula (9):

[0122]

[0123] in, represents the optimized embedding vector, ε() represents the automatic encoder of the pixel space network, I normal The target image representing good products, is the feature expression of the middle layer of the h-th U-net, d represents the matrix dimension of K and Q, and is a matrix related to the hth U-Net parameter. It can be understood that the denoising network may include T U-nets corresponding to T denoising processes, and the same mapping process is performed for each U-net.

[0124] Exemplarily, in S220-4, by mapping the optimized embedding vector The denoising network adds noise features z corresponding to the target features T After denoising, we can get the noise z predicted by the denoising network. 0 , which can also be written as

[0125] It should be noted that in the embodiments of this application and It is a matrix related to the model parameters. In the process of determining the generation of abnormal images, all the model parameters (including and ) remains unchanged and keeps the optimized embedding vector Under the condition of no change, guide the generation of target image I normal The corresponding abnormal image is generated.

[0126] In S220-5, the noise z predicted by the encoder D denoising network 0 Reconstruct and obtain the generated abnormal image corresponding to the target image

[0127] Through the solution provided by the embodiments of the present application, since the optimized embedding vector learns the distribution of real-world anomalies provided by the actual abnormal images, the optimized embedding vector is introduced into the denoising network and acts on the normal images of good products to guide the generation of abnormal images in the image generation stage, so that the generated abnormal images can contain the semantic features of real anomalies.

[0128] In an exemplary embodiment, to adjust any one of the position, size, and quantity of the abnormal regions in the generated abnormal images, a bounding box can be introduced. Specifically, the bounding box and the above-mentioned target image can be input into a pre-trained image generation model. Exemplarily, referring to Figure 10 , the image generation model adopted in the embodiment shown in this figure is the same as Figure 5 the image generation model adopted. Exemplarily, the content input into the pixel space network 51 of the model includes the target image I normal of good products and its corresponding bounding box M box , where the position of the bounding box M box can correspond to any position in the target image I normal . In the corresponding generated abnormal image , the position of the bounding box M box is the abnormal region. For example, referring to Figure 10 , for the target image I normal of hazelnuts, the inner side of the bounding box M box input into the model together with it is the hazelnut breakage region in the generated abnormal image.

[0129] In the embodiments of the present application, for each inference image (such as Figure 10 z T-1 ... z 0 ) during the denoising process, the region inside the bounding box M box will be retained, and the region outside the bounding box M box will be replaced by a noisy version of the target image I normal . In this way, the embodiments of the present application can control the generated abnormal region to be located in the specified region of the target image while keeping other regions as unchanged as possible.

[0130] Exemplarily, input the target image of a good product into the image generation model. Through the noise addition network of the image generation model, perform T steps of noise addition processing on the target image to obtain the noise-added features of the target image. The noise-added features include the noise-added features of t steps. T is a positive integer, and t + 1 is a positive integer not greater than T. And, perform T steps of denoising processing on the noise-added features of the target image through the denoising network mapped with the optimized embedded features. Wherein, the generated abnormal feature z obtained in the t-step denoising processing t can be determined according to the following formula (10):

[0131]

[0132] Wherein, represents the noise-added feature of ε(I normal ) at the t-th step, and z t ' represents the denoised feature of z t+1 , and M box represents the bounding box corresponding to the target image. It can be seen that the generated abnormal feature z obtained in the t-step denoising processing t includes: according to the first part of the features outside the bounding box in the loaded features of the t-th step, that is And, the second part of the features z t '⊙M box inside the bounding box in the denoised features corresponding to the loaded features of the (t + 1)-th step.

[0133] Furthermore, referring to Figure 10 , after decoding the generated abnormal feature z 0 through the decoder D of the pixel space network 51, the generated abnormal image corresponding to the above target image I normal is obtained

[0134] In the embodiments of the present application, based on an image generation model such as a diffusion model, an image (i.e., the above-mentioned generated abnormal image) is generated from random noise. Therefore, the generated abnormal images obtained have a certain degree of semantic diversity. Further, in order to further enhance the diversity, the above-mentioned bounding box can be introduced in the process of generating the abnormal images of the present application to adjust the position, size, quantity, etc. in the abnormal area, which is beneficial to improving the diversity of the generated abnormal images.

[0135] In an exemplary embodiment, the embodiments of the present application further provide a method for generating training samples of a defective product detection model. Wherein, the method includes: S1. Determine a positive sample image. Wherein, the positive sample image refers to an image of a good product (normal product) without abnormal areas. S2. Generate a generated abnormal image corresponding to the positive sample image through the abnormal image generation method for defective products provided in the above embodiments to obtain a negative sample image.

[0136] As described above, it is difficult to obtain a large number of negative samples of defective products. In the related art, there is a significant distribution gap between the synthetic abnormal images and the abnormal images or features in the real world, which is not conducive to training an abnormal detection model in practical applications. In the sample generation scheme provided in the embodiments of the present application, the defective product image features in the actual abnormal images are learned through embedding vectors, and the generated abnormal images are obtained based on the embedded features. Therefore, the embodiments of the present application can not only solve the problem that it is difficult to obtain a large number of negative samples of defective products, but also reduce the semantic gap between the generated abnormal images and the abnormal images (actual abnormal images) in the real world compared with the above synthetic abnormal images, which is beneficial to improving the detection performance of the abnormal detection algorithm, thereby making the trained model more robust.

[0137] The above provides an overall introduction to the method for generating abnormal images of defective products and the method for generating training samples of a defective product detection model provided in the embodiments of the present application. The following further introduces the method for generating abnormal images of defective products provided in the embodiments of the present application through specific embodiments.

[0138] Figure 11 It is a schematic diagram for generating negative samples of a defective product detection model according to a small number of actual abnormal images provided in an embodiment of the present application.

[0139] Since it is difficult to obtain a large number of negative samples of defective products, in the embodiments of the present application, negative samples including the defective product image features in the actual abnormal images can be generated based on a small number of actual abnormal images. Refer to Figure 11 , and only 3 actual abnormal images 111 are used to achieve the above goal. Specifically, based on the specific implementation manner of S210 in Figure 2 , the embedding vectors for learning the defective product image features in the actual abnormal images can be determined through 3 actual abnormal images 111.

[0140] Further, for the randomly obtained images and bounding boxes 112 of the normal products of the product, based on the specific implementation manner of S220 in Figure 2 , the generated abnormal image 114 can be obtained. For the randomly obtained images and bounding boxes 113 of the normal products of the product, based on the specific implementation manner of S220 in Figure 2 , the generated abnormal image 115 can be obtained. It can be understood that the image 112 of the normal product of the product can be a positive sample of the abnormal detection algorithm, and the generated abnormal image 114 obtained based on the specific implementation manner of S220 in Figure 2 can be used as a negative sample of the abnormal detection algorithm.

[0141] It can be seen that the embodiment of the present application proposes a solution that can drive the generation of more realistic and diverse abnormal images with a small number of actual abnormal images. Specifically, based on the embedding vectors of some given real abnormalities, the real abnormal distribution is learned. Further, the optimized embedding vectors and the given bounding boxes are used to guide the generation model to generate real and diverse abnormalities on specific objects, thereby solving the technical problem of difficult large-scale acquisition of abnormal samples. At the same time, compared with the synthetic abnormal images in the solutions provided by the related art, the solution provided by the embodiment of the present application can determine the generated abnormal images that can reflect the real abnormal semantic features. Therefore, the intelligent detection algorithm trained with the generated abnormal images in the embodiment of the present application as negative samples has higher detection performance.

[0142] In an exemplary embodiment, the embodiment of the present application further provides a defective product detection method. Specifically, the method includes: detecting a to-be-detected object through a trained defective product detection model to obtain a detection result, where the detection result includes that the type of the to-be-detected object is a qualified product or a defective product, and when the detection result is that the to-be-detected object is a defective product, the detection result further includes the positioning result of the abnormal area. Wherein, the defective product detection model is obtained by training with samples determined by the training sample generation method of the defective product detection model in the above embodiment. Since the training sample generation method of the defective product detection model can provide rich negative samples and positive samples, and the negative samples are generated abnormal images containing real abnormal features, compared with the synthetic abnormal images in the solutions provided by the related art, the intelligent detection algorithm trained with the generated abnormal images in the embodiment of the present application as negative samples has higher detection performance.

[0143] As described above in conjunction with Figures 1 to 11 ,the method embodiments of the present application have been described in detail. Below in conjunction with Figures 12 to 14 ,the device embodiments of the present application will be described in detail.

[0144] Figure 12 FIG. 14 is a schematic structural diagram of an abnormal image generation device 1200 for defective products provided by an embodiment of the present application.

[0145] Refer to Figure 12, the abnormal image generation device 1200 for defective products includes: a first determination module 1210 and a first generation module 1220; wherein, the first determination module 1210 is configured to determine an optimized embedding vector, where the optimized embedding vector is obtained by inputting an actual abnormal image into a pre-trained image generation model and optimizing the initialized embedding vector while keeping the parameters of the image generation model unchanged, and the optimized embedding vector represents the image feature distribution of the defective products; and the first generation module 1220 is configured to input a target image of non-defective products and the optimized embedding feature into the image generation model to obtain a generated abnormal image corresponding to the target image.

[0146] In an exemplary embodiment, based on the foregoing solution, the abnormal image generation device 1200 for defective products further includes: an embedding vector acquisition module; wherein, before the first determination module 1210 determines the optimized embedding vector, the embedding vector acquisition module is configured to: input the initialized embedding vector and the actual abnormal image into the pre-trained image generation model, and optimize the embedding vector while keeping the parameters of the image generation model unchanged.

[0147] In an exemplary embodiment, based on the foregoing solution, the image generation model includes a pixel space network and a latent space network;

[0148] The embedding vector acquisition module includes: a first input unit, a denoising unit, a first determination unit, and a second determination unit;

[0149] Wherein, the first input unit is configured to: input the actual abnormal image into the encoder of the pixel space network to convert the actual abnormal image into an actual abnormal feature, where the actual abnormal feature will be processed by the noise addition network of the latent space network to obtain a corresponding noise-added feature; the denoising unit is configured to: the denoising network in the latent space network performs denoising processing on the noise-added feature based on the initialized embedding vector to obtain the noise predicted by the denoising network, where the initialized embedding vector is mapped to the denoising network in a cross-attention manner; the first determination unit is configured to: determine a first objective function according to the noise predicted by the denoising network; the second determination unit is configured to: determine the optimized embedding vector according to the first objective function.

[0150] In an exemplary embodiment, based on the foregoing solution, the denoising network includes T U-nets corresponding to T denoising processes, where T is a positive integer; the embedding vector acquisition module further includes a mapping unit; wherein, the mapping unit is configured to: determine a j-th key vector according to the initialized embedding vector and the first parameter of the j-th U-net, where j is a positive integer less than or equal to T; determine a j-th value vector according to the initialized embedding vector and the second parameter of the j-th U-net; and map the initialized embedding vector to the j-th U-net based on the j-th key vector and the j-th value vector.

[0151] In an exemplary embodiment, based on the foregoing solution, the first determination unit is specifically configured to: determine a first objective function according to the noise predicted by the denoising network and the noise sampled from a Gaussian distribution; the second determination unit is specifically configured to: determine an optimized embedding vector by determining the minimum value of the first objective function.

[0152] In an exemplary embodiment, based on the foregoing solution, the embedding vector acquisition module includes: a second input unit, a third determination unit, and a fourth determination unit; wherein the second input unit is configured to: input an actual abnormal image and its corresponding mask image into a pre-trained image generation model; the third determination unit is configured to: based on the features of the actual abnormal image inside the boundary of the mask while keeping the parameters of the image generation model unchanged, determine a second objective function; the fourth determination unit is configured to: determine an optimized embedding vector according to the second objective function; wherein, the inside of the boundary of the mask corresponds to the abnormal region in the actual abnormal image.

[0153] In one implementation manner, based on the foregoing solution, the third determination unit is configured to: determine the difference between the noise sampled from a Gaussian distribution and the noise predicted by the denoising network; and determine a second objective function according to the difference and the features of the actual abnormal image inside the boundary of the mask; the fourth determination unit is specifically configured to: determine an optimized embedding vector by determining the minimum value of the second objective function.

[0154] In an exemplary embodiment, based on the foregoing solution, the first generation module 1220 is specifically configured to: input a target image of a good product, the bounding box corresponding to the target image, and the optimized embedding feature into the image generation model to obtain a generated abnormal image corresponding to the target image; wherein, the inside of the bounding box in the generated abnormal image is the abnormal region.

[0155] In an exemplary embodiment, based on the foregoing solution, the first generation module 1220 includes: a first generation unit and a second generation unit; wherein, the first generation unit is configured to: input a target image of a good product into the image generation model, and perform T-step noise addition processing on the target image through the noise addition network of the image generation model to obtain the noise-added features of the target image, the noise-added features including the noise-added features at t steps, T being a positive integer, and t + 1 being a positive integer not greater than T; perform T-step denoising processing on the noise-added features of the target image through the denoising network mapped with the optimized embedded features; wherein, the generated abnormal features obtained by the denoising processing at the t-th step include: according to the first part of the features located outside the bounding box in the loaded features at the t-th step, and the second part of the features located inside the bounding box in the denoised features corresponding to the loaded features at the (t + 1)-th step; and, the second generation unit is configured to: decode the generated abnormal features after the T-step denoising processing through the decoder of the pixel space network to obtain the generated abnormal image corresponding to the target image.

[0156] In the abnormal image generation solution provided by the embodiments of the present application, an actual abnormal image is input into a pre-trained image generation model, and while keeping the parameters of the image generation model unchanged, the initialized embedding vector is optimized to obtain an optimized embedding vector representing the image feature distribution of the defective product. Since the embodiments of the present application do not optimize the image generation model with more parameters, but optimize the embedding vector with fewer parameters, it is convenient to meet the sample requirements in the abnormal image generation process without a large number of actual abnormal images and is easy to implement. On the other hand, inputting the target image of the good product and the above-mentioned optimized embedded features into the image generation model, since the optimized embedding vector contains the image feature distribution of the defective product learned according to the actual abnormal image, the generated abnormal image corresponding to the target image can be obtained. The generated abnormal image can be used as a negative sample for training an abnormal detection algorithm, which is beneficial to improving the detection performance of the abnormal detection algorithm.

[0157] It should be understood that the embodiment of the abnormal image generation device for defective products and the embodiment of the abnormal image generation method for defective products can correspond to each other, and similar descriptions can refer to the method embodiment. To avoid repetition, it will not be elaborated here. Specifically, Figure 12 The shown abnormal image generation device for defective products can execute the embodiment of the abnormal image generation method for defective products, and the foregoing and other operations and / or functions of each module in the device are respectively for implementing the embodiment of the abnormal image generation method for defective products. For the sake of brevity, it will not be elaborated here.

[0158] Figure 13 It is a structural schematic diagram of a training sample generation device 1300 for a defective product detection model provided by the embodiments of the present application.

[0159] Reference Figure 13 , the training sample generation device 1300 of the defective product detection model includes: a second determination module 1310 and a second generation module 1320; wherein, the second determination module 1310 is configured to determine positive sample images; and, the second generation module 1320 is configured to generate a generated abnormal image corresponding to the positive sample image by the abnormal image generation method for defective products provided in the foregoing embodiments, so as to obtain negative sample images.

[0160] It is difficult to obtain a large number of negative samples of defective products on a large scale, and there is a significant distribution gap between the abnormal images obtained by synthesis in related technologies and the abnormal images or features in the real world, which is not conducive to training an abnormal detection model in practical applications. In the sample generation solution provided in the embodiments of the present application, the defective product image features in the actual abnormal images are learned through an embedding vector, and the generated abnormal images are obtained based on the embedding features. Therefore, the embodiments of the present application can not only solve the problem that it is difficult to obtain a large number of negative samples of defective products on a large scale, but also reduce the semantic gap between the generated abnormal images and the abnormal images (actual abnormal images) existing in the real world compared with the above-mentioned synthetic abnormal images, which is beneficial to improving the detection performance of the abnormal detection algorithm, thereby making the trained model more robust.

[0161] It should be understood that the embodiments of the training sample generation device of the defective product detection model and the embodiments of the training sample generation method of the defective product detection model can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, it will not be elaborated here. Specifically, Figure 13 The shown training sample generation device of the defective product detection model can execute the embodiments of the training sample generation method of the defective product detection model, and the foregoing and other operations and / or functions of each module in the device are respectively for implementing the embodiments of the training sample generation method of the defective product detection model. For the sake of brevity, it will not be elaborated here.

[0162] Figure 14 FIG. 15 is a schematic structural diagram of a defective product detection device 1400 provided in an embodiment of the present application.

[0163] Reference Figure 14 , the defective product detection device 1400 includes: a detection module 1410; wherein, the detection module 1410 is configured to detect a to-be-detected object through a trained defective product detection model to obtain a detection result; wherein, the detection result includes that the type of the to-be-detected object is a qualified product or a defective product, and in the case that the detection result is that the to-be-detected object is a defective product, the detection result further includes a positioning result of an abnormal area, and the defective product detection model is obtained by training with samples determined according to the training sample generation method of the defective product detection model.

[0164] It should be understood that the embodiments of the defective product detection device and the embodiments of the defective product detection method can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they will not be elaborated here. Specifically, Figure 14 The illustrated defective product detection device can execute the embodiments of the above-mentioned defective product detection method, and the foregoing and other operations and / or functions of each module in the device respectively implement the embodiments of the defective product detection method. For the sake of brevity, they will not be elaborated here.

[0165] In the foregoing, the device of the embodiment for generating a building information model of the present application has been described from the perspective of functional modules. It should be understood that the functional module can be implemented in the form of hardware, can also be implemented by instructions in software form, and can also be implemented by a combination of hardware and software modules. Specifically, each step of the method embodiment in the embodiments of the present application can be completed by the integrated logic circuit in hardware in the processor and / or instructions in software form. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiment.

[0166] Figure 15 is a schematic block diagram of the electronic device 1500 provided by the embodiments of the present application. Figure 15 The electronic device 1500 can be used to execute the above-mentioned method for generating abnormal images of defective products, or execute the above-mentioned method for generating training samples of the defective product detection model, or execute the above-mentioned defective product detection method. As Figure 15 shown, the electronic device 1500 may include:

[0167] A memory 1510 and a processor 1520. The memory 1510 is used to store a computer program 1530 and transmit the program code 1530 to the processor 1520. In other words, the processor 1520 can call and run the computer program 1530 from the memory 1510 to implement the method in the embodiments of the present application.

[0168] For example, the processor 1520 can be used to execute the steps in the above method according to the instructions in the computer program 1530.

[0169] In some embodiments of the present application, the processor 1520 may include but is not limited to:

[0170] General-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.

[0171] In some embodiments of the present application, the memory 1510 includes, but is not limited to:

[0172] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0173] In some embodiments of the present application, the computer program 1530 can be divided into one or more modules, which are stored in the memory 1510 and executed by the processor 1520 to complete the above-mentioned abnormal image generation method for defective products provided by the present application, or execute the above-mentioned training sample generation method for the defective product detection model, or execute the above-mentioned defective product detection method. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 1530 in the electronic device.

[0174] As shown Figure 15 the electronic device 1500 may further include:

[0175] a transceiver 1540, which may be connected to the processor 1520 or the memory 1510.

[0176] Among them, the processor 1520 may control the transceiver 1540 to communicate with other devices. Specifically, it may send information or data to other devices, or receive information or data sent by other devices. The transceiver 1540 may include a transmitter and a receiver. The transceiver 1540 may further include an antenna, and the number of antennas may be one or more.

[0177] It should be understood that the various components in the electronic device 1530 are connected through a bus system. Among them, the bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus.

[0178] According to one aspect of the present application, there is provided a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer can execute the methods in the above method embodiments. Or, the embodiments of the present application further provide a computer program product including instructions. When the instructions are executed by a computer, the computer executes the methods in the above method embodiments.

[0179] According to another aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods in the above method embodiments.

[0180] In other words, when implemented using software, it can be implemented in the form of a computer program product in whole or in part. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0181] Those of ordinary skill in the art will realize that the modules and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0182] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.

[0183] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of the present application, the various functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0184] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for generating an abnormal image of a defective product, characterized in that: The method comprises: Determine an optimized embedding vector, wherein the optimized embedding vector is obtained by inputting an actual abnormal image into a pre-trained image generation model and optimizing an initialized embedding vector while keeping parameters of the image generation model unchanged, and the optimized embedding vector represents the image feature distribution of the defective product; Inputting the target image of a good product and the optimized embedding vector into the image generation model to obtain a generated abnormal image corresponding to the target image; Before determining the optimized embedding vector, the method further includes: Inputting the initialized embedding vector and the actual abnormal image into a pre-trained image generation model, and optimizing the embedding vector while keeping the parameters of the image generation model unchanged; The image generation model includes a pixel space network and a latent space network; the step of inputting the initialized embedding vector and the actual abnormal image into the pre-trained image generation model includes: Inputting the actual abnormal image into the encoder of the pixel space network to convert the actual abnormal image into actual abnormal features, wherein the actual abnormal features are processed by the noise adding network of the latent space network to obtain corresponding noise adding features; The denoising network in the latent space network denoises the noisy features based on the initialized embedding vector to obtain the noise predicted by the denoising network, wherein the initialized embedding vector is mapped to the denoising network in a cross-attention manner; The method further comprises: Determining a first objective function based on the noise predicted by the denoising network; An optimized embedding vector is determined according to the first objective function.

2. The method according to claim 1, characterized in that: The denoising network includes T U-Nets corresponding to T denoising processes, where T is a positive integer; the method further includes: Determine a j-th key vector according to the initialized embedding vector and a first parameter of the j-th U-net, where j is a positive integer less than or equal to T; Determine a j-th value vector according to the initialized embedding vector and a second parameter of the j-th U-net; Mapping the initialized embedding vector to the j-th U-net based on the j-th key vector and the j-th value vector.

3. The method according to claim 1 or 2, characterized in that: The determining of a first objective function according to the noise predicted by the denoising network comprises: Determining a first objective function according to the noise predicted by the denoising network and the noise sampled by the Gaussian distribution; The step of determining the optimized embedding vector according to the first objective function includes: An optimized embedding vector is determined by determining the minimum value of the first objective function.

4. The method according to claim 1, characterized in that: The step of inputting the initialized embedding vector and the actual abnormal image into a pre-trained image generation model, and optimizing the embedding vector while keeping the parameters of the image generation model unchanged, comprises: Input the actual abnormal image and its corresponding mask image into the pre-trained image generation model; Under the condition that the parameters of the image generation model remain unchanged, determining a second objective function based on the features in the actual abnormal image that are inside the boundary of the mask; Determine an optimized embedding vector according to the second objective function; The inner side of the boundary of the mask corresponds to the abnormal area in the actual abnormal image.

5. The method according to claim 4, characterized in that The determining of the second objective function based on the features in the actual abnormal image that are inside the boundary of the mask while keeping the parameters of the image generation model unchanged comprises: Determining the difference between the noise sampled from the Gaussian distribution and the noise predicted by the denoising network in the image generation model; Determining a second objective function according to the difference and the features in the actual abnormal image that are inside the boundary of the mask; Determining an optimized embedding vector according to the second objective function includes: The optimized embedding vector is determined by determining the minimum value of the second objective function.

6. The method according to claim 4 or 5, characterized in that: The step of inputting the target image of a good product and the optimized embedding vector into the image generation model to obtain a generated abnormal image corresponding to the target image includes: Inputting a target image of a good product, a bounding box corresponding to the target image, and the optimized embedding vector into the image generation model to obtain a generated abnormal image corresponding to the target image; The abnormal area in the generated abnormal image is located inside the boundary box.

7. The method according to claim 4 or 5, characterized in that: The step of inputting the target image of a good product and the optimized embedding vector into the image generation model to obtain a generated abnormal image corresponding to the target image includes: Inputting a target image of a good product into the image generation model, performing T-step noise processing on the target image through the noise adding network of the image generation model, and obtaining a noise adding feature of the target image, wherein the noise adding feature includes the noise adding feature of t steps, where T is a positive integer and t+1 is a positive integer not greater than T; The denoising network to which the optimized embedding vector is mapped performs T-step denoising on the denoised features of the target image; wherein the generated abnormal features obtained in the t-step denoising process include: a first part of features located outside the bounding box corresponding to the target image in the loaded features of the t-step, and a second part of features located inside the bounding box in the denoised features corresponding to the loaded features of the t+1-step; After decoding the abnormal features generated after T-step denoising processing through the decoder of the pixel space network in the image generation model, the generated abnormal image corresponding to the target image is obtained.

8. A method for generating training samples for a defective product detection model, characterized in that: The method comprises: Determine positive sample images; The abnormal image corresponding to the positive sample image is generated by the method described in any one of claims 1 to 7 to obtain a negative sample image.

9. A method for detecting defective products, characterized in that: The method comprises: The trained defective product detection model is used to detect the object to be tested and obtain the detection result; Wherein, the detection result includes whether the type of the object to be detected is a good product or a defective product. When the detection result is that the object to be detected is a defective product, the detection result also includes a positioning result of the abnormal area; The defective product detection model is obtained by training according to the samples determined according to claim 8.

10. A device for generating an abnormal image of a defective product, characterized in that: The device comprises: A first determination module is used to determine an optimized embedding vector, wherein the optimized embedding vector is obtained by inputting an actual abnormal image into a pre-trained image generation model and optimizing an initialized embedding vector while keeping parameters of the image generation model unchanged, and the optimized embedding vector represents the image feature distribution of the defective product; A first generation module, configured to input a target image of a good product and the optimized embedding vector into the image generation model to obtain a generated abnormal image corresponding to the target image; The abnormal image generation device for defective products further includes: an embedding vector acquisition module; wherein the embedding vector acquisition module is used to: before the first determination module determines the optimized embedding vector, input the initialized embedding vector and the actual abnormal image into a pre-trained image generation model, and optimize the embedding vector while keeping the parameters of the image generation model unchanged; The image generation model includes a pixel space network and a potential space network; the embedding vector acquisition module includes: a first input unit, a denoising unit, a first determination unit and a second determination unit; Wherein, the first input unit is used to: input the actual abnormal image into the encoder of the pixel space network to convert the actual abnormal image into an actual abnormal feature, wherein the actual abnormal feature is processed by the denoising network of the latent space network to obtain a corresponding denoised feature; the denoising unit is used to: the denoising network in the latent space network denoises the denoised feature based on the initialized embedding vector to obtain the noise predicted by the denoising network, wherein the initialized embedding vector is mapped to the denoising network in a cross-attention manner; the first determination unit is used to: determine a first objective function according to the noise predicted by the denoising network; the second determination unit is used to: determine an optimized embedding vector according to the first objective function.

11. A training sample generating device for a defective product detection model, characterized in that: The device comprises: A second determination module, used to determine a positive sample image; The second generating module is used to generate an abnormal image corresponding to the positive sample image by the method described in any one of claims 1 to 7 to obtain a negative sample image.

12. A defective product detection device, characterized in that: The device comprises: A detection module, used to detect the object to be detected by using the trained defective product detection model to obtain the detection result; Wherein, the detection result includes whether the type of the object to be detected is a good product or a defective product. When the detection result is that the object to be detected is a defective product, the detection result also includes a positioning result of the abnormal area; The defective product detection model is obtained by training according to the samples determined according to claim 8.

13. An electronic device comprising a processor and a memory; The memory is used to store computer programs; The processor is used to execute the computer program to implement the method for generating abnormal images of defective products as described in any one of claims 1 to 7, or the method for generating training samples of a defective product detection model as described in claim 8, or the method for detecting defective products as described in claim 9.

14. A computer-readable storage medium, characterized in that: For storing computer programs; The computer program enables the computer to execute the method for generating abnormal images of defective products as described in any one of claims 1 to 7, or the method for generating training samples of a defective product detection model as described in claim 8, or the method for detecting defective products as described in claim 9.

Citation Information

Patent Citations

  • Auto-encoder anomaly detection method based on comparative learning

    CN114724043A