Defect picture generation method and device, electronic equipment and storage medium

By using the defect image generation model to extract features from the workpiece image and text prompt information, and combining the shape information in the mask image to generate the target defect image, the problem of the inability to accurately control the shape and position of the defect generated in the prior art is solved, and the precise control of the defect image and the actual needs are achieved.

CN120198549AActive Publication Date: 2025-06-24SHENZHEN XINRUN FULIAN DIGITAL TECH CO LTD

Patent Information

Application Number
CN202510683309.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

The prior art cannot accurately control the shape and position of defects generated, resulting in the generated defect images that do not meet actual needs and cannot be used as ideal negative samples.

Method used

By acquiring the workpiece image, text prompt information and mask image, a pre-trained defect image generation model is used to extract the generated features from the workpiece image and text prompt information, shape features are extracted from the mask image, and target defect images are generated based on these features, so that they satisfy the constraints of the mask information on the defect shape and position.

Benefits of technology

Accurate control of defect shape and position is achieved, so that the generated defect image meets actual needs and can be used as an ideal negative sample.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198549A_ABST
    Figure CN120198549A_ABST
Patent Text Reader

Abstract

The invention relates to a defect picture generation method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a workpiece image, text prompt information and a mask image, the mask image is generated based on the workpiece image and mask information, and the mask information is used for representing the shape and position of a to-be-generated defect on the workpiece image; the workpiece image, the text prompt information and the mask image are input into a pre-trained defect image generation model, a target defect image is obtained, the defect image generation model is used for extracting a first generation feature from the workpiece image and the text prompt information, extracting a first shape feature from the mask image, and generating a second generation feature; based on the first generation feature and the first shape feature, a target defect image is generated, and the defect generated in the target defect image meets the constraint of the mask information on the defect shape and the defect position. In this way, accurate control over the defect shape and the defect position can be achieved, and the target defect image meets the actual requirement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of industrial vision technology, and particularly to a method, device, electronic device, and storage medium for generating defective pictures. Background Art

[0002] In the field of industrial vision technology, vision detection models need to balance issues such as detection speed and edge deployment, and usually lightweight models are required. However, these models often need a large number of samples for training, so a large number of defective images are required as negative samples.

[0003] In related technologies, usually a Generative Adversarial Networks (GAN) is used to generate scarce negative samples. However, with this method, the shape and position of the generated defects cannot be accurately controlled, resulting in the finally generated defective images not meeting the actual needs and unable to be used as ideal negative samples. Therefore, how to accurately control the shape and position of defect generation has become a technical problem to be urgently solved. Summary of the Invention

[0004] This application provides a method, device, electronic device, and storage medium for generating defective pictures to solve the problem in related technologies that the shape and position of defect generation cannot be accurately controlled, resulting in the finally generated defective images not meeting the actual needs and unable to be used as ideal negative samples.

[0005] In a first aspect, an embodiment of this application provides a method for generating a defective picture. The method includes: Obtain a workpiece image, text prompt information, and a mask image, where the mask image is generated based on the workpiece image and mask information, and the mask information is used to characterize the shape and position of the defect to be generated on the workpiece image; Input the workpiece image, the text prompt information, and the mask image into a pre-trained defective image generation model to obtain a target defective image; Wherein, the defective image generation model is used to extract a first generation feature from the workpiece image and the text prompt information, extract a first shape feature from the mask image, and generate the target defective image based on the first generation feature and the first shape feature. The generated defect in the target defective image satisfies the constraints of the mask information on the defect shape and defect position.

[0006] Optionally, the defective image generation model includes a shape constraint sub-module and a StableDiffusion sub-model, and the StableDiffusion sub-model includes a text encoder, a latent space diffuser, an image encoder, and an image decoder; Inputting the workpiece image, the text prompt information, and the mask image into a pre-trained defect image generation model to obtain a target defect image includes: Inputting the workpiece image into the image encoder to obtain a latent vector corresponding to the workpiece image; Inputting the text prompt information into the text encoder to obtain a semantic embedding vector corresponding to the text prompt information; Inputting the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information into the latent space diffuser to obtain the first generated feature; Inputting the mask image into the shape constraint sub-module to obtain the first shape feature; Performing feature fusion on the first generated feature and the first shape feature, and performing blurred mixing on the fused feature and the mask image; Inputting the mixed feature into the image decoder to generate the target defect image.

[0007] Optionally, inputting the mask image into the shape constraint sub-module to obtain the first shape feature includes: Encoding the mask image using the shape constraint sub-module to obtain a first latent vector corresponding to the mask image, and performing multiple downsampling operations on the mask image to obtain a second latent vector corresponding to the mask image, where the first latent vector corresponding to the mask image, the second latent vector corresponding to the mask image, and the latent vector corresponding to the workpiece image have the same dimension; Performing noise mixing on the first latent vector corresponding to the mask image and the second latent vector corresponding to the mask image, and performing multi-level feature extraction on the mixed latent vector to obtain the first shape feature.

[0008] Optionally, inputting the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information into the latent space diffuser to obtain the first generated feature includes: Using the latent space diffuser to fuse the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information, and gradually denoising the fused vector to obtain the first generated feature.

[0009] Optionally, the latent space diffuser includes a plurality of cross-attention layers, and some or all of the plurality of cross-attention layers include a frozen original weight matrix, a first low-rank matrix, and a second low-rank matrix introduced by low-rank adaptation. The first low-rank matrix and the second low-rank matrix corresponding to each cross-attention layer are pre-trained based on a plurality of first training samples, and each of the first training samples includes a first workpiece sample image, a first defect sample image, and a first text annotation; The fusing of the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information by using the latent space diffuser includes: Based on the first low-rank matrix and the second low-rank matrix corresponding to each cross-attention layer, determine the weight increment corresponding to each cross-attention layer, where the weight increment corresponding to each cross-attention layer is the calculation result of the product of the first low-rank matrix and the second low-rank matrix corresponding to each cross-attention layer; Based on the weight increment corresponding to each cross-attention layer and the original weight matrix corresponding to each cross-attention layer, fuse the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information to obtain a fused vector.

[0010] Optionally, before inputting the workpiece image, the text prompt information, and the mask image into a pre-trained defect image generation model to obtain a target defect image, the method further includes: Obtain a plurality of second training samples, where each of the second training samples includes a second workpiece sample image, a second text annotation, and a mask annotation; Use the plurality of second training samples to iteratively train the model to be trained in sequence, and calculate the loss value between the inference image generated after each iteration and the input second workpiece sample image, where the model to be trained includes a stable diffusion sub-model to be trained and a shape constraint sub-module to be trained. The stable diffusion sub-model to be trained is used to extract second generated features from the second workpiece sample image and the second text annotation; the shape constraint sub-module to be trained is used to extract second shape features from the mask annotation, perform feature fusion on the second generated features and the second shape features, and perform fuzzy mixing on the fused features and the mask annotation; the stable diffusion sub-model to be trained is further used to generate an inference image based on the features after fuzzy mixing; When a preset condition is satisfied, stop the iterative training to obtain the defect image generation model, where the preset condition is that the loss value gradually decreases and then stabilizes or the number of iterations reaches a preset number.

[0011] Optionally, the calculating the loss value between the inference image generated after each iteration and the input second workpiece sample image includes: Calculate the first sub-loss value, the second sub-loss value, and the third sub-loss value between the generated inference image and the input second workpiece sample image after each iteration respectively. Wherein, the first sub-loss value is used to characterize the absolute difference of pixels between the generated inference image and the input second workpiece sample image after each iteration, the second sub-loss value is used to characterize the semantic similarity between the generated inference image and the input second workpiece sample image after each iteration, and the third sub-loss value is used to characterize the comprehensive similarity in brightness, contrast, and structure between the generated inference image and the input second workpiece sample image after each iteration; Perform weighted summation on the first sub-loss value, the second sub-loss value, and the third sub-loss value to obtain the loss value.

[0012] In a second aspect, an embodiment of the present application further provides a defective picture generation device, and the device includes: A first acquisition module, configured to acquire a workpiece image, text prompt information, and a mask image. Wherein, the mask image is generated based on the workpiece image and mask information, and the mask information is used to characterize the shape and position of the defect to be generated on the workpiece image; An input module, configured to input the workpiece image, the text prompt information, and the mask image into a pre-trained defective image generation model to obtain a target defective image; Wherein, the defective image generation model is configured to extract first generation features from the workpiece image and the text prompt information, extract first shape features from the mask image, and generate the target defective image based on the first generation features and the first shape features. The generated defect in the target defective image satisfies the constraints of the mask information on the defect shape and defect position.

[0013] In a third aspect, an embodiment of the present application further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store a computer program; The processor is configured to implement the defective picture generation method described in the first aspect when executing the program stored on the memory.

[0014] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the defective picture generation method described in the first aspect is implemented.

[0015] The above technical solution provided by the embodiments of the present application has the following advantages compared with the prior art: In the method provided by the embodiments of the present application, by obtaining a workpiece image, text prompt information, and a mask image, where the mask image is generated based on the workpiece image and mask information, and the mask information is used to characterize the shape and position of the defect to be generated on the workpiece image; inputting the workpiece image, the text prompt information, and the mask image into a pre-trained defect image generation model to obtain a target defect image, where the defect image generation model is used to extract first generation features from the workpiece image and the text prompt information, extract first shape features from the mask image, and generate the target defect image based on the first generation features and the first shape features, and the generated defect in the target defect image satisfies the constraints of the mask information on the defect shape and defect position. In the above manner, the defect image generation model can be used to extract first generation features from the workpiece image and text prompt information, extract first shape features from the mask image, and then generate a target defect image based on the first generation features and the first shape features, so that the generated defect in the target defect image can satisfy the constraints of the mask information on the defect shape and defect position, thereby achieving precise control of the defect shape and defect position, making the target defect image meet the actual needs, and can be used as an ideal negative sample. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] One or more embodiments are exemplarily illustrated by the pictures in the corresponding accompanying drawings. These exemplary illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the figures do not constitute a proportional limitation.

[0019] Figure 1 It is a schematic flow chart of a method for generating a defect picture provided by an embodiment of the present application; Figure 2 It is a schematic diagram of a target defect image provided by an embodiment of the present application; Figure 3 It is a schematic structural diagram of a defect image generation model provided by an embodiment of the present application; Figure 4 A schematic diagram of the generation process of a target defect image provided by an embodiment of the present application; Figure 5 A schematic diagram of a low-rank adaptation fine-tuning process provided by an embodiment of the present application; Figure 6 A schematic diagram of the training process of a defect image generation model provided by an embodiment of the present application; Figure 7 A schematic diagram of the structure of a defect picture generation device provided by an embodiment of the present application; Figure 8 A schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0021] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0022] To solve the problem in the related art that the shape and position of defect generation cannot be accurately controlled, resulting in the finally generated defect image not meeting the actual needs and being unable to be used as an ideal negative sample, the present application provides a defect picture generation method, device, electronic device, and storage medium, which can accurately control the shape and position of defect generation.

[0023] See Figure 1 , Figure 1 A schematic flowchart of a defect picture generation method provided by an embodiment of the present application. As Figure 1 shown, the defect picture generation method may include the following steps: Step S101, obtain a workpiece image, text prompt information, and a mask image, where the mask image is generated based on the workpiece image and mask information, and the mask information is used to characterize the shape and position of the defect to be generated on the workpiece image.

[0024] It should be noted that the above-mentioned defective image generation method can be implemented independently by a user terminal (such as a personal computer, a laptop, a smart phone, etc.), or jointly implemented by a user terminal and a server. The embodiments of the present application do not make specific limitations. When implemented independently by a user terminal, the user terminal can obtain the workpiece image, text prompt information, and mask image input by the user, and then use the pre-trained defective image generation model thereof to perform feature extraction and image generation on them to obtain the target defective image. When jointly implemented by a user terminal and a server, the user terminal can obtain the workpiece image, text prompt information, and mask image input by the user, and send them to the server. The server uses the pre-trained defective image generation model thereof to perform feature extraction and image generation on them to obtain the target defective image, and returns it to the user terminal. For example, the user can load the workpiece image to be generated with defects in the user interface of the user terminal, then manually input the mask and text prompt information on the workpiece image, and then send this information to the server. The server uses the defective image generation model to generate the required target defective image to be used as a workpiece negative sample. The target defective image is as Figure 2 shown.

[0025] Specifically, the above-mentioned workpiece image refers to the overall image or partial image of a certain workpiece (such as an automotive wheel hub, a steering knuckle, etc.) to be generated with defects. The above-mentioned text prompt information refers to the prompt text used to prompt the model to generate defects according to this information, such as the prompt text describing the workpiece model, defect type, etc. The above-mentioned mask image refers to the image generated after the user inputs a mask based on the workpiece image. The mask image has the same size as the workpiece image. The pixel points at the mask positions in the mask image are black, and the pixel points at non-mask positions are white.

[0026] Step S102: Input the workpiece image, text prompt information, and mask image into the pre-trained defective image generation model to obtain the target defective image; Among them, the defective image generation model is used to extract the first generation features from the workpiece image and text prompt information, extract the first shape features from the mask image, and generate the target defective image based on the first generation features and the first shape features. The generated defects in the target defective image satisfy the constraints of the mask information on the defect shape and defect position.

[0027] Specifically, the above-mentioned defect image generation model can be a generative model for industrial quality inspection constructed by using the Stable Diffusion sub-model as the base model and combining a shape constraint sub-module to achieve shape-guided control; it can also be a generative model for industrial quality inspection constructed by using the Stable Diffusion sub-model as the base model, implementing high-precision product generation and defect generation in the industrial vision detection scenario through the Low-Rank Adaptation (LORA) fine-tuning technique, and combining the shape constraint sub-module to achieve shape-guided control; of course, it can also be a generative model constructed by other means, which is not specifically limited in the embodiments of the present application. It should be noted that the defect image generation model can extract the first generation feature from the workpiece image and the text prompt information, extract the first shape feature from the mask image, and generate the target defect image based on the first generation feature and the first shape feature. In this way, the generated defects in the target defect image can satisfy the constraints of the mask information on the defect shape and defect position.

[0028] Through the above method, the defect image generation model can extract the first generation feature from the workpiece image and the text prompt information, and extract the first shape feature from the mask image, and then generate the target defect image based on the first generation feature and the first shape feature, so that the generated defects in the target defect image can satisfy the constraints of the mask information on the defect shape and defect position, thereby realizing precise control of the defect shape and defect position, making the target defect image meet the actual needs, and can be used as an ideal negative sample.

[0029] In an alternative embodiment, the above step S102, inputting the workpiece image, the text prompt information, and the mask image into the pre-trained defect image generation model to obtain the target defect image, includes: Inputting the workpiece image into the image encoder to obtain the latent vector corresponding to the workpiece image; Inputting the text prompt information into the text encoder to obtain the semantic embedding vector corresponding to the text prompt information; Inputting the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information into the latent space diffuser to obtain the first generation feature; Inputting the mask image into the shape constraint sub-module to obtain the first shape feature; Performing feature fusion on the first generation feature and the first shape feature, and performing blurred mixing on the fused feature and the mask image; Inputting the mixed feature into the image decoder to generate the target defect image.

[0030] Please refer to Figure 3 , Figure 3The figure is a schematic structural diagram of a defect image generation model provided by an embodiment of the present application. As Figure 3 shown, the defect image generation model may include a shape constraint sub-module and a stable diffusion sub-model. Among them, the core idea of the stable diffusion sub-model is to significantly reduce the computational complexity by performing a diffusion process in a low-dimensional latent space (i.e., Latent Space), and at the same time combine text conditional control to generate content. The stable diffusion sub-model may include a text encoder, a latent space diffuser, an image encoder, and an image decoder. Among them, the image encoder is mainly responsible for compressing the input image (i.e., the workpiece image in the above text) into a low-dimensional latent vector, while the image decoder can restore the denoised latent vector to a high-resolution image. The text encoder can convert natural language text (i.e., the text prompt information in the above text) into a semantic embedding vector to guide image generation. The latent space diffuser fuses the semantic embedding vector output by the text encoder and the latent vector output by the image encoder, and gradually denoises in the latent space to obtain a denoised latent vector. In the text-to-image mode, the input of the stable diffusion sub-model can be only the text prompt information. At this time, the image encoder does not work. The latent space diffuser can start from random Gaussian noise and, under the control of text conditions, gradually generate a clear latent vector by predicting the noise residual through multiple iterations, which is decoded by the image decoder into a generated image. In the image inpainting mode, different from the text-to-image mode, an input image needs to be input into the image encoder to encode the input image into the latent vector of the image, and then the noise residual is predicted through multiple iterations to gradually generate the latent vector of the repaired image, which is decoded by the image decoder into a generated image.

[0031] The stable diffusion sub-model generates data by simulating a "noise addition - denoising" process. During the model training process, Gaussian noise is gradually added to the original training image until it finally becomes pure noise. Simply put, the image is damaged by gradually adding noise, and the loss between the model-predicted noise and the real noise is calculated to train the model to learn image generation. During the model usage process, the image is restored from the noise in a reverse diffusion manner. Among them, the forward diffusion process can be represented by formula (1), which represents the state of the image at each time step t only depends on the state of the previous time step and a noise term . By gradually adding noise, the original data will gradually become a random noise independent of the data distribution. is a learnable function so that the model can adaptively adjust the noise ratio. is a random variable obeying the standard normal distribution, introducing randomness into the diffusion process, so that gradually deviates from the original data, ensuring that different data samples can be generated: ; (1) Among them, represents the state corresponding to time step t, represents the state corresponding to time step t - 1, represents a learnable function, represents a random variable subject to the standard normal distribution. represents obeys the standard normal distribution.

[0032] Among them, the reverse diffusion process can be represented by formula (2), which represents the data generation process of the reverse diffusion process from time step to time step . By predicting the noise and combining the random noise , the model gradually removes the noise and generates samples consistent with the original data distribution.

[0033] ; (2) Among them, represents the state corresponding to time step t, represents the state corresponding to time step t - 1, represents a learnable function, represents the predicted noise at time step t, represents the random noise at time step t, represents the noise standard deviation at time step t, which is used to control the amount of random noise added to the generated samples, represents the cumulative multiplication noise scaling factor from time step 1 to time step t.

[0034] The training process of this stable diffusion sub - model can be divided into two stages. The first stage is to train the image encoder and image decoder to learn the mutual conversion between images and the latent space. The second stage is to train the latent space diffuser, add random noise to the image, generate the noisy latent vector and the time step label t, and train it by minimizing the mean square error between the predicted noise and the actual noise of the latent space diffuser. The loss function for training the latent space diffuser is shown in formula (3), and its core is to let the latent space diffuser learn to predict the standard Gaussian noise added at each step in the forward diffusion process, and minimize the mean square error between the predicted noise and the true noise .

[0035] ; (3) Among them, L represents the loss value, represents the state corresponding to time step t, Represents the state corresponding to time step 0, represents the true noise that follows a standard normal distribution, represents the text prompt information, represents the predicted noise at time step t. E is the expectation operation, which is for the state , time step t, and random noise to take the average of the joint distribution. ||*|| 2 is the square of the L2 norm.

[0036] Specifically, when using a pre-trained defect image generation model to generate a target defect image, the workpiece image can be input into the image encoder of the stable diffusion sub-model to obtain the latent vector corresponding to the workpiece image, and the text prompt information can be input into the text encoder of the stable diffusion sub-model to obtain the semantic embedding vector corresponding to the text prompt information. Then, the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information are input into the latent space diffuser of the stable diffusion sub-model to obtain the first generated feature. At the same time, the mask image can be input into the shape constraint sub-module to obtain the first shape feature, and then the first generated feature and the first shape feature are feature-fused, and the fused feature is blurred and mixed with the mask image. Then, the mixed feature is input into the image decoder in the stable diffusion sub-model to generate the target defect image, as Figure 4 shown.

[0037] In this way, the single stable diffusion sub-model can only support the input of the original image and text. After introducing the shape constraint sub-module, the mask input with shape information constraint is added, enabling the model to have the ability to generate objects guided by shape. Therefore, by introducing the shape constraint sub-module into the stable diffusion sub-model to guide the image restoration process of the stable diffusion sub-model, the function of adding defects to a normal workpiece to generate a defect image can be realized.

[0038] In an optional embodiment, as Figure 4 shown, the above step of inputting the mask image into the shape constraint sub-module to obtain the first shape feature includes: Encoding the mask image using the shape constraint sub-module to obtain the first latent vector corresponding to the mask image, and performing multiple downsampling operations on the mask image to obtain the second latent vector corresponding to the mask image, where the first latent vector corresponding to the mask image, the second latent vector corresponding to the mask image, and the latent vector corresponding to the workpiece image have the same dimension; Mixing the first latent vector corresponding to the mask image and the second latent vector corresponding to the mask image with noise, and performing multi-level feature extraction on the mixed latent vector to obtain the first shape feature.

[0039] Specifically, the above-mentioned shape constraint sub-module is a shape-guided generation framework based on the StableDiffusion sub-model. The core implementation logic is to achieve precise geometric control through a two-branch architecture and multi-level feature fusion. Its two-branch structure mainly includes a generation branch and a shape branch. The generation branch is implemented based on the StableDiffusion sub-model and is responsible for overall image generation and semantic control. The shape branch processes the mask image, maps the mask image to the latent space, and generates latent features aligned with the StableDiffusion sub-model. The mask image is downsampled to the same resolution as the latent space, and Gaussian blur is applied to the edges to reduce generation defects and the texture tearing feeling between the generated image and the real workpiece. By using the image generation and repair function of the StableDiffusion sub-model and the shape-guided embedding of the shape constraint sub-module, the generation of negative samples of workpiece defects with precise control of the defect shape position is realized. When using the shape constraint sub-module to extract the first shape feature, the shape constraint sub-module can be used to encode the mask image to obtain the first latent vector corresponding to the mask image, and perform multiple downsampling operations on the mask image to obtain the second latent vector corresponding to the mask image. Then, the first latent vector corresponding to the mask image and the second latent vector corresponding to the mask image are mixed with noise, and multi-level feature extraction is performed on the mixed latent vector to obtain the first shape feature.

[0040] It should be noted that the first latent vector corresponding to the mask image, the second latent vector corresponding to the mask image, and the latent vector corresponding to the workpiece image have the same dimension. This is beneficial for aligning the first latent vector corresponding to the mask image and the second latent vector corresponding to the mask image with the latent space distribution of the StableDiffusion sub-model, facilitating data processing. After obtaining the first shape feature, the first shape feature can be input into the zero convolution layer. The zero convolution layer can balance the fusion ratio of the first shape feature and the first generation feature through learnable parameters (i.e., the dynamic weights in Figure 4 ). The learnable parameters work in cooperation with the zero convolution layer to effectively fuse the first generation feature and the first shape feature by dynamically adjusting the feature weights. The initialization strategy of the zero convolution layer ensures the stability of the model in the initial stage of training, while the learnable parameters are automatically optimized through the training process to improve the final performance of the model. After the first shape feature and the first generation feature are fused, they are then blurred and mixed with the original mask image to ensure a natural transition between the synthesized area and the background. This architecture works through a two-path cooperation, ensuring both the semantic accuracy of the generated content and the rationality of the object shape.

[0041] In this embodiment, the shape constraint sub-module can capture the defect shape information and position information in the mask image, providing shape and position constraints for synthesizing the defect image.

[0042] In an alternative embodiment, the above step of inputting the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information into the latent space diffuser to obtain the first generated feature includes: The latent space diffuser is used to fuse the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information, and the fused vector is gradually denoised to obtain the first generated feature.

[0043] Specifically, when using the latent space diffuser to extract the first generated feature, the latent space diffuser can be used to fuse the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information, and the fused vector is gradually denoised to obtain the first generated feature. The process of generating an image by the latent space diffuser through the reverse diffusion process has been introduced in the foregoing embodiments and will not be elaborated here.

[0044] In this embodiment, the latent space diffuser in the stable diffusion sub-model can capture the feature information in the workpiece image and the text prompt information, providing a basis for synthesizing the defect image.

[0045] In an alternative embodiment, the latent space diffuser includes a plurality of cross-attention layers, and some or all of the cross-attention layers in the plurality of cross-attention layers include a frozen original weight matrix and a first low-rank matrix and a second low-rank matrix introduced through low-rank adaptation. The first low-rank matrix and the second low-rank matrix corresponding to each cross-attention layer are pre-trained based on a plurality of first training samples, and each first training sample includes a first workpiece sample image, a first defect sample image, and a first text annotation; The above step of using the latent space diffuser to fuse the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information includes: Based on the first low-rank matrix and the second low-rank matrix corresponding to each cross-attention layer, the weight increment corresponding to each cross-attention layer is determined, where the weight increment corresponding to each cross-attention layer is the calculation result of the product of the first low-rank matrix and the second low-rank matrix corresponding to each cross-attention layer; Based on the weight increment corresponding to each cross-attention layer and the original weight matrix corresponding to each cross-attention layer, the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information are fused to obtain a fused vector.

[0046] In this embodiment, since the original Stable Diffusion sub-model cannot learn the features of the workpiece and the defect well, the first low-rank matrix A and the second low-rank matrix B can be introduced into some or all of the cross-attention layers of the latent space diffuser through the low-rank adaptation fine-tuning technique, and the output result of the cross-attention layer can be fine-tuned through the first low-rank matrix A and the second low-rank matrix B, so as to ensure that the Stable Diffusion sub-model can learn the features of the workpiece and the defect. The low-rank adaptation fine-tuning process is as follows Figure 5 shown. It mainly adds a bypass composed of the first low-rank matrix A and the second low-rank matrix B beside each cross-attention layer of the latent space diffuser. The original weight matrix w0 of the cross-attention layer in the Stable Diffusion sub-model is frozen and remains unchanged. Only the first low-rank matrix A and the second low-rank matrix B are updated during training. When generating an image, the trained first low-rank matrix A and the second low-rank matrix B can be multiplied matrix-wise, and the calculation result is used as the weight increment △w. In this way, the final weight of the model w = w0 + △w, that is, the model output is the superposition of the original output of the cross-attention layer and the output of the LORA bypass. The reason for designing the cross-attention layer of the latent space diffuser in this way is that this training method can greatly reduce the number of trainable parameters. Only the first low-rank matrix A and the second low-rank matrix B need to be trained, rather than the parameters of the entire cross-attention layer. This not only reduces the demand for computing resources, but also enables the model to be more flexible in adapting to multiple tasks, because the LORA bypass can be independently trained for different tasks while sharing the cross-attention layer in the latent space diffuser. In practical applications, workpiece and defect image data of the workpiece can be collected for text annotation to form the first training sample, and the Stable Diffusion sub-model can be fine-tuned using LORA training so that the Stable Diffusion sub-model can generate workpieces and generate defects.

[0047] In an alternative embodiment, before the step S102 of inputting the workpiece image, the text prompt information, and the mask image into the pre-trained defect image generation model to obtain the target defect image, the method further includes: Obtaining a plurality of second training samples, where each second training sample includes a second workpiece sample image, a second text annotation, and a mask annotation; Iteratively train the model to be trained using multiple second training samples, and calculate the loss value between the generated inference image and the input second workpiece sample image after each iteration. Among them, the model to be trained includes a stable diffusion sub-model to be trained and a shape constraint sub-module to be trained. The stable diffusion sub-model to be trained is used to extract second-generation features from the second workpiece sample image and the second text annotation; the shape constraint sub-module to be trained is used to extract second shape features from the mask annotation, perform feature fusion on the second-generation features and the second shape features, and perform fuzzy mixing on the fused features and the mask annotation; the stable diffusion sub-model to be trained is also used to generate an inference image based on the features after fuzzy mixing. When the preset conditions are met, stop the iterative training to obtain a defect image generation model. Among them, the preset conditions are that the loss value gradually decreases and then stabilizes, or the number of iterations reaches the preset number of times.

[0048] Specifically, before using the defect image generation model to generate the target defect image, it is also necessary to train the defect image generation model. When training the defect image generation model, it is necessary to obtain multiple second training samples. Among them, each second training sample can include a second workpiece sample image, a second text annotation, and a mask annotation. Then, the second training samples are sequentially input into the model to be trained for iterative training. Among them, the second workpiece sample image and the second text annotation can be input into the stable diffusion sub-model to be trained to extract second-generation features, and the mask annotation can be input into the shape constraint sub-module to be trained to extract second shape features. Then, perform feature fusion on the second-generation features and the second shape features, and perform fuzzy mixing on the fused features and the mask annotation. Then, use the stable diffusion sub-model to generate an inference image based on the features after fuzzy mixing, and calculate the loss value between the inference image and the input second workpiece sample image. Repeat this process to iteratively train the model to be trained until the loss value gradually decreases and then stabilizes, or the number of iterations reaches the preset number of times, and then stop the iteration to obtain the defect image generation model. It should be noted that after fusing the second-generation features and the second shape features and before performing fuzzy mixing on the fused features and the mask annotation, image decoding can be used to improve the resolution of the generated image. Here, through the encoding and decoding process in the latent space, the details and texture information in the image are restored, thereby indirectly improving the resolution of the generated image, which is convenient for generating a clearer defect image in the image decoder later.

[0049] In a practical application, the training process of the defect image generation model is as follows Figure 6As shown, first, various workpiece images to be detected in the project are collected (negative samples must be included). Mask annotation and text annotation are performed on the workpieces and the defects on the workpieces to serve as the second training samples. During model training, the workpiece sample images and text annotations are input into the generation branch. The shape branch performs noise mixing processing on the mask annotation. The output results of the two branches are fused with dynamic weights to guide the generation of inferences. The obtained inference images are constrained by the original workpiece sample images and mask shapes, the multi-modal loss is calculated, and the unfrozen neural network parameters are optimized according to the loss gradient backpropagation, so as to obtain a defect image generation model, which is convenient for subsequently using the defect image generation model to generate target defect images.

[0050] In an optional embodiment, the above step of calculating the loss value between the inference image generated after each iteration and the input second workpiece sample image includes: Calculate the first sub-loss value, the second sub-loss value, and the third sub-loss value between the inference image generated after each iteration and the input second workpiece sample image respectively. Among them, the first sub-loss value is used to characterize the absolute difference of pixels between the inference image generated after each iteration and the input second workpiece sample image. The second sub-loss value is used to characterize the similarity at the semantic level between the inference image generated after each iteration and the input second workpiece sample image. The third sub-loss value is used to characterize the comprehensive similarity in brightness, contrast, and structure between the inference image generated after each iteration and the input second workpiece sample image. Perform weighted summation on the first sub-loss value, the second sub-loss value, and the third sub-loss value to obtain the loss value.

[0051] Specifically, when performing weighted summation on the first sub-loss value, the second sub-loss value, and the third sub-loss value, the weight of each sub-loss value can be set artificially according to actual needs, and the embodiments of the present application do not make specific limitations.

[0052] The above first sub-loss value is obtained by calculating the absolute difference between the second workpiece sample image vector and each pixel in the inference image, which directly constrains the reconstruction accuracy of the model for the entire image. The above first sub-loss value can be expressed by the following formula: ; (4) Wherein, represents the first sub-loss value, represents the second workpiece sample image vector, represents the inference image vector.

[0053] The second sub-loss value is used to characterize the similarity at the semantic level between the inference image generated after each iteration and the input second workpiece sample image, which mainly constrains the high-level semantic information such as texture and edges of the model. The above second sub-loss value can be expressed by the following formula: ; (5) wherein, represents the second sub-loss value, represents the th layer of the training network, and respectively represent the height and width of the feature map of the th layer, and respectively represent the spatial positions of the feature map, represents the weight vector of the th layer, represents the dot product of two vectors. represents the feature representation of the second workpiece sample image at the th layer,

[0054] The above third sub-loss value reduces the loss of high-frequency details (such as edge sharpness) through multi-scale structural similarity constraints. Its core idea is to strengthen the detail retention and alignment accuracy of the boundary region through similarity evaluation in three dimensions: brightness, contrast, and structure. The above third sub-loss value can be expressed by the following formula: ; (6) wherein, represents the third sub-loss value, represents the second workpiece sample image vector, represents the inference image vector, is the brightness comparison term, is the contrast comparison term, is the structure comparison term. , and are weight parameters used to adjust the importance of each term.

[0055] Among them, the brightness comparison term is shown in formula (7): ; (7) wherein, and are the means of the image vector x and the image vector y respectively, and C is a constant, usually taken as 1 to avoid division by zero.

[0056] The contrast comparison term is shown in formula 8: ; (8) wherein, and are the standard deviations of the image vector x and the vector y respectively. C is a constant, usually taken as 1 to avoid division by zero.

[0057] Structure comparison term is shown in formula (9): ; (9) where and are the standard deviations of the image vector x and the vector y respectively. C is a constant, usually taken as 1 to avoid division by zero. is the covariance of the image vector x and the vector y respectively.

[0058] Through the above method, the loss value between the inference image and the input sample image can be calculated from three dimensions, so that the trained model can ensure the image reconstruction accuracy, text semantics and details of the boundary region.

[0059] See Figure 7 , Figure 7 which is a schematic structural diagram of a defective picture generation device provided by an embodiment of the present application. As Figure 7 shown, the defective picture generation device 700 includes: A first acquisition module 701, configured to acquire a workpiece image, text prompt information, and a mask image, where the mask image is generated based on the workpiece image and mask information, and the mask information is used to characterize the shape and position of the defect to be generated on the workpiece image; An input module 702, configured to input the workpiece image, text prompt information, and mask image into a pre-trained defective picture generation model to obtain a target defective picture; wherein, the defective picture generation model is configured to extract a first generation feature from the workpiece image and text prompt information, extract a first shape feature from the mask image, and generate a target defective picture based on the first generation feature and the first shape feature, and the generated defect in the target defective picture satisfies the constraints of the mask information on the defect shape and defect position.

[0060] Further, the defective picture generation model includes a shape constraint sub-module and a stable diffusion sub-model, and the stable diffusion sub-model includes a text encoder, a latent space diffuser, an image encoder, and an image decoder; the input module 702 includes: A first input sub-module, configured to input the workpiece image into the image encoder to obtain a latent vector corresponding to the workpiece image; A second input sub-module, configured to input text prompt information into a text encoder to obtain a semantic embedding vector corresponding to the text prompt information; A third input sub-module, configured to input a latent vector corresponding to a workpiece image and a semantic embedding vector corresponding to text prompt information into a latent space diffuser to obtain a first generated feature; A fourth input sub-module, configured to input a mask image into a shape constraint sub-module to obtain a first shape feature; A fusion and mixing sub-module, configured to perform feature fusion on the first generated feature and the first shape feature, and perform blurred mixing on the fused feature and the mask image; A fifth input sub-module, configured to input the mixed feature into an image decoder to generate a target defect image.

[0061] Further, the fourth input sub-module includes: An encoding and downsampling unit, configured to use the shape constraint sub-module to encode the mask image to obtain a first latent vector corresponding to the mask image, and perform multiple downsampling operations on the mask image to obtain a second latent vector corresponding to the mask image, wherein the first latent vector corresponding to the mask image, the second latent vector corresponding to the mask image, and the latent vector corresponding to the workpiece image have the same dimension; A mixing unit, configured to perform noise mixing on the first latent vector corresponding to the mask image and the second latent vector corresponding to the mask image, and perform multi-level feature extraction on the mixed latent vector to obtain a first shape feature.

[0062] Further, the third input sub-module includes: A fusion and denoising unit, configured to use the latent space diffuser to fuse the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information, and gradually denoise the fused vector to obtain a first generated feature.

[0063] Further, the latent space diffuser includes a plurality of cross-attention layers, and some or all of the cross-attention layers in the plurality of cross-attention layers include a frozen original weight matrix and a first low-rank matrix and a second low-rank matrix introduced through low-rank adaptation. The first low-rank matrix and the second low-rank matrix corresponding to each cross-attention layer are pre-trained based on a plurality of first training samples. Each first training sample includes a first workpiece sample image, a first defect sample image, and a first text annotation. The fusion and denoising unit is specifically configured to: Determine a weight increment corresponding to each cross-attention layer based on the first low-rank matrix and the second low-rank matrix corresponding to each cross-attention layer, wherein the weight increment corresponding to each cross-attention layer is the calculation result of the product of the first low-rank matrix and the second low-rank matrix corresponding to each cross-attention layer; Based on the weight increment corresponding to each cross-attention layer and the original weight matrix corresponding to each cross-attention layer, fuse the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information to obtain a fused vector.

[0064] Furthermore, the defect image generation device 700 further includes: A second acquisition module, configured to acquire a plurality of second training samples, where each second training sample includes a second workpiece sample image, a second text annotation, and a mask annotation; A calculation module, configured to iteratively train the model to be trained in turn using a plurality of second training samples, and calculate the loss value between the generated inference image and the input second workpiece sample image after each iteration, where the model to be trained includes a stable diffusion sub-model to be trained and a shape constraint sub-module to be trained. The stable diffusion sub-model to be trained is used to extract second generation features from the second workpiece sample image and the second text annotation; the shape constraint sub-module to be trained is used to extract second shape features from the mask annotation, perform feature fusion on the second generation features and the second shape features, and perform fuzzy mixing on the fused features and the mask annotation; the stable diffusion sub-model to be trained is further used to generate an inference image based on the features after fuzzy mixing; A stop module, configured to stop the iterative training when a preset condition is met to obtain a defect image generation model, where the preset condition is that the loss value gradually decreases and then stabilizes or the number of iterations reaches a preset number.

[0065] Furthermore, the calculation module includes: A calculation sub-module, configured to calculate a first sub-loss value, a second sub-loss value, and a third sub-loss value between the generated inference image and the input second workpiece sample image after each iteration, where the first sub-loss value is used to represent the absolute difference of pixels between the generated inference image and the input second workpiece sample image after each iteration, the second sub-loss value is used to represent the similarity at the semantic level between the generated inference image and the input second workpiece sample image after each iteration, and the third sub-loss value is used to represent the comprehensive similarity in brightness, contrast, and structure between the generated inference image and the input second workpiece sample image after each iteration; A weighted summation sub-module, configured to perform weighted summation on the first sub-loss value, the second sub-loss value, and the third sub-loss value to obtain the loss value.

[0066] It should be noted that the defect image generation device 700 provided in the embodiment of the present application can implement the steps of the defect image generation method in the above Figure 1 shown method embodiment and achieve the same technical effect, which will not be elaborated here.

[0067] Such as Figure 8As shown in the figure, an embodiment of the present application further provides an electronic device, including a processor 811, a communication interface 812, a memory 813, and a communication bus 814. Among them, the processor 811, the communication interface 812, and the memory 813 complete mutual communication through the communication bus 814. The memory 813 is used to store computer programs. In an embodiment of the present application, when the processor 811 is used to execute the program stored on the memory 813, it implements the defect picture generation method provided in any of the foregoing method embodiments.

[0068] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the defect picture generation method provided in any of the foregoing method embodiments.

[0069] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0070] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0071] It should be understood that the terms used in this document are only for the purpose of describing specific example embodiments and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" used in this document may also include the plural form. The terms "include", "comprise", "contain", and "have" are inclusive and thus specify the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described in this document are not to be construed as necessarily requiring them to be executed in the specific order described or illustrated, unless the execution order is clearly indicated. It should also be understood that additional or alternative steps can be used.

[0072] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for generating defective pictures, characterized in that, The method includes: Obtaining a workpiece image, text prompt information, and a mask image, where the mask image is generated based on the workpiece image and mask information, and the mask information is used to characterize the shape and position of the defect to be generated on the workpiece image; Inputting the workpiece image, the text prompt information, and the mask image into a pre-trained defect image generation model to obtain a target defect image; Wherein, the defect image generation model is used to extract a first generation feature from the workpiece image and the text prompt information, extract a first shape feature from the mask image, and generate the target defect image based on the first generation feature and the first shape feature, and the generated defect in the target defect image satisfies the constraints of the mask information on the defect shape and defect position.

2. The method for generating defective pictures according to claim 1, wherein The defect image generation model includes a shape constraint sub-module and a stable diffusion sub-model, and the stable diffusion sub-model includes a text encoder, a latent space diffuser, an image encoder, and an image decoder; The step of inputting the workpiece image, the text prompt information, and the mask image into a pre-trained defect image generation model to obtain a target defect image includes: Inputting the workpiece image into the image encoder to obtain a latent vector corresponding to the workpiece image; Inputting the text prompt information into the text encoder to obtain a semantic embedding vector corresponding to the text prompt information; Inputting the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information into the latent space diffuser to obtain the first generation feature; Inputting the mask image into the shape constraint sub-module to obtain the first shape feature; Performing feature fusion on the first generation feature and the first shape feature, and performing a blur mix on the fused feature and the mask image; Inputting the mixed feature into the image decoder to generate the target defect image.

3. The method for generating defective pictures according to claim 2, wherein, The step of inputting the mask image into the shape constraint sub-module to obtain the first shape feature includes: Encoding the mask image by using the shape constraint sub-module to obtain a first latent vector corresponding to the mask image, and performing multiple downsampling operations on the mask image to obtain a second latent vector corresponding to the mask image, where the first latent vector corresponding to the mask image, the second latent vector corresponding to the mask image, and the latent vector corresponding to the workpiece image have the same dimension; Performing noise mixing on the first latent vector corresponding to the mask image and the second latent vector corresponding to the mask image, and performing multi-level feature extraction on the mixed latent vector to obtain the first shape feature.

4. The method for generating a defect picture according to claim 2, wherein The step of inputting the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information into the latent space diffuser to obtain the first generation feature includes: The latent space diffuser is used to fuse the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information, and gradually denoise the fused vector to obtain the first generated feature.

5. The defect picture generation method according to claim 4, wherein The latent space diffuser includes a plurality of cross-attention layers, and some or all of the cross-attention layers in the plurality of cross-attention layers include a frozen original weight matrix and a first low-rank matrix and a second low-rank matrix introduced by low-rank adaptation. The first low-rank matrix and the second low-rank matrix corresponding to each cross-attention layer are pre-trained based on a plurality of first training samples, and each of the first training samples includes a first workpiece sample image, a first defect sample image, and a first text annotation. The using the latent space diffuser to fuse the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information includes: Based on the first low-rank matrix and the second low-rank matrix corresponding to each cross-attention layer, determine the weight increment corresponding to each cross-attention layer, where the weight increment corresponding to each cross-attention layer is the calculation result of the product of the first low-rank matrix and the second low-rank matrix corresponding to each cross-attention layer. Based on the weight increment corresponding to each cross-attention layer and the original weight matrix corresponding to each cross-attention layer, fuse the latent vector corresponding to the workpiece image and the semantic embedding vector corresponding to the text prompt information to obtain a fused vector.

6. The method for generating a defective picture according to claim 1, wherein Before inputting the workpiece image, the text prompt information, and the mask image into a pre-trained defect image generation model to obtain a target defect image, the method further includes: Obtain a plurality of second training samples, where each of the second training samples includes a second workpiece sample image, a second text annotation, and a mask annotation. Use the plurality of second training samples to iteratively train the model to be trained in sequence, and calculate the loss value between the generated inference image and the input second workpiece sample image after each iteration. The model to be trained includes a stable diffusion sub-model to be trained and a shape constraint sub-module to be trained. The stable diffusion sub-model to be trained is used to extract a second generated feature from the second workpiece sample image and the second text annotation. The shape constraint sub-module to be trained is used to extract a second shape feature from the mask annotation, perform feature fusion on the second generated feature and the second shape feature, and perform fuzzy mixing on the fused feature and the mask annotation. The stable diffusion sub-model to be trained is further used to generate an inference image based on the feature after fuzzy mixing. When a preset condition is satisfied, stop the iterative training to obtain the defect image generation model, where the preset condition is that the loss value gradually decreases and then stabilizes or the number of iterations reaches a preset number.

7. The method for generating a defect picture according to claim 6, characterized in that, The calculating the loss value between the generated inference image and the input second workpiece sample image after each iteration includes: Calculate the first sub-loss value, the second sub-loss value, and the third sub-loss value between the generated inference image and the input second workpiece sample image after each iteration respectively, where the first sub-loss value is used to represent the absolute difference of pixels between the generated inference image and the input second workpiece sample image after each iteration, the second sub-loss value is used to represent the semantic similarity between the generated inference image and the input second workpiece sample image after each iteration, and the third sub-loss value is used to represent the comprehensive similarity in brightness, contrast, and structure between the generated inference image and the input second workpiece sample image after each iteration; Perform weighted summation on the first sub-loss value, the second sub-loss value, and the third sub-loss value to obtain the loss value.

8. A defective picture generation device, characterized in that, The device includes: A first acquisition module, configured to acquire a workpiece image, text prompt information, and a mask image, where the mask image is generated based on the workpiece image and mask information, and the mask information is used to represent the shape and position of the defect to be generated on the workpiece image; An input module, configured to input the workpiece image, the text prompt information, and the mask image into a pre-trained defect image generation model to obtain a target defect image; Wherein, the defect image generation model is configured to extract a first generation feature from the workpiece image and the text prompt information, extract a first shape feature from the mask image, and generate the target defect image based on the first generation feature and the first shape feature, and the generated defect in the target defect image satisfies the constraints of the mask information on the defect shape and defect position.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store a computer program; The processor is configured to implement the defect picture generation method according to any one of claims 1-7 when executing the program stored on the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by the processor, implements the defect picture generation method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Image generation method and device, electronic equipment and storage medium

    CN117292020A

  • Defect image generation method and device, computer equipment and storage medium

    CN117953321A

  • Defect image generation method and device

    CN118298045A

  • Defect data generation method and device, electronic equipment and storage medium

    CN119516299A

  • Defect image generation method and device, equipment and storage medium

    CN119625474A

Cited By

  • Steel structure welding quality detection method and storage medium

    CN120219374A

  • Steel structure welding quality inspection method and storage medium

    CN120219374B

  • Defect image generation method and device, storage medium and electronic device

    CN120495455A

  • Image generation method and device, equipment, storage medium and product

    CN120997628A

  • Chip defect sample generation method and device

    CN121027136A