Abnormal image generation method and device, electronic equipment and storage medium

By acquiring abnormal images and masks, using the contrasting language image pre-training model to extract features and generate optimal text prompts, and combining graphics processing and data augmentation parameters to generate predicted image sets, the problem of poor accuracy in abnormal image generation in the prior art is solved, and more efficient abnormal image generation is achieved.

CN120259485APending Publication Date: 2025-07-04CHINA MOBILE ZIJIN INNOVATION INST CO LTD +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411747556.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the accuracy of abnormal image generation is poor, especially when cropping or pasting existing abnormal images in a model-free method, it is easy to add unrelated textures or existing exceptions. The method based on generating an adversarial network requires a large amount of abnormal data training.

Method used

By acquiring abnormal images and masks, using the contrasting language image pretrained model to extract features and generate optimal text prompts, a second sample pair is generated by combining graphics processing and data augmentation parameters, a preset model is input to generate a set of predicted images, and the model is optimized using prior knowledge and embedding similarity loss function.

Benefits of technology

It improves the accuracy and authenticity of abnormal image generation, reduces dependence on a large amount of abnormal data, and shortens the model training time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259485A_ABST
    Figure CN120259485A_ABST
Patent Text Reader

Abstract

The invention provides an abnormal image generation method, and relates to the technical field of image generation, and the method comprises the steps: obtaining an abnormal image and a mask, and enabling the mask to be obtained through the marking of the abnormal image; extracting features matched with the abnormal image based on a comparison language image pre-training model, and generating an optimal text prompt; based on the abnormal image and the mask, a preset number of first sample pairs and data enhancement parameters are obtained through graphic processing, a second sample pair is generated according to the first sample pairs, the data enhancement parameters and the optimal text prompt, the first sample pair is a combination of the abnormal image and the mask, and the second sample pair is a combination of the abnormal image and the mask. The second sample pair is a combination of the abnormal image, the mask and the optimal text prompt; the second sample pair serves as input, a prediction image set matched with the abnormal image is obtained based on a preset model, and the prediction image set comprises at least one predicted abnormal image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image generation, and in particular, to a method, apparatus, electronic device, and storage medium for generating abnormal images. Background Art

[0002] With the development of science and technology, industrial production lines have entered an automated flow production mode. It has become increasingly difficult to manually screen a large number of workpieces produced. Therefore, supervised anomaly detection algorithms have been implemented and applied.

[0003] Supervised anomaly detection algorithms require a large amount of training data. In the prior art, generally two major categories of methods, model-free and model-based, are used to generate abnormal images. Among them, the model-free method generates new images by cropping and pasting existing abnormal images onto normal images. Some of these methods add irrelevant textures or existing anomalies to normal samples. Some other methods based on Generative Adversarial Networks (GANs) can generate anomalies on normal images. Among them, training them requires a large amount of abnormal data.

[0004] Therefore, there is a problem of poor generation accuracy in the prior art for abnormal image generation technology. Summary of the Invention

[0005] Embodiments of the present invention provide a method, apparatus, electronic device, and storage medium for generating abnormal images to solve the problem of poor generation accuracy in the prior art for abnormal image generation technology.

[0006] In a first aspect, an embodiment of the present invention provides a method for generating an abnormal image, the method including:

[0007] Obtain an abnormal image and a mask, where the mask is obtained by annotating the abnormal image;

[0008] Extract features matching the abnormal image based on a contrastive language-image pre-training model and generate an optimal text prompt, where the optimal text prompt is a text description in a preset text description library;

[0009] Based on the abnormal image and the mask, obtain a preset number of first sample pairs and data augmentation parameters through graphics processing, and generate second sample pairs according to the first sample pairs, the data augmentation parameters, and the optimal text prompt, where the first sample pair is a combination of the abnormal image and the mask, and the second sample pair is a combination of the abnormal image, the mask, and the optimal text prompt;

[0010] Using the second sample pair as input, a set of predicted images matching the abnormal image is obtained based on a preset model, and the set of predicted images includes at least one predicted abnormal image.

[0011] In a second aspect, an embodiment of the present invention provides a device for generating abnormal images, the device includes:

[0012] An acquisition module, configured to acquire an abnormal image and a mask, where the mask is obtained by annotating the abnormal image;

[0013] A first processing module, configured to extract features matching the abnormal image based on a contrastive language-image pre-trained model and generate an optimal text prompt, where the optimal text prompt is a text description in a preset text description library;

[0014] A second processing module, configured to obtain a preset number of first sample pairs and data augmentation parameters through graphics processing based on the abnormal image and the mask, and generate second sample pairs according to the first sample pairs, the data augmentation parameters, and the optimal text prompt, where the first sample pair is a combination of the abnormal image and the mask, and the second sample pair is a combination of the abnormal image, the mask, and the optimal text prompt;

[0015] A third processing module, configured to use the second sample pair as input to obtain a set of predicted images matching the abnormal image based on a preset model, where the set of predicted images includes at least one predicted abnormal image.

[0016] In a third aspect, an embodiment of the present invention provides an electronic device, including a transceiver and a processor, the transceiver is configured to acquire an abnormal image and a mask, where the mask is obtained by annotating the abnormal image;

[0017] The processor is configured to extract features matching the abnormal image based on a contrastive language-image pre-trained model and generate an optimal text prompt, where the optimal text prompt is a text description in a preset text description library;

[0018] The processor is further configured to obtain a preset number of first sample pairs and data augmentation parameters through graphics processing based on the abnormal image and the mask, and generate second sample pairs according to the first sample pairs, the data augmentation parameters, and the optimal text prompt, where the first sample pair is a combination of the abnormal image and the mask, and the second sample pair is a combination of the abnormal image, the mask, and the optimal text prompt;

[0019] The processor is further configured to use the second sample pair as input to obtain a set of predicted images matching the abnormal image based on a preset model, where the set of predicted images includes at least one predicted abnormal image.

[0020] Fourthly, an embodiment of the present invention provides an electronic device, including: a processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, it implements the steps of the method for generating an abnormal image as described in the first aspect.

[0021] Fifthly, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the method for generating an abnormal image as described in the first aspect.

[0022] In an embodiment of the present invention, first, an abnormal image and a mask are obtained. Then, based on a contrastive language-image pre-training model, features matching the abnormal image are extracted, and an optimal text prompt is generated. The optimal text prompt is a text description in a preset text description library. Then, based on the abnormal image and the mask, a preset number of first sample pairs and data augmentation parameters are obtained through graphics processing, and second sample pairs are generated according to the first sample pairs, the data augmentation parameters, and the optimal text prompt. Finally, the second sample pairs are used as inputs, and based on a preset model, a set of predicted images matching the abnormal image is obtained. The set of predicted images includes at least one predicted abnormal image. Through this method, many text prompts are constructed by expanding the keywords of abnormal semantics. Then, these prompts are compared with the images for similarity in the semantic embedding space to ensure the generation of the best text prompt applicable to the real abnormal image, making full use of prior knowledge, accelerating the convergence of the model, making the generated image more realistic, and thus improving the accuracy in the generation of abnormal image generation technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 is a flowchart of the method for generating an abnormal image provided by an embodiment of the present invention;

[0025] Figure 2 is a flowchart of another method for generating an abnormal image provided by an embodiment of the present invention;

[0026] Figure 3 is a parameter diagram provided by an embodiment of the present invention;

[0027] Figure 4It is a schematic flowchart of the steps for fine-tuning a preset model provided by an embodiment of the present invention;

[0028] Figure 5 It is a schematic diagram of some abnormal image samples provided by an embodiment of the present invention;

[0029] Figure 6 It is a schematic structural diagram of a device for generating abnormal images provided by an embodiment of the present invention;

[0030] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0032] Terms such as "first" and "second" in the embodiments of the present invention are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0033] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of a method for generating abnormal images provided by an embodiment of the present invention. As Figure 1 shown, the method includes the following steps:

[0034] Step 101, obtain an abnormal image and a mask, where the mask is obtained by annotating the abnormal image;

[0035] Step 102, extract features matching the abnormal image based on a contrastive language-image pre-training model, and generate an optimal text prompt, where the optimal text prompt is a text description in a preset text description library;

[0036] Step 103: Based on the abnormal image and the mask, obtain a preset number of first sample pairs and data augmentation parameters through graphic processing, and generate second sample pairs according to the first sample pairs, the data augmentation parameters, and the optimal text prompt. The first sample pair is a combination of the abnormal image and the mask, and the second sample pair is a combination of the abnormal image, the mask, and the optimal text prompt.

[0037] Step 104: Use the second sample pair as input and obtain a set of predicted images that match the abnormal image based on a preset model. The set of predicted images includes at least one predicted abnormal image.

[0038] It should be noted that steps 101, 102, 103, and 104 included in the simultaneous interpretation method provided by the embodiments of the present invention can be executed by an electronic device or other devices, such as a computer, etc. The embodiments of the present invention do not make any limitations in this regard.

[0039] In step 101, the above abnormal images can be a small number of abnormal images of multiple categories and with representativeness, and the abnormal regions in the above abnormal images are obtained through annotation, so as to obtain the mask of the abnormal region.

[0040] Among them, the above annotation method can be carried out by manual annotation to ensure the accuracy of obtaining the above mask.

[0041] In step 102, the above contrastive language-image pre-training model can select the CLIP model. Then, the image encoder in the pre-trained CLIP model can be used to extract the features of the above abnormal image, so as to generate the optimal text prompt.

[0042] Specifically, after using the image encoder in the pre-trained CLIP model to extract the features of the above abnormal image, the encoder in the pre-trained CLIP model is used again to extract the features of each text description in the above preset text description library. By calculating the similarity between the image features and the text features one by one, the text description that is most aligned with the above abnormal image in the embedding space is selected as the above optimal text prompt.

[0043] It should be noted that the similarity calculation between the image features and the text features can be carried out by calculating the cosine similarity.

[0044] In step 103, the above-mentioned first sample pair can be understood as the combination of the above-mentioned abnormal image and the above-mentioned mask, and the above-mentioned second sample pair can be understood as the training data set to be input into the model. Then, the features of each piece of text in the text description library are extracted by pre-training the CLIP text encoder, the features of the abnormal image are extracted by pre-training the CLIP image encoder, and after selecting the optimal text prompt with the highest similarity to the abnormal image through similarity comparison, simple data augmentation operations are performed on the abnormal image and the mask, and the data augmentation parameters are saved; the data augmentation parameters are added to the optimal text prompt to form some abnormal image-mask-optimal text prompt sample pairs; the same operations are performed on other abnormal image-mask sample pairs to obtain the training data set, that is, the above-mentioned second sample pair.

[0045] In step 104, the second sample pair is used as the input, and a set of predicted images matching the abnormal image is obtained based on a preset model. Among them, after obtaining the above-mentioned set of predicted images, manual review can be performed to judge the quality of the generated images. If there are cases with poor quality, the generation of the set of predicted images can be redone.

[0046] In the embodiment of the present invention, first, an abnormal image and a mask are obtained, then the features matching the abnormal image are extracted based on a contrastive language-image pre-training model, and an optimal text prompt is generated. The optimal text prompt is a text description in a preset text description library. Then, based on the abnormal image and the mask, a preset number of first sample pairs and data augmentation parameters are obtained through graphics processing, and a second sample pair is generated according to the first sample pair, the data augmentation parameters, and the optimal text prompt. Finally, the second sample pair is used as the input, and a set of predicted images matching the abnormal image is obtained based on a preset model. The set of predicted images includes at least one predicted abnormal image. Through this method, many text prompts are constructed by expanding the keywords of abnormal semantics. Then, these prompts are compared with the images for similarity in the semantic embedding space to ensure the generation of the best text prompts applicable to real abnormal images, making full use of prior knowledge, accelerating the convergence of the model, making the generated images more realistic, and thus improving the accuracy in the generation of abnormal image generation technology.

[0047] Optionally, the extracting the features matching the abnormal image based on a contrastive language-image pre-training model and generating an optimal text prompt includes:

[0048] Extracting the features matching the abnormal image based on the contrastive language-image pre-training model;

[0049] Obtaining the text descriptions in the preset text description library, and extracting the features matching the text descriptions through the contrastive language-image pre-training model;

[0050] Calculate the cosine similarity between the features matching the abnormal image and the features matching the text description, and generate the optimal text prompt based on the cosine similarity. The optimal text prompt is the text description with the highest matching degree to the abnormal image.

[0051] In this embodiment, first, obtain the features matching the above abnormal image and the features matching the above text description. Then, calculate the cosine similarity between the features matching the abnormal image and the features matching the text description. Finally, generate the above optimal text prompt based on the above cosine similarity. Through this method, the text prompt matching the above abnormal image can be accurately obtained.

[0052] In the embodiment of the present invention, the above optimal text prompt can be obtained by calculating the similarity in the embedding space, such as through the following expression:

[0053]

[0054] where Clip img (.) and Clip txt (.) respectively represent the pre-trained CLIP image and text encoders. I is the input abnormal image, cos(.,.) is the cosine similarity between two embedding vectors, j is the index of the text prompt in the text description library, and j * is the index corresponding to the maximum similarity.

[0055] Optionally, the preset text description library is constructed by the following method:

[0056] Obtain a number of abnormal keywords and abnormal object categories;

[0057] Based on the abnormal keywords, construct a preset word library through the lexical database WordNet;

[0058] Screen through the contrastive language-image pre-training model with the visualization database ImageNet to obtain a number of sentence backbones;

[0059] Combine the sentence backbones, the preset word library, and the abnormal object categories to construct the preset text description library.

[0060] In the embodiments of the present invention, by constructing the above-mentioned preset text description library and screening out the optimal text prompt corresponding to the abnormal image through the encoder of the pre-trained contrastive language-image pre-training model (e.g., CLIP), the prior knowledge of the diffusion model (StableDiffusion) based on the latent variable model is fully utilized; during the training process, the abnormal appearance and shape are decoupled by embedding the similarity loss and augmenting the dataset, improving the authenticity and diversity of the generated anomalies; during the inference process, the alignment with the input mask is improved by increasing the local control intensity.

[0061] For ease of understanding, the following shows an example of construction: First, some abnormal keywords [wo], such as "crack", can be selected, and a thesaurus [wt] is constructed through the WordNet lexical database. The pre-trained contrastive language-image pre-training model (CLIP) is used to screen out 85 sentence backbones [wx] from the ImageNet visualization database. By combining [wx], [wt], and the category of the abnormal object, such as "capsule", a text description library [s] is formed, thereby obtaining the above-mentioned preset text description library.

[0062] Optionally, based on the abnormal image and the mask, a preset number of first sample pairs and data augmentation parameters are obtained through graphics processing, and second sample pairs are generated based on the first sample pairs, the data augmentation parameters, and the optimal text prompt, including:

[0063] Based on the abnormal image and the mask, a preset number of the first sample pairs and the data augmentation parameters are obtained through graphics processing, and the image processing includes at least one of the following: cropping, rotation, and translation;

[0064] The data augmentation parameters are added to the optimal text prompt to form an abnormal image, mask, and optimal text prompt sample pair to obtain the second sample pair.

[0065] In this implementation, based on the above abnormal image and the above mask, a preset number of the above first sample pairs and the data augmentation parameters are obtained through graphics processing, and then the data augmentation parameters are added to the above optimal text prompt to form an abnormal image, mask, and optimal text prompt sample pair to obtain the second sample pair. The data augmentation parameters obtained by this method can be added to the above optimal text prompt, so that the above preset model can learn the position and orientation information, and then the above second sample pair is obtained.

[0066] Specifically, the obtained abnormal image-mask sample pairs can be randomly cropped, rotated, and translated. One sample pair can be augmented to 100 (for example only), and the data augmentation parameters can be recorded, including cropx, cropy, angle, dx, and dy. Then, the above data augmentation parameters are added after the above optimal text prompt, so that the above preset model can learn the position and orientation information, forming an abnormal image-mask-optimal text prompt sample pair, that is, the above second sample pair.

[0067] It should be noted that by expanding the keywords of the abnormal semantics to construct many text prompts, comparing the similarity of these prompts with the images in the semantic embedding space, ensuring the generation of the best text prompts applicable to real abnormal images, making full use of the prior knowledge, accelerating the convergence of the model, and making the generated images more realistic.

[0068] Optionally, before using the preset model to take the second sample pair as input to obtain a set of predicted images matching the abnormal image, the method further includes:

[0069] Obtain a set of third sample pairs from the second sample pair. The third sample pair is a combination of a target optimal text prompt and a target mask. The target optimal text prompt is one of the text prompts in the optimal text prompts, and the target mask is one of the masks;

[0070] Erode the target mask to obtain a new masked mask;

[0071] Based on the target optimal text prompt and the new mask, obtain a set of target predicted images through the preset model;

[0072] Extract the features matching the set of target predicted images based on the contrastive language-image pre-training model, calculate the cosine similarity of the set of target predicted images, and construct an embedding similarity loss function to optimize and adjust the preset model;

[0073] Among them, the set of target predicted images, the new mask, and the target optimal text prompt can be combined into a fourth sample pair and added to the training set of the preset model to expand the data set.

[0074] In some special scenarios, the materials of specific parts of the same type of object are determined. Therefore, there is a strong correlation between the abnormal appearance and the abnormal position in the special scenario. The abnormal appearances with different shapes but the same position are usually similar, while the abnormal appearances at different positions are usually different. Therefore, a reasonable abnormal generation model should have the characteristic that the abnormal appearance is mainly affected by the abnormal position and less affected by the abnormal shape.

[0075] In this embodiment, an embedding similarity loss L is introduced es , which is used between the generated abnormal image and the optimal text prompt to constrain abnormal regions with the same position but different shapes to have similar appearances. Specifically, during the training process, first, a text prompt-mask pair is selected from the training set (which can be the second sample pair). Then, the mask is randomly eroded and occluded to simulate the situation where the masks have the same position but different shapes. As described above, the abnormal images generated with these new masks and text prompts should have similar appearances. Therefore, features are extracted from the text prompt and the generated abnormal image using the CLIP text and image encoders, and their cosine similarity is calculated to optimize and adjust the above preset model.

[0076] Among them, the construction of the embedding similarity loss function can be expressed by the following formula:

[0077]

[0078] Among them, n is the number of generated images, T is the text prompt, I is the real abnormal image, and I' is the generated abnormal image. In addition, L es can be regarded as a kind of distillation loss, making the distance between the generated image and the text prompt similar to the distance between the real image and the text prompt. By this method, overfitting that may be caused by directly enhancing the generated image to be close to the text prompt in the embedding space can be avoided. The new masks obtained by eroding and occluding the original mask and their generated images are used to increase the diversity of training data and improve the generalization ability of the model.

[0079] It should be noted that L es needs to be used in combination with the diffusion training loss function L train and can be expressed by the following formula:

[0080]

[0081] Among them, c v represents the image condition, c t represents the text prompt, t represents the time step, x0 is the input data, and ∈ is the noise sampled from the Gaussian distribution. The prediction error is minimized by optimizing the model parameters θ, making the predicted value ∈ θ (x t ,t,c t ,c v ) as close as possible to the real noise ∈.

[0082] Moreover, the overall objective function is the weighted sum of the diffusion training loss and the embedding similarity loss and can be expressed by the following formula:

[0083] L = L train + α·Les ;

[0084] Among them, α is a loss weight, usually assigned a value of 0.01.

[0085] It should be noted that by optimizing the embedding similarity loss between the generated abnormal images with the same abnormal location but different shapes and the generated optimal text prompts, the abnormal appearance and shape are decoupled, and the training data is increased to make the model more stable, improving the authenticity of the generated images while ensuring diversification.

[0086] Optionally, before using the preset model to take the second sample pair as input to obtain a set of predicted images matching the abnormal image, the method further includes:

[0087] Performing multi-scale sampling on the mask, setting the first region within the mask as the first control ratio, and setting the second region within the mask as the second control ratio to achieve enhanced local control of the second region;

[0088] Among them, the first region is the normal region, the second region is the abnormal region, and the first control ratio is less than the second control ratio.

[0089] In this implementation, it can be understood as a step of local control intensity, a local control enhancement strategy for achieving stronger control over the abnormal region during the inference process; downsampling the input mask to multiple scales, giving a lower control ratio to the normal region, and assigning a higher control ratio to the abnormal region. Through this method, the mask is shrunk to different sizes as the local control ratio during inference to strengthen the control over the abnormal region and improve the alignment.

[0090] For ease of understanding, the following gives an example: Using local control intensity during the inference process, assigning a lower control ratio of 1.0 to the normal region in the generated image, and assigning a higher control ratio to the abnormal region, which can be expressed by the following formula:

[0091] m′ k = β·m k + 1.0;

[0092] Among them, m′ k i.e., the local control intensity, m k refers to the multi-scale mask obtained by downsampling the input mask, k represents the index of different layers, and β is a hyperparameter, usually selected between 0.0 and 1.0.

[0093] As an alternative implementation, please refer to Figure 2 , Figure 2 which is a schematic flowchart of another method for generating abnormal images provided by an embodiment of the present invention, asFigure 2 As shown, this embodiment is applicable to the generation of abnormal industrial images. S11 is the image acquisition and annotation step, which is mainly used to acquire a small number of industrial abnormal images and manually annotate the abnormal areas on the workpiece to obtain a mask. Before this, the S12 step, that is, the step of constructing a text description library, can also be carried out. It mainly consists of a large number of text descriptions composed of abnormal keywords and sentence backbones. It should be noted that after the S12 step, the subsequent generation of other abnormal images no longer needs to repeat the S12 step.

[0094] Then it comes to the S13 step and the S14 step. The S13 step is about how to obtain the optimal text prompt. It mainly uses a pre-trained CLIP encoder to extract the features of abnormal images and text descriptions, and selects the one with the highest similarity as the optimal text prompt for abnormal images. The S14 step is data augmentation, which mainly performs data augmentation on the abnormal image-mask sample pairs and records the data augmentation parameters. The S13 step and the S14 step complete the S15 step, that is, constructing a training set. Specifically, the data augmentation parameters are added to the optimal text prompt to form an abnormal image-mask-optimal text prompt sample pair.

[0095] The abnormal image-mask-optimal text prompt sample pair in the S15 step is input into the fine-tuning model in the S4 step to obtain the final industrial abnormal image.

[0096] It should be noted that the fine-tuning of the preset model can be divided into the S3 step and the S2 step. Among them, the S2 step is to decouple the abnormal appearance and shape, mainly embed the similarity loss and perform dataset expansion, while the S3 step is to control the intensity locally. Specifically, the input mask is downsampled to multiple scales to give a stronger control intensity to the abnormal area.

[0097] Please continue to refer to Figure 3 , Figure 3 which is the parameter diagram provided by the embodiment of the present invention. As Figure 3 shown, abnormal keywords can construct a thesaurus through WordNet, and then a text description library is constructed by the sentence backbone. The text description library can extract text features through a pre-trained CLIP text encoder.

[0098] The initially acquired industrial abnormal images obtain a mask through manual annotation, and the industrial abnormal images also obtain image features through a pre-trained CLIP image encoder.

[0099] The text features and image features are compared through cosine similarity to obtain the optimal text prompt.

[0100] Through the data augmentation step, the data-augmented image-mask sample pairs and data augmentation parameters can be obtained, providing a data basis for finally obtaining the training dataset.

[0101] As an optional implementation manner, continuing with the industrial scenario as an example, please refer to Figure 4 , Figure 4 which is a schematic flowchart of the step of fine-tuning a preset model provided by an embodiment of the present invention. As shown in Figure 4 , based on the ControlNet model as the basic architecture, a decoupled anomaly manifestation and shape step and a local control intensity step are added to form a fine-tuning model. The Stable Diffusion v1.5 is fine-tuned using the fine-tuning model and a training set, and industrial anomaly images are generated using the fine-tuned model. Among them, the data augmentation parameters cropx, cropy, angle, dx, and dy in the input text prompt are all set to 0, so that the generated images will not have random rotation, translation, and cropping situations.

[0102] To facilitate understanding of the set of predicted images finally generated by the embodiment of the present invention, please refer to Figure 5 , Figure 5 which is a schematic diagram of some anomaly image samples provided by an embodiment of the present invention. As shown in Figure 5 , where the first row is the real anomaly image, the second row is the labeled mask, and the third row is the generated image obtained through the embodiment of the present invention.

[0103] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a device for generating anomaly images provided by an embodiment of the present invention. As shown in Figure 6 , the device 600 for generating anomaly images includes:

[0104] An acquisition module 601, configured to acquire an anomaly image and a mask, where the mask is obtained by labeling the anomaly image;

[0105] A first processing module 602, configured to extract features matching the anomaly image based on a contrastive language-image pre-training model and generate an optimal text prompt, where the optimal text prompt is a text description in a preset text description library;

[0106] A second processing module 603, configured to obtain a preset number of first sample pairs and data augmentation parameters through graphics processing based on the anomaly image and the mask, and generate a second sample pair according to the first sample pair, the data augmentation parameters, and the optimal text prompt. The first sample pair is a combination of the anomaly image and the mask, and the second sample pair is a combination of the anomaly image, the mask, and the optimal text prompt;

[0107] A third processing module 604, configured to use the second sample pair as an input and obtain a set of predicted images matching the anomaly image based on a preset model. The set of predicted images includes at least one predicted anomaly image.

[0108] Optionally, the first processing module 602 includes:

[0109] A first extraction unit, configured to extract features matching the abnormal image based on the contrastive language-image pre-training model;

[0110] A second extraction unit, configured to obtain a text description in a preset text description library, and extract features matching the text description through the contrastive language-image pre-training model;

[0111] A first generation unit, configured to calculate a cosine similarity between the features matching the abnormal image and the features matching the text description, and generate the optimal text prompt according to the cosine similarity, where the optimal text prompt is the text description with the highest matching degree to the abnormal image.

[0112] Optionally, the preset text description library is constructed by the following method:

[0113] Obtain a number of abnormal keywords and abnormal object categories;

[0114] Based on the abnormal keywords, construct a preset word library through the vocabulary database WordNet;

[0115] Filter through the contrastive language-image pre-training model using the visualization database ImageNet to obtain a number of sentence backbones;

[0116] Combine the sentence backbones, the preset word library, and the abnormal object categories to construct the preset text description library.

[0117] Optionally, the second processing module 603 includes:

[0118] A first processing unit, configured to perform image processing on the abnormal image and the mask to obtain a preset number of the first sample pairs and the data augmentation parameters, where the image processing includes at least one of the following: cropping, rotating, and translating;

[0119] An adding unit, configured to add the data augmentation parameters to the optimal text prompt to form an abnormal image, mask, and optimal text prompt sample pair, so as to obtain the second sample pair.

[0120] Optionally, the generating device 600 of the abnormal image further includes:

[0121] A fourth processing module, configured to obtain a group of third sample pairs from the second sample pairs, where the third sample pair is a combination of a target optimal text prompt and a target mask, the target optimal text prompt is one of the text prompts in the optimal text prompts, and the target mask is one of the masks;

[0122] A fifth processing module, configured to perform erosion processing on a target mask to obtain a new masked mask;

[0123] A generation module, configured to obtain a set of target prediction images through the preset model based on the target optimal text prompt and the new mask;

[0124] A first optimization module, configured to extract features matching the set of target prediction images based on the contrastive language-image pre-trained model, calculate the cosine similarity of the set of target prediction images, and construct an embedding similarity loss function to optimize and adjust the preset model;

[0125] Wherein, the set of target prediction images, the new mask, and the target optimal text prompt can be combined into a fourth sample pair and added to the training set of the preset model to expand the data set.

[0126] Optionally, the abnormal image generation device 600 further includes:

[0127] A setting module, configured to perform multi-scale sampling on the mask, set a first area in the mask as a first control ratio, and set a second area in the mask as a second control ratio to enhance local control of the second area;

[0128] Wherein, the first area is a normal area, the second area is an abnormal area, and the first control ratio is less than the second control ratio.

[0129] Specifically, referring to Figure 7 As shown, an embodiment of the present invention further provides an electronic device, including a bus 701, a transceiver 702, an antenna 703, a bus interface 704, a processor 705, and a memory 706.

[0130] The transceiver 702 is configured to obtain an abnormal image and a mask, wherein the mask is obtained by annotating the abnormal image;

[0131] The processor 705 is configured to extract features matching the abnormal image based on the contrastive language-image pre-trained model and generate an optimal text prompt;

[0132] The processor 705 is further configured to obtain a preset number of first sample pairs and data augmentation parameters through graphics processing based on the abnormal image and the mask, and generate a second sample pair based on the first sample pair, the data augmentation parameters, and the optimal text prompt, where the first sample pair is a combination of the abnormal image and the mask, and the second sample pair is a combination of the abnormal image, the mask, and the optimal text prompt;

[0133] The processor 705 is further configured to use the second sample pair as an input, and based on a preset model, obtain a set of predicted images that match the abnormal image, where the set of predicted images includes at least one predicted abnormal image.

[0134] In Figure 7 it, the bus architecture (represented by bus 701), bus 701 may include any number of interconnected buses and bridges, and bus 701 links together various circuits including one or more processors represented by processor 705 and a memory represented by memory 706. Bus 701 may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and thus will not be further described herein. The bus interface 704 provides an interface between bus 701 and transceiver 702. The transceiver 702 may be one element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices over a transmission medium. The data processed by the processor 705 is transmitted over a wireless medium via the antenna 703. Further, the antenna 703 also receives data and transmits the data to the processor 705.

[0135] The processor 705 is responsible for managing the bus 701 and general processing, and may also provide various functions including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory 706 may be used to store data used by the processor 705 when performing operations.

[0136] Optionally, the processor 705 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or a complex programmable logic device (CPLD).

[0137] An embodiment of the present invention further provides an electronic device, including: a processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, it implements each process of the above embodiment of the abnormal image generation method and can achieve the same technical effect. To avoid repetition, it will not be elaborated herein.

[0138] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned embodiment of the policy management method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0139] An embodiment of the present application further provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement each process of the above-mentioned method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0140] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0141] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment method can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0142] The above describes the embodiments of the present application in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all belong to the protection scope of the present application.

Claims

1. A method for generating an abnormal image, characterized in that, Including: Obtain an abnormal image and a mask, where the mask is obtained by annotating the abnormal image; Extract features matching the abnormal image based on a contrastive language-image pre-trained model and generate an optimal text prompt, where the optimal text prompt is a text description in a preset text description library; Based on the abnormal image and the mask, obtain a preset number of first sample pairs and data augmentation parameters through image processing, and generate second sample pairs based on the first sample pairs, the data augmentation parameters, and the optimal text prompt. The first sample pair is a combination of the abnormal image and the mask, and the second sample pair is a combination of the abnormal image, the mask, and the optimal text prompt; Use the second sample pair as an input and obtain a set of predicted images matching the abnormal image based on a preset model. The set of predicted images includes at least one predicted abnormal image.

2. The method for generating an abnormal image according to claim 1, wherein The step of extracting features matching the abnormal image based on a contrastive language-image pre-trained model and generating an optimal text prompt includes: Extract features matching the abnormal image based on the contrastive language-image pre-trained model; Obtain text descriptions in a preset text description library and extract features matching the text descriptions through the contrastive language-image pre-trained model; Calculate the cosine similarity between the features matching the abnormal image and the features matching the text description, and generate the optimal text prompt based on the cosine similarity. The optimal text prompt is the text description with the highest matching degree to the abnormal image.

3. The method for generating an abnormal image according to claim 2, wherein The preset text description library is constructed by the following method: Obtain a number of abnormal keywords and abnormal object categories; Based on the abnormal keywords, construct a preset word library through the WordNet vocabulary database; Use the contrastive language-image pre-trained model to screen with the ImageNet visualization database to obtain a number of sentence backbones; Combine the sentence backbones, the preset word library, and the abnormal object categories to construct the preset text description library.

4. The method for generating an abnormal image according to claim 3, wherein The step of obtaining a preset number of first sample pairs and data augmentation parameters through image processing based on the abnormal image and the mask, and generating second sample pairs based on the first sample pairs, the data augmentation parameters, and the optimal text prompt includes: Based on the abnormal image and the mask, obtain the preset number of first sample pairs and the data augmentation parameters through image processing. The image processing includes at least one of the following: cropping, rotation, and translation; Add the data augmentation parameters to the optimal text prompt to form a sample pair of the abnormal image, the mask, and the optimal text prompt to obtain the second sample pair.

5. The method for generating an abnormal image according to claim 4, wherein Before using the second sample pair as an input through the preset model to obtain a set of predicted images matching the abnormal image, the method further includes: Obtain a set of third sample pairs from the second sample pairs. The third sample pair is a combination of a target optimal text prompt and a target mask. The target optimal text prompt is one of the text prompts in the optimal text prompt, and the target mask is one of the masks. Erode the target mask to obtain a new masked mask. Based on the target optimal text prompt and the new mask, obtain a set of target predicted images through the preset model. Extract features matching the set of target predicted images based on the contrastive language-image pre-training model, calculate the cosine similarity of the set of target predicted images, and construct an embedding similarity loss function to optimize and adjust the preset model. Among them, the set of target predicted images, the new mask, and the target optimal text prompt can be combined into a fourth sample pair and added to the training set of the preset model to expand the data set.

6. The method for generating an abnormal image according to claim 4, wherein Before obtaining, through the preset model, the set of predicted images that match the abnormal image with the second sample pair as the input, the method further includes: Perform multi-scale sampling on the mask, set the first region in the mask as the first control ratio, and set the second region in the mask as the second control ratio to enhance the local control of the second region. Among them, the first region is a normal region, the second region is an abnormal region, and the first control ratio is less than the second control ratio.

7. An abnormal image generation device, characterized in that, Includes: An acquisition module for acquiring an abnormal image and a mask, where the mask is obtained by annotating the abnormal image. A first processing module for extracting features matching the abnormal image based on a contrastive language-image pre-training model and generating an optimal text prompt, where the optimal text prompt is a text description in a preset text description library. A second processing module for obtaining a preset number of first sample pairs and data augmentation parameters through graphics processing based on the abnormal image and the mask, and generating a second sample pair according to the first sample pair, the data augmentation parameters, and the optimal text prompt. The first sample pair is a combination of the abnormal image and the mask, and the second sample pair is a combination of the abnormal image, the mask, and the optimal text prompt. A third processing module for using the second sample pair as the input and obtaining, based on a preset model, a set of predicted images that match the abnormal image, where the set of predicted images includes at least one predicted abnormal image.

8. An electronic device, characterized in that, Includes a transceiver and a processor; The transceiver is used to acquire an abnormal image and a mask, where the mask is obtained by annotating the abnormal image. The processor is used to extract features matching the abnormal image based on a contrastive language-image pre-training model and generate an optimal text prompt, where the optimal text prompt is a text description in a preset text description library. The processor is further used to obtain a preset number of first sample pairs and data augmentation parameters through graphics processing based on the abnormal image and the mask, and generate a second sample pair according to the first sample pair, the data augmentation parameters, and the optimal text prompt. The first sample pair is a combination of the abnormal image and the mask, and the second sample pair is a combination of the abnormal image, the mask, and the optimal text prompt. The processor is further configured to use the second sample pair as an input, and based on a preset model, obtain a set of predicted images that match the abnormal image, where the set of predicted images includes at least one predicted abnormal image.

9. An electronic device, characterized in that, Comprising: A processor, a memory, and a program stored on the memory and executable on the processor, where when the program is executed by the processor, the steps of the method for generating an abnormal image according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the method for generating an abnormal image according to any one of claims 1 to 6 are implemented.