Image generation method and device and readable storage medium

By combining feature encoding module, feature decoding module and conditional diffusion model, the feature domain diffusion model is established, which solves the problem of insufficient reliability of image generation models in the prior art when the input samples are limited, and achieves higher feature consistency and reliability of image generation.

CN120070642APending Publication Date: 2025-05-30BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510240120.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When the input samples are limited, the existing image generation models are insufficient in the reliability of the generated images, making it difficult to maintain the correlation between the characteristics of the model generated samples and the characteristics of the target samples.

Method used

By combining the feature encoding module, feature decoding module and conditional diffusion model, a feature domain diffusion model is established, and an autoencoder is used to convert the image from pixel space to feature space, and then a generation model is established based on the feature space for sample generation. The generated feature image is decoded into pixel space through the feature decoding module to generate the target image.

Benefits of technology

The feature consistency of the model is improved, the correlation between the features of the model generated sample and the features of the target sample is enhanced, and the reliability of image generation is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070642A_ABST
    Figure CN120070642A_ABST
Patent Text Reader

Abstract

The invention provides an image generation method and device and a readable storage medium, belongs to the technical field of data processing, and aims to improve the reliability of generated images, the method comprises the following steps: obtaining random noise, a background category, a target object category and a mask pattern, the mask graph is used for determining the position of an image background in an image to be generated and the position of an image target object in the image to be generated; based on a feature coding module, coding the random noise and the mask pattern to generate a first feature pattern; based on a conditional diffusion model, generating a second feature map according to the first feature map, the background category and the target object category; wherein the conditional diffusion model comprises a backbone network and a branch network, the backbone network is used for carrying out feature description on an image background, and the branch network is used for carrying out feature description on an image target object; and based on a feature decoding module, decoding the second feature map to obtain a target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of data processing, and particularly relates to an image generation method, device, and readable storage medium. Background Art

[0002] The image generation model is a currently widely used and stable image augmentation method. However, it mostly generates data augmentation based on the characteristics of the entire input sample. When the input sample is limited, the reliability of the generated image is insufficient, and it is difficult to maintain the correlation between the features of the generated sample and the target sample. Summary of the Invention

[0003] This application provides an image generation method, device, and readable storage medium to solve the problem of insufficient reliability of the existing image generation model.

[0004] In the first aspect of the embodiments of this application, an image generation method is provided. The method includes:

[0005] Obtain random noise, background category, target object category, and a mask image, where the mask image is used to determine the position of the image background and the position of the image target object in the image to be generated;

[0006] Based on a feature encoding module, encode the random noise and the mask image to generate a first feature map;

[0007] Based on a conditional diffusion model, generate a second feature map according to the first feature map, the background category, and the target object category; wherein, the conditional diffusion model includes a backbone network and a branch network, the backbone network is used to describe the features of the image background, and the branch network is used to describe the features of the image target object;

[0008] Based on a feature decoding module, decode the second feature map to obtain the target image.

[0009] In the second aspect of the embodiments of this application, an industrial defect image generation method is further provided. The method includes:

[0010] Obtain random noise, background category, defect target category, and a mask image, where the mask image is used to determine the position of the image background and the position of the defect target in the image to be generated;

[0011] Based on a feature encoding module, encode the random noise and the mask image to generate a first feature map;

[0012] Based on the conditional diffusion model, a second feature map is generated according to the first feature map, the background category, and the defective target category; wherein, the conditional diffusion model includes a backbone network and a branch network, the backbone network is used to describe the features of the image background, and the branch network is used to describe the features of the defective target;

[0013] Based on the feature decoding module, the second feature map is decoded to obtain an industrial defect image, and the industrial defect image includes: defective targets belonging to the defective target category.

[0014] In the third aspect of the embodiments of the present application, an electronic device is further provided, including a processor and a memory, the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the image generation method described in the first aspect or the steps of the industrial defect image generation method described in the second aspect are implemented.

[0015] In the fourth aspect of the embodiments of the present application, a readable storage medium is further provided, and a program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the image generation method described in the first aspect or the steps of the industrial defect image generation method described in the second aspect are implemented.

[0016] The beneficial effect of the present application lies in: The image generation method proposed in the present application combines a feature encoding module, a feature decoding module, and a conditional diffusion model (i.e., a feature generation module including a backbone network and a branch network) to establish a feature domain diffusion model. An autoencoder (feature encoding module) is used to convert the image from the pixel space to the feature space, and then a generation model (conditional diffusion model) is established based on the feature space for sample generation. Then, the generated feature image (i.e., the second feature map) is decoded to the pixel space through the feature decoding module to generate the target image, so as to improve the feature consistency of the model (i.e., the correlation between the features of the samples generated by the model and the features of the target samples).

[0017] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the accompanying drawings required for the description of the embodiments or related technologies. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings. It should be noted that the ratios in the drawings are only for illustration and do not represent the actual ratios.

[0019] Figure 1 is a flowchart of the steps of an image generation method in an embodiment of the present application;

[0020] Figure 2 is a schematic structural diagram of a conditional control diffusion model in an embodiment of the present application;

[0021] Figure 3 is a schematic structural diagram of a feature encoding module and a feature decoding module in an embodiment of the present application;

[0022] Figure 4 is a schematic structural diagram of a feature encoding module and a feature decoding module after the model is extended in an embodiment of the present application;

[0023] Figure 5 is a schematic processing flow diagram of an attention fusion module in an embodiment of the present application;

[0024] Figure 6 is a schematic structural diagram of a backbone network in an embodiment of the present application;

[0025] Figure 7 is a schematic architecture diagram of a conditional control diffusion model in an embodiment of the present application;

[0026] Figure 8 is a schematic diagram of an electronic device in an embodiment of the present application. Detailed implementation manners

[0027] To make the above objects, features, and advantages of the present application more obvious and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0028] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same category, and the number of objects is not limited. For example, the first object can be one or at least two. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / " generally means that the related objects before and after are in an "or" relationship.

[0029] The image generation model is a widely adopted and stable image augmentation method at present. However, most of them generate data augmentation based on the characteristics of the entire input sample. When the input sample is relatively limited, the diversity and difference of the generated results are relatively limited, and the detail generation ability of specific targets is relatively limited. Therefore, a generation network with a certain diversity expansion ability is required to generate reliable augmented data to improve the performance of downstream detection and other tasks and avoid overfitting and other phenomena. The conditional control diffusion model can perform differential expansion for samples with different characteristics, but it requires a large amount of reliable and detailed text descriptions as conditional guidance for data generation. Therefore, the current image generation model still has room for improvement. Especially when the input sample is relatively limited, the reliability of the generated image is insufficient, and it is difficult to maintain the correlation between the characteristics of the samples generated by the model and the characteristics of the target samples.

[0030] In view of the above problems, this application provides an image generation method, device and readable storage medium. By combining a feature encoding module, a feature decoding module and a conditional diffusion model (i.e., a feature generation module including a backbone network and a branch network), a feature domain diffusion model is established. The autoencoder (feature encoding module) is used to convert the image from the pixel space to the feature space, and then a generation model (conditional diffusion model) is established based on the feature space for sample generation. Then, the generated feature image (i.e., the second feature map) is decoded to the pixel space through the feature decoding module to generate the target image, so as to improve the feature consistency of the model (i.e., the correlation between the characteristics of the samples generated by the model and the characteristics of the target samples).

[0031] The first aspect of this application proposes an image generation method. Below, the image generation method proposed in the embodiments of this application will be specifically described in Sections 1.1-1.5.

[0032] 1.1. General description of the overall solution of the image generation method.

[0033] The image generation method is applied to a conditional controlled diffusion model, and the conditional controlled diffusion model at least includes: a feature encoding module, a conditional diffusion model, and a feature decoding module. Refer to Figure 1 , Figure 1 shows a flowchart of steps of an image generation method. As Figure 1 shown, the method includes:

[0034] Step S101, obtain random noise, a background category, a target object category, and a mask map, where the mask map is used to determine the position of the image background in the image to be generated and the position of the image target object in the image to be generated.

[0035] Specifically, the random noise is randomly generated noise, and the embodiments of the present application do not limit the manner of generating random noise, and it can be any one of existing noise generation methods. The background category refers to the category to which the background in the image to be generated belongs, such as forest, grassland, etc. The target object category refers to the category to which the target object located in the background in the image to be generated (i.e., the target image) belongs, such as dog, vehicle, house, etc. The mask map, that is, the mask map, is used to represent the position of the image background in the target image and the position information of the target object in the target image. Exemplarily, this method can be applied to generate industrial defect images, that is, the generated target image can also be a defect image under an industrial background. The background in the defect image is an industrial product (the background category is the product type), and the target object in the defect image is a defect target, that is, defects, flaws, etc. on the surface of the industrial product (the target object category is the type of the defect target).

[0036] Step S102, based on the feature encoding module, encode the random noise and the mask map to generate a first feature map.

[0037] Specifically, refer to Figure 2 , Figure 2 shows a schematic structural diagram of a conditional controlled diffusion model. As Figure 2 shown, first input the random noise and the mask map into the feature encoding module for feature encoding to obtain a first feature map. Among them, the feature encoding module can be an AE encoding module (that is, the encoding module of an autoencoder). In this embodiment, an autoencoder is used to convert the image from the pixel space to the feature space.

[0038] Step S103, based on the conditional diffusion model, generate a second feature map according to the first feature map, the background category, and the target object category; wherein, the conditional diffusion model includes a backbone network and a branch network, the backbone network is used to describe the features of the image background, and the branch network is used to describe the features of the image target object.

[0039] Specifically, the conditional diffusion model includes a backbone network and a branch network. The backbone network can be used to describe the features of the image background based on the first feature map and the background category, so as to generate background-related features; the branch network can be used to describe the features of the image target object based on the first feature map and the target object category, so as to generate target object-related features. The features generated by the backbone network are fused with the features generated by the branch network to obtain a second feature map, so that the second feature map contains features related to the background and features related to the target object.

[0040] Step S104: Decode the second feature map based on the feature decoding module to obtain the target image.

[0041] Specifically, this embodiment adopts a diffusion generation scheme based on the feature domain, that is, both the input and output of the diffusion model are feature maps of the image. The input is the first feature map, and the output is the second feature map. Therefore, it is necessary to perform feature encoding and decoding on the original input, that is, first use the feature encoding module to perform feature encoding to obtain the first feature map, and finally use the feature decoding module to decode the output second feature map to obtain the target image. Among them, the feature encoding module and the feature decoding module can be the encoding module and the decoding module of an AE model (i.e., an autoencoder).

[0042] The image generation method proposed in this application combines the feature encoding module, the feature decoding module, and the conditional diffusion model (i.e., the feature generation module including the backbone network and the branch network) to establish a feature domain diffusion model. The autoencoder (feature encoding module) is used to convert the image from the pixel space to the feature space, and then a generation model (conditional diffusion model) is established based on the feature space for sample generation. Then, the generated feature image (i.e., the second feature map) is decoded to the pixel space through the feature decoding module to generate the target image, so as to improve the feature consistency of the model (i.e., the correlation between the features of the samples generated by the model and the features of the target samples).

[0043] 1.2. A scheme for training the feature encoding module and the feature decoding module.

[0044] This embodiment proposes to use the training data set to train the feature encoding module and the feature decoding module in the conditional control diffusion model separately.

[0045] In a possible implementation manner, before inputting the first feature map and the background category into the feature conversion module, the method further includes:

[0046] Step S201: Obtain a training data set, and each training sample data in the training data set includes: a sample image, and the sample background category, sample target object category, and sample mask map corresponding to the sample image.

[0047] Specifically, each training sample data in the training dataset used for training the feature encoding module and the feature decoding module mainly includes three elements: a sample image (including a sample background and a sample target object), the sample background category, the sample target object category, and the sample mask map corresponding to the sample image.

[0048] Step S202: Construct an extension module, which is used to perform classification prediction on the feature map output by the feature encoding module, and determine the background category and the target object category of the image represented by the feature map output by the feature encoding module.

[0049] Specifically, referring to Figure 3 , Figure 3 shows a schematic structural diagram of a feature encoding module and a feature decoding module. As Figure 3 shown, the structure of the feature encoding module is x cascaded downsampling convolutional modules, and the feature decoding module is x cascaded upsampling convolutional modules. During the training process, the input of the feature encoding module is the sample image. After passing through x cascaded downsampling convolutional modules, the output of the encoder is a feature map of c*ww*hh. The input of the feature decoding module is a feature map with a size of c*ww*hh. After passing through x cascaded upsampling convolutional modules, the output of the feature decoding module is a reconstructed sample image with the same size as the sample image. Among them, c is the dimension of the finally encoded feature image, and ww and hh are the sizes of the image length and width after being downsampled by 2x times respectively. Exemplarily, x can be 4. Each downsampling convolutional module is composed of a convolutional layer with a convolutional kernel size of 3*3 and a stride of 2 and two convolutional layers with convolutional kernel sizes of 3*3 and a stride of 1 cascaded together or connected using a residual structure; each upsampling convolutional module is composed of a transposed convolutional layer with a convolutional kernel size of 4*4 and a stride of 2, or composed of a 2-fold interpolation structure and two convolutional layers with convolutional kernel sizes of 3*3 and a stride of 1 cascaded together. The downsampling ratio of the feature encoding module in the above example is 16 times.

[0050] During the training process, it is necessary to expand the model and construct an encoding-decoding feature guidance structure (i.e., the extension module) to assist in training. In this embodiment, a multi-level cascaded linear structure is used to perform feature classification on the feature image. Through the loss guidance of the classification vector, the effect of assisting training is achieved. The structure after completing the model expansion is as Figure 4 shown, Figure 4 shows a schematic structural diagram of a feature encoding module and a feature decoding module after model expansion. As Figure 4As shown, an extension module is added after the feature encoding module. The input of this extension module is the output of the feature encoding module. This extension module is used to perform classification prediction based on the output of the feature encoding module (i.e., the feature map) to obtain a classification result, which characterizes the background type and target object type corresponding to the feature map.

[0051] Step S203: Based on the training data set, train the feature encoding module and the feature decoding module, and use the extension module to calculate the loss for each training.

[0052] In this embodiment, the gradient descent method can be used to train the model parameters (feature encoding module and feature decoding module) through the loss function.

[0053] In a possible implementation manner, the step S203: Based on the training data set, train the feature encoding module and the feature decoding module, and use the extension module to calculate the loss for each training, includes:

[0054] Step S203-1: Input the sample image and the sample mask image into the feature encoding module to be trained to obtain a first sample feature map.

[0055] Step S203-2: Input the first sample feature map into the extension module to obtain a classification result, which represents the background category and target object category of the image characterized by the first sample feature map.

[0056] Step S203-3: Input the first sample feature map into the feature decoding module to be trained to obtain a reconstructed sample image.

[0057] Step S203-4: Calculate the training loss according to the reconstructed sample image, the classification result, and the sample background category and sample target object category corresponding to the sample image.

[0058] In a possible implementation manner, the training loss includes: a feature classification loss and a decoded image loss; step S203-4: Calculate the training loss according to the reconstructed sample image, the classification result, and the sample background category and sample target object category corresponding to the sample image, includes:

[0059] Calculate the cross entropy between the feature classification vector representing the classification result and the label classification vector representing the sample background category and sample target object category corresponding to the sample image to obtain the feature classification loss;

[0060] Calculate the cosine similarity between the reconstructed sample image and the sample image to obtain the decoded image loss.

[0061] Specifically, the model loss consists of two parts: feature classification loss and decoded image loss.

[0062] Among them, the feature classification loss is calculated from the first sample feature map output by the feature encoding module through the feature classification vector (representing the classification result) and the label classification vector (representing the sample background category and the sample target object category) output by the expansion module. The calculation method is the cross-entropy loss between the output classification result of the expansion module and the class label (i.e., the sample background category and the sample target object category corresponding to the sample image). In this embodiment, the class label used in the training process at this stage is the combination of the sample background category and the sample target object category, and its corresponding label classification vector has a score of 1 only at the corresponding background category and target category, and the scores of other categories are 0.

[0063] The decoded image loss is calculated from the sample image input to the feature encoding module and the reconstructed sample image output by the feature decoding module. The calculation method is the pixel-level cosine similarity loss between the sample image and the reconstructed sample image.

[0064] Step S203-5: Update the parameters of the feature encoding module to be trained and the feature decoding module to be trained according to the training loss (repeat the above steps until the preset number of training times is reached or the loss value converges, and end the training), and obtain the trained feature encoding module and feature decoding module.

[0065] Each time steps S203-1 to S203-5 are executed, it is regarded as one training. By repeating the above process, iterative training of the feature encoding module and the feature decoding module is realized until the preset number of training times is reached, or the value of the loss function gradually decreases until it stabilizes and no longer changes, and then the training ends, and the trained feature encoding module and feature decoding module are obtained. Thus, using the feature encoding module, the feature decoding module, and the conditional diffusion model, a feature domain diffusion model is constructed. The image is converted from the pixel space to the feature space by using an autoencoder, and then a generative model is established based on the feature space for sample generation, and then the generated feature image is decoded into the pixel space, thereby improving the correlation between the features of the samples generated by the model and the features of the target samples. Aiming at the data limitations in the industrial scenario, the effect of data augmentation is achieved through the image generation network.

[0066] 1.3. The multi-level fusion module in the conditional control diffusion model.

[0067] In this embodiment, the conditional diffusion model includes a backbone network, a branch network, and a multi-level fusion module between the two. As Figure 1As shown, the backbone network and the branch network perform feature fusion through a multi-level fusion module. In this embodiment, a conditional diffusion model is used to generate images of specific categories or characteristics under the guidance of input conditions (such as category information).

[0068] In a possible implementation, the conditional diffusion model further includes: a multi-level fusion module; step S103, based on the conditional diffusion model, generating a second feature map according to the first feature map, the background category, and the target object category, including:

[0069] Step S1031, input the first feature map and the background category into the backbone network, and input the first feature map and the target object category into the branch network, so that the backbone network and the branch network respectively perform multi-level feature extraction according to the input. After each feature extraction is completed in the backbone network and the branch network, through the multi-level fusion module, fuse the backbone feature vector extracted by the backbone network, the branch feature vector extracted by the branch network, and the mask map to obtain a fused feature vector.

[0070] Step S1032, send the fused feature vector to the backbone network for the next level of feature extraction.

[0071] Step S1033, generate the second feature map according to the first feature transformation vector obtained by the backbone network through multi-level feature extraction and the second feature transformation vector obtained by the branch network through multi-level feature extraction.

[0072] In this embodiment, as Figure 1 shown, making the backbone network and the branch network respectively perform multi-level feature extraction according to the input in step S1031 means that the backbone network performs multi-level feature extraction according to the input first feature map and the background category, and the branch network performs multi-level feature extraction according to the input first feature map and the target object category. The backbone feature vector extracted at each level in the backbone network will be fused with the branch feature vector extracted at the corresponding level in the branch network through the multi-level fusion module to obtain a fused feature vector. Exemplarily, for the backbone feature vector A extracted at the 3rd level in the backbone network and the branch feature vector B extracted at the 3rd level in the branch network, they are fused through the multi-level fusion module to obtain a fused feature vector C. Then the fused feature vector C is sent to the backbone network for the next level of feature extraction (i.e., the backbone network performs the 4th level of feature extraction).

[0073] The backbone network obtains the first feature transformation vector through multi-level feature extraction, and the branch network obtains the second feature transformation vector through multi-level feature extraction. By fusing the first feature transformation vector and the second feature transformation vector, the second feature map is decoded.

[0074] In a possible implementation manner, the multi-level fusion module includes multiple layers of attention fusion modules, and each attention fusion module includes: a cascaded first self-attention sub-module, a first cross-attention sub-module, a second self-attention sub-module, and a second cross-attention sub-module.

[0075] Specifically, the multi-level fusion module includes multiple cascaded attention fusion modules. Refer to Figure 5 , Figure 5 shows a schematic diagram of the processing flow of an attention fusion module. As Figure 5 shown, each attention fusion module includes: a first self-attention sub-module, a first cross-attention sub-module, a second self-attention sub-module, and a second cross-attention sub-module cascaded in the following order. The conditional control diffusion model disclosed in this embodiment is constructed based on the backbone network and the branch network. In the process of using the output features (branch feature vectors) of the branch network as conditions to control the backbone network, feature fusion needs to be performed through multiple attention fusion modules (i.e., multi-level fusion modules). The attention fusion module is initialized with 0 parameters, and its parameters are continuously learned and updated through iteration and weighted into the backbone network.

[0076] Through the multi-level fusion module, fusing the backbone feature vector extracted by the backbone network, the branch feature vector extracted by the branch network, and the mask graph to obtain a fused feature vector, includes:

[0077] Step S301, obtain the backbone feature vector extracted by the backbone network, input the backbone feature vector into the first self-attention sub-module, and obtain the first self-attention feature vector.

[0078] Step S302, using the downsampled mask graph as the K value and the V value, using the branch feature vector extracted by the branch network as the Q value, input the first self-attention feature vector and the backbone feature vector into the first cross-attention sub-module together, and obtain the first cross-attention feature vector.

[0079] Step S303, input the first cross-attention feature vector and the first self-attention feature vector into the second self-attention sub-module together, and obtain the second self-attention feature vector.

[0080] Step S304: Using the branch feature vectors extracted by the branch network as the K value and V value, and using the second self-attention feature vector as the Q value, input the second self-attention feature vector and the first cross-attention feature vector into the second cross-attention sub-module to obtain the second cross-attention feature vector.

[0081] Step S305: Fuse the second self-attention feature vector and the second cross-attention feature vector to obtain the fused feature vector.

[0082] The embodiment of the present application discloses an improved feature fusion structure as an attention fusion module. This structure fuses the mask graph, branch feature vectors, and backbone feature vectors by constructing multiple self-attention-cross-attention hierarchical structures, and is guided based on the mask (i.e., the mask graph) while realizing conditional control.

[0083] 1.4. Specific architectures of the backbone network and the branch network in the conditional control diffusion model.

[0084] Most of the related generative networks (diffusion models) are established based on CNN. Considering that the ViT model has also achieved certain results in the field of images, especially in the field of large models, its performance may exceed that of CNN. The embodiment of the present application proposes to use the architecture of the Vision Transformer (ViT) model to replace CNN to further improve the image generation performance of the model. In this embodiment, the performance advantages of the transformer model are exerted through the transformer vision transformer model, and the model feature extraction and expression capabilities are improved.

[0085] 1.4.1. Architecture of the backbone network.

[0086] In a possible implementation manner, the backbone network includes: a first conditional encoding module, a first feature serialization module, and a multi-level cascaded first feature conversion module.

[0087] Specifically, referring to Figure 6 , Figure 6 shows a schematic structural diagram of a backbone network, such as Figure 6As shown in the figure, the backbone network includes: a first conditional encoding module, a first feature serialization module, and a multi-level cascaded first feature transformation module. Among them, the first conditional encoding module is used to encode the input conditional information (the background category and the random noise reduction steps) to obtain a first conditional feature vector; the first feature serialization module is used to perform feature serialization on the input first feature map to obtain a first serialized feature vector. The multi-level cascaded first feature transformation module is a multi-level transformer feature transformation module, which is used to perform multi-level feature extraction on the first serialized feature vector. During the feature extraction at each level, the first conditional encoding module sends the encoded first conditional feature vector to the first feature transformation module at the corresponding level to perform conditional control on the feature transformation according to the conditional information (the first conditional feature vector).

[0088] Step S1031, the backbone network performs multi-level feature extraction according to the input, including:

[0089] Step S1031-1, input the first feature map into the first feature serialization module to obtain a first serialized feature vector. Specifically, the first feature serialization module converts the input first feature map into a feature map downsampled by x times through a convolutional layer with a convolutional kernel size of x*x and a stride of x, flattens it into a feature vector, and obtains the feature serialization result of the input first feature map, that is, the first serialized feature vector, by superimposing position encoding parameters on the feature vector.

[0090] Step S1031-2, input the background category and the random noise reduction steps into the first conditional encoding module to obtain a first conditional feature vector.

[0091] In a possible implementation manner, the first conditional encoding module includes: a noise reduction step encoding module and a category encoding module; Step S1031-2, input the background category and the random noise reduction steps into the first conditional encoding module to obtain a first conditional feature vector, including:

[0092] Input the background category into the category encoding module, and convert the background category into a category feature vector through text encoding. Specifically, the category encoding module converts the category vocabulary (background category) into a set of vector representations, that is, the category feature vector.

[0093] Perform sequence conversion on the random noise reduction steps through the noise reduction step encoding module, and convert the sequence-converted step sequence into a random noise reduction step feature vector; the first conditional feature vector includes: the category feature vector and the random noise reduction step feature vector. Specifically, the noise reduction step encoding module performs vector representation on the random noise reduction steps: embed t= t * exp(-log(maxp) * range(0, feat_dim)); The noise reduction step encoding module converts the steps into a sequence through the above formula, and then converts the step sequence into a vector representation of a specified dimension through cascaded linear layers, that is, the random noise reduction step feature vector. Among them, embed t represents the random noise reduction step feature vector output by the noise reduction step encoding module, t represents the random noise reduction step, exp() represents the exponential operation, -log() represents the logarithmic operation, maxp represents the preset maximum noise reduction step limit, range() represents an equally spaced sequence, and feat_dim represents the dimension of the feature vector.

[0094] Step S1031-3: Input the first serialized feature vector into the first feature conversion module with multi-level cascading for multi-level feature extraction to obtain the first feature conversion vector; among them, for each level of the first feature conversion module, based on the first conditional feature vector, perform feature extraction on the fused feature vector output by the attention fusion module of the previous level, output the backbone feature vector, and use the backbone feature vector output by the first feature conversion module of the last level as the first feature conversion vector.

[0095] Specifically, for example Figure 6 as shown, input the first serialized feature vector into the first feature conversion module with multi-level cascading for multi-level feature extraction. Specifically, for the first feature conversion module of the first level, based on the first conditional feature vector, perform feature extraction according to the input first serialized feature vector to obtain the backbone feature vector output by the first feature conversion module of the first level. This backbone feature vector is used for the first feature conversion module of the next level (i.e., the second level) to perform feature extraction.

[0096] For each subsequent level of the feature transformation module (i.e., the second level to the last level), there are two inputs: one is the first conditional feature vector sent by the first conditional encoding module to the first feature transformation modules at each level; the other is the attention fusion module, which fuses the backbone feature vector and the branch feature vector of the previous level (according to the above steps S301-305) to obtain a fused feature vector. Exemplarily, the first feature transformation module of the 3rd level in the backbone network and the second feature transformation module of the 3rd level in the branch network are input into the attention fusion module of the 3rd level in the multi-level fusion module for fusion (according to the above steps S301-305) to obtain a fused feature vector. Then, it is used as the input of the first feature transformation module of the 4th level in the backbone network, and feature extraction is performed based on the first conditional feature vector to obtain the backbone feature vector output at the 4th level. The backbone feature vector output by the first feature transformation module of the last level in the multi-level cascaded first feature transformation modules in the backbone network can be used as the first feature transformation vector finally output by the multi-level cascaded first feature transformation modules.

[0097] In a possible implementation manner, each of the first feature transformation modules includes: a cascaded first linear layer, a first attention layer, and a first multi-level linear transformation layer. Steps S1031-3, the feature extraction of the fused feature vector output by the attention fusion module of the previous level based on the first conditional feature vector, and the output of the backbone feature vector, include:

[0098] Input the first conditional feature vector into the first linear layer to obtain an attention coefficient and a linear coefficient, where the attention coefficient includes: an attention weight s_a, an attention bias b_a, and an attention percentage g_a; the linear coefficient includes: a linear weight s_m, a linear bias b_m, and a linear percentage g_m. Specifically, the first linear layer divides the input first conditional feature vector into 6 coefficients, and uses these coefficients to control subsequent feature transformation.

[0099] Input the fused feature vector output by the attention fusion module of the previous level into the first attention layer, and enable the first attention layer to perform attention transformation based on the attention coefficient to obtain a first transformed attention feature vector.

[0100] Specifically, the first attention layer performs feature transformation according to the following formula to obtain a first transformed attention feature vector:

[0101] x' = x + g_a * (layer_attention(norm(x) * s_a + b_a)); where x' is the first transformed attention feature vector output by the first attention layer, x is the input of the first attention layer (i.e., the fused feature vector output by the previous-level attention fusion module), s_a is the attention weight, b_a is the attention bias, g_a is the attention percentage, norm() is the batch normalization layer, and layer_attention() is the first attention layer.

[0102] Input the first transformed attention feature vector into the first multi-level linear transformation layer, and make the first multi-level linear transformation layer perform a linear transformation based on the linear coefficients to obtain the backbone feature vector. Specifically, the first multi-level linear transformation layer is cascaded after the first attention layer and performs a linear transformation on the first transformed attention feature vector to obtain the backbone feature vector.

[0103] Specifically, the first multi-level linear transformation layer performs a linear transformation based on the linear coefficients according to the following formula:

[0104] x' = x + g_m * (layer_mlp(norm(x) * s_m + b_m)); where x' is the backbone feature vector output by the first multi-level linear transformation layer, x is the input of the first multi-level linear transformation layer (i.e., the first transformed attention feature vector output by the first attention layer), s_m is the linear weight, b_m is the linear bias, g_m is the linear percentage, norm() is the batch normalization layer, and layer_mlp() is the first multi-level linear transformation layer.

[0105] As Figure 6 shown, the backbone network may further include a linear decoding module and a feature reconstruction module. The linear decoding module is cascaded after the multi-level cascaded first feature transformation module to decode and reconstruct the extracted feature vector to obtain a second feature map. Specifically, in a possible implementation manner, the backbone network further includes: a linear decoding module and a feature reconstruction module.

[0106] Step S1033, generating the second feature map according to the first feature transformation vector obtained by multi-level feature extraction by the backbone network and the second feature transformation vector obtained by multi-level feature extraction by the branch network, includes:

[0107] Step S1033-1, input the first feature transformation vector and the second feature transformation vector into the attention fusion module of the last level for feature fusion to obtain an attention fusion feature vector. Refer to Figure 7 , Figure 7 shows a schematic architecture diagram of a conditional control diffusion model. AsFigure 7 As shown, the first feature transformation vectors output by the multi-level cascaded first feature transformation modules in the backbone network, and the second feature transformation vectors output by the multi-level cascaded second feature transformation modules in the branch network are fused through the attention fusion module (the last-level attention fusion module in the multi-level fusion module) (according to the above steps S301-305) to obtain the attention fusion feature vector.

[0108] In step S1033-2, the attention fusion feature vector and the first conditional feature vector are input into the linear decoding module for decoding to obtain a linear decoding result.

[0109] In a possible implementation manner, in step S1033-2, inputting the attention fusion feature vector and the first conditional feature vector into the linear decoding module for decoding to obtain a linear decoding result includes:

[0110] Encoding the first conditional feature vector through the first linear layer of the linear decoding module to obtain a first linear weight and a first linear bias;

[0111] Using the first linear weight and the first linear bias, decoding the attention fusion feature vector through the second linear layer of the linear decoding module to obtain the linear decoding result.

[0112] Among them, the linear decoding module includes two linear layers. One linear layer (the second linear layer) is a linear conversion layer for decoding the transformed feature vector (the attention fusion feature vector), and one linear layer (the first linear layer) encodes the input condition (the first conditional feature vector) into two coefficients, namely the first linear weight s_l and the first linear bias b_l, which act on the linear conversion layer (i.e., the second linear layer) according to the following formula:

[0113] x' = layer_linear(norm(x)*s_l + b_l); where x' represents the output of the second linear layer, that is, the linear decoding result, x is the input of the second linear layer, that is, the attention fusion feature vector, norm() is the batch normalization layer, layer_linear() is the second linear layer, s_l is the first linear weight, and b_l is the first linear bias.

[0114] In step S1033-3, the linear decoding result is input into the feature reconstruction module to obtain the second feature map. Specifically, the feature reconstruction module reconstructs the feature vector with an output of bs*hw*ppc (i.e., the linear decoding result) into a feature map of bs*c*hp*wp (i.e., the second feature map). As Figure 7As shown in the figure, the backbone network outputs the second feature map to the feature decoding module through the feature reconstruction module, and the feature decoding module decodes it to obtain the target image. In this embodiment, the target image and the category information are used as conditions to jointly control the backbone network, so as to achieve the effect of diverse data augmentation for generating the whole image with the target separated from the background.

[0115] 1.4.2 Architecture of the branch network.

[0116] The branch network disclosed in this embodiment takes the architecture of the backbone network as a reference, and is composed of a feature serialization module and a multi-level feature transformation module. Its network structure is the same as that of the backbone network. The object of action of this model is the defective target, and the purpose is to describe the features of the defective target.

[0117] In a possible implementation manner, the branch network includes: a second conditional encoding module, a second feature serialization module, and a multi-level cascaded second feature transformation module.

[0118] Specifically, as Figure 7 shown in the figure, the branch network includes: a second conditional encoding module, a second feature serialization module, and a multi-level cascaded second feature transformation module. Among them, the second conditional encoding module is used to encode the input conditional information (the target object category and the random noise reduction steps) to obtain a second conditional feature vector; the second feature serialization module is used to serialize the features of the input first feature map to obtain a second serialized feature vector. The multi-level cascaded second feature transformation module is a multi-level transformer feature transformation module, which is used to perform multi-level feature extraction on the second serialized feature vector. When performing feature extraction at each level, the second conditional encoding module sends the encoded second conditional feature vector to the corresponding level of the second feature transformation module to conditionally control the feature transformation according to the conditional information (the second conditional feature vector).

[0119] When using the specified mask map, background category, target object category, 2 independent random noise maps, and the corresponding random noise reduction steps as inputs, correspondingly, the random noises are respectively input into the feature encoder to obtain random noise feature maps as subsequent inputs. The mask map is downsampled and flattened and then input into each attention fusion module respectively. The background category, random noise feature map, and random noise reduction steps are input into the backbone network, and the target category, random noise feature map, and random noise reduction steps are input into the branch. In this way, through the conditional control diffusion model disclosed in this embodiment, the target image is output, that is, a sample is generated in the specified mask area under a specific background of a specific target object.

[0120] Step S1031, the branch network performs multi-level feature extraction according to the input, including:

[0121] Step S1031-4: Input the first feature map into the second feature serialization module to obtain a second serialized feature vector. Specifically, the second feature serialization module uses a convolutional layer with a convolutional kernel size of x*x and a stride of x to convert the input first feature map into a feature map downsampled by x times, and flattens it into a feature vector. By superimposing position encoding parameters on the feature vector, the feature serialization result of the input first feature map, that is, the second serialized feature vector, is obtained. The structure of this second feature serialization module is the same as that of the first feature serialization module.

[0122] Step S1031-5: Input the target object category and the random noise reduction steps into the second conditional encoding module to obtain a second conditional feature vector.

[0123] Specifically, the architecture of the second conditional encoding module is the same as that of the first conditional encoding module, and includes: a noise reduction step encoding module and a category encoding module; Step S1031-5 includes: Input the target object category into the category encoding module, and convert the target object category into a category feature vector through text encoding. Specifically, the category encoding module converts category vocabulary (target object category) into a set of vector representations, that is, category feature vectors. The random noise reduction steps are subjected to sequence conversion by the noise reduction step encoding module, and the sequence-converted step sequence is converted into a random noise reduction step feature vector; the second conditional feature vector includes: a category feature vector and a random noise reduction step feature vector. The processing process of the second conditional encoding module in this embodiment is the same as that of the first conditional encoding module, and will not be repeated in this embodiment.

[0124] Step S1031-6: Input the second serialized feature vector into the multi-level cascaded second feature conversion module for multi-level feature extraction to obtain the second feature conversion vector; where, for each level of the second feature conversion module, based on the second conditional feature vector, feature extraction is performed on the output of the second feature conversion module of the previous level to obtain the branch feature vector of the second feature conversion module of this level, and the branch feature vector is sent to the attention fusion module of the corresponding level, and the branch feature vector output by the second feature conversion module of the last level is used as the second feature conversion vector.

[0125] Specifically, for example Figure 7As shown, the second serialized feature vector is input into the multi-level cascaded second feature conversion module for multi-level feature extraction. Specifically, for the second feature conversion module of the first level among them, based on the second conditional feature vector, feature extraction is performed according to the input second serialized feature vector to obtain the branch feature vector output by the second feature conversion module of the first level. This branch feature vector is used for the second feature conversion module of the next level (i.e., the second level) to perform feature extraction.

[0126] For each subsequent level of the second feature conversion module (i.e., from the second level to the last level), there are two inputs: one is the second conditional feature vector sent by the second conditional encoding module to the second feature conversion modules of each level; the other is the output (branch feature vector) of the second feature conversion module of the previous level. Exemplarily, the branch feature vector output by the second feature conversion module of the 3rd level in the branch network is obtained, and then it is used as the input of the second feature conversion module of the 4th level of the backbone network, enabling it to perform feature extraction based on the second conditional feature vector to obtain the branch feature vector output by the 4th level. The branch feature vector output by the second feature conversion module of the last level in the multi-level cascaded second feature conversion module in the branch network can be used as the second feature conversion vector finally output by the multi-level cascaded second feature conversion module.

[0127] In a possible implementation manner, each of the second feature conversion modules includes: a cascaded second linear layer, a second attention layer, and a second multi-level linear conversion layer;

[0128] The feature extraction of the output of the second feature conversion module of the previous level based on the second conditional feature vector to obtain the branch feature vector of the second feature conversion module of this level includes:

[0129] Input the second conditional feature vector into the second linear layer to obtain an attention coefficient and a linear coefficient, where the attention coefficient includes: an attention weight, an attention bias, and an attention percentage; the linear coefficient includes: a linear weight, a linear bias, and a linear percentage;

[0130] Input the branch feature vector output by the second feature conversion module of the previous level into the second attention layer, and enable the second attention layer to perform attention conversion based on the attention coefficient to obtain a second transformed attention feature vector;

[0131] Input the second transformed attention feature vector into the second multi-level linear conversion layer, and enable the second multi-level linear conversion layer to perform linear conversion based on the linear coefficient to obtain the branch feature vector of the second feature conversion module of this level.

[0132] Specifically, the second feature transformation module has the same structure as the first feature transformation module. The second multi-level linear transformation is cascaded after the second attention layer to linearly transform the second transformed attention feature vector to obtain a branch feature vector. In this embodiment, the ViT model (the architecture of the backbone network and the branch network) is used to globally reconstruct the CNN-based structure and the conditional control diffusion model, so as to utilize the advantages of the attention module in the ViT model and further improve the reliability of the generated samples.

[0133] 1.5. Training scheme for the conditional control diffusion model.

[0134] 1.5.1. Two-stage training scheme.

[0135] In this embodiment, a two-stage training scheme is adopted. In the first stage, the backbone network is trained first, and in the second stage, the branch network and the multi-level fusion module are trained, so that the model can achieve a stable and effective convergence result.

[0136] In a possible implementation manner, after the feature encoding module and the feature decoding module are trained, the backbone network, the branch network, and the multi-level fusion module (i.e., the conditional diffusion model) are trained according to the following steps:

[0137] Step S401: Obtain a training data set. Each training sample data in the training data set includes: a sample image, and the sample background category, sample target object category, and sample mask map corresponding to the sample image. Specifically, the same as step S201, each training sample data in the training data set used to train the model in this embodiment mainly includes three elements: a sample image (including a sample background and a sample target object), the sample background category, sample target object category, and sample mask map corresponding to the sample image.

[0138] Step S402: Use the training data set to perform the first-stage training on the backbone network.

[0139] Step S403: Use the training data set to perform the second-stage training on the branch network and the multi-level fusion module.

[0140] In a possible implementation manner, step S402, using the training data set to perform the first-stage training on the backbone network, includes:

[0141] Combine the feature encoding module, the backbone network, and the feature decoding module to obtain a first training model.

[0142] Under the condition of freezing the parameters of the feature encoding module and the feature decoding module, the first training model is trained using the training dataset; in each training process of iterative training, the parameters of the backbone network are updated according to the calculated loss function.

[0143] Among them, in each training process, the input of the first training model is the sample image, the sample background category corresponding to the sample image, and the sample mask image in the training dataset; the output of the first training model is the reconstructed sample background image.

[0144] Specifically, when training the backbone network in the first stage, first form the first training model with the feature encoding module, the backbone network, and the feature decoding module, where the feature encoding module and the feature decoding module are trained and completed according to steps S201 - S203. The training process of the first stage is as follows:

[0145] Step S501, input the sample image and the sample mask image into the feature encoding module to obtain the first sample feature map.

[0146] Step S502, input the first sample feature map into the first feature serialization module of the backbone network to obtain the first sample serialized feature vector, and input the sample background category and the random noise reduction steps into the first conditional encoding module of the backbone network to obtain the first sample conditional feature vector.

[0147] Step S503, input the first sample serialized feature vector into the multi - level first feature conversion module of the backbone network for multi - level feature extraction. And input the first sample conditional feature vector into each first feature conversion module respectively. So that the first feature conversion module, based on the first sample conditional feature vector, performs feature extraction on the sample backbone feature vector output by the previous - level first feature conversion module to obtain the sample backbone feature vector of this level. Finally, the multi - level first feature conversion module outputs the first sample feature conversion vector.

[0148] Step S504, input the first sample feature conversion vector into the linear decoding module of the backbone network, so that the linear decoding module decodes the first sample feature conversion vector based on the first sample conditional feature vector to obtain the sample linear decoding result.

[0149] Step S505, input the sample linear decoding result into the feature reconstruction module of the backbone network to obtain the second sample feature map.

[0150] Step S506, input the second sample feature map into the feature decoding module to generate the reconstructed sample background image.

[0151] Step S507: Calculate the loss based on the similarity between the image regions belonging to the background in the input sample image and the reconstructed sample background image, and update the parameters of each module of the backbone network according to the loss.

[0152] Step S508: Repeat the above steps (S501 - S507) until the preset number of training times is reached, or the loss value gradually decreases to a stable value, then end the training to obtain the trained backbone network.

[0153] In the first stage, perform noise learning on the feature map of the noisy background image output by the feature encoding module, aiming to enable the backbone network to have the ability to generate different-step random noises, and in the inference stage, be able to obtain the image corresponding to the category according to the initial noise, denoising steps, and category by gradually denoising the initial noise.

[0154] In a possible implementation manner, the second-stage training of the branch network and the multi-level fusion module using the training dataset includes:

[0155] Combine the feature encoding module, the backbone network, the multi-level fusion module, the branch network, and the feature decoding module to obtain a conditional control diffusion model. The architecture of this conditional control diffusion model is as Figure 2 shown and will not be elaborated here.

[0156] Under the condition of freezing the parameters of the feature encoding module, the feature decoding module, and the backbone network after the first-stage training, use the training dataset to train the conditional control diffusion model; in each training process of the iterative training, update the parameters of the branch network and the multi-level fusion module according to the calculated loss function;

[0157] Among them, in each training process, the input of the conditional control diffusion model is the training sample data; the output of the conditional control diffusion model is the reconstructed sample image.

[0158] Specifically, when training the branch network and the multi-level fusion module (each level of attention fusion module in it) in the second stage, use the conditional control diffusion model (feature encoding module, backbone network, branch network, attention fusion module, feature decoding module), where the feature encoding module and the feature decoding module are trained and completed according to steps S201 - S203, and the backbone network is the backbone network after the first-stage training. The training process of the second stage is as follows:

[0159] Step S601: Input the sample image and the sample mask image into the feature encoding module to obtain the first sample feature map.

[0160] Step S602: Input the first sample feature map into the first feature serialization module of the backbone network to obtain the first sample serialized feature vector. Input the sample background category and the random noise reduction steps into the first conditional encoding module of the backbone network to obtain the first sample conditional feature vector.

[0161] Meanwhile, input the first sample feature map into the second feature serialization module of the branch network to obtain the second sample serialized feature vector. Input the sample target object category and the random noise reduction steps into the second conditional encoding module of the branch network to obtain the second sample conditional feature vector.

[0162] Step S603: Input the first sample serialized feature vector into the multi-level first feature conversion module of the backbone network for multi-level feature extraction to obtain the first sample feature conversion vector. Input the second sample serialized feature vector into the multi-level second feature conversion module of the branch network for multi-level feature extraction to obtain the second sample feature conversion vector. Among them, the sample branch feature vectors extracted at each level of the branch network are fused with the sample backbone feature vectors extracted at the corresponding levels of the backbone network through the attention fusion module at the corresponding levels to obtain the sample fusion feature vector.

[0163] Step S604: Fuse the first sample feature conversion vector and the second sample feature conversion vector through the attention fusion module to obtain the attention fusion feature vector. Input the attention fusion feature vector into the linear decoding module of the backbone network, and enable the linear decoding module to decode the attention fusion feature vector based on the first sample conditional feature vector to obtain the sample linear decoding result.

[0164] Step S605: Input the sample linear decoding result into the feature reconstruction module of the backbone network to obtain the second sample feature map.

[0165] Step S606: Input the second sample feature map into the feature decoding module to generate the reconstructed sample image.

[0166] Step S607: Calculate the loss according to the similarity between the input sample image and the reconstructed sample image, and update the parameters of the branch network and the attention fusion module according to the loss.

[0167] Step S608: Repeat the above steps (S601 - S607) until the preset number of training times is reached, or the loss value gradually decreases to reach a stable value, then end the training to obtain the trained conditional control diffusion model.

[0168] The embodiments of this application can adopt a two-stage training scheme to train the conditional control diffusion model, that is, the second feature conversion module of the branch network and the multi-level attention fusion module are trained simultaneously in the second stage. However, considering that the input of each round of model training is two independent random noises and the output of model training is the fused noise map, the model needs to learn the conditional feature extraction ability and the backbone-conditional feature fusion ability simultaneously, and the training difficulty is relatively high. To reduce the training difficulty, this application also proposes a three-stage training scheme.

[0169] 1.5.2. Three-stage training scheme.

[0170] In this embodiment, a three-stage training scheme is adopted. In the first stage, the backbone network is trained first. In the second stage, the branch network is trained. In the third stage, the multi-level fusion module is trained, so that the model can achieve a stable and effective convergence result. In this three-stage training scheme, the training method of the backbone network in the first stage is the same as that in the above (steps S501-508), and will not be elaborated here.

[0171] In a possible implementation manner, the first-stage training of the backbone network using the training data set includes:

[0172] Combining the feature encoding module, the backbone network and the feature decoding module to obtain a first training model;

[0173] Under the condition of freezing the parameters of the feature encoding module and the feature decoding module, the first training model is trained using the training data set; in each training process of iterative training, the parameters of the backbone network are updated according to the calculated loss function; wherein, in each training process, the input of the first training model is the sample image, the sample background category corresponding to the sample image, and the sample mask map in the training data set; the output of the first training model is the reconstructed sample background image. The training in the first stage is the same as the first stage in 1.5.1 above, and will not be elaborated here.

[0174] Using the training data set to perform second-stage training on the branch network and the multi-level fusion module includes:

[0175] Using the training data set to train the branch network;

[0176] After the training of the branch network is completed, the multi-level fusion module is trained using the training data set.

[0177] In a possible implementation, the backbone network includes: a first conditional encoding module, a first feature serialization module, a multi-level cascaded first feature conversion module, a linear decoding module, and a feature reconstruction module; training the branch network using the training data set includes:

[0178] Combining the feature encoding module, the branch network, the linear decoding module, the feature reconstruction module, and the feature decoding module to obtain a second training model;

[0179] Under the condition of freezing the parameters of the feature encoding module, the feature decoding module, the linear decoding module, and the feature reconstruction module, training the second training model using the training data set; in each training process of iterative training, updating the parameters of the branch network according to the calculated loss function;

[0180] Wherein, in each training process, the input of the second training model is the sample image, the sample target object category corresponding to the sample image, the sample mask map, and random noise in the training data set; the output of the second training model is the reconstructed sample target object image.

[0181] Specifically, when training the branch network in the second stage, a second training model is first constructed (cascaded in the following order: the feature encoding module, the branch network, the linear decoding module, the feature reconstruction module, and the feature decoding module), wherein the feature encoding module and the feature decoding module are trained and completed according to steps S201 - S203, and the backbone network is the backbone network after the first stage of training. The training process of the second stage is as follows:

[0182] Step S701, input the sample image and the sample mask map into the feature encoding module to obtain a first sample feature map.

[0183] Step S702, input the first sample feature map into the second feature serialization module of the branch network to obtain a second sample serialized feature vector, and input the sample target object category and the random noise reduction steps into the second conditional encoding module of the branch network to obtain a second sample conditional feature vector.

[0184] Step S703, input the second sample serialized feature vector into the multi-level second feature conversion module of the branch network for multi-level feature extraction to obtain a second sample feature conversion vector. Among them, each level of the second feature conversion module of the branch network extracts features from the output of the second feature conversion module of the previous level based on the second sample conditional feature vector to obtain a sample branch feature vector. The sample branch feature vector output by the second feature conversion module of the last level is determined as the second sample feature conversion vector.

[0185] Step S704: Input the second sample feature transformation vector into the linear decoding module, and enable the linear decoding module to decode the second sample feature transformation vector based on the second sample conditional feature vector to obtain the sample linear decoding result.

[0186] Step S705: Input the sample linear decoding result into the feature reconstruction module to obtain the second sample feature map.

[0187] Step S706: Input the second sample feature map into the feature decoding module to generate the reconstructed sample target object image.

[0188] Step S707: Calculate the loss according to the similarity between the image region belonging to the target object in the input sample image and the reconstructed sample target object image, and update the parameters of the branch network according to the loss.

[0189] Step S708: Repeat the above steps (S701 - S707) until the preset number of training times is reached, or the loss value gradually decreases to reach a stable value, then end the training to obtain the trained branch network.

[0190] In summary, in the second stage, this embodiment trains some parameters of the second feature transformation module of the branch network of the model. The method is similar to the training in the first stage, but in the training stage, an auxiliary decoding structure needs to be constructed according to the decoding structure of the backbone network to complete the training of some parameters of the branch network.

[0191] It should be noted that in the three - stage training scheme disclosed in this embodiment, the loss calculation in the first and second stage trainings only considers the loss of the corresponding image region: that is, in the first stage training, only the background region of the noise image output by the model is supervised through the mask map; while in the second stage training, only the target object region of the noise image output by the model is supervised through the mask map.

[0192] In a possible implementation manner, training the multi - level fusion module by using the training data set includes:

[0193] Combining the feature encoding module, the backbone network, the multi - level fusion module, the branch network and the feature decoding module to obtain a conditional control diffusion model; this conditional control diffusion model is as Figure 2 shown and will not be elaborated here.

[0194] Under the condition of freezing the parameters of the feature encoding module, the feature decoding module, the backbone network, and the branch network, the conditional control diffusion model is trained using the training dataset; in each training process of iterative training, the parameters of the multi-level fusion module are updated according to the calculated loss function.

[0195] Among them, in each training process, the input of the conditional control diffusion model is the training sample data; the output of the conditional control diffusion model is the reconstructed sample image.

[0196] Specifically, when training all levels of the attention fusion module in the multi-level fusion module in the third stage, the conditional control diffusion model (feature encoding module, backbone network, branch network, attention fusion module, feature decoding module) is used. Among them, the feature encoding module and the feature decoding module are trained according to steps S201 - S203, the backbone network is the backbone network after the first stage of training, and the branch network is the branch network after the second stage of training. The training process in the third stage is the same as steps S601 - S608 above and will not be elaborated here.

[0197] In the third stage, this embodiment trains some parameters of the multi-level fusion module of the model. The method is to freeze the parameters of the backbone network and the branch network (optionally, release the backbone network for synchronous fine-tuning), and use two random noises to train the attention fusion module. The target noise output is: extracting the background area of the random noise through the background mask (i.e., the mask map), and extracting the defect area (i.e., the target object area) of the random noise through the defect mask (i.e., the mask map), and combining them into the target noise for model supervision reference.

[0198] In summary, the embodiment of the present application discloses an image generation method based on a conditional control diffusion model. This method generates images based on a feature-domain diffusion model by specifying categories (background category, target object category) and specifying mask templates (i.e., mask maps). This embodiment is first established based on the feature-domain diffusion model, combines the feature encoding module, the feature decoding module, and the diffusion model (i.e., the feature generation module including the backbone network and the branch network), takes feature consistency as the main goal of generating samples, improves the reliability and diversity of the generated samples, and increases the usability of the augmented data in downstream tasks. Secondly, the diffusion process of the backbone network is controlled by the joint guidance of the conditional image and category prompt words (background category, target object category) to implement a customized image generation model; finally, the conditional control diffusion model based on the CNN structure is reconstructed by the ViT model (the architecture of the backbone network), so as to utilize the advantages of the attention module in the ViT model and improve the reliability and self-learning ability of the generated samples.

[0199] In the second aspect of the embodiments of the present application, a conditional control diffusion model is further provided. The conditional control diffusion model at least includes: a feature encoding module, a backbone network, a branch network, and a feature decoding module; the conditional control diffusion model is used to generate a target image according to the image generation method described in the first aspect of the embodiments of the present application. Specifically, the architecture of this conditional control diffusion model is as Figure 7 shown.

[0200] In the third aspect of the embodiments of the present application, an industrial defect image generation method is further provided. The method includes:

[0201] Obtaining random noise, a background category, a defect target category, and a mask image, where the mask image is used to determine the position of the image background in the image to be generated and the position of the defect target in the image to be generated;

[0202] Based on the feature encoding module, encoding the random noise and the mask image to generate a first feature map;

[0203] Based on the conditional diffusion model, generating a second feature map according to the first feature map, the background category, and the defect target category; wherein, the conditional diffusion model includes a backbone network and a branch network, the backbone network is used to describe the features of the image background, and the branch network is used to describe the features of the defect target;

[0204] Based on the feature decoding module, decoding the second feature map to obtain an industrial defect image, and the industrial defect image includes: defect targets belonging to the defect target category.

[0205] Specifically, this method can be applied to generate industrial defect images, that is, the generated target image can also be a defect image in an industrial background. The background in the defect image is an industrial product (the background category is the product type), and the target object in the defect image is a defect target, that is, defects, flaws, etc. on the surface of the industrial product (the target object category is the type of the defect target). The specific generation method is the same as the image generation method provided in the first aspect of the present application and will not be elaborated here.

[0206] The embodiments of the present application also provide an electronic device. Referring to Figure 8 , Figure 8 is a schematic diagram of the electronic device proposed in the embodiments of the present application. As Figure 4As shown in the figure, the electronic device 100 includes: a memory 110 and a processor 120. The memory 110 and the processor 120 are communicatively connected via a bus. A computer program is stored in the memory 110, and the computer program can run on the processor 120, thereby implementing the steps in the image generation method described in the first aspect or the steps in the industrial defect image generation method described in the second aspect disclosed in the embodiments of the present application.

[0207] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, the steps in the image generation method described in the first aspect or the steps in the industrial defect image generation method described in the second aspect disclosed in the embodiments of the present application are implemented.

[0208] The embodiments of the present application further provide a computer program product. When the computer program product runs on an electronic device, it causes the processor to execute the steps in the image generation method disclosed in the embodiments of the present application. Each embodiment in this specification is described in a progressive manner, and the key point of each embodiment is the difference from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0209] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the said element.

[0210] The above has introduced in detail an image generation method, device and readable storage medium provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

[0211] Other embodiments of the present application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only illustrative, and the true scope and spirit of the present application are pointed out by the following claims.

[0212] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

[0213] As used herein, the terms "one embodiment", "an embodiment", or "one or more embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. In addition, note that the examples of the phrase "in one embodiment" herein do not necessarily all refer to the same embodiment.

[0214] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and technologies are not shown in detail so as not to obscure the understanding of this specification.

[0215] In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.

[0216] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An image generation method, characterized in that: The method comprises: Obtaining random noise, background category, target object category, and a mask map, wherein the mask map is used to determine the position of the image background in the image to be generated and the position of the image target object in the image to be generated; Based on a feature encoding module, encoding the random noise and the mask image to generate a first feature image; Based on a conditional diffusion model, a second feature map is generated according to the first feature map, the background category and the target object category; wherein the conditional diffusion model includes a trunk network and a branch network, the trunk network is used to perform feature description of the image background, and the branch network is used to perform feature description of the image target object; Based on the feature decoding module, the second feature map is decoded to obtain a target image.

2. The image generation method according to claim 1, characterized in that: Before inputting the first feature map and the background category into the feature conversion module, the method further includes: Acquire a training data set, wherein each training sample data in the training data set includes: a sample image, and a sample background category, a sample target object category, and a sample mask image corresponding to the sample image; Constructing an extension module, the extension module is used to perform classification prediction on the feature map output by the feature encoding module, and determine the background category and target object category of the image represented by the feature map output by the feature encoding module; Based on the training data set, the feature encoding module and the feature decoding module are trained, and the loss of each training is calculated using the expansion module.

3. The image generation method according to claim 2, characterized in that: The step of training the feature encoding module and the feature decoding module based on the training data set and calculating the loss of each training using the expansion module includes: Inputting the sample image and the sample mask map into a feature encoding module to be trained to obtain a first sample feature map; Inputting the first sample feature map into the expansion module to obtain a classification result, wherein the classification result represents a background category and a target object category of the image represented by the first sample feature map; Inputting the first sample feature map into a feature decoding module to be trained to obtain a reconstructed sample image; Calculating training loss according to the reconstructed sample image, the classification result, and the sample background category and the sample target object category corresponding to the sample image; According to the training loss, parameters of the feature encoding module to be trained and the feature decoding module to be trained are updated to obtain the trained feature encoding module and feature decoding module.

4. The image generation method according to any one of claims 1 to 3, characterized in that: Based on the conditional diffusion model, generating a second feature map according to the first feature map, the background category and the target object category, including: Input the first feature map and the background category into the main network, input the first feature map and the target object category into the branch network, so that the main network and the branch network respectively perform multi-level feature extraction according to the input, and after the main network and the branch network complete feature extraction once, fuse the main feature vector extracted by the main network, the branch feature vector extracted by the branch network, and the mask map through a multi-level fusion module to obtain a fused feature vector; Sending the fused feature vector to the backbone network for feature extraction at the next level; The second feature map is generated according to the first feature conversion vector obtained by the trunk network through multi-level feature extraction and the second feature conversion vector obtained by the branch network through multi-level feature extraction.

5. The image generation method according to claim 4, characterized in that: The multi-level fusion module includes multiple layers of attention fusion modules, each of which includes: a cascaded first self-attention submodule, a first cross-attention submodule, a second self-attention submodule, and a second cross-attention submodule; The multi-level fusion module is used to fuse the trunk feature vector extracted by the trunk network, the branch feature vector extracted by the branch network, and the mask map to obtain a fused feature vector, including: Obtaining a backbone feature vector extracted by the backbone network, and inputting the backbone feature vector into the first self-attention submodule to obtain a first self-attention feature vector; The mask image sampled below is used as the K value and the V value, and the branch feature vector extracted by the branch network is used as the Q value, and the first self-attention feature vector and the trunk feature vector are input into the first cross-attention submodule together to obtain a first cross-attention feature vector; Inputting the first cross-attention feature vector and the first self-attention feature vector into the second self-attention submodule to obtain a second self-attention feature vector; Taking the branch feature vector extracted from the branch network as K value and V value, taking the second self-attention feature vector as Q value, inputting the second self-attention feature vector and the first cross-attention feature vector into the second cross-attention submodule together, to obtain a second cross-attention feature vector; The second self-attention feature vector is fused with the second cross-attention feature vector to obtain the fused feature vector.

6. The image generation method according to claim 5, characterized in that: The backbone network includes: a first conditional encoding module, a first feature serialization module, and a multi-layer cascaded first feature conversion module; the backbone network performs multi-level feature extraction according to the input, including: Inputting the first feature map into the first feature serialization module to obtain a first serialized feature vector; Inputting the background category and the number of random noise reduction steps into the first conditional encoding module to obtain a first conditional feature vector; The first serialized feature vector is input into the first feature conversion module of the multi-layer cascade, and multi-level feature extraction is performed to obtain the first feature conversion vector; wherein the first feature conversion module of each level, based on the first conditional feature vector, performs feature extraction on the fused feature vector output by the attention fusion module of the previous level, outputs a backbone feature vector, and uses the backbone feature vector output by the first feature conversion module of the last level as the first feature conversion vector.

7. The image generation method according to claim 6, characterized in that: The branch network includes: a second conditional encoding module, a second feature serialization module, and a multi-layer cascaded second feature conversion module; the branch network performs multi-level feature extraction according to the input, including: Inputting the first feature map into the second feature serialization module to obtain a second serialized feature vector; Inputting the target object category and the number of random noise reduction steps into the second conditional encoding module to obtain a second conditional feature vector; The second serialized feature vector is input into the multi-layer cascaded second feature conversion module to perform multi-layer feature extraction to obtain the second feature conversion vector; wherein the second feature conversion module at each level, based on the second conditional feature vector, performs feature extraction on the output of the second feature conversion module at the previous level to obtain the branch feature vector of the second feature conversion module at this level, sends the branch feature vector to the attention fusion module at the corresponding level, and uses the branch feature vector output by the second feature conversion module at the last level as the second feature conversion vector.

8. The image generation method according to claim 7, characterized in that: Each of the second feature conversion modules includes: a cascaded second linear layer, a second attention layer, and a second multi-level linear conversion layer; The step of extracting features from the output of the second feature conversion module at the previous level based on the second conditional feature vector to obtain a branch feature vector of the second feature conversion module at the current level includes: Inputting the second conditional feature vector into the second linear layer to obtain an attention coefficient and a linear coefficient, wherein the attention coefficient includes: an attention weight, an attention bias, and an attention percentage; the linear coefficient includes: a linear weight, a linear bias, and a linear percentage; Inputting the branch feature vector output by the second feature conversion module of the previous level into the second attention layer, so that the second attention layer performs attention conversion based on the attention coefficient to obtain a second converted attention feature vector; The second converted attention feature vector is input into the second multi-level linear transformation layer, so that the second multi-level linear transformation layer performs linear transformation based on the linear coefficient to obtain the branch feature vector of the second feature transformation module at this level.

9. The image generation method according to claim 6, characterized in that: The backbone network also includes: a linear decoding module and a feature reconstruction module; The generating the second feature graph according to the first feature conversion vector obtained by the trunk network through multi-level feature extraction and the second feature conversion vector obtained by the branch network through multi-level feature extraction includes: Inputting the first feature conversion vector and the second feature conversion vector into the attention fusion module of the last level for feature fusion to obtain an attention fusion feature vector; Inputting the attention fusion feature vector and the first conditional feature vector into the linear decoding module for decoding to obtain a linear decoding result; The linear decoding result is input into the feature reconstruction module to obtain the second feature map.

10. The image generation method according to claim 9, characterized in that: Inputting the attention fusion feature vector and the first conditional feature vector into the linear decoding module for decoding to obtain a linear decoding result, including: Encoding the first conditional feature vector through a first linear layer of the linear decoding module to obtain a first linear weight and a first linear bias; The attention fusion feature vector is decoded by a second linear layer of the linear decoding module using the first linear weight and the first linear bias to obtain the linear decoding result.

11. The image generation method according to claim 4, characterized in that: After the feature encoding module and the feature decoding module are trained, the backbone network, the branch network and the multi-level fusion module are trained according to the following steps: Acquire a training data set, wherein each training sample data in the training data set includes: a sample image, and a sample background category, a sample target object category, and a sample mask image corresponding to the sample image; Performing a first phase of training on the backbone network using the training data set; The training data set is used to perform a second stage of training on the branch network and the multi-level fusion module.

12. The image generation method according to claim 11, characterized in that: The first stage of training the backbone network using the training data set includes: Combining the feature encoding module, the backbone network and the feature decoding module to obtain a first training model; Under the condition of freezing the parameters of the feature encoding module and the feature decoding module, the first training model is trained using the training data set; in each training process of iterative training, the parameters of the backbone network are updated according to the calculated loss function; wherein, in each training process, the input of the first training model is the sample image in the training data set, the sample background category corresponding to the sample image, and the sample mask map; the output of the first training model is the reconstructed sample background image; The second stage of training the branch network and the multi-level fusion module using the training data set includes: The feature encoding module, the backbone network, the multi-level fusion module, the branch network and the feature decoding module are combined to obtain a conditional controlled diffusion model; Under the condition of freezing the parameters of the feature encoding module, the feature decoding module and the backbone network after the first stage of training, the conditional controlled diffusion model is trained using the training data set; in each training process of the iterative training, the parameters of the branch network and the multi-level fusion module are updated according to the calculated loss function; wherein, in each training process, the input of the conditional controlled diffusion model is the training sample data; and the output of the conditional controlled diffusion model is the reconstructed sample image.

13. The image generation method according to claim 11, characterized in that: The first stage of training the backbone network using the training data set includes: Combining the feature encoding module, the backbone network and the feature decoding module to obtain a first training model; Under the condition of freezing the parameters of the feature encoding module and the feature decoding module, the first training model is trained using the training data set; in each training process of iterative training, the parameters of the backbone network are updated according to the calculated loss function; wherein, in each training process, the input of the first training model is the sample image in the training data set, the sample background category corresponding to the sample image, and the sample mask map; the output of the first training model is the reconstructed sample background image; Using the training data set, the branch network and the multi-level fusion module are trained in the second stage, including: Using the training data set, training the branch network; After the training of the branch network is completed, the multi-level fusion module is trained using the training data set.

14. The image generation method according to claim 13, characterized in that: The backbone network includes: a first conditional encoding module, a first feature serialization module, a multi-layer cascaded first feature conversion module, a linear decoding module and a feature reconstruction module; the training of the branch network using the training data set includes: Combining the feature encoding module, the branch network, the linear decoding module, the feature reconstruction module and the feature decoding module to obtain a second training model; Under the condition of freezing the parameters of the feature encoding module, the feature decoding module, the linear decoding module, and the feature reconstruction module, the second training model is trained using the training data set; in each training process of iterative training, the parameters of the branch network are updated according to the calculated loss function; Wherein, in each training process, the input of the second training model is the sample image in the training data set, the sample target object category corresponding to the sample image, the sample mask image and random noise; the output of the second training model is the reconstructed sample target object image; The step of training the multi-level fusion module by using the training data set includes: The feature encoding module, the backbone network, the multi-level fusion module, the branch network and the feature decoding module are combined to obtain a conditional controlled diffusion model; Under the condition of freezing the parameters of the feature encoding module, the feature decoding module, the backbone network and the branch network, the conditional controlled diffusion model is trained using the training data set; in each training process of iterative training, the parameters of the multi-level fusion module are updated according to the calculated loss function; In each training process, the input of the conditional controlled diffusion model is the training sample data; and the output of the conditional controlled diffusion model is the reconstructed sample image.

15. A method for generating an industrial defect image, characterized in that: The method comprises: Obtaining random noise, background category, defect target category, and a mask map, wherein the mask map is used to determine the position of the image background in the image to be generated and the position of the defect target in the image to be generated; Based on a feature encoding module, encoding the random noise and the mask image to generate a first feature image; Based on a conditional diffusion model, a second feature map is generated according to the first feature map, the background category and the defect target category; wherein the conditional diffusion model includes a trunk network and a branch network, the trunk network is used to perform feature description of the image background, and the branch network is used to perform feature description of the defect target; Based on the feature decoding module, the second feature map is decoded to obtain an industrial defect image, wherein the industrial defect image includes defect targets belonging to the defect target category.

16. An electronic device, characterized in that: It includes a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the image generation method according to any one of claims 1 to 14 or the steps of the industrial defect image generation method according to claim 15 are implemented.

17. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the image generation method according to any one of claims 1 to 14 or the steps of the industrial defect image generation method according to claim 15 are implemented.