A latent diffusion model training method and device based on text inversion, and a defect image synthesis method and device based on a latent diffusion model

CN122549519BActive Publication Date: 2026-09-11HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611032530.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-11
Estimated Expiration
2046-07-13

AI Technical Summary

Technical Problem

AnoGen等工作将掩膜引导的潜在扩散用于异常生成,但在多类别缺陷统一建模、背景保真、以及正常底图+随机掩膜的大规模配对生成方面仍存在不足:

Benefits of technology

[0034] (1) The embodiments of this application construct shared placeholder tags for multiple defect categories and maintain corresponding embedding vectors for each category in the same potential diffusion model, thereby enabling a single model to learn visual representations of multiple defect types simultaneously without training multiple models for each category, which significantly reduces training costs and deployment complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122549519B_ABST
    Figure CN122549519B_ABST
Patent Text Reader

Abstract

The application provides a latent diffusion model training method and device based on text inversion, and a defect image synthesis method and device based on the latent diffusion model. The model training method comprises: constructing a corresponding text prompt for each training sample; the text prompt comprises a main prompt phrase and a placeholder mark, the placeholder mark is used to update a plurality of embedding vectors corresponding thereto in the training process, the plurality of embedding vectors correspond to different defect categories respectively, and each embedding vector reflects defect appearance characteristics of the corresponding defect category; and the latent diffusion model is trained by using the training sample and the text prompt thereof, in the training process, the text prompt of the current training sample is replaced with a null text prompt at a preset probability, the denoising network learns based on unconditional text characteristics in the iteration process, competition between the text condition and the image / mask condition is alleviated, and defect semantic drift or unintended rewriting of the background area caused by too strong text condition in the reasoning stage is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, specifically to a method and apparatus for training a latent diffusion model based on text inversion, and a method and apparatus for synthesizing defective images based on a latent diffusion model. Background Technology

[0002] Industrial visual anomaly detection typically relies on training with a large number of labeled defect samples. However, defects in real production lines occur infrequently, are of many types, and have complex shapes, leading to a severe imbalance in the training set and insufficient recall and generalization ability of the detection model for rare defects.

[0003] Traditional data augmentation methods, such as rotation, cropping, and color dithering, struggle to generate new defect patterns with realistic texture and semantic consistency. Defect generation methods based on Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs) often suffer from pattern collapse, unnatural boundaries with normal backgrounds, and difficulties in transferring across defect categories.

[0004] In recent years, diffusion models have demonstrated superior performance in image generation and editing. Textual Inversion, by introducing learnable placeholder (also known as token) embeddings into a pre-trained diffusion model, can learn visual representations of specific concepts using a small number of defect samples. Works such as AnoGen have used mask-guided latent diffusion for anomaly generation, but still have shortcomings in unified modeling of multi-class defects, background fidelity, and large-scale pairing generation of normal background images + random masks.

[0005] First, if multiple defect categories are trained separately using multiple models, the cost is high and it is difficult to share representations.

[0006] Secondly, forcibly pasting the background throughout the sampling process will inhibit the formation of details in defective areas;

[0007] Third, when text conditions compete with image / mask conditions, defective semantic drift or background rewriting can easily occur.

[0008] Therefore, there is an urgent need for a defect synthesis technology solution that is suitable for industrial quality inspection, and can be trained and inferred in batches. Summary of the Invention

[0009] In view of this, this application proposes a method and apparatus for training a latent diffusion model based on text inversion, and a method and apparatus for synthesizing defect images based on a latent diffusion model. Specifically, this application is implemented through the following technical solutions:

[0010] According to a first aspect of the embodiments of this specification, a method for training a latent diffusion model based on text inversion is provided. The latent diffusion model includes a variational autoencoder, a denoising network, and a text encoder, wherein the text encoder is configured with an embedding manager. During training, parameters in the variational autoencoder, the denoising network, and the text encoder that are not related to placeholders are frozen; the embedding vectors in the embedding manager corresponding to the placeholders and the parameters in the text encoder related to the placeholders are optimized. The method includes the following steps:

[0011] Step S1: Obtain a training set, which includes training samples of multiple defect categories. Each training sample includes a defect image and a defect mask and defect category identifier corresponding to the defect image. Each defect category is pre-configured with a main cue phrase describing the visual semantics of that type of defect.

[0012] Step S2: For each training sample, a template is randomly selected from a preset natural language template library. The main prompt phrase and placeholder markers corresponding to the training sample are filled into the template to generate the text prompt corresponding to the training sample. The placeholder markers are maintained by the embedding manager and are used to update their corresponding multiple embedding vectors during the training process. The multiple embedding vectors correspond to different defect categories, and each embedding vector reflects the defect appearance features of the corresponding defect category.

[0013] Step S3: Input the training samples and their text prompts into the latent diffusion model for training. During the training process, replace the text prompts of the current training samples with empty text prompts with a preset probability, and set the corresponding text conditions as unconditional text features, so that the denoising network learns based on the unconditional text features in the corresponding iteration process.

[0014] According to a second aspect of the embodiments of this specification, a defect image synthesis method based on a latent diffusion model is provided, wherein the latent diffusion model is trained by the method described in the first aspect, and the defect image synthesis method includes the following steps:

[0015] Step T1: Obtain the defect-free base map and the defect mask corresponding to the target defect category, and obtain the embedding vector corresponding to the target defect category from the trained embedding management. Construct a text prompt based on the main prompt phrase corresponding to the target defect category and the embedding vector.

[0016] Step T2: Input the defect-free base image, the defect mask corresponding to the target defect category, and the text prompt into the trained latent diffusion model to obtain the defect-synthesized image; wherein:

[0017] The encoder of the variational autoencoder encodes the defect-free base map into a latent representation;

[0018] The text encoder encodes the text prompt into text conditions, wherein the text conditions contain the semantic information of the embedding vector;

[0019] The denoising network, under the spatial constraints of the defect mask, iteratively denoises the latent representation according to a preset number of sampling steps, and in each denoising step, it combines the classifier's free guidance and the text conditions to generate a latent representation of the synthesized image. When the current sampling step index is less than the preset background pasting start step index, iterative denoising is performed only within the defect area defined by the defect mask, based on the text conditions. When the current sampling step index is not less than the preset background pasting start step index, the background area of ​​the current latent variable is replaced with the noise component corresponding to the latent representation of the defect-free base map in that sampling step.

[0020] The decoder of the variational autoencoder decodes the latent representation of the synthesized image into the final defective synthesized image.

[0021] According to a third aspect of the embodiments of this specification, a training apparatus for a latent diffusion model based on text inversion is provided. The latent diffusion model includes a variational autoencoder, a denoising network, and a text encoder, wherein the text encoder is configured with an embedding manager. During training, parameters in the variational autoencoder, the denoising network, and the text encoder that are not related to placeholders are frozen, and the embedding vectors in the embedding manager corresponding to the placeholders and the parameters in the text encoder related to the placeholders are optimized. The apparatus includes:

[0022] The training set acquisition unit is used to acquire a training set, which includes training samples of multiple defect categories. Each training sample includes a defect image and a defect mask and a defect category identifier corresponding to the defect image. Each defect category is pre-configured with a main cue phrase describing the visual semantics of that type of defect.

[0023] The prompt construction unit is used to randomly select a template from a preset natural language template library for each training sample, and fill the template with the main prompt phrase corresponding to the training sample and the placeholder mark to generate the text prompt corresponding to the training sample; wherein, the placeholder mark is maintained by the embedding manager and is used to update its corresponding multiple embedding vectors during the training process. The multiple embedding vectors correspond to different defect categories, and each embedding vector reflects the defect appearance features of the corresponding defect category.

[0024] The model training unit is used to input the training samples and their text prompts into the latent diffusion model for training. During the training process, the text prompts of the current training samples are replaced with empty text prompts with a preset probability, and the corresponding text conditions are set as unconditional text features, so that the denoising network learns based on the unconditional text features in the corresponding iteration process.

[0025] According to a fourth aspect of the embodiments of this specification, a defect image synthesis apparatus based on a latent diffusion model is provided, wherein the latent diffusion model is trained by the method described in the first aspect, and the apparatus includes:

[0026] The data acquisition unit is used to acquire a defect-free base map and a defect mask corresponding to the target defect category, and to acquire an embedding vector corresponding to the target defect category from the trained embedding management, and to construct a text prompt based on the main prompt phrase corresponding to the target defect category and the embedding vector.

[0027] The defect synthesis unit is used to input the defect-free base image, the defect mask corresponding to the target defect category, and the text prompt into the trained latent diffusion model to obtain a defect-synthesized image; wherein:

[0028] The encoder of the variational autoencoder encodes the defect-free base map into a latent representation;

[0029] The text encoder encodes the text prompt into text conditions, wherein the text conditions contain the semantic information of the embedding vector;

[0030] The denoising network, under the spatial constraints of the defect mask, iteratively denoises the latent representation according to a preset number of sampling steps, and in each denoising step, it combines the classifier's free guidance and the text conditions to generate a latent representation of the synthesized image. When the current sampling step index is less than the preset background pasting start step index, iterative denoising is performed only within the defect area defined by the defect mask, based on the text conditions. When the current sampling step index is not less than the preset background pasting start step index, the background area of ​​the current latent variable is replaced with the noise component corresponding to the latent representation of the defect-free base map in that sampling step.

[0031] The decoder of the variational autoencoder decodes the latent representation of the synthesized image into the final defective synthesized image.

[0032] According to a fifth aspect of the embodiments of this specification, an electronic device is provided, comprising: a processor; and a computer-readable storage medium storing computer program instructions that, when executed by the processor, cause the processor to perform the method as described in the first aspect.

[0033] The embodiments of this application have at least the following technical effects:

[0034] (1) The embodiments of this application construct shared placeholder tags for multiple defect categories and maintain corresponding embedding vectors for each category in the same potential diffusion model, thereby enabling a single model to learn visual representations of multiple defect types simultaneously without training multiple models for each category, which significantly reduces training costs and deployment complexity.

[0035] (2) During the model training phase, a conditional discarding mechanism is introduced to replace the text prompts with empty strings with a preset probability, so that the denoising network can learn by relying only on images, masks and potential noise trajectories in some iterations, thereby alleviating the competition between text conditions and image / mask conditions and avoiding defective semantic drift or unexpected rewriting of background areas due to overly strong text conditions during the inference phase.

[0036] (3) In the model inference stage, the background pasting strategy is adopted: in the first sampling step, the background area is not forcibly constrained, so that the defect area can obtain sufficient space for morphological formation; in the second sampling step, the background area of ​​the current latent variable is replaced with the latent representation of the defect-free base map corresponding to the noise component in the sampling step, so as to strengthen the consistency between the background and the original base map at the end of the sampling period, and take into account the realism of defect generation and background fidelity.

[0037] In summary, the embodiments of this application provide a scalable and reproducible technical path for the synthesis of defect samples in industrial visual quality inspection, effectively improving the training sample quality and generalization ability of downstream anomaly detection models. Attached Figure Description

[0038] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Some specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings in an exemplary and non-limiting manner. The same reference numerals in the drawings indicate the same or similar parts or components. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:

[0039] Figure 1 This is a schematic flowchart illustrating an exemplary embodiment of the present application of a latent diffusion model training method based on text inversion;

[0040] Figure 2 This is a schematic flowchart illustrating a defect image synthesis method based on a latent diffusion model, as shown in an exemplary embodiment of this application.

[0041] Figure 3 This is a schematic diagram of an original, defect-free base map shown in an exemplary embodiment of this application;

[0042] Figure 4 This application illustrates an exemplary embodiment based on... Figure 3 Synthesized renderings of defects related to missing gold spheres;

[0043] Figure 5 This application illustrates an exemplary embodiment based on... Figure 3 Synthesized renderings of defects related to different types of dirt;

[0044] Figure 6 This application illustrates an exemplary embodiment based on... Figure 3 Synthetic rendering of defects related to gold wire breakage;

[0045] Figure 7 This application illustrates an exemplary embodiment based on... Figure 3 Synthesized rendering of defects related to gold wire solder joint leakage;

[0046] Figure 8 This is a block diagram illustrating an electronic device according to an exemplary embodiment of this application;

[0047] Figure 9 This is a block diagram illustrating a latent diffusion model training device based on text inversion, as shown in an exemplary embodiment of this application.

[0048] Figure 10 This is a block diagram of a defect image synthesis apparatus based on a latent diffusion model, as illustrated in an exemplary embodiment of this application. Detailed Implementation

[0049] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0050] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0051] This application provides a method for training a latent diffusion model based on text inversion. The latent diffusion model (LDM) includes a variational autoencoder (VAE), a denoising network, and a text encoder (Contrastive Language-Image Pre-training, CLIP). The text encoder is configured with an embedding manager.

[0052] A variational autoencoder is used for encoding and decoding between the image space and the latent space. A U-Net network can be used for the denoising network to perform diffusion denoising in the latent space. A text encoder encodes text prompts into text conditions, which serve as the text condition input to the denoising network. The embedding manager can be an auxiliary module at the input of the text encoder, used to map the placeholder markers into trainable embedding vectors. After training, the embedding manager is fed into the text encoder for subsequent encoding.

[0053] During training, the latent diffusion model encodes the defective image into the latent space using a variational autoencoder, adds noise at relevant random time steps according to the forward diffusion process, and then the denoising network predicts the denoised target under textual constraints.

[0054] It is worth noting that during model training, parameters unrelated to placeholders in the variational autoencoder, denoising network, and text encoder are frozen. The embedding vectors corresponding to the placeholders in the embedding manager and the parameters related to the placeholders in the text encoder are optimized. In other words, during training, the main parameters of the denoising network and variational autoencoder are frozen, while most parameters of the text encoder remain unchanged. Only the embedding vectors corresponding to the placeholders and a few related parameters in the embedding manager are optimized to achieve text inversion.

[0055] Figure 1 This is a flowchart illustrating an exemplary embodiment of the latent diffusion model training method based on text inversion, as shown in this application. Figure 1 As shown, the latent diffusion model training method includes the following steps:

[0056] Step S1: Obtain a training set, which includes training samples of multiple defect categories. Each training sample includes a defect image and a defect mask and defect category identifier corresponding to the defect image. Each defect category is pre-configured with a main cue phrase describing the visual semantics of that type of defect.

[0057] This step first constructs a training set for multiple defect categories. For example, an image subdirectory is created for each defect category under the data root directory, along with a corresponding mask subdirectory. The mask subdirectory is named "category_name_mask". For instance, for the "broken gold wire" category, its image subdirectory could be named "broken_wire", and its mask subdirectory "broken_wire_mask". Each training sample consists of a defect image, a corresponding binary or soft-edge mask (hereinafter referred to as the defect mask), and a defect category identifier. In the defect mask, white areas with a pixel value of 1 represent the defect area to be edited, and black areas with a pixel value of 0 represent the normal background to be retained. The defect category identifier can be a category name string or a numeric index, used to index the corresponding main cue phrase and embedding vector during training.

[0058] In some embodiments, the main prompt phrase is generated through a mapping table, which maps defect category identifiers to text describing the visual semantics of that type of defect. Specifically, in this embodiment, the mapping table is pre-configured for each defect category, mapping the category identifier of each defect to a short English phrase describing the basic visual semantics of that type of defect. For example, the "gold ball detached" category can be mapped to "chip pad"; the "gold wire broken" category can be mapped to "broken wire".

[0059] Step S2: For each training sample, a template is randomly selected from a preset natural language template library. The main prompt phrase and placeholder markers corresponding to the training sample are filled into the template to generate the text prompt corresponding to the training sample. The placeholder markers are maintained by the embedding manager and are used to update their corresponding multiple embedding vectors during the training process. The multiple embedding vectors correspond to different defect categories, and each embedding vector reflects the defect appearance features of the corresponding defect category.

[0060] The natural language template library contains natural language templates with various sentence structures, which are randomly selected during training to increase the diversity of text prompts.

[0061] The placeholder marker is a special symbol, such as "*" or other predefined symbols, used to represent a learnable embedding vector during text inversion. This embedding vector can learn fine-grained appearance features of this type of industrial defect from defect samples.

[0062] For each defect category, this embodiment also pre-configures a diverse natural language template library. The natural language template library contains multiple natural language description templates with different sentence structures, such as "a photo of ...", "a close-up photo of ...", "an industrial image of ...", etc., where "..." indicates the position to be filled.

[0063] During the training phase, for each training sample, the corresponding main prompt phrase is first retrieved from the mapping table based on its defect category identifier. This main prompt phrase is then combined with a placeholder marker "*" to form a phrase, such as "chip pad *" or "broken wire *". Next, a template is randomly selected from a diverse natural language template library, and the combined phrase is filled into the "..." position of the template to form the final natural language text prompt for that sample. For example, if the selected template is "a photo of ...", the resulting text prompt is "a photo of chip pad *"; if the selected template is "an industrial image of ...", the resulting text prompt is "an industrial image of broken wire *". By randomly selecting templates, the diversity of text expression can be increased, avoiding the model only memorizing fixed sentence patterns, thereby improving the model's generalization ability under different prompts.

[0064] In practical applications, for industrial datasets with small sample sizes, this step also applies data augmentation operations to the images and masks before training. Specifically, this includes: geometric augmentation, such as random rotation, translation, and horizontal flipping, to ensure pixel-level alignment between the mask and the image; and mild photometric augmentation, such as brightness and contrast perturbations, to improve the model's robustness to defect scale and imaging conditions.

[0065] Step S3: Input the training samples and their text prompts into the latent diffusion model for training. During the training process, replace the text prompts of the current training samples with empty text prompts with a preset probability, and set the corresponding text conditions as unconditional text features, so that the denoising network learns based on the unconditional text features in the corresponding iteration process.

[0066] During model training, the latent diffusion model is loaded first, which mainly consists of three parts:

[0067] Variational autoencoders (VAEs) are used for encoding and decoding between image space and latent space. The encoder of a VAE maps the input image to a low-dimensional latent representation, and the decoder of a VAE reconstructs the latent representation into an image.

[0068] A denoising network is used to perform a diffusion denoising process in a latent space. The denoising network receives noisy latent variables, time steps, and text conditions, and predicts the denoising target, such as noise or the original latent variables.

[0069] The CLIP text encoder encodes text cues into text conditional vectors, which are then fed into various layers of the U-Net via a cross-attention mechanism to guide the denoising process in generating image content that conforms to the semantics of the text.

[0070] In this embodiment, the text encoder is also configured with an embedding manager for maintaining and updating the learnable embedding vectors corresponding to the placeholder tokens. Since the training set contains multiple defect categories, the embedding manager maintains a separate embedding vector for each category, and each embedding vector corresponds to a placeholder token. Although the token symbols are all "*", they are internally distinguished by different token indices.

[0071] During training, a strategy of freezing most parameters and optimizing only the embedding vectors is adopted. Specifically:

[0072] Freeze all parameters of the VAE encoder and decoder.

[0073] Freeze all parameters of the denoising network.

[0074] Freeze all parameters in the CLIP text encoder except those related to placeholder tokens. Typically, the CLIP text encoder contains a token embedding matrix, where each token corresponds to an embedding vector. This step only allows updating the row (or rows) vector in the embedding matrix that corresponds to the placeholder token "*", while the embeddings of the remaining tokens and other encoder weights remain unchanged from their pre-training state.

[0075] Simultaneously optimize the embedding vectors in the embedding manager.

[0076] With the above freezing strategy, the training process only needs to update a very small number of parameters (only the number of categories multiplied by the embedding dimension), which can achieve rapid convergence with a small number of defective samples, while avoiding overfitting.

[0077] It is worth noting that, in order to alleviate interference between different types of text prompts and enhance the controllability of the inference stage under normal background and mask conditions, a conditional discarding mechanism is introduced in step S3 of this embodiment. This mechanism is implemented during the training data reading stage so as not to change the network structure. The specific operation is as follows:

[0078] For each training sample, after constructing the text prompt, the text prompt for that sample is replaced with an empty string or an empty prompt token with a preset probability p_drop (e.g., 0.1 or 0.2). For example, the text prompt "a photo of brokenwire" has a certain probability of being replaced with an empty string. The replaced empty string is also processed by the text encoder to obtain unconditional text features. If it is not replaced, the original text prompt is retained and encoded as conditional text features. During training, the text conditions received by the denoising network may be unconditional or conditional text features, depending on the result of the random replacement.

[0079] Thus, in some iterations, the denoising network only receives empty text conditions. At this time, the model is forced to rely more on defect images, masks, and potential noise trajectories for learning. This helps the model generate reasonable defect content even when the text prompt does not perfectly match the mask area during the inference stage.

[0080] Furthermore, the classifier free guidance (CFG) technique of the latent diffusion model in this embodiment requires simultaneous training of both conditional and unconditional models (or the same model supporting both conditional and unconditional inputs). By discarding conditions, the latent diffusion model in this embodiment naturally learns the ability to generate unconditionally during training, laying the foundation for CFG sampling in the inference phase.

[0081] In some embodiments, step S3 further includes training the potential diffusion model using a preset loss function, wherein the preset loss function includes a diffusion reconstruction loss, the expression of which is as follows:

[0082] ;

[0083] in, For the spread and reconstruction losses; The denoising network; The network parameters of the denoising network; The defect images in the training samples are encoded into original latent variables by the variational autoencoder. Then, following the forward diffusion process at random time steps The latent noise variables obtained by adding noise; The text conditions are obtained by encoding the text prompts corresponding to the training samples using the text encoder. The target for denoising the diffusion process; This is a downsampled mask weight map of the defect mask in the training samples, used to ensure that the loss calculation only applies to the defect region, while the loss for the background region is set to zero.

[0084] To force the model to focus its learning on defect regions and avoid background noise interference, this embodiment introduces a diffusion reconstruction loss. First, the defect mask is downsampled to the same spatial resolution as the latent space using nearest neighbor or bilinear interpolation, resulting in a mask weight map. This mask weight map retains its binary property, with a value of 1 for defect regions and 0 for background regions. Then, the pixel-wise mean squared error between the predicted noise and the actual noise is calculated and multiplied element-wise with the mask weight map, with backpropagation performed only on the loss within the defect regions. This mask weighting strategy ensures that the gradient primarily affects defect semantic learning, preventing the model from generating meaningless updates in background regions.

[0085] To further enhance the background fidelity of the model, in some embodiments, the preset loss function further includes a background fidelity loss, which is calculated through the following steps:

[0086] The reference latent variables obtained by encoding the defect images in the training samples using the variational autoencoder, and the original latent variables predicted by the denoising network are also obtained.

[0087] Within the background region indicated by the defect mask in the training samples, the weighted mean square error between the original latent variable and the reference latent variable is used as the background fidelity loss; the background region is the region outside the defect region in the defect mask, and the background fidelity loss is used to constrain the denoising network to maintain a background texture consistent with the defect image in the background region.

[0088] This application also provides a defect image synthesis method based on a latent diffusion model, wherein the latent diffusion model is trained through the foregoing embodiments, such as... Figure 2 As shown, the defect image synthesis method includes the following steps:

[0089] Step T1: Obtain the defect-free base map and the defect mask corresponding to the target defect category, and obtain the embedding vector corresponding to the target defect category from the trained embedding management. Construct a text prompt based on the main prompt phrase corresponding to the target defect category and the embedding vector.

[0090] In practical applications, a defect-free base image I_good can be randomly selected from the normal sample library. This base image is an image of a defect-free industrial product, such as a chip gold wire package image. The size and format of the base image should be consistent with those used in the training phase.

[0091] In obtaining a defect-free base map Next, the target defect category is determined. The target defect category refers to the type of defect that this inference aims to generate, such as "missing gold ball," "dirt," "broken gold wire," or "missing gold wire solder joint." Based on the target defect category, a defect mask is randomly selected from the mask library for that category. The mask is a binary image where white areas with a pixel value of 1 represent defective areas to be edited, and black areas with a pixel value of 0 represent normal background that should be retained. The shape, size, and position of the mask can be randomly selected from pre-collected mask samples as needed to increase the diversity of generated samples.

[0092] Obtain the embedding vector corresponding to the target defect category from the trained embedding manager. Specifically, in Figure 1 During the text inversion fine-tuning process in step S3, the embedding manager stores a learnable embedding vector for each defect category, which is bound to the placeholder marker "*". During inference, the corresponding embedding vector is read from the embedding manager based on the identifier of the target defect category (such as the category name "broken_wire"). This embedding vector carries the fine-grained visual appearance features of that type of defect.

[0093] Next, text prompts are constructed based on the target defect category. Each text prompt consists of a main prompt phrase and placeholder tokens. The main prompt phrase is a short English phrase that describes the basic visual semantics of the defect category and is pre-stored in a mapping table. For example, for the "broken gold wire" category, the main prompt phrase is "broken_wire"; for the "gold ball detached" category, the main prompt phrase is "chip pad". The main prompt phrase is concatenated with the embedding vector (also known as the text token) corresponding to the placeholder token "*" to form the text prompt string. This string can be further filled into a natural language template to form a more complete sentence, such as "a photo of broken wire *", but the simplest form is sufficient to guide the generation.

[0094] Finally, the defect-free base map Defect mask The constructed text prompts serve as input to the model during the inference phase. Additionally, relevant hyperparameters can be preset, such as the total number of sampling steps S, the classifier's free-guide scale s, the initial background pasting ratio r (e.g., 0.6), and a random seed, for use in subsequent sampling.

[0095] Step T2: Input the defect-free base image, the defect mask corresponding to the target defect category, and the text prompt into the trained latent diffusion model to obtain the defect-synthesized image; wherein:

[0096] The encoder of the variational autoencoder encodes the defect-free base map into a latent representation;

[0097] The text encoder encodes the text prompt into text conditions, wherein the text conditions contain the semantic information of the embedding vector;

[0098] The denoising network, under the spatial constraints of the defect mask, iteratively denoises the latent representation according to a preset number of sampling steps, and in each denoising step, it combines the classifier's free guidance and the text conditions to generate a latent representation of the synthesized image. When the current sampling step index is less than the preset background pasting start step index, iterative denoising is performed only within the defect area defined by the defect mask, based on the text conditions. When the current sampling step index is not less than the preset background pasting start step index, the background area of ​​the current latent variable is replaced with the noise component corresponding to the latent representation of the defect-free base map in that sampling step.

[0099] The decoder of the variational autoencoder decodes the latent representation of the synthesized image into the final defective synthesized image.

[0100] Specifically, after receiving the input data, the potential diffusion model first generates a defect-free base map. The input is fed into the encoder of the variational autoencoder to obtain its representation x0 in the latent space. x0 is a low-dimensional feature tensor with a spatial resolution lower than that of the original image.

[0101] At the same time, for defect masks Preprocessing: Downsample it to the same spatial resolution as x0 to obtain a binary mask. The background area has a value of 1, and the editable defect area has a value of 0. Further define the edit mask. ,Right now The value of the defect area is 1, and the value of the background area is 0.

[0102] The text prompt string is input into the text encoder. Before encoding, the string is first tokenized to obtain a token sequence. Each placeholder token corresponds to a special token. The pre-trained embedding vector of this token is replaced with the target category embedding vector read from the embedding manager, while the embeddings of the other tokens remain unchanged. Then, the replaced token embedding sequence is fed into the Transformer network of the text encoder, outputting a text conditional vector c. This vector contains semantic information about the target defect category and fine-grained appearance features learned through text inversion.

[0103] Next, a deterministic sampler using Denoising Diffusion Implicit Models (DDIM) is employed for iterative denoising. The total number of sampling steps is preset to S (e.g., S = 50 or 100 steps), and the initial background pasting ratio r is set (0 ≤ r < 1, typical value 0.6). The starting step index is calculated. , This is the floor function. Additionally, a classifier-free guidance (CFG) scale s is set, typically s ≥ 1, for example, s = 2.0, to balance the directions of conditional and unconditional generation.

[0104] Initialize the current latent variables The noise is standard Gaussian noise, and then denoising is performed stepwise from t=T to 1, with the following operations performed at each step:

[0105] 1) Prediction noise: This refers to the noise generated by the current latent variables. Inputting time step t and text condition c into the denoising network yields conditional noise predictions. Simultaneously, replacing the text condition with an empty string yields unconditional noise predictions, and the final noise prediction is calculated using the classifier's free-guided formula.

[0106] 2) DDIM Sampling: Based on the DDIM update formula, sampling is performed using the current latent variables. The final noise prediction calculates the latent variables corresponding to time step t-1. .

[0107] 3) Pasting background at the end: Determine the index of the current step i = T - t + 1 (i.e., the number of steps completed so far, counting from 0) and... The relationship.

[0108] if i< : Do not perform background replacement operation; the current step will be processed using traditional image restoration methods.

[0109] If i≥ : Perform background replacement. Specifically, calculate the defect-free basemap reference latent variables for the current step. , It is the forward noise function of the diffusion process, representing the latent variable from noise addition at x0 to time step t. Then, the current latent variable... Force the background area to be replaced with The corresponding background portion retains the defect area unchanged. In this way, through this operation, the background is gradually pulled back to the trajectory consistent with the original base image x0 in the later stage of sampling, ensuring the high fidelity of the background of the final generated image.

[0110] After background replacement, this embodiment further performs adaptive instance normalization (AdaIN) alignment on the defect area to smooth the boundary between the defect and the background.

[0111] Specifically, after replacing the background region of the current latent variable with the latent representation of the defect-free base map corresponding to the noise component in the sampling step, the denoising network further aligns the latent features within the defect region defined by the defect mask in the current latent variable. The aligned latent features... The expression is as follows:

[0112] ;

[0113] in, The latent variable after background region replacement for the current latent variable, The potential representation of the defect-free base map is the noise component corresponding to this sampling step. The preset alignment weight coefficient, This is an edit mask that is downsampled from the defect mask. To Zhongyou Calculated by channel for the specified area and The mean and variance are calculated and then subjected to adaptive instance normalization operation of affine transformation.

[0114] This embodiment effectively mitigates abrupt changes in tone and contrast between newly generated defects and the original background through the alignment step, thereby enhancing the industrial realism of the synthesized image.

[0115] In some embodiments, the denoising network includes a cross-attention layer, and the denoising network performs the following attention enhancement operation within a preset sampling step interval, such as a step ratio interval [0.3S, 0.8S]:

[0116] The edit mask is downsampled to the same spatial resolution as the attention map of the cross-attention layer to obtain spatial gating;

[0117] Based on the spatial gating, the positions corresponding to the defect semantics in the attention weights of the cross-attention layer are enhanced to obtain the enhanced attention weights. The positions corresponding to the defect semantics are determined based on the token index corresponding to the main prompt phrase and the token index corresponding to the placeholder mark in the text prompt.

[0118] Among them, the enhanced attention weight The expression is as follows:

[0119] ;

[0120] in, The preset gain coefficient, For spatial gating, These are the original attention weights of the cross-attention layer.

[0121] This embodiment uses this enhancement to make the model pay more attention to the semantics of defects in the text prompts when generating defects, thereby improving the accuracy of the generated defect types.

[0122] After completing all sampling steps, the latent representation of the synthesized image is obtained. ,Will The input to the VAE decoder yields the final synthesized defect image. This synthesized defect image has the same characteristics as the defect-free base map. With the same size and background texture, realistic defects that conform to the semantics of the target defect category are generated within the specified defect mask area.

[0123] In some embodiments, to facilitate subsequent downstream tasks, such as training anomaly detection models, the generated defect synthesis image, the base map path used, the mask path, the target defect category, the text prompt, the random seed, and all hyperparameters are recorded for each instance. A hierarchical balanced scheduling strategy can be adopted to ensure that the number of generated defects for each category is approximately balanced. Synthetic samples can be packaged into a directory structure of "normal / defect / mask" for training industrial anomaly detection models (such as DRAEM).

[0124] Figures 3 to 7 This demonstrates the synthesis effect of this embodiment on the gold thread dataset. Figure 3 This is the original, defect-free base map. Figures 4 to 7 Four types of defects were synthesized: missing gold balls, dirt, broken gold wires, and incomplete gold wire soldering. The labeled areas are the synthesized defect regions. The results show that this embodiment can generate realistic and diverse defect morphologies without changing the background.

[0125] Figure 8 This is a schematic diagram of an electronic device illustrated in this specification according to an exemplary embodiment. Please refer to... Figure 5 At the hardware level, the device includes a processor 810, an internal bus 820, a network interface 830, memory 840, a hardware acceleration device 850, and non-volatile memory 860, and may also include other hardware required for its functions. One or more embodiments of this application can be implemented in software, for example, the processor 810 reads the corresponding computer program from the non-volatile memory 860 into the memory 840 and then runs it. Of course, in addition to software implementation, one or more embodiments of this application do not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the above processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0126] Figure 9 This is a structural block diagram illustrating an exemplary embodiment of a defect image synthesis model training device based on a latent diffusion model and text inversion. The defect image synthesis model training device can be applied to, for example... Figure 8 The electronic device shown implements the technical solution of this application. The defect image synthesis model includes a variational autoencoder, a denoising network, and a text encoder. The text encoder is configured with an embedding manager. During training, parameters in the variational autoencoder, the denoising network, and the text encoder that are not related to placeholder markers are frozen. The embedding vectors in the embedding manager corresponding to the placeholder markers, and the parameters in the text encoder related to the placeholder markers, are optimized.

[0127] The defect image synthesis model training device includes: a training set acquisition unit 110, a consensus verification unit 620, a prompt construction unit 120, and a model training unit 130, wherein:

[0128] The training set acquisition unit 110 is used to acquire a training set, which includes training samples of multiple defect categories. Each training sample includes a defect image and a defect mask and a defect category identifier corresponding to the defect image. Each defect category is pre-configured with a main cue phrase describing the visual semantics of that type of defect.

[0129] The prompt construction unit 120 is used to randomly select a template from a preset natural language template library for each training sample, and fill the template with the main prompt phrase corresponding to the training sample and the placeholder mark to generate the text prompt corresponding to the training sample; wherein, the placeholder mark is maintained by the embedding manager and is used to update its corresponding multiple embedding vectors during the training process. The multiple embedding vectors correspond to different defect categories, and each embedding vector reflects the defect appearance features of the corresponding defect category.

[0130] The model training unit 130 is used to input the training samples and their text prompts into the latent diffusion model for training. During the training process, the text prompts of the current training samples are replaced with empty text prompts with a preset probability, and the corresponding text conditions are set as unconditional text features, so that the denoising network learns based on the unconditional text features in the corresponding iteration process.

[0131] Figure 10 This is a structural block diagram of a defect image synthesis apparatus based on a latent diffusion model, as illustrated in an exemplary embodiment of this application. The defect image synthesis apparatus can be applied to, for example... Figure 8The electronic device shown implements the technical solution of this application. The defect image synthesis device includes: a data acquisition unit 210 and a defect synthesis unit 220; wherein:

[0132] The data acquisition unit 210 is used to acquire a defect-free base map and a defect mask corresponding to the target defect category, and to acquire an embedding vector corresponding to the target defect category from the trained embedding management, and to construct a text prompt based on the main prompt phrase corresponding to the target defect category and the embedding vector.

[0133] Defect synthesis unit 220 is used to input the defect-free base image, the defect mask corresponding to the target defect category, and the text prompt into the trained latent diffusion model to obtain a defect-synthesized image; wherein:

[0134] The encoder of the variational autoencoder encodes the defect-free base map into a latent representation;

[0135] The text encoder encodes the text prompt into text conditions, wherein the text conditions contain the semantic information of the embedding vector;

[0136] The denoising network, under the spatial constraints of the defect mask, iteratively denoises the latent representation according to a preset number of sampling steps, and in each denoising step, it combines the classifier's free guidance and the text conditions to generate a latent representation of the synthesized image. When the current sampling step index is less than the preset background pasting start step index, iterative denoising is performed only within the defect area defined by the defect mask, based on the text conditions. When the current sampling step index is not less than the preset background pasting start step index, the background area of ​​the current latent variable is replaced with the noise component corresponding to the latent representation of the defect-free base map in that sampling step.

[0137] The decoder of the variational autoencoder decodes the latent representation of the synthesized image into the final defective synthesized image.

[0138] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0139] Accordingly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any of the above embodiments.

[0140] Accordingly, embodiments of this application also provide a computer program product configured to perform the methods described in any of the above embodiments.

[0141] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which can take the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.

[0142] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0143] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0144] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random-access memory (SRAM), dynamic random-access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0145] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.

[0146] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0147] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0148] It should be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0149] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for training a latent diffusion model based on text inversion, characterized in that, The latent diffusion model includes a variational autoencoder, a denoising network, and a text encoder, wherein the text encoder is configured with an embedding manager. During training, parameters in the variational autoencoder, the denoising network, and the text encoder that are not related to the placeholders are frozen, and the embedding vectors in the embedding manager corresponding to the placeholders and the parameters in the text encoder related to the placeholders are optimized. The method includes the following steps: Step S1: Obtain a training set, which includes training samples of multiple defect categories. Each training sample includes a defect image and a defect mask and defect category identifier corresponding to the defect image. Each defect category is pre-configured with a main cue phrase describing the visual semantics of that type of defect. Step S2: For each training sample, a template is randomly selected from a preset natural language template library. The main prompt phrase and placeholder markers corresponding to the training sample are filled into the template to generate the text prompt corresponding to the training sample. The placeholder markers are maintained by the embedding manager and are used to update their corresponding multiple embedding vectors during the training process. The multiple embedding vectors correspond to different defect categories, and each embedding vector reflects the defect appearance features of the corresponding defect category. Step S3: Input the training samples and their text prompts into the latent diffusion model for training. During the training process, replace the text prompts of the current training samples with empty text prompts with a preset probability, and set the corresponding text conditions as unconditional text features, so that the denoising network can learn based on the unconditional text features in the corresponding iteration process. Step S3 further includes training the potential diffusion model using a preset loss function, wherein the preset loss function includes a diffusion reconstruction loss, the expression of which is as follows: ; in, For the spread and reconstruction losses; The denoising network; The network parameters of the denoising network; The defect images in the training samples are encoded into original latent variables by the variational autoencoder. Then, following the forward diffusion process at random time steps The latent noise variables obtained by adding noise; The text conditions are obtained by encoding the text prompts corresponding to the training samples using the text encoder. The target for denoising the diffusion process; This is a downsampled mask weight map of the defect mask in the training samples, used to ensure that the loss calculation only applies to the defect region, while the loss for the background region is set to zero.

2. The method according to claim 1, characterized in that, The preset loss function also includes a background fidelity loss, which is calculated through the following steps: The reference latent variables obtained by encoding the defect images in the training samples using the variational autoencoder, and the original latent variables predicted by the denoising network are also obtained. Within the background region indicated by the defect mask in the training samples, the weighted mean square error between the original latent variable and the reference latent variable is used as the background fidelity loss; the background region is the region outside the defect region in the defect mask, and the background fidelity loss is used to constrain the denoising network to maintain a background texture consistent with the defect image in the background region.

3. The method according to claim 1, characterized in that: The main prompt phrase in step S1 is generated through a mapping table, which is used to map the defect category identifier to text describing the visual semantics of that type of defect. In step S2, the natural language template library contains natural language templates with various sentence structures, which are randomly selected during training to increase the diversity of text prompts.

4. A defect image synthesis method based on a latent diffusion model, characterized in that, The latent diffusion model is trained by the method described in any one of claims 1 to 3, and the defect image synthesis method includes the following steps: Step T1: Obtain the defect-free base map and the defect mask corresponding to the target defect category, and obtain the embedding vector corresponding to the target defect category from the trained embedding management. Construct a text prompt based on the main prompt phrase corresponding to the target defect category and the embedding vector. Step T2: Input the defect-free base image, the defect mask corresponding to the target defect category, and the text prompt into the trained latent diffusion model to obtain the defect-synthesized image; wherein: The encoder of the variational autoencoder encodes the defect-free base map into a latent representation; The text encoder encodes the text prompt into text conditions, wherein the text conditions contain the semantic information of the embedding vector; The denoising network, under the spatial constraints of the defect mask, iteratively denoises the latent representation according to a preset number of sampling steps, and in each denoising step, it combines the classifier's free guidance and the text conditions to generate a latent representation of the synthesized image. When the current sampling step index is less than the preset background pasting start step index, iterative denoising is performed only within the defect area defined by the defect mask, based on the text conditions. When the current sampling step index is not less than the preset background pasting start step index, the background area of ​​the current latent variable is replaced with the noise component corresponding to the latent representation of the defect-free base map in that sampling step. The decoder of the variational autoencoder decodes the latent representation of the synthesized image into the final defective synthesized image.

5. The method according to claim 4, characterized in that, After replacing the background region of the current latent variable with the latent representation of the defect-free basemap corresponding to the noise component in that sampling step, the denoising network further aligns the latent features within the defect region defined by the defect mask in the current latent variable. The expression is as follows: ; in, The latent variable after background region replacement for the current latent variable, The potential representation of the defect-free base map is the noise component corresponding to this sampling step. The preset alignment weight coefficient, This is an edit mask that is downsampled from the defect mask. To Zhongyou Calculated by channel for the specified area and The mean and variance are calculated and then subjected to adaptive instance normalization operation of affine transformation.

6. The method according to claim 5, characterized in that, The denoising network includes a cross-attention layer, and performs the following attention enhancement operation within a preset sampling step interval: The edit mask is downsampled to the same spatial resolution as the attention map of the cross-attention layer to obtain spatial gating; Based on the spatial gating, the positions corresponding to the defect semantics in the attention weights of the cross-attention layer are enhanced to obtain the enhanced attention weights. The positions corresponding to the defect semantics are determined based on the token index corresponding to the main prompt phrase and the token index corresponding to the placeholder mark in the text prompt.

7. A training device for a defect image synthesis model based on a latent diffusion model and text inversion, characterized in that, The defect image synthesis model includes a variational autoencoder, a denoising network, and a text encoder, wherein the text encoder is configured with an embedding manager. During training, parameters in the variational autoencoder, the denoising network, and the text encoder that are not related to the placeholders are frozen, and the embedding vectors in the embedding manager corresponding to the placeholders and the parameters in the text encoder related to the placeholders are optimized. The device includes: The training set acquisition unit is used to acquire a training set, which includes training samples of multiple defect categories. Each training sample includes a defect image and a defect mask and a defect category identifier corresponding to the defect image. Each defect category is pre-configured with a main cue phrase describing the visual semantics of that type of defect. The prompt construction unit is used to randomly select a template from a preset natural language template library for each training sample, and fill the template with the main prompt phrase corresponding to the training sample and the placeholder mark to generate the text prompt corresponding to the training sample; wherein, the placeholder mark is maintained by the embedding manager and is used to update its corresponding multiple embedding vectors during the training process. The multiple embedding vectors correspond to different defect categories, and each embedding vector reflects the defect appearance features of the corresponding defect category. The model training unit is used to input the training samples and their text prompts into the latent diffusion model for training. During the training process, the text prompts of the current training samples are replaced with empty text prompts with a preset probability, and the corresponding text conditions are set as unconditional text features, so that the denoising network learns based on the unconditional text features in the corresponding iterations. The model training unit uses a preset loss function to train the latent diffusion model. The preset loss function includes a diffusion reconstruction loss, the expression of which is as follows: ; in, For the spread and reconstruction losses; The denoising network; The network parameters of the denoising network; The defect images in the training samples are encoded into original latent variables by the variational autoencoder. Then, following the forward diffusion process at random time steps The latent noise variables obtained by adding noise; The text conditions are obtained by encoding the text prompts corresponding to the training samples using the text encoder. The target for denoising the diffusion process; This is a downsampled mask weight map of the defect mask in the training samples, used to ensure that the loss calculation only applies to the defect region, while the loss for the background region is set to zero.

8. A defect image synthesis device based on a latent diffusion model and text inversion, characterized in that, The potential diffusion model is trained by the method according to any one of claims 1 to 3, and the apparatus comprises: The data acquisition unit is used to acquire a defect-free base map and a defect mask corresponding to the target defect category, and to acquire an embedding vector corresponding to the target defect category from the trained embedding management, and to construct a text prompt based on the main prompt phrase corresponding to the target defect category and the embedding vector. The defect synthesis unit is used to input the defect-free base image, the defect mask corresponding to the target defect category, and the text prompt into the trained latent diffusion model to obtain a defect-synthesized image; wherein: The encoder of the variational autoencoder encodes the defect-free base map into a latent representation; The text encoder encodes the text prompt into text conditions, wherein the text conditions contain the semantic information of the embedding vector; The denoising network, under the spatial constraints of the defect mask, iteratively denoises the latent representation according to a preset number of sampling steps, and in each denoising step, it combines the classifier's free guidance and the text conditions to generate a latent representation of the synthesized image. When the current sampling step index is less than the preset background pasting start step index, iterative denoising is performed only within the defect area defined by the defect mask, based on the text conditions. When the current sampling step index is not less than the preset background pasting start step index, the background area of ​​the current latent variable is replaced with the noise component corresponding to the latent representation of the defect-free base map in that sampling step. The decoder of the variational autoencoder decodes the latent representation of the synthesized image into the final defective synthesized image.

9. An electronic device, characterized in that, include: processor; as well as A computer-readable storage medium storing computer program instructions that, when executed by the processor, cause the processor to perform the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image editing method and system based on diffusion model inversion and attention optimization

    CN120635475A

  • Style migration method and system based on text inversion and self-attention injection

    CN121033214A