Multi-granularity prompting metal surface defect image synthesis method based on pre-training diffusion model

Through a multi-grained cues based on the pre-trained diffusion model, metal surface defect images with controllable location, size and categories are generated, which solves the problems of insufficient data and insufficient representativeness of rare categories in defect detection, and improves the generalization ability of the detection system and the diversity of generated samples.

CN120451071AActive Publication Date: 2025-08-08AUTOMATION RES & DESIGN INST OF METALLURGICAL IND

Patent Information

Application Number
CN202510516518.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-08
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing deep learning models lack sufficient data sets in metal surface defect detection, especially rare anomaly samples account for a low proportion in the data set, and traditional generative models cannot effectively control the location, size and category of defects, resulting in inaccurate and poor consistency of detection results.

Method used

A multi-grained cues based on pre-trained diffusion model is adopted to construct a multi-grained training set through multi-precision annotation, an accuracy encoding mask pyramid and CLIP semantic encoder are introduced, and combined with a denoised UNet network, defect images with controllable location, size and category are generated to achieve multi-level controllable generation.

Benefits of technology

It realizes flexible generation from macro layout to micro details, breaks through the dependence on pixel-level annotation, improves the generalization ability of the detection system and the diversity of generated samples, solves the representative problems of rare categories in the data set, and reduces the consumption of training resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451071A_ABST
    Figure CN120451071A_ABST
Patent Text Reader

Abstract

The invention discloses a metal surface defect image controllable synthesis method based on a pre-training diffusion model, belongs to the technical field of metal surface defect data enhancement, and is a multi-granularity prompt generation scheme aiming at the problems of insufficient defect data and high labeling cost in the prior art. The method comprises the following steps: performing multi-precision labeling on a metal surface defect image, and constructing a multi-granularity training set; the method comprises the following steps: designing a precision coding mask pyramid based on a pre-trained Stable Diffusion model, generating a hierarchical control signal through multi-stage downsampling, fusing category and position information by combining CLIP semantic coding, and performing hierarchical injection to denoise UNet so as to control defect generation; a VAE module and a CLIP module are frozen by adopting a transfer learning strategy, and only a denoising network and a mask coding layer are finely tuned; and a user generates various defect images through the coarse-grained mask and category prompt. The method breaks through the dependence of a traditional generation model on fine labeling, realizes flexible control of defect positions, shapes and categories, and improves the generalization ability of a detection system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of metal surface defect data enhancement, and specifically relates to a method and system for generating a multi-granularity prompt defect dataset based on a pre-trained diffusion model, which is used to solve the problem of insufficient data for surface defect detection tasks in industrial production. Background Art

[0002] In recent years, the emergence of deep learning algorithms such as convolutional neural networks and Transformers has greatly promoted the development of computer vision. Traditional visual inspection is inefficient and susceptible to human factors such as inspector fatigue, emotion, and experience differences, making it difficult to ensure the accuracy and consistency of inspection results. In contrast, deep learning-based visual algorithms can achieve efficient automated detection of surface defects, greatly reducing human intervention. Therefore, deep learning-based visual algorithms have great application potential in surface defect detection tasks in industrial production practices.

[0003] However, the significant results achieved by common visual algorithms such as Yolo and DETR require large datasets with excellent data distribution and high-quality annotations. However, collecting and annotating data in practice presents many challenges. For example, metal surface defect detection is difficult to collect due to the high yield of production lines. Furthermore, different defect categories have varying frequencies in actual production, and some rare, abnormal samples only account for a very small proportion of the dataset. Furthermore, a significant portion of industrial production data requires experienced frontline workers or engineers to annotate, consuming significant manpower and resources, making it difficult to meet the requirements for building powerful deep learning models in real-world applications.

[0004] Generative models such as Generative Adversarial Networks (GANs) and Denoising Diffusion Partitioning (DDPM) have made significant progress and have been applied to tasks such as style transfer, image editing, and text-prompted image generation. To alleviate the problem of insufficient dataset size, some researchers have proposed using generative models to construct virtual datasets for data augmentation. For surface defect data augmentation tasks, most studies use Generative Adversarial Networks (GANs) and their variants. However, these GAN-based defect data augmentation solutions still have the following issues:

[0005] (1) Industrial image defect generation methods with controllable defect areas can only use precise fine-grained pixel-level masks as control quantities. However, fine-grained masks are not easy to draw. Defects generated using only a small number of fine-grained masks will have similar edges and shapes, affecting data distribution.

[0006] (2) Surface defects have distinct image features depending on their category. However, current defect generation methods do not have the ability to generate defects of a specified category based on prompts, and therefore cannot solve the problem of rare abnormal samples accounting for a low proportion in the dataset. This means that subsequent detection tasks still face the problem of small-sample learning for rare categories. Summary of the Invention

[0007] To address the above problems, this study proposes a multi-granularity hint defect generation method based on a pre-trained diffusion model, which is used to generate defect images with controllable position, size, and category using arbitrary granularity masks without having to provide precise pixel-level hints.

[0008] The present invention provides a multi-granularity metal surface defect image synthesis method based on a pre-trained diffusion model, and the steps are as follows:

[0009] Step 1: Perform multi-precision annotation on the collected steel surface defect images to generate high-precision pixel-level masks, medium-precision contour masks, and low-precision rectangular frame masks. Preprocess the annotated data to construct a multi-granularity training dataset.

[0010] Step 2: Based on the pre-trained Stable Diffusion model, the precision coding mask pyramid module is introduced to perform multi-level downsampling of the user-provided binary mask to generate hierarchical control signals. The CLIP semantic encoder is combined to perform feature fusion on the category or text prompts, and the denoising UNet network is injected layer by layer to build a multi-granularity prompt defect synthesis model.

[0011] Step 3: Freeze the VAE module and CLIP module parameters of the pre-trained model, use the multi-granularity training dataset to fine-tune the denoising UNet network and mask encoding layer, and optimize the noise prediction loss function.

[0012] Step 4: The user inputs a custom mask and category prompt, and the synthetic model generates a controllable defect image to construct a balanced dataset.

[0013] The advantages of the present invention are:

[0014] 1. The present invention's multi-granularity metal surface defect image synthesis method achieves multi-level controllable generation from macroscopic layout to microscopic details by constructing a dynamic precision adaptation mechanism for a multi-level mask pyramid. Users only need to draw a rough mask outline to guide the model to generate defects of varying shapes. For example, by specifying the approximate distribution range of the defect area using a low-resolution mask, the model can autonomously derive defects with diverse edge morphologies and differentiated texture features. This design, which decouples coarse-grained prompting from fine-grained generation, overcomes the traditional method's strong reliance on pixel-level annotation and allows for a flexible balance between the diversity and controllability of generated samples by adjusting the precision level.

[0015] 2. This multi-granularity metal surface defect image synthesis method deeply integrates multimodal semantic encoding with spatial position features to construct a cross-modal defect feature expression system. By encoding category information into semantic vectors and spatially weighting them with a mask pyramid, it achieves a mapping from semantic concepts to pixel-level features, thus enabling customizable defect category synthesis.

[0016] 3. This method for synthesizing multi-granular metal surface defect images employs a modular transfer learning strategy, adapting to industrial scenarios based on a pre-trained diffusion model. By freezing the VAE encoder and CLIP text encoder, their powerful general feature extraction capabilities are retained. While fine-tuning the denoising UNet using a metal surface defect dataset, it achieves small-sample learning capabilities and significantly reduces training resource consumption while ensuring generation quality.

[0017] 4. The present invention's multi-granularity metal surface defect image synthesis method effectively addresses the issue of a limited number of rare, abnormal samples by controlling the synthetic defect categories and adjusting the defect data distribution. This method can balance the proportion of multiple defect categories in the dataset. This method allows users to prioritize the generation of rare defects based on actual needs, significantly improving the representation of rare categories in the dataset while maintaining the diversity of generated samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flow chart of the method for synthesizing multi-granularity metal surface defect images according to the present invention;

[0019] Figure 2 This is the overall structure diagram of the multi-granularity defect synthesis model of the present invention;

[0020] Figure 3 Schematic diagram of the variational autoencoder VAE principle;

[0021] Figure 4 Schematic diagram of the joint encoding principle of mask pyramid and category information;

[0022] Figure 5 Schematic diagram of the precision encoding mask pyramid layered injection denoising UNet process. DETAILED DESCRIPTION

[0023] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific implementation methods.

[0024] The present invention is based on a multi-granularity metal surface defect image synthesis method of a pre-trained diffusion model. Figure 1 As shown, the specific steps are:

[0025] Step 1: Preprocessing of collected steel surface defect images.

[0026] The collected steel surface defect images constitute the metal surface defect dataset, which includes five types of defects: plate dust, roller marks, slag inclusions, small indentations, and wrinkles. A total of 1321 original images (resolution 2048×1000), including 286 plate dust images, 409 roller marks, 419 slag inclusions, 56 small indentations, and 151 wrinkles.

[0027] 1) For each image, first use the LabelMe tool to perform multi-precision labeling:

[0028] a. High-precision annotation accurately draws pixel-level polygonal masks along the defect edge to fully cover the micro-texture features;

[0029] b. In medium-precision annotation, simplified polygons are used to approximate the main outline of the defect, retaining the shape trend but ignoring details;

[0030] c. In low-precision marking, only the area where the defect is located is marked with a rectangular frame to define the spatial distribution range.

[0031] 2) After each original image generates a corresponding mask through the above three annotation methods, a unified preprocessing process is performed respectively.

[0032] The image and mask are simultaneously scaled and padded to 512×512 pixels to adapt to the model input size, and a bilinear interpolation algorithm is used to maintain spatial continuity; finally, each original image is combined with high, medium, and low masks into three independent samples, forming a multi-granularity annotation dataset containing 3963 sets of "image-annotation pairs".

[0033] Step 2: Introduce the precision encoding mask pyramid into the pre-trained diffusion model to build a multi-granularity hint defect synthesis model, such as Figure 2 shown.

[0034] The pre-trained diffusion model includes pre-trained VAE, UNet and CLIP. Among them, VAE includes an encoder and a decoder, which are used to compress the image into a low-dimensional latent space and reconstruct the variational autoencoder, reducing the computational complexity of the diffusion process. The denoising UNet is a neural network based on the U-type encoding and decoding structure in the diffusion model. It gradually removes the latent space noise through multi-scale residual connections to generate the target image. The CLIP module is a multimodal encoder based on contrastive learning, which realizes the semantic alignment of text descriptions and image features to control the generated content. Specifically:

[0035] 1) Denoising diffusion process training method and loss function of diffusion model.

[0036] Diffusion models are a type of generative model based on random processes that can generate samples with a distribution similar to the real data. The basic idea is to map the data to a noise distribution through a forward process and then gradually recover the data from the noise through a reverse process.

[0037] The forward process is to give a data sample x0 during the training process, and gradually transform it into a Gaussian noise distribution through a series of noise weighting steps, which is modeled by the following formula:

[0038]

[0039] Where I is the identity matrix, β t is a parameter that controls the amount of noise added at each step, x t Indicates that at the t-th step, the forward process adds noise multiple times so that x T (When →∞) it is close to the standard normal distribution. In order to recover the data sample, it is necessary to recover the data from the noise by learning the reverse process. This is usually achieved through a conditional generative model (such as a neural network) to approximate p θ (x t-1 |x t ), that is, from the given noise x t Restore to x t-1 The goal of the reverse process is to train by maximizing the likelihood function. A common training strategy is to minimize the following variational lower bound:

[0040]

[0041] in, is the output of the neural network, representing the denoising operation, ∈ t is the true noise at step t. By minimizing this loss, the network can learn how to recover data from noise.

[0042] Among them, β t is a parameter that controls the amount of noise added at each step, x t Indicates that at the t-th step, the forward process adds noise multiple times so that x T (When T→∞) it is close to the standard normal distribution. In order to recover the data sample, it is necessary to recover the data from the noise by learning the reverse process. This is usually achieved through a conditional generative model (such as a neural network) to approximate p θ (x t-1 |x t ), that is, from the given noise x t Restore to x t-1 The goal of the reverse process is to train by maximizing the likelihood function. A common training strategy is to minimize the following variational lower bound:

[0043]

[0044] in, is the output of the neural network, representing the denoising operation, ∈ t is the true noise at step t. By minimizing this loss, the network can learn how to recover data from noise.

[0045] 2) The loss function of VAE and how the diffusion model works after the introduction of VAE.

[0046] The Latent Diffusion Model (LDM) combines the traditional diffusion model with the variational autoencoder (VAE). VAE is a generative model, such as Figure 3 As shown, variational inference is used for learning. In the standard VAE framework, it is assumed that the observed data x is generated by the latent variable z, and the distribution of the latent variable is p(z), and the conditional distribution of the data is p(x|z). The core goal of the variational autoencoder is to maximize the marginal likelihood p(x), that is:

[0047] logp(x)=log∫p(x|z)p(z)dz

[0048] However, since the integral cannot be calculated explicitly, VAE introduces the variational distribution q(z|x) and approximates the maximum likelihood estimate by optimizing the variational lower bound (ELBO):

[0049]

[0050] Among them, D KL [q(z|x)||p(z)] is the KL divergence, which measures the difference between the variational distribution and the prior distribution. In the Latent Diffusion Model, the primary function of the VAE is to map high-dimensional input data (such as images) into a low-dimensional latent space. By learning the distribution of the latent space, the diffusion process no longer occurs directly in the high-dimensional data space, but rather in the latent space. This approach significantly reduces computational complexity because the diffusion process occurs only in the low-dimensional latent space.

[0051] The LDM, after the introduction of VAE, first maps the input data to a latent space through the VAE encoder, then applies the diffusion model to the latent space for modeling. Finally, the VAE decoder transforms the latent variable z back into the data space to generate the final sample. The introduction of VAE enables the diffusion model to be learned efficiently in the latent space, rather than directly in the high-dimensional data space. By restricting the diffusion process to the latent space, the LDM is able to significantly reduce computational overhead while preserving the diversity of the generated data. This combination improves the efficiency of the diffusion model in tasks such as image generation, enabling the LDM to achieve high-quality results at a low computational cost when generating complex data.

[0052] Since the multi-granularity defect hint synthesis model uses a multi-granularity mask to control the location, approximate shape, size, and other features of the generated defects, a precision-encoded mask pyramid is introduced into the pre-trained diffusion model Stable Diffusion to control the denoising process. The hint code (including the mask hint of the category and precision level) given by the user is integrated into the Stable Diffusion denoising UNet through the convolution-pooling-normalization layer, resulting in a multi-granularity defect hint synthesis model that is injected into the Stable Diffusion denoising UNet. Specifically:

[0053] like Figure 2 and Figure 5 As shown, the mask pyramid contains L levels, each level corresponding to a downsampling / upsampling block of the denoising UNet, and the height and width are the same as the corresponding block input dimensions. The height and width of the binary mask provided by the user should be the same as the input of the denoising UNet. The mask is downsampled L times to obtain a mask of level l = Ln, where n is the number of downsampling times. According to the level c input by the user, the masks with level l>c are set to 0, thereby blocking the high-resolution level and retaining the mask information of the low-resolution level. This means that all detailed masks with accuracy higher than c are discarded, and only low-precision masks are retained.

[0054] In order to further control the category of synthetic defects, in addition to inputting the category information into UNet through the cross attention mechanism, the defect category information encoded by CLIP is also jointly encoded with the pyramid. l Omit level l, M x,y is the vector of the mask at (x, y); for a one-channel binary image mask with a dimension of X×Y×1, m x,y,1 is the value of the mask at (x,y,1), m x,y,1 ∈{0,1}. x,y,1 Replicate to form a tensor M with Z identical channels X ×Y×Z. Then, the pre-trained CLIP is introduced to encode the text and category prompt t provided by the user to obtain the Z-dimensional vector f(t). Figure 4 As shown, M X×Y×Z Each vector at (x, y) in [1] is element-wise multiplied by f(t) to obtain a mask pyramid P with category information, resulting in a fused multi-level feature. In P, every pixel in masks at all scales contains a complete category code, and masks with a precision higher than c are all zero, without affecting the backbone denoising UNet.

[0055] like Figure 5 As shown in the figure, the fused multi-level features are encoded into control signals that match the size of the feature maps at each level of the UNet through convolutional layers (3×3 kernels), maximum pooling layers, and batch normalization layers. High-resolution layer features are injected into the shallow layers of the UNet to control edge details, intermediate-resolution layer features are injected into the middle layers of the UNet to regulate shape trends, and low-resolution layer features are injected into the deep layers to guide macroscopic distribution. For example, 128×128 layer features are injected into the fourth residual block of the UNet, and 32×32 layer features are incorporated into the eighth attention layer.

[0056] Through the above design, the reverse denoising process can simultaneously respond to spatial and semantic cues of different granularities, synthesize high-quality defect images with controllable position, shape, category, and size, and output corresponding pixel-level annotation masks.

[0057] Step 3: Fine-tune the multi-granularity defect synthesis model.

[0058] Based on the pre-processed multi-granularity annotated dataset, we fine-tuned the multi-granularity defect synthesis model, froze the parameters of the pre-trained Stable Diffusion VAE module and CLIP module, and focused on optimizing the parameters of the denoising UNet and precision encoding mask pyramid modules. Specifically:

[0059] A transfer learning strategy was employed to freeze the encoder-decoder parameters of the VAE module (preserving its image-to-latent space mapping capability) and the parameters of the CLIP text encoder (maintaining text-image semantic alignment). Only the convolution-pooling-normalization encoding layers of the denoising UNet and the mask pyramid were trained, ensuring efficient adaptation to the metal surface defect generation task with limited industrial data. Training was performed using the AdamW optimizer (with an initial learning rate of 1e-5 and weight decay of 0.01) with a batch size of 8, and 50 epochs in a distributed training environment. To balance the strength of the multi-granularity control signal, the weight coefficient of the mask pyramid control signal was set to 0.3 for the first 10 epochs and gradually increased to 1.0 in subsequent epochs, enabling a smooth transition from a free generation mode to a strongly constrained cue-guided generation mode. The noise prediction mean squared error (MSE) loss function was used. By calculating the difference between the predicted noise output of the denoising UNet and the actual noise, the model was driven to learn the mapping relationship between multi-granularity mask cues and defect image generation. Ultimately, the model is able to synthesize defect images and their pixel-level annotations with diverse morphologies, precise spatial distribution, and adjustable category weights while preserving the consistency of the metal surface background texture based on the coarse-grained masks and category hints provided by the user.

[0060] Step 4: Controllable defect synthesis and dataset enhancement.

[0061] Users can select mask inputs of different precision levels based on actual needs: for scenarios requiring precise control (such as gray defects on board tracks with clear edges), a high-resolution mask is drawn and the corresponding category is selected; for rare categories (such as small indentations), a low-precision rectangular box is used to delineate the approximate area and increase the generation weight of the category. After receiving the input, the model parses the multi-granularity control signal based on the precision-encoded mask pyramid, combines the category features of the CLIP semantic encoding, and performs a 50-step iterative denoising process in the latent space. Finally, the model is decoded by the VAE decoder to generate a defect image with a resolution of 512×512. The resulting multi-category balanced dataset can be used to train defect detection models. Ultimately, both synthetic data and real data are used to train defect detection models, solving the problems of insufficient data and unbalanced data category distribution in defect detection tasks.

Claims

1. A multi-granularity metal surface defect image synthesis method based on a pre-trained diffusion model, characterized in that: The following steps are involved: Step 1: Perform multi-precision annotation on the collected steel surface defect images to generate high-precision pixel-level masks, medium-precision contour masks, and low-precision rectangular frame masks. Preprocess the annotated data to construct a multi-granularity training dataset. Step 2: Based on the pre-trained StableDiffusion model, the Precision Encoding Mask Pyramid module is introduced to perform multi-level downsampling of the user-provided binary mask to generate hierarchical control signals. The CLIP semantic encoder is combined to perform feature fusion on the category or text prompts, and the denoising UNet network is injected into the layers to build a multi-granularity prompt defect synthesis model. Step 3: Freeze the VAE module and CLIP module parameters of the pre-trained model, use the multi-granularity training dataset to fine-tune the denoising UNet network and mask encoding layer, and optimize the noise prediction loss function; Step 4: The user inputs a custom mask and category prompt, and the synthetic model generates a controllable defect image to construct a balanced dataset.

2. The method according to claim 1, characterized in that The multi-precision annotation in step 1 specifically includes: high-precision annotation is to accurately draw pixel-level polygonal masks along the edge of the defect, completely covering the micro-texture features; medium-precision annotation is to simplify the polygon to approximate the main outline of the defect, retaining the shape trend but ignoring the details; low-precision annotation is to draw a rectangular frame mask to define the defect range and the spatial distribution range; after generating the three masks, they are uniformly scaled to the preset resolution and the boundaries are filled.

3. The method according to claim 1, characterized in that The construction method of the precision coding mask pyramid module in step 2 is as follows: the binary mask provided by the user is subjected to multi-level bilinear downsampling to generate a pyramid structure containing multiple resolution levels; the high-resolution level is masked according to the precision level specified by the user, and the mask information of the low-resolution level is retained; the category or text prompt is converted into a semantic vector through the CLIP encoder, and multiplied with the mask of each level channel by channel to generate semantic-spatial coupling features, which are then layered and injected into the shallow, middle and deep layer features of the denoising UNet after convolution-pooling-normalization encoding.

4. The method according to claim 3, characterized in that The hierarchical injection specifically includes: injecting features of the high-resolution layer into the shallow layer of UNet to control edge details, injecting features of the intermediate resolution layer into the middle layer of UNet to adjust shape trends, and injecting features of the low-resolution layer into the deep layer of UNet to guide macroscopic distribution.

5. The method according to claim 1, characterized in that The fine-tuning strategy in step 3 is as follows: freeze the encoder-decoder parameters of the VAE module and the CLIP text encoder parameters, and only optimize the denoising UNet network and the mask encoding layer; adopt staged training, initially reducing the weight of the mask control signal and gradually increasing it to full strength to balance the generation freedom and prompt constraints.

Citation Information

Patent Citations

  • Super-resolution reconstruction and feature extraction method for defect detection

    CN115170483A

  • Bridge cable sheath surface defect detection method

    CN117670825A

  • Text video retrieval optimization method based on text-to-image technology

    CN118377933A

  • Display panel crack defect generation method based on stable diffusion model

    CN118570136A

  • Steel surface defect image generation algorithm based on multi-granularity feature guidance

    CN118747779A

Cited By

  • Training method and device of image generation model and image generation method and device

    CN120932042A

  • Training methods and apparatus for image generation models, as well as image generation methods and apparatus.

    CN120932042B

  • Small sample precast concrete defect image generation method and system

    CN121074190A

  • A small sample precast concrete defect image generation method and system

    CN121074190B

  • Multi-modal geological feature fusion method and system based on discrete wavelet transform and CLIP-Stable Diffusion model

    CN121167609A