Stable diffusion method for generating defect image of few-sample meter

Through the DreamBooth method, the meter knowledge is embedded, and crack feature modeling and hypernetwork control are combined to generate high-quality and diverse meter defect images, solving the problem of scarce meter defect data in substations and improving detection accuracy.

CN120339182AActive Publication Date: 2025-07-18NORTH CHINA ELECTRIC POWER UNIV

Patent Information

Application Number
CN202510310376.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-18
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

The scarcity of defective data of substation meters leads to insufficient generation quality under the condition of few samples and lack of effective control, which makes it difficult to meet the practical application needs.

Method used

Through the DreamBooth method, a stable diffusion method is designed to generate high-quality and diverse meter defect images by embedding meter knowledge, combining crack feature modeling and hypernetwork dynamic control mechanism.

Benefits of technology

Generate high-quality and diverse meter defect images under the condition of few samples, which significantly improves the accuracy of defect detection and meets the data requirements for meter defect detection in substations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339182A_ABST
    Figure CN120339182A_ABST
Patent Text Reader

Abstract

The invention discloses a stable diffusion method for few-sample meter defect image generation, which is characterized in that a pre-training model is finely adjusted, structural features and defect knowledge of a transformer substation meter are embedded, and the similarity between a generated image and an actual meter is improved. A crack feature modeling module is innovatively designed, a line draft graph, a crack mask and a constraint graph are combined, a control image with geometric constraints is generated, and the shape and the position of the crack are accurately expressed. Meanwhile, a super network mechanism is introduced, weight distribution in the generation process is dynamically adjusted, and consistency and diversity of the generated image in form and position are ensured. According to the method, the transformer substation meter defect image with specific crack characteristics can be generated under the condition of a small number of samples, and the diversity and quality of the generated image are remarkably improved. Through application in a downstream detection task, the generated data effectively improves the precision and robustness of a defect detection model, and powerful support is provided for safe and stable operation of a power system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image analysis, and particularly to a stable diffusion method for few-shot meter defect image generation. Background Art

[0002] A substation is an important part of the power system, and the safe and stable operation of its equipment directly affects the overall reliability of the power grid. As a key monitoring device in the substation, the meter is responsible for power dispatching and operation monitoring, and its normal operation is crucial for ensuring the safety of the power grid. However, due to long-term exposure to a complex external environment, the meter is vulnerable to the influence of extreme weather such as sunlight, strong wind, heavy rain, and heavy snow, resulting in physical defects such as cracks and fractures. These defects may not only cause the meter to malfunction and be unable to monitor or record data normally, but also lead to equipment damage and even systemic hidden dangers, seriously threatening the safety and stability of the power grid.

[0003] In recent years, with the development of computer vision technology, significant progress has been made in the method of power equipment defect detection based on deep learning. However, the efficient training of deep learning models usually relies on a large-scale and accurately labeled dataset, and due to the difficulty of collecting substation meter defect data and the scarcity of samples, the construction cost of the dataset is high and the cycle is long. This data scarcity problem has become a bottleneck for further popularization and application. Especially under the condition of few samples, the generalization ability and detection accuracy of the model are severely limited.

[0004] To alleviate the data scarcity problem, data augmentation technology has become an effective means. Traditional data augmentation methods expand the dataset through operations such as geometric transformation (such as flipping, rotating, cropping, scaling) or adding noise, thereby improving the robustness of the model. However, these methods usually ignore the specific semantic content of the image and have limited performance in dealing with complex scenarios or fine-grained tasks, and it is difficult to effectively improve the diversity and complexity of the data distribution.

[0005] In recent years, generative adversarial networks (GANs) and diffusion models have gradually become the core technologies in the field of data augmentation. GANs generate high-quality images through adversarial training and significantly improve sample diversity when the sample size is limited. However, GANs have problems such as mode collapse and unstable training during the generation process. Especially under the condition of few samples, the generated images are prone to distortion, blurring and other phenomena, making it difficult to meet the actual application requirements. Diffusion models, with their significant advantages in detail control and generation quality, have gradually surpassed GANs and become the mainstream method for generating high-quality images. However, the existing generation methods still face the following challenges in the task of substation meter defect generation:

[0006] 1) There are significant differences in the scale of the dataset: Most existing generation methods are based on large-scale datasets (such as MS-COCO, CC12M). However, due to the difficulty in collecting substation meter defect data and the scarcity of samples, the performance of existing generation methods on small-sample datasets often fails to meet expectations, and the generated results are difficult to meet the actual application requirements.

[0007] 2) The generation process lacks effective control: Most existing generation methods rely on the probability mechanism in the denoising process, and the generated results are random, making it difficult to achieve quantitative control and meet specific requirements. On the other hand, editing methods with strong control capabilities rely too much on the original image and can only make forced modifications to the image, restricting the diversity of the generated images. Summary of the Invention

[0008] Therefore, in the above context, the present invention proposes a stable diffusion method for generating few-shot meter defect images, which solves the problem of insufficient substation meter defect data, expands the defect dataset, and at the same time helps to improve the accuracy of defect detection, that is, solves the problems of insufficient substation meter defect data and the poor effect of downstream detection tasks caused by it. By introducing meter knowledge embedding, crack feature modeling, and a hypernetwork dynamic control mechanism, a novel method for generating meter defect images is designed, which can controllably generate high-quality and diverse meter defect images under few-shot conditions, solve the problem of insufficient substation meter defect data, and further improve the accuracy of meter defect detection.

[0009] To achieve the above object, the present invention provides the following solutions:

[0010] A stable diffusion method for generating few-shot meter defect images, comprising the following steps:

[0011] S1, construct a dataset containing normal substation meter images and meter images with crack defects, covering various meter types, including: ammeters, voltmeters, pressure gauges, wattmeters, etc. All images are preprocessed, uniformly cropped to a resolution of 512×512, and the defect categories and location information are labeled. The dataset is divided into a training set and a validation set according to an 8:2 ratio to ensure the rationality of the data distribution;

[0012] S2, fine-tune the pre-trained Stable Diffusion model through the DreamBooth method, bind the unique Chinese identifier [meter] to the meter image, different from using English identifiers in traditional methods, to achieve the embedding of meter knowledge, and establish a strong association between the Chinese identifier and the meter image features. Design a dual-objective loss function to balance new feature learning and prior knowledge retention, ensuring that the model retains general generation capabilities when generating meter defect images.

[0013] S3. Propose a crack feature modeling module. Extract the meter line drawing through an adaptive edge detection algorithm, retain the structural features of the dial contour and key components, and ensure the integrity and rationality of the generated image in terms of geometric structure. Based on SAM, perform high-precision dial area segmentation on the input image to generate a dynamic constraint map. Further fuse the crack mask images provided by expert hand-drawn and industrial datasets, and precisely limit the cracks to only distribute in the effective area of the dial through the constraint image, avoiding the abnormal phenomenon of cracks appearing in non-functional areas.

[0014] S4. Adopt a hypernetwork mechanism to dynamically adjust the weight distribution during the generation process, ensuring the consistency and diversity of the generated images in terms of shape and position. Freeze the weights of the Stable Diffusion main model, construct a trainable hypernetwork branch, encode the control image into a latent representation, and inject it into the diffusion process layer by layer to constrain the crack morphology of the generated image.

[0015] Among them, S2 specifically includes:

[0016] First, by fine-tuning Stable Diffusion, endow the unique Chinese identifier [meter] with a new meaning and inject it as an input prompt into the training process, enabling the existing model to adapt to the substation meter feature knowledge, thereby enhancing the model's ability to generate meter defect images.

[0017] Next, in order to prevent the model from forgetting the prior knowledge learned from large-scale data, an optimization objective containing the main body loss and the prior preservation loss is used to achieve a balance between new concept learning and prior knowledge preservation. The main body loss term L sbj is used to optimize the model to generate images containing new features under specific conditions. By minimizing the gap between the model prediction and the target distribution, guide the model to learn specific meter defect features. The mathematical expression is as follows:

[0018]

[0019] Among them, x represents the real image, c is the text condition, ∈ is the random noise, t is the time step in the diffusion process, ω t is a weight term for a time step used to balance the loss contributions of different time steps, represents the predicted image of the model, α t and σ t are diffusion noise scheduling parameters, Perform the expectation calculation on x, c, ∈, t, that is, take the average over the data distribution, ‖·‖ 2 is the L2 norm, which measures the gap between the predicted image and the real image. The prior preservation loss term L prAim to constrain the generation results of the model so that they do not deviate from the original data distribution, thus avoiding the decline in the generalization ability of the model caused by few-shot overfitting. The loss function is expressed as follows:

[0020]

[0021] where x pr represents the image generated by the model previously, c pr is the corresponding text condition, ∈′ is another random noise term, and t' is the time step in the diffusion process. By introducing the generated samples of large-scale unlabeled data, the model's forgetting of the general generation ability is effectively avoided.

[0022] Finally, the two loss functions are dynamically balanced in a weighted manner, and the combined form is expressed as follows:

[0023] L = L sbj + λL pr (3)

[0024] where L is the overall loss function, λ ∈ [0, 1] is the weight parameter, which is used to adjust the trade-off between the main loss and the prior preservation loss, and flexibly optimize the generation ability according to the task. Fine-tuning the base model improves the ability to generate specific images while retaining the generalization ability of the base model.

[0025] Among them, S3 specifically includes:

[0026] First, the defect-free original image I o ∈ R H×W×3 (where R represents that the image data belongs to the real number space, that is, the value of each pixel is a real number, H and W respectively represent the height and width of the image, and 3 represents the RGB three channels) undergoes preprocessing such as edge detection, image smoothing, and style enhancement, and the line drawing image I l ∈ R H×W×1 is extracted. Subsequently, the line drawing image and the crack mask I d ∈ R H×W×1 are fused pixel by pixel to generate the control image I c ′ ∈ R H×W×1 . The fusion process is expressed as:

[0027] I c ′(x, y) = max(I d (x, y), I l (x, y)) (4)

[0028] where I c ′(x, y), I d (x, y) and I l(x, y) represent the gray values of the control image, crack image, and line drawing image at the pixel point (x, y), respectively, with the range limited to [0, 255]. The max(·) operation represents taking the maximum gray value at each pixel point.

[0029] However, due to the particularity of crack distribution, directly combining the line drawing image with the crack image may lead to abnormal phenomena where crack features appear outside the dial, making it impossible to accurately model the dial area. Therefore, SAM is used to limit the modeling range of cracks and improve its authenticity. For the input original image I o and the hint condition c provided by the user, SAM first maps I o to the embedding space through the image encoder to generate the feature representation f; at the same time, encodes the hint condition c through the hint encoder to generate the hint feature g; subsequently, the features f and g are fused in the mask decoder to generate the dial constraint map I r . The formula is as follows:

[0030] I r = D(f, g), I r ∈R H×W×1 (5)

[0031] where D is the mask decoder. After obtaining the dial constraint, it constrains the pixel-by-pixel fusion process of the line drawing image and the mask image to generate the final control image. The fusion process can be updated as:

[0032]

[0033] where I c (x, y) and I r (x, y) represent the gray values of the final control image and the constraint image at the pixel point (x, y), respectively. The gray value of I r (x, y) only takes two values, 0 and 255. The former represents the area allowed for combination, and the latter represents retaining the original gray value of the line drawing image. By introducing the constraint map, more precise regional control can be imposed during the crack feature modeling process, which helps to model a more reasonable control image.

[0034] Among them, S4 specifically includes:

[0035] First, freeze the weights of the Stable Diffusion main model to retain its original image prior knowledge. On this basis, construct a trainable model copy as a hypernetwork, and dynamically adjust the generation process of the main model by sharing some feature representations.

[0036] Secondly, after the input conditional control image is processed by the encoder, a corresponding latent representation is generated. This latent representation is combined with the dynamic weights generated by the hypernetwork to effectively adjust the overall contour and crack details of the generated image. At the same time, the basic generation model generates a preliminary image, and the two work together to complete the generation of the final image, thus ensuring that the generation result strictly meets the set conditions and achieves the effect of precise control.

[0037] Finally, the constraint image is used as the input condition to show the rough contour of the overall image and the position and shape of the cracks. At this stage, a specific loss function is designed to optimize the generation process, which is defined as follows:

[0038]

[0039] where x represents the original image, c, c', c+ are conditional information, ∈ is the standard normal distribution noise, t is the time step in the diffusion process, represents the denoising model parameterized by the neural network, which is used to recover the original image from the noisy data, α t and σ t are time-dependent scaling factors that control the degree of noise addition, The expectation calculation for x, c, c', ∈, t is to take the average over the data distribution, ‖·‖ 2 is the L2 norm, which measures the difference between the predicted image and the real image. In the diffusion generation process, the control signal is injected into each feature layer of the model multiple times, which affects both the initial generation stage and provides structural constraints in the middle and late stages, thus ensuring that the generation result meets the set conditions.

[0040] Among them, it also includes: for the generation of the crack mask, a crack mask generation method based on the combination of expert hand-drawn and existing industrial datasets is proposed to ensure the integrity and accuracy of the crack features. Specifically, it includes:

[0041] The sources of the crack mask include the crack defect images in the existing industrial datasets and the crack images hand-drawn by industry experts. Among them, the crack mask hand-drawn by experts provides a reliable supplement for crack modeling. Especially in the case where SAM segmentation may fail, it can effectively make up for the deficiency of the dial area segmentation, ensure the integrity and accuracy of the crack features, thus reducing the risk of relying on segmentation and further improving the robustness and generalization performance of the generation model. By combining expert hand-drawn and existing industrial datasets, the generated crack mask can more accurately reflect the morphology and distribution characteristics of the actual meter cracks, providing high-quality input data for subsequent crack feature modeling.

[0042] This application also provides a stable diffusion device for generating few-shot meter defect images, and the device includes:

[0043] The S1 module is used to construct a dataset containing normal substation meter images and meter images with crack defects, covering various types of meters, including ammeters, voltmeters, pressure gauges, wattmeters, etc. All images are preprocessed, uniformly cropped to a resolution of 512×512, and labeled with defect categories and location information. The dataset is divided into a training set and a validation set in an 8:2 ratio to ensure the rationality of data distribution.

[0044] The S2 module is used to fine-tune the pre-trained Stable Diffusion model through the DreamBooth method, bind the unique Chinese identifier [meter] to the meter image, which is different from using English identifiers in traditional methods, realize the embedding of meter knowledge, and establish a strong association between the Chinese identifier and the meter image features. A dual-objective loss function is designed to balance new feature learning and prior knowledge retention, ensuring that the model retains general generation ability when generating meter defect images.

[0045] The S3 module is used to propose a crack feature modeling module. The meter line drawing is extracted through an adaptive edge detection algorithm, and the structural features of the dial contour and key components are retained to ensure the integrity and rationality of the generated image in terms of geometric structure. Based on SAM, high-precision dial area segmentation is performed on the input image to generate a dynamic constraint map, and further fuse the crack mask images provided by expert hand-drawn and industrial datasets. The crack is accurately restricted to the effective area of the dial through the constraint image, avoiding the abnormal phenomenon of cracks appearing in non-functional areas.

[0046] The S4 module adopts a hypernetwork mechanism to dynamically adjust the weight distribution during the generation process, ensuring the consistency and diversity of the generated images in terms of shape and position. Freeze the weights of the Stable Diffusion main model, construct a trainable hypernetwork branch, encode the control image into a latent representation, and inject it layer by layer into the diffusion process to constrain the crack morphology of the generated image.

[0047] This application also provides an electronic device, which includes: at least one memory and at least one processor; the memory stores a program, and the processor calls the program stored in the memory, and the program is used to implement the stable diffusion method for generating few-shot meter defect images.

[0048] This application also provides a storage medium, in which computer-executable instructions are stored, and the computer-executable instructions are used to execute the stable diffusion method for generating few-shot meter defect images.

[0049] The above technical solutions of the present invention can achieve the following technical effects:

[0050] A stable diffusion method, device, electronic device and storage medium for few-shot meter defect image generation. Specifically, a meter knowledge embedding method based on DreamBooth is designed. By fine-tuning the pre-trained model, the structural features and defect knowledge of substation meters are injected into the generation model, significantly improving the similarity between the generated images and actual meters. A crack feature modeling module is proposed, which combines line drawings, crack masks and constraint diagrams to generate control images with geometric constraints, accurately expressing the shape and position of cracks. The hypernetwork mechanism is introduced to dynamically adjust the weight distribution during the generation process, ensuring the consistency and diversity of the generated images in terms of shape and position. This meter defect generation method can controllably generate high-quality and diverse meter defect images under the condition of few samples, and assist in improving the accuracy of the meter defect detection network from the data level. The present invention applies the Stable Diffusion model to few-shot substation meter defect image generation. By combining meter knowledge embedding, crack feature modeling and hypernetwork dynamic control, it effectively improves the quality and controllability of defect image generation, meets the data requirements of substation meter defect detection tasks under few-sample conditions, and provides strong data support for the safe and stable operation of the power system. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0052] Figure 1 It is a schematic flowchart of a stable diffusion method for few-shot meter defect image generation according to an embodiment of the present invention;

[0053] Figure 2 It is a schematic structural diagram of crack feature modeling according to an embodiment of the present invention;

[0054] Figure 3 It is a schematic structural diagram of hypernetwork control generation according to an embodiment of the present invention;

[0055] Figure 4 It is a schematic overall structural diagram according to an embodiment of the present invention;

[0056] Figure 5 It is an effect diagram of generating meter defects according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0058] Existing methods for generating meter defect images have limitations such as insufficient generation quality, lack of effective control, and poor adaptability to few samples. For example, traditional GANs (Generative Adversarial Networks) are prone to mode collapse and unstable training, with significantly blurred and ghosting phenomena in the generated images, and a similarity too high with the original defect images, resulting in limited diversity; diffusion models rely on a random denoising process and are difficult to constrain the crack morphology and position, leading to generated results deviating from actual requirements; while existing methods rely on large-scale pre-trained data and are difficult to capture the fine-grained features of meter defects under few-sample conditions, with noise or invalid defects appearing in the generated images.

[0059] The purpose of the present invention is to provide a stable diffusion method for few-sample meter defect image generation, to solve problems such as insufficient substation meter defect data and poor effects of downstream detection tasks caused by it. By introducing meter knowledge embedding, crack feature modeling, and a hypernetwork dynamic control mechanism, a novel method for generating meter defect images is designed, which can controllably generate high-quality and diverse meter defect images under few-sample conditions, thereby achieving the purpose of improving the meter defect detection ability.

[0060] The method of the present invention can significantly improve the generation quality and controllability of meter defect images under few-sample conditions. Its core advantages are as follows: First, the DreamBooth fine-tuning technique solves the problem of the pre-trained Stable Diffusion model's misinterpretation of meters, and the generated images are highly consistent with actual meters in terms of overall contour; second, the crack feature modeling module uses the constraint map generated by SAM (Segment Anything Model) to limit the crack distribution area, achieving pixel-level precise control of crack morphology and position; finally, the hypernetwork mechanism injects the features of the control image into the diffusion process layer by layer through dynamic control without modifying the weights of the Stable Diffusion main model, ensuring both the consistency of crack generation.

[0061] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0062] As Figure 1 shown, a stable diffusion method for few-sample meter defect image generation provided by the present invention includes the following steps:

[0063] S1. Construct a dataset containing normal substation meter images and meter images with crack defects, covering various types of meters, including ammeters, voltmeters, pressure gauges, wattmeters, etc. All images are preprocessed, uniformly cropped to a resolution of 512×512, and the defect categories and location information are labeled. The dataset is divided into a training set and a validation set in an 8:2 ratio to ensure the rationality of data distribution.

[0064] S2. Fine-tune the pre-trained Stable Diffusion model using the DreamBooth method, bind the unique Chinese identifier [meter] to the meter image, different from using English identifiers in traditional methods, to achieve the embedding of meter knowledge and establish a strong association between the Chinese identifier and the meter image features. Design a dual-objective loss function to balance new feature learning and prior knowledge retention, ensuring that the model retains general generation ability when generating meter defect images.

[0065] S3. Propose a crack feature modeling module. Extract the meter line drawing through an adaptive edge detection algorithm, retain the structural features of the dial contour and key components, ensuring the integrity and rationality of the generated image in terms of geometric structure. Based on the Segment Anything Model (SAM), perform high-precision dial area segmentation on the input image to generate a dynamic constraint map, and further fuse the crack mask images provided by expert hand-drawn and industrial datasets. Precisely limit the cracks to only be distributed in the valid area of the dial through the constraint image, avoiding the abnormal phenomenon of cracks appearing in non-functional areas.

[0066] S4. Adopt a hypernetwork mechanism to dynamically adjust the weight distribution during the generation process, ensuring the consistency and diversity of the generated images in terms of morphology and position. Freeze the weights of the Stable Diffusion main model, construct a trainable hypernetwork branch, encode the control image into a latent representation, and inject it layer by layer into the diffusion process to constrain the crack morphology of the generated image.

[0067] S5. Use the Stable Diffusion method for generating few-shot meter defect images to assist in improving the accuracy of the meter defect recognition network, and further verify the effectiveness of the few-shot substation meter defect generation method.

[0068] The basic flowchart of the present invention is as Figure 1 shown.

[0069] Considering that training large models usually requires large-scale labeled datasets and high computing resources, while the number of substation meter datasets is scarce, the features are complex, and there are significant differences in the data distribution from natural images, an effective method is needed to inject the feature knowledge of substation meters into existing large models. This method fine-tunes the pre-trained Stable Diffusion model through the DreamBooth method, binds the unique Chinese identifier [meter] to the meter image, differentiates from the traditional method using English identifiers, realizes the embedding of meter knowledge, and establishes a strong association between the Chinese identifier and the meter image features. A dual-objective loss function is designed to balance new feature learning and prior knowledge retention, ensuring that the model retains its general generation ability when generating meter defect images. The specific steps in step S2 include:

[0070] First, by fine-tuning Stable Diffusion, the unique Chinese identifier [meter] is given a new meaning and injected into the training process as an input prompt, enabling the existing model to adapt to the feature knowledge of substation meters, thereby enhancing the model's ability to generate meter defect images;

[0071] Next, to prevent the model from forgetting the prior knowledge learned from large-scale data, this method uses an optimization objective that includes a main loss and a prior preservation loss to achieve a balance between new concept learning and prior knowledge retention. The main loss term L sbj is used to optimize the model to generate images containing new features under specific conditions. By minimizing the gap between the model prediction and the target distribution, the model is guided to learn specific meter defect features. The mathematical expression is as follows:

[0072]

[0073] where x represents the real image, c is the text condition, ∈ is the random noise, t is the time step in the diffusion process, ω t is a weight term for a time step to balance the loss contributions of different time steps, represents the predicted image of the model, α t and σ t are diffusion noise scheduling parameters, performs the expectation calculation on x, c, ∈, t, that is, taking the average over the data distribution, ‖·‖ 2 is the L2 norm, which measures the gap between the predicted image and the real image. The prior preservation loss term L pr is designed to constrain the generation results of the model so that they do not deviate from the original data distribution, thereby avoiding the decline in the model's generalization ability caused by overfitting with few samples. The loss function is expressed as follows:

[0074]

[0075] where xpr represents the image generated by the model previously, c pr is the corresponding text condition, ∈′ is another random noise term, and t' is the time step in the diffusion process. By introducing the generated samples of large-scale unlabeled data, the model's general generation ability is effectively prevented from being forgotten.

[0076] Finally, the two loss functions are dynamically balanced in a weighted manner, and the combined form is expressed as follows:

[0077] L = L sbj + λL pr (3)

[0078] where L is the overall loss function, λ ∈ [0, 1] is the weight parameter, which is used to adjust the trade-off between the main loss and the prior preservation loss, and flexibly optimizes the generation ability according to the task. Fine-tuning the base model improves the ability to generate specific images while retaining the generalization ability of the base model.

[0079] In the present invention, considering the complex and irregular distribution of the crack morphology of the substation meter, it is necessary to closely combine the geometric features of the crack with the meter structure information. At the same time, due to the diversity and randomness of the crack distribution, a model-based method needs to be adopted to accurately model the crack features. This method proposes a crack feature modeling module, which generates a control image with geometric constraints by combining the line drawing, crack mask, and constraint map to ensure the accurate expression of the crack morphology and position. By using edge detection and image smoothing techniques to extract the meter line drawing, combining the constraint map generated by SAM to limit the distribution range of the crack, and generating a control image through pixel-level fusion, reliable geometric constraint support is provided for subsequent generation tasks. The specific steps in step S3 include:

[0080] First, the defect-free original image I o ∈ R H×W×3 (where R represents that the image data belongs to the real number space, that is, the value of each pixel is a real number, H and W respectively represent the height and width of the image, and 3 represents the RGB three channels) undergoes preprocessing of edge detection, image smoothing, and style enhancement, and the line drawing image I l ∈ R H×W×1 is obtained. Subsequently, the line drawing image and the crack mask I d ∈ R H×W×1 are fused pixel by pixel to generate a control image I c ′ ∈ R H×W×1 . The fusion process is expressed as:

[0081] I c ′(x, y) = max(I d (x, y), I l (x, y)) (4)

[0082] Among them, I c ′(x,y), I d (x,y) and I l (x,y) are the gray values of the control image, crack image, and line drawing image at the pixel point (x,y) respectively, with the range limited to [0, 255]. The max(·) operation represents taking the maximum gray value at each pixel point.

[0083] However, due to the particularity of crack distribution, directly combining the line drawing image with the crack image may lead to abnormal phenomena where crack features appear outside the dial, making it impossible to accurately model the dial area. Therefore, SAM is used to limit the modeling range of cracks and improve its authenticity. For the input original image I o and the prompt condition c provided by the user, SAM first maps I o to the embedding space through the image encoder to generate the feature representation f; at the same time, the prompt condition c is encoded by the prompt encoder to generate the prompt feature g; subsequently, the features f and g are fused in the mask decoder to generate the dial constraint map I r . The formula is as follows:

[0084] I r = D(f, g), I r ∈R H×W×1 (5)

[0085] Among them, D is the mask decoder. After obtaining the dial constraint, it constrains the pixel-by-pixel fusion process of the line drawing image and the mask image to generate the final control image. The fusion process can be updated as:

[0086]

[0087] Among them, I c (x,y) and I r (x,y) represent the gray values of the final control image and the constraint image at the pixel point (x,y) respectively. The gray value of I r (x,y) only takes two values, 0 and 255. The former represents the area allowed for combination, and the latter represents retaining the original gray value of the line drawing image. By introducing the constraint map, more precise area control can be imposed during the crack feature modeling process, which helps to model a more reasonable control image.

[0088] Regarding the generation of the crack mask, the present invention proposes a crack mask generation method based on the combination of expert hand-drawing and existing industrial data sets to ensure the integrity and accuracy of crack features. Specifically, it includes:

[0089] The sources of crack masks include crack defect images in existing industrial datasets and crack images hand-drawn by industry experts. Among them, the crack masks hand-drawn by experts provide a reliable supplement for crack modeling. Especially in cases where SAM segmentation may fail, it can effectively make up for the deficiency in the segmentation of the dial area, ensure the integrity and accuracy of crack features, thereby reducing the risk of relying on segmentation and further improving the robustness and generalization performance of the generation model. By combining hand-drawn by experts with existing industrial datasets, the generated crack masks can more accurately reflect the morphology and distribution characteristics of actual meter cracks, providing high-quality input data for subsequent crack feature modeling.

[0090] The structural schematic diagram of crack feature modeling in the present invention is as Figure 2 shown.

[0091] In the present invention, considering the high-precision control requirements for the crack morphology and position in the substation meter defect generation task, it is necessary to closely combine geometric constraints with the generation process. At the same time, due to the limitations of traditional generation methods in control capabilities, a dynamic weight adjustment mechanism based on a hypernetwork is required. This method proposes a hypernetwork control generation module. By introducing a trainable hypernetwork layer, the weight distribution in the generation process is dynamically adjusted to ensure that the generated images strictly follow the preset crack morphology and position conditions. Among them, in the step S4, a hypernetwork mechanism is adopted to encode the control image into a latent representation and inject it layer by layer into the diffusion process to constrain the crack features of the generated images, specifically including:

[0092] First, freeze the weights of the Stable Diffusion main model to retain its original image prior knowledge. On this basis, construct a trainable model copy as a hypernetwork, and dynamically adjust the generation process of the main model by sharing some feature representations.

[0093] Second, after being processed by the encoder, the input conditional control image generates a corresponding latent representation. This latent representation is combined with the dynamic weights generated by the hypernetwork to effectively adjust the overall contour and crack details of the generated image. At the same time, the basic generation model generates a preliminary image, and the two work together to jointly complete the generation of the final image, thereby ensuring that the generation result strictly meets the set conditions and achieving the effect of precise control.

[0094] Finally, use the constraint image as the input condition to show the rough contour of the overall image and the position and morphology of the cracks. At this stage, a specific loss function is designed to optimize the generation process, which is defined as follows:

[0095]

[0096] Among them, \(x\) represents the original image, \(c\), \(c'\), \(c^+\) are conditional information, \(\epsilon\) is standard normal distribution noise, and \(t\) is the time step in the diffusion process. denotes the denoising model parameterized by a neural network, which is used to recover the original image from the noisy data, \(\alpha\) t and \(\sigma\) t are time-dependent scaling factors that control the degree of noise addition. Performing the expectation calculation on \(x\), \(c\), \(c'\), \(\epsilon\), \(t\) means taking the average over the data distribution. \(\|\cdot\|\) 2 is the L2 norm, which measures the gap between the predicted image and the real image. During the diffusion generation process, the control signal is injected into each feature layer of the model multiple times, which affects both the initial stage of generation and provides structural constraints in the middle and late stages, thus ensuring that the generated results meet the set conditions.

[0097] The schematic structural diagram of the supernetwork control generation in the present invention is as Figure 3 shown.

[0098] In the step S5, the method for generating few-shot substation meter defect images is used to assist in improving the accuracy of the meter defect recognition network, and further verify the effectiveness of the few-shot substation meter defect generation method, which specifically includes:

[0099] To verify the effectiveness of the present invention in generating few-shot meter defect images, experiments were conducted on the defect detection task. The experimental process is divided into the following steps: First, use the method proposed in the present invention to generate a large number of substation meter images containing crack defects; then, mix the generated defect images with the real defect data set in proportion to construct an expanded training set; finally, use the augmented data set to train the defect detection model and evaluate the detection performance of the model on an independent validation set. The experimental results show that the generated data significantly improves the accuracy of the detection model and indicators such as mAP50, which proves the effectiveness of the method proposed in the present invention and provides reliable data support for substation meter defect detection.

[0100] A stable diffusion method for generating few-shot meter defect images according to the present invention, the network structure of this generation method is as Figure 4 shown.

[0101] The generation effect of the meter defect images of the method of the present invention is as Figure 5As shown in the figure. The present invention designs a method for embedding meter knowledge based on DreamBooth. By fine-tuning the pre-trained model, the structural features and defect knowledge of substation meters are injected into the generation model, significantly improving the similarity between the generated images and the actual meters. A crack feature modeling module is proposed, which combines the line drawing, crack mask, and constraint map to generate a control image with geometric constraints, accurately expressing the shape and position of the crack. The hypernetwork mechanism is introduced to dynamically adjust the weight distribution during the generation process, ensuring the consistency and diversity of the generated images in terms of shape and position. This method for generating meter defects can controllably generate high-quality and diverse meter defect images under the condition of a small number of samples, assisting in improving the accuracy of the meter defect detection network at the data level. The present invention applies the StableDiffusion model to the generation of few-shot substation meter defect images. By combining meter knowledge embedding, crack feature modeling, and hypernetwork dynamic control, it effectively improves the quality and controllability of the generated defect images, meeting the data requirements for the substation meter defect detection task under the condition of few samples.

[0102] The present application also provides a Stable Diffusion device for few-shot meter defect image generation, and the device includes:

[0103] Module S1, constructing a data set containing normal substation meter images and meter images with crack defects, covering various meter types, including: ammeters, voltmeters, pressure gauges, wattmeters, etc. All images are preprocessed, uniformly cropped to a resolution of 512×512, and the defect category and location information are labeled. The data set is divided into a training set and a validation set according to a ratio of 8:2 to ensure the rationality of the data distribution.

[0104] Module S2, fine-tuning the pre-trained Stable Diffusion model by the DreamBooth method, binding the unique Chinese identifier [meter] to the meter image, different from the traditional method using English identifiers, to achieve the embedding of meter knowledge and establish a strong association between the Chinese identifier and the meter image features. Design a dual-objective loss function to balance new feature learning and prior knowledge retention, ensuring that the model retains the general generation ability when generating meter defect images.

[0105] Module S3, proposing a crack feature modeling module, extracting the meter line drawing through an adaptive edge detection algorithm, retaining the structural features of the dial contour and key components, ensuring the integrity and rationality of the generated image in terms of geometric structure. Based on SAM, high-precision dial area segmentation is performed on the input image to generate a dynamic constraint map, and further fuse the crack mask images provided by expert hand-drawn and industrial data sets. The crack is accurately restricted to the effective area of the dial through the constraint image, avoiding the abnormal phenomenon of cracks appearing in non-functional areas.

[0106] The S4 module adopts a hypernetwork mechanism to dynamically adjust the weight distribution during the generation process, ensuring the consistency and diversity of the generated images in terms of morphology and position. Freeze the weights of the Stable Diffusion main model, construct a trainable hypernetwork branch, encode the control image into a latent representation, and inject it layer by layer into the diffusion process to constrain the crack morphology of the generated images.

[0107] This application also provides an electronic device, which includes: at least one memory and at least one processor; the memory stores a program, and the processor calls the program stored in the memory, and the program is used to implement the stable diffusion method for generating few-shot meter defect images.

[0108] This application also provides a storage medium, in which computer-executable instructions are stored, and the computer-executable instructions are used to execute the stable diffusion method for generating few-shot meter defect images.

[0109] In this article, specific examples are used to elaborate on the principles and implementation methods of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0110] It should be noted that the embodiments in this specification are all described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same and similar parts between the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the method part.

[0111] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements inherent to the process, method, article or device, but also other identical elements inherent to these process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the said element.

[0112] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A stable diffusion method for few-shot meter defect image generation, characterized in that, It includes the following steps: S1. Construct a dataset containing normal substation meter images and meter images with crack defects, covering various meter types. All images are preprocessed, uniformly cropped to a resolution of 512×512, and the defect categories and location information are labeled; S2. Fine-tune the pre-trained Stable Diffusion model by the DreamBooth method, bind the unique Chinese identifier [meter] to the meter image, which is different from using English identifiers in traditional methods, realize the embedding of meter knowledge, establish a strong association between the Chinese identifier and the meter image features, design a dual-objective loss function, balance the learning of new features and the retention of prior knowledge, and ensure that the model retains the general generation ability when generating meter defect images; S3. Propose a crack feature modeling module. Extract the meter line drawing through an adaptive edge detection algorithm, retain the structural features of the dial contour and key components, ensure the integrity and rationality of the generated image in terms of geometric structure, perform high-precision dial area segmentation on the input image based on SAM, generate a dynamic constraint map, and further fuse the crack mask images provided by expert hand-drawn and industrial datasets. Precisely limit the cracks to only be distributed in the effective area of the dial through the constraint image, and avoid the abnormal phenomenon of cracks appearing in non-functional areas; S4. Adopt a hypernetwork mechanism to dynamically adjust the weight distribution during the generation process, ensure the consistency and diversity of the generated images in terms of morphology and position, freeze the weights of the Stable Diffusion main model, construct a trainable hypernetwork branch, encode the control image into a latent representation, and inject it layer by layer into the diffusion process to constrain the crack morphology of the generated image.

2. The stable diffusion method according to claim 1, characterized in that S2 specifically includes: First, by fine-tuning Stable Diffusion, endow the unique Chinese identifier [meter] with a new meaning and inject it as an input prompt into the training process, so that the existing model adapts to the substation meter feature knowledge to enhance the model's ability to generate meter defect images; Next, in order to prevent the model from forgetting the prior knowledge learned from large-scale data, an optimization objective including the main body loss and the prior preservation loss is used to achieve a balance between new concept learning and prior knowledge preservation. The main body loss term L sbj is used to optimize the model to generate images containing new features under specific conditions. By minimizing the gap between the model prediction and the target distribution, the model is guided to learn specific meter defect features. The mathematical expression is as follows: where \(x\) represents the real image, \(c\) is the text condition, \(\epsilon\) is the random noise, \(t\) is the time step in the diffusion process, and \(\omega\) t is the weight term for a time step to balance the loss contributions of different time steps, represents the predicted image of the model, \(\alpha\) t and \(\sigma\) t are the diffusion noise schedule parameters, taking the expectation over \(x\), \(c\), \(\epsilon\), \(t\) means averaging over the data distribution. \(\|\cdot\|\) 2 is the L2 norm, which measures the gap between the predicted image and the real image. The prior preservation loss term \(L\) pr aims to constrain the generation results of the model so that they do not deviate from the original data distribution, thereby avoiding the decline in the generalization ability of the model due to overfitting with few samples. The loss function is expressed as follows: where x pr represents the image generated by the model in the previous step, c pr is the corresponding text condition, ∈ ′ is another random noise term, and t' is the time step in the diffusion process. By introducing the generated samples of large-scale unlabeled data, the model's general generation ability can be effectively prevented from being forgotten; Finally, dynamically balance the two loss functions in a weighted manner, and the combined form is expressed as follows: L = L sbj + λL pr (3) Among them, L is the overall loss function, λ∈[0,1] is the weight parameter, which is used to adjust the trade-off between the main loss and the prior preservation loss, flexibly optimize the generation ability according to the task, fine-tuning the basic model improves the ability to generate specific images while retaining the generalization ability of the basic model.

3. The stable diffusion method according to claim 1, characterized in that, S3 specifically includes: First, the defect-free original image I o ∈R H×W×3 After preprocessing of edge detection, image smoothing, and style enhancement, the line drawing image I l ∈R H×W×1 is obtained. Here, R represents that the image data belongs to the real number space, i.e., the value of each pixel is a real number, H and W respectively represent the height and width of the image, and 3 represents the RGB three channels. Subsequently, the line drawing image and the crack mask I d ∈R H×W×1 are fused pixel by pixel to generate the control image I c ′ ∈R H×W×1 . The fusion process is expressed as: I c ′ (x,y) = max(I d (x,y), I l (x,y)) (4) Among them, I c ′ (x, y), I d (x, y) and I l (x, y) are the gray values of the control image, the crack image, and the line drawing image at the pixel point (x, y) respectively, with the range limited to [0, 255], and the max(·) operation represents taking the maximum gray value at each pixel point; However, due to the particularity of crack distribution, directly combining the line drawing image with the crack image may lead to abnormal phenomena where crack features appear outside the dial, making it impossible to accurately model the dial area. Therefore, SAM is used to limit the modeling range of cracks and improve its authenticity. For the input original image I o and the prompt condition c provided by the user, SAM first maps I o to the embedding space through the image encoder to generate the feature representation f; at the same time, encodes the prompt condition c through the prompt encoder to generate the prompt feature g; subsequently, the features f and g are fused in the mask decoder to generate the dial constraint graph I r , and the formula is as follows: I r = D(f,g), I r ∈R H×W×1 (5) Among them, D is the mask decoder. After obtaining the dial constraint, it constrains the pixel-by-pixel fusion process of the line drawing image and the mask image to generate the final control image. The fusion process can be updated as: Among them, I c (x, y) and I r (x, y) respectively represent the gray values of the final control image and the constraint image at the pixel point (x, y). The gray value of I r (x, y) only takes two values, 0 and 255. The former represents the area allowed for combination, and the latter represents retaining the original gray value of the line drawing image. By introducing the constraint map, more precise area control can be imposed during the crack feature modeling process, which helps to model a control image with higher rationality.

4. The stable diffusion method according to claim 1, wherein S4 specifically includes: First, freeze the weights of the Stable Diffusion main model to retain its original image prior knowledge. On this basis, construct a trainable model copy as a hypernetwork, and dynamically adjust the generation process of the main model by sharing some feature representations; Secondly, after the input conditional control image is processed by the encoder, a corresponding latent representation is generated. This latent representation is combined with the dynamic weights generated by the hypernetwork to effectively adjust the overall contour and crack details of the generated image. At the same time, the basic generation model generates a preliminary image, and the two work together to complete the generation of the final image, thus ensuring that the generation result strictly conforms to the set conditions and achieving the effect of precise control; Finally, the constraint image is used as the input condition to show the rough contour of the overall image and the position and shape of the cracks. At this stage, a specific loss function is designed to optimize the generation process, which is defined as follows: Among them, x represents the original image, c, c', c+ are conditional information, ∈ is standard normal distribution noise, and t is the time step in the diffusion process. represents the denoising model parameterized by the neural network, which is used to recover the original image from the noisy data, and α t and σ t are time-dependent scaling factors that control the degree of noise addition. The expectation calculation for x, c, c', ∈, t is to take the average over the data distribution. ‖·‖ 2 is the L2 norm, which measures the gap between the predicted image and the real image. During the diffusion generation process, the control signal is injected into each feature layer of the model multiple times, which affects both the initial stage of generation and provides structural constraints in the middle and late stages, thereby ensuring that the generation result meets the set conditions.

5. The stable diffusion method according to claim 1, characterized in that It also includes: Regarding the generation of the crack mask, a crack mask generation method based on the combination of expert hand-drawn and existing industrial datasets is proposed to ensure the integrity and accuracy of crack features. Specifically, it includes: The sources of the crack mask include crack defect images in the existing industrial dataset and crack images hand-drawn by industry experts. Among them, the crack mask hand-drawn by experts provides a reliable supplement for crack modeling. Especially in the case where SAM segmentation may fail, it can effectively make up for the deficiency of dial area segmentation, ensure the integrity and accuracy of crack features, thereby reducing the risk of relying on segmentation, and further improving the robustness and generalization performance of the generation model. By combining expert hand-drawn and existing industrial datasets, the generated crack mask can more accurately reflect the morphology and distribution characteristics of actual meter cracks, providing high-quality input data for subsequent crack feature modeling.

6. A stable diffusion device for few-shot meter defect image generation, characterized in that, The device includes: The S1 module is used to construct a dataset containing normal substation meter images and meter images with crack defects, covering various meter types. All images are preprocessed, uniformly cropped to a resolution of 512×512, and the defect category and location information are labeled; The S2 module is used to fine-tune the pre-trained Stable Diffusion model by the DreamBooth method, bind the unique Chinese identifier [meter] to the meter image, different from the traditional method using English identifiers, realize the embedding of meter knowledge, establish a strong association between the Chinese identifier and the meter image features, design a dual-objective loss function, balance the learning of new features and the retention of prior knowledge, and ensure that the model retains the general generation ability when generating meter defect images; The S3 module proposes a crack feature modeling module. It extracts the meter line drawing through an adaptive edge detection algorithm, retains the structural features of the dial contour and key components, ensures the integrity and rationality of the generated image in terms of geometric structure, performs high-precision dial area segmentation on the input image based on SAM, generates a dynamic constraint map, and further fuses the crack mask images provided by expert hand-drawn and industrial datasets. The crack is accurately restricted to the effective area of the dial through the constraint image, avoiding the abnormal phenomenon of cracks appearing in non-functional areas; The S4 module adopts a hypernetwork mechanism to dynamically adjust the weight distribution during the generation process, ensuring the consistency and diversity of the generated images in terms of morphology and position. It freezes the weights of the Stable Diffusion main model, constructs a trainable hypernetwork branch, encodes the control image into a latent representation, injects it layer by layer into the diffusion process, and constrains the crack morphology of the generated images.

7. An electronic device, characterized in that, The electronic device includes: at least one memory and at least one processor; the memory stores a program, and the processor calls the program stored in the memory, and the program is used to implement the stable diffusion method for few-shot meter defect image generation according to any one of claims 1-5.

8. A storage medium, characterized in that, The computer-executable instructions are stored in the storage medium, and the computer-executable instructions are used to execute the stable diffusion method for few-shot meter defect image generation according to any one of claims 1-5.

Citation Information

Patent Citations

  • Unmanned aerial vehicle visual inspection and recognition method for crane complex steel structure surface defects

    CN113744270A

  • Power transformation defect data small sample expansion method and system based on diffusion model

    CN117115818A

  • Part surface defect generation and embedding method based on diffusion model

    CN117671429A

  • Small sample surface defect image generation method and system based on feedback reinforcement learning

    CN117710349A

  • Display panel crack defect generation method based on stable diffusion model

    CN118570136A

Cited By

  • Steel rail surface defect detection method based on diffusion model driving

    CN121169825A