Overhead transmission line fault image sample expansion method based on Diffusion technology
The generated overhead transmission line fault image samples are solved through Diffusion technology, which solves the problem of insufficient samples in the inspection database, improves the training effect and adaptability of the deep learning model, and the generated sample diversity and clarity meets the needs of specific scenarios.
Patent Information
- Application Number
- CN202510344821.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
The lack of sufficient high-quality overhead transmission line failure samples in the existing inspection databases, resulting in poor training of deep learning models, and traditional data augmentation methods fail to effectively introduce new fault features or simulate complex scenarios.
Using the image generation method based on Diffusion technology, the image generation model is designed through U-Net network architecture, label embedding module, time-step embedding module and multi-scale feature fusion technology, and GPU accelerated computing is used to perform forward diffusion simulation and reverse diffusion optimization to generate faulty image samples that meet the conditions.
The generated faulty image samples have high diversity and clarity, which can effectively improve the training effect of deep learning models, enhance the robustness and flexibility of the model, adapt to specific scenario needs, and provide more training data support.
Smart Images

Figure CN120279354A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of power inspection and computer vision, and specifically to a method for augmenting fault image samples of overhead transmission lines based on Diffusion technology. Background Art
[0002] Overhead transmission lines are a key component of the power transmission network, and their stability and safety are directly related to the normal operation of the power grid and the reliable power supply for social production and life. In recent years, with the rapid development of drone technology, high-definition imaging technology, and computer vision technology, using drones to inspect overhead transmission lines has gradually become an efficient fault monitoring method.
[0003] However, in actual operation, the occurrence probability of transmission line faults is relatively low, especially for some special or rare defect types, which account for a very small proportion in the collected inspection images. This phenomenon results in a lack of sufficient high-quality samples in the existing inspection database to support the training of deep learning models, and the effectiveness of deep learning algorithms usually depends on large-scale, accurately labeled, and diverse datasets. In the prior art, to solve the problem of insufficient samples, researchers have tried sample augmentation methods based on data enhancement, including image rotation, scaling, flipping, color adjustment, etc. However, these traditional data enhancement methods only perform simple transformations on the original images and fail to introduce new fault features or simulate complex scenarios, so the augmentation effect is relatively limited. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a method for augmenting fault image samples of overhead transmission lines based on Diffusion technology, which solves the problem of insufficient fault image sample quantity and resulting differences in recognition.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for augmenting fault image samples of overhead transmission lines based on Diffusion technology, which specifically includes the following steps:
[0006] Perform image enhancement, filtering, denoising, and picture screening on the captured pictures, calibrate the processed pictures, and simultaneously screen out training data and test data;
[0007] Based on the U-Net network architecture for analysis, by introducing a label embedding module, a time step embedding module, and a multi-scale feature fusion technology, design an image generation model to obtain a Difusion model architecture;
[0008] Deploy a training script on a high-performance computing platform, use GPU to accelerate the calculation process, perform forward diffusion simulation and reverse diffusion optimization, and simultaneously optimize the loss function;
[0009] When the model reaches the expected performance, stop training, save the best weight parameters for the subsequent sample generation stage, and customize the generation of new images with specific attributes according to requirements to obtain image samples;
[0010] Evaluate the obtained image samples for sharpness, diversity, and conditional consistency, determine the image samples that meet the standards, and integrate them into the existing dataset to generate augmented information.
[0011] As a further solution of the present invention, the specific method for screening the training data and test data is as follows:
[0012] Obtain the taken pictures and perform image enhancement, filtering, denoising, and picture screening operations to obtain a dataset. Then, calibrate the defect types and scenes of the dataset, and randomly select 80% of the prepared dataset as training data, and the remaining 20% as test data.
[0013] As a further solution of the present invention, the specific method for designing the image generation model to obtain the Difusion model architecture is as follows:
[0014] In the first step, determine the overall framework of the model, specifically including determining the input part, the core network part, the embedding module part, and the output part;
[0015] In the second step, construct the U-Net architecture, specifically including determining the encoder part, the decoder part, and the skip connection;
[0016] In the third step, determine the label embedding module, convert the label information into a high-dimensional necklace through the embedding layer transformation, and adopt the embedding network to map the discrete label to the continuous vector space through non-linear transformation;
[0017] In the fourth step, determine the time embedding module.
[0018] As a further solution of the present invention, the specific method for performing the forward diffusion simulation is as follows:
[0019] The forward diffusion process is represented in a mathematical form where x0 represents the original clear image, x t is the image at time step t, ε is the random noise of the standard normal distribution, and α t is the parameter controlling the noise ratio.
[0020] As a further solution of the present invention, the specific method for performing the reverse diffusion optimization is as follows:
[0021] The reverse denoising process is represented in a mathematical form where ò θ (x t,t,y) is the noise value predicted by the model, indicating the image x at time step t t The noise component included, ò is the actual noise value added, x t is the input image at the current time step, t is the time step, and y is the conditional label.
[0022] As a further solution of the present invention, the specific method of optimizing the loss function is:
[0023] The formula of MSE loss function is
[0024] L MES is the mean square error (MSE) loss function, E represents expectation, which is an operation to average the variables that meet certain conditions. By optimizing the loss function, Represented as the average treatment of the variable in the time interval, ∈-∈ θ represents the model prediction function, It is the square of the L2 norm, which is used to calculate the square of the distance between the predicted value and the true value.
[0025] As a further solution of the present invention, the specific method of obtaining the image sample is:
[0026] When the training phase of the model is completed and the expected performance is achieved, the optimal weight parameters are saved, and the model is comprehensively evaluated through the indicators on the validation set. If the indicators meet the set standards, the parameter weights are saved and samples are generated. The mathematical expression of the generation process is
[0027] Among them, x t represents the image at the current time step, θ (x t ,t,y) is the noise component predicted by the model, α t is a parameter that controls the noise ratio.
[0028] As a further solution of the present invention, the specific method of generating the extended information is:
[0029] The image samples are evaluated for clarity, diversity, and conditional consistency. For clarity, PSNR and SSIM are calculated by automated tools to ensure that the generated images are visually indistinguishable from the real images. For diversity, the feature distribution of the generated samples is statistically analyzed to verify whether the model can cover various changes in actual scenes. For conditional consistency, the semantic matching between the generated samples and the input labels is compared to ensure that the generated images accurately reflect the conditions specified by the user.
[0030] Screen image samples that meet the criteria through clarity, diversity, and conditional consistency evaluations, and integrate them into the existing dataset to generate augmented information.
[0031] The present invention provides a method for augmenting fault image samples of overhead transmission lines based on Diffusion technology. Compared with the prior art, it has the following beneficial effects:
[0032] The present invention conducts strict screening and annotation work on the data. Each image is equipped with accurate labels, including information such as the type and location of the fault. Subsequently, the data is standardized. The standardization process not only improves the consistency of the data but also provides a stable data distribution for the input layer of the model, ensuring the convergence and robustness of model training.
[0033] Also, through the powerful generation ability of the diffusion model, it has the performance of diversity control in the generation of defect pictures. A user interaction mechanism can be introduced during the process of generating samples. According to specific requirements, the conditional labels or generation strategies can be dynamically adjusted. Users can also further refine their requirements based on the preliminary effects of the generated images to generate samples that are more in line with expectations. This interactive generation process greatly enhances the flexibility and practicality of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a flowchart of the method steps of the present invention;
[0035] Figure 2 It is a structural diagram of the U-net model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0037] Please refer to Figure 1 and Figure 2 , the present application provides a method for augmenting fault image samples of overhead transmission lines based on Diffusion technology. The method specifically includes the following steps:
[0038] Step S1, production of the dataset.
[0039] Since the present invention is applied to fault maps of overhead transmission lines with specific defects, aerial images of different scenes are taken at this location;
[0040] After shooting, preprocess the captured images, including operations such as image enhancement, filtering, denoising, and image screening, to ensure the consistency and quality of the input data; finally, calibrate the defect types and scenarios of the dataset, and randomly select 80% of the prepared dataset as the training data for the overhead line defect sample generation algorithm based on diffusion, and the remaining 20% as the test data for testing and verifying the defect sample generation part in the present invention.
[0041] Step S2: Construct a Difusion model architecture. Adopt a generation method based on a label-conditioned diffusion model. Based on the U-Net network architecture, design an efficient and controllable image generation model by introducing a label embedding module, a time step embedding module, and a multi-scale feature fusion technology; specifically, this architecture can generate high-quality, diverse, and highly consistent fault images with labels from noise through a gradual diffusion and denoising process.
[0042] The first step is to determine the overall framework of the model
[0043] The core of this model architecture is the U-Net network. Based on a symmetric encoder-decoder structure, with a conditional control mechanism of label embedding and time step embedding, the model can flexibly adjust the output features during the image generation process to meet the requirements of specific scenarios and conditions.
[0044] Determine the input part
[0045] Noise image x t : Represents the image at a certain time step during the diffusion process. Time step t: The time step in the current diffusion process, used to guide the generation process. Label y: Represents the specific conditions (such as fault type, location, etc.) of the target generated image.
[0046] Determine the core network part
[0047] Encoder: Gradually extract the multi-scale features of the image, and compress the input image into a feature representation with a lower resolution through downsampling (such as convolution and pooling operations).
[0048] Decoder: Gradually restore the image, restore the resolution of the image through upsampling operations, and at the same time use skip connections to fuse the features extracted by the encoder.
[0049] Residual module: Each layer of the network contains a residual module, which is used to extract fine-grained features and enhance the expressive ability of the model.
[0050] Determine the embedding module part.
[0051] Label embedding: Convert discrete label information into a high-dimensional vector and inject it into the network to ensure that the generated image matches the target conditions.
[0052] Time step embedding: Convert the time step t into a vector to provide state information in the diffusion process and guide the generation process.
[0053] Determine the output part.
[0054] Finally, the restored image x is output through the decoder t-1 , as the next input for the reverse diffusion process until a clear image x0 is generated.
[0055] The second step is to construct a U-Net architecture
[0056] U-Net is a deep learning network widely used in image generation and segmentation, characterized by a symmetric encoder-decoder structure. The encoder gradually extracts the features of the input image, and the decoder gradually restores the image. In the present invention, each residual block of U-Net is enhanced by label embedding technology, and the embedded label information can effectively guide the model to generate qualified images.
[0057] Determine the encoder part
[0058] The encoder consists of multiple convolutional layers and pooling layers, which are used to gradually reduce the resolution of the image and extract multi-scale features. Each layer of the encoder contains two convolutional operations (Conv2D) and one pooling operation (MaxPooling), and the ReLU activation function is used to increase the non-linear expression ability. During the encoding process, the number of channels of the feature map increases layer by layer to adapt to more complex feature representations.
[0059] Determine the decoder part
[0060] The decoder gradually restores the spatial resolution of the image through upsampling layers and combines skip connections to obtain the multi-scale features extracted in the encoder. Each layer of the decoder corresponds to a specific time step of the reverse process, and the features are further refined through convolutional operations.
[0061] Determine the skip connections
[0062] The skip connections directly transfer the feature maps of the encoder to the corresponding layers of the decoder, avoiding information loss caused by feature dimensionality reduction. This design can effectively retain the detailed information of the input image and enhance the model's ability to capture image edges and complex features.
[0063] The third step is to determine the label embedding module.
[0064] The label information is converted into a high-dimensional vector through the embedding layer, denoted as y. In each time step t of the diffusion model, the label embedding y is added to the current feature map to ensure that the generated image is consistent with the specified label.
[0065] The label embedding method uses an embedding network to map discrete labels to a continuous vector space through non-linear transformation.
[0066] Step 4: Determine the time step embedding module.
[0067] The time step in the diffusion process is encoded as an embedding vector, denoted as t. This vector is generated through positional encoding and can reflect the state information of the current time step.
[0068] In this network, time step embedding is injected into the feature extraction process of the U-Net network to ensure that the model can accurately capture the dynamic changes in the diffusion process.
[0069] Step S3: Model training process. The specific training process of the model is based on constructing an efficient script that meets the target task. When running the training script on a high-performance computing platform, making full use of the advantages of GPU parallel computing can significantly improve the training efficiency. The GPU has powerful matrix computing capabilities and can quickly process convolution operations and gradient updates, which is particularly crucial in complex deep learning models. The design of the training script needs to consider the multi-stage optimization process of the model, including forward diffusion simulation and reverse diffusion optimization. Forward diffusion simulation gradually adds Gaussian noise to transform a clear image into a random noise image, and this process provides multi-stage intermediate data samples for the model. Reverse diffusion optimization is the focus of model training, and the goal is to gradually denoise and restore the image to generate high-quality samples that meet the input conditions. The specific training process is as follows:
[0070] Forward diffusion process (During the forward diffusion process of model training, the process of gradually adding noise to a clear image until it is completely noised is simulated. The main purpose of forward diffusion is to generate a series of intermediate image data to provide sample support for the training of the reverse process), which is represented mathematically as:
[0071]
[0072] where x0 represents the original clear image, x t is the image at time step t, ε is the random noise of the standard normal distribution, and α t is the parameter that controls the proportion of noise. In this formula, as the time step t increases, the value of α t gradually decreases, and the noise component in the image gradually increases until x t completely becomes random noise. Specifically, the forward diffusion process is implemented by means of step-by-step sampling to ensure that the generated intermediate images are reasonably distributed and meet the target characteristics. In actual training, these intermediate images x t along with their corresponding time steps t and conditional labels y are used as input data for the model to guide the model to learn the evolution law of image features during the diffusion process.
[0073] The reverse denoising process (the reverse diffusion process is the core part of model training, aiming to gradually denoise a completely noisy image and finally generate a clear image that meets the conditions), which is represented in mathematical form as follows:
[0074]
[0075] where, ò θ (x t , t, y) is the noise value predicted by the model, representing the noise component contained in the image x at time step t t , ò is the actually added noise value, x t is the input image at the current time step, t is the time step, and y is the conditional label. By gradually reducing the noise, the reverse diffusion process can gradually restore the details of the image.
[0076] Specifically, in the reverse process, the inputs to the model include the time step t, the noisy image x t and the conditional label y. Among them, the time step information is transformed into a vector through an embedding module, and the label information generates a feature representation through a label embedding network. These information are fused with the multi-scale features of the network, enabling the model to accurately predict the noise at the current time step and generate intermediate images that meet the conditions.
[0077] During the training process, the design of the loss function directly affects the learning ability and generation performance of the model. The present invention adopts the mean squared error (MSE) loss function as the main optimization objective, aiming to minimize the difference between the noise value predicted by the model and the actual noise. The formula of the MSE loss function is:
[0078]
[0079] L MES is the mean squared error (MSE) loss function, E represents the expectation, which is an operation of taking the average of variables that meet certain conditions here. By optimizing this loss function, represents the variable average processing over the time interval, ∈ - ∈ θ represents the model prediction function, is the square of the L2 norm, used to calculate the square of the distance between the predicted value and the true value. The model can accurately learn the noise distribution characteristics, thereby improving the quality of the denoising process;
[0080] An important part of the training process is to dynamically monitor the changing trend of the loss function. By recording and analyzing the training loss in real time, the training dynamics of the model can be discovered, such as the convergence speed, signs of overfitting, or gradient explosion problems. When it is found that the training loss tends to stabilize or no longer decreases significantly after multiple rounds of iteration, the training strategy can be adjusted appropriately. For example, the model parameters can be finely tuned by reducing the learning rate to make it closer to the optimal value. Adopting a segmented learning rate scheduling strategy (such as cosine annealing or exponential decay) is one of the common optimization methods in the present invention, which can stabilize the convergence process in the later stage of training.
[0081] During the training process, every several epochs, the images generated by the model on the validation set are compared with the real images, and the peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) are calculated. PSNR measures the overall clarity of the generated images, while SSIM evaluates the similarity of the images at the structural level. Through the changing trends of these metrics, it can be judged whether the current generation performance of the model meets the expectations, and the training strategy can be adjusted accordingly.
[0082] In the specific implementation, the present invention adopts an early stopping strategy (EarlyStopping) to prevent the model from overfitting. When the performance on the validation set no longer improves significantly after multiple rounds of iteration, the training process automatically stops, and the current best model parameters are saved. In addition, to further improve the training efficiency, the model training adopts a distributed computing framework, which distributes the data and computing tasks to multiple GPUs for parallel processing. This not only speeds up the training speed but also enables handling larger-scale datasets.
[0083] Step S4, generate new samples. "Generating new samples" is a key link in generating high-quality fault image samples for specific requirements based on the trained diffusion model. The goal of this process is to generate images with specific attributes through conditional control to meet the requirements of sample diversity and specificity in different application scenarios. The generation process is not only a test of the model performance but also an important manifestation of the practical application value of the present invention.
[0084] When the training stage of the model is completed and the expected performance is achieved, the best weight parameters are saved to ensure that the optimal effect during training can be reproduced during the subsequent sample generation process. At the end of the training stage, the model is comprehensively evaluated through metrics on the validation set (such as peak signal-to-noise ratio PSNR, structural similarity index SSIM, and semantic consistency score between the generated image and the real image).
[0085] If these metrics meet the set standards, it indicates that the model already has stable generation ability. At this time, the state of the model is fixed, and its parameter weights are saved as the initial basis for the generation stage.
[0086] The input in the sample generation stage includes three main parts: an initial noise image, a specified conditional label, and a time step control parameter. The initial noise image is a random noise image generated by randomly sampling from a standard normal distribution, which provides the starting point for the sample generation process. This randomness ensures the diversity of the generated samples. Even under the same conditional label, each generated image still has subtle differences, enriching the content of the final sample library. The conditional label is provided by the user and is used to define specific attributes of the generated image, such as the fault type, fault location, or other environmental characteristics. This conditional label is converted into a high-dimensional feature vector through an embedding network and fused with the multi-scale features of the model during the generation process to ensure that the generated image meets the specified conditions. The time step control parameter is used to guide the gradual restoration of the diffusion reverse process, ensuring the stability and logical consistency of the generation process.
[0087] The core of the generation stage is the reverse diffusion process of the model. In this process, the initial noise image is gradually denoised to restore a clear fault image sample. Through the conditional control mechanism of the model, the generated image is highly consistent with the input label. At each step of the model's prediction, according to the current time step, conditional label, and input features, the noise component is estimated and removed from the image, gradually restoring the meaningful image content. The mathematical expression of the generation process is as follows:
[0088]
[0089] where x t represents the image at the current time step, ò θ (x t , t, y) is the noise component predicted by the model, and α t is a parameter that controls the proportion of noise. Through this formula, the noise in the image is gradually removed, and the content of the defect image to be generated gradually appears.
[0090] With the powerful generation ability of the diffusion model, the present invention has the performance of diversity control in defect image generation. In practical applications, although the conditional label can specify the overall characteristics of the image, the randomness in the generation process provides the ability to diversify the samples.
[0091] For example, for the "broken wire" fault type, the user can observe changes in different environmental backgrounds, wire break positions, or fracture morphologies among multiple generation results. This diversity is achieved through two ways in the model design: First, the randomness of the initial noise image ensures the basic changes in the generation results; Second, the random perturbations during the diffusion process can introduce differences in details.
[0092] In the present invention, a user interaction mechanism can also be introduced during the process of generating samples to dynamically adjust the conditional label or generation strategy according to specific requirements.
[0093] For example, when the user needs a rare fault scenario under a specific background, it can be achieved by modifying specific attributes of the condition tags. In addition, the user can further refine the requirements based on the preliminary effect of the generated image to generate samples that better meet the expectations. This interactive generation process greatly enhances the flexibility and practicality of the present invention.
[0094] Step S5, Effect verification and integration
[0095] The generated image samples need to undergo strict quality assessment to ensure that they can be used for subsequent artificial intelligence model training or power inspection auxiliary decision-making. The evaluation metrics include image clarity, diversity, and condition consistency. Among them, clarity is calculated by automated tools for PSNR and SSIM to ensure that the generated images are not significantly different from real images visually. Diversity is verified by statistically analyzing the feature distribution of the generated samples to check whether the model can cover various variations of the actual scenarios. Condition consistency is ensured by comparing the semantic matching degree between the generated samples and the input tags to ensure that the generated images accurately reflect the conditions specified by the user.
[0096] When the generated samples pass the quality verification, they can be integrated into the existing dataset for enhancing the training of the deep learning model. Since the generated samples are of high diversity and clarity, they can significantly improve the model's performance and generalization ability in dealing with complex scenarios. Especially in rare fault types or marginal scenarios, the contribution of these samples is particularly significant.
[0097] In addition, the generated samples can also be used for the auxiliary decision-making of power inspection personnel. For example, through a rich fault image library, it can help the inspection personnel quickly identify and judge the on-site fault types.
[0098] For some data in the above formula, only their numerical values are taken for calculation without substituting parameter units, and the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.
[0099] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.
Claims
1. An overhead transmission line fault image sample augmentation method based on Diffusion technology, characterized in that The method specifically includes the following steps: Perform image enhancement, filtering, denoising, and image screening on the captured images, calibrate the processed images, and simultaneously screen to obtain training data and test data; Based on the U-Net network architecture for analysis, by introducing a label embedding module, a time step embedding module, and a multi-scale feature fusion technology, design an image generation model to obtain the Difusion model architecture; Deploy the training script on a high-performance computing platform, utilize GPU to accelerate the calculation process, and perform forward diffusion simulation and reverse diffusion optimization, while optimizing the loss function; Stop training when the model reaches the expected performance, save the best weight parameters for the subsequent sample generation stage, and customize and generate new images with specific attributes according to requirements to obtain image samples; Evaluate the clarity, diversity, and conditional consistency of the obtained image samples, determine the image samples that meet the standards, and integrate them into the existing dataset to generate augmented information.
2. The method for expanding the fault image samples of the overhead transmission line based on the Diffusion technology according to claim 1, characterized in that The specific method for screening to obtain training data and test data is as follows: Obtain the captured images and perform image enhancement, filtering, denoising, and image screening operations to obtain a dataset. Then, calibrate the defect types and scenarios of the dataset, and randomly select 80% of the prepared dataset as training data, and the remaining 20% as test data.
3. The method for expanding the fault image samples of the overhead transmission line based on the Diffusion technology according to claim 1, characterized in that, The specific method for designing the image generation model to obtain the Difusion model architecture is as follows: In the first step, determine the overall model framework, specifically including determining the input part, the core network part, the embedding module part, and the output part; In the second step, construct the U-Net architecture, specifically including determining the encoder part, the decoder part, and the skip connection; In the third step, determine the label embedding module, convert the label information into a high-dimensional necklace through an embedding layer transformation, and adopt an embedding network to map the discrete label to a continuous vector space through a non-linear transformation; In the fourth step, determine the time embedding module.
4. The method for expanding the fault image samples of overhead transmission lines based on the Diffusion technology according to claim 1, wherein, The specific method for performing forward diffusion simulation is as follows: The forward diffusion process is represented in a mathematical form where x0 represents the original clear image, and x t is the image at time step t, ε is the random noise of the standard normal distribution, and α t is the parameter that controls the proportion of noise.
5. The method for expanding the fault image samples of overhead transmission lines based on Diffusion technology according to claim 1, wherein The specific method for performing reverse diffusion optimization is as follows: The reverse denoising process is represented in a mathematical form where, ò θ (x t , t, y) is the noise value predicted by the model, representing the noise component contained in the image x t at time step t, ò is the actually added noise value, x t is the input image at the current time step, t is the time step, and y is the conditional label.
6. The method for expanding the fault image samples of the overhead transmission line based on the Diffusion technology according to claim 1, wherein, The specific method for optimizing the loss function is as follows: The formula for the MSE loss function is L MES is the mean squared error (MSE) loss function. E represents the expectation, which is an operation of taking the average of variables satisfying certain conditions here. By optimizing this loss function, represents the variable averaging process over a time interval, ∈ - ∈ θ represents the model prediction function, is the square of the L2 norm, which is used to calculate the squared distance between the predicted value and the true value.
7. The method for expanding the fault image samples of the overhead transmission line based on the Diffusion technology according to claim 1, characterized in that The specific method for obtaining the image samples is as follows: After the training phase of the model is completed and the expected performance is achieved, the best weight parameters are saved. The model is comprehensively evaluated through the metrics on the validation set. If the metrics meet the set standards, the parameter weights are saved and sample generation is performed. At the same time, the mathematical expression for the generation process is where x t represents the image at the current time step, and ò θ (x t , t, y) is the noise component predicted by the model, and α t is the parameter that controls the proportion of noise.
8. The method for expanding the fault image samples of overhead transmission lines based on the Diffusion technology according to claim 1, characterized in that The specific method for generating augmented information is as follows: Evaluate the clarity, diversity, and conditional consistency of the image samples. Among them, clarity is calculated by an automated tool for PSNR and SSIM to ensure that the generated images have no significant difference from real images visually. Diversity is verified by statistically analyzing the feature distribution of the generated samples to verify whether the model can cover various variations of the actual scenario. Conditional consistency is ensured by comparing the semantic matching degree between the generated samples and the input labels to ensure that the generated images accurately reflect the conditions specified by the user; Screen the image samples that meet the standards through clarity, diversity, and conditional consistency evaluation, and integrate them into the existing dataset to generate augmented information.
Citation Information
Cited By
Diffusion model reasoning acceleration method and system
CN121706992A