Image restoration method, system, device and medium

A single-step denoising and restoration method that integrates a multi-task degradation estimation network and a visual language model solves the problem of poor image restoration in outdoor power systems under severe weather conditions, achieving efficient and accurate image restoration and equipment status recognition.

CN121504769APending Publication Date: 2026-02-10GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511871542.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

In existing technologies for outdoor power systems, image restoration is poor due to a combination of factors, including image degradation caused by severe weather, which affects the accuracy of power equipment status identification.

Method used

A multi-task degradation estimation network and a visual language model are fused together. Image restoration is performed through a single-step denoising and inpainting model. Conditional embedding vectors guide feature modulation. The encoder and decoder are combined to perform efficient dimensionality reduction and feature fusion, thereby restoring image details.

Benefits of technology

It achieves accurate perception and efficient restoration of various complex weather degradations, improves image restoration speed and quality, meets real-time requirements, and enhances the accuracy and reliability of power equipment status identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121504769A_ABST
    Figure CN121504769A_ABST
Patent Text Reader

Abstract

The invention discloses an image restoration method, system and device and a medium, and belongs to the field of power grids, and the method comprises the steps: inputting a to-be-restored image into a degradation estimation network, so as to estimate the degradation type and degradation degree of each power device under the current weather condition, and obtaining a multi-dimensional degradation vector; inputting the to-be-restored image into the visual language model to obtain a text prompt of the weather condition; fusing the text prompt and the degradation vector to obtain a conditional embedding vector; inputting the to-be-restored image into an encoder for dimension reduction processing to obtain an initial encoding feature and an intermediate feature; inputting the initial coding feature and the conditional embedding vector into a single-step denoising repair model, and performing feature modulation according to the conditional embedding vector to obtain a target coding feature; and inputting the target coding features into a decoder, and performing feature fusion through the intermediate feature information to obtain a target image, so that by implementing the method, the restoration effect and the restoration speed can be improved, and the safety of a power grid is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of power grid, and particularly to an image restoration method, system, device and medium. BACKGROUND

[0002] In the intelligent inspection and safety monitoring of outdoor power systems, images of power equipment are collected by using unmanned aerial vehicles or fixed cameras, and the state of the power equipment is identified according to the images, which is a key means to ensure the stable operation of the power grid. However, heavy rain, heavy fog, snow and other bad weather can cause serious mixed degradation of the images, which directly affects the image recognition and analysis steps, resulting in a significant decrease in recognition accuracy.

[0003] Existing methods usually use a rain removal model or a de-fogging model specially trained to restore the image. However, for composite degraded images in which multiple degradation factors exist simultaneously, the existing method can cause the degradation factors such as rain and snow stripes, fog and haze scattering to be mixed with the structural features of the equipment, resulting in problems such as artifacts, loss of details or residual degradation in the restored image, poor image restoration effect, and further affecting the accuracy of subsequent power equipment state identification. SUMMARY

[0004] The present application provides an image restoration method, system, device and medium, which can solve the problem of poor image restoration effect.

[0005] The present application provides an image restoration method, comprising: Obtaining a to-be-restored image in a power scene, inputting the to-be-restored image into a preset multi-task degradation estimation network to estimate the image degradation type and image degradation degree of each power equipment under the current weather condition, obtaining a multi-dimensional degradation vector, and inputting the to-be-restored image into a preset visual language model to obtain text prompt information for representing the current weather condition; Fusing the text prompt information and the multi-dimensional degradation vector to obtain a conditional embedding vector; inputting the to-be-restored image into a preset encoder for dimension reduction processing to obtain an initial encoding feature and intermediate feature information generated in the encoding process; Inputting the initial encoding feature and the conditional embedding vector into a preset single-step denoising restoration model to modulate the feature processing process of each layer according to the conditional embedding vector to obtain a target encoding feature; Inputting the target encoding feature into a preset decoder to fuse the features through the intermediate feature information to obtain a target image.

[0006] The embodiment of the present application features unified and accurate perception of multiple composite weather degradation; the to-be-repaired image is input into a visual language model to obtain text prompt information, high-level semantic guidance is provided, and the repair process is more in line with semantic logic; the text prompt information and the multi-dimensional degradation vector are fused to obtain a conditional embedding vector, which provides a strong guidance signal for the repair model; the to-be-repaired image is input into an encoder to obtain initial encoding features and intermediate feature information, the encoder compresses the high-dimensional image into a low-dimensional latent space, greatly reduces the subsequent core repair operation amount, reduces the computational complexity, improves the processing speed, and at the same time, the intermediate features provide a basis for subsequent detail recovery; the initial encoding features and the conditional embedding vector are input into a single-step denoising repair model to realize efficient and accurate repair, the single-step denoising greatly improves the repair speed, meets the real-time requirement, the conditional modulation ensures that the repair process is strictly guided by the composite condition, and the repair effect is improved; the target encoding features are input into a decoder, the features are fused through the intermediate feature information to obtain a target image, the texture and edge details that may be lost in the compression and repair process are effectively recovered, high-quality reconstructed image details are realized, and the high fidelity of the repaired image is ensured.

[0007] Overall, the embodiment provides an efficient, accurate and adaptive outdoor power scene bad weather image repair method, which can uniformly perceive and quantify multiple composite weather degradation, and accurately guide an efficient single-step diffusion repair model using the fused composite condition, greatly improve the repair speed to meet the real-time requirement, ensure the fidelity of the key details of the repaired image through the dynamic connection mechanism, improve the image repair effect, and significantly improve the accuracy and reliability of subsequent power equipment intelligent inspection and state recognition.

[0008] Further, the single-step denoising repair model includes a conditional projection module and a denoising repair network, and the feature modulation of the feature processing process of each layer is specifically as follows: In the feature processing process of each residual unit of the denoising repair network, the intermediate encoding features are obtained by performing convolution processing on the current input features through the convolution layer in the current residual unit; The intermediate encoding features are standardized to obtain standardized features; The conditional projection module determines the current modulation parameter of the current residual unit according to the conditional embedding vector; The standardized features are calculated according to the current modulation parameter to obtain the initial output features of the current residual unit, and the current input features and the initial output features are connected in residual to obtain the current output features of the current residual unit.

[0009] In this way, through layer-by-layer conditional normalization and residual learning, top-down and fine-grained feature control is realized, which ensures that the condition information can effectively affect each calculation of the restoration network, thereby realizing high controllability and precision of the restoration process.

[0010] Further, the multi-task degradation estimation network comprises convolutional layers, multi-layer residual units and multiple regression heads connected in sequence, and the degraded image is input into the preset multi-task degradation estimation network to estimate the image degradation type and image degradation degree of each power equipment under the current weather condition, and a multi-dimensional degradation vector is obtained, specifically: Through the convolutional layers and the residual units, feature extraction is performed, and in the last two residual units, channel attention mechanism and spatial attention mechanism are combined for feature processing to improve the extraction ability of local degradation features, and multi-scale degradation features are obtained; The multi-scale degradation features are input into the parallel regression heads to obtain rain streak intensity, fog concentration, blur degree, illumination degree and contrast degree, respectively; The rain streak intensity, the fog concentration, the blur degree, the illumination degree and the contrast degree are combined to obtain the multi-dimensional degradation vector.

[0011] Thus, an efficient and accurate multi-task degradation perception network is disclosed, which enhances local perception ability by introducing attention mechanism and realizes multi-degradation unified quantization through parallel regression heads, providing reliable and fine degradation state input for the entire restoration system.

[0012] Further, the current modulation parameter includes a scaling parameter and an offset parameter, and the standardized feature is calculated according to the current modulation parameter to obtain the initial output feature of the current residual unit, specifically: The standardized feature is multiplied by the scaling parameter and then added to the offset parameter to obtain the initial output feature.

[0013] Thus, the specific mathematical form of conditional modulation is affine transformation, which is a simple and powerful feature transformation method, and the scaling parameter and the offset parameter control the amplitude and deviation of the feature, respectively, which can flexibly and effectively integrate the condition information into the standardized feature.

[0014] Further, the current modulation parameter of the current residual unit is determined by the condition projection module according to the condition embedding vector, specifically: The condition embedding vector and the position encoding of the current residual unit are combined and input into the condition projection module for nonlinear transformation processing by the condition projection module to obtain the current modulation parameter.

[0015] This makes the generated conditional parameters layer-specific. Position coding distinguishes residual units at different depths, ensuring that unique modulation parameters are generated for each layer, enabling finer-grained layer-level control.

[0016] Furthermore, the image restoration method further includes: Acquire actual degraded image data and simulated degraded image data; The single-step denoising and restoration model is trained based on the actual degraded image data and the simulated degraded image data.

[0017] By combining real and simulated data, this approach effectively overcomes the challenge of obtaining large quantities of high-quality training data for image restoration in outdoor power scenarios under adverse weather conditions. The single-step denoising and restoration model trained using this method not only achieves restoration results close to reality but also possesses excellent generalization performance. It can stably handle various complex and even complex adverse weather conditions not fully covered in the training data, thereby ensuring the reliability and universality of the entire image restoration method in practical applications.

[0018] Furthermore, the image restoration method further includes: Acquire multiple raw images of a power scene; For each of the original images, at least one type of image degradation is simulated by an image processing algorithm, and brightness and contrast are adjusted by data augmentation to obtain simulated degradation image data.

[0019] By simulating degradation and data augmentation, a high-quality and highly diverse training dataset is systematically constructed, which enables the training of complex composite degradation models with better inference performance.

[0020] Another embodiment of the present invention provides an image restoration system, including: a degradation estimation module, a feature encoding module, a denoising restoration module, and a feature decoding module; The degradation estimation module is used to acquire the image to be repaired in the power scenario, input the image to be repaired into a preset multi-task degradation estimation network to estimate the image degradation type and degree of each power device under the current weather conditions, obtain a multi-dimensional degradation vector, and input the image to be repaired into a preset visual language model to obtain text prompt information used to represent the current weather conditions. The feature encoding module is used to fuse the text prompt information with the multidimensional degradation vector to obtain a conditional embedding vector; and to input the image to be repaired into a preset encoder for dimensionality reduction processing to obtain initial encoding features and intermediate feature information generated during the encoding process. The denoising and repair module is used to input the initial encoded features and the conditional embedding vector into a preset single-step denoising and repair model, so as to perform feature modulation on each layer of feature processing according to the conditional embedding vector to obtain the target encoded features; The feature decoding module is used to input the target encoded features into a preset decoder to perform feature fusion through the intermediate feature information to obtain the target image.

[0021] Another embodiment of the present invention provides a terminal device, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements the steps of the image restoration method of the present invention.

[0022] Another embodiment of the present invention provides a computer-readable storage medium item, including: a stored computer program, which, when the computer program is running, controls the device where the computer-readable storage medium is located to perform the steps of the image restoration method of the present invention. Attached Figure Description

[0023] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0024] Figure 1 This is a schematic flowchart of an image restoration method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of an image restoration system provided in an embodiment of the present invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.

[0027] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.

[0028] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0029] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.

[0030] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).

[0031] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.

[0032] See Figure 1 To address the problem of poor image restoration results in existing technologies, an embodiment of the present invention provides an image restoration method, comprising: Step S101: Obtain the image to be repaired in the power scenario, input the image to be repaired into a preset multi-task degradation estimation network to estimate the image degradation type and degree of each power device under the current weather conditions, obtain a multi-dimensional degradation vector, and input the image to be repaired into a preset visual language model to obtain text prompt information used to represent the current weather conditions.

[0033] In this embodiment, a multi-task degradation estimation network analyzes the degraded image to identify noise and blurring features caused by weather, such as raindrop streaks in heavy rain or overall blurring in dense fog, obtaining quantitative representations of various degradation information related to weather. The degradation estimation network employs a lightweight but efficient regressive convolutional neural network, with its overall structure modified based on the ResNet-18 backbone to adapt to the rapid degradation quantization of severe weather images. Simultaneously, the severe weather image is input into a visual language model, which describes the weather conditions in the image to be restored, such as heavy rain or dense fog. The current weather conditions can also be combined with the power scene for description, such as "transmission towers in heavy rain, image blurred," or "transmission towers in heavy rain and fog, image blurred and insufficiently lit," etc.

[0034] Step S102: Fuse the text prompt information with the multidimensional degradation vector to obtain a conditional embedding vector; input the image to be repaired into a preset encoder for dimensionality reduction processing to obtain initial encoding features and intermediate feature information generated during the encoding process.

[0035] In this embodiment, the degradation vector of the image to be repaired and the text prompts obtained from the visual language model are fused. The specific fusion process is as follows: the degradation vector is converted into a conditional vector, while the text prompts are converted into embedding vectors through a pre-trained text encoder. Subsequently, these two vectors are simply concatenated and then integrated through a fully connected layer to obtain the conditional embedding vector. Simultaneously, the image to be repaired is input into the encoder to project it into a low-dimensional latent space, obtaining initial encoded features. The dimension of the latent space is much smaller than the original dimension of the image, thus greatly accelerating the image repair speed of the denoising and repair network. Specifically, the encoder is a four-layer convolutional neural network. Each layer of the convolutional neural network performs a specific convolution operation on the input of the previous layer. This convolution operation continuously reduces the dimension of the input, ultimately projecting the input image to be repaired into the low-dimensional latent space. Furthermore, the intermediate information obtained from each convolution operation is temporarily stored in memory to prepare for the dynamic connection of the subsequent encoder and decoder.

[0036] Step S103: Input the initial encoded features and the conditional embedding vector into a preset single-step denoising and repair model to perform feature modulation on each layer of feature processing according to the conditional embedding vector to obtain the target encoded features.

[0037] In this embodiment, the initial encoded features projected into the latent space are used as the main input to a single-step diffusion denoising and inpainting network. This network is an inpainting network based on a diffusion model, but unlike the traditional multi-step diffusion process, this invention adopts a single-step denoising strategy. A pre-trained noise predictor estimates and removes noise in one step, significantly improving inpainting efficiency and real-time performance. Specifically, the core of the single-step diffusion denoising and inpainting network is a U-Net architecture neural network that combines attention mechanisms and residual connections to handle high-dimensional features in the latent space. Furthermore, based on the guidance parameters generated by the conditional embedding vector, the inpainting network adaptively adjusts its internal processing mechanism to remove degraded noise and reconstruct content in the initial encoded features, obtaining the target encoded features.

[0038] Step S104: Input the target encoded features into a preset decoder to perform feature fusion through the intermediate feature information to obtain the target image.

[0039] In this embodiment, the repaired target encoded features are input into the decoder, which is a four-layer deconvolutional neural network symmetrical to the encoder. Each layer gradually restores the dimension and uses the intermediate information previously stored in memory to perform feature fusion. Finally, the representation of the latent space is mapped back to the original image dimension, and a clear and non-degraded severe weather image repair result is output.

[0040] As an example of an embodiment of the present invention, the multi-task degradation estimation network includes convolutional layers, multiple residual units, and multiple regression heads connected in sequence. The step of inputting the image to be repaired into the preset multi-task degradation estimation network to estimate the image degradation type and degree of each power device under the current weather conditions, and obtaining a multi-dimensional degradation vector, specifically involves: feature extraction through the convolutional layers and each residual unit; and feature processing in the last two residual units by combining channel attention and spatial attention mechanisms to improve the extraction capability of local degradation features, thereby obtaining multi-scale degradation features; inputting the multi-scale degradation features into each parallel regression head to obtain rain stripe intensity, fog concentration, blur degree, illumination level, and contrast degree, respectively; and combining the rain stripe intensity, fog concentration, blur degree, illumination level, and contrast degree to obtain the multi-dimensional degradation vector.

[0041] In this embodiment, the image to be repaired first passes through an initial convolutional layer, for example, using a 7x7 large convolutional kernel with a stride of 2, to quickly downsample and extract the low-frequency degradation trend and basic structural information of the image. Subsequently, the feature map undergoes deep processing through a series of stacked residual unit groups. Each stage gradually expands the receptive field through residual learning, abstracting higher-level semantic features. For example, it enters four residual stages (corresponding to downsampling ratios of ×2, ×4, ×8, and ×16, respectively). The residual blocks in this application differ from the standard ResNet-18; the number of channels in each residual block is compressed to 1 / 2 to 1 / 4 of the original. Furthermore, to improve sensitivity to severe local degradation, a lightweight dual attention module is integrated into the processing flow of the residual units in the deeper stages of the network (e.g., the last two downsampling stages). This module executes in parallel: Channel attention: Through a squeeze-excitation operation, it calculates the weights of each feature channel, adaptively recalibrates channel importance, and enhances feature channels sensitive to specific degradation types (such as rain and fog); Spatial attention: By aggregating channel information, it generates a spatial weight map, adaptively highlighting spatial regions in the image with severe degradation (such as insulators obscured by dense fog or wires covered by dense rain streaks). Features modulated by attention can more concentratedly and strongly represent the degradation of key local areas, thereby significantly improving the network's sensitivity to spatially uneven degradation. To enable the network to perceive multiple degradation types simultaneously, the multi-scale degradation features obtained after the above deep encoding and attention enhancement are used as shared inputs and fed into multiple parallel, structurally identical regression heads to predict the following five consecutive degradation intensity values ​​(all values ​​are normalized to [0,1]): 1. Rain streaks / raindrop intensity; 2. Fog / haze concentration; 3. Blur level (mainly caused by rain / fog + motion); 4. Insufficient illumination (low light index); 5. Global contrast reduction. Each regression head consists of only global average pooling, two fully connected layers, and sigmoid activation. The five heads share backbone features but have independent parameters. Finally, the scalar values ​​output by all the regression heads (e.g., corresponding to rain stripe intensity, fog density, blur level, low light level, and global contrast reduction level, respectively) are assembled in a predetermined order to form a one-dimensional, fixed-length numerical vector. For example, the degradation estimation network outputs a 5-dimensional degradation vector [0.87, 0.92, 0.81, 0.76, 0.68], representing "heavy rain, extremely dense fog, severe blur, moderate low light, and moderate contrast reduction," respectively.

[0042] As an example of an embodiment of the present invention, the single-step denoising and repair model includes a conditional projection module and a denoising and repair network. The feature modulation of each layer's feature processing according to the conditional embedding vector specifically involves: in the feature processing of each residual unit of the denoising and repair network, the current input feature is convolved through the convolutional layer in the current residual unit to obtain intermediate encoded features; the intermediate encoded features are standardized to obtain standardized features; the current modulation parameters of the current residual unit are determined by the conditional projection module according to the conditional embedding vector; the standardized features are calculated according to the current modulation parameters to obtain the initial output features of the current residual unit; and the initial output features are residually concatenated with the current input features to obtain the current output features of the current residual unit.

[0043] In this embodiment, a unified conditional embedding vector is received as input through a conditional projection head. This global conditional signal is decoded and mapped into a set of discrete control parameters, each corresponding to a residual block within the repair network. These parameters aim to provide differentiated control instructions for residual blocks at different locations. Within each residual block, standard forward propagation is performed: first, the input features are convolved to extract new intermediate encoded features; then, these new features are standardized, for example, through instance normalization, to eliminate internal covariate bias and obtain a stable distribution of standardized features; finally, the control parameters of the current residual block are invoked to perform a learnable affine transformation (i.e., scaling and shifting) on ​​the standardized features to accurately inject external conditional information into the feature representation of the current unit, obtaining the initial output features of the current residual unit. The initial output features obtained after conditional modulation are regarded as the main output of the current residual block. In order to ensure that the network does not lose the basic input information during the depth transformation, the main output is added to the original input features of the current residual block element by element. This addition operation forms a short-circuit connection to obtain the final output of the current residual block. This output contains both the new features that are conditionally modulated to repair degradation and the original input information that has not been destroyed, for further processing by subsequent residual blocks.

[0044] As an example of an embodiment of the present invention, the current modulation parameters include scaling parameters and offset parameters. The step of calculating the normalized features based on the current modulation parameters to obtain the initial output features of the current residual unit specifically involves multiplying the normalized features by the scaling parameters and then adding them to the offset parameters to obtain the initial output features.

[0045] In this embodiment, the core of this step is to perform an element-wise affine transformation. Specifically, it involves receiving two data streams: one is a normalized feature map, which has a tensor form of [batch size, number of channels, height, width]; the other is a pair of modulation parameters generated by the conditional projection module, specific to the current residual unit, including a scaling parameter (γ) and an offset parameter (β). Typically, the dimensions of γ and β are the same as the number of channels in the normalized feature map, which is equivalent to independently and conditionally dynamically adjusting the global amplitude of each feature channel. The scaling parameter vector γ is multiplied channel-wise with the normalized feature map; that is, each scalar value in γ is multiplied by all elements (all spatial positions) of the corresponding channel in the normalized feature map. The offset parameter vector β is added channel-wise with the scaled feature map from the previous step; that is, each scalar value in β is added to all spatial elements of the corresponding channel, which is equivalent to independently and conditionally dynamically shifting the global mean of each feature channel. The result of the above multiplication and addition operations is the initial output feature map.

[0046] As an example of an embodiment of the present invention, the step of determining the current modulation parameters of the current residual unit by the conditional projection module based on the conditional embedding vector specifically involves: combining the conditional embedding vector and the position code of the current residual unit, and inputting them together into the conditional projection module, so as to obtain the current modulation parameters by performing nonlinear transformation processing through the conditional projection module.

[0047] In this embodiment, a vector identifier is pre-assigned to each residual unit in the network. This identifier reflects the hierarchical position of the residual unit in the network (e.g., which layer or stage). The conditional embedding vector and the positional encoding vector are concatenated along the feature dimension to obtain a joint input vector. The joint input vector is input into a conditional projection head comprising two fully connected layers. In the first fully connected layer, information from different sources is deeply fused and nonlinearly transformed in a high-dimensional space to extract intermediate features related to the current condition and position. These intermediate features are then input into the subsequent second fully connected layer. The number of neurons in this layer matches the total number of modulation parameters to be predicted. For example, for simultaneously predicting scaling γ and offset β, if the number of feature channels is N, then this layer outputs 2N values. The values ​​of the network output layer are segmented and recombined. For example, the first N values ​​of the output vector are parsed into the scaling parameter vector γ, and the last N values ​​are parsed into the offset parameter vector β. These two vectors constitute the current modulation parameters specific to the current residual unit.

[0048] As an example of an embodiment of the present invention, the image restoration method further includes: acquiring actual degraded image data and simulated degraded image data; and training the single-step denoising and restoration model based on the actual degraded image data and the simulated degraded image data.

[0049] When constructing the training dataset, a large number of real degraded images were first collected from actual outdoor power scenarios, such as photos of substations in heavy rain taken by drones and images of transmission lines covered in heavy snow, to ensure that the data reflects real application scenarios. To increase the diversity of the data, more degraded images were generated through simulation methods, such as adding raindrop, fog, or snowflake effects to clear images of power equipment to simulate different weather conditions. In addition, an open-source visual language model with image understanding capabilities was used to generate descriptive prompts for each image, such as "transmission tower in dense fog, image blurred." These prompts serve as additional guidance information to help the model be more targeted in the restoration process. The model was trained based on all the data to obtain the single-step denoising and restoration model.

[0050] As an example of an embodiment of the present invention, the image restoration method further includes: acquiring multiple original images of a power scene; for each of the original images, simulating the addition of at least one type of image degradation through an image processing algorithm, and adjusting the brightness and contrast through data enhancement processing to obtain simulated degradation image data.

[0051] In this embodiment, firstly, a high-definition, clear image database for power scenarios is constructed. These images are taken by drones or high-definition cameras under good weather conditions, capturing images of various power equipment (such as transmission towers, insulators, substation equipment, and conductors). For each original clear image in the database, a controllable and programmed degradation synthesis operation is performed to generate a paired simulated degradation image. Specifically, by calling a preset image degradation physics model library, one or more combinations of degradation types are randomly selected from a predefined set of degradation types (such as {rain, fog, snow, motion blur, out-of-focus blur, low light, low contrast}) for each base image; for each selected degradation type, a specific intensity parameter is randomly sampled from a preset intensity range. Based on the selected degradation type and its intensity parameter, the corresponding image processing algorithm or generative model is invoked to perform calculations on the original clear image, synthesizing a simulated image that visually conforms to the degradation physics. For example, an atmospheric scattering model formula is used to synthesize a fog image, and a rain pattern overlay algorithm is used to synthesize a rain image. To further enhance the diversity of generated data and the robustness of the model, basic data augmentation transformations are applied to the simulated degraded images (and sometimes the original sharp images). The brightness (e.g., multiplying global pixel values ​​by a factor close to 1) and contrast (e.g., using gamma correction or linear contrast stretching) are randomly fine-tuned to simulate imaging differences under different times and lighting conditions. Small-scale rotations, cropping, and horizontal flipping are also performed randomly to enhance the model's adaptability to changes in target viewpoint and position. Finally, the simulated degraded image is used as input, and its corresponding original sharp image is used as the restoration target, forming an image pair. Simultaneously, textual prompts describing the content of the degraded image are recorded or generated post-processed using a visual language model. All generated image pairs and textual descriptions are systematically organized and stored, forming a simulated degraded image dataset used to train the image restoration model.

[0052] like Figure 2 As shown, based on the above-described method embodiments, an embodiment of the present invention provides an image restoration system 200, including: a degradation estimation module 201, a feature encoding module 202, a denoising restoration module 203, and a feature decoding module 204; The degradation estimation module 201 is used to acquire the image to be repaired in the power scenario, input the image to be repaired into a preset multi-task degradation estimation network to estimate the image degradation type and degree of each power device under the current weather conditions, obtain a multi-dimensional degradation vector, and input the image to be repaired into a preset visual language model to obtain text prompt information used to represent the current weather conditions. The feature encoding module 202 is used to fuse the text prompt information with the multidimensional degradation vector to obtain a conditional embedding vector; and to input the image to be repaired into a preset encoder for dimensionality reduction processing to obtain initial encoding features and intermediate feature information generated during the encoding process. The denoising and repair module 203 is used to input the initial encoded features and the conditional embedding vector into a preset single-step denoising and repair model, so as to perform feature modulation on each layer of feature processing according to the conditional embedding vector to obtain the target encoded features; The feature decoding module 204 is used to input the target encoded features into a preset decoder to perform feature fusion through the intermediate feature information to obtain the target image.

[0053] It is understood that the above system embodiments correspond to the method embodiments of the present invention, and can implement the image restoration method provided by any of the above method embodiments of the present invention.

[0054] It should be noted that the system embodiments described above are merely illustrative, and some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the system embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0055] For ease of description and brevity, the system embodiments of the present invention include all the implementation methods described in the above-described image restoration method embodiments, and will not be repeated here.

[0056] Based on the above-described embodiments of the image restoration method, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the image restoration method of any embodiment of the present invention.

[0057] For example, in this embodiment, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the terminal device.

[0058] The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.

[0059] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.

[0060] Based on the above-described method embodiments, another embodiment of the present invention provides a computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the image restoration method described in any of the above-described method embodiments of the present invention.

[0061] The modules / units integrated in the device / terminal equipment, if implemented as software functional units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0062] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. An image restoration method, characterized in that, include: The image to be repaired in the power scene is acquired, and the image to be repaired is input into a preset multi-task degradation estimation network to estimate the image degradation type and degree of each power device under the current weather conditions, so as to obtain a multi-dimensional degradation vector. The image to be repaired is then input into a preset visual language model to obtain text prompt information to represent the current weather conditions. The text prompt information and the multidimensional degradation vector are fused to obtain a conditional embedding vector; the image to be repaired is input into a preset encoder for dimensionality reduction processing to obtain initial encoding features and intermediate feature information generated during the encoding process. The initial encoded features and the conditional embedding vector are input into a preset single-step denoising and repair model to perform feature modulation on each layer of feature processing according to the conditional embedding vector, so as to obtain the target encoded features. The target encoded features are input into a preset decoder to perform feature fusion using the intermediate feature information, thereby obtaining the target image.

2. The image restoration method as described in claim 1, characterized in that, in, The single-step denoising and restoration model includes a conditional projection module and a denoising and restoration network. Specifically, the feature modulation of each layer's feature processing based on the conditional embedding vector is as follows: In the feature processing of each residual unit of the denoising and repair network, the current input features are convolved by the convolutional layer in the current residual unit to obtain intermediate encoded features; The intermediate encoded features are standardized to obtain standardized features; The conditional projection module determines the current modulation parameters of the current residual unit based on the conditional embedding vector. The normalized features are calculated based on the current modulation parameters to obtain the initial output features of the current residual unit. The initial output features are then residually concatenated with the current input features to obtain the current output features of the current residual unit.

3. The image restoration method as described in claim 1, characterized in that, The multi-task degradation estimation network comprises sequentially connected convolutional layers, multi-layer residual units, and multiple regression heads. The image to be repaired is input into the preset multi-task degradation estimation network to estimate the image degradation type and degree of each power device under the current weather conditions, obtaining a multi-dimensional degradation vector, specifically: Feature extraction is performed through the convolutional layers and each residual unit, and feature processing is performed in the last two residual units by combining channel attention mechanism and spatial attention mechanism to improve the extraction capability of local degradation features and obtain multi-scale degradation features. The multi-scale degradation features are input into each of the parallel regression heads to obtain rain stripe intensity, fog concentration, blurring degree, illumination degree and contrast degree, respectively. The multidimensional degradation vector is obtained by combining the rain stripe intensity, the fog concentration, the blurring degree, the illumination degree, and the contrast degree.

4. The image restoration method as described in claim 2, characterized in that, in, The current modulation parameters include scaling parameters and offset parameters. The calculation of the normalized features based on the current modulation parameters to obtain the initial output features of the current residual unit is specifically as follows: The standardized feature is multiplied by the scaling parameter and then added to the offset parameter to obtain the initial output feature.

5. The image restoration method as described in claim 4, characterized in that, The step of determining the current modulation parameters of the current residual unit by the conditional projection module based on the conditional embedding vector specifically involves: The conditional embedding vector and the position code of the current residual unit are combined and input into the conditional projection module to perform nonlinear transformation processing to obtain the current modulation parameters.

6. The image restoration method as described in claim 1, characterized in that, The image restoration method further includes: Acquire actual degraded image data and simulated degraded image data; The single-step denoising and restoration model is trained based on the actual degraded image data and the simulated degraded image data.

7. The image restoration method as described in claim 6, characterized in that, The image restoration method further includes: Acquire multiple raw images of a power scene; For each of the original images, at least one type of image degradation is simulated by an image processing algorithm, and brightness and contrast are adjusted by data augmentation to obtain simulated degradation image data.

8. An image restoration system, characterized in that, include: The system includes a degradation estimation module, a feature encoding module, a denoising and restoration module, and a feature decoding module. The degradation estimation module is used to acquire the image to be repaired in the power scenario, input the image to be repaired into a preset multi-task degradation estimation network to estimate the image degradation type and degree of each power device under the current weather conditions, obtain a multi-dimensional degradation vector, and input the image to be repaired into a preset visual language model to obtain text prompt information to represent the current weather conditions. The feature encoding module is used to fuse the text prompt information with the multidimensional degradation vector to obtain a conditional embedding vector; and to input the image to be repaired into a preset encoder for dimensionality reduction processing to obtain initial encoding features and intermediate feature information generated during the encoding process. The denoising and repair module is used to input the initial encoded features and the conditional embedding vector into a preset single-step denoising and repair model, so as to perform feature modulation on each layer of feature processing according to the conditional embedding vector to obtain the target encoded features; The feature decoding module is used to input the target encoded features into a preset decoder to perform feature fusion through the intermediate feature information to obtain the target image.

9. A terminal device, characterized in that, The method includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements the image restoration method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, include: A stored computer program, wherein, when the computer program is executed, it controls the device containing the computer-readable storage medium to perform the image restoration method as described in any one of claims 1-7.

Citation Information

Cited By

  • Time sequence data few-sample generation method, device and equipment based on power time sequence large model, medium and product

    CN122045827A

  • Methods, devices, equipment, media, and products for generating few-sample time-series data based on large-scale power time-series models.

    CN122045827B