A super-resolution image processing method and system based on improved multi-scale reality guidance
By improving the multi-scale reality-guided super-resolution image processing method, using pre-trained vision encoder and artifact detector combined with diffusion model, the problems of noise and artifacts in power inspection images are solved, stable reconstruction of high-quality images is achieved, and the efficiency and accuracy of power system monitoring are improved.
Patent Information
- Application Number
- CN202510740424.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-05
AI Technical Summary
Existing diffusion models are prone to noise and artifacts when processing super-resolution of power inspection images, resulting in excessive smoothing of images and unstable quality, affecting inspection efficiency and accuracy.
The super-resolution image processing method based on improved multi-scale reality guidance is adopted, and the image reconstruction process is optimized by combining pre-training vision encoder, artifact detector and diffusion model.
It significantly improves the quality and detail retention ability of the image, solves noise and artifact problems, improves the visual quality and detail richness of the image, and enhances the clarity and stability of the monitoring image of the power system.
Smart Images

Figure CN120259083B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a super-resolution image processing method and system based on improved multi-scale reality guidance. Background Art
[0002] As power system construction continues to expand and its structure becomes increasingly complex, power inspections are placing increasingly high demands on image quality. As a critical infrastructure, the power system encompasses multiple links, including power generation, transmission, transformation, and distribution. Its safe and stable operation is crucial to all aspects of the economy and daily life. During power inspections, workers can remotely check for damaged transmission lines, aging insulators, and the normal operation of substation equipment using images captured by drones or surveillance cameras. Therefore, obtaining high-quality images of equipment and the environment is crucial for promptly identifying potential safety hazards and assessing equipment operating conditions. However, the complex and diverse power inspection environment, such as lighting changes in the natural environment, interference from objects, and the complexity and distance of the equipment itself, often results in images with insufficient resolution and noise interference, which seriously impacts the efficiency and accuracy of inspections.
[0003] In recent years, diffusion-based super-resolution technology has provided a potential solution for improving the quality of power inspection images. This technology processes low-resolution images to generate high-resolution, high-quality images, enabling inspectors to more clearly observe equipment status. In practical applications, diffusion models generate high-quality images through a gradual denoising process, theoretically effectively improving image resolution and restoring detailed information. However, traditional diffusion models are prone to generating noise and artifacts when processing power inspection image super-resolution. These noise and artifacts can obscure the true status of equipment, greatly complicating inspectors' judgment and even leading to unnecessary repairs and wasted resources. This has hindered the widespread application of super-resolution technology in power inspection.
[0004] There are two major problems in the processing of existing technologies: on the one hand, the true latent representation obtained from the lower-quality image in the early stage often causes the final output image to be over-smoothed. This is because when using the latent representation for image reconstruction, the high-frequency detail information in the image cannot be fully captured, resulting in the image losing its due clarity and texture details, making the image appear blurred and smooth, seriously affecting the image quality and usability; on the other hand, the diffusion model super-resolution strategy currently used in the optimization iteration process is unstable. Since the model's processing of noise and image structure is not accurate enough, especially when processing complex scenes or images with specific features, it is easy to produce instability, resulting in large fluctuations in the quality of the generated image, and even image distortion, artifacts and other problems, seriously restricting the effective application of super-resolution technology in the field of power inspection.
[0005] Currently, no effective solutions have been proposed for the problems in related technologies. Summary of the Invention
[0006] In response to the problems in the related art, the present invention proposes a super-resolution image processing method and system based on improved multi-scale reality guidance, which has the advantages of enhancing reality scoring, achieving stable sampling and improving image quality stability, thereby solving the problems of excessive image smoothing, unstable super-resolution quality, and the presence of noise and artifacts in the existing technology.
[0007] To this end, the specific technical solutions adopted in the present invention are as follows:
[0008] According to one aspect of the present invention, a super-resolution image processing method based on improved multi-scale reality guidance is provided, and the super-resolution image processing method based on improved multi-scale reality guidance includes:
[0009] S1. Use pre-trained visual encoder to convert the input image into an initial latent representation;
[0010] S2, based on the artifact detector, identifies the artifact area in the input image and combines it with the multi-scale true latent representation for reality-guided correction;
[0011] S3. Establish a stochastic differential equation based on the forward process of the diffusion model, and solve the diffusion ordinary differential equation of the reverse process in combination with the optimal boundary conditions to obtain the final optimized potential representation;
[0012] S4. The final optimized latent representation is decoded through the decoder, and the super-resolution image is reconstructed and post-processed and output.
[0013] Furthermore, using the pre-trained visual encoder, the input image is converted into an initial latent representation including:
[0014] S11, acquiring low-resolution images or video frames based on different acquisition devices to obtain input images;
[0015] S12, preprocessing the input image to improve the efficiency and accuracy of subsequent processing; wherein the preprocessing includes denoising and normalization;
[0016] S13. Use the pre-trained visual encoder to convert the pre-processed input image into the initial latent representation.
[0017] Furthermore, the artifact detector is used to identify artifact areas in the input image, and the multi-scale true latent representation is combined to perform reality-guided correction, including:
[0018] S21. Perform multi-scale decomposition on the initial latent representation to construct true latent representations at different scales.
[0019] S22, determining defects in the latent space based on the artifact detector, and marking areas in the image where artifacts exist by generating a binary mask;
[0020] S23, based on the generated mask, modify the current potential representation using the real potential representation at different scales;
[0021] S24. Based on the preset step size, starting from the total diffusion step, iteratively execute steps S22 and S23, and perform weighted fusion on the latent representations after correction at different scales to preserve image details.
[0022] Furthermore, the artifact detector determines the defects in the latent space and generates a binary mask to mark the areas in the image where artifacts exist, including:
[0023] In each iterative step, the current potential representation is converted to the image domain based on the decoder to obtain the iteratively generated image;
[0024] An artifact detector is used to detect artifacts on the iteratively generated image and generate a binary mask.
[0025] Furthermore, the expression for converting the current potential representation to the image domain based on the decoder is:
[0026]
[0027] Where, is the image generated at the t-1 iteration; is the decoder; α t is the noise scheduling coefficient at iteration t times; is the potential representation at scale s; x is the noisy latent variable; ε θ is the Gaussian noise prediction network controlled by parameter θ; σ θ is the noise intensity adjustment function controlled by parameter θ; ε is the standard Gaussian noise.
[0028] Furthermore, based on the generated mask, the current potential representation is modified using the true potential representation at different scales, including:
[0029] S231, based on the generated binary mask, modify the potential representation at each scale;
[0030] S232, perform Fourier transform on the potential representation and calculate the energy proportion of the high-frequency region to determine the high-frequency details contained in different scales;
[0031] S233. Determine weights based on the high-frequency details contained in different scales, and perform weighted fusion on the modified latent representations of each scale.
[0032] Furthermore, a stochastic differential equation is established according to the forward process of the diffusion model, and combined with the optimal boundary conditions, the diffusion ordinary differential equation of the reverse process is solved to obtain the final optimized potential representation including:
[0033] S31. Establishing a stochastic differential equation based on the forward process of the diffusion model to describe image changes; wherein the stochastic differential equation includes a drift coefficient for controlling the noise evolution direction and a diffusion coefficient for controlling the noise intensity;
[0034] S32. By combining the probability distribution of the forward process with the partial differential equation, the dual problem of the partial differential equation is solved to obtain the diffusion ordinary differential equation of the reverse process;
[0035] S33. Based on the diffusion ordinary differential equation of the inverse process, the optimal boundary condition constraints are introduced, and the functional is constructed and solved through the Lagrange multiplier method to obtain the final optimized potential representation.
[0036] Furthermore, by combining the probability distribution of the forward process with the partial differential equation and solving the dual problem of the partial differential equation, the diffusion ordinary differential equation of the reverse process is obtained, including:
[0037] Set the probability density function of the forward process and establish the corresponding partial differential equation;
[0038] By solving the adjoint equation of the partial differential equation, the diffusion ordinary differential equation of the reverse process is obtained;
[0039] Among them, the expression of the diffusion ordinary differential equation of the reverse process is:
[0040]
[0041] Where x t-1 is the weighted fusion result of the potential representation after correction at each scale; x t is the initial potential representation at time t; μ(x t ,t) is the drift function of the reverse process; σ(x t ,t) is the diffusion function of the reverse process; p(x t ) is the probability density function of the forward process at time t.
[0042] Furthermore, based on the diffusion ordinary differential equation of the inverse process, the optimal boundary condition constraints are introduced, and the functional is constructed and solved through the Lagrange multiplier method to obtain the final optimized potential representation including:
[0043] S331. Define the boundary conditions of the image in the super-resolution process and substitute them into the diffusion ordinary differential equation of the inverse process to obtain an ordinary differential equation system with boundary constraints;
[0044] S332. Construct functionals based on a system of ordinary differential equations with boundary constraints using the Lagrange multiplier method;
[0045] S333. By taking the variation of the functional and setting it to zero, the Euler-Lagrange equation is derived; the optimal solution that satisfies the boundary constraints is obtained by solving the equation as the final optimized potential representation.
[0046] According to another aspect of the present invention, a super-resolution image processing system based on improved multi-scale reality guidance is provided. The super-resolution image processing system based on improved multi-scale reality guidance includes:
[0047] The latent representation encoding unit is used to convert the input image into an initial latent representation using a pre-trained visual encoder;
[0048] A multi-scale correction unit for identifying artifact regions in the input image based on an artifact detector and performing reality-guided correction in combination with multi-scale true latent representations;
[0049] The inverse process solving unit is used to establish a stochastic differential equation according to the forward process of the diffusion model, and solve the diffusion ordinary differential equation of the inverse process in combination with the optimal boundary conditions to obtain the final optimized potential representation;
[0050] The super-resolution reconstruction unit is used to decode the final optimized latent representation through the decoder, reconstruct the super-resolution image, and perform post-processing and output.
[0051] The beneficial effects of the present invention are:
[0052] (1) The present invention aims to address the problem of excessive smoothing in the final output image caused by the initial true potential representation, and introduces multi-scale guidance to enhance the reality score; by solving the diffusion ordinary differential equation and using the optimal boundary condition method, the instability of the image super-resolution quality caused by the diffusion model in the optimization process is improved, thereby comprehensively improving the noise and artifact problems in the super-resolution technology based on the diffusion model, and significantly improving the quality and detail retention ability of the distribution network acquisition image.
[0053] (2) The present invention can effectively utilize information at different scales to correct artifacts through a multi-scale guided reality-guided refinement process. When processing natural scene images, it can clearly restore the texture details of the target image, thereby improving the visual quality and detail richness of the image.
[0054] (3) The present invention adopts the method of solving the diffusion ordinary differential equation and combining it with the optimal boundary conditions to achieve stable sampling. It can enhance the pre-trained super-resolution model based on the diffusion model without additional training, effectively solving the instability problem in the traditional diffusion model super-resolution strategy and has strong versatility.
[0055] (4) The improved method of the present invention can be applied to many fields, especially in the field of power systems. It can improve the clarity of monitoring images, meet users' demand for high-quality images, help to quickly identify abnormal situations, ensure the stable operation of the power system, and enhance user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0057] Figure 1 1 is a flow chart of a super-resolution image processing method based on improved multi-scale reality guidance according to an embodiment of the present invention;
[0058] Figure 2 2 is a specific implementation diagram of a super-resolution image processing method based on improved multi-scale reality guidance according to an embodiment of the present invention;
[0059] Figure 3 This is a principle block diagram of a super-resolution image processing system based on improved multi-scale reality guidance according to an embodiment of the present invention. DETAILED DESCRIPTION
[0060] To further illustrate each embodiment, the present invention provides drawings, which are part of the disclosure of the present invention. They are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. By referring to these contents, ordinary technicians in this field should be able to understand other possible implementation methods and advantages of the present invention. The components in the figures are not drawn to scale, and similar component symbols are generally used to represent similar components.
[0061] According to an embodiment of the present invention, a super-resolution image processing method and system based on improved multi-scale reality guidance are provided.
[0062] The present invention will now be further described with reference to the accompanying drawings and specific embodiments. Figure 1 and Figure 2 As shown, according to one embodiment of the present invention, a super-resolution image processing method based on improved multi-scale reality guidance is provided, and the super-resolution image processing method based on improved multi-scale reality guidance includes:
[0063] S1. Use pre-trained visual encoder to convert the input image into an initial latent representation;
[0064] S2, based on the artifact detector, identifies the artifact area in the input image and combines it with the multi-scale true latent representation for reality-guided correction;
[0065] S3. Establish a stochastic differential equation based on the forward process of the diffusion model, and solve the diffusion ordinary differential equation of the reverse process in combination with the optimal boundary conditions to obtain the final optimized potential representation;
[0066] S4. The final optimized latent representation is decoded through the decoder, and the super-resolution image is reconstructed and post-processed and output.
[0067] Specifically, the present invention comprises the following steps:
[0068] Step 1: Get the input image and generate the initial latent representation;
[0069] Step 2: Realistic guidance mechanism based on multi-scale guidance;
[0070] Step 3, stabilization based on solving diffusion ODEs and optimal boundary conditions (BCs);
[0071] Step 4: Final latent representation decoding, image reconstruction and output.
[0072] In one embodiment, using a pre-trained visual encoder, converting an input image into an initial latent representation includes:
[0073] S11, acquiring low-resolution images or video frames based on different acquisition devices to obtain input images;
[0074] S12, preprocessing the input image to improve the efficiency and accuracy of subsequent processing; wherein the preprocessing includes denoising and normalization;
[0075] S13. Use the pre-trained visual encoder to convert the pre-processed input image into the initial latent representation.
[0076] Specifically, using the pre-trained visual encoder, the pre-processed input image is converted into the initial latent representation including:
[0077] S131. Initialize the encoder using the model weights pre-trained on a large-scale dataset.
[0078] S132. Directly extract the initial features of the input image through the initialized encoder to obtain an initial latent representation.
[0079] Specifically, step 1, obtain the input image and generate the initial latent representation:
[0080] First, obtain low-resolution images or video frames from various safety monitoring and inspection collection equipment as the input image I of the model LR, the input image is preprocessed including but not limited to denoising, normalization and other operations to improve the accuracy and efficiency of subsequent processing.
[0081] Then, the pre-trained visual encoder E is used to transform the low-resolution image I LR Convert to the initial latent representation x t =E(I LR ), the visual encoder effectively extracts the key information in the image by learning the features of a large amount of image data and encodes it into a potential representation for subsequent processing.
[0082] Specifically, the input image is the low-resolution image or video frame after preprocessing mentioned in the previous paragraph, which is compressed into a low-dimensional potential representation x using the visual encoder. t .
[0083] Finally, the pre-trained visual variational encoder uses the model weights pre-trained on a large-scale dataset to initialize the encoder E, and directly extracts the initial features of the image by migrating to the encoder (no need to train it separately for this task, because it has been verified by a large dataset, so the initial potential features extracted are effective and critical), that is, the initial potential representation x here t , and then fine-tuned to adapt to specific tasks (such as super-resolution of power inspection images).
[0084] In one embodiment, identifying artifact regions in an input image based on an artifact detector and performing reality-guided correction in combination with a multi-scale true latent representation includes:
[0085] S21. Perform multi-scale decomposition on the initial latent representation to construct true latent representations at different scales.
[0086] S22, determining defects in the latent space based on the artifact detector, and marking areas in the image where artifacts exist by generating a binary mask;
[0087] S23, based on the generated mask, modify the current potential representation using the real potential representation at different scales;
[0088] S24. Based on the preset step size, starting from the total diffusion step, iteratively execute steps S22 and S23, and perform weighted fusion on the latent representations after correction at different scales to preserve image details.
[0089] In one embodiment, determining defects in a latent space based on an artifact detector and marking regions in an image where artifacts are present by generating a binary mask includes:
[0090] In each iterative step, the current potential representation is converted to the image domain based on the decoder to obtain the iteratively generated image;
[0091] An artifact detector is used to detect artifacts on the iteratively generated image and generate a binary mask.
[0092] In one embodiment, based on the generated mask, modifying the current latent representation using the true latent representation at different scales includes:
[0093] S231, based on the generated binary mask, modify the potential representation at each scale;
[0094] S232, perform Fourier transform on the potential representation and calculate the energy proportion of the high-frequency region to determine the high-frequency details contained in different scales;
[0095] S233. Determine weights based on the high-frequency details contained in different scales, and perform weighted fusion on the modified latent representations of each scale.
[0096] Specifically, step 2, realistic guidance mechanism based on multi-scale guidance:
[0097] The present invention uses diffusion-based super-resolution technology to iterate low-quality images, deriving high-quality images through mathematical function relationships. Each high-quality image in the iterative process is obtained by refining the current estimate. However, artifacts may appear in the generated high-quality images due to the influence of the diffusion model during the inference process. To overcome the appearance of artifacts, an artifact detector is used to identify unreasonable pixels and create a binary mask to highlight the artifacts. Then, by introducing a multi-scale reality-guided mechanism, the mask is combined with the real potential representation to perform reality-guided correction, thereby improving the alignment of the original image, greatly improving detection accuracy, and enhancing the representation of image details and textures.
[0098] Specifically, multi-scale guidance is a technique that applies guidance mechanisms at different scales to ensure that features at different levels of detail are preserved during the generation process. First, the initial latent representation x t Perform multi-scale decomposition to construct sub-potential representations of different scales Where s represents different scales. Based on the artifact detector (PAL) to determine the defects in the latent space, in the iterative process, a decoder is needed to convert the latent variables into the image domain, which is expressed as:
[0099]
[0100] Where, is the high-quality image generated at the t-1 iteration; is the decoder, used to convert the latent variables to the image domain; α t is the noise scheduling coefficient at iteration t times; is the potential representation at scale s; x is the noisy latent variable; εθ is the Gaussian noise prediction network controlled by parameter θ; σ θ is a noise intensity adjustment function controlled by parameter θ, which controls the amount of additional random noise introduced during the generation process; ε is standard Gaussian noise, which introduces randomness in the denoising step to ensure the diversity of the generated results.
[0101] Specifically, in each iteration step t, the PAL artifact detector is used to perform artifact detection on the current latent representation. Through the decoder Convert to image space Then, an artifact detector is used to generate a binary mask to mark the areas in the image where artifacts are present:
[0102]
[0103] Where, represents the binary mask generated by the artifact detector at iteration t-1, which is used to mark the areas in the image where artifacts are present; A represents the artifact detector (i.e., the PAL artifact detector), which inputs the image and outputs a binary mask marking the artifact areas; Represents the image at iteration t-1, which serves as the input of artifact detector A.
[0104] Specifically, based on the generated mask, the current latent representation is corrected using the initial true latent representations at different scales. Starting from the total diffusion step T, with each step length of t = T, .., 1, the above artifact detection, mask generation, and multi-scale reality-guided correction are iteratively performed. The latent representations after correction at different scales are weighted fused, and the weights are determined based on the contribution of different scales to image details and overall structure. Scales rich in high-frequency details are given higher weights to highlight the image details. This corrects artifacts based on the information of the true latent representations at different scales, preserving the details of the image. For each scale s, the correction formula is:
[0105]
[0106] Where, is the true potential representation corresponding to the s scale; E A It is a binary mask of the artifact area, with the value 0 (no artifact) or 1 (artifact exists).
[0107] Specifically, the potential representations after correction at each scale are weighted fused:
[0108]
[0109] Where, ω s ∈[0,1] and is the weight coefficient under scale s; xt-1 is the revised latent representation obtained through multi-scale reality-guided correction and weighted fusion, which is the final output of step 2 and serves as the input for solving subsequent diffusion ODEs.
[0110] Specifically, a Fourier transform is performed on the latent representation to calculate the energy fraction of high-frequency regions. A higher energy fraction indicates a richer high-frequency detail at that scale. High-frequency detail is weighted with a higher value, which is dynamically adjusted based on the image content and the results of multi-scale analysis. There is no fixed absolute value. For example, when there are three scales, the weight of the high-frequency detail scale is set between 0.4 and 0.8, while the sum of the weights of the remaining scales is between 0.2 and 0.6.
[0111] In one embodiment, a stochastic differential equation is established according to the forward process of the diffusion model, and the diffusion ordinary differential equation of the reverse process is solved in combination with the optimal boundary conditions to obtain the final optimized potential representation, including:
[0112] S31. Establishing a stochastic differential equation based on the forward process of the diffusion model to describe image changes; wherein the stochastic differential equation includes a drift coefficient for controlling the noise evolution direction and a diffusion coefficient for controlling the noise intensity;
[0113] S32. By combining the probability distribution of the forward process with the partial differential equation, the dual problem of the partial differential equation is solved to obtain the diffusion ordinary differential equation of the reverse process;
[0114] S33. Based on the diffusion ordinary differential equation of the inverse process, the optimal boundary condition constraints are introduced, and the functional is constructed and solved through the Lagrange multiplier method to obtain the final optimized potential representation.
[0115] In one embodiment, by analyzing the probability distribution of the forward process in combination with the partial differential equation, solving the dual problem of the partial differential equation, the diffusion ordinary differential equation of the reverse process is obtained, including:
[0116] Set the probability density function of the forward process and establish the corresponding partial differential equation;
[0117] By solving the adjoint equation of the partial differential equation, the diffusion ordinary differential equation of the inverse process is obtained.
[0118] In one embodiment, based on the diffusion ordinary differential equation of the inverse process, the optimal boundary condition constraints are introduced, and the functional is constructed and solved by the Lagrange multiplier method to obtain the final optimized potential representation including:
[0119] S331. Define the boundary conditions of the image in the super-resolution process and substitute them into the diffusion ordinary differential equation of the inverse process to obtain an ordinary differential equation system with boundary constraints;
[0120] S332. Construct functionals based on a system of ordinary differential equations with boundary constraints using the Lagrange multiplier method;
[0121] S333. By taking the variation of the functional and setting it to zero, the Euler-Lagrange equation is derived; the optimal solution that satisfies the boundary constraints is obtained by solving the equation as the final optimized potential representation.
[0122] Specifically, step 3 is based on the stabilization process of solving diffusion ODEs and optimal boundary conditions (BCs):
[0123] Since the optimization process uses a diffusion model, the super-resolution images generated in this way are prone to instability and over-smoothing. Therefore, the present invention solves the diffusion ordinary differential equations (ODEs) and combines them with optimal boundary conditions to stably sample high-quality super-resolution (SR) images from a pre-trained diffusion model. This method overcomes the performance fluctuation problem of traditional methods caused by randomness in the inverse process, thereby improving the performance of the traditional methods.
[0124] Specifically, when the diffusion model is used for super-resolution image generation, the basic principle is to gradually add noise to the image and then learn the inverse process to restore the high-resolution image. Let the diffusion process be a Markov chain. At time t, the image x t The variation of follows the following stochastic differential equation:
[0125] dx t =f(x t ,t)dt+g(x t ,t)dW t ;
[0126] Where, f(x t ,t) is the drift coefficient, g(x t ,t) is the diffusion coefficient, dW t is the increment of the Wiener process. By solving the adjoint equation of the forward partial differential equation (Fokker-Planck equation), we can obtain:
[0127] σ(x t ,t)=g(x t ,t);
[0128]
[0129] Drift function μ(x t ,t) and diffusion function σ(x t ,t) is described here.
[0130] It should be noted that the drift coefficient and diffusion coefficient are the core parameters that control the direction and intensity of noise evolution in the diffusion model. Using preset fixed coefficients, they are associated with the partial differential equations of the forward process through mathematical derivation and are ultimately used to construct a stable solution framework for the inverse process.
[0131] Specifically, in the inverse process, from a given low-resolution image x LR Restore to high resolution image x HR , but this process will fluctuate due to randomness.
[0132] In order to stabilize the super-resolution process, the present invention solves the diffusion ordinary differential equation, and the variational form of the forward diffusion process is:
[0133]
[0134] In the formula, q(x t |x t-1 ) is the transition probability distribution. By analyzing and deducing this distribution, we can get the corresponding distribution p(x t-1 |x t ).
[0135] By using the Feynman-Katz formula, the probability distribution of the forward process is combined with the partial differential equation, and the deterministic path of the reverse process is obtained by solving the dual problem of the partial differential equation. Specifically, let p(x t-1 ) is the probability density function of the forward process at time t, which satisfies the following partial differential equation:
[0136]
[0137] Solving the adjoint equation of the partial differential equation, we obtain the diffusion ordinary differential equations (ODEs) of the inverse process as the deterministic path, which is expressed as:
[0138]
[0139] Where x t-1 is the weighted fusion result of the potential representation after correction at each scale; x t is the initial potential representation at time t; μ(x t ,t) is the drift function of the reverse process, which is derived from the forward process coefficient; σ(x t ,t) is the diffusion function of the reverse process, which is related to the diffusion characteristics of the forward process; ▽ is the divergence operator, which acts on the vector field; f(x t ,t) is the drift term of the forward process, describing the drift direction of the probability density; g(x t ,t) is the diffusion coefficient of the forward process, which controls the intensity of random perturbations; x i 、x jare the coordinate components of different dimensions in multidimensional space; p(x t ) is the probability density function of the forward process at time t.
[0140] Specifically, in the process of solving ODEs, the stability of the super-resolution image is further improved by combining the optimal boundary conditions. Assume that in the super-resolution process, the boundary conditions of the image are Where, is the boundary of the image; h(t) is the given boundary function.
[0141] By substituting the boundary conditions into the ODEs of the inverse process, a system of ordinary differential equations with boundary constraints is obtained. The present invention uses the Lagrange multiplier method to construct a functional to solve this system and obtain the optimal solution that satisfies the boundary conditions. The expression is:
[0142]
[0143] Where, is the Lagrangian function; x t is the potential variable that changes in iteration step t; is x t The derivative of the iteration step t; t is the iteration step; λ(t) is the Lagrange multiplier; J is the functional, and the optimal solution is found by calculating the variational function; x LR:HR Represents the variable function in the process from low resolution to high resolution, and the independent variable in the super-resolution process; is the integral symbol, and the integral range is from o to HR; is the variable value on the image boundary θΩ at iteration step t, that is, the image state or potential representation at the boundary; h(t) is a given boundary function, which stipulates the conditions that the image boundary should meet at iteration step t to ensure the stability and rationality of the super-resolution image boundary.
[0144] It should be noted that the meaning of t is consistent. In the diffusion model, t refers to the parameterized representation of the evolutionary progress of the diffusion process (it has the dual meaning of time and number of iterations).
[0145] Specifically, by varying this functional and setting it to zero, we obtain the optimal solution with boundary conditions:
[0146] Assume that the optimal solution of x(t) is x * (t), introduce a small perturbation δx(t), then the function after the perturbation is:
[0147] x(t)=x * (t)+ιδx(t);
[0148] Where ι is a small parameter and δx(t) satisfies the boundary condition (i.e. the perturbation is zero at the image boundaries).
[0149] For the functional J[x LR:HR ] Take the first-order variation, defined as: That is, it measures the rate of change of the functional due to disturbance near the optimal solution.
[0150] After expansion, the variation of the main integral term is expressed as:
[0151]
[0152] Where, Reflects the Lagrangian function L on x t The dependency relationship, that is, x t The disturbance δx directly leads to the change of L; Reflect L to x t derivative Dependence, Rate of change It will also affect L.
[0153] The variation of the boundary condition term represents the interaction between the Lagrange multiplier and the perturbation at the boundary:
[0154]
[0155] The partial integration deals with the derivative term, and the main integral term Perform integration by parts:
[0156]
[0157] Since the perturbation δx satisfies at the boundary The boundary item disappears.
[0158] Combine the variational terms and set them to zero:
[0159]
[0160] Since δx t is arbitrary, and δx t =0 on the boundary. In order to make the integral equal to zero, the coefficient of the integral term must be zero, and the Euler-Lagrange equation is obtained:
[0161]
[0162] This equation is a necessary condition for the functional to take extreme values. Combined with the boundary conditions, solving the above Euler-Lagrange equation can obtain the optimal solution that meets the boundary constraints.
[0163] It should be noted that the above-mentioned Euler-Lagrange equations are classic tools in the calculus of variations, used to solve functional extreme value problems. These formulas are used to calculate the variation of the functional and set it to zero, thereby obtaining the mathematical derivation process of the optimal solution with boundary conditions.
[0164] Specifically, step 4, final latent representation decoding, image reconstruction and output:
[0165] The decoder converts the optimized latent representation into a high-resolution visual image by learning the mapping relationship between the latent representation and the image space. The decoder adopts a convolutional neural network architecture to restore the low-dimensional latent representation to a high-resolution image by learning the mapping relationship between the latent representation and the image space. Specifically, the latent representation x final Input decoder After that, the spatial dimension is first expanded through deconvolution or upsampling operations, the multi-layer convolution operation refines the features, adjusts the number of channels, and finally the high-resolution image I is output through the activation function. HR , that is, use the decoder to decode the final optimized potential representation obtained after the above processing to obtain a high-resolution image
[0166] The final optimized latent representation is the optimal solution with boundary conditions. In step 3, by substituting the boundary conditions into the ODEs of the inverse process, the Lagrange multiplier method is used to construct the functional and calculate the variation to obtain the optimal solution that satisfies the boundary constraints. This is the input to the decoder in step 4.
[0167] Specifically, the reconstructed high-resolution image undergoes post-processing, including color correction and contrast adjustment, to further improve its visual quality. Finally, the processed high-resolution image is output. This image maintains the details of the original image while offering higher resolution and better visual quality, making it applicable to a variety of fields requiring high-quality images.
[0168] To facilitate understanding of the above technical solutions of the present invention, the following is a partial code implementation and corresponding instructions based on the improved multi-scale reality-guided super-resolution image processing method. It should be noted that the code examples provided here are only used to illustrate the core concept of the present invention, and the actual implementation can be adjusted according to specific needs:
[0169] Code ①, obtains the input image and generates the initial latent representation. This part of the code mainly implements step 1, that is, obtaining the input image and generating the initial latent representation. By constructing a class called PreprocessAndEncode, using the pre-trained ResNet as the visual encoder, the input low-resolution image is preprocessed and feature extracted, and the initial latent representation is finally generated. The specific code is as follows:
[0170]
[0171]
[0172] Code ②, Multi-scale guided reality guidance mechanism. This part of the code focuses on step 2, that is, the reality guidance mechanism based on multi-scale guidance. We designed the MultiScaleGuidance class; this class guides and corrects the latent representation through operations such as multi-scale decomposition (the example uses a spatial pyramid method), artifact detection (using a simple convolutional network as an artifact detector), and correction (mixing the current representation with the true representation), thereby enhancing the reality score. The following is the corresponding code implementation:
[0173]
[0174]
[0175] Code ③, Solving the Diffusion ODE and Optimizing Boundary Conditions. This section of the code is primarily used to stabilize the optimization process by solving the diffusion ordinary differential equation and using optimal boundary conditions. We have written the solve_diffusion_ode function; this function defines the diffusion coefficient and drift coefficient (the example uses VP-SDE), solves the ODE using the Euler method, and applies boundary conditions during the optimization process to address the instability of image super-resolution quality caused by the diffusion model. The specific code description is as follows:
[0176]
[0177]
[0178] Code ④, decoding and image reconstruction. This part of the code is mainly used for step 4, namely, decoding the final latent representation, reconstructing the image, and outputting it. We have built a Decoder class; it decodes the final optimized latent representation after stabilization into a high-resolution image and outputs it through a series of deconvolution operations. The specific code implementation is as follows:
[0179]
[0180]
[0181] like Figure 3 According to another embodiment of the present invention, a super-resolution image processing system based on improved multi-scale reality guidance is provided. The super-resolution image processing system based on improved multi-scale reality guidance includes:
[0182] Latent representation encoding unit 1, used to convert the input image into an initial latent representation using a pre-trained visual encoder;
[0183] a multi-scale correction unit 2 for identifying artifact regions in the input image based on an artifact detector and performing reality-guided correction in combination with a multi-scale true latent representation;
[0184] The inverse process solving unit 3 is used to establish a stochastic differential equation according to the forward process of the diffusion model, and solve the diffusion ordinary differential equation of the inverse process in combination with the optimal boundary conditions to obtain the final optimized potential representation;
[0185] The super-resolution reconstruction unit 4 is used to decode the final optimized potential representation through a decoder, reconstruct the super-resolution image, and perform post-processing and output.
[0186] In summary, with the help of the above technical solutions of the present invention, the present invention addresses the phenomenon that the initial true potential representation causes excessive smoothing in the final output image, and introduces multi-scale guidance to enhance the reality score; by solving the diffusion ordinary differential equation and using the optimal boundary condition method, the image super-resolution quality instability caused by the diffusion model in the optimization process is improved, thereby comprehensively improving the noise and artifact problems in the super-resolution technology based on the diffusion model, and significantly improving the quality and detail retention ability of the distribution network collected images; the present invention can effectively utilize information at different scales to correct artifacts through the multi-scale guided reality-guided refinement process. When processing natural scene images, it can clearly restore the texture details of the target image, thereby improving the visual quality and detail richness of the image; the present invention adopts the method of solving the diffusion ordinary differential equation and combining the optimal boundary condition to achieve stable sampling, and can enhance the pre-trained super-resolution model based on the diffusion model without additional training, effectively solving the instability problem in the traditional diffusion model super-resolution strategy, and has strong versatility; the improved method of the present invention can be applied to many fields, especially in the power system field, which can improve the clarity of monitoring images, meet the user's demand for high-quality images, help quickly identify abnormal situations, ensure the stable operation of the power system, and enhance user experience.
[0187] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A super-resolution image processing method based on improved multi-scale reality guidance, characterized in that: include: S1. Use pre-trained visual encoder to convert the input image into an initial latent representation; S2, based on the artifact detector, identifies the artifact area in the input image and combines it with the multi-scale true latent representation for reality-guided correction; S3. Establish a stochastic differential equation based on the forward process of the diffusion model and solve the diffusion ordinary differential equation of the reverse process in combination with the optimal boundary conditions to obtain the final optimized potential representation. Specifically, it includes: Establishing a stochastic differential equation based on the forward process of the diffusion model to describe image changes; wherein the stochastic differential equation includes a drift coefficient for controlling the direction of noise evolution and a diffusion coefficient for controlling the noise intensity; By combining the probability distribution of the forward process with the partial differential equation, the dual problem of the partial differential equation is solved and the diffusion ordinary differential equation of the reverse process is obtained. Define the boundary conditions of the image during the super-resolution process and substitute them into the diffusion ordinary differential equation of the inverse process to obtain a system of ordinary differential equations with boundary constraints. Based on this system of ordinary differential equations with boundary constraints, construct a functional using the Lagrange multiplier method. Derivate the Euler-Lagrange equation by varying the functional and setting it to zero. Solve the equation to obtain the optimal solution that satisfies the boundary constraints, which serves as the final optimized potential representation. S4. The final optimized latent representation is decoded through the decoder, and the super-resolution image is reconstructed and post-processed and output.
2. The super-resolution image processing method based on improved multi-scale reality guidance according to claim 1, characterized in that: The method of converting the input image into an initial latent representation using a pre-trained visual encoder includes: S11, acquiring low-resolution images or video frames based on different acquisition devices to obtain input images; S12, preprocessing the input image to improve the efficiency and accuracy of subsequent processing; wherein the preprocessing includes denoising and normalization; S13. Use the pre-trained visual encoder to convert the pre-processed input image into the initial latent representation.
3. The super-resolution image processing method based on improved multi-scale reality guidance according to claim 1, characterized in that: The method of identifying artifact regions in an input image based on an artifact detector and performing reality-guided correction in combination with a multi-scale true potential representation includes: S21. Perform multi-scale decomposition on the initial latent representation to construct true latent representations at different scales. S22, determining defects in the latent space based on the artifact detector, and marking areas in the image where artifacts exist by generating a binary mask; S23, based on the generated mask, modify the current potential representation using the real potential representation at different scales; S24. Based on the preset step size, starting from the total diffusion step, iteratively execute steps S22 and S23, and perform weighted fusion on the latent representations after correction at different scales to preserve image details.
4. The super-resolution image processing method based on improved multi-scale reality guidance according to claim 3, characterized in that: The artifact detector is used to determine defects in the latent space and mark the areas in the image where artifacts are present by generating a binary mask. In each iterative step, the current potential representation is converted to the image domain based on the decoder to obtain the iteratively generated image; An artifact detector is used to detect artifacts on the iteratively generated image and generate a binary mask.
5. The super-resolution image processing method based on improved multi-scale reality guidance according to claim 4, characterized in that: The expression for converting the current potential representation into the image domain based on the decoder is: Where, is the image generated at the t-1 iteration; is the decoder; α t is the noise scheduling coefficient at iteration t times; is the potential representation at scale s; x is the noisy latent variable; ε θ is the Gaussian noise prediction network controlled by parameter θ; σ θ is the noise intensity adjustment function controlled by parameter θ; ε is the standard Gaussian noise.
6. The super-resolution image processing method based on improved multi-scale reality guidance according to claim 3, characterized in that: The modifying of the current potential representation using the true potential representations of different scales according to the generated mask includes: S231, based on the generated binary mask, modify the potential representation at each scale; S232, perform Fourier transform on the potential representation and calculate the energy proportion of the high-frequency region to determine the high-frequency details contained in different scales; S233. Determine weights based on the high-frequency details contained in different scales, and perform weighted fusion on the modified latent representations of each scale.
7. The super-resolution image processing method based on improved multi-scale reality guidance according to claim 1, characterized in that: By combining the probability distribution of the forward process with the partial differential equation and solving the dual problem of the partial differential equation, the diffusion ordinary differential equation of the reverse process is obtained, which includes: Set the probability density function of the forward process and establish the corresponding partial differential equation; By solving the adjoint equation of the partial differential equation, the diffusion ordinary differential equation of the inverse process is obtained; The diffusion ordinary differential equation of the reverse process is expressed as follows: Where x t-1 is the weighted fusion result of the potential representation after correction at each scale; x t is the initial potential representation at time t; μ(x t ,t) is the drift function of the reverse process; σ(x t ,t) is the diffusion function of the reverse process; p(x t ) is the probability density function of the forward process at time t.
8. A super-resolution image processing system based on improved multi-scale reality guidance, used to implement the super-resolution image processing method based on improved multi-scale reality guidance according to any one of claims 1 to 7, characterized in that: The system includes: The latent representation encoding unit is used to convert the input image into an initial latent representation using a pre-trained visual encoder; A multi-scale correction unit for identifying artifact regions in the input image based on an artifact detector and performing reality-guided correction in combination with multi-scale true latent representations; The inverse process solving unit is used to establish a stochastic differential equation based on the forward process of the diffusion model and solve the diffusion ordinary differential equation of the inverse process in combination with the optimal boundary conditions to obtain the final optimized potential representation. Specifically, it includes: Establishing a stochastic differential equation based on the forward process of the diffusion model to describe image changes; wherein the stochastic differential equation includes a drift coefficient for controlling the direction of noise evolution and a diffusion coefficient for controlling the noise intensity; By combining the probability distribution of the forward process with the partial differential equation, the dual problem of the partial differential equation is solved and the diffusion ordinary differential equation of the reverse process is obtained. Define the boundary conditions of the image during the super-resolution process and substitute them into the diffusion ordinary differential equation of the inverse process to obtain a system of ordinary differential equations with boundary constraints. Based on this system of ordinary differential equations with boundary constraints, construct a functional using the Lagrange multiplier method. Derivate the Euler-Lagrange equation by varying the functional and setting it to zero. Solve the equation to obtain the optimal solution that satisfies the boundary constraints, which serves as the final optimized potential representation. The super-resolution reconstruction unit is used to decode the final optimized latent representation through the decoder, reconstruct the super-resolution image, and perform post-processing and output.
Citation Information
Patent Citations
Single image super-resolution reconstruction method based on improved diffusion model
CN117575907A
Image super-resolution method and system based on gradient optimization diffusion model, and medium
CN117974450A