Low-quality image super-resolution reconstruction method based on generative adversarial network

By acquiring optical parameters to generate dynamic point diffusion function and wavefront aberration dynamic compensation mechanism, combined with physical constraint adversarial training strategies and multi-scale diffraction consistency verification, the problem of poor reconstruction of low-quality images in existing methods is solved, and high-resolution images with rich details, visually real and optically consistent are generated.

CN120339072AActive Publication Date: 2025-07-18GUANGZHOU YUANHAO DIGITAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510463384.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-18
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The existing super-resolution method based on generative adversarial networks lacks effective modeling and joint processing of various degradation factors such as noise, blur, compression distortion when processing low-quality images, resulting in artifacts and texture distortions in the generated high-resolution images, and fails to fully consider the characteristics of the optical system and imaging mechanism, affecting the physical authenticity and consistency of the image.

Method used

By acquiring optical parameters, generating a dynamic point diffusion function, combining the dynamic compensation mechanism of wavefront aberration, using physical constraint adversarial training strategies and multi-scale diffraction consistency verification, the generator and discriminator are optimized to ensure that the generated high-resolution images conform to optical physics laws and consistency.

Benefits of technology

The generated high-resolution images are rich in details and are visually realistic, and can effectively process the complex degradation characteristics of low-quality images, improve reconstruction effects, and keep optically consistent with real images, eliminate artifacts, and ensure physical authenticity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339072A_ABST
    Figure CN120339072A_ABST
Patent Text Reader

Abstract

The invention discloses a low-quality image super-resolution reconstruction method based on a generative adversarial network, and the method comprises the steps: generating a dynamic point spread function through a differentiable optical diffraction calculation layer, building a wavefront aberration dynamic compensation mechanism based on the dynamic point spread function, predicting and compensating wavefront aberration, and obtaining optimized low-resolution input features; according to the optimization features, training a generator and a discriminator by adopting a physical constraint adversarial training strategy, and generating a high-resolution image conforming to an optical physical rule; the diffraction characteristic of a high-resolution image is optimized through a multi-scale diffraction consistency verification module, and the physical authenticity and consistency of the image under different resolutions are ensured; according to the invention, the complex degradation characteristic of the low-quality image can be effectively processed, and the detail quality and optical rationality of the generated image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of low-quality image super-resolution processing, and in particular, to a method for super-resolution reconstruction of low-quality images based on a generative adversarial network. Background Art

[0002] With the rapid development of fields such as computer vision, remote sensing imaging, medical image analysis, and security monitoring, image super-resolution reconstruction technology, as an important means to improve image quality, has received extensive attention in recent years. In practical application scenarios, due to factors such as the performance limitations of imaging devices, environmental interference, transmission compression, and physical defects of the optical system itself, the acquired images often exhibit low-quality characteristics, such as low resolution, severe noise, blurred distortion, insufficient contrast, and missing detail information. These low-quality images not only seriously affect the accuracy and robustness of subsequent image processing tasks but also pose great challenges to traditional super-resolution reconstruction methods.

[0003] In recent years, with the rapid development of deep learning technology, significant progress has been made in super-resolution methods based on convolutional neural networks (CNNs). In particular, the introduction of generative adversarial networks (GANs) has made the generated high-resolution images closer to real images in visual perception. However, most of the existing super-resolution methods based on generative adversarial networks are designed for ideal or mildly degraded images, ignoring the complex degradation characteristics of low-quality images in real scenarios and lacking an effective modeling and joint processing mechanism for multiple degradation factors such as noise, blur, and compression distortion. As a result, the reconstruction effect is poor when dealing with low-quality images, and the generated high-resolution images often have obvious artifacts, missing details, and texture distortion. In addition, the existing methods lack an effective physical constraint mechanism and do not fully consider the optical system characteristics and imaging mechanism in the formation process of low-quality images, resulting in deficiencies in physical authenticity and cross-scale consistency of the generated images. Therefore, there is an urgent need for a super-resolution reconstruction method for low-quality images to solve the technical problems existing in the current field. Summary of the Invention

[0004] The purpose of this part is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Simplifications or omissions may be made in this part, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this part, the abstract, and the title. However, such simplifications or omissions shall not be used to limit the scope of the present invention.

[0005] In view of the above existing problems, the present invention is proposed. Therefore, the present invention provides a method for super-resolution reconstruction of low-quality images based on a generative adversarial network to solve the problems raised in the background art.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: A method for super-resolution reconstruction of low-quality images based on a generative adversarial network, comprising:

[0007] Obtain optical parameters to generate a dynamic point spread function, and establish a wavefront aberration dynamic compensation mechanism based on the dynamic point spread function to predict and compensate for wavefront aberration, obtaining an optimized low-resolution input feature;

[0008] According to the optimized low-resolution input feature, train a generator and a discriminator using a physical constraint adversarial training strategy to obtain a high-resolution image that conforms to the laws of optical physics;

[0009] Optimize the diffraction characteristics of the high-resolution image to ensure the physical authenticity and consistency of the image at different resolutions.

[0010] As a preferred embodiment of the method for super-resolution reconstruction of low-quality images based on a generative adversarial network according to the present invention, wherein: the obtaining of optical parameters to generate a dynamic point spread function includes:

[0011] Encode the aperture shape, focal length, and wavelength in an optical system into a three-dimensional learnable tensor, and generate a dynamic point spread function based on the improved Fraunhofer diffraction principle.

[0012] As a preferred embodiment of the method for super-resolution reconstruction of low-quality images based on a generative adversarial network according to the present invention, wherein: the improved Fraunhofer diffraction principle includes:

[0013] Define a dynamic aperture mask on the aperture plane, perform a Fourier transform on the dynamic aperture mask, and modulate it in combination with a wavefront aberration function to obtain a complex optical field;

[0014] Perform an inverse Fourier transform on the complex optical field and calculate the square of its amplitude.

[0015] As a preferred embodiment of the method for super-resolution reconstruction of low-quality images based on a generative adversarial network according to the present invention, wherein: establishing a wavefront aberration dynamic compensation mechanism based on the dynamic point spread function to predict and compensate for wavefront aberration, obtaining an optimized low-resolution input feature, includes:

[0016] Obtain a low-resolution image, predict a three-dimensional Zernike coefficient matrix through a lightweight neural network, linearly combine the Zernike coefficient matrix with the Zernike basis function in normalized polar coordinates, and calculate a wavefront aberration function;

[0017] Through a generator, perform a complex-domain convolution on the input feature map generated based on the low-resolution image and the phase modulation term of the wavefront aberration function to generate a compensated input feature map.

[0018] As a preferred solution of the low-quality image super-resolution reconstruction method based on the adversarial generative network according to the present invention, it further includes:

[0019] Design a shift-variable convolution kernel, where each convolution kernel corresponds to the local point spread function of the image block, and decompose each convolution kernel into the product of three matrices through matrix factorization.

[0020] As a preferred solution of the low-quality image super-resolution reconstruction method based on the adversarial generative network according to the present invention, it adopts a physical constraint adversarial training strategy to train the generator and the discriminator, including:

[0021] Input the optimized low-resolution input features into the generator, and through multi-layer convolution and upsampling operations, map the low-resolution features to a high-resolution image;

[0022] Construct a multi-scale discriminator, perform Fourier transform on the high-resolution image output by the generator in each scale branch, calculate the ratio of its amplitude to the amplitude of the Fourier transform of the low-resolution image as the modulation transfer function of the generated image;

[0023] Calculate the sum of the absolute values of the differences between the modulation transfer function of the image and the pre-computed modulation transfer function of the real image as the physical consistency adversarial loss value.

[0024] As a preferred solution of the low-quality image super-resolution reconstruction method based on the adversarial generative network according to the present invention, it further includes:

[0025] During the training process of the generator and the discriminator, apply orthogonality constraints to the gradients of the optical parameters and the network parameters, and by minimizing the loss of the network parameters and maximizing the loss of the optical parameters, add the trace of the product of their gradients as a regularization term.

[0026] As a preferred solution of the low-quality image super-resolution reconstruction method based on the adversarial generative network according to the present invention, optimize the diffraction characteristics of the high-resolution image, including:

[0027] Extract diffraction ring features of different scales through an annular Gabor filter, where each filter consists of a Gaussian attenuation term and a phase oscillation term based on the ring spacing;

[0028] Calculate the KL divergence between the generated image and the real high-resolution image in the diffraction ring feature distributions of different scales, and add up the KL divergence values of all scales as the cross-scale consistency loss value.

[0029] As a preferred solution of the low-quality image super-resolution reconstruction method based on the adversarial generative network according to the present invention, it further includes:

[0030] If the difference between the mean value of the diffraction ring features of the generated image and the mean value of the features of the real image exceeds the set threshold, it is determined that a non-physical diffraction ring pattern is detected;

[0031] Calculate the absolute value of the difference between the mean value of the features of the generated image and the real image, and adjust the weights of the features at each scale according to the normalized ratio of this difference to the mean value of the features of the real image.

[0032] Compared with the prior art, the beneficial effects of the invention are as follows:

[0033] 1. By obtaining optical parameters to generate a dynamic point spread function and combining a wavefront aberration dynamic compensation mechanism, various degradation factors such as blur, noise, and distortion introduced by optical system defects in low-quality images are modeled. Compared with the limitations of traditional super-resolution methods that are difficult to handle multiple degradations in real scenes, the present invention can jointly optimize these degradation effects to generate high-resolution images with rich details and visual reality, thereby improving the reconstruction effect of low-quality images;

[0034] 2. Adopting a physical constraint adversarial training strategy, evaluating the optical transfer characteristics of the generated image through the modulation transfer function, and optimizing the generator with the physical consistency adversarial loss value, ensuring that the generated high-resolution image conforms to the laws of optical physics. Compared with the problems of non-physical artifacts and texture distortion that are prone to occur in existing methods based on generative adversarial networks, the proposed solution of the present invention optimizes the generator with the physical consistency adversarial loss value, making the generated image consistent with the real image in terms of optical characteristics, thereby eliminating the artifacts of the generated image and ensuring the physical authenticity of the generated image. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:

[0036] Figure 1 is the overall flowchart of the low-quality image super-resolution reconstruction method based on an adversarial generative network according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0038] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0039] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an embodiment that is separate or selectively exclusive of other embodiments.

[0040] The present invention is described in detail in conjunction with schematic diagrams. When detailing the embodiments of the present invention, for ease of explanation, the cross-sectional views showing the device structure will be enlarged locally out of the general scale, and the schematic diagrams are only examples and should not limit the scope of protection of the present invention herein. In addition, in actual production, three-dimensional spatial dimensions including length, width, and depth should be included.

[0041] At the same time, in the description of the present invention, it should be noted that the orientation or positional relationships indicated by terms such as "upper, lower, inner, and outer" are based on the orientation or positional relationships shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the terms "first, second, or third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0042] Unless otherwise clearly defined and limited in the present invention, the terms "mounted, connected, and coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can also be a mechanical connection, an electrical connection, or a direct connection, or can be indirectly connected through an intermediate medium, or can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0043] Embodiment 1

[0044] Refer to Figure 1, which is the first embodiment of the present invention. This embodiment provides a low-quality image super-resolution reconstruction method based on a generative adversarial network, including:

[0045] S1. Obtain optical parameters to generate a dynamic point spread function, and establish a wavefront aberration dynamic compensation mechanism based on the dynamic point spread function to predict and compensate wavefront aberration, so as to obtain an optimized low-resolution input feature;

[0046] It should be noted that since traditional optical systems usually assume that optical parameters are fixed quantities, while the solution of the present invention represents optical parameters in tensor form, enabling them to be dynamically adjusted during model training, thereby allowing the model to adaptively optimize optical characteristics according to the content of the input image;

[0047] Furthermore, by encoding the aperture shape D, focal length f, and wavelength λ in the optical system into a three-dimensional learnable tensor and generating a dynamic point spread function based on the improved Fraunhofer diffraction principle;

[0048] where H×W represents the spatial resolution of the image, and each spatial position corresponds to local optical characteristics;

[0049] Specifically, each dimension in the three-dimensional learnable tensor corresponds to the aperture shape D, focal length f, and wavelength λ respectively;

[0050] It should be explained that traditional Fraunhofer diffraction is a mathematical description of far-field diffraction, and its core assumption is that the observation plane is far enough from the aperture and the incident light is a monochromatic plane wave; if applied to the super-resolution reconstruction of images, there are two major defects: First, since traditional optical systems assume that optical system parameters are globally fixed, they cannot model the spatially variant aberrations in actual imaging. Second, it ignores some dynamic degradation scenarios (such as the deformation of mobile phone lenses with temperature / pressure, multi-spectral mixed illumination, etc.);

[0051] Specifically, the traditional Fraunhofer diffraction formula is expressed as follows:

[0052]

[0053] where I(x,y) represents the light intensity distribution in the image plane (output result), which is the final imaging result after the light wave passes through the optical system; represents the Fourier transform, representing the transformation from the aperture plane to the image plane; P(u,v) represents the aperture mask; represents the amplitude scaling factor of far-field diffraction, describing the energy attenuation law when the light wave propagates to the far field; A is the normalized amplitude coefficient; z is the propagation distance; x and y represent the coordinates of the x-axis and y-axis respectively;

[0054] It should be noted that u and v are two-dimensional spatial coordinates (unit: meter) of the aperture plane in the optical system; in the dynamic aperture mask, it represents the light-transmitting area of the aperture, while in the wavefront aberration function, it represents the distribution position of the aberration on the aperture plane;

[0055] Specifically, the process based on the improved Fraunhofer diffraction principle is as follows: First, a dynamic aperture mask is defined on the aperture plane, which is dynamically generated by a three-dimensional learnable tensor and used to modulate the light field distribution; then, this mask is combined with a wavefront aberration function, where the wavefront aberration function describes the phase distortion of the light wave and is predicted by subsequent steps; next, the Fourier transform is performed on the modulated light field distribution to calculate its complex light field on the image plane; finally, the square of the magnitude of this complex light field is taken to obtain the dynamic point spread function representing the light intensity distribution, and its formula is as follows:

[0056]

[0057] where, PSF dynamic (x, y) is the generated dynamic point spread function, which encodes the physical degradation characteristics of the optical system; P θ (i, v) represents the dynamic aperture mask, W(u, v) represents the wavefront aberration function, and ⊙ represents element-wise multiplication, whose role is to combine the frequency-domain response of the aperture with the phase modulation term caused by the aberration to form a complex light field, simulating the propagation process of the light wave through the actual optical system, and j is the imaginary unit; represents the inverse Fourier transform, which is used to convert the dynamic aperture mask and the phase modulation caused by the aberration from the spatial domain to the frequency domain, and then through the inverse transform, it is converted into the dynamic point spread function; |·| 2 represents the light intensity distribution;

[0058] Specifically, P θ (u, v) is expressed as:

[0059]

[0060] where, Sigmoid(·) is the activation function, whose role is to constrain the dynamic aperture mask in [0, 1], simulating the light transmittance of the aperture (0 means completely blocked, 1 means completely transparent); B k (u, v) represents the B-spline basis function, which is used to construct a continuous aperture shape; θ k is the learnable parameter that controls the weight of each B-spline basis; K represents the number of B-spline basis functions;

[0061] It should be noted that since the traditional optical system assumes that the aperture in the optical parameters is of a fixed shape (such as circular or square), while the aperture of the actual optical system (such as a mobile phone lens) may change dynamically due to mechanical deformation or diaphragm adjustment, it is necessary to simulate the irregular and dynamically changing aperture shape through the linear combination of B-spline bases;

[0062] Furthermore, a low-resolution image is obtained, and through a lightweight neural network, a three-dimensional Zernike coefficient matrix is predicted. The Zernike coefficient matrix is linearly combined with the Zernike basis function in normalized polar coordinates to calculate the wavefront aberration function;

[0063] It should be explained that the lightweight neural network is an efficient model in the field of deep learning. It aims to reduce the model complexity and the number of parameters by techniques such as streamlining the structure, adopting low-rank approximation, and depthwise separable convolution, while ensuring the model performance. Commonly used lightweight neural networks include: SqueezeNet series, ShuffleNet series, MnasNet, etc.;

[0064] Specifically, the predicted three-dimensional Zernike coefficient matrix is expressed as:

[0065] U = Ψ φ (I LR ; θ φ )

[0066] where Ψ φ is expressed as the lightweight neural network, I LR is expressed as the input low-resolution image, H′×W′ is expressed as the image spatial resolution output by the lightweight neural network; θ φ is expressed as the set of learnable parameters, including the weights and biases of all layers in the lightweight neural network;

[0067] It should be noted that 36 is the channel of the three-dimensional Zernike coefficient matrix, and its definition comes from ISO 10110-5. Among them, ISO 10110-5 defaults to using 36 Zernike polynomials, and each term of the Zernike polynomial corresponds to a specific aberration mode;

[0068] Specifically, the wavefront aberration function is expressed as:

[0069]

[0070] where Z k is the k-th Zernike basis function, which is used to describe the aberration of a specific mode; (ρ,θ) is the normalized polar coordinate;

[0071] Further, through the generator, a complex-domain convolution is performed on the input feature map generated based on the low-resolution image and the phase modulation term of the wavefront aberration function to generate a compensated input feature map;

[0072] Specifically, in the residual block of the generator, the input feature map is compensated:

[0073]

[0074] Among them, F in is the input feature map, * represents complex-domain convolution, and F comp represents the compensated input feature map;

[0075] It should be noted that by considering the dynamic coupling of the aberration pattern and the image content, the limitation of global unified compensation in the traditional method is overcome, and different aberration compensation strategies are applied to different regions of the image;

[0076] S2. According to the optimized low-resolution input features, a physical constraint adversarial training strategy is used to train the generator and the discriminator to obtain a high-resolution image that conforms to the optical physical laws;

[0077] Further, a shift-variable convolution kernel is designed, each convolution kernel corresponds to the local point spread function of the image block, and each convolution kernel is decomposed into the product of three matrices through matrix decomposition;

[0078] Specifically, the shift-variable convolution kernel {E m,n} is expressed as:

[0079]

[0080] Among them, E m,n represents the local convolution kernel (sub-kernel) at the position (m, n) of the image block, corresponding to the point spread function at this position; L m,n is the left matrix in the matrix decomposition, usually called the left singular vector matrix, and is used to represent the row space basis of E m,n ; r is the rank of the matrix; ∑ m,n is the diagonal matrix in the matrix decomposition, called the singular value matrix, and its diagonal elements are non-negative singular values, which are used to control the accuracy and complexity of the matrix decomposition; is the transpose of the right matrix in the matrix decomposition, usually called the transpose of the right singular vector matrix, and is used to represent the column space basis of E m,n ; rank represents the constraint condition;

[0081] Exemplarily, assume that the parameter quantity of the original shift-variable convolution kernel E m,n is k×k. After decomposition, the obtained parameter quantities are: L m,n : k×r, ∑ m,n : r×r, The total is expressed as: r×(2k + r). It can be seen that when r = 3, it is much smaller than k 2 ;

[0082] It should be noted that the low-rank constraint (whose constraint can vary according to the value of r) reduces the number of parameters, thereby reducing the computational and storage costs while retaining the main information of the convolutional kernel;

[0083] Furthermore, the optimized low-resolution input features are input into the generator. Through multi-layer convolution and upsampling operations, the low-resolution feature map is mapped into a high-resolution image;

[0084] Specifically, the upsampling operation includes operations such as transposed convolution or pixel rearrangement, which magnify the spatial dimension of the low-resolution input features into the target high-resolution features to generate a high-resolution image;

[0085] Even further, by constructing a multi-scale discriminator, the high-resolution image output by the generator is Fourier-transformed in each scale branch, and the ratio of its amplitude to the amplitude of the Fourier transform of the low-resolution image is calculated as the modulation transfer function of the generated image;

[0086] Specifically, the modulation transfer function is expressed as:

[0087]

[0088] where s represents the downsampling scale, and the image is evaluated at different resolutions (such as the original size, 1 / 2 size, 1 / 4 size...), ensuring that the generated image has physical consistency at each scale; l represents the spatial frequency, and G θ is the generator; represents the modulation transfer function of the generated image at the downsampling scale and spatial frequency l, which is used to measure the performance of the generated image in terms of optical characteristics;

[0089] Even further, by calculating the sum of the absolute values of the difference between the modulation transfer function of the said image and the pre-computed modulation transfer function of the real image as the physical consistency adversarial loss value;

[0090] Specifically, the physical consistency adversarial loss value L phys is calculated as follows:

[0091]

[0092] where S represents the number of downsampling scales, and ‖·‖1 represents the L1 norm (the sum of absolute differences); represents the modulation transfer function of the real high-resolution image at the downsampling scale;

[0093] It should be noted that by comparing the modulation transfer functions of the generated image and the real image, the optical characteristics of the generator output in the frequency domain can be constrained to be consistent with those of the real image, thereby reducing non-physical artifacts;

[0094] Furthermore, during the training process of the generator and the discriminator, imposing an orthogonality constraint on the gradients of the optical parameters and the network parameters, and by minimizing the loss of the network parameters and maximizing the loss of the optical parameters, while adding the trace of the product of the gradients of the two as a regularization term, parameter disentanglement can be achieved;

[0095] Specifically, the formula for realizing parameter disentanglement is as follows:

[0096]

[0097] where θ is the set of the generator and the lightweight neural network; represents the gradient of the loss function L with respect to θ; represents the gradient of the loss function L with respect to Θ, where Θ here refers to the three-dimensional learnable tensor;

[0098] represents the gradient orthogonality term, whose function is to calculate and the trace of the matrix product of; L adv represents the adversarial loss, which is the standard loss form based on the generative adversarial network;

[0099] It should be noted that minimizing θ is to make the generator generate more realistic images, while maximizing Θ is to enhance the exploration of the optical parameters to ensure that the constructed three-dimensional learnable tensor can cover a variety of optical scenarios;

[0100] S3. By optimizing the diffraction characteristics of the high-resolution image, ensure the physical authenticity and consistency of the image at different resolutions;

[0101] It should be explained that through the generative adversarial network, visually realistic images are generated in the super-resolution task, but these images are not necessarily in line with the optical physical laws; for example, if the generated image contains non-physical textures or artifacts, then in scenarios that require high-precision calculations, it is unacceptable; in addition, low-quality images are usually strongly affected by complex optical degradations (such as diffraction, scattering), and even after being processed by technology, the obtained high-resolution images still need to be consistent with these degradation mechanisms; therefore, it is necessary to optimize the diffraction characteristics of the obtained high-resolution image;

[0102] Furthermore, diffraction ring features at different scales are extracted through an annular Gabor filter, where each filter consists of a Gaussian attenuation term and a phase oscillation term based on the ring spacing;

[0103] Specifically, the design of the circular Gabor filter is as follows:

[0104]

[0105] Among them, r controls the ring spacing, represents the frequency of the circular oscillation, also known as the scale index, and is used to control the density of the diffraction rings; σ represents the control bandwidth; represents the Gaussian attenuation term, indicating the amplitude distribution of the filter in the spatial domain. As the distance (x 2 +y 2 ) increases, its amplitude distribution shows exponential decay; represents the phase oscillation term based on the ring spacing;

[0106] Furthermore, calculate the KL divergence between the generated image and the true high-resolution image in the characteristic distributions of diffraction rings at different scales, and sum up the KL divergence values at all scales as the cross-scale consistency loss value;

[0107] Specifically, the cross-scale consistency loss value L diff is expressed as:

[0108]

[0109] Among them, R represents the scale set, which contains all the Gabor filter scales used for feature extraction; is the probability distribution of the diffraction ring features of the generated image at the r-th scale; is the probability distribution of the diffraction ring features of the true high-resolution image at the r-th scale; represents the KL divergence, which is used to measure the difference between the generated image and the true image in the feature distribution at the r-th scale, and can be expanded as:

[0110]

[0111] Even further, if the difference between the mean value of the diffraction ring features of the generated image and the mean value of the true image features exceeds the set threshold, it is determined that a non-physical diffraction ring pattern is detected;

[0112] Specifically, the set threshold is expressed as:

[0113]

[0114] Among them, is the mean value of the diffraction ring features of the generated image at the r-th scale, is the mean value of the diffraction ring features of the true high-resolution image at the r-th scale; τ is the set threshold, which is usually determined by experimental parameter tuning, or based on the feature distribution results of model training, or obtained by statistical analysis of the diffraction features of the true image;

[0115] Further, calculate the absolute value of the difference between the mean features of the generated image and the real image, and adjust the weights of the features at each scale according to the normalized ratio of the difference to the mean features of the real image;

[0116] Specifically, the process of adjusting the weights of the features at each scale is as follows:

[0117]

[0118] Among them, w (r) represents the weight of the diffraction ring feature at the r-th scale, represents the absolute difference (subject to the L1 norm) between the mean features of the generated image and the real image at the r-th scale; δ represents a very small positive number approaching 0;

[0119] It should be noted that since the cross-scale consistency loss value optimizes and adjusts the feature distribution through global (each scale) features, it will lead to the inability to completely eliminate local anomalies (for example, extreme deviations that occur at certain scales). Then, local correction can be performed by adjusting the weights of the features at each scale to enhance the fineness of feature optimization.

[0120] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.

[0121] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0122] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the function specified in one or more processes and / or blocks Figure 1 in the process or processes and / or blocks Figure 1 in the block or blocks.

[0123] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more processes and / or blocks Figure 1 in the process or processes and / or blocks Figure 1 in the block or blocks.

[0124] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.

[0125] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A low-quality image super-resolution reconstruction method based on a generative adversarial network, characterized in that Including: Obtain optical parameters to generate a dynamic point spread function, and based on the dynamic point spread function, establish a wavefront aberration dynamic compensation mechanism to predict and compensate for wavefront aberration, obtaining optimized low-resolution input features; According to the optimized low-resolution input features, adopt a physical constraint adversarial training strategy to train the generator and discriminator to obtain a high-resolution image that conforms to optical physical laws; Ensure the physical authenticity and consistency of the image at different resolutions by optimizing the diffraction characteristics of the high-resolution image.

2. The super-resolution reconstruction method for low-quality images based on the adversarial generative network according to claim 1, characterized in that The obtaining of optical parameters to generate a dynamic point spread function includes: Encode the aperture shape, focal length, and wavelength in the optical system into a three-dimensional learnable tensor, and based on the improved Fraunhofer diffraction principle, generate a dynamic point spread function.

3. The low-quality image super-resolution reconstruction method based on the adversarial generative network according to claim 1, characterized in that, The improved Fraunhofer diffraction principle includes: Define a dynamic aperture mask on the aperture plane, perform a Fourier transform on the dynamic aperture mask, and modulate it in combination with the wavefront aberration function to obtain a complex optical field; Perform an inverse Fourier transform on the complex optical field and calculate the square of its amplitude.

4. The method for super-resolution reconstruction of low-quality images based on generative adversarial networks according to claim 1 or 2, characterized in that, Based on the dynamic point spread function, establishing a wavefront aberration dynamic compensation mechanism to predict and compensate for wavefront aberration, obtaining optimized low-resolution input features, includes: Obtain a low-resolution image, predict a three-dimensional Zernike coefficient matrix through a lightweight neural network, linearly combine the Zernike coefficient matrix with the Zernike basis function in normalized polar coordinates, and calculate the wavefront aberration function; Through the generator, perform a complex domain convolution on the input feature map generated based on the low-resolution image and the phase modulation term of the wavefront aberration function to generate a compensated input feature map.

5. The method for super-resolution reconstruction of low-quality images based on a generative adversarial network according to claim 4, wherein Also including: Design a shift-variable convolution kernel, each convolution kernel corresponding to the local point spread function of the image block, and decompose each convolution kernel into the product of three matrices through matrix decomposition.

6. The method for super-resolution reconstruction of low-quality images based on a generative adversarial network according to claim 1 or 4, characterized in that, Adopting a physical constraint adversarial training strategy to train the generator and discriminator includes: Input the optimized low-resolution input features into the generator, and through multi-layer convolution and upsampling operations, map the low-resolution feature map to a high-resolution image; Construct a multi-scale discriminator, perform a Fourier transform on the high-resolution image output by the generator in each scale branch, calculate the ratio of its amplitude to the amplitude of the Fourier transform of the low-resolution image as the modulation transfer function of the generated image; Calculate the sum of the absolute values of the differences between the modulation transfer function of the image and the pre-computed modulation transfer function of the real image as the physical consistency adversarial loss value.

7. The super-resolution reconstruction method of low-quality images based on the adversarial generative network according to claim 6, characterized in that Also including: During the training process of the generator and discriminator, impose an orthogonality constraint on the gradients of the optical parameters and network parameters, and by minimizing the loss of the network parameters and maximizing the loss of the optical parameters, add the trace of the product of their gradients as a regularization term.

8. The low-quality image super-resolution reconstruction method based on a generative adversarial network according to claim 1, wherein Optimizing the diffraction characteristics of the high-resolution image includes: Extract diffraction ring features of different scales through an annular Gabor filter, where each filter consists of a Gaussian attenuation term and a phase oscillation term based on the ring spacing; Calculate the KL divergence between the generated image and the true high-resolution image in the diffraction ring feature distributions at different scales, and sum up the KL divergence values at all scales as the cross-scale consistency loss value.

9. The super-resolution reconstruction method for low-quality images based on the adversarial generative network according to claim 8, characterized in that, It also includes: If the difference between the mean value of the diffraction ring features of the generated image and the mean value of the true image features exceeds the set threshold, it is determined that a non-physical diffraction ring pattern is detected; Calculate the absolute value of the difference between the mean value of the features of the generated image and the true image, and adjust the weights of the features at each scale according to the normalized ratio of this difference to the mean value of the true image features.

Citation Information

Patent Citations

  • Single image super-resolution reconstruction method based on conditional generative adversarial network

    CN110136063A

  • Image processing model training method, image processing method, medium and terminal

    CN113379610A

  • Mass spectrum image super-resolution reconstruction method based on transfer learning

    CN116523756A

  • Multi-modal high-resolution light field reconstruction method based on deep learning

    CN117078850A

  • Characteristic aggregation image super-resolution reconstruction method and system

    CN117788293A

Cited By

  • Imaging device based on scene features

    CN121385883A