An image super-resolution reconstruction method based on deep learning
Through the multi-dimensional feature representation, cross-scale feature alignment, incremental mapping learning and adversarial training of deep learning networks, combined with the gradient backpropagation mechanism, the features inconsistency and information distortion problems in image super-resolution reconstruction are solved, and high-quality image reconstruction effect is achieved.
Patent Information
- Application Number
- CN202411886480.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-12-20
AI Technical Summary
The existing image super-resolution reconstruction methods have shortcomings in feature extraction, cross-scale feature alignment and training efficiency, which leads to insufficient detail fidelity and naturalness of the reconstruction image, making it difficult to meet the needs of high-precision image reconstruction.
Using a deep learning-based method, high-quality super-resolution reconstruction is achieved through multi-dimensional feature representation, cross-scale feature alignment, incremental mapping learning, adversarial training and comprehensive loss function optimization, combined with gradient backpropagation mechanism.
It significantly improves the reconstruction quality of low-resolution images, solves the problems of feature dimension inconsistency and information distortion, outputs high-quality super-resolution images, has high objective indicators and excellent subjective visual quality.
Smart Images

Figure CN119809933B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of image signal processing, and more specifically, to an image super-resolution reconstruction method based on deep learning. Background Art
[0002] With the advent of the digital age, images are increasingly widely used in various fields. However, the quality problems of low-resolution images are particularly prominent in many application scenarios. Traditional image super-resolution reconstruction methods, such as bilinear interpolation and bicubic interpolation, although simple and easy to implement, often introduce blur and distortion while increasing the image resolution, making it difficult to meet the requirements of high-precision image reconstruction.
[0003] The rapid development of deep learning technology provides a new solution for image super-resolution reconstruction. Deep learning methods can extract deep feature information from low-resolution images by constructing complex neural networks, and through training a large amount of data, accurately predict and reconstruct high-resolution images. These methods have shown great potential in restoring image details and improving visual effects. However, existing deep learning methods still face many challenges in feature extraction, cross-scale feature alignment, and training efficiency, resulting in insufficient detail fidelity and naturalness of the reconstructed images.
[0004] Currently, the demand for image data and its application scope are constantly expanding. From medical imaging, satellite remote sensing to video surveillance and advanced visual applications, high-quality image data support is required. Therefore, designing an efficient and accurate deep learning model to achieve the reconstruction of low-resolution images to high-resolution images has become an important research direction in the field of image processing.
[0005] In summary, how to achieve image super-resolution reconstruction based on deep learning technology and improve the image resolution and detail fidelity has become a technical problem that urgently needs to be solved. Summary of the Invention
[0006] In order to overcome a series of defects existing in the prior art, the purpose of this application is to provide an image super-resolution reconstruction method based on deep learning for the above problems, including the following steps:
[0007] S1. Use a deep learning network to achieve multi-dimensional feature representation of a low-resolution input image;
[0008] S2. Achieve accurate alignment and lossless mapping of cross-scale features, effectively solve the problems of inconsistent feature dimensions and information distortion;
[0009] S3. Explicitly learn the incremental mapping relationship between the low-resolution input and the reconstructed high-resolution image, and alleviate the gradient disappearance problem of the deep learning network structure;
[0010] S4, enhance the naturalness and detail fidelity of the reconstructed image through adversarial training;
[0011] S5, comprehensively combine pixel-level reconstruction loss, perceptual loss, and structural similarity loss to effectively balance the objective metrics and subjective visual quality of the reconstructed image;
[0012] S6, through a unified gradient backpropagation mechanism, achieve joint optimization of the parameters of the deep learning network, and finally output a high-quality super-resolution reconstructed image.
[0013] Furthermore, S1 includes the following steps:
[0014] By designing convolutional layers with different kernel sizes, capture the local and global texture features of the image simultaneously in the shallow and deep layers of the deep learning network, and enhance the understanding of low-resolution images;
[0015] By stacking convolutional blocks and up / downsampling modules, extract multi-level features, retain texture details while learning deep semantic information;
[0016] Fuse the shallow texture features and deep semantic features through concatenation and element-wise addition in a cross-scale manner to ensure the retention of details while capturing high-level semantic information;
[0017] Introduce deformable convolution to dynamically adjust the receptive field and sampling position of the convolution kernel, more accurately capture the local structural features of the image, and enhance the texture detail recognition ability;
[0018] Through multi-scale feature extraction and fusion, construct a feature representation with strong semantics and rich details, providing more comprehensive feature information for super-resolution reconstruction.
[0019] Furthermore, the cross-scale fusion through concatenation and element-wise addition is expressed by the formula: F f = α·F s +(1 - α)·(W p *Cpncat(F s , F d ) + b p ), where F f is the result obtained by concatenating shallow and deep features, containing the detail features and semantic information of the image; α is a dynamically learned weight coefficient used to control the relative importance of shallow and deep features in the final fusion; F s represents the shallow texture feature; W p is the convolution kernel weight for transforming features; * represents the convolution operation; Concat(F s , F d ) represents the concatenation operation, concatenating the shallow texture feature F sand the deep semantic feature F d Concatenate in the channel dimension; b p is the bias term after the convolution operation; F d represents the deep semantic feature.
[0020] Furthermore, S2 includes the following steps:
[0021] Map features of different scales to a unified query-key-value space, realize the non-linear association between features through the calculation of attention weights, and effectively capture the complex dependence relationships between feature dimensions;
[0022] Through a learnable channel transformation layer and normalization operations, dynamically adjust the channel numbers and scales of features of different scales, and use channel attention and spatial attention mechanisms to adaptively balance and reconstruct feature information;
[0023] Gradually refine the feature alignment strategy at different abstraction levels, and gradually eliminate semantic biases and scale distortions in the feature representation through adaptive weight learning and cross-scale feature interaction;
[0024] By introducing an additional cross-scale consistency loss, further optimize the semantic matching of features of different scales to ensure more accurate feature alignment;
[0025] By introducing instance normalization and conditional normalization techniques, dynamically adjust the distribution characteristics of features, eliminate the statistical differences between different scales and channels, and enhance the stability and consistency of feature maps;
[0026] Through feature consistency constraints, gradually optimize the quality of feature maps, making the generated feature representations closer to the ideal cross-scale consistency expression.
[0027] Furthermore, S3 includes the following steps:
[0028] Construct an incremental mapping learning mechanism through a deep residual learning network and a multi-branch convolution architecture, accumulate residual information layer by layer, capture the subtle changes from low-resolution to high-resolution images, and avoid information loss;
[0029] Introduce a learnable channel attention mechanism, dynamically adjust the importance of different residual branches, optimize residual propagation, alleviate the problem of gradient disappearance, and enhance the ability of image detail reconstruction;
[0030] Through a multi-scale residual fusion strategy and skip connections, retain the mutual transmission of underlying texture features and high-level semantic information, and enhance the understanding of the global and local structures of images;
[0031] Adopt a residual gradient normalization and adaptive gradient clipping strategy to control the range of residual propagation, prevent gradient explosion and disappearance, and at the same time introduce a sparsity constraint to improve the compactness and efficiency of the incremental mapping representation;
[0032] Construct a gated residual unit based on a recurrent neural network, iteratively refine the residual information, gradually enhance the reconstruction of high-frequency details, and accurately capture the non-linear mapping relationship from low-resolution to high-resolution images;
[0033] Construct an adaptive residual generator based on meta-learning, and dynamically adjust the network structure and parameters through a residual learning strategy to achieve an adaptive match with the feature distribution of the input image.
[0034] Further, S4 includes the following steps:
[0035] Design a generator and a discriminator network. The generator is responsible for reconstructing the image, and the discriminator is used to distinguish the reconstructed image from the real image;
[0036] Design a feature matching loss at the intermediate feature layer of the generator. By freezing the parameters of the discriminator, calculate the mean and variance differences between the generated image and the real image at the feature layer, and prompt the generator to generate a reconstruction result closer to the real image distribution;
[0037] Through the min-max game optimization strategy, the generator minimizes the discriminative loss of the discriminator, and the discriminator maximizes the probability of correct classification, realizing a mutually competitive training process;
[0038] Alternately optimize the generator and the discriminator, gradually improve the image quality, and make the generator approximate the real image in terms of visual quality, detail fidelity, and perceptual consistency.
[0039] Further, the mean and variance differences between the image and the real image at the feature layer are expressed by the formula: where μ G is the mean of the generated image at the intermediate feature layer; μ R is the mean of the real image at the intermediate feature layer; λ1 is the weight controlling the mean difference (μ G -μ R ) in the loss; λ2 controls the weight of the variance difference (σ G -σ R ) in the loss; σ G is the variance of the generated image at the intermediate feature layer; σ R is the variance of the real image at the intermediate feature layer; is the square of the Euclidean distance, representing the sum of the squares of the element differences of the calculated vector; L feat is the feature matching loss, used to measure the statistical difference between the generated image and the real image at the feature layer.
[0040] Further, S5 includes the following steps:
[0041] Calculate the pixel-level difference between the reconstructed image and the original image to quantify the accuracy of the reconstruction;
[0042] Extract the feature representations of the original and reconstructed images at the intermediate layer, calculate their mean square error, and evaluate the similarity of the images at the perceptual and semantic levels;
[0043] Introduce the structural similarity index as the loss function, focusing on evaluating the structure, contrast, and brightness information of the reconstructed image, and effectively balancing the texture details and overall structural integrity of the image;
[0044] Design an adaptive loss weight strategy to dynamically balance the relative importance of pixel-level reconstruction loss, perceptual loss, and similarity loss;
[0045] By evaluating the reconstruction quality at both the coarse-grained and fine-grained scales of the image, comprehensively capture the local texture details and global structural features of the image, and further improve the objective metrics and subjective visual effects of the reconstructed image;
[0046] By adding a regularization term to the loss function to form a comprehensive loss function, reduce reconstruction artifacts and improve the overall reconstruction quality.
[0047] Furthermore, the comprehensive loss function is expressed by the formula: L total = β1·MSE(L orig , L recom ) + β2·L perceptual + β3·(1 - SSIM(I orig , I recon )) + β4·L reg , where L total is the comprehensive loss function, which combines pixel-level differences, perceptual loss, structural similarity loss, and regularization terms, and is used to comprehensively evaluate the quality of the reconstructed image; β1 is the weight coefficient of the pixel-level reconstruction loss; β2 is the weight coefficient of the perceptual loss term; β3 is the weight coefficient of the structural similarity loss term; β4 is the weight coefficient of the regularization term, which adjusts the contribution of the regularization term to the total loss function; MSE(I orig , L recon ) is the mean square error between the original image I orig and the reconstructed image I recon , which is used to quantify the difference between images at the pixel level and measure the accuracy of the reconstructed image. The formula is: where N is the number of pixels in the image; I orig (i) is the pixel value of the original image I orig at the i-th pixel position; I recon (u) is the pixel value of the reconstructed image I recon at the i-th pixel position; L perceptual is the perceptual loss, calculated based on the feature representation at the intermediate layer, and is used to quantify the perceptual difference between the original image and the reconstructed image. The formula is: Among them, M is the selected middle layer number, which is used to calculate the perceptual loss; φ m (I orig ) is the feature representation of the original image I orig at the m-th layer; φ m (I recon ) is the feature representation of the reconstructed image I recon at the m-th layer; SSIM(I orig ,I recon ) is the structural similarity index, which is used to evaluate the similarity between the original image I orig and the reconstructed image I recon in terms of structure, contrast, and brightness. The formula is: where μ orig is the mean value of the original image I orig ; μ recon is the mean value of the reconstructed image I recon ; c1 is the first constant in the structural similarity index, which is used to prevent the denominator from being zero and usually takes a small constant value; σ orig is the variance of the original image I orig ; σ recon is the variance of the reconstructed image I recon ; c2 is the second constant in the structural similarity index, which is usually used to stabilize the denominator and prevent the case where the variance is very small; L reg is the regularization loss, which is used to control the smoothness of the network parameters, prevent artifacts, and improve the reconstruction quality of the image. The formula is: where is the gradient of the reconstructed image I recon at the i-th pixel point.
[0048] Furthermore, S6 includes the following steps:
[0049] Construct a corresponding error signal based on the current output situation of the deep learning network and the preset target;
[0050] Based on the constructed error signal, use the gradient backpropagation algorithm to calculate the error gradients of each parameter, and clarify the influence of the parameters on the output deviation;
[0051] Transmit the error gradient information layer by layer from the subsequent layers to ensure that each layer receives gradient feedback and guides the direction and amplitude of parameter adjustment;
[0052] According to the transmitted gradient information, jointly optimize and adjust all relevant parameters to make the network gradually approach the output of a high-quality super-resolution image;
[0053] Repeatedly perform gradient calculation, information transmission, and joint parameter optimization to gradually narrow the gap between the output and the requirements of the high-quality super-resolution reconstructed image and improve the reconstruction quality of the image;
[0054] When the output image reaches the expected high-quality level, it is output as the final super-resolution reconstructed image, completing the entire super-resolution reconstruction process.
[0055] Compared with the prior art, the present application has the following beneficial effects:
[0056] The present application significantly improves the reconstruction quality of low-resolution images through a deep learning network and multi-dimensional feature representation. It solves the problems of inconsistent feature dimensions and information distortion, enhances the naturalness and detail fidelity of the image through adversarial training, balances various loss functions, and finally outputs a high-quality super-resolution reconstructed image, with high objective indicators and excellent subjective visual quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a schematic flowchart of a method for image super-resolution reconstruction based on deep learning disclosed in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be described in more detail below with reference to the accompanying drawings in the embodiments of the present invention. In the drawings, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The described embodiments are some, but not all, of the embodiments of the present invention.
[0059] All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0060] The embodiments described below with reference to the accompanying drawings and directional terms are exemplary only and are intended to explain the present invention and should not be construed as limiting the present invention.
[0061] As Figure 1 shown, a method for image super-resolution reconstruction based on deep learning includes the following steps:
[0062] S1. Use a deep learning network to implement multi-dimensional feature representation of a low-resolution input image;
[0063] S2. Achieve precise alignment and lossless mapping of cross-scale features, effectively solving the problems of inconsistent feature dimensions and information distortion;
[0064] S3. Explicitly learn the incremental mapping relationship between the low-resolution input and the reconstructed high-resolution image, alleviating the problem of gradient disappearance in the deep learning network structure;
[0065] S4, enhancing the naturalness and detail fidelity of reconstructed images through adversarial training;
[0066] S5, integrates pixel-level reconstruction loss, perceptual loss and structural similarity loss to effectively balance the objective indicators and subjective visual quality of the reconstructed image;
[0067] S6, through a unified gradient back-propagation mechanism, realizes the joint optimization of deep learning network parameters and finally outputs high-quality super-resolution reconstructed images.
[0068] In step S1, the role of the deep learning network is to extract multi-dimensional features from the low-resolution input image, specifically by extracting features from the image through deep learning models such as convolutional neural networks (CNN). Compared with high-resolution images, low-resolution images usually contain fewer details and information, and the information distribution is relatively sparse, so it is necessary to use deep learning algorithms to mine the potential features of the image. The technical effect of this step is that the deep learning model can extract feature information of different scales and levels in the image layer by layer through multi-layer convolution operations, including multi-dimensional features such as edges, textures, and color distribution. These features not only help the network understand the basic structure of the image, but also provide strong feature support for subsequent super-resolution reconstruction. At this stage, the deep learning network can extract higher-dimensional representations from complex low-resolution images, thereby laying the foundation for subsequent image reconstruction and enhancement. The success or failure of this step determines the effect of all subsequent reconstruction operations, so its importance cannot be ignored.
[0069] Low-resolution images suffer from a loss in spatial resolution, so it is necessary to accurately align multi-scale features to ensure that important information is not lost during the reconstruction process. Step S2 uses the alignment mechanism designed in the deep learning model to achieve accurate alignment and lossless mapping of information between multi-scale levels. Especially when processing low-resolution images, features of different scales often have different spatial information, so alignment operations are required to ensure the effective fusion of cross-scale information. In terms of technical effects, by aligning cross-scale features in the feature extraction stage, the deep learning network can avoid information loss or distortion problems, ensuring that all kinds of details of low-resolution images are fully preserved and restored during high-resolution reconstruction. This process can effectively solve the problem of information distortion caused by scale differences, so that features at different levels remain consistent in the network, thereby enhancing the quality and detail fidelity of the reconstructed image.
[0070] The core objective of step S3 is to explicitly learn the incremental mapping relationship between the low-resolution input and the high-resolution reconstructed image. Traditional super-resolution methods usually only focus on the reconstruction of the image itself, while ignoring the subtle differences between the low-resolution image and the high-resolution image. By designing an incremental mapping mechanism in the deep learning network, the network can learn the parts that need to be enhanced when the low-resolution image is converted into a high-resolution image. This incremental mapping can effectively alleviate the problem of vanishing gradients, which is a common challenge in the training process of deep learning models. The incremental mapping mechanism enables the network to more accurately learn the specific details of the low-resolution image and reconstruct them by strengthening the gradual enhancement of information during the process of detail restoration, thus avoiding the phenomena of vanishing gradients and information degradation and ensuring high-quality output of the reconstruction results. This process improves the reconstruction accuracy of the image, especially in terms of restoring image details and textures.
[0071] Adversarial Training is an advanced technique in deep learning and is widely used in image generation and reconstruction tasks. In super-resolution reconstruction, using a Generative Adversarial Network (GAN) for adversarial training can significantly enhance the naturalness and detail fidelity of the image. Step S4 trains the generator to generate more natural and realistic high-resolution images by constructing a game process between the generator and the discriminator. The generator is responsible for generating high-resolution images, while the discriminator judges whether the generated images are realistic enough. The two compete and promote each other, and finally the generated images can approximate the distribution of real images. This training method effectively avoids the problem of image distortion that may exist in traditional reconstruction methods, enhances the details and texture performance of the reconstructed image, and makes the visual effect of the super-resolution image closer to that of a real high-resolution image.
[0072] Step S5 optimizes the quality of the reconstructed image by integrating different loss functions. Specifically, the pixel-level reconstruction loss (such as mean squared error), the perceptual loss (such as the feature loss in the VGG network), and the structural similarity loss (SSIM) are combined to measure and optimize the quality of the reconstructed image from multiple levels. The pixel-level reconstruction loss mainly focuses on the pixel-level accuracy of the image, ensuring that the reconstructed image is as close as possible to the real image in terms of pixels. The perceptual loss focuses on the high-level features of the image, such as textures and structures, to improve the perceptual quality of the image. The structural similarity loss optimizes the global consistency of the image by measuring the structural information of the image. Through the combination of these three, the model can not only ensure the detail and texture fidelity of the image, but also improve the visual effect of the image, avoiding the visually unnatural phenomenon caused by over-optimizing pixel-level details. This combined loss function can balance the objective quality (such as PSNR) and subjective quality (such as the perceptual quality of the human eye) of the reconstructed image, thus improving the overall effect of image reconstruction.
[0073] In step S6, the entire deep learning network jointly optimizes all parameters through a unified gradient backpropagation mechanism, ensuring that each component in the entire super-resolution reconstruction process can work in coordination to output an optimal high-quality super-resolution image. Through backpropagation, the network can adjust the parameters of each layer according to the feedback information of each loss function, enabling every detail in the image reconstruction process to be optimized. This process enables the network to simultaneously optimize the low-level features (such as pixel accuracy) and high-level features (such as visual effects) of the image, ensuring that the reconstructed image achieves the best results in multiple aspects. The technical effect of this step is that the joint optimization of all network parameters can eliminate potential training errors, ensuring that the final output of the network can comprehensively consider all optimization objectives, thereby generating a clear, natural, and detail-rich super-resolution image.
[0074] In summary, the deep learning-based image super-resolution reconstruction method conducts multi-level technical optimizations. From the feature extraction of the low-resolution input image, to the precise alignment of cross-scale features of the image, then to the explicit learning of incremental mapping, and finally to the enhancement of image naturalness and detail fidelity through adversarial training, each step focuses on improving the image reconstruction quality and visual effects. By comprehensively using pixel-level reconstruction loss, perceptual loss, and structural similarity loss, and finally jointly optimizing the network parameters through the gradient backpropagation mechanism, it ensures the high-quality output of super-resolution images. This multi-dimensional technical integration not only improves the objective quality of the images but also enhances the subjective visual effects of the images, providing strong theoretical and practical support for the application of image super-resolution reconstruction technology.
[0075] Furthermore, S1 includes the following steps:
[0076] By designing convolutional layers with different kernel sizes, simultaneously capture the local and global texture features of the image in the shallow and deep layers of the deep learning network to enhance the understanding of low-resolution images;
[0077] By stacking convolutional blocks and up / downsampling modules, extract multi-level features, retain texture details while learning deep semantic information;
[0078] Cross-scale fuse the shallow texture features and deep semantic features through concatenation and element-wise addition to ensure the retention of details while capturing high-level semantic information;
[0079] Introduce deformable convolution to dynamically adjust the receptive field and sampling position of the convolution kernel, more precisely capture the local structural features of the image, and enhance the texture detail recognition ability;
[0080] Through multi-scale feature extraction and fusion, construct a semantically powerful and detail-rich feature representation to provide more comprehensive feature information for super-resolution reconstruction.
[0081] By introducing a variety of technical means, this deep learning-based image super-resolution reconstruction method has achieved remarkable improvements in detail restoration and semantic understanding. From capturing local and global texture features in convolutional layers with different kernel sizes, to the extraction and fusion of multi-level features, and then to the introduction of deformable convolutions, the technical effects of these steps complement each other, jointly constructing a powerful feature representation system. These techniques not only effectively enhance the texture details of the image but also improve the semantic understanding ability of the image, making the super-resolution reconstructed image more natural, delicate, and realistic. Through multi-scale feature extraction and fusion, the network can comprehensively utilize information at various levels to provide more comprehensive support for image reconstruction and finally output high-quality super-resolution images. The innovation of this method not only improves the image quality but also provides new ideas for the application of deep learning in image reconstruction.
[0082] Furthermore, the cross-scale fusion through cascading and element-wise addition is expressed by the formula: F f = α·F s +(1 - α)·(W p *Concat(F s , F d ) + b p ), where F f is the result obtained by cascading shallow and deep features, containing the detail features and semantic information of the image; α is the dynamically learned weight coefficient used to control the relative importance of shallow and deep features in the final fusion; F s represents the shallow texture features; W p is the convolutional kernel weight for transforming features; * represents the convolution operation; Concat(F s , F d ) represents the cascading operation, concatenating the shallow texture features F s and the deep semantic features F d in the channel dimension; b p is the bias term after the convolution operation; F d represents the deep semantic features.
[0083] In summary, through cross-scale feature fusion by cascading and element-wise addition, combined with dynamically learned weight coefficients, the deep learning network can effectively balance detail restoration and semantic preservation when processing low-resolution images. The fusion of shallow texture features and deep semantic features not only improves the accuracy of image details but also enhances the overall structural understanding of the image. The dynamically adjusted weight coefficients enable the network to have strong adaptability under different input conditions and can flexibly handle various image characteristics. The innovation of this technical method improves the quality of image super-resolution reconstruction, provides a more efficient and accurate solution for image reconstruction tasks, and the finally output image has both rich details and good semantic consistency.
[0084] Furthermore, S2 includes the following steps:
[0085] Map features of different scales to a unified query-key-value space, and achieve non-linear association between features through attention weight calculation to effectively capture complex dependencies between feature dimensions;
[0086] Dynamically adjust the number of channels and scales of features of different scales through a learnable channel transformation layer and normalization operations, and use channel attention and spatial attention mechanisms to adaptively balance and reconstruct feature information;
[0087] Gradually refine the feature alignment strategy at different abstraction levels, and gradually eliminate semantic biases and scale distortions in feature representations through adaptive weight learning and cross-scale feature interaction;
[0088] By introducing an additional cross-scale consistency loss, further optimize the semantic matching of features of different scales to ensure more accurate feature alignment;
[0089] By introducing instance normalization and conditional normalization techniques, dynamically adjust the distribution characteristics of features, eliminate statistical differences between different scales and channels, and enhance the stability and consistency of feature maps;
[0090] Through feature consistency constraints, gradually optimize the quality of feature maps to make the generated feature representations closer to the ideal cross-scale consistency expression.
[0091] In summary, the technical design and implementation of step S2 demonstrate how to address the cross-scale feature mismatch and information loss problems in image super-resolution reconstruction through fine-grained feature alignment and optimization. Through the attention mechanism, dynamically adjusting channels and scales, cross-scale consistency loss, normalization techniques, and feature consistency constraints, the network can effectively learn and optimize the relationships between different scales, enabling the reconstructed image to enhance the overall image quality while maintaining detail and semantic consistency. The introduction of these technical solutions not only enhances the detail recovery ability of the image but also improves the stability and adaptability of the network in multi-scale feature fusion. Ultimately, through the optimization of these strategies, the generated super-resolution image has higher realism and visual quality, providing strong support for image processing tasks in practical applications.
[0092] Furthermore, S3 includes the following steps:
[0093] Construct an incremental mapping learning mechanism through a deep residual learning network and a multi-branch convolutional architecture, accumulate residual information layer by layer, capture the subtle changes from low-resolution to high-resolution images, and avoid information loss;
[0094] Introduce a learnable channel attention mechanism to dynamically adjust the importance of different residual branches, optimize residual propagation, alleviate the vanishing gradient problem, and enhance the image detail reconstruction ability;
[0095] Through a multi-scale residual fusion strategy and skip connections, retain the mutual transmission of underlying texture features and high-level semantic information, and enhance the understanding of the global and local structures of the image;
[0096] Adopt a residual gradient normalization and adaptive gradient clipping strategy to control the range of residual propagation, prevent gradient explosion and vanishing, and at the same time introduce sparsity constraints to improve the compactness and efficiency of the incremental mapping representation;
[0097] Construct a gated residual unit based on a recurrent neural network to iteratively refine residual information, gradually enhance high-frequency detail reconstruction, and accurately capture the non-linear mapping relationship from low-resolution to high-resolution images;
[0098] Construct an adaptive residual generator based on meta-learning, dynamically adjust the network structure and parameters through a residual learning strategy, and achieve an adaptive match with the feature distribution of the input image.
[0099] In summary, through a series of innovative technical means, step S3 establishes an incremental mapping learning mechanism, which solves problems such as gradient disappearance, information loss, and detail restoration faced in the process of reconstructing low-resolution images into high-resolution images. The deep residual learning network and the multi-branch convolution architecture ensure that the details of the image can be restored layer by layer, while the channel attention mechanism and residual optimization further enhance the network's detail reconstruction ability. The introduction of multi-scale residual fusion and skip connections enables effective communication between the underlying texture features and the high-level semantic information, improving the understanding of the global and local structures of the image. In addition, through residual gradient normalization, adaptive gradient clipping, and the RNN-based gated residual unit, S3 optimizes the network training process and avoids common problems in traditional deep networks. The introduction of meta-learning provides the network with the ability of adaptive adjustment, enabling the network to dynamically adjust its structure and parameters according to the feature distribution of different inputs, thereby improving the accuracy and generalization ability of image reconstruction.
[0100] Furthermore, S4 includes the following steps:
[0101] Design the generator and discriminator networks. The generator is responsible for reconstructing the image, and the discriminator is used to distinguish the reconstructed image from the real image;
[0102] Design a feature matching loss at the intermediate feature layer of the generator. By freezing the parameters of the discriminator, calculate the mean and variance differences between the generated image and the real image at the feature layer, and prompt the generator to generate a reconstruction result closer to the real image distribution;
[0103] Through the min-max game optimization strategy, the generator minimizes the discriminator's discrimination loss, and the discriminator maximizes the probability of correct classification, realizing a mutually competitive training process;
[0104] Alternately optimize the generator and the discriminator, gradually improve the image quality, and make the generator approximate the real image in terms of visual quality, detail fidelity, and perceptual consistency.
[0105] In summary, step S4 significantly improves the quality of super-resolution image reconstruction through an adversarial training mechanism. The mutual game between the generator and the discriminator, by minimizing the discriminator's discrimination loss, promotes the generator to generate more realistic images; the feature matching loss enhances the detail fidelity and naturalness of the image by reducing the differences between the generated image and the real image at the feature layer; the min-max game optimization strategy forms a healthy competition between the generator and the discriminator, further improving the image quality; the alternating optimization strategy ensures that the generator and the discriminator are continuously improved during the game process, and finally generate high-quality images. Through this series of adversarial training methods, S4 not only improves the details and clarity of the reconstructed image, but also enhances the visual perception quality of the image, providing stronger technical support for the super-resolution image reconstruction task.
[0106] Furthermore, the differences in mean and variance between the generated image and the real image at the feature level are expressed by the formula: where μ G is the mean of the generated image at the intermediate feature level; μ R is the mean of the real image at the intermediate feature level; λ1 is the weight that controls the difference in mean (μ G - μ R ) in the loss; λ2 controls the weight of the variance difference (σ G - σ R ) in the loss; σ G is the variance of the generated image at the intermediate feature level; σ R is the variance of the real image at the intermediate feature level; is the square of the Euclidean distance, representing the sum of the squares of the differences in the elements of the calculated vector; L feat is the feature matching loss, which is used to measure the statistical difference between the generated image and the real image at the feature level.
[0107] Furthermore, S5 includes the following steps:
[0108] Calculate the pixel-level difference between the reconstructed image and the original image to quantify the accuracy of the reconstruction;
[0109] Extract the feature representations of the original and reconstructed images at the intermediate layer, calculate their mean squared error, and evaluate the similarity between the images at the perceptual and semantic levels;
[0110] Introduce the structural similarity index as a loss function, focusing on evaluating the structure, contrast, and brightness information of the reconstructed image, and effectively balancing the texture details and overall structural integrity of the image;
[0111] Design an adaptive loss weight strategy to dynamically balance the relative importance of pixel-level reconstruction loss, perceptual loss, and similarity loss;
[0112] By evaluating the reconstruction quality simultaneously at the coarse-grained and fine-grained scales of the image, comprehensively capturing the local texture details and global structural features of the image, and further improving the objective metrics and subjective visual effects of the reconstructed image;
[0113] By adding a regularization term to the loss function to form a comprehensive loss function, reducing reconstruction artifacts and improving the overall reconstruction quality.
[0114] In summary, by comprehensively introducing pixel-level loss, perceptual loss, structural similarity loss, and an adaptive loss weight strategy, S5 optimizes the image reconstruction quality at multiple levels. By precisely evaluating the differences at the pixel level, perceptual level, and structural level, S5 can generate more natural and delicate high-resolution images. Especially in terms of detail restoration and structural consistency, S5 demonstrates powerful capabilities. The introduction of the adaptive loss weight strategy and multi-scale evaluation balances the quality of the reconstructed images at all levels. The finally generated images not only perform excellently in objective quality but also possess strong visual effects and a sense of realism.
[0115] Furthermore, the comprehensive loss function is expressed by the formula: L total = β1·MSE(I orig , I recon ) + β2·L perceptuakl + β3·(1 - SSIM(I orig , I recon )) + β4·L reg , where L total is the comprehensive loss function, which combines pixel-level differences, perceptual loss, structural similarity loss, and a regularization term to comprehensively evaluate the quality of the reconstructed image; β1 is the weight coefficient of the pixel-level reconstruction loss; β2 is the weight coefficient of the perceptual loss term; β3 is the weight coefficient of the structural similarity loss term; β4 is the weight coefficient of the regularization term, which adjusts the contribution of the regularization term to the total loss function; MSE(I orig , I recon ) is the mean square error between the original image I orig and the reconstructed image I recon , which is used to quantify the differences at the pixel level of the image and measure the accuracy of the reconstructed image. The formula is: where N is the number of pixels in the image; I orig (i) is the pixel value of the original image I orig at the i-th pixel position; I recon (i) is the pixel value of the reconstructed image I recon at the i-th pixel position; L perceptual is the perceptual loss, which is calculated based on the feature representations of the intermediate layers and is used to quantify the perceptual differences between the original image and the reconstructed image. The formula is: where M is the selected number of intermediate layers for calculating the perceptual loss; φ m (I orig ) is the feature representation of the original image I orig at the m-th layer; φ m (I recon ) is the feature representation of the reconstructed image I recon at the m-th layer; SSIM(I orig , I recon) is the structural similarity index, used to evaluate the original image I orig and the reconstructed image I recon in terms of similarity in structure, contrast, and brightness. The formula is: where μ orig is the mean of the original image I orig ; μ recon is the mean of the reconstructed image I recon ; c1 is the first constant in the structural similarity index, preventing the denominator from being zero, usually taking a small constant value; σ orig is the variance of the original image I orig ; σ recon is the variance of the reconstructed image I recon ; c2 is the second constant in the structural similarity index, usually used to stabilize the denominator and prevent the case of very small variance; L reg is the regularization loss, used to control the smoothness of network parameters, prevent artifacts, and improve the reconstruction quality of the image. The formula is: where is the gradient of the reconstructed image I recon at the i-th pixel point.
[0116] Furthermore, S6 includes the following steps:
[0117] Construct a corresponding error signal based on the current output situation of the deep learning network and the preset target;
[0118] Based on the constructed error signal, use the gradient backpropagation algorithm to calculate the error gradients of each parameter, and clarify the influence of the parameters on the output deviation;
[0119] Transmit the error gradient information layer by layer from the subsequent layers to ensure that each layer receives gradient feedback, guiding the direction and amplitude of parameter adjustment;
[0120] According to the transmitted gradient information, jointly optimize and adjust all relevant parameters to make the network gradually approach the output of a high-quality super-resolution image;
[0121] Repeatedly perform gradient calculation, information transmission, and joint parameter optimization to gradually narrow the gap between the output and the requirements of the high-quality super-resolution reconstructed image, and improve the reconstruction quality of the image;
[0122] When the output image reaches the expected high-quality level, it is output as the final super-resolution reconstructed image, completing the entire super-resolution reconstruction process.
[0123] In summary, S6 achieves step-by-step optimization in the process of super-resolution image reconstruction through the backpropagation algorithm and the gradient transfer mechanism. It can guide the adjustment of network parameters through error feedback and gradient calculation, thereby gradually narrowing the gap between the generated image and the target image. Repeated gradient calculations and information transfers ensure that all levels of the deep learning model are finely optimized, enabling the finally generated super-resolution image to reach a high-quality level in terms of detail restoration and structural consistency. Finally, through multiple rounds of iteration and optimization, S6 ensures that the quality of the generated image meets the expectations and completes the entire super-resolution reconstruction task.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image super-resolution reconstruction method based on deep learning, characterized in that, It includes the following steps: S1. Use a deep learning network to achieve multi-dimensional feature representation of a low-resolution input image; S2. Achieve precise alignment and lossless mapping of cross-scale features, effectively solving the problems of inconsistent feature dimensions and information distortion; S3. Explicitly learn the incremental mapping relationship between the low-resolution input and the reconstructed high-resolution image, alleviating the problem of gradient disappearance in the deep learning network structure; S4. Enhance the naturalness and detail fidelity of the reconstructed image through adversarial training; S5. Integrate pixel-level reconstruction loss, perceptual loss, and structural similarity loss to effectively balance the objective metrics and subjective visual quality of the reconstructed image; S6. Through a unified gradient backpropagation mechanism, achieve joint optimization of the parameters of the deep learning network, and finally output a high-quality super-resolution reconstructed image; S2 includes the following steps: Map features of different scales to a unified query-key-value space, and achieve non-linear association between features through attention weight calculation, effectively capturing the complex dependence relationship between feature dimensions; Through a learnable channel transformation layer and normalization operation, dynamically adjust the number of channels and scales of features of different scales, and use channel attention and spatial attention mechanisms to adaptively balance and reconstruct feature information; Gradually refine the feature alignment strategy at different abstraction levels, and gradually eliminate semantic bias and scale distortion in the feature representation through adaptive weight learning and cross-scale feature interaction; By introducing an additional cross-scale consistency loss, further optimize the semantic matching of features of different scales to ensure more accurate feature alignment; By introducing instance normalization and conditional normalization techniques, dynamically adjust the distribution characteristics of features, eliminate statistical differences between different scales and channels, and enhance the stability and consistency of feature mapping; Through feature consistency constraints, gradually optimize the quality of feature mapping, making the generated feature representation closer to the ideal cross-scale consistency expression; S3 includes the following steps: Through a deep residual learning network and a multi-branch convolution architecture, construct an incremental mapping learning mechanism, accumulate residual information layer by layer, capture the subtle changes from low-resolution to high-resolution images, and avoid information loss; Introduce a learnable channel attention mechanism to dynamically adjust the importance of different residual branches, optimize residual propagation, alleviate the problem of gradient disappearance, and enhance the ability to reconstruct image details; Through a multi-scale residual fusion strategy and skip connections, retain the mutual transmission of underlying texture features and high-level semantic information, and enhance the understanding of the global and local structures of the image; Adopt a residual gradient normalization and adaptive gradient clipping strategy to control the range of residual propagation, prevent gradient explosion and disappearance, and at the same time introduce a sparsity constraint to improve the compactness and efficiency of the incremental mapping representation; Based on a recurrent neural network, construct a gated residual unit to iteratively refine the residual information, gradually enhance the reconstruction of high-frequency details, and accurately capture the non-linear mapping relationship from low-resolution to high-resolution images; Based on meta-learning, construct an adaptive residual generator, and dynamically adjust the network structure and parameters through a residual learning strategy to achieve adaptive matching with the feature distribution of the input image; S4 includes the following steps: designing a generator and a discriminator network, the generator is responsible for reconstructing the image, and the discriminator is used to distinguish the reconstructed image from the real image; designing a feature matching loss in the intermediate feature layer of the generator, and by freezing the parameters of the discriminator, calculating the mean and variance differences between the generated image and the real image at the feature layer, prompting the generator to generate a reconstruction result that is closer to the real image distribution; through the minimum-maximum game optimization strategy, the generator minimizes the discriminant loss of the discriminator, and the discriminator maximizes the probability of correct classification, realizing a mutually competitive training process; alternately optimizing the generator and the discriminator, gradually improving the image quality, so that the generator is close to the real image in terms of visual quality, detail fidelity and perceptual consistency; S5 includes the following steps: calculating the pixel-level difference between the reconstructed image and the original image to quantify the accuracy of the reconstruction; extracting the feature representations of the original and reconstructed images at the intermediate layer, calculating their mean square error, and evaluating the similarity of the images at the perceptual and semantic levels; introducing the structural similarity index as a loss function, focusing on evaluating the structure, contrast, and brightness information of the reconstructed image, and effectively balancing the texture details and overall structural integrity of the image; designing an adaptive loss weight strategy to dynamically balance the relative importance of pixel-level reconstruction loss, perceptual loss, and similarity loss; by simultaneously evaluating the reconstruction quality at both the coarse-grained and fine-grained scales of the image, comprehensively capturing the local texture details and global structural features of the image, and further improving the objective indicators and subjective visual effects of the reconstructed image; by adding a regularization term to the loss function to form a comprehensive loss function, reducing reconstruction artifacts and improving the overall reconstruction quality.
2. The method for image super-resolution reconstruction based on deep learning according to claim 1, characterized in that S1 includes the following steps: By designing convolutional layers with different kernel sizes, local and global texture features of images are captured simultaneously in the shallow and deep layers of the deep learning network, improving the understanding of low-resolution images; By stacking convolution blocks and up / down sampling modules, multi-level features are extracted to retain texture details while learning deep semantic information; The shallow texture features and deep semantic features are fused across scales through cascading and element-by-element addition to ensure that details are preserved while capturing high-level semantic information; Introducing deformable convolution to dynamically adjust the receptive field and sampling position of the convolution kernel, more accurately capture the local structural features of the image, and improve the ability to recognize texture details; Through multi-scale feature extraction and fusion, a semantically powerful and detail-rich feature representation is constructed to provide more comprehensive feature information for super-resolution reconstruction.
3. The method for image super-resolution reconstruction based on deep learning according to claim 2, wherein, Cross-scale fusion in a cascaded and element-wise addition manner is expressed by the formula: F f = α·F s +(1 - α)·(W p *Concat(F s , F d ) + b p ), where F f is the result obtained by cascading shallow and deep features, containing the detailed features and semantic information of the image; α is the dynamically learned weight coefficient used to control the relative importance of shallow and deep features in the final fusion; F s represents the shallow texture features; W p is the convolutional kernel weight for transforming features; * represents the convolution operation; Concat(F s , F d ) represents the concatenation operation, which concatenates the shallow texture features F s and the deep semantic features F d in the channel dimension; b p is the bias term after the convolution operation; F d represents the deep semantic features.
4. A method for image super-resolution reconstruction based on deep learning according to claim 1, characterized in that The mean and variance differences between the generated image and the real image in the feature layer are expressed by the formula: where μ G is the mean of the generated image in the intermediate feature layer; μ R is the mean of the real image in the intermediate feature layer; λ1 is the weight controlling the mean difference (μ G - μ R ) in the loss; λ2 controls the weight of the variance difference (σ G - σ R ) in the loss; σ G is the variance of the generated image in the intermediate feature layer; σ R is the variance of the real image in the intermediate feature layer; is the square of the Euclidean distance, representing the sum of the squares of the element differences of the calculated vector; L feat is the feature matching loss, which is used to measure the statistical difference between the generated image and the real image in the feature layer.
5. A method for image super-resolution reconstruction based on deep learning according to claim 1, characterized in that, The comprehensive loss function is expressed by the formula: L total = β1·MSE(I orig , I recon ) + β2·L perceptual + β3·(1 - SSIM(I orig , I recon )) + β4·L reg , where L total is the comprehensive loss function, which combines pixel-level differences, perceptual loss, structural similarity loss, and regularization terms to comprehensively evaluate the quality of the reconstructed image; β1 is the weight coefficient of the pixel-level reconstruction loss; β2 is the weight coefficient of the perceptual loss term; β3 is the weight coefficient of the structural similarity loss term; β4 is the weight coefficient of the regularization term, which adjusts the contribution of the regularization term to the total loss function; MSE(I orig , I recon ) is the mean square error between the original image I orig and the reconstructed image I recon , which is used to quantify the differences in the pixel level of the image and measure the accuracy of the reconstructed image. The formula is: where N is the number of pixels in the image; I orig (i) is the pixel value of the original image I orig at the i-th pixel position; I recon (i) is the pixel value of the reconstructed image I recon at the i-th pixel position; L perceptual is the perceptual loss, which is calculated based on the feature representations of the intermediate layers and is used to quantify the perceptual differences between the original image and the reconstructed image. The formula is: where M is the selected number of intermediate layers for calculating the perceptual loss; φ m (I orig ) is the feature representation of the original image I orig at the m-th layer; φ m (I recon ) is the feature representation of the reconstructed image I recon at the m-th layer; SSIM(I orig , I recon ) is the structural similarity index, which is used to evaluate the similarity in structure, contrast, and brightness between the original image I orig and the reconstructed image I recon . The formula is: where μ orig is the mean of the original image I orig ; μ recon is the mean of the reconstructed image I recon ; c1 is the first constant in the structural similarity index to prevent the denominator from being zero, taking a small constant value; σ orig is the original image I orig is the variance of; σ recon is the reconstructed image I recon is the variance of; c2 is the second constant in the structural similarity index, used to stabilize the denominator and prevent the case where the variance is very small; L reg is the regularization loss, used to control the smoothness of network parameters, prevent artifacts and improve the reconstruction quality of the image. The formula is: where is the gradient of the reconstructed image I recon at the i-th pixel point.
6. A method for image super-resolution reconstruction based on deep learning according to claim 1, characterized in that, S6 includes the following steps: Construct the corresponding error signal based on the current output of the deep learning network and the preset goal; Based on the constructed error signal, the gradient back propagation algorithm is used to calculate the error gradient of each parameter to clarify the impact of the parameter on the output deviation; Pass the error gradient information from subsequent layers layer by layer to ensure that each layer receives gradient feedback to guide the direction and magnitude of parameter adjustment; According to the transmitted gradient information, all relevant parameters are jointly optimized and adjusted so that the network gradually approaches the output of high-quality super-resolution images; Repeatedly perform gradient calculation, information transmission, and joint parameter optimization to gradually narrow the gap between the output and the requirements of high-quality super-resolution reconstructed images, and improve the reconstruction quality of the images; When the output image reaches the expected high-quality level, it is output as the final super-resolution reconstructed image to complete the entire super-resolution reconstruction process.
Citation Information
Patent Citations
Three-dimensional ultrasonic simulation method and device using generative adversarial network
CN111260741A
Image super-resolution reconstruction method based on generative adversarial network
CN113781311A