Image Reconstruction Method, System, and Storage Medium
By random cropping and comparative learning of the image data set, combined with adaptive graph construction and aggregation blocks, the limitations of image degradation processing in the prior art are solved, efficient image structure information utilization and task imbalance solution are achieved, and image reconstruction quality is improved.
Patent Information
- Application Number
- CN202510617879.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The prior art has limitations in dealing with unknown and complex and variable image degradation, making it difficult to effectively utilize image structure information and solve task imbalance problem.
By randomly cropping the image dataset, a comparison learning task is constructed, and the encoder is used to extract the feature vectors of query samples, positive samples and negative samples, combined with adaptive graph construction and aggregation blocks, alternately using local graphs and global graphs for information aggregation to generate high-quality high-resolution images.
It improves the generalization performance of the model in real and complex degradation scenarios, enhances the degradation robustness, realizes adaptive feature extraction and multi-scale information fusion, and improves the reconstruction quality.
Smart Images

Figure CN120147137B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image processing, computer vision, and deep learning, and particularly to an image reconstruction method, system, and storage medium. Background Art
[0002] Image super-resolution (SR) aims to recover a high-resolution (HR) image from a low-resolution (LR) image, and this technology has a wide range of applications in fields such as video surveillance, medical imaging, and remote sensing image analysis.
[0003] Traditional SR methods mainly include interpolation-based methods and reconstruction-based methods. Interpolation-based methods, such as nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation, estimate unknown pixels in the HR image by weighted averaging of known pixels in the LR image. These methods are computationally simple and easy to implement, but the reconstructed HR images are usually too smooth and lack high-frequency details. To overcome the limitations of interpolation methods, reconstruction-based methods such as iterative back projection and maximum a posteriori probability have been proposed. These methods mainly attempt to model the imaging process to recover the HR image, but the models are usually complex, computationally intensive, and sensitive to noise.
[0004] With the rise of deep learning technology, this technology has also been widely applied to SR tasks. SRCNN first introduced CNN into SR tasks and adopted a relatively simple three-layer convolutional network structure. Subsequently, a series of CNN-based SR methods have been proposed successively and have achieved significant performance improvements. For example, VDSR significantly increased the depth of the network by introducing residual learning and gradient clipping strategies. EDSR further explored network structure design by removing the batch normalization (BN) layer and expanding the model scale. RCAN proposed a channel attention mechanism that enables the network to adaptively focus on more important feature channels. Compared with traditional methods, these CNN-based methods show significant advantages when processing images with known and fixed degradations (such as bicubic downsampling), but their inherent technical limitations still restrict further performance improvements. Convolutional operations are essentially a local feature extraction mechanism, and the size of their convolutional kernels directly restricts the network's ability to understand global image semantics. Under this limitation, it is difficult for the network to effectively capture long-range dependencies in the image; secondly, as the depth of the network continues to increase, the problem of gradient vanishing becomes increasingly prominent, seriously affecting the network's learning efficiency and performance; at the same time, the processing method of sharing the same convolutional kernel for all pixels ignores the diversity of image content and the imbalance of SR tasks.
[0005] To overcome the problem of limited receptive fields in CNNs, subsequent researchers have attempted to introduce the Transformer architecture and alleviate the above problems by introducing self-attention mechanisms. SwinIR introduced the idea of Swin Transformer into SR, calculating self-attention within local windows and adopting a shifted window mechanism to expand the receptive field. Although the Transformer architecture theoretically breaks through the receptive field limitation of convolutional networks, in practical applications, methods based on Transformer also face many challenges. For example, the computational complexity increases exponentially, raising the requirements for computing resources. The standard self-attention mechanism treats all parts of an image equally and does not take into account the difference in reconstruction difficulty among different regions of the image.
[0006] Currently, graph neural networks (GNNs) have also been applied to the super-resolution (SR) task due to their ability to flexibly handle irregular structures and effectively model complex relationships between nodes. DLGNN uses graph neural networks to mine non-local image dependencies through block self-similarity modeling and adopts a dual learning mechanism to optimize the reconstruction process. IGNN proposed an iterative graph neural network with multi-scale feature fusion for the super-resolution task. However, many existing super-resolution methods based on graph neural networks usually use the k-nearest neighbor method for graph construction, and the node degrees of all image regions are fixed. This approach is difficult to adapt to the long-tailed distribution characteristics inherent in the super-resolution task. In addition, existing methods usually only construct one of the local graphs or global graphs. Although IPG improves the flexibility of graph construction by introducing a detail-aware metric and a multi-scale graph aggregation scheme, it does not explicitly consider the image degradation process. In real-world scenarios, the image degradation process is often unknown and complex, which is called blind super-resolution (Blind SR). The actual degradation process may include various combinations of blur kernels and noises. When there is a deviation between the actual degradation and the assumption, the performance of the above methods will decrease significantly.
[0007] To solve the blind SR problem, existing techniques are mainly divided into two categories: explicit kernel estimation methods and implicit degradation learning methods. Explicit kernel estimation methods first attempt to estimate the degradation model of the LR image and regard it as a system identification problem. Once the degradation model is obtained, it can be used in traditional non-blind SR methods or a specific network structure can be designed to handle the estimated degradation. For example, KernelGAN uses a generative adversarial network (GAN) to estimate the blur kernel of a single LR image; IKC proposed an iterative correction method, alternating between blur kernel estimation and SR reconstruction. However, these methods usually have a high computational complexity, require optimization for each image, and the accuracy of degradation estimation directly affects the quality of SR reconstruction. Estimation errors will cause artifacts in the reconstruction results.
[0008] Implicit degradation learning methods do not directly estimate the degradation kernel. Instead, they learn a representation of degradation through a neural network and then use this representation to guide the subsequent super-resolution reconstruction process. One category is domain adaptation-based methods, which attempt to adapt a pre-trained SR model (usually trained on known degradations) to a new, unknown degradation domain. For example, ZSSR performs online fine-tuning using specific LR images during the testing phase to adapt the model to the current degradation. However, this method requires online optimization for each LR image, resulting in a large computational overhead and is not suitable for real-time applications. Another category is unsupervised learning-based methods, which do not require paired LR-HR images for training. Instead, they learn degradation information from unlabeled data through contrastive learning or other unsupervised / self-supervised methods. For example, DASR uses contrastive learning for unsupervised degradation representation learning, improving the generalization ability of the model. Based on DASR, DSAT changes the deep feature extraction module from CNN to Transformer to utilize the global modeling ability of Transformer. However, the feature extraction ability of DASR is limited by the local receptive field of CNN, while DSAT faces the problems of high computational complexity of Transformer and difficulty in capturing pixel-level fine features. In addition, neither of them fully utilizes the structural information of the image nor considers the imbalance of the super-resolution task itself. Summary of the Invention
[0009] The technical problem to be solved by the present invention is to provide an image reconstruction method, system and storage medium to improve the utilization efficiency of image structural information and solve the task imbalance problem in view of the deficiencies of the prior art.
[0010] To solve the above technical problem, the technical solution adopted by the present invention is: an image reconstruction method, comprising the following steps:
[0011] Randomly crop the low-resolution images in the image dataset to obtain image patches; extract the shallow features of the low-resolution images;
[0012] Taking a certain image patch as a query sample, define other image patches from the same low-resolution image as the query sample as positive samples, and define image patches from different low-resolution images as negative samples; extract the feature vectors of the query sample, positive samples, and negative samples respectively;
[0013] Extract the deep features of the shallow features and the feature vectors, enhance the deep features, and obtain enhanced features;
[0014] Convert the enhanced features into high-resolution images.
[0015] The present invention provides an image reconstruction method, aiming to overcome the limitations existing in the prior art when dealing with unknown and complex image degradation (i.e., blind super-resolution tasks), and improve the utilization efficiency of image structure information and solve the problem of task imbalance. Specifically, the method of the present invention randomly crops image patches from the low-resolution images in the image dataset, and extracts the shallow features of the low-resolution images. On this basis, taking a certain image patch as a query sample, a contrastive learning task is constructed. The other image patches from the same low-resolution image as the query sample are defined as positive samples, and the image patches from different low-resolution images from the query sample are defined as negative samples. The encoder is used to extract the feature vectors of the query sample, positive sample, and negative sample respectively. Through the contrastive learning mechanism, the present invention enables the network to learn the feature representations under different degradation modes, making the image patches from the same degradation source have similar features, while pulling away the feature distances of the image patches from different degraded images. This strategy enables the network to learn a more robust feature representation for different degradations, thus significantly improving the generalization performance of the model in real complex degradation scenarios without explicitly estimating the degradation kernel or performing time-consuming online fine-tuning. Subsequently, the deep features of the shallow features and the feature vectors (providing degradation-aware information enhanced by contrastive learning) are extracted, and the deep features are enhanced. At this stage, by alternately combining the global graph structure and the local structure, different levels of information are effectively fused to generate a richer and more robust deep feature representation, laying a solid foundation for the final conversion of high-quality high-resolution images.
[0016] The encoder is used to extract the feature vectors of the query sample, positive sample, and negative sample respectively; the encoder includes a plurality of cascaded encoding units, where the first encoding unit is connected to the second encoding unit through an encoding residual activation module; the encoding residual activation module includes a plurality of cascaded encoding residual activation units; the last encoding unit is connected to the perceptron. The multi-cascaded structure of the present invention enables the encoder to learn hierarchical features from simple to complex. The encoding unit is used for simple feature extraction. The encoding residual activation unit helps to optimize the transformation and transmission of features, ensuring the effectiveness and robustness of the extracted features. The perceptron maps the high-dimensional features output by the last encoding unit to the required fixed-dimensional feature vector, which is used as the input of the contrastive learning task, ensuring that the output features are suitable for the subsequent loss calculation and optimization process.
[0017] The encoded residual activation unit includes two convolutional layers, and the two convolutional layers are connected by an activation function layer. This forms a basic local feature extraction and transformation unit, where the convolutional layer is used to capture spatial patterns, and the activation function layer introduces non-linearity, enabling the unit to learn complex non-linear mapping relationships, thereby effectively extracting non-linear features in image patches. The residual connection provides a direct channel for gradient propagation, greatly alleviating the problem of vanishing gradients that may occur when stacking multiple such units. Thanks to this design, a single encoded residual activation unit can be stably trained and can serve as a robust building block. By cascading, deeper and more complex encoded residual activation modules can be formed, which in turn support the construction of the entire deep encoder, thereby learning complex, robust, and discriminative image patch features that are crucial for the blind super-resolution task.
[0018] The specific implementation process of extracting the shallow features and the deep features of the feature vector and enhancing the deep features to obtain the enhanced features includes:
[0019] Taking the shallow features and the feature vector as the input of the first graph construction and aggregation block to obtain a first output;
[0020] Taking the first output and the feature vector as the input of the second graph construction and aggregation block to obtain a second output;
[0021] And so on;
[0022] Taking the output of the last graph construction and aggregation block as the input of a convolutional layer, concatenating the output of the convolutional layer with the shallow features, taking the concatenated features as the input of a pixel recombination layer, and taking the output of the pixel recombination layer as the input of an output convolutional layer to obtain the enhanced features.
[0023] The graph construction and aggregation block performs a series of processes on the input, aiming to adaptively construct a graph structure and perform information aggregation based on the feature vector (providing degradation-aware information strengthened by contrastive learning). This way of constructing a graph based on the feature vector and performing information aggregation, different from fixed-structure graphs (such as simple kNN graphs), can better reflect the diversity of image content and degradation patterns, making relationship modeling more flexible and targeted, thereby effectively handling the differences in reconstruction difficulty in different regions of the super-resolution task. The design of concatenating the output of the last graph construction and aggregation block with the original shallow features after passing through a convolutional layer aims to combine the features after complex graph processing and deep transformation with the original shallow features, effectively retaining the details and edge information of the image and avoiding detail loss that may be caused by deep processing.
[0024] The graph construction and aggregation block processes the input through the following steps:
[0025] Perform a linear transformation on the feature vector to obtain modulation coefficients α and β;
[0026] Multiply the normalized shallow feature / the output of the i-th graph construction and aggregation block by α, add the result of the multiplication to β, and then add the obtained addition result to the shallow feature / the output of the i-th graph construction and aggregation block to obtain an addition feature;
[0027] Add the addition feature to the normalized shallow feature / the output of the i-th graph construction and aggregation block to obtain a first output feature;
[0028] Perform bicubic downsampling on the first output feature and then perform bicubic upsampling to obtain a second output feature;
[0029] Calculate the absolute value sum DS of the corresponding pixel differences between the first output feature and the second output feature on all feature channels, and construct a local graph and a global graph according to the absolute value sum DS;
[0030] For each pixel node of the local graph, search for the k1 pixels with the most similar features to the pixel node within the spatial neighborhood of the pixel node, and connect the k1 pixels to the pixel node; for the global graph, sample all pixels in the global graph at a set step size. For each sampled pixel node, calculate the k2 pixels with the most similar features to the pixel node among all sampled pixels, and connect the k2 pixels to the pixel node;
[0031] Input the first output feature, the global graph, and the local graph into the graph aggregation unit module, and splice the output of the graph aggregation unit module after passing through a convolutional layer with the shallow feature to obtain the output of the graph construction and aggregation block.
[0032] Bicubic downsampling is to calculate and select some pixel points according to the bicubic interpolation function for the original feature map to construct a new image, and the size of the new image becomes smaller.
[0033] Bicubic upsampling also uses the bicubic interpolation function. Based on the pixels of the original image, the values of newly generated pixels are estimated by calculating weights, thereby increasing the resolution of the image and enlarging the size of the original image.
[0034] The image obtained by performing bicubic downsampling and then bicubic upsampling has the same size as the original image, and can be used to calculate DS.
[0035] The process of obtaining modulation coefficients through linear transformation of the feature vector and using these coefficients to modulate the output of the aggregation block for the normalized shallow features / the i-th graph to obtain the first output feature incorporates degradation-aware information into the construction of the shallow features / the i-th graph and the output of the aggregation block, enhancing the feature representation ability. Based on the first output feature and the second output feature, the absolute value of the pixel difference and DS are calculated, and a local graph and a global graph are constructed simultaneously. The local graph focuses on the similar pixel relationships within the spatial neighborhood, capturing fine structures; the global graph models the similarity between pixels at a greater distance through sampling, capturing long-range dependencies. Utilizing both the local graph and the global graph can comprehensively model the structural and relational information of the image, overcoming the limitations of existing methods that use only local graphs or only global graphs.
[0036] The graph input graph aggregation unit module includes multiple cascaded graph aggregation units. The input of the first graph aggregation unit is the first output feature and the global graph; the input of the second graph aggregation unit is the output of the first graph aggregation unit and the local graph; the input of the third graph aggregation unit is the output of the second graph aggregation unit and the global graph; and so on. Among them, the graph aggregation unit includes an edge aggregator and a first normalization layer connected to the edge aggregator; one-channel attention module has its input side connected to the input side of the edge aggregator. The output of the channel attention module, the output of the first normalization layer, and the input of the graph aggregation unit are concatenated and then input into a convolutional feed-forward network, which is connected to a second normalization layer. The output of the second normalization layer and the input of the convolutional feed-forward network concatenated together are the output of the graph aggregation unit. The edge aggregator and the normalization layer are used to perform information aggregation based on graph nodes and ensure numerical stability. The channel attention module allows the network to adjust the importance of feature channels according to the input features during the information aggregation process, making the aggregation process more focused on the most useful feature channels for the current situation. The convolutional feed-forward network combines the characteristics of convolutional and fully connected networks and is used to perform more complex non-linear transformations on node features after graph aggregation. The residual connection inside the graph aggregation unit helps to construct deeper graph aggregation units and modules, ensuring the effective propagation of gradients and improving training stability. Through this complex unit structure including edge aggregation, channel attention, convolutional feed-forward network, and internal residual connection, as well as the alternating input strategy of the global graph and the local graph, the graph aggregation unit can efficiently and hierarchically aggregate information on the graph, extract and continuously enhance deep feature representations.
[0037] The specific implementation process of converting the enhanced features into a high-resolution image includes:
[0038] Upsample the enhanced features;
[0039] Perform a convolution operation on the upsampled feature map and map it to the HR image space to obtain a high-resolution image.
[0040] As an inventive concept, the present invention also provides an image reconstruction system, including a memory, a processor, and a computer program stored on the memory; the processor executes the computer program to implement the steps of the above method.
[0041] As an inventive concept, the present invention also provides a computer-readable storage medium, on which a computer program / instructions are stored; when the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0042] As an inventive concept, the present invention also provides a computer program product, including a computer program / instructions; when the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0044] (1) Enhanced degradation robustness: DAG-BSR uses unsupervised contrast learning to obtain degradation representations, without relying on prior knowledge of image degradation types and without explicit degradation estimation. Therefore, this method can adapt to various unknown and complex degradation types and has stronger degradation robustness.
[0045] (2) Achieved adaptive feature extraction: In the Adaptive Graph Construction and Aggregation Block (DGCAB), the adaptive graph construction unit can dynamically adjust the connection number of each pixel node according to the image details (DS) to construct content-adaptive local and global graphs. This design enables the network to more flexibly handle the reconstruction requirements of different regions in the image. Regions with rich details obtain more information aggregation, while smooth regions have less information aggregation, thus more effectively utilizing computing resources.
[0046] (3) Integrated multi-scale information: The graph aggregation unit in DGCAB effectively integrates multi-scale information of the image by alternately using local and global graphs for information aggregation. The local graph focuses on details within the pixel neighborhood and helps with the reconstruction of textures and edges; the global graph captures long-range dependencies in the image and helps with the recovery of the overall structure.
[0047] (4) Achieved effective degradation-aware fusion: DAG-BSR effectively fuses degradation information into the process of feature extraction and image reconstruction through degradation-aware feature modulation in the adaptive graph construction unit and the processing of the modulated features in GAU. This enables the network to adaptively adjust the reconstruction strategy according to different degradation situations, thereby improving the adaptability to various degradation types.
[0048] (5) Improved reconstruction quality: Benefiting from the above advantages, DAG-BSR can extract more discriminative features and effectively fuse them with the learned degradation representations, thus generating high-quality reconstructed images. Description of the Drawings
[0049] Figure 1 It is a flowchart of the method according to the embodiment of the present invention;
[0050] Figure 2 It is a structural block diagram of DAG-BSR according to the embodiment of the present invention;
[0051] Figure 3 It is a structural block diagram of the degradation encoder according to the embodiment of the present invention;
[0052] Figure 4 It is a structural diagram of the coding unit according to the embodiment of the present invention;
[0053] Figure 5 It is a structural diagram of the coding residual activation unit according to the embodiment of the present invention;
[0054] Figure 6 It is a structural block diagram of the degradation-aware graph construction and aggregation block according to the embodiment of the present invention;
[0055] Figure 7 It is a structural diagram of the adaptive graph construction unit and the graph aggregation unit;
[0056] Figure 8 It is the PSNR↑ result with anisotropic Gaussian kernel and different levels of noise (×4 experiment on the Set14 dataset). Detailed Embodiment
[0057] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0058] In the embodiments of the present invention, the degradation representations obtained by unsupervised contrastive learning are incorporated into the construction process of the dynamic pixel graph through feature modulation; the dynamic graph construction strategy based on image details is combined with the alternating use of local graphs and global graphs to achieve the adaptive fusion of multi-scale information; and the above mechanism is repeatedly applied at multiple levels of the graph neural network to form an organic whole. In addition, for the blind super-resolution task, the DGCAB block is designed, in which the AGCU combines degradation-aware feature modulation with adaptive graph construction, and the GAU aggregates local graph and global graph features.
[0059] Example 1
[0060] An embodiment of the present invention proposes a blind image super-resolution method based on graph neural network (GNN) and unsupervised degradation-aware learning, called Degradation-Aware Graph-Network for Blind Image Super-Resolution (DAG-BSR). The overall process is as Figure 1 shown. It mainly includes the following steps:
[0061] (1) Obtain training data and perform data augmentation: The present invention uses the publicly available and widely used DF2K (DIV2K + Flickr2K) dataset as training data. At the same time, in order to further improve the generalization ability and robustness of the model, enhancement means such as cropping, rotation, and flipping are performed on the obtained training data.
[0062] (2) Unsupervised degradation representation learning: The present invention uses a contrastive learning framework to train a degradation encoder to obtain the degradation representation of image patches. First, randomly crop image patches from the low-resolution (LR) images obtained in step (1); then, construct positive and negative sample pairs: take one image patch as the query sample Q, and define other image patches from the same LR image as the positive sample Q + , and define image patches from different LR images as the negative sample Q - . Q, Q + and Q - are respectively passed through an encoder (the specific structure is as Figure 3 shown, mainly composed of an encoding unit, an encoding residual activation unit, and a perceptron) to obtain the corresponding feature vectors F dpQ , F dpQ+ and F dpQ- , where F dpQ- is saved in a queue. The present invention constrains the InfoNCE contrastive loss to make the feature vector F dpQ of the query sample as similar as possible to the feature vector F dpQ+ of its positive sample in the feature space, and as dissimilar as possible to the feature vector F dpQ- of the negative samples in the queue. After training, the degradation encoder can extract feature representations that are independent of the image content and only related to the degradation type.
[0063] (3) Shallow feature extraction: Use convolutional layer 1 to map the LR image input in step (1) to a higher-dimensional feature space and extract the shallow feature F0.
[0064] (4)Deep Feature Extraction and Fusion: In the embodiments of the present invention, a series of Degradation-aware Graph Construction and Aggregation Blocks (DGCAB) are constructed to effectively extract the deep features of the image and fuse the degradation information. As Figure 3 shown, this module is mainly stacked by an Adaptive Graph Construction Unit (AGCU) and a Graph Aggregation Unit (GAU). The input of the i-th DGCAB i is the output F i-1 of the previous DGCAB i-1 and the degradation representation F dp . The input of the first DGCAB1 is the shallow feature F0 obtained in steps 2 and 3 or the previous degradation representation F dp . In the DGCAB i+1 , first, the Adaptive Graph Construction Unit AGCU transforms the degradation representation F dp through a linear layer to obtain two modulation coefficients α and β. Then, the normalized input feature F Li is multiplied by α, added to β, and then added to the feature F Li to achieve the fusion of the degradation information and the input feature F i , and the feature F FAi0 is obtained. Subsequently, based on the feature F FAi0 , the Detail Significance Index DS is calculated: F FAi0 is bicubic downsampled and then bicubic upsampled to obtain the feature F FAi0↓↑ ; the absolute value sum of the corresponding pixel differences between F FAi0 and F FAi0↓↑ on all feature channels is calculated, that is, DS = Σ|F FAi0 -F FAi0↓↑|. Next, AGCU constructs a local graph and a global graph according to DS simultaneously. For the local graph, AGCU will search for the k1 pixels with the most similar features within its spatial neighborhood for each pixel node and connect these pixels to the current pixel node. For the global graph, AGCU first samples all pixels on the feature map with a step size of s, and then for each sampled pixel node, calculates the k2 pixels with the most similar features among all sampled pixels and connects them. The numbers of k1 and k2 will be dynamically adjusted according to the DS value (the numbers of k1 and k2 will be dynamically adjusted according to the DS value. In regions with rich details (points with high DS values), the number of connections (k1, k2) is more, and in smooth regions (points with low DS values), the number of connections (k1, k2) is less. The k1 and k2 of each pixel node are proportional to the DS value), with more connections in regions with rich details and fewer connections in smooth regions. Next, the feature F FAi0 along with the local graph and the global graph will be input into a series of graph aggregation units (GAU ij ). In the first GAU i1 , the edge aggregator Grapher and the channel attention module will determine the set of neighbor nodes for each pixel node in the input feature F FAi0 based on the local graph / global graph, and then perform edge aggregation and important channel feature enhancement; subsequently, the convolutional feed-forward network (ConvFFN) module will perform further feature activation extraction and output the feature F FAi1 to the next graph aggregation unit GAU i2 . Multiple GAUs are stacked in an alternating manner of receiving the local graph and the global graph, enabling full interaction between local and global information.
[0065] (5) High-resolution image reconstruction: Convert the enhanced feature finally output in step (4) into the final high-resolution image. First, use a pixel recombination layer to upsample the feature map output in step (4). Subsequently, use a convolutional layer to map the upsampled feature map to the HR image space.
[0066] (6) Construct the loss function: The training objective of the DAG-BSR model is to make the reconstructed HR image as similar as possible to the real HR image while ensuring that the degradation encoder can extract effective degradation representations. To achieve this goal, the present invention constructs an overall loss function including a reconstruction loss and a contrast loss. For the super-resolution reconstruction part, the L1 loss is used as the reconstruction loss. For the degradation representation learning part, the contrast loss mentioned in step 2 is used, and the overall loss function of the model is the sum of the reconstruction loss and the contrast loss.
[0067] (7)Model training and testing: In the present invention, hyperparameters such as the initial learning rate, weight decay coefficient, batch size, etc. are set, and the learning rate decay strategy and Adam optimizer are used for model training. The process is as follows: The LR image is input into the network to obtain the reconstructed HR image; the L1 loss and contrast loss are calculated; the gradient is calculated by backpropagation to update the network parameters; this process is repeated until convergence or a predetermined number of epochs is reached. Finally, the test LR image is input into the trained network to obtain the reconstructed HR image, and metrics such as PSNR and SSIM are calculated to evaluate the model performance.
[0068] Experimental results:
[0069] To verify the effectiveness of the DAG-BSR method, experiments were conducted on multiple commonly used benchmark datasets, and the results were compared with some existing super-resolution methods. Table 1 shows the PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity) metrics of different methods on the Set5, Set14, B100, and Urban100 datasets.
[0070] From Table 1 to Table 2, Figure 8 it can be seen that whether under the condition of isotropic Gaussian kernel and noiseless degradation or under the condition of anisotropic Gaussian kernel and different levels of noise degradation, compared with the existing methods, DAG-BSR achieved better results on all datasets. Especially on the Urban100 dataset, the PSNR improvement was more obvious. The Urban100 dataset contains a large number of images with rich texture and structural information, indicating that DAG-BSR has stronger advantages in processing complex images.
[0071] Table 1 PSNR↑ results with isotropic Gaussian kernel and noiseless degradation (×2 and ×3 experiments on standard datasets)
[0072]
[0073] Table 2 PSNR↑ and SSIM↑ results with isotropic Gaussian kernel and noiseless degradation (×2, ×3, and ×4 experiments on the Manga109 dataset)
[0074]
[0075] Example 2
[0076] The second embodiment of the present invention provides a terminal device corresponding to the first embodiment above. The terminal device can be a processing device for a client, such as a mobile phone, laptop computer, tablet computer, desktop computer, etc., to execute the method of the above embodiment.
[0077] The terminal device in this embodiment includes a memory, a processor, and a computer program stored on the memory; the processor executes the computer program on the memory to implement the steps of the method in Embodiment 1 above.
[0078] In some implementations, the memory may be a high-speed random access memory (RAM: Random Access Memory), and may also include non-volatile memory, such as at least one disk memory.
[0079] In other implementations, the processor may be a general-purpose processor of various types such as a central processing unit (CPU) or a digital signal processor (DSP), which is not limited here.
[0080] Embodiment 3
[0081] Embodiment 3 of the present invention provides a computer-readable storage medium corresponding to Embodiment 1 above, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of the method in Embodiment 1 above are implemented.
[0082] A computer-readable storage medium may be a tangible device that holds and stores instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination of the above.
[0083] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The solutions in the embodiments of the present application can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.
[0084] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementation in the processFigure 1 means for the functions specified in one process or multiple processes and / or boxes Figure 1 or multiple boxes.
[0085] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the functions specified in one process Figure 1 or multiple processes and / or boxes Figure 1 or multiple boxes.
[0086] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0087] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. An image reconstruction method, characterized in that, It includes the following steps: Randomly crop the low-resolution images in the image dataset to obtain image patches; Extract the shallow features of the low-resolution images; Take a certain image patch as a query sample, define other image patches from the same low-resolution image as the query sample as positive samples, and define image patches from different low-resolution images as negative samples; Extract the feature vectors of the query sample, positive samples, and negative samples respectively; extract the deep features of the shallow features and the feature vectors, enhance the deep features to obtain enhanced features; transform the enhanced features into high-resolution images; The specific implementation process of extracting the deep features of the shallow features and the feature vectors, and enhancing the deep features to obtain enhanced features includes: taking the shallow features and the feature vectors as the input of the first graph construction and aggregation block to obtain a first output; taking the first output and the feature vector as the input of the second graph construction and aggregation block to obtain a second output; and so on; taking the output of the last graph construction and aggregation block as the input of a convolutional layer, splicing the output of the convolutional layer and the shallow features, taking the spliced features as the input of a pixel recombination layer, and taking the output of the pixel recombination layer as the input of an output convolutional layer to obtain enhanced features.
2. The image reconstruction method according to claim 1, wherein Use an encoder to extract the feature vectors of the query sample, positive samples, and negative samples respectively; the encoder includes multiple cascaded encoding units, where the first encoding unit is connected to the second encoding unit through an encoding residual activation module; the encoding residual activation module includes multiple cascaded encoding residual activation units; the last encoding unit is connected to a perceptron.
3. The image reconstruction method according to claim 2, wherein The encoding residual activation unit includes two convolutional layers, and the two convolutional layers are connected through an activation function layer.
4. The image reconstruction method according to claim 1, wherein The graph construction and aggregation block processes the input through the following steps: perform a linear transformation on the feature vector to obtain modulation coefficients α and β; multiply the normalized shallow features or the output of the i-th graph construction and aggregation block by α, add the multiplication result to β, and then add the added result to the shallow features or the output of the i-th graph construction and aggregation block to obtain a first output feature; perform bicubic downsampling on the first output feature and then perform bicubic upsampling to obtain a second output feature; calculate the absolute value sum DS of the corresponding pixel differences between the first output feature and the second output feature on all feature channels, and construct a local graph and a global graph according to the absolute value sum DS; for each pixel node of the local graph, search for the k1 pixels with the most similar features to the pixel node in the spatial neighborhood of the pixel node, and connect the k1 pixels to the pixel node; For the global graph, all pixels in the global graph are sampled at a set step size. For each sampled pixel node, k2 pixels with the most similar features to this pixel node are calculated among all sampled pixels, and these k2 pixels are connected to the pixel node; the first output feature, the global graph, and the local graph are input into the graph aggregation unit module, and the output of the graph aggregation unit module is concatenated with the shallow feature after passing through a convolutional layer to obtain the output of the graph construction and aggregation block.
5. The image reconstruction method according to claim 4, wherein The graph aggregation unit module includes multiple cascaded graph aggregation units. The input of the first graph aggregation unit is the first output feature and the global graph; the input of the second graph aggregation unit is the output of the first graph aggregation unit and the local graph; the input of the third graph aggregation unit is the output of the second graph aggregation unit and the global graph; and so on. Among them, the graph aggregation unit includes an edge aggregator and a first normalization layer connected to the edge aggregator; one channel attention module is connected to the input side of the edge aggregator. The output of the channel attention module, the output of the first normalization layer, and the input of the graph aggregation unit are concatenated and then input into a convolutional feed-forward network, which is connected to a second normalization layer. The output of the second normalization layer and the input of the convolutional feed-forward network after concatenation are the output of the graph aggregation unit.
6. The image reconstruction method according to claim 1, wherein The specific implementation process of converting the enhanced feature into a high-resolution image includes: performing upsampling on the enhanced feature; performing a convolutional operation on the upsampled feature map to map it to the HR image space, that is, obtaining a high-resolution image.
7. An image reconstruction system, comprising a memory, a processor, and a computer program stored on the memory; characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program / instructions stored thereon; characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer program product, comprising a computer program / instructions; characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Image super-resolution reconstruction method based on DRRN network
CN118247143A