Image reconstruction method and system and storage medium
By combining comparative learning and adaptive graph construction with aggregation blocks, the problems of low efficiency of image structure information utilization and task imbalance in blind super-resolution tasks are solved, and stronger degradation robustness and high-quality reconstruction effects are achieved.
Patent Information
- Application Number
- CN202510617879.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The prior art has limitations in dealing with unknown and complex and variable image degradation (blind super-resolution tasks), making it difficult to effectively utilize image structure information and solve task imbalance problem.
Image blocks are obtained by randomly cropping the low-resolution images in the image dataset, shallow features are extracted, and a contrast learning task is constructed to learn feature representations under different degradation modes. Combining adaptive graph construction and aggregation blocks, information aggregation is alternately used to aggregate local graphs and global graphs to generate richer and more robust deep feature representations.
It significantly improves the generalization performance of the model in real and complex degradation scenarios, enhances the degradation robustness, realizes adaptive feature extraction and multi-scale information fusion, and improves the reconstruction quality.
Smart Images

Figure CN120147137A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image processing, computer vision, and deep learning, and particularly to an image reconstruction method, system, and storage medium. Background Art
[0002] Image super-resolution (SR) aims to recover a high-resolution (HR) image from a low-resolution (LR) image, and this technology has a wide range of applications in fields such as video surveillance, medical imaging, and remote sensing image analysis.
[0003] Traditional SR methods mainly include interpolation-based methods and reconstruction-based methods. Interpolation-based methods, such as nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation, estimate unknown pixels in the HR image by weighted averaging of known pixels in the LR image. These methods are computationally simple and easy to implement, but the reconstructed HR images are usually too smooth and lack high-frequency details. To overcome the limitations of interpolation methods, reconstruction-based methods such as iterative back projection and maximum a posteriori probability have been proposed. These methods mainly attempt to model the imaging process to recover the HR image, but the models are usually complex, computationally intensive, and sensitive to noise.
[0004] With the rise of deep learning technology, this technology has also been widely applied to SR tasks. SRCNN first introduced CNN into SR tasks and adopted a relatively simple three-layer convolutional network structure. Subsequently, a series of CNN-based SR methods have been proposed successively and have achieved significant performance improvements. For example, VDSR significantly increased the depth of the network by introducing residual learning and gradient clipping strategies. EDSR further explored network structure design by removing the batch normalization (BN) layer and expanding the model scale. RCAN proposed a channel attention mechanism that enables the network to adaptively focus on more important feature channels. Compared with traditional methods, these CNN-based methods show significant advantages when dealing with images with known and fixed degradations (such as bicubic downsampling), but their inherent technical limitations still restrict further performance improvement. Convolutional operations are essentially a local feature extraction mechanism, and the size of their convolutional kernels directly restricts the network's ability to understand global image semantics. Under this limitation, it is difficult for the network to effectively capture long-range dependencies in the image; secondly, as the depth of the network continues to increase, the problem of gradient disappearance becomes increasingly prominent, seriously affecting the learning efficiency and performance of the network; at the same time, the processing method of sharing the same convolutional kernel for all pixels ignores the diversity of image content and the imbalance of SR tasks.
[0005] To overcome the problem of limited receptive fields in CNNs, subsequent researchers have attempted to introduce the Transformer architecture to alleviate the above issues by introducing self-attention mechanisms. SwinIR incorporates the idea of Swin Transformer into SR, computing self-attention within local windows and adopting a shifted window mechanism to expand the receptive field. Although the Transformer architecture theoretically breaks through the receptive field limitation of convolutional networks, in practical applications, methods based on Transformer also face numerous challenges. For example, the computational complexity increases exponentially, raising the requirements for computing resources. The standard self-attention mechanism treats all parts of the image equally and does not consider the difference in reconstruction difficulty across different regions of the image.
[0006] Currently, graph neural networks (GNNs) have also been applied to super-resolution (SR) tasks due to their ability to flexibly handle irregular structures and effectively model complex relationships between nodes. DLGNN uses graph neural networks to exploit non-local image dependencies through patch self-similarity modeling and adopts a dual learning mechanism to optimize the reconstruction process. IGNN proposes an iterative graph neural network with multi-scale feature fusion for super-resolution tasks. However, many existing graph neural network-based super-resolution methods typically use the k-nearest neighbor method for graph construction, and the node degrees of all image regions are fixed. This approach is difficult to adapt to the inherent long-tail distribution characteristics in super-resolution tasks. Additionally, existing methods usually construct only one of a local graph or a global graph. Although IPG improves the flexibility of graph construction by introducing a detail-aware metric and a multi-scale graph aggregation scheme, it does not explicitly consider the image degradation process. In real-world scenarios, the image degradation process is often unknown and complex, which is known as blind super-resolution (Blind SR). The actual degradation process may involve various combinations of blur kernels and noise. When there is a deviation between the actual degradation and the assumption, the performance of the above methods will significantly decline.
[0007] To address the blind SR problem, existing techniques can be mainly divided into two categories: explicit kernel estimation methods and implicit degradation learning methods. Explicit kernel estimation methods first attempt to estimate the degradation model of the LR image and treat it as a system identification problem. Once the degradation model is obtained, it can be used in traditional non-blind SR methods or a specific network structure can be designed to handle the estimated degradation. For example, KernelGAN uses a generative adversarial network (GAN) to estimate the blur kernel of a single LR image; IKC proposes an iterative correction method that alternates between blur kernel estimation and SR reconstruction. However, these methods generally have a high computational complexity, require optimization for each image, and the accuracy of degradation estimation directly affects the quality of SR reconstruction. Estimation errors can lead to artifacts in the reconstruction results.
[0008] Implicit degradation learning methods do not directly estimate the degradation kernel. Instead, they learn a representation of degradation through a neural network and then use this representation to guide the subsequent super-resolution reconstruction process. One type is domain adaptation-based methods, which attempt to adapt a pre-trained SR model (usually trained on known degradations) to a new, unknown degradation domain. For example, ZSSR performs online fine-tuning using specific LR images during the testing phase to adapt the model to the current degradation. However, this method requires online optimization for each LR image, resulting in a large computational overhead and is not suitable for real-time applications. Another type is unsupervised learning-based methods, which do not require paired LR-HR images for training. Instead, they learn degradation information from unlabeled data through contrastive learning or other unsupervised / self-supervised methods. For example, DASR uses contrastive learning for unsupervised degradation representation learning, improving the model's generalization ability. Based on DASR, DSAT changes the deep feature extraction module from CNN to Transformer to utilize the global modeling ability of Transformer. However, the feature extraction ability of DASR is limited by the local receptive field of CNN, while DSAT faces the problems of high computational complexity of Transformer and difficulty in capturing pixel-level fine features. In addition, neither of them fully utilizes the structural information of the image nor considers the imbalance of the super-resolution task itself. Summary of the Invention
[0009] The technical problem to be solved by the present invention is to provide an image reconstruction method, system, and storage medium to improve the utilization efficiency of image structural information and solve the task imbalance problem in view of the deficiencies of the prior art.
[0010] To solve the above technical problems, the technical solution adopted by the present invention is: An image reconstruction method, comprising the following steps: Randomly crop the low-resolution images in the image dataset to obtain image patches; extract the shallow features of the low-resolution images; Taking a certain image patch as a query sample, define other image patches from the same low-resolution image as the query sample as positive samples, and define image patches from different low-resolution images as negative samples; extract the feature vectors of the query sample, positive samples, and negative samples respectively; Extract the deep features of the shallow features and the feature vectors, enhance the deep features, and obtain enhanced features; Convert the enhanced features into high-resolution images.
[0011] The present invention provides an image reconstruction method, aiming to overcome the limitations existing in the prior art when dealing with unknown and complex image degradation (i.e., blind super-resolution tasks), and improve the utilization efficiency of image structure information and solve the problem of task imbalance. Specifically, the method of the present invention randomly crops image patches from the low-resolution images in the image dataset, and extracts the shallow features of the low-resolution images. On this basis, taking a certain image patch as a query sample, a contrastive learning task is constructed. The other image patches from the same low-resolution image as the query sample are defined as positive samples, and the image patches from different low-resolution images from the query sample are defined as negative samples. The encoder is used to extract the feature vectors of the query sample, positive sample, and negative sample respectively. Through the contrastive learning mechanism, the present invention enables the network to learn the feature representations under different degradation modes, making the image patches from the same degradation source have similar features, while pulling apart the feature distances of the image patches from different degraded images. This strategy enables the network to learn a more robust feature representation for different degradations, thereby significantly improving the generalization performance of the model in real complex degradation scenarios without explicitly estimating the degradation kernel or performing time-consuming online fine-tuning. Subsequently, the deep features of the shallow features and the feature vectors (providing degradation-aware information strengthened by contrastive learning) are extracted, and the deep features are enhanced. At this stage, by alternately combining the global graph structure and the local structure, different levels of information are effectively fused to generate a richer and more robust deep feature representation, laying a solid foundation for the final conversion of high-quality high-resolution images.
[0012] The encoder is used to extract the feature vectors of the query sample, positive sample, and negative sample respectively; the encoder includes multiple cascaded coding units, where the first coding unit is connected to the second coding unit through a coding residual activation module; the coding residual activation module includes multiple cascaded coding residual activation units; the last coding unit is connected to the perceptron. The multi-cascaded structure of the present invention enables the encoder to learn hierarchical features from simple to complex. The coding unit is used for simple feature extraction. The coding residual activation unit helps to optimize the transformation and transmission of features, ensuring the effectiveness and robustness of the extracted features. The perceptron maps the high-dimensional features output by the last coding unit to the required fixed-dimensional feature vector, and this vector is used as the input of the contrastive learning task, ensuring that the output features are suitable for subsequent loss calculation and optimization processes.
[0013] The encoded residual activation unit includes two convolutional layers, and the two convolutional layers are connected by an activation function layer. This forms a basic local feature extraction and transformation unit, where the convolutional layer is used to capture spatial patterns, and the activation function layer introduces non-linearity, enabling the unit to learn complex non-linear mapping relationships, thereby effectively extracting non-linear features in image patches. The residual connection provides a direct channel for gradient propagation, greatly alleviating the problem of vanishing gradients that may occur when stacking multiple such units. Thanks to this design, a single encoded residual activation unit can be stably trained and can serve as a robust building block. By cascading, deeper and more complex encoded residual activation modules can be formed, which in turn support the construction of the entire depth encoder, thereby learning complex, robust, and discriminative image patch features that are crucial for the blind super-resolution task.
[0014] The specific implementation process of extracting the shallow features and the deep features of the feature vector and enhancing the deep features includes: Taking the shallow features and the feature vector as the input of the first graph construction and aggregation block to obtain a first output; Taking the first output and the feature vector as the input of the second graph construction and aggregation block to obtain a second output; And so on; Taking the output of the last graph construction and aggregation block as the input of a convolutional layer, concatenating the output of the convolutional layer with the shallow features, taking the concatenated features as the input of a pixel recombination layer, and taking the output of the pixel recombination layer as the input of an output convolutional layer to obtain the enhanced features.
[0015] The graph construction and aggregation block performs a series of processes on the input, aiming to adaptively construct a graph structure and perform information aggregation based on the feature vector (providing degradation-aware information strengthened by contrastive learning). This way of constructing a graph based on the feature vector and performing information aggregation, different from a fixed-structure graph (such as a simple kNN graph), can better reflect the diversity of image content and degradation patterns, making the relationship modeling more flexible and targeted, thereby effectively handling the reconstruction difficulty differences in different regions of the super-resolution task. The design of concatenating the output of the last graph construction and aggregation block with the original shallow features after passing through a convolutional layer aims to combine the features after complex graph processing and deep transformation with the original shallow features, effectively retaining the detail and edge information of the image and avoiding detail loss that may be caused by deep processing.
[0016] The graph construction and aggregation block processes the input including the following steps: Performing a linear transformation on the feature vector to obtain modulation coefficients α and β; Multiply the normalized shallow feature / the output of the i-th graph construction and aggregation block by α, add the result of the multiplication to β, and then add the obtained addition result to the shallow feature / the output of the i-th graph construction and aggregation block to obtain an addition feature; Add the addition feature to the normalized shallow feature / the output of the i-th graph construction and aggregation block to obtain a first output feature; Perform bicubic downsampling on the first output feature and then perform bicubic upsampling to obtain a second output feature; Calculate the absolute value sum DS of the corresponding pixel differences between the first output feature and the second output feature on all feature channels, and construct a local graph and a global graph according to the absolute value sum DS; For each pixel node of the local graph, search for the k1 pixels with the most similar features to this pixel node within the spatial neighborhood of this pixel node, and connect these k1 pixels to this pixel node; for the global graph, sample all pixels in the global graph at a set step size. For each sampled pixel node, calculate the k2 pixels with the most similar features to this pixel node among all sampled pixels, and connect these k2 pixels to this pixel node; Input the first output feature, the global graph, and the local graph into the graph aggregation unit module, and splice the output of the graph aggregation unit module after passing through a convolutional layer with the shallow feature to obtain the output of the graph construction and aggregation block.
[0017] Bicubic downsampling is to calculate and select some pixel points according to the bicubic interpolation function for the original feature map to construct a new image, and the size of the new image becomes smaller.
[0018] Bicubic upsampling also uses the bicubic interpolation function. Based on the pixels of the original image, calculate the weights to estimate the values of the newly generated pixels, thereby increasing the resolution of the image and enlarging the size of the original image.
[0019] The image obtained by performing bicubic upsampling after bicubic downsampling has the same size as the original image, and can be used to calculate DS.
[0020] The process of performing a linear transformation on the feature vector to obtain modulation coefficients, and using these coefficients to modulate the output of the aggregation block with the normalized shallow features / the i-th graph to obtain the first output feature incorporates degradation-aware information into the construction of the shallow features / the i-th graph and the output of the aggregation block, enhancing the feature representation ability. Based on the first output feature and the second output feature, the absolute value sum of pixel differences DS is calculated, and based on DS, a local graph and a global graph are constructed simultaneously. The local graph focuses on the similar pixel relationships within the spatial neighborhood, capturing fine structures; the global graph models the similarity between pixels at a greater distance through sampling, capturing long-range dependencies. Utilizing both the local graph and the global graph can comprehensively model the structural and relational information of the image, overcoming the limitations of existing methods that use only local graphs or only global graphs.
[0021] The graph input graph aggregation unit module includes multiple cascaded graph aggregation units. The input of the first graph aggregation unit is the first output feature and the global graph; the input of the second graph aggregation unit is the output of the first graph aggregation unit and the local graph; the input of the third graph aggregation unit is the output of the second graph aggregation unit and the global graph; and so on. Among them, the graph aggregation unit includes an edge aggregator and a first normalization layer connected to the edge aggregator; one-channel attention module's input side is connected to the input side of the edge aggregator. The output of the channel attention module, the output of the first normalization layer, and the input of the graph aggregation unit are concatenated and then input into a convolutional feed-forward network, which is connected to a second normalization layer. The output of the second normalization layer and the input of the convolutional feed-forward network concatenated together is the output of the graph aggregation unit. The edge aggregator and the normalization layer are used to perform information aggregation based on graph nodes and ensure numerical stability. The channel attention module allows the network to adjust the importance of feature channels according to the input features during the information aggregation process, making the aggregation process more focused on the most useful feature channels for the current situation. The convolutional feed-forward network combines the characteristics of convolutional and fully connected networks and is used to perform more complex non-linear transformations on node features after graph aggregation. The residual connection inside the graph aggregation unit helps to construct deeper graph aggregation units and modules, ensuring the effective propagation of gradients and improving training stability. Through this complex unit structure including edge aggregation, channel attention, convolutional feed-forward network, and internal residual connection, as well as the alternating input strategy of the global graph and the local graph, the graph aggregation unit can efficiently and hierarchically aggregate information on the graph, extract and continuously enhance deep feature representations.
[0022] The specific implementation process of converting the enhanced features into a high-resolution image includes: Performing upsampling on the enhanced features; Performing a convolution operation on the upsampled feature map to map it to the HR image space, that is, obtaining a high-resolution image.
[0023] As an inventive concept, the present invention also provides an image reconstruction system, including a memory, a processor, and a computer program stored on the memory; the processor executes the computer program to implement the steps of the above method.
[0024] As an inventive concept, the present invention also provides a computer-readable storage medium, on which a computer program / instructions are stored; when the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0025] As an inventive concept, the present invention also provides a computer program product, including a computer program / instructions; when the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0026] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) Enhanced degradation robustness: DAG-BSR uses unsupervised contrast learning to obtain degradation representations, without relying on prior knowledge of image degradation types and without the need for explicit degradation estimation. Therefore, this method can adapt to various unknown and complex degradation types and has stronger degradation robustness. (2) Achieved adaptive feature extraction: In the Adaptive Graph Construction and Aggregation Block (DGCAB), the adaptive graph construction unit can dynamically adjust the number of connections of each pixel node according to the image details (DS) to construct content-adaptive local and global graphs. This design enables the network to more flexibly handle the reconstruction requirements of different regions in the image. Regions with rich details obtain more information aggregation, while smooth regions perform less information aggregation, thus more effectively utilizing computing resources.
[0027] (3) Integrated multi-scale information: The graph aggregation unit in DGCAB effectively integrates multi-scale information of the image by alternately using local and global graphs for information aggregation. The local graph focuses on details within the pixel neighborhood, contributing to the reconstruction of textures and edges; the global graph captures long-range dependencies in the image, contributing to the restoration of the overall structure.
[0028] (4) Achieved effective degradation-aware fusion: DAG-BSR effectively fuses degradation information into the process of feature extraction and image reconstruction through degradation-aware feature modulation in the adaptive graph construction unit and the processing of the modulated features in GAU. This enables the network to adaptively adjust the reconstruction strategy according to different degradation situations, thereby improving the adaptability to various degradation types.
[0029] (5) Improved reconstruction quality: Thanks to the above advantages, DAG-BSR can extract more discriminative features and effectively fuse them with the learned degradation representations, thereby generating high-quality reconstructed images. Brief Description of the Drawings
[0030] Figure 1 is the flowchart of the method according to the embodiment of the present invention; Figure 2 is the block diagram of the DAG - BSR structure according to the embodiment of the present invention; Figure 3 is the block diagram of the degradation encoder structure according to the embodiment of the present invention; Figure 4 is the structure diagram of the encoding unit according to the embodiment of the present invention; Figure 5 is the structure diagram of the encoding residual activation unit according to the embodiment of the present invention; Figure 6 is the structure diagram of the degradation - aware graph construction and aggregation block according to the embodiment of the present invention; Figure 7 is the structure diagram of the adaptive graph construction unit and the graph aggregation unit; Figure 8 is the PSNR↑ result with anisotropic Gaussian kernel and different levels of noise (×4 experiment on the Set14 dataset). Detailed Description of the Embodiments
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] In the embodiments of the present invention, the degradation representations obtained by unsupervised contrastive learning are incorporated into the process of constructing a dynamic pixel graph through feature modulation; the dynamic graph construction strategy based on image details is combined with the alternating use of local graphs and global graphs to achieve the adaptive fusion of multi - scale information; and the above mechanism is repeatedly applied in multiple layers of the graph neural network to form an organic whole. In addition, for the blind super - resolution task, the DGCAB block is designed, in which the AGCU combines degradation - aware feature modulation with adaptive graph construction, and the GAU aggregates local graph and global graph features.
[0033] Embodiment 1 The embodiments of the present invention propose a blind image super - resolution method based on graph neural network (GNN) and unsupervised degradation - aware learning, called Degradation - Aware Graph - Network for Blind Image Super - Resolution (DAG - BSR). The overall process is as Figure 1As shown in the figure. It mainly includes the following steps: (1) Obtain training data and perform data augmentation: The present invention uses the publicly available and widely used DF2K (DIV2K + Flickr2K) dataset as the training data. At the same time, in order to further improve the generalization ability and robustness of the model, enhancement means such as cropping, rotation, and flipping are performed on the obtained training data.
[0034] (2) Unsupervised degradation representation learning: The present invention uses a contrastive learning framework to train a degradation encoder to obtain the degradation representation of image patches. First, randomly crop image patches from the low-resolution (LR) images obtained in step (1); then, construct positive and negative sample pairs: Take an image patch as the query sample Q, and define other image patches from the same LR image as the positive sample Q + , and define image patches from different LR images as the negative sample Q - . Q, Q + and Q - are respectively passed through the encoder (the specific structure is as Figure 3 shown, mainly composed of an encoding unit, an encoding residual activation unit, and a perceptron) to obtain the corresponding feature vectors F dpQ , F dpQ+ and F dpQ- , where F dpQ- is saved in a queue. The present invention constrains the InfoNCE contrastive loss to make the feature vector F dpQ of the query sample as similar as possible to the feature vector F dpQ+ of its positive sample in the feature space, and as dissimilar as possible to the feature vector F dpQ- of the negative samples in the queue. After training, the degradation encoder can extract feature representations that are independent of the image content and only related to the degradation type.
[0035] (3) Shallow feature extraction: Use convolutional layer 1 to map the LR image input in step (1) to a higher-dimensional feature space and extract the shallow feature F 0 .
[0036] (4) Deep feature extraction and fusion: The embodiment of the present invention constructs a series of degradation-aware graph construction and aggregation blocks (DGCAB) to effectively extract the deep features of the image and fuse the degradation information. As Figure 3As shown, this module is mainly composed of a stacked Adaptive Graph Construction Unit (AGCU) and Graph Aggregation Unit (GAU). The input of the i-th DGCAB i is the output F i-1 of the previous DGCAB i-1 and the degraded representation F dp . The input of the first DGCAB 1 is the shallow feature F 0 obtained in Steps 2 and 3 or the previous degraded representation F dp . In the DGCAB i+1 , first, the Adaptive Graph Construction Unit AGCU transforms the degraded representation F dp through a linear layer to obtain two modulation coefficients α and β. Then, the normalized input feature F Li is multiplied by α, added to β, and then added to the feature F Li to achieve the fusion of the degraded information and the input feature F i , obtaining the feature F FAi0 . Subsequently, based on the feature F FAi0 , the Detail Significance Index DS is calculated: F FAi0 is bicubic downsampled and then bicubic upsampled to obtain the feature F FAi0↓↑ ; the sum of the absolute values of the corresponding pixel differences between F FAi0 and F FAi0↓↑ is calculated on all feature channels, i.e., DS = Σ|F FAi0 - F FAi0↓↑ |. Next, the AGCU constructs both a local graph and a global graph according to DS. For the local graph, the AGCU searches for the k1 pixels with the most similar features to the current pixel node within its spatial neighborhood for each pixel node and connects these pixels to the current pixel node. For the global graph, the AGCU first samples all pixels on the feature map with a stride of s, and then for each sampled pixel node, calculates the k2 pixels with the most similar features to it among all sampled pixels and connects them. The numbers of k1 and k2 are dynamically adjusted according to the DS value (the numbers of k1 and k2 are dynamically adjusted according to the DS value. In regions with rich details (points with high DS values), the number of connections (k1, k2) is more, and in smooth regions (points with low DS values), the number of connections (k1, k2) is less. The k1 and k2 of each pixel node are proportional to the DS value), with more connections in regions with rich details and fewer connections in smooth regions. Next, the feature F FAi0 as well as the local graph and the global graph are input into a series of Graph Aggregation Units (GAU ij ). In the first GAU i1In it, the edge aggregator Grapher and the channel attention module determine the input feature F according to the local graph / global graph FAi0 for each pixel node in the neighbor node set, and then perform edge aggregation and important channel feature enhancement; subsequently, the convolutional feed-forward network (ConvFFN) module will perform further feature activation extraction and output the feature F FAi1 to the next graph aggregation unit GAU i2 。Multiple GAUs are stacked in an alternating manner to receive local and global graphs, enabling full interaction between local and global information.
[0037] (5)High-resolution image reconstruction: Convert the enhanced features finally output in step (4) into the final high-resolution image. First, use a pixel reorganization layer to upsample the feature map output in step (4). Subsequently, use a convolutional layer to map the upsampled feature map to the HR image space.
[0038] (6)Construct the loss function: The training objective of the DAG-BSR model is to make the reconstructed HR image as similar as possible to the real HR image, while ensuring that the degradation encoder can extract effective degradation representations. To achieve this goal, the present invention constructs an overall loss function including a reconstruction loss and a contrast loss. For the super-resolution reconstruction part, the L1 loss is used as the reconstruction loss. For the degradation representation learning part, the contrast loss mentioned in step 2 is used, and the overall loss function of the model is the sum of the reconstruction loss and the contrast loss.
[0039] (7)Model training and testing: The present invention sets hyperparameters such as the initial learning rate, weight decay coefficient, batch size, etc., and uses a learning rate decay strategy and an Adam optimizer for model training. The process is as follows: The LR image is input into the network to obtain the reconstructed HR image; calculate the L1 loss and the contrast loss; backpropagate to calculate the gradient and update the network parameters; repeat this process until convergence or reach a predetermined number of epochs. Finally, input the test LR image into the trained network to obtain the reconstructed HR image, and calculate metrics such as PSNR and SSIM to evaluate the model performance.
[0040] Experimental results: To verify the effectiveness of the DAG-BSR method, experiments were carried out on multiple commonly used benchmark datasets, and the results were compared with some existing super-resolution methods. Table 1 shows the PSNR (peak signal-to-noise ratio) and SSIM (structural similarity) metrics of different methods on the Set5, Set14, B100, and Urban100 datasets.
[0041] From Table 1 to Table 2, Figure 8It can be seen that whether under the condition of isotropic Gaussian kernel and noiseless degradation or under the condition of anisotropic Gaussian kernel and different levels of noise degradation, compared with existing methods, DAG-BSR has achieved better results on all datasets. Especially on the Urban100 dataset, the PSNR improvement is more obvious. The Urban100 dataset contains a large number of images with rich texture and structural information, which indicates that DAG-BSR has stronger advantages in processing complex images.
[0042] Table 1 PSNR↑ results with isotropic Gaussian kernel and noiseless degradation (×2 and ×3 experiments on standard datasets)
[0043] Table 2 PSNR↑ and SSIM↑ results with isotropic Gaussian kernel and noiseless degradation (×2, ×3, and ×4 experiments on the Manga109 dataset)
[0044] Example 2 Example 2 of the present invention provides a terminal device corresponding to Example 1 above. The terminal device can be a processing device for a client, such as a mobile phone, a laptop computer, a tablet computer, a desktop computer, etc., to execute the method of the above example.
[0045] The terminal device of this example includes a memory, a processor, and a computer program stored on the memory; the processor executes the computer program on the memory to implement the steps of the method of Example 1 above.
[0046] In some implementations, the memory can be a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory.
[0047] In other implementations, the processor can be a central processing unit (CPU), a digital signal processor (DSP), or various other types of general-purpose processors, which are not limited here.
[0048] Example 3 Example 3 of the present invention provides a computer-readable storage medium corresponding to Example 1 above, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, the steps of the method of Example 1 above are implemented.
[0049] A computer-readable storage medium can be a tangible device that retains and stores instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination of the foregoing.
[0050] Those skilled in the art will appreciate that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.
[0051] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0052] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0053] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0054] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to cover these changes and modifications.
Claims
1. An image reconstruction method, characterized in that: The following steps are involved: Randomly cropping low-resolution images in an image data set to obtain image blocks; extracting shallow features of the low-resolution images; Taking a certain image block as a query sample, defining other image blocks from the same low-resolution image as the query sample as positive samples, and defining image blocks from low-resolution images different from the query sample as negative samples; extracting feature vectors of the query sample, positive samples, and negative samples respectively; Extracting the shallow features and the deep features of the feature vector, enhancing the deep features, and obtaining enhanced features; The enhanced features are converted into a high-resolution image.
2. The image reconstruction method according to claim 1, characterized in that: The encoder is used to extract the feature vectors of the query sample, positive sample and negative sample respectively; the encoder includes a plurality of cascaded encoding units, wherein the first encoding unit is connected to the second encoding unit through an encoding residual activation module; the encoding residual activation module includes a plurality of cascaded encoding residual activation units; and the last encoding unit is connected to the perceptron.
3. The image reconstruction method according to claim 2, characterized in that: The coding residual activation unit includes two convolutional layers, and the two convolutional layers are connected through an activation function layer.
4. The image reconstruction method according to claim 1, characterized in that: The specific implementation process of extracting the shallow features and the deep features of the feature vector, enhancing the deep features, and obtaining the enhanced features includes: Using the shallow features and the feature vector as input to a first graph construction and aggregation block to obtain a first output; Using the first output and the feature vector as inputs of a second graph construction and aggregation block to obtain a second output; And so on; The output of the last graph construction and aggregation block is used as the input of a convolutional layer, the output of the convolutional layer is concatenated with the shallow features, the concatenated features are used as the input of a pixel reconstruction layer, and the output of the pixel reconstruction layer is used as the input of an output convolutional layer to obtain enhanced features.
5. The image reconstruction method according to claim 4, characterized in that: The graph construction and aggregation block processes the input including the following steps: Performing a linear transformation on the feature vector to obtain modulation coefficients α and β; Multiply the normalized shallow feature / the output of the i-th graph construction and aggregation block with α, add the multiplication result with β, and then add the summed result with the output of the shallow feature / the i-th graph construction and aggregation block to obtain the summed feature; Adding the added feature and the normalized shallow feature / the output of the i-th graph construction and aggregation block to obtain a first output feature; Performing bicubic downsampling and then bicubic upsampling on the first output feature to obtain a second output feature; Calculate the absolute values and DS of the pixel differences corresponding to the first output feature and the second output feature on all feature channels, and construct a local map and a global map according to the absolute values and DS; For each pixel node of the local graph, search for k1 pixels that are most similar to the features of the pixel node in the spatial neighborhood of the pixel node, and connect the k1 pixels to the pixel node; For a global graph, all pixels in the global graph are sampled with a set step size, and for each sampled pixel node, k2 pixels that are most similar to the pixel node feature are calculated from all sampled pixels, and the k2 pixels are connected to the pixel node; The first output feature, the global graph, and the local graph are input into a graph aggregation unit module, and the output of the graph aggregation unit module is concatenated with the shallow feature after passing through a convolution layer to obtain the output of the graph construction and aggregation block.
6. The image reconstruction method according to claim 5, characterized in that: The graph input graph aggregation unit module includes multiple cascaded graph aggregation units, wherein the input of the first graph aggregation unit is the first output feature and the global graph; the input of the second graph aggregation unit is the output of the first graph aggregation unit and the local graph; the input of the third graph aggregation unit is the output of the second graph aggregation unit and the global graph; and so on; wherein the graph aggregation unit includes an edge aggregator and a first normalization layer connected to the edge aggregator; a channel attention module input side is connected to the edge aggregator input side, the channel attention module output, the first normalization layer output, and the graph aggregation unit input are spliced and input into a convolutional feed-forward network, the convolutional feed-forward network is connected to the second normalization layer, and the output of the second normalization layer and the input of the convolutional feed-forward network are spliced to be the output of the graph aggregation unit.
7. The image reconstruction method according to claim 1, characterized in that: The specific implementation process of converting the enhanced features into a high-resolution image includes: Upsampling the enhanced features; The upsampled feature map is convolved and mapped to the HR image space to obtain a high-resolution image.
8. An image reconstruction system, comprising a memory, a processor, and a computer program stored in the memory; characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program / instruction stored thereon; characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program / instructions; characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Image reconstruction system and method based on CRC-SAN network
CN112330542A
Remote sensing image super-resolution reconstruction method based on dual learning graph network
CN113643182A
Image super-resolution reconstruction method based on DRRN network
CN118247143A
Image super-resolution method and system, computer equipment and storage medium
CN119784593A
Frequency domain adaptive super-resolution reconstruction method
CN119809936A