A Blind Super-Resolution Method for Stereo Images Based on Hybrid Degradation and Content Awareness
By designing a stereoscopic image blind super-resolution network based on hybrid degradation and content perception, the complex degradation processing problem in the real world is solved, and high-performance stereoscopic image super-resolution reconstruction is realized, filling the gap in the missing scene data set.
Patent Information
- Application Number
- CN202411564646.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-05
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-11-05
AI Technical Summary
Existing stereo image super-resolution methods are difficult to effectively deal with complex degradation situations in the real world, resulting in performance degradation and artifacts, and lack of stereo image data sets in real scenes.
A stereoscopic image blind super-resolution method based on hybrid degradation and content perception is proposed. By designing a stereoscopic image blind super-resolution network (HDCAnet), a synthetic data set is generated using a stereoscopic image high-order gated degradation model, combining a hybrid degradation-content representation learning mechanism and a cross-viewpoint parallax attention module, it adaptively processes complex degradation and fuses prior information.
The performance of super-resolution reconstruction of stereo images is significantly improved, and the optimal reconstruction effect can be obtained on multiple synthetic and real test data sets is proved, proving the effectiveness and robustness of the method.
Smart Images

Figure CN119515684B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a stereo image blind super-resolution method based on hybrid degradation and content awareness. Background Art
[0002] Images and videos have become important carriers of visual information due to their vividness, intuitiveness, etc. In order to pursue a clearer and more realistic visual experience, digital image acquisition technology and processing technology have developed rapidly. With the wide application of robot navigation, virtual reality, and augmented reality technologies, stereo images with a sense of depth and three-dimensionality have received increasing attention from academia and industry. However, in the actual acquisition process, high-resolution stereo images are usually difficult to obtain due to limitations in shooting equipment and environment. Compared with high-cost hardware devices, implementing stereo image super-resolution (Stereo Image Super-Resolution, Stereo SR) using software algorithms is more flexible and economical. This technology utilizes the spatial information of the image and prior knowledge under the original hardware conditions, and strives to recover the high-frequency detail information lost during the image degradation process, thereby reconstructing a high-resolution stereo image at low cost.
[0003] As a professional technology with great practical value, stereo image super-resolution reconstruction technology has been widely applied in many fields. In the medical field, image super-resolution technology can enhance the resolution of tumor images, nuclear magnetic resonance images, etc. With the help of the reconstructed high-resolution medical stereo images, the lesions can be displayed more vividly, helping doctors more accurately locate the position of diseased cells and improving the accuracy of disease diagnosis; in the field of reconnaissance and remote sensing, applying image super-resolution technology on reconnaissance satellites and unmanned aerial vehicle satellite systems at different positions can recover the detail information of small target objects captured from high altitudes, thereby clearly identifying objects such as vehicles, roads, and buildings; in the field of culture and entertainment, mobile phones can use image super-resolution technology to improve the image quality, saving hardware costs while enhancing the user experience and increasing corporate profits; in the field of public safety and services, stereo image super-resolution technology can process surveillance videos, which helps to further identify specific targets in the surveillance videos and is of great significance to the public safety cause. In addition, super-resolution is also fully applied in aspects such as intelligent security, the animation industry, astronomy, seismology, and biometrics. Stereo image super-resolution reconstruction technology can not only improve the visual effect of stereo images but also serve as a preprocessing for some high-level visual tasks, such as detection and recognition understanding tasks of video images, to further improve the performance of high-level visual tasks. Therefore, studying stereo image super-resolution reconstruction technology has very important theoretical significance and application value.
[0004] The research on image super-resolution reconstruction technology by researchers originated in the 1960s. According to different specific usage scenarios, the super-resolution reconstruction problem can be divided into three categories, namely single-image super-resolution (SISR) applied to single-image magnification, video super-resolution (VSR) applied to continuous-frame magnification, and stereo image super-resolution applied to stereo image magnification.
[0005] Stereo image super-resolution (Stereo SR) utilizes the spatial information (single-viewpoint information) of the left and right views themselves, as well as the complementary information (cross-viewpoint information) between stereo images to restore the details lost in the low-resolution image. One of the key issues in stereo image super-resolution is to effectively utilize cross-view information. Jeon et al. were the first to use deep learning methods to solve the stereo image super-resolution reconstruction problem. By training a two-stage designed CNN, the reconstruction process of the left-viewpoint image was achieved. To alleviate the impact brought by the disparity between the left and right viewpoints, this method takes the left-viewpoint image and 64 right-viewpoint images after horizontal translation as the input of the network. The first-stage CNN is responsible for reconstructing the high-resolution luminance image, and the second-stage CNN is responsible for restoring the color information in the high-resolution image. However, this method can only handle stereo image super-resolution tasks with a horizontal disparity within 64 pixels. To adapt to the disparity changes between different stereo images, Wang et al. proposed a parallax attention module (PAM) to handle stereo images with large horizontal disparity changes. By calculating the similarity of pixel points at the epipolar line position, PAM can capture the corresponding relationship of pixel points in the left and right views, thereby aligning the features of the two viewpoints and realizing the effective utilization of cross-viewpoint information. In addition to using a module similar to PAM to alleviate the disparity problem, Yan et al. proposed to embed a stereo matching network in the stereo image super-resolution network, and the two networks share shallow features. Subsequently, the stereo matching network will calculate the disparity between the left and right viewpoints based on the shallow features, so as to perform spatial alignment on the two images and alleviate the impact brought by the disparity in stereo images. Inspired by the gated recurrent unit mechanism, Lei et al. designed an interaction unit to give full play to the complementary information in stereo images. This unit includes four gates, namely the filter gate, the reset gate, the screen gate, and the update gate. Through these gates, the interaction unit can filter, reset, and update the features of the left and right viewpoints, thereby obtaining enhanced features containing cross-viewpoint information.
[0006] Although the above-mentioned CNN-based stereo image super-resolution reconstruction algorithms have shown remarkable performance, they all only consider a single type of degradation, that is, simply using bicubic downsampling to generate a low-resolution (LR) image pair from a high-resolution (HR) image pair. Due to the various complex degradations in the real world, these stereo super-resolution methods are not suitable for directly processing real-world images, often resulting in performance degradation and artifacts. However, there is currently no dedicated stereo image super-resolution method for real scenarios. In addition, to perform super-resolution reconstruction on stereo images in real scenarios, LR-HR paired data captured from the real world is required. However, since the capture of stereo images is more complex and challenging than that of single images, there is currently no stereo image dataset in real scenarios.
[0007] Based on the above, the present invention proposes a blind stereo image super-resolution method based on hybrid degradation and content awareness to solve the above problems. Summary of the Invention
[0008] 1. Technical problems to be solved by the present invention
[0009] The purpose of the present invention is to propose a blind stereo image super-resolution method based on hybrid degradation and content awareness to solve the problems proposed in the background technology. The present invention provides a research solution for stereo image super-resolution methods in real scenarios, and to a certain extent promotes the development of high-level vision tasks such as image detection and segmentation.
[0010] 2. Technical solutions
[0011] To achieve the above object, the present invention provides the following technical solutions:
[0012] A blind stereo image super-resolution method based on hybrid degradation and content awareness, which performs super-resolution reconstruction on stereo images in real scenarios by proposing a blind stereo image super-resolution network (HDCAnet). The blind stereo image super-resolution network includes:
[0013] A stereo image high-order gated degradation (SHG) model for generating a stereo image synthesis dataset to simulate complex degradations in real-world scenarios;
[0014] A stereo graphics hybrid degradation-content representation (HDCR) learning mechanism, which extracts degradation information in stereo images through unsupervised contrast learning, jointly extracts explicit content information in the left-viewpoint image and the right-viewpoint image, and uses the learned hybrid degradation-content representation as prior information to assist the blind stereo image super-resolution network in reconstruction;
[0015] Hybrid Prior Fusion Block (HPFB) is used to adaptively process various different complex degradations and efficiently fuse the learned prior information into the stereo image blind super-resolution network;
[0016] Cross-Viewpoint Disparity Attention Module (CVPAM) is used to accurately capture cross-viewpoint information;
[0017] Based on the stereo image blind super-resolution network, the method includes the following steps:
[0018] S1. Input the low-resolution left and right viewpoint images into a 3×3 convolutional layer to extract shallow features;
[0019] S2. Together with the prior information, the shallow features extracted in S1 are passed through N cascaded Hybrid Prior Transformation Modules (HPTM) and Cross-Viewpoint Disparity Attention Modules to fuse the prior information and extract deep features;
[0020] S3. In the feature reconstruction stage, the extracted deep features are passed through a 1×1 convolutional layer, a sub-pixel convolutional layer, and a 3×3 convolutional layer for high-resolution image reconstruction, and the high-resolution residual images of the left and right viewpoints obtained are added to the upsampled low-resolution images to obtain the final high-resolution reconstructed image
[0021] Preferably, the training of the stereo image blind super-resolution network is divided into two stages:
[0022] In the first stage, the Hybrid Prior Encoder (HPE) is trained so that the Hybrid Prior Encoder initially extracts the Hybrid Degradation-Content Representation (HDCR);
[0023] In the second stage, the obtained Hybrid Degradation-Content Representation (HDCR) is used as the prior information P and fed into the stereo image blind super-resolution network (HDCAnet) for efficient fusion of the prior information; in the second stage, the stereo image blind super-resolution network (HDCAnet) and the Hybrid Prior Encoder (HPE) are jointly trained end-to-end.
[0024] Preferably, the stereo image high-order gated (SHG) degradation model is achieved by adding a gating mechanism to the high-order degradation model and is implemented by a random gate controller. Under the guidance of the random gate controller, various basic degradations are randomly shuffled and combined, thereby expanding the degradation space to cover different subsets of the basic degradations;
[0025] The stereo image high-order gated degradation model applies the same degradation operation to a pair of left and right viewpoint images and performs different degradation operations on different stereo image pairs; the formula representation of the stereo image high-order gated degradation model is as follows:
[0026] I LR = (G(D 1 ))·G(D 2 ))·G(D 3 ))……G(D n ))(I HR ))
[0027] Wherein, I LR represents the low-resolution image; I HR represents the high-resolution image; represents the basic degradation type; G represents the gate controller.
[0028] Preferably, the 3D graphic Hybrid Degradation-Content Representation (HDCR) learning mechanism specifically includes the following:
[0029] For the i-th pair of images, first randomly cut the input 3D images and to obtain the left and right view image patches and in the same pair of 3D images, as well as the image patches
[0030] from other 3D images; and Then, the obtained
[0031] The loss function L cl adopted by the 3D graphic Hybrid Degradation-Content Representation (HDCR) learning mechanism is:
[0032]
[0033] Wherein, E(·) represents the Hybrid Prior Encoder (HPE), N que represents the number of samples in the queue, τ represents the temperature hyperparameter, and B represents the batch size.
[0034] Preferably, the Hybrid Prior Encoder (HPE) consists of 6 convolutional layers, 1 average pooling layer and 1 multi-layer perceptron; the Hybrid Prior Encoder (HPE) uses the result after the intermediate convolutional layer as the prior information to retain the spatial context information;
[0035] The Hybrid Prior Transformation Module (HPTM) consists of a Hybrid Prior Fusion Block (HPFB) and a convolutional layer, and each Hybrid Prior Fusion Block (HPFB) consists of a Spatial Feature Transformation Layer (SFT), a Deformable Convolutional Layer (DCN) and a residual connection.
[0036] Preferably, the processing flow of the Hybrid Prior Fusion Block (HPFB) specifically includes the following:
[0037] A deformable convolutional layer (DCN) is adopted to dynamically adjust the sampling positions and receptive fields of the convolutional kernels by learning modulation and masks. The functional representation of the processing flow of the deformable convolutional layer (DCN) is:
[0038]
[0039] where Φ DCN represents the DCN layer, F in represents the input feature map of the left or right view point, K represents the number of sampling positions, ω n represents the weight, p 0 represents the central position, p n represents the predefined offset at the nth position; P represents the prior information, Δp n and Δm n represent the learnable offset and modulation scalar, which are learned through a convolutional layer after being concatenated by the prior information P and F in , and the specific calculation formula is as follows:
[0040] (Δp n , Δm n ) = conv(concat(F in , P))
[0041] where concat(·) represents the concatenation operator and conv(·) represents the convolutional layer;
[0042] A Spatial Feature Transformation layer (SFT) is adopted to adaptively fuse the image features with the prior information P. The Spatial Feature Transformation layer (SFT) performs an affine transformation through scaling and shifting operations to assist in adjusting the features of the reconstruction network; the functional representation of the Spatial Feature Transformation layer (SFT) is:
[0043] Φ SFT (F in |P) = γeF in +β = (conv(P))eF in +(conv(P))
[0044] where Φ SFT represents the SFT layer and e represents the element-wise multiplication;
[0045] By using the SFT layer, the Stereo Image Blind Super-Resolution Network (HDCAnet) adjusts the input feature F according to the prior information P inPerform scaling and shifting operations to adaptively modulate the left and right view feature maps, and achieve the fusion guidance of the prior information P for the reconstructed features. The functional representation of the above process is as follows:
[0046] F out = Φ DCN (F in | P) + Φ SFT (F in | P) + F in
[0047] Among them, F out represents the output feature map of the left or right view of the HPFB module.
[0048] Preferably, the processing flow of the cross-view disparity attention module (CVPAM) specifically includes the following contents:
[0049] Normalize the input features F L and F R through layer normalization (LN), and then send them to the residual block and the linear layer to obtain and The functional representation of which is as follows:
[0050]
[0051] Among them, Res represents a shared transition residual block used to alleviate training conflicts, and represent projection matrices implemented by 1×1 convolutional layers;
[0052] Pass and through the whitening layer to generate robust stereo correspondences and obtain the normalized features and The functional representation of the above process is as follows:
[0053]
[0054] Pass and the transposed through batch multiplication to generate a score map S ∈ R H×W×W ;
[0055] Apply the softmax function to the score map S, transpose the score map S along its last dimension, and finally multiply the output by the features obtained by passing F L and F R The functional representation of the above process is as follows:
[0056]
[0057] Among them, softmax represents the softmax function, T represents the transpose operation, and W L and W R represent the projection matrices obtained from the 1×1 convolutional layer;
[0058] Finally, global average pooling is used to compress the spatial information of F L and F R to the channel dimension, and then a linear layer is used to obtain the weights for modulating the cross-viewpoint features F L→R and F R→L Then, the obtained weights are multiplied element-wise with F L→R and F R→L to achieve the adaptive modulation process, and the modulation result is added element-wise with F L and F R to perform the fusion of cross-viewpoint features and within-viewpoint features, obtaining the final output process and The functional representation of the above process is:
[0059]
[0060] Among them, pool represents the global average pooling operation.
[0061] 3. Beneficial Effects
[0062] The present invention proposes a blind super-resolution method for stereoscopic images based on hybrid degradation and content awareness. More specifically, the present invention designs a blind super-resolution network for stereoscopic images based on hybrid degradation and content awareness (HDCAnet). First, aiming at the problem that there is currently no real-scene stereoscopic image dataset, the present invention proposes a stereoscopic image high-order gated degradation model (SHG degradation model) to generate a stereoscopic image synthesis dataset, thereby simulating the complex degradation in the real-world scene. Second, the present invention proposes a hybrid degradation-content representation (HDCR) learning mechanism for stereoscopic images, which extracts the degradation information in the stereoscopic images and the explicit content information in the left-viewpoint image and the right-viewpoint image through unsupervised contrast learning, and uses them as priors to assist the network in better reconstruction. Then, the present invention designs a hybrid prior fusion module (HPFB) to adaptively process various different complex degradations and efficiently fuse the learned prior information into the network. Finally, the present invention proposes a cross-viewpoint disparity attention module (CVPAM) to more accurately capture cross-viewpoint information.
[0063] The experimental results of the present invention show that the proposed method significantly improves the performance of the reconstruction network and obtains the optimal reconstruction effect on multiple synthetic and real test datasets, fully demonstrating the effectiveness of the proposed method. Description of the Drawings
[0064] Figure 1 Schematic diagram of the SHG degradation model architecture mentioned in Embodiment 1 of the present invention;
[0065] Figure 2 Network architecture diagram of the stereo image blind super-resolution method based on hybrid degradation and content awareness mentioned in Embodiment 1 of the present invention;
[0066] Figure 3 Schematic diagram of the HDCR learning mechanism mentioned in Embodiment 1 of the present invention;
[0067] Figure 4 Specific structure diagram of HPE mentioned in Embodiment 1 of the present invention;
[0068] Figure 5 Specific structure diagrams of HPTM and HPFB mentioned in Embodiment 1 of the present invention, where (a) represents a schematic diagram of the hybrid prior transformation module (HPTM), and (b) represents a schematic diagram of the hybrid prior fusion block (HPFB);
[0069] Figure 6 Specific structure diagram of CVPAM mentioned in Embodiment 1 of the present invention;
[0070] Figure 7 Visualization result diagram of the pre-trained Stereo SR method on a synthetic dataset with 4-fold super-resolution mentioned in Embodiment 2 of the present invention;
[0071] Figure 8 Visualization result of the Stereo SR method through re-training on a synthetic dataset with 4-fold super-resolution mentioned in Embodiment 2 of the present invention. Detailed implementation manners
[0072] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention in conjunction with the accompanying drawings of the specification.
[0073] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0074] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an individual or alternative embodiment that mutually excludes other embodiments.
[0075] First, the abbreviations and key terms mentioned in the present invention are defined as follows:
[0076] Stereo SR: Stereo Image Super-Resolution, which refers to super-resolution of stereo images
[0077] SISR: Single Image Super Resolution, which refers to super-resolution of a single image
[0078] VSR: Video Super-Resolution, which refers to super-resolution of videos
[0079] PAM: Parallax Attention Module
[0080] HR: High-Resolution
[0081] LR: Low-Resolution
[0082] HDCAnet: hybrid degradation-content aware stereo super-resolution network
[0083] SHG degradation model: Stereo-image High-order Gated Degradation Model
[0084] HDCR: Hybrid Degradation-Content Representation
[0085] HPFB: Hybrid Prior Fusion Block
[0086] CVPAM: Cross-View Parallax Attention Module
[0087] HPE: hybrid prior encoder
[0088] HPTM: Hybrid Prior Transform Module
[0089] SFT: spatial feature transform, spatial feature transformation
[0090] DCN: deformable convolution, deformable convolution
[0091] The following will describe a blind super - resolution method for stereo images based on hybrid degradation and content awareness proposed by the present invention with specific examples and attached drawings. The specific content is as follows.
[0092] Example 1:
[0093] The present invention proposes a blind super - resolution method for stereo images based on hybrid degradation and content awareness. First, to address the problem that there is currently no real - world stereo image dataset, the present invention proposes an SHG degradation model to generate a stereo image synthesis dataset, thereby simulating the complex degradation in the real - world scenario. At the same time, the present invention introduces a blind super - resolution network for stereo images based on hybrid degradation and content awareness (HDCAnet), which can learn degradation information and jointly learn the explicit content information of the left and right views.
[0094] To solve the problem that the simple combination of basic degradation types in classical degradation cannot adapt to the complex degradation in real scenarios and the problem that different subsets of basic degradation may be ignored in high - order gated degradation, the present invention proposes an SHG degradation model. Its specific structure is as Figure 1 shown. Specifically, the SHG degradation model adds a gating mechanism on the basis of the high - order degradation model, which is realized by a random gate controller. Under the guidance of the random gate controller, various basic degradations are randomly shuffled and combined, thereby further expanding the degradation space to cover different subsets of basic degradation. In addition, the same degradation operation is applied to a pair of left - view and right - view images, and different degradation operations are performed on different stereo image pairs. The model is expressed by the following formula:
[0095] I LR =(G(D 1 )·G(D 2 )·G(D 3 )…G(D n ))(I HR )
[0096] where, I LR and I HR represent the low - resolution image and the high - resolution image respectively, represents the basic degradation type, and G represents the gate controller. By adjusting the probability of the random gate controller, the SHG degradation model greatly expands the degradation space, and without additional data acquisition, it greatly simulates the complex degradation in the real scenario and enhances the robustness of HDCAnet.
[0097] The overall framework of HDCAnet is as shown in Figure 2 the figure. Its training is mainly divided into two stages: In the first stage, the Hybrid Prior Encoder (HPE) is trained first, as shown in Figure 3 the figure, so that the HPE can initially extract HDCR; in the second stage, the obtained HDCR is used as the prior information P (represented by P in the present invention for HDCR), and it is sent into HDCAnet for efficient fusion of prior information, as shown in Figure 2 the figure. In the second stage, end-to-end joint training of HDCAnet and HPE is carried out.
[0098] As shown in Figure 2 the figure, first, HDCAnet first passes the input low-resolution left and right view images into a 3×3 convolutional layer to extract shallow features. Subsequently, the extracted shallow features and P are passed through N cascaded stacked Hybrid Prior Transform Modules (HPTM) and CVPAM modules to fuse prior information and extract deep features. The specific structure of HPTM is as shown in Figure 5 (a). It mainly consists of two HPFBs, two 3×3 convolutional layers, and residual connections. The specific structure of HPFB is as shown in Figure 5 (b). It is used to efficiently fuse its own view information with prior information. The specific structure of CVPAM is as shown in Figure 6 the figure. It is used for the extraction of cross-view features. Finally, in the feature reconstruction stage, the features are passed through a 1×1 convolutional layer, a sub-pixel convolutional layer, and a 3×3 convolutional layer for the reconstruction of high-resolution images, and the obtained high-resolution residual images of the left and right views are added to the upsampled low-resolution images to obtain the final high-resolution reconstructed images and This network simultaneously performs super-resolution reconstruction of left and right view stereo images and shares the weights of the left and right branches.
[0099] The degradation of images in real scenes is diverse and complex, resulting in a serious impact on the reconstruction performance of existing networks. To solve this problem, the present invention proposes an HDCR learning mechanism, as shown in Figure 3 the figure. Specifically, for the i-th pair of images, first, the input stereo images and are randomly cut into pieces to obtain the left and right view image patches and in the same pair of stereo images, as well as the image patches from other stereo images. Then, the obtained as well as After being encoded by HPE, query samples, positive samples, and negative samples are obtained respectively. Taking the left-viewpoint image patch in the red frame as an example, it is regarded as query Q and also used as HDCR. The corresponding right-viewpoint image patch can be regarded as a positive sample, while the patches from other images are regarded as negative samples. The loss function adopted in this contrastive learning is as follows:
[0100]
[0101] where E(·) represents HPE, N que represents the number of samples in the queue, τ is the temperature hyperparameter, and B is the batch size.
[0102] The basic structure of HPE proposed by the present invention is as Figure 4 shown. Specifically, the basic structure of HPE consists of 6 convolutional layers, 1 average pooling layer, and a multi-layer perceptron. In order to be able to more fully retain the spatial context information, HPE selects the result after the intermediate convolutional layer as P, so that the obtained P and other features Figure 1 also retain the four-dimensional tensor of the spatial structure, rather than the two-dimensional vector lacking the spatial structure, so as to be able to further realize the efficient utilization of the image content information.
[0103] The HPTM structure consists of HPFB and convolutional layers, and the specific structure is as Figure 5 (a) shown. Each HPFB block consists of a Spatial Feature Transform (SFT) layer, a Deformable Convolution (DCN) layer, and a residual connection, and the specific structure is as Figure 5 (b) shown.
[0104] Specifically, first, DCN is adopted in HPFB, and the sampling position and receptive field of the convolutional kernel are dynamically adjusted by learning the modulation offset and mask. The whole process of DCN can be expressed by the formula as follows:
[0105]
[0106] where Φ DCN represents the DCN layer, F in is the input feature map of the left or right viewpoint, K is the number of sampling positions, ω n is the weight, p 0 is the center position, p n is the predefined offset of the nth position. Δp n and Δm nare learnable offset and modulation scalars, both determined by priors P and F in They are obtained through a concatenation operation followed by a convolutional layer, and the calculation formula is as follows:
[0107] (Δp n , Δm n ) = conv(concat(F in , P))
[0108] Among them, concat(·) is the concatenation operator, and conv(·) represents the convolutional layer. Using DCN, HDCAnet can dynamically adjust the receptive field according to the image content. In addition, the SFT layer is adopted in HPFB to adaptively fuse the image features with P. The SFT layer performs an affine transformation by using scaling and shifting operations to assist in adjusting the features in the reconstruction network. The SFT layer can be expressed by the formula as follows:
[0109] Φ SFT (F in |P) = γeF in + β = (conv(P))eF in + (conv(P))
[0110] Among them, Φ SFT represents the SFT layer, and e represents element-wise multiplication. By using the SFT layer, HDCAnet can perform scaling and shifting operations on the input feature F in , adaptively modulate the left and right view feature maps, and realize the fusion guidance of the prior P on the reconstruction features. The whole process can be expressed by the formula as follows:
[0111] F out = Φ DCN (F in |P) + Φ SFT (F in |P) + F in
[0112] Among them, F out is the output feature map of the left or right view of the HPFB module. Generally speaking, through the HPFB block, HDCAnet can adaptively process different degradations and effectively fuse prior information.
[0113] While fusing prior information, in order to better obtain cross-view features from the corresponding views, the present invention proposes CVPAM, and the specific structure is as Figure 6 shown.
[0114] First, the input features F L and F RNormalized by layer normalization (LN) and then sent to residual blocks and linear layers (implemented by 1×1 convolutions) to obtain and The formula can be expressed as follows:
[0115]
[0116] where Res represents the shared transition residual block used to alleviate training conflicts, and are projection matrices implemented by 1×1 convolutional layers. After that, and are passed through a whitening layer to generate robust stereo correspondences and obtain normalized features and The process can be expressed by the formula as follows:
[0117]
[0118] and after transposing perform batch multiplication to generate the score map S ∈ R H×W×W . Then apply the softmax function to S, transpose S along its last dimension, and finally multiply the output with the features passed through the projection matrix F L and F R The process can be expressed by the formula as follows:
[0119]
[0120] where softmax is the softmax function, T represents the transpose operation, W L and W R are projection matrices obtained from 1×1 convolutional layers. Finally, use global average pooling to compress the spatial information of F L and F R to the channel dimension, and then obtain the weights for modulating the cross-viewpoint features F L→R and F R→L through a linear layer, and then use these weights to perform element-wise multiplication with F L→R and F R→L to implement the adaptive modulation process, and add the modulation results element-wise with F L and F R to fuse the cross-viewpoint features and the in-viewpoint features to obtain the final output process and The process can be expressed by the formula as follows:
[0121]
[0122] Among them, "pool" represents the global average pooling operation. After each HPTM, a CVPAM is inserted in HDCAnet to interact the left and right view features, so as to fully extract and accurately fuse complementary information.
[0123] Example 2:
[0124] Based on Example 1 but with differences. Next, a blind super-resolution method for stereoscopic images based on hybrid degradation and content awareness proposed by the present invention will be described in combination with specific experiments. The specific content is as follows.
[0125] (1) Datasets
[0126] Training dataset: Given the relative scarcity of datasets in the field of blind super-resolution reconstruction of stereoscopic images, the present invention preprocesses the commonly used Flickr1024 dataset and generates a new training dataset suitable for blind super-resolution reconstruction of stereoscopic images through the proposed SHG degradation model.
[0127] Test datasets: First, in order to conduct appropriate tests in the field of blind super-resolution reconstruction of stereoscopic images, new test datasets are generated for the four publicly available benchmark datasets KITTI2012, KITTI2015, Middlebury, and Flickr1024 using the SHG degradation model. The specific division is as follows: The KITTI 2012 dataset contains 20 pairs of stereoscopic images; the KITTI 2015 dataset contains 20 pairs of stereoscopic images; the Middlebury dataset contains 5 pairs of stereoscopic images, mainly for testing indoor scenes; the Flickr1024 dataset contains 112 pairs of stereoscopic images, covering images of multiple categories, such as buildings, street scenes, indoors, and people, etc. Second, the dataset of Track 3 in the NTIRE2023 stereoscopic image super-resolution reconstruction competition (Track3-Flickr1024) is additionally adopted. This dataset contains 112 pairs of stereoscopic images (×4), from the validation dataset of Flickr1024, and is an artificial synthetic dataset generated using the classical degradation model.
[0128] (2) Experimental settings
[0129] During the training phase, for simplicity, the probability of each gate G is set to 0.5. Then, the obtained left and right images are cropped into small blocks of size 30×90. The training data is randomly flipped horizontally and vertically for data augmentation. The training is completed in two stages. First, in the first stage, the contrast loss in formula (3-4) is used to train the encoder network for 30K iterations, with the initial learning rate set to 2×10 -3 , and it drops to 2×10 after 18K iterations -4, N in formula (3-4) que and τ are set to 0.07 and 8192 respectively. Then, in the second stage, the entire network is jointly trained for 100K iterations. During the training process, the AdamW optimizer is used, and the parameters β 1 = 0.9 and β 2 = 0.9 are set to the default values, the batch size is set to 32, and the initial learning rate is set to 5×10 -4 , the cosine annealing strategy is used, and its minimum learning rate is set to 1×10 -7 . The model proposed in the present invention is implemented based on the Python language, and the experimental environment is shown in Table 1:
[0130] Table 1: Experimental environment
[0131]
[0132] (3) Experimental results
[0133] The method proposed in the present invention and the previous scene text detection methods are compared on five datasets, and the performance of the present invention is evaluated from two indicators: peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM).
[0134] Table 2 and Table 3 compare the proposed HDCAnet in the present invention with several latest super-resolution methods, including two single-image super-resolution methods (EDSR and RCAN), four blind super-resolution methods (DANv2, DCLS, DASR, and CDSR), and five stereo-image super-resolution methods (PASSRnet, iPASSR, SSRDE-FNet, NAFSSR, and swinFIRSSR). The optimal evaluation indicators are marked in bold in the table, and the sub-optimal indicators are marked with an underline. It can be seen that when facing complex degraded stereo images, the performance of the previous single-image super-resolution methods and stereo-image super-resolution methods drops significantly, and the blind super-resolution methods are also not satisfactory when dealing with stereo images due to the lack of cross-view information. The method of the present invention is significantly superior to the previous methods in terms of PSNR and SSIM indicators.
[0135] Table 2: Experimental results on the KITTI2012 dataset and the KITTI2015 dataset
[0136]
[0137]
[0138] Table 3: Experimental Results on Middlebuy Dataset, Flickr1024 Dataset, and Track3-Flickr1024 Dataset
[0139]
[0140]
[0141] In addition, for fair comparison, we also selected some existing stereo image super-resolution methods, namely PASSRnet, iPASSR, NAFSSR, swinFIRSSR, and retrained them from scratch using the same training dataset as the present invention (referred to as restricted conditions in the table). As can be seen from Table 4 and Table 5, the optimal evaluation metrics are marked in bold and the sub-optimal metrics are underlined in the table. It can be seen that the performance of the retrained stereo image super-resolution methods has been greatly improved compared to before (about 0.2dB - 0.6dB on the Track 3-Flickr 1024 dataset), which proves the key role of the SHG degradation model proposed by the present invention. In addition, compared with these retrained stereo image super-resolution networks, the present invention reaches the optimal in most results, demonstrating the superiority of the network design of the present invention.
[0142] Table 4: Experimental Results on KITTI2012 Dataset and KITTI2015 Dataset under Restricted Conditions
[0143]
[0144] Table 5: Experimental Results on Middlebuy Dataset, Flickr1024 Dataset, and Track3-Flickr1024 Dataset under Restricted Conditions
[0145]
[0146] To more intuitively prove the effectiveness of the method proposed by the present invention, a visual comparison was made between the HDCAnet proposed by the present invention and several super-resolution methods. Figure 7 The visual results of HDCAnet and pre-trained super-resolution methods in the 4x super-resolution experiment are shown. It can be observed that all other methods cannot remove noise or reconstruct clear lines, and some super-resolution results of these pre-trained methods are even close to the original LR. However, HDCAnet can not only effectively eliminate noise and blur, but also restore clearer patterns and textures.
[0147] Figure 8Shows the visualization results of HDCAnet and the retrained super-resolution method in the 4x super-resolution experiment. It can be seen that when dealing with severe noise and blur, the zebra crossings (in KITTI_2015_0001L) and stripes (in Flickr1024_0106_L) restored by other methods are more blurred and distorted, with artifacts. In contrast, the images reconstructed by HDCAnet are clearer, and the restored zebra crossings and stripes are straighter, which demonstrates the superiority and robustness of the method of the present invention.
[0148] Table 6 shows the results of ablation experiments for each module in HDCAnet. Variant 1 is constructed by removing L in the loss function during the second training phase without changing the network. cl In addition, the pre-trained HPE was removed in the first training phase, and the entire network was jointly trained 100,000 times to obtain Variant 2. Finally, HPE was converted into the degradation encoder used in DASR, and HPFB was converted into the DA module used in DASR, thus constructing Variant 3. The experimental results are shown in Table 6, with the best performance in bold. The results show that the HDCR learning mechanism and HPE are both essential for the network, because HDCR learning enhances the ability of HPE to obtain discriminative features, and pre-training HPE in the first training phase can obtain more accurate prior information in advance, which is beneficial to subsequent joint training. In addition, compared with DASR that extracts global degradation information, HDCAnet has obvious advantages in using content information by retaining spatial context information.
[0149] Table 6: Ablation experiment results of 4x super-resolution under different settings
[0150]
[0151]
[0152] Table 7 shows the results of ablation experiments for the extracted explicit content information. The prior P extracted by the method of the present invention contains not only degradation information but also explicit content information. Variant 4 adopts different training methods. Specifically, two non-overlapping patches are randomly selected in the same LR image to obtain positive samples with the same degradation, while different degradations from other LR images are used as negative samples, so that the extracted P contains degradation information and implicit content information. It can be seen from Table 7 that when the explicit content information is removed, the indicators on the four datasets decrease significantly. The experimental results fully prove the rationality and effectiveness of using explicit content information, as well as its importance for the HDCAnet stereo image blind super-resolution reconstruction method.
[0153] Table 7: Ablation experiment results of explicit content information
[0154]
[0155] To prove the effectiveness of the HPFB proposed in the present invention, ablation experiments were respectively carried out by removing the DCN layer, the SFT layer, and the entire HPFB. The results are shown in Table 8 and Table 9, and the best performance is in bold. Taking Real-KITTI 2012 as an example, experiments were carried out with HDCAnet, and the best result of 21.89 / 0.6003 was obtained. When the HPFB was removed, there was a significant performance drop in the PSNR and SSIM values, from 21.89 / 0.6003 to 21.64 / 0.5953. When the DCN layer and the SFT layer were removed respectively, the PSNR of the network decreased by 0.09 dB and 0.07 dB, which also proves that both the DCN layer and the SFT layer are important for improving the performance of HDCAnet.
[0156] Table 8: Results of HPFB ablation experiments for 4x super-resolution on the KITTI2012 dataset and the Middlebuy dataset
[0157]
[0158] Table 9: Results of HPFB ablation experiments for 4x super-resolution on the Flickr1024 dataset and the Track3-Flickr102 dataset
[0159]
[0160] As mentioned above, it is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its improved concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A blind super-resolution method for stereo images based on hybrid degradation and content awareness, characterized in that: A stereo image blind super-resolution network is proposed to reconstruct a stereo image in a real scene with super-resolution. The stereo image blind super-resolution network includes: A high-order gated degradation model for stereo images, which is used to generate synthetic datasets of stereo images and simulate the complex degradation in real-world scenes; The mixed degradation-content representation learning mechanism for stereo graphics extracts degradation information from stereo images through unsupervised contrast learning, jointly extracts explicit content information from left-view and right-view images, and uses the learned mixed degradation-content representation as prior information to assist the stereo image blind super-resolution network in reconstruction. The hybrid prior fusion block is used to adaptively handle various complex degradations and efficiently fuse the learned prior information into the stereo image blind super-resolution network; Cross-viewpoint disparity attention module to accurately capture cross-viewpoint information; Based on the above stereo image blind super-resolution network, the method comprises the following steps: S1, input low-resolution left and right viewpoint images Pass it into the 3×3 convolutional layer to extract shallow features; S2, the shallow features extracted in S1 are combined with the prior information through N cascaded hybrid prior transformation modules and cross-viewpoint parallax attention modules to fuse the prior information and extract deep features; S3. In the feature reconstruction stage, the extracted deep features are reconstructed into a high-resolution image through a 1×1 convolution layer, a sub-pixel convolution layer, and a 3×3 convolution layer. The high-resolution residual images of the left and right viewpoints are added to the upsampled low-resolution image to obtain the final high-resolution reconstructed image.
2. The method for blind super-resolution of stereo images based on hybrid degradation and content perception according to claim 1, characterized in that: The training of the stereo image blind super-resolution network is divided into two stages: In the first stage, the hybrid prior encoder is trained so that the hybrid prior encoder initially extracts the hybrid degradation-content representation; In the second stage, the obtained mixed degradation-content representation is used as the prior information P and sent to the stereo image blind super-resolution network for efficient fusion of the prior information; In the second stage, the stereo image blind super-resolution network and the hybrid prior encoder are jointly trained end-to-end.
3. The method for blind super-resolution of stereo images based on hybrid degradation and content perception according to claim 1, characterized in that: The stereoscopic image high-order gated degradation model is implemented by adding a gating mechanism on the basis of the high-order degradation model, and is implemented by a random gate controller. Under the guidance of the random gate controller, multiple basic degradations are randomly shuffled and combined, thereby expanding the degradation space to cover different subsets of basic degradations. The stereo image high-order gated degradation model applies the same degradation operation to a pair of left viewpoint images and right viewpoint images, and performs different degradation operations on different stereo image pairs; the formula of the stereo image high-order gated degradation model is as follows: I LR =(G(D1)·G(D2)·G(D3)……G(D n ))(I HR ) Among them, I LR Represents a low-resolution image; I HR Represents a high-resolution image; Indicates the basic degradation type; G represents the gate controller.
4. The method for blind super-resolution of stereo images based on hybrid degradation and content perception according to claim 1, characterized in that: The three-dimensional graphics mixed degradation-content representation learning mechanism specifically includes the following contents: For the i-th pair of images, first the input stereo image and Perform random cutting to obtain the left and right viewpoint image blocks in the same pair of stereo images and and patches from other stereo images Then the obtained as well as After encoding by the hybrid prior encoder, query samples, positive samples and negative samples are obtained respectively; The loss function L used by the stereographic hybrid degradation-content representation learning mechanism is cl for: Where E(·) represents the hybrid prior encoder, N que represents the number of samples in the queue, τ represents the temperature hyperparameter, and B represents the batch size.
5. The method for blind super-resolution of stereo images based on hybrid degradation and content perception according to claim 4, characterized in that: The hybrid a priori encoder consists of 6 convolutional layers, 1 average pooling layer and 1 multi-layer perceptron; the hybrid a priori encoder uses the result after the intermediate convolutional layer as a priori information to retain spatial context information; The hybrid prior transformation module consists of a hybrid prior fusion block and a convolution layer, and each of the hybrid prior fusion blocks consists of a spatial feature transformation layer, a deformable convolution layer and a residual connection.
6. The method for blind super-resolution of stereo images based on hybrid degradation and content perception according to claim 1 or 5, characterized in that: The processing flow of the hybrid prior fusion block specifically includes the following contents: A deformable convolution layer is used to dynamically adjust the sampling position and receptive field of the convolution kernel by learning modulation and mask. The function of the deformable convolution layer processing flow is expressed as: Among them, Φ DCN Indicates the DCN layer, F in represents the input feature map of the left or right viewpoint, K represents the number of sampling positions, ω n represents the weight, p0 represents the center position, p n represents the predefined offset of the nth position; P represents the prior information, Δp n and Δm n represents the learnable offset and modulation scalar, which are determined by the prior information P and F in It is obtained through a convolutional layer after the splicing operation. The specific calculation formula is as follows: (Δp n ,Δm n )=conv(concat(F in ,P)) Among them, concat(·) represents the cascade operator, conv(·) represents the convolution layer; A spatial feature transformation layer is used to adaptively fuse the image features with the prior information P. The spatial feature transformation layer performs affine transformation through scaling and shifting operations to assist in adjusting the features of the reconstruction network. The function of the spatial feature transformation layer is expressed as: Φ SFT (F in ∣P)=γe F in +β=(conv(P))e F in +(conv(P)) Among them, Φ SFT represents the SFT layer, and e represents element-wise multiplication; By using the SFT layer, the stereo image blind super-resolution network uses the prior information P to calculate the input feature F in Scaling and shifting operations are performed to adaptively modulate the left and right viewpoint feature maps to achieve the fusion guidance of the prior information P on the reconstruction features. The function of the above process is expressed as follows: F out =Φ DCN (F in ∣P)+Φ SFT (F in ∣P)+F in Among them, F out Represents the output feature map of the HPFB module for the left or right viewpoint.
7. The method for blind super-resolution of stereo images based on hybrid degradation and content perception according to claim 1, characterized in that: The processing flow of the cross-viewpoint disparity attention module specifically includes the following contents: The input feature F L and F R The normalization operation is performed through layer normalization and then sent to the residual block and linear layer to obtain and Its function is expressed as follows: Among them, Res represents the shared transition residual block used to alleviate training conflicts. and represents the projection matrix implemented by the 1×1 convolutional layer; Will and Pass the whitening layer to generate robust stereo correspondence and obtain normalized features and The function of the above process is as follows: Will and after transposition Perform batch multiplication to generate a score map S∈R H×W×W ; The softmax function is applied to the score map S, and the score map S is transposed along its last dimension, and finally the output is combined with the projection matrix F. L and F R The obtained features are multiplied, and the function of the above process is expressed as follows: Among them, softmax represents the softmax function, T represents the transposition operation, and W L and W R Represents the projection matrix obtained by the 1×1 convolutional layer; Finally, global average pooling is used to transform F L and F R The spatial information of is compressed to the channel dimension, and then a linear layer is used to modulate the cross-view feature F L→R and F R→L The weights are then used to calculate the L→R and F R→L Perform element-by-element multiplication to implement the adaptive modulation process, and compare the modulation result with F L and F R Add element by element to fuse the cross-viewpoint features with the intra-viewpoint features to obtain the final output process and The function of the above process is expressed as: Among them, pool represents the global average pooling operation.
Citation Information
Patent Citations
Three-dimensional image super-resolution reconstruction method based on iterative interaction guidance
CN117952830A
Video blind super-resolution reconstruction method and system based on self-supervised learning
WO2022155990A1