Iterative Interactive Reference Stereo Image Super-Resolution Reconstruction Method and System
Through the iterative interactive reference stereo image super-resolution reconstruction method, using high-resolution views as a reference, combining pixel and tile matching, the problems of loss of details and insufficient matching accuracy in the existing methods are solved, and efficient stereo image super-resolution reconstruction is achieved.
Patent Information
- Application Number
- CN202410442888.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-12
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-04-12
AI Technical Summary
Existing stereoscopic image super-resolution methods focus on the correspondence between views in low-resolution spaces, ignoring the reference guidance between high-quality super-segment images, resulting in loss of detail and matching accuracy affected by noise, and additional depth estimation increases network complexity.
The iterative interactive reference stereoscopic image super-resolution reconstruction method is adopted to extract features in the view through the information perception module, combine pixel matching and tile matching to simulate the dependencies between views, use high-resolution views as reference for information interaction, and weighted matching is achieved through the supervisory side output modulator.
It realizes efficient information interaction across views and resolutions, improves the accuracy and consistency of super-resolution of stereo images, and improves the detail recovery ability.
Smart Images

Figure CN118537227B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of super-resolution image reconstruction, and particularly to an iterative interactive reference-based stereo image super-resolution reconstruction method and system. Background Art
[0002] For depth estimation tasks, binocular stereo images are still important information carriers in the actual application process. How to obtain high-resolution binocular stereo images has become an important issue concerned by the academic and industrial communities, and stereo image super-resolution (SSR) has also attracted more and more attention. Most existing SSR methods mainly focus on the view-to-view correspondence relationship in the low-resolution (LR) space, while ignoring the reference guiding role between high-quality super-resolved images.
[0003] Generally speaking, existing SSR work can be divided into two main lines: 1) Exploring pixel-level correspondence matching in the LR space through a disparity attention mechanism, the disadvantage of which is that little attention is paid to high-resolution (HR) correspondence. This may lead to the loss of a large amount of details in the LR image, thus limiting the utilization of cross-view information. In addition, the degradation of the LR image itself (such as noise and blur) will also seriously affect the matching accuracy. 2) Introducing an additional disparity map as a prior, which mainly relies on additional depth estimation, resulting in a complex and bloated network. At the same time, inaccurate prior depth estimation will have an adverse effect on the correspondence matching. Summary of the Invention
[0004] To solve the deficiencies of the prior art, the present invention provides an iterative interactive reference-based stereo image super-resolution reconstruction method and system;
[0005] On the one hand, an iterative interactive reference-based stereo image super-resolution reconstruction method is provided, including:
[0006] Obtaining two images to be reconstructed; the two images to be reconstructed include a low-resolution left-view stereo image and a low-resolution right-view stereo image;
[0007] Inputting the two images to be reconstructed into a trained image reconstruction model to obtain a high-resolution left-view stereo image and a high-resolution right-view stereo image;
[0008] Among them, the trained image reconstruction model uses a number of information perception modules to extract in-view features from the two images to be reconstructed; the model also simulates the inter-view dependence relationship based on a pixel matching module and a patch matching module; the pixel matching module forms internal cross and internal iteration; the patch matching module generates a matching dictionary; projects the matching dictionary into the high-resolution space, and performs high-resolution view information interaction by using the feature maps of the two views as mutual references; the model also re-weights the feature maps of the views through a supervised side output modulator to achieve patch-level matching.
[0009] On the other hand, an iterative interactive reference-based stereoscopic image super-resolution reconstruction system is provided, including:
[0010] An acquisition module, configured to: acquire two images to be reconstructed; the two images to be reconstructed include: a low-resolution left-view stereoscopic image and a low-resolution right-view stereoscopic image;
[0011] A reconstruction module, configured to: input the two images to be reconstructed into the trained image reconstruction model to obtain a high-resolution left-view stereoscopic image and a high-resolution right-view stereoscopic image;
[0012] Among them, the trained image reconstruction model uses a number of information perception modules to extract in-view features from the two images to be reconstructed; the model also simulates the inter-view dependence relationship based on a pixel matching module and a patch matching module; the pixel matching module forms internal cross and internal iteration; the patch matching module generates a matching dictionary; projects the matching dictionary into the high-resolution space, and performs high-resolution view information interaction by using the feature maps of the two views as mutual references; the model also re-weights the feature maps of the views through a supervised side output modulator to achieve patch-level matching.
[0013] On yet another aspect, an electronic device is further provided, including:
[0014] A memory for non-temporarily storing computer-readable instructions; and
[0015] A processor for running the computer-readable instructions,
[0016] Among them, when the computer-readable instructions are run by the processor, the method described in the first aspect above is executed.
[0017] On yet another aspect, a storage medium is further provided, which non-temporarily stores computer-readable instructions, wherein when the non-temporary computer-readable instructions are executed by a computer, the instructions for executing the method described in the first aspect are executed.
[0018] On the other hand, a computer program product is also provided, including a computer program which, when running on one or more processors, is used to implement the method described in the first aspect above.
[0019] The above technical solutions have the following advantages or beneficial effects:
[0020] (1) The present invention proposes a stereo image super-resolution model based on iterative interactive reference (Reference-based Iterative Interaction for Stereo Image Super-Resolution, RIISSR), which uses reference-based pixel and patch iterative matching (referred to as P 2 -Matching) to establish cross-view and cross-resolution correspondences for SSR. Specifically, the present invention first designs parallel cascaded information perception blocks (IPBs) to extract hierarchical context features of different views. Pixel matching is embedded between two parallel IPBs to utilize cross-view interaction in the low-resolution space. Then, using the super-resolution stereo image pair as mutual reference, iterative patch matching is performed, and the cross-scale patch recursive feature is used to learn high-resolution (HR) correspondences to achieve SSR performance. In addition, the present invention also introduces a supervised side output modulator (SSOM) to re-weight local intra-view features and generate intermediate super-resolution images, thus seamlessly connecting the two matching mechanisms. Experimental results on four datasets demonstrate the superior performance of the proposed RIISSR network.
[0021] (2) To implement the SSR task, a stereo image super-resolution network based on iterative interactive reference (RIISSR) is proposed, providing a new paradigm for powerful cross-view interaction in the mutual iterative reference mode for stereo image super-resolution reconstruction.
[0022] (3) The P 2 -Matching mechanism is proposed to iteratively implement pixel-level matching and patch-level matching to learn cross-view and cross-resolution correspondences. At the same time, a cross-view matching dictionary is innovatively constructed, and the cross-scale patch recursive feature is used for patch matching.
[0023] (4) The information perception module IPB is proposed to unify channel attention, large kernel attention, and feature reallocation to extract hierarchical features within information-rich views. In addition, an SSOM with high-resolution stereo image ground truth supervision function is designed, which bridges the gap between pixel-level matching and patch-level matching, thus improving the super-resolution performance. Description of the Drawings
[0024] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments and descriptions thereof of the invention are used to explain the invention and shall not unduly limit the invention.
[0025] Figure 1 The stereo image super-resolution network based on iterative interactive reference for Embodiment 1;
[0026] Figure 2 The internal structure schematic diagram of the first iterative information perception unit for Embodiment 1;
[0027] Figure 3 The internal structure schematic diagram of the pixel matching module for Embodiment 1;
[0028] Figure 4 The internal structure schematic diagram of the first tile matching module for Embodiment 1;
[0029] Figure 5 The internal structure schematic diagram of the supervision side output modulator for Embodiment 1;
[0030] Figure 6 The internal structure schematic diagram of the fusion module for Embodiment 1;
[0031] Figure 7 The visualization results of different image super-resolution methods for Embodiment 1. Detailed implementation manners
[0032] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those of ordinary skill in the technical field to which the invention belongs.
[0033] Embodiment 1
[0034] This embodiment provides an iterative interactive reference-based stereo image super-resolution reconstruction method;
[0035] The iterative interactive reference-based stereo image super-resolution reconstruction method includes:
[0036] S101: Obtain two images to be reconstructed; the two images to be reconstructed include a low-resolution left-view stereo image and a low-resolution right-view stereo image;
[0037] S102: Input the two images to be reconstructed into the trained image reconstruction model to obtain a high-resolution left-view stereo image and a high-resolution right-view stereo image;
[0038] Among them, the trained image reconstruction model uses several information perception modules to extract in-view features from two images to be reconstructed; the model also simulates the inter-view dependency based on a pixel matching module and a patch matching module; the pixel matching module forms internal crossovers and internal iterations; the patch matching module generates a matching dictionary; projects the matching dictionary into the high-resolution space, and performs high-resolution view information interaction by using the feature maps of the two views as mutual references; the model also re-weights the feature maps of the views through a supervised side output modulator to achieve patch-level matching.
[0039] Further, the low resolution specifically refers to an image with a resolution less than or equal to 400 pixels × 450 pixels;
[0040] The high resolution specifically refers to an image with a resolution less than or equal to 1600 pixels × 1800 pixels.
[0041] The low resolution and the high resolution are relative.
[0042] The stereoscopic image refers to an image pair with different perspectives obtained by two cameras shooting at two locations parallel to the scene and having only a horizontal offset.
[0043] The left-view stereoscopic image refers to the image taken by the left-view camera in the obtained stereoscopic image.
[0044] The right-view stereoscopic image refers to the image taken by the right-view camera in the obtained stereoscopic image. Further, the trained image reconstruction model includes:
[0045] Two parallel branches: the first branch and the second branch;
[0046] The first branch includes: a first convolutional layer, M information perception modules, a first upsampling layer, a first supervised side output modulator, a first patch matching module, and a first adder connected in sequence;
[0047] The second branch includes: a second convolutional layer, M information perception modules, a second upsampling layer, a second patch matching module, a second supervised side output modulator, and a second adder connected in sequence;
[0048] The first convolutional layer is used to input the low-resolution left-view stereoscopic image;
[0049] The second convolutional layer is used to input the low-resolution right-view stereoscopic image;
[0050] The output end of the i-th information perception module in the first branch is connected to the input end of the corresponding i-th pixel matching module; the output end of the i-th pixel matching module is connected to the input end of the (i + 1)-th information perception module in the first branch;
[0051] The output end of the i-th information perception module of the second branch is connected to the input end of the corresponding i-th pixel matching module; the output end of the i-th pixel matching module is connected to the input end of the (i + 1)-th information perception module of the second branch;
[0052] The output end of the first supervised-side output modulator is connected to the input end of the second tile matching module;
[0053] The output end of the second supervised-side output modulator is connected to the input end of the first tile matching module;
[0054] The output ends of the first tile matching module and the second supervised-side output modulator are both connected to the input end of the fusion module;
[0055] The output end of the fusion module is respectively connected to the input ends of the first adder and the second adder;
[0056] The input end of the first convolutional layer is connected to the input end of the first adder through the first upsampling layer;
[0057] The input end of the second convolutional layer is connected to the input end of the second adder through the second upsampling layer;
[0058] The first adder outputs a high-resolution left-view stereo image;
[0059] The second adder outputs a high-resolution right-view stereo image.
[0060] Further, as Figure 2 shown, the information perception module includes:
[0061] A perception information extractor and a refinement feed-forward network connected in sequence;
[0062] The perception information extractor includes a first normalization layer, a first pointwise convolutional layer, a first depthwise separable convolutional layer, a first channel division gating unit, a channel attention layer, a second pointwise convolutional layer, and a third adder connected in sequence; the channel attention layer is connected in parallel with a large kernel attention layer; the input end of the third adder is further connected to the input end of the first normalization layer;
[0063] The refinement feed-forward network includes a second normalization layer, a third pointwise convolutional layer, a second depthwise separable convolutional layer, a second channel division gating unit, a fourth pointwise convolutional layer, and a fourth adder connected in sequence; the input end of the fourth adder is further connected to the input end of the second normalization layer;
[0064] The input end of the first normalization layer is the input end of the information perception module, and the output end of the fourth adder is the output end of the information perception module; the input end of the second normalization layer is connected to the output end of the third adder.
[0065] Further, the channel attention layer includes: an average pooling layer, a fifth pointwise convolution layer, and a first multiplier connected in sequence, and the input end of the average pooling layer is connected to the input end of the first multiplier.
[0066] Further, the large kernel attention module includes: a 5*5 convolution layer, a 7*7 convolution layer, a sixth pointwise convolution layer, and a second multiplier connected in sequence, and the input end of the 5*5 convolution layer is connected to the input end of the second multiplier.
[0067] Further, the internal structures of the first channel division gating unit and the second channel division gating unit are the same. The first channel division gating unit includes:
[0068] A channel division layer that divides each input feature map into several slices, performs element-wise multiplication on adjacent slices, and concatenates the obtained products to obtain an output feature map.
[0069] It should be understood that the information perception module (IPB) is a transformer-style convolutional block composed of a perception information extractor and a refinement feed-forward network. Each of the two modules contains two channel division gatings for redistributing features between channels.
[0070] Further, the information perception module is used for:
[0071] Assume that the input feature of the perception information extractor is The number of channels is C; the perception information extractor first extracts local features from the feature X after layer normalization using 1×1 pointwise convolution and 3×3 depthwise separable convolution The formula is:
[0072]
[0073] Among them, Conv PW (·) represents pointwise convolution, that is, convolution with a 1×1 convolution kernel, and Conv DW (·) is depthwise separable convolution, a convolution operation in which each channel performs convolution independently, and LN(·) is layer normalization.
[0074] Then, use the first channel division gating unit (CS) to Divide it into n slices along the channel dimension.
[0075] Pair the features between channels to form n / 2 groups, and use a non-linear gating mechanism to integrate information:
[0076]
[0077] Among them, represents the (n / 2)-th output obtained by operating on the (n - 1)-th and n-th channels and ; is the GELU activation function, and ⊙ is the element-wise product; Merge and reorder the channel features of all groups to obtain the redistributed features
[0078] Then, use the large kernel attention layer (LKA) and the channel attention layer (CA) to capture long-term dependencies and global spatial distributions respectively. The refined feed-forward network controls the information flow through channel division gating, allowing each channel to focus on restoring fine details complementary to other channels.
[0079] It should be understood that pointwise convolution is an ordinary convolution with a 1×1 convolution kernel, and depthwise separable convolution is to independently assign convolution kernels to each channel of the feature for convolution operation, and the number of parameters is greatly reduced compared to ordinary convolution.
[0080] It should be understood that the channel attention layer is used to calculate the global statistics of the feature map to strengthen the attention to important features. At the same time, the large kernel attention layer uses a larger convolution kernel to capture the long-range dependencies of the images within the view, thereby strengthening the attention to local information. The combination of these two attention mechanisms enables the perception information extractor to effectively simulate the hierarchical information contained in the input image and accurately restore the texture details.
[0081] It should be understood that through the information perception module and the overall integration and stacking, the present invention improves the performance of the model while greatly improving the processing speed of the model. The information perception modules are deployed in parallel and are used to extract the intra-view features of the left and right views. Therefore, through the information perception module, the corresponding intra-view features and
[0082] It should be understood that most of the previous methods for super-resolution reconstruction of stereoscopic images perform stereo matching by calculating pixel-level feature similarities in low-resolution images. However, the resolution limitation of the low-resolution spatial domain itself and the impact of image degradation make it difficult for single pixel-level matching to fully transmit complementary information, which hinders the recovery of high-resolution fine-grained details. Although some methods propose to improve stereo correspondence by estimating a high-resolution disparity map, the additional disparity estimation network will increase the computational cost and significantly increase the number of network parameters. In addition, inaccurate prior disparity estimation will, in turn, have an adverse impact on the correspondence relationship between the left and right views.
[0083] The present invention rethinks stereo image correspondence from the perspective of reference-based reconstruction and proposes the P 2 -Matching method, regarding the stereo information supplement between the left and right views as a mutual reference mode, and learning cross-view and cross-resolution stereo correspondence relationships through reference-based pixel matching and patch-level matching.
[0084] Furthermore, as Figure 3 shown, the pixel matching module is used for:
[0085] Assume and are the low-resolution left-view and right-view stereo paired features extracted by the i-th information perception module, respectively regarded as the reference features of the other view.
[0086] First, perform layer normalization on the stereo paired features, and then send them into pointwise convolutions respectively. The output of the current view is used as the query (query), and the features of the other view are used as the key value (key). Therefore, for the reconstruction of the left view, there is the following representation:
[0087]
[0088]
[0089]
[0090] Among them, Conv PW (·) represents pointwise convolution, LN(·) represents layer normalization, Q L represents the generated left-view query, K R represents the generated right-view key value, V R represents the value generated by pointwise convolution.
[0091] For each pixel in the right view, calculate its similarity score with all possible pixels in the left view to generate a cross-attention map A R→L :
[0092]
[0093] Among them, Softmax is a normalization activation function, is matrix multiplication, (K R ) T is the transpose of the right-view key value. After calculating the scores by multiplying with the right-view eigenvalue V R matrix, the supplementary information of the left view is used to update the right-side features
[0094]
[0095] Symmetrically, for the left-side features A R→L only needs to be transposed to the cross-attention map A L→R , thereby generating the supplementary features of the left view for the right view
[0096]
[0097] Among them, (A R→L ) T is the transpose of the right-side cross-attention, that is, A L→R , V L represents the value generated by pointwise convolution.
[0098] It should be understood that in order to overcome the loss of fine-grained details in the low-resolution feature space domain, in addition to pixel-level matching, the present invention also designs patch-level matching, using the feature representation of the high-resolution space to provide higher accuracy in the high-frequency region to guide high-resolution reconstruction. The whole process includes matching dictionary construction and super-resolution patch transfer.
[0099] Furthermore, as Figure 4 shown, the internal working processes of the first patch matching module and the second patch matching module are the same. The first patch matching module is used for:
[0100] Unfolding the left and right view features into feature patches;
[0101] Based on the matching dictionary, querying the similarity score and patch position of the current view, and transferring the super-resolution reference features to the current view;
[0102] Therefore, the first iteration from the left view to the right view is expressed as:
[0103]
[0104]
[0105] Among them, denotes the projection matrix learned through sequential 1×1 point convolution and 3×3 depth convolution, is the high-resolution feature of the left view after projection, γ R is a trainable channel scaling parameter, is based on for the query operation of rearrangement, is the similarity score obtained by querying, is the feature after upsampling the right view to high resolution, is the optimized right view feature.
[0106] In the second iteration from the right view to the left view, the guiding relationship is reversed, and the improved right view feature is used as a reference. According to formulas (9) and (10), an optimized feature with more information than the original left view feature
[0107] It should be understood that by combining the pixel-level matching and tile-level matching mechanisms, the model proposed by the present invention can achieve sufficient cross-view and cross-resolution interaction, improving the accuracy of complementary information exchange. Modeling the inter-view correlation based on the cross-scale tile reproduction characteristic, because the disparity between two views and their similar tiles often reproduce multiple times at different scales. For example, given a pair of low-resolution images, their differences in the low-resolution or high-resolution space are consistent, and similar tiles can also be found. Therefore, an LR-LLR (downsampled low-resolution image) mapping can be constructed to learn the correspondence between cross-scale views, and then similar tiles are aggregated in the high-resolution space. Since most calculations are performed between the low-resolution image and the LLR image, this method is both efficient and effective.
[0108] Furthermore, the construction process of the dictionary includes:
[0109] For the low-resolution input image First, it is downsampled to an even lower resolution Then for each tile in calculate the similarity score map
[0110]
[0111] where i and j represent that the center of the tile is in the i-th row (i corresponds to the h dimension) and the j-th column, l and r correspond to the left and right views respectively, and Norm represents the normalization operation, and represent the tile sets of the left view and the right view respectively.
[0112] Then, according to the constructed similarity score map search for its similar patch in the left low-resolution image along the epipolar line The values and positions of the elements in the left dictionary are as follows:
[0113]
[0114]
[0115] Among them, and represent the maximum similarity score and position of the patch , represents finding the maximum value, represents the position of finding the maximum value.
[0116] Transpose the matching dictionary of the left view to construct another dictionary for the right view, denoted by and .
[0117] To provide high-quality super-resolution features for patch matching, after sub-pixel convolution, the upsampled features are fed into the supervised side output modulation module (SSOM) to enhance the feature representation.
[0118] Furthermore, as Figure 5 shown, the internal structures of the first supervised side output modulator and the second supervised side output modulator are the same. The first supervised side output modulator is used for:
[0119] The upsampled left features are first mapped to the image layer through channel point convolution.
[0120] Then, add the left features to the image after bicubic interpolation to restore an initial super-resolution rough image
[0121]
[0122] Among them, represents bicubic interpolation upsampling, is the left low-resolution image, Conv 1×1 is 1×1 convolution. Remap it to the feature space through 3×3 convolution, perform a non-linear gating operation on the obtained features, and generate features which are used as reference features in patch matching to facilitate the generation of right view features.
[0123]
[0124] Among them, Sigmoid is the activation function, ⊙ is the pixel-level dot product, and Conv 3×3 is a 3×3 convolution.
[0125] For the prediction Provide clear optimization supervision information according to the feature structure similarity and pixel distance with the high-resolution ground truth image
[0126] Project and into the same feature space, and the structural similarity score (Structural Similarity Index Measurement, SSIM) is used as the screening criterion:
[0127]
[0128] Subsequently, retain the top K channels with the highest similarity scores:
[0129]
[0130] Among them, represents the TOP-K selection of the structural similarity score between the high-resolution feature based on the ground truth and . To optimize the first supervision side output modulator, pointwise convolution and residual connection are used to restore a new image It should be understood that based on the parallel first supervision side output modulator and the second supervision side output modulator, the proposed RIISSR network generates optimized super-resolution stereo features and to perform tile matching as the super-resolution reference for each other.
[0131] Furthermore, as Figure 6 shown, the fusion module includes:
[0132] First, cascade the super-resolution feature with the high-resolution reference feature on its right along the channel dimension, and then send it to the convolutional layer to generate β and γ with learnable parameters of the same size as ;
[0133] Subsequently, perform layer normalization on the right view reference feature , where and are respectively the mean and standard deviation of
[0134]
[0135]
[0136] Updated γ R and β R are merged into the normalized features as weight parameters from obtained. Finally, the fusion module outputs
[0137] Furthermore, the internal functions of the first upsampling module and the second upsampling module are the same. The first upsampling module is used to increase the feature resolution by bilinear interpolation.
[0138] Furthermore, for the trained image reconstruction model, the training process includes:
[0139] Construct a training set, which is a stereo image of the left view and the right view of the known reconstructed image;
[0140] Input the training set into the image reconstruction model to train the model. When the total loss function value of the model no longer decreases or the number of iterations exceeds the set number, stop training to obtain the trained image reconstruction model.
[0141] Furthermore, the total loss function L total of the model has the following expression:
[0142] L total = L SO + L MSE + λ3·L FC (20)
[0143] where λ3 is set to 0.01;
[0144] Express the supervision of the supervision side output modulator as L SO , and its formula is:
[0145]
[0146] where and represent the ground truth of the high-resolution stereo image, and represent the stereo image pair after super-resolution of the entire network, ‖·‖ 2 is to calculate the second norm, λ1 and λ2 are set to 0.01 and 0.02 respectively, and N is the number of pixel values.
[0147] L MSE is the MSE loss:
[0148]
[0149] Among them, represents the outputs of the left and right views, representing the high-resolution stereo image ground truth pairs;
[0150] The frequency loss L FC is defined as:
[0151]
[0152] where FFT(·) is the fast Fourier transform, and the constant ε in all experiments is set to 1×10 -6 .
[0153] In fact, the stereo information supplementation between the left and right views can be regarded as a mutual reference mode. Using the reconstructed SR image as the reference image for another view helps to better reconstruct the image of the other view because the high-resolution features can provide richer texture details, which is beneficial to the utilization of cross-view information. In view of this, the present invention realizes the stereo image super-resolution SSR task from a new perspective, that is, modeling it as a reference-based image super-resolution reconstruction (Ref-SR) task, and proposes a reference-based iterative interaction for stereo image super-resolution reconstruction method (RIISSR).
[0154] RIISSR regards the stereo images of the left and right views as reference images for each other, and learns the inter-view correspondence relationship in the left and right view spaces to achieve sufficient interaction. To this end, RIISSR proposes the P 2 -Matching method, that is, simulating the inter-view dependence relationship through reference-based pixel matching and patch matching. Specifically, when performing pixel-level matching between two different views, the present invention first designs an information perception block (IPB) with parallel weight sharing to extract the intra-view information features of the two views. IPB unifies channel attention, large kernel attention, and feature reallocation, thus achieving a powerful feature representation.
[0155] Between two parallel IPBs, the present invention calculates the feature similarity of each pixel in the current view by referring to all possible differences in the other view, so as to learn the low-resolution stereo correspondence relationship. By stacking parallel IPBs and pixel matching, the model can form an internal cross and internal iteration mechanism.
[0156] Regarding patch-level matching, the present invention utilizes the cross-scale patch recursion property of stereo images (i.e., similar image structure information can be well preserved in images of different resolutions), measures their patch similarity, and thus generates a matching dictionary for querying.
[0157] Subsequently, the matching dictionary will be further projected into the high-resolution space. The network performs high-resolution view information interaction by using the super-resolution results of the two views as mutual references. The information transfer form of super-resolution to low-resolution breaks the resolution limitation of traditional SSR methods that can only interact in the low-resolution space, greatly enhancing the effective information volume.
[0158] In addition, the present invention introduces a supervised side output modulator (SSOM), which can re-weight the super-resolution local features of the views and provide high-resolution stereo image ground truth supervision for the super-resolution images to achieve patch-level matching. With the support of SSOM, pixel-level matching and patch-level matching are realized, thereby improving the accuracy and stereo consistency of SSR.
[0159] The overall framework of the stereo image super-resolution network based on iterative interactive reference is as Figure 1 shown. It is a two-stream architecture, consisting of three key parts: an information perception module (IPB) for extracting features of the left and right views; a P 2 -Matcing matching module, consisting of pixel-level matching and patch-level matching, for learning the corresponding relationships between stereo images; a supervised side output modulator (SSOM) for improving the super-resolution feature representation within the views and restoring the high-resolution reference views. Given a pair of low-resolution stereo images The goal of RIISSR is to super-resolve them into high-quality high-resolution versions In the SSR task, both intra-view information and inter-view information play crucial roles in super-resolution reconstruction. To extract the spatial features of the left and right views, after mapping the image features to 64 channels by performing 3×3 convolution, M IPBs will be cascaded in parallel and stacked to extract the low-resolution features of the corresponding views with a deeper network. Due to the particularity of stereo images, the IPBs of the two views adopt a simple weight sharing strategy and are strictly symmetric. As analyzed above, the left and right view features in the low-resolution spatial domain are not accurate enough. Therefore, both their feature extraction and cross-view interaction need to be iteratively updated multiple times to continuously correct the errors caused by resolution and original degradation and provide more accurate reference features for subsequent cross-resolution patch-level matching.
[0160] To this end, pixel-level matching is embedded between any two parallel IPBs in the present invention to perform cross-view interaction and aggregate complementary information within the view. By iteratively deploying IPBs with pixel matching capabilities, the network can form an internal cross and internal iteration mechanism, thereby improving the network's feature exchange ability and generating enhanced low-resolution stereo features. and
[0161] After that, the present invention uses sub-pixel convolution (Pixel-Shuffle) operation on and for super-resolution processing to obtain paired super-resolution features and
[0162] The supervised side output modulator (SSOM) can re-weight the super-resolution features of the left and right views and provide high-resolution stereo image ground truth supervision based on feature structural similarity and pixel distance to emphasize more discriminative features and generate intermediate layer SR images. Subsequent tile-level matching uses the super-resolution results and as mutual references to learn high-resolution stereo correspondences, and promotes matching by leveraging the property that similar tiles can repeatedly appear across scales.
[0163] To generate the final super-resolution stereo image, the present invention adopts a spatial adaptive module similar to that in MASA as the network's fusion module (see Figure 1 ), which can remap the distribution of high-resolution reference features in one view to another view. As Figure 2As shown in the figure, the RIISSR model of the present invention makes an intuitive comparison of different 4× super-resolution methods on the Flickr1024 and Middlebury validation datasets. In the first group of complex outdoor scenes, the proposed RIISSR of the present invention accurately restores the texture of the window sill, showing clearer edges and details. The single-image super-resolution algorithm SwinIR and the binocular super-resolution algorithm in the first row can basically not restore the texture details and present a completely blurred state. Although the single-image SwinIR algorithm has better quantitative performance than the binocular algorithm PASSRnet, its visual reconstruction details are inferior. The binocular super-resolution methods in the second row restore certain rough details. In the reconstructed indoor "sword" scene, the RIISSR also significantly generates fine details at the indentation of the basketball hoop shown, and the sharpness in the lower right corner even has better visual performance than the HR. Similar to the outdoor scene, the methods in the first row cannot effectively reconstruct the scene details while the second row has better visual effects. The results show that compared with the existing state-of-the-art SSR methods, RIISSR has higher reconstruction performance and excellent indoor and outdoor scene restoration effects. Figure 7 Visualization results of different image super-resolution methods in Example 1.
[0164] Example 2
[0165] This embodiment provides an iterative interactive reference-based stereoscopic image super-resolution reconstruction system;
[0166] The iterative interactive reference-based stereoscopic image super-resolution reconstruction system includes:
[0167] An acquisition module, which is configured to: acquire two images to be reconstructed; the two images to be reconstructed include a low-resolution left-view stereoscopic image and a low-resolution right-view stereoscopic image;
[0168] A reconstruction module, which is configured to: input the two images to be reconstructed into a trained image reconstruction model to obtain a high-resolution left-view stereoscopic image and a high-resolution right-view stereoscopic image;
[0169] Among them, the trained image reconstruction model uses a number of information perception modules to extract the features within the view from the two images to be reconstructed; the model also simulates the inter-view dependence based on a pixel matching module and a patch matching module; the pixel matching module forms internal intersections and internal iterations; the patch matching module generates a matching dictionary; projects the matching dictionary into the high-resolution space, and performs high-resolution view information interaction by using the feature maps of the two views as mutual references; the model also re-weights the feature maps of the views through a supervised side output modulator to achieve patch-level matching.
[0170] It should be noted here that the above acquisition module and reconstruction module correspond to steps S101 to S102 in the first embodiment. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the first embodiment above. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.
[0171] In the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0172] The proposed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the above module division is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0173] Embodiment Three
[0174] This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, the above one or more computer programs are stored in the memory, and when the electronic device runs, the processor executes the one or more computer programs stored in the memory so that the electronic device executes the method described in the first embodiment above.
[0175] It should be understood that in this embodiment, the processor can be a central processing unit CPU, and the processor can also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0176] The memory can include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory can also include a non-volatile random memory. For example, the memory can also store information about the device type.
[0177] In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software.
[0178] The method in the first embodiment can be directly embodied as being executed by a hardware processor, or by a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0179] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0180] Embodiment Four
[0181] This embodiment also provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first embodiment is completed.
[0182] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. Iterative interactive reference-based stereoscopic image super-resolution reconstruction method, characterized in that including: Obtain two images to be reconstructed; The two images to be reconstructed include a low-resolution left-view stereo image and a low-resolution right-view stereo image; Input the two images to be reconstructed into the trained image reconstruction model to obtain a high-resolution left-view stereo image and a high-resolution right-view stereo image; Among them, the trained image reconstruction model uses a number of information perception modules to extract in-view features from the two images to be reconstructed; the model also simulates the inter-view dependency relationship based on the pixel matching module and the patch matching module; the pixel matching module forms internal cross and internal iteration; the patch matching module generates a matching dictionary; project the matching dictionary into the high-resolution space, and perform high-resolution view information interaction by using the feature maps of the two views as mutual references; the model also re-weights the feature maps of the views through the supervised side output modulator to achieve patch-level matching; the supervised side output modulator is used for: Left feature after upsampling First, it is mapped to the image layer through channel point convolution; Then, the left-side features are added to the image after bicubic interpolation to restore an initial super-resolution rough image : (14) Among them, represents bicubic interpolation upsampling, is the left low-resolution image, is convolution; through convolution, it is remapped to the feature space, a non-linear gate operation of residuals is performed on the obtained features, and is generated. The feature is used as a reference feature in patch matching to facilitate the generation of right-view features; (15) Among them, is an activation function, is a pixel-level dot product, is convolution; For the prediction ,explicit optimization supervision information is provided according to the feature structure similarity and pixel distance from the high-resolution ground truth image ; Project and onto the same feature space, and the structural similarity score is used as the screening criterion: (16) Subsequently, the top channels with the highest similarity scores are retained: (17) Among them, represents the TOP-K selection of the structural similarity score between the high-resolution features based on the ground truth and ; Pointwise convolution and residual connection are used to recover a new image .
2. The iterative interactive reference-based stereoscopic image super-resolution reconstruction method according to claim 1, characterized in that, The information perception module is used for: Suppose the input features of the perception information extractor are , and the number of channels is ; The perception information extractor first uses pointwise convolution and depthwise separable convolution to extract local features from the features after layer normalization. The formula is: (1) Among them, represents pointwise convolution, that is, convolution with a convolution kernel of ; is depthwise separable convolution, a convolution operation in which each channel performs convolution independently, is layer normalization; Then, use the first-channel partitioning gating unit to be divided along the channel dimension into n slices; Pair the features between channels to form groups, and use a non-linear gating mechanism to integrate information: (2) Among them, represents the th and n channel and to obtain the th output. is the GELU activation function, is the element-wise product; merge and reorder the channel features of all groups to obtain the redistributed feature ; then, use the large kernel attention layer and the channel attention layer to capture long-term dependencies and global spatial distributions respectively; the refined feed-forward network controls the information flow through channel division gating.
3. The iterative interactive reference type three-dimensional image super-resolution reconstruction method according to claim 1, characterized in that The pixel matching module is used for: Hypothesis and are the low-resolution left and right view stereo matching features extracted by the i th information perception module, regarded as the reference features of another view respectively; First, perform layer normalization on the stereo paired features, and then send them into pointwise convolution respectively. The output of the current view is used as the query and the features of the other view are used as the key; for the reconstruction of the left view, there is the following representation: (3) (4) (5) Among them, represents pointwise convolution, represents layer normalization, represents the generated left view query, represents the generated right view key-value, represents the one produced by pointwise convolution value; for each pixel in the right view, calculate its similarity score with all possible pixels in the left view to generate a cross-attention map : (6) Among them, is a normalization activation function, is matrix multiplication, is the transpose of the right view key value; after calculating the score by multiplying with the right view eigenvalue matrix, the supplementary information of the left view is used to update the right side feature : (7) Symmetrically, for the left - hand features , only need to be transposed into the cross - attention map , thereby generating the supplementary features of the left view to the right view : (8) Among them, is the transpose of the right cross-attention , represents the value generated by pointwise convolution.
4. The iterative interactive reference type three-dimensional image super-resolution reconstruction method according to claim 1, characterized in that The patch matching module is used for: expanding the left and right view features into feature patches; Based on the matching dictionary, query the similarity score and patch position of the current view, and transfer the reference feature SR to the current view; therefore, the first iteration from the left view to the right view is expressed as: (9) (10) Among them, represents the projection matrix learned through sequential point convolution and depth convolution, is the high-resolution feature of the left view after projection, is a trainable channel scaling parameter, is based on to perform a rearrangement query operation on is the similarity score obtained from the query, is the feature of the right view after upsampling to high resolution, is the optimized right view feature; In the second iteration from the right view to the left view, the guiding relationship is reversed, and the improved right view features are utilized as a reference; according to formulas (9) and (10), optimized features with greater information content than the original left view features are obtained . 5. The iterative interactive reference type three-dimensional image super-resolution reconstruction method according to claim 4, characterized in that, The construction process of the dictionary includes: For a low-resolution input image , first downsample it to an even lower resolution , then for each tile in , calculate the similarity score map : (11) Among them, indicates that the center of the patch is located at the th row and the th column, corresponding to the left and right views respectively, Norm represents the normalization operation, and represent the patch sets of the left and right views respectively; then, according to the constructed similarity score map , search for its similar patch in the left low-resolution image . Then, the values and positions of the elements in the left dictionary are: (12) (13) Among them, and represent the maximum similarity score and position of the tile , represents finding the maximum value, represents the position of finding the maximum value; transpose the matching dictionary of the left view to construct another dictionary for the right view, and use and to represent.
6. Iterative interactive reference-based stereoscopic image super-resolution reconstruction system, characterized in that, including: An acquisition module configured to: obtain two images to be reconstructed; The two images to be reconstructed include a low-resolution left-view stereo image and a low-resolution right-view stereo image; A reconstruction module configured to: input the two images to be reconstructed into the trained image reconstruction model to obtain a high-resolution left-view stereo image and a high-resolution right-view stereo image; Among them, the trained image reconstruction model uses a number of information perception modules to extract in-view features from the two images to be reconstructed; the model also simulates the inter-view dependency relationship based on the pixel matching module and the patch matching module; the pixel matching module forms internal cross and internal iteration; the patch matching module generates a matching dictionary; project the matching dictionary into the high-resolution space, and perform high-resolution view information interaction by using the feature maps of the two views as mutual references; the model also re-weights the feature maps of the views through the supervised side output modulator to achieve patch-level matching; the supervised side output modulator is used for: Left feature after upsampling First, it is mapped to the image layer through channel point convolution; Then, add the left-side features to the image after bicubic interpolation to restore an initial super-resolution rough image : (14) Among them, represents bicubic interpolation upsampling, is the left low-resolution image, is convolution; through convolution, it is remapped to the feature space, a nonlinear gate operation of residuals is performed on the obtained features, and is generated, and the feature is used as a reference feature in patch matching to facilitate the generation of right-view features; (15) Among them, is an activation function, is a pixel-level dot product, is a convolution; For the prediction of , clear optimization supervision information is provided according to the feature structure similarity and pixel distance with the high-resolution ground truth image ; Project and onto the same feature space, and the structural similarity score is used as the screening criterion: (16) Subsequently, the top channels with the highest similarity scores are retained: (17) Among them, represents the TOP-K selection of the structural similarity score between and high-resolution features based on the ground truth; pointwise convolution and residual connections are used to recover a new image .
7. An electronic device, characterized in that it includes: A memory for non-temporarily storing computer-readable instructions; and A processor for running the computer-readable instructions, wherein, when the computer-readable instructions are run by the processor, the method according to any one of claims 1-5 above is executed.
8. A storage medium, characterized in that it is non-transitory Stored computer-readable instructions which, when executed by a computer, execute the instructions of the method according to any one of claims 1-5.
9. A computer program product, characterized in that, Comprising a computer program which, when run on one or more processors, is used to implement the method according to any one of claims 1-5 above.
Citation Information
Patent Citations
Quick image rectification method in presence of translation and rotation at same time
CN101950419A
Image acquisition method, electronic equipment and computer readable storage medium
CN117177052A