Binocular image super-resolution reconstruction method and system based on hierarchical Transform

Through the hierarchical window design and improved space-channel correlation mechanism, combined with the disparity-aware window expansion, the problem of failing to make full use of stereo matching and disparity information in the existing methods is solved, and efficient and accurate binocular image super-resolution reconstruction is achieved, improving the reconstruction quality and resolution of the image.

CN120259078APending Publication Date: 2025-07-04CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510313390.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing Transformer-based binocular image super-resolution method fails to fully utilize stereo matching and parallax information, resulting in poor reconstruction results when processing binocular images.

Method used

The hierarchical window design and improved space-channel correlation (SCC) mechanism are adopted, combined with the disparity-aware window expansion, feature extraction and fusion are performed through the hierarchical Transformer module, and efficient and accurate super-resolution reconstruction is carried out using the stereo matching and disparity information of binocular images.

Benefits of technology

It significantly improves the reconstruction quality and resolution of binocular images, can better capture multi-scale features and global information, and improves the detailed performance and accuracy of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259078A_ABST
    Figure CN120259078A_ABST
Patent Text Reader

Abstract

The invention provides a binocular image super-resolution reconstruction method and a binocular image super-resolution reconstruction system based on hierarchical Transform. According to the method, through layered window design and an improved space-channel correlation (SCC) module, parallax information of a binocular image is fully utilized, and efficient and accurate super-resolution reconstruction is realized. The method comprises the following specific steps: carrying out normalization and denoising processing on an input low-resolution binocular image; extracting multi-scale features by using a convolutional layer; performing parallax estimation based on a layered window and an attention mechanism, and aligning left and right image features; an improved SCC module is used for extracting multi-scale features and carrying out cross-view fusion; the fused features are input into a super-resolution reconstruction network based on Transform, and a high-resolution image is generated; and optimizing the model by adopting a joint optimization strategy. The method has important application value in the fields of automatic driving, virtual reality and the like, and the resolution and detail expression capability of binocular images can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to a binocular image super-resolution reconstruction method and system based on hierarchical Transformer, which is used to improve the resolution of binocular images, enhance the details and clarity of images, and provide higher-quality image data for fields such as computer vision, autonomous driving, and virtual reality. Background Art

[0002] In the field of computer vision, the importance of binocular image super-resolution reconstruction technology has become increasingly prominent. It aims to recover high-resolution images from low-resolution binocular images, providing a high-quality data foundation for subsequent vision tasks. In the scenario of autonomous driving, accurate binocular image super-resolution reconstruction enables vehicles to more clearly identify targets such as pedestrians, vehicles, and traffic signs on the road, thereby significantly improving the safety and reliability of the autonomous driving system; in virtual reality and augmented reality applications, high-resolution binocular images can create a more realistic and immersive experience environment for users; in medical image analysis, this technology helps doctors more clearly observe the details of human tissues and organs, thereby improving the accuracy of disease diagnosis.

[0003] Most traditional binocular image super-resolution methods are based on convolutional neural networks (CNNs). With its characteristics of local connection and weight sharing, CNN shows certain advantages in processing local features of images and can effectively learn information such as texture and edges in local regions of images. However, the inherent limitations of CNN also pose challenges when dealing with long-range dependencies and multi-scale features. The long-range dependency problem means that in an image, the relationship between regions that are far apart is difficult to be effectively captured by CNN. For example, in a binocular image containing distant mountains and nearby trees, CNN may not be able to well associate the features of these two objects at different distances. When dealing with multi-scale features, CNN usually needs to stack multiple convolutional layers of different scales to achieve this, but this method often has difficulty in balancing the weights of different-scale features and information fusion, resulting in incomplete and inaccurate feature extraction for objects of different sizes.

[0004] In recent years, the Transformer architecture has made breakthrough progress in the fields of natural language processing and computer vision. Its self-attention mechanism can capture the dependencies between information globally, which has unique advantages in dealing with long-range dependencies and global context information. Through the self-attention mechanism, Transformer can dynamically allocate weights, focus on features at different positions in the image, and thus better understand the overall structure and content of the image. Nevertheless, existing Transformer-based image super-resolution methods mainly focus on monocular images. When dealing with binocular images, due to the unique stereo matching and disparity information in binocular images, existing Transformer-based methods lack effective mechanisms to fully utilize this information. Stereo matching refers to the process of finding corresponding points in the left and right images. Accurate stereo matching is crucial for super-resolution reconstruction, but the accuracy and efficiency of existing methods in this regard need to be improved. Disparity information reflects the position difference of objects in the left and right views and contains rich depth information. However, existing methods have not fully exploited and utilized this information, resulting in the inability to fully unleash the potential of Transformer in the binocular image super-resolution reconstruction task. Summary of the Invention

[0005] The present invention innovatively proposes a binocular image super-resolution reconstruction method and system based on hierarchical Transformer. By introducing an improved hierarchical window and spatial-channel correlation (SCC) mechanism, it realizes efficient and accurate binocular image super-resolution reconstruction and solves the problems existing in the prior art.

[0006] The hierarchical Transformer module, as the core component of the present invention, includes the following key designs:

[0007] Hierarchical window design: In the Transformer layer, the present invention adopts a strategy of gradually expanding the window size. Starting from an initial small window, as the network depth increases, the window gradually expands to comprehensively capture the multi-scale features of the image. Specifically, the window size is set to h i = b i h B and, where b i represents the expansion ratio of the i-th layer, h B and X ch= Linear(X) is the base window size. A smaller initial window can focus on the local details of the image, capturing fine information such as textures and small objects; as the window gradually expands, the model can obtain more extensive context information, including global features such as the overall shape of the object and the positional relationship. This design enables the model to conduct in-depth analysis of the image at different scales, fully integrating local and global information, and providing rich and comprehensive feature support for subsequent super-resolution reconstruction.

[0008] Improve spatial-channel correlation (SCC): To efficiently aggregate spatial and channel information, the present invention improves SCC, which consists of the following key parts:

[0009] Dual feature extraction (DFE): DFE extracts features from the spatial and channel dimensions respectively through carefully designed convolutional branches and linear branches. Given the input features, the output of the improved DFE is calculated by the formula: where X ch = Linear(X) represents extracting channel features through linear transformation, and X sp = Conv(X) represents extracting spatial features using convolutional operations. ⊙ represents element-wise multiplication. In this way, spatial and channel features are fused, while retaining the original convolutional branch and linear branch features, enabling the extracted features to contain both spatial structure and inter-channel correlation information, effectively enhancing the expressive power of the features.

[0010] Spatial self-correlation (S-SC): S-SC uses a self-correlation mechanism with linear complexity to aggregate spatial information. During the calculation, given the query and values, the improved formula for S-SC is:

[0011] where D is the normalization constant. Compared with the traditional self-attention mechanism, this design can significantly reduce the computational complexity while maintaining information accuracy, enabling the model to still operate efficiently when processing large-sized images and being able to better capture the feature dependency relationships in the spatial dimension.

[0012] Channel self-correlation (C-SC): C-SC also uses a self-correlation mechanism with linear complexity to aggregate channel information. Given the query Q and the value V, its improved calculation formula is:

[0013] By increasing the fine-tuning of D, the C-SC can more deeply explore the correlation between channels, capture the interaction between different channel features, further optimize the feature representation, and provide more discriminative channel information for image reconstruction.

[0014] In order to fully exploit and utilize the left - right disparity information of binocular images, the present invention designs a complete feature fusion module, which specifically includes the following steps:

[0015] Step 1: Feature extraction operations are respectively performed on the left and right images. Through a carefully designed convolutional layer and Transformer layer, representative aligned features F L A and F R A are generated. In this process, the characteristics of binocular images are fully considered, and the network structure and parameters are optimized to ensure that the extracted features can accurately reflect the information in the images, laying a solid foundation for subsequent feature fusion and disparity processing.

[0016] Step 2: With the help of a specially designed feature interaction module, the features F L A and F R A of the left and right images are deeply fused. This module realizes the efficient interaction and fusion of left - right features F F through a series of operations such as convolution and attention mechanisms, generating fused features. During the fusion process, not only can the respective feature advantages of the left and right images be retained, but also the depth clues contained in the disparity information can be fully exploited, making the fused features more representative and complementary.

[0017] Step 3: In the key process of window expansion, the present invention dynamically adjusts the window size according to the disparity map. The disparity map reflects the position differences of corresponding points in the left and right images. By analyzing the disparity map, the model can intelligently adjust the window size to ensure more accurate capture of the corresponding relationship between the left and right images in different regions. In regions with object edges or large disparity changes, the window is appropriately reduced to improve the matching accuracy; in regions with similar textures or small disparities, the window is enlarged to obtain richer context information. This disparity - aware window expansion mechanism can significantly improve the accuracy and effectiveness of feature fusion, further enhancing the model's ability to understand and process binocular images.

[0018] Using the fused feature F F , a high - resolution image is reconstructed through a carefully designed convolutional layer and upsampling layer. The convolutional layer further extracts the detailed information in the fused features, and the upsampling layer gradually restores the low - resolution feature map to a high - resolution image. During the upsampling process, advanced upsampling algorithms such as sub - pixel convolution are used, which can effectively avoid problems such as jaggedness and blurring in traditional upsampling methods, thus reconstructing a high - quality, detail - rich high - resolution image. Brief Description of the Drawings

[0019] Figure 1Binocular Image Super-Resolution Framework Based on Hierarchical Transformer;

[0020] Figure 2 Schematic Diagram of Hierarchical Window Design;

[0021] Figure 3 Structural Diagram of Spatial-Channel Correlation (SCC) Module;

[0022] Figure 4 Schematic Diagram of Binocular Image Feature Fusion Module; Specific Implementation Scheme

[0023] The following further elaborates on the specific implementation of the present invention with reference to the accompanying drawings.

[0024] First, obtain the low-resolution left and right images. These two images serve as the input for the entire super-resolution reconstruction process, and their quality and resolution directly affect the final reconstruction effect. In practical applications, these low-resolution images may come from various acquisition devices such as cameras, sensors, etc. After obtaining the images, preliminary preprocessing operations such as cropping and normalization can be performed on the images as needed to ensure that the images meet the requirements of subsequent processing.

[0025] Feature Extraction: Use a carefully designed convolutional layer to separately perform shallow feature extraction on the input low-resolution left and right images to obtain shallow features and. These shallow features contain basic information of the images, such as edges, textures, etc. In the design of the convolutional layer, different sizes of convolutional kernels and numbers of convolutional layers can be selected to adapt to different image features and task requirements. Through reasonable setting of convolutional layer parameters, key information in the images can be effectively extracted, providing a basis for subsequent deep feature extraction and processing.

[0026] Hierarchical Transformer Module: Input the shallow features and into the hierarchical Transformer module for deep feature extraction to generate deep features and

[0027] Hierarchical Window Design: In the Transformer layer, strictly set the window size according to the rules of h i = b i h B and ω i = b i ω i In practical applications, the basic window sizes h B , ω B , and the expansion ratio b i can be reasonably selected according to the characteristics of the images and task requirements.For images with complex textures, the basic window size can be appropriately reduced to better capture local details; for scenarios where global information needs to be focused on, the expansion ratio can be increased to enable the window to expand faster to obtain global features. Through this hierarchical window design, the model can effectively capture multi-scale features at different levels, enhancing the ability to understand and analyze images.

[0028] Spatial-Channel Correlation (SCC): In the hierarchical Transformer module, the SCC method is optimized and improved to process features, achieving efficient aggregation of spatial and channel information.

[0029] Dual Feature Extraction (DFE): For the input feature X, perform dual feature extraction according to the formula In actual calculations, use linear layers and convolutional layers to extract the channel feature X ch and the spatial feature X sp respectively, and fuse the two by element-wise multiplication and then add the channel feature X ch and the spatial feature X sp to obtain more complete hierarchical features. During implementation, the parameters of the linear layer and convolutional layer can be optimized to improve the accuracy and efficiency of feature extraction.

[0030] Spatial Self-Correlation (S-SC): Use the formula:

[0031] to perform spatial self-correlation calculation. In actual operations, select the normalization constant reasonably according to the window size and feature dimension to ensure the stability and effectiveness of the calculation results. Through the S-SC operation, it is possible to effectively capture the feature dependence relationship in the spatial dimension while keeping the computational complexity relatively low, providing strong support for subsequent feature fusion and image reconstruction.

[0032] Channel Self-Correlation (C-SC): According to the formula:

[0033] to perform channel self-correlation calculation. In actual applications, adjust the normalization constant according to the number of channels of the features and the data distribution to optimize the aggregation effect of the channel features. The C-SC operation can deeply explore the correlation between channels, enhance the feature expression ability, and further improve the model's ability to understand and process images.

[0034] Binocular Image Feature Fusion: Fuse the left and right image features processed by the hierarchical Transformer module to generate the fused feature F F .

[0035] Left-Right Feature Extraction: Further process the deep features of the left and right images, and generate aligned features through a specially designed network structure and In this process, an attention mechanism-based method can be adopted to dynamically adjust the weights of features according to the content and disparity information of the images, so that the extracted aligned features can more accurately reflect the correspondence between the left and right images.

[0036] Feature interaction: Use a feature interaction module to fuse the aligned left and right features and together. In the feature interaction module, technologies such as convolutional neural networks and attention mechanisms can be used to achieve deep fusion of left and right features. Through multiple convolutional operations and attention calculations, the fused features can fully contain the information of the left and right images and can better utilize the depth cues contained in the disparity information.

[0037] Disparity-aware window expansion: During the window expansion process, dynamically adjust the window size according to the calculated disparity map. By analyzing the disparity map, determine the disparity magnitude and variation in different regions. In regions with large disparities, appropriately reduce the window size to improve the accuracy of feature matching; in regions with small disparities, expand the window size to obtain more abundant context information. Through this disparity-aware window expansion mechanism, the effect of feature fusion can be significantly improved, providing more accurate feature information for subsequent image reconstruction.

[0038] Image reconstruction: Use the fused features to reconstruct a high-resolution image through carefully designed convolutional layers and upsampling layers. In the convolutional layer, multiple transposed convolution operations can be adopted to further extract the detailed information in the fused features. The upsampling layer uses advanced upsampling algorithms such as sub-pixel convolution to gradually restore the low-resolution feature map to a high-resolution image. During the reconstruction process, the parameters of the transposed convolution layer and the upsampling layer can be optimized as needed to improve the quality and resolution of the reconstructed image. At the same time, some post-processing techniques, such as image enhancement and denoising, can be combined to further enhance the visual effect of the reconstructed image.

Claims

1. A binocular image super-resolution reconstruction method based on hierarchical Transformer, Its features include the following steps: Step S1: Normalize and denoise the input low-resolution binocular images to improve image quality and lay a foundation for subsequent feature extraction; Step S2: Use convolutional layers with different kernel sizes to extract features from the preprocessed binocular images, obtaining feature maps of different scales to capture local and global information of the images; Step S3: Perform disparity estimation based on a hierarchical window and attention mechanism, align the features of the left and right images, and ensure the spatial consistency of the left and right image features to provide an accurate correspondence for subsequent feature fusion; Step S4: Apply an improved spatial-channel correlation (SCC) module in the left and right image branches to extract multi-scale features and perform cross-view feature fusion, making full use of the left and right disparity information of the binocular images to enhance the expressive power of the features; Step S5: Input the fused features into a Transformer-based super-resolution reconstruction network to generate high-resolution binocular images, and utilize the long-range dependence capture ability of the Transformer to further enhance the details and textures of the images; Step S6: Adopt a joint optimization strategy, considering the super-resolution reconstruction loss, disparity smoothness loss, and perceptual loss simultaneously to optimize the model and improve the quality and visual effect of the reconstructed images.

2. The method according to claim 1, wherein The setting method of the hierarchical window is as follows: Given the basic window size, the window size is gradually expanded through the hierarchical ratio b i By gradually expanding the window size, the model can capture features of different scales, thereby improving the effect of image super-resolution. In addition, the expansion strategy of the hierarchical window also includes dynamically adjusting the window size according to the complexity of the image content to better meet the feature extraction requirements in different scenarios.

3. The method according to claim 1, wherein In the parallax estimation step, the parallax offset is obtained by calculating the correlation matrix of the features of the left and right images. The calculation formula is where Q is the query of the left image and K is the key value of the right image. In this way, the parallax is accurately estimated, providing an accurate basis for feature alignment.

4. The method according to claim 1, wherein The improved SCC module includes a dual feature extraction (DFE) layer, which extracts spatial and channel features through a convolutional branch and a linear branch respectively, and incorporates disparity information into the feature extraction process to enhance the expressive power of the features and the sensitivity to disparity.

5. The method according to claim 1, wherein The super-resolution reconstruction network adopts a Transformer-based structure and combines a deconvolution layer and a sub-pixel convolution layer for image reconstruction, where the Transformer structure uses a multi-head attention mechanism to capture the long-range dependence relationship between features and further improve the quality of the reconstructed images.