Light field image super-resolution reconstruction method and device, terminal and storage medium
By extracting features through EPI-Mamba and SA-Mamba subspace scanning, and combining multi-scale interactive fusion and dual attention modulation, the light field image is reconstructed using a diffusion model. This solves the problems of edge detail loss and high computational cost in existing technologies, and achieves efficient super-resolution reconstruction of light field images and restoration of high-frequency details.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-03-17
AI Technical Summary
Existing light field image super-resolution reconstruction techniques are prone to loss of edge details or excessive smoothing, have high computational overhead and slow processing speed, fail to fully utilize the multidimensional information of light field images, and are difficult to simultaneously take into account global-local dependencies and high-frequency detail recovery.
Features are extracted using EPI-Mamba and SA-Mamba subspace scanning. Reconstruction is performed in the residual domain and frequency domain by multi-scale interactive fusion and dual attention modulation, combined with a diffusion model, to ensure the structural consistency and detail fidelity of the light field image.
It achieves efficient super-resolution reconstruction of light field images, improves the visual quality of images, effectively restores high-frequency details and complex structures, and solves the problem of dimensional collaborative optimization in traditional methods.
Smart Images

Figure CN121685262A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image super-resolution reconstruction, and in particular to a light field image super-resolution reconstruction method and device, a terminal and a storage medium. BACKGROUND
[0002] Light field image super-resolution reconstruction technology has important significance in the field of computer vision and image processing. Light field cameras can capture both spatial position information and propagation direction information of light, providing rich depth information for subsequent image processing and analysis. With the continuous development of computational photography, three-dimensional reconstruction and computer vision, the demand for spatial resolution of light field images is increasing. Light field image super-resolution reconstruction technology can generate higher resolution images using the information in existing low-resolution light field images without increasing hardware costs, which has important practical significance for high-speed moving object tracking, multi-agent target recognition, complex scene detail recovery and image quality improvement in low-light environments.
[0003] In the task of light field super-resolution reconstruction, there are currently traditional methods and deep learning-based methods. In traditional methods, interpolation and reconstruction algorithms improve image resolution through simple mathematical operations. In recent years, convolutional neural network (CNN)-based methods have been widely used. Some solutions use a single network to process spatial information of light field images by stacking sub-aperture images as input. Techniques using multi-scale network structures extract features from different scales to enhance detail recovery. More advanced deep learning methods, such as generative adversarial networks (GANs) and light field-specific networks (such as SAR-Net), can more accurately capture spatial and angular information, improving image reconstruction quality and detail recovery.
[0004] However, the existing technology has obvious defects. Existing deep learning methods are prone to edge detail loss or excessive smoothing, and cannot accurately recover image structure. Large computational overhead, slow processing speed and hardware performance limitations are also significant problems, limiting the real-time performance and application range of the technology. Many methods fail to fully utilize the multi-dimensional information such as space and angle in light field images, resulting in incomplete recovery of image details and depth information. Traditional methods often rely on convolution or attention mechanisms for direct modeling in the image domain, making it difficult to simultaneously consider global-local dependencies and high-frequency detail recovery. There is a trade-off between spatial and angular resolution, resulting in insufficient detail recovery. SUMMARY
[0005] The purpose of the present application is to overcome the above technical problems, and a light field image super-resolution reconstruction method, device, terminal and storage medium are provided, which ensures the structural consistency and detail fidelity of high-resolution light field images.
[0006] Firstly, this application provides a method for super-resolution reconstruction of light field images, which adopts the following technical solution: A method for super-resolution reconstruction of light field images includes the following steps: Based on the EPI-Mamba subspace scan, feature extraction is performed on the received original light field image to obtain EPI structural features, and based on the SA-Mamba subspace scan, feature extraction is performed on the original light field image to obtain spatial-angle correlation features; Based on the EPI structural features and the spatial-angle correlation features, multi-scale interactive fusion across subspaces is performed, and global dependencies are modeled through bidirectional scanning to obtain multi-scale fusion features; The multi-scale fusion features are subjected to dual attention modulation of space and angle to obtain comprehensive enhanced features; Under the pre-built diffusion model framework, guided by the enhanced features, the predicted residual image is reconstructed under the dual constraints of the residual domain and frequency domain through stepwise denoising. The predicted residual image is then compensated to obtain a high-resolution light field image.
[0007] By employing the aforementioned techniques, leveraging EPI-Mamba's precise capture of the epipolar structure of the light field and SA-Mamba's global modeling advantages in spatial-angular correlation, the difficulties caused by long-term dependencies in long-sequence modeling are alleviated, maintaining the integrity of 4D global information. Multi-scale interactive fusion enables deep complementarity and coupling of information from different dimensions at the global level. Dual attention modulation of space and angle strengthens the feature representation of key regions, enhancing overall representational capabilities. The diffusion model reconstruction process under dual constraints in the residual and frequency domains allows the model to focus more on texture edges and high-frequency details, reducing generation difficulty, accelerating training convergence, ensuring cross-view consistency and global structural stability, ultimately improving the resolution of the light field image, effectively restoring high-frequency details and complex structures, suppressing blurring and artifacts, and enhancing the visual quality of the image. This solves the problem of traditional methods struggling to simultaneously optimize the spatial and angular dimensions of the light field.
[0008] Preferably, the step of extracting features from the received original light field image based on EPI-Mamba subspace scanning to obtain EPI structural features, and extracting features from the original light field image based on SA-Mamba subspace scanning to obtain spatial-angle correlation features, specifically includes the following steps: Receive a low-resolution raw light field image, perform preliminary feature extraction on the raw light field image to obtain initial features, fuse the raw light field image and the initial features to obtain light field features, the light field features are specifically 4D features composed of spatial dimension (u,v) and angular dimension (h,w), denoted as I(u,v,h,w). The optical field features are unfolded into the EPI-Mamba subspace and the SA-Mamba subspace for bidirectional scanning, wherein the EPI-Mamba subspace includes the EPI-H subspace and the EPI-W subspace; specifically, In the EPI-Mamba subspace, the light field features are fixed in the angle dimension in the EPI-H subspace and the EPI-W subspace, and unfolded along the spatial dimension to obtain horizontal EPI slices and vertical EPI slices. After the obtained slices are flattened into a one-dimensional sequence, bidirectional subspace scanning is performed and bidirectional dependency modeling is executed to obtain EPI structural features. The EPI structural features are used to represent the interaction relationship between space and angle. In the SA-Mamba subspace, scanning is performed along the spatial or angular dimensions to obtain spatial and angular sequences. After flattening the obtained sequences into one-dimensional sequences, bidirectional subspace scanning is performed and bidirectional dependency modeling is executed to obtain spatial-angular correlation features that characterize the association between different viewpoints and spatial locations.
[0009] By employing the aforementioned techniques, the original light field is first transformed into a 4D feature tensor, which is then bidirectionally scanned in both the EPI-Mamba and SA-Mamba subspaces. In the EPI-Mamba subspace, EPI structural features representing the interaction between space and angle can be obtained, while in the SA-Mamba subspace, spatial-angle correlation features representing the relationship between different viewpoints and spatial positions can be obtained. This allows EPI-Mamba and SA-Mamba to focus on the specific extraction of epipolar geometric features and spatial-angle correlation features, respectively. In principle, this achieves the separation and precise capture of information in different dimensions of the light field, avoiding information redundancy or loss caused by single feature extraction, and providing high-purity feature input for subsequent fusion.
[0010] Preferably, in the EPI-Mamba subspace, after flattening the obtained slices into a one-dimensional sequence, bidirectional subspace scanning and bidirectional dependency modeling are performed, specifically including the following steps: The horizontal EPI slice is flattened into a one-dimensional first sequence, and the vertical EPI slice is flattened into a one-dimensional second sequence. The first sequence and the second sequence are respectively input into the bidirectional subspace scanning module, and LayerNorm normalization processing, forward-backward bidirectional scanning modeling, secondary LayerNorm normalization processing, and channel attention interaction processing are performed in sequence. A convolutional layer is connected in series after the bidirectional subspace scanning module to enhance the features of the sequence after the bidirectional dependency modeling, wherein the bidirectional subspace scanning module and the convolutional layer corresponding to the horizontal EPI slice and the vertical EPI slice share weights. The features output from the convolutional layer are reshaped into a 4D feature tensor to obtain the EPI structural features.
[0011] By employing the above-mentioned techniques, LayerNorm normalization is used to eliminate feature distribution differences, forward-backward bidirectional scanning is used to model long-distance dependencies, and channel attention is used to enhance key feature interactions. At the same time, the symmetry of horizontal and vertical EPI slices is used to design weight sharing, which not only ensures the global correlation and extraction stability of epipolar structure features, but also improves network efficiency through parameter reuse, thus solving the problem of balancing long sequence modeling and computational cost in EPI feature extraction.
[0012] Preferably, the step of performing multi-scale interactive fusion across subspaces based on the EPI structural features and the spatial-angle correlation features, and obtaining multi-scale fusion features by modeling global dependencies through bidirectional scanning, specifically includes the following steps: The EPI structural features and the spatial-angle correlation features are combined to obtain the features to be fused. The features to be fused are expanded into the first token sequence of the EPI-H subspace. The first token sequence is then globally dependently modeled using a bidirectional subspace scanning module to obtain the EPI-H enhanced sequence. The EPI-H enhanced sequence is reshaped into the second token sequence of the EPI-W subspace. The second token sequence is then globally dependently modeled using a bidirectional subspace scanning module to obtain the EPI-W enhanced sequence. The EPI-W enhanced sequence is reshaped into a 4D feature tensor to obtain and output multi-scale fused features.
[0013] By employing the aforementioned techniques, EPI structural features and spatial-angular correlation features are merged to obtain the features to be fused. This integrates information from different dimensions, and then global dependency modeling across subspaces is achieved through bidirectional scanning of the EPI-H / EPI-W subspaces. This deeply complements and couples information from different dimensions. The geometric constraints of the EPI structure guide the alignment and fusion of spatial-angular features, ensuring that the fused features retain the regularity of the epipolar structure while containing rich spatial-angular details. This avoids feature misalignment or information dilution caused by cross-subspace fusion, which helps improve the effect of light field image super-resolution reconstruction, more accurately restores image details and textures, and improves the quality of the reconstructed image.
[0014] Preferably, the step of performing dual attention modulation of space and angle on the multi-scale fused features to obtain comprehensive enhanced features specifically includes the following steps: The multi-scale fusion features are rearranged into a subspace of the spatial dimension, and a spatial attention map is generated through a convolutional layer and a sigmoid activation function. Spatial attention modulation is then performed on the multi-scale fusion features based on the spatial attention map and element-wise multiplication to obtain spatially enhanced features. The spatial augmentation features are rearranged into a subspace of the angular dimension, and an angular attention map is generated through a convolutional layer and a sigmoid activation function. An angular attention modulation is then performed on the spatial augmentation features based on the angular attention map and element-wise multiplication to obtain the comprehensive augmentation features.
[0015] By employing the aforementioned techniques, attention modulation is applied to multi-scale fusion features at both spatial and angular levels. Based on the dynamic weighting principle of the attention mechanism, key regional details are highlighted in the spatial dimension and perspective consistency is strengthened in the angular dimension. This achieves precise screening and enhancement of multi-scale fusion features, further improving the discriminativeness and effectiveness of the features. It allows the subsequent diffusion model to focus on residual reconstruction guided by high-value features, solving the problem of redundant information in fusion features interfering with reconstruction accuracy or uneven feature distribution, and improving the overall representation capability.
[0016] Preferably, the step of reconstructing a predicted residual image under the dual constraints of the residual domain and frequency domain through stepwise denoising within a pre-built diffusion model framework and guided by the comprehensive enhancement features, and then compensating the predicted residual image to obtain a high-resolution light field image, specifically includes the following steps: The target residual image is acquired, and a noise coefficient is generated using a cosine scheduling function. Then, based on the noise coefficient, a preset number of forward diffusion steps are performed on the target residual image within a pre-built diffusion model framework. For each time step of the forward diffusion, a noisy residual image is generated according to a Gaussian distribution, and the noisy residual images corresponding to all time steps are summarized to obtain the forward diffusion sequence. The denoised residual image at the current time step is obtained from the forward diffusion sequence as the main input, and the integrated enhancement features are injected into the diffusion denoising network to obtain the predicted noise at the current time step. Based on the predicted noise at the current time step, reverse updates are performed step by step according to the preset reverse diffusion formula. At the same time, the consistency between the predicted residual and the calculated true residual in the residual domain and frequency domain is constrained by the preset global loss function until the time step is 0, and the predicted residual image is obtained. The predicted residual image is superimposed pixel by pixel with the result of bicubic upsampling of the original light field image to obtain a high-resolution light field image.
[0017] By employing the aforementioned techniques, a noise coefficient that conforms to the characteristics of the light field residual is generated using a cosine scheduling function, enabling the forward diffusion process to smoothly simulate residual degradation. Under the dual constraints of the residual domain and frequency domain, the diffusion model framework and comprehensive enhancement features guide the gradual denoising process, allowing the model to gradually strip away noise and approximate the true residual under feature guidance, thereby improving the ability to restore high-frequency details, ensuring cross-view consistency and global structural stability, suppressing blurring and artifact phenomena, and at the same time, leveraging the randomness of the diffusion model to ensure detail diversity. This solves the problem of detail blurring or over-smoothing that is prone to occur in traditional residual reconstruction, ultimately resulting in a high-resolution light field image.
[0018] Preferably, the step of constraining the consistency between the predicted residual and the calculated true residual in the residual domain and frequency domain by using a preset global loss function specifically includes the following steps: The high-resolution light field image in the training dataset corresponding to the original light field image is subtracted pixel-by-pixel from the result of bicubic upsampling of the original light field image to obtain the true residual. The global loss function includes a residual domain loss term and a frequency domain loss term, wherein the residual domain loss term is the L1 norm between the predicted residual image and the true residual; Fast Fourier Transform is performed on the predicted residual image and the true residual respectively to obtain the predicted frequency domain features and the true frequency domain features respectively. The frequency domain loss term is the L1 norm between the predicted frequency domain features and the true frequency domain features. The residual domain loss term and the frequency domain loss term are weighted by preset weight coefficients to obtain a global loss value. The parameters of the diffusion denoising network are updated by backpropagation algorithm with the goal of minimizing the global loss value.
[0019] By employing the aforementioned techniques, the true residual is obtained by subtracting the high-resolution light field image from the upsampled original light field image, which can accurately measure the difference between the predicted residual and the actual situation. Based on the principle that the L1 norm in the residual domain can constrain pixel-level errors and the L1 norm in the frequency domain can enhance the consistency of high-frequency components, a dual-domain collaborative global loss function is constructed. This allows the model to learn both the pixel-level matching of the residual and accurately fit the frequency domain distribution of high-frequency details, thus making up for the deficiency of single residual loss in constraining high-frequency details and ensuring that the reconstruction result achieves the best in terms of visual texture and detail fidelity.
[0020] Secondly, this application provides a light field image super-resolution reconstruction device, which adopts the following technical solution: A light field image super-resolution reconstruction device includes the following modules: The subspace feature extraction module is used to extract features from the received original light field image based on EPI-Mamba subspace scanning to obtain EPI structural features, and to extract features from the original light field image based on SA-Mamba subspace scanning to obtain space-angle correlation features. The cross-space multi-scale interactive fusion module is used to perform cross-subspace multi-scale interactive fusion based on the EPI structural features and the space-angle correlation features, and obtain multi-scale fusion features by modeling global dependencies through bidirectional scanning. A dual attention modulation module is used to perform spatial and angular dual attention modulation on the multi-scale fusion features to obtain comprehensive enhanced features; The diffusion residual reconstruction module is used to reconstruct the predicted residual image under the dual constraints of the residual domain and frequency domain by stepwise denoising under the guidance of the enhanced features within a pre-built diffusion model framework, and to compensate the predicted residual image to obtain a high-resolution light field image.
[0021] By adopting the above technical solutions, the subspace feature extraction module can efficiently extract the EPI structural features and spatial-angle correlation features of the light field using EPI-Mamba and SA-Mamba subspace scanning, respectively, thereby enhancing the global spatial-angle dependence; the cross-space multi-scale interactive fusion module can perform cross-subspace multi-scale interactive fusion of features from different subspaces and model global dependencies through bidirectional scanning, achieving deep-level information complementarity and coupling; the dual attention modulation module performs dual attention modulation of the multi-scale fused features in both space and angle, strengthening key dependencies and suppressing irrelevant or noisy information, thereby improving feature discriminability; the diffusion residual reconstruction module, guided by enhanced features, gradually denoises and reconstructs the predicted residual image under dual constraints in the residual domain and frequency domain, and compensates to obtain a high-resolution light field image, which can effectively improve the ability to restore high-frequency details, ensure cross-view consistency and global structural stability, suppress blurring and artifacts, and improve the visual quality of the image.
[0022] In summary, this application includes at least one of the following beneficial technical effects: (1) This application uses bidirectional scanning of EPI-Mamba and SA-Mamba subspaces to extract features, which can effectively capture the EPI structure and spatial-angle correlation of the light field, obtain stronger global spatial-angle dependence, achieve accurate separation of heterogeneous features, alleviate the difficulties caused by long-term dependence in long sequence modeling, and maintain the integrity of 4D global information; (2) This application utilizes a multi-scale interactive fusion module and a space-angle modulation module to fuse and modulate features, thereby achieving deep information complementarity enhancement and coupling, strengthening key features, suppressing noise and redundant information, and improving feature discriminability; (3) This application achieves accurate reconstruction of high-frequency residuals through a diffusion model with dual-domain constraints. Diffusion modeling is performed simultaneously in the residual domain and frequency domain. Guided by the global-local consistency features extracted by Mamba, the model focuses more on texture edges and high-frequency details during reconstruction, effectively alleviating blurring and artifact problems, improving the ability to restore high-frequency details, and ensuring cross-view consistency and global structural stability. Attached Figure Description
[0023] Figure 1 This is an overall structural diagram of the light field image super-resolution reconstruction network in this embodiment; Figure 2 This is the overall flowchart of the method in this embodiment; Figure 3 This is an internal structure diagram of the bidirectional subspace scanning module in the method of this embodiment; Figure 4 This is a structural diagram of EPI-Mamba in the method of this embodiment; Figure 5 This is a structural diagram of the space-angle modulation module in the method of this embodiment; Figure 6 This is a flowchart of the network structure for super-resolution reconstruction of light field images in this embodiment; Figure 7 This is a visual comparison of the 4× super-resolution results between the method of this embodiment and other methods; Figure 8 This embodiment describes the visual impact map of removing the MCI module in a 2× upsampling task of the ISO Chart 1 scenario in the ablation experiment. Figure 9 This embodiment describes the method for removing the visual impact map of the SAM module in a 2× upsampling task of the Tarot Cards S scene in the ablation experiment. Figure 10 This embodiment describes the method used to remove the visual impact map of the FRDiff module in the 2× upsampling task of the Sculpture_Decoded scene in the ablation experiment. Figure 11 This is an architectural diagram of a light field image super-resolution reconstruction device. Detailed Implementation
[0024] The technical solutions in the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments are only possible technical implementations of the present invention, but are not limited thereto. Other embodiments obtained by those skilled in the art in conjunction with the embodiments of the present invention without creative effort are also within the protection scope of the present invention.
[0025] This application mainly adopts the combination of Mamba subspace scanning and diffusion model to achieve super-resolution of light field images, thereby improving the resolution of light field images and the ability to restore high-frequency details. The following is a further detailed description of this application.
[0026] The light field image super-resolution reconstruction method provided in this application includes feature extraction, feature fusion, feature modulation, and image generation steps, such as... Figure 1 The diagram shows the overall framework of the reconstruction network. EPI structural features and spatial-angle correlation features of the light field image are extracted by EPI-Mamba subspace scanning and SA-Mamba subspace scanning, respectively. These features are input into the multi-scale interactive fusion module (MCI) for feature complementarity and coupling. Spatial-angle dual modulation (SAM) is applied to the fused features. Finally, the frequency domain enhanced residual diffusion generation module (FRDiff) is used for high-frequency compensation and combined with the diffusion process to gradually denoise and generate a high-resolution light field image, thereby improving the resolution of the light field image and the ability to restore high-frequency details.
[0027] S1. Based on the EPI-Mamba subspace scan, feature extraction is performed on the received original light field image to obtain EPI structural features, and based on the SA-Mamba subspace scan, feature extraction is performed on the original light field image to obtain spatial-angle correlation features. Specifically, the steps are as follows: S11. Receive a low-resolution original light field image, perform preliminary feature extraction on the original light field image to obtain initial features, and stitch the pixel features of the original light field image and the initial features channel by channel to obtain light field features. The light field features are specifically 4D features composed of spatial dimension (u, v) and angular dimension (h, w), denoted as I(u, v, h, w).
[0028] S12. Expand the light field features into the EPI-Mamba subspace and SA-Mamba subspace respectively for bidirectional scanning. The EPI-Mamba subspace includes the EPI-H subspace and the EPI-W subspace.
[0029] Specifically, S121. In the EPI-Mamba subspace, fix the angular dimension (h or w) of the light field features in the EPI-H and EPI-W subspaces, and expand them along the spatial dimension (u or v) to obtain horizontal EPI slices and vertical EPI slices.
[0030] If we fix the angular dimension w, for example, let w=w0, where w0 is a fixed viewpoint, then the light field data is reduced from 4D to 3D. Then, we slice along the spatial dimension v, fixing v=v0, and finally obtain a 2D horizontal EPI slice EPI_H(u,h)=I(u,v0,h,w0). The horizontal EPI slice is the epipolar image formed by the horizontal dimension u and the angular dimension h in space, which reflects the light distribution under different viewpoints in the same spatial column.
[0031] If we fix the angular dimension h, for example, let h=h0, and slice along the spatial dimension u, and fix u=u0, we get a 2D vertical EPI slice EPI_W(v,w)=I(u0,v,h0,w).
[0032] S122. Flatten the obtained slices into a one-dimensional sequence. Flatten the horizontal EPI slices into a first one-dimensional sequence with dimensions BVW×UH×C; flatten the vertical EPI slices into a second one-dimensional sequence with dimensions BUH×VW×C.
[0033] S123. Input the first sequence and the second sequence into independent bidirectional subspace scanning modules (BiSS) for bidirectional subspace scanning; like Figure 3 As shown, the BiSS module sequentially performs LayerNorm normalization, forward-backward bidirectional scan modeling, secondary LayerNorm normalization, and channel attention (CA) interaction to achieve global dependency modeling across tokens.
[0034] S124, such as Figure 4 As shown, a convolutional layer is connected in series after the bidirectional subspace scanning module to enhance the features of the sequence after bidirectional dependency modeling. The bidirectional subspace scanning module and convolutional layer corresponding to the horizontal EPI slice and the vertical EPI slice share weights to transmit implicit information along the light field EPI line and improve network efficiency. By enhancing local features through convolutional layers, the features output by the convolutional layers are reshaped into 4D feature tensors, resulting in EPI structural features. EPI structural features are used to represent the interaction between space and angle.
[0035] S13. In the SA-Mamba subspace, expand along the spatial dimension (u or v) or the angular dimension (h or w) into a spatial sequence of dimension B×UV×HW×C and an angular sequence of dimension B×HW×UV×C.
[0036] In this embodiment, it is only necessary to replace the horizontal and vertical sequences of EPI used in EPI-Mamba with spatial sequences, respectively. and angle sequence .
[0037] S14. After flattening the obtained spatial sequence and angle sequence into a one-dimensional sequence, input them into another set of independent bidirectional subspace scanning modules to perform bidirectional subspace scanning. The BiSS module performs forward and backward bidirectional dependency modeling on the one-dimensional sequence, and strengthens local features through convolutional layers to capture image texture details and cross-view consistency information, thereby obtaining spatial-angle correlation features that characterize the association between different viewpoints and spatial locations.
[0038] The bidirectional subspace scanning module (BiSS) of the EPI-Mamba subspace and the SA-Mamba subspace does not share parameters, and focuses on the extraction of specific features of epipolar structure and space-angle correlation, respectively.
[0039] S2. Based on the EPI structural features and spatial-angle correlation features, perform multi-scale interactive fusion across subspaces. Global dependencies are modeled through bidirectional scanning to obtain multi-scale fusion features. This includes the following steps: S21. Merge the EPI structural features and spatial-angle correlation features along the channel dimension to obtain the features to be fused.
[0040] The feature to be fused is the concatenation result of the EPI structural features output by EPI-Mamba and the spatial-angular correlation features output by SA-Mamba, with dimensions of B×U×V×H×W×C, where B is the batch size, U and V are the spatial dimensions, H and W are the angular dimensions, and C is the total number of channels for the two types of features.
[0041] S22. Perform the first-way fusion on the features to be fused: expand the features to be fused into the first token sequence of the EPI-H subspace, input it into the bidirectional subspace scanning module (BiSS) for global dependency modeling, and output the EPI-H enhanced sequence through the convolutional layer.
[0042] The internal structure of the bidirectional scanning module in the EPI-Mamba subspace satisfies the following: two bidirectional subspace scanning (BiSS) modules are executed separately in the EPI-H and EPI-W subspaces, and the corresponding convolutional layers are formed after each scan. A weight-sharing design is adopted, allowing the two BiSS modules and the convolutional layers to share weights. Through weight sharing, the implicit information of linear features in the light field EPI slices is transmitted along the EPI linear direction, and the number of network parameters is reduced to improve operating efficiency.
[0043] The EPI-H subspace is a horizontal polar feature space formed by a fixed angular dimension w, a fixed spatial dimension v, and a spatial dimension u and an angular dimension h.
[0044] Specifically, assuming the current input feature to be fused is the first... The output of the space-angle modulation module First, expand it into the first token sequence of EPI-H. .
[0045] The BiSS module and convolutional layer operations are performed on the first token sequence, and then the residual of the first token sequence is concatenated to calculate the EPI-H enhanced sequence. The formula is: ; in, This is the first token sequence. For convolution operations, This is an EPI-H enhanced sequence; S23. Perform second-path fusion on the features to be fused: assemble the EPI-H enhanced sequence. The second token sequence reshaped into the EPI-W subspace The second token sequence is composed of It was achieved through reconstructive surgery.
[0046] The second token sequence is input into the same dedicated Bidirectional Subspace Scanning (BiSS) module for fusion. The BiSS module and convolutional layers perform global dependency modeling, and the sequence is concatenated with the residuals of the second token sequence. The resulting EPI-W enhanced sequence is then output through the convolutional layers. The formula is: ; in, This is the second token sequence.
[0047] S24. Enhance the EPI-W sequence Reconstructing it into a 4D feature tensor of dimensions B×U×V×H×W×C, achieving cross-subspace deep information complementarity and coupling between EPI structural features and spatial-angle correlation features, resulting in multi-scale fused features. , and will As the first The output of the Multi-Scale Fusion Feature (MCI) module.
[0048] In this step, the bidirectional subspace scanning module does not share parameters with the BiSS module used for feature extraction in step S1, and is only used for cross-subspace fusion modeling.
[0049] The reshaped 4D tensor fully preserves the multidimensional correlation of light field space, angle, and EPI, avoiding the loss of dimensional information caused by sequence morphology, and laying the foundation for cross-view consistency modeling and high-frequency detail recovery in subsequent frequency domain enhancement (FFT) and diffusion reconstruction (FRDiff) modules.
[0050] Ultimately, the combination of EPI-Mamba and SA-Mamba enables the full fusion of complementary information from the spatial-angular dimension and the EPI dimension, thereby achieving efficient 4D global relation modeling.
[0051] S3. After completing the cross-dimensional interactive fusion of EPI-Mamba and SA-Mamba features, the resulting global representation may still contain redundancy or imbalanced feature distribution. To further improve the discriminativeness and effectiveness of the features, the fused features are subjected to dual calibration in both spatial and angular dimensions. Specifically, the multi-scale fused features are input into the Space-Angle Modulation (SAM) module to strengthen key dependencies and suppress irrelevant or noisy information.
[0052] This involves performing dual attention modulation (spatial and angular) on multi-scale fused features to obtain comprehensive enhanced features, specifically including the following steps: S31, Fusing features across multiple scales Rearranged to a subspace of the spatial dimension.
[0053] Rearrangement to a spatial dimension subspace refers to fixing the angular dimensions H and W of the multi-scale fusion features, retaining the spatial dimensions U and V and the channel dimension C, and converting the multi-scale fusion features from a 4D tensor form of B×U×V×H×W×C into a spatial dimension feature tensor of B×U×V×C to adapt to the input format of subsequent spatial attention modulation.
[0054] S32, such as Figure 4 The diagram shows the structure of the SAM module. It uses convolutional layers to perform dimensionality transformation and detail extraction on spatial features, resulting in a spatial feature map. The Sigmoid activation function is then used to compress the numerical range of the spatial feature map to the [0,1] interval, generating a spatial attention map. ,Right now, .
[0055] In the attention map, the closer the value is to 1, the more prominent the key features of the corresponding spatial region, such as edges and textures, are. Positions with values close to 0 are redundant or noisy regions.
[0056] S33. Multi-scale fusion features based on spatial attention maps and element-wise multiplication. Spatial attention modulation is performed to obtain spatial enhancement features, which enhance key regions and suppress redundant regions of the spatial features. ; in, For spatial enhancement features, it represents the result after adaptive weighting of spatial dimensions.
[0057] S34. Enhance spatial features Dimensional rearrangement is performed to the subspace of the angle dimension. Spatial dimensions U and V are fixed, while angle dimensions H and W and channel C are retained. The spatial enhancement features are converted from B×U×V×C to angle dimension feature tensors (B×H×W×C), thus completing the dimensional adaptation to the angle subspace.
[0058] S35. Following the same operations as in the spatial dimension, generate an angular attention map using convolutional layers and a sigmoid activation function. ,Right now, ; S36. Based on the angle attention map and element-wise multiplication, angle attention modulation is performed on the spatial augmentation features to obtain the comprehensive augmentation features, i.e.,
[0059] ; in, To comprehensively enhance features.
[0060] Through the steps described above, the SAM module can apply attention modulation at both the spatial and angular levels, thereby further enhancing the feature representation of key regions and improving the overall representation capability.
[0061] The entire process employs a design that separates and modulates spatial and angular dimensions to optimize feature quality for the two core requirements of spatial detail integrity and angular viewpoint consistency in light field images. This provides crucial support for detail restoration and cross-viewpoint consistency assurance in subsequent super-resolution reconstruction.
[0062] It should be noted that if the current process is iterative feature processing, the resulting comprehensive enhanced features, as the output of the i-th SAM, can be used as the input for the next round of multi-scale interactive fusion (MCI module).
[0063] S4. Under the pre-built diffusion model framework, guided by enhanced features, the predicted residual image is reconstructed under the dual constraints of the residual domain and frequency domain through stepwise denoising. The predicted residual image is then compensated to obtain a high-resolution light field image. The specific steps include the following: S41. Establish a diffusion model framework, including forward diffusion and backward diffusion.
[0064] During forward diffusion, the acquired target residual image needs to be gradually denoised according to a preset number of diffusion steps. Forward diffusion is the data preparation stage of the diffusion model. Its core is to gradually denoise the target residual image according to the set rules, providing labeled training samples for subsequent inverse denoising.
[0065] During the reverse diffusion process, the UNet denoising network, combined with multi-dimensional conditions, is used to infer the real noise from the noisy image and gradually denoise it to restore the noise-free residual image.
[0066] S41. Forward Diffusion Process: First, the target residual image is obtained from the preset training dataset. The target residual image is the difference between the high-resolution light field image and the upsampled low-resolution light field image, i.e., ; in, The target residual image contains only the details missing after low-resolution upsampling, such as edges and high-frequency textures, and is the core content that the diffusion model needs to learn to generate. This is the original high-resolution light field image. It is a low-resolution light field.
[0067] In one specific implementation, the high-resolution light field image is downsampled (e.g., 4× downsampling) to obtain a low-resolution light field, and then the low-resolution light field is upsampled (e.g., bicubic upsampling). The difference between the low-resolution light field and the original high-resolution light field is used to obtain the target residual, which is then used for training the diffusion model.
[0068] S42. Under the diffusion model framework, perform forward diffusion of the target residual image with a preset number of diffusion steps based on the calculated cumulative noise coefficient.
[0069] The specific steps for generating the noise coefficients are as follows: ; in, To preset the number of diffusion steps, , For the first The noise coefficients of the first step are generated by the cosine scheduling function, representing the noise coefficients of the second step. When noise is added to a step separately, the proportion of original information retained in the signal; For the first The cumulative noise coefficient of the forward diffusion step represents the noise level from the initial noiseless state to the nth step. The proportion of the signal that is not contaminated by noise during the step.
[0070] S43. For each time step of the forward diffusion, generate a noisy residual image according to a Gaussian distribution. Summarize the noisy residual images corresponding to all time steps to obtain the forward diffusion sequence, i.e., ,in For the target residual image, For the first Step-by-step noise-added residual image; The forward diffusion process follows a fixed Gaussian distribution. generate ; Forward diffusion sequence representation based on known initial target residual image Under the conditions, the first Step-by-step noise-added residual image The conditional probability distribution; As an identity matrix, it ensures that noise is independently distributed across each pixel dimension, and so on. Increase As the noise level gradually approaches 1, the noise intensity gradually increases, eventually reaching t=T. Approximately obey The pure noise distribution.
[0071] The generation of the diffusion sequence follows the principle that noise intensity increases with... The increasing pattern, that is The larger the residual image, the noisier it becomes. The higher the proportion of noise, The lower the information content, the more crucial this characteristic becomes for the gradual denoising of the reverse process.
[0072] S44. Reverse Diffusion Process: The diffusion denoising network uses the noisy residual image of the current time step obtained from the forward diffusion sequence as the main input, and integrates the enhanced features... As conditional features, these features are injected into each layer of the diffusion denoising network UNet through feature concatenation or cross-attention mechanisms, thereby introducing both frequency domain and spatial-angular conditions into the denoising process to obtain the predicted noise at the current time step. In this embodiment, time coding is incorporated into the network to guide noise suppression and detail recovery strategies at different time steps.
[0073] Specifically, the predicted noise is, ; in, This represents the noise predicted by the diffusion denoising network. Indicates the first step in the diffusion process The noisy residual image of the step, derived from the target residual image The result was obtained by gradually adding noise. For the number of diffusion steps, The parameterized UNet diffusion denoising network is essentially a network that learns from noisy residual images. To real noise The mapping function, i.e., the result of progressively adding noise to the target residual, is denoted by the network parameter set. During training, the noise is updated by minimizing the difference between the predicted noise and the actual noise. For time step The embedding maps discrete diffusion steps to continuous vector representations, which is achieved through positional encoding, enabling the network to learn denoising patterns related to the diffusion stage.
[0074] S45. Based on the predicted noise at the current time step, perform reverse updates step by step according to the preset reverse diffusion formula.
[0075] The reverse diffusion process from a pure noise image Starting at (t=T), the noise at each step is predicted using a diffusion denoising network. and based on generate (Noise intensity decreases), repeat the iteration until t=1, and the result is recovered. This is the final predicted residual image.
[0076] Specifically, the reverse diffusion formula is: ; in, It follows a standard Gaussian distribution. random noise, It is the noise standard deviation related to the diffusion time step. Determined by the cosine scheduling strategy.
[0077] S46. Simultaneously, the consistency between the predicted residual and the calculated true residual in the residual domain and frequency domain is constrained by a preset global loss function until the time step is 0, thus obtaining the predicted residual image.
[0078] S461. Take the high-resolution light field image corresponding to the original light field image in the preset training dataset and perform pixel-by-pixel subtraction operation on the result of the original light field image after bicubic upsampling to obtain the true residual.
[0079] This process is related to calculating the target residual image. The process is the same. For each training sample (HR, LR), LR is first upsampled to the resolution of HR by bicubic upsampling, and the difference is directly calculated as the true residual of the sample, which remains unchanged throughout the training process.
[0080] The true residual and the target residual are essentially the same concept. In the forward diffusion process, it serves as the starting point for adding noise to the diffusion model, and in the reverse diffusion process, it serves as the benchmark for supervised optimization.
[0081] S462. The global loss function includes a residual domain loss term and a frequency domain loss term. The residual domain loss term is the L1 norm between the predicted residual image and the true residual. The L1 norm difference between the true and predicted residuals in the spatial domain is calculated, forcing the model-predicted residuals to conform to the spatial structure of the true residuals. This ensures the accuracy of macroscopic spatial details in the predicted residuals, such as the position of object edges and the distribution of textures, avoiding the problem of the predicted residuals deviating from the spatial structure of the true residuals.
[0082] S463. Perform Fast Fourier Transform on the predicted residual image and the true residual respectively to obtain the predicted frequency domain features. and true frequency domain characteristics The frequency domain loss term is the L1 norm between the predicted and actual frequency domain features. After transforming the residuals from the spatial domain to the frequency domain using FFT, the difference in the L1 norm between the actual and predicted residual frequency domain features is calculated, forcing the model to match the frequency domain characteristics of the actual residuals.
[0083] It can ensure that the high-frequency components of the predicted residual are consistent with the true residual, avoiding the problem of blurred high-frequency details, such as blurred object edges and lost textures, when reconstructing traditional diffusion models.
[0084] The global loss value is obtained by weighting the residual domain loss term and the frequency domain loss term by pre-setting weighting coefficients.
[0085] The global loss function is: ; in, This is the global loss value, which is the objective function that needs to be minimized during model training; It is a true residual image, containing the true detail information that was lost after low-resolution upsampling; The residual image predicted by the model, i.e., the final output of the inverse generation process of the diffusion denoising network (UNet), is... The goal is to approximate the result as closely as possible. Consistent; For the Fast Fourier Transform (FFT) operation, The frequency domain characteristics of the true residual. To predict the frequency domain characteristics of the residuals.
[0086] L1 norm loss is used to calculate the difference between two features (spatial domain or frequency domain).
[0087] and These are the weighting coefficients for the loss, all positive numbers, used to balance the contributions of spatial and frequency domain losses. They need to be adaptively adjusted according to the texture complexity of the light field image; for example, they can be increased when emphasizing high-frequency details. .
[0088] Used to control the proportion of space loss in the total loss, if If the value is too large, the model will focus excessively on spatial domain consistency and may ignore high-frequency details in the frequency domain; if the value is too small, spatial structure consistency will be difficult to guarantee.
[0089] The contribution of frequency domain loss is used to enhance high-frequency details, therefore Typically, a reasonable value (non-zero and large enough) needs to be set to ensure that the consistency of frequency domain features is fully learned by the model.
[0090] S464. With the goal of minimizing the global loss, update the parameters of the diffusion denoising network using the backpropagation algorithm.
[0091] S47. The predicted residual image and the original light field image after bicubic upsampling are superimposed pixel by pixel to obtain the high-resolution light field image ISR, as shown. Figure 6 This is a flowchart of the network structure of this method.
[0092] The core of this step, the Frequency Domain Enhanced Residual Diffusion Generation Module (FRDiff), is to achieve high-resolution light field image reconstruction in the residual domain through a diffusion model under the guidance of the frequency domain enhancement mechanism, so that the model pays more attention to texture edges and high-frequency details during the reconstruction process.
[0093] By modeling only the residuals, the generation difficulty is reduced, accelerating training convergence. Simultaneously, spatial and frequency domain constraints enhance the ability to reproduce high-frequency details. Global-local consistency features extracted by Mamba are used as diffusion conditions to ensure cross-view consistency and global structural stability. After obtaining enhanced features through the aforementioned feature extraction and multi-scale frequency domain fusion modules, this invention performs progressive denoising and reconstruction within the diffusion model framework to achieve the gradual generation of high-quality, high-resolution light fields from a high-noise initial input.
[0094] To test the usability of the method in this application, in one specific implementation, training and testing were performed using five publicly available training datasets widely used in current light field super-resolution research.
[0095] In the five training sets, HCInew and HCIold are synthetic datasets, while EPFL, INRIA, and STFgantry are real datasets.
[0096] The input LF image is converted to the YCbCr color space, and only the Y channel of the super-resolution image is used. The Cb and Cr channel images are bicubic upsampled.
[0097] Light field images with a center resolution of 5×5 and an angular resolution of 9×9 were extracted from these datasets for training and testing. Specifically, for ease of training, the light field images were cropped into light field image patches with a size of 128×128 and a stride of 64, and 4×bicubic downsampling was performed to obtain light field image patches for the 4×SR experiment.
[0098] The model was pre-trained using the Adam optimizer, with a batch size of 4 and a learning rate of 2×10⁻⁶. -4The number of iterations is halved every 100k iterations. This paper performs data augmentation on the cropped light field image patches using horizontal flipping (×2), vertical flipping (×2), and 90° rotation (×3) to enable the model to learn richer texture and structural features, improving its adaptability in different scenes. The model was trained using an NVIDIA RTX 4090 GPU for 200,000 iterations. The examples use Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM) as quantitative evaluation metrics, with a focus on the quality performance of the Y channel.
[0099] This embodiment conducts experiments on multiple light field training datasets to compare the quantitative results of different super-resolution reconstruction methods at 2× and 4× magnifications. The quantitative results are shown in Table 1: Table 1. Quantitative results of different methods on 2× and 4× super-resolution tasks.
[0100] As shown in Table 1, Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM) were used as evaluation metrics. The experiments covered multiple light field datasets, comprehensively reflecting the performance of each method under different scenes and parallax conditions, and verifying the stability and generalization ability of the proposed method.
[0101] The results show that the proposed method (OURS) achieves best or near-best performance on multiple datasets, exhibiting particularly strong robustness in scenes with large parallax and complex structures. Compared to existing mainstream methods, the proposed model maintains high PSNR and SSIM scores even at more complex 4× magnification levels, demonstrating excellent overall performance and showcasing its application potential in light field super-resolution.
[0102] In another specific implementation, to more comprehensively evaluate the method proposed in this application, this embodiment also utilizes multiple different training datasets for qualitative analysis. By displaying and comparing the reconstruction results under different datasets and magnification levels, the advantages of the method in this embodiment in detail restoration, texture reconstruction, and complex scene processing are intuitively demonstrated.
[0103] The method in this embodiment also performs exceptionally well in restoring complex textures and details in more challenging 4× super-resolution tasks. Figure 7 As shown, this embodiment compares the reconstruction results of different methods in the Hublas Decoded scene of the INRIA dataset and the Bicycle scene of the HCInew dataset. To more clearly compare the differences in detail recovery of each method, local magnified images are used to highlight the texture and edge information of key areas, so as to intuitively evaluate the performance of different methods in the reconstruction of complex structures and fine textures.
[0104] Depend on Figure 7 As can be seen, in the Hublas Decoded scene of the INRIA dataset, the method proposed in this embodiment performs outstandingly in the recovery of complex textures and details, accurately reconstructing fine structures such as characters and effectively avoiding edge blurring and detail loss. This result fully demonstrates the advantages of the proposed model in high-frequency information reconstruction, making the super-resolution image visually closer to the original high-resolution image. In contrast, methods such as VDSR, EDSR, RCAN, and resLF generally exhibit blurring and artifacts in the reconstruction results, especially at character edges and in complex texture areas, where incomplete detail recovery leads to decreased image sharpness and insufficient realism.
[0105] In the Bicycle scene of the HCInew dataset, the method proposed in this embodiment also demonstrates excellent detail restoration capabilities, stably preserving the texture structure of objects. Especially in areas with rich edge and texture features, such as telescopes and books, the method in this embodiment accurately reconstructs their clear outlines and surface details. Other methods, however, perform poorly in this scene, often resulting in reconstructed images with blurred edges, texture distortion, or missing details, leading to a decline in overall visual quality and difficulty in reproducing the true structure of the scene.
[0106] The qualitative analysis results above demonstrate that the method described in this embodiment can accurately reproduce high-frequency details and complex structures in multiple scenarios, significantly improving the visual quality of images. Compared to traditional single-image super-resolution methods and existing light field super-resolution methods, the model has advantages in detail preservation and image sharpness, while effectively suppressing blurring and artifacts.
[0107] In another specific implementation, the performance of various methods in terms of computational overhead and reconstruction quality was compared in 2× and 4× super-resolution tasks.
[0108] The metrics include parameter size (Params), floating-point operations (FLOPs), and PSNR / SSIM.
[0109] All experiments were conducted using the PyTorch framework on an AMD EPYC 9754 (128 cores, 2.25GHz), 512GB of RAM, and an NVIDIA GeForce RTX 4090 D GPU. Table 2 below summarizes the specific comparative data, where computational cost is based on an input light field image patch size of 5×5×32×32, and PSNR and SSIM are the average values of the five test datasets.
[0110] Table 2. Parameter quantity, computational cost, and average PSNR / SSIM value of different methods in 2× and 4× super-resolution tasks.
[0111] In the 2× super-resolution experiment, the proposed method has only 2.08M parameters and a computational cost of 503.8G FLOPs, while achieving the highest average PSNR / SSIM (39.43dB / 0.987) among the comparative models.
[0112] Under the 4× task, the parameter size increased to 2.49M and the complexity was 548.5G FLOPs, but the best average PSNR / SSIM (33.70dB / 0.945) was still achieved.
[0113] In another specific implementation, this embodiment also conducted ablation experiments to analyze the effectiveness of each module.
[0114] A. Validity of the MCI module: To verify the effectiveness of the Multi-Scale Interactive Fusion (MCI) module, this module was removed while keeping the spatial-angle modulation module and frequency domain enhanced residual diffusion generation module unchanged. The features extracted by EPI-Mamba and SA-Mamba were simply concatenated and then input into the subsequent network.
[0115] Experimental results show that, in the absence of MCI, the model exhibits a significant decrease in both PSNR and SSIM metrics, as shown in Table 3.
[0116] Table 3. Quantitative results of network variants with MCI modules removed in the 2× upsampling task.
[0117] Table 3 above illustrates that MCI plays a key role in promoting deep interaction and complementarity between spatial and angular dual-dimensional features, effectively avoiding model bias towards single-dimensional features, thereby enhancing the discriminativeness and robustness of features.
[0118] like Figure 8 The image shows the visual effect of removing MCI in a 2× upsampling task in the ISO Chart1 scene. Figure 8 Subjective visual results further demonstrate that the reconstruction after removing MCI suffers from lost texture details and blurred edges. This verifies the importance of this module in high-frequency information recovery and cross-viewpoint consistency maintenance.
[0119] B. Validity of the SAM module: To evaluate the effectiveness of the Space-Angle Modulation (SAM) module, this module was removed from the complete model, and only the features fused by MCI were directly input into the subsequent network for reconstruction. The experimental results are shown in Table 4.
[0120] Table 4. Quantitative results of network variants with SAM module removed in the 2× upsampling task.
[0121] Table 4 shows that the model lacking SAM has decreased PSNR and SSIM, especially in complex scenes where it shows significant performance degradation. Removing SAM leads to weakened cross-view consistency and deviations in the geometric structure of the reconstructed image.
[0122] like Figure 9 The image shows the removal of the visual effects of SAM in a 2× upsampling task in the Tarot Cards S scene. Figure 9 The subjective visual results also show that without SAM, the reconstruction results are prone to noise enhancement and texture aliasing problems.
[0123] This indicates that the SAM module plays a key role in enhancing key features and suppressing redundant information through a dual spatial and angular attention mechanism, thereby improving the model's performance in detail recovery and cross-view consistency.
[0124] C. Validity of the FRDiff module: With the FRDiff module removed, the features output by the SAM module are upsampled by performing a pixel shuffle operation. Then, the upsampled results are regressed pixel by pixel using regular residual learning to compensate for some missing high-frequency information.
[0125] While this approach can restore the basic structural outline, it lacks the diverse generation capabilities offered by diffusion modeling and cannot explicitly compensate for high-frequency components through frequency domain constraints. Specific indicators are shown in Table 5.
[0126] Table 5. Quantitative results of network variants with the FRDiff module removed in the 2× upsampling task.
[0127] like Figure 10 The image shows the visual effects of removing FRDiff in the 2× upsampling task of the Sculpture_Decoded scene. Therefore, the visual effect shows that it is significantly lacking in detail reproduction, texture clarity and cross-view consistency.
[0128] Experimental results show that the proposed light field image super-resolution reconstruction network exhibits superior performance on publicly available real-world light field datasets. In the 4× super-resolution reconstruction scenario, it achieves the highest PSNR scores on the EPFL, HCInew, INRIA, and STFgantry datasets, and the highest SSIM scores on HCInew and STFgantry datasets. Furthermore, its average scores for both metrics are the highest compared to other methods. Qualitative analysis also demonstrates good visual performance, accurately reproducing high-frequency details and complex structures in multiple scenes, significantly improving image visual quality. Compared to traditional single-image super-resolution methods and existing light field super-resolution methods, the model demonstrates superior detail preservation and image sharpness while effectively suppressing blurring and artifacts.
[0129] Based on the same inventive concept described above, this application also discloses a light field image super-resolution reconstruction device, the structure of which is as follows: Figure 11 As shown, the device includes the following modules: The subspace feature extraction module is used to extract features from the received original light field image based on EPI-Mamba subspace scanning to obtain EPI structural features, and to extract features from the original light field image based on SA-Mamba subspace scanning to obtain space-angle correlation features. The cross-space multi-scale interactive fusion module is used to perform cross-subspace multi-scale interactive fusion based on EPI structural features and spatial-angle correlation features. It obtains multi-scale fusion features by modeling global dependencies through bidirectional scanning. The dual attention modulation module is used to perform spatial and angular dual attention modulation on multi-scale fused features to obtain comprehensive enhanced features; The diffusion residual reconstruction module is used to reconstruct the predicted residual image under the dual constraints of the residual domain and frequency domain by stepwise denoising under the guidance of enhanced features within a pre-built diffusion model framework. The predicted residual image is then compensated to obtain a high-resolution light field image.
[0130] In one specific implementation scheme, the subspace feature extraction module includes the following units: The first subspace feature extraction unit is used to receive the low-resolution original light field image, perform preliminary feature extraction on the original light field image to obtain initial features, and fuse the original light field image and the initial features to obtain light field features. The light field features are specifically 4D features composed of spatial dimension (u,v) and angular dimension (h,w), denoted as I(u,v,h,w). The second subspace feature extraction unit is used to unfold the light field features into the EPI-Mamba subspace and SA-Mamba subspace for bidirectional scanning. The EPI-Mamba subspace includes the EPI-H subspace and the EPI-W subspace. Specifically, In the EPI-Mamba subspace, the light field features are fixed in the angle dimension in the EPI-H and EPI-W subspaces. They are then unfolded along the spatial dimension to obtain horizontal and vertical EPI slices. After flattening the slices into a one-dimensional sequence, bidirectional subspace scanning and bidirectional dependency modeling are performed to obtain EPI structural features. These EPI structural features are used to represent the interaction between space and angle. In the SA-Mamba subspace, scanning along the spatial or angular dimensions yields spatial and angular sequences. After flattening the obtained sequences into one-dimensional sequences, bidirectional subspace scanning is performed and bidirectional dependency modeling is executed to obtain spatial-angular correlation features that characterize the association between different viewpoints and spatial locations.
[0131] In one specific implementation scheme, the second subspace feature extraction unit includes the following subunits: The first subspace feature extraction subunit is used to flatten the horizontal EPI slice into a one-dimensional first sequence and the vertical EPI slice into a one-dimensional second sequence. The second subspace feature extraction subunit is used to input the first sequence and the second sequence into the bidirectional subspace scanning module respectively, and sequentially perform LayerNorm normalization processing, forward-backward bidirectional scanning modeling, secondary LayerNorm normalization processing, and channel attention interaction processing. The third subspace feature extraction subunit is used to connect convolutional layers after the bidirectional subspace scanning module to enhance the features of the sequence after bidirectional dependency modeling. The bidirectional subspace scanning modules and convolutional layers corresponding to the horizontal EPI slice and the vertical EPI slice share weights. The features output by the convolutional layer are reshaped into 4D feature tensors to obtain EPI structural features.
[0132] In a specific feasible implementation, the cross-space multi-scale interactive fusion module includes the following units: The first cross-space multi-scale interactive fusion unit is used to merge EPI structural features and spatial-angle correlation features to obtain the features to be fused. The second cross-space multi-scale interactive fusion unit is used to expand the features to be fused into the first token sequence of the EPI-H subspace, and to perform global dependency modeling on the first token sequence through the bidirectional subspace scanning module to obtain the EPI-H enhanced sequence. The third cross-space multi-scale interactive fusion unit is used to reshape the EPI-H enhanced sequence into the second token sequence of the EPI-W subspace. The second token sequence is then globally dependently modeled by the bidirectional subspace scanning module to obtain the EPI-W enhanced sequence. The fourth cross-space multi-scale interactive fusion unit is used to reshape the EPI-W enhanced sequence into a 4D feature tensor, and obtain and output multi-scale fused features.
[0133] In one specific implementation, the dual attention modulation module includes the following units: The first dual attention modulation unit is used to rearrange the multi-scale fused features to a subspace of the spatial dimension. It generates a spatial attention map through a convolutional layer and a sigmoid activation function. Based on the spatial attention map and element-wise multiplication, it performs spatial attention modulation on the multi-scale fused features to obtain spatially enhanced features. The second dual attention modulation unit is used to rearrange the spatial augmentation features to a subspace of the angular dimension. An angular attention map is generated through a convolutional layer and a sigmoid activation function. An angular attention modulation is then performed on the spatial augmentation features based on the angular attention map and element-wise multiplication to obtain the comprehensive augmentation features.
[0134] In one specific implementation scheme, the diffusion residual reconstruction module includes the following units: The first diffusion residual reconstruction unit is used to acquire the target residual image, generate noise coefficients using a cosine scheduling function, and perform forward diffusion of the target residual image with a preset number of diffusion steps based on the noise coefficients under the pre-built diffusion model framework. The second diffusion residual reconstruction unit is used to generate a noisy residual image according to a Gaussian distribution for each time step of the forward diffusion, and to summarize the noisy residual images corresponding to all time steps to obtain the forward diffusion sequence. The third diffusion residual reconstruction unit is used to obtain the noisy residual image of the current time step from the forward diffusion sequence as the main input, and inject the comprehensive enhancement features into the diffusion denoising network to obtain the predicted noise of the current time step. The fourth diffusion residual reconstruction unit is used to perform reverse updates step by step according to the preset reverse diffusion formula based on the prediction noise at the current time step. At the same time, the preset global loss function constrains the consistency between the predicted residual and the calculated real residual in the residual domain and frequency domain until the time step is 0, and the predicted residual image is obtained. The fifth diffusion residual reconstruction unit is used to superimpose the predicted residual image and the original light field image after bicubic upsampling pixel by pixel to obtain a high-resolution light field image.
[0135] In one specific implementation scheme, the fourth diffusion residual reconstruction unit includes the following sub-units: The first diffusion residual reconstruction subunit is used to perform pixel-by-pixel subtraction between the high-resolution light field image in the training dataset that corresponds to the original light field image and the result of the original light field image after bicubic upsampling, to obtain the true residual. The global loss function includes a residual domain loss term and a frequency domain loss term. The residual domain loss term is the L1 norm between the predicted residual image and the true residual. The second diffusion residual reconstruction subunit is used to perform fast Fourier transform on the predicted residual image and the real residual respectively to obtain the predicted frequency domain features and the real frequency domain features respectively. The frequency domain loss term is the L1 norm between the predicted frequency domain features and the real frequency domain features. The third diffusion residual reconstruction subunit is used to weight the residual domain loss term and the frequency domain loss term with preset weight coefficients to obtain the global loss value. With the goal of minimizing the global loss value, the parameters of the diffusion denoising network are updated through the backpropagation algorithm.
[0136] Based on the same inventive concept described above, this application also discloses a smart terminal, including a memory and a processor. The memory stores at least one instruction, at least one program, code set, or instruction set. The processor loads and executes the at least one instruction, at least one program, code set, or instruction set to implement the light field image super-resolution reconstruction method described above.
[0137] Based on the same inventive concept described above, this application also discloses a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set. The at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the light field image super-resolution reconstruction method described above.
[0138] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware, or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, and other media capable of storing program code.
[0139] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method of light field image super-resolution reconstruction, characterized in that, The method comprises the following steps: According to the EPI-Mamba subspace scanning, the original light field image is extracted to obtain the EPI structure feature, and according to the SA-Mamba subspace scanning, the original light field image is extracted to obtain the space-angle correlation feature; According to the EPI structure feature and the space-angle correlation feature, a multi-scale interactive fusion across the subspace is performed, a global dependence relationship is modeled through bidirectional scanning, and a multi-scale fusion feature is obtained; A spatial and angular double attention modulation is performed on the multi-scale fusion feature to obtain a comprehensive enhanced feature; Under the framework of a pre-built diffusion model, the enhanced feature is taken as a guide condition, a predicted residual image is reconstructed under the dual constraints of the residual domain and the frequency domain through step-by-step denoising, and a high-resolution light field image is obtained through compensation.
2. The light field image super-resolution reconstruction method of claim 1, wherein, According to the EPI-Mamba subspace scanning, the original light field image is extracted to obtain the EPI structure feature, and according to the SA-Mamba subspace scanning, the original light field image is extracted to obtain the space-angle correlation feature, specifically comprising the following steps: A low-resolution original light field image is received, preliminary feature extraction is performed on the original light field image to obtain an initial feature, the original light field image and the initial feature are fused to obtain a light field feature, and the light field feature is specifically a 4D feature composed of spatial dimensions (u, v) and angle dimensions (h, w), denoted as I(u, v, h, w); The light field feature is expanded to the EPI-Mamba subspace and the SA-Mamba subspace for bidirectional scanning, wherein the EPI-Mamba subspace comprises an EPI-H subspace and an EPI-W subspace; specifically, In the EPI-Mamba subspace, the light field feature is expanded along the spatial dimension to obtain a horizontal EPI slice and a vertical EPI slice while fixing the angle dimension in the EPI-H subspace and the EPI-W subspace, the obtained slices are flattened into one-dimensional sequences, bidirectional subspace scanning is performed respectively, and bidirectional dependence modeling is performed to obtain an EPI structure feature, wherein the EPI structure feature is used to represent the interaction between space and angle; In the SA-Mamba subspace, scanning is performed along the spatial dimension or the angle dimension to obtain a spatial sequence and an angle sequence, the obtained sequences are flattened into one-dimensional sequences, bidirectional subspace scanning is performed respectively, and bidirectional dependence modeling is performed to obtain a space-angle correlation feature representing the correlation between different viewing angles and spatial positions.
3. The light field image super-resolution reconstruction method of claim 2, wherein, In the EPI-Mamba subspace, after the obtained slices are flattened into one-dimensional sequences, bidirectional subspace scanning is performed respectively, and bidirectional dependence modeling is performed, specifically comprising the following steps: The horizontal EPI slice is flattened into a first one-dimensional sequence, and the vertical EPI slice is flattened into a second one-dimensional sequence; The first sequence and the second sequence are input into a bidirectional subspace scanning module, and a LayerNorm normalization process, a forward-backward bidirectional scanning modeling, a secondary LayerNorm normalization process, and a channel attention interaction process are sequentially performed; A convolution layer is connected in series after the bidirectional subspace scanning module, and the sequence subjected to the bidirectional dependence modeling is subjected to feature enhancement, wherein the bidirectional subspace scanning modules and the convolution layers corresponding to the horizontal EPI slice and the vertical EPI slice share weights; The features output by the convolution layer are reshaped into a 4D feature tensor to obtain the EPI structure feature.
4. The light field image super-resolution reconstruction method of claim 2, characterized in that, The EPI structure feature and the space-angle correlation feature are combined to obtain a to-be-fused feature, and the to-be-fused feature is unfolded into a first token sequence of the EPI-H subspace, and a bidirectional subspace scanning module is used to model the global dependence of the first token sequence to obtain an EPI-H enhanced sequence. The EPI-H enhanced sequence is reshaped into a second token sequence of the EPI-W subspace, and a bidirectional subspace scanning module is used to model the global dependence of the second token sequence to obtain an EPI-W enhanced sequence. The EPI-W enhanced sequence is reshaped into a 4D feature tensor to obtain and output a multi-scale fusion feature. The multi-scale fusion feature is rearranged into a subspace of a spatial dimension, a spatial attention map is generated through a convolution layer and a Sigmoid activation function, and spatial attention modulation is performed on the multi-scale fusion feature according to the spatial attention map and element-wise multiplication to obtain a spatial enhanced feature. The spatial enhanced feature is rearranged into a subspace of an angle dimension, an angle attention map is generated through a convolution layer and a Sigmoid activation function, and angle attention modulation is performed on the spatial enhanced feature according to the angle attention map and element-wise multiplication to obtain a comprehensive enhanced feature.
5. The light field image super-resolution reconstruction method of claim 2, wherein, Under the pre-built diffusion model framework, the comprehensive enhanced feature is taken as a guide condition, a predicted residual image is reconstructed under the dual constraints of the residual domain and the frequency domain through step-by-step denoising, and a high-resolution light field image is obtained by compensating the predicted residual image, and the method comprises the following steps: A target residual image is obtained, a cosine scheduling function is used to generate a noise addition coefficient, and a forward diffusion of a preset diffusion step number is performed on the target residual image according to the noise addition coefficient under the pre-built diffusion model framework; For each time step of the forward diffusion, a noise-added residual image is generated according to a Gaussian distribution, and the noise-added residual images corresponding to all the time steps are summarized to obtain a forward diffusion sequence.
6. The light field image super-resolution reconstruction method of claim 2, wherein, obtaining a current time step of a noise-added residual image from the forward diffusion sequence as a main input, injecting the integrated enhanced feature into a diffusion denoising network to obtain a predicted noise of the current time step; according to the predicted noise of the current time step, performing inverse update step by step according to a preset inverse diffusion formula, and simultaneously constraining the consistency of the predicted residual and the calculated real residual in the residual domain and the frequency domain through a preset global loss function until the time step is 0, to obtain a predicted residual image; superimposing the predicted residual image and the result of bicubic upsampling of the original light field image pixel by pixel to compensate and obtain a high-resolution light field image.
7. The light field image super-resolution reconstruction method of claim 6, characterized in that, The step of constraining the consistency of the predicted residual and the calculated real residual in the residual domain and the frequency domain through the preset global loss function specifically includes the following steps: performing pixel-by-pixel subtraction operation on the high-resolution light field image corresponding to the original light field image in the training data set and the result of bicubic upsampling of the original light field image to obtain the real residual; The global loss function includes a residual domain loss term and a frequency domain loss term, and the residual domain loss term is the L1 norm between the predicted residual image and the real residual; performing fast Fourier transform on the predicted residual image and the real residual respectively to obtain predicted frequency domain features and real frequency domain features respectively, and the frequency domain loss term is the L1 norm between the predicted frequency domain features and the real frequency domain features; weighting the residual domain loss term and the frequency domain loss term through a preset weight coefficient respectively to obtain a global loss value, and updating the parameters of the diffusion denoising network through a back propagation algorithm with the goal of minimizing the global loss value.
8. An optical field image super-resolution reconstruction apparatus, characterized in that, The method comprises the following modules: a subspace feature extraction module configured to extract EPI structure features from the received original light field image according to EPI-Mamba subspace scanning, and extract space-angle correlation features from the original light field image according to SA-Mamba subspace scanning; a cross-space multi-scale interactive fusion module configured to perform multi-scale interactive fusion across subspaces according to the EPI structure features and the space-angle correlation features, model global dependency through bidirectional scanning, and obtain multi-scale fusion features; a dual attention modulation module configured to perform spatial and angular dual attention modulation on the multi-scale fusion features to obtain integrated enhanced features; a diffusion residual reconstruction module configured to reconstruct a predicted residual image under the dual constraints of residual domain and frequency domain through step-by-step denoising in a pre-built diffusion model framework with the enhanced features as a guide condition, and compensate the predicted residual image to obtain a high-resolution light field image.
9. A smart terminal, characterized by The method comprises a memory and a processor, the memory stores at least one instruction, at least one program, a code set or an instruction set, and the processor loads and executes the at least one instruction, at least one program, code set or instruction set to realize the light field image super-resolution reconstruction method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, which are loaded and executed by the processor to implement the light field image super-resolution reconstruction method according to any one of claims 1 to 7.