A light field image super-resolution reconstruction method, device, equipment and medium

By decomposing light field images into spatial, angular, and polar plane subspaces, and using convolutional blocks for feature extraction and fusion, combined with network parameter adjustments, efficient and lightweight super-resolution reconstruction of light field images is achieved. This solves the problem of high model complexity in existing methods and improves reconstruction performance and interpretability.

CN121582066BActive Publication Date: 2026-04-21NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NAT UNIV OF DEFENSE TECH
Filing Date
2026-01-26
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing deep learning-based light field super-resolution methods are large in scale, complex in structure, and computationally expensive. Furthermore, they lack systematic research on the importance of subspace branches and the balance between network depth and width, which limits the interpretability and scalability of the models.

Method used

The residual module is used to decompose the basic feature map into three subspaces: spatial, angular and epipolar. Convolutional blocks are used to extract features, and weighted fusion is used to generate global features. Combined with network channels and parameter adjustments, efficient and lightweight super-resolution reconstruction of light field images is achieved.

Benefits of technology

It reduces model complexity, improves the performance of light field image super-resolution reconstruction, enhances spatial resolution and detail restoration capabilities, and strengthens geometric consistency and depth perception capabilities. It is suitable for fields such as light field camera image enhancement, virtual reality, and 3D scene reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121582066B_ABST
    Figure CN121582066B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and medium for super-resolution reconstruction of light field images, relating to the field of image processing technology. The method includes preprocessing a light field image at a first resolution, and extracting feature maps from the processed light field image. The first resolution is a resolution that meets a preset low-resolution judgment condition. The basic feature map is decomposed into subspaces, and features are extracted from the basic feature map using convolutional blocks to obtain features for each subspace. The features of each subspace are fused to generate a light field image at a second resolution. The second resolution is a resolution that meets a preset high-resolution judgment condition. Parameters in the convolutional blocks or residual modules are adjusted, and the process jumps to the preprocessing flow. Target light field images are screened, and the optimal parameters in the corresponding convolutional blocks or residual modules are determined. Based on the optimal parameters, efficient and lightweight super-resolution reconstruction of light field images is achieved, reducing model complexity and improving the performance of super-resolution reconstruction of light field images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method, apparatus, device, and medium for super-resolution reconstruction of light field images. Background Technology

[0002] Light field imaging is an imaging technique that can simultaneously record spatial and angular information. Existing deep learning-based light field super-resolution methods generally introduce redundant modules or complex feature coupling and attention mechanisms, resulting in large model size, complex structure, and high computational cost. This not only increases the difficulty of network implementation but also limits its interpretability and scalability. In addition, existing methods still lack systematic research on fundamental issues such as the relative importance of each subspace branch and the balance between network depth and width, which restricts the further optimization and generalization of subspace decoupling mechanisms.

[0003] As can be seen from the above, how to achieve efficient and lightweight super-resolution reconstruction of light field images, reduce model complexity, and improve the performance of super-resolution reconstruction of light field images is a problem to be solved in this field. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a method, apparatus, device, and medium for super-resolution reconstruction of light field images, which can achieve efficient and lightweight super-resolution reconstruction of light field images, reduce model complexity, and improve the performance of super-resolution reconstruction of light field images. The specific solution is as follows:

[0005] In a first aspect, this application discloses a method for super-resolution reconstruction of light field images, comprising:

[0006] A light field image at a first resolution is acquired, the light field image at the first resolution is preprocessed, and a feature map is extracted from the processed light field image to obtain a basic feature map; the first resolution is a resolution that meets a preset low-resolution determination condition; the preprocessing includes normalization processing and view arrangement processing.

[0007] The basic feature map is decomposed into subspaces using a residual module, and a convolutional block is set for each subspace. The convolutional block is used to extract features from the basic feature map to obtain the features of each subspace.

[0008] The features of each subspace are fused to generate global features, and a light field image at a second resolution is generated based on the global features; the second resolution is a resolution that meets a preset high-resolution determination condition.

[0009] The super-resolution reconstruction method is determined according to business requirements. The parameters in the convolution block or the residual module are adjusted according to the super-resolution reconstruction method. Then, the process jumps to the preprocessing of the light field image at the first resolution.

[0010] Filter out the target light field image from all the light field images at the second resolution, and determine the optimal parameters in the convolution block or the residual module corresponding to the target light field image;

[0011] Super-resolution reconstruction of the light field image at the first resolution is achieved based on the optimal parameters.

[0012] Optionally, the step of acquiring the light field image at a first resolution, preprocessing the light field image at the first resolution, and extracting feature maps from the processed light field image to obtain a basic feature map includes:

[0013] Acquire the light field image at the first resolution, represented by a four-dimensional data structure;

[0014] The light field image at the first resolution is normalized and its view is arranged using a two-dimensional convolutional layer.

[0015] The processed light field image is used to extract feature maps using a convolutional network to obtain basic feature maps.

[0016] Optionally, the step of decomposing the basic feature map into subspaces using the residual module includes:

[0017] Using a residual module and based on the principle of light field imaging, the basic feature map is decomposed into subspaces. The subspaces include a spatial subspace for extracting spatial texture and local detail information, an angular subspace for modeling the geometric parallax relationship between different viewpoints, and a polar plane subspace for characterizing the parallax slope and depth-related features.

[0018] Optionally, the step of using the convolutional block to extract features from the basic feature map includes:

[0019] The basic feature map is extracted using deep convolutions in the convolution block;

[0020] Alternatively, the basic feature map can be refined using the lightweight feature extraction unit in the convolutional block.

[0021] Optionally, the fusion of the features of each of the subspaces includes:

[0022] A weighted fusion mechanism is used to fuse the spatial, angular, and polar plane features of each subspace.

[0023] Optionally, generating the light field image at a second resolution based on the global features includes:

[0024] We utilize multi-layer convolution and residual structures to generate light field images at a second resolution based on global features.

[0025] Optionally, the step of adjusting the parameters in the convolutional block or the residual module according to the super-resolution reconstruction method, and then jumping to the process of preprocessing the light field image at the first resolution; selecting the target light field image from all the light field images at the second resolution, and determining the optimal parameters in the convolutional block or the residual module corresponding to the target light field image, includes:

[0026] If the super-resolution reconstruction method is network channel adjustment, the network channel allocation ratio in the convolution block is adjusted, and then the process of preprocessing the light field image at the first resolution is resumed. The target light field image is selected from all the light field images at the second resolution, and the optimal network channel allocation ratio corresponding to the target light field image is determined.

[0027] If the super-resolution reconstruction method is network parameter adjustment, then the network width and network depth in the residual module are adjusted, and then the process of preprocessing the light field image at the first resolution is resumed. The target light field image is selected from all the light field images at the second resolution, and the optimal network width and network depth corresponding to the target light field image are determined.

[0028] Secondly, this application discloses a light field image super-resolution reconstruction apparatus, comprising:

[0029] The feature map extraction module is used to acquire a light field image at a first resolution, preprocess the light field image at the first resolution, and extract a feature map from the processed light field image to obtain a basic feature map; the first resolution is a resolution that meets a preset low-resolution judgment condition; the preprocessing includes normalization processing and view arrangement processing.

[0030] The feature extraction module is used to decompose the basic feature map into subspaces using the residual module, and set convolutional blocks for each subspace. The convolutional blocks are used to extract features from the basic feature map to obtain features for each subspace.

[0031] The fusion module is used to fuse the features of each subspace to generate global features, and generate a light field image at a second resolution based on the global features; the second resolution is a resolution that meets a preset high resolution determination condition.

[0032] The parameter adjustment module is used to determine the super-resolution reconstruction method according to business requirements, adjust the parameters in the convolution block or the residual module according to the super-resolution reconstruction method, and then jump to the process of preprocessing the light field image at the first resolution.

[0033] The optimal parameter determination module is used to filter out the target light field image from all the light field images at the second resolution, and determine the optimal parameters in the convolution block or the residual module corresponding to the target light field image.

[0034] The super-resolution reconstruction module is used to achieve super-resolution reconstruction of the light field image at the first resolution based on the optimal parameters.

[0035] Thirdly, this application discloses an electronic device, including:

[0036] Memory, used to store computer programs;

[0037] A processor is used to execute the computer program to implement the aforementioned light field image super-resolution reconstruction method.

[0038] Fourthly, this application discloses a computer storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the aforementioned disclosed light field image super-resolution reconstruction method.

[0039] As can be seen, this application provides a method for super-resolution reconstruction of light field images, including acquiring a light field image at a first resolution, preprocessing the light field image at the first resolution, and extracting a feature map from the processed light field image to obtain a basic feature map; the first resolution is a resolution that meets a preset low-resolution judgment condition; the preprocessing includes normalization processing and view arrangement processing; decomposing the basic feature map into subspaces using a residual module, setting convolutional blocks for each subspace, using the convolutional blocks to refine features from the basic feature map to obtain features for each subspace; and fusing the features of each subspace to generate a global feature map. The system generates a light field image at a second resolution based on the global features. The second resolution is a resolution that meets a preset high-resolution judgment condition. A super-resolution reconstruction method is determined according to business requirements. The parameters in the convolutional block or the residual module are adjusted according to the super-resolution reconstruction method. Then, the process jumps to the preprocessing flow of the light field image at the first resolution. A target light field image is selected from all the light field images at the second resolution, and the optimal parameters in the convolutional block or the residual module corresponding to the target light field image are determined. Super-resolution reconstruction of the light field image at the first resolution is achieved based on the optimal parameters. This application preprocesses the light field image at a first resolution and extracts feature maps from the processed light field image to obtain a basic feature map. A residual module is used to decompose the basic feature map into subspaces, and convolutional blocks are used to refine the features of the basic feature map, improving the spatial resolution and detail restoration capability during the light field reconstruction process. This effectively captures geometric differences between multiple views, enhances the geometric consistency and depth perception capability of the light field image, and fuses the features of each subspace to generate global features. Based on these global features, a light field image at a second resolution is generated, reducing the number of parameters and computational complexity. This is done according to the super-resolution reconstruction method. The parameters in the convolutional block or residual module are adjusted, and then the process jumps to the preprocessing of the light field image at the first resolution. The target light field image is selected from all the light field images at the second resolution, and the optimal parameters in the convolutional block or residual module corresponding to the target light field image are determined. Based on the optimal parameters, efficient and lightweight super-resolution reconstruction of the light field image at the first resolution is achieved, reducing model complexity and improving the performance of light field image super-resolution reconstruction. It can be applied to fields such as light field camera image enhancement, virtual reality, 3D scene reconstruction, and stereoscopic display, and has good versatility and promotion value. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0041] Figure 1 This is a flowchart of a light field image super-resolution reconstruction method disclosed in this application;

[0042] Figure 2 This is a network framework diagram of a scalable light field image spatial super-resolution method disclosed in this application;

[0043] Figure 3 This is a comparison chart of visualization results on a light field dataset disclosed in this application;

[0044] Figure 4 This is a diagram illustrating the effect of network width and depth on super-resolution performance as disclosed in this application.

[0045] Figure 5 This application discloses a ternary parameter space diagram composed of three subspaces: space, angle, and polar plane.

[0046] Figure 6 This is a performance distribution diagram corresponding to a total number of channels of 48 as disclosed in this application;

[0047] Figure 7 This is a performance distribution diagram corresponding to a total number of channels of 96 as disclosed in this application;

[0048] Figure 8 This is a performance distribution diagram corresponding to a total number of channels of 192 as disclosed in this application;

[0049] Figure 9 This is a schematic diagram of the structure of a light field image super-resolution reconstruction device disclosed in this application;

[0050] Figure 10 This application provides a structural diagram of an electronic device. Detailed Implementation

[0051] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0052] Light field super-resolution is an imaging technique capable of simultaneously recording spatial and angular information. Existing deep learning-based light field super-resolution methods generally introduce redundant modules or complex feature coupling and attention mechanisms, resulting in large model size, complex structure, and high computational cost. This not only increases the difficulty of network implementation but also limits its interpretability and scalability. Furthermore, existing methods still lack systematic research on fundamental issues such as the relative importance of each subspace branch and the balance between network depth and width, limiting further optimization and generalization of subspace decoupling mechanisms. Therefore, achieving efficient and lightweight light field image super-resolution reconstruction, reducing model complexity, and improving the performance of light field image super-resolution reconstruction are problems that need to be solved in this field.

[0053] See Figure 1 As shown in the figure, this invention discloses a method for super-resolution reconstruction of light field images, which may specifically include:

[0054] Step S11: Obtain a light field image at a first resolution, preprocess the light field image at the first resolution, and extract a feature map from the processed light field image to obtain a basic feature map; the first resolution is a resolution that meets the preset low resolution judgment condition; the preprocessing includes normalization processing and view arrangement processing.

[0055] In this embodiment, a light field image at a first resolution is obtained, represented by a four-dimensional data structure; the light field image at the first resolution is normalized and its view is arranged using a two-dimensional convolutional layer; and a feature map is extracted from the processed light field image using a convolutional network to obtain a basic feature map.

[0056] The light field image at the first resolution in this application is represented by a four-dimensional data structure, which can be expressed as follows: ,in, Indicates angle or dimension. Indicates spatial dimension.

[0057] Furthermore, if a macro-pixel image is obtained, it is converted into four-dimensional light field data, which is represented by L.

[0058] In this embodiment, after normalizing and arranging the light field image at the first resolution using a two-dimensional convolutional layer with a kernel size of 3×3, a stride of 1, and padding of 1, an initial feature extractor is used to extract the local spatial features of the processed light field image and generate an initial feature representation. Then, a convolutional network is used to extract the feature map of the processed light field image to obtain the basic feature map. The feature extractor is a convolution operator.

[0059] Step S12: Decompose the basic feature map into subspaces using the residual module, set convolutional blocks for each subspace, and use the convolutional blocks to extract features from the basic feature map to obtain features for each subspace.

[0060] In this embodiment, a residual module is used to decompose the basic feature map into subspaces based on the principle of light field imaging. Convolutional blocks are set for each subspace. The deep convolution in the convolutional blocks is used to refine the features of the basic feature map, or the lightweight feature extraction unit in the convolutional blocks is used to refine the features of the basic feature map to obtain the features of each subspace. The subspaces include a spatial subspace for extracting spatial texture and local detail information, an angular subspace for modeling the geometric parallax relationship between different viewpoints, and an epipolar subspace for characterizing the parallax slope and depth-related features.

[0061] This application decomposes the basic feature map into three independent subspaces and sets corresponding convolutional blocks for the three subspaces, with each convolutional block receiving the basic feature map. Features are extracted using deep convolutions or lightweight feature extraction units to obtain their respective subspace features; the convolutional blocks include spatial convolutional blocks, horizontal epipolar convolutional blocks, angular convolutional blocks, and vertical epipolar convolutional blocks. The specific steps are as follows:

[0062] (1) The basic feature map The input spatial convolutional block extracts local spatial texture and detail features sequentially through several layers of convolutional units. Specifically, the spatial convolutional block uses a convolution operator with a kernel size of 3×3, a stride of 1, and padding of 1, and feature enhancement can be performed through residual structures. Each residual unit consists of two convolutional layers and a nonlinear activation function to achieve stable extraction of local details. After multiple layers of convolution and residual stacking, the features are channel normalized and dimension mapped to obtain spatial subspace features. These subspace features are mainly used to restore the high-frequency texture and edge structure of the image, thereby improving the spatial resolution and detail restoration capability in the light field reconstruction process.

[0063] (2) Based on the arrangement rules of the light field perspective, the basic feature map is arranged... The data is reorganized into a view sequence based on different perspectives, and a shared convolutional encoding operation is performed on each view to extract its basic features. Subsequently, the multi-view features are stacked along the angular dimension, and convolutional blocks are refined through cross-view convolution or lightweight angular feature extraction to model the geometric relationships and disparity variations between different perspectives. To further improve angular consistency, several angular residual units can be superimposed on the convolutional blocks to enhance the expressive power of the features. Finally, an angular subspace feature is obtained through a channel mapping convolution layer. This feature effectively captures the geometric differences between multiple views, providing feature support for subsequent multi-view consistency reconstruction.

[0064] (3) Based on the principle of light field imaging, from the basic feature map Epiplane slices are extracted along the horizontal or vertical viewing direction. Each epiplane slice corresponds to a two-dimensional image with varying viewing angles at a fixed spatial position, which can characterize the disparity slope and depth structure characteristics. Subsequently, feature extraction is performed on the epiplane slices using a convolutional block composed of consecutive 3×3 convolutional layers and nonlinear activation functions to refine depth-sensitive texture features. After feature refinement, all epiplane encoding results are rearranged to their corresponding spatial positions, and feature aggregation is performed through channel fusion convolution to finally obtain epiplane subspace features. These features are mainly used to supplement depth-related information and disparity structure description, enhancing the geometric consistency and depth perception capability of the light field image.

[0065] Among them, subspace features include spatial subspace features, spatial subspace features, and polar plane subspace features.

[0066] Step S13: Fuse the features of each subspace to generate global features, and generate a light field image at a second resolution based on the global features; the second resolution is a resolution that meets the preset high resolution judgment conditions.

[0067] In this embodiment, a weighted fusion mechanism is used to fuse the spatial, angular, and polar plane features in each of the subspace features to generate global features. Multi-layer convolution and residual structures are then used to generate a light field image at a second resolution based on the global features.

[0068] In this embodiment, a weighted fusion mechanism is used to fuse the spatial, angular, and polar plane features in each of the subspace features. In other words, a weighted fusion mechanism is used to fuse the spatial subspace features. Features of spatial subspace Polar plane subspace characteristics Fusion to generate global features , Then, using multi-layer convolution and residual structures, and based on global features, a high-resolution light field image is generated, namely the light field image at the second resolution.

[0069] Step S14: Determine the super-resolution reconstruction method according to business requirements, adjust the parameters in the convolution block or the residual module according to the super-resolution reconstruction method, and then jump to the process of preprocessing the light field image at the first resolution.

[0070] In this embodiment, if the super-resolution reconstruction method is network channel adjustment, the network channel allocation ratio in the convolution block is adjusted, and then the process jumps to the preprocessing flow of the light field image at the first resolution; if the super-resolution reconstruction method is network parameter adjustment, the network width and network depth in the residual module are adjusted, and then the process jumps to the preprocessing flow of the light field image at the first resolution.

[0071] Step S15: Select the target light field image from all the light field images at the second resolution, and determine the optimal parameters in the convolution block or the residual module corresponding to the target light field image.

[0072] In this embodiment, if the super-resolution reconstruction method is network channel adjustment, the target light field image is selected from all the light field images at the second resolution, and the optimal network channel allocation ratio corresponding to the target light field image is determined; if the super-resolution reconstruction method is network parameter adjustment, the target light field image is selected from all the light field images at the second resolution, and the optimal network width and network depth corresponding to the target light field image are determined.

[0073] In other words, if the super-resolution reconstruction method is network channel adjustment, then while keeping the network parameters constant, the ratio of the three branches of the angle space and the polar plane is adjusted (i.e., the network channel allocation ratio is adjusted) to find the network channel allocation pattern and determine the optimal network channel allocation ratio; if the super-resolution reconstruction method is network parameter adjustment, then the channel allocation ratio is kept constant, and the network width and network depth are adjusted to find the distribution pattern of network width and network depth under the condition of constant parameters and determine the optimal network width and network depth.

[0074] Step S16: Perform super-resolution reconstruction of the light field image at the first resolution based on the optimal parameters.

[0075] This application constructs a scalable network framework for light field image spatial super-resolution methods, the specific structure of which is as follows: Figure 2As shown, a 3×3 convolutional layer is used to preprocess the light field image at the first resolution. Feature maps are extracted from the processed light field image to obtain a basic feature map. A residual module is used to decompose the basic feature map into subspaces. Convolutional blocks are used to refine the features of each subspace, and an upsampling module is used to fuse the features of each subspace to generate global features. A light field image at the second resolution is generated based on these global features. The parameters in the convolutional blocks or the residual module are adjusted according to the super-resolution reconstruction method. Then, the process jumps to the preprocessing of the light field image at the first resolution. A target light field image is selected from all the light field images at the second resolution, and the optimal parameters in the convolutional blocks or residual modules corresponding to the target light field image are determined. Based on the optimal parameters, super-resolution reconstruction of the light field image at the first resolution is achieved. This application achieves faithful reconstruction of light field images through joint learning of space, angle, and epiplanet. Utilizing a modular design with a multi-branch structure, it effectively improves reconstruction accuracy and model scalability while maintaining network lightweightness and structural simplicity.

[0076] To verify the effectiveness and universality of the proposed method, experiments were conducted using multiple publicly available light field datasets, including EPFL, HCInew, HCIold, INRIA Lytro, and STF Gantry. The experiments used low-resolution light fields as input, performed super-resolution reconstruction using the model of this invention, and compared it with several existing representative methods to evaluate its performance in terms of spatial reconstruction accuracy, angular consistency, and model complexity.

[0077] To verify the effectiveness of the proposed invention, this application is compared with algorithms currently in the field, namely: LF-DFnet, MEG-Net, LF-IINet, DistgSSR, and HLFSSR. Visualization results on the STFGantry and HCI_new light field datasets are compared, for example... Figure 3 As shown in Table 1, the PSNR (Peak Signal-to-Noise Ratio) results of this application and current algorithms in the field are analyzed.

[0078] Table 1. Analysis of Peak Signal-to-Noise Ratio Results between This Application and Current Field Algorithms

[0079]

[0080] Experimental results show that the method of this invention achieves a peak signal-to-noise ratio (PSNR) comparable to or better than existing methods on various datasets, while maintaining a model parameter count of 3.52M, significantly lower than HLFSR's 13.9M. Specifically, the method of this invention achieves a PSNR of 29.23dB on the EPFL dataset, higher than LF-DFNet and DistgSSR; and a PSNR of 37.61dB on the HCIold dataset, only 0.16dB different from HLFSR, but with only about one-quarter of the parameter count. This indicates that the method of this invention can achieve high-quality light field image reconstruction while ensuring a lightweight network, effectively overcoming the problems of model redundancy and structural complexity in existing methods.

[0081] This application analyzes the impact of network width and depth on super-resolution performance, such as... Figure 4 As shown in the figure, experiments demonstrate that increasing either network width or depth improves the accuracy of light field reconstruction, but performance gains tend to plateau as model size increases. With a fixed model parameter budget, increasing network depth is more beneficial for performance improvement than increasing network width, indicating that deep feature extraction is more critical for light field reconstruction. For networks with specific parameter values, the optimal depth for a 3.57M parameter network is between 16 and 32, while the optimal depth for a 1.18M parameter network is between 24 and 40, providing a reference range for network design.

[0082] This application also explores the impact of the three-branch subspace allocation on performance, in a ternary parameter space composed of the spatial, angular, and polar plane subspaces, such as... Figure 5 As shown. The performance distribution corresponding to a total of 48 channels is as follows. Figure 6 As shown, the performance distribution corresponding to a total of 96 channels is as follows: Figure 7 As shown, the performance distribution corresponding to a total of 192 channels is as follows: Figure 8 As shown, in the ternary parameter space composed of the spatial, angular, and polar plane subspaces, the performance of the interior region is generally higher than that of the edges or vertices, verifying the effectiveness of the three-branch design. Good performance can still be obtained while maintaining a relative balance between the spatial and polar plane subspaces, indicating that the angular subspace is relatively less important under these conditions. As the model parameter budget increases, the optimal allocation point of the three-branch subspace gradually shifts from the lower right region towards the center, indicating the trend of the optimal allocation ratio with the budget.

[0083] In summary, the method of this invention not only outperforms existing methods in average performance, but also achieves an optimal trade-off between performance and efficiency under different parameter budgets through reasonable design of network width, depth, and three-branch subspace allocation. This fully demonstrates the flexibility and practicality of the method in structural design, providing a reliable basis for the lightweight, high-performance, and practical application of light field image super-resolution networks.

[0084] The purpose of this invention is to overcome the shortcomings of existing light field spatial super-resolution techniques. Therefore, a simple, modular, and scalable light field spatial super-resolution method is proposed, which reduces the number of model parameters and computational complexity while achieving reconstruction performance comparable to or even better than existing complex networks.

[0085] In this embodiment, a light field image at a first resolution is acquired, preprocessed, and feature map extraction is performed on the processed light field image to obtain a basic feature map. The first resolution is a resolution that meets a preset low-resolution judgment condition. The preprocessing includes normalization and view arrangement. The basic feature map is decomposed into subspaces using a residual module, and convolutional blocks are set for each subspace. The convolutional blocks are used to refine the features of the basic feature map to obtain features for each subspace. The features of each subspace are fused to generate global features. Based on the global features... A light field image at a second resolution is generated; the second resolution is a resolution that meets a preset high-resolution judgment condition; a super-resolution reconstruction method is determined according to business requirements, and the parameters in the convolution block or the residual module are adjusted according to the super-resolution reconstruction method, and then the process jumps to the preprocessing of the light field image at the first resolution; a target light field image is selected from all the light field images at the second resolution, and the optimal parameters in the convolution block or the residual module corresponding to the target light field image are determined; super-resolution reconstruction of the light field image at the first resolution is achieved based on the optimal parameters. This application preprocesses the light field image at a first resolution and extracts feature maps from the processed light field image to obtain a basic feature map. A residual module is used to decompose the basic feature map into subspaces, and convolutional blocks are used to refine the features of the basic feature map, improving the spatial resolution and detail restoration capability during the light field reconstruction process. This effectively captures geometric differences between multiple views, enhances the geometric consistency and depth perception capability of the light field image, and fuses the features of each subspace to generate global features. Based on these global features, a light field image at a second resolution is generated, reducing the number of parameters and computational complexity. This is done according to the super-resolution reconstruction method. The parameters in the convolutional block or residual module are adjusted, and then the process jumps to the preprocessing of the light field image at the first resolution. The target light field image is selected from all the light field images at the second resolution, and the optimal parameters in the convolutional block or residual module corresponding to the target light field image are determined. Based on the optimal parameters, efficient and lightweight super-resolution reconstruction of the light field image at the first resolution is achieved, reducing model complexity and improving the performance of light field image super-resolution reconstruction. It can be applied to fields such as light field camera image enhancement, virtual reality, 3D scene reconstruction, and stereoscopic display, and has good versatility and promotion value.

[0086] See Figure 9 As shown in the figure, an embodiment of the present invention discloses a light field image super-resolution reconstruction device, which may specifically include:

[0087] The feature map extraction module 11 is used to acquire a light field image at a first resolution, preprocess the light field image at the first resolution, and extract a feature map from the processed light field image to obtain a basic feature map; the first resolution is a resolution that meets a preset low-resolution judgment condition; the preprocessing includes normalization processing and view arrangement processing.

[0088] Feature extraction module 12 is used to decompose the basic feature map into subspaces using the residual module, set convolution blocks for each subspace, and use the convolution blocks to extract features from the basic feature map to obtain features of each subspace.

[0089] The fusion module 13 is used to fuse the features of each subspace to generate global features, and generate a light field image at a second resolution based on the global features; the second resolution is a resolution that meets the preset high resolution judgment conditions.

[0090] The parameter adjustment module 14 is used to determine the super-resolution reconstruction method according to business requirements, adjust the parameters in the convolution block or the residual module according to the super-resolution reconstruction method, and then jump to the process of preprocessing the light field image at the first resolution.

[0091] The optimal parameter determination module 15 is used to filter out the target light field image from all the light field images at the second resolution and determine the optimal parameters in the convolution block or the residual module corresponding to the target light field image.

[0092] The super-resolution reconstruction module 16 is used to perform super-resolution reconstruction of the light field image at the first resolution based on the optimal parameters.

[0093] In this embodiment, a light field image at a first resolution is acquired, preprocessed, and feature map extraction is performed on the processed light field image to obtain a basic feature map. The first resolution is a resolution that meets a preset low-resolution judgment condition. The preprocessing includes normalization and view arrangement. The basic feature map is decomposed into subspaces using a residual module, and convolutional blocks are set for each subspace. The convolutional blocks are used to refine the features of the basic feature map to obtain features for each subspace. The features of each subspace are fused to generate global features. Based on the global features... A light field image at a second resolution is generated; the second resolution is a resolution that meets a preset high-resolution judgment condition; a super-resolution reconstruction method is determined according to business requirements, and the parameters in the convolution block or the residual module are adjusted according to the super-resolution reconstruction method, and then the process jumps to the preprocessing of the light field image at the first resolution; a target light field image is selected from all the light field images at the second resolution, and the optimal parameters in the convolution block or the residual module corresponding to the target light field image are determined; super-resolution reconstruction of the light field image at the first resolution is achieved based on the optimal parameters. This application preprocesses the light field image at a first resolution and extracts feature maps from the processed light field image to obtain a basic feature map. A residual module is used to decompose the basic feature map into subspaces, and convolutional blocks are used to refine the features of the basic feature map, improving the spatial resolution and detail restoration capability during the light field reconstruction process. This effectively captures geometric differences between multiple views, enhances the geometric consistency and depth perception capability of the light field image, and fuses the features of each subspace to generate global features. Based on these global features, a light field image at a second resolution is generated, reducing the number of parameters and computational complexity. This is done according to the super-resolution reconstruction method. The parameters in the convolutional block or residual module are adjusted, and then the process jumps to the preprocessing of the light field image at the first resolution. The target light field image is selected from all the light field images at the second resolution, and the optimal parameters in the convolutional block or residual module corresponding to the target light field image are determined. Based on the optimal parameters, efficient and lightweight super-resolution reconstruction of the light field image at the first resolution is achieved, reducing model complexity and improving the performance of light field image super-resolution reconstruction. It can be applied to fields such as light field camera image enhancement, virtual reality, 3D scene reconstruction, and stereoscopic display, and has good versatility and promotion value.

[0094] In some specific embodiments, the feature map extraction module 11 may specifically include:

[0095] The light field image acquisition module is used to acquire a light field image at a first resolution, represented by a four-dimensional data structure.

[0096] The normalization and view arrangement module is used to perform normalization and view arrangement processing on the light field image at the first resolution using a two-dimensional convolutional layer.

[0097] The basic feature map determination module is used to extract feature maps from the processed light field image using a convolutional network to obtain basic feature maps.

[0098] In some specific embodiments, the feature extraction module 12 may specifically include:

[0099] The subspace decomposition module is used to decompose the basic feature map into subspaces based on the principle of light field imaging using the residual module. The subspaces include a spatial subspace for extracting spatial texture and local detail information, an angular subspace for modeling the geometric parallax relationship between different viewpoints, and a polar plane subspace for characterizing the parallax slope and depth-related features.

[0100] In some specific embodiments, the feature extraction module 12 may specifically include:

[0101] The first feature extraction module is used to extract features from the basic feature map using deep convolution in the convolution block;

[0102] The second feature extraction module is used to extract features from the basic feature map by utilizing the lightweight feature extraction unit in the convolution block.

[0103] In some specific embodiments, the fusion module 13 may specifically include:

[0104] The fusion module is used to fuse the spatial, angular, and polar plane features of each of the subspace features using a weighted fusion mechanism.

[0105] In some specific embodiments, the fusion module 13 may specifically include:

[0106] The second-resolution light field image generation module is used to generate a second-resolution light field image based on global features by utilizing multi-layer convolution and residual structures.

[0107] In some specific embodiments, the optimal parameter determination module 15 may specifically include:

[0108] The optimal network channel allocation ratio determination module is used to adjust the network channel allocation ratio in the convolution block if the super-resolution reconstruction method is network channel adjustment, and then jump to the process of preprocessing the light field image at the first resolution, filter out the target light field image from all the light field images at the second resolution, and determine the optimal network channel allocation ratio corresponding to the target light field image.

[0109] The optimal network width and network depth determination module is used to adjust the network width and network depth in the residual module if the super-resolution reconstruction method is network parameter adjustment, and then jump to the process of preprocessing the light field image at the first resolution, to filter out the target light field image from all the light field images at the second resolution, and to determine the optimal network width and network depth corresponding to the target light field image.

[0110] Figure 10 This is a schematic diagram of an electronic device provided in an embodiment of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the light field image super-resolution reconstruction method performed by the electronic device disclosed in any of the foregoing embodiments.

[0111] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0112] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon include operating system 221, computer program 222 and data 223, etc., and the storage method can be temporary storage or permanent storage.

[0113] The operating system 221 manages and controls the various hardware devices and computer programs 222 on the electronic device 20 to enable the processor 21 to perform calculations and processing on the data 223 in the memory 22. The operating system 221 can be Windows, Unix, Linux, etc. The computer program 222, in addition to including a computer program capable of performing the light field image super-resolution reconstruction method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, may further include computer programs capable of performing other specific tasks. The data 223 may include data received by the light field image super-resolution reconstruction device from external devices, as well as data collected by its own input / output interface 25.

[0114] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0115] Furthermore, this application also discloses a computer-readable storage medium storing a computer program. When the computer program is loaded and executed by a processor, it implements the steps of the light field image super-resolution reconstruction method disclosed in any of the foregoing embodiments.

[0116] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0117] The above provides a detailed description of the light field image super-resolution reconstruction method, apparatus, device, and storage medium provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for super-resolution reconstruction of light field images, characterized in that, include: A light field image at a first resolution is acquired, the light field image at the first resolution is preprocessed, and a feature map is extracted from the processed light field image to obtain a basic feature map. The first resolution is the resolution that meets the preset low resolution determination condition; The preprocessing includes normalization and view arrangement. The basic feature map is decomposed into subspaces using a residual module, and a convolutional block is set for each subspace. The convolutional block is used to extract features from the basic feature map to obtain the features of each subspace. The features of each subspace are fused to generate global features, and a light field image at a second resolution is generated based on the global features; the second resolution is a resolution that meets a preset high-resolution determination condition. The super-resolution reconstruction method is determined according to business requirements. The parameters in the convolution block or the residual module are adjusted according to the super-resolution reconstruction method. Then, the process jumps to the preprocessing of the light field image at the first resolution. Filter out the target light field image from all the light field images at the second resolution, and determine the optimal parameters in the convolution block or the residual module corresponding to the target light field image; Super-resolution reconstruction of the light field image at the first resolution is achieved based on the optimal parameters. The process involves adjusting the parameters in the convolutional block or the residual module according to the super-resolution reconstruction method, and then proceeding to the preprocessing step for the light field image at the first resolution. The next step involves selecting a target light field image from all the light field images at the second resolution and determining the optimal parameters in the convolutional block or the residual module corresponding to the target light field image. This includes: if the super-resolution reconstruction method is network channel adjustment, adjusting the network channel allocation ratio in the convolutional block, and then proceeding to the preprocessing step for the light field image at the first resolution, selecting the target light field image from all the light field images at the second resolution, and determining the optimal network channel allocation ratio corresponding to the target light field image; if the super-resolution reconstruction method is network parameter adjustment, adjusting the network width and network depth in the residual module, and then proceeding to the preprocessing step for the light field image at the first resolution, selecting the target light field image from all the light field images at the second resolution, and determining the optimal network width and network depth corresponding to the target light field image.

2. The light field image super-resolution reconstruction method according to claim 1, characterized in that, The process of acquiring a light field image at a first resolution, preprocessing the light field image at the first resolution, and extracting a feature map from the processed light field image to obtain a basic feature map includes: Acquire the light field image at the first resolution, represented by a four-dimensional data structure; The light field image at the first resolution is normalized and its view is arranged using a two-dimensional convolutional layer. The processed light field image is used to extract feature maps using a convolutional network to obtain basic feature maps.

3. The method for super-resolution reconstruction of light field images according to claim 1, characterized in that, The process of decomposing the basic feature map into subspaces using the residual module includes: Using a residual module and based on the principle of light field imaging, the basic feature map is decomposed into subspaces. The subspaces include a spatial subspace for extracting spatial texture and local detail information, an angular subspace for modeling the geometric parallax relationship between different viewpoints, and a polar plane subspace for characterizing the parallax slope and depth-related features.

4. The method for super-resolution reconstruction of light field images according to claim 1, characterized in that, The step of extracting features from the base feature map using the convolutional block includes: The basic feature map is extracted using deep convolutions in the convolution block; Alternatively, the basic feature map can be refined using the lightweight feature extraction unit in the convolutional block.

5. The method for super-resolution reconstruction of light field images according to claim 1, characterized in that, The fusion of the features of each of the subspaces includes: A weighted fusion mechanism is used to fuse the spatial, angular, and polar plane features of each subspace.

6. The method for super-resolution reconstruction of light field images according to claim 1, characterized in that, The generation of the light field image at the second resolution based on the global features includes: We utilize multi-layer convolution and residual structures to generate light field images at a second resolution based on global features.

7. A device for super-resolution reconstruction of light field images, characterized in that, include: The feature map extraction module is used to acquire a light field image at a first resolution, preprocess the light field image at the first resolution, and extract a feature map from the processed light field image to obtain a basic feature map; the first resolution is a resolution that meets a preset low-resolution judgment condition; the preprocessing includes normalization processing and view arrangement processing. The feature extraction module is used to decompose the basic feature map into subspaces using the residual module, and set convolutional blocks for each subspace. The convolutional blocks are used to extract features from the basic feature map to obtain features for each subspace. The fusion module is used to fuse the features of each subspace to generate global features, and generate a light field image at a second resolution based on the global features; the second resolution is a resolution that meets a preset high resolution determination condition. The parameter adjustment module is used to determine the super-resolution reconstruction method according to business requirements, adjust the parameters in the convolution block or the residual module according to the super-resolution reconstruction method, and then jump to the process of preprocessing the light field image at the first resolution. The optimal parameter determination module is used to filter out the target light field image from all the light field images at the second resolution, and determine the optimal parameters in the convolution block or the residual module corresponding to the target light field image. The super-resolution reconstruction module is used to achieve super-resolution reconstruction of the light field image at the first resolution based on the optimal parameters. The process involves adjusting the parameters in the convolutional block or the residual module according to the super-resolution reconstruction method, and then proceeding to the preprocessing step for the light field image at the first resolution. The next step involves selecting a target light field image from all the light field images at the second resolution and determining the optimal parameters in the convolutional block or the residual module corresponding to the target light field image. This includes: if the super-resolution reconstruction method is network channel adjustment, adjusting the network channel allocation ratio in the convolutional block, and then proceeding to the preprocessing step for the light field image at the first resolution, selecting the target light field image from all the light field images at the second resolution, and determining the optimal network channel allocation ratio corresponding to the target light field image; if the super-resolution reconstruction method is network parameter adjustment, adjusting the network width and network depth in the residual module, and then proceeding to the preprocessing step for the light field image at the first resolution, selecting the target light field image from all the light field images at the second resolution, and determining the optimal network width and network depth corresponding to the target light field image.

8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the light field image super-resolution reconstruction method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when the computer program is executed by a processor, it implements the light field image super-resolution reconstruction method as described in any one of claims 1 to 6.