Multi-scale feature fusion light field image dense decoupling reconstruction method

Through the dense decoupling reconstruction method of multi-scale feature fusion, the problem of global connection loss of light field information and insufficient use of subspace information in the existing light field angle super-resolution technology is solved, and the light field image reconstruction effect with higher accuracy is achieved.

CN119942277APending Publication Date: 2025-05-06SHANGHAI NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411768297.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing light field angle super-resolution technology is prone to lose the global connection of light field information when reconstructing high-angle resolution images using sparse sampling, and insufficient extraction and utilization of light field subspace information, making it difficult to improve the robustness and generation effect of the model.

Method used

The dense decoupling reconstruction method of light field images with multi-scale feature fusion is adopted. By designing a multi-layer dense feature extractor, including a spatial feature extractor, an angle feature extractor, a dense spatial feature extractor and an extreme plane feature fusion feature extractor, the deep fusion of space and angle information is achieved, and a multi-scale attention mechanism is adopted in the extreme plane feature fusion feature extractor.

Benefits of technology

It realizes a higher precision light field image reconstruction effect, which can effectively process complex details and depth information in the light field image, and improves the robustness and generation effect of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942277A_ABST
    Figure CN119942277A_ABST
Patent Text Reader

Abstract

The invention relates to a light field image dense decoupling reconstruction method based on multi-scale feature fusion, and the method comprises the steps: carrying out the sampling of a light field image, obtaining a sub-aperture image set, obtaining a sparse sub-aperture image set through sparse sampling, and converting the sparse sub-aperture image set into a first macro-pixel image; then, feature extraction is carried out on the first macro pixel image through a convolutional layer, and a first feature map is output; and then, inputting the feature map into a plurality of dense feature extractors for processing to obtain a second feature map, outputting a second macro-pixel image through an angle feature extractor and an up-sampling module, and finally converting the second macro-pixel image into a reconstructed sub-aperture image set. According to the method, through multi-level feature fusion, the precision of decoupling reconstruction of the light field image is remarkably improved, and the method is suitable for high-quality light field image reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and in particular relates to a light field image dense decoupling reconstruction method with multi-scale feature fusion. Background Art

[0002] Light field imaging technology can record the intensity and direction information of light and is an imaging method that contains four-dimensional data. The current light field angle super-resolution technology faces two main problems: first, when using sparsely sampled light fields to reconstruct high-angle resolution images, the global connection of light field information is easily lost; second, the existing methods do not extract and utilize light field subspace information sufficiently, making it difficult to improve the robustness and generation effect of the model. Therefore, an improved light field angle super-resolution model is proposed, which can extract the features of the five dimensions of the light field and effectively fuse them by using dense decoupling. Summary of the invention

[0003] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and to provide a light field image dense decoupling reconstruction method with multi-scale feature fusion.

[0004] The purpose of the present invention can be achieved by the following technical solutions:

[0005] The present invention provides a light field image dense decoupling reconstruction method with multi-scale feature fusion, comprising the following steps:

[0006] Step S1: sampling the light field image to obtain a sub-aperture image set, and sparsely sampling the sub-aperture image set to obtain a sparse sub-aperture image set;

[0007] Step S2: converting the sparse sub-aperture image set into a first macro-pixel image;

[0008] Step S3: inputting the first macro-pixel image into the first convolutional layer for feature extraction, and outputting a first feature map;

[0009] Step S4: input the first feature map into a plurality of identical and sequentially connected dense feature extractors for processing, and output a second feature map;

[0010] Step S5: passing the second feature map through an angle feature extractor and an upsampling module to output a second macro-pixel image;

[0011] Step S6: Convert the second macro-pixel image into a set of reconstructed sub-aperture images.

[0012] Furthermore, the first convolutional layer is a convolutional layer with a convolution kernel of 3×3, a stride of 1, an expansion of 2, and a padding of 2.

[0013] Furthermore, the dense feature extractor includes a spatial feature extractor, an angle feature extractor, a dense spatial feature extractor, a polar plane feature fusion feature extractor, and a first fusion convolution.

[0014] Further, the step S4 includes the following steps:

[0015] The first feature map is input into a plurality of dense feature extractors connected in sequence for processing, each dense feature extractor outputs a fused feature map, and the fused feature map is input into a plurality of subsequent dense feature extractors, and the second feature map is outputted through the last dense feature extractor.

[0016] Furthermore, the step of inputting the first feature map into a plurality of sequentially connected dense feature extractors for processing, wherein each dense feature extractor outputs a fused feature map, comprises the following steps:

[0017] Input the first feature map into the spatial feature extractor of the dense feature extractor to extract spatial features and output the spatial feature map. Input the spatial feature map into the angle feature extractor and output the angle feature map. Input the spatial feature map into the dense spatial feature extractor and output the local spatial feature map. Input the spatial feature map into the polar plane feature fusion feature extractor and output the polar plane feature map. Input the spatial feature map, the angle feature map, the local spatial feature map and the polar plane feature map into the first fused convolution and output the first fused feature map. Input the first fused feature map into the spatial feature extractor. Perform a residual connection between the feature map processed by the spatial feature extractor and the first fused feature map and output the fused feature map.

[0018] Furthermore, the spatial feature extractor includes two layers of identical dilated convolutions, wherein the dilated convolutions have a convolution kernel size of 3×3, a stride of 1, a dilation rate of 2, and a padding of 2.

[0019] Furthermore, the step of inputting the spatial feature map into the angle feature extractor and outputting the angle feature map specifically includes the following steps:

[0020] The spatial feature map is input into a convolution with an output channel number of Channel, a convolution kernel size of 2×2, a stride of 2, and a padding of 0 for dimensionality reduction processing to extract angle-related local features. The processed feature map is processed by a convolution with an output channel number of 4×Channel, a convolution kernel size of 1×1, and a stride of 1. The output feature map is processed by the PixelShuffle operation to output the angle feature map.

[0021] Furthermore, the dense spatial feature extractor includes two convolutions with the same output channel number of channels, a convolution kernel size of 5×5, a stride of 1, a padding of 2, and a dilation rate of 1. Nonlinear changes are introduced between the two convolution layers through the LeakyReLU activation function.

[0022] Furthermore, the polar plane feature fusion feature extractor includes a horizontal polar plane feature extractor, a vertical polar plane feature extractor and a second fusion convolution, and the spatial feature map is input into the polar plane feature fusion feature extractor, and the polar plane feature map is output, including the following steps:

[0023] The spatial feature map is input into the horizontal polar plane feature extractor, the disparity information in the horizontal direction is extracted through the convolution operation, the disparity information is decoded to the horizontal direction through PixelShuffle1D, and the horizontal disparity information is output;

[0024] The spatial feature map is input into the vertical polar plane feature extractor, the vertical polar plane feature extractor transposes the spatial feature map and extracts the vertical disparity information through convolution operation, decodes the disparity information to the vertical direction through PixelShuffle1D, and outputs the vertical disparity information;

[0025] The horizontal disparity information is spliced ​​with the vertical disparity information, and the spliced ​​features are channel-fused through the second fusion convolution with an output channel number of 2×EpiChannel, a convolution kernel size of 1×1, a stride of 1, and a padding of 0 to output the epiplane feature map.

[0026] Furthermore, the extreme plane feature fusion feature extractor adopts a multi-scale attention mechanism.

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] (1) The present invention achieves deep fusion of spatial and angular information by designing a multi-layer dense feature extractor, including a spatial feature extractor, an angular feature extractor, a dense spatial feature extractor, and an extreme plane feature fusion feature extractor. In particular, the multi-scale feature fusion mechanism introduced in the dense feature extractor enables the model to capture both detail information and global information in the light field image, thereby achieving a higher-precision reconstruction effect.

[0029] (2) The present invention adopts a multi-scale attention mechanism in the polar plane feature fusion feature extractor, which combines channel attention and spatial attention, extracts global information through global average pooling and maximum pooling, and uses multi-scale convolution kernels to extract features, thereby enhancing the expressiveness of features. This makes the reconstruction process more detailed and can effectively process complex details and depth information in light field images.

[0030] (3) The present invention adopts residual connection between multiple feature extractors. This design can avoid the gradient vanishing problem and help the network maintain stability during training. By performing residual connection between the fused feature map and the original feature map, the low-level features are effectively retained, while the expression ability of the model is improved, further improving the reconstruction accuracy of the light field image.

[0031] (4) The present invention gradually extracts spatial features and angular features through multi-layer convolution operations, and fuses these features in the subsequent processing process. This deep decoupling design enables the model to understand the relationship between spatial and angular dimensions more accurately, thereby achieving better light field image reconstruction effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is a flow chart of the method of the present invention;

[0033] Figure 2 It is a flowchart of the light field image angle super-resolution method of the present invention;

[0034] Figure 3 This is a structural diagram of the EPI fusion module of the present invention;

[0035] Figure 4 This is a schematic diagram of SSIM comparison of experimental results of the embodiment;

[0036] Figure 5 Schematic diagram of PSNR comparison of experimental results in the embodiment. DETAILED DESCRIPTION

[0037] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0038] Embodiment 1:

[0039] This embodiment provides a light field image dense decoupling reconstruction method with multi-scale feature fusion, such as Figure 1 As shown, the following steps are included:

[0040] Step S1: sampling the light field image to obtain a sub-aperture image set, and sparsely sampling the sub-aperture image set to obtain a sparse sub-aperture image set;

[0041] Step S2: converting the sparse sub-aperture image set into a first macro-pixel image;

[0042] Step S3: inputting the first macro-pixel image into the first convolutional layer for feature extraction, and outputting a first feature map;

[0043] Step S4: input the first feature map into a plurality of identical and sequentially connected dense feature extractors for processing, and output a second feature map;

[0044] Step S5: passing the second feature map through an angle feature extractor and an upsampling module to output a second macro-pixel image;

[0045] Step S6: converting the second macro-pixel image into a reconstructed sub-aperture image set, where the reconstructed sub-aperture image set is the same as the sub-aperture image set obtained by sampling the light field image.

[0046] Among them, the first convolutional layer is a convolutional layer with a convolution kernel of 3×3, a stride of 1, an expansion of 2, and a padding of 2.

[0047] Among them, the dense feature extractor includes a spatial feature extractor, an angle feature extractor, a dense spatial feature extractor, a polar plane feature fusion feature extractor, and a first fusion convolution.

[0048] Wherein, step S4 comprises the following steps:

[0049] The first feature map is input into a plurality of dense feature extractors connected in sequence for processing, each dense feature extractor outputs a fused feature map, and the fused feature map is input into a plurality of subsequent dense feature extractors, and the second feature map is outputted through the last dense feature extractor.

[0050] The first feature map is input into a plurality of sequentially connected dense feature extractors for processing, and each dense feature extractor outputs a fused feature map, including the following steps:

[0051] Input the first feature map into the spatial feature extractor of the dense feature extractor to extract spatial features and output the spatial feature map. Input the spatial feature map into the angle feature extractor and output the angle feature map. Input the spatial feature map into the dense spatial feature extractor and output the local spatial feature map. Input the spatial feature map into the polar plane feature fusion feature extractor and output the polar plane feature map. Input the spatial feature map, the angle feature map, the local spatial feature map and the polar plane feature map into the first fused convolution and output the first fused feature map. Input the first fused feature map into the spatial feature extractor. Perform a residual connection between the feature map processed by the spatial feature extractor and the first fused feature map and output the fused feature map.

[0052] Among them, the spatial feature extractor includes two layers of identical dilated convolutions, with the convolution kernel size of 3×3, stride of 1, dilation rate of 2, and padding of 2.

[0053] The spatial feature map is input into the angle feature extractor, and the angle feature map is output, which specifically includes the following steps:

[0054] The spatial feature map is input into a convolution with an output channel number of Channel, a convolution kernel size of 2×2, a stride of 2, and a padding of 0 for dimensionality reduction processing to extract angle-related local features. The processed feature map is processed by a convolution with an output channel number of 4×Channel, a convolution kernel size of 1×1, and a stride of 1. The output feature map is processed by the PixelShuffle operation to output the angle feature map.

[0055] Among them, the dense spatial feature extractor includes two convolutions with the same output channel number of channels, convolution kernel size of 5×5, stride of 1, padding of 2, and expansion rate of 1. Nonlinear changes are introduced between the two convolution layers through the LeakyReLU activation function.

[0056] The polar plane feature fusion feature extractor includes a horizontal polar plane feature extractor, a vertical polar plane feature extractor and a second fusion convolution, inputs the spatial feature map into the polar plane feature fusion feature extractor, and outputs the polar plane feature map, including the following steps:

[0057] The spatial feature map is input into the horizontal polar plane feature extractor, the disparity information in the horizontal direction is extracted through the convolution operation, the disparity information is decoded to the horizontal direction through PixelShuffle1D, and the horizontal disparity information is output;

[0058] The spatial feature map is input into the vertical polar plane feature extractor, the vertical polar plane feature extractor transposes the spatial feature map and extracts the vertical disparity information through convolution operation, decodes the disparity information to the vertical direction through PixelShuffle1D, and outputs the vertical disparity information;

[0059] The horizontal disparity information is spliced ​​with the vertical disparity information, and the spliced ​​features are channel-fused through the second fusion convolution with an output channel number of 2×EpiChannel, a convolution kernel size of 1×1, a stride of 1, and a padding of 0 to output the epiplane feature map.

[0060] Among them, the extreme plane feature fusion feature extractor adopts a multi-scale attention mechanism.

[0061] Embodiment 2:

[0062] The parts not mentioned in this embodiment are the same as those in Embodiment 1.

[0063] This embodiment proposes a light field dense decoupling reconstruction method of multi-scale EPI fusion, including the following steps:

[0064] Step S1: sampling the original light field image into a sub-aperture image and performing sparse sampling, specifically comprising the following steps:

[0065] First, a sub-aperture image I is sampled from the original light field image L(u,v,s,t) S (x, y), where (u, v) represents the angular dimension of the light field and (s, t) represents the spatial dimension. S (x, y) is sparsely sampled to obtain a sparsely sampled light field sub-aperture image S spanse is a set of sparsely sampled view indices.

[0066] Step S2: Convert the acquired sparsely sampled light field sub-aperture image SAI into a macro-pixel image MacPI.

[0067] The specific process of step S2 is as follows:

[0068] Sub-aperture image I S (x, y), each s corresponds to an angular viewpoint. By arranging multiple sub-aperture images according to their viewing angle information (u, v), they can be converted into a macro-pixel image M(x, y, u, v) containing an angular dimension. The pixel value of each macro-pixel image M(x, y, u, v) comes from the value of the corresponding position (x, y) in the sub-aperture image set, where (u, v) = s. The conversion process integrates the angular information from the discrete sub-aperture images into the angular dimension of the macro-pixel image, forming a four-dimensional representation M(x, y, u, v).

[0069] Step S3: The macro pixel image information is passed through a convolution Inconv with a convolution kernel of 3×3, a step size of 1, an expansion of 2, and a padding of 2, and the first feature map is output;

[0070] Step S4: Input the first feature map into 16 identical dense feature extractors DFCG connected in sequence for processing. The specific steps are as follows:

[0071] like Figure 2 As shown, after the first feature map of the input dense feature extractor passes through a spatial feature extractor SFE, it passes through the angle feature extractor AFE, the dense spatial feature extractor DSSFE, and the polar plane feature fusion feature extractor respectively. The polar plane feature fusion feature extractor includes the horizontal polar plane feature extractor EPI-S and the vertical polar plane feature extractor EPI-T. The four extracted features are subjected to a fusion convolution, and the obtained feature map passes through a spatial feature extractor SFE. Finally, the feature map that passes through the spatial feature extractor SFE is residually connected with the original feature map. The final output result is input into the next dense feature extractor DFCG.

[0072] The specific process of each extractor is as follows:

[0073] The spatial feature extractor SFE consists of two layers of dilated convolutions. The first layer is convolved with a kernel size of 3×3, stride of 1, dilation rate of 2, and padding of 2. The second layer further processes the features with the same convolution operation parameters as the first layer. Both layers of convolution do not change the number of channels or resolution of the features. Finally, the features processed by the two layers of convolution are added to the fused features of the input through the residual connection SpaConv(buffer)+buffer, and the enhanced spatial features are output.

[0074] The angle feature extractor consists of three steps. In the first step, a convolution with an output channel number of Channel, a convolution kernel size of 2×2, a stride of 2, and padding of 0 is used to reduce the dimension of the input feature x in the angle dimension with a stride of 2 to extract angle-related local features. In the second step, the number of channels is increased by a convolution with an output channel number of 4×Channel, a convolution kernel size of 1×1, and a stride of 1, so that the feature dimension adapts to the angle decoding requirements. Finally, the angle information in the feature is reconstructed back to the angle dimension through the PixelShuffle operation, and the angle feature is output.

[0075] The dense spatial feature extraction module consists of two consecutive convolution operations. The first layer extracts local spatial features through a convolution with an output channel number of channels, a convolution kernel size of 5×5, a stride of 1, a padding of 2, and a dilation rate of 1, keeping the input and output feature dimensions consistent. The second layer uses a convolution operation with the same parameters as the first layer to process the features again for further extraction of local spatial information. Nonlinear changes are introduced between the two layers of convolution through the LeakyReLU activation function. The output of this module is dense local spatial features.

[0076] Extreme plane feature fusion feature extractor, such as Figure 3 As shown, it includes a horizontal polar plane feature extractor EPI-S and a vertical polar plane feature extractor EPI-T, and the specific structure is as follows:

[0077] In the horizontal direction, the input feature x is first convolved with an output channel number of Channel, a convolution kernel size of [1,4], a stride of [1,2], and a padding of [0,1] to extract the horizontal disparity information. The horizontal features are expanded through a convolution with an output channel number of 2×EpiChannel, a convolution kernel size of 1×1, a stride of 1, and a padding of 0, and then the disparity information is decoded to the horizontal direction using PixelShuffle1D, and finally a multi-scale attention mechanism is used. After transposing the input feature x in the vertical direction, the same operation process is repeated to finally output the vertical disparity information. Finally, the features extracted in the horizontal and vertical directions are spliced ​​through Concat, and the spliced ​​features are channel-fused through the convolution FuseEPI with an output channel number of 2×EpiChannel, a convolution kernel size of 1×1, a stride of 1, and a padding of 0.

[0078] PixelShuffle1D is a one-dimensional pixel rearrangement operation that is mainly used to upsample the last spatial dimension of the feature while keeping the overall feature structure consistent by reducing the channel dimension.

[0079] The multi-scale attention mechanism used in the extreme plane feature fusion feature extractor has the following structure:

[0080] The multi-scale attention mechanism is a multi-scale attention mechanism that combines channel attention and spatial attention to enhance the expressiveness of features. It first extracts the global information of the input features through global average pooling and maximum pooling, and uses three different scale convolution kernels (1×1, 3×3, 5×5) to extract multi-scale features in the channel and spatial dimensions respectively. In the channel attention part, the extracted multi-scale features are spliced ​​by channel, and channel weights are generated after two 1×1 convolutions, which are multiplied with the original features channel by channel to strengthen important channels. In the spatial attention part, the results of average pooling and maximum pooling are spliced, and multi-scale spatial features are extracted by convolution of three scales respectively. Then, spatial weights are generated by 1×1 convolution, and multiplied with features position by position to highlight important areas. Finally, the enhanced features of the channel and space are combined as the output, which is consistent with the input shape.

[0081] Step S5: Pass the second feature map output from step S4 through the output angle feature extractor, and adjust the number of output channels based on the angle feature extractor. The size of the downsampled feature map is one-fourth of the original size, and the number of channels is adjusted to 7x7xchannel1 to facilitate subsequent upsampling. Upsample the feature-processed image to obtain a restored macro-pixel image;

[0082] Step S6: Convert the restored macro-pixel image into a sub-aperture image to obtain a complete reconstructed image.

[0083] This embodiment compares the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) with LFEPICNN, ShearedEPI, Kalantari, Yeung, LFASR-geo, P4DCNN, FS-GAF, DistgASR and EASR respectively. All methods except LFEPICNN in the table are the results of retraining on the same dataset. Among them, LFEPICNN, ShearedEPI, and P4DCNN only use EPI to extract features for estimation, Kalantari, LFASR-geo, and FS-GAF use disparity-related methods for estimation, and Yeung, DistgASR, and EASR use disparity-independent methods for estimation.

[0084] Compared with other experiments, the PSNR Figure 4 As shown. The bold font represents the best effect, the underlined font represents the second best effect, and all data in the table are rounded to three decimal places. It can be seen that the best results were achieved in the HCIold, 30scenes, Occlusion, and Reflective test scenes, and it was only second to FS-GAF and EASR in the HCInew test scene. Compared with EASR, which had the best results before, most test scenes have been improved, among which the PSNR index in the HCIold scene has increased by 1.77dB at most. The HCI mean is second only to FS-GAF, and the STFlytro mean reaches 41.20, which is the highest mean. The above results show that the angle super-resolution method proposed in this embodiment has the best overall reconstruction effect.

[0085] Compared with the SSIM of other experiments Figure 5 As shown in the figure, the bold font is the best result, the underline is the second best result, and all the data in the table are rounded to three decimal places. The best results are achieved in the HCInew, 30scenes, and Occlusion test scenes. Compared with the EASR with the best results before, all scenes except Reflective are improved. This shows that the angle super-resolution method proposed in this embodiment reconstructs the view closest to the original image in structure and details.

[0086] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program code.

[0087] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A light field image dense decoupling reconstruction method with multi-scale feature fusion, characterized in that: The following steps are involved: Step S1: sampling the light field image to obtain a sub-aperture image set, and sparsely sampling the sub-aperture image set to obtain a sparse sub-aperture image set; Step S2: converting the sparse sub-aperture image set into a first macro-pixel image; Step S3: inputting the first macro-pixel image into the first convolutional layer for feature extraction, and outputting a first feature map; Step S4: input the first feature map into a plurality of identical and sequentially connected dense feature extractors for processing, and output a second feature map; Step S5: passing the second feature map through an angle feature extractor and an upsampling module to output a second macro-pixel image; Step S6: Convert the second macro-pixel image into a set of reconstructed sub-aperture images.

2. The method for dense decoupling and reconstruction of light field images with multi-scale feature fusion according to claim 1, characterized in that: The first convolutional layer is a convolutional layer with a convolution kernel of 3×3, a stride of 1, an expansion of 2, and a padding of 2.

3. The method for dense decoupling and reconstruction of light field images with multi-scale feature fusion according to claim 1, characterized in that: The dense feature extractor includes a spatial feature extractor, an angle feature extractor, a dense spatial feature extractor, an extreme plane feature fusion feature extractor, and a first fusion convolution.

4. The method for dense decoupling and reconstruction of light field images with multi-scale feature fusion according to claim 1, characterized in that: The step S4 comprises the following steps: The first feature map is input into a plurality of dense feature extractors connected in sequence for processing, each dense feature extractor outputs a fused feature map, and the fused feature map is input into a plurality of subsequent dense feature extractors, and the second feature map is outputted through the last dense feature extractor.

5. The method for dense decoupling and reconstruction of light field images with multi-scale feature fusion according to claim 1 or 3, characterized in that: The step of inputting the first feature map into a plurality of sequentially connected dense feature extractors for processing, wherein each dense feature extractor outputs a fused feature map, comprises the following steps: Input the first feature map into the spatial feature extractor of the dense feature extractor to extract spatial features and output the spatial feature map. Input the spatial feature map into the angle feature extractor and output the angle feature map. Input the spatial feature map into the dense spatial feature extractor and output the local spatial feature map. Input the spatial feature map into the polar plane feature fusion feature extractor and output the polar plane feature map. Input the spatial feature map, the angle feature map, the local spatial feature map and the polar plane feature map into the first fused convolution and output the first fused feature map. Input the first fused feature map into the spatial feature extractor. Perform a residual connection between the feature map processed by the spatial feature extractor and the first fused feature map and output the fused feature map.

6. The method for dense decoupling and reconstruction of light field images with multi-scale feature fusion according to claim 5, characterized in that: The spatial feature extractor includes two layers of identical dilated convolutions, wherein the dilated convolutions have a convolution kernel size of 3×3, a stride of 1, a dilation rate of 2, and a padding of 2.

7. The method for dense decoupling and reconstruction of light field images with multi-scale feature fusion according to claim 5, characterized in that: The step of inputting the spatial feature map into the angle feature extractor and outputting the angle feature map specifically includes the following steps: The spatial feature map is input into a convolution with an output channel number of Channel, a convolution kernel size of 2×2, a stride of 2, and a padding of 0 for dimensionality reduction processing to extract angle-related local features. The processed feature map is processed by a convolution with an output channel number of 4×Channel, a convolution kernel size of 1×1, and a stride of 1. The output feature map is processed by the PixelShuffle operation to output the angle feature map.

8. The method for dense decoupling and reconstruction of light field images with multi-scale feature fusion according to claim 5, characterized in that: The dense spatial feature extractor includes two convolutions with the same output channel number of channels, a convolution kernel size of 5×5, a stride of 1, a padding of 2, and a dilation rate of 1. Nonlinear changes are introduced between the two convolution layers through the LeakyReLU activation function.

9. The method for dense decoupling and reconstruction of light field images by multi-scale feature fusion according to claim 5, characterized in that: The polar plane feature fusion feature extractor includes a horizontal polar plane feature extractor, a vertical polar plane feature extractor and a second fusion convolution, and the spatial feature map is input into the polar plane feature fusion feature extractor to output the polar plane feature map, including the following steps: The spatial feature map is input into the horizontal polar plane feature extractor, the disparity information in the horizontal direction is extracted through the convolution operation, the disparity information is decoded to the horizontal direction through PixelShuffle1D, and the horizontal disparity information is output; The spatial feature map is input into the vertical polar plane feature extractor, the vertical polar plane feature extractor transposes the spatial feature map and extracts the vertical disparity information through convolution operation, decodes the disparity information to the vertical direction through PixelShuffle1D, and outputs the vertical disparity information; The horizontal disparity information is spliced ​​with the vertical disparity information, and the spliced ​​features are channel-fused through the second fusion convolution with an output channel number of 2×EpiChannel, a convolution kernel size of 1×1, a stride of 1, and a padding of 0 to output the epiplane feature map.

10. The method for dense decoupling and reconstruction of light field images with multi-scale feature fusion according to claim 9, characterized in that: The polar plane feature fusion feature extractor adopts a multi-scale attention mechanism.