Method for compressing light field quality enhancement based on spatial angle deformable convolution network

By combining spatially deformable convolutional networks and dense residual networks, the problem of distortion in light field compression is solved, and the quantification and visualization of light field quality are improved.

CN116934647BActive Publication Date: 2026-05-29SHANGHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI UNIV
Filing Date
2023-08-08
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing light field imaging technologies suffer from distortion and artifacts when compressed under hardware or transmission bandwidth limitations, making it difficult to maintain high-quality light field compression effects in low bit rate environments.

Method used

A spatial angle deformable convolutional network-based approach is adopted. By acquiring a compressed distortion light field dataset, local and global viewpoints are determined. Local and global features are extracted using the spatial angle deformable convolutional network, and after fusion processing, they are input into a dense residual network to improve the quality of the light field.

Benefits of technology

It improves the quantitative and visual indicators of optical field compression quality, enhances the quality of compressed optical fields, and overcomes distortion problems during transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116934647B_ABST
    Figure CN116934647B_ABST
Patent Text Reader

Abstract

The disclosure provides a compression light field quality enhancement method based on a spatial angle deformable convolutional network, which comprises: acquiring a compressed distorted light field dataset; determining a local viewpoint and a global viewpoint of the compressed distorted light field dataset according to the viewpoint position of the compressed distorted light field dataset; determining the local feature of the compressed distorted light field dataset according to the local viewpoint and the spatial angle deformable convolutional network; performing feature extraction processing on the global viewpoint to determine the global feature of the compressed distorted light field dataset; and inputting the local feature and the global feature after fusion processing into a preset dense residual network to determine a quality-enhanced compressed light field generated image. Through the disclosure, the spatial features and angle features of the compressed light field are implicitly aggregated by using the deformable convolution, and the global residual learning is introduced by using the dense residual network, so that the quality enhancement of the compressed light field is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of light field processing technology, and more specifically, to a method, system, medium, and electronic device for enhancing compressed light field quality based on spatially angular deformable convolutional networks. Background Technology

[0002] Light field imaging is a technique for enriching visual information. Existing light field imaging technologies employ a four-dimensional light field model, reconstructing immersive 3D scenes by recording spatial and angular information within the scene. Compared to traditional imaging, angular information is utilized. In light field imaging, a higher visual dimension indicates a greater ability to understand the scene from the light field data. On one hand, the light field can perform fundamental tasks such as depth / visual estimation and refocusing, and can also be applied to fields such as 3D reconstruction, image super-resolution, image segmentation, and image enhancement. However, on the other hand, because the light field captures additional ray direction information, the amount of light field data surges dramatically.

[0003] Currently, lossy compression of light fields can achieve a bit rate reduction of at least 80%, while lossless compression can achieve at least 60%. However, in practical applications such as security monitoring and real-time video, due to hardware or transmission bandwidth limitations, square signals require compression with larger quantization parameters, leading to compression distortion and artifacts. Furthermore, in lossy compression frameworks, the encoding and decoding of residuals distort the input 4D light field structure to some extent, affecting subsequent light field processing, such as depth estimation or super-resolution. Therefore, improving the compression quality of light fields in low bit rate environments and further reducing the bit rate without changing the current compression coefficient is of significant research importance. Moreover, overcoming the distortion generated during light field compression and transmission is crucial for the development of light field technology. Summary of the Invention

[0004] In view of the deficiencies in the prior art, the purpose of this disclosure is to provide a method, system, medium and electronic device for compressing optical field quality enhancement based on spatial angle deformable convolutional networks.

[0005] To achieve the above objectives, according to a first aspect of the present invention, a method for enhancing compressed optical field quality based on a spatially angularly deformable convolutional network is provided, comprising:

[0006] Obtain a compressed distortion light field dataset, which includes compressed light fields with multiple distortion levels;

[0007] Based on the viewpoint position of the compressed distortion light field dataset, determine the local viewpoint and global viewpoint of the compressed distortion light field dataset;

[0008] Based on the local viewpoint and spatial angle deformable convolutional network, the local features of the compressed distortion light field dataset are determined;

[0009] The global viewpoint is subjected to feature extraction processing to determine the global features of the compressed distortion light field dataset;

[0010] The local features and global features are fused and then input into a preset dense residual network to determine a quality-enhanced compressed light field generated image.

[0011] Optionally, determining the local and global viewpoints of the compressed distortion light field dataset based on the viewpoint positions of the compressed distortion light field dataset includes:

[0012] The viewpoints located above, below, to the left, to the right, to the upper left, to the upper right, to the lower left, and to the lower right of the viewpoint to be enhanced are identified as local viewpoints.

[0013] The viewpoints located at the middle row, middle column, diagonal line from the top left to the bottom right, and diagonal line from the top right to the bottom left in the compressed distorted light field dataset are determined as global viewpoints.

[0014] Optionally, determining the local features of the compressed distortion light field dataset based on the local viewpoint and spatial angle deformable convolutional network includes:

[0015] The local viewpoint is input into a spatially separable convolutional network to determine the prediction offset and modulation coefficients.

[0016] The predicted offset, the modulation coefficient, and the local viewpoint are input into the spatial angle deformable convolutional network to determine the local features of the compressed distortion light field dataset.

[0017] Optionally, the step of performing feature extraction processing on the global viewpoint to determine the global features of the compressed distortion light field dataset includes:

[0018] The global viewpoint located on the middle row of the compressed distorted light field dataset is input into the preset residual learning and residual network to determine the first global branch feature;

[0019] The global viewpoint located in the middle column of the compressed distortion light field dataset is input into the preset residual learning and residual network to determine the second global branch feature;

[0020] The global viewpoint located on the diagonal line from the top left to the bottom right corner of the compressed distortion light field dataset is input into the preset residual learning and residual network to determine the third global branch feature;

[0021] The global viewpoint located on the diagonal line from the upper right corner to the lower left corner of the compressed distortion light field dataset is input into the preset residual learning and residual network to determine the fourth global branch feature;

[0022] The first global branch feature, the second global branch feature, the third global branch feature, and the fourth global branch feature are fused to determine the global features of the compressed distortion light field dataset.

[0023] Optionally, the step of fusing the local features and the global features and then inputting the result into a preset dense residual network to determine a quality-enhanced compressed light field generated image includes:

[0024] The local features and global features located in the same channel dimension are fused to determine the fused features;

[0025] The fusion features are input into the preset dense residual network to determine the quality-enhanced compressed light field generated image.

[0026] Optionally, obtaining the compressed distortion light field dataset includes:

[0027] Based on multiple preset quantization parameters and the random access mode of the HEVC encoder, the light field is subjected to first encoding compression processing using an efficient video encoding method to determine the compressed distortion light field dataset.

[0028] Optionally, obtaining the compressed distortion light field dataset further includes:

[0029] Based on a set of multiple Lagrange multipliers, a multidimensional light field encoder is used to perform a second encoding compression process on the light field to determine the compressed distortion light field dataset.

[0030] According to a second aspect of this disclosure, a compressed optical field quality enhancement system based on a spatially angularly deformable convolutional network is provided, comprising:

[0031] The acquisition module is used to acquire a compressed distortion light field dataset, which includes compressed light fields with multiple distortion levels.

[0032] The viewpoint determination module is used to determine the local viewpoint and global viewpoint of the compressed distortion light field dataset based on the viewpoint position of the compressed distortion light field dataset.

[0033] The local feature determination module is used to determine the local features of the compressed distortion light field dataset based on the local viewpoint and the spatial angle deformable convolutional network.

[0034] A global feature determination module is used to perform feature extraction processing on the global viewpoint to determine the global features of the compressed distortion light field dataset;

[0035] The light field quality enhancement module is used to fuse the local features and the global features and then input them into a preset dense residual network to determine the quality-enhanced compressed light field generated image.

[0036] According to a third aspect of this disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the spatial angle deformable convolutional network quality enhancement method provided in the first aspect of this disclosure.

[0037] According to a fourth aspect of this disclosure, an electronic device is provided, comprising:

[0038] A memory on which computer programs are stored;

[0039] A processor is configured to execute the computer program in the memory to implement the steps of the spatial angle deformable convolutional network quality enhancement method provided in the first aspect of this disclosure.

[0040] Compared with the prior art, the embodiments of the present invention have at least one of the following beneficial effects:

[0041] The above technical solution involves acquiring a compressed distortion light field dataset, which includes compressed light fields with various distortion levels; determining the local and global viewpoints of the compressed distortion light field dataset based on its viewpoint location; determining the local features of the compressed distortion light field dataset based on the local viewpoints and a spatial angle deformable convolutional network; extracting features from the global viewpoints to determine the global features of the compressed distortion light field dataset; and fusing the local and global features and inputting them into a preset dense residual network to generate a quality-enhanced compressed light field image. This disclosure employs a spatially deformable convolutional network to implicitly aggregate the spatial and angular features of the light field, i.e., to aggregate the spatial and angular contextual information of the light field, so as to make full use of the spatial and angular information of the light field. Furthermore, based on global features, it prevents interference noise from being introduced into the position of the target viewpoint during the extraction of local features, and fuses local and global features to maintain the spatial structure of the compressed viewpoint through a multi-angle global feature network. It also employs a dense residual network to introduce global residual learning, realize network parameter sharing and feature reuse, retain and transmit fine-grained feature information, improve the network's understanding and expressive capabilities, enhance the quality of the compressed light field, and improve the quantitative and visual indicators of the compressed light field quality. Attached Figure Description

[0042] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0043] Figure 1 This is a flowchart illustrating a compressed light field quality enhancement method based on a spatially angularly deformable convolutional network, according to an exemplary embodiment.

[0044] Figure 2 This is a block diagram illustrating a compressed optical field quality enhancement system based on a spatial angle deformable convolutional network according to an exemplary embodiment.

[0045] Figure 3 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation

[0046] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention. These all fall within the scope of protection of the present invention.

[0047] Figure 1This is a flowchart illustrating a compressed optical field quality enhancement method based on a spatially angularly deformable convolutional network, according to an exemplary embodiment. Figure 1 As shown, a compressed optical field quality enhancement method based on spatial angle deformable convolutional networks includes S11 to S15.

[0048] S11, Obtain the compressed distortion light field dataset.

[0049] The compressed distortion light field dataset includes compressed light fields with various distortion levels. Obtaining the compressed distortion light field dataset involves encoding the light field dataset and determining the compressed distortion light field dataset corresponding to the encoding and its corresponding distortion type.

[0050] In one possible embodiment, a method for obtaining a compressed distortion light field dataset includes:

[0051] For High Efficiency Video Coding (HEVC), based on multiple preset quantization parameters and the random access mode of the HEVC encoder, the light field is subjected to the first encoding compression process using an efficient video coding method to determine the compressed distortion light field dataset.

[0052] Specifically, by setting four quantization parameters for the HEVC (High-Efficiency Video Coding) encoding light field, four distortion levels can be determined to encode the original light field and form a compressed distortion light field dataset. The preset quantization parameters can use QP = {32, 37, 42, 47}, and the random access mode of the HEVC encoder can be employed to determine the compressed light field at four distortion levels, thus forming the compressed distortion light field dataset.

[0053] In another possible embodiment, the method for obtaining the compressed distortion light field dataset further includes:

[0054] The Multidimensional Light Field Encoder (MULE) uses a multidimensional light field encoder to perform a second encoding and compression process on the light field based on multiple preset Lagrange multipliers, thereby determining the compressed and distorted light field dataset.

[0055] Specifically, for the Multidimensional Light Field Encoder (MULE), in the Rate-Distortion Optimization (RDO) process, four Lagrange multipliers can be set to determine four distortion levels, encoding the original light field and thus defining the compressed distortion light field dataset. The preset Lagrange multipliers can be λ = {10} 3 10 4 10 5 10 6This allows us to determine the compressed light fields at four distortion levels, forming a compressed distortion light field dataset.

[0056] In this disclosure, the light field dataset can be the light field dataset EPFL.

[0057] S12, Based on the viewpoint position of the compressed distortion light field dataset, determine the local viewpoint and global viewpoint of the compressed distortion light field dataset.

[0058] The compressed distortion light field dataset includes the viewpoint V to be enhanced. LQ (p) Local viewpoint and global viewpoint For each viewpoint to be enhanced, the compressed distortion light field dataset is divided into local viewpoints and global viewpoints according to the viewpoint location.

[0059] Local viewpoints are used to determine the local features of the compressed light field, including angular information specific to the compressed light field. Global viewpoints are used to determine the global features of the compressed light field and also to supplement spatial features.

[0060] S13. Based on the local viewpoint and spatial angle deformable convolutional network, determine the local features of the compressed distortion light field dataset.

[0061] Among them, the deformable convolution in the spatial angle deformable convolutional network is used to implicitly aggregate angular information in the light field to extract local features in the local viewpoint.

[0062] S14, perform feature extraction processing on the global viewpoint to determine the global features of the compressed distortion light field dataset.

[0063] The feature extraction process includes global branch feature extraction and global branch feature fusion. The global viewpoint can include multiple global viewpoint arrays. These arrays are input into a pre-defined residual learning and residual network to determine the global branch features corresponding to each array. The global branch features of each array are then fused to determine the global features of the compressed distortion light field dataset.

[0064] S15: After fusing local and global features, input the results into a preset dense residual network to determine the quality-enhanced compressed light field generated image.

[0065] The process involves fusing local and global features to form shallow features of the light field. Based on these shallow features, the light field is then input into a pre-defined dense residual network to perform refined quality enhancement. This process determines the quality-enhanced compressed light field, which in turn generates the image.

[0066] The above technical solution involves acquiring a compressed distortion light field dataset, which includes compressed light fields with various distortion levels; determining the local and global viewpoints of the compressed distortion light field dataset based on its viewpoint location; determining the local features of the compressed distortion light field dataset based on the local viewpoints and a spatial angle deformable convolutional network; extracting features from the global viewpoints to determine the global features of the compressed distortion light field dataset; and fusing the local and global features and inputting them into a preset dense residual network to generate a quality-enhanced compressed light field image. This disclosure employs a spatially deformable convolutional network to implicitly aggregate the spatial and angular features of the light field, i.e., to aggregate the spatial and angular contextual information of the light field, so as to make full use of the spatial and angular information of the light field. Furthermore, based on global features, it prevents interference noise from being introduced into the position of the target viewpoint during the extraction of local features, and fuses local and global features to maintain the spatial structure of the compressed viewpoint through a multi-angle global feature network. It also employs a dense residual network to introduce global residual learning, realize network parameter sharing and feature reuse, retain and transmit fine-grained feature information, improve the network's understanding and expressive capabilities, enhance the quality of the compressed light field, and improve the quantitative and visual indicators of the compressed light field quality.

[0067] In some possible embodiments, determining the local and global viewpoints of the compressed distortion light field dataset based on the viewpoint position of the compressed distortion light field dataset includes steps S21 to S22.

[0068] S21, the viewpoints located above, below, to the left, to the right, to the upper left, to the upper right, to the lower left, and to the lower right of the viewpoint to be enhanced are identified as local viewpoints.

[0069] As an example, the viewpoints of eight adjacent positions of the viewpoint to be enhanced are taken as local viewpoints. The eight adjacent positions include the upper adjacent position, the lower adjacent position, the left adjacent position, the right adjacent position, the upper left adjacent position, the upper right adjacent position, the lower left adjacent viewpoint position, and the lower right adjacent viewpoint position.

[0070] As another example, if the viewpoint to be enhanced is located at the edge of the light field, and the viewpoints at its eight adjacent locations may be missing to varying degrees, then the viewpoint to be enhanced will be added to the adjacent locations where viewpoints are missing as local viewpoints.

[0071] S22, the viewpoints located in the middle row, middle column, diagonal line from the top left corner to the bottom right corner, and diagonal line from the top right corner to the bottom left corner of the compressed distorted light field dataset are determined as global viewpoints.

[0072] The global viewpoint can include multiple global viewpoint arrays. The global viewpoints in the middle row can be used as one global viewpoint array, the global viewpoints in the middle column can be used as one global viewpoint array, the global viewpoints on the diagonal from the top left corner to the bottom right corner can be used as one global viewpoint array, and the global viewpoints on the diagonal from the top right corner to the bottom left corner can be used as one global viewpoint array.

[0073] As an example, if the total number of rows and columns of the viewpoint array in the compressed distortion light field dataset is odd, then the row containing the median of the total number of rows in the compressed distortion light field dataset is taken as the middle row, and the column containing the median of the total number of columns is taken as the middle column.

[0074] As another example, if the total number of rows and columns of the viewpoint array in the compressed distortion light field dataset is even, then the average of the light field viewpoints in the two rows whose median total number of rows is determined is used as the viewpoint in the middle row, and the average of the light field viewpoints in the two columns whose median total number of columns is determined is used as the viewpoint in the middle column. For example, if the total number of rows and columns of the light field viewpoints in the compressed distortion light field dataset is 8, then the average of the light field viewpoints in the fourth and fifth rows is used as the viewpoint in the middle row, and the average of the light field viewpoints in the fourth and fifth columns is used as the viewpoint in the middle column.

[0075] As another example, if the total number of rows in the viewpoint array of the compressed distortion light field dataset is even and the total number of columns is odd, then the average of the two rows of light field viewpoints with a median of a certain total number of rows in the compressed distortion light field dataset is taken as the viewpoint on the middle row, and the column containing the median of the total number of columns is taken as the middle column.

[0076] As another example, if the total number of rows in the viewpoint array of the compressed distortion light field dataset is odd and the total number of columns is even, then the row containing the median of the total number of rows in the compressed distortion light field dataset is taken as the middle row, and the average of the two light field viewpoints that determine the median of the total number of columns is taken as the viewpoint in the middle column.

[0077] The above technical solution, based on multi-angle global feature extraction, can overcome the problem of introducing noise from adjacent positions to the viewpoint to be enhanced into the location of the viewpoint during the extraction of local features, and maintain the spatial structure of the compressed viewpoint through a multi-angle global feature network.

[0078] In some possible embodiments, determining the local features of the compressed distortion light field dataset based on the local viewpoint and spatial angle deformable convolutional network includes steps S31 to S32.

[0079] S31 inputs the local viewpoint into the spatially separable convolutional network to determine the prediction offset and modulation coefficients.

[0080] The prediction offset represents the offset information of the compressed light field at each spatial location. The prediction offset is used to adjust the sampling position of the convolution kernel on the input feature image, and the modulation coefficient is used to adjust the sampling weight of the convolution kernel at different sampling positions.

[0081] The constructed local viewpoint is input into a spatially and angularly separable convolutional network to expand the receptive field and capture high-angle dynamics. Spatial features are extracted through spatial convolutional units, and angular features are extracted through angular convolutional units. Based on the features resulting from the interaction of spatial and angular features, offset prediction processing is performed to determine the predicted offset Δp of the deformable convolution. k and modulation coefficient Δm k .

[0082] S32 determines the local features of the compressed distortion light field dataset by predicting the offset, modulation coefficients, and local viewpoint input spatial angle deformable convolutional network.

[0083] Based on the predicted offset Δp k and modulation coefficient Δm k The position and size of the convolution kernels of a spatially angular deformable convolutional network are dynamically adjusted to obtain the angular changes and local features of a compressed distortion light field dataset.

[0084] As an example, local features of a compressed distortion light field dataset are determined using the following formula:

[0085]

[0086] Among them, F L (p) represents a local feature. p represents a local viewpoint k w represents a sampling grid with K sampling points. k Δp represents the weight of each sampling position p. k Indicates the predicted offset, Δm k This represents the modulation coefficient.

[0087] The above technical solution employs spatially deformable convolution to implicitly aggregate the spatial and angular information of the light field, thereby fully utilizing the spatial and angular information of the light field to obtain local features.

[0088] In some possible embodiments, the step of performing feature extraction processing on the global viewpoint to determine the global features of the compressed distortion light field dataset includes S41 to S45.

[0089] S41, input the global viewpoint located on the middle row of the compressed distortion light field dataset into the preset residual learning and residual network to determine the first global branch feature.

[0090] S42, input the global viewpoint located on the middle column of the compressed distortion light field dataset into the preset residual learning and residual network to determine the second global branch feature.

[0091] S43, input the global viewpoint located on the diagonal line from the top left to the bottom right corner of the compressed distortion light field dataset into the preset residual learning and residual network to determine the third global branch feature.

[0092] S44. Input the global viewpoint located on the diagonal line from the upper right corner to the lower left corner of the compressed distortion light field dataset into the preset residual learning and residual network to determine the fourth global branch feature.

[0093] Through the above steps S41 to S44, each set of global viewpoint arrays is input into a preset residual learning and residual network to extract the global branch features corresponding to each set of global viewpoint arrays.

[0094] S45, perform global branch feature fusion processing on the first global branch feature, the second global branch feature, the third global branch feature, and the fourth global branch feature to determine the global features of the compressed distortion light field dataset.

[0095] Following the example above, the global branch features corresponding to each set of global viewpoint arrays are subjected to global branch feature fusion processing to determine the global features F of the compressed distortion light field dataset. G (p).

[0096] The above technical solution enables the construction of a multi-angle global feature network based on a determined global viewpoint, so as to maintain the spatial structure of the compressed viewpoint using the global feature network.

[0097] In some possible embodiments, the step of fusing the local features with the global features and then inputting the result into a preset dense residual network to determine a quality-enhanced compressed light field generated image includes steps S51 to S52.

[0098] S51 performs feature fusion processing on local and global features located in the same channel dimension to determine the fused features.

[0099] Local features F L (p) and global feature F G (p) performs feature fusion processing along the channel dimension to integrate local features F located in the same channel dimension. L (p) and global feature F G (p), perform feature fusion, and determine the fused feature F(p), which is the shallow feature.

[0100] S52, input the fusion features into a preset dense residual network to determine the quality-enhanced compressed light field generated image.

[0101] Following the example above, for the fusion feature F(p), i.e., the shallow feature, it is input into a preset dense residual network for fine-grained quality enhancement of the compressed light field, and the quality-enhanced compressed light field is determined to generate the image V. EH (p).

[0102] In one possible embodiment, the refined quality enhancement of the compressed optical field based on fusion features includes:

[0103] Based on dense residual networks, and combining dense connections and residual networks, hierarchical features of fused features are obtained, and hierarchical features are fused to adaptively determine the local hierarchical features of fused features.

[0104] Local hierarchical features are fused to adaptively preserve hierarchical features globally.

[0105] Through the above technical solutions, dense links can ensure the full transmission and sharing of network features, and residual networks can promote the rapid propagation of gradients and network training, thereby introducing global residual learning, realizing network parameter sharing and feature reuse, preserving and transmitting fine-grained feature information, improving the network's understanding and expressive capabilities, thereby enhancing the quality of the light field, and further improving the quantitative and visual indicators of the light field quality.

[0106] Based on the same concept, this disclosure also provides a compressed optical field quality enhancement system based on a spatially angle-deformable convolutional network. Figure 2 This is a block diagram illustrating a compressed optical field quality enhancement system based on a spatially angularly deformable convolutional network, according to an exemplary embodiment. (Refer to...) Figure 2 The compressed light field quality enhancement system 100 based on spatial angle deformable convolutional networks includes: an acquisition module 110, a viewpoint determination module 120, a local feature determination module 130, a global feature determination module 140, and a light field quality enhancement module 150.

[0107] The acquisition module 110 is used to acquire a compressed distortion light field dataset, which includes compressed light fields with multiple distortion levels.

[0108] The viewpoint determination module 120 is used to determine the local viewpoint and global viewpoint of the compressed distortion light field dataset based on the viewpoint position of the compressed distortion light field dataset.

[0109] The local feature determination module 130 is used to determine the local features of the compressed distortion light field dataset based on the local viewpoint and spatial angle deformable convolutional network.

[0110] The global feature determination module 140 is used to perform feature extraction processing on the global viewpoint to determine the global features of the compressed distortion light field dataset;

[0111] The light field quality enhancement module 150 is used to fuse the local features and the global features and then input them into a preset dense residual network to determine the quality-enhanced compressed light field generated image.

[0112] The above technical solution involves acquiring a compressed distortion light field dataset, which includes compressed light fields with various distortion levels; determining the local and global viewpoints of the compressed distortion light field dataset based on its viewpoint location; determining the local features of the compressed distortion light field dataset based on the local viewpoints and a spatial angle deformable convolutional network; extracting features from the global viewpoints to determine the global features of the compressed distortion light field dataset; and fusing the local and global features and inputting them into a preset dense residual network to generate a quality-enhanced compressed light field image. This disclosure employs a spatially deformable convolutional network to implicitly aggregate the spatial and angular features of the light field, i.e., to aggregate the spatial and angular contextual information of the light field, so as to make full use of the spatial and angular information of the light field. Furthermore, based on global features, it prevents interference noise from being introduced into the position of the target viewpoint during the extraction of local features, and fuses local and global features to maintain the spatial structure of the compressed viewpoint through a multi-angle global feature network. It also employs a dense residual network to introduce global residual learning, realize network parameter sharing and feature reuse, retain and transmit fine-grained feature information, improve the network's understanding and expressive capabilities, enhance the quality of the compressed light field, and improve the quantitative and visual indicators of the compressed light field quality.

[0113] Optionally, the viewpoint determination module 120 includes:

[0114] The local viewpoint determination submodule is used to determine the viewpoints located above, below, to the left, to the right, to the top left, to the top right, to the bottom left, and to the bottom right of the viewpoint to be enhanced as local viewpoints.

[0115] The global viewpoint determination submodule is used to determine the viewpoints located in the middle row, middle column, diagonal line from the upper left corner to the lower right corner, and diagonal line from the upper right corner to the lower left corner of the compressed distorted light field dataset as global viewpoints.

[0116] Optionally, the local feature determination module 130 includes:

[0117] The first determining submodule is used to input the local viewpoint into a spatially separable convolutional network to determine the prediction offset and modulation coefficients.

[0118] The local feature determination submodule is used to input the predicted offset, the modulation coefficient, and the local viewpoint into the spatial angle deformable convolutional network to determine the local features of the compressed distortion light field dataset.

[0119] Optionally, the global feature determination module 140 includes:

[0120] The first global branch feature determination submodule is used to input the global viewpoint located on the middle row of the compressed distortion light field dataset into the preset residual learning and residual network to determine the first global branch feature;

[0121] The second global branch feature determination submodule is used to input the global viewpoint located on the middle column of the compressed distortion light field dataset into the preset residual learning and residual network to determine the second global branch feature;

[0122] The third global branch feature determination submodule is used to input the global viewpoint located on the diagonal line from the upper left corner to the lower right corner of the compressed distortion light field dataset into the preset residual learning and residual network to determine the third global branch feature;

[0123] The fourth global branch feature determination submodule is used to input the global viewpoint located on the diagonal line from the upper right corner to the lower left corner of the compressed distortion light field dataset into the preset residual learning and residual network to determine the fourth global branch feature;

[0124] The global feature determination submodule is used to perform global branch feature fusion processing on the first global branch feature, the second global branch feature, the third global branch feature, and the fourth global branch feature to determine the global features of the compressed distortion light field dataset.

[0125] Optionally, the light field quality enhancement module 150 includes:

[0126] The feature fusion determination submodule is used to perform feature fusion processing on the local features and the global features located in the same channel dimension to determine the fused features;

[0127] The light field quality enhancement submodule is used to input the fusion features into the preset dense residual network to determine the quality-enhanced compressed light field generated image.

[0128] Optionally, the acquisition module 110 is further configured to perform a first encoding compression process on the light field using an efficient video encoding method based on multiple preset quantization parameters and the random access mode of the HEVC encoder, thereby determining the compressed distortion light field dataset.

[0129] Optionally, the acquisition module 110 is further configured to perform a second encoding compression process on the light field using a multidimensional light field encoder based on a preset plurality of Lagrange multipliers, thereby determining the compressed distortion light field dataset.

[0130] Figure 3 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Figure 3 As shown, the electronic device 300 may include a processor 301 and a memory 302. The electronic device 300 may also include one or more of a multimedia component 303, an input / output interface 304, and a communication component 305.

[0131] The processor 301 controls the overall operation of the electronic device 300 to complete all or part of the steps in the spatial angle deformable convolutional network quality enhancement method of the first aspect described above. The memory 302 stores various types of data to support the operation of the electronic device 300. This data may include, for example, instructions for any application or method operating on the electronic device 300, and application-related data such as contact data, sent and received messages, pictures, audio, video, etc. The memory 302 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. Multimedia component 303 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 302 or transmitted via communication component 305. The audio component also includes at least one speaker for outputting audio signals. Input / output interface 304 provides an interface between processor 301 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 305 is used for wired or wireless communication between the electronic device 300 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IoT, eMTC, or other 5G technologies, or combinations thereof, is not limited here. Therefore, the corresponding communication component 305 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.

[0132] In another exemplary embodiment, a non-transitory computer-readable storage medium including program instructions is also provided. When executed by a processor, these program instructions implement the steps of the spatial angle-deformable convolutional network quality enhancement method of the first aspect described above. For example, the computer-readable storage medium may be the aforementioned memory including program instructions, which can be executed by a processor of an electronic device to perform the aforementioned spatial angle-deformable convolutional network quality enhancement method.

[0133] In another exemplary embodiment, a computer program product is also provided, the computer program product comprising a computer program executable by a programmable device, the computer program having a code portion for performing the above-described spatial angle deformable convolutional network quality enhancement method when executed by the programmable device.

[0134] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention. The above preferred features can be used in any combination without conflict.

Claims

1. A method for enhancing compressed optical field quality based on spatially angular deformable convolutional networks, characterized in that, include: Obtain a compressed distortion light field dataset, which includes compressed light fields with multiple distortion levels; Based on the viewpoint position of the compressed distortion light field dataset, determine the local viewpoint and global viewpoint of the compressed distortion light field dataset; Based on the local viewpoint and spatial angle deformable convolutional network, the local features of the compressed distortion light field dataset are determined; The global viewpoint is subjected to feature extraction processing to determine the global features of the compressed distortion light field dataset; The local features and global features are fused and then input into a preset dense residual network to determine a quality-enhanced compressed light field generated image. The step of determining the local features of the compressed distortion light field dataset based on the deformable convolutional network of local viewpoint and spatial angle includes: The local viewpoint is input into a spatially separable convolutional network to determine the prediction offset and modulation coefficients. The predicted offset, the modulation coefficient, and the local viewpoint are input into the spatial angle deformable convolutional network to determine the local features of the compressed distortion light field dataset. The step of performing feature extraction processing on the global viewpoint to determine the global features of the compressed distortion light field dataset includes: The global viewpoint located on the middle row of the compressed distorted light field dataset is input into the preset residual learning and residual network to determine the first global branch feature; The global viewpoint located in the middle column of the compressed distortion light field dataset is input into the preset residual learning and residual network to determine the second global branch feature; The global viewpoint located on the diagonal line from the top left to the bottom right corner of the compressed distortion light field dataset is input into the preset residual learning and residual network to determine the third global branch feature; The global viewpoint located on the diagonal line from the upper right corner to the lower left corner of the compressed distortion light field dataset is input into the preset residual learning and residual network to determine the fourth global branch feature; The first global branch feature, the second global branch feature, the third global branch feature, and the fourth global branch feature are fused to determine the global features of the compressed distortion light field dataset.

2. The method according to claim 1, characterized in that, The step of determining the local and global viewpoints of the compressed distortion light field dataset based on the viewpoint positions of the dataset includes: The viewpoints located above, below, to the left, to the right, to the upper left, to the upper right, to the lower left, and to the lower right of the viewpoint to be enhanced are identified as local viewpoints. The viewpoints located at the middle row, middle column, diagonal line from the top left to the bottom right, and diagonal line from the top right to the bottom left in the compressed distorted light field dataset are determined as global viewpoints.

3. The method according to claim 1, characterized in that, The step of fusing the local features with the global features and then inputting the result into a preset dense residual network to determine a quality-enhanced compressed light field generated image includes: The local features and global features located in the same channel dimension are fused to determine the fused features; The fusion features are input into the preset dense residual network to determine the quality-enhanced compressed light field generated image.

4. The method according to claim 1, characterized in that, The acquisition of the compressed distortion light field dataset includes: Based on multiple preset quantization parameters and the random access mode of the HEVC encoder, the light field is subjected to first encoding compression processing using an efficient video encoding method to determine the compressed distortion light field dataset.

5. The method according to claim 4, characterized in that, The process of obtaining the compressed distortion light field dataset also includes: Based on a set of multiple Lagrange multipliers, a multidimensional light field encoder is used to perform a second encoding compression process on the light field to determine the compressed distortion light field dataset.

6. A compressed optical field quality enhancement system based on a spatially angularly deformable convolutional network, characterized in that, include: The acquisition module is used to acquire a compressed distortion light field dataset, which includes compressed light fields with multiple distortion levels. The viewpoint determination module is used to determine the local viewpoint and global viewpoint of the compressed distortion light field dataset based on the viewpoint position of the compressed distortion light field dataset. The local feature determination module is used to determine the local features of the compressed distortion light field dataset based on the local viewpoint and the spatial angle deformable convolutional network. A global feature determination module is used to perform feature extraction processing on the global viewpoint to determine the global features of the compressed distortion light field dataset; The light field quality enhancement module is used to fuse the local features and the global features and then input them into a preset dense residual network to determine the quality-enhanced compressed light field generated image. The local feature determination module includes: The first determining submodule is used to input the local viewpoint into a spatially separable convolutional network to determine the prediction offset and modulation coefficients. The local feature determination submodule is used to input the predicted offset, the modulation coefficient, and the local viewpoint into the spatial angle deformable convolutional network to determine the local features of the compressed distortion light field dataset. The global feature determination module includes: The first global branch feature determination submodule is used to input the global viewpoint located on the middle row of the compressed distortion light field dataset into the preset residual learning and residual network to determine the first global branch feature; The second global branch feature determination submodule is used to input the global viewpoint located on the middle column of the compressed distortion light field dataset into the preset residual learning and residual network to determine the second global branch feature; The third global branch feature determination submodule is used to input the global viewpoint located on the diagonal line from the upper left corner to the lower right corner of the compressed distortion light field dataset into the preset residual learning and residual network to determine the third global branch feature; The fourth global branch feature determination submodule is used to input the global viewpoint located on the diagonal line from the upper right corner to the lower left corner of the compressed distortion light field dataset into the preset residual learning and residual network to determine the fourth global branch feature; The global feature determination submodule is used to perform global branch feature fusion processing on the first global branch feature, the second global branch feature, the third global branch feature, and the fourth global branch feature to determine the global features of the compressed distortion light field dataset.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-5.

8. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Image enhancement progressive fusion method for infrared light field equipment

    CN114757862A

  • Multi-mode aggregation low-light environment behavior recognition method and system based on feature guidance

    CN115565248A