Light field image perceptual coding method based on intra coding tree unit level code rate allocation
By selecting corner sub-aperture images for compression in light field image coding and using a light field angle super-resolution reconstruction network to synthesize images at other locations, combined with a new intra-frame coding tree unit-level bitrate allocation strategy, the problem of perceptual redundancy in light field image coding is solved, coding efficiency is improved, and visual effects and structural consistency of significant regions are maintained.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NINGBO UNIV
- Filing Date
- 2023-06-01
- Publication Date
- 2026-07-14
AI Technical Summary
Existing light field image coding methods are insufficient in removing perceptual redundancy and maintaining the visual quality and structural consistency of salient regions, especially in maintaining the perceptual quality and structural consistency of salient regions, where there is still room for improvement.
A light field image sensing coding method based on intra-frame coding tree unit-level bitrate allocation is adopted. The sub-aperture images at the four corners of the light field image are selected for compression, and the sub-aperture images at other positions are synthesized by the light field angle super-resolution reconstruction network at the decoding end. At the same time, a new intra-frame coding tree unit-level bitrate allocation strategy is designed to allocate the bitrate considering the characteristics of human eye perception.
It improves coding efficiency, reduces perceptual redundancy, and maintains visual quality and structural consistency in salient regions. Experimental results show that it saves a significant amount of bitrate while maintaining the same visual quality.
Smart Images

Figure CN116781933B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a light field image coding method, and more particularly to a light field image perceptual coding method based on intra-frame coding tree unit-level bitrate allocation. Background Technology
[0002] Light is one of the most important mediums for humans to perceive the natural world. Light field imaging, as an emerging imaging technology, can simultaneously capture both the intensity and direction of light, and is attracting widespread attention from industry and academia. Meanwhile, many industrial and commercial applications of light field images are being developed, such as post-capture refocusing, virtual reality displays, and depth estimation. However, the rich scene information makes the data volume of light field images far greater than that of ordinary images at the same resolution. Therefore, efficient compression of light field images is necessary.
[0003] Existing light field image coding methods can be divided into two categories:
[0004] The first category is compression methods based on traditional encoders. Some of these methods directly utilize traditional encoders to compress light field images, such as directly encoding lens images using existing image compression techniques, or representing the light field image as a series of 2D sub-aperture images (SAIs), arranging the SAIs into a pseudo-video sequence, and then compressing it using a video encoder. However, existing encoders do not consider the unique 4D structure of light field images. Therefore, subsequent researchers have improved existing encoders, such as introducing new prediction methods into existing video encoders to make them more suitable for light field image compression. Although these methods can remove most of the redundancy, encoding all light field data limits compression efficiency.
[0005] The second category is viewpoint synthesis-based methods. These methods compress only a subset of viewpoints (SAIs) and synthesize the remaining viewpoints at the decoder to reconstruct the complete light field image. Bakir et al. used the next-generation video coding standard VVC to compress sparsely sampled SAIs, and the decoder used an adversarial generative network with dual discriminators to reconstruct the complete light field image. Huang et al. only compressed the depth maps of the selected and unselected SAIs, and the decoder used the depth maps to draw the unselected SAIs. Liu et al. compressed eight selected SAIs and proposed a multi-disparity geometry for learning the multi-stream reconstruction network, achieving good reconstruction results. These methods further remove angular redundancy in the light field image, improving coding efficiency. However, the selected SAIs are usually encoded using video coding techniques. Existing video encoders' intra-frame coding tree-based unit-level bitrate allocation algorithms do not fully consider visual perception characteristics, leading to perceptual redundancy in the compressed SAI subset. Since these SAIs will be used as references on the decoder side, this perceptual redundancy will be further transmitted to the synthesized SAIs.
[0006] In summary, although current research has achieved good light field image compression results, there are still some shortcomings in removing perceptual redundancy. In particular, there is still room for improvement in maintaining the perceptual quality and structural consistency of salient regions. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a light field image perceptual coding method based on intra-frame coding tree unit-level bit rate allocation, which can improve coding efficiency, remove perceptual redundancy, and maintain the visual quality and structural consistency of significant regions.
[0008] The technical solution adopted by this invention to solve the above-mentioned technical problems is: a light field image sensing coding method based on intra-frame coding tree unit-level bitrate allocation, characterized by including the following steps:
[0009] Step 1: At the encoding end, select only a portion of the sub-aperture images from the sub-aperture image array used to represent the light field image; then arrange all the selected sub-aperture images to form a pseudo-video sequence, with each sub-aperture image in the pseudo-video sequence serving as the sub-aperture image to be encoded.
[0010] Step 2: Select the central sub-aperture image from the sub-aperture image array used to represent the light field image; then calculate the initial bitrate allocation weight of each coding tree unit in each sub-aperture image to be encoded in the pseudo-video sequence, and denote the initial bitrate allocation weight of the i-th coding tree unit in any sub-aperture image to be encoded as T. i , Where 1≤i≤Num, the total number of coding tree units in the sub-aperture image to be coded is the same as the total number of coding tree units in the central sub-aperture image, both being Num units. The length of the coding tree units in both the sub-aperture image to be coded and the central sub-aperture image is M, and the width of both the sub-aperture image to be coded and the central sub-aperture image is N. (x,y) represents the coordinate position of a pixel in the coding tree unit of the central sub-aperture image within its corresponding coding tree unit. G i (x,y) represents the gradient value of the pixel at coordinate (x,y) in the i-th coding tree unit of the central sub-aperture image, where c is a constant greater than 0;
[0011] Step 3: Estimate the depth map of the central sub-aperture image using a depth estimation network; then binarize the depth map of the central sub-aperture image to obtain the foreground mask of the depth map of the central sub-aperture image, denoted as mask. f ; then calculate the mask f The foreground density of each coding tree unit in the mask f The foreground density of the i-th coding tree unit is denoted as in, Indicates mask f The number of pixels with a value of 1 in the i-th coding tree unit, mask f The length of the coding tree unit in the mask is also M. f The width of the coding tree unit in the mask is also N. f The number of pixels in any coding tree unit is M×N;
[0012] A saliency detection network is used to detect the saliency map of the centroidal aperture image. Then, the saliency map of the centroidal aperture image is binarized to obtain a salient object mask of the saliency map of the centroidal aperture image, denoted as mask. s ; then calculate the mask s The salient object density of each coding tree unit in the mask s The salient object density of the i-th coding tree unit is denoted as in, Indicates mask s The number of pixels with a value of 1 in the i-th coding tree unit, mask s The length of the coding tree unit in the mask is also M. s The width of the coding tree unit in the mask is also N. s The number of pixels in any coding tree unit is also M×N;
[0013] Step 4: Calculate the final bitrate allocation weight of each coding tree unit in each sub-aperture image to be coded in the pseudo-video sequence. Let W be the final bitrate allocation weight of the i-th coding tree unit in any sub-aperture image to be coded. i , Where α>1, β>1;
[0014] Step 5: Intra-frame coding of the pseudo-video sequence using a standard video encoder. For any sub-aperture image to be encoded in the pseudo-video sequence, treat it as the current frame; traverse each coding tree unit in the current frame, and for the i-th coding tree unit in the current frame, use the final bitrate of that coding tree unit to assign weight W. i Assign a target bitrate to this coding tree unit, and denote the target bitrate of the i-th coding tree unit in the current frame as R. i , The target bitrate allocation for each coding tree unit in all sub-aperture images to be encoded in the pseudo-video sequence is completed in the same manner as described above, and the remaining encoding steps are completed by a standard video encoder; where R p R represents the total number of target bits in the current frame. h R represents the actual number of bits encoded in the frame header information of the current frame. c W represents the number of bits consumed by the encoded coding tree units in the current frame. k This represents the final bitrate allocation weight of the k-th coding tree unit in the current frame;
[0015] Step 6: At the decoding end, a decoding terminal aperture image array is obtained by decoding all the sub-aperture images. The decoding terminal aperture image array is obtained by combining two parts. The first part is all the sub-aperture images obtained by decoding, and the second part is all the sub-aperture images at other positions except for all the sub-aperture images obtained by decoding, which are synthesized by inputting all the sub-aperture images obtained by decoding into the optical field angle super-resolution reconstruction network.
[0016] In step 1, the selected sub-aperture images are four sub-aperture images located at the four corners of the sub-aperture image array.
[0017] In step 1, the four selected sub-aperture images are arranged in the order of the top left sub-aperture image, the top right sub-aperture image, the bottom left sub-aperture image, and the bottom right sub-aperture image to form a pseudo-video sequence.
[0018] In step 2, M = 64 and N = 64.
[0019] In step 2, G i (x,y)=|p i (x,y)-pi (x+1,y)|+|p i (x,y)-p i (x,y+1)|, where p i (x,y) represents the Y component value of the pixel at coordinate (x,y) in the i-th coding tree unit of the central sub-aperture image, p i (x+1,y) represents the Y component value of the pixel at coordinate (x+1,y) in the i-th coding tree unit of the central sub-aperture image, p i (x,y+1) represents the Y component value of the pixel at coordinate position (x,y+1) in the i-th coding tree unit of the central sub-aperture image, and the symbol "||" is the absolute value symbol.
[0020] In step 3, the binarization method used is the Otsu binarization method.
[0021] In step 6, the decoded images are the sub-aperture images at the four corner positions. The decoded sub-aperture images at the four corner positions are input into the optical field angle super-resolution reconstruction network to synthesize all sub-aperture images at the positions other than the four corner positions.
[0022] The position index numbers of the coding tree units at the same position in the foreground mask of the sub-aperture image to be encoded, the central sub-aperture image, and the salient object mask of the depth map are consistent.
[0023] Compared with the prior art, the advantages of the present invention are as follows:
[0024] This invention takes into account the unique 4D structure of light field images, selecting only the sub-aperture images at the four corners for compression and transmission. At the decoding end, a light field angle super-resolution network is used to synthesize sub-aperture images at other locations to improve coding efficiency. Simultaneously, this invention designs a novel intra-frame coding tree unit-level bitrate allocation strategy. Using the central sub-aperture image, a bitrate allocation weight that considers human visual perception characteristics is calculated for each coding tree unit in the selected four corner sub-aperture images, and applied during the encoding process to remove perceptual redundancy in the light field image, further improving coding efficiency. Experimental results show that this invention improves coding efficiency while maintaining better visual effects and structural consistency in salient areas of the image. Attached Figure Description
[0025] Figure 1 This is a flowchart illustrating the overall implementation process of the method of the present invention;
[0026] Figure 2a This is a schematic diagram showing the positions of the four selected sub-aperture images (shown in gray) in the sub-aperture image array;
[0027] Figure 2b A schematic diagram showing the position of the central sub-aperture image (shown in gray) within the sub-aperture image array;
[0028] Figure 3a The image is the original centroidal aperture image of the Caution_Bees light field image, with a non-salient region and a salient region outlined.
[0029] Figure 3b.1 for Figure 3a Enlarged view of the non-significant regions in the image;
[0030] Figure 3b.2 for Figure 3a Enlarged view of the significant areas in the image;
[0031] Figure 3c.1 This is a magnified view of the non-significant region in the decoded central sub-aperture image after encoding the Caution_Bees light field image using the HEVC-AI method at a bit rate of 0.198 bpp.
[0032] Figure 3c.2 This is a magnified view of a significant region in the decoded central sub-aperture image after encoding the Caution_Bees light field image using the HEVC-AI method at a bit rate of 0.198 bpp.
[0033] Figure 3d.1 This is a magnified view of the non-significant region in the decoded central sub-aperture image after encoding the Caution_Bees light field image using the HEVC-ASR method at a bit rate of 0.016 bpp.
[0034] Figure 3d.2 This is a magnified view of a significant region in the decoded central sub-aperture image after encoding the Caution_Bees light field image using the HEVC-ASR method at a bit rate of 0.016 bpp.
[0035] Figure 3e.1 This is a magnified view of the non-significant region in the decoded central sub-aperture image after encoding the Caution_Bees light field image at a bit rate of 0.016 bpp using the method of this invention.
[0036] Figure 3e.2 This is a magnified view of a significant region in the decoded central sub-aperture image after encoding the Caution_Bees light field image at a bit rate of 0.016 bpp using the method of this invention.
[0037] Figure 4a Comparison of EPI consistency in significant regions of Danger_de_Mort light field images using the HEVC-AI method, HEVC-ASR method, and the method of this invention;
[0038] Figure 4b Comparison of EPI consistency in significant regions of Fountain_&_Vincent_1 light field images using the HEVC-AI method, HEVC-ASR method, and the method of this invention;
[0039] Figure 4c A comparison of EPI consistency in salient regions of Stone_Pillars_Outside light field images using the HEVC-AI method, the HEVC-ASR method, and the method of this invention. Detailed Implementation
[0040] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0041] The present invention proposes a light field image sensing coding method based on intra-frame coding tree unit-level bitrate allocation, the overall implementation flowchart of which is as follows: Figure 1 As shown, it includes the following steps:
[0042] Step 1: At the encoding end, select only a portion of the sub-aperture images from the sub-aperture image array used to represent the light field image; then arrange all the selected sub-aperture images to form a pseudo-video sequence, with each sub-aperture image in the pseudo-video sequence serving as the sub-aperture image to be encoded.
[0043] In this embodiment, in step 1, the selected sub-aperture images are four sub-aperture images located at the four corners of the sub-aperture image array. Figure 2a The positions of the four selected sub-aperture images (shown in gray) are shown in the sub-aperture image array.
[0044] In this embodiment, in step 1, the four selected sub-aperture images are arranged in the order of the top left sub-aperture image, the top right sub-aperture image, the bottom left sub-aperture image, and the bottom right sub-aperture image to form a pseudo-video sequence. In actual implementation, the order in which these four sub-aperture images are arranged has little impact on the coding performance of the method of the present invention.
[0045] Step 2: Select the central sub-aperture image from the sub-aperture image array used to represent the light field image. Figure 2b The position of the central sub-aperture image (shown in gray) in the sub-aperture image array is shown; then, the initial bitrate allocation weight of each coding tree unit in each sub-aperture image to be encoded in the pseudo-video sequence is calculated, and the initial bitrate allocation weight of the i-th coding tree unit in any sub-aperture image to be encoded is denoted as T. i , Where 1≤i≤Num, the total number of coding tree units in the sub-aperture image to be coded is the same as the total number of coding tree units in the central sub-aperture image, both being Num units. The length of the coding tree units in both the sub-aperture image to be coded and the central sub-aperture image is M, and the width of both the sub-aperture image to be coded and the central sub-aperture image is N. (x,y) represents the coordinate position of a pixel in the coding tree unit of the central sub-aperture image within its corresponding coding tree unit. G i (x,y) represents the gradient value of the pixel at coordinate (x,y) in the i-th coding tree unit of the central sub-aperture image, where c is a constant greater than 0, such as c = 100.
[0046] In this embodiment, in step 2, M = 64 and N = 64.
[0047] In this embodiment, in step 2, G i (x,y)=|p i (x,y)-p i (x+1,y)|+|p i (x,y)-p i (x,y+1)|, where p i (x,y) represents the Y component value, i.e., the brightness value, of the pixel at coordinate (x,y) in the i-th coding tree unit of the central sub-aperture image. i (x+1,y) represents the Y component value, i.e., the brightness value, of the pixel at coordinate (x+1,y) in the i-th coding tree unit of the central sub-aperture image. i (x, y+1) represents the Y component value, i.e., the brightness value, of the pixel at coordinate position (x, y+1) in the i-th coding tree unit of the central sub-aperture image. The symbol "||" is the absolute value symbol. The formula for calculating the Y component value is Y = 0.299 × R + 0.587 × G + 0.114 × B, where R, G, and B represent the R component value, G component value, and B component value of the pixel, respectively.
[0048] Step 3: Estimate the depth map of the central sub-aperture image using an existing depth estimation network; then binarize the depth map of the central sub-aperture image to obtain the foreground mask of the depth map of the central sub-aperture image, denoted as mask. f ; then calculate the mask f The foreground density of each coding tree unit in the mask f The foreground density of the i-th coding tree unit is denoted as in, Indicates mask fThe number of pixels with a value of 1 in the i-th coding tree unit, mask f The length of the coding tree unit in the mask is also M. f The width of the coding tree unit in the mask is also N. f The number of pixels in any coding tree unit is M×N.
[0049] The saliency map of the centroid aperture image is detected using an existing saliency detection network. Then, the saliency map of the centroid aperture image is binarized to obtain a salient object mask, denoted as mask. s ; then calculate the mask s The salient object density of each coding tree unit in the mask s The salient object density of the i-th coding tree unit is denoted as in, Indicates mask s The number of pixels with a value of 1 in the i-th coding tree unit, mask s The length of the coding tree unit in the mask is also M. s The width of the coding tree unit in the mask is also N. s The number of pixels in any coding tree unit is also M×N.
[0050] In practice, any mature network structure from existing depth estimation networks can be selected for depth map estimation, and similarly, any mature network structure from existing saliency detection networks can be selected for saliency map detection.
[0051] In this embodiment, the binarization method used in step 3 is the Otsu binarization method.
[0052] Step 4: Calculate the final bitrate allocation weight of each coding tree unit in each sub-aperture image to be coded in the pseudo-video sequence. Let W be the final bitrate allocation weight of the i-th coding tree unit in any sub-aperture image to be coded. i , Where α > 1 and β > 1, in this embodiment, α = 1.1 and β = 1.5 are taken.
[0053] Step 5: Intra-frame coding of the pseudo-video sequence using a standard video encoder. For any sub-aperture image to be encoded in the pseudo-video sequence, treat it as the current frame; traverse each coding tree unit in the current frame, and for the i-th coding tree unit in the current frame, use the final bitrate of that coding tree unit to assign weight W. iAssign a target bitrate to this coding tree unit, and denote the target bitrate of the i-th coding tree unit in the current frame as R. i , The target bitrate allocation for each coding tree unit in all sub-aperture images to be encoded in the pseudo-video sequence is completed in the same manner as described above, and the remaining encoding steps are completed by a standard video encoder; where R p R represents the total number of target bits in the current frame. h R represents the actual number of bits encoded in the frame header information of the current frame. c W represents the number of bits consumed by the encoded coding tree units in the current frame. k This represents the final bitrate allocation weight of the k-th coding tree unit in the current frame. The standard video encoder used in this implementation is the HEVC official encoder.
[0054] Step 6: At the decoding end, a decoding terminal aperture image array is obtained by decoding all the sub-aperture images. The decoding terminal aperture image array is obtained by combining two parts. The first part is all the sub-aperture images obtained by decoding, and the second part is all the sub-aperture images at other positions except for all the sub-aperture images obtained by decoding, which are synthesized by inputting all the sub-aperture images obtained by decoding into the existing optical field angle super-resolution reconstruction network.
[0055] In this embodiment, in step 6, the sub-aperture images at the four corner positions are decoded; the decoded sub-aperture images at the four corner positions are input into the existing optical field angle super-resolution reconstruction network to synthesize all sub-aperture images at the positions other than the four corner positions.
[0056] In this embodiment, the position index numbers of the coding tree units at the same position in the foreground mask of the sub-aperture image to be encoded, the central sub-aperture image and its depth map and the salient object mask of the salience map are consistent, that is: the i-th coding tree unit in the sub-aperture image to be encoded, the i-th coding tree unit in the central sub-aperture image, the i-th coding tree unit in the foreground mask of the depth map of the central sub-aperture image, and the i-th coding tree unit in the salient object mask of the salience map of the central sub-aperture image are at the same position.
[0057] To further illustrate the performance of the method of the present invention, the method of the present invention was tested.
[0058] To evaluate the effectiveness of the method of this invention, experiments were conducted in the HEVC reference software HM16.20. Light field images from the EPFL light field database, namely “Caution_Bees”, “Danger_de_Mort”, “Fountain_&_Vincent_1”, “Stone_Pillars_Outside”, “Sophie_&_Vincent_on_a_Bench”, and “Sophie_Krios_&_Vincent”, were selected as standard test images. The experiment used only a 7×7 sub-aperture image array at the center, and the spatial resolution was cropped to 432×624. Code tree unit-level bitrate control was enabled, and the bitrates at HEVC coding method QP values of 22, 27, 32, and 37 were collected as target bitrates.
[0059] Table 1 lists the specific information of the light field images “Caution_Bees”, “Danger_de_Mort”, “Fountain_&_Vincent_1”, “Stone_Pillars_Outside”, “Sophie_&_Vincent_on_a_Bench”, and “Sophie_Krios_&_Vincent” used in the experiment.
[0060] Table 1 shows the specific information of the tested light field images.
[0061] Light field image Angular resolution Spatial resolution Caution_Bees 7×7 432×624 Danger_de_Mort 7×7 432×624 Fountain_&_Vincent_1 7×7 432×624 Stone_Pillars_Outside 7×7 432×624 Sophie_&_Vincent_on_a_Bench 7×7 432×624 Sophie_Krios_&_Vincent 7×7 432×624
[0062] To illustrate the performance of the method of the present invention, it is compared with two light field image coding methods. The first light field image coding method arranges all sub-aperture images into a pseudo-video sequence and encodes them using HEVC full intra-frame mode, abbreviated as HEVC-AI. The second light field image coding method is an encoding method that removes the new intra-frame coding tree unit-level bitrate allocation strategy in the method of the present invention and uses the HEVC bitrate allocation strategy instead, abbreviated as HEVC-ASR.
[0063] Table 2 lists the coding performance comparison of the light field images listed in Table 1 encoded using the method of this invention, as well as HEVC-AI and HEVC-ASR. Image quality is objectively evaluated using Y-PPSNR (Perceptual PSNR) and VSI (Visual Saliency-Induced Index), and coding performance is evaluated using BD-BR (…). Measured by Delta Bit Rate.
[0064] Table 2 compares the coding performance of the light field images listed in Table 1 using the method of this invention, as well as HEVC-AI and HEVC-ASR.
[0065]
[0066] As shown in Table 2, the novel intra-coding tree unit-level bitrate allocation strategy proposed in this invention achieves average BD-BR savings of 13.676% and 2.045% respectively when Y-PPSNR and VSI are used as quality evaluation metrics. Compared with the HEVC-AI method, this invention achieves over 90% average BD-BR savings across both quality evaluation metrics. This demonstrates that this invention can save a significant amount of bitrate while maintaining the same visual quality.
[0067] Figure 3a The original centroidal aperture image of the Caution_Bees light field image is given, and a non-salient region and a salient region are outlined. Figure 3b.1 Given Figure 3a Enlarged view of the non-significant regions in the image. Figure 3b.2 Given Figure 3a Enlarged view of the significant areas in the image; Figure 3c.1 Enlarged images of insignificant regions in the decoded central sub-aperture image after encoding the Caution_Bees light field image using the HEVC-AI method at a bit rate of 0.198 bpp are presented. Figure 3c.2 Enlarged views of significant regions in the decoded central sub-aperture image after encoding the Caution_Bees light field image using the HEVC-AI method at a bit rate of 0.198 bpp are presented. Figure 3d.1 Enlarged images of insignificant regions in the decoded central sub-aperture image after encoding the Caution_Bees light field image using the HEVC-ASR method at a bit rate of 0.016 bpp are presented. Figure 3d.2 Enlarged views of significant regions in the decoded central sub-aperture image after encoding the Caution_Bees light field image using the HEVC-ASR method at a bit rate of 0.016 bpp are presented. Figure 3e.1 Enlarged images of insignificant regions in the decoded central sub-aperture image after encoding the Caution_Bees light field image using the method of this invention at a bit rate of 0.016 bpp are provided. Figure 3e.2 Enlarged views of significant regions in the decoded central sub-aperture image after encoding the Caution_Bees light field image at a bit rate of 0.016 bpp using the method of this invention are provided. Figure 3c.1 , Figure 3c.2 , Figure 3d.1 , Figure 3d.2 , Figure 3e.1 , Figure 3e.2 The numbers in the table represent the PSNR for the corresponding region. From... Figure 3e.1 and Figure 3e.2As can be seen, the method of the present invention maintains better details in salient regions. Correspondingly, the quality in non-salient regions decreases, but the non-salient regions receive less attention and have a smaller impact on the overall perceived quality. In addition, in the method of the present invention, the central sub-aperture image is not encoded and compressed; it is synthesized at the decoding end. This indicates that the bitrate allocation strategy proposed in the method of the present invention not only affects the encoded sub-aperture images but also further influences the sub-aperture images synthesized at the decoding end, thereby improving the visual quality of their salient regions.
[0068] Figure 4a A comparison of EPI (Epipolar Plane Image) consistency in salient regions of Danger_de_Mort light field images is presented using the HEVC-AI method, the HEVC-ASR method, and the method of this invention. Figure 4b A comparison of EPI consistency in salient regions of Fountain_&_Vincent_1 light field images is presented using the HEVC-AI method, the HEVC-ASR method, and the method of this invention. Figure 4c A comparison of EPI consistency in salient regions of Stone_Pillars_Outside light field images using the HEVC-AI method, HEVC-ASR method, and the method of this invention is presented. Figure 4a , Figure 4b , Figure 4c As can be seen, the method of the present invention increases the bit rate of the salient region, improves the visual quality of the salient region, and thus enhances the structural consistency of the salient region of the decoded light field.
Claims
1. A light field image sensing coding method based on intra-frame coding tree unit-level bit rate allocation, characterized in that... Includes the following steps: Step 1: At the encoding end, select only a portion of the sub-aperture images in the sub-aperture image array used to represent the light field image. The selected portion of the sub-aperture images consists of four sub-aperture images located at the four corners of the sub-aperture image array. Then, arrange all the selected sub-aperture images to form a pseudo-video sequence. Each sub-aperture image in the pseudo-video sequence is used as the sub-aperture image to be encoded. Step 2: Select the central sub-aperture image from the sub-aperture image array used to represent the light field image; then calculate the initial bitrate allocation weight for each coding tree unit in each sub-aperture image to be encoded in the pseudo-video sequence, and assign the first bitrate weight to any sub-aperture image to be encoded. The initial code rate allocation weights for each coding tree unit are denoted as follows: , ;in, The total number of coding tree units in the sub-aperture image to be coded is the same as the total number of coding tree units in the central sub-aperture image, and both are... The lengths of the coding tree units in the sub-aperture image to be coded and the coding tree units in the central sub-aperture image are both [number missing]. The width of the coding tree unit in the sub-aperture image to be coded and the coding tree unit in the central sub-aperture image are both... , This indicates the coordinate position of a pixel within a coding tree unit in the central sub-aperture image, within that coding tree unit. The first sub-aperture image represents the central sub-aperture. The coordinate position in each coding tree unit is The gradient value of the pixel, A constant greater than 0; Step 3: Estimate the depth map of the central sub-aperture image using a depth estimation network; then binarize the depth map of the central sub-aperture image to obtain the foreground mask of the depth map of the central sub-aperture image, denoted as . ; then calculate The foreground density of each coding tree unit in the code will The first in The foreground density of each coding tree unit is denoted as... , ;in, express The first in The number of pixels with a value of 1 in each coding tree unit The length of the coding tree unit in the code is also... , The width of the coding tree unit in the code is also... , The number of pixels in any coding tree unit is ; A saliency detection network is used to detect the saliency map of the centroidal aperture image. Then, the saliency map of the centroidal aperture image is binarized to obtain a salient object mask, denoted as . ; then calculate The significant object density of each coding tree unit in the data will be... The first in The salient object density of a coding tree unit is denoted as , ;in, express The first in The number of pixels with a value of 1 in each coding tree unit The length of the coding tree unit in the code is also... , The width of the coding tree unit in the code is also... , The number of pixels in any coding tree unit is also ; Step 4: Calculate the final bitrate allocation weight for each coding tree unit in each sub-aperture image to be coded in the pseudo-video sequence, and assign the final bitrate allocation weight to any sub-aperture image to be coded. The final code rate allocation weights for each coding tree unit are denoted as follows: , ;in, , ; Step 5: Perform intra-frame coding on the pseudo-video sequence using a standard video encoder. For any sub-aperture image to be encoded in the pseudo-video sequence, treat it as the current frame; traverse each coding tree unit in the current frame, and for the first sub-aperture image in the current frame... Each coding tree unit is used to assign weights based on the final bitrate of that coding tree unit. Assign a target bitrate to this coding tree unit, and select the first bitrate in the current frame. The target code rate of each coding tree unit is denoted as . , The target bitrate allocation for each coding tree unit in all sub-aperture images to be encoded in the pseudo-video sequence is completed in the same manner as described above, and the remaining encoding steps are completed by a standard video encoder; among which, Indicates the total number of target bits in the current frame. This indicates the actual number of bits encoded in the frame header information of the current frame. This indicates the number of bits consumed by the encoded code tree units in the current frame. Indicates the first in the current frame The final code rate allocation weights for each coding tree unit; Step 6: At the decoding end, a decoding terminal aperture image array is obtained by decoding all the sub-aperture images. The decoding terminal aperture image array is obtained by combining two parts. The first part is all the sub-aperture images obtained by decoding, and the second part is all the sub-aperture images at other positions except for all the sub-aperture images obtained by decoding, which are synthesized by inputting all the sub-aperture images obtained by decoding into the optical field angle super-resolution reconstruction network.
2. The light field image sensing coding method based on intra-frame coding tree unit-level bit rate allocation according to claim 1, characterized in that... In step 1, the four selected sub-aperture images are arranged in the order of the top left sub-aperture image, the top right sub-aperture image, the bottom left sub-aperture image, and the bottom right sub-aperture image to form a pseudo-video sequence.
3. The light field image sensing coding method based on intra-frame coding tree unit-level bit rate allocation according to claim 1 or 2, characterized in that... In step 2, .
4. The light field image sensing coding method based on intra-frame coding tree unit-level bit rate allocation according to claim 3, characterized in that... In step 2, ,in, The first sub-aperture image represents the central sub-aperture. The coordinate position in each coding tree unit is The Y component value of the pixel, The first sub-aperture image represents the central sub-aperture. The coordinate position in each coding tree unit is The Y component value of the pixel, The first sub-aperture image represents the central sub-aperture. The coordinate position in each coding tree unit is The Y component value of the pixel, symbol " " is the absolute value symbol.
5. The light field image sensing coding method based on intra-frame coding tree unit-level bit rate allocation according to claim 1, characterized in that... In step 3, the binarization method used is the Otsu binarization method.
6. The light field image sensing coding method based on intra-frame coding tree unit-level bit rate allocation according to claim 1, characterized in that... In step 6, the decoded images are the sub-aperture images at the four corner positions. The decoded sub-aperture images at the four corner positions are input into the optical field angle super-resolution reconstruction network to synthesize all sub-aperture images at the positions other than the four corner positions.
7. The light field image perceptual coding method based on intra-frame coding tree unit-level bit rate allocation according to claim 1, characterized in that... The position index numbers of the coding tree units at the same position in the foreground mask of the sub-aperture image to be encoded, the central sub-aperture image, and the salient object mask of the depth map are consistent.
Citation Information
Patent Citations
CN110191359A
CN113630619A