View Packing for Image or Video Coding - Patent application

By rearranging and transforming additional views into contiguous blocks and encoding metadata, the method addresses inefficiencies in 3DoF+ video encoding, achieving substantial reductions in bit rate and pixel rate.

JP7768218B2Active Publication Date: 2025-11-12KONINKLIJKE PHILIPS NV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023504715
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-31
Filing Date
2021-07-26
Publication Date
2025-11-12
Estimated Expiration
2041-07-26

AI Technical Summary

Technical Problem

Existing encoding methods for 3DoF+ video data are inefficient in terms of computational effort, energy consumption, and data rate, leading to high bandwidth requirements and decoder complexity due to redundancy in multi-view images.

Method used

The method involves dividing and rearranging additional views into contiguous blocks, applying transformations, and encoding metadata to reduce redundancy, using lossy and lossless compression algorithms to minimize bit rate and pixel rate.

Benefits of technology

The proposed method significantly reduces bit rate and pixel rate by up to 82% and 61%, respectively, while maintaining image quality, thus optimizing storage and transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007768218000002
    Figure 0007768218000002
  • Figure 0007768218000003
    Figure 0007768218000003
  • Figure 0007768218000004
    Figure 0007768218000004
Patent Text Reader

Abstract

An encoder, decoder, encoding method, and decoding method for 3DoF+ video are disclosed. The encoding method includes receiving 110 multiview image or video data comprising a base view and at least a first additional view of a scene. The method proceeds to identifying 220 pixels in the first additional view that need to be encoded because they contain scene content not visible in the base view. The first additional view is divided into first blocks of pixels (230). First blocks that include at least one of the identified pixels are retained (240), and first blocks that do not include any of the identified pixels are discarded. The retained blocks are rearranged (250) to be contiguous in at least one dimension. Packed additional views are generated from the rearranged first retained blocks (260) and encoded (264).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the coding of multi-view image or video data, and in particular to a method and apparatus for encoding and decoding video sequences for virtual reality (VR) or immersive video applications. [Background technology]

[0002] Encoding schemes for several different types of immersive media content have been studied in the art. One type is 360° video, also known as three degrees of freedom (3DoF) video. This allows a view of a scene to be reconstructed for a viewpoint that has any orientation (selected by the content consumer) but is only at a fixed point in space. In 3DoF, the degrees of freedom are angles: pitch, roll, and yaw. 3DoF video supports head rotation; in other words, a user consuming video content can look in any direction within a scene but cannot move to a different location within the scene.

[0003] As the name suggests, "3DoF+" represents an extension of 3DoF video. The "+" reflects the additional support for limited translational changes of viewpoint within a scene. This allows the occupant to, for example, move their head slightly up and down, left and right, and forward and backward. This enhances the experience by allowing the user to experience the parallax effect and, to some extent, "look around" objects in the scene.

[0004] Unconstrained translation is the goal of six degrees of freedom (6DoF) video. This enables a fully immersive experience, whereby the observer can move freely within the virtual scene and look in any direction from any point within the scene. 3DoF+ does not support these large translations.

[0005] 3DoF+ is a key enabling technology for virtual reality (VR) applications, which are gaining increasing interest. Typically, VR 3DoF+ content is recorded by using multiple cameras to capture a scene, looking in a variety of different directions from various (slightly) different viewing positions. Each camera generates a respective "view" of the scene, which includes image data (also called "texture" data) and depth data. For each pixel, the depth data represents the depth at which the corresponding image pixel data is observed.

[0006] Because the views all depict the same scene from slightly different positions and angles, there is typically a high degree of redundancy in the content of different views. In other words, much of the visual information captured by each camera is also captured by one or more other cameras. It is desirable to reduce this redundancy in order to store and / or transmit content in a bandwidth-efficient manner and to encode and decode content in a computationally efficient manner. Because content is generated (and encoded) once but may be consumed (and therefore decoded) multiple times by multiple users, it is particularly desirable to minimize decoder complexity.

[0007] Among the views, one view can be designated as the "base" or "center" view, while others can be designated as "additional" or "side" views. Summary of the Invention

[0008] [Problem to be solved by the invention]

[0009] It is desirable to efficiently encode and decode the base and additional views in terms of computational effort, energy consumption, and data rate (bandwidth). It is desirable to increase coding efficiency in terms of both bit rate and the number of pixels that need to be processed (pixel rate). The bit rate affects the bandwidth required to store and / or transmit the encoded views as well as the complexity of the decoder. The pixel rate affects the complexity of the decoder.

[0010] [Means for solving the problem]

[0011] The invention is defined by the claims.

[0012] According to an example according to an aspect of the present invention, there is provided a method for encoding multi-view image or video data as claimed in claim 1.

[0013] Here, "contiguous in at least one dimension" means (i) scanning left-to-right or right-to-left along each row of the blocks, with no gaps between the retained first blocks; or (ii) scanning top-to-bottom or bottom-to-top along all columns of the blocks, with no gaps between the retained first blocks; or (iii) the retained first blocks are contiguous in two dimensions. Case (i) means the blocks are connected along rows, with all retained first blocks adjacent to another retained first block on the left and right, except for the leftmost and rightmost blocks in each row. However, there may be one or more rows without a retained block. Case (ii) means the blocks are connected along columns, with all retained first blocks adjacent to another retained first block on the top and bottom, except for the topmost and bottommost blocks in each column. However, there may be one or more columns without a retained block.

[0014] In case (iii), "contiguous in two dimensions" means that every first block retained is adjacent to at least one other such block (above, below, left, or right). Thus, there are no isolated blocks or groups of blocks. Preferably, as described above for the two one-dimensional cases, there are no gaps along any of the columns and no gaps along any of the rows.

[0015] Reordering the retained first blocks may include shifting each retained first block in one dimension, and in particular positioning it immediately adjacent along that dimension to its nearest neighboring retained first block.

[0016] The shifting may include shifting horizontally along rows of blocks or shifting vertically along columns of blocks. Horizontal shifting may be preferred. In some instances, blocks may be shifted both horizontally and vertically. For example, blocks may be shifted horizontally to generate consecutive rows of blocks. The consecutive rows may then be shifted vertically so that the blocks are consecutive in two dimensions.

[0017] The shifting may include shifting the retained first block in the same direction, for example, shifting the block to the left.

[0018] In a packed additional view, the first retained block may be contiguous with one edge of the view, which may be the left edge of the packed additional view.

[0019] The blocks may all have the same size.

[0020] The method may further include, before encoding the packed additional view, dividing the packed additional view into a first portion and a second portion, transforming the second portion relative to the first portion to generate a transformed packed view, and encoding the transformed packed view into a video bitstream. That is, the transformed packed view is encoded instead of the original packed additional view. The transformation may be selected such that the transformed packed view has a reduced size in at least one dimension. In particular, the transformed packed view may have a reduced horizontal size (i.e., the number of columns of pixels is reduced).

[0021] The transformation optionally includes one or more of: horizontally flipping the second portion; vertically flipping the second portion; transposing the second portion; circularly shifting the second portion along the horizontal direction; and circularly shifting the second portion along the vertical direction.

[0022] Inverting produces a mirror image (left to right) of the row. Reversing means flipping the column upside down. Transposing means swapping the columns with the rows, so that the first row replaces the first column of the original and the second row replaces the second column of the original (and vice versa).

[0023] The retained blocks in at least one of the first and second portions may be rearranged by shifting them to the left. This left shifting may be performed before and / or after transforming the second portion relative to the first portion. This approach may work well when compressing the transformed packed additional views later. As many compression standards work, this approach may help reduce the bit rate after compression.

[0024] The method may further include encoding into the metadata bitstream a description of how the second portion was transformed relative to the first portion.

[0025] The method may further include encoding into the metadata bitstream a description of the order in which the additional views are packed into the packed additional view.

[0026] The metadata bitstream may be encoded using lossless compression, optionally with error detection and / or correction codes.

[0027] The packed additional views may have the same size along at least one dimension as each additional view, in particular, they may have the same size along the vertical dimension (i.e., the same number of rows of pixels).

[0028] The method may further comprise compressing the base view and the packed additional views using a video compression algorithm, optionally using a standardized video compression algorithm that may employ lossy compression. Examples include, but are not limited to, H.265 and High Efficiency Video Coding (HEVC), also known as MPEG-H Part 2. The bitstream may have the compressed base view and the compressed packed additional views.

[0029] The compressed block size of a video compression algorithm may be larger in at least one dimension than the size of the first and second blocks in that dimension. This may allow multiple smaller blocks (or slices of blocks) to be combined into a single compressed block for video compression. This may help improve the coding efficiency of the retained blocks.

[0030] Each view may have an image (texture) value and a depth value.

[0031] A method for decoding multi-view image or video data as claimed in claim 10 is also provided.

[0032] Arranging the first blocks can include shifting them in one dimension according to the description in the first packing metadata. In particular, the first blocks can be shifted to spaced positions along the dimension. In some examples, the arranging can include shifting the first blocks in two dimensions.

[0033] The views in the video bitstream may have been compressed using a video compression algorithm, optionally using a standardized video compression algorithm, and the method may include, when decoding the views, decompressing the views according to the video compression algorithm.

[0034] The method may include inverse transforming a second portion of the packed additional view relative to the first portion. The inverse transform may be based on a description, decoded from the metadata bitstream, of how the second portion was transformed relative to the first portion during encoding.

[0035] There is also provided a computer program according to claim 12, which may be provided on a computer readable medium, preferably a non-transitory computer readable medium.

[0036] An encoder according to claim 13, a decoder according to claim 14 and a bitstream according to claim 15 are also provided.

[0037] This bitstream can be encoded and decoded using the methods summarized above. It can be embodied on a computer-readable medium or as a signal modulated onto an electromagnetic carrier wave. The computer-readable medium can be a non-transitory computer-readable medium.

[0038] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief explanation of the drawings]

[0039] For a better understanding of the present invention and to show more clearly how the same may be carried into effect, reference will now be made, by way of example only, to the accompanying drawings in which: [Figure 1] 1 illustrates a video encoding and decoding system operating in accordance with one embodiment. [Figure 2] 1 is a block diagram of an encoder according to one embodiment. [Figure 3] FIG. 3 illustrates the components of the block diagram of FIG. 2 in more detail. [Figure 4] 2 is a flowchart illustrating an encoding method performed by the encoder of FIG. 1; [Figure 5A-C] FIG. 10 illustrates rearrangement of retained blocks of pixels according to one embodiment. [Figure 6] 10 is a flow chart illustrating further steps for rearranging blocks of pixels. [Figure 7A-D] 7 illustrates the transformation of some of the packed additional views using the process shown in FIG. 6. [Figure 8] 1 is a block diagram of a decoder according to one embodiment; [Figure 9] 9 is a flowchart illustrating a decoding method performed by the decoder of FIG. 8; DETAILED DESCRIPTION OF THE INVENTION

[0040] The present invention will now be described with reference to the drawings.

[0041] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of the devices, systems, and methods, are intended for purposes of illustration only and are not intended to limit the scope of the invention. These and other features, aspects, and advantages of the devices, systems, and methods of the present invention will become better understood from the following description, appended claims, and accompanying drawings. It should be understood that the drawings are merely schematic and are not drawn to scale. It should also be understood that the same reference numerals are used throughout the drawings to indicate the same or similar parts.

[0042] As used herein, a "view" refers to an image of a scene (this image may be a still image or a frame of video). An image consists of a two-dimensional array of pixels made up of rows and columns. Within this array, rows run horizontally and columns run vertically. The directions "left" and "right" refer to the horizontal (i.e., row) dimension. The directions "top" / "above" and "bottom" / "below" refer to the vertical (i.e., column) dimension. The leftmost pixel is the first pixel in each row. The topmost pixel is the first pixel in each column. When an image is divided into blocks of pixels that all have the same height (in terms of number of pixels), this results in rows of blocks. When an image is divided into blocks of pixels that all have the same width (again measured as number of pixels), this results in columns of blocks. When an image is divided into blocks that all have the same height and width, this results in a regular array of blocks made up of rows and columns of blocks.

[0043] The base (or "center") view may be coded in its entirety, but to the extent that it contains extra visual content, i.e., visual content that is already represented accurately enough by the base view, the additional views may be "pruned." This results in pruned additional views that contain relatively less visual content. The inventors have recognized that before compressing the additional views, it may be advantageous to divide these additional views into blocks and rearrange these blocks to pack them more efficiently.

[0044] FIG. 1 illustrates an overall system according to one embodiment. FIG. 1 illustrates, in simplified form, a system for encoding and decoding 3DoF+ video. An array of cameras 10 is used to capture multiple views of a scene. Each camera captures a conventional image of its front view (referred to herein as a texture map) and a depth map. The set of views, including texture and depth data, is provided to an encoder 100. The encoder encodes both the texture and depth data into a conventional video bitstream, e.g., a High Efficiency Video Coding (HEVC) bitstream. This is accompanied by a metadata bitstream to inform a decoder 400 of the meaning of each portion of the video bitstream. For example, the metadata tells the decoder which portions of the video bitstream correspond to the texture map and which portions correspond to the depth map. Depending on the complexity and flexibility of the encoding scheme, more or less metadata may be required. For example, a very simple scheme could very strictly dictate the structure of the bitstream, requiring little or no metadata to unpack it at the decoder side. The more possible options there are in a bitstream, the more metadata is needed.

[0045] The decoder 400 decodes the coded views (texture and depth) and renders at least one view of the scene. It passes the rendered views to a display device, such as a virtual reality headset 40. The headset 40 uses the decoded views to request the decoder 400 to render a particular view of the 3D scene according to the headset 40's current position and orientation.

[0046] The advantage of the system shown in Figure 1 is that texture and depth data can be encoded and decoded using conventional 2D video codecs. However, the disadvantage is that a large amount of data is required to encode, transmit, and decode. Therefore, it is desirable to reduce the bit rate and / or pixel rate while minimizing the loss in quality of the reconstructed views.

[0047] 2 is a block diagram of an encoding device 100 according to this embodiment. The encoder 100 includes an input unit 110 configured to receive video data, a pruning unit 120, a packing unit 130, a video encoder 140, and a metadata encoder 150. An output of the pruning unit 120 is connected to an input of the packing unit 130. An output of the packing unit 130 is connected to an input of the video encoder 140 and the metadata encoder 150, respectively. The video encoder 140 outputs a video bitstream, and the metadata encoder 150 outputs a metadata bitstream.

[0048] FIG. 3 shows the pruning unit 120 and the packing unit 130 in more detail. The pruning unit 120 has a set of pixel identification units 122a, b,..., one for each side view of the scene. In the example of FIG. 1, there were eight views in total: one base view and seven side views. FIG. 3 shows only two side views for ease of explanation. It will be understood that other side views can be processed similarly. The pruning unit 120 also has a set of block alignment mutators 124a, b, again one per side view. The packing unit 130 has a corresponding set of left shift units 132a, b, etc. It further has a view combiner 134 for combining the side views into packed additional views.

[0049] The method performed by encoder 100 will now be described with reference to FIG. 4. In step 210, input unit 110 receives video data including a base view and additional (side) views. For the purposes of this description, it is assumed that the base view is coded and compressed separately, which is beyond the scope of this disclosure and will not be described further herein. The side views are passed to pruning unit 120. In particular, the first side view is passed to pixel identifier 122a and block alignment mutator 124a. The second side view is passed to pixel identifier 122b and block alignment mutator 124b.

[0050] In step 220, each pixel identifier 122 identifies pixels in the respective side view that need to be coded because they contain scene content that is not visible in the base view. This can be done in one of many different ways. In one example, each pixel identifier is configured to examine the gradient magnitude of the depth map. Pixels whose gradient exceeds a predetermined threshold are identified as needing to be coded. These identified pixels capture depth discontinuities. Visual information at depth discontinuities needs to be coded because they appear different in different views of the scene, for example, due to parallax effects. In this way, identifying pixels with large gradient magnitudes provides one way of identifying regions of the image that need to be coded because they are not visible in the base view.

[0051] In another example, the encoder may be configured to construct a test viewport based on certain pixels that have been discarded (i.e., not encoded). This may be compared to a reference viewport constructed while retaining these pixels. The pixel identifier may be configured to calculate the difference (e.g., the sum of squared differences between pixel values) between the test viewport and the reference viewport. If the absence of the selected pixels would not significantly affect the rendering of the test viewport (i.e., the difference is not greater than a predetermined threshold), the tested pixels may be discarded from the encoding process. Otherwise, if discarding them would significantly affect the rendered test viewport, the pixel identifier 122 should mark them for retention. The encoder may experiment with different sets of pixels proposed for discard and select the setting that provides the highest quality and / or lowest bitrate or pixel rate.

[0052] The output of pixel identifier 122 is a binary flag for each pixel indicating whether the pixel should be kept or discarded. This information is passed to the respective block alignment mutator 124. In step 230, block alignment mutator 124a divides the first side view into a plurality of first pixel blocks. In parallel, block alignment mutator 124b divides the second side view into a plurality of second pixel blocks. In step 240, block alignment mutator 124a retains the first blocks containing one or more pixels identified by pixel identifier 122a as needing to be coded. These blocks are passed to left shift unit 132a of packing unit 130. Blocks that do not contain any of the identified pixels are discarded (i.e., not passed to the packing unit). In this embodiment, this is achieved by replacing all of the discarded blocks in the side view with black pixels. This replacement with black pixels is referred to herein as "muting." Corresponding steps are performed by block alignment mutator 124b on the second side view. The retained second pixel block is passed to left shift unit 132b.

[0053] In step 250, the left shift unit 132a rearranges the retained first pixel blocks so that they are contiguous in at least one dimension. This is done by shifting the blocks left until they are all adjacent to each other along their respective rows, with the leftmost block in each row adjacent to the left edge of the image. This procedure is illustrated in Figures 5A-C. Figure 5A shows a side view 30 with individual retained blocks 32. Figure 5B illustrates the process of shifting blocks 32 left. Figure 5C shows the blocks after being moved to the left edge of the image. Each row of blocks is contiguous along the row dimension; that is, there are no gaps between blocks along each row. In this example, the blocks are also contiguous in the column direction, but this is not necessarily the case when shifting blocks along the rows. Some rows may not have retained blocks in them, in which case there will be gaps between some rows of blocks in the rearranged image. Blocks other than the retained blocks 32 shown in Figures 5A-C are blacked out. Note that Figures 5A-C show a small number of blocks in a small region of an exemplary side view. In practice, there are typically many more blocks. The inventors have found that good results are obtained with blocks that are rectangular rather than square, i.e., have a vertical height that is different from their horizontal width. In particular, better results can be achieved with blocks that have a horizontal width that is smaller than their vertical height. A vertical height of 32 pixels and a horizontal width of either 1 pixel or 4 pixels have been found to give good results.

[0054] In step 260, the view combiner adds the rearranged first retained block (from left shift unit 132a) to the packed additional view. After the single side view is added, the packed additional view is identical to FIG. 5C. In step 270, left shift unit 132a generates first packing metadata describing how the first retained block has been rearranged. Left shift unit 132b performs a similar rearrangement process on the second retained block of the second side view and generates second packing metadata describing how it has been rearranged. The rearranged blocks are passed to view combiner 134 to be added to the packed additional view. They can be added in various ways. In this example, each row of the retained block from the second side view is appended to the corresponding row of the retained block from the first side view. This procedure can be repeated for each of the side views until the packed additional view is complete. Note that because the side views have relatively sparsely spaced retained blocks after the muting stage, the retained blocks of all side views can be packed into an image with a smaller number of pixels and the total number of pixels of all side views. In particular, in this example, the packed additional views can have the same number of rows (i.e., the same vertical dimension) as each of the original side views, but a smaller number of columns (i.e., a smaller horizontal dimension). This facilitates a reduction in the pixel rate to be coded / transmitted.

[0055] In step 264, video encoder 140 receives the packed additional views from packing unit 130 and encodes the packed additional views and the base view into a video bitstream. The base view and the packed additional views may be encoded using a video compression algorithm, which may be a lossy video compression algorithm. In step 274, metadata encoder 150 encodes the first packing metadata and the second packing metadata into the metadata bitstream. Metadata encoder 150 may also encode into the metadata bitstream a definition of the sequence in which the additional views were added / packed into the packed additional views. This should be done, particularly if the additional views were not added / packed in a predetermined fixed order. The metadata is encoded using lossless compression and, optionally, error detection and / or correction codes. This is because errors in the metadata are likely to have a more significant impact on the decoding process if not received correctly by the decoder. Suitable error detection and / or correction codes are known in the art of communications theory.

[0056] An optional additional encoding stage will now be described with reference to Figure 6 and Figures 7A-D. Figure 6 is a flowchart illustrating the process steps, illustrated by the example graphs of Figures 7A-D. The process of Figure 6 may be performed by packing unit 130. This may be performed for each side view separately, or for a combination of side views included in the packed additional view. Figure 6 assumes the latter case.

[0057] In step 136, packing unit 130 divides the packed additional view into two parts. In the example shown in FIG. 7A, the packed additional view is divided into a left part 30a (part 1) and a right part 30b (part 2). The blocks in the right part 30b are shaded gray for clarity. Next, the right part 30b of the packed additional view is transformed to more uniformly distribute the number of muted (discarded) blocks on each row. In step 137, the right part 30b is flipped from left to right. This replaces the right part 30b with its mirror image, as shown in FIG. 7B. In step 138, packing unit 130 shifts the retained blocks in the right part 30b vertically and circularly (so that the top row moves to the bottom row when shifted vertically "up" by one row). In the example shown in FIG. 7C, the blocks are shifted up by four rows. As shown in FIG. 7C, each transformed row contains a similar number of muted (discarded) blocks. Conversely, each row can be said to contain a similar number of retained blocks. This allows the retained blocks of the transformed right portion (shown in gray) to be shifted left to move closer to the retained blocks of the left portion. In step 139, packing unit 130 recombines the transformed right portion 30b with the left portion 30a. In the recombination process, the retained blocks of the transformed right portion are shifted left to generate the transformed compressed view 30c, as shown in FIG. 7D. The left shifting can be performed in various ways. In the example shown in FIG. 7D, all retained blocks are shifted left by the same number of blocks (i.e., by the same number of columns), so that at least one retained block of the transformed right portion is adjacent to at least one block of the left portion along a given row. Alternatively, each row of the transformed right portion 30b may be shifted left by a row-specific number of blocks until each row of blocks in the transformed right portion 30b is contiguous with a respective row of blocks in the left portion 30a. The metadata encoder 150 encodes into the metadata bitstream a description of how the retained blocks of the right portion (portion 2) were manipulated when generating the transformed packed view.Note that the size of this description, and therefore the amount of metadata, depends to some extent on the complexity of the transformation. For example, if all rows in the right-hand part are shifted left by the same number of columns, then only one value needs to be encoded in the metadata to describe this part of the transformation. On the other hand, if each row is shifted by a different number of columns, then a metadata value is generated for each row.

[0058] The complexity of the transform (and the corresponding size of the metadata) will be traded off against the bit rate and / or pixel rate reduction resulting from the transform. As is clear from the above discussion, there are several variables when selecting a transform for the right-hand portion (portion 2). These can be selected in a variety of different ways. For example, the encoder can try different choices of transform and measure the bit rate and / or pixel rate reduction for each different choice. The encoder can then select the combination of transform parameters that results in the greatest reduction in bit rate and / or pixel rate.

[0059] Figure 8 shows a decoder 400 configured to decode the video and metadata bitstreams produced by the encoder of Figure 2. Figure 9 shows a corresponding method performed by decoder 400.

[0060] In step 510, a video bitstream is received at the first input 410. In step 520, a metadata bitstream is received at a second input, which may be the same as or different from the first input. In this example, the second input is the same as the first input 410. In step 530, the video decoder 420 decodes the video bitstream to obtain the base view and the packed additional views. This may include decoding according to a standard video compression codec. In step 540, the metadata decoder 430 decodes the metadata bitstream to obtain first packing metadata describing how the first additional (side) view was added to the packed additional view and second packing metadata describing how the second additional (side) view was added to the packed additional view. This includes metadata describing the block rearrangement and optional transformation of portions described above with reference to Figures 5A-C and 7A-D.

[0061] The decoded packed additional views and the decoded metadata are passed to a reconstruction unit 440. In step 550, the reconstruction unit 440 arranges blocks from the decoded packed additional views into individual side views. This is done by using the decoded metadata to reverse the operations performed in the encoder. The decoded base view and the reconstructed side views are then passed to a renderer 450, which in step 560 renders a view of the scene based on this input.

[0062] The above encoding (and decoding) method was tested against the current state-of-the-art MPEG solutions for multi-view 3DoF+ coding (see ISO / IEC JTC 1 / SC 29 / WG 11 N18464: Working Draft 1 of Metadata for Immersive Media (Video); ISO / IEC JTC 1 / SC 29 / WG 11 N18470: Test Model for Immersive Video) using MPEG test sequences. The results are shown in Table 1 below. The results show that the method of the present embodiment achieves pixel rates between 34% and 61% of the current state-of-the-art algorithms and bit rates between 27% and 82% of the state-of-the-art algorithms, depending on the test sequence and block size. In the right column, 4x32 means 4 pixels wide horizontally and 32 pixels high vertically, and 1x32 means 1 pixel wide horizontally and 32 pixels high vertically. [Table 1]

[0063] Those skilled in the art will appreciate that the above-described embodiment is merely an example within the scope of the present disclosure. Many variations are possible. For example, rearrangement of the retained blocks is not limited to left shifting. Blocks may be shifted right instead of left. They may be shifted vertically along columns instead of horizontally along rows. In some embodiments, vertical and horizontal shifting may be combined to achieve better packing of the retained blocks. Without wishing to be bound by theory, it is believed that coding efficiency may be improved (and therefore bitrate may be reduced) if blocks are rearranged such that retained blocks adjacent to each other in the packed representation contain similar visual content. This may enable standard video compression algorithms to achieve the best coding efficiency because standard video compression algorithms are typically designed to exploit spatial redundancy in such image content. As a result, different rearrangements and transformations of blocks may work better for different types of scenes. In some embodiments, the encoder may test a variety of different reorderings and transformations and choose the combination of reorderings and / or transformations that results in the greatest reduction in bitrate and / or pixel rate for that scene while maintaining the highest quality (i.e., accuracy of reproduction).

[0064] The encoding and decoding methods of Figures 4 and 9 and the encoders and decoders of Figures 2 and 8 may be implemented in hardware or software, or a mixture of both (e.g., as firmware executing on a hardware device). To the extent that an embodiment is implemented partially or entirely in software, the functional steps illustrated in the process flowcharts may be performed by appropriately programmed physical computing devices, such as one or more central processing units (CPUs) or graphics processing units (GPUs). Each process, and its individual component steps illustrated in the flowcharts, may be performed by the same or different computing devices. According to an embodiment, a computer-readable storage medium stores a computer program including computer program code configured to cause one or more physical computing devices to perform an encoding or decoding method as described above when the program is executed on the one or more physical computing devices.

[0065] The storage medium may include volatile and non-volatile computer memory such as RAM, PROM, EPROM, and EEPROM. Various storage media may be installed in a mobile computing device or may be transportable such that one or more programs stored on the storage medium may be read by a processor.

[0066] Metadata according to an embodiment may be stored on a storage medium. A bitstream according to an embodiment may be stored on the same storage medium or a different storage medium. The metadata may be embedded in the bitstream, but this is not required. Similarly, the metadata and / or the bitstream (with the metadata in the bitstream or separate from it) may be transmitted as a signal modulated onto an electromagnetic carrier wave. The signal may be defined according to a standard for digital communication. The carrier wave may be an optical carrier wave, a radio frequency wave, a millimeter wave, or a short-range communication wave. It may be wired or wireless.

[0067] To the extent that an embodiment is implemented partially or entirely in hardware, the blocks shown in the block diagrams of Figures 2 and 8 may be separate physical components, logical subdivisions of a single physical component, or all integrated into one physical component. The functionality of a block shown in the figures may be split among multiple components in implementation, or the functionality of multiple blocks shown in the figures may be combined into a single component in implementation. Hardware components suitable for use in embodiments of the present invention include, but are not limited to, conventional microprocessors, application-specific integrated circuits (ASICs), and field-programmable gate arrays (FPGAs). One or more blocks may be implemented as a combination of dedicated hardware to perform some functions and one or more programmed microprocessors and associated circuitry to perform other functions.

[0068] Variations to the disclosed embodiments can be understood and implemented by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single processor or other unit may fulfill the functions of several items recited in the claims. The mere fact that certain means are recited in mutually different dependent claims does not indicate that a combination of these means cannot be used to advantage. Where a computer program is described above, the computer program can be stored or distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but it may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems. When the term "adapted for" is used in the claims or the description, it has the same meaning as the term "configured to." Any reference signs in the claims should not be construed as limiting the scope.

Claims

1. 1. A method for encoding multi-view image data or multi-view video data having a base view and at least a first additional view of a scene, each view comprising an array of pixels, the method comprising: receiving the multi-view image data or the multi-view video data; identifying pixels in the first additional view that need to be coded because they contain scene content that is not visible in the base view; dividing the first additional view into a first plurality of blocks of pixels; retaining a first block containing at least one of the identified pixels; discarding a first block that does not contain the identified pixel; rearranging the retained first block of pixels so that they are contiguous in at least one dimension; generating packed additional views from the rearranged retained first blocks; generating first packing metadata describing how the retained first blocks have been rearranged; encoding the base view and the packed additional views into a video bitstream; encoding the first packing metadata into a metadata bitstream.

2. 2. The method of claim 1, wherein the step of rearranging the retained first blocks shifts each retained first block in one dimension to position it immediately adjacent to a nearest retained first block along that dimension.

3. The method of claim 1 or 2, wherein the first block is a rectangular block having a width in pixels and a height in pixels, the width being different from the height.

4. The multi-view image data or the multi-view video data further comprises a second additional view, the method further comprising: identifying pixels in the second additional view that need to be coded because they contain scene content that is not visible in the base view; dividing the second additional view into a second plurality of blocks of pixels; retaining a second block containing at least one of the identified pixels; discarding second blocks that do not contain the identified pixel; rearranging the retained second block of pixels so that they are contiguous in the at least one dimension; generating second packing metadata that describes how the retained second blocks have been rearranged; adding the relocated second block to the packed additional view; and encoding the second packing metadata into the metadata bitstream.

5. The method of claim 4 , further comprising encoding into the metadata bitstream a description of the order in which the additional views were added to the packed additional views.

6. Before encoding the packed additional views, dividing the packed additional view into a first portion and a second portion; transforming the second portion relative to the first portion to generate a transformed packed view; and encoding the transformed packed views into the video bitstream.

7. The conversion is horizontally flipping the second portion; vertically inverting said second portion; Transposing, circularly shifting the second portion along a horizontal direction; and Circulating and shifting the second portion along a vertical direction; 7. The method of claim 6, comprising one or more of:

8. 8. The method of claim 6 or 7, wherein the retained blocks in at least one of the first and second parts are rearranged by shifting them to the left.

9. 9. The method of claim 1, wherein the packed additional view has at least the same size along at least one dimension as the first additional view.

10. 1. A method for decoding multi-view image data or multi-view video data representing a scene, comprising: receiving a video bitstream in which a base view and packed additional views are encoded, each view having an array of pixels; receiving a metadata bitstream having first packing metadata including a description of how a first block of pixels of a first additional view have been rearranged into the packed additional view; decoding the video bitstream to obtain the base view and the packed additional views; decoding the first packing metadata from the metadata bitstream; reconstructing the first additional view from the packed additional views using the first packing metadata to generate a reconstructed first additional view; rendering at least one view of the scene based on the base view and the reconstructed first additional view; the reconstruction of the first additional view arranges the first block according to a description in the first packing metadata.

11. the packed additional view comprises a second block of pixels belonging to a second additional view, the metadata bitstream comprises second packing metadata comprising a description of how the second block of pixels have been rearranged into the packed additional view, the method further comprising: decoding the second packing metadata from the metadata bitstream; reconstructing the second additional view from the packed additional view using the second packing metadata to generate a reconstructed second additional view; and rendering at least one view of the scene based on the base view and the reconstructed second additional view; The method of claim 10 , wherein the reconstruction of the second additional view arranges the second blocks according to descriptions in the second packing metadata.

12. A computer program which, when executed by a computer, causes the computer to carry out the method of any one of claims 1 to 9.

13. A computer program that, when executed by a computer, causes the computer to perform the method of claim 10 or 11.

14. 1. An encoder configured to encode multi-view image data or multi-view video data having a base view and at least a first additional view of a scene, each view comprising an array of pixels, the encoder comprising: an input configured to receive the multi-view image data or the multi-view video data; identifying pixels in the first additional view that need to be coded because they contain scene content that is not visible in the base view; Dividing the first additional view into a first plurality of blocks of pixels; retaining a first block containing at least one of the identified pixels; discarding a first block that does not contain the identified pixel; rearranging the retained first block of pixels so that they are contiguous in at least one dimension; generating packed additional views from the rearranged retained first blocks; a pruning unit configured to generate first packing metadata describing how the retained first blocks have been rearranged; a video encoder configured to encode the base view and the packed additional views into a video bitstream; a metadata encoder configured to encode the first packing metadata into a metadata bitstream.

15. 1. A decoder for multi-view image data or video data, comprising: a first input configured to receive a video bitstream in which base views and packed additional views have been encoded, each view having an array of pixels; a second input configured to receive a metadata bitstream having first packing metadata including a description of how a first block of pixels of a first additional view has been rearranged into the packed additional view; and a video decoder configured to decode the video bitstream to obtain the base view and the packed additional views; a metadata decoder configured to decode the first packing metadata from the metadata bitstream; and a reconstruction unit configured to reconstruct the first additional view from the packed additional views using the first packing metadata to generate a reconstructed first additional view; a renderer configured to render at least one view of a scene based on the base view and the reconstructed first additional view; the reconstruction unit is configured to arrange the first block according to a description in the first packing metadata when reconstructing the first additional view.

Citation Information

Patent Citations

  • Processing video data for a video player apparatus

    EP3672251A1

  • Method for generating and reconstructing a three-dimensional video stream, based on the use of the occlusion map, and corresponding generating and reconstructing device

    US20150092845A1