Method for encoding and decoding multi-view image or video data, encoder and decoder using the method
By rearranging and transforming blocks of multi-view image or video data, the problems of high redundancy and low coding efficiency in 3DoF+ video are solved, achieving more efficient encoding and decoding, reducing bit rate and pixel rate, while maintaining view reconstruction quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- KONINKLIJKE PHILIPS NV
- Filing Date
- 2021-07-26
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies for encoding and decoding multi-view image or video data, especially 3DoF+ video, suffer from high redundancy, low encoding efficiency, and high decoding complexity, making it difficult to effectively reduce redundancy between views and improve encoding efficiency.
By rearranging and transforming blocks of multi-view image or video data, especially by arranging and shifting blocks continuously along one or two dimensions, and combining this with standard video compression algorithms, basic and additional views are encoded and decoded, reducing redundancy and improving encoding efficiency.
It achieves the goal of maintaining view reconstruction quality while reducing bit rate and pixel rate, improving encoding efficiency and reducing decoding complexity.
Smart Images

Figure CN116158075B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to the encoding of multi-view image or video data. The invention particularly relates to methods and apparatus for encoding and decoding video sequences for Virtual Reality (VR) or immersive video applications. BACKGROUND
[0002] Encoding schemes for several different types of immersive media content have been investigated in the art. One type is 360° video, also known as three degrees of freedom (3DoF) video. This allows the reconstruction of a view of a scene for a viewpoint with an arbitrary orientation (chosen by the consumer of the content) but at a fixed point in space only. In 3DoF, the degrees of freedom are angular - i.e. pitch, roll and yaw. 3DoF video supports head rotation - in other words, a user consuming the video content can look in any direction in the scene, but cannot move to a different location in the scene.
[0003] As the name suggests, "3DoF+" denotes an enhancement of 3DoF video. The "+" reflects the fact that it additionally supports limited translational changes of the viewpoint in the scene. For example, this can allow a seated user to move their head up, down, left and right, forward and backward by a small distance. This enhances the experience, as it allows the user to experience parallax effects, and to some extent, to look at "surrounding" objects in the scene.
[0004] Unconstrained translation is the goal of six degrees of freedom (6DoF) video. This allows a fully immersive experience whereby a viewer can move freely around a virtual scene, and can look in any direction from any point in the scene. 3DoF+ does not support these large translations.
[0005] 3DoF+ is an important enabling technology for Virtual Reality (VR) applications, where there is increasing interest. Typically, VR 3DoF+ content is recorded by capturing a scene using multiple cameras, looking in a range of different directions from a range of (slightly) different viewing positions. Each camera generates a respective "view" of the scene, comprising image data (sometimes also referred to as "texture" data) and depth data. For each pixel, the depth data represents the depth at which the corresponding image pixel data is observed.
[0006] Because the views all represent the same scene from slightly different positions and angles, there is typically a high degree of redundancy in the content of the different views. In other words, much of the visual information captured by each camera is also captured by one or more of the other cameras. To store and / or transmit the content in a bandwidth-efficient manner, and to encode and decode it in a computationally efficient manner, it is desirable to reduce this redundancy. It is particularly desirable to minimize the complexity of the decoder, because the content can be produced (and encoded) once, but can be consumed (and thus decoded) many times by multiple users.
[0007] Among the views, one view can be designated as a "base" view or a "central" view. The other views can be designated as "additional" views or "side" views. SUMMARY
[0008] It would be desirable to efficiently encode and decode the base and additional views in terms of computational effort, energy consumption, and data rate (bandwidth). It would be desirable to improve the coding efficiency in terms of bit rate and number of pixels that need to be processed (pixel rate). The bit rate affects the bandwidth required to store and / or transmit the encoded views and the complexity of the decoder. The pixel rate affects the complexity of the decoder.
[0009] The invention is defined by the claims.
[0010] According to an example of an aspect of the invention, there is provided a method of encoding multi-view image or video data.
[0011] Here, "contiguous in at least one dimension" means (i) scanning along each row of blocks from left to right or from right to left with no gaps between the retained first blocks, or (ii) scanning along all columns of blocks from top to bottom or from bottom to top with no gaps between the retained first blocks, or (iii) the retained first blocks are contiguous in two dimensions. Case (i) means that the blocks are connected along the rows: each retained first block is adjacent to another retained first block to its left and right, except for the blocks at the left and right ends of each row. However, there can be one or more rows with no retained blocks. Case (ii) means that the blocks are connected along the columns: each retained first block is adjacent to another retained first block above and below, except for the blocks at the top and bottom of each column. However, there can be one or more columns with no retained blocks.
[0012] In case (iii), "contiguous in two dimensions" means that each retained first block is adjacent (above, below, left, or right) to at least one other such block. Thus, there are no isolated blocks or groups of blocks. Preferably, there are no gaps along any column, and no gaps along any row, as described above for the two one-dimensional cases.
[0013] Re-arranging the reserved first blocks can comprise moving each reserved first block in one dimension, in particular positioning it to be directly adjacent to its nearest neighbouring reserved first block along that dimension.
[0014] The shifting can comprise horizontal shifting along rows of the blocks, or vertical shifting along columns of the blocks. Horizontal shifting can be preferred. In some examples, the blocks can be shifted horizontally as well as vertically. For example, the blocks can be shifted horizontally to produce consecutive rows of blocks. The consecutive rows can then be shifted vertically so that the blocks are consecutive in two dimensions.
[0015] The shifting can comprise shifting the reserved first blocks in the same direction. For example, shifting the blocks to the left.
[0016] In the packed additional view, the reserved first blocks can abut one edge of the view. This can be the left edge of the packed additional view.
[0017] The blocks can all have the same size.
[0018] The method can further comprise, prior to encoding the packed additional view: splitting the packed additional view into a first part and a second part; transforming the second part relative to the first part to generate a transformed packed view; and encoding the transformed packed view into the video bitstream. That is, the transformed packed view rather than the original packed additional view is encoded. The transformation can be selected so that the transformed packed view has a reduced size in at least one dimension. In particular, the transformed packed view can have a reduced horizontal size (i.e. a reduced number of columns of pixels).
[0019] The transformation can optionally comprise one or more of: reversing the second part in a horizontal direction; reversing the second part in a vertical direction; transposing the second part; cyclically shifting the second part along the horizontal direction, and cyclically shifting the second part along the vertical direction.
[0020] Reversing produces a mirror image of the rows (left-right). Reversing means inverting the columns. Transposing means swapping the rows for columns (and vice versa) so that the first row is replaced by the original first column, the second row is replaced by the original second column, and so on.
[0021] The reserved blocks in at least one of the first part and the second part can be re-arranged by shifting them to the left. This shifting to the left can be performed before and / or after the transformation of the second part relative to the first part. This approach can work well when the transformed packed additional view is subsequently compressed. This approach can help to reduce the bit rate after compression due to the way many compression standards work.
[0022] The method can further comprise encoding into the metadata bitstream a description of how the second portion is to be transformed relative to the first portion.
[0023] The method can further comprise encoding into the metadata bitstream a description of the order in which the additional views are packed into the packed additional views.
[0024] The metadata bitstream can be encoded using lossless compression, optionally with error detection and / or correction codes.
[0025] The packed additional views can have the same size as each additional view along at least one dimension. In particular, they can have the same size (i.e. the same number of rows of pixels) along a vertical dimension.
[0026] The method can further comprise compressing the base view and the packed additional views using a video compression algorithm (optionally a standardised video compression algorithm, which can employ lossy compression). Examples include but are not limited to High Efficiency Video Coding (HEVC), also known as H.265 and MPEG-H Part 2. The bitstream can comprise the compressed base view and the compressed packed additional views.
[0027] The compression block size of the video compression algorithm can be larger than the size of the first and second blocks in at least one dimension. This can allow multiple smaller blocks (or slices of blocks) to be gathered together into a single compression block for video compression. This can help improve the coding efficiency of the preserved blocks.
[0028] Each view can comprise image (texture) values and depth values.
[0029] A method of decoding multi-view image or video data is also provided.
[0030] Arranging the first blocks can comprise shifting them in one dimension according to the description in the first packing metadata. In particular, the first blocks can be shifted along the dimension to spaced apart positions. In some examples, the arranging can comprise moving the first blocks in two dimensions.
[0031] The views in the video bitstream can have been compressed using a video compression algorithm (optionally a standardised video compression algorithm). The method can comprise, when decoding the views, decompressing the views according to the video compression algorithm.
[0032] The method can comprise inversely transforming the second portion of the packed additional views relative to the first portion. The inverse transformation can be based on the description of how the second portion was transformed relative to the first portion during encoding, decoded from the metadata bitstream.
[0033] A computer program is also provided, which can be provided on a computer readable medium, preferably a non-transitory computer readable medium.
[0034] An encoder; a decoder; and a bitstream are also provided.
[0035] The bitstream can be encoded and decoded using the method as outlined above. It can be embodied on a computer readable medium or as a signal modulated onto an electromagnetic carrier. The computer readable medium can be a non-transitory computer readable medium.
[0036] These and other aspects of the application will be apparent from and elucidated with reference to the embodiments described hereinafter. BRIEF DESCRIPTION OF DRAWINGS
[0037] For a better understanding of the present application, and to show how it can be implemented in practice, reference will now be made, purely by way of example, to the accompanying drawings, in which:
[0038] Figure 1 a video encoding and decoding system operating in accordance with an embodiment is illustrated;
[0039] Figure 2 is a block diagram of an encoder according to an embodiment;
[0040] Figure 3 components of the block diagram of Figure 2 are shown in more detail;
[0041] Figure 4 is a flowchart illustrating an encoding method performed by the encoder of Figure 1
[0042] Figure 5A -C illustrates a rearrangement of reserved pixel blocks according to an embodiment;
[0043] Figure 6 is a flowchart illustrating further steps for rearranging pixel blocks;
[0044] Figure 7A -D illustrates a transformation of a part of the additional view using the process illustrated in Figure 6
[0045] Figure 8 is a block diagram of a decoder according to an embodiment;
[0046] Figure 9 is a flowchart illustrating a decoding method performed by the decoder of Figure 8 DETAILED DESCRIPTION
[0047] The application will be described with reference to the Figures.
[0048] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of apparatuses, systems and methods, are intended for purposes of illustration only and are not intended to limit the scope of the present application. These and other features, aspects, and advantages of the apparatuses, systems and methods of the present application will become better understood from the following description, appended claims, and accompanying drawings. It should be understood that the Figures are merely schematic and are not drawn to scale. It should also be understood that the same reference numerals are used throughout the Figures to indicate the same or similar parts.
[0049] As used herein, a "view" refers to an image of a scene. (The image can be a still image or a video frame) The image comprises an array of two-dimensional pixels consisting of rows and columns. In the array, the rows extend horizontally, and the columns extend vertically. The directions "left" and "right" refer to the horizontal (i.e., row) dimension. The directions "up" / "upward" and "down" / "downward" refer to the vertical (i.e., column) dimension. The leftmost pixel is the first pixel on each row. The topmost pixel is the first pixel in each column. When the image is divided into blocks of pixels that all have the same height (in terms of number of pixels), this results in rows of blocks. When the image is divided into blocks of pixels that all have the same width (again, measured in number of pixels), this results in columns of blocks. When the image is divided into blocks that have the same height and width, this results in a regular array of blocks consisting of rows and columns of blocks.
[0050] While the base (or "central") view can be encoded in its entirety, the additional views can be "pruned" to the extent that they contain redundant visual content (i.e., visual content that has already been represented accurately enough by the base view). This results in pruned additional views that are relatively sparse in visual content. The inventors have recognized that it can be advantageous to divide these additional views into blocks and to rearrange these blocks to more efficiently pack them together before compressing the additional views.
[0051] Figure 1 The overall system is illustrated in accordance with an embodiment. Figure 1A system for encoding and decoding 3DoF+ video is illustrated in simplified form. A camera array 10 is used to capture multiple views of a scene. Each camera captures a regular image (referred to herein as a texture map) and a depth map of the view in front of it. A set of views comprising texture and depth data is provided to an encoder 100. The encoder encodes both the texture data and the depth data into a regular video bitstream - for example, a High Efficiency Video Coding (HEVC) bitstream. This is accompanied by a metadata bitstream to inform a decoder 400 of the meaning of different parts of the video bitstream. For example, the metadata tells the decoder which parts of the video bitstream correspond to texture maps and which to depth maps. Depending on the complexity and flexibility of the encoding scheme, more or less metadata can be required. For example, a very simple scheme can very rigidly dictate the structure of the bitstream such that little or no metadata is required at the decoder end to unpack it. With more optional possibilities for the bitstream, a greater amount of metadata will be required.
[0052] The decoder 400 decodes the encoded views (texture and depth) and renders at least one view of the scene. It passes the rendered view to a display device, such as a virtual reality headset 40. The headset 40 requests the decoder 400 to render a particular view of the 3D scene using the decoded views, depending on the current position and orientation of the headset 40.
[0053] Figure 1 An advantage of the illustrated system is that it is able to use regular 2-D video codecs to encode and decode the texture and depth data. However, a disadvantage is that there is a large amount of data to encode, transmit and decode. It is therefore desirable to reduce the bitrate and / or the pixel rate, while impairing the quality of the reconstructed view as little as possible.
[0054] Figure 2 is a block diagram of an encoder 100 according to the present embodiment. The encoder 100 comprises an input 110 configured to receive video data, a pruning unit 120, a packing unit 130, a video encoder 140 and a metadata encoder 150. An output of the pruning unit 120 is connected to an input of the packing unit 130. Outputs of the packing unit 130 are connected to inputs of the video encoder 140 and the metadata encoder 150, respectively. The video encoder 140 outputs a video bitstream; the metadata encoder 150 outputs a metadata bitstream.
[0055] Figure 3 The pruning unit 120 and the packing unit 130 are shown in more detail. The pruning unit 120 comprises a set of pixel identifier units 122a, b,... - one for each side view of the scene. In the example shown, there are eight views in total, i.e. one base view and seven side views. For ease of explanation, the pixel identifier units 122a, b,... are shown as separate units, but they can be implemented as a single unit or as a set of units. Figure 1 Figure 1 The packing unit 130 comprises a set of packing units 132a, b,... - one for each side view of the scene. In the example shown, there are eight views in total, i.e. one base view and seven side views. For ease of explanation, the packing units 132a, b,... are shown as separate units, but they can be implemented as a single unit or as a set of units.Figure 3 Only two side views are shown. It will be appreciated that the other side views can be processed similarly. The pruning unit 120 also comprises a set of block-aligned muter units 124a, 124b,... - again, one for each side view Figure 1 The packing unit 130 comprises a respective set of left-shift units 132a, b, etc. It also comprises a view combiner 134 for combining the side views into a packed additional view.
[0056] The method performed by the encoder 100 will now be described with reference to Figure 4 In step 210, the input 110 receives video data comprising a base view and additional (side) views. For the purposes of this description, it is assumed that the base view is encoded and compressed separately - this is outside the scope of the present disclosure and will not be discussed further herein. The side views are passed to the pruning unit 120. In particular, the first side view is passed to the pixel identifier 122a and the block-aligned muter 124a. The second side view is passed to the pixel identifier 122b and the block-aligned muter 124b.
[0057] In step 220, each pixel identifier 122 identifies the pixels in the respective side view that need to be encoded because they contain scene content that is not visible in the base view. This can be done in one of a number of different ways. In one example, each pixel identifier is configured to examine the magnitude of the gradient of the depth map. Pixels for which this gradient is above a predetermined threshold are identified as needing to be encoded. These identified pixels will capture depth discontinuities. The visual information at depth discontinuities needs to be encoded because it will appear differently in different views of the scene - for example, due to the parallax effect. In this way, identifying pixels for which the gradient is large provides a way of identifying image regions that need to be encoded because they are not visible in the base view.
[0058] In another example, the encoder can be configured to construct a test viewport based on certain pixels being discarded (i.e. not encoded). This can be compared to a reference viewport constructed while those pixels are retained. The pixel identifier can be configured to compute the difference (e.g. sum of squared differences between pixel values) between the test viewport and the reference viewport. If the absence of the selected pixels does not affect the appearance of the test viewport too much (i.e. if the difference is not greater than a predetermined threshold), then the test pixels can be discarded from the encoding process. Otherwise, if discarding them has a significant impact on the rendered test viewport, the pixel identifier 122 should flag them for retention. The encoder can experiment with different sets of pixels to discard and select the configuration that provides the highest quality and / or lowest bitrate or pixel rate.
[0059] The output of the pixel identifier 122 is a binary flag for each pixel, indicating whether the pixel is to be retained or discarded. This information is passed to the respective block-aligned silencer 124. In step 230, the block-aligned silencer 124a divides the first side view into a plurality of first blocks of pixels. In parallel, the block-aligned silencer 124b divides the second side view into a plurality of second blocks of pixels. In step 240, the block-aligned silencer 124a retains those first blocks that contain at least one of the pixels identified by the pixel identifier 122a as requiring encoding. These blocks are passed to the left-shifting unit 132a of the packing unit 130. Blocks that do not contain any identified pixels are discarded (i.e. they are not passed to the packing unit). In the present embodiment, this is achieved by replacing all discarded blocks in the side view with black pixels. This replacement with black pixels is referred to herein as "silencing". Corresponding steps are performed by the block-aligned silencer 124b on the second side view. The retained second blocks of pixels are passed to the left-shifting unit 132b.
[0060] In step 250, the left-shifting unit 132a re-arranges the retained first blocks of pixels so that they are contiguous in at least one dimension. This is achieved by shifting the blocks to the left until they are adjacent to each other along the respective rows of blocks, where the leftmost block in each row is adjacent to the left edge of the image. This process is illustrated in Figure 5A -C. Figure 5A A side view 30 is shown, in which the individual blocks 32 are to be retained. Figure 5B The process of shifting the blocks 32 to the left is illustrated. Figure 5C The blocks are shown after they have all been shifted to the left-hand edge of the image. Each row of blocks is contiguous along the row dimension, i.e. there are no gaps between the blocks along each row. In this example, the blocks are also contiguous in the column direction; however, this is not necessarily always the case when shifting the blocks along the rows. Some rows can have no retained blocks, in which case there will be gaps between some rows of blocks in the re-arranged image. In addition to the retained blocks 32 shown in Figure 5A The blocks other than the retained blocks 32 shown in -C are coloured black. Note that Figure 5A -C shows a small number of blocks in a small region of an exemplary side view. In practice, there will typically be many more blocks. The inventors have found that good results can be obtained with rectangular rather than square blocks, i.e. blocks having a vertical height that is different from its horizontal width. In particular, better results can be achieved using blocks having one or more blocks. The horizontal width is less than the vertical height. A vertical height of 32 pixels has been found to give good results, with a horizontal width of 1 pixel or 4 pixels.
[0061] In step 260, the view combiner adds the re-arranged first reserved blocks (from left-shift unit 132a) to the packed additional view. After adding a single side view, the packed additional view is identical to Figure 5C In step 270, left-shift unit 132a generates first packing metadata describing how the first reserved blocks are re-arranged. Left-shift unit 132b performs a similar re-arrangement operation on the second reserved blocks of the second side view and generates second packing metadata describing how these blocks are re-arranged. The re-arranged blocks are passed to view combiner 134 to be added to the packed additional view. They can be added in various ways. In the present example, each row of reserved blocks from the second side view is appended to the corresponding row of reserved blocks from the first side view. This flow can be repeated for each side view until the packed additional view is complete. Note that because the side views are relatively sparsely populated with reserved blocks, after the silent phase, the reserved blocks of all side views can be packed into an image having a smaller number of pixels and a total number of pixels of all side views. In particular, in the present example, although the packed additional view has the same number of rows as each of the original side views (i.e., the same vertical dimension), it can have a smaller number of columns (i.e., a smaller horizontal dimension). This helps to reduce the pixel rate to be encoded / transmitted.
[0062] In step 264, video encoder 140 receives the packed additional view from packing unit 130 and encodes the packed additional view and the base view into a video bitstream. The base view and the packed additional view can be encoded using a video compression algorithm, which can be a lossy video compression algorithm. In step 274, metadata encoder 150 encodes the first packing metadata and the second packing metadata into a metadata bitstream. Metadata encoder 150 can also encode the definition of the sequence in which the additional views are added / packed into the packed additional view into the metadata bitstream. In particular, if the additional views are not added / packed in a predetermined fixed order, this should be done. The metadata is encoded using lossless compression, optionally using error detection and / or correction codes. This is because an error in the metadata can have a more significant impact on the decoding process if it is not received correctly at the decoder. Suitable error detection and / or correction codes are known in the field of communication theory.
[0063] Reference will now be made to Figure 6 and 7A -D describes optional additional encoding stages. Figure 6 is a flowchart showing the processing steps, which are illustrated in the graphical examples in Figure 7A -D. Figure 6The process can be performed by the packing unit 130. It can be performed separately for each side view, or it can be performed on a combination of side views included in the packed additional view. In the present case, the latter case is assumed. Figure 6
[0064] In step 136, the packing unit 130 splits the packed additional view into two parts. In the example illustrated in Figure 7A , the packed additional view is split into a left part 30a (part 1) and a right part 30b (part 2). For clarity of illustration, the blocks of the right part 30b are represented with a gray shading. Next, the right part 30b of the packed additional view is transformed to make the number of silent (discarded) blocks on each row more uniform. In step 137, the right part 30b is flipped from left to right. This replaces the right part 30b with its mirror image, as shown in Figure 7B . In step 138, the packing unit 130 moves the kept blocks of the right part 30b in a circular vertical movement (whereby when moving vertically "up" by one row, the top row moves to the bottom row). In the example illustrated in Figure 7C , the blocks are shifted up by 4 rows. As shown in Figure 7C , each row of the transformed now includes a similar number of silent (discarded) blocks. Conversely, it can be said that each row contains a similar number of kept blocks. This allows the kept blocks of the transformed right part (shown in gray) to be shifted to the left, to be closer to the kept blocks of the left part. In step 139, the packing unit 130 recombines the transformed right part 30b with the left part 30a. During the recombination, the kept blocks of the transformed right part are shifted to the left to produce a transformed packed view 30c, as shown in Figure 7D . The shifting to the left can be performed in various ways. In the example illustrated in Figure 7D , each kept block is shifted to the left by the same number of blocks (i.e. the same number of columns), such that at least one kept block of the transformed right part is adjacent to at least one block of the left part along a given row. Alternatively, each row of the transformed right part 30b can be shifted to the left by a row-specific number of blocks, until each row of blocks of the transformed right part 30b is contiguous with the corresponding row of blocks of the left part 30a. The metadata encoder 150 will encode a description of how the kept blocks of the right part (part 2) were manipulated when generating the transformed packed view into the metadata bitstream. It should be noted that the size of this description, and thus the amount of metadata, will depend to some extent on the complexity of the transformation. For example, if all rows of the right part are shifted to the left by the same number of columns, only one value needs to be encoded into the metadata to describe this part of the transformation. On the other hand, if each row is shifted by a different number of columns, each row will generate a metadata value.
[0065] The complexity of the transform (and the corresponding size of the metadata) can be traded off against the reduction in bitrate and / or pixel rate caused by the transform. As will be apparent from the foregoing description, when selecting a transform for the correct part (part 2), there are several variables. These can be chosen in various different ways. For example, the encoder can experiment with different transform choices, and can measure the reduction in bitrate and / or pixel rate for each different choice. The encoder can then select the combination of transform parameters that results in the largest reduction in bitrate and / or pixel rate.
[0066] Figure 8 A decoder 400 is shown that is configured to decode a video and metadata bitstream produced by an encoder. Figure 2 A corresponding method performed by the decoder 400 is shown. Figure 9 A corresponding method performed by the decoder 400 is shown.
[0067] In step 510, a video bitstream is received at a first input 410. In step 520, a metadata bitstream is received at a second input, which can be the same or different from the first input. In the present example, the second input is the same as the first input 410. In step 530, the video bitstream is decoded by a video decoder 420 to obtain a base view and packed additional views. This can include decoding according to a standard video compression codec. In step 540, the metadata bitstream is decoded by a metadata decoder 430 to obtain first packing metadata describing how to add a first additional (side) view to the packed additional views, and second packing metadata describing how to add a second additional (side) view to the packed additional views. This includes metadata describing the rearrangement of blocks and optional transforms of parts described above with reference to Figure 5A - C and 7A-D describe the rearrangement of blocks and optional transforms of parts.
[0068] The decoded packed additional views and the decoded metadata are passed to a reconstruction unit 440. In step 550, the reconstruction unit 440 arranges the blocks from the decoded packed additional views into separate side views. It does this by reversing the manipulations performed at the encoder using the decoded metadata. Then, in step 560, the decoded base view and the reconstructed side views are passed to a Tenderer 450, which renders views of the scene based on the input.
[0069] The MPEG test sequence has been used to evaluate the multi-view encoding and decoding described above. Figure 3The above-mentioned encoding (and decoding) method was tested by the prior art MPEG solution for DoF+ encoding (see ISO / IEC JTC 1 / SC 29 / WG 11 N18464: Working Draft 1 of Metadata for Immersive Media (Video); ISO / IEC JTC 1 / SC 29 / WG 11 N18470: Test Model for Immersive Video). The results are in Table 1 below. The results show that the method of the present embodiments achieves between 34% and 61% of the pixel rate and between 27% and 82% of the bitrate of the prior art algorithm, depending on the test sequence and block size. In the right-hand column, 4x32 means that the block is 4 pixels wide horizontally and 32 pixels high vertically; 1x32 means that the block is 1 pixel wide horizontally and 32 pixels high vertically.
[0070] Table 1: Experimental results on MPEG test sequences relative to MPEG Working Draft for Immersive Video
[0071] Bit rate Pixel rate blkh x blkv sa 82% 61% 4x32 sb 62% 41% 4x32 sc 40% 34% 4x32 sd 80% 52% 4x32 Bit rate Pixel rate blkh x blkv sa 69% 43% 1x32 sb 41% 37% 1x32 sc 27% 34% 1x32 sd 64% 52% 1x32
[0072] Those skilled in the art will appreciate that the above embodiments are merely one example within the scope of the present disclosure. Many variations are possible. For example, the rearrangement of the preservation blocks is not limited to shifting to the left. The blocks can be shifted to the right instead of to the left. They can be shifted vertically along the columns instead of horizontally along the rows. In some embodiments, vertical shifting and horizontal shifting can be combined to achieve better packing of the preservation blocks. Without wishing to be bound by theory, it is believed that encoding efficiency (and thus bitrate reduction) can be improved if the rearranged blocks are such that similar visual content is contained in preservation blocks that are adjacent to each other in the packed representation. This can allow standard video compression algorithms to achieve optimal encoding efficiency, as they are typically designed to exploit spatial redundancy in image content like this. Thus, different rearrangements and transformations of the blocks can work better for different types of scenes. In some embodiments, the encoder can test various different rearrangements and transformations, and can select the combination of rearrangement and / or transformation that results in the greatest reduction in bitrate and / or pixel rate for that scene, while maintaining the highest quality (i.e. accuracy of the reproduction).
[0073] Figure 4 and 9 encoding and decoding method and Figure 2 and 8The encoder and decoder of the present application can be implemented in hardware or software, or a hybrid thereof (e.g., as firmware running on a hardware device). To the extent the embodiments are implemented partly or entirely in software, the functional steps illustrated in the process flow diagrams can be performed by a physically- implemented computer device such as one or more central processing units (CPUs) or graphics processing units (GPUs) suitably programmed. Each process - and the individual component steps thereof as illustrated in the flow diagrams - can be performed by the same or different computing devices. According to an embodiment, a computer-readable storage medium stores a computer program comprising computer program code configured to cause one or more physical computing devices to perform the encoding or decoding method as described above when the program is run on the one or more physical computing devices.
[0074] The storage media can include volatile and nonvolatile computer memory such as RAM, PROM, EPROM, and EEPROM. The various storage media can be fixed within a computing device or can be transportable, such that the one or more programs stored thereon can be loaded into a processor.
[0075] The metadata according to an embodiment can be stored on the storage medium. The bitstream according to an embodiment can be stored on the same storage medium or on a different storage medium. The metadata can be embedded in the bitstream, but this is not required. Likewise, the metadata and / or the bitstream (with the metadata in the bitstream or separate from it) can be transmitted as a signal modulated on an electromagnetic carrier. The signal can be defined according to a standard for digital communication. The carrier can be an optical carrier, a radio frequency wave, a millimeter wave, or a near-field communication wave. It can be wired or wireless.
[0076] To the extent the embodiments are implemented partly or entirely in hardware, Figure 2 and 8 The blocks shown in the block diagrams of the present application can be individual physical components or logical subdivisions of a single physical component, or can all be implemented in an integrated manner in one physical component. In implementations, the functionality of one block shown in the drawings can be divided among multiple components, or in implementations, the functionality of multiple blocks shown in the drawings can be combined into a single component. Hardware components suitable for use in the embodiments of the present application include, but are not limited to, conventional microprocessors, application specific integrated circuits (ASICs), and field programmable gate arrays (FPGAs). One or more blocks can be implemented as a combination of dedicated hardware to perform some functions and one or more programmed microprocessors and associated circuitry to perform other functions.
[0077] Variations to the disclosed embodiments can become apparent to those of ordinary skill in the art from a reading of the drawings, the disclosure, and the claims. In the claims, the term comprising does not exclude other elements or steps, and the terms a or an do not exclude a plurality. A single processor or other unit can fulfil the functions of several items recited in the claims. Although certain measures are recited in dependent claims, dependency does not occur only if other independent claims are present. If a computer program is discussed, it can be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state storage medium supplied together with or as part of other hardware, but can also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems. If the term "comprises" is used in the claims or the specification, it is noted that the term "comprises" is intended to be equivalent to the term "consists of" in the context of the claims. Any reference signs in the claims should not be construed as limiting the scope.
Claims
1. A method of encoding multi-view image or video data, the multi-view image or video data comprising a base view and at least a first additional view of a scene, wherein, Each view is captured by a respective different camera, each view comprising an array of pixels, the method comprising: receiving (110) the multi-view image or video data; identifying (220) pixels in the first additional view that need to be encoded, the pixels needing to be encoded because they contain scene content that is not visible in the base view; dividing (230) the first additional view into a plurality of first blocks of pixels; retaining (240) first blocks that contain at least one of the identified pixels; discarding first blocks that do not contain any of the identified pixels; rearranging (250) the retained first blocks of pixels so that the retained first blocks are contiguous in at least one dimension; generating (260) a packed additional view from the rearranged first retained blocks; generating (270) first packing metadata describing how the retained first blocks were rearranged; encoding (264) the base view and the packed additional view into a video bitstream; and encoding (274) the first packing metadata into a metadata bitstream.
2. The method of claim 1, wherein, Rearranging (250) the retained first blocks comprises shifting each retained first block in one dimension to position each retained first block directly adjacent to its nearest neighbouring retained first block along that dimension.
3. The method of claim 1 or 2, wherein, The blocks are rectangular blocks having a width in pixels and a height in pixels, wherein the width is different from the height.
4. The method of claim 1 or 2, wherein, The multi-view image or video data further comprises a second additional view, the method further comprising: identifying (220) pixels in the second additional view that need to be encoded, the pixels needing to be encoded because they contain scene content that is not visible in the base view; dividing (230) the second additional view into a plurality of second blocks of pixels; retaining (240) second blocks that contain at least one of the identified pixels; discarding second blocks that do not contain any of the identified pixels; rearranging (250) the retained second blocks of pixels so that the retained first blocks are contiguous in the at least one dimension; generating (270) second packing metadata describing how the retained second blocks were rearranged; adding the rearranged second blocks to the packed additional view; and encoding (274) the second packing metadata into the metadata bitstream.
5. The method of claim 4, further comprising encoding into the metadata bitstream a description of an order in which the additional views were added to the packed additional view.
6. The method of claim 1 or 2, further comprising, prior to encoding the packed additional view: splitting (136) the packed additional view into a first portion and a second portion; transforming (137, 138) the second portion relative to the first portion to generate a transformed packed view; and encoding the transformed packed view into the video bitstream. The transforming comprises one or more of:
7. The method of claim 6, wherein, inverting (137) the second portion in a horizontal direction; inverting (138) the second portion in a vertical direction; and inverting (138) the second portion in a diagonal direction. inverting the second portion in a vertical direction; transposing; circularly shifting the second portion along the horizontal direction; and circularly shifting the second portion along the vertical direction (138).
8. The method of claim 6, wherein, rearranging the reserved blocks in at least one of the first portion and the second portion by shifting the reserved blocks to the left.
9. The method of claim 1 or 2, wherein, the packed additional view has the same size along at least one dimension as at least the first additional view.
10. A method of decoding multi-view image or video data depicting a scene, the method comprising: receiving (510) a video bitstream in which a base view and a packed additional view are encoded, wherein each view is captured by a respective different camera, each view comprising an array of pixels; receiving (520) a metadata bitstream comprising first packing metadata containing a description of how a first block of pixels of a first additional view is rearranged into the packed additional view; decoding (530) the video bitstream to obtain the base view and the packed additional view; decoding (540) the first packing metadata from the metadata bitstream; reconstructing (550) the first additional view from the packed additional view using the first packing metadata to generate a reconstructed first additional view; and rendering (560) at least one view of the scene based on the base view and the reconstructed first additional view, wherein reconstructing the first additional view comprises arranging (550) the first block according to the description in the first packing metadata.
11. The method of claim 10, wherein, the packed additional view comprises a second block of pixels belonging to a second additional view, and the metadata bitstream comprises second packing metadata containing a description of how the second block of pixels is rearranged into the packed additional view, the method further comprising: decoding (540) the second packing metadata from the metadata bitstream; reconstructing (550) the second additional view from the packed additional view using the second packing metadata to generate a reconstructed second additional view; and rendering (560) at least one view of the scene based on the base view and the reconstructed second additional view, wherein reconstructing the second additional view comprises arranging (550) the second block according to the description in the second packing metadata.
12. A computer program product comprising a computer program comprising computer code for causing a processing system to implement a method according to any one of claims 1 to 11 when said computer program is run on the processing system.
13. An encoder (100) configured to encode multi-view image or video data comprising a base view and at least a first additional view of a scene, wherein, each view is captured by a respective different camera, each view comprising an array of pixels, the encoder comprising: an input (110) configured to receive (210) the multi-view image or video data; a pruning unit (120) configured to: identifying (220) pixels in the first additional view that need to be encoded because they contain scene content that is not visible in the base view; dividing (230) the first additional view into a plurality of first blocks of pixels; keeping (240) first blocks that contain at least one of the identified pixels; and discarding first blocks that do not contain any of the identified pixels; and a packing unit (130) configured to: re-arrange (250) the kept first blocks of pixels so that the kept first blocks are contiguous in at least one dimension; generate (260) a packed additional view from the re-arranged first kept blocks; and generate (270) first packing metadata describing how the kept first blocks are re-arranged; a video encoder (140) configured to encode (264) the base view and the packed additional view into a video bitstream; and a metadata encoder (150) configured to encode (274) the first packing metadata into a metadata bitstream.
14. A decoder (400) for depicting a multi-view image or video data of a scene, the decoder comprising: a first input (410) configured to receive (510) a video bitstream in which a base view and a packed additional view are encoded, wherein each view is captured by a respective different camera, each view comprising an array of pixels; a second input (410) configured to receive (520) a metadata bitstream comprising first packing metadata containing a description of how first blocks of pixels of a first additional view are re-arranged into the packed additional view; a video decoder (420) configured to decode (530) the video bitstream to obtain the base view and the packed additional view; a metadata decoder (430) configured to decode (540) the first packing metadata from the metadata bitstream; a reconstruction unit (440) configured to reconstruct (550) the first additional view from the packed additional view using the first packing metadata to generate a reconstructed first additional view; and a Tenderer (450) configured to render (560) at least one view of the scene based on the base view and the reconstructed first additional view, wherein the reconstruction unit is configured to arrange (550) the first blocks according to the description in the first packing metadata when reconstructing the first additional view.
15. A computer program product comprising computer-executable instructions for causing a processor to perform the method of any of claims 1 to 13.
Citation Information
Patent Citations
Processing video data for a video player apparatus
CN111355954A
An apparatus, a method and a computer program for omnidirectional video
EP3422724A1