Symbolization and decoding method and apparatus

The new VSP mode in video encoding technologies addresses the inefficiencies of existing methods by using orthographic projection and depth information to enhance inter-view prediction, improving the encoding efficiency and quality of MVD content captured by camera arrays.

JP7717902B2Active Publication Date: 2025-08-04INTERDIGITAL VC HOLDINGS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024083084
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-19
Filing Date
2024-05-22
Publication Date
2025-08-04
Estimated Expiration
2040-11-30

AI Technical Summary

Technical Problem

Existing video encoding technologies, such as MV-HEVC and 3D-HEVC, are inadequate for efficiently encoding multi-view plus depth (MVD) content captured by camera arrays, as they primarily utilize inter-view prediction based on disparity estimation between adjacent views, which is not suitable for complex camera configurations.

Method used

A new view synthesis prediction (VSP) mode is introduced, utilizing orthographic projection methods to generate intermediate prediction images by projecting pixels from a reference view to a current view, incorporating depth information for improved prediction.

Benefits of technology

Enhances the encoding efficiency of MVD content by leveraging depth information for more accurate inter-view prediction, improving the overall compression ratio and quality of reconstructed images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007717902000036
    Figure 0007717902000036
  • Figure 0007717902000037
    Figure 0007717902000037
  • Figure 0007717902000038
    Figure 0007717902000038
Patent Text Reader

Abstract

To provide a method and a device for video encoding or decoding of MVD (Multiview+Depth) data.SOLUTION: A decoding method includes obtaining view parameters for a set of views including at least one reference view and a current view of a multiview video content, applying a forward projection method to pixels of the reference view for at least one pair of reference and current views of the set of views to generate intermediate predicted images, and projecting these pixels from the camera coordinate system of the reference view to the camera coordinate system of the current view. The predicted images include information that allows image data to be reconstructed. The method also includes storing at least one final predicted image obtained from the at least one intermediate predicted image in a buffer of reconstructed images of the current view, and reconstructing a current image of the current view from the images stored in the buffer including the at least one final predicted image.SELECTED DRAWING: Figure 14A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] At least one of the embodiments generally relates to methods and apparatuses for video encoding or decoding, and more particularly to methods and apparatuses for video encoding or decoding of MVD (Multi-View + Depth) data.

Background Art

[0002] The compression of multi-view images or video content (multiple views from multiple cameras) has been investigated by image and video experts for several years. Two types of content are generally referred to as multi-view (MV) content, including synchronized images, where each image corresponds to a different viewpoint on the same scene, and content where the MV content is complemented by depth information of the scene, which is called multi-view + depth content (MVD).

[0003] In 2015, to improve the encoding efficiency of multi-view content, two extensions of HEVC (ISO / IEC 23008-2-MPEG-H Part 2, High Efficiency Video Coding / ITU-T H.265) were adopted: · MV-HEVC for MV content, · 3D-HEVC for MVD content.

[0004] In MV-HEVC, in addition to HEVC spatial intra picture prediction (i.e., intra prediction) and temporal inter picture prediction (i.e., inter prediction), an inter-view prediction mode that utilizes similarities between views is introduced. The first view is selected as a reference, and at least the second view is encoded with respect to this reference view using disparity-based motion prediction. FIG. 11 shows an example of the interdependence between images of MV content in both temporal and inter-view directions. View View0 represents a reference view that can be decoded without inter-view prediction to maintain backward compatibility with HEVC. At time T0, views View1 and View2 are not encoded as all intra pictures (I pictures), but are encoded / decoded using the reconstructed picture of view View0 at T0 as a reference picture for prediction. The picture of view View1 at time T1 is encoded / decoded using not only the picture of view View1 as a reference picture, but also the picture of view View0 at time T1.

[0005] The content on which MV-HEVC was tested is only stereo content or multi-view content, but only uses the "3" views obtained by aligned cameras. However, according to this MV-HEVC approach, inter-view prediction only utilizes the redundancy with adjacent views based on disparity estimation between adjacent views. This approach is not suitable for content captured by a camera array.

[0006] In 3D-HEVC, the same approach as MV-HEVC is adopted, but also considers the transmission of high-density depth information (i.e., depth information per pixel of each view). Since the content is the same, the same inter-view approach as with adjacent views is adopted. More complex combinations of inter-view prediction, including the use of additional depth information, are introduced.

[0007] To use the depth information for prediction mode selection, a view synthesis prediction mode (VSP) is introduced. The basic VSP mode uses, for the current block, the disparity motion vector (DMV) information corresponding to the depth information of the blocks adjacent to the current block. The depth information is used to obtain a texture block from the reference view as a predictor of the current block. Since the depth of the current block is decoded after the texture, the depth information used is one of the already reconstructed adjacent blocks. The depth values of the reconstructed adjacent blocks are generally regarded as sub-optimal depth values for inter-view prediction.

[0008] It is desirable to propose a solution that enables the provision of an improved VSP mode.

SUMMARY OF THE INVENTION

[0009] In a first aspect, one or more of the present embodiments provide a method for decoding, the method comprising: obtaining view parameters for a set of views including at least one reference view and a current view of multi-view video content, each view including a texture layer and a depth layer; for at least one pair of the reference view and the current view of the set of views, applying an orthographic projection method to the pixels of the reference view to generate an intermediate prediction image and projecting these pixels from the camera coordinate system of the reference view to the camera coordinate system of the current view, the prediction image including information that enables the reconstruction of the image data; storing at least one final prediction image obtained from at least one intermediate prediction image in a buffer of the reconstructed image of the current view; reconstructing the current image of the current view from the images stored in the buffer, the buffer including the at least one final prediction image.

[0010] In a second aspect, one or more of the present embodiments provide a method for encoding, the method comprising: Obtaining view parameters for a set of views including at least one reference view and a current view of multi-view video content, each view including a texture layer and a depth layer, For at least one pair of the reference view and the current view of the set of views, applying a forward projection method to the pixels of the reference view to generate an intermediate prediction image and projecting these pixels from the camera coordinate system of the reference view to the camera coordinate system of the current view, wherein the prediction image includes information enabling reconstruction of the image data, Storing at least one final prediction image obtained from at least one intermediate prediction image in a buffer of the reconstructed image of the current view, Reconstructing the current image of the current view from the images stored in the buffer, the buffer including the at least one final prediction image,

[0011] In a third aspect, one or more of the present embodiments provide a device for decoding, the device comprising: Means for obtaining view parameters for a set of views including at least one reference view and a current view of multi-view video content, each view including a texture layer and a depth layer, Means for, for at least one pair of the reference view and the current view of the set of views, applying a forward projection method to the pixels of the reference view to generate an intermediate prediction image and projecting these pixels from the camera coordinate system of the reference view to the camera coordinate system of the current view, wherein the prediction image includes information enabling reconstruction of the image data, Means for storing at least one final prediction image obtained from at least one intermediate prediction image in a buffer of the reconstructed image of the current view, Means for reconstructing the current image of the current view from the images stored in the buffer, the buffer including the at least one final prediction image,

[0012] In a fourth aspect, one or more of the present embodiments provide a device for encoding, the device being to obtain view parameters regarding a set of views including at least one reference view and a current view of multi-view video content, each view including a texture layer and a depth layer; for at least one pair of a reference view and a current view of the set of views, to apply an orthographic projection method to pixels of the reference view to generate an intermediate prediction image and project these pixels from the camera coordinate system of the reference view to the camera coordinate system of the current view, the prediction image including information that enables reconstruction of image data; to store at least one final prediction image obtained from at least one intermediate prediction image within a buffer of a reconstructed image of the current view; to reconstruct a current image of the current view from the images stored in the buffer, the buffer including the at least one final prediction image.

[0013] In a fifth aspect, one or more of the present embodiments provide an apparatus including the device according to the third and / or fourth aspects.

[0014] In a sixth aspect, one or more of the present embodiments provide a signal including data generated according to the encoding method according to the second aspect or by the encoding device according to the fourth aspect.

[0015] In a seventh aspect, one or more of the present embodiments provide a computer program including program code instructions for implementing the method according to the first or second aspect.

[0016] In an eighth aspect, one or more of the present embodiments provide information storage means for storing program code instructions for implementing the method according to the first or second aspect.

[0017] In a ninth aspect, one or more embodiments also provide a method and apparatus for transmitting or receiving a signal according to the sixth aspect.

[0018] In a tenth aspect, one or more embodiments also provide a computer program product including instructions for performing at least a part of any of the methods described above.

[0019] In any embodiment of any of the foregoing aspects, the information enabling the reconstruction of the image data includes texture data and depth data.

[0020] In any embodiment of any of the foregoing aspects, the forward projection method is to apply back-projection from the camera coordinate system of the reference view to the current pixel of the reference view in the world coordinate system to obtain a back-projected pixel, where the back-projection uses the pose matrix of the camera that obtains the reference view called the reference camera, the inverse intrinsic matrix of the reference camera, and the depth value associated with the current pixel, project the back-projected pixel into the coordinate system of the current view and obtain a forward-projected pixel using the intrinsic matrix and the extrinsic matrix of the camera that obtains the current view called the current camera, where each matrix is obtained from the view parameters, and if the obtained forward-projected pixel does not correspond to a pixel on the grid of the pixels of the current camera, select the pixel of the grid that is closest to the forward-projected pixel to obtain a corrected forward-projected pixel.

[0021] In any embodiment of any of the foregoing aspects, the method includes filling in missing pixels separated within each intermediate projection image or within the final projection image.

[0022] In any embodiment of any of the foregoing aspects, the information enabling the reconstruction of the image data includes motion information.

[0023] In any embodiment of any of the foregoing aspects, the forward projection method Applying back-projection from the camera coordinate system of the reference view to the current pixel in the world coordinate system to obtain a back-projected pixel, where the back-projection uses the pose matrix of the camera that obtains the reference view, called the reference camera, the inverse intrinsic matrix of the reference camera, and the depth value associated with the current pixel, Projecting the back-projected pixel into the coordinate system of the current view to obtain a forward-projected pixel using the intrinsic matrix and the extrinsic matrix of the camera that obtains the current view, called the current camera, where each matrix is obtained from the view parameters, If the obtained forward-projected pixel does not correspond to a pixel on the grid of the pixels of the current camera, selecting the pixel of the grid that is closest to the forward-projected pixel to obtain a corrected forward-projected pixel, Calculating a motion vector representing the displacement between the forward-projected pixel or the corrected forward-projected pixel and the current pixel of the reference view. The method includes the above steps.

[0024] In any of the embodiments of the foregoing aspects, the method includes filling in missing motion information separated within each intermediate projection image or within the final projection image.

[0025] In any of the embodiments of the foregoing aspects, at least one final projection image is an intermediate projection image.

[0026] In any of the embodiments of the foregoing aspects, at least one final projection image results from the aggregation of at least two intermediate prediction images.

[0027] In any of the embodiments of the foregoing aspects, the method includes subsampling the depth layer of the reference view before applying the forward-projection method to the pixels of the reference view for at least one pair of the reference view and the current view.

[0028] In any embodiment of the foregoing aspects, the method is to reconstruct the current block of the current image from a bidirectional predictor block calculated as a weighted sum of two one - direction predictor blocks, wherein each one - direction predictor block is extracted from one image stored in the buffer of the reconstructed image of the current view, and at least one of the one - direction predictor blocks is extracted from the final predicted image stored in the buffer.

[0029] In any embodiment of the foregoing aspects, at least one weight used in the weighted sum is modified as a function of the reliability of the pixels of the one - direction predictor block.

[0030] In any embodiment of the foregoing aspects, the view parameters of each view are provided by an SEI message.

[0031] In any embodiment of the foregoing aspects, a syntax element representing information that enables the reconstruction of each final predicted image of the current view is included within a slice header or a sequencing header, or an image header, or at a synchronization point or at the picture level.

[0032] In any embodiment of the foregoing aspects, multi - view video content is encoded within an encoded video stream or decoded from an encoded video stream, called the VSP mode, when the current block is encoded according to a prediction mode using the final predicted image to generate a prediction block for the current block, the encoding of the current block in the VSP mode is explicitly signaled by a flag within a part of the encoded video stream corresponding to the current block, or implicitly signaled by a syntax element representing the index of the final predicted image within a list of reconstructed images stored in the buffer of the reconstructed images of the current view.

[0033] In any embodiment of the foregoing aspects, a part of the encoded video stream corresponding to the current block includes a syntax element representing motion information.

[0034] In any of the embodiments of the foregoing aspects, the motion information represents an index of the final predicted image in the list of final predicted images stored in the motion vector refinement and / or the buffer of the reconstructed image of the current view.

[0035] In any of the embodiments of the foregoing aspects, when the current block encoded in the merge mode or the skip mode inherits the encoding parameters from the block encoded in the VSP mode, the current block also inherits the VSP parameters.

Brief Description of the Drawings

[0036]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13A

Figure 13B

Figure 14A

Figure 14B

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22A

Figure 22B

Figure 23

Figure 24

Figure 25

DETAILED DESCRIPTION OF THE INVENTION

[0037] In the following description, some embodiments use tools developed in the context of the Versatile Video Coding (VVC), an international standard under the development of a joint working team of ITU-T and ISO / IEC experts known as the Joint Video Expert Team (JVET), or in the context of HEVC, MV-HEVC, or 3D-HEVC. However, these embodiments are not limited to video encoding / decoding methods corresponding to VVC, HEVC, MV-HEVC, or 3D-HEVC, but are also applicable to other video encoding / decoding methods as well as other image encoding / decoding methods adapted to MVD content.

[0038] In the embodiments described below, a new VSP mode is proposed.

[0039] Hereinafter, FIGS. 24, 6, 7, and 8 illustrate basic embodiments that enable the introduction of some terms.

[0040] FIG. 24 schematically shows the typical encoding structure and picture dependencies of the MV-HEVC and 3D-HEVC codecs.

[0041] MV and 3D-HEVC are known to use a multi-layer approach in which layers are multiplexed into one bitstream and may be dependent on each other. In MV and 3D-HEVC, a layer can represent the texture, depth, or other auxiliary information of a scene related to a specific camera. All layers belonging to the same camera are referred to as views, and layers holding the same type of information (e.g., texture or depth) are usually called components within the scope of 3D video.

[0042] Figure 24 shows a typical coding structure that includes two views: view "0" (also called the base view) 2409 and view 1 2410. Each view includes two layers. View 0 2409 includes a first layer consisting of texture images 2401 and 2405, and a second layer including depth images 2403 and 2407. View 1 2410 includes a first layer consisting of texture images 2402 and 2406, and a second layer including depth images 2404 and 2408.

[0043] Two consecutive times are shown, and by design choice, all images associated with the same capture time or display time instance are included in one access unit (AU). Images 2401, 2402, 2403, and 2404 are in the same AU 0 2411. Images 2405, 2406, 2407, and 2408 are in the same AU 1 2412. The base layer is generally required to conform to the HEVC single-layer profile and is thus the texture component of the base view.

[0044] The layer of images following the base layer image of an AU is referred to as the enhancement layer, and views other than the base view are referred to as enhancement views. In an AU, the order of the views must be the same for all components. To facilitate coding combinations, in 3D-HEVC, it is further required that the depth component of a particular view directly follows its texture component. An overview of the dependency relationships between different layers and images within an AU is shown in Figure 24 and is discussed further below.

[0045] However, in MV-HEVC, beyond conventional temporal inter-prediction (represented by the arrows associated with the acronym TIIP in FIG. 24) that uses images of the same view and component, MV-HEVC enables prediction from images within the same AU and component but different views. Hereinafter, this prediction is referred to as inter-view prediction (represented by the arrows associated with the acronym IVP in FIG. 24). For inter-view prediction, decoded images from other views can be used as the reference images for the current image.

[0046] The motion vectors associated with the current block of the current image can be temporal (hereinafter referred to as TMV) when related to the temporal reference image of the same view, or can be disparity MVs (hereinafter referred to as DMV) when related to the inter-view reference image. Regardless of whether the MV is TMV or DMV, an existing block-level HEVC motion compensation module that operates in the same way can be used.

[0047] For improved compression performance, 3D-HEVC extends MV-HEVC by enabling a new type of inter-layer prediction. As shown in FIG. 24, the new prediction types are as follows. · A combination of temporal and inter-view prediction (represented by the arrow with the acronym TII+IVP in FIG. 24) that references images within the same component but different AUs and different views; · Intra-component prediction (represented by the arrow with the acronym ICP in FIG. 24) that references images within the same AU and view but different components; · A combination of intra-component and inter-view prediction (represented by the arrow with the acronym ICIP in FIG. 24) that references images within the same AU but different views and components.

[0048] Further design changes compared to MV-HEVC are that in addition to sample and motion information, residuals, disparities, and split information can also be predicted or inferred. A detailed overview of texture and depth coding tools is provided in the document "Overview of the Multiview and 3D Extensions of High Efficiency Video Coding, IEEE Transactions on circuits and systems for video technology, Vol. 26, No. 1, January 2016, G. Tech; Y. Chen; K. Muller; J-R. Ohm; A. Vetro; Y-K. Wang".

[0049] Due to the similarity between HEVC and VVC, it should be possible to adapt the compression tools defined in the context of MV and 3D-HEVC to the context of VVC to obtain a codec that can handle multi-view content (regardless of the presence or absence of depth information). As in the context of HEVC, in the context of VVC, the base layer of multi-view content containing only texture information should be fully compatible with VVC.

[0050] Figures 6, 7, and 8 recall some of the main features of the basic compression methods that can be used to encode the base layer of multi-view content.

[0051] Figure 6 shows an example of a split by the image of pixel 11 of the original video 10. Here, the pixel is considered to consist of three components corresponding to the base layer of multi-view content, namely the luminance component and two chrominance components. The same split can be applied to all layers of multi-view content, namely the texture layer and the depth layer. In addition, the same split can be applied to a different number of components or layers, for example, a texture layer and a depth layer containing four components (luminance component, two chrominance components, and transparency component).

[0052] The image is divided into a plurality of encoding entities. First, as shown by reference numeral 13 in FIG. 6, the image is divided into a grid of blocks called coding tree units (CTUs). A CTU consists of an N×N block of luminance samples and two corresponding blocks of chrominance samples. N is generally a power of 2 having a maximum value of, for example, "128". Second, the image can be divided into one or more tile rows and tile columns, where a tile is a sequence of CTUs covering a rectangular region of the image. In some video compression schemes, a tile can be divided into one or more bricks, each of which consists of at least one CTU row within the tile. Above the concepts of tiles and bricks, there exists another encoding entity called a slice that can contain at least one tile of the image or at least one brick of a tile.

[0053] In the example of FIG. 6, as shown by reference numeral 12, the image 11 is divided into three slices S1, S2, and S3, and each slice contains a plurality of tiles (not shown).

[0054] As shown by reference numeral 14 in FIG. 6, a CTU can be divided into a hierarchical tree of one or more sub-blocks called coding units (CUs). A CTU is the root of the hierarchical tree (i.e., the parent node) and can be divided into a plurality of CUs (i.e., child nodes). Each CU becomes a leaf of the hierarchical tree if it is not further divided into smaller CUs, and becomes the parent node of smaller CUs (i.e., child nodes) if it is further divided. Different types of hierarchical trees can be used, where the CTU or CU is a quadtree divided into four square CUs or CUs of equal size, and in a binary tree, the CTU (CU) can be divided horizontally or vertically into "two" rectangular CUs of equal size. In a ternary tree, the CTU (CU) can be divided horizontally or vertically into "three" rectangular CUs.

[0055] In the example of FIG. 6, the CTU 14 is first divided into “4” square CUs using a quadtree type of division. The upper left CU is not further divided and is thus a leaf of the hierarchical tree, i.e., not the parent node of other CUs. The upper right CU is further divided into “4” smaller square CUs using again a quadtree type of division. The lower right CU is vertically divided into “2” rectangular CUs using a binary tree type of division. The lower left CU is vertically divided into “3” rectangular CUs using a ternary tree type of division.

[0056] During the encoding of the image, the division is adaptive and each CTU is divided so as to optimize the compression efficiency criteria.

[0057] In some video compression schemes, the concepts of prediction unit (PU) and transform unit (TU) have emerged. In fact, in this case, the encoding entities used for prediction (i.e., PU) and transform (i.e., TU) can be parts of the CU. For example, as shown in FIG. 6, a CU of size 2N×2N can be divided into PUs 1411 of size N×2N or size 2N×N. Further, the said CU can be divided into “4” TUs 1412 of size N×N or

[0058]

Number

[0059] In this application, the terms “block” or “image block” can be used to refer to any one of CTU, CU, PU, and TU. Further, the terms “block” or “image block” can be used to refer to macroblock, partition, and sub-block, and more generally, can be used to refer to an array of samples of many sizes.

[0060] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image" and "picture" may be used interchangeably. In the context of MVD data, similar to the AU of FIG. 24, at time T, a frame is considered an entity that includes an image (texture and depth) corresponding to time T for each view. Usually, the term "reconstruction" is used on the encoder side, and "decoding" is used on the decoder side.

[0061] FIG. 7 schematically shows a method for encoding a video stream executed by an encoding module. Although variations of this method for encoding are conceivable, for the sake of clarity, the method for encoding in FIG. 7 will be described below without describing all the conceivable variations. In particular, the described method for encoding is applied to the base layer of multi-view content, and each pixel of the base layer includes a luminance component and two chrominance components. Specific encoding tools adapted for encoding multi-view content, particularly depth layers, will not be described further.

[0062] The encoding of the current original image 501 starts from the division of the current original image 501 during step 502, as described in relation to FIG. 6. Thus, the current image 501 is divided into CTUs, CUs, PUs, TUs, etc. For each block, the encoding module determines an encoding mode between intra prediction and inter prediction.

[0063] Intra prediction consists of predicting the pixels of the current block from a predicted block derived from the pixels of a reconstructed block located in the causal neighborhood of the current block being encoded, according to an intra prediction method, during step 503. The result of intra prediction is a prediction direction indicating which pixels of the neighboring block are used and a residual block resulting from the calculation of the difference between the current block and the predicted block.

[0064] Inter prediction consists of predicting the pixels of the current block from a block of pixels of an image before or after the current image, called the reference block, and this image is called the reference image. During the encoding of the current block by the inter prediction method, the block of the reference image that is closest to the current block according to a similarity criterion is determined by the motion estimation step 504. During step 504, a motion vector indicating the position of the reference block within the reference image specified by the index is determined. The motion vector and the index of the reference image are used during the motion compensation step 505, during which the residual block is calculated in the form of the difference between the current block and the reference block. Note that only one prediction, inter prediction, is described here. Also, there is a two-prediction inter prediction (or B mode) in which the current block is associated with two motion vectors, specifying two reference blocks in two different images (each specified by a reference image index), and the residual block of this block is the average of the two residual blocks.

[0065] Note that intra prediction and inter prediction are general terms that include many modes based on the general principles of spatial and temporal prediction.

[0066] During selection step 506, among the tested prediction modes, the prediction mode that optimizes compression performance according to the rate / distortion criterion is selected by the encoding module. Once the prediction mode is selected, the residual block is transformed during step 507 and quantized during step 509. Note that the encoding module can skip the transformation and apply quantization directly to the untransformed residual signal. When the current block is encoded according to intra prediction, the prediction direction and the residual block that is transformed and quantized are encoded by the entropy encoder during step 510. When the current block is encoded according to inter prediction, the motion vector of the block is predicted from a prediction vector selected from a set of motion vectors corresponding to reconstructed blocks located near the block to be encoded. The motion information (including motion vector residual, index of motion vector predictor, index of reference image) is then encoded by the entropy encoder during step 510 in the form of an index for identifying the motion residual and the prediction vector. The transformed and quantized residual block is encoded by the entropy encoder during step 510. Note that the encoding module can bypass both the transformation and the quantization, i.e., the entropy encoding is applied to the residual without applying the transformation process or the quantization process. The result of the entropy encoding is inserted into the encoded video stream 511.

[0067] After quantization step 509, the current block is reconstructed so that the pixels corresponding to that block can be used for future prediction. This reconstruction stage is also called the prediction loop. Thus, inverse quantization is applied to the residual block that was transformed and quantized during step 512, and inverse transformation is applied during step 513. Depending on the prediction mode used for the block obtained during step 514, the predicted block of the block is reconstructed. If the current block is encoded according to inter prediction, the encoding module applies motion compensation that uses the motion vector of the current block to identify the reference block of the current block during step 516. If the current block is encoded according to intra prediction, the prediction direction corresponding to the current block is used to reconstruct the reference block of the current block during step 515. To obtain the reconstructed current block, the reference block and the reconstructed residual block are added.

[0068] After reconstruction, during step 517, in-loop post-filtering intended to reduce encoding artifacts is applied to the reconstructed block. This post-filtering is performed in the prediction loop to obtain the encoding of the same reference picture as the decoder and thus avoid drift between encoding and decoding, and is therefore called in-loop post-filtering. For example, in-loop post-filtering includes deblocking filtering and SAO (Sample Adaptive Offset) filtering. A parameter indicating activation or deactivation of the in-loop deblocking filter, and the characteristics of the in-loop deblocking filter if activated, are introduced into the encoded video stream 511 during the entropy encoding step 510.

[0069] When a block is reconstructed, the block is inserted into the reconstructed image stored in the reconstructed image memory 519 during step 518, which memory is also referred to as a reference image memory, a reference image buffer, or a decoded picture buffer (DPB). The reconstructed image stored in this way can function as a reference image for other images to be encoded.

[0070] FIG. 8 schematically shows a method for decoding an encoded video stream 511 encoded according to the method described in relation to FIG. 7, performed by a decoding module. Although variants of this method for decoding are conceivable, for the sake of clarity, the method for decoding in FIG. 8 will be described below without describing all the conceivable variants.

[0071] Decoding is performed block by block. For the current block, this starts with the entropy decoding of the current block during step 610. Entropy decoding makes it possible to obtain the prediction mode of the block.

[0072] When the block is encoded according to inter prediction, entropy decoding makes it possible to obtain a prediction vector index, a motion residual, an index of a reference image, and a residual block. During step 608, a motion vector is reconstructed for the current block using the prediction vector index and the motion residual.

[0073] When a block is encoded according to intra prediction, entropy decoding enables obtaining the prediction direction and the residual block. Steps 612, 613, 614, 615, 616, and 617 implemented by the decoding module are all identical to steps 512, 513, 514, 515, 516, and 517 implemented by the encoding module, respectively. The decoded block is stored in the decoded image, and the decoded image is stored in the DPB 619 in step 618. When the decoding module decodes a given image, the image stored in the DPB 619 is identical to the image stored in the DPB 519 by the encoding module during the encoding of the given image. The encoded image can also be output by the decoding module, for example, for display.

[0074] FIG. 1 schematically shows an example of a camera array adapted to obtain MVD content.

[0075] FIG. 1 represents a camera array 10 including "16" cameras 10A to 10P positioned on a 4×4 grid. Each camera of the camera array 10 is focused on the same scene and can obtain, for example, an exemplary image of an image in which pixels include a luminance component and two chrominance components. For example, computing means or measuring means (not shown) connected to the camera array 10 are used to generate a depth map of each image generated by the cameras of the camera array 10. In the example of FIG. 1, each depth map associated with an image has the same resolution as the image (i.e., the depth map includes the depth value of each pixel of the image). Thus, the camera array 10 generates MVD content including "16" texture layers and "16" depth layers. In such a type of camera array, the overlap between the captured views is important. Improving the overall compression ratio achievable with such multi-view content is the goal of the following embodiments.

[0076] Each camera of the camera array 10 is associated with internal and external camera parameters. As will be described later in this specification, these parameters are required by the decoder to create a predicted image. In one embodiment, the internal and external parameters are provided to the decoder in the form of an SEI (Supplemental Enhancement Information) message. The SEI message is defined in H.264 / AVC and HEVC for transmitting metadata.

[0077] Table TAB1 describes the syntax of the SEI message adapted to transmit the internal and external parameters of the camera array. This syntax is the same as the syntax of the multi-view acquisition information SEI message syntax in HEVC (Section G.14.2.6).

[0078] [Table 1]

[0079] One of the goals of the embodiments described below is to improve the prediction of one view based on at least one other view. To target the multi-view content captured by the camera array as described above, any of the cameras can provide a good prediction for some or all of the neighboring views. To create a predicted image of the current image of the current view, previously decoded views and their associated camera parameters as well as the camera parameters associated with the current view are used. Note that either the texture layer or the depth layer within the view can use the new vsp mode.

[0080] a camera calibrated as a planar pinhole and

[0081] [Number] Consider the internal matrix of the camera. ·f refers to the distance from the exit pupil to the camera's sensor, which is expressed in pixels and is often mistakenly called the "focal length" in the literature. In Table TAB1, this information is described by the following set of parameters.

[0082]

Table 2

[0083]

Number

[0084]

Table 3

[0085]

Table 4

[0086]

Table 5

[0087]

Number

[0088]

Number

[0089]

Mathematics

[0090]

Mathematics

[0091]

Mathematics

[0092]

Mathematics

[0093]

Mathematics

[0094] For each camera, in Table TAB1, the R and T matrices are described as follows.

[0095]

Table 6

[0096]

Mathematics

[0097]

Mathematics

[0098] Here, assume that a given camera c provides the current view. The camera c is associated with an internal matrix K c and a pose matrix P c .

[0099]

Number

[0100]

Number

[0101]

Number

[0102]

Number

[0103] Figure 9 schematically shows an example of a method for encoding an encoded video stream representing MVD content.

[0104] The method of FIG. 9 is a method that enables encoding of a first view 501 and a second view 501B. In this example, as shown in FIG. 24, the first view includes a base layer (layer "0") containing texture data and a layer "1" containing depth data. The second view includes a layer "2" containing texture data and a layer "3" containing depth data. For simplicity of representation, only two views are shown as being encoded, but more views can be encoded by the method of FIG. 9. For example, "16" views generated by the camera array 10 can be encoded by the method for encoding of FIG. 9.

[0105] In embodiment (9a), the first view 501 is considered a root view from which all other views are directly or indirectly predicted. The first view 501 is encoded without any inter-view or inter-layer prediction. In one embodiment, layer "0" and layer "1" are encoded separately either in parallel or sequentially. In one embodiment, layer "0" and layer "1" are encoded using the same steps 502, 503, 504, 505, 506, 507, 508, 509, 510, 512, 513, 514, 515, 516, 517, 518, and 519 described with reference to FIG. 7. In other words, the texture and depth data of the first view 501 are encoded using the method of FIG. 7 (corresponding to arrow TIIP in FIG. 24).

[0106] In embodiment (9b), layer "0" is encoded using the method of FIG. . However, the method is slightly modified for layer "1" to incorporate the modes defined in 3D HEVC and predict the depth layer of the view from the texture layer of the view (corresponding to arrow ICP in FIG. 24).

[0107] It should be noted that there seems to be a missing reference number in the description of embodiment (9b) where it says "layer '0' is encoded using the method of FIG. ". Please check and correct if necessary.In Embodiment (9c), the texture layer (layer "2") of the second view 501B is encoded by a process including steps 502, 503, 504, 505B, 507, 508, 509, 510, 512, 513, 514, 515, 516, 517, 518, and 519 and the same steps 502B, 503B, 504B, 505B, 507B, 508B, 509B, 510B, 512B, 513B, 515B, 516B, 517B, 518B, and 519B respectively.

[0108] A new predicted image is generated by processing module 20 in step 521 and introduced into DPB 519B. This new predicted image is used by processing module 20 in step 522 to determine a predictor, called the signal of the VSP predictor, for the current block of the current image of the texture layer of the second view 501B. The prediction by the VSP predictor corresponds to a new VSP mode, hereinafter simply also called the VSP mode.

[0109] The VSP mode is very similar to the conventional inter mode. In fact, when introduced into DPB 519B, the new predicted image generated during step 521 is treated as a normal reference image for temporal prediction (even if the predicted image generated during step 521 is placed at the same temporal location as the current image of the texture layer of the second view 501B). Thus, the new VSP mode can be considered as an inter mode that uses a specific reference image generated by inter-view prediction. Step 522 includes a motion estimation step and a motion compensation step. The blocks encoded using the VSP mode are encoded in the form of motion information and residuals, and the motion information includes an identifier of the predicted image generated during step 521.

[0110] During step 506B, processing module 20 performs steps different from step 506 in that the VSP predictor generated during step 522 is considered in addition to the normal intra and inter predictors. Similarly, processing module 20 performs a different step 514B from step 514 in that the new VSP mode belongs to the set of prediction modes that can potentially be applied to the current block. If a new VSP mode is selected for the current block during step 506B, processing module 20 reconfigures the corresponding VSP predictor during step 523.

[0111] In embodiment (9d), the depth layer (layer "3") of the second view 501B is encoded using the same steps 502B, 503B, 504B, 505B, 506B, 507B, 508B, 509B, 510B, 512B, 513B, 514B, 515B, 516B, 517B, 518B, 519B, 521, 522, and 523. Thus, the new VSP mode is applied to the depth layer (layer "3") of the second view 501B. More generally, the VSP mode can be applied to the depth layer of a view predicted from another view.

[0112] In embodiment (9e), the encoding of layer "3" incorporates the modes defined in 3D HEVC and predicts the depth layer of the view from that view (corresponding to the arrow ICP in FIG. 24).

[0113] In the example of FIG. 9, the second view 501B is encoded from the first view 501 in which at least the texture layer "0" is encoded without any inter-view prediction. When two or more views are encoded by the method of FIG. 9, any third view can be encoded from a view in which at least the texture layer is encoded without any inter-view prediction (e.g., from the first view 501), or from a view in which the texture layer is encoded with inter-view prediction (e.g., from the second view 501B).

[0114] As shown, the encoding method of FIG. 9 includes two encoding layers, one for each view. Of course, when three or more views are encoded, the encoding method of FIG. 9 is assumed to include the same number of encoding layers as the number of views. In the example of FIG. 9, each encoding layer has its own DPB. In other words, each view is associated with its own DPB.

[0115] As described below, the images in the DPB are indexed by a plurality of reference indices: · ref_idx: Index of the reference image used in the DPB · ref_idx_l0: Index of the reference image used in list l0 of the reference images stored in the DPB. To decode view i at frame T, the list referred to by ref_idx_l0 includes views of view i from different frames. · ref_idx_l1: Index of the reference image used in list l1 of the reference images stored in the DPB. To decode view i at frame T, the list referred to by ref_idx_l1 includes views of view i from different frames. · ref_idx2: Index of the reference image to be used among the reference images that temporally correspond to the current image. The index ref_idx2 refers only to the images generated by forward projection. To decode view i at frame T, the list referred to by ref_idx2 includes the reference images corresponding to frame T.

[0116] FIG. 10 schematically shows an example of a method for decoding an encoded video stream representing multi-view content.

[0117] In an embodiment (10a) corresponding to embodiment (9a), layer "0" and layer "1" are decoded separately either in parallel or sequentially. In one embodiment, "0" and layer "1" are decoded using the same steps 608, 610, 612, 613, 614, 615, 616, 617, 618, 619 described in connection with FIG. 8. In other words, the texture and depth data of the first view 501 are encoded using the method of FIG. 8 (corresponding to arrow TIIP in FIG. 24).

[0118] In an embodiment (10b) corresponding to embodiment (9b), layer "0" is decoded using the method of FIG. 8, which is slightly modified for layer "1" to incorporate the modes defined in 3D HEVC (corresponding to arrow ICP in FIG. 24).

[0119] In an embodiment (10c) corresponding to embodiment (9c), the texture layer (layer "2") of the second view 501B is decoded by a process including steps 608B, 610B, 612B, 613B, 615B, 616B, 617B, 618B, 619B that are the same as steps 608, 610, 612, 613, 615, 616, 617, 618, 619 respectively. In step 621, the processing module 20 generates a new prediction image identical to the image generated during step 521 and introduces this image into the DPB 619B. The processing module 20 performs a step 614B that is different from step 614 in that the new VSP mode belongs to a set of prediction modes that can potentially be applied to the current block. If a new VSP mode is selected for the current block during step 506B, the processing module 20 reconfigures the corresponding VSP predictor during step 623.

[0120] In an embodiment (10d) corresponding to embodiment (9d), the depth layer (layer "3") of the second view 501B is decoded using the same steps 608B, 610B, 612B, 613B, 614B, 615B, 616B, 617B, 618B, 619B, 621, and 623.

[0121] As shown, the decoding method of FIG. 10 includes two encoding layers, one for each view. Of course, if three or more views are decoded, the decoding method of FIG. 9 shall include the same number of decoding layers as the number of views. In the example of FIG. 9, each decoding layer has its own DPB. In other words, each view is associated with its own DPB.

[0122] FIG. 2 schematically represents a processing module adapted to encode MVD content provided by a camera array.

[0123] In FIG. 2, a simplified representation of a camera array 10 including only two cameras 10A and 10B is shown. Each camera of the camera array 10 communicates with the processing module 20 using a communication link that can be wired or wireless. In FIG. 2, the processing module 20 encodes the multi-view content generated by the camera array 10 using the new VSP mode described below within the encoded video stream.

[0124] FIG. 3 schematically represents a processing module adapted to decode an encoded video stream representing MVD content.

[0125] In FIG. 3, the processing module 20 decodes the encoded video stream. The processing module 20 is connected by a communication link that can be wired or wireless to a display device 26 capable of displaying the images resulting from the decoding. The display device is, for example, a virtual reality headset, a 3D TV, or a computer display.

[0126] FIG. 4 schematically shows an example of the hardware architecture of a processing module 20 that can implement an encoding module or a decoding module capable of implementing various embodiments described below. The processing module 20 includes, by way of non-limiting example, a processor or CPU (Central Processing Unit) 200 that includes one or more microprocessors, general-purpose computers, dedicated computers, and processors based on a multi-core architecture connected by a communication bus 205, a random access memory (RAM) 201, a read-only memory (ROM) 202, and a storage device 203 that may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash, magnetic disk drive, and / or optical disk drive, or a storage medium reader such as an SD (Secure Digital) card reader and / or a hard disk drive (HDD) and / or a network-accessible storage device, and at least one communication interface 204 for exchanging data with other modules, devices, or apparatuses. The communication interface 204 can include, but is not limited to, a transceiver configured to transmit and receive data via a communication channel. The communication interface 204 can include, but is not limited to, a modem or a network card.

[0127] When the processing module 20 implements a decoding module, the communication interface 204 enables, for example, the processing module 20 to receive an encoded video stream and provide a decoded video stream.

[0128] When the processing module implements an encoding module, the communication interface 204 enables, for example, the processing module 20 to receive original image data and encode and provide an encoded video stream.

[0129] The processor 200 can execute instructions loaded from the ROM 202, an external memory (not shown), a storage medium, or a communication network into the RAM 201. When the processing module 20 is powered on, the processor 200 can read instructions from the RAM 201 and execute them. These instructions form, for example, a computer program that causes the processor 200 to implement the decoding method described in connection with FIG. 9 or the encoding method described in connection with FIG. 10, and the decoding method and the encoding method include various aspects and embodiments described later in this specification.

[0130] All or part of the algorithms and steps of the encoding or decoding method may be implemented in software form by the execution of an instruction set by a programmable machine such as a DSP (Digital Signal Processor) or a microcontroller, or may be implemented in hardware form by a machine or a dedicated component such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0131] FIG. 5 shows a block diagram of an example of a system 2 in which various aspects and embodiments are implemented. System 2 can be embodied as a device that includes various components described below and is configured to execute one or more of the aspects and embodiments described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, virtual reality headsets, and servers. The elements of system 2 can be embodied alone or in combination as a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, system 2 includes one processing module 20 that implements a decoding module or an encoding module. However, in another embodiment, system 2 can include one processing module 20 that implements a decoding module and one processing module 20 that implements a decoding module, or one processing module 20 that implements a decoding module and an encoding module. In various embodiments, system 2 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or via dedicated input and / or output ports. In various embodiments, system 2 is configured to implement one or more of the aspects and embodiments described in this document.

[0132] In one embodiment, system 2 includes at least one processing module 20 that can implement one or both of an encoding module or a decoding module.

[0133] The input to processing module 20 can be provided through various input modules as shown in block 22. Such input modules include, but are not limited to, (i) a radio frequency (RF) module that receives, for example, an RF signal transmitted wirelessly from a broadcasting station, (ii) a component (COMP) input module (or a set of COMP input modules), (iii) a universal serial bus (USB) input module, and / or (iv) a high definition multimedia interface (HDMI) input module. Other embodiments, not shown in FIG. 5, include composite video.

[0134] In various embodiments, the input module of block 22 has respective input processing elements known in the art. For example, an RF module may be associated with appropriate elements to (i) select a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band), (ii) down-convert the selected signal, (iii) in certain embodiments, band-limit to a narrower frequency band again to select a signal frequency band, which may be referred to as a channel for example, (iv) demodulate the down-converted and band-limited signal, (v) perform error correction, and (vi) demultiplex to select a desired stream of data packets. The RF modules of various embodiments may include one or more elements to perform these functions, such as a frequency selector, signal selector, band limiter, channel selector, filter, down-converter, demodulator, error corrector, and demultiplexer. The RF portion may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency closer to baseband) or to baseband. In one embodiment of a set-top box, the RF module and its associated input processing elements perform frequency selection by receiving an RF signal transmitted via a wired (e.g., cable) medium and filtering, down-converting, and re-filtering to a desired frequency band. In various embodiments, the order of the above-described (and other) elements may be rearranged, some of these elements may be deleted, and / or other elements performing similar or different functions may be added. Adding elements may include, for example, inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF module includes an antenna.

[0135] Furthermore, the USB module and / or the HDMI module can each include an interface processor for connecting System 2 to other electronic devices via a USB connection and / or an HDMI connection. It should be understood that various aspects of the input processing, such as Reed-Solomon error correction, can be implemented, for example, in a separate input processing IC or within the processing module 20 as needed. Similarly, aspects of the USB or HDMI interface processing can be implemented, as needed, within a separate interface IC or within the processing module 20. The demodulated, error-corrected, and de-multiplexed stream is provided to the processing module 20.

[0136] The various elements of System 2 can be provided within an integrated housing. Within the integrated housing, the various elements are interconnected and can transmit data between them using an internal bus known in the art, including a suitable connection arrangement such as an Inter-IC (I2C) bus, wiring, and a printed circuit board. For example, in System 2, the processing module 20 is interconnected to the other elements of the system by a bus 205.

[0137] The communication interface 204 of the processing module 20 enables System 2 to communicate over a communication channel 21. The communication channel 21 can be implemented, for example, within a wired and / or wireless medium.

[0138] In various embodiments, data is streamed to system 2 or otherwise provided using a Wi-Fi network, such as a wireless network like IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals of these embodiments are received via communication channel 21 and communication interface 204 adapted for Wi-Fi communication. The communication channel 21 of these embodiments is generally connected to an access point or router that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. In another embodiment, a set-top box that distributes data via the HDMI connection of input block 22 is used to provide streaming data to system 2. In yet another embodiment, the RF connection of input block 22 is used to provide streaming data to system 2. As noted above, various embodiments provide data in ways other than streaming. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks. The data provided to system 2 includes, for example, the MVD signal provided by the array of cameras 10.

[0139] System 2 can provide output signals to various output devices, including a display 26 via a display interface 23, a speaker 27 via an audio interface 24, and other peripheral devices 28 via an interface 25. The display 26 in various embodiments includes, for example, one or more of a touch screen display, an organic light emitting diode (OLED) display, a curved display, and / or a foldable display. The display 26 can be for a television, a tablet, a laptop, a mobile phone, a smartphone, a virtual reality headset, or other devices. The display 26 may also be integrated with other components (such as in a smartphone) or separate (such as an external monitor for a laptop). Examples of other peripheral devices 28 include, in various examples of the embodiments, one or more of a stand-alone digital video disc (or digital versatile disc) (abbreviated as DVR for both terms), a disc player, a stereo system, and / or an illumination system. Various embodiments use one or more peripheral devices 28 that provide functions based on the output of system 2. For example, a disc player performs a function of playing back the output of system 2.

[0140] In various embodiments, the control signal is communicated between the system 2 and the display 26, speaker 27, or other peripheral device 28 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable control between devices regardless of the presence or absence of user intervention. The output devices can be communicatively coupled to the system 2 via dedicated connections through their respective interfaces 23, 24, and 25. Alternatively, the output devices can be connected to the system 2 using the communication channel 21 via the communication interface 204. The display 26 and speaker 27 can be integrated into a single unit with other components of the system 2 within an electronic device such as a television, for example. In various embodiments, the display interface 23 includes a display driver, such as a timing controller (T Con) chip, for example.

[0141] Alternatively, for example, if the RF module of the input 22 is part of an individual set-top box, the display 26 and speaker 27 can be separated from one or more of the other components. In various embodiments where the display 26 and speaker 27 are external components, the output signal can be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.

[0142] Various implementations include decoding. As used in this application, "decoding" can include all or part of the process performed on a received encoded video stream to produce a final output suitable for display, for example. In various embodiments, such processing includes one or more of the processes commonly performed by a decoder, such as entropy decoding, inverse quantization, inverse transformation, and prediction, for example. In various embodiments, such a process also includes or alternatively includes, for example, processes performed by the decoders of various implementations described in this application, such as for decoding a new VSP mode.

[0143] Whether the phrase "decryption process" is intended to specifically refer to a subset of operations or to generally refer to a broader decryption process will become apparent based on the specific context of the description and is considered well understood by those skilled in the art.

[0144] Various implementations include encoding. Similar to the above considerations regarding "decryption", as used in this application, "encoding" may include, for example, all or part of the process executed on an input video sequence to generate an encoded video stream. In various embodiments, such processing may include, for example, one or more of the processes generally performed by an encoder, such as splitting, prediction, transformation, quantization, and entropy encoding. In various embodiments, such processing may, in addition to or alternatively, include the processing performed by the encoders of various implementations for decryption according to, for example, the new VSP mode described in this application.

[0145] Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or to generally refer to a broader encoding process will become apparent based on the specific context of the description and is considered well understood by those skilled in the art.

[0146] Note that syntax elements used in this specification, such as the flag VSP and the index ref_idx2, are descriptive terms. Therefore, they do not exclude the use of other syntax element names.

[0147] When a figure is presented as a flowchart, it should be understood that the figure also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that the figure also provides a flowchart of the corresponding method / process.

[0148] Various embodiments refer to rate distortion optimization. In particular, during the encoding process, the balance or trade-off between rate and distortion is typically considered, in many cases, to impose constraints on the computational complexity. Rate distortion optimization is typically formulated to minimize a rate distortion function that is a weighted sum of rate and distortion. There are various approaches to solving the rate distortion optimization problem. For example, these approaches can be based on an extensive test of all encoding options that include all considered modes or encoding parameter values, involving a complete evaluation of their encoding cost and the associated distortion of the reconstructed signal after encoding and decoding. Also, to reduce the encoding complexity, faster approaches can be used, in particular, using the calculation of approximate distortion based on the predicted or prediction residual signal rather than the reconstructed signal. These two approaches can also be used in combination. For example, approximate distortion can be used for only some of the possible encoding options, while complete distortion can be used for other encoding options. In another approach, only a subset of the possible encoding options is evaluated. More generally, many approaches employ any of various techniques to perform the optimization, but the optimization is not necessarily a complete evaluation of both the encoding cost and the associated distortion.

[0149] The implementation forms and modes described in this specification can be implemented, for example, in a method or process, an apparatus, a software program, a data stream, or a signal. Even if described only in the context of a single form of implementation (for example, described only as a method), the implementation form of the described features can also be implemented in other forms (for example, an apparatus or a program). The apparatus can be implemented, for example, in appropriate hardware, software, and firmware. The method can be implemented, for example, in a processor, and the processor refers to a general processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor can also include, for example, a communication device such as a computer, a mobile phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate the communication of information between the end user.

[0150] References to "one embodiment" or "an embodiment" or "one implementation form" or "an implementation form", and other variations thereof, mean that the specific features, structures, characteristics, etc. described in connection with the embodiment are included in at least one embodiment. Accordingly, the occurrences of the phrases "in one embodiment" or "in an embodiment" or "in one implementation form" or "in an implementation form" that appear in various places in this specification, as well as any other variations, do not necessarily all refer to the same embodiment.

[0151] In addition, this application may refer to "determining" various information. Determining information can include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.

[0152] Furthermore, this application may refer to "accessing" various information. Accessing information can include, for example, one or more of receiving information, obtaining information (e.g., from memory), storing information, moving information, copying information, computing information, determining information, predicting information, or estimating information.

[0153] In addition, this application may refer to "receiving" various information. Receiving is intended to be a broad term, similar to "accessing". Receiving information can include, for example, one or more of accessing information or obtaining information (e.g., from memory). Further, "receiving" generally involves in some form during operations such as storing information, processing information, transmitting information, moving information, copying information, deleting information, computing information, determining information, predicting information, or estimating information.

[0154] The use of any of " / ", "and / or", "at least one of", "one or more", for example, in the case of "A / B", "A and / or B", "at least one of A and B", "one or more of A and B", is intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", "one or more of A, B, and C", such phrases are intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be apparent to those of ordinary skill in the art and the relevant art, this can be extended to any number of listed items.

[0155] Also, as used herein, the term "signaling" specifically means indicating something to the corresponding decoder. For example, in certain embodiments, the encoder signals information representing a new VSP mode. Thus, in some embodiments, the same parameters are used on both the encoder side and the decoder side. Accordingly, for example, the encoder can send (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has those specific parameters and other parameters, implicit signaling can be used, i.e., no transmission is made simply to enable the decoder to recognize and select those specific parameters. By avoiding the transmission of actual functionality, bit savings are achieved in various embodiments. It should be understood that signaling can be achieved in various ways. For example, one or more syntax elements, flags, etc. are used in various embodiments to signal information to the corresponding decoder. The above explanation pertains to the verb form of the word "signal", but the word "signal" can also be used as a noun herein.

[0156] As will be apparent to those skilled in the art, in an implementation form, for example, various signals can be generated that are formatted to convey information that can be stored or transmitted. Such information can include, for example, instructions for executing a method or data generated by one of the described implementation forms. For example, a signal can be formatted to convey the encoded video stream of the described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding the encoded video stream and modulating a carrier wave with the encoded video stream. The information conveyed by the signal can be analog information or digital information, for example. The signal can be transmitted via various different wired or wireless links, as is known. The signal can be stored in a processor-readable medium.

[0157] FIG. 13A schematically shows an example of a forward projection method used in a predicted image generation process. FIG. 13B is another representation of the forward projection method of FIG. 13A. The forward projection processes of FIGS. 13A and 13B are used during steps 521 and 621. The forward projection process is applied to the pixels of the first view acquired by the first view of camera m to project these pixels from the camera coordinate system of the first view to the camera coordinate system of the second view acquired by camera n. Each pixel is considered to include texture information and depth information.

[0158] In step 130, the processing module 20 back-projects the current pixel P(u, v) of the first view from the camera coordinate system of the first view to a reference coordinate system (i.e., the world coordinate system) to obtain the back-projected pixel P as shown in FIG. 13B. w to obtain. The back-projection is performed using the pose matrix P of camera m m , the inverse internal matrix of camera m

[0159]

Number

[0160] In step 131, the processing module 20 projects the back-projected pixel P w into the coordinate system of the second view using the internal matrix K n of camera n and the external matrix Q n Again, the internal matrix K n and the external matrix Q n are defined using the camera parameters obtained by the processing module 20 from, for example, the SEI message described in Table TAB1. When this projection does not fall within the camera n region, this projection is rejected. When it falls within the camera n region, this projection almost certainly does not fall into the actual pixel and is between the "4" pixels.

[0161] w In step 132, the processing module 20 selects the pixel P'(u', v') of the grid of the pixels of camera n that is closest to the projection of the back-projected pixel P w For example, the closest pixel is, for example, a pixel of the grid of the pixels of camera n and minimizes the distance to the projection of the back-projected pixel P w This distance is calculated as the square root of the sum of the squares (or the sum of the absolute differences) between the coordinates of the grid pixel and the projection of the back-projected pixel P

[0162] The pixel P'(u', v') obtained by the forward projection process of FIGS. 13A and 13B holds the texture value and the depth value of the projected pixel P(u, v). The set of pixels P'(u', v') forms a projected image.

[0163] FIG. 14A shows a first embodiment of the predicted image generation process.

[0164] In the embodiment of FIG. 14A, also referred to as Embodiment (14A), one view is signaled to be used as a possible predictor for reconfiguring the current view.

[0165] The process described in FIG. 14A is executed during step 521 of the encoding method of FIG. 9 and during step 621 of the decoding method of FIG. 10, and includes steps 140 to 143 of generating a reference image from a first view 501 and encoding a current image of a second view 501B.

[0166] In step 140, the processing module 20 acquires camera parameters (i.e., view parameters) for a reference view (e.g., the first view 501) and a current view (e.g., the second view 501B). When the process of FIG. 14A is executed during step 521, the processing module 20 acquires these parameters directly from the cameras of the camera array 10 or from the user. When the process of FIG. 14A is executed during step 621, the processing module 20 acquires these parameters from an SEI message (e.g., the SEI message described in Table TAB1) or from the user.

[0167] In step 141, the processing module 20 applies the forward projection method described in FIGS. 13A and 13B between the reference view (e.g., the first view 501) and the current view (e.g., the second view 501B) to generate a predicted image G(k). The predicted image G(k) is introduced into the DPB of the current view (e.g., DPB519B(619B)) and is intended to be one of the Kth predicted images for the current image of the current view. th of.

[0168] Each pixel of the predicted image G(k) resulting from a successful prediction holds the texture value and depth value of the corresponding projected pixel of the reference view. After forward projection, the separated missing pixels may remain (due to failed projections that do not enter the second view area).

[0169] In step 142, the processing module 20 fills in the isolated missing pixels. In one embodiment, the isolated missing pixels are filled with the average of the adjacent pixel values. In another embodiment, the isolated missing pixels are filled with the median of the adjacent pixel values. In another embodiment, the isolated missing pixels are filled with a default value (typically 128 for values encoded in 8 bits).

[0170] In step 143, the processing module 20 stores the predicted image G(k) in the DPB of the current view.

[0171] In step 144, the processing module 20 reconstructs the current image of the current view using the reference images included in the DPB of the current view, and the DPB includes the predicted image G(k).

[0172] When the process of FIG. 14A is applied to the encoding method of FIG. 9, the generation includes steps 502B, 503B, 504B, 505B, 506B, 507B, 508B, 509B, 510B, 512B, 513B, 514B, 515B, 516B, 517B, 518B, 519B, 522, and 523.

[0173] When the process of FIG. 14A is applied to the decoding method of FIG. 10, the generation includes steps 608B, 610B, 612B, 613B, 614B, 615B, 616B, 617B, 618B, 619B, and 623.

[0174] FIG. 14B shows details of a second embodiment of the predicted image generation process.

[0175] In the embodiment of FIG. 14B, also referred to as embodiment (14B), several figures are available for predicting the current view. For example, returning to FIG. 9, at time point T, the images (texture and depth) of the first view and the images of the second view are encoded and reconstructed, and the image of the third view is ready to be encoded using the predicted image generated from the reconstructed images of the first and second views. Similarly, when the inventors return to FIG. 10, at time point T, the images of the first view (texture and depth) and the second view are reconstructed, and the image of the third view is ready to be decoded using the predicted image generated from the reconstructed images of the first and second views.

[0176] In embodiment (14B), multiple views are used to generate one aggregated predicted image for reconstructing the current view. More precisely, in embodiment (14B), a predicted image is generated for each view of the multiple views, and the aggregated predicted image is generated from the multiple predicted images.

[0177] In step 140, the processing module 20 acquires the camera parameters (i.e., view parameters) of each view of the multiple views and the current view.

[0178] Compared with embodiment (14A), step 141 is replaced by steps 1411 - 1415.

[0179] In step 1411, the processing module 20 initializes the variable j to "0". The variable j is used to enumerate all the views of the multiple views.

[0180] In step 1412, the processing module 20 applies the forward projection method described in FIGS. 13A and 13B between view j and the current view to generate the predicted image G(k) j For example, view j is the first view 501 or the second view 501B, and the current view is the third view.

[0181] In step 1413, the processing module 20 compares the value of variable j with the number of views Nb_views in the plurality of views. If j < Nb_views, step 1414 follows step 1413, and j is incremented by 1 unit.

[0182] After step 1414, step 1412 follows, and a new predicted image G(k) j is generated.

[0183] If j = Nb_views, step 1415 follows step 1413, and the predicted image G(k) j is aggregated to generate an aggregated predicted image G(k) intended to be stored in the DPB of the current view.

[0184] In one embodiment of the aggregation process, the pixel values (texture values and depth values) of the first predicted image G(k) j of the plurality of predicted images are retained to aggregate the plurality of predicted images G(k) j The first predicted image is, for example, the predicted image G(k) j=0 .

[0185] In one embodiment of the aggregation process, the predicted image G(k) j is aggregated by retaining the pixel values (texture values and depth values) of G(k) j generated from the view j closest to the current view. If several views are at the same distance from the current view (i.e., there are several closest views), one of the several closest views is randomly selected. For example, in the camera array 10, assume that only the first view generated by camera 10A and the second view generated by camera 10C are available for predicting the current view generated by camera 10B. Then, the first and second views are the views closest to the current view and are at the same distance to the current view. One of the first and second views is randomly selected and the pixel values are provided to the aggregated predicted image G(k).

[0186] In one embodiment of the aggregation process, the size G(k) of the predicted image j is aggregated by retaining the pixel values (texture values and depth values) of the predicted image G(k) with the best quality. For example, the information representing the quality of a pixel is the value of the quantization parameter applied to the transform block containing the pixel. The quality of the pixels of the predicted image G(k) j depends on the quality of the pixels (i.e., quantization parameters) of the image to which the forward projection is applied to obtain the predicted image G(k). j of the predicted image G(k) j is aggregated by retaining the pixel values (texture values and depth values) of the predicted image G(k)

[0187] In one embodiment of the aggregation process, the predicted image G(k) j is aggregated by retaining the pixel values (texture values and depth values) of the predicted image G(k) with the closest depth value (z-buffer algorithm). j is aggregated by retaining the pixel values (texture values and depth values) of the predicted image G(k).

[0188] In one embodiment of the aggregation process, the predicted image G(k) j is aggregated by retaining the pixel values (texture values and depth values) of the predicted image G(k) when the adjacent pixels of the aggregated predicted image G(k) j have already been predicted from the predicted image. j is aggregated by retaining the pixel values (texture values and depth values) of the predicted image G(k).

[0189] In one embodiment of the aggregation process, the predicted image G(k) j is aggregated by calculating the average, weighted average, or median of the pixel values (texture values and depth values) of the predicted image G(k) j to calculate the dimensions of the predicted image.

[0190] After step 1415, step 142 follows, and the processing module 20 fills in the separated missing pixels in the aggregated predicted image G(k).

[0191] In step 143, the aggregated predicted image G(k) is stored in the DPB of the current view.

[0192] In step 144, the processing module 20 reconstructs the current image of the current view using the reference image included in the DPB of the current view, and the DPB includes the aggregated prediction image G(k).

[0193] FIG. 15 shows a third embodiment of the prediction image generation process.

[0194] In the embodiment of FIG. 15, also referred to as embodiment (15), similar to embodiment (14B), several figures are available for predicting the current view.

[0195] In embodiment (15), a prediction image is generated for each view of a plurality of views. However, instead of generating an aggregated prediction image and inserting the aggregated prediction image into the DPB of the current view as in embodiment 14B, in embodiment (15), each generated prediction image is inserted into the DPB.

[0196] In step 140, the processing module 20 acquires the camera parameters (i.e., view parameters) of each view of the plurality of views and the current view.

[0197] In step 1501, the processing module 20 initializes the variable j to "0". The variable j is used to enumerate all views of the plurality of views.

[0198] In step 1502, the processing module 20 applies the forward projection method described in FIGS. 13A and 13B between view j and the current view to generate the prediction image G(k). j to generate.

[0199] In step 1503, the processing module 20 fills in the separated missing pixels in the aggregated prediction image G(k).

[0200] In this step 1504, the processing module 20 stores the prediction image G(k) in the DPB of the current view. j to store.

[0201] In step 1505, the processing module 20 compares the value of variable j with the number of views Nb_views in the plurality of views. If j < Nb_views, step 1506 follows step 1505, and j is incremented by 1 unit.

[0202] After step 1506, step 1502 follows, and a new predicted image G(k) j is generated.

[0203] If j = Nb_views, step 144 follows step 1505. In step 144, the processing module 20 reconstructs the current image of the current view using the reference images included in the DPB of the current view, and the DPB includes a plurality of predicted images G(k) j including.

[0204] In a variant of Embodiment 15, in addition to the predicted image G(k) j a aggregated predicted image generated from the predicted image G(k) j and / or an aggregated predicted image generated from a subset of the predicted image G(k) j is inserted into the DPB of the current view.

[0205] In a variant of Embodiment 15, instead of the predicted image G(k) j an aggregated predicted image generated from the predicted image G(k) j and an aggregated predicted image generated from a subset of the predicted image G(k) j are inserted into the DPB of the current view, or only an aggregated predicted image generated from a subset of the predicted image G(k) j is inserted into the DPB of the current view.

[0206] Figure 16 shows a fourth embodiment of the predicted image generation process.

[0207] The objective of the embodiment (16), also referred to as Embodiment (16), of FIG. 16 is to reduce the complexity of generating a predicted image. In Embodiment (16), each image of the reference view used to generate a predicted image for the image of the current view is divided into blocks. Next, the depth layer of the image of the reference view is subsampled in order to retain only one depth value per block. As a result, all the pixels of the block use the same depth value for forward projection. A policy for selecting the depth value associated with the block is defined. It can include one of the following approaches. · The depth value of one specific pixel of the block represents the block, for example, the top - left block or one of the middle ones. · The average or central depth value of the block (having an average position) represents the block. · The more frequent depth value (having an associated position or central position) represents the block.

[0208] Embodiment (16) starts from step 140 where the processing module 20 acquires the camera parameters (i.e., view parameters) of the reference view and the current view.

[0209] In step 161, the processing module 20 initializes the variable n to "1".

[0210] In step 162, the processing module 20 checks the value of the variable N_sub to determine whether subsampling is applied to the image of the reference view. If N_sub = 1, subsampling is not applied to the depth layer of the reference view. In that case, step 141 follows step 162, and the processing module 20 generates a predicted image G(k) by applying the forward projection method described in FIGS. 13A and 13B between the reference view and the current view.

[0211] In step 143, the predicted image G(k) is stored in the DPB of the current view.

[0212] After step 143, step 142 follows, and the processing module 20 fills in the separated missing pixels.

[0213] In step 144, the processing module 20 reconstructs the current image of the current view using the reference image included in the DPB of the current view, and the DPB includes the predicted image G(k).

[0214] When N_sub > 1, sub-sampling is applied to the depth layer of the reference view. In one embodiment, when N_sub > 1, N_sub is a multiple of 2. If this reference view image has a width w and a height h, the image is, for example, of the same size

[0215]

Number

[0216] After step 162, step 163 follows, and the processing module 20 selects a depth value for the block number n of the reference view image using a policy defined for selection.

[0217] In step 164, the processing module 20 generates a predicted block Gblock(n,k) for which the forward projection method described in FIGS. 13A and 13B is applied between the reference view and the current view for the block number n of the reference view image.

[0218] In step 165, the processing module 20 stores the predicted block Gblock(n,k) in the DPB of the current view at the same position as the position of the block number n of the reference view image.

[0219] In step 166, the processing module 20 compares the value of the variable n with the number of blocks in the image of the reference view NB_Blocks. If n < Nb_blocks, after step 166, step 167 follows and the variable n is incremented by 1 unit. After step 167, step 162 follows and the forward projection is applied to the new block.

[0220] When n = Nb_blocks, steps 142 and 144 described above follow step 166. Note that at the end of the loop on NB_Blocks of the reference view image, the combination of blocks Gblock(n, k) forms the predicted image G(k).

[0221] In a modification of embodiment (16), similar to embodiments (14B) and (15), embodiment (16) can be applied to images of multiple reference views to obtain multiple predicted images.

[0222] In the embodiment of this modification, the multiple predicted images are stored in the DPB of the current view.

[0223] In the embodiment of this modification, the multiple predicted images are aggregated to form an aggregated predicted image, and the aggregated predicted image is stored in the DPB of the current view.

[0224] In the embodiment of this modification, at least one subset of the multiple predicted images is aggregated to form an aggregated predicted image, and each aggregated predicted image is stored in the DPB of the current view.

[0225] In the embodiment of this modification, at least one subset of the multiple predicted images is aggregated to form an aggregated predicted image, and each aggregated predicted image is stored in the DPB of the current view in addition to the multiple predicted images and the aggregated predicted image that aggregates all of the multiple predicted images.

[0226] In a modification of embodiment (16), the reference view image is divided into blocks of unequal sizes. For example, the image is divided into large blocks (128×128, 64×64, 32×32, 16×16, or 8×8) where the depth values are uniform (e.g., the difference between the minimum depth value and the maximum depth value does not exceed ±10% of the minimum depth value), and small blocks (4×4 or 2×2) where the depth values are non-uniform (e.g., the difference between the minimum depth value and the maximum depth value exceeds ±10% of the minimum depth value).

[0227] In one embodiment, referred to as the bidirectional embodiment, at least one predicted image of embodiments (14A), (14B), (15), and (16) is used to provide a reference block (i.e., the VSP predictor block) to the current block of the current image predicted using bi-prediction (i.e., inter bi-prediction). In that case, the current block is associated with two motion information, specifying two reference blocks in two different images, and the residual block of this block, i.e., the first reference block weighted by weight w0 = 1 / 2 and the second reference block weighted by weight w1 = 1 / 2, is the average of the two residual blocks. The sample S of this current block curr is obtained as follows.

[0228]

Number

[0229] In one embodiment, referred to as embodiment WP, at least one predicted image of embodiments (14A), (14B), (15), and (16) is used to provide a reference block (i.e., the VSP predictor block) to the current block of the current image predicted using weighted prediction (WP). In that case, the current block is associated with two motion information, specifying two reference blocks in two different images, and the residual block of this block, i.e., the first reference block weighted by weight w0 and the second reference block weighted by weight w1, is the weighted average of the two residual blocks. Again, the sample S of the current block curr is obtained as follows.

[0230]

Number

[0231] Note that the WP embodiment can be generalized to all modes by weighting the samples, for example using the triangular mode.

[0232] As described above, forward prediction can generate a predicted image including isolated missing pixels. So far, the hole filling process has been used to fill the isolated missing pixels. However, the hole filling process only provides an approximation of the actual pixels.

[0233] In a variant of the embodiment, in a variant of the WP embodiment called a bidirectional embodiment and an embodiment with modified weighting, the weighting process is modified to take into account a value representing the reliability of the samples (i.e., pixels) of the predicted image. In this variant, the sample S of the current block curr is obtained as follows.

[0234]

Equation

[0235] In the first variant of the embodiment with modified weighting, Mask0 (Mask1) is equal to zero if sample S0 (S1) is obtained by hole filling, and equal to "1" otherwise. When Mask0 = Mask1 = 0, the processing module 20 provides a default value to S curr .

[0236] In the second variant of the embodiment with modified weighting, Mask0 (Mask1) is set to a low positive value (e.g., "1") if sample S0 (S1) is obtained by hole filling, and a high positive value (e.g., "10000") otherwise. In other words, the value of Mask0 (Mask1) when sample S0 (S1) is obtained by hole filling is lower than that of Mask0 (Mask1) when sample S0 (S1) is obtained by forward projection.

[0237] In a third variant of the embodiment with modified weighting, the reliability of a sample depends on the similarity between the sample and samples in its vicinity. For example, a sample S0(S1) similar to the samples in the vicinity is associated with a value of Mask0(Mask1) that is higher than the value of Mask0(Mask1) associated with a sample S0(S1) that is different from the samples in the vicinity. The difference between two samples is calculated, for example, as the square root of the difference between the values of the two samples.

[0238] In a fourth variant of the embodiment with modified weighting, the reliability of a sample depends on the similarity between the sample and samples in its vicinity, and on the process applied to obtain the sample (hole filling or direct forward projection).

[0239] In a fifth variant of the embodiment with modified weighting, the reliability of samples in the predicted image is calculated at the block level (typically size 4×4) instead of the pixel level. The value Mask0(Mask1) associated with the samples of a block depends on the average reliability of the samples of the block.

[0240] In a sixth variant of the embodiment with modified weighting, the reliability of samples in the predicted image depends on the agreement of its depth value with other depth maps. The forward projection of the sample position Pn at depth Dn in view n to view m corresponds to the sample position Pm at depth Dm. The depth Dn at the sample position Pn is considered to be the matching depth if the forward projection of the sample position Pm at depth Dm to view n reaches the sample position Pn. Otherwise, the depth Dn at the sample position is not considered to be in agreement with view m. The same process is applied to other views, and then a score for the depth Dn at the same position Pn can be established between non - match and perfect match. The reliability of the sample is proportional to the agreement of its depth.

[0241] In some cases, a block predicted using bidirectional inter prediction mode or weighted prediction can use one reference block from the predicted picture and one reference block from a picture not obtained by forward prediction. In the third, fourth, fifth, and sixth variations of the embodiment with modified weighting, samples of the picture not obtained by forward prediction are considered samples with the highest possible reliability. For example, when the possible values of Mask0 (Mask1) are · In the case of "0" and "1", samples of the picture not obtained by forward prediction are associated with a value Mask0 (Mask1) equal to "1", · In the case of "1" and "10000", samples of the picture not obtained by forward prediction are associated with a value Mask0 (Mask1) equal to "10000", · In the case between "0" and "1", samples of the picture not obtained by forward prediction are associated with a value Mask0 (Mask1) equal to "1", · In the case between "1" and "10000", samples of the picture not obtained by forward prediction are associated with a value Mask0 (Mask1) equal to "10000".

[0242] To reduce the burden on the decoder side and determine the upper limit of the maximum complexity of a compliant decoder using view - to - view prediction, the relationship between the current view and the view used to generate the predicted picture is signaled in the encoded video stream corresponding to the encoded MVD data (e.g., encoded video streams 511 and / or 511B). From this signaling, the decoder can advantageously pre - calculate the predicted picture. The advantage of such an approach is that it enables the use of a legacy decoder with very few changes since only the reference picture buffer filling (i.e., DPB filling) is modified.

[0243] In the following, a syntax element called view_parameter is proposed that enables the reconstruction of a predicted image or an aggregated predicted image and represents information adapted to the above-described embodiments. In one embodiment, the syntax element view_parameter is inserted into the encoded video stream at the slice header level. In another embodiment, the syntax element view_parameter is inserted into the sequence header (i.e., the sequence parameter set (SPS)), the picture header (i.e., the picture parameter set (PPS)) shared by one or more pictures, or at the synchronization point or picture level into the encoded video stream (e.g., the header of an IDR (Instantaneous Decoding Refresh) picture). Each time the syntax element is received, the decoder can update its knowledge of the relationship between views.

[0244]

Table 7

[0245] Table TAB2 represents the first version of the syntax element view_parameter that is adapted to embodiments (14A), (14B), (16) in which only one predicted image or only one aggregated predicted image and only one predicted image is inserted into the DPB of the current view (typically, embodiments in which only one predicted image is generated or only one aggregated predicted image is generated).

[0246] The first version of the syntax element view_parameter includes the parameter view_id that represents the unique identifier of the current view. If the current view is not the first view decoded for the frame, the flag vsp_flag indicates whether the vsp mode is used for the current view. The parameter number_of_inter_view_predictor_used represents the maximum number of (already decoded) views used to decode the current view. The parameter predictor_id[view_id][i] provides the identifier of each view used to create the predicted picture of the current view. In one embodiment, the maximum number of views used to decode the current view is fixed at "8". In that case, "3" bits are required to encode the parameter predictor_id[view_id].

[0247] Naturally, the inter-view prediction between the first view and the second view is possible only if the camera parameters of the two views are available on the decoder side, that is, if the SEI message described in Table TAB1 is received and decoded by the decoder.

[0248]

Table 8

[0249] Table TAB3 represents the second version of the syntax element view_parameter that is adapted to embodiments where multiple predicted pictures and / or aggregated predicted pictures are inserted into the DPB of the current view (typically, embodiments (15) and (16) where multiple predicted pictures or aggregated predicted pictures are generated).

[0250] In that case, the current view can be associated with a plurality of reference views. In the second version of the syntax element view_parameter, the parameter number_inter_view_predictor_minus1 specifies the number of predicted images or aggregated predicted images used for the inter-view prediction of the current view specified by the parameter view_id. The parameter number_inter_view_predictor_used_minus1 specifies, for each predicted image or aggregated predicted image, the number of reference views used to generate the predicted image or aggregated predicted image. In the case of a predicted image, the parameter number_inter_view_predictor_used_minus1 is set to 1. The parameter predictor_id specifies which view is used to generate the predicted image or aggregated predicted image.

[0251] As can be seen from Tables TAB2 and TAB3, the VSP mode can be activated at the slice header level of the syntax element view_parameter by the flag vsp_flag.

[0252] Signaling at the slice level makes it possible to tell the decoder whether the blocks included in this slice can potentially use the VSP mode. However, it does not specify which blocks within the slice actually use the VSP mode.

[0253] In one embodiment, when activated at the slice level, the actual use of the VSP mode is signaled at the block level.

[0254] FIG. 17 schematically shows a basic embodiment of the syntax analysis process of a video compression method that does not use the VSP mode.

[0255] The basic embodiment of FIG. 17 is based on the syntax of the blocks (also referred to as prediction units (PUs)) described in Table TAB4. This basic embodiment is executed by the decoder when decoding the current block. However, the encoder encodes a syntax compliant with what the decoder can decode.

[0256] In step 1700, processing module 20 determines whether the current block is encoded in skip mode. If so, processing module 20 decodes an identifier merge_idx for the current block. The identifier merge_idx identifies which candidate block in the neighborhood of the current block provides information for decoding the current block. After decoding the identifier merge_idx, the decoding of the current block continues by applying a decoding process adapted to the skip mode.

[0257] If the current block is not encoded in skip mode, in step 1701, processing module 20 determines whether the current block is encoded in intra mode. If so, in step 1702, the current block is decoded using the intra mode decoding process.

[0258] If the current block is not encoded in intra mode, in step 1703, the processing module determines whether the current block is encoded in merge mode. If the current block is encoded in merge mode, the processing module decodes an identifier merge_idx for the current block in step 1704. After decoding the identifier merge_idx, the decoding of the current block continues by applying a decoding process adapted to the merge mode.

[0259] If the current block is not encoded in merge mode, in step 1705, processing module 20 determines whether the current block is encoded in bi - directional or uni - directional inter - prediction mode.

[0260] If the current block is encoded in the uni-directional inter prediction mode, step 1712 follows step 1705, and the processing module 20 decodes one index (ref_idx_l0 or ref_idx_l1) within the list of reference images stored in the DPB. This index indicates which reference image provides the predictor block for the current block.

[0261] In step 1713, the processing module 20 decodes the motion vector refinement mvd for the current block.

[0262] In step 1714, the processing module 20 decodes the motion vector predictor index that specifies the motion vector predictor. Using this motion information, the processing module 20 decodes the current block.

[0263] When the current block is encoded in the bi-directional prediction mode, step 1706 follows step 1705, during which the processing module 20 decodes the first index (ref_idx_l0) within the list of reference images stored in the DPB.

[0264] In step 1707, the processing module 20 decodes the first motion vector refinement mvd for the current block.

[0265] In step 1708, the processing module 20 decodes the first motion vector predictor index that specifies the first motion vector predictor.

[0266] In step 1709, the processing module 20 decodes the second index (ref_idx_l1) within the list of reference images stored in the DPB.

[0267] In step 1710, the processing module 20 decodes the second motion vector refinement mvd for the current block.

[0268] In step 1711, processing module 20 decodes a second motion vector predictor index that specifies a second motion vector predictor.

[0269] Using this motion information, processing module 20 generates two predictors and decodes the current block using these two predictors.

[0270] [Table 9]

[0271] FIG. 18 schematically shows a first embodiment of a syntax analysis process of a video compression method using a new VSP mode.

[0272] The embodiment of FIG. 18, hereinafter referred to as embodiment (18), is based on the syntax of the blocks described in Table TAB5. The differences between the syntaxes of Table TAB4 and Table TAB5 are represented in bold. This embodiment is executed by the decoder when decoding the current block. However, the encoder encodes a syntax compliant with what the decoder can decode.

[0273] [Table 10]

[0274] As described below, in Embodiment (18), the use of the VSP mode is signaled at the block level by the flag VSP. When the flag VSP = 1, the VSP mode is activated for the current block. Otherwise, the VSP mode is deactivated. Further, in Embodiment (18), when a block is encoded in the VSP mode, the predictor block is placed in the same location as the current block. As a result, no motion vector is required to obtain the block predictor from the reference image (the predicted image or the aggregated predicted image in that case). Further, as will be described later in connection with FIG. 18, when two predictor blocks are extracted from the same predicted image or aggregated predicted image, the combination of the VSP mode and the bidirectional inter prediction is not possible. In fact, in the VSP mode, each predictor is placed in the same position as the current block, so that in the case of bidirectional inter prediction, the two predictor blocks are the same.

[0275] The syntax and parsing method of Embodiment (18) are adapted to Embodiments (14A), (14B), and (16) when only one predicted image or aggregated predicted image is inserted into the DPB of the current layer.

[0276] In step 1800, the processing module 20 determines whether the current block is encoded in the skip mode. If so, in step 1804, the processing module 20 decodes the identifier merge_idx for the current block. After decoding the identifier merge_idx, the current block is decoded by applying the decoding process adapted to the skip mode.

[0277] If the current block is not encoded in the skip mode, in step 1801, the processing module 20 determines whether the current block is encoded in the intra mode. If so, in step 1802, the current block is decoded using the intra mode decoding process.

[0278] If the current block is not encoded in intra mode, in step 1803, the processing module determines whether the current block is encoded in merge mode. If the current block is encoded in merge mode, in step 1806, the processing module decodes the identifier merge_idx for the current block. After decoding the identifier merge_idx, the decoding of the current block continues by applying a decoding process adapted to the merge mode.

[0279] If the current block is not encoded in merge mode, in step 1807, the processing module 20 determines whether the current block is encoded in bidirectional or unidirectional inter prediction mode.

[0280] If the current block is encoded in unidirectional inter prediction mode, step 1808 follows step 1807, and the processing module 20 decodes the flag VSP to determine whether the current block is encoded in VSP mode. If the current block is encoded in VSP mode, the processing module 20 decodes the current block according to the VSP mode decoding process. In other words, the current block is predicted from the block of the predicted image (or aggregated predicted image) stored in the DPB of the current view. In that case, the position of the predicted image (or aggregated predicted image) in the DPB is implicit and known to the decoder (i.e., the predicted image (or aggregated predicted image) is systematically at the same position in the DPB).

[0281] If the current block is not encoded in VSP mode, step 1810 follows step 1808, and the processing module 20 decodes one index (ref_idx_l0 or ref_idx_l1) from the list of reference images stored in the DPB.

[0282] In step 1811, the processing module 20 decodes the motion vector refinement mvd for the current block.

[0283] In step 1812, the processing module 20 decodes a motion vector predictor index that specifies a motion vector predictor. Using this motion information, the processing module 20 decodes the current block.

[0284] When the current block is encoded in the bidirectional prediction mode, step 1813 follows step 1807, and the processing module 20 decodes the flag VSP to determine whether the first prediction block of the current block is obtained from the prediction image (or from the aggregated prediction image). If the first prediction block of the current block is obtained from the prediction image (or from the aggregated prediction image), at the same step 1814 as step 1809, the first predictor block is obtained. Step 1819 follows step 1814, and the processing module 20 decodes an index (ref_idx_l1) within the list of reference images stored in the DPB.

[0285] In step 1820, the processing module 20 decodes the motion vector refinement mvd for the current block.

[0286] In step 1821, the processing module 20 decodes a motion vector predictor index that specifies a motion vector predictor. Using the motion information obtained in steps 1819, 1820, and 1821, the processing module 20 determines the second predictor block. Using these two predictors, the processing module 20 determines the bidirectional predictor a for decoding the current block.

[0287] If the first prediction block of the current block is not obtained from the prediction image (or from the aggregated prediction image), the processing module executes steps 1815, 1816, and 1817, which are the same as steps 1810, 1811, and 1812 respectively, to obtain the first predictor.

[0288] In step 1818, the processing module 20 decodes the flag VSP to determine whether the second predicted block of the current block is obtained from the predicted image (or from the aggregated predicted image). If the second predicted block of the current block is obtained from the predicted image (or from the aggregated predicted image), in step 1822 which is the same as step 1809, the second predictor block is obtained. Using the first and second predictor blocks, the processing module decodes the current block.

[0289] If the second predicted block of the current block is not obtained from the predicted image (or from the aggregated predicted image), in step 1819, the processing module 20 decodes the second index (ref_idx_l1) in the list of reference images stored in the DPB.

[0290] In step 1820, the processing module 20 decodes the second motion vector refinement mvd for the current block.

[0291] In step 1821, the processing module 20 decodes the second motion vector predictor index that specifies the second motion vector predictor.

[0292] Based on the motion information obtained in steps 1815, 1816, 1817, 1819, 1820, and 1821, the processing module 20 decodes the current block.

[0293] In a modification of embodiment (18), in step 1804, the processing module 20 decodes the VSP flag of the current block. If the VSP mode is activated for the current block, the processing module 20 executes step 1805 which is the same as step 1809. If the VSP mode is not activated for the current block, the processing module 20 executes step 1806.

[0294] FIG. 19 schematically shows a second embodiment of the syntax analysis process of a video compression method using a new VSP mode.

[0295] The following Embodiment (19) of FIG. 19, which is referred to as "Embodiment (19)", is based on the syntax of the blocks described in Table TAB6. The differences between the syntaxes of Table TAB4 and Table TAB6 are represented in bold. This embodiment is executed by the decoder when decoding the current block. However, the encoder encodes a syntax compliant with what the decoder can decode.

[0296] [Table 11]

[0297] Embodiment (19) is very similar to Embodiment (18). Embodiment (19) differs from Embodiment (18) in that the syntax of the blocks encoded in VSP mode includes a syntax element representing the motion vector difference mvd. The result of this feature is that the combination of VSP mode and bi - directional inter - prediction is now possible when two predictor blocks are extracted from the same predicted image or aggregated predicted image. In fact, in Embodiment (19), the presence of the motion vector difference mvd makes it possible to obtain two different predictor blocks.

[0298] The syntax and parsing method of Embodiment (19) are adapted to Embodiments (14A), (14B), and (16) when only one predicted image or aggregated predicted image is inserted into the DPB of the current layer.

[0299] Embodiment (19) includes steps 1900 - 1908, 1910 - 1913, 1915 - 1921 which are the same as steps 1800 - 1808, 1810 - 1813, 1815 - 1821 respectively.

[0300] When the VSP mode is activated for the current block, step 1909 follows step 1908, and the motion vector difference mvd is calculated for the current block. This motion vector difference mvd makes it possible to specify a predictor block in the predicted image or the aggregated predicted image. Then, the predictor is used to decode the current block.

[0301] If the VSP flag specifies that the first predictor of the current block is generated from the predicted image or the aggregated predicted image, step 1914 follows step 1913, and the motion vector difference mvd is calculated for the current block. This motion vector difference mvd makes it possible to specify the first predictor block in the predicted image or the aggregated predicted image.

[0302] Step 1918 follows step 1914. If the VSP flag specifies that the second predictor of the current block is generated from the predicted image or the aggregated predicted image, step 1922 follows step 1918, and the motion vector difference mvd is calculated for the current block. This motion vector difference mvd makes it possible to specify the second predictor block in the predicted image or the aggregated predicted image. The current block is decoded from the first and second predictors as a bi - directional prediction mode.

[0303] Note that step 1905 is the same as step 1909.

[0304] FIG. 20 schematically shows a third embodiment of the syntax analysis process of a video compression method using the new VSP mode.

[0305] The embodiment of FIG. 20, hereinafter referred to as "Embodiment (20)", is based on the syntax of the blocks described in Table TAB7. The differences between the syntaxes of Table TAB4 and Table TAB7 are shown in bold. This embodiment is executed by the decoder when decoding the current block. However, the encoder encodes a syntax compliant with what the decoder can decode.

[0306] Embodiment (20) is very similar to embodiment (19). Embodiment (20) differs from embodiment (19) in that the syntax of the block encoded in VSP mode does not include a syntax element representing the motion vector difference mvd, but includes at least one index (ref_idx2_l0 or ref_idx2_l1) in the list of reference images stored in the DPB. This index indicates which predicted image provides the predictor block for the current block. As a result of this feature, a combination of VSP mode and bi-directional inter prediction becomes possible. In fact, in embodiment (20), the presence of two indexes in the case of bi-directional inter prediction that specifies two different predicted images (or aggregated predicted images) enables the acquisition of two different predictor blocks.

[0307] The syntax and analysis method of embodiment (20) are adapted to embodiments (15) and (16) when a plurality of predicted images or aggregated predicted images are inserted into the DPB of the current layer.

[0308] Embodiment (20) includes steps 2000 to 2008, 2010 to 2013, 2015 to 2021, which are respectively the same as steps 1900 to 1908, 1910 to 1913, 1915 to 1921.

[0309] Step 1909 of Embodiment (19) is replaced by step 2009 of Embodiment (20). In step 2009, processing module (20) decodes a syntax element ref_idx2_l0 (or ref_idx2_l1) representing an index within a list l0 (or l1) of reference images in a predicted image or an aggregated predicted image that temporally corresponds (i.e., within the same frame) to the image including the current block. The processing module 20 extracts a predictor block that is spatially located at the same position as the current block from the predicted image (or aggregated predicted image) specified by the index ref_idx2_l0 (or ref_idx2_l1). Next, the processing module decodes the current block using the obtained predictor block.

[0310] Step 1914 of Embodiment (19) is replaced by step 2014 of Embodiment (20). In step 2014, processing module (20) decodes a syntax element ref_idx2_l0 representing an index within a first list l0 of reference images to be used among the predicted images or the aggregated predicted images that temporally correspond (i.e., within the same frame) to the image including the current block. The processing module 20 extracts a first predictor block that is spatially located at the same position as the current block from the predicted image (or aggregated predicted image) specified by the index ref_idx2_l0.

[0311] Step 1922 of Embodiment (19) is replaced by step 2022 of Embodiment (20). In step 2022, processing module (20) decodes a syntax element ref_idx2_l1 representing an index within a second list l1 of reference images to be used among the predicted images or the aggregated predicted images that temporally correspond (i.e., within the same frame) to the image including the current block. The processing module 20 extracts a second predictor block that is spatially located at the same position as the current block from the predicted image (or aggregated predicted image) specified by the index ref_idx2_l1.

[0312] After step 2021 or 2022, processing module 20 decodes the current block in dual prediction inter mode using the first and second predictors.

[0313] Note that step 2005 is the same as step 2009.

[0314] [Table 12]

[0315] FIG. 21 schematically shows a fourth embodiment of a syntax analysis process of a video compression method using a new VSP mode.

[0316] Hereinafter, the embodiment of FIG. 21, referred to as embodiment (21), is based on the syntax of the blocks described in Table TAB8. The differences between the syntaxes of Table TAB4 and Table TAB8 are shown in bold. This embodiment is executed by the decoder when decoding the current block. However, the encoder encodes a syntax compliant with what the decoder can decode.

[0317] Embodiment (21) is very similar to embodiment (18). However, in embodiment (21), the use of the VSP mode at the block level is inferred from the index of the reference picture (ref_idx_l0 or ref_idx_l1) instead of being explicitly specified by the flag VSP.

[0318] The syntax and analysis method of embodiment (21) are adapted to embodiments (14A), (14B), and (16) when only one predicted picture or aggregated picture is inserted into the DPB of the current layer.

[0319] Embodiment (21) includes steps 1800 to 1803, 1805 to 1807, 1809 to 1812, 1814 to 1817, 1819 to 1822 and steps 2100 to 2103, 2105 to 2107, 2109 to 2112, 2114 to 2117, 2119 to 2122 which are the same as those respectively.

[0320] In step 2104, when the index ref_idx2_l0 or the index ref_idx2_l1 that designates a reference image in the list of reference images corresponding to the predicted image or the aggregated predicted image is inherited from the candidate block designated by the identifier merge_idx, the processing module 20 considers that the VSP mode is activated for the current block.

[0321] In step 2108, when the index ref_idx_l0 that designates a reference image in the list l0 designates the predicted image or the aggregated predicted image, the VSP mode considers that it is activated for the current block. For example, ref_idx_l0 = 0 indicates a reference image corresponding to the predicted image or the aggregated predicted image.

[0322] In step 2113, when the index ref_idx_l0 that designates a reference image in the list l0 of reference images designates the predicted image or the aggregated predicted image, the processing module 20 considers that the first predictor of the current block is obtained from the predicted image or the aggregated predicted image.

[0323] In step 2118, when the index ref_idx_l1 that designates a reference image in the list l1 of reference images designates the predicted image or the aggregated predicted image, the processing module 20 considers that the second predictor of the current block is obtained from the predicted image or the aggregated predicted image. For example, ref_idx_l1 = 0 designates a reference image corresponding to the predicted image or the aggregated predicted image.

[0324] [Table 13]

[0325] In the syntax of Table TAB8, the motion vector difference mvd and the motion vector predictor index mvp are decoded only when the function is_vsp_generated returns false. This function is_vsp_generated is defined as follows.

[0326] is_vsp_generated(idx) returns · true if the reference index idx points to a frame generated from a view within the same frame, and · false otherwise.

[0327] In the following, in a variant of Embodiment (18) called Embodiment (18bis), if the current block is encoded in merge mode or skip mode, the flag VSP is not encoded for the current block. In that case, the processing module 20 first decodes the identifier merge_idx and determines whether the candidate block specified by the identifier merge_idx is encoded in VSP mode. If the candidate block is encoded in VSP mode, the current block inherits the VSP parameters from the candidate block and the current block is decoded using these parameters. Otherwise, the current block is decoded by applying the normal merge mode decoding process. This Embodiment (18bis) is based on the syntax of the blocks described in Table TAB9.

[0328] Additional embodiments can be obtained by combining Embodiments (18), (18bis), (19), (20), and (21).

[0329] For example, the syntax of the current block encoded in VSP mode decodes the motion vector difference mvd and the syntax elements ref_idx2_l0 and / or ref_idx2_l1 that represent the index in the first list l0 and / or the second list l1 of the reference image to be used among the predicted image or the aggregated predicted image that temporally corresponds to the image including the current block (i.e., within the same frame). This corresponds to the combination of embodiments (19) and (20).

[0330] In another example, the syntax of the current block encoded in VSP mode can include the motion vector difference mvd, and the use of the VSP mode can be inferred from the syntax elements ref_idx_l0 and / or ref_idx_l1 instead of being indicated by the flag display VSP. This corresponds to the combination of embodiments (19) and (21).

[0331] In another example, the syntax of the current block encoded in VSP mode can include the syntax elements ref_idx2_l0 and / or ref_idx2_l1, and the use of the VSP mode can be inferred from the syntax elements ref_idx_l0 and / or ref_idx_l1 instead of being indicated by the flag display VSP. This corresponds to the combination of embodiments (20) and (21).

[0332] In other examples, · Embodiment (18) can be combined with embodiments (19), (20), and (21). · Embodiments (19), (20), and (21) can be combined. · Embodiments (19), (20), (21), and (22) can be combined. · And so on.

[0333]

Table 14

[0334] Up to now, the projected image (or aggregated projected image) used for view - to - view prediction has been considered to include per - pixel texture data and depth data. In another embodiment called the MI (motion information) - based VSP embodiment, the predicted image (and aggregated predicted image) is replaced by an image called an MI (motion information) predicted image (or MI aggregated predicted image) that includes only the motion information of each pixel or a subset of pixels.

[0335] In the MI - based VSP embodiment, the forward projection process of FIG. 13A includes an additional step 133. In step 133, the processing module 20 calculates a motion vector MV that represents the displacement between the pixel P’(u’, v’) obtained by the forward projection in steps 130 - 132 and the projected pixel P(u, v). This motion vector MV is intended to be stored in the MI predicted image.

[0336] Hereinafter, the impact of the MI - based VSP embodiment on embodiments (14A), (14B), (15), and (16) will be described.

[0337] In the MI - based VSP embodiment, embodiment (14A) is modified to become embodiment (14A_MI). Embodiment (14A_MI) is represented in FIG. 22A.

[0338] The first step of embodiment (14A_MI) is step 140 which has already been described in relation to embodiment (14A).

[0339] In step 141_MI, the processing module 20 applies the forward projection of steps 130 - 133 between the reference view (e.g., the first view 501) and the current view (e.g., the second view 501B) to generate an MI predicted image MI(k). The MI predicted image MI(k) is introduced into the DPB (e.g., DPB519B(619B)) of the current view and is intended to be one of the Kth predicted images for the current image of the current view. th of the predicted images.

[0340] In step 142_MI, the processing module 20 fills the separated missing motion information. In one embodiment, the separated missing motion information is filled with adjacent pixel motion information. In another embodiment, the separated missing motion information is filled with a default value (typically, motion vector = (0,0)). In another embodiment, the separated missing motion information is considered invalid, and a flag indicating the validity of the motion information is associated with each motion information.

[0341] In step 143_MI, the processing module 20 stores the MI prediction image MI(k) in the DPB of the current view.

[0342] In step 144_MI, the processing module 20 reconstructs the current image of the current view using the reference images included in the DPB of the current view, and the DPB includes the MI prediction image MI(k). The MI prediction image MI(k) is used by the processing module 20 to generate the prediction image G(k). In fact, the motion information included in the MI prediction image MI(k) is used to apply motion compensation to each pixel specified by this motion information in the reference image.

[0343] Optionally, embodiment (14A_MI) includes step 220 which consists of reducing the amount of motion information in the MI prediction image MI(k). In fact, having motion information for each pixel position of the image represents a large amount of data. In one embodiment, the MI prediction image MI(k) is divided into blocks of size N×M, where N and M are multiples of 2 and are smaller than the width and height of the MI prediction image MI(k). For each N×M block, only one motion information is retained. In other words, the motion information is subsampled by the factor N×M. In one embodiment, N = M = 4, and one motion information is excluded from "16" motion informations.

[0344] In one embodiment, the subsampling consists of selecting one specific motion information within the N×M motion information for each block.

[0345] In one embodiment, sub-sampling consists of selecting the median of the N×M motion information for each block (the median is calculated using the norm of the motion vector).

[0346] In one embodiment, sub-sampling consists of selecting the motion information that most frequently appears within the N×M motion information.

[0347] In one embodiment, sub-sampling consists of selecting the motion information corresponding to the minimum depth of the view of the entire sub-block (z-buffer algorithm).

[0348] In one embodiment, sub-sampling consists of retaining the first projection value among the N×M motion information.

[0349] In the MI-based VSP embodiment, embodiment (14B) is modified to become embodiment (14B_MI). Embodiment (14B_MI) is shown in FIG. 22B.

[0350] Compared with embodiment (14B), in embodiment (14B_MI), step 1412 is replaced by step 1412_MI, and step 1415 is replaced by step 1415_MI.

[0351] In step 1412, processing module 20 applies the forward projection method of steps 130 to 133 between view j and the current view to generate the MI prediction image MI(k) j thereby.

[0352] In step 1415_MI, processing module 20 calculates the aggregated MI prediction image MI(k) from the prediction image MI(k) j and this MI prediction image MI(k) is intended to be stored in the DPB of the current view.

[0353] In one embodiment of the aggregation process, the prediction image MI(k) j is the first MI prediction image MI(k) among a plurality of MI prediction images jis aggregated by retaining the motion information of. The first MI prediction image is, for example, the prediction image MI(k) j = 0.

[0354] In one embodiment of the aggregation process, the MI prediction image MI(k) j is aggregated by retaining the motion information of the MI(k) generated from the view j closest to the current view j If several views are at the same distance from the current view (i.e., there are several closest views), one of the several closest views is randomly selected. For example, in the camera array 10, assume that only the first view generated by the camera 10A and the second view generated by the camera 10C are available for predicting the current view generated by the camera 10B. Then, the first and second views are the views closest to the current view and are at the same distance to the current view. One of the first and second views is selected to provide the motion information to the MI aggregation prediction image MI(k).

[0355] In one embodiment of the aggregation process, the MI prediction image MI(k) j is aggregated by retaining the motion information of the prediction image MI(k) having the best quality j For example, the information representing the quality of a pixel is the value of the quantization parameter applied to the transform block including the pixel.

[0356] In one embodiment of the aggregation process, the MI prediction image MI(k) j is aggregated by retaining the motion information of the prediction image MI(k) having the closest depth value (z-buffer algorithm) j In one embodiment of the aggregation process, the MI prediction image MI(k)

[0357] is aggregated by retaining the motion information of the MI(k) when the adjacent pixels of the aggregated prediction image MI(k) j are already predicted from the prediction image j the prediction image MI(k) jIt is aggregated by retaining the pixel values (texture values and depth values).

[0358] In one embodiment of the aggregation process, the MI prediction image MI(k) j is aggregated by calculating the average, weighted average, and median of the motion information of the MI prediction image MI(k) j .

[0359] Note that the motion information includes information representing motion vectors and information representing the indices of the reference images in the list of reference images (e.g., ref_idx_l0, ref_idx_l1, ref_idx2_l0, ref_idx2_l1).

[0360] As can be seen from the above, in embodiment (14B_MI), subsampling (step 220) is performed on the aggregated MI prediction image MI(k). In a variant of embodiment (14B_MI), the subsampling of step 220 is performed on each MI prediction image MI(k) j .

[0361] As can be seen from the above, in embodiment (14B_MI), subsampling (step 220) and the aggregation step 1415_MI are separate steps. In a variant of embodiment (14B_MI), subsampling is performed during the aggregation step.

[0362] In the MI-based VSP embodiment, embodiment (15) is modified to become embodiment (15_MI). Embodiment (15_MI) is shown in FIG. 23.

[0363] Compared with embodiment (15), in embodiment (15_MI), step 1502 is replaced by step 1502, step 1503 is replaced by step 1503_MI, step 1504 is replaced by step 1504_MI, and step 144 is replaced by step 144_MI.

[0364] Step 1502_MI is the same as step 1412_MI.

[0365] This step 1503_MI is the same as step 142_MI, except that the hole filling process is applied to the MI prediction image MI(k) instead of the MI prediction image MI(k). j

[0366] In step 1504_MI, the prediction image MI(k) j is stored in the DPB of the current view.

[0367] Step 144_MI in the embodiment (15_MI) is the same as step 144_MI in the embodiment (14B_MI), except that the DPB of the current view includes the number Nb_views of the MI prediction image MI(k). j

[0368] In a modification of the embodiment (15_MI), the subsampling step 220 is introduced between steps 1503_MI and 1504_MI.

[0369] In a modification of the embodiment (15_MI), in addition to the prediction image MI(k) j the aggregated prediction image generated from the prediction image MI(k) and / or the aggregated prediction image generated from a subset of the prediction image MI(k) j is inserted into the DPB of the current view. j

[0370] In a modification of the embodiment (15_MI), instead of the prediction image MI(k) j the aggregated prediction image generated from the prediction image MI(k) and the aggregated prediction image generated from a subset of the prediction image MI(k) are inserted into the DPB of the current view, or only the aggregated prediction image generated from a subset of the prediction image MI(k) is inserted into the DPB of the current view. j j j

[0371] ​In the MI-based VSP embodiment, embodiment (16) is modified to become embodiment (16_MI). Embodiment (16_MI) is shown in FIG. 25.

[0372] Compared with embodiment (16), in embodiment (16_MI), step 164 is replaced by step 164_MI, step 165 is replaced by step 165_MI, step 141 is replaced by step 141_MI, step 143 is replaced by step 143_MI, step 142 is replaced by step 142_MI, and step 144 is replaced by step 144_MI.

[0373] Step 141_MI in embodiment (16_MI) is the same as step 141_MI in embodiment (14A_MI).

[0374] Step 143_MI in embodiment (16_MI) is the same as step 143_MI in embodiment (14A_MI).

[0375] Step 142_MI in embodiment (16_MI) is the same as step 142_MI in embodiment (14A_MI).

[0376] Step 144_MI in embodiment (16_MI) is the same as step 144_MI in embodiment (14A_MI).

[0377] In step 164, for the block number n of the image of the reference view, the processing module 20 generates a predicted block MIblock(n,k) of motion information by applying the forward projection method described in steps 130 to 133 between the reference view and the current view.

[0378] In step 165_MI, the processing module 20 stores the block MIblock(n,k) in the DPB of the current view at the same position as the position of the block number n of the image of the reference view.

[0379] In a modification of the embodiment (16_MI), similar to the embodiments (14B) and (15), the embodiment (16) can be applied to images of a plurality of reference views to obtain a plurality of MI prediction images.

[0380] In the embodiment of this modification, the plurality of MI prediction images are stored in the DPB of the current view.

[0381] In the embodiment of this modification, the plurality of MI prediction images are aggregated to form an MI aggregated prediction image, and the aggregated MI prediction image is stored in the DPB of the current view.

[0382] In the embodiment of this modification, at least one subset of the plurality of MI prediction images is aggregated to form an aggregated MI prediction image, and each aggregated MI prediction image is stored in the DPB of the current view.

[0383] In the embodiment of this modification, at least one subset of the plurality of MI prediction images is aggregated to form an aggregated MI prediction image, and each aggregated MI prediction image is stored in the DPB of the current view in addition to the MI prediction image and the aggregated MI prediction image that aggregates all of the plurality of prediction images.

[0384] Embodiments with bidirection, WP, and modified weighting are all applied in the same way to all MI-based VSP embodiments (i.e., embodiments (14A), (14B), (15), and (16)).

[0385] So far, the motion information is considered to include information representing a motion vector and information representing an index of a reference image in a list of reference images (e.g., ref_idx_l0, ref_idx_l1, ref_idx2_l0, ref_idx2_l1). In a modification of the MI-based VSP embodiment, when the MI prediction image MI(k) is divided into blocks of size N×M, the motion information associated with each N×M block includes parameters of a pseudo-motion model that enables determining pixels of the current block in the current view from pixels of the N×M block of the reference view instead of information representing a motion vector.

[0386] The syntax and analysis methods of embodiments (18) and (18bis) are adapted to embodiments (14A_MI), (14B_MI), and (16_MI) when only one MI prediction image or an aggregated MI prediction image is inserted into the current layer's DPB.

[0387] The syntax and analysis methods of embodiment (19) are adapted to embodiments (14A_MI), (14B_MI), and (16_MI) when only one MI prediction image or an aggregated MI prediction image is inserted into the current layer's DPB.

[0388] The syntax and analysis methods of embodiment (20) are adapted to embodiments (15_MI) and (16_MI) when multiple MI prediction images or aggregated MI prediction images are inserted into the current layer's DPB.

[0389] The syntax and analysis methods of embodiment (21) are adapted to embodiments (14A_MI), (14B_MI), and (16_MI) when only one MI prediction image or only an aggregated MI prediction image is inserted into the current layer's DPB.

[0390] Embodiments combining the features of embodiments (18), (18bis), (19), (20), and (21) are also applicable to MI-based VSP embodiments.

[0391] Some embodiments have been described above. The features of these embodiments can be provided alone or in any combination. Further, an embodiment can include one or more of the following features, devices, or aspects alone or in combination across various claim categories and types. · A bitstream or signal including one or more of the described syntax elements or variations thereof. · Creating and / or transmitting and / or receiving and / or decoding a bitstream or signal including one or more of the described syntax elements or variations thereof. · A television, set-top box, mobile phone, tablet, or other electronic device that performs MVD encoding or decoding according to any of the described embodiments. · A television, set-top box, mobile phone, tablet, or other electronic device that performs MVD encoding according to any of the described embodiments and displays the acquired image (e.g., using a monitor, screen, or other type of display). · A television, set-top box, mobile phone, tablet, or other electronic device that tunes a channel (e.g., using a tuner) to receive a signal including an encoded video stream and performs multi-view decoding according to any of the described embodiments. · A television, set-top box, mobile phone, tablet, or other electronic device that receives a wireless signal including an encoded video stream (e.g., using an antenna) and performs MVD decoding according to any of the described embodiments.

Claims

1. A method for decoding, comprising: obtaining a first camera parameter associated with at least one reference view of multi-view video content and a second camera parameter associated with a current view, each view including a texture layer and a depth layer; applying a forward projection method to pixels of the reference view to project the pixels from the camera coordinate system of the reference view defined by the first camera parameter to the camera coordinate system of the current view defined by the second camera parameter, thereby generating an intermediate predicted image, wherein each projected pixel of the intermediate predicted image is associated with a motion information value representing a displacement between the projected pixel and the pixel of the reference view from which the pixel is projected; storing at least one final predicted image obtained from at least one intermediate predicted image in a decoded picture buffer of a reconstructed image used for temporal prediction of the current view; reconstructing a current image of the current view from the image stored in the decoded picture buffer.

2. The forward projection method comprises: obtaining back-projected pixels by applying back-projection from the camera coordinate system of the reference view to a world coordinate system to current pixels of the reference view, wherein the back-projection uses a pose matrix of a reference camera for obtaining the reference view, an inverse intrinsic matrix of the reference camera, and a depth value associated with the current pixels, and the pose matrix and the inverse intrinsic matrix of the reference camera are obtained from the first camera parameter; obtaining forward-projected pixels by projecting the back-projected pixels to the camera coordinate system of the current view using an intrinsic matrix and an extrinsic matrix of a current camera for obtaining the current view, wherein the intrinsic matrix and the extrinsic matrix of the current camera are obtained from the second camera parameter; selecting a pixel of the pixel grid closest to the forward-projected pixel to obtain a corrected forward-projected pixel when the forward-projected pixel does not correspond to a pixel on the pixel grid of the current camera; calculating a motion vector representing a displacement between the forward-projected pixel and the current pixel of the reference view.

3. The method according to claim 1, wherein the method includes filling separated missing pixel motion information values in each intermediate prediction image or the final prediction image with adjacent pixel motion information values or default values.

4. The method according to claim 1, wherein the at least one final prediction image is obtained by aggregating at least two intermediate prediction images.

5. A method for encoding, obtaining a first camera parameter associated with at least one reference view of multi-view video content and a second camera parameter associated with a current view, each view including a texture layer and a depth layer, applying a forward projection method to pixels of the reference view to project the pixels from the camera coordinate system of the reference view defined by the first camera parameter to the camera coordinate system of the current view defined by the second camera parameter, thereby generating an intermediate prediction image, wherein each projected pixel of the intermediate prediction image is associated with a motion information value representing a displacement between the projected pixel and the pixel of the reference view from which the pixel is projected, storing at least one final prediction image obtained from at least one intermediate prediction image in a decoded picture buffer of a reconstructed image used for temporal prediction of the current view, reconstructing a current image of the current view from the image stored in the decoded picture buffer.

6. The forward projection method is obtaining a back-projected pixel by applying an inverse projection from the camera coordinate system of the reference view to a world coordinate system to a current pixel of the reference view, wherein the inverse projection uses a pose matrix of a reference camera for obtaining the reference view, an inverse internal matrix of the reference camera, and a depth value associated with the current pixel, and the pose matrix and the inverse internal matrix of the reference camera are obtained from the first camera parameter, obtaining a forward-projected pixel by projecting the back-projected pixel to the camera coordinate system of the current view using an internal matrix and an external matrix of a current camera for obtaining the current view, wherein the internal matrix and the external matrix of the current camera are obtained from the second camera parameter. If the forward-projected pixel does not correspond to a pixel on the pixel grid of the current camera, select the pixel of the pixel grid closest to the forward-projected pixel to obtain a corrected forward-projected pixel; calculating a motion vector representing a displacement between the forward-projected pixel and the current pixel of the reference view, the method according to claim 5, comprising: **Claim 7** The method according to claim 5, comprising filling separated missing pixel values in each intermediate prediction image or the final prediction image with an average of adjacent pixel values, a median of adjacent pixel values, or a default value. **Claim 8** The method according to claim 5, wherein at least one final prediction image is obtained by aggregating at least two intermediate prediction images. **Claim 9** A device for decoding, means for obtaining a first camera parameter associated with at least one reference view of multi-view video content and a second camera parameter associated with a current view, each view including a texture layer and a depth layer; means for generating an intermediate prediction image configured to project a pixel from a camera coordinate system of the reference view defined by the first camera parameter to a camera coordinate system of the current view defined by the second camera parameter by applying a forward projection method to a pixel of the reference view, each projected pixel of the intermediate prediction image being associated with a motion information value representing a displacement between the projected pixel and the pixel of the reference view from which the pixel is projected; means for storing at least one final prediction image obtained from at least one intermediate prediction image in a decoded picture buffer of a reconstructed image used for temporal prediction of the current view; A device comprising means for reconstructing a current image of the current view from an image stored in the decoded picture buffer. **Claim 10** The forward projection method is Obtaining inverse projection pixels by applying the inverse projection from the camera coordinate system of the reference view to the world coordinate system to the current pixels of the reference view, wherein the inverse projection uses the pose matrix of the reference camera for obtaining the reference view, the inverse internal matrix of the reference camera, and the depth value associated with the current pixels, and the pose matrix and the inverse internal matrix of the reference camera are obtained from the first camera parameters, Obtaining forward projection pixels by projecting the inverse projection pixels into the camera coordinate system of the current view using the internal matrix and the external matrix of the current camera for obtaining the current view, wherein the internal matrix and the external matrix of the current camera are obtained from the second camera parameters, When the forward projection pixels do not correspond to pixels on the pixel grid of the current camera, selecting the pixel of the pixel grid closest to the forward projection pixels to obtain corrected forward projection pixels, Calculating a motion vector representing the displacement between the forward projection pixels and the current pixels of the reference view. The device according to claim 9, comprising: **Claim 11** The device according to claim 9, wherein the device is configured to fill the separated missing pixel values in each intermediate prediction image or the final prediction image with the average of adjacent pixel values, the median of adjacent pixel values, or a default value. **Claim 12** The device according to claim 9, wherein the at least one final prediction image is obtained by aggregating at least two intermediate prediction images. **Claim 13** A device for encoding, Means for obtaining a first camera parameter associated with at least one reference view of multi-view video content and a second camera parameter associated with a current view, wherein each view includes a texture layer and a depth layer. Means for generating an intermediate prediction image configured to project the pixel from the camera coordinate system of the reference view defined by the first camera parameter to the camera coordinate system of the current view defined by the second camera parameter by applying an orthographic projection method to the pixels of the reference view, wherein each projected pixel of the intermediate prediction image is associated with a motion information value representing a displacement between the projected pixel and the pixel of the reference view from which the pixel is projected. Means for storing at least one final prediction image obtained from at least one intermediate prediction image in a decoded picture buffer of a reconstructed image used for temporal prediction of the current view. Means for reconstructing the current image of the current view from the image stored in the decoded picture buffer. A device comprising the same.

14. The orthographic projection method is Obtaining a back-projected pixel by applying a back-projection from the camera coordinate system of the reference view to the world coordinate system to the current pixel of the reference view, wherein the back-projection uses a pose matrix of a reference camera for obtaining the reference view, an inverse intrinsic matrix of the reference camera, and a depth value associated with the current pixel, and the pose matrix and the inverse intrinsic matrix of the reference camera are obtained from the first camera parameter. Obtaining a forward-projected pixel by projecting the back-projected pixel onto the camera coordinate system of the current view using an intrinsic matrix and an extrinsic matrix of a current camera for obtaining the current view, wherein the intrinsic matrix and the extrinsic matrix of the current camera are obtained from the second camera parameter. When the forward-projected pixel does not correspond to a pixel on the pixel grid of the current camera, selecting the pixel of the pixel grid closest to the forward-projected pixel to obtain a corrected forward-projected pixel. Calculating a motion vector representing a displacement between the forward-projected pixel and the current pixel of the reference view. The device according to claim 13, comprising the same.

15. The device according to claim 13, wherein the device is configured to fill separated missing pixel values in each intermediate prediction image or the final prediction image with an average of adjacent pixel values, a median of adjacent pixel values, or a default value.

16. The device according to claim 13, wherein the at least one final prediction image is obtained by aggregating at least two intermediate prediction images.

Citation Information

Patent Citations

  • Image encoding device and method, image decoding device and method, and programs therefor

    WO2015141613A1

  • A method and an apparatus for reducing an amount of data representative of a multi-view plus depth content

    WO2019211429A1