Multiple perspectives signal codec

By utilizing redundant information between viewpoints in multi-view signal coding and employing coding parameter prediction and adaptive techniques, the compression rate and coding efficiency of multi-view signals are improved, and the problem of excessive data volume in multi-view signals is solved.

JP2026032148APending Publication Date: 2026-02-25DOLBY VIDEO COMPRESSION LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025203929
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2010-08-11
Filing Date
2025-11-26
Publication Date
2026-02-25

AI Technical Summary

Technical Problem

Existing multi-view signal coding technologies suffer from insufficient data compression rates when transmitting and processing multi-view signals, especially intermediate-view signals, due to the large amount of data. This makes it impossible to effectively utilize redundant information between viewpoints.

Method used

By employing interview-based encoding parameter prediction and adaptive techniques, and utilizing redundant information between viewpoints, the encoding process is optimized, including edge detection and interview prediction, thereby reducing the redundancy of encoding parameters.

Benefits of technology

It improves the compression rate and coding efficiency of multi-view signals, reduces the bandwidth requirements for data transmission, and maintains the quality of the view signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026032148000001_ABST
    Figure 2026032148000001_ABST
Patent Text Reader

Abstract

To provide a multiple perspectives signal codec which enables a higher compression rate or a better rate / distortion ratio.SOLUTION: The encoder 20 achieves a higher compression rate or a better rate / distortion ratio by adopting or predicting the second encoding parameters to be used for encoding the second view 122 of the multiple perspectives signal from the first encoding parameters 461, 481, 501 used when encoding the first view 121 of the multiple perspectives signal 10. The redundancy between the viewpoints of the multiple perspectives signal is not limited to the viewpoints themselves such as the video information of these viewpoints, but is a point that the coding parameters when coding these viewpoints in parallel also indicate similarity, and if this point is used, the coding rate can be further improved.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to coding of multi-view signals. [Background technology]

[0002] Multi-view signals are relevant for many applications, such as stereoscopic video applications, including binocular and multi-view displays, and free-viewpoint video applications. For stereoscopic and multi-view video content, the MVC (Multi-View Video Coding) standard [1, 2] has been developed. In this standard, video sequences from several nearby cameras are compressed. The MVC decoding process simply reconstructs these camera views at their original camera positions. However, different multi-view displays require different numbers of viewpoints with different spatial locations, e.g., additional viewpoints between the original camera positions. Summary of the Invention [Problem to be solved by the invention]

[0003] A problem with handling multi-view signals is the enormous amount of data required to transmit information about the numerous views contained within the multi-view signal. This amount of data is even greater when it comes to the amount of data required to enable extraction / combining of intermediate views, since in this case the video associated with each view may be accompanied by additional data, such as depth / disparity map data, that allows re-projection of each view onto other views, such as intermediate views. Due to this enormous amount of data, it is very important that the compression rate of the multi-view signal codec is as high as possible.

[0004] It is therefore an object of the present invention to provide a multi-view signal codec that allows a higher compression rate or a better rate / distortion ratio. [Means for solving the problem]

[0005] This object is achieved by the subject matter of the respective independent claims of the present application.

[0006] The embodiments provided by the present application utilize the finding that a higher compression rate or a better rate-distortion ratio can be achieved by adopting or predicting a second encoding parameter used to encode a second view of a multi-view signal from a first encoding parameter used to encode a first view of the multi-view signal. In other words, the inventors of the present application have discovered that the redundancy between the views of a multi-view signal is not limited to the view itself, such as the video information of the view, but also that the encoding parameters when encoding the view in parallel show similarity. By utilizing this finding, the encoding rate can be further improved.

[0007] Some embodiments of the present application are further based on the finding that the segmentation of a depth / disparity map associated with a given frame of a video of a given viewpoint used in encoding the depth / disparity map may be determined or predicted using edges detected in the video frame as clues, i.e., by determining wedgelet dividing lines to extend along the edges in the video frame. Although edge detection has the drawback of increasing the amount of computation on the decoder side, this drawback may be acceptable in application scenarios where a lower transmission rate at acceptable quality outweighs the increased computational overhead. Such scenarios may include broadcast applications where the decoder is configured as a stationary device.

[0008] In addition, some embodiments of the present application are further based on the finding that if the adaptation / prediction of coding parameters involves scaling of those coding parameters based on the ratio between spatial resolutions, then a view whose coding parameters have been adapted / predicted from coding information of other views can also be coded at a lower spatial resolution, i.e., predicted and residual-corrected, thereby saving coded bits.

[0009] Advantageous configurations of the above-mentioned aspects of the embodiments are the subject matter of the dependent claims of the present application.In particular, preferred embodiments of the present application are described below with reference to the accompanying drawings. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 2 is a block diagram of an encoder according to an embodiment of the present invention; [Figure 2] FIG. 1 is a schematic diagram showing a portion of a multi-view signal to illustrate information reuse across multiple viewpoints and image depth / disparity boundaries. [Figure 3] FIG. 2 is a block diagram of a decoder according to one embodiment of the present invention; [Figure 4] FIG. 10 is a diagram showing prediction configuration and motion / disparity vectors in an example multi-view coding with two views and two time instants. [Figure 5] FIG. 1 illustrates point correspondences between nearby viewpoints using disparity vectors. [Figure 6] FIG. 1 illustrates intermediate view synthesis with scene content projection from viewpoint 1 and viewpoint 2 using scaled disparity vectors. [Figure 7] FIG. 10 illustrates N-viewpoint extraction from separately decoded color and supplementary data to generate intermediate viewpoints at any viewpoint. [Figure 8] FIG. 10 shows an example of N-view extraction in a two-view bitstream for a nine-view display. DETAILED DESCRIPTION OF THE INVENTION

[0011] FIG. 1 is a block diagram of an encoder according to an embodiment of the present invention for encoding a multi-view signal. The multi-view signal of FIG. 1 is designated by the illustrative reference numeral 10 and includes two views 121 and 122. However, the embodiment of FIG. 1 may include more views. Furthermore, according to the embodiment of FIG. 1, each view 121 and 122 includes an image 14 and depth / disparity map data 16. However, many of the beneficial principles of the embodiment shown in FIG. 1 would remain beneficial even if used with a multi-view signal that includes views that do not include depth / disparity map data. Such generalizations of embodiments of the present invention are discussed below following the description of FIGS. 1-3.

[0012] The images 14 of each viewpoint 121 and 122 represent spatiotemporal sampling along different projection / viewpoint directions of a projection of a common scene. Preferably, the temporal sampling rates of the images 14 of each viewpoint 121 and 122 are equal to each other, but this condition does not necessarily have to be satisfied. As shown in Figure 1, each of the images 14 preferably includes a sequence of frames, each of which is associated with a respective time stamp t, t-1, t-2, ... In Figure 1, the image frames are respectively denoted by V 視点番号,時間スタンプ番号 Each frame V i,t represents the spatial sampling of scene i along each viewing direction at each timestamp t. i,t may include one or more sample arrays, such as one sample array for luma samples and two sample arrays for chroma samples, or may include only luminance samples, or may include sample arrays for other color components, such as an RGB color space or similar color components, etc. The spatial resolution of such one or more sample arrays may vary within an image 14 or within each image 14 of different viewpoints 121 and 122.

[0013] Similarly, the depth / disparity map data 16 represents spatiotemporal sampling of the depths of scene objects of a common scene, measured along the viewing direction of each of the viewpoints 121 and 122. The temporal sampling rate of this depth / disparity map data 16 may be the same as or different from the temporal sampling rate of the associated video of the same viewpoint, as shown in Figure 1. In the case of Figure 1, each of the video frames V has a depth / disparity map d associated with each of the depth / disparity map data 16 of each of the viewpoints 121 and 122. In other words, in the example of Figure 1, each video frame V of viewpoint i and timestamp t has a depth / disparity map d associated with it. i,t is the depth / disparity map d associated with it. i,t Regarding the spatial resolution of this depth / disparity map d, the same applies as described above for the video frame, i.e. the spatial resolution may differ between the depth / disparity maps of different viewpoints.

[0014] To efficiently compress the multi-view signal 10, the encoder of FIG. 1 encodes the views 121 and 122 in parallel into the data stream 18. However, the encoding parameters used to encode the first view 121 are reused to predict or adopt as the second encoding parameters to be used to encode the second view 122. In this way, the encoder of FIG. 1 takes advantage of a fact discovered by the inventors that encoding the views 121 and 122 in parallel produces similar results in determining the encoding parameters for these views within the encoder. Therefore, by efficiently utilizing the redundancy between these encoding parameters, the compression rate or rate-distortion ratio can be increased (where distortion is measured, for example, as the average distortion of both views, and rate is measured as the encoding rate of the entire data stream 18).

[0015] Specifically, the encoder of Figure 1 is generally designated by the reference numeral 20 and comprises an input for receiving the multi-view signal 10 and an output for outputting a data stream 18. As can be seen from Figure 2, the encoder 20 of Figure 1 comprises two coding branches for each of the views 121 and 122: one for the video data and one for the depth / disparity map data. That is, the encoder 20 comprises an encoding branch 22 for the video data of view 1. v,1 and coding branch 22 for the depth / disparity map data of viewpoint 1. d,1 and an encoding branch 22 for the second viewpoint video data. v,2 and an encoding branch 22 for the depth / disparity map data of the second viewpoint. d,2 Each of these encoding branches 22 is similarly constructed. For purposes of describing the construction and function of the encoder 20, the following will first describe the encoding branches 22. v,1 The structure and function of the branch 22 will be explained below, which function is common to all branches 22. The individual characteristics of the branches 22 will be explained later.

[0016] encoding branch 22 v,1 is for encoding the video 141 of the first viewpoint 121 of the multi-viewpoint signal 12, and the encoding branch 22 v,1 has an input that receives the image 141. In addition, the encoding branch 22 v,1 The subtractor 24 includes, connected in series in the order mentioned above, a subtractor 24, a quantization / transform module 26, a quantization / inverse transform module 28, an adder 30, an additional processing module 32, a decoded image buffer 34, two parallel-connected prediction modules 36 and 38, and a combiner or selector 40. The combiner or selector 40 is connected between the outputs of the prediction modules 36 and 38, on the one hand, and to an inverting input of the subtractor 24. The output of the combiner 40 is also connected to an additional input of the adder 30. The non-inverting input of the subtractor 24 receives the image 14.

[0017] encoding branch 22 v,1The elements 24 to 40 cooperate to encode the video 141. In this encoding, the video 141 is encoded in units of predetermined portions. For example, when encoding the video 141, a frame V 1,k is segmented into segments, such as blocks or other groups of samples. This segmentation may be constant in time or may vary over time. Furthermore, this segmentation may be known to the encoder and decoder by default or may be signaled in the data stream 18. This segmentation may be a regular frame-to-block segmentation, such as segmentation into non-overlapping blocks arranged in rows and columns, or a quad-tree-based segmentation into blocks of variable size. In the following description, the segment of the image 141 that is input to the non-inverting input of the subtractor 24 and is currently being coded is referred to as the current portion of the image 141.

[0018] The prediction modules 36 and 38 are intended to predict the current portion, and for this purpose, their inputs are connected to the decoded picture buffer 34. In fact, both prediction modules 36 and 38 use a previously reconstructed portion of the video 141 present in the decoded picture buffer 34 to predict the current portion / segment being input to the non-inverting input of the subtractor 24. From this perspective, the prediction module 36 acts as an intra predictor, spatially predicting the current portion of the video 141 from a spatially neighboring, already reconstructed portion of the same frame of the video 141. On the other hand, the prediction module 38 acts as an inter predictor, temporally predicting the current portion from a previously reconstructed frame of the video 141. Both prediction modules 36 and 38 perform their predictions according to predetermined prediction parameters or as described. More specifically, the latter parameters are determined by the encoder 20 within an optimization framework aimed at certain optimization goals, such as optimizing the rate / distortion ratio, subject to some constraints, such as a maximum bit rate, or without any constraints at all.

[0019] For example, intra-prediction module 36 may determine spatial prediction parameters for the current portion, such as a prediction direction, and may extend / copy the content of a neighboring, already-reconstructed portion of the same frame of video 141 to serve as a prediction for the current portion. Inter-prediction module 38 may predict the current portion from a previously reconstructed frame using motion compensation, where the relevant inter-prediction parameters may include a motion compensation vector, a reference frame index, motion compensation subdivision information for the current portion, a hypothesis number, or any combination thereof. Combiner or selector 40 may combine one or more of the predictions provided by modules 36 and 38, or simply select one. Combiner 40 transmits the resulting prediction for the current portion to an inverting input of subtractor 24 and an additional input of adder 30, respectively.

[0020] At the output of the subtractor 24, a residual of the prediction of the current part is output. The quantization / transform module 26 is configured to transform this residual signal by quantizing the transform coefficients. This transformation may be any spectrally decomposing transformation, such as a DCT. Quantization makes the result of the quantization / transform module 26 lossy, i.e., lossy coding. The output of the module 26 is a residual signal 421 to be transmitted in the data stream. The residual signal 421 is inversely quantized and inversely transformed in the module 28 so that the residual signal is reconstructed as much as possible, i.e., corresponds as much as possible to the residual signal output by the subtractor 24 despite the quantization noise. The adder 30 combines this reconstructed residual signal with the prediction of the current part by summation. Other combination methods are also possible. For example, the subtractor 24 may alternatively operate as a divider measuring the residual in ratio, and the adder may be configured as a multiplier reconstructing the current part. In this way, the output of the adder 30 represents a pre-reconstruction of the current part. Additional processing in module 32 may optionally be used to enhance the reconstruction. Such additional processing may include, for example, deblocking, adaptive filtering, and the like. All reconstructions available so far are buffered in decoded picture buffer 34. That is, decoded picture buffer 34 buffers pre-reconstructed frames of video 141 and pre-reconstructed portions of the current frame to which the current portion belongs.

[0021] To allow a decoder to reconstruct the multi-view signal from the data stream 18, the quantization / transform module 26 transmits the residual signal 421 to a multiplexer 44 of the encoder 20. At the same time, the prediction module 36 transmits intra-prediction parameters 461 to the multiplexer 44, the inter-prediction module 38 transmits inter-prediction parameters 481 to the multiplexer 44, and the additional processing module 32 transmits additional processing parameters 501 to the multiplexer 44. The multiplexer 44 multiplexes or inserts all of this information into the data stream 18.

[0022] As can be seen from the above description of the embodiment of FIG. 1, the encoding branch 22 v,1 The method for encoding video 141 according to is self-contained in that the encoding is independent of the depth / disparity map data 161 and the data of any other viewpoints 122. From a more general perspective, the encoding branch 22 v,1 Coding according to may be seen as a method of encoding a video 141 into a data stream 18 by determining coding parameters, predicting a current portion of the video 141 according to the first coding parameters from a previously coded portion of the video 141, which portion has been coded into the data stream 18 by the encoder 20 prior to the coding of the current portion, and determining a prediction error in the prediction of the current portion to obtain correction data, i.e. the aforementioned residual signal 421, which coding parameters and correction data are inserted into the data stream 18.

[0023] encoding branch 22 v,1 The above-mentioned coding parameters inserted into the data stream 18 by may include one, a combination or all of the following:

[0024] First, the coding parameters for the video 141 may define / signal a segmentation of the frames of the video 141 in a manner as briefly described above.

[0025] Furthermore, the coding parameters may include coding mode information indicating for each segment or current portion which coding mode should be used to predict each segment, such as intra prediction, inter prediction or a combination thereof.

[0026] The coding parameters may include the prediction parameters mentioned above, such as intra-prediction parameters for parts / segments predicted by intra-prediction and inter-prediction parameters for parts / segments predicted by inter-prediction.

[0027] The coding parameters may also include additional processing parameters 501 that signal to the decoder how already reconstructed portions of the video 141 should be further processed before using them to predict the current or subsequent portions of the video 141. These additional processing parameters 501 may include indices that point to respective filters, filter coefficients, or the like.

[0028] The prediction parameters 461, 481 and further processing parameters 501 may further include sub-segmentation data which defines a further sub-segmentation in comparison to the segmentation described above, which defines the granularity of the mode selection, or which defines a completely independent segmentation, such as when applying different adaptive filters for different parts of a frame in further processing.

[0029] Coding parameters may influence the determination of the residual signal and may therefore be part of the residual signal 421. For example, the spectral transform coefficient levels output by the quantization / transform module 26 may be recognized as correction data, while the quantization step size may also be signaled in the data stream 18, and the quantization step size parameter may be recognized as a coding parameter in the sense explained below.

[0030] The coding parameters may further define prediction parameters that define a second prediction stage, next to the prediction residual of the first prediction stage described above. In this regard, intra / inter prediction may be used.

[0031] To improve coding efficiency, the encoder 20 includes a coding information exchange module 52, which receives all coding parameters and additional information that affect or are affected by the processing in, for example, modules 36, 38, and 32. This is illustrated in FIG. 1 by vertical arrows extending from each module downward to the coding information exchange module 52. The role of the coding information exchange module 52 is to share coding parameters and, optionally, additional coding information among the coding branches 22, so that each coding branch 22 can predict or adopt each other's coding parameters. For this purpose, in the embodiment shown in FIG. 1, an order is defined between the data entities, i.e., video and depth / disparity map data, for the views 121 and 122 of the multi-view signal 10. Specifically, video 141 of a first viewpoint 121 precedes, followed by depth / disparity map data 161 for the first viewpoint, followed by video 142 of a second viewpoint 122, followed by depth / disparity map data 162 for the second viewpoint, and so on. Note that this strict ordering among the data entities of multi-view signal 10 need not be strictly applied when encoding the entire multi-view signal 10. However, for ease of explanation, the following assumes that this ordering is constant. Naturally, this ordering among the data entities also dictates the ordering among the branches 22 associated with them.

[0032] As already mentioned above, the coding branch 22 d,1 , encoding branch 22 v,2 and coding branch 22 d,2 Further coding branches of the coding branch 22 such as v,1 However, due to the above-mentioned order between the image and depth / disparity map data of the viewpoints 121 and 122 and the corresponding order defined between the encoding branches 22, the encoding branches 22 d,1has an additional degree of freedom when predicting the coding parameters to be used for encoding the current portion of the depth / disparity map data 161 of, for example, the first viewpoint 121. This is due to the above-mentioned ordering between the video and depth / disparity map data of the different viewpoints. That is, for example, each of these entities may be coded using its own reconstructed portion or may be coded using the entity of that data that is located before it in the above-mentioned ordering between these data entities. Therefore, when coding the depth / disparity map data 161, the coding branch 22 d,1 can use information that is known from a previously reconstructed portion of the corresponding image 141. d,1 141 can be used to predict certain characteristics of the depth / disparity map data 161, thereby increasing the compressibility of the depth / disparity map data 161. However, in addition to this, the encoding branch 22 d,1 may obtain the coding parameters for encoding the depth / disparity map data 161 by predicting / adapting the coding parameters used in encoding the video 141 as described above. In the case of adaptation, it is possible to suppress signaling of any coding parameters for the depth / disparity map data 161 in the data stream 18. In the case of prediction, it is also possible to simply signal prediction residuals / correction data for these coding parameters in the data stream 18. Examples of such prediction / adaptation of coding parameters are also described below.

[0033] For the subsequent data entities, i.e., the image 142 of the second view 122 and the depth / disparity map data 162, additional prediction possibilities exist. For these coding branches, the inter-prediction module not only performs temporal prediction, but can also perform inter-view prediction. The corresponding inter-prediction parameters contain similar information compared to the temporal prediction, i.e., for each inter-view predicted segment, a disparity vector, a view index, a reference frame index and / or an indication of the number of hypotheses, i.e., the number of inter-view predictions involved in forming the inter-view prediction, for example by summing. Such inter-view prediction is performed by the branch 22 for the image 142. v,2 In addition to being useful for the depth / disparity map data 162, d,2 1. The inter-view prediction parameters are also valid for the inter-prediction module 38 of the first embodiment. Of course, these inter-view prediction parameters also represent coding parameters that can serve as the basis for adoption / prediction for subsequent view data of a potential third view, which is not shown in FIG.

[0034] In this way, the amount of data to be inserted into the data stream 18 by the multiplexer 44 is further reduced. d,1 and branch 22 v,2 and branch 22 d,2 The amount of coding parameters for the coding branch 22 can be significantly reduced by adopting the coding parameters of the preceding coding branch or by simply inserting a prediction residual for the coding parameters of the preceding coding branch into the data stream 18 by the multiplexer 44. The ability to select between temporal prediction and inter-view prediction allows the coding branch 22 to v,2 and coding branch 22 d,2 The amount of residual data 423, 424 can also be reduced. The reduction in the amount of residual data more than compensates for the increased amount of coding operations required to differentiate between temporal and inter-view prediction modes.

[0035] For a more detailed explanation of the coding parameter adoption / prediction, please refer to Figure 2. Figure 2 shows an example portion of a multi-view signal 10. Figure 2 shows a video frame v 1,t is segmented into segments or portions 60a, 60b, and 60c. 1,t Although only three portions of a video frame v are shown, the segmentation may seamlessly and contiguously divide the frame into segments / portions. 1,t The segmentation of the image 141 may be constant or may vary over time. Furthermore, this segmentation may or may not be signaled in the data stream. In the example shown in Figure 2, portions 60a and 60b are temporally predicted using motion vectors 62a and 62b from a reconstructed version of an arbitrary reference frame of image 141, which in this case is illustratively frame v 1,t-1 As is known in the art, the coding order of frames of video 141 does not necessarily match the order of appearance of those frames, and therefore the reference frame is the current frame v 1,t In time order 64, portion 60c may be later than portion 60b in time order 64. For example, portion 60c is an intra-predicted portion, for which intra-prediction parameters are inserted into data stream 18.

[0036] encoding branch 22 d,1 is the depth / disparity map d 1,t When encoding the above possibilities may be utilized in one or more of the ways described below, by way of example, with reference to FIG.

[0037] -For example, coding branch 22 d,1 is the depth / disparity map d 1,t When encoding, the encoding branch 22 v,1 The video frame v used by 1,t Then, if a video frame v 1,tIf segmentation parameters exist in the coding parameters for 1,t Alternatively, the retransmission of the coding branch 22 can be omitted. d,1 is the video frame v 1,t Segmentation of depth / disparity map 1,t , and in that case, the video frame v of the segmentation to be used 1,t 2, the deviation from the encoding branch 22 is signaled via the data stream 18. d,1 is the video frame v 1,t Segmentation of depth / disparity map 1,t This is used as a pre-segmentation of the coding branch 22. d,1 is the video v 1,t The previous segmentation is taken from the segmentation of or predicted from it.

[0038] -Furthermore, coding branch 22 d,1 is the video frame v 1,t From the coding modes assigned to the portions 60a, 60b, and 60c in 1,t The coding mode of each portion 60a, 60b, 60c of the video frame v may be adopted or predicted. 1,t and depth / disparity map d 1,t If the segmentation is different between video frame v and 1,t The adoption / prediction of the coding mode from the video frame v 1,t The depth / disparity map d may be controlled to be obtained from co-located portions of the segmentation of 1,t The video frame v for the current part of 1,t The common position part in the depth / disparity map d 1,tIn the case of prediction of the coding mode, the coding branch 22 d,1 is the video frame v explicitly signaled in the data stream 18. 1,t Depth / disparity map d for the coding mode in 1,t The difference in the coding modes of the portions 60a to 60c may be signaled.

[0039] - As far as prediction parameters are concerned, the coding branch 22 d,1 is the same depth / disparity map d 1,t or spatially adopt / predict prediction parameters used to encode neighboring portions of video frame v 1,t For example, Fig. 2 shows the depth / disparity map d 1,t The portion 66a of the video frame v is an inter-predicted portion, and the corresponding motion vector 68a is 1,t 4 shows that the motion vector 62a of the co-located portion 60a of the data stream 18 may be taken or predicted from the motion vector 62a of the co-located portion 60a of the data stream 18. In the case of prediction, the difference in the motion vectors need only be inserted into the data stream 18 as part of the inter prediction parameters 482.

[0040] In order to improve the coding efficiency, the coding branch 22 d,1 is used to generate the depth / disparity map d 1,t 2, it is desirable to have the ability to sub-divide the segments of the previous segmentation of d and signal the location of this Wedgelet division line 70 in the data stream 18 to the decoder. In this way, in the example shown in FIG. 1,t The portion 66c is subdivided into two wedgelet-shaped portions 72a and 72b. d,1The encoder 22 may be configured to encode these sub-segments 72a and 72b separately. In the example shown in FIG. 2, both sub-segments 72a and 72b are illustratively inter-predicted using respective motion vectors 68c and 68d. If intra-prediction is used for both sub-segments 72a and 72b, each DC value for each segment may be derived by extrapolation from the DC values ​​of neighboring related segments. In this case, each of these derived DC values ​​may be optionally refined by transmitting the corresponding refined DC value to the decoder as an intra-prediction parameter. There are several possibilities for enabling the decoder to determine the Wedgelet division line used by the encoder to subdivide the previous segmentation of the depth / disparity map. d,1 Alternatively, the encoding branch 22 may use any of these possibilities exclusively. d,1 may have the freedom to select from among coding alternatives as described below and signal that choice to the decoder as side information in the data stream 18.

[0041] The wedgelet dividing line 70 may be, for example, a straight line. Signaling the location of this dividing line 70 to the decoder may include signaling a single intersection point along the boundary of segment 66c along with slope or gradient information, or along with an indication of two intersection points between the wedgelet dividing line 70 and the boundary of segment 66c. In some embodiments, the wedgelet dividing line 70 may be explicitly signaled in the data stream by indicating two intersection points between the wedgelet dividing line 70 and the boundary of segment 66c, where the granularity of the grid indicating the possible intersection points, i.e., the granularity or resolution of the intersection indication, may depend on coding parameters such as the size of segment 66c or the quantization parameter.

[0042] In an alternative embodiment, pre-segmentation is performed by a quadtree-based block partitioning, e.g., using dyadic square blocks, and the allowable set of intersection points for each block size is given by a look-up table (LUT), where the signaling of each intersection point includes the signaling of the corresponding LUT index.

[0043] However, according to yet another embodiment, the coding branch 22 d,1 is the video frame v in the decoded image buffer 34. 1,t , and signals in the data stream to the decoder any deviation in the Wedgelet dividing lines 70 that should actually be used in encoding segment 66c. 1,t 66c in the image frame v 1,t For example, the edge detection may be performed on a video frame v 1,t , where the spatial gradient of a scaled feature, such as vividness, luminance, saturation, or chrominance, or the like, exceeds a minimum threshold. Module 52 can determine a Wedgelet dividing line 70 based on the location of this edge 72, such that the Wedgelet dividing line extends along edge 72. The decoder also generates a reconstructed video frame v 1,t, the decoder can similarly determine Wedgelet dividing line 70 for subdividing portion 66c into Wedgelet-shaped portions 72a and 72b, thereby saving signaling capacity for signaling Wedgelet dividing line 70. The aspect of having a resolution that depends on the size of portion 66c for representing the position of Wedgelet dividing line 70 can also be applied to aspects of the invention in which the position of dividing line 70 is determined by edge detection, and can also be applied to signal any deviation from the expected position.

[0044] When encoding the video 142, the encoding branch 22 v,2 is the coding branch 22 v,1 In addition to the selection of coding modes available for the video, there is also the option of inter-view prediction.

[0045] FIG. 2 shows, for example, a video frame v 2,t The segmentation portion 74b of the first viewpoint image 141 corresponds to the time-corresponding image frame v 1,t 7 shows how inter-viewpoint prediction is performed using a disparity vector 76.

[0046] Apart from these differences, the coding branch 22 v,2 is the video frame v 1,t and depth / disparity map d 1,t All information available from the encoding of the sigma and sigma, in particular the encoding parameters used in these encodings, may be additionally utilized. v,2 is the video frame v 2,t For the temporally inter-predicted portion 74a of the video frame v, the temporally aligned video frame v 1,t and the motion vector 62a of the common position portion 60a of the depth / disparity map d 1,tThe motion parameters, including motion vector 78, may be adapted or predicted from either the motion vector 68a of co-located portion 66a of the first frame or a combination thereof. With respect to the inter prediction parameters for portion 74a, the prediction residual, if present, may be signaled. In this regard, it should be noted that motion vector 68a may have already been predicted / adapted from motion vector 62a itself.

[0047] Depth / Disparity Map 1,t As described above for encoding a video frame v 2,t Other possibilities for adopting / predicting coding parameters to encode , are also discussed in the coding branch 22. v,2 Video frame v 2,t However, in that case, the video frame v 1,t The encoding parameters of and the corresponding depth / disparity map d 1,t The available common data provided by module 52 is increased since both the encoding parameters of the .DELTA..DELTA. and the .DELTA. ... coding parameters are available.

[0048] Then the encoding branch 22 d,2 is the coding branch 22 d,1 Depth / disparity map by 1,t In a similar way to encoding the depth / disparity map d 2,t This means that, for example, a video frame v 2,t This is also true for all cases where the coding parameters are adopted / predicted from the coding branch 22. d,2 furthermore, the depth / disparity map d of the preceding viewpoint 121 1,t Furthermore, the coding branch 22 has the opportunity to adopt / predict the coding parameters from the coding parameters used to code the d,2 is the coding branch 22 v,2 As explained with respect to

[0047] , inter-view prediction can also be used.

[0049] Regarding the adoption / prediction of coding parameters, the coding branch 22d,2 It may be useful to restrict the possibility of adopting / predicting its coding parameters from among the coding parameters of previously coded entities of the multi-view signal 10 to the video 142 of the same view 122 and the depth / disparity map data 161 of nearby and previously coded views 121. As a result, the depth / disparity map d 2,t This reduces the signaling overhead that arises from the need to signal in the data stream 18 to the decoder side the source of the adoption / prediction for each part of the coding branch 22. d,2 is the depth / disparity map d 2,t For the inter-view predicted portion 80a of the video frame v 2,t 2. In this case, the indication of the data entity from which the prediction is made, i.e., the image 142 in FIG. 2, may be omitted, since the depth / disparity map d 2,t However, in adapting / predicting the inter-prediction parameters of the temporally inter-predicted portion 80b, the coding branch 22 d,2 may adopt / predict the corresponding motion vector 84 from any one of the motion vectors 78, 68a and 62a, thus coding branch 22 d,2 may be configured to signal the source of the adoption / prediction for the motion vectors 84 in the data stream 18, thereby reducing overhead by limiting the possible sources to video 142 and depth / disparity map data 161.

[0050] Regarding the Wedgelet dividing line, the coding branch 22 d,2 In addition to the above-mentioned methods, the following options are also available:

[0051] Depth / disparity map d of viewpoint 122 using Wedgelet dividing line 2,tTo encode the signal d, we use methods such as edge detection and the corresponding implicit derivation of Wedgelet dividing lines. 1,t where the depth correction is done by using the depth / disparity map d 1,t The detected dividing lines in the depth / disparity map d 2,t For depth correction, the depth / disparity map d 1,t The foreground depth / disparity values ​​along each detected edge in may be used.

[0052] Alternatively, the depth / disparity map d of the viewpoint 122 using the Wedgelet dividing line 2,t As a method of encoding the signal d 1,t , i.e., using a predetermined Wedgelet dividing line within the disparity-corrected portion of the signal d 1,t By using or adopting as a predictor the Wedgelet division line already used in encoding the portion at the common position of 1,t The corresponding disparity corrected portion of

[0053] Having described the encoder 20 of FIG. 1, it should be noted that the encoder can be implemented in software, hardware, or firmware, i.e., programmable hardware. While the block diagram of FIG. 1 depicts the encoder 20 as having parallel encoding branches, i.e., one encoding branch for each image and depth / disparity data in the multi-viewpoint signal 10, this is not a requirement. For example, software routines, circuitry, or programmable logic configured to perform the tasks of each of the components 24-40 may be used sequentially to perform the tasks of each of the encoding branches. In parallel processing, the operations of the parallel encoding branches may be performed in parallel processor cores, parallel operating circuits, etc.

[0054] Figure 3 shows an example of a decoder capable of decoding data stream 18, which reconstructs one or more view images corresponding to a scene represented by a multi-view signal from data stream 18. The structure and functionality of the decoder shown in Figure 3 are largely similar to the structure and functionality of the encoder in Figure 1. Therefore, the reference numerals of Figure 1 are reused as much as possible to indicate that the functional description of Figure 1 above also applies to Figure 3.

[0055] The decoder of Figure 3 is generally designated by the reference numeral 100 and has an input for a data stream 18 and an output for outputting the reconstruction 102 of one or more views as described above. The decoder 100 comprises a demultiplexer 104, a pair of decoding branches 106 for each data entity of the multi-view signal 10 (see Figure 1) represented by the data stream 18, a viewpoint extractor 108 and a coding parameter exchanger 110. As in the encoder of Figure 1, these decoding branches 106 comprise the same decoding elements in the same interconnections and are represented by a decoding branch 106 responsible for decoding the image 141 of the first view 121. v,1 Specifically, each decoding branch 106 has an input connected to the output of the multiplexer 104 and an output connected to the input of the viewpoint extractor 108. v,11, and an output for outputting a respective data entity of the multi-view signal 10, which in this case is a video 141, to the viewpoint extractor 108. Between this input and the output, each decoding branch 106 includes an inverse quantization / inverse transform module 28, an adder 30, an additional processing module 32, and a decoded picture buffer 34, which are connected in series between the multiplexer 104 and the viewpoint extractor 108. The adder 30, the additional processing module 32, and the decoded picture buffer 34, together with the parallel-connected prediction modules 36 and 38 and their subsequent combiner / selector 40, form a closed circuit. The parallel-connected prediction modules 36 and 38 and the combiner / selector 40 are connected, in the aforementioned order, between the decoded picture buffer 34 and an additional input of the adder 30. As indicated using the same reference numerals as in FIG. 1, the structure and function of the components 28-40 of the decoding branch 106 are similar to the structure and function of the corresponding elements of the encoding branch of FIG. 1. That is, each component of the decoding branch 106 performs a task similar to that in the encoding process, using the information transmitted in the data stream 18. Of course, each decoding branch 106 simply reverses the encoding procedure for each coding parameter ultimately selected by the encoder 20. On the other hand, the encoder 20 of Figure 1 must find an optimal set of coding parameters in an optimization sense, e.g., coding parameters that optimize a rate / distortion cost function, possibly subject to certain constraints, e.g., a maximum bit rate or similar constraint.

[0056] The demultiplexer 104 supplies the data stream 18 to the various decoding branches 106. For example, the demultiplexer 104 supplies residual data 421 to the quantization / inverse transform module 28, additional processing parameters 501 to the additional processing module 32, intra prediction parameters 461 to the intra prediction module 36, and inter prediction parameters 481 to the inter prediction module 38. The coding parameter exchanger 110 serves a similar role to the corresponding module 52 in Figure 1, supplying common coding parameters and other common data to the various decoding branches 106.

[0057] The viewpoint extractor 108 receives the multi-view signal reconstructed by the parallel decoding branches 106 and extracts from the multi-view signal one or more viewpoints 102 corresponding to the viewpoint angles or viewpoint directions defined by externally provided intermediate viewpoint extraction control data 112.

[0058] Since the structure of the decoder 100 is similar to that of the corresponding part of the encoder 20, the functionality up to the interface to the viewpoint extractor 108 can be easily explained in a similar manner to that described above.

[0059] In fact, the decoding branch 106 v,1 and 106 d,1 and cooperate to reconstruct the first view 121 of the multi-view signal 10 from the data stream 18, in which case the first coding parameters (e.g., scaling parameters in 421, parameters 461, 481, 501 and corresponding unadopted parameters) contained in the data stream 18 are used together with the second branch 106. d,1This reconstruction is performed by predicting the current portion of the first view 121 from a previously reconstructed portion of the multi-view signal 10, i.e., a portion reconstructed from the data stream 18 prior to the reconstruction of the current portion of the first view 121, based on the prediction residual of the coding parameters, i.e., 422, and parameters 462, 482, 502, etc., and correcting the prediction error of the prediction of the current portion of the first view 121 using first correction data, i.e., the data in 421 and 422 contained in the data stream 18. v,1 is responsible for decoding the video 141, while the decoding branch 106 d,1 is responsible for reconstructing the depth / disparity map data 161. See for example Figure 2, where the decoding branch 106 v,1 reconstructs the image 141 of the first view 121 from the data stream 18, by predicting the current portion of the image 141, such as 60a, 60b or 60c, from a previously reconstructed portion of the multi-view signal 10 based on the corresponding coding parameters read from the data stream 18, i.e., the scaling parameters in 421 and the parameters 461, 481, 501, and correcting the prediction error of that prediction using the corresponding correction data obtained from the data stream 18, i.e., from the transform coefficient levels in 421. For example, the decoding branch 106 v,1The image 141 is processed in units of segments / portions, using the coding order between video frames, and the coding order between segments of the frames is used to decode the segments within a frame in the same manner as the corresponding coding branch of the encoder. In this way, all previously reconstructed portions of the image 141 can be used for predicting the current portion. The coding parameters for the current portion may include one or more intra-prediction parameters 50, inter-prediction parameters 481, filter parameters for the additional processing module 32, etc. Correction data for correcting prediction errors may be represented by spectral transform coefficient levels in the residual data 421. Not all of these coding parameters need to be transmitted in their entirety. Some of the coding parameters may be spatially predicted from neighboring segments of the image 141. For example, motion vectors for the image 141 may be transmitted in the bitstream as vector differences between the motion vectors of neighboring portions / segments of the image 141.

[0060] Second Decoding Branch 106 d,1 In terms of this, the decoding branch 106 d,1 are signaled in the data stream 18 and sent by the demultiplexer 104 to the individual decoding branches 106, as are the residual data 422 and the corresponding prediction and filter parameters. d,1 , i.e., parameters that are not predicted across inter-view boundaries, but also via a demultiplexer 104 to a decoding branch 106. v,1 , or any information that can be derived from that information, such as provided via the coding information exchange module 110. d,1 sends the coding parameters for reconstructing the depth / disparity map data 161 to the decoding branch 106 via the demultiplexer 104 for the first viewpoint 121. v,1 and 106 d,1The decoding branch 106 determines the coding parameters from a portion of the coding parameters sent for the pair of v,1 , which partially overlaps with the portion of the coding parameters specifically selected and transmitted for the decoding branch 106. d,1 On the other hand, for example, the motion vector 68a is 1,t 481 as a motion vector difference for another neighboring portion of the image, and on the other hand, from a motion vector 62a explicitly transmitted in 481 as a motion vector difference for another neighboring portion of the image, and on the other hand, from a motion vector difference explicitly transmitted in 482. Additionally or alternatively, the decoding branch 106 d,1 may use the reconstructed portion of the video 141 as described above for predicting the Wedgelet dividing lines to predict the coding parameters for decoding the depth / disparity map data 161.

[0061] More specifically, the decoding branch 106 d,1 reconstructs depth / disparity map data 161 for the first viewpoint 121 from the data stream using the coding parameters, which are at least in part transmitted to the decoding branch 106 v,1 and / or are predicted (or adopted) from the coding parameters used by the decoding branch 106 v,1 The decoding branch 106 is predicted from the reconstructed portion of the video 141 in the decoded picture buffer 34. The coding parameter prediction residual may be obtained from the data stream 18 via the demultiplexer 104. d,1 The other coding parameters for the depth / disparity map data 161 may be transmitted in their entirety in the data stream 18 or in relation to other references, i.e. with reference to coding parameters already used to code any of the previously reconstructed parts of the depth / disparity map data 161 itself. Based on these coding parameters, the decoder branch 106 d,1may be used to decode the current portion of the depth / disparity map data 161 from a previously reconstructed portion of the depth / disparity map data 161, i.e., prior to the reconstruction of the current portion of the depth / disparity map data 161, by the decoding branch 106. d,1 and corrects the prediction error of the prediction of the current portion of the depth / disparity map data 161 using the respective correction data 422 .

[0062] Thus, data stream 18 may include, for example, for portion 66a of depth / disparity map data 161, the following items: an indication as to whether or which part the coding parameters for the current part should be taken or predicted, for example from corresponding coding parameters of a co-located and temporally aligned part of the image 141 (or from other specific data of the image 141, such as a reconstructed version of the image 141 for predicting the Wedgelet dividing line); - if it should be adopted or predicted, the residuals of the coding parameters in the prediction All coding parameters for the current portion, if they should not be adopted or predicted, which may be signaled as prediction residuals compared with the coding parameters of a previously reconstructed portion of the depth / disparity map data 161. - In the case where not all coding parameters should be predicted / adopted as described above, the remaining part of the coding parameters for the current part may be signaled as a prediction residual compared with the coding parameters of a previously reconstructed part of the depth / disparity map data 161.

[0063] For example, if the current portion is an inter-predicted portion, such as portion 66a, motion vector 68a may be signaled in data stream 18 as being derived or predicted from motion vector 62a. d,1may predict the position of the Wedgelet dividing line 70 depending on the edges 72 detected in the reconstructed portion of the image 141 as described above, and may apply this Wedgelet dividing line without signaling in the data stream 18, or may apply this Wedgelet dividing line depending on the individual application signaling in the data stream 18. In other words, the use of the Wedgelet dividing line prediction for the current portion may be suppressed or allowed due to signaling in the data stream 18. In further other words, the decoding branch 106 d,1 can effectively predict the surroundings of the currently reconstructed portion of the depth / disparity map data.

[0064] A pair of decoding branches 106 for the second viewpoint 122 v,2 , 106 d,2 The function of the pair of decoding branches 106 is similar to that of the pair of decoding branches 106 for the first view 121, as described above for the encoding. Both decoding branches use their own coding parameters to jointly reconstruct the second view 122 of the multi-view signal 10 from the data stream 18. Only the part of these coding parameters that is not adopted / predicted across the view boundary between the views 121 and 122 is used by the two decoding branches 106. v,2 and 106 d,2 , and optionally the remainder of the inter-view predicted portion may also be transmitted and distributed via the demultiplexer 104. The current portion of the second view 122 is predicted from a previously reconstructed portion of the multi-view signal 10, i.e., a portion reconstructed from the data stream 18 by one of the decoding branches 106 prior to the reconstruction of the current portion of the second view 122, and furthermore, correction data, i.e., correction data, is transmitted and distributed by the demultiplexer 104 to the pair of decoding branches 106 of these. v,2 and 106 d,2 Using the provided 423 and 424, the prediction error is corrected appropriately.

[0065] Decoding branch 106 v,2 sends its coding parameters to the decoding branch 106 v,1 and 106 d,1The following information of the coding parameters may be present for the current portion of the video 142: an indication as to whether or which coding parameters for the current portion should be taken or predicted, for example from the corresponding coding parameters of a co-located and temporally aligned portion of the video 141; - the residuals of the coding parameters in the prediction, if they should be adopted or predicted All coding parameters for the current portion, if they should not be adopted or predicted, which may be signaled as a prediction residual compared with the coding parameters of a previously reconstructed portion of the video 142. - if not all coding parameters are to be predicted / adapted as described above, the remaining coding parameters for the current portion, which may be signaled as a prediction residual compared with the coding parameters of a previously reconstructed portion of the video 142; A signal may be signaled in the data stream 18 as to whether, for the current portion 74a, the corresponding coding parameters for that portion, such as for example the motion vectors 78, should be read completely anew from the data stream, or should be predicted spatially, or should be predicted from the motion vectors of the co-located portion of the image 141 of the first viewpoint 121 or the depth / disparity map 161, and accordingly the decoding branch 106 v,2 may operate by either extracting, employing, or predicting motion vectors 78 from data stream 18, and, if predicting, may also extract prediction error data from data stream 18 relating to the coding parameters for current portion 74a.

[0066] Decoding branch 106 d,2 may operate in a similar manner. d,2 is the decoding branch 106 v,1 and 106 d,1 and 106v,2 The coding parameters may be determined by at least partially adopting / predicting coding parameters used by either the first viewpoint 121 or the second viewpoint 122, the reconstructed image 142, and / or the reconstructed depth / disparity map data 161 of the first viewpoint 121. For example, the data stream 18 may signal, for a current portion 80b of depth / disparity map data 162, whether or which coding parameters for the current portion 80b should be adopted or predicted from co-located portions of the image 141, the depth / disparity map data 161, or the image 142, or from appropriate subsets thereof. Important portions of these coding parameters may include, for example, motion vectors such as 84 or disparity vectors such as 82. Additionally, other coding parameters, such as those related to Wedgelet dividing lines, may be detected by the decoding branch 106 using edge detection in the image 142. d,2 Alternatively, edge detection may be applied to the reconstructed depth / disparity map data 161 using a predetermined reprojection, where the reprojection is the depth / disparity map d 1,t The position of the detected edges in the depth / disparity map d 2,t This serves as the basis for predicting the location of the Wedgelet dividing line.

[0067] In either case, the reconstructed portions of the multi-view signal 10 reach the viewpoint extractor 108, where the viewpoints contained in those reconstructed portions form the basis for viewpoint extraction of the new viewpoints, i.e., the images associated with the new viewpoints. Such viewpoint extraction may include or involve reprojecting the images 141 and 142 using their associated depth / disparity map data. In simple terms, when reprojecting an image to another intermediate viewpoint, portions of that image corresponding to portions of the scene closer to the viewer are shifted more along the disparity direction, i.e., the direction of the gaze direction difference vector, than portions of that image corresponding to portions of the scene farther from the viewer. An example of viewpoint extraction performed by the viewpoint extractor 108 is described below with reference to FIGS. 4-6 and 8. Occlusion compensation processing may also be performed by the viewpoint extractor.

[0068] Before describing other embodiments below, it should be noted that some modifications may be made to the above-described embodiments. For example, the multi-view signal 10 does not necessarily have depth / disparity map data for each view. It is possible that none of the views of the multi-view signal 10 has depth / disparity map data associated with it. However, the reuse and sharing of coding parameters among the multiple views as described above can improve coding efficiency. Furthermore, for some views, the depth / disparity map data transmitted in the data stream may be limited to occlusion-compensated areas, i.e., areas that should fill the occlusion-compensated areas in the view reprojected from other views of the multi-view signal. The depth / disparity map data may be set to values ​​that have no effect in the remaining areas of the map.

[0069] As mentioned above, the views 121 and 122 of the multi-view signal 10 may have different spatial resolutions, i.e., these two views may be transmitted using different resolutions in the data stream 18. In other words, in the above-mentioned order between the views, view 121 followed by view 122, the coding branch 22 v,1 and 22 d,1 The spatial resolution at which the predictive coding of the viewpoint 121 is performed is determined by the coding branch 22. v,2 and 22 d,2 The spatial resolution of the view 121 and the view 122 may be higher than the spatial resolution at which the predictive coding of the view 122 is performed. The inventors of the present invention have found that this further improves the rate-to-distortion ratio when the quality of the synthesized view 102 is taken into account. For example, the encoder of FIG. 1 may initially receive the multi-view signal 10 at the same spatial resolution for the views 121 and 122, and then downsample the image 142 and depth / disparity map data 162 of the view 122 to a lower spatial resolution before they are subjected to the predictive coding process by modules 24-40. However, the above-described method of adapting and predicting coding parameters across view boundaries can also be implemented by scaling the coding parameters that form the basis of the adaptation or prediction according to the ratio between the different resolutions of the source view and the destination view. See, for example, FIG. 2. If the encoding branch 22 v,2 However, if we want to adopt or predict motion vector 78 from either motion vector 62a or 68a, then coding branch 22 v,2 can downscale one of its motion vectors 62a and 68a by a value corresponding to the ratio between the high resolution of the source viewpoint 121 and the low resolution of the destination viewpoint 122. Of course, a similar process can be applied to the decoder and decoding branch 106. That is, the decoding branch 106 v,2 and 106 d,2 The decoding branch 106 v,1 and 106 d,1 Predictive decoding can be performed at a lower spatial resolution compared to the , in which case the decoding branch 106 is used after the reconstruction and before reaching the viewpoint extractor 108. v,2and 106 d,2Upsampling may be used to convert the reconstructed video and depth / disparity map output from the decoded image buffers 34 from a lower spatial resolution to a higher spatial resolution. Respective upsamplers may be disposed between each decoded image buffer 34 and a respective input of the viewpoint extractor 108. As described above, within a single viewpoint 121 or 122, the video and associated depth / disparity map data may have the same spatial resolution. However, additionally or alternatively, these video and depth / disparity map data pairs may have different spatial resolutions, and the above-described method may be performed across spatial resolution boundaries, i.e., between the depth / disparity map data and the video. Furthermore, according to other embodiments, there may be three viewpoints, including viewpoint 123, which is not shown in FIGS. 1-3 for illustrative purposes. In this case, the first and second viewpoints may have the same spatial resolution, while the third viewpoint 123 may have a lower spatial resolution. Furthermore, according to the above-described embodiment, subsequent viewpoints, such as viewpoint 122, may be downsampled before encoding and upsampled after decoding. Such downsampling and upsampling represent a kind of pre- or post-processing, respectively, of the encoding / decoding branch, whereby the coding parameters used to adapt / predict the coding parameters of any subsequent (destination) view are scaled according to the respective ratio of the spatial resolutions of the source and destination views. As mentioned above, the processing within the intermediate view extractor 108 ensures that the lower quality of the transmitted and predictively coded subsequent views, such as 122, does not significantly affect the quality of the intermediate view output 102 of the intermediate view extractor 108. The view extractor 108 performs a kind of interpolation / low-pass filtering on the images 141 and 142 by reprojecting to the intermediate view and necessary resampling of the reprojected image sample values ​​onto the sample grid of the intermediate view. Taking advantage of the fact that the first view 121 has been transmitted at a higher spatial resolution than the neighboring view 122, the intermediate view between them may be obtained primarily from view 121, in which case the lower spatial resolution view 122 and its image 142 may simply be used as supplementary view.That is, for example, viewpoint 122 and its image 142 may be used simply to fill the occlusion-compensated region of the reprojected version of image 141, or may be used as a small weighting factor when performing some averaging between the reprojected versions of the image with viewpoint 121 on the one hand and viewpoint 122 on the other. In this way, the lower spatial resolution of viewpoint 122 is compensated for, even if the coding rate of second viewpoint 122 is significantly reduced for transmission at the lower spatial resolution.

[0070] It should be noted that the above-described embodiments can also be modified with regard to the internal configuration of the encoding / decoding branches. For example, intra-prediction modes may not exist, i.e., spatial prediction modes may not be available. Similarly, inter-view and temporal prediction modes may not be available. Furthermore, any additional processing may be optional. On the other hand, an out-of-loop post-processing module may be present at the output of the decoding branch 106, performing, for example, adaptive filtering or other quality improvement techniques and / or the above-mentioned upsampling. Furthermore, no residual transformation may be performed. Rather, the residual may be transmitted in the spatial domain rather than the frequency domain. In a more general sense, the hybrid encoding / decoding designs shown in FIGS. 1 and 3 may be replaced by other encoding / decoding concepts, for example, based on wavelet transforms.

[0071] It should also be noted that the decoder does not necessarily need to include the viewpoint extractor 108. In fact, the extractor 108 may not be present. In that case, the decoder 100 simply reconstructs any number of viewpoints 121 and 122, e.g., one, several, or all. Even if there is no depth / disparity map data for each viewpoint 121 and 122, the viewpoint extractor 108 may perform intermediate viewpoint extraction using disparity vectors associated with corresponding portions of adjacent viewpoints. Using these disparity vectors as auxiliary disparity vectors of a disparity vector field associated with the images of the neighboring viewpoints, the viewpoint extractor 108 may form an intermediate viewpoint image by applying this disparity vector field from such images of the neighboring viewpoints 121 and 122. For example, for a video frame v 2,t Assume that 50% of the portion / segment is cross-view predicted. That is, for 50% of the portion / segment, a disparity vector exists. For the remaining portion, the disparity vectors can be determined by the viewpoint extractor 108 using interpolation / extrapolation in the spatial sense. Temporal interpolation may also be used, using disparity vectors for previously reconstructed portions / segments of frames of video 142. Next, for a video frame v 2,t and / or reference video frame v 1,t However, the disparity vectors may be distorted according to these disparity vectors to obtain intermediate viewpoints. For this purpose, the disparity vectors are scaled according to the intermediate viewpoint position between the viewpoint positions of the first viewpoint 121 and the second viewpoint 122. Details of this procedure will be explained in more detail later.

[0072] By using the above-mentioned option of determining the Wedgelet dividing lines to extend along the edges detected in the reconstructed current frame of the video, gains in coding efficiency can be obtained: as already explained, the above-mentioned Wedgelet dividing line position predictions may be used for each viewpoint, i.e., for all viewpoints or only for a suitable subset of the viewpoints.

[0073] From the above description of FIG. 3, it is clear that the decoder extracts the current frame v of video 141 from the data stream 18. 1,t A decoding branch 106 that reconstructs v,1 and the decoding branch 106 d,1 Decoding branch 106 d,1 is the current frame v from data stream 18. 1,t An edge 72 in the current frame v is detected, and a wedgelet dividing line 70 is determined to extend along the edge 72. 1,t The depth / disparity map d related to 1,t the depth / disparity map d 1,t The segmentation units 66a, 66b, 72a, and 72b are reconstructed in units of the segmentation units, and the depth / disparity map d 1,t In the segmentation, two adjacent segments 72a and 72b are separated from each other by a Wedgelet dividing line 70. The decoder decodes the current frame v of the video. 1,t The depth / disparity map d related to 1,t , or any previously decoded frame v of the video 1,t-1 The depth / disparity map d related to 1,t-1 From a pre-reconstructed segment of , a depth / disparity map d is generated using a separate set of prediction parameters for that segment. 1,t The decoder may be configured to make the Wedgelet dividing line 70 a straight line, and the decoder may further configure the depth / disparity map d 1,t The block-based prior segmentation may be configured to determine a segmentation from the prior segmentation, where block 66c of the prior segmentation is divided along wedgelet dividing line 70, so that two adjacent segments 72a and 72b are wedgelet-shaped segments that together constitute block 66c of the prior segmentation.

[0074] To summarize some of the above-described embodiments, these embodiments enable viewpoint extraction by jointly decoding multi-view video and supplemental data. "Supplemental data" is used below to mean depth / disparity map data. According to these embodiments, multi-view video and supplemental data are embedded in a single compressed representation. The supplemental data may include per-pixel depth maps, disparity data, or 3D wireframes. The extracted viewpoint 102 may differ in viewpoint number and spatial location from the viewpoints 121 and 122 contained in the compressed representation or bitstream 18. The compressed representation 18 is previously generated by an encoder 20, which can also use the supplemental data to improve the encoding of the video data.

[0075] In contrast to state-of-the-art methods, joint decoding is performed, in which the decoding of video and supplementary data may be supported and controlled by common information, such as a common set of motion or disparity vectors, which are used to decode the video and supplementary data. Finally, viewpoints are extracted from the decoded video data, supplementary data, and possibly combined data, with the number and position of the extracted viewpoints being controlled by extraction control in the receiving device.

[0076] Furthermore, the multi-view compression concept described above can be used in conjunction with disparity-based view synthesis, which provides a three-dimensional perception of scene content to the viewer if this content is captured by multiple cameras, such as images 141 and 142. To achieve this, a stereo pair with slightly different viewing directions for each eye must be provided. The shift in both views of the same content at the same time is represented by a disparity vector. Similarly, the shift in content between different times within a sequence is a motion vector, as illustrated in Figure 4 for two views at two times.

[0077] Typically, disparity is estimated directly or as scene depth, either externally provided or recorded using a special sensor or camera. Motion estimation is already performed by standard encoders. If multiple views are jointly encoded, temporal and inter-view motion estimation are processed similarly, so that motion estimation is performed in both directions during encoding. This has already been explained with reference to FIGS. 1 and 2. The estimated inter-view motion vectors are disparity vectors. Such disparity vectors are exemplarily shown in FIG. 2 with reference numerals 82 and 76. Thus, the encoder 20 also implicitly performs disparity estimation, and the disparity vectors are included in the encoded bitstream 18. These vectors can also be used for additional intermediate view synthesis on the decoder side, i.e., in the viewpoint extractor 108.

[0078] Let us assume that pixel p1(x1,y1) at position (x1,y1) in viewpoint 1 and pixel p2(x2,y2) at position (x2,y2) in viewpoint 2 have the same luminance value. p1(x1,y1)=p2(x2,y2) (1) In this case, the positions of two pixels (x1, y1) and (x2, y2) are, for example, a two-dimensional disparity vector from viewpoint 2 to viewpoint 1, that is, the component d x,21 (x2,y2) and component d y,21 (x2, y2) and d 21 They are connected by (x2, y2). Therefore, the following equation holds: (x1,y1)=(x2+d x,21 (x2,y2),y2+d y,21 (x2,y2)) (2) Combining equations (1) and (2) gives the following equation (3). p1(x2+d x,21 (x2,y2),y2+d y,21 (x2,y2))=p2(x2,y2) (3)

[0079] As shown in the bottom right of Figure 5, two points with the same content can be connected by a disparity vector. Adding this vector to the coordinates of p2 gives the position of p1 in image coordinates. If the disparity vector d 21 If (x2,y2) is scaled by a factor K=[0...1], then any intermediate position between positions (x1,y1) and (x2,y2) becomes addressable. In this way, intermediate viewpoints can be generated by shifting the image content of viewpoint 1 and / or viewpoint 2 by the scaled disparity vector. An example of an intermediate viewpoint is shown in Figure 6.

[0080] As described above, a new intermediate viewpoint can be generated for any position between viewpoint 1 and viewpoint 2.

[0081] Furthermore, viewpoint extrapolation can also be achieved by using scaling factors K<0 and K>1 for the disparity.

[0082] These scaling methods can also be applied in the temporal direction, where new frames can be extracted by scaling motion vectors, leading to the generation of higher frame rate video sequences.

[0083] Returning now to the embodiments described with reference to FIGS. 1-3, these embodiments particularly illustrate a parallel decoding structure with multiple decoding branches that process video and supplementary data such as depth maps and include a common information module, i.e., module 110. This module uses spatial information from both signals generated by the encoder. An example of common information is a set of motion vectors or disparity vectors, extracted, for example, during the encoding process of the video data, that can also be used, for example, for depth data. On the decoder side, this common information is used to drive the decoding of the video and depth data, provide the necessary information for each decoding branch, and, optionally, to extract new viewpoints. With this information, all viewpoints required, for example, for an N-view display, can be extracted in parallel from the video data. Common information or coding parameters that should be shared between the individual encoding / decoding branches include: Common motion and disparity vectors, also used for supplementary data, e.g. from video data - A common block partitioning structure, e.g. from video data partitioning, that is also used for supplementary data - Prediction mode Edge and contour data in luminance and / or chrominance information, such as lines within a luminance block, which are used to partition the complementary data into non-rectangular blocks, called wedgelets, which divide a block into two regions by a line with a given angle and position.

[0084] The common information may be used as a prediction from one decoding branch (e.g., for video) to be refined in another branch (e.g., for supplementary data), or vice versa. This common information may specifically include refining motion or disparity vectors, initializing the block structure in the supplementary data with that of the video data, extracting a line from one video block based on luminance or chrominance edge or contour information, and using this line for Wedgelet segmentation line prediction (with the same angle, but possibly a different position in the corresponding depth block while preserving that angle). The common information module may also transmit partially reconstructed data from one decoding branch to another. Finally, data from this module may be sent to a viewpoint extraction module, which extracts all viewpoints required for, for example, one display (which may be 2D, binocular stereoscopic, or N-view autostereoscopic).

[0085] One important point is that if more than one pair of video and depth / supplementary signals are coded / decoded using the coding / decoding scheme described above, then for each time t, the color viewpoint V Color_ 1(t) and V Color_ 2(t) and the corresponding depth data V Depth_ 1(t) and V Depth_ 2(t) together. The above-described embodiment proposes the following: First, the signal V Color_ 1(t) is encoded / decoded using, for example, conventional motion compensated prediction. Then, in a second step, the corresponding depth signal V Depth_ To encode / decode 1(t), we use the encoded / decoded signal V Color_ The information from 1(t) can be reused as described above. Then, V Color_ 1(t) and V Depth_ The accumulated information from 1(t) is V Color_ 2(t) and / or V Depth_2(t) can be further utilized to encode / decode 2(t). In this way, redundancy can be significantly exploited by sharing and reusing common information between different viewpoints and / or depths.

[0086] The decoding and viewpoint extraction arrangement shown in FIG. 3 will now be described with alternative reference to FIG.

[0087] The decoder architecture shown in Figure 7 is based on two parallel conventional video decoding architectures for chrominance and complementary data. In addition, this embodiment includes a common information module, which can send, process, and receive any shared information from and to any module in both decoding architectures. The decoded video and complementary data are finally combined in an image extraction module to extract the required number of views, where the common information from the novel module mentioned above can also be used. The novel decoding modules and view extraction method proposed by the present invention are highlighted by the hatched boxes in Figure 7.

[0088] The decoding process begins by receiving a common compressed representation or bitstream that contains video data and supplementary data, as well as information common to both decoding processes from one or more views, such as motion or disparity vectors, control information, block partition information, prediction modes, contour data, etc.

[0089] First, entropy decoding is applied to the bitstream to extract quantized transform coefficients for the video and supplemental data. These transform coefficients are fed into two separate decoding branches, labeled "Video Data Processing" and "Supplemental Data Processing" and highlighted by dotted boxes in Figure 7. In addition, this entropy decoding also extracts shared or common data and feeds it into a novel common information module.

[0090] After entropy decoding, both decoding branches operate similarly. The received quantized transform coefficients are scaled and an inverse transform is applied to obtain a difference signal. To this, previously decoded data from temporally or spatially neighboring viewpoints is added. The type of information to be added is controlled by special control data. In the case of intra-coded images or supplementary data, intra-frame reconstruction is applied because previous or neighboring information is not available. In the case of inter-coded images or supplementary data, previously decoded data from temporally preceding or neighboring viewpoints is available (see current switch settings in Figure 7). The previously decoded data is shifted by the associated motion vector within the motion compensation block and added to the difference signal to generate an initial frame. If the previously decoded data is from a neighboring viewpoint, the motion data represents disparity data. These initial frames or viewpoints can be further processed by a deblocking filter and possible enhancements, such as edge smoothing, to improve visual quality.

[0091] After this refinement stage, the reconstructed data is fed to a decoded image buffer, which orders the decoded data and outputs decoded images in the correct temporal order for each time instant. The stored data is also used for the next processing cycle and serves as input for scalable motion / disparity vector compensation.

[0092] In addition to the separate decoding of video and supplementary data as described above, a novel common information module is used. This module processes any data common to the video and supplementary data. Examples of common information include shared motion / disparity vectors, block partitioning information, prediction modes, contour data, and control information, as well as common transform coefficients or modes and viewpoint enhancement data. Any data processed in the individual video and supplementary data modules can also be part of the data in the common module. Thus, there may be connections back and forth between the common module and all parts of the individual decoding branches. Furthermore, the common information module may contain enough data so that all video and supplementary data can be decoded using just one of the individual decoding branches and the common module. An example of such a case is a compressed representation where some parts only have video data and all other parts have common video and supplementary data. In that case, video data is decoded in the video decoding branch, while all supplementary data is processed in the common module and output to viewpoint synthesis. Therefore, in this example, a separate supplementary data branch is not used. Furthermore, individual data from each module of the separate decoding branches may be sent back to the common information processing module, for example in the form of partially decoded data, to be used there or transmitted to other decoding branches. Examples of such data include decoded video data such as transform coefficients, motion vectors, modes or settings, etc., which are transmitted to the appropriate supplementary decoding module.

[0093] After decoding, the reconstructed video and supplementary data are transmitted from a separate decoding branch or a common information module to a viewpoint extractor. In the viewpoint extraction module, such as that shown in FIG. 3 at 110, the viewpoints required for the receiving device, e.g., a multi-view display, are extracted. This process is controlled by an intermediate viewpoint extraction control, which sets the required number and position of the viewpoint sequence. An example of viewpoint extraction is viewpoint synthesis. That is, if a new viewpoint is to be synthesized between two original viewpoints 1 and 2, as shown in FIG. 6, the data from viewpoint 1 may first be shifted to a new position. However, this disparity shift will be different for foreground and background objects because this shift is inversely proportional to the depth of the original scene (the distance in front of the camera). In this way, new background areas that were not visible in viewpoint 1 become visible in the synthesized viewpoint. Viewpoint 2 may then be used to fill in this information. Furthermore, spatially neighboring data, such as adjacent background information, may also be used.

[0094] As an example, consider the setup shown in Figure 8. In this example, the decoded data consists of color data V Color 1 and V Color 2 and depth data V Depth 1 and V Depth From this data, the view sequences V D 1. Viewpoint V D 2,...,Viewpoint V D 9, viewpoints for a 9-view display are extracted. The display signals the number and spatial location of the viewpoints via the intermediate viewpoint extraction control. Here, a spatial distance of 0.25 is required for the 9 viewpoints, and the distance between adjacent display viewpoints (e.g., V D 2 and V D3) are four times closer to each other in terms of spatial position and stereoscopic perception compared to the viewpoints in the bitstream. Therefore, the viewpoint extraction factors {K1, K2, ..., K9} are set to {-0.5, -0.25, 0, 0.25, 0.5, 0.75, 1, 1.25, 1.5}. That is, the decoded color viewpoint V Color 1 and V Color 2, at their spatial positions (since K3=0 and K7=1), the viewpoint V D 3 and V D 7. Furthermore, V D 4 and V D 5 and V D 6 means V Color 1 and V Color 2. Finally, V D 1 and V D 2 and V D 8 and V D 9 and V are bitstream pairs, respectively. Color 1 and V Color 2 on each side. Using a set of viewpoint extraction factors, the depth data V Depth 1 and V Depth 2 is converted to pixel-by-pixel displacement information in the viewpoint extraction stage and scaled appropriately to obtain nine differently shifted versions of the decoded color data.

[0095] Although some aspects have been presented in the context of describing an apparatus, it is clear that these aspects also describe a corresponding method, with the blocks or apparatus corresponding to method steps or features of the method steps. Similarly, aspects presented in the context of describing a method step also represent corresponding blocks or items or features of the corresponding apparatus. Some or all of the method steps may be performed by (using) hardware, such as, for example, a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.

[0096] The encoded multi-viewpoint signal of the present invention may be stored in a digital storage medium, or may be transmitted via a wireless transmission medium such as the Internet or a wired transmission medium.

[0097] Depending on certain implementation requirements, embodiments of the present invention may be configurable in hardware or software. This may be implemented using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, flash memory, etc., having electronically readable control signals stored therein and cooperating (or capable of cooperating) with a programmable computer system to perform the methods of the present invention. Thus, the digital storage medium may be computer-readable.

[0098] Some embodiments according to the invention may include a data carrier having electronically readable control signals cooperable with a computer system programmable to carry out one of the methods described above.

[0099] Generally, embodiments of the present invention may be configured as a computer program product, the program code being operative to perform one of the methods of the present invention when the computer program product is run on a computer, the program code being stored on, for example, a machine readable carrier.

[0100] Other embodiments of the invention comprise the computer program stored on a machine readable carrier for performing one of the methods described above.

[0101] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the above described methods, when the computer program runs on a computer.

[0102] Another embodiment of the invention is a data carrier (or digital storage medium or computer readable medium) comprising a computer program stored thereon for performing one of the methods described above, the data carrier, digital storage medium or stored medium being typically tangible and / or non-transitory.

[0103] Another embodiment of the invention is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described above, which data stream or sequence of signals may be adapted to be transmitted over a data communication connection, for example via the Internet.

[0104] Another embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described above.

[0105] Another embodiment comprises a computer having installed thereon the computer program for performing one of the methods described above.

[0106] Other embodiments of the invention include a device or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the above-described methods to a receiver, which may be, for example, a computer, a mobile device, a memory device, or the like. The device or system may include a file server for transmitting the computer program to the receiver.

[0107] In some embodiments, a programmable logic device (such as a field programmable gate array) may be used to perform some or all of the functions of the methods described above. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described above. In general, such methods may be preferably performed by any hardware apparatus.

[0108] The above-described embodiments are merely illustrative of the principles of the present invention. Modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. Therefore, the present invention is not to be limited by the specific details presented herein for purposes of illustration and description of the embodiments, but should be limited only by the scope of the appended claims. [Claim 1] means (106) for reconstructing the first view (121) of the multi-view signal (10) from the data stream (18) by predicting a current portion of the first view (121) from a first pre-reconstructed portion of the multi-view signal (10) according to first coding parameters (461, 481, 501) obtained from the data stream (18) and correcting a prediction error of the prediction of the current portion of the first view (121) using first correction data (421, 422) included in the data stream (18); v,1 ,106 d,1 ), wherein the first pre-reconstructed portion was reconstructed from the data stream (18) by a decoder prior to the reconstruction of the current portion of the first viewpoint (121). v,1 ,106 d,1 )and, means for at least partially adapting or predicting second coding parameters from said first coding parameters; means (106) for reconstructing the second view (122) of the multi-view signal (10) from the data stream (18) by predicting a current portion of a second view from a second pre-reconstructed portion of the multi-view signal (10) according to the second coding parameters and correcting a prediction error of the prediction of the current portion of the second view (122) using second correction data (423, 424) included in the data stream (18); v,2 ,106 d,2 ), wherein the second pre-reconstructed portion was reconstructed from the data stream (18) by the decoder prior to the reconstruction of the current portion of the second view (122). v,2 ,106 d,2 )and, A decoder comprising: [Claim 2] 2. The decoder of claim 1, The decoder further comprises means (108) for extracting an intermediate viewpoint (102) from the first viewpoint and the second viewpoint. [Claim 3] 3. A decoder according to claim 1 or 2, A decoder in which each of the first viewpoint (121) and the second viewpoint (122) comprises video (141, 142) captured from a separate camera position and associated depth / disparity map data (161, 162). [Claim 4] 4. The decoder of claim 3, means (106) for reconstructing the image (141) of the first viewpoint (121) from the data stream (18) by predicting a current portion of the image (141) of the first viewpoint (121) from a third pre-reconstructed portion of the multi-view signal (10) according to a first portion (461, 481, 501) of the first coding parameters and correcting a prediction error of the prediction of the current portion of the image (141) of the first viewpoint (121) using a first subset (421) of first correction data (421, 422) included in the data stream (18); v,1 ), wherein the third pre-reconstructed portion was reconstructed from the data stream (18) by the decoder prior to the reconstruction of the current portion of the image (141) of the first viewpoint (121). v,1 )and, means for at least partially adapting or predicting a second portion of the first coding parameters from the first portion of the first coding parameters; means (106) for reconstructing the depth / disparity map data (161) for the first viewpoint (121) from the data stream (18) by predicting a current portion of the depth / disparity map data (161) for the first viewpoint (121) from a fourth pre-reconstructed portion of the multi-view signal (10) according to the second portion of the first coding parameters and correcting a prediction error of the prediction of the current portion of the depth / disparity map data (161) for the first viewpoint (121) using the second subset (422) of first correction data; d,1 ), wherein the fourth pre-reconstructed portion was reconstructed from the data stream (18) by the decoder prior to the reconstruction of the current portion of the depth / disparity map data (161) for the first viewpoint (121). d,1 )and, A decoder comprising: [Claim 5] 5. The decoder of claim 4, a decoder, characterized in that the first encoding parameters define a segmentation of frames of the video (141) of the first viewpoint (121), and the decoder, when reconstructing the depth / disparity map data (161) of the first viewpoint (121), uses the segmentation of frames of the video (141) of the first viewpoint (121) as a pre-segmentation of the depth / disparity map data (161) of the first viewpoint (121). [Claim 6] 5. The decoder of claim 4, The decoder decodes the current frame (v 1,t ) using the reconstructed portion of the Wedgelet dividing line (70) to predict the position of the Wedgelet dividing line (70); When the depth / disparity map data (161) of the first viewpoint is reconstructed from the data stream (18), the current frame (v) of the image of the first viewpoint is reconstructed so as to coincide with the Wedgelet dividing line. 1,t ) of the current portion (d 1,t) boundaries. [Claim 7] 7. The decoder according to claim 4, a first pre-reconstructed portion (v) of the image of the first viewpoint (121); 1,t-1 ) to predict the current portion (60a) of the image of the first viewpoint (121) from the first pre-reconstructed portion (v 1,t-1 ) was reconstructed from the data stream (18) by the decoder prior to reconstruction of the current portion of the video of the first perspective; a first pre-reconstructed portion (d 1,t-1 ) to predict a current portion (66a) of the depth / disparity map data of the first viewpoint (121) from the first pre-reconstructed portion (d 1,t-1 ) was reconstructed from the data stream (18) by the decoder prior to reconstruction of the current portion (66a) of the depth / disparity map data for the first viewpoint (121); A decoder comprising: [Claim 8] A decoder according to any one of claims 3 to 7, means for at least partially adapting or predicting a first portion of the second coding parameters from the first coding parameters; means (106) for reconstructing the image (142) of the second view (122) from the data stream (18) by predicting a current portion of the image (142) of the second view (122) from a fifth pre-reconstructed portion of the multi-view signal (10) according to a first portion of the second coding parameters and correcting a prediction error of the prediction of the current portion of the image (142) of the second view (122) using a first subset (423) of second correction data included in the data stream (18); v,2), wherein the fifth pre-reconstructed portion was reconstructed from the data stream (18) by the decoder prior to the reconstruction of the current portion of the image (142) of the second perspective (122). v,2 )and, means for at least partly adapting or predicting a second portion of the second coding parameters from the first coding parameters and / or from the first portion of the second coding parameters; means (106) for reconstructing the depth / disparity map data (162) for the second viewpoint (122) from the data stream (18) by predicting a current portion of the depth / disparity map data (162) for the second viewpoint (122) from a sixth pre-reconstructed portion of the multi-view signal (10) according to the second portion of the second coding parameters and correcting a prediction error of the prediction of the current portion of the depth / disparity map data (162) for the second viewpoint (122) using the second subset (424) of second correction data; d,2 ), wherein the sixth pre-reconstructed portion was reconstructed from the data stream (18) by the decoder prior to the reconstruction of the current portion of the depth / disparity map data (162) for the second viewpoint (122). d,2 )and, A decoder comprising: [Claim 9] 9. The decoder of claim 8, The decoder Reconstructing (106) the depth / disparity map data (161) of the first viewpoint (121) from the data stream (18). d,1 ) using a first Wedgelet dividing line (70) in the depth / parallax map data (161) of the first viewpoint, using the first Wedgelet division line (70) as a predicted value for a second Wedgelet division line in the depth / disparity map data (162) of the second viewpoint, or adopting the first Wedgelet division line (70) as a second Wedgelet division line; When the depth / disparity map data (162) of the second viewpoint is reconstructed from the data stream (18), the current frame (v) of the image of the second viewpoint is reconstructed so as to coincide with the second Wedgelet dividing line. 2,t ) of the depth / disparity map data (162) of the second viewpoint associated with 2,t ) to set a boundary of the current portion of the [Claim 10] 10. A decoder according to claim 8 or 9, a first pre-reconstructed portion (v) of the image (142) of the second viewpoint (122); 2,t-1 ) or from a second pre-reconstructed portion (v) of the image (141) of the first viewpoint (121). 1,t ) means for predicting a current portion (74a; 74b) of the image (142) of the second viewpoint (122) from the first pre-reconstructed portion (v 2,t-1 ) was reconstructed from the data stream (18) by the decoder prior to the reconstruction of the current portion of the image (142) of the second viewpoint (122), and the second pre-reconstructed portion (v 1,t ) was reconstructed from the data stream (18) by the decoder prior to reconstruction of the current portion of the image (142) of the second perspective (122); a first pre-reconstructed portion (d 2,t-1 ) or from a second pre-reconstructed portion (d 1,t ) to predict a current portion (80a; 80b) of the depth / disparity map data (162) of the second viewpoint (122), from the first pre-reconstructed portion (d 2,t-1 ) was reconstructed from the data stream (18) by the decoder prior to the reconstruction of the current portion of the depth / disparity map data (162) for the second viewpoint (122), and the second pre-reconstructed portion (d 1,t) was reconstructed from the data stream (18) by the decoder prior to reconstruction of the current portion of the depth / disparity map data for the second viewpoint; A decoder comprising: [Claim 11] A decoder according to any one of claims 8 to 10, The current frame (v 2,t means for detecting an edge (72) in the image plane and determining a wedgelet dividing line (70) extending along the edge (72); When the depth / disparity map data (162) of the second viewpoint is reconstructed from the data stream (18), the current frame (v) of the image of the second viewpoint is reconstructed so as to coincide with the Wedgelet dividing line. 2,t ) of the current portion (d 2,t and means for setting the boundaries of [Claim 12] 12. The decoder of claim 11, a decoder, characterized in that when reconstructing the depth / disparity map data (162) of the second viewpoint from the data stream (18), prediction is performed for each segment within the unit of the segment to which the current portion belongs, using an individual set of prediction parameters for the segment. [Claim 13] 13. A decoder according to claim 11 or 12, a decoder characterized in that the Wedgelet division line (70) is a straight line, the decoder divides a block (66c) of the previous segmentation of the depth / disparity map data (162) of the second viewpoint along the Wedgelet division line (70), and two adjacent segments (72a, 72b) are Wedgelet-shaped segments that together form the block (66c) of the previous segmentation. [Claim 14] A decoder according to any one of claims 1 to 13, The decoder, wherein the first and second encoding parameters are first and second prediction parameters that control prediction of the current portion of the first viewpoint and prediction of the current portion of the second viewpoint, respectively. [Claim 15] A decoder according to any one of claims 1 to 14, A decoder, wherein the current portion is a segment of a segmentation of a frame of the video of the first perspective and the second perspective, respectively. [Claim 16] A decoder according to any one of claims 1 to 15, The decoder reconstructing the first view (121) of the multi-view signal (10) from the data stream by predicting and correcting a current portion of the first view (121) at a first spatial resolution; a decoder for reconstructing the second view (122) of the multi-view signal (10) from the data stream by predicting and correcting a current portion of the second view (122) at a second spatial resolution lower than the first spatial resolution, and then upsampling the reconstructed current portion of the second view (122) from the second spatial resolution to the first spatial resolution. [Claim 17] 17. The decoder of claim 16, 10. A decoder comprising: a decoder for scaling the first coding parameters according to a ratio between the first spatial resolution and the second spatial resolution when at least partially employing or predicting the second coding parameters from the first coding parameters. [Claim 18] means for encoding a first viewpoint of a multi-view signal into a data stream by performing the following steps (1) to (3); (1) determining first encoding parameters; (2) predicting the current portion of the first view from a first pre-encoded portion of the multi-view signal according to the first encoding parameters, and determining a prediction error of the prediction of the current portion of the first view to obtain first correction data, wherein the first pre-encoded portion is encoded into the data stream by the encoder prior to encoding of the current portion of the first view. (3) inserting the first encoding parameters and the first correction data into the data stream; means for encoding a second viewpoint of the multi-view signal into the data stream by performing the following steps (4) to (6); (4) determining the second encoding parameters by adapting or predicting the second encoding parameters from the first encoding parameters; (5) predicting the current portion of the second view from a second pre-encoded portion of the multi-view signal according to the second encoding parameters, and determining a prediction error of the prediction of the current portion of the second view to obtain second correction data included in the data stream, the second pre-encoded portion being encoded into the data stream by the encoder prior to encoding the current portion of the second view. (6) inserting the second correction data into the data stream; An encoder including: [Claim 19] a data stream including a first portion and a second portion, the first portion is a first view of a multi-view signal into which the first view is encoded, the first portion includes first correction data and first encoding parameters, according to which the current view of the first view is predictable from a first pre-encoded portion of the multi-view signal, the first pre-encoded portion being encoded into the data stream prior to encoding of the current view of the first view, and a prediction error of the prediction of the current view of the first view being correctable using the first correction data; A data stream in which the second portion is encoded with a second view of a multi-view signal, the second portion includes second correction data, the current portion of the second view is predictable from a second pre-encoded portion of the multi-view signal according to second encoding parameters that are predictable or adoptable from the first encoding parameters, the second pre-encoded portion being encoded into the data stream prior to encoding of the current portion of the second view, and a prediction error of the prediction of the current portion of the second view being correctable using the second correction data. [Claim 20] a step (106) of reconstructing the first view (121) of the multi-view signal (10) from the data stream (18) by predicting a current portion of the first view (121) from a first pre-reconstructed portion of the multi-view signal (10) according to first coding parameters (461, 481, 501) obtained from the data stream (18) and correcting a prediction error of the prediction of the current portion of the first view (121) using first correction data (421, 422) included in the data stream (18); v,1 ,106 d,1 ), wherein the first pre-reconstructed portion was reconstructed from the data stream (18) by a decoder prior to the reconstruction of the current portion of the first view (12), step (106 v,1 ,106 d,1 )and, adapting or predicting second coding parameters at least in part from said first coding parameters; a step (106) of reconstructing the second view (122) of the multi-view signal (10) from the data stream (18) by predicting a current portion of a second view from a second pre-reconstructed portion of the multi-view signal (10) according to the second coding parameters and correcting a prediction error of the prediction of the current portion of the second view (122) using second correction data (423, 424) included in the data stream (18). v,2 ,106 d,2), wherein the second pre-reconstructed portion was reconstructed from the data stream (18) by the decoder prior to the reconstruction of the current portion of the second view (122), step (106 v,2 ,106 d,2 )and, A decoding method comprising: [Claim 21] encoding a first viewpoint of a multi-view signal into a data stream by performing the following steps (1) to (3); (1) determining first encoding parameters; (2) predicting the current portion of the first view from a first pre-encoded portion of the multi-view signal according to the first encoding parameters, and determining a prediction error of the prediction of the current portion of the first view to obtain first correction data, wherein the first pre-encoded portion is encoded into the data stream by the encoder prior to encoding of the current portion of the first view. (3) inserting the first encoding parameters and the first correction data into the data stream; encoding a second view of the multi-view signal into the data stream by performing the following steps (4) to (6); (4) determining the second encoding parameters by adapting or predicting the second encoding parameters from the first encoding parameters; (5) predicting the current portion of the second view from a second pre-encoded portion of the multi-view signal according to the second encoding parameters, and determining a prediction error of the prediction of the current portion of the second view to obtain second correction data included in the data stream, the second pre-encoded portion being encoded into the data stream by the encoder prior to encoding the current portion of the second view. (6) inserting the second correction data into the data stream; An encoding method including: [Claim 22] 22. A computer program having a program code for performing the method according to claim 20 or 21, when the computer program runs on a computer.

Claims

1. 1. A decoder for decoding an encoded video data stream to generate multi-view video, comprising: The decoder includes a view decoder, the view decoder comprising: using a processor to extract from the data stream first information relating to a first coded block in a first view of a multi-view video signal, wherein the first information may indicate whether motion parameters of the first coded block (a) should be read from the data stream independently of other coding parameters, or (b) should reuse one or more coding parameters of a second coded block disposed in a second view of the multi-view video signal; extracting, using the processor, motion parameters of the first coded block from the data stream if the first information indicates that the motion parameters of the first coded block should be read from the data stream independently of other coding parameters; receiving, using the processor, one or more coding parameters including motion parameters of the second coding block when the first information indicates that the motion parameters of the first coding block should be reused; predicting, using the processor, motion parameters of the first coding block based on the motion parameters of the second coding block; and extracting prediction error data related to the motion parameters of the first coding block; using the processor to generate a prediction of the first coding block based on (i) the extracted or predicted motion parameters of the first coding block and (ii) the prediction error data if the first information indicates that the motion parameters of the first coding block should reuse one or more coding parameters; Using the processor, obtain residual data associated with the first coded block from the data stream; and configured, using the processor, to reconstruct the first coded block using the prediction of the first coded block and the residual data to generate a portion of a video frame at a first view of the multi-view video. Decoder.

2. The decoder of claim 1 , wherein the first view and the second view each include different types of information components.

3. The decoder of claim 2 , wherein the different types of information components include an image and a depth map corresponding to the image.

4. The decoder of claim 3 , wherein the first coded block contains video data and is reconstructed based on a first subset of one or more coding parameters from the second coded block.

5. The decoder of claim 4 , wherein the first coded block includes depth data and is reconstructed based on a second subset of one or more coding parameters from the second coded block.

6. The view decoder further comprises: and configured to reconstruct a first depth coding block in the first view based on a first edge associated with the first depth coding block, the decoder further comprising: a further view decoder, the further view decoder comprising: predicting a second edge associated with a second depth coding block of the second view based on the first edge; and configured to reconstruct the second depth coding block of the second viewpoint based on the second edge.

4. The decoder of claim 3.

7. The decoder of claim 1 , wherein the first coded blocks are reconstructed according to a first spatial resolution and the second coded blocks are reconstructed according to a second spatial resolution.

8. The decoder of claim 1 , wherein the decoder is further configured to generate an intermediate coded block based on the first coded block and the second coded block.

9. The decoder of claim 1 , wherein the decoder is further configured to generate an intermediate view based on the first view and the second view.

10. The first information may further indicate that motion parameters of the first coding block should be predicted from motion parameters of a pre-reconstructed part of the first view, in which case the view decoder:

2. The decoder of claim 1, configured to: use the processor to obtain motion parameters of the pre-reconstructed portion; and use the processor to predict motion parameters of the first coding block based on the motion parameters of the pre-reconstructed portion.

11. The decoder of claim 1 , wherein the one or more coding parameters include a projection parameter and a depth parameter, both of which relate to a second coded block in the second perspective.

12. 1. An encoder for encoding multi-view video into a data stream, the encoder comprising: a view encoder, the view encoder comprising: configured to use a processor to encode into the data stream first information related to a first coding block in a first view of a multi-view video signal, wherein the first information may indicate whether motion parameters of the first coding block (a) should be read from the data stream independently of other coding parameters, or (b) should reuse one or more coding parameters of a second coding block located in a second view of the multi-view video signal; using the processor to encode motion parameters of the first coded block into the data stream if the first information indicates that the motion parameters of the first coded block should be read from the data stream independently of other coding parameters; encoding, using the processor, one or more coding parameters including motion parameters of a second coding block in the second view and prediction error data related to the motion parameters of the first coding block into the data stream when the first information indicates that one or more coding parameters from a second coding block in the second view should be reused for the motion parameters of the first coding block; using the processor to generate a prediction of the first coded block based at least on motion parameters of the first coded block; using the processor to determine residual data associated with the first coded block based on a difference between the first coded block and a prediction of the first coded block; and using the processor to encode the residual data associated with the first coded block into the data stream, the first coded block being reconstructed using a prediction of the first coded block, the prediction error data, and the residual data to generate a portion of a video frame at a first view of the multi-view video; encoder.

13. The first information may further indicate that motion parameters of the first coding block should be predicted from motion parameters of a pre-reconstructed part of the first view, in which case the view encoder: The encoder of claim 12 , configured to use the processor to encode motion parameters of the pre-reconstructed portion into the data stream.

14. The encoder of claim 12 , wherein the first view and the second view each include different types of information components.

15. The encoder of claim 14 , wherein the different types of information components include a video and a depth map corresponding to the video.

16. 16. The encoder of claim 15, wherein the first encoded block contains video data and is reconstructed based on a first subset of one or more encoding parameters from the second encoded block.

17. The encoder of claim 16 , wherein the first coded block includes depth data and is reconstructed based on a second subset of one or more coding parameters from the second coded block.

18. The encoder of claim 12 , wherein the encoder is further configured to encode an intermediate view based on the first view and the second view.

19. 1. A non-transitory computer-readable storage medium configured to store information including a data stream representing encoded multi-view video, the data stream comprising: coded first information associated with a first coded block in a first view of a multi-view video signal, the first coded block representing a portion of a video frame in the first view, the first information being capable of indicating whether motion parameters of the first coded block (a) should be read from the data stream independently of other coding parameters, or (b) should reuse one or more coding parameters of a second coded block located in a second view of the multi-view video signal; - coded motion parameters of the first coded block if the first information indicates that the motion parameters of the first coded block should be read from the data stream independently of other coding parameters; When the first information indicates that the motion parameters of the first coding block should reuse one or more coding parameters from a second coding block in the second view, one or more coded coding parameters including motion parameters of a second coding block in the second view and prediction error data related to the motion parameters of the first coding block; coded residual data related to the first coded block based on a difference between the first coded block and a prediction of the first coded block, the prediction of the first coded block being determined using motion parameters of the first coded block; and Including, the first coding block is reconstructed using a prediction of the first coding block, the prediction error data, and the residual data to generate a portion of a video frame at a first viewpoint of the multi-view video. storage medium.

20. The first information may further indicate that motion parameters of the first coding block should be predicted from motion parameters of a pre-reconstructed part of the first viewpoint, in which case the data stream may include: The non-transitory computer-readable storage medium of claim 19 , further comprising encoded motion parameters of the pre-reconstructed portion.