Method and apparatus for processing multi-view video data

By flexibly selecting the acquisition mode of the composite data items, the problems of high transmission rate and redundant information in depth maps in immersive videos are solved, enabling efficient encoding and decoding on terminal devices and improving the composite quality of intermediate views.

CN115104312BActive Publication Date: 2026-04-14ORANGE SA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-04
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In existing immersive video processing solutions, the high transmission rate and redundant information of depth maps lead to increased encoding and decoding complexity, making it difficult to deploy efficiently on terminal devices such as smartphones. Furthermore, the lack of interaction between depth estimation and synthesis algorithms affects the synthesis quality of intermediate views.

Method used

A flexible synthetic data item acquisition mode is adopted, which allows the acquisition method of synthetic data items to be selected on the encoder side or the decoder side. Synthetic data items are encoded in the data stream through the first acquisition mode, or synthetic data items are estimated from the reconstructed image on the decoder side, which reduces the encoding cost and improves the synthesis quality.

Benefits of technology

It effectively reduces the encoding cost of multi-view video, improves the synthesis quality of intermediate views, adapts to the processing capabilities of different terminal devices, reduces redundant information transmission, and enhances the display effect of immersive video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115104312B_ABST
    Figure CN115104312B_ABST
Patent Text Reader

Abstract

The invention relates to a method for processing multi-view video data, in which method, for at least one block of an image of a view encoded in an encoded data stream representing a multi-view video, at least one item of information is obtained. The item of information specifies a mode for obtaining at least one item of synthesis data from a first obtaining mode and a second obtaining mode, the item of synthesis data being used for synthesizing at least one image of an intermediate view of the multi-view video, the intermediate view not being encoded in the encoded data stream. The first obtaining mode involves decoding from the encoded data stream at least one item of information representing the at least one item of synthesis data, the second obtaining mode involving obtaining the at least one item of synthesis data from at least the reconstructed encoded image. At least a portion of the image of the intermediate view is synthesized from at least the reconstructed encoded image and the obtained at least one item of synthesis data according to the specified obtaining method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to immersive videos representing scenes captured by one or more cameras, including videos used for virtual reality and free navigation. More specifically, this invention relates to the processing (encoding, decoding, and synthesis of intermediate views) of data from such videos. Background Technology

[0002] Immersive video allows viewers to view a scene from any viewpoint, even from a point not yet captured by the camera. A typical capture system is a set of cameras capturing the scene, with several cameras located outside the scene, or dispersed cameras built on a spherical platform. The video is typically displayed via a virtual reality headset (also known as a head-mounted device, or HMD), but can also be displayed on a 2D screen using additional systems for user interaction.

[0003] Free navigation within a scene requires that every user action be properly managed to avoid motion sickness. This motion is typically captured correctly by the display device (e.g., HMD headset). However, providing the correct pixels for display regardless of the user's movement (rotation or translation) is currently problematic. This necessitates the ability to capture multiple views and generate additional virtual (synthesized) views, calculated based on the decoded captured views and associated depth. The number of views to be transmitted varies depending on the use case. However, the number of views to be transmitted and the associated data volume are typically high. Therefore, view transmission is a necessary aspect of immersive video applications. Consequently, the bit rate of the information to be transmitted must be reduced as much as possible without compromising the synthetic quality of intermediate views.

[0004] In typical immersive video processing schemes, the view is physically captured or generated by a computer. In some cases, depth is also captured using dedicated sensors. However, the quality of this depth information is often poor and hinders the effective synthesis of intermediate viewpoints.

[0005] Depth maps can also be calculated from texture images of captured video. Many depth estimation algorithms exist and are used in the existing technology.

[0006] like Figure 1 As shown, the texture image and estimated depth information are encoded and sent to the user's display device. Figure 1 This illustrates examples including, for example, texture information T. x0y0 and T x1y0 An immersive video processing scheme using two captured views. With each view T x0y0 and T x1y0 Associated depth information D x0y0 and D x1y0Estimated by the estimation module FE. For example, depth information D is obtained through depth estimation software (depth estimation reference software or DERS). x0y0 and D x1y0 Then, for example, using an MV-HEVC encoder to view T x0y0 and T x1y0 And the obtained depth information D x0y0 and D x1y0 Encode (CODEC). On the client side, view T* x0y0 and T* x1y0 And the associated depth D* for each view x0y0 and D* x1y0 The decoded data is used by the synthesis algorithm (SYNTHESIS) to compute intermediate views, for example, the intermediate view S here. x0y0 and S x1y0 For example, VSRS (View Composition Reference Software) software can be used as a view composition algorithm.

[0007] Various challenges arise when calculating depth maps before encoding and transmitting immersive video data. Specifically, the transmission rates associated with various views are high. In particular, although depth maps are generally cheaper than textures, they still constitute a significant portion of the bitstream (15% to 30% of the total).

[0008] Furthermore, while a complete depth map is generated and sent, not all parts of the depth map are necessarily useful on the client side. In fact, views can contain redundant information, making some parts of the depth map unnecessary. Additionally, in some cases, viewers may only require a specific viewpoint. Without a feedback channel between the client and the server providing the encoded immersive video, the server-side depth estimator is unaware of these specific viewpoints.

[0009] Computing depth information on the server side avoids any interaction between the depth estimator and the synthesis algorithm. For example, if the depth estimator wants to inform the synthesis algorithm that it cannot correctly find the depth of a specific region, it must transmit that information in a binary stream, most likely in the form of a binary graph.

[0010] Furthermore, the configuration of the encoder that encodes the depth map is not obvious in order to achieve the best trade-off between synthesis quality and encoding cost of depth map transmission.

[0011] Finally, the decoder has to process a large number of pixels when texture and depth maps are encoded, transmitted, and decoded. This could slow down the deployment of immersive video processing solutions on devices such as smartphones.

[0012] Therefore, it is necessary to improve existing technologies. Summary of the Invention

[0013] This invention improves upon existing technology. For this purpose, the present invention relates to a method for processing multi-view video data, the method comprising:

[0014] - For at least one block of images of views encoded in an encoded data stream representing multi-view video, obtain at least one information item of a mode specifying a first acquisition mode and a second acquisition mode for obtaining at least one composite data item.

[0015] The at least one synthetic data item is used to synthesize at least one image of an intermediate view of the multi-view video, the intermediate view being not encoded in the encoded data stream.

[0016] The first acquisition mode corresponds to decoding at least one information item representing the at least one synthetic data item from the encoded data stream, and the second acquisition mode corresponds to obtaining the at least one synthetic data item from at least the reconstructed encoded image.

[0017] - Obtain the at least one synthetic data item according to the acquisition pattern specified by the at least one information item.

[0018] - Synthesize at least a portion of the image of the intermediate view from at least the reconstructed coded image and the obtained at least one synthetic data item.

[0019] By allowing the selection of the optimal mode for obtaining each synthetic data item, such as according to the encoding cost / quality of the synthetic data item or depending on the tools available on the decoder side and / or encoder side, the present invention flexibly utilizes various modes for obtaining synthetic data. This selection is flexible because it can be done at the block, image, view, or video level. Therefore, the granularity level of the mode used to obtain synthetic data can be adjusted depending on, for example, the content of multi-view video or the tools available on the client / decoder side.

[0020] According to the first acquisition mode, the synthetic data item is determined on the encoder side, encoded in the data stream, and sent to the decoder. According to this first acquisition mode, the quality of the synthetic data item can be prioritized because it is determined from, for example, the unencoded original image. The synthetic data item is unaffected by encoded artifacts of the decoded texture during its estimation.

[0021] According to the second acquisition mode, the synthesized data items are determined on the decoder side. Under this second acquisition mode, the data necessary for synthesizing the intermediate view is obtained from the decoded and reconstructed view that has been sent to the decoder. Such synthesized data can be obtained at the decoder or by a decoder-independent module that takes the decoded and reconstructed view as input. This second acquisition mode reduces the encoding cost of multi-view video data and makes decoding multi-view video easier because the decoder no longer needs to decode the data used for intermediate view synthesis.

[0022] This invention also improves the quality of intermediate view synthesis. In practice, in some cases, the synthesized data items estimated at the decoder may be more suitable for view synthesis than the encoded synthesized data items, for example, when different estimators are available on the client and server sides. In other cases, determining the synthesized data items at the encoder may be more appropriate, for example, when the decoded texture has compression artifacts, or when the texture does not include sufficient redundant information for estimating the synthesized data on the client side.

[0023] According to a specific embodiment of the present invention, the at least one synthetic data item corresponds to at least a portion of the depth map.

[0024] According to another specific embodiment of the invention, at least one information item specifying a pattern for obtaining a synthesized data item is obtained by decoding syntax elements. According to this specific embodiment of the invention, the information item is encoded in a data stream.

[0025] According to another specific embodiment of the invention, the at least one information item specifying the mode for obtaining the synthetic data item is obtained from at least one encoded data item of the reconstructed encoded image. According to this specific embodiment of the invention, the information item is not directly encoded in the data stream, but derived from the encoded data of the image in the data stream. The process for deriving the synthetic data item is the same at both the encoder and decoder.

[0026] According to another specific embodiment of the invention, an acquisition mode is selected from a first acquisition mode and a second acquisition mode, depending on the value of the quantization parameter used to encode at least the block.

[0027] According to another specific embodiment of the present invention, the method further includes, when the at least one information item specifies that a synthetic data item is obtained according to a second acquisition mode:

[0028] - Decode at least one control parameter from the encoded data stream.

[0029] - When the synthesized data item is obtained according to the second acquisition mode, the control parameters are applied.

[0030] This particular embodiment of the invention makes it possible to control the method used to obtain synthetic data items, such as controlling features of the depth estimator, like the size or precision of the search window. The control parameters can also specify which depth estimator and / or its parameters to use, or to initialize the depth map of the estimator.

[0031] The present invention also relates to an apparatus for processing multi-view video data, including a processor configured to:

[0032] - For at least one block of images of views encoded in an encoded data stream representing multi-view video, obtain at least one information item of a mode specifying a first acquisition mode and a second acquisition mode for obtaining at least one composite data item.

[0033] The at least one synthetic data item is used to synthesize at least one image of an intermediate view of the multi-view video, the intermediate view being not encoded in the encoded data stream.

[0034] The first acquisition mode corresponds to decoding at least one information item representing the at least one synthetic data item from the encoded data stream, and the second acquisition mode corresponds to obtaining the at least one synthetic data item from at least the reconstructed encoded image.

[0035] - Obtain the at least one synthetic data item according to the acquisition pattern specified by the at least one information item.

[0036] - Synthesize at least a portion of the image of the intermediate view from at least the reconstructed coded image and the obtained at least one synthetic data item.

[0037] According to a specific embodiment of the present invention, the device for processing multi-view video data is included in a terminal.

[0038] The present invention also relates to a method for encoding multi-view video data, the method comprising:

[0039] - For at least one block of an image representing a view in a coded data stream of multi-view video, determine at least one information item specifying a mode for obtaining at least one composite data item, which is one of a first acquisition mode and a second acquisition mode.

[0040] The at least one synthetic data item is used to synthesize at least one image of an intermediate view of the multi-view video, the intermediate view being not encoded in the encoded data stream.

[0041] The first acquisition mode corresponds to decoding at least one information item representing the at least one synthetic data item from the encoded data stream, and the second acquisition mode corresponds to obtaining the at least one synthetic data item from at least the reconstructed encoded image.

[0042] - The image is encoded in the encoded data stream.

[0043] According to a specific embodiment of the invention, the encoding method includes encoding in a data stream a syntax element associated with the information item, the information item specifying a pattern for obtaining a composite data item.

[0044] According to a specific embodiment of the present invention, when the information item specifies that a synthetic data item is obtained according to a second acquisition mode, the encoding method further includes:

[0045] - Encode at least one control parameter in the encoded data stream, which will be applied when the synthesized data item is obtained according to the second acquisition mode.

[0046] The present invention also relates to a multi-view video data encoding device, the device comprising a processor and a memory, configured to:

[0047] - For at least one block of an image representing a view in a coded data stream of multi-view video, determine at least one information item specifying a mode for obtaining at least one composite data item, which is one of a first acquisition mode and a second acquisition mode.

[0048] The at least one synthetic data item is used to synthesize at least one image of an intermediate view of the multi-view video, the intermediate view being not encoded in the encoded data stream.

[0049] The first acquisition mode corresponds to decoding at least one information item representing the at least one synthetic data item from the encoded data stream, and the second acquisition mode corresponds to obtaining the at least one synthetic data item from at least the reconstructed encoded image.

[0050] - The image is encoded in the encoded data stream.

[0051] The method for processing multi-view video data according to the present invention can be implemented in various ways, particularly in wired or software form. According to a specific embodiment of the invention, the method for processing multi-view video data is implemented by a computer program. The invention also relates to a computer program comprising instructions, which, when executed by a processor, are used to implement the method for processing multi-view video data according to any of the foregoing specific embodiments. Such a program can be used in any programming language. It can be downloaded from a communication network and / or recorded on a computer-readable medium.

[0052] The program can use any programming language and can be in the form of source code, object code, or intermediate code between source code and object code (e.g., partially compiled), or any other desired form.

[0053] The present invention also relates to a computer-readable storage medium or data medium comprising instructions of a computer program as described above. The aforementioned recording medium can be any entity or device capable of storing a program. For example, the medium may include storage devices such as ROMs (e.g., CD-ROMs or microelectronic circuit ROMs), USB flash drives, or magnetic recording devices (e.g., hard disk drives). Alternatively, the recording medium may correspond to a transmissible medium, such as electrical or optical signals that can be carried via cables or optical fibers, by radio, or by other means. Programs according to the present invention can be downloaded, in particular, over Internet-type networks.

[0054] Alternatively, the recording medium may correspond to an integrated circuit in which a program is embedded, which is suitable for executing or used to execute the method in question. Attached Figure Description

[0055] Other features and advantages of the invention will become clearer from the following description and accompanying drawings of specific embodiments, which are provided as simple illustrative and non-limiting examples, wherein:

[0056] [ Figure 1 ] Figure 1 A multi-view video data processing scheme based on existing technology is shown.

[0057] [ Figure 2 ] Figure 2 A multi-view video data processing scheme according to a specific embodiment of the present invention is illustrated.

[0058] [ Figure 3A ] Figure 3A The steps of a method for processing multi-view video data according to a specific embodiment of the present invention are shown.

[0059] [ Figure 3B ] Figure 3B The steps of a method for processing multi-view video data according to another specific embodiment of the present invention are shown.

[0060] [ Figure 4A ] Figure 4A The steps of a multi-view video encoding method according to a specific embodiment of the present invention are shown.

[0061] [ Figure 4B ] Figure 4B The steps of a multi-view video encoding method according to a specific embodiment of the present invention are shown.

[0062] [ Figure 5 ] Figure 5 An example of a multi-view video data processing scheme according to a specific embodiment of the present invention is shown.

[0063] [ Figure 6A ] Figure 6A The texture matrix of a multi-view video according to a specific embodiment of the present invention is shown.

[0064] [ Figure 6B ] Figure 6B The steps of a depth encoding method for the current block according to a specific embodiment of the present invention are shown.

[0065] [ Figure 7A ] Figure 7A An example of a data stream according to a specific embodiment of the present invention is shown.

[0066] [ Figure 7B ] Figure 7B An example of a data stream according to another specific embodiment of the present invention is shown.

[0067] [ Figure 8 ] Figure 8 A multi-view video encoding device according to a specific embodiment of the present invention is shown.

[0068] [ Figure 9 ] Figure 9 An apparatus for processing multi-view video data according to a specific embodiment of the present invention is shown. Detailed Implementation

[0069] As mentioned above, Figure 1 A multi-view video data processing scheme according to the prior art is illustrated. According to this embodiment, depth information is determined, encoded, and sent in a data stream to a decoder that decodes it.

[0070] Figure 2 A multi-view video data processing scheme according to a specific embodiment of the present invention is illustrated. According to this specific embodiment of the invention, depth information is not encoded into the data stream, but is determined on the client side from the reconstructed image of the multi-view video.

[0071] according to Figure 2 The scheme shown, for example, uses an MV-HEVC encoder to process the texture image T derived from a captured view. x0y0 and T x1y0 Encode the data (CODEC) and send it to the user's display device. On the client side, the view's texture T* x0y0 and T* x1y0 The decoded and estimated module FE is used to estimate the value of each view T. x0y0 and T x1y0 Associated depth information D' x0y0 and D' x1y0 For example, depth information D' x0y0and D' x1y0 Obtained through depth estimation software (DERS).

[0072] Decoded view T* x0y0 and T* x1y0 and the associated depth D' of each view x0y0 and D' x1y0 The synthesizing algorithm is used to compute intermediate views, such as the intermediate view S' here. x0y0 and S' x1y0 For example, the VSRS software mentioned above can be used as a view composition algorithm.

[0073] When estimating depth information after transmitting encoded multiview video data, the following problems may occur. Incorrect depth values ​​may be obtained due to compression artifacts (e.g., blockiness or quantization noise, especially at low rates) in the texture that is decoded and used to estimate depth information.

[0074] Furthermore, the complexity of the client-side terminal is greater than the complexity when depth information is transmitted to the decoder. This may mean that simpler depth estimation algorithms used at the encoder may fail in complex scenes.

[0075] On the client side, the texture information may not include enough redundancy or data useful for performing depth estimation, for example, because the texture information may not be encoded on the server side.

[0076] This invention proposes a method for selecting a mode for obtaining synthetic data from a first acquisition mode (M1) and a second acquisition mode (M2). According to the first acquisition mode, the synthetic data is encoded and transmitted to a decoder, and according to the second acquisition mode, the synthetic data is estimated on the client side. This method utilizes both schemes in a flexible manner.

[0077] Therefore, for each image, or each block, or for any other granularity, select the optimal path to obtain one or more synthetic data items.

[0078] Figure 3A The steps of a method for processing multi-view video data according to a specific embodiment of the present invention are illustrated. According to this specific embodiment of the invention, the selected acquisition mode is encoded and transmitted to the decoder.

[0079] Specifically, a data stream BS containing texture information of one or more views of a multi-view video is transmitted to the decoder. For example, it is assumed that two views have already been encoded in the data stream BS.

[0080] The data stream BS also includes at least one syntax element representing an information item, which specifies one of a first acquisition mode M1 and a second acquisition mode M2 ​​for acquiring at least one composite data item.

[0081] In step 30, the decoder decodes the texture information of the data stream to obtain textures T*0 and T*1.

[0082] In step 31, a syntax element representing the information item specifying the acquisition mode is decoded from the data stream. This syntax element is encoded in the data stream of at least one block of the texture image of the view. Therefore, its value can change at each texture block of the view. According to another variation, the syntax element is encoded once for all blocks of the texture image of view T0 or T1. The information item specifying the mode used to acquire the composite data item is therefore the same for all blocks of texture image T0 or T1.

[0083] In another variation, the syntax element is encoded once for all texture images of the same view, or the syntax element is encoded once for all views.

[0084] Here, a variant of the encoding syntax element for each texture image of the view is considered. Then, after step 31, the acquisition mode information item d0 associated with the decoded texture image T*0 and the acquisition mode information item d1 associated with the decoded texture image T*1 are obtained.

[0085] In step 32, for each information item d0 and d1 that specifies a mode for obtaining synthetic data associated with the decoded texture images T*0 and T*1 respectively, it is checked whether the obtained mode corresponds to the first obtained mode M1 or the second obtained mode M2.

[0086] If information item d0 and correspondingly d1 specify the first acquisition mode M1, then in step 34, the synthetic data D*0 and D*1 associated with the decoded texture images T*0 and T*1 are decoded from the data stream BS.

[0087] If information item d0 and correspondingly d1 specify the second acquisition mode M2, then in step 33, synthetic data D associated with the decoded texture images T*0 and T*1, respectively, are estimated from the reconstructed texture images of the multi-view video. + 0, D + 1. For this purpose, the estimation can be performed using the decoded texture T*0, the corresponding T*1, and possibly other previously reconstructed texture images.

[0088] In step 35, the decoded textures T*0 and T*1, and the decoded (D*0, D*1) or estimated (D*1) are... + 0, D +1) The composite information is used to composite the image of the intermediate view S0.5.

[0089] Figure 3B Steps for a method of processing multi-view video data according to another specific embodiment of the invention are illustrated. According to this other specific embodiment of the invention, the selected acquisition mode is not transmitted to the decoder. It is derived from previously decoded texture data.

[0090] Specifically, a data stream BS containing texture information of one or more views of a multi-view video is transmitted to the decoder. For example, it is assumed that two views have already been encoded in the data stream BS.

[0091] In step 30', the decoder decodes the texture information of the data stream to obtain textures T*0 and T*1.

[0092] In step 32', the decoder obtains an information item specifying one of the first acquisition mode M1 and the second acquisition mode M2 ​​to obtain at least one synthesis data item for synthesizing the image of the intermediate view. Depending on the variation, this information item can be obtained for each block of the texture image of the view. Therefore, the acquisition mode can be changed at each texture block of the view.

[0093] According to another embodiment, the information item is obtained once for all blocks of the texture image in view T*0 or T*1. Therefore, the information item specifying the mode for obtaining the composite data item is the same for all blocks of texture image T*0 or T*1.

[0094] According to another variation, information items are obtained once for all texture images in the same view, or information items are obtained once for all views.

[0095] Here, we consider a variation of obtaining information items based on each texture image for the view. Then, after step 32', we obtain the obtained mode information item d0 associated with the decoded texture image T*0 and the obtained mode information item d1 associated with the decoded texture image T*1. Here, the obtained mode information items are obtained by applying the same determination process applied at the encoder. An example of the determination process will be described later with reference to FIG4.

[0096] After step 32', if information item d0, correspondingly d1 specifies the first acquisition mode M1, then in step 34', synthetic data D*0 and D*1, respectively associated with the decoded texture images T*0 and T*1, are decoded from the data stream BS.

[0097] If information item d0 and correspondingly d1 specify the second acquisition mode M2, then in step 33', synthetic data D associated with the decoded texture images T*0 and T*1, respectively, is estimated from the reconstructed texture images of the multi-view video.+ 0, D + 1. For this purpose, the estimation can use the decoded textures T*0, T*1, and possibly other previously reconstructed texture images.

[0098] In step 35', the decoded textures T*0 and T*1, and the decoded (D*0, D*1) or estimated (D*1) are... + 0, D + 1) The composite information is used to composite the image of the intermediate view S0.5.

[0099] The method for processing multi-view video data described herein according to specific embodiments of the invention is particularly applicable to cases where the synthetic data corresponds to depth information. However, this data processing method is applicable to any type of synthetic data, such as object segmentation maps.

[0100] For a given view of a video at a given time, the above methods can be applied to several types of composite data. For example, if the compositing module is aided by depth maps and object segmentation maps, then these two types of composite data can be partially transmitted to the decoder, and partially derived by the decoder or the compositing module.

[0101] It should also be noted that a portion of the texture can be estimated, for example, through interpolation. The view corresponding to this texture estimated at the decoder is considered a synthetic data item in this case.

[0102] The examples described here include two texture views that generate two depth maps respectively, but other combinations are of course possible, including processing depth maps associated with one or more texture views at a given time.

[0103] Figure 4A The steps of a multi-view video encoding method according to a specific embodiment of the present invention are illustrated. The encoding method is described here with the two views comprising textures T0 and T1 respectively.

[0104] In step 40, each texture T0 and T1 is encoded and decoded to provide decoded textures T*0 and T*1. It should be noted that a texture here can correspond to an image of a view, or a patch of an image of a view, or any other type of granularity related to the texture information of a multi-view video.

[0105] In step 41, a depth estimator is used to estimate synthetic data, such as depth map D, from the decoded textures T*0 and T*1. + 0 and D + 1. This is the second method, M2, used to obtain synthetic data.

[0106] In step 42, synthetic data D0 and D1 are estimated, for example, using a depth estimator on the unencoded textures T0 and T1. Then, in step 43, the obtained synthetic data D0 and D1 are encoded and decoded to provide reconstructed synthetic data D*0 and D*1. This is the first mode M1 used to obtain the synthetic data.

[0107] In step 44, the acquisition mode for obtaining synthetic data at the decoder is determined from the first acquisition mode M1 and the second acquisition mode M2.

[0108] According to a specific embodiment of the invention, syntax elements are encoded in the data stream to specify the selected acquisition mode. According to this specific embodiment of the invention, various variations are possible depending on how rate and distortion are evaluated according to the criterion to be minimized, J = D + λR, where R corresponds to rate, D corresponds to distortion, and λ is the Lagrange operator used for optimization.

[0109] The first variant is based on the synthesis of intermediate views or blocks of intermediate views. In this case, the obtained pattern is encoded for each block, and to evaluate the quality of the synthesized view, two patterns used to obtain the synthesized data are considered. Therefore, the synthesized data D is estimated based on the decoded textures T*0 and T*1 and from the decoded textures T*0 and T*1. + 0 and D + 1. For the first version of the intermediate view synthesized for acquisition mode M2. Then, the rate corresponds to the encoding cost of the T*0 and T*1 textures and the encoding cost of the syntax elements specifying the selected acquisition mode. This rate can be precisely calculated using, for example, an entropy encoder (e.g., arithmetic binary encoding, variable-length encoding, with or without context adaptation).

[0110] Furthermore, based on the decoded textures T*0 and T*1 and the decoded composite data D*0 and D*1, a second version of the intermediate view is synthesized for acquisition mode M1. The rate then corresponds to the encoding cost of textures T*0 and T*1 and the composite data D*0 and D*1, wherein the encoding cost of the syntax elements specifying the selected acquisition mode is added to the encoding cost of textures T*0 and T*1 and the composite data D*0 and D*1. This rate can be calculated according to the above specifications.

[0111] In both cases, distortion can be calculated by comparing the image or block of the synthetic view with the uncoded image or block of the synthetic view from the uncoded textures T0 and T1 and the uncoded synthetic data D0 and D1.

[0112] Choose the acquisition mode that provides the lowest rate / distortion cost J.

[0113] According to another variation, distortion may be determined by applying a non-referenced metric to the synthesized image or patch to avoid using the original uncompressed texture. Such a non-referenced metric could, for example, measure the amount of noise, blur, blockiness, edge sharpness, etc., in the synthesized image or patch.

[0114] According to another variation, for example, by combining the synthetic data D0 and D1 of the uncompressed texture estimate with the synthetic data D from the encoded and decoded texture estimate... + 0 and D + 1. Comparison is used to select the appropriate acquisition mode. If the synthesized data is sufficiently close, estimating the synthesized data on the client side, according to defined criteria, will be more efficient than encoding and transmitting the synthesized data. This variant avoids the synthesis of images or blocks from intermediate views.

[0115] When the synthetic data corresponds to a depth map, it is also possible to determine other variations of the mode used to obtain the synthetic data. The choice of the acquisition mode can, for example, depend on the characteristics of the depth information item. For example, computer-generated depth information items or high-quality captured depths are more likely to be suitable for acquisition mode M1. According to this variation, the depth map can also be estimated from the decoded texture as described above and compete with the computer-generated or high-quality captured depth map. The computer-generated or high-quality captured depth map then replaces the depth map estimated from the uncompressed texture in the above method.

[0116] According to another variation, depth quality can be used to determine the pattern from which synthetic data is obtained. Depth quality, which can be measured by appropriate objective metrics, can include relevant information. For example, when depth quality is low, or when the temporal coherence of depth is low, pattern M2 may be best suited for obtaining depth information.

[0117] Once a mode for obtaining synthetic data is selected at the end of step 44, in step 45, the syntax element d representing the selected acquisition mode is encoded in the data stream. When the selected and encoded mode corresponds to the first acquisition mode M1, synthetic data D0 and D1 are also encoded in the data stream for the block or image under consideration.

[0118] According to a specific embodiment of the invention, when the selected and encoded mode corresponds to the second acquisition mode M2, additional information may also be encoded in the data stream in step 46. For example, when the synthesized data item is obtained according to the second acquisition mode, such information may correspond to one or more control parameters to be applied to the decoder or by the synthesis module. For example, these parameters may be used to control the synthesized data or the depth estimator.

[0119] For example, control parameters can control the characteristics of the depth estimator, such as increasing or decreasing the search interval, or increasing or decreasing the accuracy.

[0120] Control parameters can specify how the synthesized data items should be estimated on the decoder side. For example, control parameters specify which depth estimator to use. For instance, in step 41 for estimating the depth map, the encoder can test several depth estimators and select the one that provides the best rate / distortion tradeoff from the following: a pixel-based depth estimator, a triangle deformation-based depth estimator, a fast depth estimator, a monocular neural network depth estimator, and a neural network depth estimator using multiple references. Depending on this variation, the encoder instructs the decoder or synthesis module to use a similar synthesized data estimator.

[0121] Depending on another variant or in addition to the previous variant, the control parameters may include parameters of the depth estimator, such as disparity interval, accuracy, neural network model, optimization or aggregation method, smoothing factor of energy function, cost function (color-based, correlation-based, frequency-based), simple depth map that can be used as initialization for the client-side depth estimator, etc.

[0122] Figure 4B Steps of a multi-view video encoding method according to another specific embodiment of the present invention are illustrated. According to the specific embodiment described herein, the mode used to obtain the synthetic data is not encoded in the data stream, but rather the available encoding information is derived from the decoder.

[0123] The encoding method is described here with the two views including textures T0 and T1 respectively.

[0124] In step 40', each texture T0 and T1 is encoded and decoded to provide decoded textures T*0 and T*1. It should be noted that a texture here may correspond to an image of a view, a block of an image of a view, or any other type of granularity related to the texture information of a multi-view video.

[0125] In step 44', the acquisition mode for obtaining synthetic data at the decoder is determined from the first acquisition mode M1 and the second acquisition mode M2.

[0126] According to the specific embodiments described herein, the encoder can use the decoder to determine the acquisition mode to be applied to the block or image under consideration, based on any available information items.

[0127] Depending on the variant, the acquisition mode can be selected based on quantization parameters (e.g., QP used to encode image or texture patches). For example, a second acquisition mode is selected when the quantization parameter is greater than a given threshold, and a first acquisition mode is selected otherwise.

[0128] According to another variation, when the synthetic data corresponds to depth information, the synthetic data D0 and D1 can be computer-generated or captured in high quality. This type of synthetic data is more suitable for obtaining mode M1. Therefore, when this is the case, the selected mode for obtaining the synthetic data will be obtaining mode M1. According to this variation, metadata items must be transmitted to the decoder to specify the source of the depth (computer-generated, captured in high quality). This information item can be transmitted at the view sequence level.

[0129] At the end of step 44', if the first acquisition mode M1 is selected, then in step 42', for example, the synthetic data D0 and D1 are estimated using a depth estimator on the unencoded textures T0 and T1. This estimation is, of course, not performed if the synthetic data comes from computer generation or high-quality capture.

[0130] Then, in step 47', the obtained synthetic data D0 and D1 are encoded in the data stream.

[0131] According to a specific embodiment of the invention, when the selected acquisition mode corresponds to the second acquisition mode M2, additional information can also be encoded in the data stream in step 46'. For example, when the synthesized data item is obtained according to the second acquisition mode, such information may correspond to one or more control parameters to be applied to the decoder or by the synthesis module. Such control parameters are similar to... Figure 4A The parameters described in the document.

[0132] Figure 5 An example of a multi-view video data processing scheme according to a specific embodiment of the present invention is shown.

[0133] According to a specific embodiment of the invention, the scene is captured by a video capture system CAPT. For example, the view capture system includes one or more cameras that capture the scene.

[0134] Based on the example described here, the scene is captured by two converging cameras located outside the scene and facing it from two different positions. Therefore, the cameras are at different distances from the scene and have different angles / orientations. Each camera provides an uncompressed sequence of images. The image sequences consist of sequences of texture images T0 and T1, respectively.

[0135] Texture images T0 and T1, derived from image sequences from two cameras respectively, are encoded by an encoder COD (e.g., an MV-HEVC encoder as a multi-view video encoder). The encoder COD provides a data stream BS, which is transmitted to the decoder DEC, for example, via a data network.

[0136] During encoding, depth maps D0 and D1 are estimated from the uncompressed textures T0 and T1 using a depth estimator (e.g., a DERS estimator), and depth map D* is estimated from the decoded textures T*0 and T*1. + 0 and D + 1. Synthesize a first view T'0 located at the position captured by one of the cameras (e.g., position 0 here) using depth map D0, and use depth map D + A second view T'0 located at the same position is synthesized. For example, the quality of the two synthesized views is compared by calculating the PSNR (Peak Signal-to-Noise Ratio) between each synthesized view T'0, T'0 and the captured view T0 located at the same position. This comparison allows selection of the acquisition mode of the depth map D0 from a first acquisition mode and a second acquisition mode. According to the first acquisition mode, the depth map D0 is encoded and transmitted to the decoder; according to the second acquisition mode, the depth map D... + 0 is estimated at the decoder. The same method is repeated for the depth map D1 associated with the captured texture T1.

[0137] Figure 7A An example of a portion of a data stream BS according to a particular embodiment of the present invention is shown. The data stream BS includes encoded textures T0 and T1 and syntax elements d0 and d1, which specify the mode for obtaining depth maps D0 and D1 for each of textures T0 and T1, respectively.

[0138] If it is decided to encode and transmit depth maps D0 and D1 accordingly, and the values ​​of syntax elements d0 and d1 are, for example, 0, then the data stream BS includes the encoded depth maps D0 and D1 accordingly.

[0139] If it is decided not to encode depth maps D0 and D1, and the values ​​of syntax elements d0 and d1 are, for example, 1, then the data stream BS does not include depth maps D0 and D1. According to a variant of the embodiment, it may include the process where depth maps D0 and D1 are obtained by the decoder or synthesis module respectively. + 0, D + 1. One or more control parameters PAR to be applied.

[0140] The encoded data stream BS is then decoded by the decoder DEC. For example, the decoder DEC is included in a smartphone equipped with a free navigation decoding feature. According to this example, the user views the scene from the viewpoint provided by the first camera. The user then slowly swipes their viewpoint left to another camera. During this process, the smartphone displays an intermediate view of the scene that was not captured by the camera.

[0141] For this purpose, the data stream BS is scanned and decoded by, for example, an MV-HEVC decoder to provide two decoded textures T*0 and T*1. The syntax element d associated with each texture... k Decoded, where k = 0 or 1. If syntax element d k If the value is 0, the decoder decodes the depth map D* from the data stream BS. k .

[0142] If syntax element d k If the value is 1, then the depth map D is estimated at the decoder or by the synthesis module from the decoded textures T*0 and T*1. + k .

[0143] For example, the compositing module SYNTH, based on the VVS (Generic View Composer) compositing algorithm, uses, when appropriate, decoded textures T*0 and T*1 and decoded depth maps D*0 and D*1 or estimated depth maps D. + 0 and D + 1. Composite the intermediate view into a composite intermediate view that is included between the views corresponding to textures T0 and T1.

[0144] Figure 5 The multi-view video data processing scheme described herein is not limited to the embodiments described above.

[0145] According to another specific embodiment of the invention, the scene is captured from six different positions by six omnidirectional cameras located within the scene. Each camera provides a sequence of 2D images in an equal rectangular projection format (ERP). The six textures from the cameras are encoded using a 3D-HEVC encoder as a multi-view encoder, providing a data stream BS, for example, transmitted via a data network.

[0146] When encoding multi-view sequences, a 2×3 matrix of the source texture T (texture from the camera) is provided as input to the encoder. Figure 6A A texture matrix is ​​shown, which includes texture T. xiyj , where i = 0, 1 or 2, and j = 0, 1 or 2.

[0147] According to the embodiment described herein, a neural network-based depth estimator is used to estimate the source depth map matrix D from the uncompressed texture. The texture matrix T is encoded and decoded using a 3D-HEVC encoder to provide a decoded texture matrix T*. The decoded texture matrix T* is used to estimate the depth map matrix D using the neural network-based depth estimator. + .

[0148] According to a specific embodiment of the invention described herein, a depth map acquisition mode associated with the texture is selected for each coding block or unit (also referred to as a coding tree unit (CTU) in an HEVC encoder).

[0149] Figure 6B The current block D to be encoded is shown. x0y0 The steps of the depth coding method (x, y, t) are given, where x and y correspond to the top-left corner of a block in the image, and t corresponds to the time of the image.

[0150] Consider the first block of the first depth map encoded at time t=0 of the video sequence (subsequently encoded by D). x0y0 (0, 0, 0) identifier). When this first block D x0y0 When encoding (0, 0, 0), the depth of all other unprocessed blocks is assumed to originate from the estimated source depth D. The other unprocessed blocks belong to both the current view x0y0 and other neighboring views.

[0151] First, the current block D is evaluated by determining the optimal encoding mode from the various block depth encoding tools available to the encoder. x0y0 Depth encoding of (0, 0, 0). Such encoding tools can include any type of depth encoding tool available in multi-view encoders.

[0152] In step 60, the depth D of the current block is determined using the first encoding tool. x0y0 (0, 0, 0) is encoded to provide the encoding-decoding depth D* of the current block. x0y0 (0, 0, 0).

[0153] In step 61, a view at the location of one of the cameras is synthesized, for example, using VVS compositing software. For example, the view at location x1y0 is synthesized using views decoded at locations x0y0, x2y0, and x1y1 in the texture matrix T*. During view synthesis, the depth of all blocks of the multi-view video that have not yet been processed is derived from the estimated source depth D. The depth of all blocks of the multi-view video whose depth has been encoded is derived from the encoded-decoded or estimated depth from the decoded texture according to the acquisition mode selected for each block. Depending on the evaluated encoding tool, the depth of the current block used for synthesizing the view at location x1y0 is the encoded-decoded depth D*. x0y0 (0, 0, 0).

[0154] In step 62, the composite view at position x1y0 is compared with the source view T. x1y0 The quality of the synthesized view is evaluated using an error metric (e.g., squared error) and the encoding cost is calculated based on the current block depth of the testing tool.

[0155] In step 63, it is checked whether all depth coding tools have been tested for the current block. If not, steps 60 to 62 are repeated for the next coding tool; otherwise, the method proceeds to step 64.

[0156] In step 64, a deep coding tool that provides the best rate / distortion tradeoff is selected, such as a deep coding tool that minimizes the rate / distortion criterion J = D + λR.

[0157] In step 65, the estimated depth D of the current block is used in conjunction with the VVS software. + x0y0 (0, 0, 0), using the decoded textures at positions x0y0, x2y0, and x1y1, synthesize another view at the same positions as in step 61.

[0158] In step 66, the composite view at position x1y0 and the source view T are calculated. x1y The distortion is between 0 and 0, and the depth encoding cost is set to 0 because, according to this acquisition mode, the depth is not encoded but estimated at the decoder.

[0159] In step 67, the optimal acquisition mode is determined based on the rate / distortion cost of each mode used to obtain the depth. In other words, the mode used to obtain the depth that minimizes the rate / distortion criterion is selected from encoding the depth using the best encoding tool selected in step 64 and estimating the depth at the decoder.

[0160] In step 68, syntax elements are encoded in the data stream to specify the acquisition mode selected for the current block. If the selected acquisition mode corresponds to depth encoding, then depth encoding is performed in the data stream according to the previously selected best encoding tool.

[0161] If the size of the first block is 64×64, for example considering the next block D to be processed... x0y0 If (64, 0, 0), then repeat steps 60 to 68. Taking into account the encoded-decoded depth or estimated depth of the blocks previously processed during view composition, all blocks in the depth map associated with the view texture at position x0y0 are processed accordingly.

[0162] Depth maps for other views are processed in a similar manner.

[0163] According to this particular embodiment of the invention, the encoded data stream also includes different information for each block. If it has been decided to encode and transmit depth for a given block, then for that block, the data stream includes the encoded texture of the block, the encoded depth data block, and syntax elements specifying the mode used to obtain the depth of the block.

[0164] If it is decided not to encode the depth of the block, then for that block, the data stream includes the encoded texture of the block, a depth information block including the same grayscale values, and a syntax element specifying the mode used to obtain the depth of the block.

[0165] It should be noted that in some cases, the data stream may include the texture of all blocks being sequentially encoded, followed by the syntax elements and depth data of the blocks.

[0166] Decoding can be performed, for example, via a virtual reality headset equipped with free navigation features and worn by the user. The viewer observes the scene from a viewpoint provided by one of six cameras. The user looks around and slowly begins to move within the scene. The headset follows the user's movement and displays corresponding views of the scene not captured by the cameras.

[0167] To this end, the decoder DEC decodes the texture matrix T* from the encoded data stream. The syntax elements for each block are also decoded from the encoded data stream. Based on the values ​​of the syntax elements decoded for that block, the depth of each block is obtained either by decoding the depth data block encoded for that block or by estimating the depth data from the decoded texture.

[0168] An intermediate view is synthesized using a decoded texture matrix T* and a reconstructed depth matrix, the reconstructed depth matrix including the depth data of the block obtained according to the acquisition mode specified by the syntax element decoded for each block.

[0169] According to another specific embodiment of the present invention, Figure 5 The multi-view video data processing scheme described in the document is also applicable to situations where syntax elements are not encoded at the block level or image level.

[0170] For example, the encoder COD can apply a decision mechanism at the image level to determine whether depth should be transmitted to the decoder or estimated after decoding.

[0171] To this end, an encoder operating in variable rate mode assigns quantization parameters (QP) to blocks of a texture image in a known manner in order to achieve a target total rate.

[0172] It's possible to use a weighted average across blocks to calculate the average QP assigned to each block of the texture image. This provides the average PQ of the texture image, representing the importance level of the image.

[0173] If the obtained average value PQ is higher than a certain threshold, it indicates that the target rate is low. Then, the encoder decides to calculate the depth map of the texture image from the uncompressed texture of the multi-view video, encodes the calculated depth map, and transmits it as a data stream.

[0174] If the average value PQ is below or equal to a predetermined threshold, the target rate is high. The encoder does not calculate the depth of the texture image and moves on to the next texture image. No depth is encoded for that image, and no indicator is transmitted to the decoder.

[0175] Figure 7B An example of a portion of an encoded data stream according to a particular embodiment of the present invention is shown.

[0176] The encoded data stream specifically includes the encoded textures for each image, here T0 and T1. The encoded data stream also includes information for obtaining the average value PQ for each image. This can be encoded at the image level, or conventionally obtained from the QP encoded for each block in the data stream.

[0177] For each texture image T0 and T1, the encoded data stream also includes computed and encoded depth data D0 and / or D1, depending on the decision made at the encoder. Note that the syntax elements d0 and d1 are not encoded in the data stream here. Once the depth of the texture image has been determined at the decoder, the data stream may include the parameter PAR to be applied when estimating the depth. These parameters have been described above.

[0178] The decoder (DEC) scans the encoded data stream and decodes texture images T*0 and T*1. The decoder applies the same decision mechanism as the encoder, calculating the average value PQ for each texture image. Then, the decoder uses a predetermined threshold, either transmitted in the data stream or known to the decoder, to infer whether to decode or estimate the depth of a given texture image.

[0179] Then, the decoder uses a similar approach to... Figure 5 The operation is carried out in the manner described in the first embodiment.

[0180] Figure 8 A simplified structure of an encoding device COD is shown, which is adapted to implement the encoding method according to any particular embodiment of the invention described above, particularly regarding... Figure 2 , Figure 4A and Figure 4B The encoder COD can, for example, correspond to information about... Figure 5 The encoder COD is described.

[0181] According to a specific embodiment of the invention, the steps of the encoding method are implemented by computer program instructions. For this purpose, the encoding device COD has a standard computer architecture and specifically includes a memory MEM and a processing unit UT equipped with, for example, a processor PROC, and driven by a computer program PG stored in the memory MEM. The computer program PG includes instructions for implementing the steps of the encoding method as described above when the program is executed by the processor PROC.

[0182] During initialization, the code instructions of the computer program PG are loaded into memory, for example, before being executed by the processor PROC. Specifically, the processor PROC of the processing unit UT implements the steps of the above-described encoding method according to the instructions of the computer program PG.

[0183] Figure 9 A simplified structure of a device DTV for processing multi-view video data is shown, which is adapted to implement a method for processing multi-view data according to any specific embodiment of the invention previously described, particularly concerning... Figure 2 , Figure 3A and Figure 3B A device used for processing multi-view video data, such as a DTV, can correspond to, for example, a combination of... Figure 5 The described synthesis module SYNTH, or corresponding to the module including Figure 5 The device consists of the SYNTH synthesis module and the DEC decoder.

[0184] According to a specific embodiment of the present invention, a device DTV for processing multi-view video data has a standard computer architecture and specifically includes a memory MEM0 and a processing unit UT0, which is equipped with, for example, a processor PROC0 and driven by a computer program PG0 stored in the memory MEM0. The computer program PG0 includes instructions for implementing the steps of the method for processing multi-view video data as described above when the program is executed by the processor PROC0.

[0185] During initialization, the code instructions of the computer program PG0 are loaded into memory, for example, before being executed by the processor PROC0. Specifically, the processor PROC0 of the processing unit UT0 implements the steps of the method for processing multi-view video data described above according to the instructions of the computer program PG0.

[0186] According to a specific embodiment of the present invention, a device DTV for processing multi-view video data includes a decoder DEC adapted to decode one or more coded data streams representing multi-view video.

Claims

1. A method for processing multi-view video data, the method comprising: - For at least one block of an image representing a view encoded in an coded data stream representing a multi-view video, at least one information item is obtained, based on the granularity associated with the texture information of the multi-view video, specifying a first acquisition mode and a second acquisition mode for obtaining at least one synthetic data item. The at least one synthetic data item is used to synthesize at least one image of an intermediate view of the multi-view video, the intermediate view being not encoded in the encoded data stream. The first acquisition mode corresponds to decoding at least one information item representing the at least one synthetic data item from the encoded data stream, and the second acquisition mode corresponds to obtaining the at least one synthetic data item from at least a reconstructed encoded image, the reconstructed encoded image being derived from the encoding of the at least one image and the reconstruction of the at least one encoded image. - Obtain the at least one synthetic data item according to the acquisition pattern specified by the at least one information item. - Synthesize at least a portion of the image of the intermediate view from at least the reconstructed coded image and the obtained at least one synthetic data item.

2. The method for processing multi-view video data according to claim 1, wherein, The at least one synthetic data item corresponds to at least a portion of the depth map.

3. The method for processing multi-view video data according to claim 1, wherein, The at least one information item that specifies the pattern used to obtain the synthetic data item is obtained by decoding the syntax elements.

4. The method for processing multi-view video data according to claim 1, wherein, The at least one information item that specifies the mode for obtaining the synthetic data item is obtained from at least one data item encoded for the reconstructed coded image.

5. The method for processing multi-view video data according to claim 4, wherein, The acquisition mode is selected from a first acquisition mode and a second acquisition mode based on the value of the quantization parameter used to encode at least the block.

6. The method for processing multi-view video data according to claim 1, further comprising: when the at least one information item specifies that the synthesized data item is obtained according to the second acquisition mode: - Decode at least one control parameter from the encoded data stream. - When the synthesized data item is obtained according to the second acquisition mode, the control parameters are applied.

7. An apparatus for processing multi-view video data, comprising: Processor, the processor being configured to: - For at least one block of an image representing a view encoded in an coded data stream representing a multi-view video, at least one information item is obtained, based on the granularity associated with the texture information of the multi-view video, specifying a first acquisition mode and a second acquisition mode for obtaining at least one synthetic data item. The at least one synthetic data item is used to synthesize at least one image of an intermediate view of the multi-view video, the intermediate view being not encoded in the encoded data stream. The first acquisition mode corresponds to decoding at least one information item representing the at least one synthetic data item from the encoded data stream, and the second acquisition mode corresponds to obtaining the at least one synthetic data item from at least a reconstructed encoded image, the reconstructed encoded image being derived from the encoding of the at least one image and the reconstruction of the at least one encoded image. - Obtain the at least one synthetic data item according to the acquisition pattern specified by the at least one information item. - Synthesize at least a portion of the image of the intermediate view from at least the reconstructed coded image and the obtained at least one synthetic data item.

8. A terminal comprising the device according to claim 7.

9. A method for encoding multi-view video data, the method comprising: - For at least one block of an image representing a view in a coded data stream of multi-view video, at least one information item is determined, based on the granularity associated with the texture information of the multi-view video, specifying a mode for obtaining at least one synthetic data item, which is one of a first acquisition mode and a second acquisition mode. The at least one synthetic data item is used to synthesize at least one image of an intermediate view of the multi-view video, the intermediate view being not encoded in the encoded data stream. The first acquisition mode corresponds to decoding at least one information item representing the at least one synthetic data item from the encoded data stream, and the second acquisition mode corresponds to obtaining the at least one synthetic data item from at least a reconstructed encoded image, the reconstructed encoded image being derived from the encoding of the at least one image and the reconstruction of the at least one encoded image. - Encode the image in the encoded data stream.

10. The method for encoding multi-view video data according to claim 9, comprising: Syntax elements associated with the information item are encoded in the data stream, the information item specifying the pattern used to obtain the synthesized data item.

11. The method for encoding multi-view video data according to any one of claims 9 to 10, further comprising, when the information item specifies that the synthesized data item is obtained according to the second acquisition mode: - Encode at least one control parameter in the encoded data stream, which will be applied when the synthesized data item is obtained according to the second acquisition mode.

12. An apparatus for encoding multi-view video data, comprising: The device is configured to include a processor and memory. - For at least one block of an image representing a view in a coded data stream of multi-view video, at least one information item is determined, based on the granularity associated with the texture information of the multi-view video, specifying a mode for obtaining at least one synthetic data item, which is one of a first acquisition mode and a second acquisition mode. The at least one synthetic data item is used to synthesize at least one image of an intermediate view of the multi-view video, the intermediate view being not encoded in the encoded data stream. The first acquisition mode corresponds to decoding at least one information item representing the at least one synthetic data item from the encoded data stream, and the second acquisition mode corresponds to obtaining the at least one synthetic data item from at least a reconstructed encoded image, the reconstructed encoded image being derived from the encoding of the at least one image and the reconstruction of the at least one encoded image. - Encode the image in the encoded data stream.

13. A computer program product comprising instructions which, when executed by a processor, are configured to implement the method for processing multi-view video data according to any one of claims 1 to 6, or to implement the method for encoding multi-view video data according to any one of claims 9 to 11.

14. A computer-readable storage medium comprising instructions, which, when executed by a processor, are configured to implement the method for processing multi-view video data according to any one of claims 1 to 6, or to implement the method for encoding multi-view video data according to any one of claims 9 to 11.

Citation Information

Patent Citations

  • Method and apparatus for video coding

    US20140168362A1

  • An apparatus, a method and a computer program for video coding and decoding

    WO2013159330A1