Method and apparatus for encoding, transmitting, and decoding volumetric video

The method addresses the challenge of assessing view reliability in MVD frames by encoding metadata about depth information fidelity within the data stream, leading to improved volumetric video rendering quality by ensuring only reliable views contribute to viewport synthesis.

JP7692408B2Active Publication Date: 2025-06-13INTERDIGITALCE PATENT HLDG SAS
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022519816
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-10-02
Filing Date
2020-10-01
Publication Date
2025-06-13
Estimated Expiration
2040-10-01

AI Technical Summary

Technical Problem

There is a lack of a method to assess the reliability of information carried by different views of Multi-View+Depth (MVD) frames for synthesizing viewport frames, which affects the quality of volumetric video rendering.

Method used

A method is introduced that involves obtaining a parameter representing the fidelity of depth information for each view of a multi-view frame and encoding this metadata within the data stream. This parameter can be a boolean value or a numerical value indicating the reliability of the depth information, allowing for weighted contribution of views during synthesis.

Benefits of technology

The proposed method enhances the rendering quality of volumetric video by ensuring that only reliable views contribute to the synthesis of viewport frames, thereby improving immersion and reducing artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007692408000002
    Figure 0007692408000002
  • Figure 0007692408000003
    Figure 0007692408000003
  • Figure 0007692408000004
    Figure 0007692408000004
Patent Text Reader

Abstract

A method, device, and stream for encoding, decoding, and transmitting multi-view frames, in which some of the views are more reliable than others, is disclosed. The multi-view frames are encoded in a data stream that is associated with metadata that includes, for at least one of the views, a parameter indicating the reliability of the information carried by this view. This information is used at the decoding side to determine the view's contribution when synthesizing pixels of a viewport frame for a given field of view in 3D space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This principle generally relates to the domain of three-dimensional (3D) scenes and volumetric video content. This document is also understood in the context of the encoding, formatting, and decoding of data representing textures and the geometric shapes of 3D scenes for the rendering of volumetric content on end-user devices such as mobile devices or head-mounted displays (HMDs). Among other themes, this principle relates to pruning pixels of multi-view images to guarantee optimal bitstreams and rendering quality.

Background Art

[0002] This section is intended to introduce the reader to various technical aspects that may be related to various aspects of the principle described and / or claimed below. This discussion is thought to be useful in providing the reader with background information to facilitate a better understanding of the various aspects of the principle. Thus, it should be understood that these descriptions are to be read from this perspective and should not be read as an admission of prior art.

[0003] In recent years, there has been growth in the availability of large field-of-view content (up to 360°). Such content may not be fully visible to users viewing content on immersive display devices such as head-mounted displays, smart glasses, PC screens, tablets, smartphones, etc. This means that at any given moment, a user may only be able to view a portion of the content. However, users can typically navigate within the content by various means such as head movement, mouse movement, touch screen, voice, etc. Typically, it is desirable to encode and decode this content.

[0004] With immersive video, also known as 360° flat video, users can view everything around them by rotating their heads around a stationary point. The rotation only enables a 3 Degrees of Freedom (3DoF) experience. For example, even if a 3DoF video is sufficient for a first omnidirectional video experience using a head-mounted display device (HMD), it can quickly become frustrating for viewers who expect more degrees of freedom, such as by experiencing parallax. Additionally, 3DoF can also induce dizziness due to the translation that is not reproduced in a 3DoF video experience, where the user not only rotates their head but also translates it in three directions.

[0005] Large field of view content can be, among other things, three-dimension computer graphic imagery scene (3D CGI scene), point cloud, or immersive video. Many terms can be used to design such immersive video. For example, Virtual Reality (VR), 360, panorama, 4π steradian, immersive, omnidirectional, or large field of view.

[0006] Volumetric video (also known as 6 Degrees of Freedom (6DoF) video) is an alternative to 3DoF video. When viewing a 6DoF video, in addition to rotation, the user can also translate their head and even their body within the viewed content, experiencing parallax and even volume. Such videos greatly increase the sense of immersion and the perception of scene depth, preventing dizziness by providing consistent visual feedback during head translation. The content is created by means of dedicated sensors that enable simultaneous recording of the color and depth of the scene of interest. The use of a rig of color cameras combined with photogrammetry techniques is a way to perform such recording, even if technical difficulties remain.

[0007] 3DoF video includes a series of images resulting from the unwarping of a texture image (e.g., a spherical image encoded according to latitude / longitude projection mapping or orthographic cylindrical mapping), while 6DoF video frames embed information from several viewpoints. They can be viewed as a temporal series of point clouds resulting from 3D capture. Depending on the viewing conditions, two types of volumetric video can be considered. The first one (i.e., full 6DoF) enables full free navigation within the video content, while the second one (also known as 3DoF+) restricts the user viewing space to a limited volume called the viewing bounding box and enables a limited volume of head and parallax experience. This second context is a valuable trade-off between the free navigation of seated audience members and passive viewing conditions.

[0008] 3DoF+ content can be provided as a set of Multi-View+Depth (MVD) frames. Such content may be captured by dedicated cameras or generated from existing computer graphic (CG) content by dedicated (potentially photorealistic) rendering. The volumetric information is transmitted as a combination of color and depth patches stored in corresponding color and depth atlases, which are video encoded using a codec (e.g., HEVC). Each combination of color and depth patches represents a portion of an MVD input view, and the set of all patches is designed at the encoding stage to cover the whole.

[0009] The information carried by different views of MVD frames is variable. There is a lack of a way to take the reliability of the information carried by the views of MVD for the synthesis of viewport frames. SUMMARY OF THE INVENTION

[0010] The following presents a simplified overview of the present principle to provide a basic understanding of some aspects of the present principle. This overview is not an extensive overview of the present principle. It is not intended to identify important or critical elements of the present principle. The following overview merely presents some aspects of the present principle in a simplified form as a prelude to the more detailed explanation provided below.

[0011] The present principle relates to a method for encoding a multi-view frame. The method includes - obtaining, for a view of the multi-view frame, a parameter representing the fidelity of the depth information carried by the view; and - encoding the multi-view frame in the data stream in relation to metadata including the parameter.

[0012] In certain embodiments, the parameter representing the fidelity of the depth information of a view is determined according to the internal and external parameters of the camera that captured the view. In another embodiment, the metadata includes information indicating whether a parameter is provided for each view of the multi-view frame, and, if so, for each view, the parameter associated with the view. In a first embodiment of the present principle, the parameter representing the fidelity of the depth information of a view is a boolean value indicating whether the depth fidelity is fully reliable or partially reliable. In a second embodiment of the present principle, the parameter representing the fidelity of the depth information of a view is a numerical value indicating the reliability of the depth fidelity of the view.

[0013] The present principle also relates to a device comprising a processor configured to implement this method.

[0014] The present principle also relates to a method for decoding a multi-view frame pruned from a data stream. The method includes - decoding the multi-view frame and associated metadata from the data stream; and - Obtaining information indicating whether a parameter representing the fidelity of depth information carried by the view of the multi-view frame is provided from the metadata, and if so, obtaining the parameter for each view; - Generating a viewport frame according to the viewing pose by determining the contribution of each view of the multi-view frame as a function of the parameter associated with the view.

[0015] In one embodiment, the parameter representing the fidelity of the depth information of the view is a boolean value indicating whether the depth fidelity is completely reliable or partially reliable. In a variant of this embodiment, the contribution of a partially reliable view is ignored. In a further variant, a completely reliable view with the lowest depth information is used on the condition that multiple views are completely reliable. In another embodiment, the parameter representing the fidelity of the depth information of the view is a numerical value indicating the reliability of the depth fidelity of the view. In a variant of this embodiment, the contribution of each view during view synthesis is proportional to the numerical value of the parameter.

[0016] The present principle also relates to a device comprising a processor configured to implement this method.

[0017] The present principle also relates to a data stream comprising - Data representing a multi-view frame, and - Metadata associated with the data, the metadata including, for each view of the multi-view frame, a parameter representing the fidelity of the depth information carried by the view.

Brief Description of the Drawings

[0018] The present disclosure will be better understood and other specific features and advantages will become apparent when reading the following description. This specification refers to the accompanying drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

[0019] This principle will be more fully described below with reference to the accompanying drawings, in which examples of the principle are shown. However, the principle can be embodied in many alternative forms and should not be construed as limited to the embodiments set forth herein. Accordingly, while there is room for various modifications and alternative forms, specific examples thereof are shown by way of example in the drawings and are described in detail herein. However, it should be understood that there is no intention to limit the principle to the particular forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the principle as defined by the claims.

[0020] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the principle. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. As used herein, the terms "comprises", "comprising", "includes" and / or "including" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. Further, when an element is referred to as being "responsive to" or "connected to" another element, it can be directly responsive to or connected to the other element or intervening elements may be present. In contrast, when an element is referred to as being "directly responsive to" or "directly connected to" another element, no intervening elements are present. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated as " / ".

[0021] In this specification, terms such as first, second, etc. may be used to describe various elements, but it should be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the teachings of this principle.

[0022] Some parts of the figures include arrows on the communication path to indicate the main direction of communication, but it should be understood that communication may occur in the direction opposite to the drawn arrows.

[0023] Some examples are described with respect to block diagrams and operation flowcharts representing circuit elements, modules, or portions of code, where each block includes one or more executable instructions for implementing a specified logical function. It should also be noted that in other implementations, the functions described in the blocks may occur in the order described. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or depending on the functions involved, the blocks may be executed in the reverse order.

[0024] "According to an example" or "in an example" in this specification means that the specific features, structures, or characteristics described in connection with this example may be included in at least one implementation form of this principle. The appearance of the phrases "according to an example" or "in an example" in various places in this specification does not necessarily refer to the same example, and in separate or alternative examples, they are not necessarily mutually exclusive with other examples.

[0025] The reference numbers appearing in the claims are for illustrative purposes only and shall not have a limiting effect on the claims. Although not explicitly stated, this example and its variations may be used in any combination or partial combination.

[0026] Figure 1 shows a three-dimensional (3D) model 10 of an object and the points of a point cloud 11 corresponding to the 3D model 10. The 3D model 10 and the point cloud 11 can correspond to a potential 3D representation of an object in a 3D scene that includes, for example, other objects. The model 10 can be in a 3D mesh representation, and the points of the point cloud 11 can be the vertices of the mesh. The points of the point cloud 11 can also be points spread on the surface of the faces of the mesh. The model 10 can also be represented as a splatted version of the point cloud 11, and the surface of the model 10 is created by splatting the points of the point cloud 11. The model 10 can be represented by many different representations such as voxels or splines. Figure 1 shows the fact that the point cloud can be defined as a surface representation of a 3D object, and the surface representation of the 3D object can be generated from a cloud of points. As used herein, the projected points of a 3D object (by the extension points of the 3D scene) on an image are equivalent to projecting any representation of this 3D object, for example, a point cloud, a mesh, a spline model, or a voxel model.

[0027] The point cloud can be represented in memory, for example, as a vector-based structure, and each point has its own coordinates (e.g., three-dimensional coordinates XYZ, or the solid angle and distance (also called depth) from / to the viewpoint) within the reference frame of the viewpoint and one or more attributes also called components. Examples of components are color components that can be expressed in various color spaces, for example, RGB (red, green, and blue) or YUV (where Y is the luminance component and UV are the two chrominance components). The point cloud is a representation of a 3D scene that includes an object. The 3D scene can be viewed from a given viewpoint or a range of viewpoints. The point cloud can be obtained in many ways, for example, · from the capture of a real object photographed by a camera rig, optionally complemented by a depth active sensing device, · from the capture of a virtual / synthetic object photographed by a virtual camera rig in a modeling tool, · from a mixture of both real and virtual objects.

[0028] A 3D scene, when prepared especially for 3DoF rendering, can be represented by a Multi-View+Depth (MVD) frame. Then, the volumetric video is a sequence of MVD frames. In this approach, the volumetric information is transmitted as a combination of color and depth patches stored in the corresponding color and depth atlases, which are then video encoded using a codec (typically, HEVC). Each combination of color and depth patches typically represents a portion of an MVD input view, and the set of all patches is designed at the encoding stage to cover the entire scene while minimizing redundancy as much as possible. At the decoding stage, the atlas is first video decoded, and the patches are rendered in a view synthesis process to recover the viewport associated with the desired viewing position.

[0029] Figure 2 shows a non-limiting example of the encoding, transmission, and decoding of data representing a sequence of 3D scenes. For example, simultaneously, an encoding format that can be adapted for 3DoF, 3DoF+ and 6DoF decoding.

[0030] A sequence of 3D scenes 20 is acquired. When the sequence of pictures is a 2D video, the sequence of 3D scenes is a 3D (also called volumetric) video. The sequence of 3D scenes can be provided to a volumetric video rendering device for 3DoF, 3Dof+ or 6DoF rendering and display.

[0031] The sequence of 3D scenes 20 is provided to an encoder 21. The encoder 21 takes as input one 3D scene or a sequence of 3D scenes and provides a bitstream representing the input. The bitstream can be stored in a memory 22 and / or on an electronic data medium and can be transmitted via a network 22. The bitstream representing the sequence of 3D scenes can be read from the memory 22 and / or received from the network 22 by a decoder 23. The decoder 23 is input by the bitstream and provides, for example, the sequence of 3D scenes in point cloud format.

[0032] Encoder 21 may include several circuits that implement several steps. In a first step, encoder 21 projects each 3D scene onto at least one 2D photograph. 3D projection is any method of mapping three-dimensional points onto a two-dimensional plane. Since the latest methods for displaying graphic data are based on a planar (pixel information from several bit planes) two-dimensional medium, the use of this type of projection extends widely, especially in computer graphics, manipulation, and drafting. Projection circuit 211 provides at least one two-dimensional frame 2111 for the 3D scenes of sequence 20. Frame 2111 includes color information and depth information representing the 3D scene projected onto frame 2111. In a variant, the color information and depth information are encoded in two separate frames 2111 and 2112.

[0033] Metadata 212 is used and updated by projection circuit 211. Metadata 212 includes information regarding the projection operation (e.g., projection parameters) and the way color and depth information is organized within frames 2111 and 2112, as described in relation to FIGS. 5 - 7.

[0034] Video encoding circuit 213 encodes the sequence of frames 2111 and 2112 as a video. Photographs of 3D scenes 2111 and 2112 (or a sequence of photographs of 3D scenes) are encoded in a stream by video encoder 213. Then, the video data and metadata 212 are encapsulated within the data stream by data encapsulation circuit 214.

[0035] Encoder 213 is, for example, - JPEG, Specification ISO / CEI10918 - 1UIT - T Recommendation T.81, https: / / www.itu.int / rec / T - REC - T.81 / en; -Complies with encoders such as AVC, also known as MPEG-4 AVC or h264. UIT-TH.264 and ISO / CEI MPEG-4-Part 10 (ISO / CEI14496-10), http: / / www.itu.int / rec / T-REC-H.264 / en, HEVC (the specification of which can be found on the ITU website, T Recommendation, Series H, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en), -3D-HEVC (an extension of HEVC whose specification can be found in annex G and I of the ITU website, T Recommendation, Series H, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en), -VP9 developed by Google, -AV1 (AO Media Video 1) developed by the Alliance for Open Media or -Complies with future standards such as future versions of the Versatile Video Coder or MPEG-I or MPEG-V or encoders.

[0036] The data stream is stored in a memory accessible, for example, via network 22 by decoder 23. Decoder 23 comprises different circuits implementing different steps of decoding. Decoder 23 takes as input the data stream generated by encoder 21 and provides a sequence of 3D scenes 24 that are rendered and displayed by a volumetric video display device such as a head-mounted device (HMD). Decoder 23 acquires the stream from source 22. For example, source 22 is -Local memory such as, for example, video memory or RAM (or random access memory), flash memory, ROM (or read-only memory), hard disk, and -A storage interface such as, for example, an interface with mass storage, RAM, flash memory, ROM, optical disk or magnetic support, - For example, it belongs to a set including a communication interface such as a wired interface (e.g., a bus interface, a wide area network interface, a local area network interface) or a wireless interface (such as an IEEE802.11 interface or a Bluetooth (registered trademark) interface), and - a user interface such as a graphical user interface that enables a user to input data.

[0037] The decoder 23 includes a circuit 234 for extracting the data encoded in the data stream. The circuit 234 takes the data stream as an input and provides metadata 232 corresponding to the metadata 212 encoded in the stream and the two-dimensional video. The video is decoded by a video decoder 233 that provides a sequence of frames. The decoded frames include color and depth information. In a variant, the video decoder 233 provides a sequence of two frames, one including color information and the other including depth information. The circuit 231 uses the metadata 232 to provide a sequence of 3D scenes 24 without projecting the color and depth information from the decoded frames. The sequence of 3D scenes 24 corresponds to the sequence of 3D scenes 20 and video compression with a potentially reduced accuracy related to the encoding as a 2D video.

[0038] FIG. 3 shows an exemplary architecture of a device 30 that can be configured to implement the method described in relation to FIGS. 7 and 8. The encoder 21 and / or decoder 23 of FIG. 2 can implement this architecture. Alternatively, each circuit of the encoder 21 and / or decoder 23 can be a device with the architecture of FIG. 3, connected together, for example, via their bus 31 and / or via the I / O interface 36.

[0039] The device 30 includes the following elements connected together by a data and address bus 31: - For example, a microprocessor 32 (or CPU) which is a DSP (or digital signal processor), and - A ROM (or read-only memory) 33, and - A RAM (or random access memory) 34, and - A storage interface 35, and - An I / O interface 36 for receiving data to be transmitted from an application, and - A power source, for example, a battery, are provided.

[0040] According to one example, the power source is external to the device. In each of the memories mentioned, the word "register" as used herein can correspond to a small-capacity area (a few bits) or a very large area (e.g., an entire program or a large amount of received or decoded data). The ROM 33 includes at least programs and parameters. The ROM 33 can store algorithms and instructions for executing techniques according to this principle. When switched on, the CPU 32 uploads the program in the RAM and executes the corresponding instructions.

[0041] The RAM 34 includes, within the register, a program executed by the CPU 32 and uploaded after the device 30 is switched on, input data within the register, intermediate data of different states of the method within the register, and other variables used for the execution of the method within the register.

[0042] The implementations described in this specification may be implemented, for example, in a method or process, an apparatus, a computer program product, a data stream, or a signal. Even when considered in the context of a single form of implementation (e.g., only considered as a method or a device), the implementation of the features considered may also be implemented in other forms (e.g., a program). The apparatus may be implemented, for example, in appropriate hardware, software, and firmware. This method may be implemented, for example, in an apparatus such as a processor that generally refers to a processing device, including a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor may also include, for example, a communication device such as a computer, a mobile phone, a portable / personal digital assistant (``PDA''), and other devices that facilitate the communication of information between end users.

[0043] According to an embodiment, the device 30 is configured to implement the method described in connection with FIGS. 7 and 8, - a mobile device, and - a communication device, and - a game device, and - a tablet (or tablet computer), and - a laptop, and - a still camera, and - a video camera, and - an encoding chip, and - a server (e.g., a broadcast server, a video - on - demand server, or a web server), and belongs to a set including.

[0044] FIG. 4 shows an example of an embodiment of the syntax of a stream when data is transmitted via a packet-based transmission protocol. FIG. 4 shows an exemplary structure 4 of a volumetric video stream. The structure consists of containers that organize the stream in independent elements of the syntax. The structure may include a header portion 41 that is a set of data common to all syntax elements of the stream. For example, the header portion may include some metadata regarding the syntax elements, explaining the nature and role of each of them. The header portion may also include a part of the metadata 212 of FIG. 2, for example, the coordinates of the central viewing point used to project points of the 3D scene onto frames 2111 and 2112. The structure includes a payload that includes elements of syntax 42 and at least one element of syntax 43. The syntax element 42 includes data representing color and depth frames. The images may be compressed according to a video compression method.

[0045] The elements of syntax 43 are part of the payload of the data stream and may include metadata regarding how the frames of the elements of syntax 42 are encoded, for example, how points of the 3D scene are projected onto the frames, parameters used for packing. Such metadata may be associated with each frame of the video or a group of frames (also known as a Group of Pictures (GoP) in video compression standards).

[0046] 3DoF+ content may be provided as a set of Multi-View+Depth (MVD) frames. Such content may be captured by a dedicated camera or generated from existing computer graphics (CG) content by dedicated (potentially photorealistic) rendering.

[0047] FIG. 5 shows the process used by the view synthesizer 231 of FIG. 2 when generating an image for a given viewport from an MVD frame. When attempting to synthesize a pixel 51 for viewport 50 for synthesis, the synthesizer (e.g., circuit 231 of FIG. 2) does not project the light rays (e.g., light rays 52 and 53) passing through this given pixel, but checks the contributions of each source camera 54-57 along this light ray. As shown in FIG. 5, when some objects in the scene create occlusion from one camera to another, or when visibility cannot be ensured for camera settings, no consensus among all source cameras 54-57 regarding the characteristics of the pixel for synthesis may be found. In the example of FIG. 5, the first group of three cameras 54-56 "vote" to synthesize pixel 51 when they all "see" this object along the light ray using the color of the foreground object 58. The second group of a single camera 57 cannot see this object because it is outside its viewport. Thus, camera 57 "votes" for the background object 59 to synthesize pixel 51. A strategy for resolving the ambiguity of such a situation is to blend and / or merge the contributions of each camera by weight according to the distance to the viewport for synthesis. In the example of FIG. 5, the first group of cameras 54-56 provides the greatest contribution when they are more numerous and closer to the viewport for synthesis. Finally, pixel 51 is synthesized by using the characteristics of the foreground object 68 as expected.

[0048] Figure 6 shows the view synthesis of a set of cameras with non-uniform sampling of the 3D space. Depending on the configuration of the source camera rig, especially when the volume scene to be obtained is not optimally sampled, this weighting strategy can fail, as can be observed in Figure 6. In such a situation, the rig poorly samples the object to be captured because most of the input cameras cannot see it and a simple weighting strategy does not give the expected result. In the example of Figure 6, the foreground object 68 is captured only by camera 64. When trying to synthesize the pixel 61 for the viewport 60 for synthesis, the synthesizer does not project the rays (e.g., rays 62 and 63) passing through this given pixel and checks the contributions of each source camera 64, 66, and 67 along these rays. In the example of Figure 6, camera 64 uses the color of the foreground object 68 to synthesize pixel 61, while the group of cameras 66 and 67 votes for the background object 69 to synthesize pixel 61. Finally, the contribution of the color of the background object 69 is larger than the contribution of the color of the foreground object 68, resulting in visual artifacts.

[0049] Even if poor sampling of the scene to be obtained can be overcome at the capture stage by adapting the spatial configuration of the cameras, scenarios where the geometric shape of the scene cannot be predicted can occur, for example, in live streaming. Furthermore, in the case of natural scenes with complex motion and numerous potential occlusions, it is almost impossible to find a complete rig setup.

[0050] However, in some specific scenarios, especially when the virtual rig of the camera is used to capture computer generated (CG) 3D scenes, other weighting strategies can be envisioned beyond those previously presented where the virtual cameras are "perfect" and can be fully trusted. In fact, in a real (non-CG) context, depth information is not captured directly and, for example, needs to be pre-computed by photogrammetry, so it is necessary to estimate the MVD that functions as an input to the volumetric scene. This latter step is the source of many artifacts (especially the inconsistencies between the geometric information of remote cameras), and these need to be mitigated / are mitigated by a weighting / voting strategy similar to that described in Figure 5. Conversely, in a computer-generated scenario, the scene to be obtained is fully modeled and such artifacts cannot occur because the depth information is directly given by the model in a complete manner. If the synthesizer knows in advance that it should fully trust the information given by the source (View+Depth), then the process can be significantly accelerated and the weighting problem can be prevented as in that described in Figure 6.

[0051] According to this principle, a method for overcoming these drawbacks is proposed. The information is the inserted metadata sent to the decoder, indicating to the synthesizer that the camera used for synthesis is reliable and that alternative weighting should be assumed. The reliability of the information carried by each view of the multi-view frame is encoded in the metadata associated with the multi-view frame. The reliability is related to the faithfulness of the depth information when obtained. As detailed above, for the views captured by the virtual camera, the faithfulness of the depth information is maximum, and for the views captured by the real camera, the faithfulness of the depth information depends on the internal and external parameters of the real camera.

[0052] The implementation of such a feature can be done by inserting a flag into the camera parameter list within the metadata, as described in Table 1. This flag can be a boolean value per camera that enables a special profile of the view synthesizer, which can be considered to imply that a given camera is complete and its information should be considered fully reliable, as explained earlier.

[0053] The general flag "source_confidence_params_equal_flag" is set. This flag indicates whether to (if true) enable or (if false) disable the feature, and ii) when the latter flag is enabled, an array of boolean values "source_confidence" is inserted into the metadata, which indicates whether each component should be considered fully reliable (if true) or not (if false) per camera.

[0054]

Table 1

[0055] During the rendering stage, if a camera is identified as being fully reliable (the associated component of source_confidence is set to true), then its geometric information (depth value) overwrites all geometric information carried by other "less reliable" (i.e., normal) cameras. In that case, the weighting scheme can be advantageously replaced by a simple selection of the geometric shape (e.g., depth) information of the camera identified as reliable. In other words, in the weighting / voting scheme proposed in Figures 5 and 6, if a consensus on the position (foreground or background) of the points to be retained for the synthesis of a given pixel cannot be found between cameras whose source_confidence property is true and those whose source_confidence property is false, then those for which source_confidence is enabled are preferred.

[0056] For a given pixel to be synthesized, if this property of multiple cameras is enabled (the components associated with source_confidence are set to true), since it can be executed in the depth buffer of a normal rasterization engine, the camera with the minimum depth information is selected. Such a selection is not necessarily motivated by the fact that a given reliable camera, when viewing an object closer to a given pixel to be synthesized than other cameras, will necessarily create an occlusion for other cameras and thus carry information about the occluded further objects. In Figure 6, such a strategy will result in selecting the information carried by camera 64 for use in synthesizing pixel 61.

[0057] In another embodiment, non-binary values are used for source confidence, such as a normalized floating point number from 0 to 1, which indicates how "reliable" a camera should be considered in a rendering scheme.

[0058] In the real world environment, cameras are typically not considered to be completely reliable and complete. The terms "completely reliable" and "complete" generally refer to depth information. In a CG environment, since depth information is generated according to a model, it is known. Thus, the depth is known for all objects for all virtual cameras. Such virtual cameras are modeled as part of a virtual rig generated inside the CG environment. Thus, virtual cameras are completely reliable and complete.

[0059] In the example of FIG. 6, where the camera is part of the real-world system and depth is being estimated, the camera is not expected to be perfectly reliable and complete. Thus, if a significant weighting scheme is used for the pixels 61 of the viewport camera 60, then the resulting answer will be the background color of pixel 61. Similarly, if the camera is part of the virtual rig and is perfectly reliable and complete, most weighting schemes are still used, and then the background color is still selected for pixel 61. However, if the camera is part of the virtual rig and a perfectly reliable state is used, as a result, the minimum depth of the perfectly reliable camera is selected, and then the foreground color (from camera 64) is selected for pixel 61.

[0060] CG movies can benefit from the described embodiments. For example, a CG movie (e.g., The Lion King) can be re-shot using a virtual rig where multiple virtual cameras provide multiple views. The resulting output allows the user to have an immersive experience in the movie and select the viewing position. Rendering different viewing positions typically takes time. However, considering that the virtual cameras are perfectly reliable and complete (with respect to depth), for example, the rendering time can be reduced by having the minimum depth camera provide the color for a given pixel, or alternatively, providing the average value of the colors of closer depth values. This eliminates the processing typically required to perform the weighting operation.

[0061] The concept of trust can be extended to real-world cameras. However, relying on a single real-world camera based on estimated depth has the risk that the wrong color may be selected for any given pixel. However, if specific depth information for a given camera is more reliable, then this information can be utilized to shorten the rendering time, but it can also improve the final quality by relying on the "best" camera and thus avoiding possible artifacts.

[0062] Complementarily, in addition to complete geometric information, a "fully reliable" camera can also be used to carry the reliability of color information between different cameras of the rig. It is well known that calibrating different cameras with respect to color information is not always easy to achieve. Thus, also, the "fully reliable" camera concept can be used to identify the camera as a color reference and rely more on it at the color weighted rendering stage.

[0063] Figure 7 shows a method 70 for encoding a multi-view (MV) frame in a data stream according to a non-limiting embodiment of the present principle. In step 71, a multi-view frame is obtained from a source. In step 72, a parameter representing the reliability of the information carried by a given view of the multi-view frame is obtained. In one embodiment, the parameter is obtained for all views of the MV frame. This parameter can be a boolean value indicating whether the information of the view is fully reliable or "not fully" reliable. In a variant, the parameter is, for example, an integer between -100 and 100 or 0 and 255 or a real number, for example, a degree of reliability in the range of -1.0 to 1.0 or 0.0 to 1.0. In step 73, the MV frame is encoded in a data stream associated with metadata. The metadata includes pairs of data associating a view, for example an index, with its parameter.

[0064] Figure 8 shows a method 80 for decoding a multi-view frame from a data stream according to a non-limiting embodiment of the present principle. In step 81, the multi-view frame is decoded from a source. The metadata associated with this MV frame is also decoded from the stream. In step 82, pairs of data are obtained from the metadata, and these data associate the views of the MV frame with a parameter representing the reliability of the information carried by this view. In step 73, a viewport frame is generated for a viewing pose (i.e., a location and orientation within the 3D space of the renderer). For the pixels of the viewport frame, the weight of the contribution of each view (also called a "camera" in this application) is determined according to the reliability associated with each view.

[0065] The implementation forms described herein may be implemented, for example, in a method or process, an apparatus, a computer program product, a data stream, or a signal. Even when considered only in the context of a single form of implementation (e.g., only considered as a method or a device), the implementation forms of the features considered may also be implemented in other forms (e.g., a program). The apparatus may be implemented, for example, in appropriate hardware, software, and firmware. This method may be implemented, for example, in an apparatus such as a processor generally referring to a processing device, including a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor may also include a communication device such as a smartphone, a tablet, a computer, a mobile phone, a portable / personal digital assistant ("personal digital assistant, PDA"), and other devices that facilitate the communication of information between the end user.

[0066] The implementation of the various processes and features described herein can be embodied in a variety of different devices or applications, particularly in devices or applications associated with, for example, data encoding, data decoding, view generation, texture processing, and other processing of images and associated texture information and / or depth information. Examples of such devices include encoders, decoders, post-processors that process the output from a decoder, pre-processors that provide input to an encoder, video coders, video decoders, video codecs, web servers, set-top boxes, laptops, personal computers, mobile phones, PDAs, and other communication devices. As should be apparent, the devices can be mobile and can be installed in a mobile vehicle.

[0067] Furthermore, the method can be implemented by instructions executed by a processor, and such instructions (and / or data values generated by the implementation) can be stored, for example, on a processor-readable medium such as an integrated circuit, a software carrier, or other storage device, such as a hard disk, a compact diskette (CD), an optical disk (such as a DVD, often referred to as a digital versatile disk or digital video disk), a random access memory (RAM), or a read-only memory (ROM). The instructions can form an application program tangibly embodied on the processor-readable medium. The instructions can be, for example, hardware, firmware, software, or a combination. The instructions can be found, for example, in an operating system, a separate application, or a combination of the two. Thus, a processor can be characterized as both a device configured to execute a process and a device that includes a processor-readable medium (such as a storage device) having instructions for executing the process. Furthermore, the processor-readable medium can store data values generated by the implementation in addition to, or instead of, the instructions.

[0068] As will be apparent to those skilled in the art, the implementation forms can generate various signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for executing a method or data generated by one of the described implementation forms. For example, the signal can be formatted to carry as data rules for writing or reading the syntax of the described embodiments, or the actual syntax values written by the described embodiments. Such signals can be formatted, for example, as electromagnetic waves (e.g., using the radio frequency part of the spectrum) or as baseband signals. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog information or digital information. The signal can be transmitted via various different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.

[0069] Numerous implementation forms have been described. Nevertheless, it will be understood that various modifications can be made. For example, elements of different implementation forms can be combined, supplemented, modified, or deleted to generate other implementation forms. Further, those skilled in the art can substitute other structures and processes for those disclosed, and the resulting implementation forms will be understood to perform at least substantially the same functions in at least substantially the same ways to achieve at least substantially the same results as the disclosed implementation forms. Accordingly, these and other implementation forms are contemplated by this application.

Claims

1. A method for encoding a multi-view frame, comprising: - obtaining, for each view of the multi-view frame, a parameter representing the fidelity of the depth information carried by the view, wherein the parameter is a boolean value indicating whether the fidelity is completely reliable; - encoding the multi-view frame into a data stream in association with metadata including the parameter.

2. The method according to claim 1, wherein the parameter representing the fidelity of the depth information of the view is determined according to the internal and external parameters of the camera that captured the view.

3. The method according to claim 1, wherein the metadata includes information indicating whether a parameter is provided for each view of the multi-view frame, and when a parameter is provided for each view of the multi-view frame, encoding the parameter associated with each view for each view.

4. A device for encoding a multi-view frame, comprising: - obtaining, for each view of the multi-view frame, a parameter representing the fidelity of the depth information carried by the view, wherein the parameter is a boolean value indicating whether the fidelity is completely reliable; - a processor configured to encode the multi-view frame into a data stream in association with metadata including the parameter.

5. The device according to claim 4, wherein the processor is configured to determine the parameter representing the fidelity of the depth information of the view according to the internal and external parameters of the camera that captured the view.

6. The device according to claim 4, wherein the processor is configured to encode metadata including information indicating whether a parameter is provided for each view of the multi-view frame, and when a parameter is provided for each view of the multi-view frame, encoding the parameter associated with each view for each view.

7. A method for decoding a multi-view frame from a data stream, comprising: - decoding the multi-view frame and associated metadata from the data stream; - Obtaining information indicating whether a parameter representing the fidelity of depth information carried by a view of the multi-view frame is provided from the metadata, and when the parameter representing the fidelity is provided, obtaining the parameter for each view, where the parameter is a boolean value indicating whether the fidelity is completely reliable; - Generating a viewport frame according to the viewing pose by determining the contribution of each view of the multi-view frame as a function of the parameter associated with the view. A method comprising. **Claim 8** The method according to claim 7, wherein the contribution of a view that is not completely reliable is ignored. **Claim 9** The method according to claim 7, wherein the completely reliable view having the lowest depth information is used on the condition that a plurality of views are completely reliable. **Claim 10** The method according to claim 7, wherein the contribution of each view is proportional to a numerical value associated with the view. **Claim 11** A device for decoding a multi-view frame from a data stream, - Decoding the multi-view frame and associated metadata from the data stream; - Obtaining information indicating whether a parameter representing the fidelity of depth information carried by a view of the multi-view frame is provided from the metadata, and when the parameter representing the fidelity is provided, obtaining the parameter for each view, where the parameter is a boolean value indicating whether the fidelity is completely reliable; - A device comprising a processor configured to generate a viewport frame according to the viewing pose by determining the contribution of each view of the multi-view frame as a function of the parameter associated with the view. **Claim 12** The device according to claim 11, wherein the contribution of a view that is not completely reliable is ignored. **Claim 13** The device according to claim 11, wherein the completely reliable view having the lowest depth information is used on the condition that a plurality of views are completely reliable. **Claim 14** The device according to claim 11, wherein the contribution of each view is proportional to a numerical value associated with the view.

Citation Information

Patent Citations

  • System and method for encoding and decoding brightfield image files

    JP2014535191A

  • Encoder, Method in an Encoder, Decoder and Method in a Decoder for Providing Information Concerning a Spatial Validity Range

    US20140192148A1

  • Systems and methods for encoding light field image files having low resolution images

    US20150199794A1

  • Stereo photography device

    WO2013099169A1