Method and apparatus for encoding and decoding multi-view 3DoF+ content
By defining the viewing frame within a 3D scene and using the central view and peripheral blocks for encoding and decoding, the problem of encoding and decoding large field-of-view content is solved, achieving efficient data compression and a user-friendly experience.
Patent Information
- Application Number
- CN202080085578.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-11
- Filing Date
- 2020-11-30
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2040-11-30
AI Technical Summary
Existing technologies struggle to effectively encode and decode large field-of-view content, resulting in a poor user experience. This is especially true in 3DoF videos, where the large data volume leads to high storage and transmission requirements.
By defining a reference and a central viewing frame in a 3D scene, the reference center view, peripheral blocks, and metadata are encoded and decoded separately to generate volumetric video content, and conventional deprojection and projection techniques are used for data processing.
It significantly reduces data volume requirements, improves user experience, avoids dizziness, and keeps the complexity of encoding and decoding unchanged.
Smart Images

Figure CN114868396B_ABST
Abstract
Description
1. TECHNICAL FIELD
[0001] The present principles generally relate to the domain of three-dimensional (3D) scene and volumetric video content. The present document is also understood in the context of encoding, formatting and decoding data representative of textures and geometry of a 3D scene, to render the volumetric video content on an end-user device such as a mobile device or a Head-Mounted Display (HMD). 2. BACKGROUND
[0002] This section is intended to introduce the reader to various aspects of art that can be related to various aspects of the present principles that are described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present principles. Accordingly, it should be understood that these statements are to be read in this light, and not as admissions of prior art.
[0003] Recently, there has been a growth in available large field of view content (up to 360°). A user watching content on an immersive display device (such as a head-mounted display, smart glasses, PC screen, tablet, smartphone, etc.) can not be able to see the whole of such content. This means that at a given moment, the user can only watch a part of the content. However, the user can typically navigate within the content through various means such as head movement, mouse movement, touch screen, voice, and the like. It is generally desirable to encode and decode such content.
[0004] Immersive video (also called 360° planar video) allows a user to watch everything around him by rotating his head around a static viewpoint. The rotation only allows a 3 degrees of freedom (3DoF) experience. Even if 3DoF video is sufficient to meet the requirements of a first omnidirectional video experience (e.g. using a head-mounted display (HMD device)), 3DoF video can quickly become frustrating for a viewer expecting more freedom (e.g. by experiencing parallax). Moreover, 3DoF can also cause dizziness because a user never only rotates his head but also translates his head in three directions, which are not reproduced in a 3DoF video experience.
[0005] Among others, the large field of view content can be a three-dimensional computer graphics image scene (3D CGI scene), a point cloud or an immersive video. Many terms can be used to design such immersive video: for example, Virtual Reality (VR), 360, panoramic, 4π steradians, immersive, omnidirectional or large field of view.
[0006] Volume videos, also called 6 Degrees of Freedom (6DoF) videos, are an alternative to 3DoF videos. When watching a 6DoF video, in addition to rotations, the user can translate its head, and even its body, in the content of the watch, and experience parallax and even volume. This video significantly increases the immersion and the perception of the depth of the scene and prevents dizziness by providing a consistent visual feedback during head translation. The content is created by dedicated sensors that allow recording the color and the depth of the scene of interest simultaneously. Even if technical difficulties still exist, using a color camera equipment combined with photogrammetry techniques is one way to perform such recording.
[0007] While 3DoF videos consist in a sequence of images resulting from the de-mapping of texture images (e.g. spherical images encoded according to a latitude / longitude projection mapping or an equirectangular projection mapping), 6DoF video frames embed information from multiple viewpoints. They can be seen as a temporal sequence of point clouds resulting from a three-dimensional capture. Two kinds of volume videos can be considered depending on the viewing conditions. The first one, i.e. full 6DoF, allows a full freedom of navigation within the video content, while the second one, also called 3DoF+, limits the user viewing space to a limited volume called viewing bounding box, allowing a limited head translation and parallax experience. This second case is a valuable compromise between free navigation and passive viewing conditions for seated viewers.
[0008] The data volume of volume video content is important and requires large storage capacities, and high bitrates to transmit such data. Solutions for reducing the data volume corresponding to those volume videos for storage, transmission or decoding purposes represent a broad research object to be investigated. 3. SUMMARY
[0009] The following presents a simplified summary of the principles of the application in order to provide a basic understanding of some aspects of the application. This summary is not an extensive overview of the principles of the application. It is not intended to identify key or critical elements of the principles of the application. The following summary merely presents some aspects of the principles of the application in a simplified form as a prelude to the more detailed description provided below.
[0010] According to a first aspect, there is provided a method for encoding a 3D scene in a data stream. The method comprises:
[0011] - obtaining a reference viewing box and an intermediate viewing box defined within the 3D scene;
[0012] - encoding in the data stream a reference center view captured from a viewpoint at the center of the reference viewing box and reference peripheral patches encoding images captured from different viewpoints in the reference viewing box;
[0013] - encoding in the data stream at least one intermediate center patch, the at least one intermediate center patch encoding a difference between a view captured from a center of the intermediate viewing frame and the reference center view; and
[0014] - encoding in the data stream metadata describing the reference viewing frame and the intermediate viewing frame and the different viewpoints.
[0015] In one or more embodiments, the reference viewing frame is the closest to the intermediate viewing bounding frame among a set of reference viewing frames defined within the 3D scene (e.g. within a navigation distance inside the 3D scene). A reference peripheral patch can encode a difference between a peripheral image and the reference center view.
[0016] In an embodiment, the intermediate viewing bounding frame overlaps the reference viewing bounding frame. In another embodiment, the data stream encoding the 3D scene is transmitted to a client device, e.g. via a network.
[0017] There is also provided a device for encoding a 3D scene in a data stream. The device comprises means (e.g. a processor associated with a memory) for performing the method according to the first aspect.
[0018] According to a second aspect, there is also provided a method for retrieving a 3D scene from a data stream. The method comprises:
[0019] - decoding from the data stream:
[0020] • metadata describing a reference viewing frame and an intermediate viewing frame in the 3D scene;
[0021] • a reference center view, the reference center view being captured from a viewpoint at a center of the reference viewing frame;
[0022] • at least one intermediate center patch, the at least one intermediate center patch encoding a difference between a view captured from a center of the intermediate viewing frame and the reference center view;
[0023] - retrieving the 3D scene by de-projecting pixels of the reference center view and pixels of the at least one intermediate center patch.
[0024] In an embodiment, the method comprises:
[0025] - decoding from the data stream a reference peripheral patch, the reference peripheral patch encoding an image captured from a different viewpoint in the reference viewing frame;
[0026] - retrieving the 3D scene by de-projecting pixels of a subset of the reference peripheral patch.
[0027] The subset of reference peripheral patches can be selected according to a viewpoint located in the intermediate viewing frame. The reference viewing frame can be the closest reference viewing frame to the intermediate viewing boundary frame among a set of reference viewing boundary frames defined within the 3D scene.
[0028] In some embodiments, the method further comprises rendering a viewport image for a viewpoint located in the intermediate viewing frame.
[0029] There is also provided a device comprising means (e.g. a processor associated with a memory) for performing the method according to the second aspect.
[0030] There is also provided a non-transitory processor-readable medium having stored instructions for causing at least one processor to perform at least the steps of the method according to the first aspect or the second aspect, respectively. 4. BRIEF DESCRIPTION OF DRAWINGS
[0031] The disclosure will be better understood and other specific features and advantages will become apparent from reading the following description, given with reference to the drawings in which:
[0032] - Figure 1 A three-dimensional (3D) model of an object and points of a point cloud corresponding to the 3D model are shown, according to a non-limiting embodiment of the principles of the application;
[0033] - Figure 2 Non-limiting examples of encoding, transmitting and decoding data representative of a sequence of 3D scenes are shown, according to non-limiting embodiments of the principles of the application;
[0034] - Figure 3 An exemplary architecture of a device that can be configured to implement the methods described with respect to Figures 12 to 15 An exemplary architecture of a device that can be configured to implement the methods described with respect to
[0035] - Figure 4 An example of an implementation of the syntax of a data stream when transmitting data over a packet-based transmission protocol is shown, according to a non-limiting embodiment of the principles of the application;
[0036] - Figure 5 A spherical projection from a central viewpoint is shown, according to a non-limiting embodiment of the principles of the application;
[0037] - Figure 6 An example of an atlas comprising texture information for points of a 3D scene is shown;
[0038] - Figure 7 An example of an atlas comprising texture information for points of a 3D scene is shown; Figure 6an example of a chart of depth information of points of the same 3D scene as the 3D scene encoded in a chart of texture information;
[0039] - Figure 8 Aspects of a storage and streaming representation of volumetric video content representing a 3D scene are shown in accordance with non-limiting embodiments of the principles of the application.
[0040] - Figure 9 Aspects of steps for encoding intermediate volumetric video sub-content are shown in accordance with non-limiting embodiments of the principles of the application.
[0041] - Figure 10 Charts associated with volumetric video reference sub-content and volumetric video intermediate sub-content are shown in accordance with non-limiting embodiments of the principles of the application.
[0042] - Figure 11 Intermediate center tiles are shown in accordance with non-limiting embodiments of the principles of the application.
[0043] - Figure 12 A flowchart of a method for encoding volumetric video content related to a 3D scene is shown in accordance with non-limiting embodiments of the principles of the application.
[0044] - Figure 13 A flowchart of a method for transmitting volumetric video content related to a 3D scene is shown in accordance with non-limiting embodiments of the principles of the application.
[0045] - Figure 14 A flowchart of a method for decoding volumetric video content related to a 3D scene is shown in accordance with non-limiting embodiments of the principles of the application.
[0046] - Figure 15 A flowchart of a method for rendering volumetric video content related to a 3D scene is shown in accordance with non-limiting embodiments of the principles of the application.
[0047] - Figure 16 Aspects of a storage and streaming representation of volumetric video content representing a 3D scene are shown in accordance with non-limiting embodiments of the principles of the application. 5. DETAILED DESCRIPTION
[0048] The principles of the present application will be more fully understood in view of the detailed description and the drawings attached hereto. The principles of the present application can be embodied in many alternative forms and should not be construed as limited to the examples set forth herein. Accordingly, while the principles of the present application are susceptible to various modifications and alternative forms, specific examples thereof are shown by way of example in the drawings and will be described herein in detail. It should be understood, however, that there is no intent to limit the principles of the present application to the particular forms disclosed, but on the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the principles of the present application as defined by the claims.
[0049] The terminology used herein is for the purpose of describing particular examples only and is not intended to be limiting of the principles of the present application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Additionally, when an element is referred to as being "responsive" or "connected" to another element, it can be directly responsive or connected to the other element, or indirectly responsive or connected to the other element through one or more other elements. In contrast, when an element is referred to as being "directly responsive" or "directly connected" to another element, there are no intervening elements. As used herein the term "and / or" includes any and all combinations of one or more of the associated items and can be abbreviated as " / ".
[0050] It should be understood that, although the terms first, second, etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the teachings of the present principles.
[0051] Although some of the diagrams include arrows on communication paths to show a primary direction of communication, it is to be understood that communication can occur in the opposite direction to the depicted arrows.
[0052] Some examples are described with respect to block and operational flow diagrams that include blocks interconnecting with other blocks representing circuit elements, modules or code portions for performing specified functions. It should also be noted that in other specific implementations, the function(s) noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved.
[0053] Reference in the specification to "one example" or "an example" means that a particular feature, structure, or characteristic described in connection with the example can be included in at least one implementation of the present principles. The appearances of the phrase "one example" or "an example" in various places in the specification are not necessarily all referring to the same example, nor are they necessarily mutually exclusive, or alternative examples to one another.
[0054] The reference signs appearing in the claims are presented merely by way of illustration and have no limiting effect on the scope of the claims. Although not explicitly described, the present examples and variants can be taken in any combination or sub-combination.
[0055] The present principles will be described with reference to specific embodiments of a method for encoding volumetric video content related to a 3D scene in a data stream; a method for decoding such volumetric video content from a data stream; and a method for volumetric rendering of volumetric video content decoded according to the mentioned decoding method.
[0056] According to the present principles, from a 3D scene comprising a plurality of objects, an encoding method is implemented for generating an encoded volumetric video content carrying data representative of the 3D scene, e.g. a data stream. The encoding method generates volumetric information contained in the 3D scene. A reference 3DoF+ viewing bounding box and an intermediate 3DoF+ viewing bounding box are defined within a navigation space in the 3D scene.
[0057] For the reference 3DoF+ viewing bounding box, a volumetric video reference sub-content representative of a portion of the 3D scene is encoded using a reference central view and one or more reference peripheral patches. The reference central view encodes a central image of the volumetric video reference sub-content, while the one or more reference peripheral patches encode one or more peripheral images of the volumetric video reference sub-content.
[0058] From the central image of the volumetric video reference sub-content, it is possible to understand an image captured by a camera positioned at the center of the reference 3DoF+ viewing bounding box and oriented according to the main viewpoint of the reference 3DoF+ viewing bounding box. It is also possible to understand an image obtained by interpolation of two other images, as seen by a virtual camera positioned at the center of the reference 3DoF+ viewing bounding box and oriented according to the main viewpoint of the reference 3DoF+ viewing bounding box.
[0059] From the peripheral images, it is possible to understand an image captured by a camera having a different pose than the camera that captured the central image of the volumetric video reference sub-content and corresponding to a viewpoint comprised in the reference 3DoF+ viewing bounding box. It is also possible to understand an image obtained by interpolation of two other images, as seen by a virtual camera having a particular pose.
[0060] The term patch specifies a residual image that can result from the difference between two images.
[0061] According to the principles of the application, the encoding of the intermediate volumetric video sub-content, corresponding to the intermediate 3DoF+ viewing boundary box, is based on the differential encoding of the central image of the intermediate 3DoF+ viewing boundary box. This encoding is a relative encoding of the intermediate volumetric video sub-content with respect to a reference volumetric video sub-content. This reference volumetric video sub-content can correspond to the closest reference viewing boundary box to the intermediate viewing boundary box among a set of reference viewing boundary boxes defined within the navigation space in the 3D scene. The differential intermediate central image can be generated from a reference central view or a reference central image corresponding to the considered reference 3DoF+ viewing boundary box, for example by de-projection and re-projection of the reference central view or the reference central image.
[0062] The intermediate volumetric video sub-content is encoded by one or more residual patches, also called one or more intermediate central patches therein. For the intermediate 3DoF+ viewing boundary box, the volumetric video intermediate sub-content representing a portion of the 3D scene is encoded using at least one intermediate central patch encoding the difference between the central image of the volumetric video intermediate sub-content and the central image of the volumetric video reference sub-content, or more precisely the difference between the central image of the volumetric video intermediate sub-content and the intermediate central image of the volumetric video reference sub-content.
[0063] The central image of the volumetric video intermediate sub-content can be an image captured by a camera positioned at the center of the intermediate 3DoF+ viewing boundary box and oriented according to the main viewpoint of the intermediate 3DoF+ viewing boundary box. It can be an image obtained by interpolation of two other images, as seen by a virtual camera positioned at the center of the intermediate 3DoF+ viewing boundary box and oriented according to the main viewpoint of the intermediate 3DoF+ viewing boundary box.
[0064] The principles of the application allow to significantly reduce the amount of data to store and / or transmit and / or decode a volumetric video made of a plurality of 3DoF+ contents spatially arranged to enable a 6DoF experience.
[0065] Moreover, the encoding and decoding complexity is not increased compared to the independent encoding of the 3DoF+ contents: the central image is generated using regular de-projection and projection, for example using a graphics rendering pipeline, and the difference function used to extract the residual patches only involves basic pixel-wise comparison.
[0066] According to the principles of the present invention, a transmission method implemented in a streaming device is disclosed. A volumetric video content representative of a 3D scene encoded according to the encoding method presented above is obtained from a source, for example a memory. The position corresponding to a viewpoint within a guided airspace is considered, as well as a corresponding 3DoF+ viewing bounding box including this position. In the case where the corresponding 3DoF+ viewing bounding box is an intermediate viewing bounding box, then according to the method, at least one intermediate center tile encoding a volumetric video intermediate sub-content associated with the intermediate viewing bounding box and a reference center view encoding a center image of a volumetric video reference sub-content of the 3D scene for a given reference viewing bounding box within the guided airspace are transmitted.
[0067] According to the principles of the present invention, a decoding method implemented in a decoder is disclosed. The decoder obtains at least one intermediate center tile encoding a difference between, on the one hand, a center image of a volumetric video intermediate sub-content of a volumetric video content representative of a 3D scene and associated with an intermediate viewing bounding box within a guided airspace of the 3D scene, and, on the other hand, a center image of a volumetric video reference sub-content of the volumetric video content encoded for a given reference viewing bounding box. The decoder also obtains a reference center view encoding the center image of the volumetric video reference sub-content. Using the at least one intermediate center tile and the reference center view, the decoder generates a decoded volumetric video sub-content in the form of a point cloud.
[0068] The given reference viewing bounding box can be the closest reference viewing bounding box to the intermediate viewing bounding box among a set of reference viewing bounding boxes defined within the guided airspace in the 3D scene.
[0069] The generation of the decoded volumetric video sub-content in the form of a point cloud can be done as follows. The reference center view undergoes a 2D to 3D de-projection based on projection parameters and is de-projected into a temporary point cloud. Then, the temporary point cloud undergoes a 3D to 2D projection based on projection parameters on a corresponding volume plane in the direction of the main viewpoint of the 3DoF+ intermediate viewing bounding box. The viewpoints in the viewing bounding box are associated with volume planes. A volume plane is a 2D projection of volumetric information related to the 3D scene with respect to a virtual camera pose corresponding to a viewpoint associated with the volume plane. By 3D to 2D projection of the temporary point cloud, an intermediate center image is obtained. The corresponding pixels in the intermediate center image are then replaced using the at least one intermediate center tile obtained by the decoder to obtain a reconstructed intermediate center image. Then, the reconstructed intermediate center image undergoes a 2D to 3D de-projection based on projection parameters and is de-projected into a point cloud as seen from the main viewpoint of the 3DoF+ intermediate viewing bounding box. The projections considered herein are any type of projection or de-projection known for example in the field of graphics rendering. They transport a parameterization from 3D data to 2D data (mapping projection), or vice versa.
[0070] The decoder can also have access to one or more reference peripheral tiles encoding one or more peripheral images of the volumetric video reference sub-content. In this case, the decoder can generate a decoded volumetric video sub-content in the form of a point cloud viewed from a viewpoint different from the main viewpoint of the 3DoF+ intermediate viewing bounding box comprised in the 3DoF+ intermediate viewing bounding box. The generation of the corresponding decoded volumetric video sub-content is performed by using at least one of the intermediate central tile, the reference central view and at least one of the one or more reference peripheral tiles.
[0071] The generation of the corresponding decoded volumetric video sub-content can be performed as follows. As previously mentioned, the point cloud viewed from the main viewpoint of the 3DoF+ intermediate viewing bounding box is obtained from the reconstructed intermediate central image. Considering that at least one of the one or more reference peripheral tiles is used in order to generate another point cloud viewed from a current viewpoint comprised in the 3DoF+ intermediate viewing bounding box and different from the main viewpoint of the 3DoF+ intermediate viewing bounding box, it is related to the current viewpoint. The point cloud viewed from the main viewpoint of the 3DoF+ intermediate viewing bounding box undergoes a 3D to 2D projection based on the projection parameters in the direction of the current viewpoint to reconstruct a current central image. Then, the current central image is completed with pixels from at least one of the one or more reference peripheral tiles. Subsequently, the current central image undergoes a 2D to 3D de-projection based on the projection parameters to obtain the point cloud viewed from the current viewpoint.
[0072] According to the principles of the present invention, a method for rendering a volumetric video content representing a 3D scene is disclosed. An end user selects a viewpoint within a rendering 3D space. It is considered that the position corresponding to the viewpoint within the rendering 3D space, and the corresponding 3DoF+ viewing bounding box centered at this position. In case the corresponding 3DoF+ viewing bounding box is an intermediate viewing bounding box, then according to the method, a volumetric video intermediate sub-content of the volumetric video content representing the 3D scene and associated with the intermediate viewing bounding box is decoded according to the method presented above. The decoded volumetric video intermediate sub-content is then rendered on a rendering device.
[0073] Figure 1A three-dimensional (3D) model 10 of an object and points of a point cloud 11 corresponding to the 3D model 10 are shown. The 3D model 10 and the point cloud 11 can for example correspond to a possible 3D representation of an object of a 3D scene comprising other objects. The model 10 can be a 3D mesh representation and the points of the point cloud 11 can be the vertices of the mesh. The points of the point cloud 11 can also be points distributed on the surface of the faces of the mesh. The model 10 can also be represented as a splatted version of the point cloud 11, the surface of the model 10 being created by splatting the points of the point cloud 11. The model 10 can be represented by many different representations such as voxels or splines. Figure 1 The fact that a point cloud can be defined with a surface representation of a 3D object and that a surface representation of a 3D object can be generated from a cloud of points is shown. As used herein, projecting points of a 3D object (by extension points of a 3D scene) onto an image is equivalent to projecting any representation of this 3D object, for example a point cloud, a mesh, a spline model or a voxel model.
[0074] A point cloud can be represented in memory as for example a vector-based structure where each point has its own coordinates in the frame of reference of a viewpoint (for example three-dimensional coordinates XYZ, or an angle in space and a distance from / to the viewpoint (also called depth) and one or more attributes, also called components. One example of a component is a color component which can be represented in various color spaces, for example RGB (red, green and blue) or YUV (Y is the luminance component and UV are two chrominance components). The point cloud is a representation of a 3D scene comprising an object. The 3D scene can be seen from a given viewpoint or range of viewpoints. The point cloud can be obtained in many ways, for example:
[0075] • from a capture of a real object taken by a camera rig, optionally complemented with depth active sensing devices;
[0076] • from a capture of a virtual / synthetic object taken by a virtual camera rig in a modeling tool;
[0077] • from a mix of both real and virtual objects.
[0078] Figure 2 Non-limiting examples of encoding, transmitting and decoding data representing a sequence of 3D scenes are shown. The encoding format can for example be compatible with 3DoF, 3DoF+ and 6DoF decoding at the same time.
[0079] A sequence of 3D scenes 20 is obtained. As a sequence of pictures is a 2D video, a sequence of 3D scenes is a 3D (also called volumetric) video. The sequence of 3D scenes can be provided to a volumetric video rendering device for 3DoF, 3Dof+ or 6DoF rendering and display.
[0080] A sequence of 3D scenes 20 can be provided to an encoder 21. The encoder 21 takes as input one 3D scene or a sequence of 3D scenes and provides a data stream 22 representing the input. The data stream 22 can be stored in a memory and / or on an electronic data medium and can be transmitted over a network. The data stream 22 can be received and stored by a streaming device 26 configured to transmit the data stream 22 to a decoder 23. The data stream 22 representing a sequence of 3D scenes can be read from a memory and / or received over a network by the decoder 23 and / or on an electronic data medium. The decoder 23 is input by the data stream 22 and provides a sequence of 3D scenes in a point cloud format for example. This sequence of 3D scenes can be rendered by a rendering device 28.
[0081] The encoder 21 can comprise several circuits implementing several steps. In a first step, the encoder 21 projects each 3D scene onto at least one 2D picture. A 3D projection is any method of mapping three-dimensional points into a two-dimensional plane. Since most current methods for displaying graphical data are based on planar (pixel information from several bitplanes) two-dimensional media, the use of this type of projection is widespread, especially in computer graphics, engineering, and cartography. A projection circuit 211 provides a sequence of 3D scenes 20 with at least one two-dimensional frame 2111. The frame 2111 comprises color information and depth information representing a 3D scene projected onto the frame 2111. In variants, the color information and the depth information are encoded in two separate frames 2111 and 2112.
[0082] The metadata 212 is used and updated by the projection circuit 211. The metadata 212 comprises information on the projection operation (e.g. projection parameters) and information on the way the color and depth information are organized within the frames 2111 and 2112, as described in connection with the Figures 5 to 7
[0083] The video encoding circuit 213 encodes the sequence of frames 2111 and 2112 into a video. The frames 2111 and 2112 of a 3D scene (or a sequence of frames of a 3D scene) are encoded in a data stream by the video encoder 213. Then, the video data and the metadata 212 are encapsulated in a data stream by the data encapsulation circuit 214.
[0084] The encoder 213 is compatible with encoders such as:
[0085] - JPEG, specification ISO / CEI 10918-1 UIT-T Recommendation T.81, https: / / www.itu.int / rec / T-REC-T.81 / en;
[0086] - AVC, also called MPEG-4 AVC or h264. Specified in both UIT-T H.264 and ISO / CEI MPEG-4 Part 10 (ISO / CEI 14496-10), http: / / www.itu.int / rec / T-REC-H.264 / en, HEVC (whose specification is found on the ITU website, T recommendation, H series, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en) ;
[0087] - 3D-HEVC (extension of HEVC, whose specification is found on the ITU website, T recommendation, H series, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en annex G and I) ;
[0088] - VP9 developed by Google; or
[0089] - AV1 (AOMedia Video 1) developed by the Alliance for Open Media.
[0090] The data stream 22 is stored in a memory accessible by the decoder 23, for example through a network. The decoder 23 comprises different circuits implementing different decoding steps. The decoder 23 takes as input the data stream generated by the encoder 21 and provides a sequence of 3D scenes 24 to be rendered and displayed by a volumetric video display device such as a head-mounted device (HMD). The decoder 23 obtains the data stream from a source 22. For example, the source 22 belongs to a group comprising:
[0091] - a local memory, for example a video memory or a RAM (or Random Access Memory), a flash memory, a ROM (or Read Only Memory), a hard disk;
[0092] - a storage interface, for example an interface with a mass storage device, a RAM, a flash memory, a ROM, an optical or magnetic support;
[0093] - a communication interface, for example a wired interface (for example a bus interface, a wide area network interface, a local area network interface) or a wireless interface (such as an IEEE 802.11 interface or an interface); and
[0094] - a user interface enabling a user to input data, such as a graphical user interface.
[0095] The decoder 23 comprises a circuit 234 for extracting the data encoded in the data stream. The circuit 234 takes as input the data stream and provides metadata 232 corresponding to the metadata 212 encoded in the data stream and a two-dimensional video. The video is decoded by a video decoder 233 providing a sequence of frames. The decoded frames comprise color and depth information. In a variant, the video decoder 233 provides two sequences of frames, one containing color information and the other containing depth information. The circuit 231 uses the metadata 232 to de-project the color and depth information from the decoded frames to provide a sequence of 3D scenes 24. The sequence of 3D scenes 24 corresponds to the sequence of 3D scenes 20, possibly with a loss of precision related to the encoding as a 2D video and to the video compression.
[0096] The principles of the application disclosed herein concern a video encoding method 213, the metadata 212 generated and required for the 3D to 2D projection step 211, and the encoder 21.
[0097] They also concern a video decoding method 233, the metadata 232 received and used for the 2D to 3D de-projection step 231, and the decoder 23.
[0098] Figure 3 An exemplary architecture of a device 30 that can be configured to implement the method described with reference to any of the Figures 12 to 15 is shown. Figure 2 The encoder 21 and / or the decoder 23 can implement this architecture. Alternatively, each circuit of the encoder 21 and / or the decoder 23 and / or of the streaming device 26 and / or of the rendering device 28 can be a device according to the architecture of Figure 3 , linked together, for example via their buses 31 and / or via their I / O interfaces 36.
[0099] The device 30 comprises the following elements connected together by data and address buses 31 :
[0100] - a microprocessor 32 (or CPU), which is for example a DSP (or Digital Signal Processor);
[0101] - a ROM (or Read Only Memory) 33;
[0102] - a RAM (or Random Access Memory) 34;
[0103] - a storage interface 35;
[0104] - an I / O interface 36 for receiving data to be transmitted from an application; and
[0105] - a power supply, for example a battery.
[0106] According to an example, the power supply is external to the device. In each of the mentioned memories, the word "register" used in the description can correspond to a small capacity area (some bits) or a very large area (for example, the entire program or a large amount of received or decoded data). The ROM 33 comprises at least the program and the parameters. The ROM 33 can store, according to the principles of the application, algorithms and instructions for performing the technique. When switched on, the CPU 32 uploads the program in the RAM and executes the corresponding instructions.
[0107] The RAM 34 comprises the program in the registers executed by the CPU 32 and uploaded after the switch-on of the device 30, the input data in the registers, the intermediate data in the different states of the method in the registers, and other variables for executing the method in the registers.
[0108] The specific implementations described herein can be implemented in, for example, a method or process, an apparatus, a computer program product, a data stream, or a signal. Even if discussed in the context of a single form of implementation (for example, discussed only as a method or as an apparatus), the implementation of the discussed features can also be implemented in other forms (for example, program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The method can be implemented in, for example, an apparatus that is generally referred to as a processing device, such as, for example, a processor, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, such as, for example, a computer, a cell phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate communication of information between end users.
[0109] According to an example, the device 30 is configured to implement the method described with reference to each of Figures 12 to 15 may belong to the set comprising:
[0110] - a mobile device;
[0111] - a communication device;
[0112] - a gaming device;
[0113] - a tablet (or tablet computer);
[0114] - a laptop;
[0115] - a still picture camera;
[0116] - a video camera;
[0117] - an encoding chip;
[0118] - a server (for example, a broadcast server, a video on demand server, or a web server).
[0119] Figure 4An example of an implementation of the syntax of a data stream when transmitting data through a packet-based transmission protocol is shown. Figure 4 An exemplary structure 4 of a volumetric video data stream is shown. The structure is contained in a container that organizes the data stream in independent syntax elements. The structure can comprise a header part 41 that is a set of data common to each syntax element of the data stream. For example, the header part includes some metadata about the syntax elements, describing the nature and the role of each of them. The header part can also include Figure 2 a part of the metadata 212, for example the coordinates of the center viewpoint used to project the points of the 3D scene onto the frames 2111 and 2112.
[0120] In the principles of the application, Figure 2 The metadata 212 can include the position and the size of the reference viewing boundary box and of the intermediate viewing boundary box defined in the navigation space of the 3D scene to be encoded, transmitted, decoded and rendered. They can also include projection parameters such as 3D to 2D projection parameters or 2D to 3D de-projection parameters. The projection parameters can be called reference projection parameters when they refer to the reference viewing boundary box. The projection parameters can be called intermediate projection parameters when they refer to the intermediate viewing boundary box. The projections considered herein are any type of projection or de-projection known in the field of graphics rendering for example. They deliver a parameterization from 3D data to 2D data (mapping projection) or vice versa.
[0121] The structure comprises a payload that includes syntax elements 42 and at least one syntax element 43. The syntax elements 42 include data representative of color and depth frames. The images can have been compressed according to a video compression method.
[0122] The syntax element 43 is part of the payload of the data stream and can include metadata about how the frames of the syntax elements 42 have been encoded, for example parameters used to project and pack the points of the 3D scene onto the frames. Such metadata can be associated with each frame or group of frames of the video, also called group of pictures (GoP) in video compression standards.
[0123] Figure 5A tile atlas approach is shown, which can be used to encode 3DoF+ volumetric video content associated with a 3D scene, associated with a 3DoF+ viewing frustum. The tile atlas comprises tiles, i.e. residual images that can result from the difference between two images. The tiles encode volumetric information of the 3DoF+ volumetric video content for different areas of a portion from the 3D scene represented in the 3DoF+ volumetric video content. The tiles are obtained by a 3D to 2D projection onto a projection center. The 3D to 2D projection can be any type known in the field of graphics rendering, for example. A center view can be included in the atlas, corresponding to a 3D to 2D projection in the direction of a main viewpoint of the 3DoF+ viewing frustum, which can coincide with the center of the 3DoF+ viewing frustum. Such a center view can comprise a portion of the 3D scene visible from the main viewpoint. Small peripheral tiles can be included in the atlas, corresponding to a 3D to 2D projection in the direction of a viewpoint different from the main viewpoint of the 3DoF+ viewing frustum. The small peripheral tiles can comprise portions of the 3D scene not visible from the main viewpoint.
[0124] The center view encodes a central image (e.g. non-residual image) of the 3D scene as seen from the main viewpoint of the 3DoF+ viewing frustum.
[0125] In Figure 5 an example of 4 projection centers is shown. The 3D scene 50 comprises a person. For example, the projection center 51 is a perspective camera and the camera 53 is an orthographic camera. The cameras can also be omnidirectional cameras with e.g. spherical mapping (e.g. equirectangular mapping) or cubic mapping. According to the projection operation described in the projection data of the metadata, the 3D points of the 3D scene are projected onto a 2D plane associated with a virtual camera located at the projection center. In Figure 5 the example, the projection of the points captured by the camera 51 is mapped onto the tile 52 according to a perspective mapping and the projection of the points captured by the camera 53 is mapped onto the tile 54 according to an orthographic mapping.
[0126] The clustering of the projected pixels results in a plurality of 2D tiles, which are packed in a rectangular atlas 55. The organization of the tiles within the atlas defines an atlas layout. In embodiments, two atlases with the same layout: one for texture (i.e. color) information and one for depth information. Two tiles captured by the same camera or by two different cameras can comprise information representing the same portion of the 3D scene, like e.g. tiles 54 and 56.
[0127] The packing operation produces patch data for each generated patch. The patch data includes a reference to the projected data (e.g. an index in a table of projected data or a pointer to the projected data (i.e. an address in memory or in a data stream)) and information describing the position and size of the patch within the atlas (e.g. top-left coordinates, size and width in pixels). The patch data items are added to the metadata to be encapsulated in the data stream in association with the compressed data of one or two atlases.
[0128] Figure 6 An example of an atlas 60 comprising texture information (e.g. RGB data or YUV data) of points of a 3D scene is shown according to a non-limiting embodiment of the principles of the application. As explained with respect to Figure 5 The atlas is a packed patch of aggregated images with or without a central view. Within the atlas, the central view can also be referred to as the central patch, although the central view is typically not a residual image but a full image of the 3D scene.
[0129] In the example of Figure 6 , the atlas 60 comprises a first portion 61 comprising texture information of points of the 3D scene visible from the viewpoint and one or more second portions 62. The texture information of the first portion 61 can be obtained for example according to an equirectangular projection mapping, which is an example of a spherical projection mapping. In the example of Figure 6 , the second portions 62 are arranged at the left and right borders of the first portion 61, but the second portions can be arranged differently. The second portions 62 comprise texture information of portions of the 3D scene complementary to the portion visible from the viewpoint. The second portions can be obtained by removing from the 3D scene the points visible from the first viewpoint (whose texture is stored in the first portion) and projecting the remaining points according to the same viewpoint. The latter process can be iteratively repeated to obtain each time a hidden portion of the 3D scene. According to a variant, the second portions can be obtained by removing from the 3D scene the points visible from the viewpoint (e.g. the central viewpoint) (whose texture is stored in the first portion) and projecting the remaining points according to a viewpoint different from the first viewpoint, for example from one or more second viewpoints centered around the central viewpoint (e.g. a viewing space of a 3DoF rendering).
[0130] The first portion 61 can be seen as a first large texture patch (corresponding to the first portion of the 3D scene) and the second portions 62 comprise smaller texture patches (corresponding to the second portions of the 3D scene complementary to the first portion). Such an atlas has the advantage of being compatible with both 3DoF rendering (when only the first portion 61 is rendered) and 3DoF+ / 6DoF rendering.
[0131] Figure 7 An example of an atlas 60 comprising texture information (e.g. RGB data or YUV data) of points of a 3D scene is shown according to a non-limiting embodiment of the principles of the application. As explained with respect to Figure 6an example of an atlas 70 of depth information of points of a 3D scene. The atlas 70 can be seen as corresponding to Figure 6 a depth image of the texture image 60.
[0132] The atlas 70 comprises a first part 71 comprising depth information of points of the 3D scene visible from a central viewpoint and one or more second parts 72. The atlas 70 can be obtained in the same way as the atlas 60 but contains depth information associated with points of the 3D scene instead of texture information.
[0133] 6DoF volumetric video content can be represented by a set of multiple 3DoF+ volumetric video contents at discrete viewing positions.
[0134] Figure 8 It is shown that volumetric video content representing a 3D scene is stored and streamed for 6DoF rendering from previously represented volumetric video content. As shown at the top of Figure 8 At discrete viewing positions in a guided airspace 804 of the 3D scene, a plurality of 3DoF+ viewing bounding boxes 81 is defined and volumetric video sub-contents are encoded for each of the 3DoF+ viewing bounding boxes. In this example, the 3DoF+ viewing bounding boxes do not overlap. An end user 80 can move within the guided airspace of the 3D scene. When the end user 80 enters a 3DoF+ viewing bounding box 81 1, the viewing position is updated 801 and the encoded volumetric video sub-content associated with the newly entered 3DoF+ viewing bounding box 81 1 is obtained by a decoder. The encoded volumetric video content associated with the 3DoF+ viewing bounding box 81 is then decoded in step 802 to synthesize a view rendered on a rendering device used by the end user in step 803.
[0135] In 3DoF+ rendering, the user can move the viewpoint within the 3DoF+ viewing bounding box. This enables the experience of parallax. Data representing the part of the 3D scene visible from any viewpoint of the 3DoF+ viewing bounding box is included in the volumetric video sub-content of the volumetric video content representing the entire 3D scene and associated with the 3DoF+ viewing bounding box, including data representing the 3D scene visible from the previously mentioned main viewpoint.
[0136] Typically, the volumetric video sub-content associated with a 3DoF+ viewing bounding box is encoded in the form of an atlas having a central view and peripheral patches. The central view encodes a central image captured by a camera positioned at the center of the 3DoF+ viewing bounding box and oriented according to the so-called main viewpoint of the 3DoF+ viewing bounding box. The peripheral patches encode peripheral images captured by cameras having a different pose than the camera capturing the central image of the volumetric video reference sub-content and corresponding to viewpoints included in the 3DoF+ viewing bounding box.
[0137] The term "surrounding patches" designates the residual images that can result from the difference between two images. Indeed, for each viewing frustum, each of the surrounding patches encodes the difference between the peripheral image and the central image. In particular, the surrounding patches comprise de-occultation data or disparity information that are used when the current viewpoint is changed due to, for example, a shift of the user from the main viewpoint. The central view and the surrounding patches can be packed in one atlas (or two atlases with the same layout, one comprising texture (or color) data and the other one comprising depth data).
[0138] When staying inside a 3DoF+ viewing frustum, the end user will have access to all the accessible volumetric information representing the rendered 3D scene. When going out of the 3DoF+ viewing frustum, the volumetric information relative to the 3D scene will be lost if additional information is not available and, in particular, the central view of the exited 3DoF+ viewing frustum will no longer be appropriate. The end user will need to enter another 3DoF+ viewing frustum to recover the volumetric rendering. Therefore, interpolating data between two non-overlapping 3DoF+ viewing frustums will be of interest.
[0139] A 3DoF+ viewing frustum among an initial set of non-overlapping 3DoF+ viewing frustums defined within the navigation space (or 3D rendering space) of the 3D scene can be referred to as a reference viewing frustum. Between those reference viewing frustums, so-called intermediate viewing frustums can be defined. Volumetric video sub-contents representing volumetric video content of the 3D scene can be associated with viewing frustums (as reference viewing frustums or intermediate viewing frustums). A volumetric video sub-content associated with a reference viewing frustum will be referred to as a volumetric video reference sub-content. A volumetric video sub-content associated with an intermediate viewing frustum will be referred to as a volumetric video intermediate sub-content.
[0140] The purpose of the encoding method (respectively, the transmission method and the decoding method) according to the principles of the application is to reduce the amount of data to be encoded (respectively, transmitted and decoded) by reducing the amount of data used to encode the volumetric video sub-contents associated with each of the one or more intermediate viewing frustums.
[0141] While the volumetric video reference sub-contents related to the 3D scene are encoded using central views and surrounding patches as described above, the volumetric video intermediate sub-contents related to the 3D scene are encoded differently according to the principles of the application.
[0142] Figure 9 An aspect of performing steps for encoding intermediate volumetric video sub-contents of an intermediate viewing frustum is illustrated.
[0143] Encoding the intermediate volumetric video sub-content is based on the encoding of a central image of the intermediate volumetric video sub-content, also referred to herein as an intermediate central image. This encoding is a relative encoding of the intermediate volumetric video sub-content with respect to a reference volumetric video sub-content, for example a reference volumetric video sub-content corresponding to the reference viewing bounding box closest to the considered intermediate viewing bounding box. The intermediate volumetric video sub-content that cannot come from the reference volumetric video sub-content is encoded by residual patches.
[0144] More precisely, as illustrated in Figure 9 the at least one intermediate central image (color image 930 and / or depth image 931) is obtained from the at least one reference central image (color image 910 and / or depth image 911) by using 2D to 3D de-projection and 3D to 2D re-projection to generate at least one intermediate central image (color image 920 and / or depth image 921): the reference central image of the reference volumetric video sub-content is thus warped to be viewed from the main viewpoint of the intermediate viewing bounding box.
[0145] The reference central image can be an image captured by a real camera whose pose corresponds to the main viewpoint of the reference viewing bounding box and which is placed at the main center of the reference viewing bounding box. The reference central image can also be an interpolation of two images. The reference central image can also be a reference central view encoded in the data stream for the considered reference viewing bounding box.
[0146] The projection and de-projection of the reference central image can be performed for example by a graphics rendering pipeline. In the intermediate central image, points that are visible from the main viewpoint of the intermediate viewing bounding box but not from the main viewpoint of the reference viewing bounding box are lost.
[0147] The intermediate central image (color image 921 and / or depth image 922) is generated by projection of the 3D scene in the direction of the main viewpoint of the intermediate viewing bounding box. For example, this intermediate central image can be an image captured by a real camera whose pose corresponds to the main viewpoint of the intermediate viewing bounding box and which is placed at the center of the intermediate viewing bounding box. The intermediate central image can also be an interpolation of two images.
[0148] Only the pixels of the intermediate central image that do not match the pixels of the intermediate central image (color image 910, depth image 911) are kept and encoded by residual patches (940, 941).
[0149] As illustrated in Figure 9The 2D to 3D de-projection of the reference center image is shown from left to right in the direction of the reference primary viewpoint 900 (here the reference center viewpoint) with respect to the reference viewing bounding box, followed by a 3D to 2D projection in the direction of the intermediate primary viewpoint 901 (here the intermediate center viewpoint 901) with respect to the intermediate viewing bounding box to generate the intermediate center image (920, 921).
[0150] The difference image is then computed as the difference between the intermediate center image and the intermediate center image. On the right side, from the difference image, only one or more residual patches are kept for encoding the intermediate volumetric video sub-content. The amount of data is reduced compared to encoding the intermediate volumetric video sub-content by a patch atlas as described above for the reference volumetric video sub-content.
[0151] The difference between the intermediate center image and the intermediate center image can be performed by using a pixel-wise difference function. In a first embodiment, only the absolute difference between depth values is considered. In a second embodiment, in addition to the difference between depth values, the absolute difference between color values is considered.
[0152] In one or more embodiments, a threshold is determined and only when the pixel value of the difference image is above the defined threshold, the corresponding depth value and color value of this difference image is kept and encoded into a residual patch. These parts of the difference image correspond to (typically small) parts of the 3D scene that are not seen in the reference volumetric video sub-content or seen in different colors in case of specular reflection and / or directional lighting. These kept pixels of the difference image are clustered into residual patches that are further packed into a (small size) residual atlas. The location of the residual patch within the intermediate center patch is stored.
[0153] Further iterations of the 3D scene peeling process that capture scene parts occluded from the center viewpoint coming from an offset position (to achieve parallax) occur and can produce additional residual patches.
[0154] The considered reference viewing bounding box can be the closest to the intermediate viewing bounding box among a set of reference viewing bounding boxes defined within the 3D scene. The considered reference viewing bounding box and the intermediate viewing bounding box can overlap.
[0155] Figure 10 An example of a color patch atlas 1000 encoding the volumetric video reference sub-content for the reference viewing bounding box is shown on the left part of it, while on the right part of it, the center image 1010 viewed from a viewpoint 10 cm to the right of the center viewpoint corresponding to the intermediate viewing bounding box is shown.
[0156] Figure 11The result of the pixel-wise difference 1100 between the intermediate center image and the corresponding intermediate center image is shown on the left side. From this difference image, intermediate center patches 1110 can be extracted, as shown on the right side of the figure. Those intermediate center patches are residual patches that can be packed in a smaller size atlas, also called residual atlas in the present document. This residual atlas does not comprise any center view, only residual patches. Figure 11
[0157] Figure 12 A method for encoding volumetric video content related to a 3D scene according to non-limiting embodiments of the principles of the application is illustrated. The steps of the method can be performed by the device 30 described with reference to Figure 3 and / or the encoder 21 described with reference to Figure 2 .
[0158] In a step 1200, different parameters of the device 30 are updated. In particular, a 3D scene is obtained from a source.
[0159] In a step 1201, a reference viewing boundary box and an intermediate viewing boundary box defined within a viewing volume in the 3D scene are obtained.
[0160] In a step 1202, two sub-steps 1202A and 1202B are performed:
[0161] - in a step 1202A, encoding a volumetric video reference sub-content using a reference center view and one or more reference peripheral patches. The reference center view encodes a center image of the volumetric video reference sub-content and the one or more reference peripheral patches encode one or more peripheral images of the volumetric video reference sub-content.
[0162] - in a step 1202B, encoding a volumetric video intermediate sub-content by at least one intermediate center patch encoding a difference between a center image of the volumetric video intermediate sub-content and a center image of the volumetric video reference sub-content.
[0163] Figure 13 A method for transmitting volumetric video content related to a 3D scene according to non-limiting embodiments of the principles of the application is illustrated. The steps of the method can be performed by the device 30 described with reference to Figure 3 and / or the streaming device 26 described with reference to Figure 2 .
[0164] In a step 1300, different parameters of the device 30 are updated. In particular, a volumetric video content related to a 3D scene encoded according to the encoding method presented in the present document is obtained from a source, for example a memory.
[0165] In a step 1301, a position corresponding to a viewpoint within the 3D scene is obtained.
[0166] In step 1302, a current viewing boundary box corresponding to the intermediate viewing boundary box comprising the position obtained at step 1301 is obtained.
[0167] In step 1303, at least one intermediate center patch encoding a volumetric video intermediate sub-content associated with the intermediate viewing boundary box and a reference center view encoding a volumetric video reference sub-content for a reference viewing boundary box are transmitted.
[0168] Figure 14 A method for decoding a volumetric video intermediate sub-content representative of a volumetric video content of a 3D scene according to a non-limiting embodiment of the principles of the application is illustrated. The steps of the method can be performed by a decoder 23 as described Figure 3 The apparatus 30 and / or the decoder 23 as described are described. Figure 2 The decoder 23 as described performs.
[0169] In step 1400, different parameters of the apparatus 30 are updated. In particular, for an intermediate viewing boundary box in the 3D scene, at least one intermediate center patch encoding a difference between a center image of the volumetric video intermediate sub-content and a center image of a volumetric video reference sub-content of the volumetric video content encoded for a reference viewing boundary box is obtained. The reference viewing boundary box can be the closest to the intermediate viewing boundary box among a set of reference viewing boundary boxes defined within a navigation space of the 3D scene.
[0170] In step 1401, a reference center view encoding the center image of the volumetric video reference sub-content is obtained.
[0171] In step 1402, the decoded volumetric video sub-content in the form of a point cloud is generated from the at least one intermediate center patch and the reference center view.
[0172] Step 1402 can comprise the following sub-steps 1402A, 1402B, 1402C and 1402D:
[0173] - in sub-step 1402A, the reference center view undergoes a 2D to 3D de-projection and is de-projected onto a temporary point cloud;
[0174] - in sub-step 1402B, the temporary point cloud undergoes a 3D to 2D projection to obtain an intermediate center image;
[0175] - in sub-step 1402C, the corresponding pixels within the intermediate center image are replaced with the at least one intermediate center patch to obtain a reconstructed intermediate center image;
[0176] - in a sub-step 1402D, the reconstructed intermediate central image undergoes 2D to 3D de-projection to obtain a point cloud corresponding to a central image in the intermediate viewing bounding box.
[0177] After step 1402 can be an additional step 1402' of obtaining metadata related to the intermediate volumetric video sub-content. The metadata can comprise the position of the center of the reference viewing bounding box and the center of the intermediate viewing bounding box, the reference projection parameters and the intermediate projection parameters. In sub-step 1402A, the center of the reference viewing bounding box and the reference projection parameters are used. In sub-steps 1402B and 1402D, the center of the intermediate viewing bounding box and the intermediate projection parameters are used.
[0178] The decoding method can also comprise the following additional steps 1403 and 1404.
[0179] In step 1403, one or more reference peripheral tiles encoding one or more peripheral images of the volumetric video reference sub-content are obtained.
[0180] In step 1404, for a peripheral image in the intermediate viewing bounding box, the decoded volumetric video sub-content in the form of a point cloud is generated from at least one of the at least one intermediate central tile, the reference central view and one or more reference peripheral tiles.
[0181] Step 1404 can comprise the following sub-steps:
[0182] - sub-step 1404A corresponding to sub-step 1402A for obtaining a temporary point cloud;
[0183] - sub-step 1404B corresponding to sub-step 1402B for obtaining an intermediate central image;
[0184] - sub-step 1404C corresponding to sub-step 1402C for obtaining a reconstructed intermediate central image;
[0185] - sub-step 1404D in which the one or more reference peripheral tiles and the reconstructed intermediate central image are used to reconstruct a current central image in the intermediate viewing bounding box;
[0186] - sub-step 1404E in which the current central image undergoes 2D to 3D de-projection to obtain a point cloud corresponding to a peripheral image in the intermediate viewing bounding box.
[0187] Figure 15 A method for rendering a volumetric video content representing a 3D scene is illustrated according to a non-limiting embodiment of the principles of the application. The steps of the method can be performed by the rendering device 28 described with reference to Figure 3 the device 30 described and / or with reference to Figure 2 the rendering device 28 described.
[0188] In step 1500, different parameters of the updating device 30 are updated. In particular, the first viewpoint within the rendered 3D space is obtained.
[0189] In step 1501, the intermediate volumetric video sub-content of the volumetric video content is decoded according to the decoding method presented above.
[0190] In step 1502, the decoded intermediate volumetric video sub-content is rendered.
[0191] Figure 16 Volumetric video contents representing 3D scenes are shown. 3DoF+ viewing boundary boxes 160, 161, 162 and 163 are represented. 3DoF+ viewing boundary boxes 160 and 163 can be 3DoF+ reference viewing boundary boxes, while 3DoF+ viewing boundary boxes 161 and 162 can be 3DoF+ intermediate viewing boundary boxes. 3DoF+ viewing boundary boxes 160, 161, 162 and 163 are associated with 3DoF+ volumetric video sub-contents 1600, 1601, 1602 and 1603. 3DoF+ volumetric video sub-contents 1600 and 1603 can be reference 3DoF+ volumetric video sub-contents, while 3DoF+ volumetric video sub-contents 1601 and 1602 can be intermediate 3DoF+ volumetric video sub-contents. As previously mentioned, 3DoF+ volumetric video sub-contents 1600 and 1603 can be encoded using reference center views and peripheral patches. According to the principles presented previously, 3DoF+ volumetric video sub-contents 1601 and 1602 can be encoded using at least one intermediate center patch.
[0192] When the end-user follows the path as Figure 16The 3D path illustrated successively through positions X equal to 0, 1, 2 and 3 can possibly undergo a 6DoF rendering. From the position X equal to 0, the end user is positioned in the 3DoF+ viewing bounding box 160. The 3DoF+ volumetric video sub-content 1600 can be transmitted to the rendering device used by the end user, decoded and rendered. Then when moving to the position X equal to 1, the end user enters the 3DoF+ viewing bounding box 161. The 3DoF+ volumetric video sub-content 1601 can be transmitted to the end user's rendering device according to the transmission method disclosed above, then decoded and rendered according to the decoding method and the rendering method according to the principles presented above, respectively. Then when moving to the position X equal to 2, the end user enters the 3DoF+ viewing bounding box 162. The 3DoF+ volumetric video sub-content 1602 can be transmitted to the end user's rendering device according to the transmission method disclosed above, then decoded and rendered according to the decoding method and the rendering method according to the principles presented above. Next, when moving to the position X equal to 3, the end user enters the 3DoF+ viewing bounding box 163. The 3DoF+ volumetric video sub-content 1603 can be transmitted to the end user's rendering device, then decoded and rendered.
[0193] Some of the benefits brought by the principles concern the efficiency of the encoding of the volumetric video content and allow to reduce the amount of data to be stored or transmitted. They can be applied in particular to a tiled atlas layout comprising a main large central view (with respect to the 3DoF+ content), embedding the part of the scene visible from the main viewpoint.
[0194] The implementations described herein can be implemented in, for example, a method or a process, an apparatus, a computer program product, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method or as an apparatus), the implementation of the discussed features is applicable in other forms (for example, as a program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The method can be implemented in, for example, an apparatus (for example, a general purpose computer that is programmed to perform the method) that is generally programmed to perform the method. The apparatus includes, for example, a processor (for example, a general purpose computer, a microprocessor, an integrated circuit, or a programmable logic device) programmed to perform the method. The processor also includes, for example, a communication device (for example, a smartphone, a tablet, a computer, a mobile phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate communication of information between end-users).
[0195] Implementations of the various processes and features described herein can be embodied in a variety of different equipment or applications, particularly equipment or applications associated with data encoding, data decoding, view generation, texture processing, and other processing of images and related texture information and / or depth information. Examples of such equipment include an encoder, a decoder, a post-processor processing output from a decoder, a pre-processor providing input to an encoder, a video encoder, a video decoder, a video codec, a web server, a set-top box, a laptop, a personal computer, a cell phone, a PDA, and other communication devices. As should be apparent, the equipment can be mobile, even if implemented in a stationary location.
[0196] Additionally, the methods can be implemented by instructions executed by a processor, and such instructions (and / or data values produced by implementations) can be stored on a processor-readable medium, such as, for example, an integrated circuit, a software carrier, or other storage device, such as, for example, a hard disk, a compact diskette ("CD"), an optical disk, such as, for example, a DVD, generally referred to as a digital versatile disk or a digital video disk, a random access memory ("RAM"), or a read-only memory ("ROM"). The instructions can form an application program embodied on a processor-readable medium. The instructions can be, for example, in hardware, firmware, software, or a combination. The instructions can be found in, for example, an operating system, a separate application, or a combination of the two. The processor can therefore be characterized as, for example, a device configured to perform the process in response to instructions defining the process and including receiving a medium having the instructions on it. The processor can also be characterized as a device configured to perform the process based on stored values, including receiving a medium having stored values which configure the processor to perform the process. Further, the processor can be characterized as a device configured to perform the process based on a combination of
[0197] As will be apparent, the implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method or data produced by one of the described implementations. For example, a signal can be formatted to carry as data the rules for writing or reading the syntax of a described embodiment, or as data the actual syntax-values written by a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted by multiple different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
[0198] A number of implementations have been described. Of course, it is to be understood that many modifications can be made. For example, elements of different implementations can be combined, supplemented, modified, or removed to produce other implementations. Additionally, one of ordinary skill will understand that other structures and processes can be substituted for those disclosed and the resulting implementations will perform at least substantially the same function(s) in at least substantially the same way(s) to achieve at least substantially the same result(s). As a result, these and other implementations are contemplated by this application.
Claims
1. A method for encoding a 3D scene in a data stream, the method comprising: - Obtain the reference viewing frame and the intermediate viewing frame defined within the 3D scene; - The reference center view captured from the viewpoint at the center of the reference viewing frame and the reference perimeter blocks that encode images captured from different viewpoints in the reference viewing frame are encoded in the data stream; - At least one intermediate center block is encoded in the data stream, the at least one intermediate center block encoding the difference between the view captured from the center of the intermediate viewing frame and the reference center view; as well as - The metadata describing the reference viewing frame, the intermediate viewing frame, and the different viewpoints is encoded in the data stream.
2. The method according to claim 1, wherein the reference viewing frame is the reference viewing frame closest to the intermediate viewing frame among a set of reference viewing frames defined within the 3D scene.
3. The method according to claim 1 or 2, wherein the reference peripheral blocks encode the differences between the peripheral image and the reference center view.
4. The method according to claim 1 or 2, wherein the intermediate viewing frame overlaps with the reference viewing frame.
5. The method according to claim 1 or 2, further comprising transmitting the data stream encoding the 3D scene.
6. A method for retrieving a 3D scene from a data stream, the method comprising: - Decode the following items from the data stream: ●Metadata, which describes the reference viewing frame and intermediate viewing frame in the 3D scene; ● Reference center view, which is captured from a viewpoint at the center of the reference viewing frame; ● At least one intermediate center block, which encodes the difference between the view captured from the center of the intermediate viewing frame and the reference center view; - The 3D scene is retrieved by deprojecting the pixels of the reference center view and the pixels of the at least one intermediate center block.
7. The method of claim 6, comprising: - Decode reference perimeter blocks from the data stream, the reference perimeter blocks encoding images captured from different viewpoints within the reference viewing frame; - The 3D scene is retrieved by deprojecting the pixels of the subgroups of the reference perimeter block.
8. The method of claim 7, wherein the subgroup of the reference peripheral block is selected based on a viewpoint located in the central viewing frame.
9. The method according to claim 6 or 7, wherein the reference viewing frame is the reference viewing frame closest to the intermediate viewing frame among a set of reference viewing frames defined within the 3D scene.
10. The method of claim 6 or 7, further comprising rendering a viewport image for a viewpoint located in the intermediate viewing frame.
11. A non-transitory processor-readable medium having instructions stored therein for causing at least one processor to perform at least the steps of the method according to any one of claims 1 to 5 or any one of claims 6 to 10.
12. An apparatus for encoding a 3D scene in a data stream, the apparatus comprising a memory and a processor configured to: - Obtain the reference viewing frame and the intermediate viewing frame defined within the 3D scene; - The reference center view captured from the viewpoint at the center of the reference viewing frame and the reference perimeter blocks that encode images captured from different viewpoints in the reference viewing frame are encoded in the data stream; - At least one intermediate center block is encoded in the data stream, the at least one intermediate center block encoding the difference between the view captured from the center of the intermediate viewing frame and the reference center view; as well as - The metadata describing the reference viewing frame, the intermediate viewing frame, and the different viewpoints is encoded in the data stream.
13. The device of claim 12, wherein the reference viewing frame is the reference viewing frame closest to the intermediate viewing frame among a set of reference viewing frames defined within the 3D scene.
14. The device of claim 12 or 13, wherein the reference peripheral blocks encode the differences between the peripheral image and the reference central view.
15. The device of claim 12 or 13, wherein the intermediate viewing frame overlaps with the reference viewing frame.
16. The device of claim 12 or 13, wherein the processor is further configured to transmit the data stream encoding the 3D scene.
17. An apparatus for retrieving a 3D scene from a data stream, the apparatus comprising a memory and a processor, the processor being configured to: - Decode the following from the data stream: ●Metadata, which describes the reference viewing frame and intermediate viewing frame in the 3D scene; ● Reference center view, which is captured from a viewpoint at the center of the reference viewing frame; ● At least one intermediate center block, which encodes the difference between the view captured from the center of the intermediate viewing frame and the reference center view; - The 3D scene is retrieved by deprojecting the pixels of the reference center view subgroup and the pixels of the at least one intermediate center block.
18. The device of claim 17, wherein the subgroup of the reference peripheral blocks is selected based on a viewpoint located in the central viewing frame.
19. The device of claim 17 or 18, wherein the reference viewing frame is the reference viewing frame closest to the intermediate viewing frame among a set of reference viewing frames defined within the 3D scene.
20. The device of claim 17 or 18, wherein the processor is further configured to render a viewport image for a viewpoint located in the intermediate viewing frame.
Citation Information
Patent Citations
Video coding techniques for multi-view video
CN110313181A
Method, apparatus and stream for volumetric video format
WO2019209838A1