Volumetric Video with Auxiliary Patches
The method for encoding and decoding 3D scene data using MVD content with metadata-managed patches addresses the limitations of 3DoF by enabling efficient 6DoF experiences with consistent visual feedback and additional processing capabilities.
Patent Information
- Application Number
- JP2022536635
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-19
- Filing Date
- 2020-12-17
- Publication Date
- 2025-12-15
- Estimated Expiration
- 2040-12-17
AI Technical Summary
Existing immersive video technologies, such as 3DoF, fail to provide a seamless experience for users who expect more degrees of freedom, leading to dizziness and limited navigation due to the lack of head translation and body movement capabilities, while 6DoF video requires efficient encoding and decoding methods to manage additional geometric and color information beyond the visible bounding box.
A method for encoding and decoding 3D scene data using multi-view plus depth (MVD) content, where patches from different regions of the scene are generated, packed into an atlas with metadata indicating their use for rendering or pre/post-processing, allowing efficient management of visible and auxiliary information.
Enables seamless 6DoF experiences by efficiently encoding and decoding volumetric video, ensuring consistent visual feedback during head translation and providing additional information for relighting, collision detection, and haptic interactions without increasing data complexity.
Smart Images

Figure 0007785676000005 
Figure 0007785676000006 
Figure 0007785676000007
Abstract
Description
[Technical Field]
[0001] The present principles generally relate to the domain of three-dimensional (3D) scenes and volumetric video content. This document is also understood in the context of encoding, formatting, and decoding data representing textures and 3D scene geometry for rendering of volumetric content on end-user devices such as mobile devices or head-mounted displays (HMDs). [Background technology]
[0002] This section is intended to introduce the reader to various aspects of art that may be related to various aspects of the present principles, which are described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present principles. Accordingly, it should be understood that these statements are to be read in this light, and not as admissions of prior art.
[0003] In recent years, there has been a growth in available large field-of-view content (up to 360°). Such content may not be fully visible to users viewing the content on immersive display devices such as head-mounted displays, smart glasses, PC screens, tablets, or smartphones. This means that at any given moment, only a portion of the content is visible to the user. However, users can typically navigate within the content by various means, such as head movements, mouse movements, touchscreens, and voice. It is typically desirable to encode and decode this content.
[0004] Immersive video, also known as 360° flat video, allows users to view everything around them through head rotation around a stationary point. The rotation only allows for a three-degrees-of-freedom (3DoF) experience. For example, even if 3DoF video is sufficient for a first-order omnidirectional video experience using a head-mounted display device (HMD), it can quickly become frustrating for viewers who expect more degrees of freedom, such as by experiencing parallax. Furthermore, 3DoF can also induce dizziness because users not only rotate their head but also translate their head in three directions, a translation that is not reproduced in a 3DoF video experience.
[0005] The large field-of-view content can be, among others, a three-dimensional computer graphic imagery scene (3D CGI scene), a point cloud, or an immersive video. Many terms can be used to design such immersive video, such as Virtual Reality (VR), 360, panoramic, 4π steradian, immersive, omnidirectional, or large field-of-view.
[0006] Volumetric video (also known as 6 Degrees of Freedom (6DoF) video) is an alternative to 3DoF video. When watching 6DoF video, in addition to rotation, users can also translate their head and even their body within the viewed content, experiencing parallax and even volume. Such video significantly increases the sense of immersion and the perception of scene depth, and prevents dizziness by providing consistent visual feedback during head translation. Content is created by means of dedicated sensors that allow simultaneous recording of the color and depth of the desired scene. The use of color camera rigs combined with photogrammetry techniques is a method for performing such recording, even though technical difficulties remain.
[0007] While 3DoF video involves a series of images resulting from the unmapping of texture images (e.g., spherical images encoded according to latitude / longitude projection mapping or equirectangular mapping), 6DoF video frames embed information from several viewpoints. They can be viewed as a temporal series of point clouds resulting from three-dimensional capture. Depending on the viewing conditions, two types of volumetric video can be considered. The first (i.e., full 6DoF) allows complete free navigation within the video content, while the second (also known as 3DoF+) restricts the user's visual space to a limited volume called the visual bounding box, allowing for a limited volume of head and parallax experiences. This second context represents a valuable trade-off between free navigation for seated audience members and passive viewing conditions.
[0008] In a 3DoF+ scenario, an approach consists of transmitting only the information necessary to view the 3D scene from any point in the visible bounding box. Another approach considers transmitting additional geometric and / or color information that is not visible from the visible bounding box but is useful for performing other processes on the decoder side, such as relighting, collision detection, or haptic interaction. This additional information may be conveyed in the same format as the visible points. However, there is a need for a format and method to indicate to the decoder that some of the information is used for rendering and other parts of the information are used for other processing. Summary of the Invention
[0009] The following presents a simplified summary of the present principles to provide a basic understanding of some aspects of the present principles. This summary is not an extensive overview of the present principles. It is not intended to identify key or critical elements of the present principles. The following summary merely presents some aspects of the present principles in a simplified form as a prelude to the more detailed description provided below.
[0010] The present principles relate to a method for encoding data representing a 3D scene in a data stream, the method comprising: Generating a first set of patches from first multi-view plus depth (MVD) content acquired for rendering a 3D scene, the first MVD acquired from a first region of the 3D scene, the patches being part of one of the views of the MVD content. - Generating a second set of patches from second MVD content acquired for pre-processing or post-processing use, the second MVD acquired from a second region of the 3D scene, which may overlap or be separate from the first region. - generating an atlas with first and second patches, where the atlas is an image packing the patches according to the atlas layout, and for the patches of the atlas, associated metadata indicating whether the patch is a first patch or a second patch. encoding said atlas within said data stream; Includes:
[0011] The present principles also relate to a method for decoding data representing a 3D scene from a data stream, the method comprising: - Decoding the data stream to obtain the atlas and associated metadata. The atlas is an image packing patches according to an atlas layout. A patch is part of one view of MVD content obtained from a region of the 3D scene. The metadata includes data indicating, for a patch of the atlas, whether the patch is a first patch or a second patch, where the first patch is part of MVD content obtained from the first region of the 3D scene and the second patch is part of MVD content obtained from the second region of the 3D scene. The first and second regions may overlap or be separate. - rendering a viewport image from a viewpoint within the 3D scene by using a patch indicated as a first patch in the metadata; using the patch indicated as the second patch in the metadata to pre-process and / or post-process the viewport image; Includes:
[0012] The present principles also relate to a device including a processor configured to implement the above encoding method, as well as a device including a processor configured to implement the above decoding method.
[0013] The present principles also relate to a data stream and / or non-transitory medium carrying data representing a 3D scene. The data stream or non-transitory medium may include: an atlas image packing first and second patches according to an atlas layout, where the first patch is part of one view of MVD content acquired for rendering of a 3D scene and the second patch is part of one view of MVD content acquired for pre-processing or post-processing use; - metadata associated with the atlas, the metadata including, for a patch of the atlas, data indicating whether the patch is a first patch or a second patch; Includes: [Brief explanation of the drawings]
[0014] The present disclosure will be better understood, and other particular features and advantages will become apparent, on reading the following description, which makes reference to the accompanying drawings, in which: [Figure 1] 1 illustrates a three-dimensional (3D) model of an object and points of a point cloud corresponding to the 3D model, in accordance with a non-limiting embodiment of the present principles. [Figure 2] 1 shows a non-limiting example of encoding, transmission and decoding of data representing a sequence of 3D scenes, in accordance with a non-limiting embodiment of the present principles; [Figure 3] 10 shows an exemplary architecture of a device that may be configured to implement the method described in connection with FIGS. 8 and 9, in accordance with a non-limiting embodiment of the present principles. [Figure 4] 1 illustrates an example of one embodiment of the syntax of a stream when data is transmitted via a packet-based transmission protocol, in accordance with a non-limiting embodiment of the present principles. [Figure 5]1 illustrates a spherical projection from a central viewpoint, in accordance with a non-limiting embodiment of the present principles; [Figure 6] 1 shows an example of the generation of atlases 60 and 61 by an encoder, in accordance with a non-limiting embodiment of the present principles. [Figure 7] 10 illustrates obtaining a view of a 3DoF+ rendering and an additional view of an auxiliary patch, in accordance with a non-limiting embodiment of the present principles. [Figure 8] 8 illustrates a method 80 for encoding volumetric video content with auxiliary information, in accordance with a non-limiting embodiment of the present principles. [Figure 9] 9 illustrates a method 90 for decoding volumetric video content including auxiliary information, in accordance with a non-limiting embodiment of the present principles. DETAILED DESCRIPTION OF THE INVENTION
[0015] The present principles are more fully described below with reference to the accompanying drawings, in which examples of the present principles are shown. However, the present principles may be embodied in many alternative forms and should not be construed as limited to the embodiments set forth herein. Accordingly, while the present principles are susceptible to various modifications and alternative forms, specific examples thereof are shown by way of example in the drawings and are described in detail herein. It is to be understood, however, that there is no intention to limit the present principles to the particular forms disclosed, but on the contrary, the present disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present principles as defined by the appended claims.
[0016] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the present principles. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It will be further understood that as used herein, the terms "comprises," "comprising," "includes," and / or "including" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, when an element is referred to as "responsive to" or "connected to" another element, it may be directly responsive to or connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as "directly responsive to" or "directly connected to" another element, there are no intervening elements present. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated as " / ".
[0017] In this specification, terms such as "first," "second," etc. may be used to describe various elements, but it should be understood that these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first element can be referred to as a second element, and similarly, a second element can be referred to as a first element without departing from the teachings of the present principles.
[0018] Some of the figures include arrows on communication paths to indicate the primary direction of communication, however, it should be understood that communication may occur in the opposite direction to the depicted arrow.
[0019] Some examples are described with reference to block diagrams and operational flowcharts, in which each block represents circuit elements, modules, or portions of code, with each block including one or more executable instructions for implementing a specified logical function. It should also be noted that in other implementations, the functions noted in the blocks may occur out of the order noted. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may be executed in the reverse order, depending on the functionality involved.
[0020] As used herein, "by one example" or "in one example" means that a particular feature, structure, or characteristic described in connection with this embodiment may be included in at least one implementation of the present principles. The appearances of the phrase "by one example" or "in one example" in various places in this specification do not necessarily all refer to the same embodiment, and in separate or alternative embodiments are not necessarily mutually exclusive of other embodiments.
[0021] Reference numerals appearing in the claims are by way of example only and shall have no limiting effect on the scope of the claims. Although not expressly stated, the present embodiments and variations may be used in any combination or subcombination.
[0022] FIG. 1 illustrates a three-dimensional (3D) model 10 of points of an object and a point cloud 11 corresponding to the 3D model 10. The 3D model 10 and point cloud 11 may correspond to, for example, a potential 3D representation of an object in a 3D scene containing other objects. The model 10 may be a 3D mesh representation, and the points of the point cloud 11 may be vertices of the mesh. The points of the point cloud 11 may also be points spread on the surface of a face of the mesh. The model 10 may also be represented as a splatted version of the point cloud 11, where the surface of the model 10 is created by splatting the points of the point cloud 11. The model 10 may be represented by many different representations, such as voxels or splines. FIG. 1 illustrates the fact that a point cloud may be defined as a surface representation of a 3D object, and that a surface representation of a 3D object may be generated from a cloud of points. As used herein, projecting a point of a 3D object (by an elongation point of the 3D scene) onto an image is equivalent to projecting any representation of this 3D object, for example a point cloud, a mesh, a spline model or a voxel model.
[0023] A point cloud can be represented in memory, for example, as a vector-based structure, with each point having its own coordinates in the reference frame of the viewpoint (e.g., three-dimensional coordinates XYZ, or solid angle and distance (also called depth) from / to the viewpoint) and one or more attributes, also called components. Examples of components are color components that can be expressed in various color spaces, for example, RGB (red, green, and blue) or YUV (Y is the luminance component and UV are the two color difference components). A point cloud is a representation of a 3D scene containing objects. A 3D scene can be viewed from a given viewpoint or range of viewpoints. Point clouds can be represented in many ways, for example, from capturing real objects photographed by a camera rig, optionally complemented by a depth active sensing device; From capturing virtual / synthetic objects photographed by virtual camera rigs in modeling tools, · It can be obtained from a mixture of both real and virtual objects.
[0024] 2 shows a non-limiting example of encoding, transmission and decoding of data representing a sequence of 3D scenes, e.g., a coding format that can accommodate 3DoF, 3DoF+ and 6DoF decoding simultaneously.
[0025] A sequence of 3D scenes 20 is captured. Whereas the sequence of photographs is a 2D video, the sequence of 3D scenes is a 3D (also called volumetric) video. The sequence of 3D scenes can be provided to a volumetric video rendering device for 3DoF, 3Dof+ or 6DoF rendering and display.
[0026] A sequence of 3D scenes 20 is provided to an encoder 21. The encoder 21 takes as input a 3D scene or a sequence of 3D scenes and provides a bitstream representing the input. The bitstream may be stored in a memory 22 and / or on an electronic data medium and may be transmitted over a network 22. The bitstream representing the sequence of 3D scenes may be read from the memory 22 and / or received from the network 22 by a decoder 23. The decoder 23 is input with the bitstream and provides the sequence of 3D scenes, for example in point cloud format.
[0027] The encoder 21 may include several circuits that implement several steps. In a first step, the encoder 21 projects each 3D scene onto at least one 2D photograph. 3D projection is any method of mapping three-dimensional points onto a two-dimensional plane. Because modern methods for displaying graphic data are based on planar (pixel information from several bit planes) two-dimensional media, the use of this type of projection is widespread, especially in computer graphics, manipulation, and drafting. The projection circuit 211 provides at least one two-dimensional frame 2111 for each 3D scene in the sequence 20. The frame 2111 includes color and depth information representing the 3D scene projected onto the frame 2111. In a variant, the color and depth information are encoded in two separate frames 2111 and 2112.
[0028] The metadata 212 is used and updated by the projection circuitry 211. The metadata 212 includes information about the projection operation (e.g., projection parameters) and how color and depth information is organized within frames 2111 and 2112, as described in connection with Figures 5-7.
[0029] A video encoding circuit 213 encodes the sequence of frames 2111 and 2112 as a video. The pictures of the 3D scene 2111 and 2112 (or a sequence of pictures of the 3D scene) are encoded in a stream by the video encoder 213. The video data and metadata 212 are then encapsulated in a data stream by a data encapsulation circuit 214.
[0030] The encoder 213 may, for example, -JPEG, specification ISO / CEI10918-1UIT-T Recommendation T.81, https: / / www.itu.int / rec / T-REC-T.81 / en; -Compliant with encoders such as AVC, also known as MPEG-4 AVC or h264, ITU-TH.264 and ISO / CEI MPEG-4-Part 10 (ISO / CEI14496-10), http: / / www.itu.int / rec / T-REC-H.264 / en, HEVC (whose specifications can be found on the ITU website, T Recommendation, H Series, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en), -3D-HEVC (an extension of HEVC whose specification can be found on the ITU website, T Recommendation, H Series, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en annex G and I), -VP9, developed by Google, or -AV1 (AO Media Video 1) developed by the Alliance for Open Media.
[0031] The data stream is stored by a decoder 23 in a memory accessible, for example, via a network 22. The decoder 23 comprises different circuits that implement different steps of the decoding. The decoder 23 takes as input the data stream generated by the encoder 21 and provides a sequence of 3D scenes 24 to be rendered and displayed by a volumetric video display device, such as a head-mounted device (HMD). The decoder 23 obtains the stream from a source 22. For example, the source 22 may be - local memory, such as video memory or RAM (or Random Access Memory), flash memory, ROM (or Read Only Memory), hard disk, etc. a storage interface, such as an interface to a mass storage, RAM, flash memory, ROM, optical disk or magnetic support; a communication interface, such as a wired interface (e.g., a bus interface, a wide area network interface, a local area network interface) or a wireless interface (e.g., an IEEE 802.11 interface or a Bluetooth interface), - a user interface, such as a graphical user interface, that allows a user to input data.
[0032] The decoder 23 includes a circuit 234 for extracting data encoded within the data stream. The circuit 234 takes the data stream as input and provides metadata 232 corresponding to the metadata 212 encoded in the stream and the two-dimensional video. The video is decoded by a video decoder 233, which provides a sequence of frames. The decoded frames include color and depth information. In a variant, the video decoder 233 provides two sequences of frames, one containing color information and the other containing depth information. The circuit 231 does not use the metadata 232 to project the color and depth information from the decoded frames, but instead provides a sequence of 3D scene 24. The sequence of 3D scene 24 corresponds to the potentially reduced accuracy associated with encoding the sequence of 3D scene 20 as 2D video and video compression.
[0033] Other circuits and functionality may be added, for example, before the backprojection step by circuit 231 or in post-processing steps after backprojection. For example, circuitry may be added for relighting the scene from another light located somewhere in the scene. Collision detection may be performed for depth compositing, such as adding new objects to the 3DoF+ scene in a consistent and realistic manner, or for path planning. Such circuitry may require geometric and / or color information about the 3D scene that is not used for the 3DoF+ rendering itself. The meaning of the different types of information must be indicated by the bitstream representing the 3DoF+ scene.
[0034] Figure 3 shows an example of the architecture of a device 30 that may be configured to perform the methods described in relation to Figures 8 and 9. The encoder 21 and / or decoder 23 of Figure 2 may implement this architecture. Alternatively, each circuit of the encoder 21 and / or decoder 23 may be a device according to the architecture of Figure 3, coupled together, for example, via their bus 31 and / or via an I / O interface 36.
[0035] The device 30 includes the following elements coupled together by a data and address bus 31: a microprocessor 32 (or CPU), for example a DSP (or Digital Signal Processor), -ROM (or read-only memory) 33; a RAM (or Random Access Memory) 34; a storage interface 35; an I / O interface 36 for receiving data to be transmitted from an application; - a power source, for example a battery;
[0036] According to one example, the power source is external to the device. In each of the mentioned memories, the word "register" as used herein may correspond to a small area (a few bits) or a very large area (e.g., an entire program or a large amount of received or decoded data). The ROM 33 contains at least programs and parameters. The ROM 33 can store algorithms and instructions for performing techniques according to the present principles. When switched on, the CPU 32 uploads the program in the RAM and executes the corresponding instructions.
[0037] The RAM 34 contains in registers the programs executed by the CPU 32 and uploaded after switching on the device 30, input data in registers, intermediate data for different states of the methods in registers, and other variables used for the execution of the methods in registers.
[0038] The implementations described herein may be implemented in, for example, a method or process, an apparatus, a computer program product, a data stream, or a signal. Even when discussed only in the context of a single form of implementation (e.g., discussed only as a method or device), the implementation of the discussed features may also be implemented in other forms (e.g., a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The method may be implemented in an apparatus such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include, for example, communication devices such as computers, mobile phones, handheld / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.
[0039] According to an embodiment, the device 30 is configured to implement the method described in relation to Figures 8 and 9, -Mobile devices and a communication device; -A gaming device, -a tablet (or tablet computer); -Laptop and -A still camera, -Video camera and - an encoding chip; - a server (for example a broadcast server, a video-on-demand server or a web server), and belongs to the set including:
[0040] FIG. 4 illustrates an example of a syntax embodiment of a stream when data is transmitted via a packet-based transmission protocol. FIG. 4 illustrates an exemplary structure 4 of a volumetric video stream. The structure consists of containers that organize the stream into independent elements of the syntax. The structure may include a header section 41, which is a set of data common to all syntax elements of the stream. For example, the header section includes some of the metadata about the syntax elements, describing their respective nature and role. The header section may also include some of the metadata 212 of FIG. 2, such as the coordinates of a central viewpoint used to project points of the 3D scene onto frames 2111 and 2112. The structure includes a payload that includes an element of syntax 42 and at least one element of syntax 43. Syntax element 42 includes data representing color and depth frames. The images may be compressed according to a video compression method.
[0041] Elements of syntax 43 are part of the payload of the data stream and may contain metadata about how frames of elements of syntax 42 are encoded, for example, parameters used to project or pack points of a 3D scene onto the frame. Such metadata may be associated with each frame of video or a group of frames (also known as a Group of Pictures (GoP) in video compression standards).
[0042] FIG. 5 illustrates a patch atlas approach with four example centers of projection. 3D scene 50 includes features. For example, center of projection 51 is a perspective projection camera, and camera 53 is an orthographic projection camera. The cameras may also be omnidirectional cameras with spherical mapping (e.g., equirectangular mapping) or cubic mapping. 3D points of the 3D scene are projected onto a 2D plane associated with a virtual camera located at the center of projection according to a projection operation described in the projection data of the metadata. In the example of FIG. 5, the projection of points captured by camera 51 is mapped onto patch 52 according to a perspective mapping, and the projection of points captured by camera 53 is mapped onto patch 54 according to an orthogonal mapping.
[0043] Clustering the projected pixels results in a number of 2D patches, which are packed into a rectangular atlas 55. The organization of the patches within the atlas defines the atlas layout. In one embodiment, two atlases have the same layout: one for texture (i.e., color) information and one for depth information. Two patches captured by the same camera or two separate cameras may contain information representing the same portion of a 3D scene, such as patches 54 and 56.
[0044] The packing operation generates patch data for each generated patch. The patch data includes a reference to the projection data (e.g., an index into a table of projection data or a pointer to the projection data (an address in memory or the data stream)) and information describing the location and size of the patch within the atlas (e.g., a pixel's top-left corner coordinate, size, and width). The patch data items are associated with the compressed data of one or two atlases and are added to the metadata encapsulated in the data stream.
[0045] Figure 6 shows an example of the generation of atlases 60 and 61 by an encoder. Atlases 60 and 61 contain texture information (e.g., RGB or YUV data) of points in a 3D scene, according to a non-limiting embodiment of the present principles. As explained in relation to Figure 5, an atlas is an image packing of patches. For example, the encoder takes as input a multiview+depth video containing three views 62, 63, and 64 in the example of Figure 6. The encoder removes inter-view redundancy (pruning step) and packs selected patches of texture and depth into one or more atlases. Thus, the bitstream consists of multiple video streams (e.g., HEVC video streams) carrying atlases of texture and depth patches, along with metadata describing the camera parameters of the input views and the atlas layout.
[0046] A patch atlas consists of a pair of texture and depth atlas components with the same photo size and the same layout (same packing) for texture and depth. In one approach, the atlas carries only the information necessary for 3DoF+ rendering of the scene from any point within the visible bounding box. In another approach, the atlas may carry additional geometry and / or color information useful for other processing, such as scene relighting or collision detection. For example, this additional information may be the geometry of the backsides of objects in the 3D scene. Such patches are called auxiliary patches. They are not rendered by the decoder, but are instead intended to be used by pre- or post-processing circuitry in the decoder.
[0047] FIG. 7 illustrates the acquisition of views for 3DoF+ rendering and additional views of auxiliary patches. On the encoder side, the generation of auxiliary patches may be performed by different means. For example, in acquiring scene 70, a first group of real or virtual cameras 71 may be positioned pointing at the front of scene 70. A second group of real or virtual cameras 72 may be positioned to view the back and sides of the volumetric scene. In one embodiment, camera 72 captures a view at a lower resolution than camera 71. Camera 72 obtains the geometry and / or color of hidden portions of objects. Patches acquired from views captured by camera 71 are patches for 3DoF+ rendering, while patches acquired from views captured by camera 72 are auxiliary patches for completing the description of geometry and / or color information for pre- or post-processing use. Metadata associated with the atlas patches may be formatted to signal the meaning of each patch. On the decoder side, the viewport renderer must skip patches that are invalid for rendering. The metadata also indicates which modules in the decoder may use these invalid rendering patches. For example, the relighting circuitry uses this auxiliary information to update its geometry map from a lighting perspective and to vary the lighting texture of the entire scene accordingly to produce appropriate shadows.
[0048] A means is added to the encoder to generate auxiliary patches associated with the camera 72 capturing images from the rear and side, which describe the geometry of the rear portion of the object, typically at a lower resolution. Additional depth views from the rear and side must first be acquired, which can be done in a variety of ways. For synthetically generated objects, depth images associated with a virtual camera placed at an arbitrary location are acquired directly from the 3D model. In natural 3D capture, additional color and / or active depth cameras can be added to the capture stage, with the depth camera providing the depth view directly, while a photogrammetry algorithm estimates depth from the color view. When neither a 3D model nor additional capture is available, a convex shape completion algorithm can be used to generate a plausible closed shape from the open-form geometry received from the front camera. Inter-view redundancy is then removed by pruning in a similar manner to that performed for the views from camera 71. In one embodiment, pruning is performed independently for the two groups of views. Therefore, potential redundancies between the regular patches and the extra patches are not removed. The resulting auxiliary depth patches are packed together with the regular patches in a depth patch atlas.
[0049] In another embodiment, if auxiliary patches are defined for depth only, and the same layout is used for the texture and depth atlas, the texture portion on the atlas will be left empty. This results in a loss of headroom in the texture atlas, even though these auxiliary patches are likely to be defined at a lower resolution. In such an embodiment, different layouts may be used for the depth and texture atlases, and this difference will be indicated in the metadata associated with the atlas.
[0050] A possible syntax for the metadata describing the atlas may include a high-level concept called "entity_id", which allows attaching a group of patches to high-level semantic processes such as object filtering or compositing. A possible syntax for the atlas parameters metadata is shown in the table below:
[0051] [Table 1]
[0052] According to one embodiment of the present principles, auxiliary patches are identified as specific entities called auxiliary entities. Then, some entities and their functions (i.e., whether they are auxiliary entities or not) are described in metadata as shown in the table below.
[0053] [Table 2]
[0054] auxiliary_flag equal to 1 indicates that an auxiliary description is present for each entity structure.
[0055] auxiliary_entity_flag[e] equal to 1 indicates that the patch associated with entity e is not for viewport rendering.
[0056] According to another embodiment of the present principles, auxiliary patches are signaled at the patch level by modifying the syntax of the atlas parameters as shown in the table below.
[0057] [Table 3]
[0058] An auxiliary_flag equal to 1 indicates that an auxiliary description is present for each patch structure.
[0059] auxiliary_patch_flag[a][p] equal to 1 indicates that patch p of atlas a is not for viewport rendering.
[0060] In another embodiment, the patch information data syntax defines auxiliary patch flags as shown in the table below.
[0061] [Table 4]
[0062] On the decoding side, auxiliary_patch_flag is used to determine if the patch contains information for rendering and / or for another module.
[0063] FIG. 8 illustrates a method 80 for encoding volumetric video content with auxiliary information, in accordance with a non-limiting embodiment of the present principles. In step 81, patches used for 3DoF+ rendering are generated, e.g., by pruning redundant information from multi-view plus depth content acquired by a first group of cameras. In step 82, auxiliary patches are generated from views captured by cameras capturing parts of the scene not intended to be rendered. Steps 81 and 82 may be performed in parallel or sequentially. The views used to generate the auxiliary patches are captured by a second group of cameras, e.g., located at the back and sides of the 3D scene. The auxiliary patches are generated, e.g., by pruning redundant information contained in views captured by the first and second groups of cameras. In another embodiment, the auxiliary patches are generated, e.g., by pruning redundant information contained in views captured only by the second group of cameras. In this embodiment, redundancy may exist between the 3DoF+ patches and the auxiliary patches. In step 83, an atlas is generated by packing the 3DoF+ and auxiliary patches into the same image. In one embodiment, the packing layout differs for depth and color components of the atlas. Metadata describing the atlas parameters and patch parameters is generated according to the syntax as set forth in the table above. For each patch, the metadata includes information indicating whether the patch is a 3DoF+ patch to be rendered or an auxiliary patch to be used for pre-processing and / or post-processing. In step 84, the generated atlas and associated metadata are encoded in a data stream.
[0064] FIG. 9 shows a method 90 for decoding volumetric video content including auxiliary information, according to a non-limiting embodiment of the present principles. In step 91, a data stream representing the volumetric content is obtained from the stream. The data stream is decoded to obtain an atlas and associated metadata. The atlas is an image packing at least one patch according to a packing layout. The patch is a photograph containing depth and / or color information representing a portion of a 3D scene. The metadata contains information for backprojecting the patch and obtaining the 3D scene. In step 92, the patches are unpacked from the atlas, and properties are attributed to each patch according to the information contained in the metadata. The patch may be a 3DoF+ patch to be used to render a viewport image in step 93, or an auxiliary patch to be used for pre-processing or post-processing operations in step 94. Steps 93 and 94 may be performed in parallel or sequentially.
[0065] The implementations described herein may be implemented in, for example, a method or process, an apparatus, a computer program product, a data stream, or a signal. Even when discussed in the context of only a single form of implementation (e.g., discussed only as a method or device), the implementation of the discussed features may also be implemented in other forms (e.g., a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The method may be implemented in an apparatus such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices such as, for example, smartphones, tablets, computers, mobile phones, portable / personal digital assistants ("personal digital assistants," or "PDAs"), and other devices that facilitate communication of information between end users.
[0066] Implementations of the various processes and features described herein may be embodied in a variety of different devices or applications, particularly, for example, devices or applications associated with data encoding, data decoding, view generation, texture processing, and other processing of images and associated texture and / or depth information. Examples of such devices include encoders, decoders, post-processors that process output from decoders, pre-processors that provide input to encoders, video coders, video decoders, video codecs, web servers, set-top boxes, laptops, personal computers, mobile phones, PDAs, and other communication devices. As should be clear, the devices may be mobile and installed in mobile vehicles.
[0067] Furthermore, a method may be implemented by instructions executed by a processor, and such instructions (and / or data values produced by an implementation) may be stored on a processor-readable medium, such as, for example, an integrated circuit, a software carrier, or other storage device, e.g., a hard disk, a compact diskette ("CD"), an optical disk (e.g., a DVD, often referred to as a digital versatile disk or digital video disk), a random access memory ("RAM"), or a read-only memory ("ROM"). The instructions may form an application program tangibly embodied on the processor-readable medium. The instructions may be, for example, hardware, firmware, software, or a combination. The instructions may be found, for example, in an operating system, a separate application, or a combination of the two. A processor may therefore be characterized, for example, as both a device configured to execute a process and a device that includes a processor-readable medium (e.g., a storage device) having instructions for executing a process. Furthermore, a processor-readable medium may store data values produced by an implementation in addition to, or in place of, instructions.
[0068] As will be apparent to those skilled in the art, implementations may generate various signals formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry, as data, rules for writing or reading syntax of a described embodiment, or to carry, as data, the actual syntax value written by a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The signal it carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.
[0069] Many implementations have been described. Nevertheless, it will be understood that various modifications may be made. For example, elements of different implementations may be combined, supplemented, modified, or deleted to produce other implementations. Moreover, those skilled in the art will understand that other structures and processes may be substituted for those disclosed, with the resulting implementation performing at least substantially the same function in at least substantially the same way to achieve at least substantially the same results as the disclosed implementations. Accordingly, these and other implementations are contemplated by this application.
Claims
1. 1. A method for encoding a 3D scene in a data stream, comprising: generating a first set of patches from first Multiview Plus Depth (MVD) content acquired for rendering the 3D scene, the first patches being part of one of the views of the first MVD content and being used to render a viewport image from a viewpoint within the 3D scene; generating a second set of patches from second MVD content obtained for use in pre-processing of rendering of the 3D scene and post-processing of rendering of the 3D scene, the second patches being part of one of the views of the second MVD content and being used to pre-process and post-process the viewport image; - generating a first atlas for packing the first patches and a second atlas for packing the second patches, where the atlases are images of packing patches according to an atlas layout; - encoding said first and second atlases in said data stream; A method comprising:
2. The method of claim 1 , wherein the second MVD is acquired at a resolution lower than a resolution of the first MVD.
3. The method of claim 1 or 2, wherein the patch is a portion of one view of the MVD obtained by removing information redundancy between views of the MVD.
4. 1. A method for decoding a 3D scene from a data stream, comprising: - decoding the data stream to obtain first and second atlases, where the atlas is an image packing first or second patches according to an atlas layout, the first patch being part of one view of a first MVD content obtained for rendering the 3D scene, and the second patch being part of one view of a second MVD content obtained for pre-processing or post-processing use; - using the second patch to preprocess the viewport image; - rendering the viewport image from a viewpoint within the 3D scene by using a first patch; - using a second patch to post-process the viewport image; A method comprising:
5. The method of claim 4 , wherein the second MVD has a resolution lower than a resolution of the first MVD.
6. 1. A device for encoding a 3D scene in a data stream, said device comprising: generating a first set of patches from first Multiview Plus Depth (MVD) content acquired for rendering the 3D scene, the first patches being part of one of the views of the first MVD content and being used to render a viewport image from a viewpoint within the 3D scene; generating a second set of patches from second MVD content obtained for use in pre-processing the rendering of the 3D scene and post-processing the rendering of the 3D scene, the second patches being part of one of the views of the second MVD content and being used to pre-process and post-process the viewport image; - generating a first atlas for packing the first patch and a second atlas for packing the second patch, where the atlases are images for packing the patches according to an atlas layout; - encoding said first and second atlases in said data stream, A device that includes a processor and associated memory configured to:
7. The device of claim 6 , wherein the second MVD is acquired at a resolution lower than that of the first MVD.
8. 8. The device of claim 6 or 7, wherein the patch is part of one view of the MVD obtained by removing information redundancy between views of the MVD.
9. 1. A device for decoding a 3D scene from a data stream, comprising: - decoding the data stream to obtain first and second atlases, where the atlas is an image packing first or second patches according to an atlas layout, the first patch being part of one view of a first MVD content obtained for rendering the 3D scene, and the second patch being part of one view of a second MVD content obtained for pre-processing or post-processing use; - using a second patch to preprocess the viewport image; - rendering the viewport image from a viewpoint within the 3D scene by using a first patch; - Use a second patch to post-process the viewport image 1. A device including a processor configured to:
10. The device of claim 9 , wherein the second MVD has a resolution lower than a resolution of the first MVD.
11. 2. The method of claim 1, wherein the first and second patches are packed into a common atlas, and the common atlas is associated with metadata for each patch indicating whether the patch is a first patch or a second patch.
12. 7. The device of claim 6, wherein the first and second patches are packed into a single atlas, and the single atlas is associated with metadata for each patch indicating whether the patch is a first patch or a second patch.
13. 5. The method of claim 4, wherein the first atlas and the second atlas are a single atlas, and the single atlas is associated with metadata for each patch indicating whether the patch is a first patch or a second patch.
14. 10. The device of claim 9, wherein the first atlas and the second atlas are a single atlas, and the single atlas is associated with metadata for each patch indicating whether the patch is a first patch or a second patch.
Citation Information
Patent Citations
Method, apparatus and stream for immersive video format
WO2018130491A1
A method and apparatus for encoding a point cloud representing three-dimensional objects
WO2019110405A1
Processing video patches for three-dimensional content
WO2019202207A1