Different atlas packing for volumetric video

By generating and encoding a combination of the first and second images, the problem of viewing large field-of-view content on immersive display devices is solved, reducing the consumption of encoding and transmission resources and improving the efficiency and quality of immersive displays.

CN115462088BActive Publication Date: 2025-12-23INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180030985.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-07
Filing Date
2021-04-01
Publication Date
2025-12-23
Estimated Expiration
2041-04-01

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively encode and decode large field-of-view content, resulting in users being unable to fully view the content on immersive display devices or experiencing dizziness. Furthermore, existing methods are inefficient in transmitting 3D scene point attributes, leading to a waste of high bitrate and pixel rate.

Method used

A method is adopted to generate a viewport image by generating a combination of a first image and a second image, respectively encoding different attributes of the 3D scene, and using metadata to indicate the location and existence of the blocks, thereby reducing the bit rate and pixel rate of encoding and transmission.

Benefits of technology

It improves encoding efficiency, reduces the consumption of transmission and storage resources, prevents dizziness and enhances immersion, and achieves efficient 3D scene rendering and display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115462088B_ABST
    Figure CN115462088B_ABST
Patent Text Reader

Abstract

Methods, devices and streams for encoding and decoding a scene, such as a point cloud, in the context of a tile-based transmission of volumetric video content are disclosed. Attributes of points of the scene are projected onto tiles. Each point has a geometry attribute. Some points can not have a value for other attributes, such as transparency for displacement attributes. According to the principles of the invention, each attribute is encoded in a different atlas with its own layout. This allows to save the pixel rate in the memory of the renderer.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present principles generally relate to the field of three-dimensional (3D) scene and volumetric video content. The present document is also understood in the context of encoding, formatting and decoding data representative of attributes of points of a 3D scene, to render volumetric content on an end-user device such as a mobile device or a head-mounted display (HMD). BACKGROUND

[0002] This section is intended to introduce the reader to various aspects of art that can be related to various aspects of the present principles that are described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present principles. Accordingly, it should be understood that these statements are to be read in this light, and not as admissions of prior art.

[0003] Recently, there has been a growth in available large field of view content (up to 360°). A user watching content on an immersive display device (such as a head-mounted display, smart glasses, PC screen, tablet, smartphone, etc.) can not be able to see the whole of such content. This means that at a given moment, the user can only watch a part of the content. However, the user can typically navigate within the content through various means such as head movement, mouse movement, touch screen, voice, and the like. It is generally desirable to encode and decode such content.

[0004] Immersive video (also called 360° planar video) allows a user to watch everything around him by rotating his head around a static viewpoint. The rotation only allows a 3 degrees of freedom (3DoF) experience. Even if 3DoF video is sufficient to meet the requirements of a first omnidirectional video experience (e.g. using a head-mounted display (HMD device)), 3DoF video can quickly become frustrating for a viewer who expects more freedom (e.g. by experiencing parallax). Moreover, 3DoF can also cause dizziness because a user never only rotates his head but also translates his head in three directions, which are not reproduced in a 3DoF video experience.

[0005] Amongst others, the large field of view content can be a three-dimensional computer graphics image scene (3D CGI scene), a point cloud or an immersive video. Many terms can be used to design such immersive video: for example, virtual reality (VR), 360, panoramic, 4π steradians, immersive, omnidirectional or large field of view.

[0006] Volume videos, also called 6 Degrees of Freedom (6DoF) videos, are an alternative to 3DoF videos. When watching a 6DoF video, in addition to rotations, the user can translate his head, and even his body, in the content of the watch, and experience parallax and even volume. This video significantly increases the immersion and the perception of the depth of the scene and prevents dizziness by providing a consistent visual feedback during head translation. The content is created with dedicated sensors that allow recording simultaneously the color and the depth of the scene of interest. Even if technical difficulties still exist, using a color camera equipment combined with photogrammetry techniques is one way to perform such recording.

[0007] While 3DoF videos consist in a sequence of images resulting from the de-mapping of texture images (e.g. spherical images encoded according to a latitude / longitude projection mapping or an equirectangular projection mapping), 6DoF video frames embed information from multiple viewpoints. They can be seen as a temporal sequence of point clouds resulting from a three-dimensional capture. Two kinds of volume videos can be considered depending on the viewing conditions. The first one, i.e. full 6DoF, allows a full freedom of navigation within the video content, while the second one, also called 3DoF+, restricts the user viewing space to a limited volume called viewing bounding box, allowing a limited head translation and parallax experience. This second case is a valuable compromise between free navigation and passive viewing conditions for seated viewers.

[0008] In a 3DoF+ scenario, one approach consists in sending only the geometry and color information needed to observe the 3D scene from any point of the viewing bounding box. Another approach considers sending additional information, i.e. other attributes of the points of the 3D scene than the color attribute, whether visible or not from the viewing bounding box, but usable to perform higher quality viewport rendering or other processes at the decoder side, like re-illumination, collision detection or haptic interaction. This additional information can be conveyed in the same format as the color attribute of the pixels of the image resulting from the projection of the points of the 3D scene. However, each point of the 3D scene does not share the same number of attributes. For example, a transparency attribute does not need to be transmitted for each point of the scene, as the vast majority of points have a default transparency value, i.e. an opaque value. Other attributes can be more diffuse and do not require the fine resolution of the projected map. Therefore, there is a need for a format and a method for carrying each attribute of the points of the 3D scene while limiting the bit rate and the pixel rate of the encoded bitstream, the transmitted bitstream and the decoded bitstream. SUMMARY

[0009] The following presents a simplified summary of the principles of the application in order to provide a basic understanding of some aspects of the application. This summary is not an extensive overview of the principles of the application. It is not intended to identify key or critical elements of the principles of the application. The following summary merely presents some aspects of the principles of the application in a simplified form as a prelude to the more detailed description provided below.

[0010] The principles of the present invention relate to a method comprising decoding from a data stream a first image, a second image and associated metadata. The metadata comprises a list of data items. A data item comprises:

[0011] - a position and a size of a region of the first image corresponding to the current tile;

[0012] - a flag indicating whether the current tile is present in the second image; and

[0013] - on condition that the flag indicates that the current tile is present in the second image, a position of a region of the second image corresponding to the current tile.

[0014] In one embodiment, the pixels of the first image encode a first attribute of a portion of points of a 3D scene and the pixels of the second image encode a second attribute of the same portion of the 3D scene. The first attribute is different from the second attribute.

[0015] In another embodiment, the pixels of the first image and of the second image are back-projected according to the decoded metadata to generate a 3D scene and to generate a viewport image to render volumetric content from a viewpoint within the 3D scene.

[0016] The principles of the present invention also relate to a device comprising a processor configured to implement the steps of the above method.

[0017] The principles of the present invention also relate to a data stream encoded according to the above method.

[0018] The principles of the present invention also relate to a method comprising:

[0019] - obtaining a set of first tiles and second tiles. The first tiles encode projections of a first attribute of a portion of points of a 3D scene; the second tiles encode projections of a second attribute and of the first attribute of the portion of points of the 3D scene;

[0020] - encoding the first attribute in a first image by packing the first tiles and the second tiles of the set into the first image, and encoding the second attribute in a second image by packing the second tiles of the set into the second image;

[0021] - generating a data stream comprising the first image, the second image and associated metadata. For a current tile of the set of tiles, the metadata comprises:

[0022] • a position and a size of a region of the first image corresponding to the current tile;

[0023] • a flag indicating whether the current tile is a second tile; and

[0024] • the position of the region of the second image corresponding to the current tile, on condition that the current tile is a second tile.

[0025] The principles of the application also relate to a device comprising a processor configured to implement the steps of the above-described method. BRIEF DESCRIPTION OF DRAWINGS

[0026] The present disclosure will be better understood and other specific features and advantages will emerge upon reading the following description, the description making reference to the drawings in which:

[0027] - Figure 1 A three-dimensional (3D) model of an object and points of a point cloud corresponding to this 3D model are shown, according to a non-limiting embodiment of the principles of the application;

[0028] - Figure 2 Non-limiting examples of encoding, transmitting and decoding data representative of a sequence of 3D scenes are shown, according to a non-limiting embodiment of the principles of the application;

[0029] - Figure 3 An exemplary architecture of a device that can be configured to implement the methods described in relation to Figure 8 and Figure 9 are shown, according to a non-limiting embodiment of the principles of the application;

[0030] - Figure 4 An example of an implementation of the syntax of a stream when transmitting data through a packet-based transmission protocol is shown, according to a non-limiting embodiment of the principles of the application;

[0031] - Figure 5 A tile atlas approach with 4 projection centers is shown, according to a non-limiting embodiment of the principles of the application;

[0032] - Figure 6 Generating an atlas for a set of tiles storing one or two attributes is shown, according to a non-limiting embodiment of the principles of the application;

[0033] - Figure 7 Determination of a displacement attribute is shown, according to a non-limiting embodiment of the principles of the application;

[0034] - Figure 8 A method 80 for decoding a data stream representative of volumetric content is shown, according to a non-limiting embodiment of the principles of the application;

[0035] - Figure 9 A method 90 for encoding volumetric content in a data stream is shown, according to a non-limiting embodiment of the principles of the application. DETAILED DESCRIPTION

[0036] The principles of the present application will be more fully understood in view of the detailed description and the drawings attached hereto. The principles of the present application can be embodied in many alternative forms and should not be construed as limited to the examples set forth herein. Accordingly, while the principles of the present application are susceptible to various modifications and alternative forms, specific examples thereof are shown by way of example in the drawings and will be described herein in detail. It should be understood, however, that there is no intent to limit the principles of the present application to the particular forms disclosed, but on the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the principles of the present application as defined by the claims.

[0037] The terminology used herein is for the purpose of describing particular examples only and is not intended to be limiting of the principles of the present application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Additionally, when an element is referred to as being "responsive" or "connected" to another element, it can be directly responsive or connected to the other element, or indirectly responsive or connected to the other element through one or more other elements. In contrast, when an element is referred to as being "directly responsive" or "directly connected" to another element, there are no intervening elements. As used herein the term "and / or" includes any and all combinations of one or more of the associated items and can be abbreviated as " / ".

[0038] It will be understood that, although the terms first, second, etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the teachings of the present principles.

[0039] Although some of the diagrams include arrows on communication paths to show a primary direction of communication, it is to be understood that communication can occur in the opposite direction to the depicted arrows.

[0040] Some examples are described with respect to block and operational flow diagrams in which each block represents a circuit element, a module, or a portion of code that comprises one or more executable instructions for implementing the specified logical function. It should also be noted that in other implementations the function(s) noted in the blocks can occur out of the order noted in the figure. For example, two blocks shown in succession can in fact be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved.

[0041] Reference in this document to "one example" or "an example" means that a particular feature, structure, or characteristic described in connection with the example is included in at least one implementation of the present principles. The appearances of the phrase "one example" or "an example" in various places in the specification are not necessarily all referring to the same example, nor are they necessarily mutually exclusive, or alternative examples to one another.

[0042] The reference signs in the claims are presented purely for purposes of illustration and do not limit the scope of the claims. Although not explicitly described, the present examples and variants can be taken in any combination or sub-combination.

[0043] Figure 1 A three-dimensional (3D) model 10 of an object and points of a point cloud 11 corresponding to the 3D model 10 are shown. The 3D model 10 and the point cloud 11 can for example correspond to a possible 3D representation of an object of a 3D scene comprising other objects. The model 10 can be a 3D mesh representation and the points of the point cloud 11 can be the vertices of the mesh. The points of the point cloud 11 can also be points distributed on the surface of the faces of the mesh. The model 10 can also be represented as a splatted version of the point cloud 11, the surface of the model 10 being created by splatting the points of the point cloud 11. The model 10 can be represented by many different representations such as voxels or splines. Figure 1 The fact that a point cloud can be defined with a surface representation of a 3D object and that a surface representation of a 3D object can be generated from a cloud of points is shown. As used herein, projecting points of a 3D object (by extension points of a 3D scene) onto an image is equivalent to projecting any representation of this 3D object, for example a point cloud, a mesh, a spline model or a voxel model.

[0044] A point cloud can be represented in memory as for example a vector-based structure where each point has its own coordinates in the frame of reference of a viewpoint (for example three-dimensional coordinates XYZ, or a solid angle and a distance from / to the viewpoint (also called depth) and one or more attributes, also called components. One example of a component is a color component which can be represented in various color spaces, for example RGB (red, green and blue) or YUV (Y is the luminance component and UV are two chrominance components). The point cloud is a representation of a 3D scene comprising an object. The 3D scene can be seen from a given viewpoint or range of viewpoints. The point cloud can be obtained in many ways, for example:

[0045] • from a capture of a real object taken by a camera rig, optionally complemented with depth active sensing devices;

[0046] • from a capture of a virtual / synthetic object taken by a virtual camera rig in a modeling tool;

[0047] •• from a mix of both real and virtual objects.

[0048] Figure 2Non-limiting examples of encoding, transmitting and decoding data representative of a sequence of 3D scenes are shown. The encoding format can be compatible with 3DoF, 3DoF+ and 6DoF decoding, for example, simultaneously.

[0049] A sequence of 3D scenes 20 is obtained. As a sequence of pictures is a 2D video, a sequence of 3D scenes is a 3D (also called volumetric) video. The sequence of 3D scenes can be provided to a volumetric video rendering device for 3DoF, 3Dof+ or 6DoF rendering and display.

[0050] The sequence of 3D scenes 20 can be provided to an encoder 21. The encoder 21 takes as input one 3D scene or a sequence of 3D scenes and provides a bitstream representative of the input. The bitstream can be stored in a memory 22 and / or on an electronic data medium and can be transmitted over a network 22. The bitstream representative of the sequence of 3D scenes can be read from the memory 22 and / or received from the network 22 by a decoder 23. The decoder 23 takes as input the bitstream and provides a sequence of 3D scenes in a point cloud format, for example.

[0051] The encoder 21 can comprise several circuits implementing several steps. In a first step, the encoder 21 projects each 3D scene onto at least one 2D picture. A 3D projection is any method of mapping three-dimensional points into a two-dimensional plane. As most current methods for displaying graphical data are based on planar (pixel information from several bitplanes) two-dimensional media, the use of this type of projection is widespread, especially in computer graphics, engineering and cartography. The projection circuit 211 provides a sequence of 3D scenes 20 with at least one two-dimensional frame 2111. The frame 2111 comprises color information and depth information representative of the 3D scene projected onto the frame 2111. In a variant, the color information and the depth information are encoded in two separate frames 2111 and 2112. In one embodiment, the points of a 3D scene carry not only geometry attributes and color attributes. For example, the points of a scene can have normal attributes, transparency attributes, diffuse or specular reflection attributes. Other attributes not directly related to the position and color of a point can be part of a 3D model of the scene, for example, semantic attributes associating a point with an object (e.g. a person, a tree, a wall, a floor, etc.) or a part of an object (e.g. a head, an arm, a leaf, etc.). In this embodiment, these attributes are projected onto several frames, one frame per attribute or in a variant, one frame with several attributes per pixel.

[0052] The metadata 212 is used and updated by the projection circuit 211. The metadata 212 comprises information on the projection operation (e.g. projection parameters) and information on the way the color and depth information is organized within the frames 2111 and 2112, as described in connection with the Figures 5 to 7

[0053] ​The video encoding circuit 213 encodes the sequence of frames 2111 and 2112 into a video. The pictures 2111 and 2112 of the 3D scene (or the sequence of pictures of the 3D scene) are encoded by the video encoder 213 in a stream. The video data and the metadata 212 are then encapsulated by the data encapsulation circuit 214 in a data stream.

[0054] The encoder 213 is for example compatible with encoders such as:

[0055] - JPEG, specification ISO / CEI 10918-1 UIT-T Recommendation T.81, https: / / www.itu.int / rec / T-REC-T.81 / en;

[0056] - AVC, also known as MPEG-4 AVC or h264. Specified in both UIT-T H.264 and ISO / CEI MPEG-4 Part 10 (ISO / CEI 14496-10), http: / / www.itu.int / rec / T-REC-H.264 / en, HEVC (whose specification is found on the ITU website, T recommendation, H series, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en);

[0057] - 3D-HEVC (extension of HEVC, whose specification is found on the ITU website, T recommendation, H series, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en annex G and I);

[0058] - VP9 developed by Google; or

[0059] - AV1 (AOMedia Video 1) developed by the Alliance for Open Media.

[0060] - AV1 (AOMedia Video 1) developed by the Alliance for Open Media.

[0061] The data stream is stored in a memory accessible by the decoder 23, for example through the network 22. The decoder 23 comprises different circuits implementing different decoding steps. The decoder 23 takes as input the data stream generated by the encoder 21 and provides a sequence of 3D scenes 24 to be rendered and displayed by a volumetric video display device such as a head-mounted device (HMD). The decoder 23 obtains the stream from the source 22. For example, the source 22 belongs to a group comprising:

[0062] - local memory, such as video memory or RAM (or Random Access Memory), flash memory, ROM (or Read Only Memory), hard disk;

[0063] - a storage interface, such as an interface with a mass storage device, RAM, flash memory, ROM, optical or magnetic support;

[0064] - a communication interface, such as a wired interface (for example a bus interface, a wide area network interface, a local area network interface) or a wireless interface (such as an IEEE 802.11 interface or a Bluetooth® interface);

[0065] and

[0066] - a user interface enabling a user to input data, such as a graphical user interface.

[0067] The decoder 23 comprises a circuit 234 for extracting the data encoded in the data stream. The circuit 234 takes as input the data stream and provides metadata 232 corresponding to the metadata 212 encoded in the stream and a two-dimensional video. The video is decoded by a video decoder 233 providing a sequence of frames. The decoded frames comprise color and depth information. In a variant, the video decoder 233 provides two sequences of frames, one containing color information and the other containing depth information. In one embodiment, other properties than depth and color are encoded in the frames. In this embodiment, the pixels of the frames have more than two components. In a variant of this embodiment, the video decoder 233 provides more than two sequences of frames, one for each property. The circuit 231 uses the metadata 232 to back-project the color and depth information from the decoded frames to provide a sequence of 3D scenes 24. The sequence of 3D scenes 24 corresponds to the sequence of 3D scenes 20, possibly with a loss of precision related to the encoding as a 2D video and to the video compression.

[0068] For example, other circuits and functionalities can be added, for example before the back-projection step by the circuit 231 or in a post-processing step after the back-projection. For example, a circuit can be added to re-illuminate the scene from another light located at any position in the scene. Collision detection can be performed on the depth composition, for example to add new objects to the 3DoF+ scene in a consistent realistic way or for path planning. Such circuits can require geometry information and / or color information on the 3D scene that are not used for the 3DoF+ rendering itself. The semantics of the different kinds of information must be indicated by the bitstream representing the 3DoF+ scene.

[0069] Figure 3 An exemplary architecture of a device 30 that can be configured to implement the methods described with respect to Figure 8 and Figure 9 is shown. Figure 2 ​The encoder 21 and / or the decoder 23 can implement this architecture. Alternatively, each circuit in the encoder 21 and / or the decoder 23 can be a device according to the architecture of Figure 3 linked together, for example, via its bus 31 and / or via the I / O interface 36.

[0070] The device 30 comprises the following elements linked together by a data and address bus 31 :

[0071] - a microprocessor 32 (or CPU), which is, for example, a DSP (or Digital Signal Processor);

[0072] - a ROM (or Read Only Memory) 33;

[0073] - a RAM (or Random Access Memory) 34;

[0074] - a storage interface 35;

[0075] - an I / O interface 36 for receiving data to be transmitted from an application; and

[0076] - a power supply, for example a battery.

[0077] According to one example, the power supply is external to the device. In each of the mentioned memories, the word "register" used in the description can correspond to an area of small capacity (a few bits) or to a very large area (for example, the entire program or a large amount of received or decoded data). The ROM 33 comprises at least the program and the parameters. The ROM 33 can store, according to the principles of the application, algorithms and instructions for performing the techniques. When switched on, the CPU 32 uploads the program in the RAM and executes the corresponding instructions.

[0078] The RAM 34 comprises the program in the registers executed by the CPU 32 and uploaded after the switch-on of the device 30, the input data in the registers, the intermediate data in the different states of the method in the registers, and other variables for executing the method in the registers.

[0079] The specific implementations described herein can be implemented in, for example, a method or a process, an apparatus, a computer program product, a data stream, or a signal. Even if implementations are discussed in a specific form of implementation (for example, only as a method or device), implementations of the discussed features can be implemented in other forms (for example, a program). The device can be implemented, for example, in appropriate hardware, software, and firmware. The method can be implemented, for example, in an apparatus, generally referred to as a processing device, such as, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices, such as, for example, a computer, a cell phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate communication of information between end users.

[0080] According to examples, the device 30 is configured to implement the method described with respect to Figure 8 and Figure 9 relating to the method, and belongs to the set comprising:

[0081] - a mobile device;

[0082] - a communication device;

[0083] - a game device;

[0084] - a tablet (or tablet computer);

[0085] - a laptop;

[0086] - a still picture camera;

[0087] - a video camera;

[0088] - an encoding chip;

[0089] - a server (for example a broadcast server, a video on demand server or a web server).

[0090] Figure 4 An example of an implementation of the syntax of a stream when transmitting data through a packet-based transmission protocol is shown. Figure 4 An exemplary structure 4 of a volumetric video stream is shown. This structure is contained in a container that organizes the stream in independent elements of syntax. This structure can comprise a header part 41, which is a set of data common to each syntax element of the stream. For example, the header part includes some metadata about the syntax elements, describing the nature and role of each of them. The header part can also include a part of the metadata 212 of the Figure 2 , for example the coordinates of the central viewpoint used to project the points of the 3D scene onto the frames 2111 and 2112. This structure comprises a payload, which includes syntax elements 42 and at least one syntax element 43. The syntax elements 42 comprise data representing depth frames, and if other attributes exist, different attribute frames (for example color, normal, transparency, specular, etc.). The images can have been compressed according to a video compression method.

[0091] The syntax element 43 is part of the payload of the data stream, and can comprise metadata about how the frames of the syntax elements 42 are encoded, for example parameters used to project and pack the points of the 3D scene onto the frames. Such metadata can be associated with each frame of the video or with groups of frames (also called Groups of Pictures (GoP) in video compression standards).

[0092] Figure 5A tile-based atlas approach is illustrated with an example of 4 projection centers. The 3D scene 50 includes a person. For example, the projection center 51 is a perspective camera and the camera 53 is an orthographic camera. The cameras can also be omnidirectional cameras with e.g. spherical mapping (e.g. equirectangular mapping) or cubic mapping. According to the projection operation described in the projection data of the metadata, the 3D points of the 3D scene are projected onto the 2D plane associated with the virtual camera located at the projection center. In Figure 5 In the example, the projection of the points captured by the camera 51 is mapped onto the tile 52 according to a perspective mapping and the projection of the points captured by the camera 53 is mapped onto the tile 54 according to an orthographic mapping. The pixels of the tile include at least the geometry attribute, typically the depth attribute (i.e. quantification of the distance between the center of the projection and the projected point). The pixels of the tile can also store other attributes such as the color or the texture coordinates or the transparency of the projected point, etc.

[0093] The aggregation of the pixels is performed and a plurality of 2D tiles is generated, which are packed in a rectangular atlas 55. The organization of the tiles within the atlas defines the atlas layout. In one embodiment, the pixels of the atlas (i.e. the pixels of each tile packed in the atlas) include a plurality of components because there are point attributes projected onto the projection map. In another embodiment, the tile-based atlas approach generates as many atlases as there are point attributes projected onto the projection map. In both embodiments, the projection geometry attribute (e.g. the depth value or the 3D coordinates). The inverse projection (also called the back projection) of the pixels of the atlas requires the geometry attribute. Each point of the 3D scene has at least the geometry attribute. Other attributes can be attributed to only a part of the points of the 3D scene. Even the color attribute can be attributed to only a part of the points, for example in a medical application where the subject has no specific color (i.e. a default color) and only one organ under study. Other attributes such as the transparency can be involved only for a part of the points of the 3D scene. In the second embodiment, each atlas has the same atlas layout and shares the same metadata, as detailed below. Two tiles captured by the same camera or by two different cameras can include information representative of the same part of the 3D scene, like e.g. the tiles 54 and 56.

[0094] The packing operation generates a tile data item for each generated tile. The tile data item includes a reference to the projection data (e.g. an index in the table of projection data or a pointer to the projection data (i.e. an address in memory or in a data stream)) and information describing the position and the size of the tile within the atlas (e.g. the top-left corner coordinates, the size and the width in pixels). The tile data item is added to the metadata to be encapsulated in the data stream in association with the compressed data of one, two or more atlases.

[0095] In the prior art, for a given number of projected attributes, the same number of atlases with the same layout is generated. For example, for two attributes A and B, a set of tiles is obtained, for example according to the method described above. In the case of two attributes A and B, two categories of tiles can be distinguished. The pixels of a first tile store the value of attribute A and do not store the value of attribute B (or a predetermined default value is omitted). The pixels of a second tile store the value of both attributes A and B. The existing method creates a first atlas image encoding the first attribute A and a second atlas image encoding the second attribute B, the first atlas packing each tile of the set, the second atlas having the same size as the first atlas and packing only the second tiles in the same packed position, orientation and size as the first tiles in the first atlas. Thus, the second atlas is partially empty: at each position where a first tile exists in the first atlas, the same position in the second atlas exists with a pixel without value (i.e. a default value, for example 0). Large rectangles of pixels with the same value can be easily compressed and slightly increase the bit rate of the generated stream. Thus, there is a difference between the pixel rate (the image size in raw pixels (width x height) that a GPU can manage per unit of time) and the bit rate (i.e. the image size in bits after compression). However, the pixel rate (i.e. the memory space and access at the renderer side) is twice the sum of the bit depth of the two attributes (i.e. the number of bits needed to encode the value of each attribute) times the size of the atlas image. The pixel rate is multiplied by the number of attributes. For attributes that are due to a small number of points of the scene (e.g. transparency), this very high pixel rate is a useless consumption of memory and processor resources.

[0096] Figure 6A set of 60 patches 61 to 67 is generated for storing one or two attributes. The principle of the invention proposes a format for encoding atlas images and associated metadata to reduce the bitrate and the pixel rate required for a data stream representing a volumetric scene. According to the principle of the invention, at least two attributes A and B of points of a scene are projected onto a projection map and gathered into patches. A set of 60 patches is obtained. In the set 60, the second patches 65(A and B) to 67(A and B) store two values in their pixels, one being a value of attribute A (65A to 67A) and one being a value of attribute B (65B to 67B). The first patches 61A to 64A store only values of attribute A (61A to 64A) and do not store values of attribute B. According to the principle of the invention, a first atlas 68 is generated by packing the pixels of attribute A of each patch 61A to 67(A and B). The packing step reorganizes the patches 61(A) to 67(A and B) so as to minimize the non-used areas of the atlas 68. To this end, the rectangles of the patches are reorganized according to a layout. They can be oriented in a direction different from their original orientation in the obtained set 60. The position, the size and the orientation of each patch are encoded in the metadata associated with the atlas 68. According to the principle of the invention, a second atlas 69 is generated by packing only the pixels of the second patches storing attribute B 65B to 67B. The second patches can be located at different coordinates and oriented in different directions than the atlas 68. Thus, the second patches are a subset of the set of 60 patches 60 and the size (= width x height) of the pixels of the atlas 69 is smaller than the size of the pixels of the atlas 68. In a variant, the proportions of the patches of the set 60 can be readjusted when packing. According to the principle of the invention, the proportions of the patches packed in the first atlas 68 can be different from the same patches in the second atlas 69. In the example of Fig. 2, the patch 66A is not readjusted in proportion for its attribute A in the first atlas 68, but is reduced in proportion for its attribute B in the second atlas 69. Figure 6

[0097] In a variant, a third category of patches can be provided; the pixels of the third patches store values of attribute B and do not store values of attribute A (or a default value). This variant is not developed in the present application. In the present text, the geometry attribute (attribute A) is mandatory and the patches of this category do not appear in the context of the principle of the invention.

[0098] ​In other variants, more than two attributes of the points of the projected scene are packed. For three attributes A, B and C, four categories are determined: (A), (A,B), (A,C), (A,B,C). According to the principles of the application, a first atlas packing the four categories of tiles (each tile storing attribute A) is generated, a second atlas packing the second and fourth tiles (for attribute B) is generated, and a third atlas packing the third and fourth tiles (for attribute C) is generated; each of these three atlases has its own atlas layout. The principles of the application are applicable to any number of attributes, without loss of generality. For four attributes, eight categories can be identified for the generation of four atlases, etc.

[0099] The metadata associated with the list of generated atlases must indicate the different layouts of the atlases. A possible syntax to indicate the specific packing for each attribute can be the following one.

[0100] The miv_atlas_sequence_params(vuh_atlas_id) element describes the atlas parameters that apply to the whole sequence for each attribute. In particular, three following syntax elements (the atlas frame horizontal and vertical dimensions and a flag enabling patch scale adjustment) are parameters that are different for each attribute and that are relevant to the described method.

[0101] masp_attr_frame_width_minus1[i] + 1 and masp_attr_frame_height_minus1[i] + 1 specify the dimensions of the atlas for the i-th attribute.

[0102] masp_attribute_per_patch_scale_enable[i] is a binary flag enabling patch scale adjustment of the i-th attribute in the atlas.

[0103]

[0104]

[0105] For each attribute, including the geometry, the patch_data_unit(p) syntax structure with index p of the tile describes how the tile is packed in the atlas of tiles: more precisely, it specifies its presence and, if present, its position and dimensions.

[0106] The syntax elements pdu_2d_pos_x[p] and pdu_2d_pos_y[p] are distributed set to pdu_geo_atlas_pos_x[p] and pdu_geo_atlas_pos_y[p] to indicate that they specify only the position of the top-left corner of the tile in the atlas of geometry.

[0107] The syntax elements pdu_2d_size_x_minus1 [p] and pdu_2d_size_y_minus1 [p] only specify the tile size in the source view with index equal to pdu_view_id[p] and in the geometry atlas (and not anymore in the attribute atlas).

[0108] For each attribute with index i:

[0109] pdu_attr_atlas_present_flag[p][i] is a binary flag indicating whether a tile is present in the atlas of the i-th attribute

[0110] pdu_attr_atlas_pos_x[p][i] and pdu_attr_atlas_pos_y[p][i] specify the position of the top-left corner of the tile in the atlas of the i-th attribute.

[0111] pdu_attr_atlas_orientation_idx[p][i] indicates the tile orientation index in the atlas of the i-th attribute

[0112] pdu_attr_atlas_size_x_minus1[p][i]+1 and pdu_attr_atlas_size_y_minus1[p][i]+1 specify the size of the tile in the atlas of the i-th attribute

[0113]

[0114]

[0115] Figure 7 The determination of the displacement attribute according to a non-limiting embodiment of the principles of the application is illustrated. Some cases of insufficient depth are present, which are transported by the geometry attribute of the atlas. Indeed, although there are video coding profiles for more depth bits, most video coding implementations work on 10 bits and this leads to the appearance of depth quantization errors, which are sensitive when the configuration of the camera fixture is not limited to a small area.

[0116] Figure 7The case is shown where the scenes 70 captured by the projection of the cameras 71 to 75 inwardly aligned to the center of the scene and their truncations therefore cross each other. This volume content is expected to be consumed in the same way, i.e. the user looks roughly in the same position and in the same direction of the fields of view 71 to 75. Potentially, the user at the position of the camera 72 will see some information delivered by the field of view 75 as seen in the profile. However, the depth quantification law 76 of the field of view in the volume content is generally based on a 1 / z law designed to minimize the depth quantification error to approach the original camera, as illustrated by the straight parallel lines away from the camera field of view 75. It is not desirable to see the depth delivered by 75 as seen in the profile by 72. This would lead to very noticeable artifacts at the contour. To overcome this problem, the solution is to allocate more bits to the depth.

[0117] According to the principles of the invention, the depth geometry information is kept at the normal bit depth quantized by the 1 / z law (e.g. 8 bits or 10 bits or 16 bits) while it is complemented by a depth quantized by uniform coding to express the small difference between the effective coded depth and the target fine depth. This "shift attribute" can be determined according to the following steps:

[0118] • Quantize the geometry information by the 1 / z law according to the real metric depth Zmetric(scene): Zquantized(geometry) ;

[0119] • Determine the quantized depth delivered to the decoder side by abstracting the depth coding error Zquantized(geometry) ;

[0120] • Obtain the value Zrecovered_metric(Zquantized(geometry)) by a dual operation of comparison with the real metric depth;

[0121] • Encode this small difference Zmetric(scene) - Zrecovered_metric(Zquantized(geometry)) generally by a linear quantization law, e.g. in gray levels in parallel with the extended attribute. A small amount of metadata is needed to derive the metric value from the quantized linear value, e.g. the metric value of the quantization unit.

[0122] This "shift attribute" only shifts the depth delivered by the geometry information of the first atlas by a small number of bits. The geometry information of the first atlas represents the most significant bits while the shift attribute represents the least significant bits. Shift is also a term used in graphics engineering to apply a small geometry deformation to the mesh geometry in a specific shader.

[0123] In a preferred embodiment, no scaling down is performed on patches carrying displacement attributes. It is relevant to apply such adjustment to patches for which rendering artifacts would be very visible, typically patches of foreground objects. Such attributes are well designed to be delivered to specific atlases according to the principles of the application.

[0124] Figure 8 A method 80 for decoding a data stream representing volumetric content according to non-limiting embodiments of the principles of the application is illustrated. At step 81, a first atlas and a second atlas are decoded from the stream. Their sizes can not be identical. Pixels of the first and second atlases store values representative of attributes of points of the volumetric content. In one embodiment, pixels of the first atlas encode a first attribute and pixels of the second atlas encode a second attribute different from the first attribute. At step 82, metadata associated with the two atlases are decoded from the data stream. The metadata comprise a list of data items including:

[0125] - a position and a size of a region of the first atlas corresponding to the current patch; the patch encoding a projection of a portion of points of the 3D scene of the volumetric content;

[0126] - a flag indicating whether the current patch is present in the second atlas; and

[0127] - on condition that the flag indicates that the current patch is present in the second atlas, a position of a region of the second atlas corresponding to the current patch.

[0128] In one embodiment, decoding the atlases and the metadata can be used to generate at step 83 the 3D scene of the volumetric content by back-projecting pixels of the two atlases according to the principles of the application.

[0129] Figure 9 A method 90 for encoding a volumetric content in a data stream according to non-limiting embodiments of the principles of the application is illustrated. At step 91, a set of patches is obtained from a source. The patches are pictures encoding a projection of a portion of points of a 3D scene of the volumetric content. The set of patches comprises a first patch encoding a projection of a first attribute of a portion of points of the 3D scene and a second patch encoding a projection of a second attribute and the first attribute of a portion of points of the 3D scene of the volumetric content. At step 92, a first atlas image is generated by packing the first patch and the second patch of the set and a second atlas image is generated by packing the second patch. At step 93, a data stream is generated. The data stream comprises the first image, the second image and associated metadata including, for a current patch of the set of patches:

[0130] - a position and a size of a region of the first image corresponding to the current patch;

[0131] - a flag indicating whether the current tile is a second tile; and

[0132] - a location of a region of the second image corresponding to the current tile, on condition that the current tile is a second tile.

[0133] The implementations described herein can be implemented in, for example, a method or a process, an apparatus, a computer program product, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method or device), the implementation of the features discussed can also be implemented in other forms (for example, a program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, an apparatus such as, for example, a processing device generally, such as for example a processor, that includes, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, smartphones, tablets, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.

[0134] Implementations of various processes and features described herein can be embodied in a variety of different equipment or applications, particularly equipment or applications associated with data encoding, data decoding, view generation, texture processing, and other processing of images and related texture information and / or depth information. Examples of such equipment include an encoder, a decoder, a post-processor processing output from a decoder, a pre-processor providing input to an encoder, a video encoder, a video decoder, a video codec, a web server, a set-top box, a laptop, a personal computer, a cell phone, a PDA, and other communication devices. As should be apparent, the equipment can be mobile, even if implemented in a stationary device.

[0135] Additionally, the methods can be implemented by instructions executed by a processor, and such instructions (and / or data values produced by implementations) can be stored on a processor-readable medium such as, for example, an integrated circuit, a software carrier, or another storage device such as, for example, a hard disk, a compact diskette ("CD"), an optical disk, such as, for example, a DVD, generally referred to as a digital universal disk or digital video disk, random access memory ("RAM"), or read-only memory ("ROM"). The instructions can form an application program embodied on a processor-readable medium. The instructions can be, for example, in hardware, firmware, software, or a combination. The instructions can be found in, for example, an operating system, a separate application, or a combination of the two. The processor can be characterized, therefore, as being configured to perform a process in response to the processor being made to execute the instructions. The processor can also be considered to be entirely hardware-based, even if the processor comprises a software part or a firmware part. A processor-based entity can comprise a processor and a memory coupled with the processor, such that the memory provides the processor with instructions the processor can execute to complete a process. A processor-based entity can also comprise a processor and a memory coupled with the processor, such that the memory provides the processor with instructions the processor can execute to complete a process, the memory and processor being further configured to produce a data stream comprising the instructions. The instructions can also be considered a processor-readable medium containing a computer program, determinable by a suitably configured processor. In addition to the instructions, the processor-readable medium can store data values produced by implementations.

[0136] It will be apparent to those skilled in the art that a specific embodiment can generate various signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method or data generated by one of the specific embodiments. For example, a signal can be formatted to carry as data the rules for writing or reading the syntax of a described embodiment, or to carry as data the actual syntax values written by a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a portion of the spectrum from radio waves to gamma rays) or a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The signal that carries the information can be, for example, analog or digital. As is known, signals can be transmitted through various types of wired and wireless media. Signals can be stored on processor-readable media.

[0137] A number of implementations have been described. Nevertheless, it will be understood that numerous modifications can be made. For example, elements of different implementations can be combined, supplemented, modified, or removed to produce other implementations. In addition, one of ordinary skill will understand that other structures and processes can be substituted for those disclosed and the resulting implementations will perform at least substantially the same function(s) in at least substantially the same way(s) to achieve at least substantially the same result(s). Accordingly, these and other implementations are contemplated by this application.

Claims

1. A method for different atlas packing of volumetric video, the method comprising obtaining a first image, a second image and associated metadata from a data stream; The metadata comprises a list of tile data items, a tile data item comprising: - a first position and size of a region of the first image corresponding to a current tile, a tile encoding a projection of a portion of points of a 3D scene; - for each tile, a flag indicating whether the tile is present in the second image; and - on condition that the flag indicates that the tile is present in the second image, a second position of a region of the second image corresponding to the tile; and, for a tile data item of the list of tile data items: - decoding the tile of the first image at the first position and of the size, and - on condition that the flag indicates that the current tile is present in the second image, decoding a corresponding tile of the second image at the second position and of the size.

2. The method of claim 1, wherein, A tile present in the first image encodes a first attribute of a portion of points of the 3D scene onto the tile, and if present in the second image, the tile encodes a second attribute of the portion of points of the 3D scene, the first attribute being different from the second attribute.

3. The method of claim 1 or 2, wherein, The tile data item comprises projection information of the current tile, and wherein the tiles of the first image encode a geometry attribute; the method further comprises generating a 3D scene from the decoded tiles.

4. An apparatus for different atlas packing of volumetric video, comprising a processor configured to decode a first image, a second image and associated metadata from a data stream; The metadata comprises a list of tile data items, a tile data item comprising: - a first position and size of a region of the first image corresponding to a current tile, a tile encoding a projection of a portion of points of a 3D scene; - for each tile, a flag indicating whether the tile is present in the second image; and - on condition that the flag indicates that the tile is present in the second image, a second position of a region of the second image corresponding to the tile; and, for a tile data item of the list of tile data items: - decoding the tile of the first image at the first position and of the size, and - on condition that the flag indicates that the current tile is present in the second image, decoding a corresponding tile of the second image at the second position and of the size.

5. The apparatus of claim 4, wherein, A tile present in the first image encodes a first attribute of a portion of points of the 3D scene onto the tile, and if present in the second image, the tile encodes a second attribute of the portion of points of the 3D scene, the first attribute being different from the second attribute.

6. The apparatus of claim 4 or 5, wherein, The tile data item comprises projection information of the current tile, and wherein the tiles of the first image encode a geometry attribute; the processor is further configured for generating a 3D scene from the decoded tiles.

7. A method for different atlas packing of volumetric videos, the method comprising: - obtaining a set of first tiles and second tiles, a first tile encoding a projection of a first attribute of a portion of points of a 3D scene; a second tile encoding a projection of a second attribute and of the first attribute of the portion of points of the 3D scene; - encoding the first attribute in the first image by packing the first and second tiles of the group into the first image, and encoding the second attribute in the second image by packing the second tile of the group into the second image; - generating a data stream comprising the first image, the second image and associated metadata, the associated metadata comprising, for a current tile of the group: • a first position and size of a region of the first image corresponding to the current tile; • information indicating whether the current tile is a second tile; and • a second position of a region of the second image corresponding to the current tile, on condition that the current tile is a second tile.

8. The method of claim 7, wherein, The first attribute is different from the second attribute.

9. The method of claim 7 or 8, wherein, The tile data item comprises projection information of the current tile, and wherein the first attribute is a geometry attribute.

10. A device for different atlas packing of volumetric video, comprising a processor configured for: - obtaining a set of first patches and second patches, the first patches encoding a projection of a first attribute of a portion of points of the 3D scene; encoding a projection of a second attribute and the first attribute of the part of points of the 3D scene for a second tile; - encoding the first attribute in the first image by packing the first and second tiles of the group into the first image, and encoding the second attribute in the second image by packing the second tile of the group into the second image; - generating a data stream comprising the first image, the second image and associated metadata, the associated metadata comprising, for a current tile of the group: • a first position and size of a region of the first image corresponding to the current tile; • information indicating whether the current tile is a second tile; and • a second position of a region of the second image corresponding to the current tile, on condition that the current tile is a second tile.

11. The apparatus of claim 10, wherein, The first attribute is different from the second attribute.

12. The apparatus of claim 10 or 11, wherein, The tile data item comprises projection information of the current tile, and wherein the first attribute is a geometry attribute.

13. A computer program product having instructions stored thereon, the instructions, when executed, causing a computing device to perform the method of any one of claims 1-3.

14. A computer program product having instructions stored thereon, the instructions, when executed, causing a computing device to perform the method of any one of claims 7-9.

Citation Information

Patent Citations

  • Method, apparatus and stream for immersive video format

    US20190371051A1