Method and apparatus for encoding and decoding volumetric content in and from a data stream

By extracting attributes and geometric atlas images from the data stream and performing back projection using projection parameters and metadata, the problems of large data volume and poor user experience in volumetric video encoding are solved, achieving efficient storage and transmission, and enhancing immersion and free navigation experience.

CN116458161BActive Publication Date: 2026-05-29INTERDIGITAL CE PATENT HOLDINGS SAS

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INTERDIGITAL CE PATENT HOLDINGS SAS
Filing Date
2021-06-14
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as large data volume, storage space, network transmission, and decoding performance when encoding and decoding large field-of-view content, especially large-volume videos, resulting in poor user experience. In particular, in 3DoF videos, it may cause dizziness and make it impossible to navigate freely.

Method used

By extracting attribute and geometric atlas images from the data stream, performing back projection using projection parameters and metadata, and combining different attribute atlas image encoding and decoding, data packaging and transmission are optimized to adapt to rendering needs with different degrees of freedom.

Benefits of technology

It effectively reduces data volume, improves storage and transmission efficiency, prevents dizziness, enhances user immersion and perception of scene depth, and realizes free navigation and parallax experience of 6DoF video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116458161B_ABST
    Figure CN116458161B_ABST
Patent Text Reader

Abstract

Methods and apparatus for encoding and decoding volumetric scenes are disclosed. A set of attribute and geometry tiles is obtained by projecting samples of a volumetric scene onto tiles according to projection parameters. If a geometry tile corresponds to a planar layer at constant depth according to the projection parameters, only the attribute tile is packed in an attribute atlas image and a depth value is encoded in metadata. Otherwise, both the attribute tile and the geometry tile are packed in an atlas. At decoding, if metadata of an attribute tile indicates that its geometry structure can be determined from the projection parameters and a constant depth, the attribute is back-projected on a planar layer. Otherwise, the attribute is back-projected according to the associated geometry tile.
Need to check novelty before this filing date? Find Prior Art

Description

1. Technical Field

[0001] The principles of this invention generally relate to the domain of three-dimensional (3D) scenes and volumetric video content. This document is also understood in the context of encoding, formatting, and decoding data representing the textures and geometry of 3D scenes for rendering volumetric content on end-user devices such as mobile devices or head-mounted displays (HMDs). 2. Background Technology

[0002] This section aims to introduce the reader to various aspects of the art that may relate to the aspects of the inventive principles described and / or claimed below. This discussion is believed to help provide the reader with background information to facilitate a better understanding of the various aspects of the inventive principles. Therefore, it should be understood that these statements should be interpreted in this light, rather than as an admission of prior art.

[0003] Recently, the availability of wide field-of-view content (up to 360°) has increased. Users viewing content on immersive display devices (such as head-mounted displays, smart glasses, PC screens, tablets, smartphones, etc.) may not be able to see the entire content. This means that at any given moment, a user can only view a portion of the content. However, users can typically navigate within the content using various means such as head movement, mouse movement, touchscreens, voice, and the like. Encoding and decoding of this content is generally required.

[0004] Immersive video (also known as 360° planar video) allows users to see everything around them by rotating their heads around a stationary viewpoint. Rotation only allows for a 3-degree-of-freedom (3DoF) experience. Even if 3DoF video is sufficient for first-time omnidirectional video experiences (e.g., using a head-mounted display (HMD device)), it can quickly become frustrating for viewers expecting more freedom (e.g., by experiencing parallax). Furthermore, 3DoF can cause dizziness because users never just rotate their heads, but also translate them in three directions, movements that are not reproduced in a 3DoF video experience.

[0005] In this context, large field-of-view content can be three-dimensional computer graphics scenes (3D CGI scenes), point clouds, or immersive videos. Many terms can be used to design such immersive videos: for example, virtual reality (VR), 360, panoramic, 4π spherical, immersive, omnidirectional, or large field of view.

[0006] Volumetric video (also known as 6DoF video) is an alternative to 3DoF video. When watching 6DoF video, in addition to rotation, users can pan their head and even their body within the content being viewed, experiencing parallax and even volume. This type of video significantly increases immersion and perception of scene depth, and prevents motion sickness by providing consistent visual feedback during head panning. The content is created using dedicated sensors, allowing for the simultaneous recording of color and depth of the scene of interest. Even though technical challenges remain, using color camera equipment incorporating photogrammetry is another way to perform this recording.

[0007] While 3DoF video comprises a sequence of images derived from the demapping of textured images (e.g., spherical images encoded according to latitude / longitude projection maps or isometric projection maps), 6DoF video frames embed information from multiple viewpoints. They can be viewed as a temporal series of point clouds generated by 3D capture. Two types of volumetric video can be considered depending on the viewing conditions. The first (i.e., full 6DoF) allows for completely free navigation within the video content, while the second (aka 3DoF+) restricts the user's viewing space to a finite volume called the viewing bounding box, thus allowing for limited head translation and parallax experience. This second case represents a valuable trade-off between free navigation and the passive viewing conditions of a seated audience.

[0008] Volumetric video (3DoF+ or 6DoF) is a sequence of 3D scenes. A solution for encoding volumetric video is to project each 3D scene of the sequence onto a projection map that is clustered into color (or other attribute) images and depth images called chunks. The chunks are packed into color and depth images stored in the video track of the video stream. This encoding has the advantage of utilizing standard image and video processing standards. During decoding, pixels of the color image are back-projected at the depth determined by information stored in the associated depth image. Such solutions are efficient. However, encoding this large amount of data into images in the video track of the video stream presents problems. The size of the bitstream raises bitrate technology issues regarding storage space, transmission over the network, and decoding performance. 3. Summary of the Invention

[0009] The following is a simplified overview of the principles of the invention to provide a basic understanding of some aspects of these principles. This summary is not a broad overview of the principles of the invention and is not intended to identify key or essential elements of the invention. The following summary presents only some aspects of the principles of the invention in a simplified form as a preface to the more detailed description that follows.

[0010] The present invention relates to a method comprising obtaining attribute atlas images and geometry atlas images from a data stream. The (attribute or geometry) atlas images are packaged into tiled frames. Each (attribute or geometry) tiled frame is a projection of a 3D scene sample. Metadata is also obtained from the data stream. The metadata for the attribute tiled frames of the attribute atlas images includes:

[0011] The projection parameters associated with the attribute-trimmed image, and

[0012] Information indicating whether an attribute tile screen is associated with a geometric tile screen of a geometry atlas image or whether an attribute tile screen is associated with a depth value encoded in metadata.

[0013] When an attribute block image is associated with a geometric block image, the method includes back-projecting pixels of the attribute block image onto positions determined by the geometric block image and projection parameters associated with the attribute block image.

[0014] Alternatively, when the attribute block image is associated with a depth value, the method includes back-projecting pixels of the attribute block image onto a location determined by the depth value and projection parameters associated with the attribute block image.

[0015] In one implementation, pixels in an attribute tile are encoded with two (or more) values ​​for different attributes (e.g., color, normal vector, illumination, heat, speed). In another implementation, each attribute receives an attribute atlas. Different attribute atlases are encoded according to the same packing layout, and metadata is applied to the corresponding tile for each attribute atlas.

[0016] The principles of the present invention also relate to a device including a processor configured to implement the above-described methods.

[0017] The principle of this invention also relates to a method, which includes:

[0018] - Obtain a set of attribute blocks associated with the geometric block blocks, which are obtained by projecting 3D scene samples according to projection parameters;

[0019] - Attribute-specific frames within a set of attribute-specific frames.

[0020] Package the attribute-segmented images into an attribute atlas image; and

[0021] If the geometric tile associated with the attribute tile corresponds to a planar layer at a location determined by the depth value and projection parameters, then metadata is generated that includes the following: projection parameters, depth value, and information indicating the association between the attribute tile and the depth value, or

[0022] In another scenario, the geometric tile images are packaged within a geometric atlas image, and metadata including projection parameters and information indicating the association between the attribute tile images and the geometric tile images is generated; and

[0023] - Encode the attribute atlas image, the geometry atlas image, and the generated metadata in the data stream.

[0024] The principles of the present invention also relate to a device including a processor configured to implement the above-described methods.

[0025] The principle of this invention also relates to a data stream, for example, generated by the method described above. This data stream includes an attribute atlas image, a geometry atlas image, and metadata. The atlas image is packaged into tiled frames, each tiled frame being a projection of a 3D scene sample. The metadata, for the attribute tiled frames of the attribute atlas image, includes:

[0026] - Projection parameters associated with the attribute tiled screen, and

[0027] Information indicating whether an attribute tile screen is associated with a geometric tile of a geometry atlas image or whether an attribute tile screen is associated with a depth value encoded in metadata.

[0028] In one implementation, pixels in an attribute tile are encoded with two (or more) values ​​for different attributes (e.g., color, normal, illumination, heat, speed). In another implementation, the data stream includes an attribute atlas for each attribute. Different attribute atlases are encoded according to the same packing layout, and metadata is applied to the corresponding tile for each attribute atlas. 4. Description of the attached drawings

[0029] This disclosure will be better understood, and further specific features and advantages will emerge after reading the following description and referring to the accompanying drawings, in which:

[0030] - Figure 1 A three-dimensional (3D) model of an object according to a non-limiting embodiment of the principles of the present invention and points of a point cloud corresponding to the 3D model are shown;

[0031] - Figure 2 Non-limiting examples of encoding, transmitting, and decoding data representing a sequence of 3D scenes according to a non-limiting embodiment of the principles of the present invention are shown;

[0032] - Figure 3 The illustration shows a non-limiting embodiment of the invention that can be configured to achieve the following: Figure 10 and Figure 11 An exemplary architecture of the device for the described method;

[0033] - Figure 4Examples of embodiments of the syntax of a stream when transmitting data via a packet-based transport protocol are shown, according to a non-limiting embodiment of the principles of the present invention;

[0034] - Figure 5 A spherical projection from a central viewpoint is shown as a non-limiting embodiment according to the principles of the present invention;

[0035] - Figure 6a An example of a texture atlas including points of a 3D scene, according to a non-limiting embodiment of the principles of the present invention, is shown;

[0036] Figure 6b shows the encoding in atlases with different layouts. Figure 6a The later color frames of a 3D scene sequence;

[0037] Figure 7a illustrates a non-limiting embodiment of the invention, including... Figure 6a An example of a 3D scene atlas containing depth information of points;

[0038] - Figure 7b shows a later depth frame of the 3D scene sequence of Figure 7a encoded in an atlas with different layouts, according to a non-limiting embodiment of the principles of the present invention.

[0039] - Figure 8 A planar cardboard scene consisting of three layers obtained from three cardboard blocks is shown as a non-limiting embodiment according to the principles of the present invention;

[0040] - Figure 9 An exemplary attribute + depth atlas layout of a non-limiting embodiment according to the principles of the present invention is shown;

[0041] - Figure 10 A method 100 for encoding volumetric content according to a non-limiting embodiment of the present invention is shown;

[0042] - Figure 11 A method 110 for decoding a data stream representing volumetric content is shown according to a non-limiting embodiment of the present invention. 5. Detailed Implementation

[0043] The principles of the invention will be described more fully below with reference to the accompanying drawings, in which examples of the principles of the invention are shown. However, the principles of the invention may be embodied in many alternative forms and should not be construed as limited to the examples set forth herein. Therefore, while the principles of the invention are susceptible to various modifications and alternatives, specific examples are shown by way of example in the drawings and will be described in detail herein. However, it should be understood that there is no intention to limit the principles of the invention to the specific forms disclosed, but rather, this disclosure is intended to cover all modifications, equivalents, and alternatives that fall within the spirit and scope of the principles of the invention as defined by the claims.

[0044] The terminology used herein is for the purpose of describing particular examples only and is not intended to limit the principles of the invention. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It will be further understood that, when used in this specification, the terms “comprising” and / or “including” specify the presence of the stated feature, integer, step, operation, element, and / or component, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, when an element is referred to as “responding” or “connected” to another element, it may directly respond to or be connected to the other element, or there may be intermediate elements present. Conversely, when an element is referred to as “directly responding” or “directly connected” to another element, there are no intermediate elements present. As used herein, the term “and / or” includes any and all combinations of one or more of the listed related items and may be abbreviated to “ / ”.

[0045] It should be understood that although the terms first, second, etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, without departing from the teachings of the principles of the invention, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element.

[0046] Although some illustrations include arrows along the communication path to show the main communication direction, it should be understood that communication can occur in the opposite direction to the arrows depicted.

[0047] Examples of block diagrams and operation flowcharts are described where each box represents a circuit element, module, or code section, the code section including one or more executable instructions for implementing a specified logical function. It should also be noted that in other specific implementations, the functions marked in the boxes may not appear in the order they are marked. For example, two boxes shown consecutively may actually be executed substantially simultaneously, or these boxes may sometimes be executed in reverse order depending on the functions involved.

[0048] The references to "according to an example" or "in an example" in this document mean that a particular feature, structure, or characteristic described in connection with the example may be included in at least one specific embodiment of the principles of the invention. The appearance of the phrases "according to an example" or "in an example" in various places in the specification does not necessarily refer to the same example in all instances, nor is it necessarily a separate or alternative example that is mutually exclusive with other examples.

[0049] The reference numerals appearing in the claims are for illustrative purposes only and do not limit the scope of the claims. Although not explicitly described, these examples and variations may be employed in any combination or sub-combination.

[0050] Figure 1 A three-dimensional (3D) model 10 of an object and points corresponding to a point cloud 11 of the 3D model 10 are shown. The 3D model 10 and point cloud 11 may, for example, correspond to possible 3D representations of objects in a 3D scene including other objects. Model 10 may be a 3D mesh representation, and the points of point cloud 11 may be vertices of the mesh. The points of point cloud 11 may also be points distributed on the surface of the mesh face. Model 10 may also be represented as a sputtered version of point cloud 11, the surface of which is created by sputtering the points of point cloud 11. Model 10 may be represented by many different representations such as voxels or splines. Figure 1 This demonstrates that a point cloud can be defined using a surface representation of a 3D object, and that a surface representation of a 3D object can be generated from cloud points. As used herein, projecting points of a 3D object (and by extension, points of a 3D scene) onto an image is equivalent to projecting any representation of that 3D object, such as a point cloud, mesh, spline model, or voxel model.

[0051] Point clouds can be represented in memory as, for example, a vector-based structure, where each point has its own coordinates (e.g., 3D coordinates XYZ, or solid angle and distance from / to the viewpoint (also called depth)) and one or more attributes, also called components, in the viewpoint's frame of reference. An example of components is color components, which can be represented in various color spaces, such as RGB (red, green, and blue) or YUV (Y is the luminance component and UV are the two chrominance components). A point cloud is a representation of a 3D scene including objects. The 3D scene can be viewed from a given viewpoint or viewpoint range. Point clouds can be obtained in various ways, such as:

[0052] • Capture of real objects from camera equipment, optionally supplemented by active depth sensing devices;

[0053] • Capture of virtual / composite objects taken by a virtual camera setup within a modeling tool;

[0054] • A mixture of real and virtual objects.

[0055] Figure 2A non-limiting example of encoding, transmitting, and decoding data representing a sequence of 3D scenes is shown. The encoding format may, for example, be compatible with 3DoF, 3DoF+, and 6DoF decoding simultaneously.

[0056] Obtain 20 3D scene sequences. Just as frame sequences are 2D video, 3D scene sequences are 3D (also known as volumetric) video. These 3D scene sequences can be provided to volumetric video rendering devices for 3DoF, 3DoF+, or 6DoF rendering and display.

[0057] A 3D scene sequence 20 can be provided to an encoder 21. The encoder 21 takes a 3D scene or a sequence of 3D scenes as input and provides a bitstream representing that input. The bitstream can be stored in a memory 22 and / or on an electronic data medium, and can be transmitted via a network 22. The bitstream representing the 3D scene sequence can be read from the memory 22 and / or received from the network 22 by a decoder 23. The decoder 23 takes the bitstream input and provides a 3D scene sequence, for example, in point cloud format.

[0058] Encoder 21 may include several circuits implementing several steps. In a first step, encoder 21 projects each 3D scene onto at least one 2D frame. 3D projection is any method of mapping three-dimensional points onto a two-dimensional plane. This type of projection is widely used, especially in computer graphics, engineering, and drafting, because most current methods for displaying graphics data are based on a planar (pixel information from several bit planes) two-dimensional medium. Projection circuitry 211 provides at least one two-dimensional frame 2111 for the sequence of 3D scenes 20. Frame 2111 includes color and depth information representing the 3D scene projected onto frame 2111. In a variant, the color and depth information are encoded in two separate frames 2111 and 2112.

[0059] Metadata 212 is used and updated by projection circuitry 211. Metadata 212 includes information about projection operations (e.g., projection parameters) and information about how color and depth information is organized within frames 2111 and 2112, such as... Figure 5 As shown in Figure 7.

[0060] The video encoding circuit 213 encodes the sequence of frames 2111 and 2112 into video. The frames 2111 and 2112 of the 3D scene (or the sequence of frames of the 3D scene) are encoded in the stream by the video encoder 213. Then, the video data and metadata 212 are encapsulated in the data stream by the data encapsulation circuit 214.

[0061] Encoder 213 is compatible with, for example, encoders such as:

[0062] -JPEG, specification ISO / CEI 10918-1UIT-T Recommendation T.81, https: / / www.itu.int / rec / T-REC-T.81 / en;

[0063] -AVC, also known as MPEG-4 AVC or h264. It is specified in both UIT-T H.264 and ISO / CEI MPEG-4 Part 10 (ISO / CEI 14496-10), http: / / www.itu.int / rec / T-REC-H.264 / en, HEVC (its specification can be found on the ITU website, T recommendation, H series, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en);

[0064] -3D-HEVC (an extension of HEVC, the specification of which can be found on the ITU website, T recommendation, H series, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en annex G and I);

[0065] - VP9 developed by Google; or

[0066] - AV1 (AOMedia Video 1) was developed by Alliance for Open Media.

[0067] The data stream is stored in a memory accessible by the decoder 23, for example, via network 22. The decoder 23 includes different circuitry implementing various decoding steps. The decoder 23 takes the data stream generated by the encoder 21 as input and provides a sequence 24 of 3D scenes to be rendered and displayed by a volumetric video display device, such as a head-mounted display (HMD). The decoder 23 obtains the stream from the source 22. For example, the source 22 belongs to a group that includes:

[0068] - Local storage, such as video storage or RAM (or random access memory), flash memory, ROM (or read-only memory), hard disk;

[0069] - Storage interfaces, such as interfaces with mass storage devices, RAM, flash memory, ROM, optical discs, or magnetic media;

[0070] - Communication interfaces, such as wired interfaces (e.g., bus interfaces, WAN interfaces, LAN interfaces) or wireless interfaces (e.g., IEEE 802.11 interfaces or...). Interface); and

[0071] - User interfaces that enable users to input data, such as graphical user interfaces.

[0072] Decoder 23 includes circuitry 234 for extracting data encoded in the data stream. Circuitry 234 takes the data stream as input and provides metadata 232 corresponding to metadata 212 encoded in the stream and two-dimensional video. The video is decoded by video decoder 233, which provides a sequence of frames. The decoded frames include color and depth information. In a variant, video decoder 233 provides two frame sequences, one containing color information and the other containing depth information. Circuitry 231 uses metadata 232 to back-project the color and depth information from the decoded frames to provide a 3D scene sequence 24. 3D scene sequence 24 corresponds to 3D scene sequence 20, potentially resulting in a loss of accuracy associated with encoding and video compression as 2D video.

[0073] Figure 3 It shows that it can be configured to implement about Figure 10 and Figure 11 An exemplary architecture of the device 30 described in the method. Figure 2 The encoder 21 and / or decoder 23 can implement this architecture. Alternatively, each circuit in the encoder 21 and / or decoder 23 can be based on... Figure 3 Devices with an architecture that are linked together, for example, via their bus 31 and / or via I / O interface 36.

[0074] Device 30 includes the following components connected together via data and address bus 31:

[0075] - Microprocessor 32 (or CPU), which is, for example, a DSP (or digital signal processor);

[0076] -ROM (or read-only memory) 33;

[0077] -RAM (or random access memory) 34;

[0078] - Storage interface 35;

[0079] -I / O interface 36, which is used to receive data to be transmitted from the application; and

[0080] - Power source, such as a battery.

[0081] According to one example, the power supply is external to the device. In each mentioned memory, the term "register" used in the specification can correspond to a small area (a few bits) or a very large area (e.g., the entire program or a large amount of received or decoded data). ROM 33 includes at least the program and parameters. ROM 33 can store algorithms and instructions for executing the technology according to the principles of the invention. When powered on, CPU 32 loads the program from RAM and executes the corresponding instructions.

[0082] RAM 34 includes a program executed by CPU 32 and uploaded after device 30 is turned on, input data in the register, intermediate data in different states of the method in the register, and other variables used to execute the method in the register.

[0083] The specific embodiments described herein may be implemented, for example, in methods or processes, apparatus, computer program products, data streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method or apparatus), the specific implementation of the discussed features may be implemented in other forms (e.g., programs). Apparatus may be implemented, for example, in suitable hardware, software, and firmware. Methods may be implemented in apparatus (such as, for example, a processor) that generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, mobile phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0084] According to the example, device 30 is configured to implement about Figure 10 and Figure 11 The described method belongs to a set that includes the following items:

[0085] -mobile device;

[0086] - Communication equipment;

[0087] -Gaming devices;

[0088] - Tablet PC (or tablet computer);

[0089] - Laptop;

[0090] - Still image camera;

[0091] -Camera;

[0092] - Encoding chip;

[0093] - Servers (such as broadcast servers, video-on-demand servers, or web servers).

[0094] Figure 4An example of an implementation of the syntax for streams is shown when data is transmitted via a packet-based transport protocol. Figure 4 An exemplary structure 4 for a volumetric video stream is shown. This structure is contained within a container that organizes the stream by syntactic, independent elements. The structure may include a header section 41, which is a set of data common to each syntactic element of the stream. For example, the header section includes metadata about the syntactic elements, describing the properties and roles of each of them. The header section may also include... Figure 2 This is part of the metadata 212, such as the coordinates of the center viewpoint used to project points of the 3D scene onto frames 2111 and 2112. The structure includes a payload comprising syntax element 42 and at least one syntax element 43. Syntax element 42 includes data representing color and depth frames. The image may have been compressed according to a video compression method.

[0095] Syntax element 43 is part of the payload of the data stream and may include metadata about how the frames of syntax element 42 are encoded, such as parameters for projecting points of the 3D scene onto the frames. Such metadata may be associated with each frame or group of frames of the video (also known as a group of frames (GoP) in video compression standards).

[0096] Figure 5 A tiled atlas method is illustrated using four projection centers as an example. The 3D scene 50 includes characters. For example, projection center 51 is a perspective camera, and camera 53 is an orthophoto camera. The camera can also be an omnidirectional camera with, for example, a spherical mapping (e.g., an isorectangular mapping) or a cubic mapping. Based on the projection operations described in the projection data of the metadata, 3D points of the 3D scene are projected onto a 2D plane associated with a virtual camera located at the projection center. Figure 5 In the example, the projection of the points captured by camera 51 is mapped onto block 52 according to perspective mapping, and the projection of the points captured by camera 53 is mapped onto block 54 according to orthophoto mapping.

[0097] Clustering of projected pixels produces multiple 2D tiles, which are packed into a rectangular atlas 55. The organization of tiles within the atlas defines the atlas layout. In one embodiment, two atlases have the same layout: one for texture (i.e., color) information and one for depth information. Two tiles captured by the same camera or by two different cameras may include information representing the same portion of the 3D scene, such as, for example, tiles 54 and 56.

[0098] The packing operation generates chunk data for each generated chunk. Chunk data includes references to the projection data (e.g., an index in the projection data table or a pointer to the projection data (i.e., an address in memory or the data stream)) and information describing the location and size of the chunk within the atlas (e.g., top-left corner coordinates, size, and width in pixels). Chunk data items are added to metadata to be encapsulated in the data stream in association with the compressed data of one or two atlases.

[0099] Figure 6a An example of an atlas 60 comprising texture information (e.g., RGB or YUV data) of points in a 3D scene, according to a non-limiting embodiment of the principles of the invention, is shown. (See attached image.) Figure 5 The atlas, as explained, is a packaged and segmented image, which is a view obtained by projecting a portion of points from a 3D scene.

[0100] In the example of Figure 6, atlas 60 includes a first portion 61 and one or more second portions 62. The first portion includes texture information of points in the 3D scene visible from the viewpoint. The texture information of the first portion 61 can be obtained, for example, according to an isometric projection map, an example of a spherical projection map. In the example of Figure 6, the second portions 62 are arranged at the left and right boundaries of the first portion 61, but the second portions can be arranged differently. The second portions 62 include texture information of a portion of the 3D scene that is complementary to the portion visible from the viewpoint. The second portions can be obtained by removing points visible from the first viewpoint (whose textures are stored in the first portion) from the 3D scene and projecting the remaining points according to the same viewpoint. The latter process can be repeated iteratively to obtain the hidden portion of the 3D scene each time. According to the variant, the second part can be obtained by removing points visible from the viewpoint (e.g., the center viewpoint) from the 3D scene (whose textures are stored in the first part) and projecting the remaining points according to viewpoints different from the first viewpoint, such as one or more second viewpoints from the viewing space centered on the center viewpoint (e.g., the viewing space rendered by 3DoF).

[0101] The first part 61 can be viewed as the first large texture tile (corresponding to the first part of the 3D scene), and the second part 62 includes smaller texture tiles (corresponding to the second part of the 3D scene that complements the first part). This type of atlas has the advantage of being compatible with both 3DoF rendering (when only the first part 61 is rendered) and 3DoF+ / 6DoF rendering.

[0102] Figure 6b shows the encoding in atlases with different layouts. Figure 6a The later color frames of the 3D scene sequence. Part 61b (which corresponds to...) Figure 6aThe first part 61 (also called the central block) is located at the bottom of the atlas image 60b, and the second part 62b (which corresponds to...) Figure 6a Part 62) is located at the top of image 60b in the atlas.

[0103] Figure 7a illustrates a non-limiting embodiment of the present invention, including... Figure 6a An example of atlas 70 containing depth information of points in a 3D scene. Atlas 70 can be viewed as corresponding to... Figure 6a The texture image is a depth image of 60.

[0104] Atlas 70 includes a first portion 71 and one or more second portions 72, the first portion including depth information of points in the 3D scene visible from a central viewpoint. Atlas 70 can be obtained in the same manner as Atlas 60, but contains depth information associated with points in the 3D scene instead of texture information.

[0105] For 3DoF rendering of a 3D scene, only one viewpoint is considered, typically a central viewpoint. The user can rotate their head with three degrees of freedom around this primary viewpoint to view different parts of the 3D scene, but the user cannot move this single viewpoint. The points in the scene to be encoded are those visible from this single viewpoint, and only texture information needs to be encoded / decoded for 3DoF rendering. For 3DoF rendering, it is not necessary to encode points in the scene that are not visible from this single viewpoint because the user cannot access them.

[0106] For 6DoF rendering, users can move their viewpoint into the scene. In this case, every point (depth and texture) of the scene in the bitstream needs to be encoded because a user who can move their viewpoint may access each point. At the encoding stage, there is no way to know in advance from which viewpoint the user will be viewing the 3D scene.

[0107] For 3DoF+ rendering, the user can move the viewpoint within a limited space around the central viewpoint. This allows for the experience of parallax. Data representing portions of the scene visible from any point in the viewing space is encoded into the stream, including data representing the 3D scene visible from the central viewpoint (i.e., the first portions 61 and 71). For example, the size and shape of the viewing space can be determined and encoded in the bitstream at the encoding step. The decoder obtains this information from the bitstream, and the renderer restricts the viewing space to the space determined by the obtained information. According to another example, the renderer determines the viewing space based on hardware constraints, such as hardware constraints related to the ability of sensors to detect user movement. In this case, if a point visible from a point within the renderer's viewing space has not yet been encoded in the bitstream at the encoding stage, that point will not be rendered. According to yet another example, data representing each point of the 3D scene (e.g., texture and / or geometry) is encoded in the stream, regardless of the rendering viewing space. To optimize the size of the stream, only a subset of the points in the scene can be encoded, such as a subset of the points visible from the rendering viewing space.

[0108] Figure 7b shows a later depth frame of the 3D scene sequence of Figure 7a encoded in atlases with different layouts. In the example of Figure 7b, the depth encoding is the inverse of the depth encoding in Figure 7. In Figure 7b, the closer the object, the brighter its projection on the tile. For example, in Figure 7, the number of depth tiles in the second part 72b is the same as the number of color tiles in the first part 62b of Figure 6. If Figure 6a The layout of the atlas in Figure 7a is different from the layout of the atlases in Figures 6b and 7b. Figure 6a Figure 7a shares the same layout as Figure 6b, and Figure 7b also shares the same layout. The metadata associated with the tiles does not indicate the differences between color tiles and depth tiles.

[0109] Figure 8A planar “cardboard” scene is illustrated, consisting of three layers obtained from three cardboard blocks. A 3D scene can be represented by several atlases encoded according to the same layout: one atlas image for color attributes, one atlas image for each of the other attributes (e.g., normal vector, lighting, heat, velocity, etc.), and one atlas image for depth (i.e., geometric attributes). However, some simple 3D samples of a scene (referred to herein as “cardboard” scenes) may not require description of such a level of complexity (i.e., block divisions of geometric attributes) due to their inherent nature. In fact, such samples consist of a set of statically stacked planar layers. A planar layer is a surface in 3D space (e.g., a flat rectangle, a sphere, a cylinder, or a combination of such surfaces forming a surface in 3D space). According to the principles of the invention, a “cardboard” scene is a part of a 3D scene represented by an ordered list of planar layers defined by the camera used to acquire the blocks. Due to this particular planar shape, transmitting depth information to represent cardboard objects is overly cumbersome, as it can be represented by simple, constant values ​​from a given camera viewpoint.

[0110] According to the principles of this invention, cardboard blocks are defined. The shape of the planar layer corresponding to the back projection of the cardboard block is determined by the projection parameters of the block itself. For example, for a block storing information obtained through orthographic projection, the planar layer of the cardboard block is a rectangular flat plane. In a variation, when the pixels of the cardboard block are obtained through spherical projection, the planar layer of the cardboard is a sphere. Any other projection mode can be used to create cardboard blocks, such as pyramidal or cubic projection modes. In each variation, the "depth" information is constant.

[0111] exist Figure 8 In the example, a first cardboard block is used to generate a planar layer 81 at a first depth z1. A second cardboard block, acquired from the same camera and therefore sharing the same projection parameters as the first cardboard block, is used to generate a planar layer 82 of the same shape at a second depth z2, which is lower than z1. A third cardboard block, having the same projection parameters as the first and second cardboard blocks, is used to generate a planar layer 83 at a depth z3, preceding the two other cardboard blocks, which is lower than z2. The cardboard blocks are encoded in a color atlas (and optionally in other attribute atlases). However, since the depth of a cardboard block is a constant associated with the block's projection parameters, the cardboard block does not have a corresponding portion in the depth atlas.

[0112] Figure 9An exemplary attribute + depth atlas layout according to the principles of the present invention is illustrated. Whether the obtained chunk is encoded in a geometry / depth atlas frame depends on whether it is to be represented as a cardboard chunk or a regular chunk. The technical effect of this method is a gain in pixel rate for transmission. In one embodiment, all chunks of a given atlas are cardboard chunks. In this embodiment, encoding of volumetric video includes only the color atlas (and optionally other attribute atlases), excluding the depth atlas.

[0113] In one implementation, cardboard tiles are collected in a single atlas, separate from the atlas of packing rule tiles. In this implementation, encoding the volumetric scene includes three atlases: a color atlas and a depth atlas with a shared layout, and a cardboard atlas without corresponding depth portions. In a variant, rule tiles and cardboard tiles can coexist, and a packing strategy is implemented to reduce the overall pixel rate, such as... Figure 9 As shown. Figure 9 The example illustrates a method that does not affect packing alignment for color and depth information. This method involves, during the encoding phase, categorizing blocks starting with regular chunking and ending with cardboard chunking. Using this MaxRect-like strategy for packing results in... Figure 4 The image depicts paired color and depth atlases. For example, regular depth blocks 91 and 92 are grouped at the bottom of the color and geometric depth frame, while cardboard blocks 93 are grouped at the top, for example. The depth atlas is cropped along line 94 to fit a useful size. This reduces the overall pixel rate.

[0114] Another implementation considers two separate packs for color and depth information. In this implementation, the color and depth / geometry atlas frames are independent. This results in better coding flexibility and efficiency. However, the packing information associated with each chunk must be copied to report both color and geometry packing.

[0115] If the cardboard section is not completely occupied (i.e., some pixels are transparent and not like in...), Figure 8 If 3D points are generated as in the example, this must be signaled. In one implementation, an occupancy map is constructed to convey the usefulness of each pixel. This information can be time-varying. Occupancy mapping is a well-known technique that can be advantageously used to convey this information. It is a binary map indicating whether a pixel in a given atlas is occupied. In an implementation where all cardboard pieces are packaged in a dedicated atlas, the atlas does not require a geometry atlas, but it does require an occupancy map. In another implementation, the transparency of pixels is encoded in a 4-channel image (e.g., an RGBA image), with the fourth channel α encoding binary or progressive transparency. According to the principles of the invention, completely transparent pixels do not generate 3D points. Partially transparent pixels generate similarly partially transparent 3D points.

[0116] According to one of the implementation schemes described above, the cardboard tiles must be signaled in the data stream so that the decoder can generate 3D points with the correct geometry. The first syntax for the metadata describing the tiles and their projection parameters (which is associated with a pair (optionally more) atlases) is to specify cardboard signaling for each tile to notify the decoder that the tile will be decoded as a regular tile or a cardboard pattern tile. In this implementation, the tile is defined by a `patch_data_unit` structure, which includes information associated with its relative position in the view from which the tile originates and its relative position in the atlas frame to which it is packaged. The `patch_data_unit` structure also includes two fields called `pdu_depth_start` and `pdu_depth_end`, which indirectly indicate the depth range of the tile. An additional flag, `pdu_cardboard_flag`, is used to signal the cardboard supplementary attributes. If the decoder is enabled, it should not attempt to retrieve the tile geometry from the depth / geometry atlas frame, but instead simply read its constant depth value within the `patch_data_unit` structure. Therefore, the pdu_depth_start information is interpreted as the actual depth of the tile (and pdu_depth_end is forced to be equal to pdu_depth_start). In another implementation, a check for equality (pdu_depth_start == pdu_depth_end) can determine whether the tile under consideration is a regular tile or a cardboard tile. In another implementation, a field pdu_cardboard_depth can be added to signal the tile depth value. An additional flag pdu_fully_occupied_flag, conditional on pdu_cardboard_flag being true, can be introduced to signal whether all pixels of the cardboard tile are valid, and then whether an occupancy map must be obtained and used, or whether the fourth channel α is used in the encoded atlas.

[0117]

[0118]

[0119]

[0120] Another implementation of the syntax signals a given camera / projection to be associated only with cardboard tiles. Each tile is created from a specific projection mode signaled within the bitstream syntax. The atlas transmits tiles from multiple projections. Tiles associated with a camera that has cardboard marking enabled must be interpreted on the decoder side as having their depth information as described in the per-tile information rather than from the depth atlas frame. When the `mvp_fully_occupied` flag equals 1, all tiles belonging to that view are fully occupied. When the `mvp_fully_occupied` flag equals 0, a tile may or may not be fully occupied.

[0121]

[0122] In another implementation, signaling for cardboard tiles is embedded at the level of the metadata describing the atlas. For example, each pair of color and depth atlases is described at a high level by a set of very general attributes, including depth scaling or grouping properties. Cardboard tags indicate to the decoder that all tiles associated with that particular atlas are cardboard tiles, and that depth information should be retrieved from per-tile information rather than from the geometry atlas frame. Cardboard tiles can be grouped within dedicated atlases, which therefore do not require any associated depth / geometry frames. When the asme_fully_occupied flag equals 1, all tiles belonging to that view are fully occupied. When the asme_fully_occupied flag equals 0, a tile may or may not be fully occupied.

[0123]

[0124] In addition to improvements in pixel rate and bit rate, one of these signaling methods utilizing the cardboard pattern can also impact decoder performance by allowing for faster strategies used for rendering. In fact, highly efficient rendering techniques such as painter algorithms can be used by the notified decoder. Such methods stack graphic primitives associated with chunks from the background to the foreground without requiring any complex visibility calculations.

[0125] Figure 10 A method 100 for encoding volumetric content is shown, according to a non-limiting embodiment of the principles of the present invention. In step 101, a volumetric scene sample to be encoded is projected onto a projection surface and, for example, according to… Figure 5The projection parameters shown are used to obtain color (or any attribute) and depth tile images. At step 102, at least one color (or any attribute) atlas is prepared. The color tiles are packaged into a color atlas. In one implementation, color tiles associated with planar layers, rather than with depth tiles, are packaged into separate color atlases. At step 103, the method determines whether the depth (or other geometric data) tiles are equivalent to the planar layers described above, i.e., whether the color tiles are cardboard tiles. For example, cardboard tiles can be created as an environment map of a background combined with regular tiles of the remainder of a volumetric scene. Scenes can be generated using regular backgrounds containing flat cardboard objects. Such flat cardboard can be used directly as cardboard tiles. Scenes can be represented as implementations in a cartoon domain, such as... Figure 8 As shown. Cardboard blocks can be created directly from such slides. If a depth block does not correspond to a planar layer, then at step 104, the depth block is packaged in a geometry atlas and associated with a color block in the metadata. Otherwise, the depth block is not encoded in any atlas, and information indicating the depth of the color block determined by projection parameters and constant depth values ​​is encoded in metadata with constant depth values. This method encodes at least one attribute atlas, a geometry atlas, and metadata for each color block in a data stream.

[0126] Figure 11 A method 110 for decoding a data stream representing volumetric content, according to a non-limiting embodiment of the invention, is shown. At step 111, a data stream representing a volumetric scene is obtained from a source. Attribute atlases (e.g., color, normals, etc.) and geometry atlases (e.g., depth 3D coordinates, etc.) and associated metadata are decoded from the data stream. At step 112, for a given attribute tile packed in the attribute atlas, information indicating whether the attribute tile is a cardboard tile is extracted from the metadata. If not, step 113 is performed, and the pixels of the attribute tile are back-projected based on the geometry tiles packed in the geometry atlas and associated with the current attribute tile in the metadata. Back-projection uses back-projection parameters associated with the attribute tile in the metadata. If the attribute tile is a cardboard tile, step 114 is performed. At step 114, a planar layer is reconstructed using projection parameters and depth values ​​stored in the metadata. The pixels of the attribute tile are back-projected onto this planar layer.

[0127] The specific implementations described herein may be implemented, for example, in methods or processes, apparatus, computer program products, data streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method or apparatus), the specific implementations of the discussed features may also be implemented in other forms (e.g., programs). Apparatus may be implemented, for example, in suitable hardware, software, and firmware. Methods may be implemented in apparatus (such as, for example, a processor) that generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, smartphones, tablets, computers, mobile phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end users.

[0128] Specific implementations of the various processes and features described herein can be found in a wide variety of devices or applications, particularly those associated with data encoding, data decoding, view generation, texture processing, and other processing of images and related texture and / or depth information. Examples of such devices include encoders, decoders, post-processors that process the output from decoders, pre-processors that provide input to encoders, video encoders, video decoders, video codecs, web servers, set-top boxes, laptops, personal computers, cellular phones, PDAs, and other communication devices. It should be understood that the devices can be mobile, even mounted in mobile vehicles.

[0129] Additionally, the method can be implemented by instructions executed by a processor, and such instructions (and / or data values ​​generated by the implementation) can be stored on a processor-readable medium, such as, for example, an integrated circuit, a software carrier, or other storage device, such as, for example, a hard disk, a compact disk (“CD”), an optical disk (such as, for example, a DVD, commonly referred to as a digital versatile optical disk or digital video optical disk), random access memory (“RAM”), or read-only memory (“ROM”). Instructions can form an application program tangibly embodied on the processor-readable medium. Instructions can be, for example, hardware, firmware, software, or a combination thereof. Instructions can be found, for example, in an operating system, a standalone application, or a combination of both. Thus, a processor can be characterized, for example, as a device configured to execute a process and a device comprising a processor-readable medium (such as a storage device) having instructions for executing the process. Furthermore, in addition to or instead of instructions, the processor-readable medium can store data values ​​generated by the implementation.

[0130] It will be apparent to those skilled in the art that the embodiments may produce various signals formatted to carry, for example, storable or transmissible information. The information may include, for example, instructions for performing a method or data generated by one of the embodiments. For example, the signal may be formatted as data carrying rules for writing or reading the syntax of the described embodiment, or as data carrying actual syntax values ​​written by the described embodiment. Such signals may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding a data stream and using a modulated carrier for the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal can be transmitted via a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.

[0131] Several specific embodiments have been described. However, it should be understood that many modifications can be made. For example, elements of different embodiments can be combined, supplemented, modified, or removed to produce other embodiments. Furthermore, those skilled in the art will understand that other structures and processes can be replaced with those disclosed, and the resulting embodiments will perform at least substantially the same function in at least substantially the same manner to achieve at least substantially the same results as the disclosed embodiments. Therefore, this application considers these and other embodiments.

Claims

1. A method, the method comprising: The attribute atlas image, geometry atlas image, and metadata are obtained from the data stream. The atlas image is packaged into tiled frames, which are projections of 3D scene samples. The metadata for the attribute tiled frames of the attribute atlas image includes: The projection parameters associated with the attribute-divided image, and Information indicating whether the attribute block view is associated with the geometric block view of the geometric atlas image or whether the attribute block view is associated with the depth value encoded in the metadata; Under the condition that the attribute block image is associated with the geometric block image, the pixels of the attribute block image are back-projected onto positions determined by the geometric block image and the projection parameters associated with the attribute block image; and Under the condition that the attribute block image is associated with the depth value, the pixels of the attribute block image are back-projected onto the position determined by the depth value and the projection parameters associated with the attribute block image.

2. The method according to claim 1, wherein, The pixels of the attribute atlas image encode two values ​​for two different attributes, which are then back-projected together.

3. The method according to claim 1, wherein, The data stream includes two attribute atlases encoded according to the same packaged layout, and the metadata is generated for a pair of attribute tile screens for each attribute atlas, the pair of attribute tile screens being back-projected together.

4. A device including a processor, said processor being configured to: The attribute atlas image, geometry atlas image, and metadata are obtained from the data stream. The atlas image is packaged into tiled frames, which are projections of 3D scene samples. The metadata for the attribute tiled frames of the attribute atlas image includes: The projection parameters associated with the attribute-divided image, and Information indicating whether the attribute block screen is associated with a geometric block of the geometry atlas image or whether the attribute block screen is associated with a depth value encoded in the metadata; Under the condition that the attribute block image is associated with the geometric block image, the pixels of the attribute block image are back-projected onto the position determined by the geometric block image and the projection parameters associated with the attribute block image. as well as Under the condition that the attribute block image is associated with the depth value, the pixels of the attribute block image are back-projected onto the position determined by the depth value and the projection parameters associated with the attribute block image.

5. The device according to claim 4, wherein, The pixels of the attribute atlas image encode two values ​​for two different attributes, which are then back-projected together.

6. The device according to claim 4, wherein, The data stream includes two attribute atlases encoded according to the same packaged layout, and the metadata is generated for a pair of attribute tile screens for each attribute atlas, the pair of attribute tile screens being back-projected together.

7. A method, the method comprising: Obtain a set of attribute blocks associated with the geometric block blocks. The attributes and geometric block blocks are obtained by projecting 3D scene samples according to projection parameters. For the attribute block screens in the set of attribute block screens, The attribute-segmented images are packaged into an attribute atlas image; and If the geometric tile associated with the attribute tile corresponds to a planar layer at a location determined by the depth value and the projection parameter, then metadata is generated including: the projection parameter, the depth value, and information indicating that the attribute tile is associated with the depth value, or In another case, the geometric tile images are packaged in a geometric atlas image, and metadata including the following items are generated: the projection parameters, and information indicating that the attribute tile images are associated with the geometric tile images; as well as The attribute atlas image, the geometry atlas image, and the metadata are encoded in a data stream.

8. The method according to claim 7, wherein, The pixels of the attribute atlas image are encoded with two values ​​for two different attributes.

9. The method according to claim 7, wherein, The set of attribute tiled screens includes pairs of attribute tiled screens for two different attributes, the pairs of attribute tiled screens are packaged in two attribute atlases according to the same packaging layout, and the metadata is generated for a pair of attribute tiled screens.

10. An apparatus including a processor, the processor being configured to: Obtain a set of attribute blocks associated with the geometric block blocks. The attributes and geometric block blocks are obtained by projecting 3D scene samples according to projection parameters. For the attribute block screens in the set of attribute block screens, The attribute-segmented images are packaged into an attribute atlas image; and If the geometric tile associated with the attribute tile corresponds to a planar layer at a location determined by the depth value and the projection parameter, then metadata is generated including: the projection parameter, the depth value, and information indicating that the attribute tile is associated with the depth value, or In another case, the geometric tile images are packaged in a geometric atlas image, and metadata including the following items are generated: the projection parameters, and information indicating that the attribute tile images are associated with the geometric tile images; as well as The attribute atlas image, the geometry atlas image, and the metadata are encoded in a data stream.

11. The device according to claim 10, wherein, The pixels of the attribute atlas image are encoded with two values ​​for two different attributes.

12. The device according to claim 10, wherein, The set of attribute tiled screens includes pairs of attribute tiled screens for two different attributes, the pairs of attribute tiled screens are packaged in two attribute atlases according to the same packaging layout, and the metadata is generated for a pair of attribute tiled screens.

13. A non-transitory computer-readable medium comprising instructions stored thereon, which, when executed by a processor, perform the method of any one of claims 1-3 and 7-9.