Method and apparatus for depth encoding and decoding

By encoding and decoding the distance in the point cloud using quantization functions, the difficulty of encoding depth information in large field of view content is solved, the immersiveness and visual stability of immersive videos are improved, and it is suitable for 3DoF and 6DoF video rendering.

CN113785591BActive Publication Date: 2025-09-23INTERDIGITAL VC HOLDINGS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080033462.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-20
Filing Date
2020-03-17
Publication Date
2025-09-23
Estimated Expiration
2040-03-17

AI Technical Summary

Technical Problem

Existing technologies have difficulty effectively encoding and decoding depth information in wide field of view content, especially when the depth value range is large and the bit depth is limited, resulting in dizziness and parallax problems in immersive video experience.

Method used

A quantization function is used to encode and decode the distance between a point in the point cloud and the first point. The distance value is quantized and metadata is encoded in the data stream by using the quantization function defined by a given angle and error value. The true distance is restored using the inverse quantization function during decoding.

Benefits of technology

It improves the immersive feeling and scene depth perception of immersive videos, reduces dizziness, ensures visual errors within a predetermined angle range, and is suitable for 3DoF and 6DoF video rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113785591B_ABST
    Figure CN113785591B_ABST
Patent Text Reader

Abstract

This document discloses methods, devices, and data stream formats for encoding, formatting, and decoding depth information representing a 3D scene. The compression and decompression of quantized values ​​by video codecs results in value errors. These value errors are particularly sensitive to depth coding. The present invention proposes encoding and decoding depth using a quantization function that minimizes angular errors when the value error of the quantized depth produces a positional difference between the projected and deprojected points. The inverse of this quantization function must be encoded in metadata associated with the 3D scene, for example as a lookup table (LUT), so that it can be retrieved during decoding, as such functions are difficult to control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present principles generally relate to the field of three-dimensional (3D) scenes and stereoscopic video content. This document can also be understood in the context of encoding, formatting, and decoding data representing the geometry of a 3D scene, for example, for rendering stereoscopic content on an end-user device such as a mobile device or a head-mounted display (HMD). Background Art

[0002] This section is intended to introduce the reader to various aspects of the art that may be related to various aspects of the present principles described and / or claimed below. It is believed that this discussion will help provide the reader with background information to facilitate a better understanding of the various aspects of the present principles. Therefore, it should be understood that these statements should be read in this light, and not as admissions of prior art.

[0003] Recently, the availability of content with a wide field of view (up to 360°) has increased. When a user views this content on an immersive display device, such as a head-mounted display, smart glasses, a PC screen, a tablet, or a smartphone, they may not be able to see the entire content. This means that at any given moment, the user may only be able to view a portion of the content. However, the user can typically navigate within the content through various means, such as head movement, mouse movement, touchscreen, voice, and so on. It is often desirable to encode and decode this content.

[0004] Immersive video, also known as 360° planar video, allows users to see everything around them by rotating their head around a stationary viewpoint. Rotation only allows for a 3-degree-of-freedom (3DoF) experience. Even though 3DoF video is sufficient for a first omnidirectional video experience, such as using a head-mounted display (HMD), it can quickly become frustrating for viewers who expect more freedom, for example due to experiencing parallax. 3DoF can also induce dizziness because users never just rotate their heads, but also translate them in three directions, a movement that cannot be reproduced in a 3DoF video experience.

[0005] Wide field of view content can be, among other things, three-dimensional computer graphics scenes (3D CGI scenes), point clouds, or immersive videos, etc. Many terms can be used to designate such immersive videos: for example, virtual reality (VR), 360, panoramic, 4π stereo, immersive, omnidirectional, or wide field of view.

[0006] Volumetric video, also known as 6 degrees of freedom (6DoF) video, is an alternative to 3DoF video. When watching 6DoF video, in addition to rotation, the user can also pan their head or even their body within the content being viewed and experience parallax or even volume. This type of video significantly increases immersion and the perception of scene depth by providing consistent visual feedback during head translation and prevents dizziness. The content is created with the help of dedicated sensors, allowing the simultaneous recording of color and depth of the scene of interest. Using a rig with a color camera combined with photogrammetry technology is one way to perform this recording, although there are still technical difficulties.

[0007] While 3DoF videos consist of a sequence of images resulting from the unmapping of texture images (e.g., spherical images encoded according to latitude / longitude projection mapping or equirectangular projection mapping), 6DoF video frames embed information from several viewpoints. They can be viewed as a time series of point clouds resulting from three-dimensional capture. Depending on the viewing conditions, two types of volumetric videos can be considered. The first (i.e., full 6DoF) allows for completely free navigation in the video content, while the second (also known as 3DoF+) restricts the user viewing space to a limited volume called the viewing bounding box, allowing limited head translation and parallax experiences. This second context is a valuable trade-off between free navigation and the passive viewing conditions of seated viewers.

[0008] Apart from the specific case of volumetric video, encoding and decoding of depth information of 3D scenes or volumetric content can be problematic, especially when the range of depth values ​​to be encoded is large and the bit depth available for encoding does not provide a sufficient number of encoded values. Summary of the Invention

[0009] The following is a simplified summary of the present principles to provide a basic understanding of some aspects of the present principles. This summary is not an extensive overview of the present principles. It is not intended to identify key or important elements of the present principles. The following summary merely presents some aspects of the present principles in a simplified form as a prelude to the more detailed description provided below.

[0010] The present principles relate to a method for encoding data representing the distance between a point of a point cloud and a first point located within the point cloud. The method comprises quantizing a value representing the distance between a first and a second point using a quantization function defined by a third point, a given angle, and an error value. The quantization function may be defined such that dequantization of the sum of the quantized value and the error value produces a fourth point; the angle between the fourth, third, and second points being less than or equal to the given angle. When a value is quantized, the method encodes the quantized value in a data stream associated with metadata representing the quantization function. In an embodiment, the quantization function is a parameterized function known to both the encoder and the decoder, whereby the metadata comprises the given angle and / or the error value and / or the coordinates of the third point, or in a variant, the distance between the first and third points. In another embodiment, the metadata comprises a lookup table responsive to the inverse of the quantization function.

[0011] The present principles also relate to a data stream generated by an apparatus implementing the method, the data stream including data representing quantized values ​​representing the distance between a point of a point cloud and a first point located within the point cloud, the distance having been quantized using a quantization function defined by a third point, a given angle, and an error value.

[0012] The present principles also relate to a method for decoding data representing the distance between a point in a point cloud and a first point within the point cloud. The method includes decoding quantized values ​​and associated metadata from a data stream. The metadata includes data representing a quantization function defined by a third point, a given angle, and an error value. The method further dequantizes the extracted quantized values ​​using the inverse of this quantization function. In an embodiment, the inverse quantization function is a parameterized function known to both the encoder and the decoder, and the metadata thus includes an error value for the given angle and / or coordinates and / or the third point, or in a variant, the distance between the first point and the third point. In this embodiment, on the decoding side, the inverse quantization function is initialized with these parameters. In a variant, some of these parameters are optional if a default value is determined for one of them. In another embodiment, the metadata includes a lookup table responsive to the inverse of the quantization function. The actual distance value is obtained by looking up the quantized value in the table.

[0013] The present principles also relate to a device comprising a processor configured to implement such a method. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The present disclosure will be better understood, and other specific features and advantages will emerge, from a reading of the following description, which refers to the accompanying drawings, in which:

[0015] - Figure 1 A three-dimensional (3D) model of an object and points of a point cloud corresponding to the 3D model are shown according to a non-limiting embodiment of the present principles;

[0016] - Figure 2 shows a non-limiting example of encoding, transmission and decoding of data representing a 3D scene sequence according to a non-limiting embodiment of the present principles;

[0017] - Figure 3 An example architecture of a device according to a non-limiting embodiment of the present principles is shown, which can be configured to implement Figure 8 and 9 Methods of description;

[0018] - Figure 4 illustrates an example of an embodiment of the syntax of a stream when data is transmitted over a packet-based transport protocol according to a non-limiting embodiment of the present invention;

[0019] - Figure 5 illustrates a spherical projection from a central viewpoint according to a non-limiting embodiment of the present principles;

[0020] - Figure 6 An example of a projection map including depth information of points of a 3D scene visible from a projection center (also referred to as a first point) according to a non-limiting embodiment of the present principles is shown;

[0021] - Figure 7 illustrates how quantization error is perceived from a second viewpoint in a 3D scene according to a non-limiting embodiment of the present principles;

[0022] - Figure 8 A method of encoding data representing the depth of a point of a 3D scene according to a non-limiting embodiment of the present principles is illustrated;

[0023] - Figure 9 A method of decoding data representing a distance between a point of a 3D scene and a first point within the 3D scene according to a non-limiting embodiment of the present principles is illustrated; DETAILED DESCRIPTION

[0024] The present principles will be described more fully hereinafter with reference to the accompanying drawings, in which examples of the present principles are shown. However, the present principles may be implemented in many alternative forms and should not be construed as limited to the examples set forth herein. Thus, while the present principles are susceptible to various modifications and alternative forms, specific examples thereof are illustrated by way of example in the accompanying drawings and will be described in detail herein. However, it should be understood that there is no intention to limit the present principles to the particular form disclosed, but rather, the present invention is intended to cover all modifications, equivalents, and alternative forms falling within the spirit and scope of the present principles as defined in the claims.

[0025] The terms used herein are only for the purpose of describing specific examples and are not intended to limit this principle. As used herein, the singular forms "one", "an" and "the" are also intended to include plural forms, unless the context clearly indicates otherwise. It will be further understood that, when used in this specification, the terms "include" and / or "comprise" specify the presence of the features, integers, steps, operations, elements and / or components. But do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. Moreover, when an element is referred to as "in response to" or "connected to" another element, it can directly respond to or be connected to another element, or there can be an intermediate element. On the contrary, when an element is referred to as "directly responding to" or "directly connected to" other elements, there is no intermediate element. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and can be abbreviated as " / ".

[0026] It will be understood that although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the teachings of the present principles.

[0027] While some diagrams include arrows on communication paths to illustrate a primary direction of communication, it should be understood that communication can occur in the opposite direction of the depicted arrows.

[0028] Some examples are described with respect to block diagrams and operational flow charts, wherein each block represents a portion of a circuit element, module, or code that includes one or more executable instructions for implementing (one or more) specified logical functions. It should also be noted that, in other embodiments, the (one or more) functions indicated in the blocks may not occur in the order indicated. For example, depending on the functions involved, two blocks shown in succession may actually be executed substantially concurrently, or may sometimes be executed in the reverse order.

[0029] References herein to "according to an example" or "in an example" mean that a particular feature, structure, or characteristic described in connection with the example can be included in at least one embodiment of the present principles. The appearances of the phrase "according to an example" or "an example" in various places in the specification are not necessarily all referring to the same example, nor are separate or alternative examples necessarily mutually exclusive of other examples.

[0030] Reference numerals appearing in the claims are for illustrative purposes only and have no limiting effect on the scope of the claims.The present examples and variations may be employed in any combination or subcombination although not explicitly described.

[0031] According to a non-limiting embodiment of the present disclosure, a method and apparatus for encoding and decoding data representing the depth of a point of a 3D scene are proposed herein. According to the present principles, the depth of a point in a 3D scene is the distance (e.g., the Euclidean distance in a Cartesian reference frame) between this point and a given point (referred to herein as the first point or projection center). A portion of the 3D scene to be encoded is projected onto an image plane relative to the projection center. The center projection operation is used in conjunction with a mapping operation, such as a sphere mapping projection (e.g., an equirectangular projection (ERP) or a Cassini projection or a sinusoidal projection) or a cube mapping projection (according to different layouts of the facets of a cube) or a pyramid mapping projection. In the context of the present disclosure, the depth data is quantized and stored as image data, i.e., as an array of 2D matrices, pixels. Such images are compressed and transmitted to a decoder, which decompresses them.

[0032] Point P0 is projected onto the image plane and its depth is stored as a real number (i.e., represented by a floating-point value). Quantization is the process of restricting the input from a continuous or otherwise large set of values ​​(such as real numbers) to a discrete set (such as an interval of integers, usually between 0 and n). Therefore, when the number of values ​​to be quantized is greater than n, quantization introduces a loss of precision. During encoding, the distance d0 is quantized to the value v0 according to the quantization function, and during decoding, the value v0 is dequantized to the distance d1 according to the inverse quantization function. Point P1 is then deprojected at a distance D1 from the center of projection.

[0033] The compression-decompression operation introduces some errors in the quantized values. For example, after compression-decompression, the quantized value v0 may store (vault) a value v2 between v0-δ and v0+δ, where δ=1 or 2 or 10. According to the inverse quantization function, this value v2 is dequantized in the distance d2. Thus, instead of point P1, point P2 is deprojected on the same line passing through the projection center at a distance d2 from the projection center. The error (dl-d2) depends on the quantization function. In addition, the error in the position between P1 and P2 is perceived differently depending on the position of the viewpoint A from which they are observed. In fact, the angle Determine the difference in position between P1 and P2 as viewed from A. For example, it is known that at an angle γ = 0.5' (half arc minute), known as human visual acuity, the human eye cannot distinguish between two different points in 3D space.

[0034] According to a non-limiting embodiment of the present principles, data representing the distance between a point of a 3D scene and a first point within the 3D scene are encoded in a data stream. These depth data are quantized using a quantization function defined based on a second point, a given angle and an error value, so that the difference in the error value between the two quantized values ​​results in an error angle at the second point being lower than the given angle. The depth data represents the distance between point P1 and the center of projection. The difference in error value between the quantized value of the depth data for point P before compression and the quantized value of the same data after decompression (for example, a difference of 1, 3 or 7) results in deprojection of this data at different distances and sets point P2 farther away from or closer to the center of projection. The error angle at a given point A is the angle formed at this given point due to the error value of the quantized value. According to the present principles, a quantization function is determined to ensure that for a given second point in the 3D scene, the error angle will not exceed a given angle (selected according to the desired accuracy) when the error in the quantized value is a given level called the error value.

[0035] Figure 1 A three-dimensional (3D) model 10 of an object and the points of a point cloud 11 corresponding to the 3D model 10 are shown. The 3D model 10 and the point cloud 11 may correspond, for example, to a possible 3D representation of an object of a 3D scene including other objects. The model 10 may be a 3D mesh representation, and the points of the point cloud 11 may be vertices of the mesh. The points of the point cloud 11 may also be points scattered on the surface of the faces of the mesh. The model 10 may also be represented as a tiled version of the point cloud 11, with the surface of the model 10 being created by tiling the points of the point cloud 11. The model 10 may be represented using many different representations, such as voxels or splines. Figure 1 This illustrates the fact that a point cloud can be defined using a surface representation of a 3D object, and that a surface representation of the 3D object can be generated from the points of the cloud. As used herein, projecting the points of a 3D object (by extension of the 3D scene) onto an image is equivalent to projecting any representation of the 3D object, such as a point cloud, mesh, spline model, or voxel model.

[0036] A point cloud can be represented in memory, for example, as a vector-based structure where each point has its own coordinates in the viewpoint's reference frame (e.g., three-dimensional coordinates XYZ), or a solid angle and distance from the viewpoint (also called depth) and one or more attributes, also called components. Examples of components are color components that can be expressed in various color spaces, such as RGB (red, green, and blue) or YUV (Y is the luminance component and UV are the two chrominance components). A point cloud is a representation of a 3D scene including objects. A 3D scene can be seen from a given viewpoint or range of viewpoints. A point cloud can be obtained in many ways, for example:

[0037] From the capture of real objects filmed by a set of cameras, optionally supplemented by depth active sensing devices;

[0038] From the capture of virtual / synthetic objects shot by a set of virtual cameras in the modeling tool;

[0039] From a mix of real and virtual objects.

[0040] Figure 2 A non-limiting example of encoding, transmission and decoding of data representing a 3D scene sequence is shown. For example and at the same time the encoding format may be compatible for 3DoF, 3DoF+ and 6DoF decoding.

[0041] A 3D scene sequence is obtained 20. Since a picture sequence is a 2D video, a 3D scene sequence is a 3D (also called volumetric) video. The 3D scene sequence can be provided to a volumetric video rendering device for 3DoF, 3DoF+ or 6DoF rendering and display.

[0042] A 3D scene sequence 20 is provided to an encoder 21. The encoder 21 takes a 3D scene or a sequence of 3D scenes as input and provides a bitstream representing the input. The bitstream can be stored in a memory 22 and / or on an electronic data medium and can be transmitted over a network 22. The bitstream representing the 3D scene sequence can be read from the memory 22 and / or received from the network 22 by a decoder 23. The decoder 23 receives the bitstream as input and provides the 3D scene sequence, for example, in a point cloud format.

[0043] The encoder 21 may include several circuits that implement several steps. In a first step, the encoder 21 projects each 3D scene onto at least one 2D picture. 3D projection is any method of mapping three-dimensional points onto a two-dimensional plane. Since most current methods for displaying graphic data are based on flat (pixel information from several bit planes) two-dimensional media, this type of projection is widely used, especially in computer graphics, engineering design, and drawing. The projection circuit 211 provides at least one two-dimensional frame 2111 for the 3D scene of the sequence 20. The frame 2111 includes depth information representing the 3D scene projected onto the frame 2111. In a variant, color information representing the color information of the points of the 3D scene is also projected and stored in the pixels of the frame 2111. In another variant, the color and depth information are encoded in two separate frames 2111 and 2112. For example, Figure 1 The points of the 3D scene 10 only include depth information. No texture is attached to the model, and the points of the 3D scene have no color components. In any case, depth information is required to be encoded in the representation of the 3D scene.

[0044] The metadata 212 is used and updated by the projection circuit 211. The metadata 212 includes information about the projection operation (e.g., projection parameters) and information about how the color and depth information is organized within the frames 2111 and 2112, such as information about the Figures 5 to 7 According to the present principles, metadata includes information representative of an inverse quantization function used to encode depth information.

[0045] The video encoding circuit 213 encodes the sequence of frames 2111 and 2112 into a video. The pictures 2111 and 2112 of the 3D scene (or a sequence of pictures of the 3D scene) are encoded in a stream by the video encoder 213. The video data and metadata 212 are then encapsulated in a data stream by the data encapsulation circuit 214.

[0046] The encoder 213 conforms to an encoder such as the following, for example:

[0047] -JPEG, specification ISO / CEI 10918-1UIT-T Recommendation T.81, https: / / www.itu.int / rec / T-REC-T.81 / en;

[0048] -AVC, also known as MPEG-4 AVC or h264. Specified in UIT-T H.264 and ISO / CEI MPEG-4 Part 10 (ISO / CEI 14496-10), http: / / www.itu.int / rec / T-REC-H.264 / en, HEVC (whose specifications can be found on the ITU website, T Recommendation, H Series, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-Een);

[0049] - 3D-HEVC (an extension of HEVC, whose specifications are found in the ITU website, T Recommendation, H series, h.265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en, Appendices G and I);

[0050] - VP9 developed by Google; or

[0051] -AV1 (AOMedia Video 1), developed by the Alliance for Open Media.

[0052] The data stream is stored in a memory accessible by a decoder 23, for example, via a network 22. The decoder 23 includes different circuits that implement different decoding steps. The decoder 23 takes as input the data stream generated by the encoder 21 and provides a 3D scene sequence 24 to be rendered and displayed by a volumetric video display device, such as a head-mounted device (HMD). The decoder 23 obtains the stream from a source 22. For example, the source 22 belongs to the set including:

[0053] - local memory, such as video memory or RAM (or random access memory), flash memory, ROM (or read-only memory), hard disk;

[0054] - Storage interfaces, such as interfaces to mass storage devices, RAM, flash memory, ROM, optical disks or magnetic media;

[0055] - a communication interface, such as a wired interface (e.g. a bus interface, a wide area network interface, a local area network interface) or a wireless interface (such as an IEEE 802.11 interface or interface); and

[0056] - A user interface, such as a graphical user interface enabling a user to input data.

[0057] Decoder 23 includes circuitry 234 for extracting data encoded in a data stream. Circuitry 234 receives the data stream as input and provides metadata 232 corresponding to metadata 212 and the two-dimensional video encoded in the stream. According to present principles, metadata 232 includes information representing an inverse quantization function for retrieving the depth of a point in a 3D scene. In the context of the present disclosure, the depth of a point corresponds to the distance between the point to be projected and the center of projection. The coordinates of the center of projection are included in metadata 232 or defined by default, for example, at the origin of a reference system for the 3D space of the 3D scene. The video is decoded by a video decoder 233 that provides a sequence of frames. The decoded frames include depth information. Due to the compression-decompression process, the quantized depth value after decoding may differ from the quantized depth value during encoding. In a variant, the decoded frames include depth information and color information. In another variant, video decoder 233 provides two sequences of frames, one including color information and the other including depth information. Circuitry 231 uses metadata 232 to retrieve an inverse quantization function and deproject depth information and ultimately color information from the decoded frame to provide a 3D scene sequence 24. The 3D scene sequence 24 corresponds to the 3D scene sequence 20, possibly with a loss of precision associated with encoding as 2D video and video compression. The 3D scene of the decoded sequence 24 is rendered from the current viewpoint by projecting the 3D scene onto the image plane of the viewport.

[0058] Figure 3 An example architecture of a device 30 is shown, which may be configured to implement Figure 8 and9 Described method. Figure 2 The encoder 21 and / or decoder 23 can implement this architecture. Alternatively, each circuit of the encoder 21 and / or decoder 23 can be based on Figure 3 The devices of the architecture are linked together, for example via their bus 31 and / or via the I / O interface 36.

[0059] The device 30 comprises the following elements, which are linked together by a data and address bus 31:

[0060] - a microprocessor 32 (or CPU), for example a DSP (or digital signal processor);

[0061] -ROM (or read-only memory) 33;

[0062] - RAM (or random access memory) 34;

[0063] - Storage interface 35;

[0064] - an I / O interface 36 for receiving data to be transmitted from an application; and

[0065] - A power source, such as a battery.

[0066] According to an example, the power supply is external to the device. In each memory reference, the word "register" used in this specification can correspond to a small area (a few bits) or a very large area (e.g., an entire program or a large amount of received or decoded data). ROM 33 includes at least programs and parameters. ROM 33 can store algorithms and instructions for implementing the technology according to the present principles. After power is applied, CPU 32 uploads the program to RAM and executes the corresponding instructions.

[0067] RAM 34 includes in registers the program executed by CPU 32 and uploaded when device 30 is powered on, input data in registers, intermediate data in registers at different states of the method, and other variables in registers used to execute the method.

[0068] The embodiments described herein can be implemented, for example, with methods or processes, devices, computer program products, data streams or signals. Even if discussed only in the context of a single form of embodiment (for example, discussed only as a method or device), the embodiments of the features discussed can also be implemented in other forms (for example, programs). Devices can be implemented, for example, with suitable hardware, software and firmware. The method can be implemented in, for example, a device, such as a processor, which refers to a processing device, generally including, for example, a computer, a microprocessor, an integrated circuit or a programmable logic device. The processor also includes a communication device, such as, for example, a computer, a cellular phone, a portable / personal digital assistant ("PDA") and other devices that promote information communication between end users.

[0069] According to an example, the device 30 is configured to implement Figure 8 and 9 describes a method, and belongs to a set that includes:

[0070] -mobile device;

[0071] -communication equipment;

[0072] -Gaming equipment;

[0073] - Tablet computer (or tablet computer);

[0074] - Laptop computer;

[0075] - Still picture camera;

[0076] - Camera;

[0077] -Encoding chip;

[0078] - Server (eg, a broadcast server, a video-on-demand server, or a web server).

[0079] Figure 4 An example of an embodiment of the syntax of a stream when data is transmitted over a packet-based transport protocol is shown. Figure 4 An example structure 4 of a volumetric video stream is shown. The structure is contained in a container that organizes the stream into individual syntax elements. The structure may include a header section 41, which is a collection of data common to each syntax element of the stream. For example, the header section includes some of the metadata about the syntax elements, describing the nature and purpose of each element. The header section may also include Figure 2 The structure includes a portion of metadata 212, such as the coordinates of a central viewpoint used to project points of a 3D scene onto frames 2111 and 2112. The structure includes a payload that includes syntax elements 42 and at least one syntax element 43 of syntax. Syntax elements 42 include data representing color and depth frames. The images may have been compressed according to a video compression method.

[0080] Syntax element 43 is part of the payload of the data stream and may include metadata about how the frames of syntax element 42 are encoded, such as parameters for projecting and packing points of a 3D scene onto frames. Such metadata may be associated with each frame of the video or a group of frames (also known as a group of pictures (GoP) in video compression standards).

[0081] Figure 5 The spherical projection from the central viewpoint 50 is shown. Figure 5 In the example of , the 3D scene includes three objects 52, 53 and 54. According to the viewpoint 50, the points of the object 52 form a surface with a front side and a back side. The back side points of the object 42 are not visible from the viewpoint 50. According to the viewpoint 50, the points of the objects 53 and 54 form a surface with a front side. The points of the object 53 are visible from the viewpoint 50, but due to occlusion by the surface of the object 53, only a part of the points of the object 54 is visible from this viewpoint. Therefore, a spherical projection (e.g., an equirectangular projection ERP) does not project every point of the 3D scene onto the frame. Many other types of projections can be used, such as perspective projection or orthographic projection. For example, the points of the point cloud visible from the viewpoint 50 are projected on the projection map 51 according to the projection method. About Figure 5 The projection method is an equirectangular projection, such as the latitude / longitude projection or the equirectangular projection (also known as ERP), so the projection mapping is in Figure 5 51 is represented above. In a variant, the projection method is a cube projection method, a pyramid projection method, or any projection method centered on viewpoint 50. Points on the front side of object 52 are projected into area 55 of the projection map. Points on the back side of object 52 are not projected because they are not visible from viewpoint 50. Every point of object 53 is visible from viewpoint 50. They are projected onto area 56 of projection map 51 according to the projection method. In an embodiment, only a portion of the points of object 54 are visible from viewpoint 50. The visible points of object 54 are projected onto area 57 of projection map 51. The information stored in the pixels of projection map 51 corresponds to the distance between the projected point in the 3D scene and the projection center 50. In a variant, the color components of the projected points are also stored in the pixels of projection map 51.

[0082] Figure 6 An example of a projection map 60 including depth information of points of a 3D scene visible from a projection center (also referred to as a first point) is shown in accordance with a non-limiting embodiment of the present principles. Figure 6 In the example of FIG. 6 , the farther away from the center of projection of a point of the 3D scene, the brighter the pixel in the image 60. The distance to be stored in the pixel of the image 60 is expressed in grayscale (ie, in grayscales between 0 and N=2). n-1; n is the encoding bit depth (i.e., the number of bits used to encode the integer value, typically 8, 10, or 12 for the HEVC codec). Figure 6 In the example, the depth is encoded as 10 bits, so pixels store values ​​between 0 and 1023. For example, in Figure 6 In the example, the depth information is from z min = 0.5 m to z max = 28 meters. There are more than 1024 different distances, and a quantization function must be used to convert the real value into a discrete value. The affine transformation Eq1 or the inverse function Eq2 is a possible quantization function for quantizing the real value z of the depth:

[0083]

[0084]

[0085] However, this quantization is not perceptually consistent but rather scene-driven, since it essentially depends on z min and z max .

[0086] Figure 7 Figure 1 illustrates how quantization error is perceived from a second viewpoint in a 3D scene. Figure 7 In the example of , point 71 is projected onto the image plane of the projection map. Sphere 72 represents the viewport, also called the viewing bounding box, from which the user can see the 3D scene in the context of a 3DoF+ scene. Sphere 72 can also represent a projection map. The distance 74 between point 71 and a first point 73 (i.e., the projection center), which we call z, is quantized by a quantization function f, and the quantized value v is stored in a pixel of the projection map. The projection map is compressed by the image codec and decompressed on the decoding side, as Figure 2 As shown. The value v may have been shifted to a value v' by this compression-decompression process, for example by plus or minus 1, 5, or 8. The distance between the first point 73 and the deprojected point 76 is determined by applying an inverse quantization function to the value v'. The difference 75 between the distance 74 and the distance between points 76 and 73 is due to the compression error. The difference 75 depends on the inverse quantization function and therefore also on the quantization function. Observed from the first point 73, the difference 75 is perceptible only in the relationship between point 76 and its neighbors surrounding the deprojected point. However, observed from another point in the field of view (for example, point 77), the difference 75 is perceived by itself because this difference between the positions of points 71 and 76 forms an angle 78 pointing to point 77. The larger the angle 78, the greater the visual artifact generated by the compression error. Angle 78 is called the error angle at point 77.

[0087] According to the present principles, the quantization function is chosen and parameterized so as to keep the angle 78 below a predetermined angle γ of a second point in 3D space. For example, a second point 77 or 79 is chosen in the viewport defined for a given 3DoF+ rendering of the 3D scene. For example, the angle γ is set to a value corresponding to a known human visual acuity of approximately half an arc minute. Depending on the expected robustness of the quantization on the expected level of compression error, the angle γ may be set to any angle value. In this context, depending on the position of the user in the viewport, the associated error angle 78, referred to herein as φ(z), may be zero when the user is standing at the center of the projection, and may be significantly larger when he is standing at point A. In Figure 7 In the example of , the viewport is a sphere, and it can be shown that the maximum of φ(z) is obtained at points at the front boundary of the sphere. In the case of area-constrained viewports (i.e., 3DoF+ rendering), this property is important to design specialized depth quantization laws (i.e., functions) that prevent the user from experiencing visual artifacts within the viewport. If we call φ max (z) is the maximum value of φ(z) over the viewing area, then the quantization law should ensure that the associated quantization error forces φ max (z) = γ, where γ is a predetermined angle for the expected robustness of the quantization function to compression error levels (e.g., human visual acuity as described above). This ensures that quantization errors due to the compression-decompression process are not perceived from the visual zone, since any associated error angle remains below the perceptibility threshold. The resulting quantization law depends only on the predetermined angle γ and the coordinates of the second point (e.g., located in the visual zone). In a variant, instead of the coordinates of the second point, the quantization function may depend on the distance from the first point. This is equivalent to calculating the distance from the first point according to any point on a sphere centered on the first point (e.g., Figure 7 Point 79) defines the quantization function.

[0088] This quantization law is in the optimal position (where φ(z) = φ max (z)) is not easily controllable analytically, but the quantization discrete table can be obtained numerically. For example, for a point 79 at a distance R from the center of projection 73, a good approximation can be obtained by using Eq. 3 for a quantization function that forces the error angle at point 79 to be below a predetermined angle γ for a value error of 1 (i.e., the decompressed value of value v is either v+1 or v-1).

[0089]

[0090] Where K is a constant that is set so that for the maximum predetermined depth value, q p (z) = 0. The definition of the quantization function following these guidelines depends on the context. Finding a good candidate for the second point depends on the shape and size of the viewing area. The function chosen is not always like Figure 7In other words, the parameterizable function must be known to the encoder and decoder, and the selected parameters at the encoder must be encoded in the metadata associated with the 3D scene in the formatted stream so as to be retrieved on the decoding side. In another embodiment, a lookup table (LUT) responsive to the inverse quantization function is constructed at the encoder and encoded in the metadata associated with the image representing the 3D scene. The lookup table is information that associates each value of the quantized value interval with an actual value of the distance. The advantage of this embodiment is that the decoder does not need to know the inverse quantization function in advance. The decoder extracts the LUT from the stream and retrieves the actual depth from the quantized depth by using this LUT. Since the quantization function has been determined to force the error angle (as defined above) at the second point to be lower than the predetermined angle γ for the value error e, the inverse quantization function encoded as the LUT also enforces the same criterion.

[0091] In the context of encoding and decoding of 3D scenes or 3D scene sequences for 3DoF+ rendered scenes, an advantage of the present principles is that the perceived quantization error is guaranteed to be below a predetermined level (e.g. below human visual acuity) for any viewpoint within the viewing area of ​​the scene.

[0092] Figure 8 A method for encoding data representing the depth of a point in a 3D scene, according to a non-limiting embodiment of the present principles, is illustrated. At step 81, depth data is obtained from a source and quantized using a quantization function determined based on a second point, a given angle, and an error value, such that the difference in the error value between two quantized values ​​results in an error angle at the second point being less than the given angle. The depth data is quantized and stored in pixels of an image compressed using an image or video codec. At step 82, the compressed image is encoded in a data stream in association with metadata representing the inverse of the quantization function. In one embodiment, the inverse quantization function is a parameterized function known to both the encoder and the decoder. In this embodiment, the metadata includes the angle of the second point and / or the error value and / or the coordinates or the distance between the first and second points. This metadata is optional if it is predetermined and known a priori by the decoder. In another embodiment, a lookup table associating each possible quantized value with a distance determined by the quantization function is generated and encoded in the data stream in association with the compressed image.

[0093] Figure 9The diagram illustrates a method for decoding data representing the distance between a point in a 3D scene and a first point within the 3D scene. At step 91, a data stream containing encoded data representing the geometry of the 3D scene is obtained from the stream. A compressed image and metadata representing an inverse quantization function are extracted from the data stream. The image is decompressed. At step 92, the true value of depth information contained in pixels of the decompressed image is retrieved by applying the quantized value to the inverse quantization function. Points of the 3D scene are deprojected—that is, positioned at a dequantized distance from the first point in a direction determined by the coordinates of the pixel in the image and the projection operation used to generate the image. In one embodiment, the inverse quantization function is a parameterized function known to both the encoder and the decoder. The metadata includes parameters required to initialize the function: an angle and / or an error value and / or the coordinates of a second point or the distance between the first and second points. This metadata is optional if it is predetermined and known a priori by the decoder. In another embodiment, the inverse quantization function is encoded as a lookup table in the metadata. The actual distance value is retrieved directly from this lookup table based on the quantized value stored in the pixel.

[0094] The embodiments described herein can be implemented in, for example, a method or process, an apparatus, a computer program product, a data stream, or a signal. Even if discussed only in the context of a single implementation (e.g., discussed only as a method or device), the embodiments of the features discussed can also be implemented in other forms (e.g., programs). The apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The method can be implemented in, for example, an apparatus such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, such as, for example, a smart phone, a tablet computer, a computer, a mobile phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between end users.

[0095] The embodiments of the various processes and features described herein can be implemented in a variety of different equipment or applications, particularly equipment or applications associated with data encoding, data decoding, view generation, texture processing, and other processing of images and related texture information and / or depth information. Examples of such equipment include encoders, decoders, post-processors for processing outputs from decoders, pre-processors for providing inputs to encoders, video encoders, video decoders, video codecs, web servers, set-top boxes, laptop computers, personal computers, phones, PDAs, and other communication devices. As should be clear, the equipment can be mobile and can even be installed in a mobile vehicle.

[0096] Furthermore, the methods may be implemented by instructions executed by a processor, and such instructions (and / or data values ​​produced by the embodiments) may be stored on a processor-readable medium, for example, an integrated circuit, a software carrier, or other storage device (such as, for example, a hard disk, a compact disk ("CD"), an optical disk (such as, for example, a DVD, often referred to as a digital versatile disk or digital video disk), a random access memory ("RAM"), or a read-only memory ("ROM"). The instructions may form an application program tangibly embodied on the processor-readable medium. The instructions may be, for example, hardware, firmware, software, or a combination thereof. The instructions may be found, for example, in an operating system, a separate application, or a combination of both. Thus, a processor may be characterized as, for example, a device configured to perform a process and a device including a processor-readable medium (such as a storage device) having instructions for performing the process. Additionally, the processor-readable medium may store data values ​​produced by the embodiments in addition to or in lieu of the instructions.

[0097] It will be apparent to those skilled in the art that embodiments may generate various signals that are formatted to carry information that can, for example, be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described embodiments. For example, a signal may be formatted to carry as data the rules for writing or reading the grammar of the described embodiment, or to carry as data the actual grammar values ​​written by the described embodiment. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal may be transmitted over a variety of different wired or wireless links. The signal may be stored on a processor readable medium.

[0098] Many embodiments have been described. However, it will be understood that various modifications can be made. For example, elements of different embodiments can be combined, supplemented, modified, or removed to produce other embodiments. In addition, it will be understood by those skilled in the art that the disclosed structures and processes can be replaced with other structures and processes, and that the resulting embodiments will perform at least substantially the same (one or more) functions in at least substantially the same (one or more) manners to achieve at least substantially the same (one or more) results as the disclosed embodiments. Thus, the present application contemplates these and other embodiments.

Claims

1. A method comprising: - quantizing a value representing a distance between a first point of a 3D point cloud representing a 3D video content scene and a second point of the 3D point cloud using a quantization function defined by a third point (77, 79) of the 3D point cloud, a given angle, and an error value to obtain a quantized value; the quantization function being defined such that a fourth point is generated by a sum obtained by dequantizing the quantized value and the error value, and such that an angle formed by the fourth point, the third point, and the second point is less than or equal to the given angle; and - encoding said quantized value (42) in a data stream (4) in association with metadata (43) representative of said quantization function. 2 . The method of claim 1 , wherein the metadata includes the given angle and the coordinates of the third point or the distance between the first point and the third point.

3. The method of claim 1, wherein the metadata comprises a lookup table corresponding to an inverse of the quantization function.

4. The method of one of claims 1 to 3, wherein the first point and the third point belong to a given zone, the third point being selected to maximize the given angle.

5. The method of one of claims 1 to 3, wherein the given angle corresponds to a measurement of human visual acuity and / or wherein the error value is equal to 1.

6. A device (30) comprising a processor (32) configured to: - quantizing a value representing a distance between a first point of a 3D point cloud representing a 3D video content scene and a second point of the 3D point cloud using a quantization function defined by a third point (77, 79) of the 3D point cloud, a given angle, and an error value to obtain a quantized value; the quantization function being defined such that a fourth point is generated by a sum obtained by dequantizing the quantized value and the error value, and such that an angle formed by the fourth point, the third point, and the second point is less than or equal to the given angle; and - encoding said quantized value (42) in a data stream (4) in association with metadata (43) representative of said quantization function. 7 . The apparatus of claim 6 , wherein the metadata includes the given angle and coordinates of the third point or a distance between the first point and the third point.

8. The apparatus of claim 6, wherein the metadata comprises a lookup table corresponding to an inverse of the quantization function.

9. The device of one of claims 6 to 8, wherein the first point and the third point belong to a given zone, the third point being selected to maximize the given angle.

10. The device of one of claims 6 to 8, wherein the given angle corresponds to a measure of human visual acuity and / or wherein the error value is equal to 1.

11. A method comprising: - decoding, from a data stream (4), a quantized value (42) representing a quantization function and associated metadata (43), the quantization function being defined by a third point (77, 79) of a 3D point cloud representing a 3D video content scene, a given angle, and an error value; the quantization function being defined such that a fourth point is generated by a sum obtained by dequantizing the quantized value and the error value, and such that an angle formed by the fourth point, the third point, and a second point of the 3D point cloud is less than or equal to the given angle; as well as - dequantizing said quantized values ​​according to the inverse of said quantization function. 12 . The method of claim 11 , wherein the metadata comprises the given angle and coordinates of the third point or a distance between the first point and the third point of the 3D point cloud.

13. The method of claim 11, wherein the metadata comprises a lookup table corresponding to an inverse of the quantization function.

14. A device comprising a processor configured to: - decoding, from a data stream (4), a quantized value (42) representing a quantization function and associated metadata (43), the quantization function being defined by a third point (77, 79) of a 3D point cloud representing a 3D video content scene, a given angle, and an error value; the quantization function being defined such that a fourth point is generated by a sum obtained by dequantizing the quantized value and the error value, and such that an angle formed by the fourth point, the third point, and a second point of the 3D point cloud is less than or equal to the given angle; and - dequantizing said quantized values ​​according to the inverse of said quantization function. 15 . The apparatus of claim 14 , wherein the metadata comprises the given angle and coordinates of the third point or a distance between a first point and a third point of the 3D point cloud.

16. The apparatus of claim 14, wherein the metadata comprises a lookup table corresponding to an inverse of the quantization function.

Citation Information

Patent Citations

  • Method for encoding / decoding a picture block

    CN106170090A

  • Method of bit allocation for image & video compression using perceptual guidance

    US20140269903A1