Method and apparatus for encoding, decoding, and rendering 6DOF content from 3DOF components

By clustering and encoding 3D scene points into separate data streams with metadata, the method addresses rendering artifacts in 3DoF+ volumetric video, ensuring seamless navigation and efficient data management.

JP7824219B2Active Publication Date: 2026-03-04INTERDIGITAL VC HOLDINGS INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-12-18
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Existing 3DoF+ volumetric video rendering experiences suffer from rendering artifacts due to zones of missing information, and increasing the number of viewpoints to reduce these artifacts increases data load, impacting storage and transmission, with potential latency issues during seamless navigation.

Method used

The method involves clustering points in a 3D scene based on criteria such as depth range, semantic classification, or motion, projecting these clusters into 2D images, and encoding them into separate data streams, along with metadata for decoding and rendering, allowing for efficient data transmission and seamless navigation.

Benefits of technology

This approach reduces rendering artifacts by ensuring all necessary data is available for smooth navigation, minimizing data load, and optimizing storage and transmission, thereby enhancing the user experience in 6DoF volumetric rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007824219000001
    Figure 0007824219000001
  • Figure 0007824219000002
    Figure 0007824219000002
  • Figure 0007824219000003
    Figure 0007824219000003
Patent Text Reader

Abstract

The volumetric content is encoded by an encoder as a set of clusters and transmitted to a decoder, which retrieves the volumetric content. Clusters common to different viewpoints are retrieved and correlated. The clusters are projected onto a 2D image and coded as an independent video stream. Visual artifacts and data for storage and streaming are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present principles generally relate to the domain of three-dimensional (3D) scenes and volumetric video content. This document is also understood in the context of encoding, formatting, and decoding data representing textures and 3D scene geometry for rendering of volumetric content on end-user devices such as mobile devices or head-mounted displays (HMDs). [Background technology]

[0002] This section is intended to introduce the reader to various aspects of art that may be related to various aspects of the present principles, which are described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present principles. Accordingly, it should be understood that these statements are to be read in this light, and not as admissions of prior art.

[0003] In recent years, there has been a growth in available large field-of-view content (up to 360°). Such content may not be fully visible to users viewing the content on immersive display devices such as head-mounted displays, smart glasses, PC screens, tablets, or smartphones. This means that at any given moment, only a portion of the content is visible to the user. However, users can typically navigate within the content by various means, such as head movements, mouse movements, touchscreens, and voice. It is typically desirable to encode and decode this content.

[0004] Immersive video, also known as 360° flat video, allows users to view everything around them through head rotation around a stationary point. The rotation only allows for a three-degrees-of-freedom (3DoF) experience. For example, even if 3DoF video is sufficient for a first-order omnidirectional video experience using a head-mounted display device (HMD), it can quickly become frustrating for viewers who expect more degrees of freedom, such as by experiencing parallax. Furthermore, 3DoF can also induce dizziness because users not only rotate their head but also translate their head in three directions, a translation that is not reproduced in a 3DoF video experience.

[0005] The large field-of-view content can be, among others, a three-dimensional computer graphic imagery scene (3D CGI scene), a point cloud, or an immersive video. Many terms can be used to design such immersive video, such as Virtual Reality (VR), 360, panoramic, 4π steradian, immersive, omnidirectional, or large field-of-view.

[0006] Volumetric video (also known as 6 Degrees of Freedom (6DoF) video) is an alternative to 3DoF video. When watching 6DoF video, in addition to rotation, users can also translate their head and even their body within the viewed content, experiencing parallax and even volume. Such video significantly increases the sense of immersion and the perception of scene depth, and prevents dizziness by providing consistent visual feedback during head translation. Content is created by means of dedicated sensors that allow simultaneous recording of the color and depth of the desired scene. The use of color camera rigs combined with photogrammetry techniques is a method for performing such recording, even though technical difficulties remain.

[0007] While 3DoF video involves a sequence of images resulting from the unmapping of texture images (e.g., spherical images encoded according to latitude / longitude projection mapping or equirectangular mapping), 6DoF video frames embed information from several viewpoints. They can be viewed as a temporal sequence of points resulting from three-dimensional capture. Depending on the viewing conditions, two types of volumetric video can be considered. The first (i.e., full 6DoF) allows complete free navigation within the video content, while the second (known as 3DoF+) restricts the user's visual space to a limited volume called the visual bounding box, allowing limited translation of the head and parallax experience. This second context represents a valuable trade-off between free navigation and passive viewing conditions for seated audience members.

[0008] However, rendering artifacts, such as zones of missing information, may appear during a 3DOF+ volumetric rendering experience. There is a need to reduce the rendering artifacts.

[0009] In a 3DoF+ rendering experience, the user can move their viewpoint within a view bounding box. This is achieved by encoding the 3D scene from multiple viewpoints within the view bounding box. For multiple viewpoints within the view bounding box, points visible within 360 degrees from these viewpoints are projected to obtain 2D projections of the 3D scene. These 2D projections are encoded and transmitted over the network using well-known video coding techniques, such as HEVC (High Efficiency Video Coding).

[0010] The quality of the user experience depends on the number of viewpoints considered when encoding the 3D scene for a given visible bounding box. Increasing the number of viewpoints can reduce artifacts.

[0011] However, increasing the number of viewpoints increases the amount of data load associated with volumetric video, impacting storage and transmission.

[0012] Furthermore, when a user moves from a visible bounding box to an adjacent visible bounding box with a large amplitude, the data associated with the adjacent visible bounding box needs to be retrieved for rendering, and if the data load is large, there is a risk that the latency for retrieving and rendering the content will be perceptible to the user.

[0013] The data load for supporting 3DoF+ volumetric video needs to be minimized while providing a seamless navigation experience for users. Summary of the Invention

[0014] The following presents a simplified summary of the present principles to provide a basic understanding of some aspects of the present principles. This summary is not an extensive overview of the present principles. It is not intended to identify key or critical elements of the present principles. The following summary merely presents some aspects of the present principles in a simplified form as a prelude to the more detailed description provided below.

[0015] According to one or more embodiments, a method and device are provided for encoding volumetric content associated with a 3D scene, the method comprising: clustering points in the 3D scene into a plurality of clusters according to at least one clustering criterion; projecting the cluster according to projection parameters to obtain a set of 2D images; encoding the set of 2D images and the projection parameters into a set of data streams.

[0016] According to one embodiment, each of the 2D images is coded in a separate data stream. In another embodiment, a viewing box is defined in the 3D scene, and 2D images obtained by projecting clusters visible from two viewpoints into the viewing box are coded in the same data stream. In another embodiment, two viewing boxes are defined in the 3D scene, and 2D images obtained by projecting clusters visible from two viewpoints into each of the two viewing boxes are coded in the same data stream.

[0017] The present disclosure also relates to a method and device for decoding a 3D scene, the method comprising: obtaining at least one 2D image from the set of data streams, the 2D image representing a projection according to the projection parameters of at least one cluster of points in the 3D scene, the points in the cluster of points satisfying at least one clustering criterion; and backprojecting at least pixels of the 2D image according to the projection parameters and the viewpoint within the 3D scene.

[0018] In one embodiment, the method comprises: Obtaining metadata, the metadata comprising: A list of visibility boxes defined in the 3D scene, and a description of the view box, the data stream encoding a 2D image representing a cluster of 3D points that are visible from the viewpoint of the view box. and decoding a 2D image from the data stream that includes the clusters of 3D points that are visible from that viewpoint.

[0019] The present disclosure also relates to a medium storing instructions for causing at least one processor to perform at least the steps of the encoding method, and / or the decoding method, and / or the rendering method, and / or the receiving method described above. [Brief explanation of the drawings]

[0020] The present disclosure will be better understood, and other particular features and advantages will become apparent, on reading the following description, which makes reference to the accompanying drawings, in which: [Figure 1] 1 illustrates a three-dimensional (3D) model of an object and points of a point cloud corresponding to the 3D model, in accordance with a non-limiting embodiment of the present principles. [Figure 2] 1 shows an example of an encoding device, transmission medium, and decoding device for encoding, transmitting, and decoding data representing a sequence of 3D scenes, in accordance with a non-limiting embodiment of the present principles; [Figure 3] 16 shows an example of the architecture of an encoding and / or decoding device that may be configured to implement the encoding and / or decoding methods described in connection with FIGS. 14 and 15, in accordance with a non-limiting embodiment of the present principles. [Figure 4] 1 illustrates an example of one embodiment of the syntax of a stream when data is transmitted via a packet-based transmission protocol, in accordance with a non-limiting embodiment of the present principles. [Figure 5] Shows a 3D scene containing several objects. [Figure 6] Regarding 3DoF+ rendering, we present the concept of a 3DoF+ visibility bounding box in the three-dimensional space in which the 3D scene takes place. [Figure 7] Demonstrates the parallax experience made possible by volumetric rendering. [Figure 8] The parallax experience and disocclusion effect are shown. [Figure 9] 1 illustrates a method for structuring volumetric information according to a non-limiting embodiment of the present principles. [Figure 10] 1 shows an example of a method used to cluster a 3D scene into clusters of points, in accordance with a non-limiting embodiment of the present principles; [Figure 11] 1 illustrates a 2D parameterization of a 3D scene, in accordance with a non-limiting embodiment of the present principles. [Figure 12] 1 shows an example of a top view of a 3D scene with clusters, in accordance with a non-limiting embodiment of the present principles; [Figure 13]1 shows an example of a top view of a 3D scene with clusters, in accordance with a non-limiting embodiment of the present principles; [Figure 14] 1 illustrates a method for encoding volumetric content associated with a 3D scene, in accordance with a non-limiting embodiment of the present principles; [Figure 15] 1 illustrates a method for decoding volumetric content associated with a 3D scene, in accordance with a non-limiting embodiment of the present principles; [Figure 16] 1 illustrates a method for rendering volumetric content associated with a 3D scene, in accordance with a non-limiting embodiment of the present principles. [Figure 17] 1 illustrates a method for receiving volumetric content associated with a 3D scene, in accordance with a non-limiting embodiment of the present principles. DETAILED DESCRIPTION OF THE INVENTION

[0021] The present principles are more fully described below with reference to the accompanying drawings, in which examples of the present principles are shown. However, the present principles may be embodied in many alternative forms and should not be construed as limited to the embodiments set forth herein. Accordingly, while the present principles are susceptible to various modifications and alternative forms, specific examples thereof are shown by way of example in the drawings and are described in detail herein. It is to be understood, however, that there is no intention to limit the present principles to the particular forms disclosed, but on the contrary, the present disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present principles as defined by the appended claims.

[0022] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the present principles. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It will be further understood that as used herein, the terms "comprises," "comprising," "includes," and / or "including" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, when an element is referred to as "responsive to" or "connected to" another element, it may be directly responsive to or connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as "directly responsive to" or "directly connected to" another element, there are no intervening elements present. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated as " / ".

[0023] In this specification, terms such as "first," "second," etc. may be used to describe various elements, but it should be understood that these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first element can be referred to as a second element, and similarly, a second element can be referred to as a first element without departing from the teachings of the present principles.

[0024] Some of the figures include arrows on communication paths to indicate the primary direction of communication, however, it should be understood that communication may occur in the opposite direction to the depicted arrow.

[0025] Some examples are described with reference to block diagrams and operational flowcharts, in which each block represents circuit elements, modules, or portions of code, with each block including one or more executable instructions for implementing a specified logical function. It should also be noted that in other implementations, the functions noted in the blocks may occur out of the order noted. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may be executed in the reverse order, depending on the functionality involved.

[0026] As used herein, "by one example" or "in one example" means that a particular feature, structure, or characteristic described in connection with this embodiment may be included in at least one implementation of the present principles. The appearances of the phrase "by one example" or "in one example" in various places in this specification do not necessarily all refer to the same embodiment, and in separate or alternative embodiments are not necessarily mutually exclusive of other embodiments.

[0027] Reference numerals appearing in the claims are by way of example only and shall have no limiting effect on the scope of the claims. Although not expressly stated, the present embodiments and variations may be used in any combination or subcombination.

[0028] The present principles are described with respect to particular embodiments of a method for encoding volumetric content relating to a 3D scene into a stream, a method for decoding such volumetric content from the stream, and a method for volumetric rendering of the volumetric content decoded in accordance with the mentioned decoding methods.

[0029] According to non-limiting embodiments, a method is disclosed for structuring volumetric information related to a 3D scene to be coded and / or transmitted (e.g., streamed) and / or decoded and / or rendered based on clustering of points in the 3D scene. To capture the 3D scene, the 3D space is organized into visibility bounding boxes, called 3DoF+ visibility bounding boxes. Clusters common to different 3DoF+ visibility bounding boxes are obtained. The volumetric content of the 3DOF+ visibility bounding boxes is coded using the clusters. A 6DoF volumetric rendering experience is achieved by a succession of 3DoF+ volumetric rendering experiences.

[0030] The advantages of the present principles for encoding, transmitting, receiving and rendering are presented in the following description with reference to the drawings.

[0031] FIG. 1 illustrates a three-dimensional (3D) model 10 of points of an object and a point cloud 11 corresponding to the 3D model 10. The 3D model 10 and point cloud 11 may correspond to, for example, a potential 3D representation of an object in a 3D scene containing other objects. The model 10 may be a 3D mesh representation, and the points of the point cloud 11 may be vertices of the mesh. The points of the point cloud 11 may also be points spread on the surface of a face of the mesh. The model 10 may also be represented as a splatted version of the point cloud 11, where the surface of the model 10 is created by splatting the points of the point cloud 11. The model 10 may be represented by many different representations, such as voxels or splines. FIG. 1 illustrates the fact that a point cloud may be defined as a surface representation of a 3D object, and that a surface representation of a 3D object may be generated from a cloud of points. As used herein, projecting a point of a 3D object (by an elongation point of the 3D scene) onto an image is equivalent to projecting any representation of this 3D object, for example a point cloud, a mesh, a spline model or a voxel model.

[0032] A point cloud can be represented in memory, for example, as a vector-based structure, with each point having its own coordinates in the reference frame of the viewpoint (e.g., three-dimensional coordinates XYZ, or solid angle and distance (also called depth) from / to the viewpoint) and one or more attributes, also called components. Examples of components are color components that can be expressed in various color spaces, for example, RGB (red, green, and blue) or YUV (Y is the luminance component and UV are the two color difference components). A point cloud is a representation of a 3D scene containing objects. A 3D scene can be viewed from a given viewpoint or range of viewpoints. Point clouds can be represented in many ways, for example, From capturing real objects photographed by a camera rig, optionally complemented by a depth active sensing device; From capturing virtual / synthetic objects photographed by virtual camera rigs in modeling tools, • It can be obtained from a mixture of both real and virtual objects.

[0033] 2 shows a non-limiting example of encoding, transmission and decoding of data representing a sequence of 3D scenes, e.g., a coding format that can accommodate 3DoF, 3DoF+ and 6DoF decoding simultaneously.

[0034] A sequence of 3D scenes 20 is captured. Whereas the sequence of photographs is a 2D video, the sequence of 3D scenes is a 3D (also called volumetric) video. The sequence of 3D scenes can be provided to a volumetric video rendering device for 3DoF, 3Dof+ or 6DoF rendering and display.

[0035] A sequence of 3D scenes 20 is provided to an encoder 21. The encoder 21 takes as input a 3D scene or a sequence of 3D scenes and provides a bitstream representing the input. The bitstream may be stored in a memory 22 and / or on an electronic data medium and may be transmitted over a network 22. The bitstream representing the sequence of 3D scenes may be read from the memory 22 and / or received from the network 22 by a decoder 23. The decoder 23 is input with the bitstream and provides the sequence of 3D scenes, for example in point cloud format.

[0036] The encoder 21 may include several circuits that implement several steps. In the first step, the encoder 21 projects each 3D scene onto at least one 2D photograph. 3D projection is any method of mapping three-dimensional points onto a two-dimensional plane. Because modern methods for displaying graphic data are based on planar (pixel information from several bit planes) two-dimensional media, the use of this type of projection is widespread, especially in computer graphics, manipulation, and drafting. The projection method selected and used can be represented and coded as a set or list of projection parameters. The projection circuit 211 provides at least one two-dimensional image 2111 for each 3D scene in the sequence 20. The image 2111 includes color and depth information representing the 3D scene projected onto the image 2111. In a variant, the color and depth information are coded in two separate images 2111 and 2112.

[0037] The metadata 212 is used and updated by the projection circuitry 211. The metadata 212 includes information about the projection operation (e.g., projection parameters) and how color and depth information is organized within the images 2111 and 2112, as described in connection with Figures 5-7.

[0038] A video encoding circuit 213 encodes the sequence of images 2111 and 2112 as a video. The images of the 3D scene 2111 and 2112 (or a sequence of images of a 3D scene) are encoded in a stream by the video encoder 213. The video data and metadata 212 are then encapsulated in a data stream by a data encapsulation circuit 214.

[0039] The encoder 213 may, for example, -JPEG, specification ISO / CEI10918-1UIT-T Recommendation T.81, https: / / www.itu.int / rec / T-REC-T.81 / en; -Compliant with encoders such as AVC, also known as MPEG-4 AVC or h264, ITU-TH.264 and ISO / CEI MPEG-4-Part 10 (ISO / CEI14496-10), http: / / www.itu.int / rec / T-REC-H.264 / en, HEVC (whose specifications can be found on the ITU website, T Recommendation, H Series, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en), -3D-HEVC (an extension of HEVC whose specification can be found on the ITU website, T Recommendation, H Series, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en annex G and I), -VP9, developed by Google, or -AV1 (AO Media Video 1) developed by the Alliance for Open Media.

[0040] The data stream is stored by a decoder 23 in a memory accessible, for example, via a network 22. The decoder 23 comprises different circuits that implement different steps of the decoding. The decoder 23 takes as input the data stream generated by the encoder 21 and provides a sequence of 3D scenes 24 to be rendered and displayed by a volumetric video display device, such as a head-mounted device (HMD). The decoder 23 obtains the stream from a source 22. For example, the source 22 may be - local memory, such as video memory or RAM (or Random Access Memory), flash memory, ROM (or Read Only Memory), hard disk, etc. a storage interface, such as an interface to a mass storage, RAM, flash memory, ROM, optical disk or magnetic support; a communication interface, such as a wired interface (e.g., a bus interface, a wide area network interface, a local area network interface) or a wireless interface (e.g., an IEEE 802.11 interface or a Bluetooth interface), - a user interface, such as a graphical user interface, that allows a user to input data.

[0041] The decoder 23 includes a circuit 234 for extracting data encoded within the data stream. The circuit 234 takes the data stream as input and provides metadata 232 corresponding to the metadata 212 encoded in the stream and the two-dimensional video. The video is decoded by a video decoder 233, which provides a sequence of images. The decoded images include color and depth information. In a variant, the video decoder 233 provides a sequence of two images, one including color information and the other including depth information. The circuit 231 does not use the metadata 232 to project the color and depth information from the decoded images, but instead provides a sequence of 3D scene 24. The sequence of 3D scene 24 corresponds to the sequence of 3D scene 20 and video compression, potentially with reduced precision associated with encoding it as 2D video.

[0042] The principles disclosed herein relate to the encoder 21, and more particularly to the projection circuitry 211 and the metadata 212. They also relate to the decoder 23, and more particularly to the backprojection circuitry 231 and the metadata 232.

[0043] Figure 3 shows an example of the architecture of a device 30 that may be configured to perform the methods described in relation to Figures 14 and 15. The encoder 21 and / or decoder 23 of Figure 2 may implement this architecture. Alternatively, the circuits of the encoder 21 and / or decoder 23 may be devices according to the architecture of Figure 3, coupled together, for example, via their bus 31 and / or via an I / O interface 36.

[0044] The device 30 includes the following elements coupled together by a data and address bus 31: a microprocessor 32 (or CPU), for example a DSP (or Digital Signal Processor), -ROM (or read-only memory) 33; a RAM (or Random Access Memory) 34; a storage interface 35; an I / O interface 36 for receiving data to be transmitted from an application; - a power source, for example a battery;

[0045] According to one example, the power source is external to the device. In each of the mentioned memories, the word "register" as used herein may correspond to a small area (a few bits) or a very large area (e.g., an entire program or a large amount of received or decoded data). The ROM 33 contains at least programs and parameters. The ROM 33 can store algorithms and instructions for performing techniques according to the present principles. When switched on, the CPU 32 uploads the program in the RAM and executes the corresponding instructions.

[0046] The RAM 34 contains in registers the programs executed by the CPU 32 and uploaded after switching on the device 30, input data in registers, intermediate data for different states of the methods in registers, and other variables used for the execution of the methods in registers.

[0047] The implementations described herein may be implemented in, for example, a method or process, an apparatus, a computer program product, a data stream, or a signal. Even when discussed only in the context of a single form of implementation (e.g., discussed only as a method or device), the implementation of the discussed features may also be implemented in other forms (e.g., a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The method may be implemented in an apparatus such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include, for example, communication devices such as computers, mobile phones, handheld / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.

[0048] According to an embodiment, the device 30 is configured to implement the method described in relation to Figures 14 and 15, -Mobile devices and a communication device; -A gaming device, -a tablet (or tablet computer); -Laptop and -A still camera, -Video camera and - an encoding chip; - a server (for example a broadcast server, a video-on-demand server or a web server), and belongs to the set including:

[0049] FIG. 4 shows an example of an embodiment of the syntax of a stream when data is transmitted via a packet-based transmission protocol. FIG. 4 shows an example of Structure 4 of a volumetric video stream for one visible bounding box. Structure 4 organizes the stream with independent elements of the syntax. In this example, Structure 4 includes three elements: Syntaxes 41, 42, and 43. The Syntax 41 element is a header that includes data common to all elements of the syntax of Structure 4. For example, Header 41 includes metadata describing the nature and role of each element of the syntax of Structure 4. Header portion 41 also includes part of the metadata 212 of FIG. 2, such as information about the position of the visible bounding box (e.g., the center viewpoint of the visible bounding box).

[0050] Structure 4 includes a payload that includes elements of syntax 42 and at least one element of syntax 43. The elements of syntax 42 include data representing encoded video data, e.g., color and depth images 2111 and 2112.

[0051] Elements of syntax 43 contain metadata about how images 2111 and 2113 are encoded, in particular the specific parameters used to project and pack points of a 3D scene onto the images. Such metadata can be associated with each image of a video or with a group of images (also known as a Group of Pictures (GoP) in video compression standards).

[0052] As mentioned above, rendering artifacts, such as zones of missing information, can appear during a volumetric rendering experience. An example of missing information is disparity information. For example, in 3DoF+ volumetric rendering, the viewing space is restricted to a limited volume called a viewing bounding box. A central viewpoint is associated with each viewing bounding box. When a user translates within the viewing bounding box from the central viewpoint of the viewing bounding box, portions of the 3D scene that were initially occluded become visible. This is called the parallax effect, and the data associated with the occluded portions is called disparity data. To render these occluded portions as the user moves, disparity data must be encoded and transmitted. Depending on how the data is encoded, some disparity data may be missing, degrading the rendering experience. The parallax effect is described in more detail with reference to Figures 5, 6, and 7.

[0053] Figure 5 shows an image representing a 3D scene. The 3D scene can be captured using any suitable technique. The exemplary 3D scene shown in Figure 5 comprises several objects: houses 51 and 52, people 54 and 55, and a well 56. A cube 53 is shown in Figure 5 to indicate a view bounding box from which a user is likely to observe the 3D scene. The central viewpoint of view bounding box 53 is referred to as 50.

[0054] FIG. 6 illustrates the concept of visibility bounding boxes in more detail when the 3D scene of FIG. 5 is rendered on an immersive rendering device (e.g., a CAVE or a head-mounted display device (HMD)). Scene point 64a in the 3D scene corresponds to the elbow of person 54. The scene point 64a is visible from viewpoint 50 because no opaque objects are placed between viewpoint 50 and scene point 64a. In contrast, scene point 65a, which corresponds to the elbow of person 55, is not visible from viewpoint 50 because it is occluded by a point on person 54. In 3DoF+ rendering, a user can change their viewpoint within the 3DoF+ visibility bounding box, as described above. For example, as illustrated in connection with FIG. 7, a user can move their viewpoint within visibility bounding box 53 and experience parallax.

[0055] FIG. 7 illustrates the parallax experience enabled by volumetric rendering of the 3D scene of FIG. 5. FIG. 7B shows a portion of the 3D scene that a user can see from a central viewpoint 50. From this viewpoint, people 54 and 55 are in a given spatial configuration; for example, the left elbow of person 55 is obscured by the body of person 54 while the head is visible. This configuration does not change when the user rotates their head with three degrees of freedom around the central viewpoint 50. When the viewpoint is fixed, the left elbow of person 55 (designated 65a in FIG. 6) is invisible. FIG. 7A shows the same 3D scene from a first peripheral viewpoint (designated 67 in FIG. 6) to the left of the visibility bounding box 53. From viewpoint 67, point 65a is visible due to the parallax effect. This is called the disocclusion effect. For example, by moving from viewpoint 50 to viewpoint 67, point 65a is disoccluded. Figure 7C shows the same 3D scene viewed from a second peripheral viewpoint (designated 68 in Figure 6) to the right of visibility bounding box 53. From viewpoint 68, person 55 is almost completely obscured by person 54, but is still visible from viewpoint 50. With reference to Figure 6, it can be seen that by moving from viewpoint 50 to viewpoint 68, point 65b is occluded.

[0056] In most cases, the deoccluded data corresponds to a small patch of data. Figure 8 illustrates the deoccluded data required for volumetric rendering. Figure 8A is a top view of a 3D scene including two objects P1 and P2 captured by three virtual cameras: a first peripheral camera C1, a central camera C2, and a second peripheral camera C3, associated with a view bounding box V. The view bounding box V is centered at the position of the central camera C2. Points visible from the virtual cameras C1, C2, and C3 are represented by lines 81, 82, and 83, respectively. Figures 8B, 8C, and 8D show renderings of the 3D scene captured as described in connection with Figure 8A. In Figures 8B and 8C, a cone F bounds the field of view and the portion of the 3D scene visible from viewpoints O0 and O1, respectively. O0 and O1 are viewpoints included in the view bounding box V. By moving from viewpoint O0 toward viewpoint O1, the user experiences parallax. The deoccluded points represent small patches within the background object.

[0057] In FIG. 8D, O2 represents a viewpoint outside the view bounding box V viewpoint. From viewpoint O2, new data that was invisible from view bounding box V, represented by segment D, is now visible and unmasked. This is a de-occlusion effect. Segment D does not belong to the volumetric content associated with view bounding box V. When a user makes a large amplitude movement, such as going from viewpoint O0 to viewpoint O2, and moves outside view bounding box V, the de-occlusion effect in different regions of the 3D scene may no longer be compensated. The unmasked portions may represent missing information for large areas of high visibility on the rendering device, resulting in a poor immersive experience.

[0058] The way in which the volumetric content information to be coded is structured affects coding efficiency, as will be seen below.

[0059] FIG. 9A shows a first method for structuring volumetric information representing a 3D scene, and FIG. 9B shows a method for structuring the same volumetric information for the 3D scene of FIG. 8 according to a non-limiting embodiment of the present principles.

[0060] According to the first method, the unique elements encompassed by the closed dotted line 910 are captured from the viewpoint O0. In fact, the only accessible data is the data represented by the thick lines 911, 912, and 913. It can be observed that the area of ​​object P2 occluded by object P1 is not accessible, i.e., the area of ​​P2 is missing.

[0061] In this principle, points in a 3D scene are clustered according to a clustering criterion. In the embodiment shown in FIG. 9B, the clustering criterion relates to the depth range of the points in the 3D scene, thus separating the 3D scene into multiple depth layers. This makes it possible, for example, to create background and foreground clusters that include parts of physical objects that contribute to the background and foreground of the scene, respectively. Alternatively, or in combination, the clustering can be based on, for example, semantic classification of the points, motion classification, and / or color segmentation. All points in a cluster share the same characteristics. In FIG. 9B, two clusters are obtained, enclosed by the closed dotted lines 921 and 922, respectively. The accessible data, represented by the bold lines 923 and 924, differs from that obtained with the first method, as shown in FIG. 9A. In FIG. 9B, all information related to object P2 is available, even information behind object P1 as seen from viewpoint O0. This is not the case with the method described in connection with the diagram in FIG. 9A. By structuring the volumetric information representing a 3D scene by clustering points according to the present principles, the information available for rendering the 3D scene can be increased. Referring again to the parallax experience mentioned above, one advantage of the above clustering method is that data related to occluded regions is accessible regardless of viewpoint.

[0062] 10 shows how to obtain clusters 921 and 922. This example refers to the case where the clustering criterion is a depth filtering criterion. One way to obtain clusters is to capture points by virtual cameras with different positions, orientations, and fields of view. Each virtual camera is optimized to capture as many points of a given cluster as possible. For example, in FIG. 10, cluster 921 is obtained by capturing points captured by virtual camera C. A_0 The virtual camera C A_0 captures all pixels within the near depth range and cuts out object P2 that does not belong to the near depth range. B_0 The virtual camera C B_0 captures all pixels within the far depth range and crops out the object P1 that does not belong to the far depth range. Advantageously, the background cluster is acquired with a virtual camera positioned at a far distance, regardless of the viewpoint and the visible bounding box, while the foreground cluster is acquired with a virtual camera positioned at a different viewpoint within the visible bounding box. The mid-depth cluster is typically acquired with a virtual camera positioned at a fewer number of viewpoints within the visible bounding box compared to the foreground cluster.

[0063] We now describe how volumetric information representing a 3D scene structured by point clustering methods such as those described above can be encoded into a video stream.

[0064] FIG. 11 illustrates a 2D atlas approach used to encode volumetric content representing a 3D scene from a given viewpoint 116. In FIG. 11, a top view 100 of the 3D scene is shown. The 3D scene includes a person 111, a flowerpot 112, a tree 113, and a wall 114. Image 117 is an image representing the 3D scene as observed from viewpoint 116. In a point clustering method, clusters represented by dotted ellipses 111c, 112c, 113c, and 114c are obtained from the volumetric content and projected toward viewpoint 116 to create a set of 2D images. The set of 2D images is then packed to form atlas 115 (an atlas is a collection of 2D images). The organization of the 2D images within the atlas defines the atlas layout. In one embodiment, two atlases with identical layouts are used: one for color (i.e., texture) information and one for depth information.

[0065] At successive points in time, a time series of 2D atlases is generated. Typically, the time series of 2D atlases is transmitted in the form of a set of coded videos, each video corresponding to a particular cluster, and each image in the video corresponding to a 2D image obtained by projecting this particular cluster at a given moment from viewpoint 116. The succession of 2D images of a particular cluster constitutes an independent video.

[0066] The point clustering method according to the present principles aims to structure the volumetric information representing a 3D scene in a way that allows this volumetric information to be coded as a set of independent videos.

[0067] In this principle, the 3D scene is not transmitted as a single video stream corresponding to a series of images 117 acquired at different times, but as a set of smaller, independent videos corresponding to the succession of 2D images in the time sequence of the 2D atlas. Each video can be transmitted independently of the others. For example, the different videos can be acquired by using virtual cameras with different fields of view. In another example, the different videos can be encoded at different image rates or different quality levels.

[0068] For example, a frequent configuration is a 3D scene in which animated foreground objects move a lot compared to the background of the scene. These animated objects have their own life cycle and can advantageously be coded at a higher image rate than the background.

[0069] Also, when volumetric content is streamed, the video quality can be adjusted for each video stream to suit the streaming environment, e.g., a video stream corresponding to the foreground may be encoded at a higher quality than a video stream corresponding to the background of a scene.

[0070] Another advantage is that it allows for the individualization of the scalable 3D scene, for example customization by inserting specific objects, for example advertisements, etc. The customization is optimized compared to volumetric content that is coded in a monolithic way.

[0071] For decoding, the 3D scene is obtained by combining the independent video streams. 2D images corresponding to different clusters in the 2D atlas are recombined to construct an image representing the 3D scene as seen from viewpoint 116. This image undergoes a 2D-to-3D backprojection step to obtain volumetric data. The volumetric data is rendered during the volumetric rendering experience from a viewpoint corresponding to viewpoint 116 in the 3D rendering space.

[0072] We now describe how a 6DOF volumetric rendering experience, which builds on the continuation of the 3DOF+ volumetric rendering experience, can benefit from using the point clustering method as described above.

[0073] A 3D scene can be rendered by successively rendering the volumetric content associated with the view bounding boxes and moving from one view bounding box to another within the 3D rendering space. For example, advantages related to data storage and transfer are highlighted below.

[0074] Figure 12 is a top view of the 3D scene of Figure 11, with the visibility bounding box represented in the form of a dotted ellipse 121. Two dotted lines 122 and 123 represent the field of view visible from the visibility bounding box 121. This field of view includes four clusters obtained by clustering points in the 3D scene of Figure 11: cluster 120a associated with the flower pot 112, cluster 120b associated with the person 111, cluster 120c associated with the tree 113, and cluster 120d associated with the wall 114.

[0075] Two viewpoints 124 and 125 contained within the visibility bounding box 121 are shown, along with their respective fields of view (represented by two cones 126 and 127). It can be observed that some clusters or parts of some clusters are common to viewpoints 124 and 125. In the example of FIG. 12, these common clusters are clusters 120c and 120d. In this particular example, they correspond to parts of the 3D scene that are far away from viewpoints 124 and 125. The 2D image resulting from the 3D-2D projection step of these common clusters is called a 2D common image. The 2D images resulting from the 3D-2D projection step of clusters other than the common cluster are called 2D patches.

[0076] The 2D common image typically contains a large number of non-empty pixels. For example, when a depth criterion is used, the common cluster often corresponds to background points of the volumetric content and contains a large number of points. 2D patches are typically small regions that are distinct from the region surrounding them. 2D patches typically contain less information than the 2D common image and thus have a smaller size, e.g., in terms of the number of pixels. For example, a cluster corresponding to foreground points of the volumetric content often contains a limited number of points, e.g., representing characters or objects placed in front of large background features.

[0077] The two atlases, each containing a set of 2D images resulting from a 3D-2D projection of the set of clusters associated with viewpoints 124 and 125, respectively, have a common 2D common image. Thus, when moving within view bounding box 121 from viewpoint 124 to viewpoint 125 or vice versa, the data corresponding to the 2D common image is already available for rendering. This improves the user's parallax experience and eliminates the latency that would otherwise be required to retrieve and render this data. Another advantage is that the amount of data transmitted is reduced.

[0078] Referring again to the 2D atlas approach, the 2D common image is transmitted in the form of one common video, while each 2D patch is transmitted as one specific video. The common information previously embedded in each image 117 is mutualized and transmitted separately in the common video. When a depth criterion is used, the common video usually corresponds to a cluster representing the background part of the 3D scene. The common video is very stable or hardly changes over a period of time, such as the wall 114 in FIG. 11. Therefore, a highly efficient codec can be used to encode the common video, for example, by temporal prediction.

[0079] Figure 13 is a top view of the 3D scene of Figure 11, with two view bounding boxes 131 and 138 represented. One viewpoint 134 within view bounding box 131 and one viewpoint 135 within view bounding box 138 are shown. The first viewpoint 134 is located within view bounding box 131, and the second viewpoint 135 is located within view bounding box 138. The views from viewpoints 134 and 135 are referenced 136 and 137, respectively. It can be seen that clusters or portions of clusters are common to both views 136 and 137. Thus, view bounding box 131 and view bounding box 138 have a common cluster or portion of a cluster.

[0080] The 2D common images corresponding to these common clusters can be correlated between several visible bounding boxes. These images can be stored, encoded, transmitted, and rendered with respect to several visible bounding boxes. This further reduces the data load for storage and transmission. Another advantage is the reduction of latent artifacts when a user makes large movements in the rendering space and goes from a first to a second visible bounding box.

[0081] 14 shows a method for encoding volumetric content associated with a 3D scene, in accordance with a non-limiting embodiment of the present principles, which method is intended for use in encoder 21 of FIG.

[0082] In step 1400, a 3D scene is acquired from a source.

[0083] In step 1401, points in the 3D scene are clustered into multiple clusters according to at least one clustering criterion. In one embodiment, the clustering criterion relates to the depth range of the points in the 3D scene, thereby separating the 3D scene into multiple depth layers. This makes it possible, for example, to create background and foreground clusters that include parts of physical objects that contribute to the background and foreground of the scene, respectively. Alternatively, or in combination, the clustering may be based, for example, on semantic classification of the points, and / or motion classification, and / or color segmentation. For a given viewpoint, the 3D scene is described as a set of clusters.

[0084] In step 1402, the clusters of the set of clusters are projected according to the projection parameters to obtain a set of 2D images. The 2D images are packed into an atlas, or into two atlases with the same layout, for example, one atlas containing color data and the other containing depth data.

[0085] In step 1403, a volumetric content is generated that holds data representing the 3D scene, the data representing the 3D scene being the atlas or atlas pair obtained in step 1402.

[0086] In one embodiment, the 3D rendering space is organized into visibility bounding boxes, each of which includes a central viewpoint, in a preferred embodiment a peripheral viewpoint. In step 1401', clusters common to different visibility bounding boxes are obtained.

[0087] When step 1401' is performed, step 1402 includes two sub-steps 1402A and 1402B. In sub-step 1402A, clusters common to different visibility bounding boxes are projected according to projection parameters to obtain a 2D common image. In sub-step 1402B, clusters other than the clusters common to different visibility bounding boxes are projected to obtain 2D patches. This is done for each visibility bounding box. For each visibility bounding box, the cluster is projected in the direction of the center point of the visibility bounding box to create a set of 2D patches. Preferably, the cluster is also projected in the direction of one or more peripheral viewpoints, so that additional sets of 2D patches are created (one for each peripheral viewpoint). As a result, each visibility bounding box is associated with a 2D common image and several sets of 2D patches.

[0088] In step 1402', metadata is generated that includes a list of visible bounding boxes contained in the 3D rendering space of the 3D scene and a list of sets of 2D common images and 2D patches to apply with respect to the visible bounding boxes in the 3D rendering space. The metadata generated in step 1402' is included in the volumetric content generated in step 1403. For example, a structure 4 as depicted in Figure 4 is used to pack information related to the visible bounding boxes, and all structures 4 of the 3D scene are packed together in a super structure that includes a header containing the metadata generated in step 1402'.

[0089] For example, the metadata generated in step 1402' may be: - A list of visible bounding boxes in the 3D rendering space, a list of common clusters in the 3D rendering space, each common cluster being characterized by a common cluster identifier and associated with a unique resource identifier used to obtain a corresponding video stream from a source; for each visible bounding box, a list of a set of clusters representing the 3D scene for this visible bounding box; For each set of clusters associated with a visible bounding box, ○ a common cluster identifier, and a list of clusters other than the common cluster with unique resource identifiers to obtain the corresponding video streams from the source; Includes:

[0090] In an advantageous embodiment, the 2D images are encoded at different levels of quality or different image rates so that several sets of 2D images are generated for the same viewpoint, which allows adapting the video quality or speed, for example, to take into account streaming environments.

[0091] 15 shows a method for decoding volumetric content associated with a 3D scene, in accordance with a non-limiting embodiment of the present principles, which method is intended for use with decoder 23 of FIG.

[0092] In step 1500, volumetric content is obtained from a source. The volumetric content includes at least one 2D image representing at least one cluster of points in a 3D scene. The points in the cluster satisfy a clustering criterion. In one embodiment, the clustering criterion relates to the depth range of the points in the 3D scene. Alternatively, or in combination, the clustering criterion relates to, for example, semantic classification, and / or motion classification, and / or color segmentation of the points.

[0093] In step 1501, at least one 2D image is predicted according to projection parameters.

[0094] In step 1502, a 3D point cloud representing the 3D scene is obtained from the backprojected 2D image.

[0095] FIG. 16 illustrates a method for rendering volumetric content associated with a 3D scene in a device configured to function as a volumetric display or rendering device, in accordance with a non-limiting embodiment of the present principles.

[0096] In step 1600, a first viewpoint in the 3D rendering space is obtained. This first viewpoint is associated with a first view bounding box in the 3D rendering space. When the rendering device is an HMD, the first viewpoint is the end user's position, for example, obtained using an IMU (Inertial Measurement Unit) of the HMD. The HMD comprises one or more display screens (e.g., LCD (Liquid Crystal Display), OLED (Organic Light Emitting Diode), or LCOS (Liquid Crystal on Silicon)) configured to measure change(s) in position of the HMD, e.g., gyroscope or IMU (Inertial Measurement Unit), according to one, two, or three axes of the real world (pitch, yaw, and / or roll axes).

[0097] In step 1601, first volumetric content related to a 3D scene is received by a rendering device, the first volumetric content including metadata associated with the 3D scene (a list of visible bounding boxes contained in the 3D rendering space and, for each visible bounding box, a list of a 2D common image and a set of 2D patches), as well as video data and metadata associated with the first visible bounding boxes, as described above in connection with step 1402′.

[0098] In step 1602, the first volumetric content is decoded using the decoding method described above to obtain a first 3D point cloud representing the 3D scene. Based on the metadata received in step 1601, a 2D common image and a set of 2D patches corresponding to the first viewpoint are selected. The 2D image is not projected according to the projection parameters transmitted in the stream. As a result, a first 3D point cloud is obtained.

[0099] In step 1603, the first 3D point cloud is rendered from a first viewpoint and displayed according to a volumetric rendering.

[0100] As mentioned above, 6DoF rendering can be enabled by successive 3DoF+ rendering of some volumetric content. To achieve this, the rendering method according to the present principles includes the following additional steps:

[0101] In step 1604, the user moves from a first viewpoint to a second viewpoint in the rendered 3D space.

[0102] In step 1605, a set of 2D images to be used for rendering from the second viewpoint is obtained based on the metadata obtained in step 1601. 2D images that are not already available for rendering are obtained from the source. Previously obtained 2D common images do not need to be obtained again.

[0103] In step 1606, the 2D image acquired from the source is unprojected to create a second 3D point cloud that is combined with points from the first 3D point cloud that correspond to 2D images that are common between the first and second visible bounding boxes.

[0104] In step 1607, the result of this combination is rendered from a second perspective and displayed according to a 3DoF+ volumetric rendering technique.

[0105] Steps 1604-1607 may be repeated as the user moves from one viewpoint to another within the 3D scene.

[0106] The rendering method described above shows how the present principles enable 6DoF volumetric rendering based on multi-viewpoint 3DoF+ rendering by using a set of volume elements in the form of clusters.

[0107]

[0013] Figure 17 illustrates a method for receiving volumetric content related to a 3D scene in a 3D rendering space in a device configured to function as a receiver in accordance with a non-limiting embodiment of the present principles. In the example of Figure 17, the volumetric rendering experience takes place in an adaptive streaming environment. Video streams are encoded at different quality levels or different picture rates. The receiver also includes an adaptive streaming player that detects the conditions of the adaptive streaming environment and selects which video stream to transmit.

[0108] In step 1700, metadata associated with the 3D scene is received by a receiver. For example, when using the DASH streaming protocol, the metadata is transmitted using a Media Presentation Description (MPD), also called a manifest. As mentioned above, the metadata includes a list of visible bounding boxes contained in the 3D rendering space and, for the visual bounding boxes / viewpoints, information about the clusters used for rendering (identification of the clusters used and information for obtaining the clusters from the source).

[0109] In step 1701, the adaptive streaming player detects the conditions of the streaming environment, for example, available bandwidth.

[0110] In step 1702, a particular view bounding box / viewpoint in the 3D rendering space is considered. The adaptive streaming player uses the conditions of the streaming environment to select a set from a list of sets of at least one 2D common image and at least one 2D patch. For example, priority is given to the foreground cluster so that high quality 2D patches are selected along with low quality 2D common images.

[0111] In step 1703, the adaptive streaming player sends a request for the selected set to the server.

[0112] The receiver receives the selected set in step 1704. The set is then decoded and rendered according to one of the methods described above.

[0113] Criteria other than depth, such as movement, can be used in addition to or instead of depth. Typically, 2D patches encoding fast-moving clusters are selected with a bandwidth priority compared to stationary clusters. Indeed, parts of a 3D scene may be static, while other objects may be moving at various speeds. This aspect is particularly noticeable for small animated objects (often in the foreground), which may have their own life cycle (position, color) that differs from other elements of the scene (often in the background). For example, by clustering such objects with respect to their movement speed, they can be transmitted according to different transmission parameters, such as frequency rates. Therefore, the advantage is a reduction in streaming costs due to content heterogeneity.

[0114] In another implementation of the present principles, the receiver includes a prediction module for predicting the user's next position in the 3D rendering space. The corresponding set is selected based on metadata. If several sets of clusters are available, one of them is selected as described above. Finally, the receiver transmits a request to obtain the corresponding video stream.

[0115] In this principle, some video streams are likely to be needed, for example, background video streams that are more stable. Advantageously, the receiver takes into account the occurrence probability and triggers acquisition of the most probable video streams first. Foreground clusters are more versatile and easier to transmit. The receiver can postpone prediction and acquisition until the last acceptable moment. As a result, the cost of misprediction is reduced.

[0116] The embodiments described herein may be implemented in, for example, a method or process, an apparatus, a computer program product, a data stream, or a signal. Even when discussed only in the context of a single form of implementation (e.g., discussed only as a method or device), the implementation of the discussed features may also be implemented in other forms (e.g., ). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. A method may be implemented in an apparatus such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices such as, for example, smartphones, tablets, computers, mobile phones, handheld / personal digital assistants ("personal digital assistants," or "PDAs"), and other devices that facilitate communication of information between end users.

[0117] Implementations of the various processes and features described herein may be embodied in a variety of different devices or applications, particularly, for example, devices or applications associated with data encoding, data decoding, view generation, texture processing, and other processing of images and associated texture and / or depth information. Examples of such devices include encoders, decoders, post-processors that process output from decoders, pre-processors that provide input to encoders, video coders, video decoders, video codecs, web servers, set-top boxes, laptops, personal computers, mobile phones, PDAs, and other communication devices. As should be clear, the devices may be mobile and installed in mobile vehicles.

[0118] Furthermore, a method may be implemented by instructions executed by a processor, and such instructions (and / or data values ​​produced by an implementation) may be stored on a processor-readable medium, such as, for example, an integrated circuit, a software carrier, or other storage device, e.g., a hard disk, a compact diskette ("CD"), an optical disk (e.g., a DVD, often referred to as a digital versatile disk or digital video disk), a random access memory ("RAM"), or a read-only memory ("ROM"). The instructions may form an application program tangibly embodied on the processor-readable medium. The instructions may be, for example, hardware, firmware, software, or a combination. The instructions may be found, for example, in an operating system, a separate application, or a combination of the two. A processor may therefore be characterized, for example, as both a device configured to execute a process and a device that includes a processor-readable medium (e.g., a storage device) having instructions for executing a process. Furthermore, a processor-readable medium may store data values ​​produced by an implementation in addition to, or in place of, instructions.

[0119] As will be apparent to those skilled in the art, implementations may generate various signals formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry, as data, rules for writing or reading syntax of a described embodiment, or to carry, as data, the actual syntax value written by a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The signal it carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.

[0120] Many implementations have been described. Nevertheless, it will be understood that various modifications may be made. For example, elements of different implementations may be combined, supplemented, modified, or deleted to produce other implementations. Moreover, those skilled in the art will understand that other structures and processes may be substituted for those disclosed, with the resulting implementation performing at least substantially the same function in at least substantially the same way to achieve at least substantially the same results as the disclosed implementations. Accordingly, these and other implementations are contemplated by this application.

Claims

1. 1. A method for encoding a 3D scene, comprising: clustering the points of the 3D scene into a plurality of clusters according to depth ranges of the points in the 3D scene, the plurality of clusters including at least a background cluster and a foreground cluster; acquiring a first set of 2D images by projecting clusters visible from a first set of viewpoints comprising at least two viewpoints according to first projection parameters, the first set of viewpoints being encompassed by a first visibility box defined on the 3D scene; acquiring a second set of 2D images by projecting the clusters visible from a second set of viewpoints according to second projection parameters, the second set of viewpoints being encompassed by a second view box defined in the 3D scene that is different from the first view box; determining a 2D common image corresponding to a cluster common to the first and second view boxes; encoding the first set of 2D images other than the 2D common image and the first projection parameters into a first data stream; encoding each 2D image of the second set of 2D images other than the 2D common image and the second projection parameters into a separate set of data streams; encoding said 2D common image only once in a data stream; A method comprising:

2. The method of claim 1 , wherein the clustering is further based on meanings associated with points of the 3D scene, on colors of the points of the 3D scene, or on movements of points of the 3D scene.

3. and encoding metadata, the metadata comprising: a list of visibility boxes defined in the 3D scene and a list of common clusters for the 3D scene; for a view box, a list of a set of clusters representing the 3D scene relative to the view box, and for each set of clusters associated with the view box, an identifier of the common cluster and the list of clusters other than the common cluster; 3. The method of claim 1 or 2, comprising:

4. A method for decoding a 3D scene of points clustered into a plurality of clusters according to depth ranges of the points in the 3D scene, the plurality of clusters including at least a background cluster and a foreground cluster, the method comprising: decoding a first set of at least two first 2D images other than a 2D common image from a first data stream, the first 2D images representing projections according to first projection parameters of at least one cluster of points of the 3D scene that are visible from a first set of viewpoints encompassed by a first view box defined for the 3D scene; decoding a second 2D image other than the 2D common image from each data stream of the set of separate data streams, the second 2D image representing a projection according to second projection parameters of at least one cluster of points in the 3D scene that is visible from a second set of viewpoints encompassed in a second view box defined for the 3D scene that is different from the first view box; decoding a 2D common image corresponding to a common cluster for the first and second view boxes only once; backprojecting pixels of the first 2D image according to first projection parameters and to a first set of viewpoints, backprojecting pixels of the second 2D image according to second projection parameters and to a second set of viewpoints, and backprojecting pixels of the 2D common image; A method comprising:

5. Obtaining metadata, the metadata comprising: a list of visibility boxes defined in the 3D scene and a list of common clusters for the 3D scene; for a view box, a list of a set of clusters representing the 3D scene relative to the view box, and for each set of clusters associated with the view box, an identifier of the common cluster and the list of clusters other than the common cluster; and decoding a 2D image from a data stream comprising clusters of 3D points that are visible from a current viewpoint; The method of claim 4 further comprising:

6. 1. A device for encoding a 3D scene, comprising a memory associated with a processor, the processor comprising: clustering the points of the 3D scene into a plurality of clusters according to depth ranges of the points in the 3D scene, the plurality of clusters including at least a background cluster and a foreground cluster; acquiring a first set of 2D images by projecting clusters visible from a first set of viewpoints comprising at least two viewpoints according to first projection parameters, the first set of viewpoints being encompassed by a first visibility box defined on the 3D scene; acquiring a second set of 2D images by projecting the clusters visible from a second set of viewpoints according to second projection parameters, the second set of viewpoints being encompassed by a second view box defined in the 3D scene that is different from the first view box; determining a 2D common image corresponding to a cluster common to the first and second view boxes; encoding the first set of 2D images other than the 2D common image and the first projection parameters into a first data stream; encoding each 2D image of the second set of 2D images other than the 2D common image and the second projection parameters into a separate set of data streams; encoding said 2D common image only once in a data stream; A device that is configured to:

7. The device of claim 6 , wherein the clustering is further based on meanings associated with points of the 3D scene, on colors of the points of the 3D scene, or on movements of points of the 3D scene.

8. The processor is further configured to encode metadata, the metadata comprising: a list of visibility boxes defined in the 3D scene and a list of common clusters for the 3D scene; for a view box, a list of a set of clusters representing the 3D scene relative to the view box, and for each set of clusters associated with the view box, an identifier of the common cluster and the list of clusters other than the common cluster; 8. The device of claim 6 or 7, comprising:

9. A device for decoding a 3D scene of points clustered into a plurality of clusters according to depth ranges of the points in the 3D scene, the plurality of clusters including at least a background cluster and a foreground cluster, the device comprising: a memory associated with a processor; and the processor: decoding a first set of at least two first 2D images other than a 2D common image from a first data stream, the first 2D images representing projections according to first projection parameters of at least one cluster of points of the 3D scene that are visible from a first set of viewpoints encompassed by a first view box defined for the 3D scene; decoding a second 2D image other than the 2D common image from each data stream of the set of separate data streams, the second 2D image representing a projection according to second projection parameters of at least one cluster of points in the 3D scene that is visible from a second set of viewpoints encompassed in a second view box defined for the 3D scene that is different from the first view box; decoding a 2D common image corresponding to a common cluster for the first and second view boxes only once; backprojecting pixels of the first 2D image according to first projection parameters and to a first set of viewpoints, backprojecting pixels of the second 2D image according to second projection parameters and to a second set of viewpoints, and backprojecting pixels of the 2D common image; A device that is configured to:

10. the processor: Obtaining metadata, the metadata comprising: a list of visibility boxes defined in the 3D scene and a list of common clusters for the 3D scene; for a view box, a list of a set of clusters representing the 3D scene relative to the view box, and for each set of clusters associated with the view box, an identifier of the common cluster and the list of clusters other than the common cluster; and decoding a 2D image from a data stream comprising clusters of 3D points that are visible from a current viewpoint; The device of claim 9 , further configured to:

Citation Information

Patent Citations

  • Method, apparatus and stream for volumetric video format

    EP3562159A1

  • Image processing apparatus, image processing method, program, and storage medium

    JP2013257843A

  • Image processing device and image processing method

    WO2017082079A1