Volumetric video supporting light effects
The method addresses the limitations of 3DoF+ video formats by encoding 3D scenes with depth, color, and reflectance atlases, ensuring accurate rendering of complex lighting effects and improving user immersion.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- INTERDIGITALCE PATENT HLDG SAS
- Filing Date
- 2026-01-21
- Publication Date
- 2026-04-21
AI Technical Summary
Existing 3DoF+ video formats fail to handle specular reflections and other complex lighting effects, leading to issues like duplicate reflections and inadequate rendering of 3D scenes from different viewpoints.
A method for encoding 3D scenes using depth, color, and reflectance atlases, along with bidirectional reflectance distribution function (BRDF) information, to enable accurate rendering of complex lighting effects and specular reflections.
Enables realistic rendering of 3D scenes with complex lighting effects by correcting reflections and ensuring consistent visual feedback during head translation, enhancing immersion and reducing dizziness.
Smart Images

Figure 2026067932000001_ABST
Abstract
Description
Technical Field
[0001] This principle generally relates to the domain of three-dimensional (3D) scenes and volumetric video content. This document is also understood in the context of the encoding, formatting, and decoding of data representing the texture and geometric shape of 3D scenes for the rendering of volumetric content on end-user devices such as mobile devices or head-mounted displays (HMDs). In particular, this document relates to the encoding of volumetric scenes in a way that enables rendering that can handle specular reflections and other complex lighting effects from various viewpoints.
Background Art
[0002] This section is intended to introduce the reader to various technical aspects that may be related to various aspects of the principle described and / or claimed hereinafter. This discussion is thought to be useful in providing the reader with background information to facilitate a better understanding of the various aspects of the principle. Therefore, it should be understood that these descriptions are to be read from this perspective and should not be read as an admission of prior art.
[0003] In recent years, there has been a growth in the availability of large-view content (up to 360°). Such content may not be fully visible to a user viewing the content on an immersive display device such as a head-mounted display, smart glasses, a PC screen, a tablet, or a smartphone. This means that at a given moment, the user can only view a portion of the content. However, the user can typically navigate within the content by various means such as head movement, mouse movement, touch screen, voice, etc. Typically, it is desirable to encode and decode this content.
[0004] Immersive video, also known as 360° flat video, allows users to view everything around them by rotating their heads around a stationary point. This rotation only enables a 3-degrees-of-freedom (3DoF) experience. For example, even if 3DoF video is sufficient for a first-generation omnidirectional video experience using a head-mounted display device (HMD), 3DoF video can be instantly frustrating for viewers who expect more degrees of freedom, such as by experiencing parallax. Furthermore, 3DoF can also induce dizziness because, in addition to rotating the head, it also involves translating the head in three directions, which is not reproduced in the 3DoF video experience.
[0005] Large field-of-view content can include, among other things, three-dimensional computer graphic imagery scenes (3D CGI scenes), point clouds, or immersive videos. Many terms can be used to design such immersive videos, for example, virtual reality (VR), 360, panoramic, 4π steradian, immersive, omnidirectional, or large field of view.
[0006] Volumetric video (also known as 6-degrees-of-freedom (6DoF) video) is an alternative to 3DoF video. When viewing 6DoF video, in addition to rotation, the user can also translate their head, and even their body, within the viewed content, experiencing parallax and even volumetricity. Such video significantly increases the sense of immersion and the perception of scene depth, and prevents dizziness by providing consistent visual feedback during head translation. The content is created by means of dedicated sensors that enable simultaneous recording of the color and depth of the scene of interest. The use of a color camera rig combined with photogrammetry techniques is a way to perform such recordings, even if technical difficulties remain.
[0007] 3DoF video contains a series of images resulting from the unmapping of a textured image (e.g., a spherical image encoded according to latitude / longitude projection mapping or equirectangular projection mapping), while 6DoF video frames embed information from several viewpoints. These can be viewed as a temporal series of points resulting from three-dimensional capture. Depending on the viewing conditions, two types of volumetric video can be considered. The first (i.e., full 6DoF) allows for complete free navigation within the video content, while the second (also known as 3DoF+) restricts the user's viewing space to a limited volume called the viewing bounding box, allowing for a limited volume of head and parallax experience. This second context represents a valuable trade-off between free navigation and passive viewing conditions for seated audience members.
[0008] In such videos, the viewport image the user sees is a composite field of view, i.e., a field of view of the scene not captured by the camera. Existing 3DoF+ video formats cannot handle specular reflection and other complex lighting effects, and assume that the 3D scene consists of Lambertian planes (i.e., only diffuse reflection). However, when specular reflection is captured by one camera of the acquisition rig, rendering the 3D scene from different virtual viewpoints, as observed from the viewpoint of this camera, requires correcting the position and appearance of the reflected content according to the new viewpoint. Furthermore, since the rendered virtual view is generated by mixing patches resulting from several input views, each input view captures a given reflection at a different position in the frame. A duplicate of the reflected object can be observed at rendering time. Therefore, there is a lack of 3DoF+ video formats that support complex lighting effects at rendering time. [Overview of the Initiative]
[0009] The following is a simplified overview of the Principle to provide a basic understanding of some aspects of it. This overview is not a comprehensive overview of the Principle. It is not intended to identify any important or significant elements of the Principle. The following overview merely presents some aspects of the Principle in a simplified form as a prelude to the more detailed explanation provided below.
[0010] This principle relates to a method for encoding 3D scenes. - For the 3D scene portion, obtain the first color patch, reflectance patch, and first depth patch, - Obtain a second color patch and a second depth patch for the portion outside the 3D scene that is reflected in at least one portion of the 3D scene, -Generating a depth atlas by packing the first and second depth patches, -Generating a color atlas by packing a second color patch with a subset of the first color patch, -Generating a reflectance atlas by packing a subset of reflectance patches, - For each reflectance patch packed in the reflectance atlas, To generate first information that encodes the parameters of the bidirectional reflectance distribution function model of light reflection on a reflectance patch, and To generate second information showing a list of color patches reflected by the reflectance patch, -In the data stream, • Depth Atlas, Color Atlas, Reflectance Atlas, and the First and This includes encoding a second piece of information within a data stream.
[0011] In the first embodiment, the subset of first color patches packed in the color atlas is empty, and the subset of reflectance patches packed in the reflectance atlas contains all reflectance patches. In the second embodiment, the subset of first color patches packed in the color atlas corresponds to the Lambertian portion of the 3D scene, and the subset of reflectance patches packed in the reflectance atlas corresponds to the non-diffuse reflective portion of the 3D scene. In the third embodiment, the subset of first color patches packed in the color atlas contains all first color patches, and the subset of reflectance patches packed in the reflectance atlas corresponds to the non-diffuse reflective portion of the 3D scene. In a modified example, the method further includes generating a surface normal atlas by packing surface normal patches corresponding to subsets of reflectance patches in the reflectance atlas.
[0012] The principle also relates to a device having a processor associated with memory, wherein the processor is configured to perform the method described above.
[0013] This principle also applies to a data stream that encodes a 3D scene, - A depth atlas that packs a first depth patch corresponding to a portion of the 3D scene and a second depth patch corresponding to a portion outside the 3D scene that is reflected in at least one portion of the 3D scene, - A color atlas that packs a first color patch corresponding to a portion of the 3D scene and a second color patch corresponding to a portion outside the 3D scene that is reflected in at least one portion of the 3D scene, - A reflectance atlas that packs reflectance patches corresponding to parts of the 3D scene, - For each reflectance patch packed in the reflectance atlas, • First information that encodes the parameters of the bidirectional reflectance distribution function model of light reflection on reflectance patches, and A data stream containing second information showing a list of color patches reflected by the reflectance patch.
[0014] This principle also relates to methods for rendering 3D scenes. This method is From the data stream, - A depth atlas that packs a first depth patch corresponding to a portion of the 3D scene and a second depth patch corresponding to a portion outside the 3D scene that is reflected in at least one portion of the 3D scene, - A color atlas that packs a first color patch corresponding to a portion of the 3D scene and a second color patch corresponding to a portion outside the 3D scene that is reflected in at least one portion of the 3D scene, - A reflectance atlas that packs reflectance patches corresponding to parts of the 3D scene, - Information signaling the rendering mode determined according to the first color patch and reflectance patch, - For each reflectance patch packed in the reflectance atlas, • First information that encodes the parameters of the bidirectional reflectance distribution function model of light reflection on reflectance patches, and • Decoding a second piece of information that shows a list of color patches reflected by the reflectance patch, The method includes rendering a 3D scene by back-projecting the first and second color patches according to the first and second depth patches, and by using ray tracing for the reflectance patch according to the first and second information and the associated color patches. [Brief explanation of the drawing]
[0015] This disclosure will be better understood by reading the following description, which will reveal other specific features and advantages, and this specification will refer to the attached drawings. [Figure 1] A non-limiting embodiment of this principle is shown, illustrating a three-dimensional (3D) model of points in a point cloud corresponding to an object and a 3D model. [Figure 2] This document presents non-limiting examples of encoding, transmitting, and decoding data representing a sequence of 3D scenes, based on non-limiting embodiments of the present principle. [Figure 3]An exemplary architecture of a device configured to implement the method described in connection with FIGS. 13 and 14, according to a non-limiting embodiment of the present principle, is shown. [Figure 4] An example of an embodiment of the syntax of a stream when data is transmitted via a packet-based transmission protocol, according to a non-limiting embodiment of the present principle, is shown. [Figure 5] A patch atlas approach with an example of four projection centers, according to a non-limiting embodiment of the present principle, is shown. [Figure 6] An example of an atlas including texture information of points of a 3D scene, according to a non-limiting embodiment of the present principle, is shown. [Figure 7] An example of an atlas including depth information of points of the 3D scene of FIG. 6, according to a non-limiting embodiment of the present principle, is shown. [Figure 8] Two of the views of the 3D scene captured by the camera array are shown. [Figure 9] A simple scene to be captured is shown. [Figure 10] A first example of encoding the 3D scene of FIG. 9 in a depth atlas, a reflectance atlas, and a color atlas, according to the first embodiment of the present principle, is shown. [Figure 11] A second example of encoding the 3D scene of FIG. 9 in a depth atlas, a reflectance atlas, and a color atlas, according to the second embodiment of the present principle, is shown. [Figure 12] A third example of encoding the 3D scene of FIG. 9 in a depth atlas, a reflectance atlas, and a color atlas, according to the third embodiment of the present principle, is shown. [Figure 13] A method for encoding a 3D scene using complex lighting effects is illustrated. [Figure 14] A method for rendering a 3D scene using complex lighting effects is illustrated.
Embodiments for Carrying Out the Invention
[0016] The principle is fully described below with reference to the accompanying drawings, which illustrate examples of the principle. However, the principle can be embodied in many alternative forms and should not be construed as being limited to the embodiments described herein. Thus, the principle is open to various modifications and alternative forms, specific examples of which are shown as examples in the drawings and described in detail herein. However, it should be understood that there is no intention to limit the principle to any particular form disclosed, on the contrary, this disclosure covers all modifications, equivalents, and alternatives that fall within the spirit and scope of the principle as defined by the claims.
[0017] The terms used herein are for the purpose of illustrating only specific embodiments and are not intended to limit the principles herein. Where used herein, the singular forms “a,” “an,” and “the” are intended to include the plural form unless the context otherwise explicitly indicates. Where used herein, the terms “comprises,” “comprising,” “includes,” and / or “including” specify the presence of the described features, integers, steps, actions, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, actions, elements, components, and / or groups thereof. Furthermore, where an element is referred to as “responding to” or “connecting to” another element, it may directly respond to or be able to connect to the other element, or an intervening element may exist. In contrast, where an element is referred to as “directly responding to” or “directly connecting to” another element, there is no intervening element. As used herein, the term "and / or" includes any and all combinations of one or more of the associated enumerated items, and may be abbreviated as " / ".
[0018] In this specification, terms such as "first," "second," etc., may be used to describe various elements, but it will be understood that these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, the first element may be called the second element, and similarly, the second element may be called the first element without deviating from the teachings of this principle.
[0019] Some parts of the diagram include arrows on the communication path to indicate the main direction of communication, but please understand that communication may occur in the opposite direction to the depicted arrows.
[0020] Some examples are described with respect to block diagrams and operation flowcharts that represent parts of circuit elements, modules, or code, where each block contains one or more executable instructions for implementing a specified logical function. Note that in other implementations, the functions described in the blocks may occur in the order they are described. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or the blocks may be executed in reverse order depending on the functions involved.
[0021] In this specification, “by example” or “in one example” means that certain features, structures, or properties described in connection with this embodiment may be included in at least one implementation of the principle. The appearance of the phrase “by example” or “in one example” in various places in this specification does not necessarily refer to the same example, and separate or alternative embodiments are not necessarily mutually exclusive with other embodiments.
[0022] Reference numerals appearing in the claims are for illustrative purposes only and shall not limit the scope of the claims. Unless expressly stated otherwise, these embodiments and modifications may be used in any combination or partial combination.
[0023] Figure 1 shows a three-dimensional (3D) model 10 of the points of a point cloud 11 corresponding to an object and a 3D model 10. The 3D model 10 and the point cloud 11 may correspond to a potential 3D representation of an object in a 3D scene, for example, including other objects. Model 10 may be a 3D mesh representation, and the points of the point cloud 11 may be the vertices of the mesh. The points of the point cloud 11 may also be points spread out on the surface of a face of the mesh. Model 10 may also be represented as a splatted version of the point cloud 11, and the surface of Model 10 is created by splatting the points of the point cloud 11. Model 10 may be represented by many different representations, such as voxels or splines. Figure 1 illustrates the fact that a point cloud can be defined as a surface representation of a 3D object, and that a surface representation of a 3D object can be generated from the points of the cloud. As used herein, the projection point of a 3D object on an image (by the elongation point of the 3D scene) is equivalent to projecting any representation of this 3D object, such as a point cloud, mesh, spline model, or voxel model.
[0024] A point cloud can be represented in memory, for example, as a vector-based structure, where each point has its own coordinates within the viewpoint's reference frame (e.g., three-dimensional coordinates XYZ, or solid angle and distance (also called depth) from / to the viewpoint) and one or more attributes called components. Examples of components are color components that can be expressed in various color spaces, e.g., RGB (red, green, and blue) or YUV (where Y is the luminance component and UV are the two color difference components). A point cloud is a representation of a 3D scene containing objects. A 3D scene can be viewed from a given viewpoint or range of viewpoints. Point clouds can be generated in many ways, for example, From the capture of real objects captured by a camera rig, optionally complemented by depth-active sensing devices, From capturing virtual / composite objects captured by a virtual camera rig in a modeling tool, It can be obtained from a mixture of both real and virtual objects.
[0025] A 3D scene corresponds to a captured scene that is part of a real (or virtual) scene. Firstly, some parts or scenes to be captured are not visible (hidden) from all cameras. These parts are outside the 3D scene. Secondly, the field of view of the camera rig may be less than 360°. In that case, parts of the real scene remain outside the captured 3D scene. Nevertheless, some parts outside the 3D scene may be reflected into parts of the 3D scene.
[0026] Figure 2 shows a non-limiting example of encoding, transmission, and decoding of data representing a sequence of 3D scenes. For example, an encoding format that can simultaneously accommodate 3DoF, 3DoF+, and 6DoF decoding.
[0027] A sequence of 3D scenes 20 is acquired. When a sequence of photos is a 2D video, a sequence of 3D scenes is a 3D (also called volumetric) video. The sequence of 3D scenes can be provided to a volumetric video rendering device for 3DoF, 3DoF+, or 6DoF rendering and display.
[0028] A sequence of 3D scenes 20 is provided to encoder 21. Encoder 21 takes one 3D scene or a sequence of 3D scenes as input and provides a bitstream representing the input. The bitstream may be stored in memory 22 and / or on an electronic data medium and transmitted over network 22. The bitstream representing the sequence of 3D scenes may be read from memory 22 and / or received from network 22 by decoder 23. Decoder 23 receives the bitstream as input and provides, for example, the sequence of 3D scenes in point cloud format.
[0029] The encoder 21 may comprise several circuits that implement several steps. In the first step, the encoder 21 projects each 3D scene onto at least one 2D image. 3D projection is any method of mapping three-dimensional points onto a two-dimensional plane. The use of this type of projection is widespread, especially in computer graphics, manipulation, and drafting, as modern methods for displaying graphic data are based on a two-dimensional medium (pixel information from some bit plane). The projection circuit 211 provides at least one 2D frame 2111 for the 3D scenes of the sequence of 3D scenes 20. The frame 2111 contains depth information representing the 3D scene projected onto the frame 2111. In variations, the frame 2111 may contain other attributes. According to this principle, the projected attributes may represent the texture (i.e., color attributes), heat, reflectivity, or other attributes of the 3D scene projected onto the frame. In the modified version, the information is encoded in separate frames, for example, in two separate frames 2111 and 2112, or in one frame for each attribute.
[0030] Metadata 212 is used and updated by the projection circuit 211. Metadata 212 includes information about the projection operation (e.g., projection parameters) and how color and depth information are organized within frames 2111 and 2112, as described in relation to Figures 5 to 7.
[0031] The video encoding circuit 213 encodes the sequence of frames 2111 and 2112 as video. The images (or sequences of images of the 3D scenes) of 3D scenes 2111 and 2112 are encoded in the stream by the video encoder 213. The video data and metadata 212 are then encapsulated in the data stream by the data encapsulation circuit 214.
[0032] Encoder 213 is, for example, -JPEG, specification ISO / CEI10918-1UIT-T recommended T.81, https: / / www.itu.int / rec / T-REC-T.81 / en; - Complies with encoders such as AVC, also known as MPEG-4AVC or h264. UIT-TH.264 and ISO / CEI MPEG-4-Part 10 (ISO / CEI14496-10), http: / / www.itu.int / rec / T-REC-H.264 / en, HEVC (its specifications can be found on the ITU website, T recommended, H series, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en), -3D-HEVC (HEVC file extension found on the ITU website, T recommended, H series, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en annex G and I), - VP9 developed by Google, or AV1 (AO Media Video 1), developed by the Alliance for Open Media.
[0033] The data stream is stored by the decoder 23 in memory accessible, for example, via the network 22. The decoder 23 comprises different circuits that implement different steps of decoding. The decoder 23 takes the data stream generated by the encoder 21 as input and provides a sequence of 3D scenes 24 that are rendered and displayed by a volumetric video display device such as a head-mounted device (HMD). The decoder 23 obtains the stream from the source 22. For example, source 22 is -For example, local memory such as video memory or RAM (or random access memory), flash memory, ROM (or read-only memory), and hard disk, - For example, storage interfaces such as mass storage, RAM, flash memory, ROM, optical disks, or magnetic support interfaces, - For example, a communication interface such as a wired interface (e.g., bus interface, wide area network interface, local area network interface) or a wireless interface (IEEE 802.11 interface or Bluetooth® interface, etc.), - Belongs to a set that includes user interfaces such as graphical user interfaces that allow users to input data.
[0034] Decoder 23 includes circuit 234 for extracting data encoded in a data stream. Circuit 234 takes the data stream as input and provides metadata 232 corresponding to the stream and metadata 212 encoded in the two-dimensional video. The video is decoded by video decoder 233, which provides a sequence of frames. The decoded frames include color and depth information. In a modified example, video decoder 233 provides a sequence of two frames, one containing color information and the other containing depth information. Circuit 231 uses metadata 232 to provide a sequence of 3D scenes 24 without projecting the color and depth information from the decoded frames. The sequence of 3D scenes 24 corresponds to a sequence of 3D scenes 20 and video compression, which have potentially reduced accuracy associated with encoding as 2D video.
[0035] In rendering, the viewport image seen by the user is a composite view, i.e., a view of the scene not captured by the camera. If specular reflection is captured by one camera of the acquisition rig as observed from the camera's viewpoint, rendering a 3D scene from a different virtual viewpoint requires correcting the position and appearance of the reflected content according to the new viewpoint. According to this principle, information for rendering complex lighting effects is carried in the data stream.
[0036] Figure 3 shows an exemplary architecture of device 30 that may be configured to implement the methods described in relation to Figures 13 and 14. The encoder 21 and / or decoder 23 of Figure 2 may implement this architecture. Alternatively, the encoder 21 and / or decoder 23 circuits may be connected together, for example, via their bus 31 and / or via the I / O interface 36, in a device according to the architecture of Figure 3.
[0037] Device 30 consists of the following elements, which are linked together by the data and address bus 31: -For example, a microprocessor 32 (or CPU) which is a DSP (Digital Signal Processor), -ROM (Read Only Memory) 33, -RAM (Random Access Memory) 34 and, -Storage interface 35, - An I / O interface 36 for receiving data to be sent from the application, - Equipped with a power source, such as a battery.
[0038] For example, the power supply is external to the device. In each of the memories mentioned herein, the term “register” as used herein may refer to a small area of capacity (a few bits) or a very large area (e.g., an entire program or a large amount of received or decoded data). ROM33 contains at least a program and parameters. ROM33 can store algorithms and instructions for performing the technology according to this principle. When switched on, CPU32 uploads the program in RAM and executes the corresponding instructions.
[0039] RAM34 contains, within its registers, a program executed by CPU32 and uploaded after device30 is switched on, input data within the registers, intermediate data for different states of methods within the registers, and other variables used for executing methods within the registers.
[0040] The implementations described herein may be implemented, for example, in methods or processes, apparatus, computer program products, data streams, or signals. Even if considered only in the context of a single implementation (e.g., considered only as a method or device), the implementations of the considered features may also be implemented in other forms (e.g., programs). Apparatus may be implemented, for example, with appropriate hardware, software, and firmware. The method may be implemented in apparatus such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices such as computers, mobile phones, and personal digital assistants ("PDAs") that facilitate the communication of information between end users.
[0041] According to the embodiment, device 30 is configured to implement the method described in relation to Figures 13 and 14, - Mobile devices and, -Communication devices and, -Game devices and, - A tablet (or tablet computer), -Laptop and, -Still camera and, -Video camera and, - Encoding chip, - Belongs to a set that includes a server (e.g., a broadcast server, a video-on-demand server, or a web server).
[0042] Figure 4 shows an example of an embodiment of the stream syntax when data is transmitted via a packet-based transmission protocol. Figure 4 shows an exemplary structure 4 of a volumetric video stream. The structure consists of a container that organizes the stream in independent elements of syntax. The structure may include a header portion 41, which is a set of data common to all syntax elements of the stream. For example, the header portion includes some metadata about the syntax elements, describing the nature and role of each of them. The header portion may also include some of the metadata 212 in Figure 2, for example, the coordinates of the central viewpoint used to project points of the 3D scene onto frames 2111 and 2112. The structure includes elements of syntax 42 and a payload containing at least one element of syntax 43. Syntax elements 42 include data representing color and depth frames. The images may be compressed according to a video compression method.
[0043] The elements of syntax 43 are part of the payload of the data stream and may contain metadata about how the frames of the elements of syntax 42 are encoded, for example, parameters used to project or pack points of a 3D scene onto the frame. Such metadata may be associated with each frame of the video or with a group of frames (also known as a Group of Pictures (GoP) in video compression standards).
[0044] Figure 5 shows a patch atlas approach with an example of four projection centers. The 3D scene 50 includes features. For example, projection center 51 is a perspective projection camera, and camera 53 is an orthographic projection camera. The cameras may also be omnidirectional cameras with, for example, spherical mapping (e.g., equirectangular projection mapping) or cubic mapping. The 3D points of the 3D scene are projected onto a 2D plane associated with a virtual camera located at the projection center, according to the projection behavior described in the projection data of the metadata. In the example in Figure 5, the projection of points captured by camera 51 is mapped onto patch 52 according to perspective mapping, and the projection of points captured by camera 53 is mapped onto patch 54 according to orthogonal mapping.
[0045] Clustering of projected pixels yields numerous 2D patches, which are packed into a rectangular atlas 55. The organization of the patches within the atlas defines the atlas layout. In one embodiment, there are two atlases having identical layouts: one for texture (i.e., color) information and one for depth information. Two patches captured by the same camera or two separate cameras may contain information representing the same part of a 3D scene, such as patches 54 and 56.
[0046] The packing operation generates patch data for each patch produced. The patch data includes references to projection data (e.g., an index in a table of projection data or a pointer to projection data (address in memory or data stream)) and information describing the location and size of the patch within the atlas (e.g., the coordinates of the top-left corner of a pixel, its size, and its width). Patch data entries are added to metadata that is associated with the compressed data of one or two atlases and encapsulated within the data stream.
[0047] Figure 6 shows an example of an atlas 60 containing attribute information, such as texture (also called color) information (e.g., RGB data or YUV data) of points in a 3D scene, according to a non-limiting embodiment of the present principle. As described in relation to Figure 5, the atlas is an image packing patch, and the patch is a photograph obtained by projecting a portion of the points in a 3D scene.
[0048] In the example in Figure 6, the atlas 60 includes a first part 61 containing texture information for points in the 3D scene visible from the viewpoint and one or more second parts 62. The texture information for the first parts 61 may be obtained, for example, according to equirectangular projection mapping, which is an example of spherical projection mapping. In the example in Figure 6, the second parts 62 are positioned at the left and right boundaries of the first part 61, although the second parts may be positioned differently. The second parts 62 contain texture information for parts of the 3D scene that are complementary to the parts visible from the viewpoint. The second parts can be obtained by removing the points visible from the first viewpoint (textures stored in the first part) from the 3D scene and by projecting the remaining points according to the same viewpoint. The latter process may be repeated iteratively so that the hidden parts of the 3D scene are obtained at each point in time. In a modified version, the second part can be obtained by removing points visible from a viewpoint, for example, a central viewpoint (a texture stored in the first part), from the 3D scene, and by projecting the remaining points from one or more second viewpoints, according to a viewpoint different from the first viewpoint, for example, a view space centered on the central viewpoint (e.g., the viewing space of a 3DoF rendering).
[0049] The first part 61 can be viewed as the first large texture patch (corresponding to the first part of the 3D scene), and the second part 62 contains a smaller texture patch (corresponding to the second part of the 3D scene, which is complementary to the first part). Such an atlas has the advantage of being compatible with both 3DoF rendering and 3DoF+ / 6DoF rendering simultaneously (when rendering only the first part 61).
[0050] Figure 7 shows an example of an atlas 70 containing depth information for points in the 3D scene of Figure 6, according to a non-limiting embodiment of the present principle. The atlas 70 can be viewed as a depth image corresponding to the texture image 60 in Figure 6.
[0051] Atlas 70 includes a first part 71 containing depth information for points in a 3D scene as seen from a central viewpoint, and one or more second parts 72. Atlas 70 can be obtained in the same way as Atlas 60, but instead of texture information, it includes depth information associated with points in a 3D scene.
[0052] In 3DoF rendering of a 3D scene, only one viewpoint, typically a central viewpoint, is considered. The user can rotate their head with three degrees of freedom around this primary viewpoint to view different parts of the 3D scene, but the user cannot move this intrinsic viewpoint. The points of the scene to be encoded are those visible from this intrinsic view, and only their texture information is needed for encoding / decoding for 3DoF rendering. There is no need to encode points of the scene that are not visible from this intrinsic viewpoint for 3DoF rendering when the user cannot access them.
[0053] In 6DoF rendering, the user can move their viewpoint throughout the scene. In this case, since every point is potentially accessible to the user who can move their viewpoint, every point in the scene (depth and texture) must be encoded within the bitstream. At the encoding stage, there is no way to know in advance from which viewpoint the user will be viewing the 3D scene.
[0054] Regarding 3DoF+ rendering, the user can move their viewpoint within a limited space around a central viewpoint. This allows them to experience parallax. Data representing a portion of the scene visible from any point in the view space should be encoded into a stream containing data representing the 3D scene as seen according to the central viewpoint (i.e., the first parts 61 and 71). The size and shape of the view space can be determined and encoded in the bitstream, for example, during the encoding step. The decoder can retrieve this information from the bitstream, and the renderer restricts the view space to the space determined by the retrieved information. In another example, the renderer determines the view space according to hardware constraints, for example, in relation to the ability of sensors to detect user movement. In such a case, if a point visible from a point in the renderer's view space is not encoded in the bitstream during the encoding stage, this point will not be rendered. In a further example, data representing all points in the 3D scene (e.g., textures and / or geometric shapes) is encoded in the stream without considering the rendering space of the view. To optimize the stream size, only a subset of the scene's points can be encoded, for example, a subset of points that can be seen according to the view's rendering space.
[0055] According to this principle, a volumetric video transmission format is proposed. This format includes signaling of non-Lambertian patches along with their light reflection properties to enable a ray-tracing based rendering engine to synthesize a visually realistic virtual view with respect to light effects.
[0056] The syntax of the format based on this principle includes the following: -For each non-Lambert patch: Reflectance attribute of patch sample, The light reflectivity characteristics (bidirectional reflectance distribution function) of the patch material, and A list of other patches reflected within the current patch. - Reflected patches found from the scene view frustum are considered as light sources, along with their geometry and texture components. - Parameters of other time-sensitive or diffuse light sources.
[0057] While existing rendering engines can render such described 3D scenes, retro-compatible embodiments that do not employ advanced lighting effects are also described.
[0058] Figure 8 shows two views of the 3D scene captured by the camera array. View 811 is a top-down view of the scene and is to the left of view 835. The 3D scene includes reflective objects 81 and 82 (the oven door reflects a giant spider on the floor). Views 811 and 835 contain information corresponding to the same points in the 3D scene. However, due to the lighting of the scene and different acquisition positions, the color information associated with these points may differ from view to view. View 811 also contains information about points in the 3D scene that cannot be seen from the viewpoint of view 835, and vice versa.
[0059] To assist in stitching during rendering, at least one atlas is generated to encode the 3D scene from the captured multi-view + depth (MVD) image by removing redundant information and preserving some overlap between the removed regions of 3D space. The atlas is assumed to be sufficient to reconstruct / composit any viewport image from any viewpoint in the 3DoF+ viewing space that the user can navigate. To do so, a compositing process is performed to stitch all the patches from the atlas together to restore the desired viewport image. However, this stitching step can be subject to strong artifacts when the scene represented in the atlas contains specular / reflective or transmissive components, as illustrated in Figure 8. Such light effects are dependent on the viewing position, and therefore the perceived color of the spatial portion in question can change from one viewpoint to another.
[0060] Figure 9 shows a simplified captured scene for illustrative purposes. This scene consists of two planes ("wall" and "floor") with diffuse reflection and one non-plane 91 ("mirror") with both specular and diffuse reflection properties. Two objects 93 located outside the camera 92's viewing frustum (i.e., outside the captured 3D scene) are reflected by the mirror 91.
[0061] Figure 10 shows an example of encoding the 3D scene of Figure 9 in a depth atlas 100a, a reflectance atlas 100b, and a color atlas 100c according to a first embodiment of the principle. The parts of the 3D scene and the parts outside the 3D scene that are reflected onto at least one part of the 3D scene are projected onto patches as described in relation to Figure 5. For each patch sample, depth values and different attribute values are obtained. According to the principle, depth patches, color patches, and reflectance patches are obtained for each of these parts.
[0062] In a first embodiment of this principle, the depth atlas 100a is generated by packing all depth patches 101a to 107a (i.e., patches 101a to 105a obtained by projecting portions of the captured 3D scene as described in relation to Figure 1, and patches 106a and 107a obtained by projecting portions of the captured 3D scene outside the scene that are reflected in at least one portion of the 3D scene). In the example of Figure 9, the mirror and the two objects reflected in the mirror are not planes. The corresponding depth patches 101a, 106a and 107a then store different depth values, represented by the gray gradient in Figure 10.
[0063] The color atlas 100c is generated by packing the color patches 106c and 107c of the parts outside the 3D scene that are reflected in at least one part of the 3D scene (in the example in Figure 9, the two objects reflected in a non-planar mirror).
[0064] The reflectance atlas 100b is generated by packing reflectance patches 101b-105b corresponding to projections of parts of the 3D scene. Reflectance attributes describing the spectral reflectance characteristics of the patch samples can be specified in three dimensions, for example, in the R, G, and B channels of the atlas frame. Reflectance patch 101b corresponding to the mirror in Figure 9 contains only the reflectance attributes of the projection of the point corresponding to the mirror. Therefore, the reflected object 93 is not visible in this patch. In all embodiments of this principle, each reflectance patch is associated with information representing a parameterized model that defines how light is reflected at its surface, also known as the bidirectional reflectance distribution function (BRDF). Several BRDF parametric models exist, among which the empirical Phong model is widely used in the art. The Phong model is defined by the following four parameters: ks, the reflectance of the specular term of incident light. kd is the reflectance of the diffuse term of incident light (Lambertian reflectance). • ka, the ambient reflectance that exists at all points in the rendered scene. ·α is the gloss constant of this material, and is larger for smoother and more mirror-like surfaces.
[0065] In rendering, deriving light reflection and incident light from the surface BRDF requires knowledge of the surface normals for each sample. Such normal values can be calculated from the depth map on the rendering side, or, in the modifications of all embodiments of this principle, an additional normal attribute patch atlas is sent along with the depth atlas, reflectance atlas, and color atlas. This modification represents a trade-off between bandwidth and computing resources on the rendering side.
[0066] In all embodiments of this principle, for each patch in the reflectance atlas, a list of color patches reflected by the current patch is added to the patch parameters (i.e., metadata associated with the patch). In the example in Figure 10, the parameters for reflectance patch 101a indicate that reflectance patches 106c and 107c in color atlas 100c are reflected by reflectance patch 101a. Without such information, the renderer would have to reconstruct and analyze the entire 3D scene geometry to extract this information.
[0067] Renderers based on ray tracing techniques utilize transmitted surface properties to synthesize realistic viewpoint-dependent lighting effects.
[0068] Figure 11 shows an example of encoding the 3D scene from Figure 9 in a depth atlas 100a, a reflectance atlas 110b, and a color atlas 110c according to a second embodiment of the present principle. The same depth, color, and reflectance patches are obtained for parts of the 3D scene and parts outside the 3D scene that are reflected to at least one part of the 3D scene. In the second embodiment, the depth atlas 100a is generated by packing each depth patch 101a to 107a.
[0069] The color atlas 110c is generated by packing color patches 102c to 105c, which correspond to the Lambertian portion (i.e., the non-reflective portion) of the 3D scene, with color patches 106c and 107c, which correspond to the portion outside the 3D scene that is reflected by at least one portion of the 3D scene.
[0070] The reflectance atlas 110b is generated by packing reflectance patches 101b that correspond to the reflective parts of the 3D scene (i.e., the non-Lambertian parts of the 3D scene). For each reflectance patch in the patch atlas 110b, BRDF information and a list of color patches reflected by the current patch are associated with the patch in the metadata.
[0071] In the modified version, a normal atlas that packs normal patches corresponding to the reflective parts of a 3D scene is associated with a depth atlas 100a, a reflectance atlas 110b, and a color atlas 110c.
[0072] Figure 12 shows an example of encoding the 3D scene of Figure 9 in depth atlas 100a, reflectance atlas 110b, and color atlas 120c according to a third embodiment of the present principle. The same depth, color, and reflectance patches are obtained for portions of the 3D scene and portions outside the 3D scene that are reflected to at least one portion of the 3D scene. In the second embodiment, the depth atlas 100a is generated by packing depth patches 101a to 107a.
[0073] The color atlas 120c is generated by packing color patches 101c-105c, which correspond to parts of the 3D scene (i.e., the Lambertian part and the reflective part), with color patches 106c and 107c, which correspond to parts outside the 3D scene that are reflected in at least one part of the 3D scene. In Figure 12, texture patch 101c, which carries the reflection as seen from the camera viewpoint, is packed into the color atlas and is only useful for retro-compatible renderers. In such rendering modes, only depth patches 101a-105a and color patches 101c-105c are decoded and supplied to the renderer.
[0074] The reflectance atlas 110b is generated by packing reflectance patches 101b that correspond to the reflective parts of the 3D scene (i.e., the non-Lambertian parts of the 3D scene). For each reflectance patch in the patch atlas 110b, BRDF information and a list of color patches reflected by the current patch are associated with the patch in the metadata.
[0075] In the modified version, a normal atlas that packs normal patches corresponding to the reflective parts of a 3D scene is associated with a depth atlas 100a, a reflectance atlas 110b, and a color atlas 120c.
[0076] Metadata is associated with the atlas that encodes the 3D scene. According to this principle, metadata allows for separate packing for each attribute (i.e., the position and orientation of patches within the atlas), and also allows for the possibility that patches may not always be present in every attribute atlas frame. Possible syntaxes for metadata can be based on the MIV standard syntax, as follows:
[0077] Atlas sequence parameters can be extended with bolded syntax elements.
[0078] [Table 1]
[0079] Patch data units may be extended with bolded elements.
[0080] [Table 2-1]
[0081] [Table 2-2]
[0082] Here, A pdu_light_source_flag[tileID][p] equal to 1 indicates that the patch with index p in the tile with ID tileID is a light source outside the scene's viewing frustum, exists within the texture atlas frame, and does not exist within the reflectance atlas frame.
[0083] A value of 1 for pdu_reflection_parameters_present_flag[tileID][p] indicates that the reflection model parameters exist within the syntax structure for the patch with index p in the tile with ID tileID, which are assumed to exist within the reflectance atlas frame.
[0084] pdu_reflection_model_id[tileID][p] specifies the ID of the reflection model for the patch with index p in the tile with ID tileID. pdu_reflection_model_id[tileID][p] equal to 1 indicates a Phong model.
[0085] pdu_specular_reflection_constant[tileID][p] specifies the specular reflection constant of the Phong model for the patch with index p in the tile with ID tileID.
[0086] pdu_diffuse_reflection_constant[tileID][p] specifies the diffuse reflectance constant of the Phong model for the patch with index p in the tile with ID tileID.
[0087] pdu_ambient_reflection_constant[tileID][p] specifies the ambient reflection constant of the Phong model for the patch with index p in the tile with ID tileID.
[0088] pdu_diffuse_reflection_constant[tileID][p] specifies the reflection constant of the Phong model for the patch with index p in the tile having ID tileID.
[0089] pdu_num_reflected_patches_minus1[tileID][p]+1 specifies the number of texture patches reflected by the patch with index p in the tile with ID tileID.
[0090] pdu_reflected_patch_idx[tileID][p]][i] specifies the index in the texture atlas frame of the i-th texture patch reflected within the patch with index p in the tile with ID tileID.
[0091] Alternatively, patch reflection properties can be interlocated to a set of "material reflection properties" (e.g., "metal", "wood", "grass", etc.), and the pdu_entity_id[tileID][p] syntax element can be used to associate each non-Lambert patch with a single material ID. In this case, the syntax elements related to the reflection model parameters are provided to the renderer via an external means (for each registered material), and only the list of reflected patches is signaled to the patch data unit MIV extension.
[0092] The common atlas sequence parameter set for MIV can be extended as follows:
[0093] [Table 3]
[0094] The `casme_miv_v1_rendering_compatible_flag` flag specifies that the atlas geometry and texture frames are compatible with rendering using the ISO / IEC 23090-12(1E) virtual rendering process. When `casme_MIV_v1_rendering_compatible_flag` is equal to 1, the bitstream compatibility requirement is that at least one subset of patches in the atlas geometry and texture frames is compatible for rendering using the ISO / IEC 23090-12(1E) virtual rendering process. If it is not present, the value of `casme_MIV_v1_rendering_compatible_flag` is inferred to be equal to 0.
[0095] Figure 13 illustrates a method 130 for encoding a 3D scene using complex lighting effects. In step 131, a first depth patch, a first color patch, and a reflectance patch are obtained by projecting a portion of the captured 3D scene. A second depth patch and a second color patch are also obtained by projecting a portion of the captured 3D scene that has been reflected to at least one portion of the 3D scene. In step 132, a depth atlas is generated by packing the first and second depth patches, and a color atlas is generated by packing the second color patch and a subset of the first color patch. According to the first embodiment, the subset of the first color patch packed in the color atlas is empty. In the second embodiment, the subset of the first color patch packed in the color atlas corresponds to the Lambertian portion of the 3D scene. In the third embodiment, the subset of the first color patch packed in the color atlas contains all of the first color patches. In step 133, the reflectance atlas is generated by packing a subset of reflectance patches. In the first embodiment, the subset of reflectance patches packed in the reflectance atlas includes all reflectance patches. In the second embodiment, the subset of reflectance patches packed in the reflectance atlas corresponds to the non-diffuse reflective portion of the 3D scene. In the third embodiment, the subset of reflectance patches packed in the reflectance atlas corresponds to the non-diffuse reflective portion of the 3D scene. In all embodiments, the reflectance atlas is associated with metadata for each reflectance patch packed in the reflectance atlas, which includes first information encoding the parameters of a bidirectional reflectance distribution function model of light reflection on the reflectance patch, and second information indicating a list of color patches reflected by the reflectance patch. In the optional step 134, the normal atlas is generated by packing normal patches corresponding to a subset of reflectance patches in the reflectance atlas. In step 135, the generated atlas and associated metadata are encoded in a data stream.
[0096] Figure 14 illustrates method 140 for rendering a 3D scene using complex lighting effects. In step 141, a data stream containing data representing the 3D scene is acquired. In step 142, a depth atlas packing depth patches and a color atlas packing color patches are decoded from the data stream. In step 143, a reflectance atlas packing reflectance patches is decoded from the data stream. Metadata associated with the reflectance atlas is also decoded. The metadata includes, for each reflectance patch packed in the reflectance atlas, first information encoding the parameters of the bidirectional reflectance distribution function model of light reflection on the reflectance patch, and second information indicating a list of color patches reflected by the reflectance patch. In an optional step 144, a normal atlas packing normal patches corresponding to a subset of reflectance patches in the reflectance atlas is decoded from the data stream.
[0097] In step 145, the pixels of the color patch are back-projected according to the pixels of the corresponding depth patch to extract points in the 3D scene. In step 146, the light effect is extracted by using ray tracing based on the reflectance patch and associated metadata, as well as the pixels of the depth patch and color patch enumerated in the metadata. In a modified example, a normal patch may be used to facilitate ray tracing.
[0098] The implementations described herein may be implemented, for example, in methods or processes, apparatus, computer program products, data streams, or signals. Even if considered only in the context of a single implementation (e.g., considered only as a method or device), the implementations of the considered features may also be implemented in other forms (e.g., programs). Apparatus may be implemented, for example, with appropriate hardware, software, and firmware. The method may be implemented in apparatus such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, smartphones, tablets, computers, mobile phones, portable / personal digital assistants ("personal digital assistants, PDAs"), and other devices that facilitate the communication of information between end users.
[0099] The various processes and features described herein may be embodied in a variety of different devices or applications, particularly in devices or applications associated with other processing of images and associated texture information and / or depth information, for example, data encoding, data decoding, view generation, texture processing, and related texture information and / or depth information. Examples of such devices include encoders, decoders, post-processors that process the output from decoders, pre-processors that provide input to encoders, video coders, video decoders, video codecs, web servers, set-top boxes, laptops, personal computers, mobile phones, PDAs, and other communication devices. As should be clear, the devices may be mobile and may be installed in mobile vehicles.
[0100] In addition, the method may be implemented by instructions executed by the processor, and such instructions (and / or data values produced by the implementation) may be stored on a processor-readable medium such as an integrated circuit, software carrier or other storage device, e.g., a hard disk, a compact diskette ("CD"), an optical disc (e.g., a DVD, often referred to as a digital multipurpose disc or digital video disc), random access memory ("RAM") or read-only memory ("ROM"). The instructions may form an application program explicitly embodied on the processor-readable medium. The instructions may be, for example, hardware, firmware, software, or a combination of the two. The instructions may be found, for example, in an operating system, a separate application, or a combination of the two. Thus, a processor may be characterized as both, for example, a device configured to execute a process and a device containing a processor-readable medium (such as a storage device) having instructions for executing the process. Furthermore, the processor-readable medium may store data values produced by the implementation in addition to, or instead of, instructions.
[0101] As will be apparent to those skilled in the art, the implementations can produce a variety of signals formatted to carry, for example, information that can be stored or transmitted. This information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry, as data, rules for writing or reading the syntax of a described embodiment, or actual syntax values written by a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The signal carried by the signal may be, for example, analog or digital information. The signal may be transmitted by a variety of different wired or wireless links, as is known. The signal may be stored in a processor-readable medium.
[0102] Many implementations are described. Nevertheless, it will be understood that various modifications are possible. For example, elements of different implementations can be combined, supplemented, modified, or deleted to create other implementations. In addition, those skilled in the art will understand that other structures and processes can be substituted for those disclosed, and that the resulting implementations will achieve at least substantially the same results as the disclosed implementations, performing at least substantially the same functions in at least substantially the same manner. Accordingly, these and other implementations are contemplated in this application.
Claims
1. It is a method, - Regarding the 3D scene portion, obtain the first color patch, reflectance patch, and first depth patch, - To obtain a second color patch and a second depth patch for the portion outside the 3D scene that is reflected in at least one portion of the 3D scene, - Generating a depth atlas by packing the first and second depth patches, - A color atlas is generated by packing the second color patch and a subset of the first color patch, - A reflectance atlas is generated by packing a subset of the reflectance patches, - For each reflectance patch packed in the reflectance atlas, To generate first information that encodes the parameters of the bidirectional reflectance distribution function model of light reflection on the reflectance patch, and To generate second information showing a list of color patches reflected by the aforementioned reflectance patch, A method comprising encoding the depth atlas, the color atlas, the reflectance atlas, the first information, and the second information into a data stream.
2. - The subset of the first color patch packed in the color atlas is empty, The method according to claim 1, wherein the subset of the reflectance patches packed in the reflectance atlas includes all of the reflectance patches.
3. - The subset of the first color patch packed in the color atlas corresponds to the Lambertian portion of the 3D scene, The method according to claim 1, wherein the subset of reflectance patches packed in the reflectance atlas corresponds to the non-diffuse reflectance portion of the 3D scene.
4. - The subset of the first color patches packed in the color atlas includes all of the first color patches, The method according to claim 1, wherein the subset of reflectance patches packed in the reflectance atlas corresponds to the non-diffuse reflectance portion of the 3D scene.
5. The method according to any one of claims 1 to 4, wherein the bidirectional reflectance distribution function model is a Phong model.
6. The method according to any one of claims 1 to 5, further comprising generating a surface normal atlas by packing surface normal patches corresponding to the subset of reflectance patches in the reflectance atlas.
7. It is a device, Processor and A non-temporary computer-readable medium for storing instructions, wherein the instructions, when executed by the processor, - For the 3D scene portion, obtain the first color patch, reflectance patch, and first depth patch. - A second color patch and a second depth patch are obtained for the portion outside the 3D scene that is reflected in at least one portion of the 3D scene. - A depth atlas is generated by packing the first and second depth patches. - A color atlas is generated by packing the second color patch and a subset of the first color patch. - A reflectance atlas is generated by packing a subset of the reflectance patches. - For each reflectance patch packed in the reflectance atlas, First information is generated to encode the parameters of the bidirectional reflectance distribution function model of light reflection on the reflectance patch, and Second information is generated showing a list of color patches reflected by the reflectance patch, and A device comprising a data stream containing a non-temporary computer-readable medium that operates to encode the depth atlas, the color atlas, the reflectance atlas, the first information, and the second information.
8. - The subset of the first color patch packed in the color atlas is empty, - The device according to claim 7, wherein the subset of the reflectance patches packed in the reflectance atlas includes all of the reflectance patches.
9. - The subset of the first color patch packed in the color atlas corresponds to the Lambertian portion of the 3D scene, - The device according to claim 7, wherein the subset of reflectance patches packed in the reflectance atlas corresponds to the non-diffuse reflectance portion of the 3D scene.
10. - The subset of the first color patches packed in the color atlas includes all of the first color patches, - The device according to claim 7, wherein the subset of reflectance patches packed in the reflectance atlas corresponds to the non-diffuse reflectance portion of the 3D scene.
11. The device according to any one of claims 7 to 10, wherein the bidirectional reflectance distribution function model is a Phong model.
12. The device according to any one of claims 7 to 11, wherein the non-temporary computer-readable medium further stores instructions for operating to generate a surface normal atlas by packing surface normal patches corresponding to the subset of reflectance patches in the reflectance atlas.
13. A data stream that encodes a 3D scene, - A depth atlas that packs a first depth patch corresponding to a portion of the 3D scene and a second depth patch corresponding to a portion outside the 3D scene that is reflected in at least one portion of the 3D scene, - A color atlas that packs a first color patch corresponding to a portion of the 3D scene and a second color patch corresponding to a portion outside the 3D scene that is reflected to at least one portion of the 3D scene, - A reflectance atlas for packing reflectance patches corresponding to the portion of the 3D scene, - For each reflectance patch packed in the reflectance atlas, - First information that encodes the parameters of the bidirectional reflectance distribution function model of light reflection on the reflectance patch, and A data stream including: a second piece of information indicating a list of color patches reflected by the reflectance patch.
14. The data stream according to claim 13, wherein the bidirectional reflectance distribution function model is a Phong model.
15. The data stream according to claim 13 or 14, further comprising a surface normal atlas that packs surface normal patches corresponding to a subset of the reflectance patches in the reflectance atlas.
16. A method for rendering a 3D scene, wherein the method is From the data stream, - A depth atlas that packs a first depth patch corresponding to a portion of the 3D scene and a second depth patch corresponding to a portion outside the 3D scene that is reflected in at least one portion of the 3D scene, - A color atlas that packs a first color patch corresponding to a portion of the 3D scene and a second color patch corresponding to a portion outside the 3D scene that is reflected to at least one portion of the 3D scene, - A reflectance atlas for packing reflectance patches corresponding to the portion of the 3D scene, - Information signaling the rendering mode determined according to the first color patch and the reflectance patch, - For each reflectance patch packed in the reflectance atlas, - First information that encodes the parameters of the bidirectional reflectance distribution function model of light reflection on the reflectance patch, and - Decoding a second piece of information that shows a list of color patches reflected by the reflectance patch, A method comprising rendering the 3D scene by back-projecting the first and second color patches according to the first and second depth patches, and by using ray tracing for reflectance patches according to the first and second information and associated color patches.
17. The method according to claim 16, wherein the bidirectional reflectance distribution function model is a Phong model.
18. The method according to claim 16 or 17, further comprising decoding a surface normal atlas from the data stream, which packs surface normal patches corresponding to a subset of the reflectance patches in the reflectance atlas, and using the surface normal patches for ray tracing.
19. It is a device, Processor and A non-temporary computer-readable medium for storing instructions, wherein the instructions, when executed by the processor, From the data stream, - A depth atlas that packs a first depth patch corresponding to a portion of the 3D scene and a second depth patch corresponding to a portion outside the 3D scene that is reflected to at least one portion of the 3D scene, - A color atlas that packs a first color patch corresponding to a portion of the 3D scene and a second color patch corresponding to a portion outside the 3D scene that is reflected to at least one portion of the 3D scene, - A reflectance atlas for packing reflectance patches corresponding to the portion of the 3D scene, - Information signaling the rendering mode determined according to the first color patch and the reflectance patch, - For each reflectance patch packed in the reflectance atlas, - First information that encodes the parameters of the bidirectional reflectance distribution function model of light reflection on the reflectance patch, and - Decode the second piece of information which shows a list of color patches reflected by the reflectance patch. and A device comprising a non-temporary computer-readable medium that operates to render the 3D scene by back-projecting the first and second color patches according to the first and second depth patches, and by using ray tracing for reflectance patches according to the color patches associated with the first and second information.
20. The device according to claim 19, wherein the bidirectional reflectance distribution function model is a Phong model.
21. The device according to claim 19 or 20, wherein the processor is further configured to decode from the data stream a surface normal atlas that packs surface normal patches corresponding to a subset of the reflectance patches in the reflectance atlas, and to use the surface normal patches for ray tracing.