A data signal including the representation of a three-dimensional scene

JP2025519644A5Pending Publication Date: 2026-06-08KONINKLIJKE PHILIPS NV

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
KONINKLIJKE PHILIPS NV
Filing Date
2023-06-08
Publication Date
2026-06-08

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The device generates a data signal including three-dimensional image data that provides a representation of a three-dimensional scene. The three-dimensional image data includes at least one image that provides visual data for the scene. The data signal further includes a view-dependency index for an image area of the image, and the view-dependency index indicates, as a function of the line-of-sight direction, the degree of variation of one or more visual characteristics with respect to the scene point of the image area. The rendering device has a receiver 201 that receives the data signal, and a renderer 203 that generates a view image of the scene from a view pose from the three-dimensional image data according to the view-dependency index. Specifically, the mixing of the contributions from different points of the scene to a given pixel of the view image depends on the view-dependency index. The view-dependency index indicates the view-dependency of the scene point represented by the image area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data signal including the representation of a three-dimensional scene, an apparatus and a method for generating such a data signal, and an apparatus and a method for rendering a view image based on such a data signal. In particular, the present invention relates to a data signal that provides a three-dimensional video signal for immersive video, for example, but not limited to this.

Background Art

[0002] In recent years, the diversity and scope of image and video applications have increased significantly, and new services and methods for using and consuming video are continuously being developed and introduced.

[0003] For example, one popular service is to provide an image sequence in such a way that an observer can actively interact with the system to change the rendering parameters. A very attractive feature in many applications is the function of changing the effective observation position and observation direction of the observer, for example, enabling the observer to move around in the displayed scene.

[0004] Such features can, in particular, make it possible to provide a virtual reality experience to the user. Thereby, the user can move around relatively freely in the virtual environment, for example, and dynamically change his position and the place he is looking at. Typically, such virtual reality (VR) applications are based on a three-dimensional model of the scene, and this model is dynamically evaluated to provide a specific requested view. This approach is well known from game applications, for example, in the category of first-person shooting games for computers and consoles. Other examples include augmented reality (AR) or mixed reality (MR) applications.

[0005] Examples of the proposed video services or applications are immersive videos where the video is played, for example, on a VR headset to provide a three-dimensional experience. In the case of immersive videos, the observer has the freedom to move around looking at the displayed scene, which can be perceived as being viewed from different viewpoints. However, in many typical approaches, the amount of movement is restricted to a relatively small area around a nominal viewpoint that typically corresponds to the viewpoint from which the video capture of the scene was performed. In such applications, three-dimensional scene information that enables high-quality viewpoint image synthesis for viewpoints relatively close to the reference viewpoint is often provided, but degrades when the viewpoint deviates too much from the reference viewpoint.

[0006] Immersive videos are also often referred to as six degrees of freedom (6DoF) or three-dimensional videos. MPEG Immersive Video (MIV) is a new standard in which metadata for enabling and standardizing immersive videos is used on top of existing video codecs.

[0007] Several different representations have been developed and standardized to enable efficient data description of the scene and to enable view images to be generated for different viewports.

[0008] An approach often used to represent the scene is known as multi-view with depth (MVD) representation and capture. In such an approach, the scene is represented by a plurality of images with associated depth data, and the images typically represent different viewports from a limited capture area. The images can actually be captured by using a camera rig equipped with a plurality of cameras and depth sensors.

[0009] Other examples of representations include, for example, point clouds, multi-plane images, multi-spherical images, and volumetric representations sampled at high density. Such representations are known as volumetric representations in which points in space are represented by their position and light characteristics at that position. Other representations may include a light irradiation field sampled at high density or other so-called light irradiation field representations. In the case of a light irradiation field representation, a given point in the scene is linked to different light characteristics corresponding to different light rays passing through that scene point (corresponding to different light rays and light ray directions). Summary of the Invention Problems to be Solved by the Invention

[0010] However, while such representations may be suitable for many different applications and scenarios, they tend not to provide ideal performance or enable complete generation of images. They may also result in a higher data rate and / or processing complexity and / or resource requirements than desirable in many scenarios.

[0011] Therefore, an improved approach for scene representation and its processing is advantageous. In particular, an approach that enables improved operation, increased flexibility, improved immersive user experience, reduced complexity, facilitated implementation, improved image quality, improved and / or accelerated rendering, improved and / or accelerated scene representation or its processing, and / or improved performance and / or operation is advantageous. Means for Solving the Problems

[0012] Therefore, the present invention preferably seeks to alleviate, reduce or eliminate one or more of the above disadvantages, either alone or in any combination.

[0013] According to one aspect of the present invention, there is provided an apparatus arranged to generate a data signal describing a three-dimensional scene, the apparatus comprising: a generator configured to generate a data signal including three-dimensional image data providing a representation of the three-dimensional scene, the three-dimensional image data including at least a first image having light characteristic values of scene points in the scene; and a processor arranged to generate a view-dependency index for an image region of the first image, the view-dependency index indicating a degree of variation of one or more visual characteristics of the scene points of the image region as a function of the line-of-sight direction, wherein the generator is configured to include the view-dependency index in the data signal.

[0014] The present invention can provide improved performance and / or operation and / or implementation in many embodiments. It typically enables an improved representation of a scene, and in particular enables the rendering of more accurate and higher-quality view images of the scene.

[0015] This approach is suitable for many different representations. In particular, since the format of the volume representation of a three-dimensional scene does not inherently allow the view-dependency to be represented, it can provide additional information and improved view adaptation for the volume representation of a three-dimensional scene.

[0016] The view-dependency index can be a Lambert index indicating the level of Lambert of the scene points of the image region, which indicates the level of view-dependency of the scene points of the image region.

[0017] The light characteristic values of the scene can be any values indicating the luminance, chrominance, chroma, luma, luminance and / or color of the scene points. The view-dependency index for the image region can indicate the degree of Lambert of the scene points for which the image region includes light characteristic values.

[0018] The first image can be any two-dimensional data structure that provides light characteristic values of scene points within a scene. The first image can be a frame of a video sequence. The first image can be a projection image for a point cloud, an image corresponding to a viewport of a scene from a capture pose, a multi-plane image that provides light characteristic values for different planes or spheres, an image of a video atlas, etc.

[0019] In many embodiments, a plurality of view-dependent metrics can be provided to reflect the Lambertian of different image regions of the first image and / or image regions of other images. In some embodiments, the view-dependent metrics may vary spatially and can provide metrics of the Lambertian for different image regions / scene regions (e.g., including image regions that are single image samples / pixels).

[0020] According to an optional feature of the present invention, the processor is configured to generate view-dependent metrics depending on three-dimensional image data.

[0021] This can provide improved performance and / or operation in many embodiments. It can provide an efficient and high-performance approach for determining view-dependent metrics in many embodiments.

[0022] According to an optional feature of the present invention, the processor is configured to determine the light characteristics in each direction for the scene region represented by the image region and determine the view-dependent metrics according to the variation of the light characteristics in each direction.

[0023] This can provide improved performance and / or operation in many embodiments. It can provide an efficient and high-performance approach for determining view-dependent metrics in many embodiments.

[0024] In some embodiments, the processor is configured to generate a view-dependency metric in response to a comparison of light output from an image region in each direction.

[0025] According to an optional feature of the present invention, the view-dependency metric indicates a variation in light emission as a function of the direction of a scene point in the image region.

[0026] This can provide improved performance and / or operation in many embodiments.

[0027] According to an optional feature of the present invention, the representation of the three-dimensional scene is a multi-view + depth representation, and the first image is one of the set of multi-view images of the multi-view + depth representation.

[0028] This approach can provide particularly advantageous operation and performance for the multi-view + depth representation. The use of the view-dependency metric with the multi-view + depth representation can provide a synergistic effect that enables, for example, an improved view image of the scene to be rendered.

[0029] According to an optional feature of the present invention, the representation of the three-dimensional scene is a multi-plane image representation, and the first image is one of the planes of the multi-plane image representation.

[0030] This approach can provide particularly advantageous operation and performance for the multi-plane image representation. The use of the view-dependency metric with the multi-plane image representation can provide a synergistic effect for a volume representation that, for example, does not inherently consider or encode any variation in view direction.

[0031] According to an optional feature of the present invention, the representation of the three-dimensional scene is a point cloud representation, and the first image includes light characteristic values for projecting at least a part of the point cloud representation onto an image plane.

[0032] This approach can provide particularly advantageous operation and performance for multi-plane image representation. The use of view-dependent metrics with multi-plane image representation can provide a synergistic effect, for example, for volume representations that do not inherently consider or encode any view direction variations.

[0033] The image region can include pixels that indicate light emission for points of the point cloud. The scene points of the image region can be points of the point cloud.

[0034] According to an optional feature of the present invention, the representation of the three-dimensional scene is a representation that includes at least a first image and projection data indicating a relationship between the light characteristic value positions of the first image and the positions of corresponding scene points within the three-dimensional scene.

[0035] This can provide improved performance and / or operation in many embodiments.

[0036] The corresponding scene points of the light characteristic value positions are scene points whose light characteristic values at those light characteristic value positions provide an indication of the light characteristic values of the scene points. The light characteristic value positions can be pixel positions within the first image, and the light characteristic values can be pixel values.

[0037] According to an optional feature of the present invention, the generator is configured to receive input three-dimensional image data for a three-dimensional scene and select a subset of the input three-dimensional image data for inclusion in the data signal, and the selection of the subset depends on a view-dependent metric.

[0038] This can provide improved performance and / or operation in many embodiments. This feature can provide an improved data signal that enables higher quality rendering at a reduced data rate in many embodiments. This feature can enable an improved trade-off between the image data for high-quality rendering and the data rate of the data signal.

[0039] In some embodiments, the generator can be configured to select a subset to have a size that monotonically increases with the degree of view-dependence indicated by the view-dependence metric.

[0040] According to an optional feature of the present invention, the generator is configured to include a view-dependence metric in the chrominance color channel of the image of the three-dimensional image data.

[0041] This can provide improved performance and / or operation in many embodiments.

[0042] According to one aspect of the present invention, there is provided an apparatus for rendering a view image of a three-dimensional scene, the apparatus comprising three-dimensional image data providing a representation of the three-dimensional scene, the three-dimensional image data including at least a first image consisting of light characteristic values of scene points in the scene, and at least one view-dependence metric for an image region of the first image, the view-dependence metric indicating a degree of variation of one or more visual characteristics of scene points in the image region as a function of the viewing direction, a first receiver configured to receive a data signal including the three-dimensional image data and the view-dependence metric, and a renderer configured to generate a view image for a viewport from the three-dimensional image data depending on the view-dependence metric.

[0043] The present invention can provide improved performance and / or operation and / or implementation in many embodiments. It can typically enable an improved representation of the scene and, in particular, enable the rendering of more accurate and higher quality view images of the scene.

[0044] According to an optional feature of the present invention, the renderer is configured to generate the pixel values of the view image by mixing contributions from a plurality of light source characteristic values of the three-dimensional image data projected onto the positions of the pixels in the view image, and the mixing of the contributions from the light characteristic values of the image region depends on the view-dependence metric.

[0045] This can provide particularly advantageous operation in many scenarios and embodiments. This approach can enable improved rendering, specifically improved / more realistic images. It can provide improved perceived image quality, including, for example, less image noise generated by the mixing operation.

[0046] The weight of the contribution from the light characteristic value of the image region in the mixing / mixing for generating the pixel value of the view image depends on the view-dependency index of the image region.

[0047] According to an optional feature of the present invention, the three-dimensional image data includes transparency values, and the renderer is configured to modify at least one transparency value according to the view-dependency index.

[0048] This can provide particularly advantageous operation in many scenarios and embodiments

[0049] According to an optional feature of the present invention, the renderer is further configured to modify at least one transparency value according to the difference between the light direction of the pixel of at least one transparency value and the direction from the position in the three-dimensional scene represented by the pixel and the view pose.

[0050] This can provide particularly advantageous operation in many scenarios and embodiments.

[0051] According to one aspect of the present invention, there is provided a method of generating a data signal for describing a three-dimensional scene (excluding a method of performing such a mental act), the method comprising: generating a data signal to include three-dimensional image data providing a representation of the three-dimensional scene, the three-dimensional image data including at least a first image including light characteristic values of scene points in the scene; generating a view-dependency index for an image region of the first image, the view-dependency index indicating a degree of variation of one or more visual characteristics of scene points in the image region as a function of a line-of-sight direction; and including the view-dependency index in the data signal.

[0052] According to one aspect of the present invention, there is provided a method of rendering a view image of a three-dimensional scene (excluding a method of performing such a mental act), the method comprising: receiving a data signal including three-dimensional image data providing a representation of the three-dimensional scene, the three-dimensional image data including at least a first image including light characteristic values of scene points in the scene, and at least one view-dependency index for an image region of the first image, the view-dependency index indicating a degree of variation of one or more visual characteristics of scene points in the image region as a function of a line-of-sight direction; receiving a view pose of the three-dimensional scene; and generating a view image for the view pose from the three-dimensional image data depending on the view-dependency index.

[0053] These and other aspects, features, and advantages of the present invention will become apparent from and will be described with reference to the embodiments described hereinafter.

Brief Description of the Drawings

[0054] Embodiments of the present invention are described by way of example only with reference to the drawings.

Figure 1

Figure 2

Figure 3

DETAILED DESCRIPTION OF THE INVENTION

[0055] The capture, distribution, and display of three-dimensional video are becoming increasingly popular and desirable in several applications and services. A particular approach is known as immersive video and typically includes providing views of real-world scenes and often real-time events, enabling small observer movements such as relatively small head movements and rotations. For example, a local client-based view generation that follows the small head movements of an observer enables a real-time video broadcast of a sports event to provide the impression that the user is sitting in the stand watching the sports event. The user can, for example, look around and have a natural experience similar to that of the spectators present at that location within the stand. In recent years, there has been an increasing spread of display devices equipped with position tracking and three-dimensional interaction support applications based on the three-dimensional capture of real-world scenes. Such display devices are very suitable for immersive video applications that provide an enhanced three-dimensional user experience.

[0056] To provide such services for real-world scenes, an appropriate data representation of the scene must be generated and communicated to an apparatus that renders the view to the end user.

[0057] Typically, scenes are captured from different positions with different camera capture poses being used. As a result, the relevance and importance of multi-camera capture and, for example, 6DoF (6 degrees of freedom) processing are increasing rapidly. Applications include live concerts, live sports, and telepresence. The freedom to choose one's own perspective enriches these applications by enhancing the sense of presence more than normal video. Furthermore, immersive scenarios can be considered where an observer can move through and interact with a live-captured scene. In the case of broadcast applications, this may require real-time depth estimation on the production side and real-time view synthesis on the client device. Both depth estimation and view synthesis introduce errors, and these errors depend on the details of the algorithm implementation.

[0058] In this field, the terms "placement" and "pose" are used as general terms related to position and / or orientation / direction. For example, a combination of the position and orientation / direction of an object, camera, head, or view may be referred to as a pose or placement. Thus, a placement or pose indicator can include six values / components / degrees of freedom, and each value / component typically describes an individual characteristic of the position / location or orientation / direction of the corresponding object. Of course, in many situations, for example, when one or more components are considered fixed or irrelevant (e.g., when all objects are at the same height and are considered to have a horizontal orientation, four components can provide a complete representation of the object's pose), the placement or pose may be considered or represented with fewer components. In the following, the term "pose" is used to refer to a position and / or orientation that can be represented by one to six values (corresponding to the maximum possible degrees of freedom). The term "pose" can be replaced by the term "placement". The term "pose" can be replaced by the terms "position and / or orientation". The term "pose" can be replaced by the terms "position and orientation" (when the pose provides information about both position and orientation), "position" (when the pose provides position information in some cases), or "orientation" (when the pose provides orientation information in some cases).

[0059] FIG. 1 shows an example of a data signal generation device configured to generate a data signal including a representation of a three-dimensional scene. FIG. 2 shows an example of a rendering device for rendering a view image of a three-dimensional scene from a data signal including a representation of the scene. Specifically, in this example, the signal generation device generates a data signal provided to the rendering device of FIG. 2, and this data signal proceeds to render a view image for the scene based on the scene representation data of the signal generated by the signal generation device of FIG. 1. This approach is described with reference to this example.

[0060] The data signal generation device includes an image data source 101 configured to provide three-dimensional image data that provides a representation of a three-dimensional scene. The image data source 101 includes, for example, a storage device or memory in which the three-dimensional image data is stored and from which it can be read. In other embodiments, the image data source 101 can receive or read the three-dimensional image data from an external (or other internal source). In many embodiments, the image data source 101 can be received, for example, from a video capture system including a video camera that captures a real-world scene in real time.

[0061] The three-dimensional image data provides an appropriate representation of the scene using an appropriate image or typically a video format / representation. In some embodiments, the image data source 101 can receive image data from different sources and / or in different formats, and then generate three-dimensional image data, for example, by converting or processing the received scene data. For example, the image and depth can be received from a remote camera and processed to generate a video representation according to a given format. In some embodiments, the image data source 101 can be configured to generate three-dimensional image data by evaluating a model of the scene.

[0062] The three-dimensional image data is provided according to an appropriate three-dimensional data format / representation. The representation can be, for example, multi-view and depth, multi-layer (multi-plane, multi-spherical), mesh model, and / or point cloud representation. Also, in the above example, the three-dimensional image data is specifically video data and can include a time component.

[0063] The three-dimensional image data is provided according to a representation that includes one or more images. In many embodiments where the three-dimensional image data represents video data, the three-dimensional image data can include time-series images. The image can be any two-dimensional structure of the light characteristic values of the scene points of the scene.

[0064] The image can include, for example, a view image of a scene, a frame of a video signal, a video atlas, a picture, a coded picture, a decoded picture, etc.

[0065] The light characteristic value can be, for example, brightness, color, chroma, luma, chrominance and / or luminance value. The light characteristic value of a scene point can provide an indicator of light radiated from the scene point in at least one direction. In some expressions, each scene point can be associated with a plurality of light source characteristic values. The scene point represented by the light characteristic value of the 3D image data may also be called a scene sample, an image sample, or simply a sample.

[0066] The image data source 101 is coupled to a generator 103 configured to generate a data signal such as a bitstream including three-dimensional image data. The data signal / bitstream is typically a (three-dimensional) video data signal / bitstream. Specifically, the generator 103 can generate a bitstream including three-dimensional image data according to an appropriate transport format.

[0067] The generator 103 is coupled to a transmitter 105 configured to transmit the data signal to a remote source. The transmitter 105 can include, for example, a network interface to enable the data signal to be transmitted to an appropriate destination via a network. The network can be the Internet or can include the Internet.

[0068] The rendering device includes a first receiver 201 configured to receive a data signal from the data signal generating device. The receiver can be any appropriate function for receiving three-dimensional image data according to the preferences and requirements of individual embodiments. Specifically, the first receiver 201 can include a network interface that enables the image data source 101 to be received via a network such as (or including) the Internet.

[0069] The first receiver 201 is configured to extract three-dimensional image data from the received data signal. The first receiver 201 is coupled to a renderer 203 configured to generate view images of different view poses from the three-dimensional image data.

[0070] The rendering device of FIG. 2 further includes a second receiver 205 configured to receive a view pose (specifically, a view pose within a three-dimensional scene) for an observer / user. The view pose represents the position and / or orientation from which a viewer views the scene, and in particular can provide the pose from which a view of the scene is to be generated.

[0071] Many different approaches for determining and providing view poses are known, and it will be understood that any suitable approach can be used. For example, the second receiver 205 can be configured to receive pose data from a VR headset or an eye tracker worn by the user. In some embodiments, relative view poses can be determined (e.g., a change from an initial pose can be determined), which is related to a reference pose such as a camera pose or the center of a capture pose region.

[0072] The second receiver 205 is coupled to a renderer 203 configured to generate view frames / images from the three-dimensional image / video data, and view images are generated to represent views of the three-dimensional scene from the view poses. Thus, the renderer 203 can generate a video stream of view images / frames of the three-dimensional scene, specifically from the received three-dimensional image data and the view poses. In the following, the operation of the renderer 203 will be described with reference to the generation of a single image. However, it will be understood that in many embodiments, the image can be part of a series of images and specifically can be a frame of a video sequence. In fact, the approach described can be applied to generate multiple, often all, frames / images of the output video sequence.

[0073] It will be appreciated that a stereoscopic video sequence including a video sequence for the right eye and a video sequence for the left eye is often generated. Thus, when an image is presented to a user, for example, via an AR / VR headset, it appears as if a three-dimensional scene is being observed from a view pose. In another example, an image can be presented to a user using a tablet having an inclinometer, and a monoscopic video sequence can be generated. In yet another example, for display on an autostereoscopic display, multiple images are woven or tiled.

[0074] The renderer 203 can perform appropriate image composition / generation operations for generating a view image in a particular representation / format of three-dimensional image data (and in fact, in some embodiments, the three-dimensional image data may be converted to different formats from which the view image is then generated).

[0075] For example, the renderer 203 can be configured to perform a view shift or projection of the received multi-view images, typically based on depth information, for a multi-view + depth representation. This typically involves techniques such as shifting pixels (changing pixel positions to reflect appropriate parallax corresponding to a parallax change), deocclusion (usually based on filling from other images), and combining pixels from different images, as is known to those skilled in the art.

[0076] Many algorithms and approaches are known for synthesizing images from different three-dimensional image data formats and representations, and it will be appreciated that any suitable approach can be used by the renderer 203.

[0077] In this way, the image synthesizing apparatus can generate a view image / video of a scene. Further, since the view pose may change dynamically in response to a user moving within the scene, the view of the scene can be continuously updated to reflect the change in the view pose.

[0078] In this approach, the three-dimensional image data is further supplemented by the inclusion of a view-dependent index that indicates the degree of variation of one or more visual characteristics of an image region as a function of the viewing direction.

[0079] Specifically, the view-dependent index can be a Lambert index that indicates the level of Lambert for at least one image region of at least one image of the three-dimensional image data. The view-dependent / Lambert index can indicate the degree of view-dependence of the visual characteristics of the image region. The view-dependent / Lambert index can indicate the degree / level of variation of one or more visual characteristics / light characteristics for the image region as a function of the viewing direction. The view-dependent / Lambert index can indicate the variation of one or more visual characteristics of the image region as a function of the viewing direction. The view-dependent / Lambert index can indicate the degree / level of variation of light emission and / or reflection with respect to the image region as a function of the viewing direction. The visual characteristics can specifically be the luminance and / or chrominance characteristics of the image region.

[0080] Hereinafter, the term Lambert index will be mainly used, but it will be understood that this term may be replaced by the term view-dependent index. For the sake of brevity, references to the view-dependent / Lambert index will be simply referred to by the term Lambert indicator. Similarly, references to Lambert or degree of Lambert may be understood to be replaced by the terms view-dependent and view-dependence as appropriate. Therefore, for the sake of brevity, the following text uses the term Lambert to include the term view-dependence. It will also be understood that the term view-dependence includes the terms view-direction dependence or viewing angle dependence.

[0081] Accordingly, the data signal generating device comprises a Lambert determiner 107 configured to generate a Lambert index. The Lambert determiner 107 is coupled to a data generator 103 supplied with the Lambert index and configured to include it in the data signal / bitstream. The data generator 103 can specifically encode the Lambert index and add it to a dedicated field of the bitstream.

[0082] The Lambert index can indicate the variation of the light output (from the scene surface / element / object / region represented by the image region) as a function of the direction with respect to the scene region represented by the image region.

[0083] Lambertian reflection is the light reflection provided by an ideal diffusing reflector. The apparent brightness of a Lambertian surface to an observer is the same regardless of the observer's viewing angle. A perfectly Lambertian surface provides the same visual impression regardless of the viewing direction, such as a perfectly matte white surface that reflects light in all directions. Similarly, a non-collimated light source provides the same light in all directions and thus has the same visual appearance in all directions. In the case of a non-collimated light source, the light is not visually dependent on the observation direction / viewing angle. For the purposes of this specification, it does not matter whether the observed light is the result of reflection or emission, and thus a non-collimated light source is effectively Lambertian. The term Lambert includes non-collimated light sources and thus is not limited to Lambertian reflection only, but also includes Lambertian light or non-collimated light (source) illumination / radiation.

[0084] For many practical reflectors and light sources, there is a certain degree of viewing direction sensitivity, and any variation in light can result in a variation in the observer's direction. Therefore, there is some view-dependence for such non-Lambertian effects. The degree of view-dependence can vary substantially. For example, the glossier a white wall is, the lower its Lambertian degree becomes, and the visual impression of the wall becomes view-dependent. Other view-dependent effects, also known as non-Lambertian effects, can include the following: · Glossy surfaces such as marble and wood. · Hard specular reflection. · Metal surfaces. · View differences due to polarization, glare, etc.

[0085] The data signal generation device is configured to generate a Lambertian index indicating the variation of the light output as a function of the direction with respect to the scene region represented by the image region.

[0086] In some embodiments, the Lambertian index can be received from a given source or can be generated by manual input, for example, by a person selecting a particular image region and manually assigning a Lambertian level to that region. In other embodiments, information about the objects in the scene, such as the material of each object in the scene, can be provided, and a corresponding Lambertian level can be assigned based on the object.

[0087] In some embodiments, the Lambert determiner 107 is configured to generate a Lambert metric according to three-dimensional image data. Specifically, in many embodiments, the Lambert metric can be generated according to a comparison of light outputs from image regions in different directions. In many embodiments, the Lambert determiner 107 is configured to determine light emissions in different directions for a scene region represented by an image region. Then, it can determine the Lambert metric according to the variation of light emissions for different directions. The light from the same scene region in different directions can be compared to evaluate the similarity. For example, differences in luminance, color, luminance and / or chrominance can be determined. The greater the determined difference, the smaller the degree of Lambertness of the corresponding image region. If the light is considered to be the same for all different directions, the image region can be considered to be completely Lambertian.

[0088] The light in different directions can be determined in different ways in different embodiments, actually for different representations of the scene used by the three-dimensional image data. For example, in the case of a multi-view + depth representation, pixel values corresponding to the same scene point in different multi-view images can be compared. For a representation including a plurality of point clouds, such as obtained by using one LiDAR for each point cloud, the scene points of the plurality of point clouds can be aligned and compared. For a point cloud captured using a photogrammetry booth having a plurality of cameras and light sources, the appearance of the scene points in the plurality of cameras and / or under a plurality of light conditions can be compared.

[0089] As a specific example, the image sample for which the Lambertian degree is calculated is associated with a plurality of other image samples that (substantially) correspond to the same scene point. From this set of image samples, a statistical value is calculated based on the color values of these image samples. Suitable statistics are variance, standard deviation, or the range of values (max - min), and for example, higher values indicate more appearance variations. Then, the Lambertian degree is determined by applying a mapping to the statistical value such that higher variations result in lower Lambertian degrees. To calculate the Lambertian degree of an image region, the Lambertian degrees of a plurality of image samples can be averaged.

[0090] In many embodiments, the image region can be an image region where the scene point represented by that image region is part of the same surface within the scene. However, it will be understood that this is not essential. For example, in some embodiments, an image can be divided into predetermined image regions / segments, and a view - dependency / Lambertian index is generated for each region / segment. In such cases, the view - dependency index can simply reflect, for example, the maximum or average degree of view - dependency for the scene points of the image region. Such an approach can typically introduce some inaccuracies and potentially degrade the rendering achieved when compared to the ideal case, but in many practical applications, it provides acceptable performance and image quality.

[0091] Each pixel of the image region can provide at least one light characteristic value for one scene point. The Lambertian index indicates the degree to which the light from this scene point depends on the direction in which it is viewed, i.e., the direction from the view - pose to the scene point. The scene point often corresponds to a point on a surface within the scene, and the Lambertian index indicates the degree of Lambertianness of that surface.

[0092] The Lambertian indicator can, in some embodiments, be a simple binary indicator that indicates whether an image region is a Lambertian surface / point or a non-Lambertian surface / point. The Lambertian indicator can provide an indication of whether the visual information provided by the light characteristic values / pixel values of the image region is view-dependent or not for a given image region of a given image.

[0093] However, in many embodiments, a more detailed indicator can be provided, and the Lambertian indicator can reflect more levels / degrees of view-dependence for the image region. For example, the Lambertian indicator can be a 2-bit indicator that indicates whether the image region represents the following surfaces: · Completely Lambertian (e.g., a white wall). · Somewhat non-Lambertian (e.g., a marble floor). · Almost non-Lambertian (e.g., fire, smoke). · Completely non-Lambertian (e.g., a hard specular reflection).

[0094] In other embodiments, more levels can be used, and the Lambertian indicator can be a 4-bit, 6-bit, or 8-bit value or more.

[0095] In some applications, only a single Lambertian indicator for a single image region of a single image can be included in the bitstream. However, in many embodiments, multiple Lambertian indicators can be included.

[0096] In many embodiments, Lambertian indicators for multiple image regions of an image can be included. In fact, in some applications, an image is divided into multiple image regions, and the Lambertian indicator is provided for each image region to indicate the level / degree of Lambertian for that image region. Each element of the image can have an associated Lambertian indicator.

[0097] In some applications, the image region is predetermined and fixed. For example, the image can be divided into small rectangles, and the Lambertian indicies can be provided for each rectangle. In fact, in some embodiments, each region can include a single pixel, and thus, in some embodiments, the Lambertian indicies can be provided for each pixel of the image.

[0098] The bitstream can be generated to include spatially varying Lambertian indicies that reflect the Lambertian degree of the surfaces represented by different parts of the image. In many embodiments, a Lambertian map can be generated that includes Lambertian indicies corresponding to / mapping to the image. For example, a two-dimensional map of Lambertian indicies having the same resolution as the image can be provided. In many embodiments, the resolution of the Lambertian map can be lower than the resolution of the image.

[0099] In some embodiments, the image can be divided into image regions by a dynamic and adaptive process that depends on the data of the image. For example, segmentation can be performed according to the image data to divide the image into regions having similar visual characteristics. Then, the Lambertian indicies can be provided for each region or for a subset of the regions. In some embodiments, the segmentation can be performed to detect and identify objects, and the Lambertian indicies can be provided for each object.

[0100] Furthermore, in many embodiments, the Lambertian indicies or Lambertian index maps are not only provided for a single image, but can also be applied to two, more or all of the images used to represent a three-dimensional scene at a given point in time.

[0101] Furthermore, in the case of a video sequence, a Lambertian metric or Lambertian metric map can be provided for images at different points in time. In some embodiments, one or more Lambertian metric maps can be provided for each image / frame of the video sequence. In such cases, the Lambertian metric can be determined, for example, taking into account temporal aspects and correlations over different instants. Similarly, the encoding of the Lambertian metric can include, for example, relative encoding. For example, the Lambertian metric map of one image / frame can be encoded as relative values with respect to the Lambertian metric map of the corresponding previous image / frame.

[0102] The renderer 203 is configured to perform image generation according to the Lambertian metric. Specifically, the generation of pixel values for the view image can take into account, for some pixels, multiple scene points and light characteristic values (i.e., multiple samples) from the three-dimensional image data. For example, when multiple samples are projected onto the same pixel position within the view image, the pixel value can be determined by combining the contributions from the multiple light characteristic values. In many applications, the renderer can perform a blending operation to combine multiple operations.

[0103] When such a combination includes samples / light characteristic values from the image region where the Lambertian metric is provided, the combination is dependent on the Lambertian metric. Specifically, in many cases, the determination of the pixel values of the view image can include a combination / merge / blending of multiple light source characteristic values of one or more samples (including samples from different images) of the three-dimensional image data. For example, when the three-dimensional image data projects scene points with light characteristic values onto the image positions of the view image, multiple scene points can be projected onto the same pixel position. The multiple scene points of the three-dimensional image data can provide visual information about the image position. And multiple pixel values / visual data points can be combined to generate the pixel values of the view image.

[0104] This combination can take into account the depth of the corresponding scene points. For example, if the scene points are at very different depths, the frontmost scene point is selected and its pixel value can be used for the view image pixel value. However, if the depth values are close to each other, all pixel values and scene points are related and the view image pixel value can be generated as a mixture that includes contributions from multiple scene points. In many cases, multiple scene points and light characteristic values may actually represent the same surface. For example, in the case of multi-view + depth representation, the scene point / pixel value may represent the same surface seen from different viewpoints.

[0105] The combination / merging / mixing of multiple pixel values further depends on the Lambertian metric. For example, if the Lambertian metric indicates that one or more of the pixel values are associated with a very high degree of Lambertianness, the combination can be a selection combination where the pixel values are selected from, for example, a preferred image. In contrast, if the Lambertian metric indicates a low degree of Lambertianness, the combination can include contributions from all pixel values. For example, the weighting of each contribution depends on how close the corresponding observation direction is to the observation direction of the multi-view image. Such an approach can provide improved view image quality in many embodiments.

[0106] In some cases, the renderer 203 can be configured to switch between a selection mix and a combination mix according to the Lambertian metric. In some embodiments, the renderer can be configured to determine a mixing weight for the contribution to the pixels being mixed from the scene points of the image region according to the Lambertian metric.

[0107] Thus, in many embodiments, the renderer 203 can be configured to generate the light characteristic value of a pixel of the view image by mixing the contributions from different light characteristic values of the three-dimensional image data projected onto the position of the pixel within the viewport of the view image. For example, the projection of a scene point results in a given position within the viewport of the image (and a given (pixel) position within the view image) for a given view pose. The light characteristic values of the scene points projected onto the same pixel are mixed together to generate the image / pixel value of that pixel within the view image. This mixing can be such that the contribution from the image values within a given image region depends on the Lambertian index for that image region.

[0108] In many embodiments, the renderer 203 can be configured to consider transparency values when compositing the view image. In many applications, an image can be associated with a transparency / alpha map that includes transparency / alpha values indicating the transparency of the associated pixels. When combining / mixing different light characteristic values, the mixing can weight the contributions of the different light characteristic values based on the transparency values. For example, for a transparency value of 1 (corresponding to complete opacity), the contribution of the light characteristic value is 100% if it is the frontmost pixel. For a transparency value of 0 (corresponding to complete transparency), the contribution of the light characteristic value is 0%. When the transparency value is between 0 and 1 (corresponding to partial transparency), the contribution degree is between 0% and 100%. Such mixing enables the representation of partially transparent objects such as cloudy glass or some plastic.

[0109] In embodiments where transparency / alpha values are provided, the renderer 203 can be configured to modify one or more of the transparency values according to the Lambertian index. For example, for a given scene point and light characteristic value of the three-dimensional image data, the Lambertian index can be used to modify the transparency value of that light characteristic value.

[0110] As a specific example, the transparency value can be used to perform the above-described mixing / selection. In the case of pixel values that are substantially in the same depth and have a high Lambertian degree, one pixel can be selected by setting the transparency value so as not to show transparency, while in the case of pixels with a low Lambertian degree, it can be set to a low level so as to enable mixing between different values.

[0111] As another example, for instance, without changing the transparency for complete Lambertianness, the transparency can be changed more gradually based on the degree of Lambertianness by increasing the transparency (i.e., decreasing the opacity) as the value of the Lambertian degree is lower and the difference between the pose of the original view and the pose of the virtual view is larger.

[0112] In the example of FIG. 2, the renderer 203 includes an alpha processor 207 that receives the Lambertian indicator and the transparency value. And it can proceed to modify the transparency value based on the Lambertian indicator. The alpha processor 207 is coupled to a blender 209, which is configured to mix a plurality of light characteristic values to generate pixel values of the view image, and the mixing depends on the transparency value. The blender 209 further has a function for performing a related projection operation to determine samples of the three-dimensional image data to include in the mixing for a given pixel of the view image. The blender is further coupled to a view image generator 211 that generates the view image by combining individual pixel values.

[0113] As a specific example suitable for many types of representations (including, for example, multi-views, video-based point clouds, meshes, etc.), the mixing of image values / samples at similar depths in the viewport of a view image can use transparency (alpha) values that can be modified according to the Lambertian metric. For example, if the Lambertian metric is provided but no transparency is indicated (completely opaque), the output of this mixing operation can be a single opaque value (R, G, B, alpha = 1). This is achieved by scaling the alpha value of the contribution. Otherwise, the transparency / alpha values are mixed, for example, using a saturation addition operation (to constrain alpha to the range [0,1]).

[0114] In many embodiments, the Lambertian metric may have little or no practical effect when only a single sample / image value is projected at a given position within the viewport of the view image. However, when multiple samples are projected onto the same scene point (sufficiently close), the contributions can typically be advantageously mixed. In many embodiments, the renderer 203 can distinguish scene points along the same ray (one in front of the other), and typically only scene points at similar depths are blended.

[0115] Specific examples of possible approaches for modifying the transparency value depending on the Lambertian metric are as follows: TIFF2025519644000002.tif1254

[0116] · x is an integer or fractional sample coordinate within the scene. · α0(x) is the value of the transparency component, or the transparency value automatically calculated based on the texture and geometric components (rendering in two layers). If no transparency information exists, α0(x) = 1. · L(x) is the value of the Lambertian metric at position x. · β is a measure of the difference in the ray direction between the sample and the viewport. A given sample of three-dimensional image data can provide light characteristic values for the scene. Such values are typically explicitly or implicitly associated with the ray direction. For example, in the case of an image in a multi-view representation, the ray direction is the direction from the scene point of the sample to the capture point. For many images, the ray direction will be perpendicular to the image plane. For some representations such as point clouds, multiple light characteristic values can be provided for each point, and each of these can be associated with a ray direction. β can be a value that reflects the difference between the ray direction of the light characteristic value / sample and the ray direction from the corresponding scene point to the view pose of the view image. · Terms for some constants c1 and c2 TIFF2025519644000003.tif718 is an exemplary formula for correcting the effect of the Lambert index on transparency. An alternative can be a piecewise linear curve where the coefficients are predetermined or transmitted. Also, this exponent can be replaced with another non-linear function, such as a scaled sigmoid, etc.

[0117] There are multiple ways to encode and derive the ray difference value β, which typically depends on the selected stereoscopic representation. The ray difference value β always depends on the viewport / view pose of the view image. Typically, the ray difference value β is a positive angle, and a value related to this can be, for example, b = sinβ.

[0118] In the case of perspective projection and / or orthographic projection representations such as MPEG immersive video (MIV), the ray difference value β can be determined as the angle between two vectors, where the first vector is from the principal point of the source view to the scene point, and the second vector is the vector from the viewport / view pose (the principal point of the virtual camera from which the view image is generated) to the scene point.

[0119] In some orthographic projections, such as MIV or some examples of point cloud representations, the ray difference value β can be determined as the angle between two vectors, where the first vector is the normal vector of the orthographic plane and the second vector is the vector from the viewport / viewpose (the principal point of the virtual camera from which the view image is generated) to the scene point.

[0120] In some embodiments, additional metadata indicating the ray direction of samples of the three-dimensional image data can be provided to the three-dimensional image data. Such metadata can include, for example, a table having either a source view position or a direction vector. And the metadata of the image region can include identification information or an index within the table. And the ray difference value β can be determined based on the ray direction of a given sample with respect to the direction to the viewpose.

[0121] This approach can provide advantageous operation and performance in many scenarios and applications. In particular, this approach can provide improved view images to be generated in many scenarios. This approach results in a lower bitrate at a given visual quality level of the view image. This approach results in a lower pixel rate at a given visual quality level of the view image. This approach reduces the rendering complexity in terms of instruction count or battery consumption at a given quality level of the view image.

[0122] This approach further enables efficient implementation and operation. This approach can often be implemented with low complexity and relatively low computational overhead. Furthermore, relatively few modifications are required for many existing devices and systems, and thus a high degree of backward compatibility and simple modifications are typically achieved.

[0123] This approach is suitable for various different representations, such as multi-view and depth representations / formats as described above.

[0124] The described approach would be well-suited for a representation of a scene that includes one or typically more images having pixels indicative of light characteristics / light emission intensity, and further includes projection data indicative of the relationship between the pixels of the image and spatial positions within the three-dimensional scene. Thus, a representation can be used in which the spatial / geometric information associating the image pixels with scene points is provided implicitly and / or explicitly. Such data can be effectively used, specifically, to determine appropriate rendering operations, such as the mixing and determination of transparency values as described above. And using the obtained transparency values, a mixing operation can be performed.

[0125] This approach is particularly suitable for volumetric representations such as multi-plane image representations or point cloud representations. A volumetric representation is, specifically, a representation that describes the geometry and attributes (e.g., color, reflectivity, material) of one or more objects within a volume. Problems with pure volumetric representations are that they typically do not represent, and in many cases cannot represent, view-dependent / non-Lambertian effects. The described approach enables the communication and rendering of such effects also for volumetric representations.

[0126] In many embodiments, a multi-plane image representation can be used, and an image including an image region that provides a Lambertian indication can be a plane of the multi-plane image representation. In particular, each plane of the multi-plane image typically has few image regions that are not completely translucent. These image regions of one or more planes can be packed together in an atlas, which can have image (or video) components for each attribute of the multi-plane image. Typically, the attributes are color and transparency. Typically, only specific regions within the scene are non-Lambertian, and only a portion of the image regions correspond to these scene regions. These specific image regions can be packed multiple times with different indications of the original viewport. Thus, this approach enables the transmission of non-Lambertian effects with only a slight increase in atlas size.

[0127] In many embodiments, a point cloud representation can be used, and an image including an image region that provides a Lambertian indication can be a (partial) projection onto an image plane of the point cloud representation. The image can be an atlas that includes patches of pixels that indicate light emission to the points of the point cloud. In particular, an atlas encoding the point cloud typically has only a few patches corresponding to non-Lambertian scene regions. These specific patches can be packed multiple times into the atlas with different indications of the original viewport. Thus, this approach enables the transmission of non-Lambertian effects with only a slight increase in atlas size.

[0128] In some embodiments, the image data source 203 can be configured to receive input three-dimensional image data of a scene and then proceed to select this subset to be included in the data signal / bitstream. The image data source 203 can be configured to select this subset based on a Lambertian metric. Specifically, for scene regions where a relatively high degree of Lambertianness is indicated and thus low view-dependency, the amount of alternative data representing that scene region in the bitstream can be significantly reduced compared to the case where that region is shown to have a low level of Lambertianness. In some embodiments, the image data source 203 can be configured to select the subset to have a size that monotonically increases with the degree of view-dependency indicated by the Lambertian metric. The greater the view-dependency, the larger the size of the subset.

[0129] As a specific example, in the case of multi-view + depth representation, a given surface can be represented by different images corresponding to different view / capture positions. If the surface is shown to be completely Lambertian, only data from a single multi-view image is included. If the surface is shown to be the opposite, data from all of the received multi-view images is included. For Lambertianness between these two extremes, for example, data from only some of the multi-view images can be included.

[0130] As another example, for specular highlights and specular surfaces, all available (or in some cases a (large) subset of) multi-view images are included, and a low Lambertian metric causes the renderer to use mainly the image of the viewport pose closest to the viewport pose for rendering the viewport.

[0131] The Lambertian metric can be included in the data signal in any suitable manner. In many embodiments, the Lambertian metric may be included as metadata within a dedicated field of the data signal. For example, a Lambertian metric map may be included for each image of the representation.

[0132] In some embodiments, the generator 103 can be configured to include a Lambertian indicator within the chrominance color channel of the three-dimensional image data. The Lambertian indicator can be packed into the channels of the video component in some embodiments. Such an approach can provide a very efficient approach for communicating the Lambertian indicator.

[0133] For example, the approach described may be well-suited for representations compliant with the currently under-development Visual Volumetric Video-based Coding (V3C) standard.

[0134] For such applications, the current data format can be modified to include the Lambertian indicator in the video component data field. In the case of V3C, the data format and the packing_information(j) syntax structure can be found in ISO / IEC DIS 23090-5(2E):2021 "V3C". The packing information explains how multiple video components are packed into a single video bitstream. The approach described can be modified to allow packing multiple components in different channels but at the same location, as shown below.

[0135] For videos other than 4:4:4, the video component that needs to have the highest resolution can be in the first component, and the other video components can be in subsequent channels. For example: · Alpha component → 0 (Y / Luma) · Lambertian indicator component → 1 (Cb / Blue) · Geometric component → 2 (Cr / Red) The proposed syntax change can leave this choice to the encoder.

[0136] The syntax to be described can be modified to include the following: TIFF2025519644000004.tif18165

[0137] The resulting complete syntax is as follows: TIFF2025519644000005.tif176170TIFF2025519644000006.tif232166TIFF2025519644000007.tif51163

[0138] The purpose of the vps_region_dimension_offset_enabled_flag is to maintain backward compatibility. This can be added to the V3C parameter set, for example, as shown below. If additional flags are introduced, there can be a vps_v3c3e_flag that enables the presence of the vps_region_dimension_offset_enabled_flag at another bitstream position instead. TIFF2025519644000008.tif147163TIFF2025519644000009.tif234164TIFF2025519644000010.tif47162

[0139] The data signal generation device and the rendering device can specifically be realized by a properly programmed processor. Examples of suitable processors are shown below.

[0140] FIG. 3 is a block diagram showing an exemplary processor 300 according to an embodiment of the disclosure. The processor 300 is used to implement one or more processors that implement the data signal generation device of FIG. 1 or the rendering device of FIG. 2. The processor 300 can be any suitable processor type including, but not limited to, a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA) programmed to form a processor, a graphics processing unit (GPU), an application specific integrated circuit (ASIC) designed to form a processor, or a combination thereof.

[0141] The processor 300 may include one or more cores 302. The core 302 can include one or more arithmetic logic units (ALUs) 304. In some embodiments, the core 302 can include, in addition to or instead of the ALU 304, a floating point logic unit (FPLU) 306 and / or a digital signal processing unit (DSPU) 308.

[0142] The processor 300 can include one or more registers 312 communicatively coupled to the core 302. The registers 312 may be implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 312 may be implemented using static memory. The registers can provide data, instructions, and addresses to the core 302.

[0143] In some embodiments, processor 300 may include one or more levels of cache memory 310 communicatively coupled to core 302. Cache memory 310 can provide computer-readable instructions to core 302 for execution. Cache memory 310 can provide data for processing by core 302. In some embodiments, the computer-readable instructions may be provided to cache memory 310 by local memory, such as local memory attached to external bus 316. Cache memory 310 can be implemented using any suitable cache memory type, such as static random access memory, dynamic random access memory, and / or any other suitable memory technology.

[0144] Processor 300 can include a controller 314 that can control input to processor 300 from other processors and / or components included in the system, and / or output from processor 300 to other processors and / or components included in the system. Controller 314 can control the data paths within ALU 304, FPLU 306, and / or DSPU 308. Controller 314 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of controller 314 can be implemented as stand-alone gates, FPGAs, ASICs, or any other suitable technology.

[0145] Registers 312 and cache memory 310 can communicate with controller 314 and core 302 via internal connections 320A, 320B, 320C, and 320D. The internal connections can be implemented as a bus, multiplexer, crossbar switch, and / or any other suitable connection technology.

[0146] Inputs and outputs for the processor 300 may be provided via a bus 316 that may include one or more conductive lines. The bus 316 can be communicatively coupled to one or more components of the processor 300, such as, for example, the controller 314, the cache memory 310, and / or the registers 312. The bus 316 can be coupled to one or more components of the system, such as the aforementioned components BBB and CCC.

[0147] The bus 316 can be coupled to one or more external memories. The external memory may include a read-only memory (ROM) 332. The ROM 332 can be a mask ROM, an electrically programmable read-only memory (EPROM), or any other suitable technology. The external memory can have a random access memory 333. The RAM 333 can be a static RAM, a battery-backed static RAM, a dynamic RAM (DRAM), or any other suitable technology. The external memory can have an electrically erasable programmable read-only memory (EEPROM) 335. The external memory can have a flash memory 334. The external memory can have a magnetic storage device, such as a disk 336. In some embodiments, the external memory may be included in the system.

[0148] The present invention can be implemented in any suitable form including hardware, software, firmware, or any combination thereof. Optionally, the present invention can be at least partially implemented as computer software executed on one or more data processors and / or digital signal processors. Elements and components of embodiments of the present invention can be physically, functionally, and logically implemented in any suitable manner. Indeed, the functions may be implemented in a single unit, in multiple units, or as part of other functional units. Accordingly, the present invention may be implemented in a single unit or may be physically and functionally distributed among different units, circuits, and processors.

[0149] In accordance with standard terminology in this field, the term "pixel" can be used to refer to characteristics associated with a pixel, such as the light intensity, depth, position, etc. of the portion / element of the scene represented by the pixel. The depth of a pixel can be understood to refer to the depth of the object represented by that pixel. Similarly, the brightness of a pixel, i.e., pixel luminance, can be understood to refer to the brightness of the object represented by that pixel.

[0150] The present invention has been described in connection with several embodiments, but it is not intended to be limited to the specific forms described herein. Rather, the scope of the present invention is limited only by the appended claims. Further, although a certain feature may appear to be described in connection with a particular embodiment, one of ordinary skill in the art will recognize that various features of the described embodiments can be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps. The methods of the claims are methods other than methods for performing mental acts. The term "image region" can be replaced (appropriately) with the terms "image segment", "image portion", "image area", or "continuous set of pixels" of the first image. An image region can, in many embodiments, include, for example, 95%, 90%, 70%, 50%, 30%, 10% or 5% or less of the pixels / regions of the entire first image.

[0151] Furthermore, although individually recited, a plurality of means, elements, circuits or method steps may be implemented, for example, by a single circuit, unit or processor. Furthermore, although individual features may be included in different claims, these may, in some cases, be advantageously combined, and being included in different claims does not mean that the combination of features is not feasible and / or not advantageous. Also, including a feature in one category of claims does not mean a limitation to this category, but rather indicates that the feature is equally applicable to other claim categories as required. Furthermore, the order of features in the claims does not mean a specific order in which the features must operate, and in particular, the order of individual steps in method claims does not mean that the steps must be performed in this order. Rather, the steps can be executed in any suitable order. Furthermore, a reference to the singular does not exclude the plural. Thus, references to "a", "an", "first", "second", etc. do not exclude the plural. The reference signs in the claims are provided merely as illustrative examples and should not be construed as limiting the scope of the claims in any way.

[0152] Generally, examples of apparatus and methods for generating data signals are shown by the following embodiments. Embodiments: Embodiment 1. An apparatus configured to generate a data signal for describing a three-dimensional scene, comprising: A generator (101) for generating a data signal to include three-dimensional image data providing a representation of the three-dimensional scene, the three-dimensional image data including at least a first image including light characteristic values of scene points in the scene, the generator; A processor (107) configured to generate a Lambert index of an image region of the first image, the Lambert index indicating a level of Lambertianness of scene points in the image region, the processor; The generator (103) is configured to include the Lambert index in the data signal. Embodiment 2. The apparatus according to Embodiment 1, wherein the processor (107) is configured to generate the Lambertian index based on the three-dimensional image data. Embodiment 3. The apparatus according to Embodiment 1 or 2, wherein the processor (107) is configured to determine light characteristics in different directions for a scene region represented by the image region, and to determine the Lambertian index according to variations in the light characteristics for the different directions. Embodiment 4. The apparatus according to any of the preceding embodiments, wherein the Lambertian index indicates a variation in light emission as a function of the direction of a scene point in the image region. Embodiment 5. The apparatus according to any of the preceding embodiments, wherein the representation of the three-dimensional scene is a multi-view + depth representation, and the first image is an image of a set of multi-view images in the multi-view + depth representation. Embodiment 6. The apparatus according to any of the preceding embodiments, wherein the representation of the three-dimensional scene is a multi-plane image representation, and the first image is a plane in the multi-plane image representation. Embodiment 7. The apparatus according to any of the preceding embodiments, wherein the representation of the three-dimensional scene is a point cloud representation, and the first image includes light characteristic values of at least a partial projection of the point cloud representation onto an image plane. Embodiment 8. The apparatus according to any of the preceding embodiments, wherein the representation of the three-dimensional scene includes at least the first image and projection data indicating a relationship between the position of the light characteristic value of the first image and the position of a corresponding scene point in the three-dimensional scene. Embodiment 9. The apparatus according to any of the preceding embodiments, wherein the generator (101) is configured to receive the input three-dimensional image data of the three-dimensional scene and to select a subset of the input three-dimensional image data for inclusion in the three-dimensional image data of the data signal, and the selection of the subset depends on the Lambertian index. Embodiment 10. The apparatus according to any of the foregoing embodiments, wherein the generator (101) is configured to include the Lambertian index in the chrominance color channel of the image of the three-dimensional image data. Embodiment 11. An apparatus for rendering a view image of a three-dimensional scene, the apparatus comprising: Three-dimensional image data providing a representation of the three-dimensional scene, the three-dimensional image data including at least a first image including light characteristic values for scene points in the scene, and at least one Lambertian index for an image region of the first image, the Lambertian index indicating a level of Lambertianness of the image region; a first receiver (201) configured to receive a data signal including the Lambertian index; A receiver (205) configured to receive a view pose for the three-dimensional scene; A renderer (203) configured to generate a view image for the view pose from the three-dimensional image data depending on the Lambertian index. Embodiment 12. The apparatus according to embodiment 11, wherein the renderer (203) is configured to generate a pixel value of the view image by mixing contributions from a plurality of light characteristic values of the three-dimensional image data projected onto the position of the pixel in the view image, and the mixing of the contribution from the light characteristic value of the image region depends on the Lambertian index. Embodiment 13. The apparatus according to embodiment 11 or 12, wherein the three-dimensional image data includes transparency values, and the renderer (203) is configured to modify at least one transparency value according to the Lambertian index. Embodiment 14. The apparatus according to embodiment 13, wherein the renderer (203) is further configured to modify the at least one transparency value according to a difference between a light direction with respect to the pixel of the at least one transparency value and a direction from the position in the three-dimensional scene represented by the pixel and the view pose. Embodiment 15. A method for generating a data signal describing a three-dimensional scene, the method comprising: A step of generating a data signal for including three-dimensional image data that provides a representation of a three-dimensional scene, the three-dimensional image data including at least a first image including light characteristic values for scene points in the scene; A step of generating a Lambertian index for an image region of the first image, the Lambertian index indicating a level of Lambertian property for scene points in the image region; including the Lambertian index in the data signal. Embodiment 16. A method of rendering a view image of a three-dimensional scene, the method comprising: Receiving a data signal including three-dimensional image data that provides a representation of a three-dimensional scene, the three-dimensional image data including at least a first image including light characteristic values for scene points in the scene, and at least one Lambertian index for an image region of the first image, the Lambertian index indicating a level of Lambertian property of the image region; Receiving a view pose of the three-dimensional scene; generating a view image for the view pose from the three-dimensional image data depending on the Lambertian index. Embodiment 17. A data signal that provides a representation of a three-dimensional scene, the data signal comprising: Three-dimensional image data that provides a representation of a three-dimensional scene, the three-dimensional image data including at least a first image that provides visual information of the scene; At least one Lambertian index in an image region of the first image, the Lambertian index indicating a level of Lambertian property of the image region.

[0153] More specifically, the present invention is defined by the appended claims.

Claims

1. A device configured to generate data signals that describe a three-dimensional scene, A generator for generating a data signal, wherein the three-dimensional image data includes three-dimensional image data that provides a representation of the three-dimensional scene, and the three-dimensional image data includes at least a first image that includes optical characteristic values ​​of scene points in the scene; A processor configured to generate a view-dependency index for an image region of the first image, wherein the view-dependency index indicates the degree of variation of one or more visual characteristics of a scene point in the image region as a function of the viewing direction, A device having, wherein the generator is configured to include the view dependency index in the data signal.

2. The apparatus according to claim 1, wherein the processor is configured to determine the optical properties of a scene region represented by the image region in different directions, and to determine the view dependency index in accordance with the variation in the optical properties in the different directions.

3. The apparatus according to claim 1, wherein the view dependency index shows variations in optical properties as a function of the direction of the scene point in the image region.

4. The apparatus according to claim 1, wherein the representation of the three-dimensional scene is a multi-view + depth representation, and the first image is an image from a set of multi-view images of the multi-view + depth representation.

5. The apparatus according to claim 1, wherein the representation of the three-dimensional scene is a multiplane image representation, and the first image is a plane of the multiplane image representation.

6. The apparatus according to claim 1, wherein the representation of the three-dimensional scene is a point cloud representation, and the first image has optical characteristic values ​​of the projection of at least a portion of the point cloud representation onto an image plane.

7. The apparatus according to claim 1, wherein the representation of the three-dimensional scene includes at least the first image and projection data showing the relationship between the optical characteristic value position of the first image and the position of a corresponding scene point in the three-dimensional scene.

8. The apparatus according to claim 1, wherein the generator is configured to receive input three-dimensional image data for the three-dimensional scene and to select a subset of the input three-dimensional image data to be included in the three-dimensional image data of the data signal, the selection of the subset depending on the view dependency index.

9. The apparatus according to claim 1, wherein the generator is configured to include the view dependency index in the chrominance color channel of the three-dimensional image data.

10. A device for rendering a view image of a three-dimensional scene, wherein the device is A first receiver configured to receive a data signal comprising three-dimensional image data and at least one view-dependent index, wherein the three-dimensional image data comprises at least a first image that provides a representation of the three-dimensional scene and includes optical characteristic values ​​of scene points of the scene, and the view-dependent index for an image region of the first image indicates the degree of variation of one or more visual characteristics of scene points in the image region as a function of the viewing direction, A receiver configured to receive a view pose for the three-dimensional scene, A renderer configured to generate a view image for the view pose from the three-dimensional image data according to the view dependency index, A device having.

11. The apparatus according to claim 10, wherein the renderer is configured to generate the pixel values ​​of the view image by mixing contributions from a plurality of optical characteristic values ​​of the three-dimensional image data projected onto the position of the pixel in the view image, and the mixing of contributions from the optical characteristic values ​​of the image region depends on the view dependency index.

12. A method for generating data signals that describe a three-dimensional scene, A step of generating a data signal that includes three-dimensional image data providing a representation of the three-dimensional scene, wherein the three-dimensional image data includes at least a first image containing optical characteristic values ​​of scene points of the scene, A step of generating a view dependency index for an image region of the first image, wherein the view dependency index indicates the degree of variation of one or more visual characteristics of a scene point in the image region as a function of the observation direction, A method comprising the step of including the view dependency index in the data signal.

13. A method for rendering a view image of a three-dimensional scene, A step of receiving a data signal comprising three-dimensional image data and at least one view-dependent index, wherein the three-dimensional image data comprises at least a first image that provides a representation of the three-dimensional scene and includes optical characteristic values ​​of scene points of the scene, and the view-dependent index for an image region of the first image indicates the degree of variation of one or more visual characteristics of scene points of the image region as a function of the viewing direction. The steps include receiving a view pose for the aforementioned three-dimensional scene, A method comprising the step of generating a view image for the view pose from the three-dimensional image data according to the view dependency index.

14. A computer program that is executed by a computer and causes the computer to perform the method described in claim 12 or claim 13.

15. A data signal providing a representation of a three-dimensional scene, comprising three-dimensional image data and at least one view-dependent index, wherein the three-dimensional image data comprises at least a first image that provides a representation of the three-dimensional scene and includes optical characteristic values ​​of scene points of the scene, and the view-dependent index for an image region of the first image indicates the degree of variation of one or more visual characteristics of scene points of the image region as a function of the viewing direction.