Video Signals and Their Processing
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-03-20
- Publication Date
- 2026-03-18
AI Technical Summary
When the prior art provides -immersive video, the observation space is limited, resulting in errors and inaccuracies in view synthesis when the observer moves, affecting the user experience.
By segmenting the first depth map into a two-dimensional geometric principle with an average area greater than twice the area of the original depth image pixel, and generating a corresponding vertex position map and a lower resolution depth map, a bit stream for the video signal is generated.
Improves the data efficiency of scene description, supports high-quality view switching and synthesis, reduces the data rate of bitstreams, improves user experience, and reduces computing complexity and data rate.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to video signals and processing thereof, and in particular, but not exclusively, to the generation and rendering of images from bitstreams based on 3D meshes. [Background technology]
[0002] 2. Description of the Related Art The variety and scope of image and video applications has increased significantly in recent years, and new services and methods for using and consuming video are continually developed and introduced.
[0003] For example, one service that is gaining popularity is the presentation of image sequences in such a way that the observer can actively interact with the system to change the parameters of the rendering.A very attractive feature in many applications is the ability to change the observer's effective viewing position and direction, e.g., to allow the observer to move and look around the displayed scene.
[0004] Such features may in particular allow a virtual reality experience to be provided to the user, whereby the user can, for example, move around in the virtual environment with (relative) freedom and dynamically change his position and where he is looking. Typically, such extended reality (XR) applications are based on a three-dimensional model of the scene, which is dynamically evaluated to provide a specific requested view. This approach is well known from gaming applications for computers and consoles, for example in the first-person shooter category. Extended reality (XR) applications include Virtual Reality (VR), Augmented Reality (AR) and Mixed Reality (MR) applications.
[0005] An example of a proposed video service or application is immersive video, where the video is played, for example, on a VR headset, to provide a three-dimensional experience. In the case of immersive video, the observer has the freedom to move around while watching the displayed scene, which can be perceived as being seen from different viewpoints. However, in many typical approaches, the amount of movement is restricted to a relatively small area around a nominal viewpoint, which typically corresponds, for example, to the viewpoint from which the video capture of the scene was performed. In such applications, three-dimensional scene information is often provided that allows high-quality viewpoint image synthesis for viewpoints relatively close to the reference viewpoint, but degrades when the viewpoint deviates too much from the reference viewpoint.
[0006] Immersive video is often referred to as six degrees of freedom (6DoF) or three-dimensional oF+ video. MPEG Immersive Video (MIV) is a new standard that uses metadata on top of existing video codecs to enable and standardize immersive video.
[0007] A problem with immersive video is that the viewing space, the three-dimensional space in which an observer has a sufficient quality 6DoF experience, is limited. As the observer moves outside the viewing space, degradation and errors due to the synthesis of view images become more and more noticeable, which can result in an unacceptable user experience. Errors, artifacts and inaccuracies in the generated view images can occur especially because the provided three-dimensional video data does not provide enough information for view synthesis (e.g., deocclusion data).
[0008] For example, immersive video data may be provided in the form of a multi-view, possibly accompanied by a depth data (MVD) representation of the scene: the scene is captured by several spatially distinct cameras, and the captured images can be provided together with a depth map.
[0009] View synthesis for applications such as immersive video is often based on depth maps that provide depth information for the source views. The depth maps may correspond directly to captured images from those views or may be accompanied by texture maps that describe textures for mapping onto three-dimensional objects and features in the scene.
[0010] Many view synthesis algorithms are based on constructing a mesh model from, for example, a depth map that is packed into the pixel data of a video stream. During rendering, based on the mesh model, some source view textures are warped, blended, and / or composited to a target viewpoint to generate an image for that target viewpoint.
[0011] Such operations require a lot of processing, including, for example, warping many vertices of a mesh model to appropriate two-dimensional positions for a particular target viewpoint. The process lends itself to parallel processing, and efficient implementations using massively parallel processing structures are widespread. For example, efficient graphics pipelines have been developed and implemented in graphics cards for standard and gaming personal computers.
[0012] In traditional graphics processing, such as games and many full virtual reality experiences, rendering may be based on a relatively static mesh representation of the scene that is updated only relatively slowly to reflect dynamic changes in the scene. In contrast, however, for immersive video applications, new depth maps are often provided with each new frame, and these need to be converted (possibly implicitly) into a mesh representation as part of the rendering.
[0013] Thus, for immersive video, very high computational resources are often required for rendering the received video. Furthermore, the data rate of the video signal containing the three-dimensional information for rendering tends to be undesirably high. Approaches that attempt to mitigate these issues tend to result in degradation of performance, user experience and / or image quality. Thus, conventional approaches for three-dimensional / immersive video tend to be suboptimal. Summary of the Invention [Problem to be solved by the invention]
[0014] Thus, improved approaches would be advantageous, particularly approaches that provide improved behavior, increased flexibility, improved immersive user experience, reduced complexity, easier implementation, improved synthetic image quality, improved rendering, improved and / or easier view synthesis for different view poses, improved computational efficiency, improved quality, reduced data rate, more accurate representation of the scene, and / or improved performance and / or behavior.
[0015] Accordingly, the Invention seeks to preferably mitigate, reduce or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]
[0016] According to one aspect of the invention, there is provided an apparatus for generating a video signal including a bitstream, the apparatus comprising: a receiver configured to receive a first depth map for at least a first video frame of a plurality of video frames; a divider configured to divide at least a portion of the first depth map into two-dimensional geometric primitives having an average area that is more than twice a pixel area of the first depth map, the divider comprising: determining vertex positions in the first depth map for vertices of the geometric primitives; a map generator configured to generate a vertex position map indicative of vertex positions in the first depth map; a depth map generator configured to generate a second depth map including depth values of the first depth map at vertex positions of the first depth map; and a bitstream generator configured to generate a video bitstream including the second depth map and the vertex position map for the first video frame, the divider configured to generate the vertex positions dependent on the depth values of the first depth map.
[0017] The present invention can allow for improved data characterizing the scene being generated. This approach allows for improved image synthesis for different view poses based on the generated bitstream. In many embodiments, data is generated that allows for efficient, high-quality view shifting and synthesis.
[0018] The present invention can provide an improved user experience in many embodiments and scenarios. This approach allows for improved immersive video or XR / AR / VR / MR applications, for example. This approach can support an improved user experience in many embodiments with the possibility of rendering improved quality view images from a video signal and bitstream.
[0019] This approach can, in many embodiments, reduce the data rate of the bitstream. This approach allows for improved spatial information at a given data rate. This approach can, in many embodiments, provide improved support for rendering and can allow for reduced complexity and / or higher computational efficiency for generation and / or rendering purposes. This approach can provide improved suitability for processing by highly efficient and optimized features, such as graphics pipelines that use a high degree of parallelism.
[0020] In some embodiments, the receiver may further receive a texture map linked to the depth map, and the bitstream generator may be configured to include the texture map in the bitstream.
[0021] The division can be a tessellation in many embodiments. The two-dimensional geometric primitives are polygons in many embodiments, often 3, 4 or 6 sided polygons. The vertex positions are the positions, often pixel positions, in the first depth map of the vertices / corners of the geometric primitives / polygons.
[0022] The two-dimensional geometric primitives may have an average area that is 3, 4, 8, 16 or more times the pixel area of the first depth map in many embodiments. The second depth map contains fewer depth values / pixel than the first depth map, often 2, 3, 4, 8, 16 or more times fewer.
[0023] The divider may be configured to generate vertex positions to reduce depth change in the geometric primitive, for example, an average depth change in the geometric primitive may be reduced.
[0024] According to an optional feature of the invention, the divider is configured to determine the vertex positions by offsetting a reference vertex position of a reference division of the depth map, the offset of a first vertex being dependent on a first depth map value of a pixel in the geometric primitive of the first vertex.
[0025] This can provide improved and / or accelerated operation and / or performance in many embodiments: it can facilitate determination of vertex positions / two-dimensional primitives and can often allow for more efficient representation.
[0026] The reference division may correspond / equivalent to a reference set of vertex positions / a reference pattern of reference vertex positions / a reference pattern of geometric primitives.
[0027] The offset of the reference vertex positions can be such that the average area of the geometric primitive remains unchanged.
[0028] The number of vertices represented by the vertex position map can be less than or equal to the number of reference vertex positions.The number of vertices represented by the vertex position map can be the same as the number of reference vertex positions.
[0029] In accordance with an optional feature of the invention, the reference vertex positions of the reference division form a regular grid.
[0030] This can provide improved and / or accelerated operation and / or performance in many embodiments. In many embodiments, the reference grid can provide an improved or optimal starting point for adapting to a particular depth.
[0031] In many embodiments, the new vertex positions may be determined by a relative shift of the vertex positions relative to a reference vertex position of the relative vertex position pattern.
[0032] In accordance with an optional feature of the invention, the vertex position map indicates vertex positions as relative positions to a reference vertex position of a reference division.
[0033] This can provide advantageous operations in many embodiments, often allowing substantially reduced data rates and allowing vertex position maps to be generated that are highly suitable for encoding and compression, including using existing video encoding techniques.
[0034] According to an optional feature of the invention, the offset includes a first offset along a first predetermined direction, the first offset being determined as an offset at which a change in depth map values along the first predetermined direction satisfies a criterion.
[0035] This can provide a particularly advantageous operation in many embodiments, which allows for lower complexity and / or improved computational efficiency in many embodiments. This approach can provide efficient and high performance adaptation to many different depth maps. The divider, in some embodiments, can determine step changes for different offsets and determine the offsets as a function of the depth variation for the different offsets.
[0036] The first predetermined direction may be a horizontal direction and / or a vertical direction.
[0037] According to an optional feature of the invention, the offset includes a second offset along a second predetermined direction, the second offset being determined as an offset for which a change in depth map values along the second predetermined direction satisfies a criterion, the map generator is configured to generate a vertex position map to indicate the offset along the first predetermined direction and further configured to generate an additional vertex position map to indicate the offset along the second predetermined direction, and the bitstream generator is configured to include the additional vertex position map in the bitstream.
[0038] This can provide a particularly advantageous operation in many embodiments, which allows for lower complexity and / or improved computational efficiency in many embodiments. This approach can provide efficient and high performance adaptation to many different depth maps. The divider, in some embodiments, can determine step changes for different offsets and determine the offsets as a function of the depth variation for the different offsets.
[0039] The first and second predetermined directions may be a horizontal direction and a vertical direction.
[0040] In accordance with an optional feature of the invention, the first predetermined direction is a horizontal direction.
[0041] In accordance with an optional feature of the invention, the first predetermined direction is a vertical direction.
[0042] In accordance with an optional feature of the invention, the offset is constrained to be below an offset threshold.
[0043] According to an optional feature of the invention, the bitstream generator comprises a video encoder configured to encode the vertex position map using video encoding.
[0044] This can provide more efficient encoding in many embodiments, and / or can allow reuse of existing functionality that is also used for other purposes.
[0045] According to an optional feature of the invention, the divider is configured to determine vertex positions to align the geometric primitives with depth transitions in the depth map.
[0046] This can provide particularly advantageous operation in many aspects and scenarios: it can, in particular, reduce artifacts and inaccuracies that can result from depth transitions.
[0047] According to another aspect of the present invention, there is provided an apparatus for rendering a bitstream of a video signal comprising: a receiver for receiving a video bitstream, the video bitstream including, for at least a first frame, a first depth map and a vertex position map indicating vertex positions of depth values of the first depth map in a second map having a higher resolution than the first depth map, the vertex positions defining two-dimensional geometric primitives forming a division of the second map; a renderer configured to render a video image from the video bitstream, the renderer comprising: a mesh generator configured to generate a mesh representation of a scene, the mesh being formed by three-dimensional geometric primitives formed by projections of the two-dimensional geometric primitives into a three-dimensional space, a projection of a first vertex of the two-dimensional geometric primitive depending on a vertex position of a first vertex in the second depth map and a depth value in the first depth map for the first vertex; and a rendering unit configured to render an image based on the mesh representation.
[0048] The present invention can enable improved image generation in many embodiments and scenarios. This approach allows for improved image synthesis for different view poses.
[0049] The present invention can provide an improved user experience in many embodiments and scenarios. This approach allows for improved immersive video or XR / AR / VR / MR applications, for example. This approach can support an improved user experience in many embodiments with the possibility of rendering improved quality view images from a video signal and bitstream.
[0050] This approach can reduce the data rate of the bitstream in many embodiments. This approach can often reduce rendering artifacts and inaccuracies, especially in relation to depth steps / transitions. This approach can allow for reduced complexity and / or increased computational efficiency in many embodiments. This approach can provide improved suitability for processing by highly efficient and optimized features, such as graphics pipelines that use a high degree of parallelism.
[0051] In some embodiments, the bitstream further includes a texture map linked to the depth map, and the rendering unit can be configured to render the image based on the texture map.
[0052] The rendering unit can be configured to render an image to correspond to a view image from a view pose. The rendering can include evaluating a mesh representation of the view pose.
[0053] According to an optional feature of the invention, the vertex position map indicates the vertex positions as relative vertex positions with respect to reference vertex positions of a reference partition, and the mesh generator is configured to determine absolute vertex positions in the second depth map as a function of the relative vertex positions and the reference vertex positions.
[0054] This can provide improved and / or accelerated operation and / or performance in many embodiments.
[0055] According to an optional feature of the invention, the bitstream includes indices of the reference division, and the mesh generator is configured to determine the reference vertex positions in response to the indices of the reference division.
[0056] This can provide improved and / or accelerated operation and / or performance in many embodiments.
[0057] According to another aspect of the present invention, there is provided a method for generating a video signal including a bitstream, the method comprising the steps of: receiving a first depth map for at least a first video frame of a plurality of video frames; dividing at least a portion of the first depth map into two-dimensional geometric primitives, the two-dimensional geometric primitives having an average area that is more than twice a pixel area of the first depth map, the dividing including determining vertex positions in the first depth map for vertices of the geometric primitives; generating a vertex position map indicating vertex positions in the first depth map; generating a second depth map including depth values of the first depth map at vertex positions in the first depth map; and generating a video bitstream for the first video frame to include the second depth map and the vertex position map, wherein generation of the vertex positions is dependent on the depth values of the depth map.
[0058] According to another aspect of the present invention, there is provided a method for rendering a bitstream of a video signal, the method comprising the steps of: receiving a video bitstream, the video bitstream including, for at least a first frame, a first depth map and a vertex position map indicating vertex positions of depth values of the first depth map in a second map having a higher resolution than the first depth map, the vertex positions defining two-dimensional geometric primitives forming a division of the second map; rendering a video image from the video bitstream, the rendering step comprising the steps of generating a mesh representation of a scene, the mesh being formed by three-dimensional geometric primitives formed by projections of the two-dimensional geometric primitives into a three-dimensional space, the projection of a first vertex of the two-dimensional geometric primitive depending on a vertex position of a first vertex in the second depth map and a depth value in the first depth map for the first vertex; and rendering an image based on the mesh representation.
[0059] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter.
[0060] Embodiments of the present invention will now be described, by way of example only, with reference to the drawings in which: [Brief description of the drawings]
[0061] [Figure 1] FIG. 1 illustrates an example of elements of a server-client system. [Diagram 2] 1 illustrates example elements of an apparatus according to some embodiments of the present invention. [Diagram 3] 1 illustrates example elements of an apparatus according to some embodiments of the present invention. [Figure 4] 1 illustrates an example of a depth map representation of a view of a scene. [Diagram 5] 1 illustrates an example of a depth map representation of a view of a scene. [Figure 6] 1 illustrates an example of a depth map representation of a view of a scene. [Figure 7] 1 illustrates an example of a depth map representation of a view of a scene. [Figure 8] 1 illustrates an example of a depth map representation of a view of a scene. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0062] The following description focuses on immersive video applications, however, it will be appreciated that in other embodiments, the concepts, implementations, techniques and principles described may be applied to other video applications.
[0063] The capture, delivery and display of three-dimensional video has become increasingly popular and desirable in several applications and services. A particular approach is known as immersive video, which typically involves the provision of views of real-world scenes, and often real-time events, and allows for small observer movements, such as relatively small head movements and rotations. For example, a real-time video broadcast of a sporting event can provide a user with the impression of sitting in the stands watching the sporting event. The user can, for example, look around and have a natural experience similar to that of a spectator present at that location in the stands. In recent years, there has been an increasing popularity of display devices with position tracking and three-dimensional interaction that support applications based on three-dimensional capture of real-world scenes. Such display devices are well suited for immersive video applications that provide an enhanced three-dimensional user experience.
[0064] The following description focuses on immersive video applications, although it will be understood that the principles and concepts described can be used in many other applications and embodiments.
[0065] In many approaches, the immersive video can be provided locally to the viewer, for example by a stand-alone device that does not use or even have access to a remote video server. In other applications, however, the immersive application can be based on data received from a remote or central server. For example, video data can be provided to a video rendering device from a remote central server and processed locally to generate the desired immersive video experience.
[0066] 1 shows an example of an immersive video system in which a video rendering client 101 cooperates with a remote immersive video server 103 over a network 105, such as the Internet. The server 103 may be configured to support a potentially large number of client devices 101 simultaneously.
[0067] The server 103 can support immersive video experiences, for example, by transmitting a video signal including three-dimensional video data describing a real-world scene. The data can specifically describe the visual features and geometric properties of the scene generated from real-time capture of the real world by a set of (possibly three-dimensional) cameras.
[0068] To provide such services for real-world scenes, the scene is typically captured from different positions and different camera capture poses are used. As a result, the relevance and importance of multi-camera capture and, for example, 6DoF (six degrees of freedom) processing is rapidly increasing. Applications include live concerts, live sports, and telepresence. The freedom to choose one's viewpoint enriches these applications by providing a greater sense of realism than regular video. Furthermore, one can consider immersive scenarios where observers can move through and interact with the live-captured scene. For broadcast applications, this would require real-time view synthesis in the client device. View synthesis introduces errors, and these errors vary depending on the implementation details of the algorithm.
[0069] A common approach to provide the relevant information to the rendering side is to provide a video signal that can provide a depth map and a texture map for each video frame. The renderer, in such an embodiment, can generate a mesh representation based on the depth map and then generate a view image by projecting the mesh into the appropriate location within the view image / port for a given view pose. The renderer can warp the texture map associated with the mesh to synthesize a visible image for a given view pose.
[0070] The server 103 of Figure 1 is configured to generate a video signal comprising data describing a scene by a sequence of frames, with depth map and texture data being provided for at least some of the frames. The client 101 is configured to receive and process this video signal to provide a view for a given view pose. The client 101 can generate an output video stream or display drive signal that dynamically reflects changes in view pose, thereby providing an immersive video experience in which the displayed view adapts to changes in viewing / user pose / position.
[0071] In the field, the terms configuration and pose are used as general terms for position and / or orientation. For example, a combination of the position and orientation of an object, a camera, a head or a view may be referred to as a pose or configuration. A configuration or pose index may thus include six values / components / degrees of freedom, each value / component typically describing an individual characteristic of the position / location or orientation / direction of the corresponding object. Of course, in many situations, a configuration or pose may be considered or represented with fewer components, for example when one or more components are considered fixed or irrelevant (e.g., four components may provide a complete representation of the pose of an object if all objects are considered to be at the same height and have a horizontal orientation). In the following, the term pose is used to refer to a position and / or orientation that can be represented by one to six values (corresponding to the maximum possible degrees of freedom). The term pose may be replaced by the term configuration. The term pose may be replaced by the term position and / or orientation. The term pose can be replaced by the terms position and orientation (if pose provides both position and orientation information), by the term position (if pose (possibly) provides position information), or by orientation (if pose (possibly) provides orientation information).
[0072] Typically, rendering a video image for a given view pose based on a depth map and a texture map involves decoding a still image or video frame into GPU memory, fetching / sampling image space vertex data from the still image or video frame in a vertex shader, constructing an image space uniform coordinate vector for each vertex, and transforming / warping each vertex into the image space of the target view. The vertices are associated with a predefined mesh topology that is typically created in an initialization phase. The mesh topology (e.g., triangles / quads or other) determines how the fragment shader following the vertex shader phase needs to warp the associated textures to the target view. Since the image or video frame is already in GPU memory, this fragment shader samples directly from it to generate the result view.
[0073] Traditionally, depth maps are required to be high resolution in order to obtain a fine mesh that accurately represents the geometric properties. Typically, the depth map has at least the same resolution as the texture map. This provides an accurate representation of the spatial properties of the scene with a detailed high-resolution mesh model. However, this approach requires high computational resources and in practice, such high accuracy typically does not allow real-time rendering due to the large number of mesh triangles (e.g., two triangles per pixel).
[0074] To address this, a coarser mesh model with larger triangles can be used, and these larger triangles can be generated by subsampling the high-resolution depth map, or by simply packing the lower-resolution depth map within the video frame in anticipation of using a lower-resolution mesh that combines the triangles of the finer mesh generated from the depth map. However, such conventional approaches can result in degradation of rendering quality, for example near depth transitions.
[0075] Below, we describe approaches that may provide improved operation and / or performance in many scenarios. In many cases, improved rendering quality may be achieved.
[0076] Figure 2 illustrates elements of an exemplary apparatus for generating a bitstream including a video signal representing a scene. The apparatus of Figure 2 corresponds specifically to the server 103 of Figure 1 and will be described in this context. Figure 3 illustrates elements of an exemplary apparatus for rendering a video bitstream of a video signal from a video bitstream generated by the apparatus of Figure 2. This apparatus can be an end-user device or a client that receives an audiovisual data stream from a remote device, such as the apparatus of Figure 2. In particular, the apparatus of Figure 3 is an example of an element of the client 101 of Figure 1 and the following description will focus on this context.
[0077] 2, and thus in a particular example, server 102, comprises a receiver 201 that receives data including a depth map and a linked texture map for one or more video frames. In some embodiments, multiple depth maps may be linked with the same texture map, e.g., a texture map may not be provided for every frame and / or a depth map may not be provided for every frame.
[0078] The server may specifically receive an input video signal having time-sequenced video frames represented by a depth map and a texture map. Such an input video signal may be received from a processing unit that generates the depth map and texture map from a live capture, such as from multiple images of a scene captured from different view poses. In other embodiments, the input data may be received from an external or internal source that is configured to generate the depth map and texture map from a model of the scene, for example. It should be understood that the generation and use of depth maps and linked texture maps are known to those skilled in the art of three-dimensional video, and therefore the description will not describe this in further detail for the sake of clarity and brevity.
[0079] The depth map may be provided at a relatively high resolution, and in particular, in many embodiments, the depth map and the texture map may be provided at the same resolution. The device of Fig. 2, hereinafter referred to as the encoding device, is configured to generate an output video signal in which one, several, and often all, video frames are represented by a depth map and a texture map. Furthermore, one, several, or all of the output depth maps are generated to have a lower resolution than the corresponding depth map. The encoding device uses a particular technique for reducing the resolution of the depth map, so that an improved representation enabling improved rendering can be achieved in many scenarios.
[0080] The depth map is provided from the receiver 201 to a divider 203 configured to divide at least a portion of the depth map into two-dimensional geometric primitives. The two-dimensional geometric primitives are typically polygons, such as triangles or squares, among others. In most, but not necessarily all, embodiments, the division is into identical geometric primitives. However, in some embodiments, the division can be into different geometric primitives.
[0081] The partitioning is typically applied to the entire depth map, although in some embodiments it may be for a portion of the depth map. In some embodiments, the partitioning allows for some gaps or holes to exist between the defined geometric primitives (e.g., the partitioning does not perfectly align and join to cover the entire depth map region, but may use a polygon or set of polygons with small gaps / holes). However, these are often too small to have any significant impact on performance.
[0082] The partitioning can correspond to spatial two-dimensional image space regions that overlap each other such that the resulting warped meshes overlap when rendered for a given view pose. Each distinct spatial two-dimensional image space region can have different primitive parameters. For example, the size of the mesh primitives can vary between regions.
[0083] In many embodiments, the division is a tessellation of at least a portion of the depth map, typically the entire depth map. The tessellation may be, for example, triangular, square, or hexagonal. The tessellation is typically a regular tessellation, but in many embodiments may be an irregular tessellation.
[0084] Divider 203 is typically configured to perform a division, specifically a tessellation, into polygons having the same number of sides, such as triangles or squares, but where the shape of each polygon varies. Indeed, in many embodiments, divider 203 may be configured to perform the division / tessellation by first performing a standard regular division / tessellation, which is then modified to provide the final division / tessellation.
[0085] The divider 203 may be configured to perform the division / tessellation by determining vertices in the depth map for a geometric primitive; in particular, (new) vertex positions in the depth map may be determined to define the geometric primitive.
[0086] Geometric primitives are typically polygons, often triangles or squares. Vertex points are usually the corners of polygons, such as the corners of a triangle or a quadrilateral. Vertices are shared between multiple polygons to provide a tessellation. For example, for a triangular geometric primitive, most vertices can be shared between six or eight triangles, and for a quadrilateral geometric primitive, most vertices can be shared between four quadrilaterals. Thus, changing the position of a vertex can change multiple geometric primitives simultaneously.
[0087] In many embodiments, the vertex positions are pixel positions of the depth map, and thus each vertex position is determined as the position of a pixel in the depth map. The position resolution of the vertices may be equal to the pixel resolution of the depth map.
[0088] The divider 203 is configured to determine geometric primitives / polygons such that these geometric primitives / polygons are on average larger than a pixel, in fact larger than two pixel areas. Thus, the vertex positions are on average more than two pixels apart for at least one side of the polygon, and often at least twice the pixel distance for at least one side. Typically, the number of geometric primitives determined for the portion of the depth map to be divided / tessellated is typically 2, 3, 5 or even 10 or more times less than the number of pixels of the portion to be divided. In many embodiments, the number of geometric primitives generated by the division is 2, 3, 5 or even 10 times less than the number of pixels in the input depth map. The number of distinct vertices of the geometric primitives is in many embodiments 2, 3, 5 or even 10 or more times less than the number of pixels in the (portion of) the input depth map to be divided / tessellated.
[0089] This division therefore provides a decimation / resolution reduction compared to the division into geometric primitives that results from using pixels as vertices.
[0090] The divider 203 is configured to generate vertex positions that depend on the depth values of the depth map. Thus, rather than generating a predetermined or fixed division into polygons / primitives, an adaptive and flexible approach is applied in which the division into polygons / primitives is made to reflect the depth changes of the depth map. In many embodiments, the divider 203 is configured to determine vertex positions to align polygons with depth transitions in the depth map. In some embodiments, the depth transition alignment can be done with sub-pixel accuracy. The divider 203 can be configured to determine a fixed number of polygons and vertices, but the size and position of the polygons are changed to align the polygons with the depth transitions.
[0091] The divider 203 is coupled to a depth map generator 205 that receives data from the divider 203 describing the determined vertex positions in the input depth map. The depth map generator 205 is further coupled to a receiver 201 that receives the input depth map therefrom.
[0092] The depth map generator 205 is configured to generate an output depth map that includes the depth values of the input depth map at the determined vertex positions in the first depth map. Thus, the depth values at the selected vertex positions are included in the output depth map, but the depth values for other positions / pixels in the input depth map are not included. Thus, the output depth map includes a subset of the depth values of the input depth map, and thus spatial decimation / resolution reduction is achieved.
[0093] The depth map generator 205 is further coupled to a position map generator 207 to which indexes of the determined new vertex positions are provided. The position map generator 207 is configured to generate a vertex position map that indicates vertex positions in the input depth map. In many embodiments, the vertex position map may include a position index for each vertex position. Thus, in many embodiments, the resolution of the output depth map and the vertex position map may be the same, and in fact, each depth value of the output depth map may be linked with a position value in the vertex position map, or vice versa. In some embodiments, the vertex position map may have vertex positions that have a higher resolution than the original vertex positions (which generally tend to correspond to discrete depth map pixel positions).
[0094] The apparatus of FIG. 2 further comprises a bitstream generator 209 configured to generate a video bitstream for the video signal.
[0095] The bitstream generator 209 may be specifically configured to generate a video stream to include, for a given frame, a texture map received from the receiver 201 (often the same directly as received by the receiver 201), an output depth map received from the depth map generator 205, and a vertex position map received from the position map generator 207. In some embodiments, all three types of data may be provided for all frames of the video signal. In other embodiments, one, more, or all of the three types of data may be provided only for some of the frames. In such instances, data for frames for which no explicit data is included may be generated on the rendering side from other frames for which data is provided. For example, in some embodiments, values from other frames may be used directly or may be used for interpolation or extrapolation.
[0096] The bitstream generator 209 can be configured to generate the bitstream to conform to any suitable structure, arrangement and / or standard. For example, the bitstream can be generated according to a known video standard, e.g., the vertex position map is included in a metadata section. Furthermore, the bitstream generator 209 can implement a suitable encoding algorithm. In particular, the texture map and the depth map can be encoded using known video encoding algorithms. In many embodiments, the vertex position map can be encoded using a video encoding scheme, potentially including a lossy video encoding scheme.
[0097] The following description, for clarity and brevity, focuses on an embodiment in which divider 203 is configured to perform a tessellation of the entire input depth map into triangles, and therefore each geometric primitive / polygon has three vertices, however, it will be understood that this is merely an example and that other polygons / primitives may also be used, only a portion of the depth map may be split and tessellated, and / or the splitting may leave gaps or holes in some cases.
[0098] In many embodiments, the determination of vertex positions, and therefore the division / tessellation, is based on a reference division of the depth map. The reference division is given by a reference pattern of vertex positions, and therefore equivalently, the determination of vertices is based on a reference pattern of vertex positions in the input depth map.
[0099] In many embodiments, the reference division / location pattern can be a regular grid. For example, the reference vertex locations can be given as locations with a predefined distance between nearest neighbors in each of the horizontal and vertical directions. The distances can be given as pixel locations.
[0100] An example of a reference division / pattern based on a regular grid is shown in FIG. 4. In this example, the grid is formed by the vertex positions of every fourth input depth map pixel, both horizontally and vertically. In this example, the reference geometric primitives are triangles, and two triangles are formed per 4×4 area. Thus, in the example of FIG. 4, the reference division corresponds to a regular mesh with two triangles per 4×4 pixel block. In the example of FIG. 4, the depth map includes values reflecting the depth steps / transitions that occur between pixels shaded in gray and pixels not shaded in gray.
[0101] The reference division and mesh thus provide an undersampling of the input depth map, thereby providing a spatial decimation of the depth map. The reference pattern can thus provide a significant reduction in the number of vertices required to represent the corresponding mesh compared to that required by using all depth values to form the depth map. Thus, a data-efficient representation of the underlying scene can be achieved. However, the coarse division and mesh results in a less accurate representation of spatial characteristics. In the apparatus of FIG. 2, the determined reference vertex positions can be offset relative to the reference vertex positions. The vertex position offset is determined based on the depth values of the input depth map. This approach can mitigate the effects of resolution reduction in many scenarios.
[0102] In many embodiments, the new vertex position is determined by a relative shift of the vertex position with respect to a reference vertex position in a reference pattern / division. The divider 203 can be configured to determine the offset for a given vertex depending on the depth values of the input depth map for pixels in at least one of the geometric primitives that contain the given vertex. In many embodiments, the divider 203 can take into account depth values in multiple primitives that contain the given vertex, and in fact can take into account at least one of the values of at least two primitives that are (partially) determined by the given vertex. In some embodiments, the offset can also take into account depth values that belong to primitives that are not formed by the given vertex that is offset. For example, if the offset of a vertex results in other vertices being shifted, this may affect depth variations in other primitives, and these variations can be taken into account.
[0103] As yet another example, in some embodiments, the positioning of the vertices can be further based on texture variations within a region. For example, regions that are reasonably flat (according to any suitable criteria) can be determined, with each region having a similar texture, but different regions having different (degrees of) texture. Thus, regions that do not have large depth transitions, but have different textures, can be determined. The vertices can be positioned such that this results in distinct vertices. This effect becomes visible after view synthesis, since typically more textured regions are more susceptible to depth artifacts.
[0104] It will be appreciated that the particular algorithms and criteria for determining the vertex positions, and specifically the offsets, in the output depth map will vary depending on the requirements and preferences of a particular implementation.
[0105] In many embodiments, the divider 203 is configured to determine offset vertex positions to align a primitive with a depth transition in the depth map. For example, a line corresponding to a depth transition in the depth map or a contour line of an area with a high depth gradient can be determined, such as by determining a line where the depth gradient exceeds a given threshold (it will be appreciated that many different techniques for determining a depth transition line in a depth map are known to those skilled in the art). A vertex offset can then be determined based on the determined depth transition line. For example, a given vertex can be offset to increase the size of a primitive under the criterion that the primitive will not extend beyond the depth transition line and will receive a maximum offset.
[0106] As another example, the offset can be determined as a function of the depth variation of the depth values within the primitive. For example, a depth change measure can be determined for a given primitive for different possible offsets. Then, the offset that results in the largest primitive with a depth change below a given threshold is selected. As another example, the depth change measure can be determined for all primitives affected by the offset. Then, a cost measure can be determined for each primitive as a monotonically increasing function of the size of the primitive and a monotonically increasing function of the depth change (e.g., for a given offset, a cost measure can be determined for each primitive by multiplying the size of the primitive by the depth change measure of the primitive). Then, the cost measure can be determined by combining the cost values of all primitives for a given offset, and the offset with the smallest cost value can be selected. Such an approach can result in many scenarios in which the depth transitions and variations are concentrated in smaller primitives, while offsets that tend to be relatively flat for larger primitives are selected.
[0107] It will be appreciated that different depth change measures may be used, for example, depth change, maximum depth difference, difference between different percentiles, etc.
[0108] Figure 5 shows an example of how the reference positions of the example of Figure 4 can be offset to generate primitives that more closely match the depth variations and depth transitions in the image. In this example, two different depth levels are indicated by the gray shades of the pixels. As can be seen, smaller primitives are generated around the depth transitions and larger primitives are generated away from the depth transitions.
[0109] The position map generator 207, in many embodiments, can be configured to indicate vertex positions in the input depth map as relative positions to a reference vertex position of the reference division. For example, the vertex position map can include an offset index that indicates the size and possibly the direction of the determined offset.
[0110] In many embodiments, vertex positions in the input depth map are indicated by pixel positions, and offsets may simply be given as an integer number of pixels, or possibly sub-pixels. Furthermore, the maximum offset is often a relatively small number of pixels, and the vertex position map may therefore be given as a map of relatively small integers. In many embodiments, a relatively large number of offsets may be zero (e.g., in larger, relatively flat regions where the benefit of shifting vertex positions is very low). The vertex position map, in many embodiments, may be a form of data that is highly suitable for efficient image or video compression as part of its encoding and inclusion in an output bitstream.
[0111] In some embodiments, the depth map generator 205 can be configured to perform an offset in a predetermined direction. For example, the depth map generator 205 can be configured to determine an offset in a horizontal or vertical direction. In some embodiments, the offset can be in the horizontal direction of the input depth map. In some embodiments, the offset can be in the vertical direction of the input depth map. For example, FIG. 5 is generated by introducing an offset only in the horizontal direction to the regular grid of FIG. 4.
[0112] By restricting the offset in a given direction, the offset value is purely one-dimensional and can therefore be represented in the vertex position map by a simple scalar value. Moreover, implementing the offset specifically in the horizontal direction can greatly reduce complexity while achieving significant benefits. Since humans typically move primarily in the horizontal plane, spatial rendering is especially beneficial in better depth step alignment in the horizontal direction.
[0113] In many embodiments where the offset is along a predefined direction, the offset can be determined as the offset for which the depth map value change meets a criterion for different offsets along the first predefined direction. For example, the depth map generator 205 can determine a measure of depth change of the current offset relative to a minimum offset to increase the offset within a range. For example, to increase (or decrease) the offset, a maximum depth step can be determined, for example, as the difference between the maximum and minimum depth values for all offsets smaller than the current offset. Then, the maximum offset for which the depth step is below a threshold can be selected. This range can be a range that includes only negative offsets, positive offsets, or both negative and positive offsets.
[0114] FIG. 6 shows an example of how such an approach can result in the vertex positions of FIG. 5. In this example, vertex position A is the initial vertex position that does not change. A new vertex position B is determined by offsetting the next reference vertex position in the horizontal direction by a certain horizontal offset. This horizontal offset is increased until a depth step exceeding a given threshold is detected. The new vertex position is then determined as the last offset before that depth step. For the next horizontal position, the offset is increased in the opposite direction, so that a new vertex position C is determined. D is at the last position of the row and is therefore not offset. A corresponding approach can be applied to the remaining vertex positions, specifically the other rows of the reference pattern.
[0115] Figure 6 also shows an example of a resulting decimated output depth map (b) that includes depth values for the 12 vertex positions A-L determined in a 9x13 example depth map (it will be understood that typical depth maps tend to be much larger and the figure is intended merely as an illustrative example). Additionally, Figure 6 shows an example of a corresponding vertex position map (c) that shows the horizontal shift of each vertex position by a simple signed integer.
[0116] Thus, Figure 6 shows vertices from A to L, and the arrows indicate how the vertices are horizontally displaced from their original positions. This shows how a lower resolution depth map is constructed by simply storing the depth values at the transformed vertex positions of the original high resolution depth map. A vertex position map / transformation map with the same resolution as the output depth map is generated, which stores the transformations / offsets applied to the vertices of a rectangular shaped mesh. Since the resulting map is a two-dimensional array, it can be efficiently encoded, packed and compressed using various video encoding algorithms.
[0117] As a particular example, divider 203 can be configured to determine the offset vertex positions using an algorithm that shifts vertices horizontally based on an initial grid. For example, the following approach can be used (where Q denotes the grid distance between vertices, e.g., Q=4 for a mesh with a 4x4 pixel grid, and D denotes the depth values of the input depth map): 1. For each vertex i,
number
number
[0118] This approach attempts to shift pixels horizontally until a new vertex position is found that is adjacent to the depth transition.
[0119] In some embodiments, the depth map generator 205 can be further configured to determine an offset (component) in different predefined directions. Thus, in many embodiments, an offset can be determined in two different predefined directions. The determination can be sequential and / or independent in many embodiments. The depth map generator 205 can be configured to first determine an offset (component) along a first predefined direction, for example, as described above. Then, typically starting from the position offset in the first direction, a second offset (component) along a second predefined direction is determined. This determination can, for example, use the same approach as the first direction, but now the offset is along the second direction. The result can be considered as a single offset with two components, or equivalently as two offsets in the two respective directions.
[0120] The two predetermined directions can typically be the horizontal and vertical directions, for example, the offset along the row direction is determined first, followed by the offset along the column direction.
[0121] An example of applying the described approach in the horizontal direction is shown in FIG. 7. In this example, the depth map generator 205 starts with the regular reference grid of FIG. 4. Then, offsets along the horizontal direction are determined for the vertices as described above, resulting in the modified vertex positions of FIG. 5. The same approach was then used to determine offsets along the vertical direction for the vertices of FIG. 5. This results in the vertex position pattern as shown in FIG. 7. As can be seen, an improved adaptation of the triangular primitives to depth variations in the depth map, especially to depth steps, is achieved. The triangles around the depth steps are small and align close to the depth transitions. This is compensated by the triangles in the flat regions being significantly larger. In fact, the resulting mesh pattern covers the entire depth map using the same number of primitives and vertices as the reference pattern / division. Thus, the average size of the primitives does not change with the deformation of the regular division / tessellation.
[0122] In many embodiments, divider 203 is configured to maintain the same number of primitives / vertex positions for the output depth map / vertex position map as the reference division / vertex position pattern.
[0123] The position map generator 207 is configured to generate information reflecting two-dimensional offsets. In some embodiments, a vertex position map can be generated that includes two offset components along two predetermined directions for each pixel.
[0124] However, in many embodiments, the position map generator 207 may be configured to generate separate vertex position maps for two predetermined directions. Thus, in addition to the vertex position map indicating the offset in a first direction, e.g., the horizontal direction, the position map generator 207 may further generate an additional vertex position map indicating the offset in a second direction, e.g., the vertical direction. Thus, rather than a single combined vertex position map, two separate vertex position maps may be generated. Each of these may include a scalar value, and typically the offset may be indicated as an integer indicating the number of pixels offset. FIG. 8 shows an example corresponding to FIG. 7, where the offset is indicated relative to the original reference pattern, with the first vertex position map (b) indicating the offset in the horizontal direction and the second vertex position map (c) indicating the offset in the vertical direction. As shown in the example, the vertex position maps may often include small integers, and in many scenarios may include a large number of zero values.
[0125] The position map generator 207 can supply both vertex position maps to the bitstream generator 209, which is configured to include both maps in the generated bitstream. A particular advantage of such an approach is that it allows for very efficient encoding of the offset information. For example, each of the vertex position maps can be encoded using an efficient video encoding algorithm or using other compression techniques optimized for efficient compression of small integer values (e.g., having many identical values). Encoding two separate such two-dimensional maps may be significantly more efficient than encoding a combined vertex position map.
[0126] Thus, in many embodiments, the bitstream generator 209 can be configured to encode the output depth map and / or vertex position map using a video encoding algorithm. Indeed, the encoding can be based on lossy video encoding techniques. More specifically, for example, assume the common YUV4:2:0 video frame format that is typically input to video compression algorithms. The YUV4:2:0 video format has a full resolution luminance channel, but the two chrominance channels are downsampled by a factor of four. In this case, the depth map can be placed in the Y component, and the horizontal and vertical displacement maps can be placed in the U and V components. Here, the video encoder can be configured with a quality level (quality=18) that corresponds to lossy compression but is still of high quality. Due to the lossy encoding, there may be some permanent errors in both the depth map and the displacement map. However, this is likely not to have a dramatic effect on the rendering quality, since it only means a slight shift in one or more vertices. Note that the video coding uses a block-based spatial-frequency transform followed by quantization and motion prediction to exploit the spatial and temporal redundancies that are very likely to be present in both the decimated depth and displacement maps.
[0127] FIG. 3 specifically illustrates elements of an exemplary apparatus for rendering a video bitstream of a video signal from the video bitstream generated by the apparatus of FIG.
[0128] The apparatus comprises a receiver 301 configured to receive a video signal from the server of Fig. 2. Thus, the receiver 301 receives the video signal including a bitstream including, for one or more, often every frame, a first depth map (output depth map of the server) and a vertex position map indicating vertex positions of depth values of the first depth map in a second depth map (input depth map of the server) having a higher resolution than the first depth map. The vertex positions of the vertex position map define two-dimensional geometric primitives that form a partition of the second depth map.
[0129] The receiver is coupled in this example to a renderer 303 configured to render a video image from the received data. The renderer 303 is configured to generate a view image of the scene depending on the observer pose input, thereby providing an immersive video experience.
[0130] The renderer 303 comprises a mesh generator 305 configured to generate a three-dimensional mesh representation of the scene based on the received information, in particular based on the first depth map and the vertex position map.
[0131] The mesh generator 305 is configured to generate a mesh formed by three-dimensional geometric primitives formed by projecting the two-dimensional geometric primitives of the second depth map onto the three-dimensional representation.
[0132] The mesh generator 305 can specifically be configured to generate the mesh by first determining vertex positions in a second map, i.e., in a map having a higher spatial resolution than the received first depth map, the determination being based on the vertex position map.
[0133] As an example, the vertex position map may include vertex position information provided as position offsets relative to a reference pattern of vertex positions, as described above. The mesh generator 305 may determine vertex positions corresponding to depth values in the first depth map as positions in the higher resolution map corresponding to the reference positions offset by values indicated in the vertex position map.
[0134] For example, the mesh generator 305 may start from the reference positions shown in Figure 4 and then offset each position by the offset indicated by the vertex position map. Thus, in the example where a vertex position map such as that shown in Figures 6(c) and 8(b)(c) is received, the mesh generator 305 applies these values to recreate the vertex positions determined by the encoder as shown in Figures 6(a) and 8(a).
[0135] The mesh generator 305 is then configured to project the vertex positions from their 2D positions in the high resolution depth map to positions in the 3D space of the mesh. The projection is given as a function of the 2D positions and the depth map values for the vertex positions. The projection of the 2D vertex positions into the 3D space also corresponds to the projection of the 2D primitives onto the 3D primitives and thus to the 3D mesh that is generated.
[0136] It will be appreciated that the generation of meshes from depth maps and (vertex) positions in a depth map is known to those skilled in the art and will not be described further for the sake of brevity and clarity.
[0137] The renderer 303 further comprises a rendering unit configured to generate an output image based on the generated mesh and texture map. The rendering unit 307 can in particular receive a view pose for which a view image of the scene is desired. The rendering unit 307 can then perform warping of the mesh to a two-dimensional viewport / image for the observer pose and generate the image from the texture map.
[0138] The rendering unit 307 may, for example, implement a graphics pipeline including a vertex shader, a geometry shader, a pixel shader, etc., as known to those skilled in the art. The renderer 303 may, for example, use a rendering implementation in which a map is sampled in a vertex shader of the graphics pipeline. First, the (reference) vertex coordinates may be shifted according to a vertex position map in the vertex shader. Then, the shifted coordinates may be passed to a fragment shader and used to calculate a vertex 3D position (which is also the output of the vertex shader pass).
[0139] It will be appreciated that the rendering unit 307 may use any suitable algorithm or approach to generate an output image based on the mesh, and many such approaches will be known to those skilled in the art.
[0140] In many embodiments, the reference pattern of vertices, and therefore the reference division or tessellation, is a predetermined fixed pattern that is known at both the sending / encoding / server end and the receiving / decoding / client end, for example, the reference pattern can be a standardized pattern that is applicable at both ends.
[0141] However, in some embodiments, there may be more than one option for the reference pattern, and the bitstream may include an index of the reference pattern for which the vertex position map is generated. For example, in some embodiments, the divider 203 may select between a predefined set of possible reference patterns and include an index of that pattern in the bitstream. For example, an identification may be assigned (e.g., by a standard) to each possible pattern, and the bitstream may be generated to include identification information for the predefined pattern, or, as another example, it may be predetermined that the reference pattern is a regular grid, but the divider 203 may dynamically select the size of the grid. For example, the bitstream may be generated to include a variable indicating the distance (as a number of pixels in the input depth map) between the reference vertex positions.
[0142] In embodiments where the video bitstream is generated to include an indication of the selected reference division / pattern, the renderer 303 may be configured to generate a local copy of the reference pattern based on this indication, for example it may extract the pattern corresponding to the received identification information, or it may generate the position of the pattern by applying a function or rule indicated in the bitstream (such as the distance between points in a regular grid).
[0143] In some embodiments, the reference pattern may be selected / indicated for the entire video signal, while in other embodiments it may be dynamically adapted and changed within the video signal. Indeed, in some embodiments, an indication of the reference pattern may be provided with each new vertex position map being transmitted / received.
[0144] Such adaptive use of reference positions can provide improved performance in many embodiments. It can, for example, allow better adaptation to particular characteristics of the depth map and / or video signal. In some cases, the reference pattern can be selected based on, for example, a desired / required maximum bitrate or based on a given quality level.
[0145] The described approach may provide improved operation in many scenarios and embodiments. It allows for a significant reduction in the depth map resolution of the depth map included in the bitstream / video signal. This approach may allow for closer adaptation to the particular scene being represented and may provide improved information of spatial characteristics to the rendering side. The adaptation of primitives to depth changes allows for improved characterization of the scene. This may often reduce and mitigate errors, artifacts and inaccuracies. In particular it may often reduce or mitigate errors, artifacts and inaccuracies around depth transitions or steps.
[0146] This approach can provide improved computational efficiency: in many scenarios, the number of vertices that need to be considered and processed can be significantly reduced, enabling, for example, real-time processing and three-dimensional adaptive video rendering in devices with relatively low computational resources in many scenarios.
[0147] This approach enables, for example, improved immersive video, where the displayed / rendered view image is adapted according to the observer's pose, and specifically adapted to reflect the user's head movements.
[0148] This approach often allows for an improved user experience, for example often allowing for a more immersive three-dimensional perception by the user.
[0149] While the above description has focused on a generated bitstream that includes a texture map and on rendering based on this texture map, it will be understood that this is not essential or necessary for operation, particularly for communication of information about the mesh or creation of same on the client side. For example, in other embodiments, the mesh can be generated without considering the texture map. For example, in some embodiments, a client application can assign material properties to the mesh, introduce light sources into the scene, and render a shaded mesh.
[0150] The invention may be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention may optionally be implemented at least partly as computer software running on one or more data processors and / or digital signal processors. Elements and components of embodiments of the invention may be physically, functionally and logically implemented in any suitable way. Indeed functionality may be implemented in a single unit, in multiple units or as part of other functional units. Thus, the invention may be implemented in a single unit or may be physically and functionally distributed between different units, circuits and processors.
[0151] In accordance with standard terminology in the field, the term pixel may be used to refer to pixel-related properties such as light intensity, depth, position, etc. of the portion / element of a scene represented by the pixel. For example, the depth of a pixel, i.e., pixel depth, may be understood to refer to the depth of the object represented by that pixel. Similarly, the brightness of a pixel, i.e., pixel luminosity, may be understood to refer to the brightness of the object represented by that pixel.
[0152] Although the present invention has been described in relation to several embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Moreover, although certain features may appear to be described in relation to a particular embodiment, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.
[0153] Furthermore, although individually recited, a plurality of means, elements, circuits or method steps may be implemented by, for example, a single circuit, unit or processor. Moreover, although individual features may be included in different claims, these may in some cases be advantageously combined, and their inclusion in different claims does not imply that the combination of features is not feasible and / or advantageous. Moreover, the inclusion of a feature in one category of claims does not imply limitation to this category, but rather indicates that the feature is equally applicable to other claim categories, as appropriate. Moreover, the order of features in the claims does not imply a particular order in which the features must be operated, and in particular the order of individual steps in a method claim does not imply that the steps must be performed in this order. Rather, the steps may be performed in any suitable order. Moreover, a reference to the singular does not exclude a plurality. Thus, reference to "a", "an", "first", "second", etc. does not exclude a plurality. Reference signs in the claims are provided merely as a clarifying example and shall not be construed as limiting the scope of the claims in any manner.
Claims
1. A device for generating a video signal including a bitstream, A receiver configured to receive a first depth map for at least a first video frame among multiple video frames, A divider configured to divide at least a portion of the first depth map into two-dimensional geometric primitives, wherein the two-dimensional geometric primitives have an average area greater than twice the pixel area of the first depth map, and the division includes determining the vertex positions in the first depth map for the vertices of the geometric primitives. A map generator configured to generate a vertex position map showing the vertex positions in the first depth map, A depth map generator configured to generate a second depth map that includes the depth values of the first depth map at the vertex positions in the first depth map, A bitstream generator configured to generate a video bitstream including the second depth map and the vertex position map for the first video frame, It has, A device wherein the divider is configured to generate the vertex positions depending on the depth values of the first depth map.
2. The apparatus according to claim 1, wherein the divider is configured to determine the vertex position by offsetting the reference vertex position of a reference division of the first depth map, the offset to the first vertex depends on the depth map value for the pixels in the geometric primitive of the first vertex.
3. The apparatus according to claim 2, wherein the reference vertex positions of the reference division form a regular grid.
4. The apparatus according to claim 2, wherein the vertex position map shows the vertex positions as relative positions to the reference vertex positions of the reference division.
5. The apparatus according to claim 2, wherein the offset includes a first offset along a first predetermined direction, and the first offset is determined as an offset where the change in the depth map value along the first predetermined direction satisfies a predetermined criterion.
6. The apparatus according to claim 2, wherein the offset includes a second offset along a second predetermined direction, the second offset is determined as an offset where the change in depth map value along the second predetermined direction satisfies a predetermined criterion, the map generator is configured to generate the vertex position map to indicate the offset along the first predetermined direction, and to further generate an additional vertex position map to indicate the offset along the second predetermined direction, and the bitstream generator is configured to include the additional vertex position map in the bitstream.
7. The apparatus according to claim 2, wherein the first predetermined direction is the horizontal direction.
8. The apparatus according to claim 2, wherein the first predetermined direction is the vertical direction.
9. The apparatus according to claim 2, wherein the offset is constrained to be below an offset threshold.
10. The apparatus according to claim 1, wherein the bitstream generator has a video encoder configured to encode the vertex position map using video encoding.
11. The apparatus according to claim 1, wherein the divider is configured to determine the vertex positions such that the geometric primitives are aligned with the depth transitions in the depth map.
12. A device for rendering a bitstream of a video signal, A receiver for receiving a video bitstream, wherein the video bitstream includes a first depth map for at least a first frame, and a vertex position map indicating the vertex positions of depth values of the first depth map in a second map having a higher resolution than the first depth map, wherein the vertex positions define two-dimensional geometric primitives that form a division of the second map. A renderer configured to render video images from the aforementioned video bitstream, The renderer has, A mesh generator configured to generate a mesh representation of a scene, wherein the mesh is formed by three-dimensional geometric primitives formed by projecting the two-dimensional geometric primitives into three-dimensional space, and the projection of the first vertex of the two-dimensional geometric primitive depends on the vertex position of the first vertex in the second depth map and the depth value in the first depth map for the first vertex, A rendering unit configured to render an image based on the aforementioned mesh representation, A device having.
13. The apparatus according to claim 12, wherein the vertex position map shows the vertex positions as relative vertex positions with respect to the reference vertex positions of the reference division, and the mesh generator is configured to determine absolute vertex positions in the second depth map according to the relative vertex positions and the reference vertex positions.
14. The apparatus according to claim 13, wherein the bitstream includes the index of the reference division, and the mesh generator is configured to determine the reference vertex positions according to the index of the reference division.
15. A method for generating a video signal including a bitstream, The steps include receiving a first depth map for at least a first video frame among a plurality of video frames, A step of dividing at least a portion of the first depth map into two-dimensional geometric primitives, wherein the two-dimensional geometric primitives have an average area greater than twice the pixel area of the first depth map, and the division includes determining the vertex positions in the first depth map for the vertices of the geometric primitives. The steps include generating a vertex position map showing the vertex positions in the first depth map, A step of generating a second depth map that includes the depth values of the first depth map at the vertex positions in the first depth map, The steps of generating a video bitstream to include the second depth map and the vertex position map for the first video frame, It has, A method for generating the vertex positions, wherein the generation of the vertex positions depends on the depth values of the depth map.
16. A method for rendering a bitstream of a video signal, A step of receiving a video bitstream, wherein the video bitstream includes, for at least a first frame, a first depth map and a vertex position map indicating the vertex positions of depth values of the first depth map in a second map having a higher resolution than the first depth map, wherein the vertex positions define two-dimensional geometric primitives that form a division of the second map. The steps include rendering a video image from the aforementioned video bitstream, The rendering step includes, A step of generating a mesh representation of a scene, wherein the mesh is formed by three-dimensional geometric primitives formed by projecting the two-dimensional geometric primitives into three-dimensional space, and the projection of the first vertices of the two-dimensional geometric primitives depends on the vertex position of the first vertex in the second depth map and the depth value in the first depth map for the first vertex; The steps include rendering an image based on the aforementioned mesh representation, Methods that include...
17. A computer program that runs on a computer and causes the computer to perform the method described in claim 15 or 16.
18. A video signal including a video bitstream, wherein the video bitstream includes a first depth map for at least a first frame, and a vertex position map indicating the vertex positions of depth values of the first depth map in a second map having a higher resolution than the first depth map, wherein the vertex positions define two-dimensional geometric primitives that form a division of the second map.