Methods and apparatuses for volumetric video encoding and decoding

By performing patch division and surface parameter coding methods on volume video, the problem of large amount of encoded data and high bandwidth consumption of volume video is solved, and more efficient coding and rendering effects are achieved.

CN113853796BActive Publication Date: 2025-06-13NOKIA TECHNOLOGIES OY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080037271.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-22
Filing Date
2020-04-15
Publication Date
2025-06-13
Estimated Expiration
2040-04-15

AI Technical Summary

Technical Problem

Volume video encoding has problems with large data volume and high bandwidth consumption, especially when capturing and transmitting data from multiple 3D cameras.

Method used

By receiving video presentation frames, generating patches, dividing them into blocks, surface parameters are determined, encoded into bitstreams, and stored for transmission to rendering devices. During decoding, the three-dimensional block data is decoded from the compressed bitstream, surface parameters are determined, bounding boxes are generated, ray direction is calculated, ray projection is performed, and the three-dimensional data is reconstructed for rendering.

Benefits of technology

It improves the efficiency of volume video encoding, reduces data volume and bandwidth consumption, and enhances the rendering ability of 6DOF.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113853796B_ABST
    Figure CN113853796B_ABST
Patent Text Reader

Abstract

A decoding method, comprising: receiving a compressed bitstream related to a video presentation; decoding data related to one or more three-dimensional blocks of a video frame from the received bitstream; for each block of the video frame, determining information about surface parameters; generating a bounding box of the three-dimensional block according to the surface parameters; for each pixel of the three-dimensional block, calculating a ray direction from a viewpoint to the coordinates of the pixel; determining at least two points according to the intersection of the ray and the generated bounding box; performing ray casting on the points between the determined at least two points until a condition for completing ray casting is satisfied; reconstructing three-dimensional data from a geometric image and a texture image according to information about one or more surface parameters of the block; reconstructing the video presentation for rendering according to the reconstructed three-dimensional data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This solution generally relates to volumetric video coding. In particular, the solution relates to point cloud compression. Background Art

[0002] Since the beginning of photography and videography, the most common types of image and video content have been captured by cameras with a relatively narrow field of view and are displayed as rectangular scenes on flat displays. Cameras are mainly oriented so that they only capture a limited angular field of view (the field of view they are pointing at).

[0003] Recently, new image and video capture devices have become available. These devices are capable of capturing the visual and audio content around them, i.e., they can capture the entire angular field of view, sometimes referred to as a 360-degree field of view. More precisely, they can capture a spherical field of view (i.e., 360 degrees in all spatial directions). In addition, new output technologies, such as head-mounted displays, have been invented and produced. These devices allow people to see the visual content around him / her, giving a feeling of "immersing" into the scene captured by a 360-degree camera. In the case of a spherical field of view, the new capture and display paradigm is generally referred to as virtual reality (VR) and is considered a common way for people to experience media content in the future.

[0004] For volumetric video, one or more 3D (three-dimensional) cameras can be used to capture the scene. These cameras are in different positions and orientations in the scene. One issue to consider is that compared to 2D (two-dimensional) video content, volumetric 3D video content has more data, so viewing it requires a large amount of bandwidth (whether it is transferred from a storage location to a viewing device): disk I / O, network traffic, memory bandwidth, GPU (graphics processing unit) upload. Capturing volumetric content also generates a large amount of data, especially when multiple capture devices are used in parallel. Summary of the Invention

[0005] Now, an improved method and a technical device for implementing the method have been invented to provide improvements for volumetric video coding. Respective aspects include a method, an apparatus, and a computer-readable medium including a computer program stored therein, characterized by what is described in the independent claims. Various embodiments are disclosed in the dependent claims.

[0006] In a first aspect, there is provided an encoding method, including: receiving a video presentation frame, where the video presentation represents three-dimensional data; generating one or more patches from the video presentation frame; dividing the patches of the video frame into one or more blocks; determining surface parameters for each patch according to information related to the one or more blocks of the patch; encoding the determined surface parameters into a bitstream; and storing the encoded bitstream for sending to a rendering device.

[0007] According to a second aspect, a method for decoding is provided, including: receiving a compressed bitstream related to a video presentation, the bitstream including at least a geometry image and a texture image; decoding data related to one or more three-dimensional blocks of a video frame from the received bitstream; for each block of the video frame, determining information about one or more surface parameters; generating a bounding box of the three-dimensional block according to the one or more surface parameters; for each pixel of the three-dimensional block, calculating a ray direction from a viewpoint to the coordinates of the pixel; determining at least two points according to the intersection of the ray and the generated bounding box; performing ray casting on the points between the determined at least two points until a condition for completing ray casting is satisfied; reconstructing three-dimensional data from the geometry image and the texture image according to the information about one or more surface parameters of the block; and reconstructing the video presentation for rendering according to the reconstructed three-dimensional data.

[0008] According to a third aspect, an apparatus for encoding a bitstream is provided, including: a component for receiving a video presentation frame, where the video presentation represents three-dimensional data; a component for generating one or more patches from the video presentation frame; a component for dividing the patches of the video frame into one or more blocks; a component for determining surface parameters for each patch according to information related to the one or more blocks of the patch; a component for encoding the determined surface parameters into the bitstream; and a component for storing the encoded bitstream for sending to a rendering device.

[0009] According to a fourth aspect, an apparatus for decoding a bitstream is provided, including: a component for receiving a compressed bitstream related to a video presentation, the bitstream including at least a geometry image and a texture image; a component for decoding data related to one or more three-dimensional blocks of a video frame from the received bitstream; a component for determining information about one or more surface parameters for each block of the video frame; a component for generating a bounding box of the three-dimensional block according to the one or more surface parameters; a component for calculating a ray direction from a viewpoint to the coordinates of each pixel of the three-dimensional block; a component for determining at least two points according to the intersection of the ray and the generated bounding box; a component for performing ray casting on the points between the determined at least two points until a condition for completing ray casting is satisfied; a component for reconstructing three-dimensional data from the geometry image and the texture image according to the information about one or more surface parameters of the block; and a component for reconstructing the video presentation for rendering according to the reconstructed three-dimensional data.

[0010] According to a fifth aspect, there is provided an apparatus including at least one processor and a memory including computer program code, the memory and the computer program code being configured to, with the at least one processor, cause the apparatus to at least perform the following operations: receive a video presentation frame, wherein the video presentation represents three-dimensional data; generate one or more patches from the video presentation frame; divide the patches of the video frame into one or more blocks; determine surface parameters for each patch according to information related to the one or more blocks of the patch; encode the determined surface parameters into a bitstream; and store the encoded bitstream for transmission to a rendering device.

[0011] According to a sixth aspect, there is provided an apparatus including at least one processor and a memory including computer program code, the memory and the computer program code being configured to, with the at least one processor, cause the apparatus to at least perform the following operations: receive a compressed bitstream related to a video presentation, the bitstream including at least a geometry image and a texture image; decode data related to one or more three-dimensional blocks of a video frame from the received bitstream; determine information about one or more surface parameters for each block of the video frame; generate a bounding box of the three-dimensional block according to the one or more surface parameters; calculate a ray direction from a viewpoint to the coordinates of each pixel of the three-dimensional block; determine at least two points according to the intersection of the ray and the generated bounding box; perform ray casting on points between the determined at least two points until a condition for completing ray casting is satisfied; reconstruct three-dimensional data from the geometry image and the texture image according to information about one or more surface parameters of the block; and reconstruct the video presentation for rendering according to the reconstructed three-dimensional data.

[0012] According to a seventh aspect, there is provided a computer program product including computer program code configured to, when executed on at least one processor, cause an apparatus or a system to: receive a video presentation frame, wherein the video presentation represents three-dimensional data; generate one or more patches from the video presentation frame; divide the patches of the video frame into one or more blocks; determine surface parameters for each patch according to information related to the one or more blocks of the patch; encode the determined surface parameters into a bitstream; and store the encoded bitstream for transmission to a rendering device.

[0013] According to an eighth aspect, there is provided a computer program product comprising computer program code configured to, when executed on at least one processor, cause an apparatus or system to: receive a compressed bitstream related to video rendering, the bitstream including at least a geometry image and a texture image; decode data related to one or more three-dimensional blocks of a video frame from the received bitstream; for each block of the video frame, determine information about one or more surface parameters; generate a bounding box of the three-dimensional block according to the one or more surface parameters; for each pixel of the three-dimensional block, calculate a ray direction from a viewpoint to the coordinates of the pixel; determine at least two points according to the intersection of the ray and the generated bounding box; perform ray casting on the points between the determined at least two points until a condition for completing ray casting is satisfied; reconstruct three-dimensional data from the geometry image and the texture image according to the information about one or more surface parameters of the block; and reconstruct the video rendering for rendering according to the reconstructed three-dimensional data.

[0014] According to an embodiment, the information about one or more surface parameters is decoded from the bitstream.

[0015] According to an embodiment, the information about one or more surface parameters is determined from the pixels of various depth layers.

[0016] According to an embodiment, the condition for completing ray transmission is determined by: determining the depth values of two depth layers at the position of the pixel, calculating the depth values from the viewing ray, and comparing the depth values from the viewing ray with the determined depth values.

[0017] According to an embodiment, the condition for completing ray casting is determined from another bounding box formed by the depth difference and the pixel coordinates.

[0018] According to an embodiment, the surface parameter is the depth or depth difference of the patch.

[0019] According to an embodiment, the surface parameter is a rendering thickness parameter determined according to the depth difference.

[0020] According to an embodiment, based on the ray direction, at least two points are determined, and ray casting is performed on each pixel between the two points in two-dimensional coordinates.

[0021] According to an embodiment, for each two-dimensional pixel between the two points, the depth values of the first depth layer and the second depth layer are obtained.

[0022] According to an embodiment, it is determined whether there is an intersection between the point cloud content and the ray.

[0023] According to an embodiment, the rendering thickness parameter is encoded into a supplementary enhancement information (SEI) message or decoded from the supplementary enhancement information message.

[0024] According to an embodiment, rendering parameters are encoded / decoded for each block.

[0025] According to an embodiment, the rendering parameters are encoded into an occupancy map or decoded from the occupancy map.

[0026] According to an embodiment, color interpolation between depth layers is encoded into a bitstream or decoded from the bitstream.

[0027] According to an embodiment, a computer program product is embodied on a non-transitory computer-readable medium. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Hereinafter, various embodiments will be described in more detail with reference to the accompanying drawings, in which:

[0029] Figure 1 An example of a compression process is shown;

[0030] Figure 2 An example of a layer projection structure is shown.

[0031] Figure 3 An example of a decompression process is shown;

[0032] Figure 4 An example of a patch in a frame is shown;

[0033] Figure 5 An example of a block in 3D is shown;

[0034] Figure 6 An example of a block in 2D is shown;

[0035] Figure 7 An example of a decoded signal with potential rendering surface thickness values is shown;

[0036] Figure 8 is a flowchart showing a method according to an embodiment.

[0037] Figure 9 is a flowchart showing a method according to another embodiment.

[0038] Figure 10 A system according to an embodiment is shown.

[0039] Figure 11 An encoding process according to an embodiment is shown;

[0040] Figure 12 A decoding process according to an embodiment is shown. DETAILED DESCRIPTION

[0041] In the following, several embodiments will be described in the context of digital volumetric video. In particular, these embodiments enable the encoding and decoding of digital volumetric video material. The proposed embodiments are applicable to, for example, point cloud coding (V-PCC) based on MPEG video.

[0042] Volumetric video can be captured using one or more three-dimensional (3D) cameras. When multiple cameras are used, the captured footage is synchronized so that the cameras provide different viewpoints of the same world. Compared to traditional two-dimensional / three-dimensional (2D / 3D) video, volumetric video describes a 3D model of the world, where the viewer can freely move and observe different parts of the world.

[0043] Volumetric video enables the viewer to move in six degrees of freedom (DOF): compared to common 360° video, where the user has 2 to 3 degrees of freedom (yaw, pitch, and possibly roll), volumetric video represents the 3D volume of the shape rather than a flat image plane. Volumetric video frames contain a large amount of data because they model the content of the 3D volume rather than just a 2D plane. However, only a relatively small portion of the volume changes over time. Therefore, the total data volume can be reduced by encoding only the information about the initial state and the possible changes that occur between frames. Volumetric video can be rendered, for example, from synthetic 3D animations, reconstructed from multi-view video using 3D reconstruction techniques such as structure from motion, or captured with a combination of cameras and depth sensors such as LiDAR.

[0044] Volumetric video data represents a three-dimensional scene or object and can be used as input for augmented reality (AR), virtual reality (VR), and mixed reality (MR) applications. Such data describes the geometric structure (shape, size, position in 3D space) and corresponding properties (e.g., color, opacity, reflectivity...). Additionally, volumetric video data can define any possible temporal changes of the geometric structure and properties at a given time instance (such as a frame in 2D video). Volumetric video can be generated from 3D models, i.e., computer-generated imagery (CGI), or can be captured from real-world scenes using various capture schemes (e.g., a combination of multiple cameras, raster scanning, video, and dedicated depth sensors, etc.). A combination of CGI and real-world data is also possible. Examples of representation formats for such volumetric data include triangle meshes, point clouds, or voxels. The temporal information about the scene can be included in the form of individual capture instances, i.e., the "frames" in 2D video, or in other ways, for example, as the position of an object as a function of time.

[0045] Since volumetric video describes a 3D scene (or object), this data can be viewed from any viewpoint. Thus, volumetric video is an important format for any AR, VR, or MR application, especially for providing 6DOF viewing capabilities.

[0046] Increasing computational resources and the development of 3D data acquisition devices have made it possible to reconstruct highly detailed volumetric video representations of natural scenes. Infrared, lasers, time-of-flight, and structured light are example devices that can be used to build 3D video data. The representation of 3D data depends on how the 3D data is used. Dense voxel arrays have been used to represent volumetric medical data. In 3D graphics, polygon meshes are widely used. On the other hand, point clouds are well-suited for applications such as capturing 3D scenes of the real world where the topology does not have to be a 2D manifold. Another way to represent 3D data is to encode the 3D data as a set of texture and depth maps, as in the case of multi-view plus depth. Closely related to the techniques used in multi-view plus depth is the use of elevation maps and multi-level surface maps.

[0047] In a 3D point cloud, each point in each 3D surface is described as a 3D point with color and / or other attribute information such as surface normal or material reflectance. A point cloud is a set of data points (i.e., positions) in a coordinate system (e.g., a three-dimensional coordinate system defined by X, Y, and Z coordinates). These points can represent the outer surface of an object in screen space (e.g., 3D space). The points can be associated with a vector of attributes. Point clouds can be used to reconstruct an object or scene as a combination of points. Multiple cameras and depth sensors can be used to capture point clouds. A dynamic point cloud is a series of static point clouds, where each static point cloud is in its own "point cloud frame".

[0048] In a dense point cloud or voxel array, the reconstructed 3D scene can contain tens of millions or even hundreds of millions of points. If such a representation is to be stored or exchanged between entities, efficient compression is required. Volumetric video representation formats such as point clouds, meshes, and voxels do not have sufficient temporal compression performance. Identifying motion-compensated correspondences in 3D space is an ill-defined problem because both the geometric structure and the corresponding attributes can change. For example, temporally consecutive point cloud frames do not have to have the same number of meshes, points, or voxels. Thus, the compression of dynamic 3D scenes is inefficient. Methods for compressing volumetric data based on 2D video (i.e., multi-view and depth) have better compression efficiency but cover only a limited portion of the scene. Thus, they provide only limited 6DOF capabilities.

[0049] Instead of the methods mentioned above, 3D scenes represented as meshes, points, and / or voxels can be projected onto one or more geometric structures. These geometric structures can be "unrolled" onto a 2D plane (two planes per geometric structure: one for texture and one for depth), and then they can be encoded using standard 2D video compression techniques. The relevant projection geometry information can be sent to the decoder along with the encoded video file. The decoder can decode the video and perform inverse projection to regenerate the 3D scene in any desired representation format (not necessarily the starting format).

[0050] Projecting a volume model onto a 2D plane allows the use of standard 2D video encoding tools with efficient temporal compression. Thus, the encoding efficiency is greatly improved. Using geometric projection instead of known 2D video-based methods (i.e., multi-view and depth) provides better scene (or object) coverage. Thus, the 6DOF capability is improved. Using several geometric structures for each object further improves the scene coverage. In addition, standard video encoding hardware can be used for real-time compression / decompression of the projected planes. The projection and inverse projection steps have low complexity.

[0051] An overview of the compression process is briefly discussed next. Such a process can be applied to, for example, V-PCC. In the encoding stage, the input point cloud frame is processed as follows: First, the volumetric 3D data can be represented as a set of 3D projections with different components. In the separation stage, the image is decomposed into the near-far components of the geometric structure and the corresponding attribute components, and an occupancy Figure 2 D image can be created to indicate the parts of the image that will be used. The 2D projections consist of independent patches based on the geometric characteristics of the input point cloud frame. After the patches have been generated and the 2D frames for video encoding have been created, the occupancy map, geometric information, and auxiliary information can be compressed. At the end of the process, the separated bitstreams are multiplexed into the output compressed binary file.

[0052] Figure 1 The encoding process is shown in a more detailed manner.

[0053] The process starts with an input frame representing a point cloud frame 101, which is provided for patch generation 102, geometric image generation 104, and texture image generation 105. Each point cloud frame 101 represents a dataset of points with unique coordinates and attributes within a 3D volume space.

[0054] The patch generation 102 process decomposes the point cloud frame 101 by converting 3D samples into 2D samples on a given projection plane using a strategy that provides optimal compression. According to an example, the patch generation 102 process aims to decompose the point cloud frame 101 into the minimum number of patches with smooth boundaries while also minimizing the reconstruction error.

[0055] At the initial stage of patch generation 102, the normal per each point is estimated. Based on the closest neighbors m of the points within a predefined search distance, the tangent plane and its corresponding normal are defined per each point. A k-d tree can be used to separate the data and find the neighbors within the vicinity of point p i and the centroid of the point set is used to define the normal. The centroid c can be calculated as follows:

[0056]

[0057] The normal is estimated from the eigen decomposition of the features for the defined point cloud as:

[0058]

[0059] Based on this information, each point is associated with the corresponding plane of the point cloud bounding box. Each plane is defined by the corresponding normal with the following values defined as:

[0060] -(1.0, 0.0, 0.0),

[0061] -(0.0, 1.0, 0.0),

[0062] -(0.0, 0.0, 1.0),

[0063] -(-1.0, 0.0, 0.0),

[0064] -(0.0, -1.0, 0.0),

[0065] -(0.0, 0.0, -1.0)

[0066] More precisely, each point can be associated with the plane having the closest normal (i.e., maximizing the dot product of the point normal and the plane normal ).

[0067]

[0068] The sign of the normal is defined based on the position of the point relative to the "center".

[0069] Furthermore, based on the normal of each point and the clustering index of its closest neighbors, the initial clustering can be refined by iteratively updating the clustering index associated with each point. The final step of patch generation 102 can include extracting the patches by applying a connected component extraction process.

[0070] The patch information determined at patch generation 102 for the input point cloud frame 101 is transmitted to patch packing 103, geometric image generation 104, texture image generation 105, attribute smoothing (3D) 109, and auxiliary patch information compression 113. Patch packing 103 aims to generate geometric and texture maps by appropriately considering the generated patches and attempting to efficiently place the geometric and texture data corresponding to each patch onto a 2D grid of size WxH. This placement also takes into account a user-defined minimum size block TxT (e.g., 16x16), which specifies the minimum distance between different patches as placed on this 2D grid. The parameter T can be encoded in the bitstream and sent to the decoder.

[0071] The packing process 103 can iteratively attempt to insert patches into the WxH grid. W and H are user-defined parameters that correspond to the resolution of the geometric / texture image to be encoded. The patch positions can be determined by an exhaustive search, which can be performed in raster scan order. Initially, the patches are placed on the 2D grid in a way that will guarantee non-overlapping insertion. Samples belonging to a patch (values rounded to be a multiple of T) are considered to occupy a block. Additionally, the guard between adjacent patches is forced to be at least one block (a multiple of T) distance. The patches are processed in an ordered manner based on a patch index list. Each patch from this list is iteratively placed on the grid. The grid resolution depends on the size of the original point cloud, and its width (W) and height (H) are sent to the decoder. In the case where there is no available empty space for the next patch, the height value of the grid is initially doubled, and the insertion of this patch is evaluated again. If all patch insertions are successful, the height will be trimmed to the minimum required value. However, it is not allowed to set this value lower than the value initially specified in the encoder. The final values of W and H correspond to the frame resolution used to encode the texture and geometric video signals using a suitable video codec.

[0072] Geometric image generation 104 and texture image generation 105 are configured to generate geometric images and texture images. The image generation process can utilize the 3D-to-2D mapping computed during the packing process to store the geometric structure and texture of the point cloud as images. To better handle the case where multiple points are projected onto the same pixel, each patch can be projected onto two images (referred to as layers). For example, let H(u,y) be the set of points in the current patch that are projected onto the same pixel (u,v). Figure 2 An example of the layer projection structure is shown. The first layer (also referred to as the near layer) stores the points in H(u,v) with the lowest depth D0. The second layer (referred to as the far layer) captures the points in H(u,v) with the highest depth within the interval [D0, D0+Δ], where Δ is a user-defined parameter that describes the surface thickness. The generated video can have the following characteristics:

[0073] ● Geometric structure: WxH YUV420 - 8 bits

[0074] ● Texture: WxH YUV420 - 8 bits

[0075] Note that the geometric video is monochromatic. Additionally, the texture generation process utilizes the reconstructed / smoothed geometric structure in order to calculate the color to be associated with the resampled points.

[0076] When there are stacks of multiple different surfaces in a connected component, a surface separation method is applied to prevent the mixing of different surfaces in that connected component. One of the surface separation methods is to use the difference in the MSE values of points in the RGB color gamut:

[0077] If

[0078] MSE(R 1 -R 2 ,G 1 -G 2 ,B 1 -B 2 ) > threshold;

[0079] Threshold = 20

[0080] Then separate the patch, where R 1 ,G 1 ,B 1 are the attribute values belonging to T0, and R 2 ,G 2 ,B 2 are the attribute values belonging to T1.

[0081] The geometric image and the texture image can be provided to the inpainting 107. The inpainting 107 can also receive an occupancy map (OM) 106 to be used together with the geometric image and the texture image as input. The occupancy map 106 can include a binary map that indicates for each cell of the grid whether it belongs to empty space or to the point cloud. In other words, the occupancy map (OM) can be a binary image with binary values, where the occupied pixels and the unoccupied pixels are distinguished and depicted separately. Alternatively, the occupancy map can include a non - binary image that allows additional information to be stored therein. Thus, the representation value of the DOM can include binary values or other values, such as integer values. Note that during the image generation process, one cell of the 2D grid can produce one pixel.

[0082] The filling process 107 aims to fill the empty spaces between the patches to generate a piecewise smooth image suitable for video compression. For example, in a simple filling strategy, each TxT (e.g., 16x16) pixel block is compressed independently. If the block is empty (i.e., unoccupied, meaning all its pixels belong to the empty space), the pixels of the block are filled by copying the last row or the last column of the previous TxT block in raster order. If the block is full (i.e., occupied, meaning there are no empty pixels), nothing is done. If the block has both empty pixels and filled pixels (i.e., an edge block), the empty pixels are filled iteratively with the average of its non-empty neighbors.

[0083] A filled geometric image and a filled texture image can be provided for video compression 108. The generated image / layer can be stored as a video frame and compressed using, for example, the HM video codec according to the High Efficiency Video Coding (HEVC) Test Model 16 (HM) configuration provided as a parameter. Video compression 108 also generates a reconstructed geometric image to be provided for smoothing 109, where the smoothed geometric structure is determined based on the reconstructed geometric image and the patch information from patch generation 102. The smoothed geometric structure can be provided to texture image generation 105 to adapt the texture image.

[0084] Patches can be associated with auxiliary information as metadata that is encoded / decoded for each patch. The auxiliary information can include the index of the projection plane, the 2D bounding box, and the 3D position of the patch represented by the depth δ0, the tangential offset s0, and the bitangential offset r0.

[0085] The following metadata can be encoded / decoded for each patch:

[0086] ● Index of the projection plane

[0087] ○ Index 0 for the planes (1.0, 0.0, 0.0) and (-1.0, 0.0, 0.0)

[0088] ○ Index 1 for the planes (0.0, 1.0, 0.0) and (0.0, -1.0, 0.0)

[0089] ○ Index 2 for the planes (0.0, 0.0, 1.0) and (0.0, 0.0, -1.0)

[0090] ● 2D bounding box (u0, v0, u1, v1)

[0091] ● 3D position (x0, y0, z0) of the patch represented by the depth δ0, the tangential offset s0, and the bitangential offset r0. Depending on the selected projection plane, (δ0, s0, r0) is calculated as follows:

[0092] ○ Index 0, δ0 = x0, s0 = z0, r0 = y0

[0093] ○ Index 1, δ0 = y0, s0 = z0, r0 = x0

[0094] ○ Index 2, δ0 = z0, s0 = x0, r0 = y0

[0095] In addition, the mapping information for the patch indices associated with each TxT block can be encoded as follows:

[0096] ● For each TxT block, let L be the ordered list of patch indices such that their 2D bounding boxes enclose the block. The order in this list is the same as the order used to encode the 2D bounding box. L is called the candidate patch list.

[0097] ● The empty space between patches is treated as a patch and assigned a specific index 0, which is added to the candidate patch list for all blocks.

[0098] ● Let I be the index of the patch to which the current TxT block belongs, and let J be the position of I in L. Instead of explicitly encoding the index I, its position J is arithmetically encoded, which results in better compression efficiency.

[0099] The compression process can include one or more of the following example operations:

[0100] ● A binary value can be associated with the B0xB0 sub-blocks belonging to the same TxT block. If the sub-block contains at least one unfilled pixel, the value 1 is associated with the sub-block, otherwise 0. If the value of the sub-block is 1, it is considered full, otherwise it is an empty sub-block.

[0101] ● If all sub-blocks of a TxT block are full (i.e., the value is 1), then the block is considered full.

[0102] Otherwise, the block is considered non-full.

[0103] ● Binary information can be encoded for each TxT block to indicate whether it is full.

[0104] ● If the block is non-full, additional information indicating the positions of the full / empty sub-blocks can be encoded as follows:

[0105] ○ Different traversal orders can be defined for the sub-blocks, for example, horizontally, vertically, or diagonally starting from the upper right or upper left corner

[0106] ○ The encoder selects one of the traversal orders and can explicitly signal its index in the bitstream

[0107] ○ The binary values associated with the sub-blocks can be encoded using a run-length encoding strategy

[0108] ■ The binary values of the initial sub-blocks are encoded.

[0109] ■ Runs of 0s and 1s are detected while following the traversal order selected by the encoder.

[0110] ■ The number of detected runs is encoded.

[0111] ■ The length of each run, except for the last one, is also encoded.

[0112] In encoding an occupancy map (lossy condition) for a two-dimensional binary image of resolution (Width / B0) x (Height / B1), where "Width" and "Height" are the width and height of the geometric and texture images intended to be compressed. A sample equal to 1 means that at decoding, one or more corresponding / co-located samples in the geometric and texture images should be treated as point cloud points, while samples equal to 0 should be ignored (usually including padding information). The resolution of the occupancy map does not have to be the same as that of the geometric and texture images, but the occupancy map can be encoded with the precision of B0xB1 blocks. To achieve lossless encoding, B0 and B1 are chosen to be equal to 1. In practice, B0 = B1 = 2 or B0 = B1 = 4 can give visually acceptable results while significantly reducing the number of bits required to encode the occupancy map. The generated binary image only covers a single color plane. However, given the popularity of 4:2:0 codecs, it may be desirable to extend the image with "neutral" or fixed-value chrominance planes (e.g., adding chrominance planes with all sample values equal to 0 or 128, assuming an 8-bit codec).

[0113] A video codec with support for lossless encoding tools (e.g., AVC, HEVC RExt, HEVC-SCC) can be used to compress the obtained video frames.

[0114] The occupancy map can be simplified by detecting empty and non-empty blocks of resolution TxT in the occupancy map, and only for non-empty blocks, we encode their patch indices as follows:

[0115] ○ For each TxT block, create a candidate patch list by considering all patches containing that block.

[0116] ○ The candidate list is sorted in reverse order of the patches.

[0117] ○ For each block,

[0118] 1. If the candidate list has one index, no encoding is done.

[0119] 2. Otherwise, arithmetic coding is done on the indices of the patches in this list.

[0120] The point cloud geometry reconstruction process utilizes occupancy map information to detect non-empty pixels in the geometry / texture image / layer. The 3D positions of the points associated with these pixels are calculated by using the auxiliary patch information and the geometry image. More precisely, let P be the point associated with the pixel (u, v), let (δ0, s0, r0) be the 3D position of the patch to which it belongs, and let (u0, v0, u1, v1) be its 2D bounding box. P can be expressed as follows in terms of the depth δ(u, v), the tangential displacement s(u, v), and the bi-tangential displacement r(u, v):

[0121] δ(u, v) = δ0 + g(u, v)

[0122] s(u, v) = s0 – u0 + u

[0123] r(u, v) = r0 – v0 + v

[0124] where g(u, v) is the luminance component of the geometry image.

[0125] The attribute smoothing procedure 109 is designed to mitigate potential discontinuities that may occur at patch boundaries due to compression artifacts. The method implemented moves the boundary points to the centroid of their closest neighbors.

[0126] The multiplexer 112 can receive the compressed geometry video and the compressed texture video from the video compression 108, and optionally receive the compressed auxiliary patch information from the auxiliary patch information compression 111. The multiplexer 112 uses the received data to generate a compressed bitstream.

[0127] Figure 3 An overview of the decompression process for MPEG Point Cloud Coding (PCC) is shown. The demultiplexer 201 receives the compressed bitstream and, after demultiplexing, provides the compressed texture video and the compressed geometry video to the video decompression 202. Additionally, the demultiplexer 201 sends the compressed occupancy map to the occupancy map decompression 203. It can also send the compressed auxiliary patch information to the auxiliary patch information decompression 204. The decompressed geometry video from the video decompression 202 is transmitted to the geometry reconstruction 205, as are the decompressed occupancy map and the decompressed auxiliary patch information. The point cloud geometry reconstruction 205 process utilizes the occupancy map information to detect non-empty pixels in the geometry / texture image / layer. The 3D positions of the points associated with these pixels can be calculated by using the auxiliary patch information and the geometry image.

[0128] A reconstructed geometric image can be provided for smoothing 206, with the aim of alleviating potential discontinuities that may occur at the patch boundaries due to compression artifacts. The method implemented moves the boundary points to the centroid of their closest neighbors. The smoothed geometry can be sent to texture reconstruction 207, which also receives the decompressed texture video from video decompression 202. The texture values for texture reconstruction are read directly from the texture image. Texture reconstruction 207 outputs a reconstructed point cloud for color smoothing 208, and color smoothing 208 further provides the reconstructed point cloud.

[0129] Occupancy information can be encoded using a geometric image. A specific depth value (e.g., 0) or a specific depth value range can be reserved to indicate that a pixel is inpainted and does not exist in the source material. The specific depth value or the specific depth value range can be, for example, predefined in a standard, or the specific depth value or the specific depth value range can be encoded into or with the bitstream, and / or the specific depth value or the specific depth value range can be decoded from or with the bitstream. This way of multiplexing occupancy information in the depth sample array creates sharp edges in the image, which can be affected by additional bitrate as well as compression artifacts around the sharp edges.

[0130] A method for compressing a time-varying volumetric scene / object is to project the 3D surface onto a number of predefined 2D planes. Subsequently, conventional 2D video compression algorithms can be used to compress various aspects of the projected surface. For example, a time-varying 3D point cloud with spatial and texture coordinates can be mapped into a sequence of at least two sets of planes, where one of the two sets carries texture data and the other set carries the distance of the mapped 3D surface points to the projection plane.

[0131] For accurate 2D-to-3D reconstruction on the receiving side, the decoder must know which 2D points are "valid" and which points result from interpolation / filling. This requires the transmission of additional data. The additional data can be encapsulated in the geometric image as a predefined depth value (e.g., 0) or a predefined depth value range. This will only improve the coding efficiency for the texture image because the geometric image is not blurred / filled. Additionally, coding artifacts at the object boundaries of the geometric image can produce severe artifacts, which require post-processing and may not be hidden.

[0132] Currently, a simple V-PCC decoder implementation is configured to render all decoded and reconstructed data as points. However, current mobile devices are not designed to render millions of points. Since games and similar applications use triangles as rendering primitives, points such as rendering primitives are not optimized for mobile graphics processing units (GPUs). The quality of rendering dense point clouds can also be affected by visual artifacts because these points can overlap with adjacent points. When looking at point cloud content from a close distance, this can result in an unpleasant visual quality. In the best case, each point can be rendered with a cube, which would result in better visual quality. However, in this case, each cube would consist of 12 triangles. It should be realized that this is 6 times more complex than rendering a point with 2 triangles (quadrilateral), and thus is not practical on mobile devices with limited battery in any case.

[0133] The proposed embodiments are directed to the rendering performance and visual quality of point clouds by introducing a fast and high-quality rendering pipeline. The proposed embodiments provide a method for encoding and a method for rendering with a corresponding apparatus.

[0134] In the method for encoding, according to an embodiment, a volumetric video frame is received as input, where the volumetric video frame is represented as a set of 3D samples. A patch generation process ( Figure 1 ; 102) converts each 3D sample into a number of 2D samples related to different projections. In patch packing ( Figure 1 ; 103), based on the generated patches, geometric and texture images ( Figure 1 ; 104, 105) are generated. Each patch is further projected onto two images (i.e., two layers), where the first layer is called the "near layer" or depth 0 layer, and the second layer is called the "far layer" or depth 1 layer. The first layer stores the lowest depth values of the point set of the patch, and the second layer stores the highest depth values of the point set of the patch. Thus, the frame is decomposed into the near and far components of the geometric structure and the corresponding attribute components. Additionally, an occupancy map ( Figure 1 ; 106) is created to indicate the occupied and unoccupied parts of the frame.

[0135] The encoder determines surface parameters for each patch. For example, the surface parameter can be the "rendering_thickness" parameter, which is calculated by the encoder based on the patch depth difference. The encoder encodes depth frames and decodes them to find the decoded depth values used by the decoder. The encoder uses the original maximum depth values from depth 0 and depth 1 layers and subtracts the decoded depth minimum from depth 0 and depth 1 layers. This value is the maximum depth difference between the original depth value and the decoded depth value, and the rendering thickness parameter is the maximum difference for all pixels for a given patch. The encoder can detect whether there are many differences between the thickness values for all pixels for a given patch. The encoder can detect whether there are many differences between the thickness values and can split patches with high differences into smaller patches.

[0136] Pre-computed surface parameters such as the rendering thickness mentioned above are stored in the encoded bitstream. This parameter can be signaled by the rendering thickness parameter, which is provided for each patch, for each layer, or for a single layer.

[0137] The rendering thickness parameter (8 bits) can be signaled in the encoded bitstream or signaled together with an additional SEI message. Instead, the rendering thickness parameter can be signaled at the block (e.g., 16x16) level, which provides a more accurate thickness parameter. The rendering thickness parameter can also be stored in the occupancy map instead of being signaled in the SEI message. An occupancy map value of 0 means the pixel is not occupied, and any other value is the direct rendering thickness value. Alternatively, 0 means unoccupied, 1 means the fill value between depth 0 and depth 1, and values of 2 or above give the actual rendering thickness. This will optimize small rendering thickness values and provide better compression. The encoder can also filter adjacent thickness values. For example, if the thickness values are 3, 2, 3, 3, the encoder can decide to set all values to 3.

[0138] It is also possible that some patches can use the rendering thickness value per pixel, while other patches can use the value per patch. There can be an additional 1-bit signal per patch to use the value per patch or the occupancy map value per pixel. However, as a further alternative, the rendering thickness can be signaled separately for each layer or only for a single layer.

[0139] According to another embodiment, color interpolation between layers can also or instead be signaled in the bitstream. Mode 0 does not interpolate color between layer 0 and layer 1, but mode 1 can interpolate color between layers.

[0140] According to another embodiment, the difference between the original depth0 and the decoded depth0 is calculated, and this difference (delta) is encoded into the occupancy map. This allows for lossless depth0 coordinates, as the encoded depth0 values are lossy while the delta values encoded into the occupancy map are lossless.

[0141] A method for rendering a point cloud may include the following steps:

[0142] The method begins with receiving a bitstream that has been encoded according to the foregoing embodiments. The bitstream includes at least a geometry image and a texture image. The method includes decoding corresponding patches from the bitstream. Figure 4 An example of a frame 410 including a patch 420 is shown. The decoded patch 420 is segmented into blocks 425 having a size of, for example, 16x16 pixels. For each block 425, a depth minimum and a depth maximum are determined. According to a first alternative, the depth minimum and the depth maximum are calculated in real time from the pixels of depth layer 0 and depth layer 1. According to a second alternative, the depth minimum and the depth maximum may be decoded from the patch metadata that has been stored by the encoder.

[0143] The determined depth minimum and depth maximum will be used to form a 3D bounding box (AABB) (also referred to as the "first bounding box") of the block (i.e., the entire pixel block) onto which the content of the patch is projected. Optionally, at this stage, the formed bounding box may be checked against the viewing frustum and culled if not visible. In the case where the point cloud data formed by the pixel block is visible, the block may be drawn into the frame buffer as a three-dimensional cube (12 triangles). Figure 5 A frame 510 including a block 525 in 3D is shown, where the block 525 includes rendered pixels 530.

[0144] In a second stage, each pixel 530 in the 3D block 525 is rendered into the frame buffer. For each pixel 530, a ray direction from the user's viewpoint to the coordinates of the pixel in 3D may be calculated. The viewpoint may be detected by a head orientation detector of the viewing device. Based on the ray, an "entry point" and an "exit point" are determined according to the intersection of the ray with the generated three-dimensional bounding box (AABB) of the block.

[0145] Figure 6 A block 625 of a patch in 2D is shown. In the block 625, the ray is indicated by 640. In Figure 6In it, the entry point is indicated by 641, and the exit point is indicated by 645. These two points 641, 645 will determine the 3D coordinates where the light will enter the bounding box (AABB) and the 3D coordinate points where the light exits the bounding box (AABB). These two points 641, 645 correspond to the 2D coordinates in block 625. Between these 2D coordinates, ray casting is performed so that each pixel 650 that intersects the ray between the "entry point" 641 and the "exit point" 645 is accessed in sequence starting from the "entry point" 641. These pixels are shown filled with the pattern indicated by 650. In ray casting (also known as "ray marching" or "ray tracing"), each pixel between the entry point 641 and the exit point 645 is evaluated to obtain a depth value.

[0146] Ray casting continues from one pixel to another along the viewing ray until the intersection point of the ray with the point cloud or the "exit point" is reached. To determine the intersection point, the depth values of depth layer 0 and depth layer 1 are obtained for each 2D block pixel. For each step of the ray casting, the depth value from the ray is also calculated. Now, if the depth value from the ray is between depth layer 0 and depth layer 1, it is determined as the intersection point between the point cloud content and the ray. This intersection point can also be calculated more accurately by forming another bounding box (AABB) (also known as the "second bounding box") from the depth difference (depth0 - depth1) and the two-dimensional coordinates of the pixel being processed. The second bounding box is the span of the point positions represented by a single pixel. The size of the second bounding box can be 1x1x(depth1 - depth0 + thickness), that is, the size of a single pixel with the depth of the surface projected onto that one pixel. The depth difference can be increased by rendering the surface thickness parameter.

[0147] If the intersection point is not determined, the next 2D pixel is continued to be processed.

[0148] After the intersection point has been reached, the GPU will compare the intersection point depth value with the depth value written to the depth buffer. In the case where the intersection point is close to the viewer, the new depth value is written to the depth buffer, and the newly calculated color value will be written. The color value of texture layer 0 or texture layer 1 (depending on which layer is closest to the intersection point) is thus obtained. Alternatively, an interpolated color between texture layer 0 and texture layer 1 can be used.

[0149] Due to the depth compression artifact, the values of depth 0 and depth 1 can be changed so that holes can be created when using lossy compression. This can be detected in the encoder, and the surface thickness value can be increased in those areas to effectively fill these holes when finally rendering the model. Figure 7 The case of a simple one depth layer is shown, where the depth value is changed due to compression. Figure 7Shows the original surface 710 and the decoded single layer 720. These holes and the new surface rendering thickness values evaluated by the encoder are calculated according to depth pixels so that more surface thickness is added to repair the holes. Figure 7 Shows the decoded signal with potential rendering surface thickness values calculated as (0,0,2,1,0,0,0,1,2,0) 730. Reconstructing the point cloud with the rendering surface thickness parameter allows the surface to be extended so that new points are created to make the surface thicker. The new points generated by the rendering surface thickness parameter are shown as dark pixels 725. There are various ways to signal the rendering surface thickness parameter, ranging from per-pixel values to per-patch values. Note that if point reconstruction is used, new points need to be generated based on the rendering thickness parameter. However, this is not the case if the previously described ray-tracing process is used, because the size of the bounding box (AABB) can be adjusted by the rendering surface thickness. Therefore, if more points are constructed, there will not be an additional performance cost for adjusting the box size as before.

[0150] In lossless and lossy encoded frames, there can be additional specific patches (missed points patched, PCM) that contain only XYZ coordinates. Since the points are in random positions, it is not possible to perform ray casting on these types of patches. However, it is possible to render these special patches by rendering each point with a cube, because the count of missed points is low. This will result in the same quality as ray-tracing those points (a single cube per point).

[0151] Figure 8 Is a flowchart showing a method for encoding according to an embodiment. A method includes receiving 810 a video presentation frame, where the video presentation represents three-dimensional data; generating 820 one or more patches from the video presentation frame; partitioning 830 the patches of the video frame into one or more blocks; determining 840 surface parameters for each patch according to information related to the one or more blocks of the patch; encoding 850 the determined surface parameters into a bitstream; and storing 860 the encoded bitstream for transmission to a rendering device.

[0152] Figure 9FIG. is a flowchart showing a method for decoding according to another embodiment. A method includes: receiving 910 a compressed bitstream related to a video presentation, the bitstream including at least a geometric image and a texture image; decoding 920 data related to one or more three-dimensional blocks of a video frame from the received bitstream; for each block of the video frame, determining 930 information about one or more surface parameters; generating 940 a bounding box of the three-dimensional block according to the one or more surface parameters; for each pixel of the three-dimensional block, calculating 950 a ray direction from a viewpoint to the coordinates of the pixel; determining 960 at least two points according to the intersection of the ray and the generated bounding box; performing ray casting 970 on the points between the determined at least two points until a condition for completing ray casting is satisfied; reconstructing 980 three-dimensional data from the geometric image and the texture image according to the information about the one or more surface parameters of the block; and reconstructing 990 the video presentation for rendering according to the reconstructed three-dimensional data.

[0153] The three-dimensional data in the above example may be a point cloud.

[0154] The apparatus according to an embodiment includes at least: a component for receiving a video presentation frame, wherein the video presentation represents three-dimensional data; a component for generating one or more patches from the video presentation frame; a component for dividing the patches of the video frame into one or more blocks; a component for determining surface parameters for each patch according to information related to the one or more blocks of the patch; a component for encoding the determined surface parameters into a bitstream; and a component for storing the encoded bitstream for transmission to a rendering device.

[0155] These components include at least one processor and a memory including computer program code, wherein the processor may further include processor circuitry. The memory and the computer program code are configured to, together with the at least one processor, cause the apparatus to execute Figure 8 the method of the flowchart in

[0156] An apparatus according to another embodiment comprises at least: means for receiving a compressed bitstream related to a video presentation, the bitstream comprising at least a geometry image and a texture image; means for decoding data related to one or more three-dimensional blocks of a video frame from the received bitstream; means for determining, for each block of the video frame, information regarding one or more surface parameters; means for generating a bounding box of the three-dimensional block according to the one or more surface parameters; means for calculating, for each pixel of the three-dimensional block, a ray direction from a viewpoint to the coordinates of the pixel; means for determining at least two points according to the intersection of the ray and the generated bounding box; means for performing ray casting on points between the at least two determined points until a condition for completing ray casting is satisfied; means for reconstructing three-dimensional data from the geometry image and the texture image according to the information regarding one or more surface parameters of the block; means for reconstructing the video presentation for rendering according to the reconstructed three-dimensional data.

[0157] These means comprise at least one processor and a memory including computer program code, wherein the processor may also comprise processor circuitry. The memory and the computer program code are configured to, with the at least one processor, cause the apparatus to perform Figure 9 the method of the flowchart in

[0158] The three-dimensional data in the above example may be a point cloud.

[0159] Figure 10 A system and apparatus for viewing volumetric video according to the proposed embodiments are shown. The task of the system is to capture sufficient visual and auditory information from a specific location so that one or more viewers physically located at different positions can and optionally at some subsequent time in the future achieve a convincing reproduction of the experience or presentation at that location. Such a reproduction requires more information than can be captured by a single camera or microphone so that viewers can use their eyes and ears to determine the distance and position of objects within the scene. To create a pair of images with disparity, two camera sources are used. In a similar way, the human auditory system needs to use at least two microphones to be able to perceive the direction of sound (it is well known that stereo is created by recording two audio channels). The human auditory system can detect cues such as in the timing difference of the audio signals to detect the direction of sound.

[0160] Figure 10The system can include three parts: image sources 1001, 1003, server 1005, and rendering device 1007. The image sources can be video capture device 1001, which includes two or more cameras with overlapping fields of view such that the view area around the video capture device is captured from at least two cameras. The video capture device 1001 can include multiple microphones (not shown in the figures) to capture the timing and phase differences of audio originating from different directions. The video capture device 1001 can include high-resolution orientation sensors such that the orientations (viewing directions) of the multiple cameras can be detected and recorded. The video capture device 1001 includes or is functionally connected to computer processor PROCESSOR1 and memory MEMORY1, which includes computer program code for controlling the video capture device 1001. The image stream captured by the video capture device 1001 can be stored in memory MEMORY1 and / or removable memory MEMORY9 for use in another device (e.g., viewer), and / or transmitted to server 1005 using communication interface COMMUNICATION1.

[0161] Alternatively, or in addition to the video capture device 1001 that creates an image stream or multiple such image streams, one or more image source devices 1003 for synthetic images can exist in the system. Such image source devices 1003 for synthetic images can use a computer model of a virtual world to calculate the various image streams they send. For example, image source 703 can calculate N video streams corresponding to N virtual cameras located at virtual viewing positions. When viewing using such a set of synthetic video streams, the viewer can see a three-dimensional virtual world. The image source device 1003 includes or is functionally connected to computer processor PROCESSOR3 and memory MEMORY3, which includes computer program code for controlling the image source device 1003. In addition to the video capture device 1001, there can also be a storage, processing, and data stream service network. For example, there can be server 1005 or multiple servers that store the output from the video capture device 1001 or image source device 1003. Server 705 includes or is functionally connected to computer processor PROCESSOR5 and memory MEMORY5, which includes computer program code for controlling server 1005. Server 1005 can be connected to sources 1001 and / or 1003 via a wired or wireless network connection or both, and can be connected to viewer device 1009 via communication interface COMMUNICATION5.

[0162] To view the captured or created video content, one or more viewer devices 1009 (also referred to as playback devices) may exist. These viewer devices 1009 may have one or more displays and may include or be functionally connected to a rendering module 1007. The rendering module 1007 includes a computer processor PROCESSOR7 and a memory MEMORY7, and the memory includes computer program code for controlling the viewer device 1009. The viewer device 1009 may include a data stream receiver for receiving a video data stream from a server and for decoding the video data stream. The data stream may be received over a network connection via a communication interface or from a memory device 1011, such as a memory card. The viewer device 1009 may have a graphics processing unit for processing the data into a format suitable for viewing. The viewer device 709 may be a high-resolution stereoscopic image head-mounted display for viewing the rendered volumetric video sequence. The head-mounted display may have an orientation detector 1013 and stereophonic audio headphones. According to an embodiment, the viewer device 1009 is a 3D technology-enabled display (for displaying stereoscopic video), and the rendering device 1007 may have a head orientation detector 1015 connected thereto. Alternatively, the viewer device 1009 may include a 2D display, as volumetric video rendering may be done in 2D by rendering viewpoints from a single eye rather than a stereoscopic eye pair. Any one of the devices 1001, 1003, 1005, 1007, 1009 may be a computer or a portable computing device, or may be connected to such a device. Such a device may have computer program code for performing the methods according to the various examples described herein.

[0163] As described above, the viewer device may be a head-mounted display (HMD). The head-mounted display includes two screen portions or two screens for displaying left-eye and right-eye images. The display is close to the eyes, and thus lenses are used to make the images easy to view and to spread the images out to cover as much of the eye's field of view as possible. The device is attached to the user's head so that it remains in place even when the user turns their head. The device may have an orientation detection module for determining head movement and the direction of the head. The head-mounted display provides the user with a three-dimensional (3D) perception of the recorded / streamed content.

[0164] Video material captured or generated by any image source may be provided to an encoder, which transforms the input video into a compressed representation suitable for storage / transmission. The compressed video is provided for a decoder, which may decompress the compressed video representation back into a viewable form. The encoder may be located in the image source or in a server. The decoder may be located in a server or in a viewer, such as an HMD. The encoder may discard some information in the original video sequence in order to represent the video in a more compact form (i.e., at a lower bitrate). InFigure 11 An example of the encoding process is shown. Figure 11 The image to be encoded (I n ); the predicted representation of the image block (P’ n ); the prediction error signal (D n ); the reconstructed prediction error signal (D’ n ); the preliminary reconstructed image (I’ n ); the final reconstructed image (R’ n ); the transform (T) and the inverse transform (T -1 ); the quantization (Q) and the inverse quantization (Q -1 ); the entropy encoding (E); the reference frame memory (RFM); the inter-frame prediction (P inter ); the intra-frame prediction (P intra ); the mode selection (MS) and the filtering (F). In Figure 12 An example of the decoding process is shown. Figure 12 The predicted representation of the image block (P’ n ); the reconstructed prediction error signal (D’ n ); the preliminary reconstructed image (I’ n ); the final reconstructed image (R’ n ); the inverse transform (T -1 ); the inverse quantization (Q -1 ); the entropy decoding (E -1 ); the reference frame memory (RFM); the prediction (inter-frame or intra-frame) (P); and the filtering (F).

[0165] Various embodiments can provide several advantages. For example, compared to point rendering and automatic hole filling with zero cost, the proposed embodiments provide excellent visual quality. In addition, since the rendering speed is faster than point rendering, energy-efficient or more complex point clouds can be rendered. Additionally, various embodiments provide flexible thickness signaling per patch or per pixel.

[0166] Various embodiments can be implemented with the help of computer program code that resides in a memory and causes a related device to execute a method. For example, a device can include circuitry and electronics for processing, receiving, and transmitting data, computer program code in a memory, and a processor that causes the device to execute the features of the embodiment when the computer program code is run. Further, a network device (such as a server) can include circuitry and electronics for processing, receiving, and transmitting data, computer program code in a memory, and a processor that causes the network device to execute the features of the embodiment when the computer program code is run. The computer program code includes one or more operational characteristics. The operational characteristics are defined by configuration by the computer based on the type of the processor, wherein the system can be connected to the processor via a bus, and wherein the programmable operational characteristics of the system according to the embodiment at least include steps defined by the Figure 8 or Figure 9 flowchart as described.

[0167] If desired, the different functions discussed herein can be performed in a different order and / or concurrently with other functions. Additionally, if desired, one or more of the above functions and embodiments can be optional or can be combined.

[0168] Although aspects of the embodiments are set forth in the independent claims, other aspects include other combinations of features from the described embodiments and / or dependent claims with the features of the independent claims, rather than only the combinations explicitly set forth in the claims.

[0169] It should also be noted here that although several example embodiments are described above, these descriptions should not be regarded as restrictive. Instead, various variations and modifications can be made thereto without departing from the scope of the present disclosure as defined in the appended claims.

Claims

1. A method for encoding, comprising: receiving a video presentation frame, wherein the video presentation represents three-dimensional data; generating one or more patches from the video presentation frame; dividing the patches of the video frame into one or more blocks; determining a surface thickness parameter for each patch according to the patch depth difference; and encoding the determined surface thickness parameter into a bitstream.

2. A method for decoding, comprising: receiving a compressed bitstream related to a video presentation, the compressed bitstream at least including a geometric image and a texture image; decoding data related to one or more blocks of a video frame from the compressed bitstream; determining information about one or more surface thickness parameters for the blocks of the video frame; generating a three-dimensional bounding box of the block according to the one or more surface thickness parameters; calculating a ray direction from a viewpoint to the coordinates of the pixel for the pixels of the block; determining at least two points according to the intersection of the ray and the generated three-dimensional bounding box; performing ray casting on the points between the determined at least two points until a condition for completing the ray casting is satisfied; reconstructing three-dimensional data from the geometric image and the texture image according to the information about one or more surface thickness parameters of the block; reconstructing the video presentation for rendering according to the reconstructed three-dimensional data.

3. An apparatus for encoding a bitstream, comprising: a component for receiving a video presentation frame, wherein the video presentation represents three-dimensional data; a component for generating one or more patches from the video presentation frame; a component for dividing the patches of the video frame into one or more blocks; a component for determining a surface thickness parameter for each patch according to the patch depth difference; and a component for encoding the determined surface thickness parameter into a bitstream.

4. The apparatus according to claim 3, further comprising: a component for encoding the surface thickness parameter into a supplementary enhancement information (SEI) message.

5. The apparatus according to claim 3 or 4, wherein the surface thickness parameter is determined for each block.

6. The apparatus according to claim 3 or 4, wherein the surface thickness parameter is encoded into an occupancy map.

7. An apparatus for decoding a bitstream, comprising: a component for receiving a compressed bitstream related to a video presentation, the compressed bitstream at least including a geometric image and a texture image; a component for decoding data related to one or more blocks of a video frame from the compressed bitstream; a component for determining information about one or more surface thickness parameters for the blocks of the video frame; a component for generating a three-dimensional bounding box of the block according to the one or more surface thickness parameters; a component for calculating a ray direction from a viewpoint to the coordinates of the pixel for the pixels of the block; a component for determining at least two points according to the intersection of the ray and the generated three-dimensional bounding box; a component for performing ray casting on the points between the determined at least two points until a condition for completing the ray casting is satisfied; and A component for reconstructing three-dimensional data from the geometric image and the texture image according to the information on one or more surface thickness parameters of a block; A component for reconstructing a video presentation for rendering according to the reconstructed three-dimensional data.

8. The apparatus according to claim 7, wherein, the information on one or more surface thickness parameters is decoded from the compressed bitstream.

9. The apparatus according to claim 7, wherein, the information on one or more surface thickness parameters is determined from pixels of various depth layers.

10. The apparatus according to any one of claims 7 to 9, wherein, the condition for completing the ray casting is determined by: determining the depth value of the depth layer at the position of the pixel, calculating the depth value from the viewing ray, and comparing the depth value from the viewing ray with the determined depth value.

11. The apparatus according to any one of claims 7 to 9, wherein, the condition for completing the ray casting is determined from another bounding box formed by the depth difference and the pixel coordinates.

Citation Information

Patent Citations

  • Method and System for Rendering 3D Distance Fields

    US20100085357A1