Video-Based Mesh Compression

By projecting mesh surface data and encoding connectivity using video-based techniques, the method addresses the lack of point connectivity transmission in current standards, achieving efficient and detailed 3D mesh compression compatible with existing systems.

JP7672629B2Active Publication Date: 2025-05-08SONY GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023521420
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-17
Filing Date
2021-09-29
Publication Date
2025-05-08
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

Current standards for encoding 3D point clouds do not have a mechanism for transmitting point connectivity, which is essential for 3D mesh compression, and existing methods for extending video-based point cloud compression to meshes are inefficient and can result in data loss.

Method used

The method involves projecting mesh surface data and encoding connectivity data using video-based compression techniques, encapsulating this data in a vertex video component structure that allows for hierarchical encoding of mesh connectivity detail.

Benefits of technology

This approach extends the functionality of existing standards to efficiently encode 3D meshes by enabling progressive mesh coding and preserving detailed connectivity information, while also being compatible with existing video-based compression systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007672629000012
    Figure 0007672629000012
  • Figure 0007672629000013
    Figure 0007672629000013
  • Figure 0007672629000014
    Figure 0007672629000014
Patent Text Reader

Abstract

This document describes a method for compressing 3D mesh data using a video representation of mesh surface data projections and connectivity data. The method utilizes 3D surface patches to represent sets of connected triangles on a mesh surface. The projected surface data is stored in patches (mesh patches) and encoded into atlas data. The mesh connectivity, i.e., the vertices and triangles of the surface patches, is coded using video-based compression techniques. This data is encapsulated in a new video component called vertex video data, and the disclosed architecture enables hierarchical mesh coding by separating vertex sets into layers to create levels of detail for the mesh connectivity. This approach extends the capabilities of the V3C (volumetric video-based) standard currently used for coding point clouds and multiview+depth content.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority under 35 U.S.C. § 119 to U.S. Provisional Patent Application No. 63 / 088,705, filed October 7, 2020, entitled “VIDEO BASED MESH COMPRESSION,” and U.S. Provisional Patent Application No. 63 / 087,958, filed October 6, 2020, entitled “VIDEO BASED MESH COMPRESSION,” each of which is incorporated herein by reference in its entirety for all purposes.

[0002] The present invention relates to three-dimensional graphics, and more particularly to coding of three-dimensional graphics. [Background technology]

[0003] Recently, a new method for compressing volumetric content such as point clouds based on a 3D to 2D projection has been standardized. This method, also known as V3C (visual volumetric video-based compression), maps 3D volumetric data into multiple 2D patches, which are then further organized into an atlas image that is then encoded by a video encoder. The atlas image corresponds to the geometry of the points, their respective textures, and an occupancy map that indicates which locations should be considered for point cloud reconstruction.

[0004] In 2017, MPEG published a call for proposals (CfP) for point cloud compression. Currently, after evaluating several proposals, MPEG is considering two different point cloud compression techniques: 3D native encoding techniques (based on octree and similar encoding methods) or 3D to 2D projection followed by traditional video encoding. For dynamic 3D scenes, MPEG uses a test model software (TMC2) based on patch surface modeling, projection of the patches from 3D to a 2D image, and encoding of the 2D image using a video encoder such as HEVC. This method has been proven to be more efficient than native 3D encoding and to achieve competitive bitrates with acceptable quality. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] U.S. Patent Application Publication No. 17 / 161,300 Summary of the Invention [Problem to be solved by the invention]

[0006] Following the success of the projection-based method (also known as video-based method or V-PCC) for 3D point cloud coding, future versions of this standard are expected to include additional 3D data, such as 3D meshes. However, the current version of the standard is only suitable for transmitting a set of unconnected points, and therefore there is no mechanism for transmitting point connectivity as required for 3D mesh compression.

[0007] Methods have also been proposed to extend the functionality of V-PCC to meshes. One possible method is to use V-PCC to encode the vertices, and then use a mesh compression method such as TFAN or Edgebreaker to encode the connectivity. A limitation of this method is that the original mesh must be dense so that the point cloud generated from the vertices can be efficiently encoded after projection without becoming sparse. Furthermore, the order of the vertices affects the connectivity encoding, so different methods have been proposed to reorganize the mesh connectivity. Another method to encode sparse meshes is to use the raw patch data to encode the vertex positions in 3D. Since raw patches directly encode (x,y,z), in this method all vertices are encoded as raw data, while the connectivity is encoded by a similar mesh compression method as described above. Note that raw patches can send vertices in any preferred order, and therefore the order resulting from the connectivity encoding can be used. While this method can encode sparse point clouds, raw patches are less efficient at encoding 3D data, and additional data such as triangular facet attributes may be lost from this approach. [Means for solving the problem]

[0008] Described herein is a method for compressing 3D mesh data using a projection of mesh surface data and a video representation of connectivity data. The method utilizes 3D surface patches to represent sets of connected triangles on a mesh surface. The projected surface data is stored in patches (mesh patches) and encoded into the atlas data. The mesh connectivity, i.e., the vertices and triangles of the surface patches, is coded using a video-based compression technique. This data is encapsulated in a new video component named vertex video data, and the disclosed structure enables progressive mesh coding by separating the vertex sets in layers to create levels of detail of the mesh connectivity. This approach extends the capabilities of the V3C (volumetric video-based) standard currently used for coding point clouds and multiview+depth content.

[0009] In one aspect, a method includes performing mesh voxelization on an input mesh, performing patch generation to segment the mesh into patches including a rasterized mesh surface and vertex position and connectivity information, generating a visual volumetric video-based compression (V3C) image from the rasterized mesh surface, performing video-based mesh compression using the vertex position and connectivity information, and generating a V3C bitstream based on the V3C image and the video-based mesh compression. The vertex position and connectivity information includes triangulation information of the surface patch. Data resulting from performing video-based mesh compression using the vertex position and connectivity information is encapsulated in a vertex video component structure. The vertex video component structure enables hierarchical mesh coding by separating sets of vertices in layers to generate mesh connectivity levels of detail. If only one layer is implemented, the video data is embedded in the occupancy map. The connectivity information is generated using a surface reconstruction algorithm including Poisson surface reconstruction or ball pivoting. Generating the V3C image from the rasterized mesh surface includes combining untracked mesh information and tracked mesh information. The method further includes performing an edge removal filter on the two-dimensional projected patch region.The method further includes performing a patch-based surface subdivision of the connectivity information.

[0010] In another aspect, an apparatus includes a non-transitory memory that stores an application for performing mesh voxelization on an input mesh, performing patch generation to segment the mesh into patches including a rasterized mesh surface and vertex position and connectivity information, generating a visual volumetric video-based compression (V3C) image from the rasterized mesh surface, performing video-based mesh compression using the vertex position and connectivity information, and generating a V3C bitstream based on the V3C image and the video-based mesh compression, and a processor coupled to the memory and configured to process the application. The vertex position and connectivity information includes triangulation information of the surface patch. Data resulting from performing video-based mesh compression using the vertex position and connectivity information is encapsulated in a vertex video component structure. The vertex video component structure enables hierarchical mesh encoding by separating sets of vertices into layers to generate mesh connectivity levels of detail. If only one layer is implemented, the video data is embedded in the occupancy map. The connectivity information is generated using a surface reconstruction algorithm including Poisson surface reconstruction or ball pivoting. Generating the V3C image from the rasterized mesh surface includes combining the untracked mesh information and the tracked mesh information. The application is further configured to perform an edge removal filter on the two-dimensional projected patch region. The application is further configured to perform a patch-based surface subdivision of the connectivity information.

[0011] In another aspect, a system includes one or more cameras that capture three-dimensional content, and an encoder that encodes the three-dimensional content by performing mesh voxelization on the input mesh, performing patch generation to segment the mesh into patches including a rasterized mesh surface and vertex position and connectivity information, generating a visual volumetric video-based compression (V3C) image from the rasterized mesh surface, performing video-based mesh compression using the vertex position and connectivity information, and generating a V3C bitstream based on the V3C image and the video-based mesh compression. The vertex position and connectivity information includes triangulation information of the surface patch. Data resulting from performing video-based mesh compression using the vertex position and connectivity information is encapsulated in a vertex video component structure. The vertex video component structure enables hierarchical mesh encoding by separating sets of vertices into layers to generate mesh connectivity levels of detail. If only one layer is implemented, the video data is embedded in the occupancy map. The connectivity information is generated using a surface reconstruction algorithm including Poisson surface reconstruction or ball pivoting. Generating the V3C image from the rasterized mesh surface includes combining untracked mesh information and tracked mesh information. The encoder is configured to perform an edge removal filter on the two-dimensional projected patch region. The encoder is configured to perform a patch-based surface subdivision of the connectivity information. [Brief description of the drawings]

[0012] [Figure 1] FIG. 2 illustrates a flowchart of a method for performing V3C mesh coding according to some embodiments. [Diagram 2] FIG. 1 is an illustration of mesh voxelization according to some embodiments. [Diagram 3] FIG. 1 is an illustration of patch generation according to some embodiments. [Figure 4] FIG. 1 is a diagram of patch rasterization according to some embodiments. [Diagram 5]FIG. 1 is an illustration of V3C image generation according to some embodiments. [Figure 6] FIG. 2 illustrates an image of vertex video data according to some embodiments. [Figure 7] 1 illustrates an example level of detail generation image in accordance with some embodiments. [Figure 8] FIG. 1 is an illustration of a mesh according to some embodiments. [Figure 9] FIG. 1 is an illustration of a mesh reconstruction according to some embodiments. [Figure 10] FIG. 1 is an illustration of a high-level syntax and images that enable sending a mixture of point cloud patches and mesh patches, according to some embodiments. [Figure 11A] FIG. 2 is a diagram of combined untracked and tracked mesh information according to some embodiments. [Figure 11B] FIG. 2 is a diagram of combined untracked and tracked mesh information according to some embodiments. [Figure 12] 1A-1C are diagrams illustrating example images of patch-based edge removal in accordance with some embodiments. [Figure 13] FIG. 1 illustrates an example image of patch-based clustering decimation according to some embodiments. [Figure 14] FIG. 1 illustrates an example image of patch-based surface subdivision according to some embodiments. [Figure 15] FIG. 1 illustrates an example image of patch-based surface reconstruction in accordance with some embodiments. [Figure 16] FIG. 1 is an illustration of triangle edge detection according to some embodiments. [Figure 17] FIG. 13 is an illustration of segmented edges of triangles separated based on color, according to some embodiments. [Figure 18] FIG. 13 is an illustration of segmented edges of a joining triangle, according to some embodiments. [Figure 19] FIG. 13 is an illustration of edge resizing and rescaling according to some embodiments. [Figure 20] FIG. 13 is an illustration of edge resizing and rescaling according to some embodiments. [Figure 21] FIG. 13 is an illustration of edge resizing and rescaling according to some embodiments. [Figure 22] FIG. 13 is an illustration of edge resizing and rescaling according to some embodiments. [Figure 23] FIG. 13 is an illustration of edge resizing and rescaling according to some embodiments. [Figure 24] FIG. 1 is a block diagram of an example computing device configured to implement a video-based mesh compression method, according to some embodiments. [Diagram 25] FIG. 1 is a diagram of a system configured to perform video-based mesh compression in accordance with some embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] Described herein is a method for compressing 3D mesh data using a projection of mesh surface data and a video representation of connectivity data. The method utilizes 3D surface patches to represent sets of connected triangles on a mesh surface. The projected surface data is stored in patches (mesh patches) and encoded into the atlas data. The mesh connectivity, i.e., the vertices and triangles of the surface patches, is encoded using a video-based compression technique. This data is encapsulated in a new video component named vertex video data, and the disclosed structure enables hierarchical mesh coding by separating the vertex sets in layers to create levels of detail of the mesh connectivity. This approach extends the capabilities of the V3C (volumetric video-based) standard currently used for coding point clouds and multiview+depth content.

[0014] In 3D point cloud coding using video encoders, 3D to 2D projection is important to generate videos representing the point cloud. The most efficient way to generate these videos is to segment the surface of the object by using 3D patches and use orthogonal projection to generate segmented depth images that are bundled together and used as input for the video encoder. Additionally, points not captured by the projection step can also be directly encoded in the video signal. Current point cloud standards cannot code 3D meshes because there is no prescribed way to code mesh connectivity. Furthermore, the standard does not work well when vertex data is sparse, as it cannot exploit correlation between vertices.

[0015] This document describes how to perform mesh encoding using the V3C standard for volumetric data encoding. It describes how to segment the mesh surface and propose joint surface sampling and 2D patch generation. Local connectivity and vertex positions projected onto the 2D patch are encoded for each patch. It also describes how to signal connectivity and vertex positions to allow reconstruction of the original input mesh. It also describes how to map vertices and connectivity into video frames and use video encoding tools to encode the mesh connectivity data into a video sequence called vertex video data.

[0016] 1 shows a flowchart of a method for performing V3C mesh encoding according to some embodiments. In step 100, an input mesh is received or acquired. For example, the input mesh is downloaded (e.g., from a network device) or acquired / captured by a device (e.g., a camera or an autonomous vehicle).

[0017] In step 102, we perform mesh voxelization. Meshes can have vertex positions in floating point, so we transform these positions to integer space. V-PCC and V3C assume a voxelized point cloud.

[0018] In step 104, patch generation (or creation) is performed. Patch generation includes normal calculation, adjacency calculation, initial segmentation, refinement, patch projection, and patch rasterization. Normal calculation involves calculating the normal (cross product of the triangle edges) for each triangle. Adjacency calculation involves calculating the adjacency of each triangle (e.g., which triangles in the mesh are adjacent or touching the current triangle or other triangles). Initial segmentation includes classifying normals according to their orientation. For example, a triangle's normal can point up, down, left, right, front or back, and can be classified based on direction / orientation. In some embodiments, triangles are colored based on the orientation of their normal (e.g., all triangles whose normals point up are colored green). Refinement involves identifying outliers (e.g., a single red triangle surrounded by blue triangles) and smoothing the outliers (e.g., modifying the single red triangle to match its neighbors that are blue). Refinement is performed by analyzing the neighborhood and smoothing the orientation (e.g., adjusting the orientation of normals). Once a smooth surface is obtained, patch projection is performed, which projects patches of a particular triangle class (e.g., based on the orientation). In this projection, the vertices and connectivity are shown on the patch. For example, the body and face in this example are separate projections since there are triangles of different classes that separate them. However, V3C and V-PCC do not understand this, but rather points, and therefore rasterize this projection (e.g., sample the points on the surface, including the distance of the points, to generate a shape image and surface attributes). The rasterized mesh surface is very similar to the V3C image.

[0019] The patch generation results in a rasterized mesh surface, vertex positions and connectivity. In step 106, the rasterized mesh surface is used in V3C image generation / creation. In step 108, the vertex positions and connectivity are used for mesh encoding (e.g., video-based mesh compression). In step 110, a V3C bitstream is generated from the generated V3C image and the base mesh encoding. In some embodiments, the mesh encoding does not involve further encoding, and the vertex positions and connectivity go directly to the V3C bitstream.

[0020] The V3C bitstream enables point cloud reconstruction in step 112 and / or mesh construction in step 114. The point cloud and / or mesh can be extracted from the V3C bitstream, which allows great flexibility. In some embodiments, fewer or more steps are performed. In some embodiments, the order of the steps is changed.

[0021] The methods described herein are related to U.S. patent application Ser. No. 17 / 161,300, entitled “PROJECTION-BASED MESH COMPRESSION,” filed Jan. 28, 2021, which is incorporated by reference in its entirety for all purposes.

[0022] To support voxelization, mesh scaling and offset information is sent in the Atlas Adaptation Parameter Set (AAPS). Available camera parameters can be used. Alternatively, a new syntax element for voxelization (using only scaling and offset) is introduced. Below is an example syntax: TIFF0007672629000001.tif41141 aaps_voxelization_parameters_present_flag equal to 1 specifies that voxelization parameters are present in the current atlas adaptation parameter set. aaps_voxelization_parameters_present_flag equal to 0 specifies that there are no voxelization parameters for the current adaptation parameter set. vp_scale_enabled_flag equal to 1 indicates the presence of a scale parameter for the current voxelization. vp_scale_enabled_flag equal to 0 indicates that there are no scale parameters for the current voxelization. vp_scale_enabled_flag is inferred to be equal to 0 if not present. vp_offset_enabled_flag equal to 1 indicates the presence of offset parameters for the current voxelization. vp_offset_enabled_flag equal to 0 indicates that there are no offset parameters for the current voxelization. vp_offset_enabled_flag is inferred to be equal to 0 if not present. vp_scale is the current scale for voxelization, set by 2 -16 vp_scale is specified in increments of 1 to 2, inclusive. 32 vp_scale is in the range -1. If not present, vp_scale is 2. 16 The value of Scale is calculated as follows: Scale=vp_scale÷2 16 vp_offset_on_axis[d] is the offset along the d axis for the current voxelization, Offset[d], multiplied by 2. -16 The value of vp_offset_on_axis[d] is inclusive of -2. 31 ~2 31It is in the range of -1, and d is from 0 to 2. When the value of d is equal to 0, 1, and 2, it corresponds to the X, Y, and Z axes respectively. vp_offset_on_axis[d] is presumed to be equal to 0 if it does not exist. Offset[d]=vp_offset_on_axis[d])2 16

[0023] This process specifies the inverse voxelization process from voxelized decoded vertex values to floating-point reconstruction values. The following applies. In the case of (n = 0; n < VertexCnt; n++), In the case of (k = 0; k < 3; k++), vertexReconstructed[n][k]=Scale * (decodedVertex[n][k])+Offset[k]

[0024] Figure 2 shows diagrams of mesh voxelization according to some embodiments. Each frame has a different bounding box. Obtain the bounding boxes of each frame (for example, frames 200 at t = 1, t = 16, and t = 32). Next, calculate the sequence bounding box 202 as sequenceBB=(minPoint,maxPoint) from many bounding boxes. The sequence bounding box 202 contains all vertices regardless of the frame. To fit the maximum range within the range determined by bitdepth, maxRange=max(maxPoint16[0..2]-minPoint16[0..2]), scale=(2 bitdepthCalculate the minimum value as voxelizedpoint=floor((originalPoint-minPoint16) / scale16). Scale and shift the result by the minimum as voxelizedpoint=floor((originalPoint-minPoint16) / scale16). The scale and shift amounts can be user-defined or computer-generated based on a learning algorithm (e.g., by analyzing bounding boxes and automatically calculating the scale and shift amounts). These values ​​are stored in the AAPS as offset=minPoint16, scale=scale16. The input parameters (modelScale) are as follows: (-1): Automatically calculate scale per frame (0): Automatically calculate sequence scale (>1): User-defined scale

[0025] In some embodiments, mesh voxelization involves converting floating point values ​​of the positions of points of the input mesh to integers. The precision of the integers can be set by the user or automatically. In some embodiments, mesh voxelization involves shifting the values ​​so that there are no negative numbers.

[0026] For example, negative numbers occur when the original mesh lies below an axis. The mesh is shifted and / or scaled to avoid negative and non-integer values ​​throughout the mesh voxelization. In one embodiment, once the lowest vertex value is found that is less than zero, the value may be shifted so that the lowest vertex value is greater than zero. In some embodiments, the range of values ​​is fitted (e.g., by scaling) within a specified bit range, such as 11 bits.

[0027] Voxelized mesh 210 is the original mesh after scaling and shifting, for example, voxelized mesh 210 is the original mesh after growing and shifting to be positive values ​​only, which in some cases is advantageous for encoding.

[0028] Voxelization may result in degenerate triangles (vertices that occupy the same position), but the degenerate vertices are removed by the encoding procedure, resulting in an increase in the number of vertices due to mesh segmentation (e.g., a remove duplicates vertices filter can be used to reduce the number of vertices). As an example, the original mesh had 20,692 vertices, 39,455 faces, and 20,692 points in the voxelized vertices, while the reconstructed mesh had 27,942 vertices, 39,240 faces, and the reconstructed point cloud had 1,938,384 points.

[0029] FIG. 3 shows a diagram of patch generation according to some embodiments. As described, patch generation involves normal calculation, adjacency calculation, initial segmentation (or normal categorization), and segmentation refinement (or category refinement). Normal calculation for each triangle involves a cross product between the edges of the triangle. Adjacency calculation determines whether the triangles share a vertex, and if so, determines that the triangles are adjacent. Initial segmentation and segmentation refinement are performed similarly to V-PCC by analyzing the normal orientation, classifying the normal orientation (e.g., up, down, left, right, front, back), determining whether the normal orientation is classified differently compared to neighboring normals that are all similarly classified (e.g., the initial patch was classified as up-facing, whereas most or all patches are classified as front-facing), and then changing the classification of the patch's normal to match the orientation of the neighboring normals (e.g., changing the classification of the initial patch to front-facing).

[0030] As described, patch generation is performed to segment the mesh into patches. In patch generation, 1) a rasterized mesh surface and 2) vertex position and connectivity information are also generated. The rasterized mesh surface is a set of points that undergoes V3C image generation or V-PCC image generation and is encoded as a V3C image or V-PCC image. The vertex position and connectivity information is received for base mesh encoding.

[0031] The patch generation described here is similar to that in V-PCC. However, instead of calculating normals per point, we calculate normals per triangle. We use the cross product between edges to calculate normals per triangle to determine normal vectors. We then categorize the triangles according to their normals. For example, we divide the normals into n (e.g., 6) categories, such as front, back, top, bottom, left and right. The normals are shown in different colors to indicate the initial segmentation. In FIG. 3, different colors, such as black and light gray, are shown in grayscale to indicate different normals. It may be difficult to see, but the top surface (e.g., the top of the person's head, the top of the ball and the top of the sneakers, etc.) is one color (e.g., green), the first surface of the person / ball is very dark and represents another color (e.g., red), the bottom of the ball is another color (e.g., purple), and the front of the person and ball, which is mostly light gray, represents another color (e.g., cyan).

[0032] The main direction can be found by multiplying the product of the normals by the direction. A smoothing / refinement process can be performed by looking at neighboring triangles. For example, a triangle that has more than a threshold number of neighbors that are all blue will be classified as blue even if there is an anomaly that would initially be shown as red.

[0033] Generate connected components of triangles to identify which triangles have the same color (eg, triangles of the same category that share at least one vertex).

[0034] Connectivity information describes how points are connected in 3D. These connections (more specifically, three different connections sharing three points) combine to generate triangles, which result in a face (represented by a group of triangles). Triangles are described herein, but other geometric shapes (e.g., rectangles) are also possible.

[0035] Color is used to encode connectivity by distinguishing between triangles of different colors: each triangle that is distinguished by three connectivity is coded with a unique color.

[0036] FIG. 4 shows a diagram of patch rasterization, one of the components of the patch generation process, according to some embodiments. Patch generation also includes generating connected components of triangles (triangles of the same category that share at least one vertex). If the bounding box of the connected components is smaller than a predefined area, the triangle is moved to a separate list for independent triangle coding. These unprojected triangles are not rasterized, but are coded as vertices with associated color vertices. Otherwise, each triangle is projected into a patch. If the projected position of the vertex is already occupied, the triangle is coded in another patch and goes into a missing triangles list to be processed again later. Alternatively, the map can be used to identify overlapping vertices and further represent triangles with overlapping vertices. The triangles are rasterized to generate points for point cloud representation.

[0037] Shown are the original voxelized vertices 400. The rasterized surface points 402 (which are added to the point cloud representation) follow the structure of the mesh, so the point cloud geometry can be as sparse as the underlying mesh. However, this geometry can be improved by sending additional positions for each rasterized pixel.

[0038] Figure 5 shows a diagram of V3C image generation according to some embodiments. Once a mesh is projected, an occupancy map, a geometry map and a texture map are generated. In V3C image generation, the occupancy map and the geometry map are generated traditionally. The attribute maps (textures) are generated from the uncompressed geometry.

[0039] Once the patches are created, information is added that indicates where the patches are located on the 2D image, as well as which are the vertex locations and how they are connected. These tasks are performed with the following syntax: TIFF0007672629000002.tif244165 If mpdu_binary_object_present_flag[tileID][p] is equal to 1, it specifies that the syntax elements mpdu_mesh_binary_object_size_bytes[tileID][p] and mpdu_mesh_binary_object[tileID][p][i] for the current patch p of the current atlas tile with tile ID equal to tileID are present. If mpdu_binary_object_present_flag[tileID][p] is equal to 0, the syntax elements mpdu_mesh_binary_object_size_bytes[tileID][p] and mpdu_mesh_binary_object[tileID][p][i] for the current patch are not present. If mpdu_binary_object_present_flag[tileID][p] is not present, its value is inferred to be equal to 0. mpdu_mesh_binary_object_size_bytes[tileID][p] specifies the number of bytes used to represent the mesh information in binary form. mpdu_mesh_binary_object[tileID][p][i] specifies the i byte of the binary representation of the mesh of the pth patch. mpdu_vertex_count_minus3[tileID][p]+3 specifies the number of vertices present in the patch. mpdu_face_count[tileID][p] specifies the number of triangles present in the patch. If there are no triangles present, mpdu_face_count[tileID][p] is set to 0. mpdu_face_vertex[tileID][p][i][k] specifies the kth value of the index of the vertex of the i-th triangle or quad of the current patch p of the current atlas tile whose tile ID is equal to tileID. The value of mpdu_face_vertex[tileID][p][i][k] is in the range of 0 to mpdu_vert_count_minus3[tileID][p]+2. mpdu_vertex_pos_x[tileID][p][i] specifies the value of the X coordinate of the i-th vertex of the current patch p in the current atlas tile with tile ID equal to tileID. The value of mpdu_vertex_pos_x[p][i] is within the range of 0 to mpdu_2d_size_x_minus1[p], inclusive. mpdu_vertex_pos_y[tileID][p][i] specifies the value of the y coordinate of the i-th vertex of the current patch p in the current atlas tile with tile ID equal to tileID. The value of mpdu_vertex_pos_y[tileID][p][i] is within the range of 0 to mpdu_2d_size_y_minus1[tileID][p], inclusive.

[0040] Some elements of the mesh patch data are controlled by parameters defined in the Atlas Sequence Parameter Set (ASPS). New extensions to the ASPS for meshes are available. TIFF0007672629000003.tif88150 When asps_mesh_extension_present_flag is equal to 1, it specifies the presence of an asps_mesh_extension() syntax within the atlas_sequence_parameter_set_rbsp syntax structure. asps_mesh_extension_present_flag equal to 0 specifies that this syntax construct is not present. If not present, the value of asps_mesh_extension_present_flag is inferred to be equal to 0. asps_extension_6bits equal to 0 specifies that the asps_extension_data_flag syntax element is not present in the ASPS RBSP syntax structure. When present, asps_extension_6bits is assumed to be equal to 0 in bitstreams conforming to this version of this document. A value of asps_extension_6bits equal to 0 is reserved for future use by ISO / IEC. Decoders shall allow values ​​of asps_extension_6bits not equal to 0 and shall ignore all asps_extension_data_flag syntax elements in ASPS NAL units. If not present, the value of asps_extension_6bits is inferred to be equal to 0. TIFF0007672629000004.tif36135 TIFF0007672629000005.tif26153 asps_mesh_binary_coding_enabled_flag, when equal to 1, indicates that the vertex and connectivity information associated with the patch is present in binary format. asps_mesh_binary_coding_enabled_flag equal to 0 indicates that mesh vertex and connectivity data is not present in binary format. If not present, asps_mesh_binary_coding_enabled_flag is inferred to be 0. asps_mesh_binary_codec_id indicates the identifier of the codec used to compress the vertex and connectivity information of the patch. asps_mesh_binary_codec_id shall be in the range 0 to 255 inclusive. This codec may be identified through a profile defined in Annex A, or through means outside this document. asps_mesh_quad_face_flag equal to 1 indicates that quads should be used for polygon representation. asps_mesh_quad_face_flag equal to 0 indicates that triangles should be used in the polygonal representation of the mesh. If not present, asps_mesh_quad_flag is inferred to be equal to 0. A value of asps_mesh_vertices_in_vertex_map_flag equal to 1 indicates the presence of vertex information in the vertex video data. A value of asps_mesh_vertices_in_vertex_flag equal to 0 indicates the presence of vertex information in the patch data. If not present, the value of asps_mesh_vertices_in_vertex_map_flag is inferred to be equal to 0.

[0041] The syntax allows four different types of vertex / connectivity encodings: 1. Send vertex and connectivity information directly in the patch. 2. Vertex information is sent in a vertex map and face connectivity is sent in a patch. 3. Send the vertex information in a vertex map and derive the face connectivity at the decoder side (eg, using ball pivoting or Poisson reconstruction). 4. Encode the vertex and connectivity information using an external mesh encoder (e.g., SC3DM, Draco). Unlike a normal 3D mesh, a 2D mesh is encoded.

[0042] The new V3C video data unit carries vertex position information. It can contain either a binary value indicating the projected vertex position or it can contain multi-level information to be used for connectivity reconstruction. It can enable level of detail reconstruction by coding the vertices in multiple layers. It can use VPS extensions to define additional parameters. TIFF0007672629000006.tif130154 TIFF0007672629000007.tif62159 vuh_lod_index indicates the lod index of the current vertex stream, if present. If not present, the lod index of the current vertex sub-bitstream is derived based on the type of the sub-bitstream and the operations described in subclause XX of the vertex video sub-bitstream, respectively. The value of vuh_lod_index, if present, shall be in the range 0 to vms_lod_count_minus1[vuh_atlas_id], inclusive. vuh_reserved_zero_13bits, if present, shall be equal to 0 in bitstreams conforming to this version of this document. Other values ​​of vuh_reserved_zero_13bits are reserved for future use by ISO / IEC. Decoders shall ignore values ​​of vuh_reserved_zero_13bits. TIFF0007672629000008.tif119152 vme_lod_count_minus1[k]+1 indicates the number of lods used to encode the vertex data of the atlas with atlas ID k. vme_lod_count_minus1[j] is in the range of 0 to 15, inclusive. vme_embed_vertex_in_occupancy_flag[k] equal to 1 specifies that the vertex information of the atlas with atlas ID k is derived from the occupancy map of XX term. vme_embed_vertex_in_occupancy_flag[k] equal to 0 indicates that vertex information is not derived from the occupancy video. vme_embed_vertex_in_occupancy_flag[k] is inferred to be equal to 0 if not present. vme_multiple_lod_streams_present_flag[k] equal to 0 indicates that all lods for the atlas with atlas ID k are placed in a single vertex video stream, vme_multiple_lod_streams_present_flag[k] equal to 1 indicates that all lods for the atlas with atlas ID k are placed in separate video streams, vme_multiple_lod_streams_present_flag[k] is inferred to have a value equal to 0 if not present. When vme_lod_absolute_coding_enabled_flag[j][i] is equal to 1, it indicates that lod with index i of atlas with atlas ID k is coded without any form of map prediction. vme_lod_absolute_coding_enabled_flag[k][i] equal to 0 indicates that lod with index i of atlas with atlas ID k is predicted first before encoding from another previously encoded map. vme_lod_absolute_coding_enabled_flag[j][i] is inferred to have a value equal to 1 if not present. vme_lod_predictor_index_diff[k][i] is used to calculate the predictor for lod with index i of atlas with atlas ID k when vps_map_absolute_coding_enabled_flag[j][i] is equal to 0. Specifically, LodPredictorIndex[i], the map predictor index for lod i, shall be calculated as follows: LodPredictorIndex[i]=(i-1)-vme_lod_predictor_index_diff[j][i] The value of vme_lod_predictor_index_diff[j][i] is in the range 0 to i-1, inclusive. If vme_lod_predictor_index_diff[j][i] is not present, its value is inferred to be equal to 0. When vme_vegrtex_video_present_flag[k] is equal to 0, it indicates that the atlas with ID k does not have vertex data. vms_vertex_video_present_flag[k] equal to 1 indicates that the atlas with ID k has vertex data. vms_vertex_video_present_flag[j] is inferred to be equal to 0 if not present. TIFF0007672629000009.tif31134 vi_vertex_codec_id[j] indicates the identifier of the codec used to compress the vertex information of the atlas with atlas ID j. vi_vertex_codec_id[j] shall be in the range 0 to 255 inclusive. This codec may be identified through a profile, a component codec mapping SEI message, or through means outside this document. vi_lossy_vertex_compression_threshold[j] indicates the threshold used to derive binary vertices from the decoded vertex video of the atlas with atlas ID j. vi_lossy_vertex_compression_threshold[j] is in the range of 0 to 255, inclusive. vi_vertex_2d_bit_depth_minus1[j]+1 represents the nominal 2D bit depth to which the vertex video of the atlas with atlas ID j should be converted. vi_vertex_2d_bit_depth_minus1[j] ranges from 0 to 31, inclusive. vi_vertex_MSB_align_flag[j] indicates how the decoded vertex video samples associated with the atlas with atlas ID j are converted to samples at the nominal vertex bit depth.

[0043] Figure 6 shows an image of vertex video data according to some embodiments. The vertex video data uses multiple layers to provide a level of detail useful for hierarchical mesh coding. If only one layer is used, the video data can be embedded in an occupancy map to avoid generating multiple decoded instances. A surface reconstruction algorithm (e.g., Poisson surface reconstruction and ball pivoting) can be used to generate connectivity information.

[0044] Vertex video data appears as dots / points in the image. These points indicate where the vertices are located in the projected image. Vertex video data can be transmitted directly within another video as shown in diagram 600 or embedded in the occupancy image 602. Diagram 604 shows an example of ball pivoting.

[0045] FIG. 7 shows images of an exemplary level of detail generation according to some embodiments. Vertices can be combined in layers to generate a multi-level representation of the mesh. Level of detail generation is specified. Image 700 shows an original point cloud with 762 vertices and 1,245 faces. Instead of sending all the data, only 10% of the vertices / connectivity is sent, which is 10% clustering decimation, sending 58 vertices and 82 faces, as shown in image 702. Image 704 shows 5% clustering decimation, sending 213 vertices and 512 faces. Image 706 shows 2.5% clustering decimation, sending 605 vertices and 978 faces. By separating the layers, multiple layers can be sent in stages (e.g., 1st layer 10% decimation, 2nd layer 5%, 3rd layer 2.5%) to improve quality. The layers can be combined to obtain the original (or near) vertex and face count of the mesh.

[0046] FIG. 8 illustrates a diagram of a mesh according to some embodiments. In some embodiments, SC3DM (MPEG) can be used to encode mesh information for each patch. SC3DM can be used to encode connectivity information and (u,v) information. In one embodiment, Draco is used to encode mesh information for each patch. Draco can be used to encode connectivity information and (u,v) information.

[0047] 9 is a diagram of mesh reconstruction according to some embodiments. Patches can be added together, although connectivity uses new vertex numbering. Vertices at seams may not match due to compression. To address this issue, mesh smoothing or a zippering algorithm can be used.

[0048] Figure 10 shows a diagram of high level syntax and images that can send a mixture of point cloud and mesh patches according to some embodiments. Because meshes are described at the patch level, it is possible to mix and match patches of an object. For example, one can use point cloud only patches for the head or hair, and mesh patches for flat areas such as the body.

[0049] The tracked mesh patch data unit can use patches to indicate that connectivity has not changed between patches. This is particularly useful for tracked meshes, since only delta positions are transmitted. Tracked meshes experience global motion that can be captured by bounding box position and rotation (a newly introduced syntax element using quaternions) and cosmetic motion that can be captured by vertex motion. Vertex motion can be transmitted explicitly in the patch information, or derived from video data if the reference patch uses V3C_VVD data. The number of bits to transmit delta vertex information can be transmitted in the Atlas Frame Parameter Set (AFPS). Alternatively, motion information can be transmitted as homography transforms. TIFF0007672629000010.tif140157 tmpdu_vertices_changed_position_flag specifies whether the vertices have changed position. tmpdu_vertex_delta_pos_x[p][i] specifies the difference between the x-coordinate value of the i-th vertex of patch p and the x-coordinate value of the matching patch indicated by tmpdu_ref_index[p]. The value of tmpdu_vertex_pos_x[p][i] is in the range of 0 to pow2(afps_num_bits_delta_x)-1, inclusive. tmpdu_vertex_delta_pos_y[p][i] specifies the difference between the y coordinate value of the i-th vertex of patch p and the y coordinate value of the matching patch indicated by tmpdu_ref_index[p]. The value of tmpdu_vertex_pos_x[p][i] is in the range of 0 to pow2(afps_num_bits_delta_y)-1, inclusive. tmpdu_rotation_present_flag specifies whether a rotation value is present. tmpdu_3d_rotation_qx specifies the x component, qX, of the geometry rotation of the current patch using quaternion representation. The value of tmpdu_3d_rotation_qx is -2 inclusive. 15 ~2 15 It is assumed to be in the range -1. If tmpdu_3d_rotation_qx is not present, its value is inferred to be equal to 0. The value of qX is calculated as follows: qX=tmpdu_3d_rotation_qx÷2 15 tmpdu_3d_rotation_qy specifies the y component, qY, of the geometric rotation of the current patch using quaternion representation. The value of tmpdu_3d_rotation_qy is -2 inclusive. 15 ~2 15 It is assumed to be in the range -1. If tmpdu_3d_rotation_qy is not present, its value is inferred to be equal to 0. The value of qY is calculated as follows: qY=tmpdu_3d_rotation_qy÷2 15 tmpdu_3d_rotation_qz specifies the z component, qZ, of the geometry rotation of the current patch using quaternion representation. The value of tmpdu_3d_rotation_qz is -2 inclusive. 15 ~2 15 It is assumed to be in the range -1. If tmpdu_3d_rotation_qz is not present, its value is inferred to be equal to 0. The value of qZ is calculated as follows: qZ=tmpdu_3d_rotation_qz÷2 15 The fourth component of the geometric rotation of the current point cloud image using the quaternion representation, qW, is calculated as follows: qW = Sqrt(1-(qX 2 +qY 2 +qZ 2 )) The unit quaternion can be expressed as a rotation matrix R as follows: TIFF0007672629000011.tif30157

[0050] 11A-11B show diagrams of combining untracked and tracked mesh information according to some embodiments. To avoid tracking issues, some algorithms segment the mesh into tracked and untracked parts. The tracked part is time-consistent and can be represented by the proposed tracked_mesh_patch_data_unit(), whereas the untracked part is new each frame and can be represented by mesh_patch_data_unit(). This notation also allows blending point clouds with geometry, improving surface representation (e.g., keeping the original mesh and inserting point clouds on the mesh to hide defects).

[0051] FIG. 12 shows an example image of patch-based edge collapse according to some embodiments. To reduce the number of coded triangles, an edge-clearing filter can be applied to the patch data. To improve rendering, geometry and texture information can be left intact. Mesh simplification can be reversed by using fine geometry data. Meshlab has an option to apply edge-clearing even when preserving boundaries. However, this algorithm works in 3D space and does not use the projection properties of the mesh. The new idea is to perform an "edge-clearing filter in 2D projected patch domain", i.e., to apply edge-clearing to the patch data while considering the 2D properties of the edges.

[0052] Image 1200 shows the original dense mesh with 5,685 vertices and 8,437 faces. Image 1202 shows a patch edge erase (preserving the boundary) with 2,987 vertices and 3,041 faces. Image 1204 shows a patch edge erase with 1,373 vertices and 1,285 faces. Image 1206 shows the complete mesh edge erase with 224 vertices and 333 faces.

[0053] Figure 13 shows an example image of patch-based clustering decimation according to some embodiments. Meshlab has the option to perform decimation (clustering decimation) based on a 3D grid. Since patches are projected data in 2D, decimation can be performed in 2D space instead. Furthermore, the number of decimated vertices can be kept to reconstruct the faces (using fine geometry data and face subdivision). This information can be transmitted in an occupancy map.

[0054] Image 1300 shows an original dense mesh with 5,685 vertices and 8,437 faces. Image 1302 shows a clustering decimation (1%) with 3,321 vertices and 4,538 faces. Image 1304 shows a clustering decimation (2.5%) with 730 vertices and 870 faces. Image 1306 shows a clustering decimation (5%) with 216 vertices and 228 faces. Image 1308 shows a clustering decimation (10%) with 90 vertices and 104 faces.

[0055] Figure 14 shows an example image of patch-based surface subdivision according to some embodiments. Meshlab has several filters to generate finer meshes, but these are somewhat heuristic (e.g., splitting triangles at their midpoints). Better results can be obtained if geometry information is used to guide where the triangles should be split. For example, a high-resolution mesh can be generated from a low-resolution mesh. The upsampling of the mesh information is guided by the geometry.

[0056] Image 1400 shows a mesh with 224 vertices and 333 faces. Image 1402 shows a subdivision surface: midpoint (1 iteration) with 784 vertices and 1332 faces. Image 1404 shows a subdivision surface: midpoint (2 iterations) with 2,892 vertices and 5,308 faces. Image 1406 shows a subdivision surface: midpoint (3 iterations) with 8,564 vertices and 16,300 faces.

[0057] Figure 15 shows an example image of patch-based surface reconstruction according to some embodiments. Meshlab has filters (screen Poisson and ball pivoting) that reconstruct mesh surfaces from point clouds. Using these algorithms, mesh connectivity can be reconstructed at the patch level at the decoder side (e.g., available vertex list is signaled via an occupancy map). In some embodiments, connectivity information is not transmitted and Poisson or ball pivoting can be used to regenerate connectivity information.

[0058] Image 1500 shows a dense vertex cloud with 5,685 vertices. Image 1502 shows a Poisson surface reconstruction with 13,104 vertices and 26,033 faces. Image 1504 shows a ball pivoting with 5,685 vertices and 10,459 faces.

[0059] In some embodiments, the vertex positions are obtained from an occupancy map. The color information embedded in the occupancy map can be used. At the encoder side, paint the face region associated with each triple vertex set with a fixed color. The paint color is unique per face and is chosen to facilitate color segmentation. An m-ary level occupancy map is used. At the decoder side, decode the occupancy map. Derive face information based on the segmented colors.

[0060] In some embodiments, the vertex positions are obtained from the occupancy. A new attribute is assigned that carries face information. At the encoder side, an attribute rectangle is generated with size (width x height) equal to the number of faces. This attribute has three dimensions, while each dimension carries the index of one vertex of the triple vertex. At the decoder side, the attribute video is decoded. The face information is derived from the decoded attribute video.

[0061] In some embodiments, the vertex positions are obtained from the occupancy map using Delaunay triangulation. At the decoder side, decode the occupancy map video. Triangulate the vertices obtained from the decoded occupancy map. Use the triangulated points to obtain face information.

[0062] Figure 16 shows a diagram of triangle edge detection according to some embodiments. An original image 1600 has blue, red and yellow triangles. A segmented blue triangle image 1602 shows the blue triangle, a segmented red triangle image 1604 shows the red triangle, and a segmented yellow triangle image 1606 shows the yellow triangle.

[0063] These triangles can be grouped based on color, and segmented to show where the triangles lie, where the triangle edges lie, and even where the vertices lie based on the edges that intersect.

[0064] 17 shows a diagram of segmented edges of triangles separated based on color, according to some embodiments. Original image 1600 has blue, red, and yellow triangles. Segmented blue edge image 1700 shows the edges of the blue triangles, segmented red edge image 1702 shows the edges of the red triangles, and segmented yellow edge image 1704 shows the edges of the yellow triangles.

[0065] 18 shows an illustration of segmented edges of joined triangles according to some embodiments. Original image 1600 has blue, red and yellow triangles. Triangle edges image 1800 shows the joined edges.

[0066] 19-23 show diagrams of edge resizing and rescaling according to some embodiments. Images, triangles and / or edges can be resized (e.g., reduced), after which the resized triangles can be detected, rescaled edges can be determined, and resized edges can be determined / generated.

[0067] As described herein, by performing segmentation, finding edges, triangles and vertices, and determining which locations are connected, color can be used to encode triangle locations and triangle connectivity.

[0068] FIG. 24 illustrates a block diagram of an exemplary computing device configured to implement a video-based mesh compression method, according to some embodiments. The computing device 2400 can be used for acquiring, storing, computing, processing, communicating, and / or displaying information, such as images and videos, including 3D content. The computing device 2400 can implement any of the encoding / decoding aspects. In general, a hardware structure suitable for implementing the computing device 2400 includes a network interface 2402, a memory 2404, a processor 2406, an I / O device 2408, a bus 2410, and a storage device 2412. The selection of the processor is not critical as long as a suitable processor of sufficient speed is selected. The memory 2404 can be any conventional computer memory known in the art. The storage device 2412 can include a hard drive, CDROM, CDRW, DVD, DVDRW, high definition disk / drive, ultra HD drive, flash memory card, or any other storage device. The computing device 2400 can include one or more network interfaces 2402. An example of a network interface includes a network card connected to an Ethernet or other type of LAN. The I / O device(s) 2408 may include one or more of a keyboard, mouse, monitor, screen, printer, modem, touch screen, button interface, and other devices. The storage device 2412 and memory 2404 likely store a video-based mesh compression application(s) 2430 used to execute the implementation of the video-based mesh compression and process it as an application would normally be processed. The computing device 2400 may also include more or fewer components than those shown in FIG. 24. In some embodiments, dense mesh compression hardware 2420 is included.Although the computing device 2400 of Figure 24 includes an application 2430 and hardware 2420 for implementing video-based mesh compression, the video-based mesh compression method may also be implemented on the computing device in hardware, firmware, software, or any combination thereof. For example, in some embodiments, the dense mesh compression application 2430 is programmed into memory and executed using a processor. As another example, in some embodiments, the video-based mesh compression hardware 2420 is programmed hardware logic that includes gates specifically designed to implement the video-based mesh compression method.

[0069] In some embodiments, the video-based mesh compression application(s) 2430 include multiple applications and / or modules. In some embodiments, a module also includes one or more sub-modules. In some embodiments, fewer or additional modules may be included.

[0070] Examples of suitable computing devices include a personal computer, a laptop computer, a computer workstation, a server, a mainframe computer, a handheld computer, a personal digital assistant, a cellular / mobile phone, a smart appliance, a game console, a digital camera, a digital camcorder, a camera phone, a smart phone, a portable music player, a tablet computer, a mobile device, a video player, a video disc writer / player (such as a DVD writer / player, a high definition disc writer / player, an ultra high definition disc writer / player, etc.), a television, a home entertainment system, an augmented reality device, a virtual reality device, smart jewelry (e.g., a smart watch), a vehicle (e.g., a self-driving vehicle), or any other suitable computing device.

[0071] 25 illustrates a diagram of a system configured to implement video-based mesh compression according to some embodiments. The encoder 2500 is configured to perform an encoding process. Any encoding, such as video-based mesh compression, may be performed as described herein. The mesh and other information may be communicated to the decoder 2504 directly or via a network 2502. The network may be any type of network, such as a local area network (LAN), the Internet, a wireless network, a wired network, a cellular network, and / or any other network or combination of networks. The decoder 2504 decodes the encoded content.

[0072] To utilize the video-based mesh compression method, a device acquires or receives 3D content (e.g., point cloud content). The video-based mesh compression method can be performed with user assistance or automatically without user involvement.

[0073] In operation, the video-based mesh compression method enables more efficient and accurate 3D content encoding than previous implementations.

[0074] Some embodiments of video-based mesh compression 1. A method comprising: performing mesh voxelization on an input mesh; performing patch generation to segment the mesh into patches comprising a rasterized mesh surface and vertex position and connectivity information; generating a visual volumetric video-based compression (V3C) image from the rasterized mesh surface; performing video-based mesh compression using the vertex position and connectivity information; and generating a V3C bitstream based on the V3C image and the video-based mesh compression.

[0075] 2. The method of claim 1, wherein the vertex position and connectivity information includes triangulation information of surface patches.

[0076] 3. The method of claim 1, wherein data resulting from performing video-based mesh compression using the vertex position and connectivity information is encapsulated in a vertex video component structure.

[0077] 4. The method of claim 3, wherein the vertex video component structure enables hierarchical mesh encoding by separating sets of vertices into layers to generate levels of detail of mesh connectivity.

[0078] 5. The method of clause 1, where if only one layer is implemented, the video data is embedded in the occupancy map.

[0079] 6. The method of claim 1, wherein the connectivity information is generated using a surface reconstruction algorithm including Poisson surface reconstruction or ball pivoting.

[0080] 7. The method of claim 1, wherein generating the V3C image from the rasterized mesh surface includes combining untracked mesh information and tracked mesh information.

[0081] 8. The method of claim 1, further comprising performing an edge removal filter on the two-dimensional projected patch region.

[0082] 9. The method of claim 1, further comprising performing a patch-based surface subdivision of the connectivity information.

[0083] 10. An apparatus comprising: a non-transitory memory storing an application for performing mesh voxelization on an input mesh, performing patch generation to segment the mesh into patches including a rasterized mesh surface and vertex position and connectivity information, generating a visual volumetric video-based compressed (V3C) image from the rasterized mesh surface, performing video-based mesh compression using the vertex position and connectivity information, and generating a V3C bitstream based on the V3C image and the video-based mesh compression; and a processor coupled to the memory and configured to process the application.

[0084] 11. The apparatus of claim 10, wherein the vertex position and connectivity information includes triangle information for a surface patch.

[0085] 12. The apparatus of clause 10, wherein data resulting from performing video-based mesh compression using the vertex position and connectivity information is encapsulated in a vertex video component structure.

[0086] 13. The apparatus of clause 12, wherein the vertex video component structure enables hierarchical mesh coding by separating sets of vertices into layers to generate levels of detail of mesh connectivity.

[0087] 14. The apparatus of clause 10, wherein if only one layer is implemented, the video data is embedded in the occupancy map.

[0088] 15. The apparatus of claim 10, wherein the connectivity information is generated using a surface reconstruction algorithm including Poisson surface reconstruction or ball pivoting.

[0089] 16. The apparatus of clause 10, wherein generating the V3C image from the rasterized mesh surface includes combining untracked mesh information and tracked mesh information.

[0090] 17. The apparatus of claim 10, wherein the application is further configured to perform an edge removal filter on the two-dimensional projected patch region.

[0091] 18. The apparatus of clause 10, wherein the application is further configured to perform patch-based surface subdivision of the connectivity information.

[0092] 19. A system comprising one or more cameras for acquiring three-dimensional content, and an encoder for encoding the three-dimensional content by performing mesh voxelization on an input mesh, performing patch generation to segment the mesh into patches comprising a rasterized mesh surface and vertex positions and connectivity information, generating a visual volumetric video-based compressed (V3C) image from the rasterized mesh surface, performing video-based mesh compression using the vertex positions and connectivity information, and generating a V3C bitstream based on the V3C image and the video-based mesh compression.

[0093] 20. The system of clause 19, wherein the vertex position and connectivity information includes triangulation information for surface patches.

[0094] 21. The system of clause 19, wherein data resulting from performing video-based mesh compression using the vertex position and connectivity information is encapsulated in a vertex video component structure.

[0095] 22. The system of clause 21, wherein the vertex video component structure enables hierarchical mesh encoding by separating sets of vertices into layers to generate levels of detail of mesh connectivity.

[0096] 23. The system of claim 19, wherein if only one layer is implemented, the video data is embedded in the occupancy map.

[0097] 24. The system of claim 19, wherein the connectivity information is generated using a surface reconstruction algorithm including Poisson surface reconstruction or ball pivoting.

[0098] 25. The system of clause 19, wherein generating the V3C image from the rasterized mesh surface includes combining untracked mesh information and tracked mesh information.

[0099] 26. The system of claim 19, wherein the encoder is configured to perform an edge removal filter on the two-dimensional projected patch region.

[0100] 27. The system of clause 19, wherein the encoder is configured to perform patch-based surface subdivision of the connectivity information.

[0101] The present invention has been described in terms of specific embodiments containing details to facilitate an understanding of the principles of construction and operation of the invention. Reference herein to such specific embodiments and details of these embodiments is not intended to limit the scope of the claims appended hereto. It will be readily apparent to those skilled in the art that various other modifications can be made in the embodiments selected for illustration without departing from the spirit and scope of the invention as defined by the claims. [Explanation of symbols]

[0102] 100 input meshes 102 Mesh Voxelization 104 Patch Generation 106 V3C Image Generation 108 Base Mesh Coding (Uncoded) 110 V3C bitstream 112 Point cloud reconstruction 114 Mesh Reconstruction

Claims

1. performing mesh voxelization on the input mesh; performing patch generation to segment the mesh into patches containing mesh surfaces and vertex positions and connectivity information, where points on the surface are sampled to generate a shape image and surface attributes; generating a visual volumetric video-based compressed (V3C) image from the generated mesh surface; performing video-based mesh compression using the vertex position and connectivity information; and generating a V3C bitstream based on the V3C image and the video-based mesh compression; Including, data resulting from performing video based mesh compression using the vertex position and connectivity information is encapsulated into a vertex video component structure configured to enable hierarchical mesh encoding that encodes mesh information by separating the set of vertices into layers to generate a mesh connectivity level of detail including a number of vertices and a number of faces of the mesh; generating the V3C image from the generated mesh surface includes combining mesh information in which the movements of the vertices of the patches are not tracked with mesh information in which the movements of the vertices of the patches are tracked; A method comprising:

2. the vertex position and connectivity information includes triangulation information of surface patches; The method of claim 1.

3. If only one layer of the vertex set is implemented in the vertex video component structure, video data resulting from performing video-based mesh compression using the vertex position and connectivity information is embedded in an occupancy map. The method of claim 1.

4. The connectivity information is generated using a surface reconstruction algorithm including Poisson surface reconstruction or ball pivoting. The method of claim 1.

5. applying edge erasure to the patch data while taking into account two-dimensional characteristics of the edges; The method of claim 1.

6. performing a subdivision of a mesh surface on a patch-by-patch basis using the connectivity information. The method of claim 1.

7. Perform mesh voxelization on the input mesh, performing patch generation to segment the mesh into patches containing mesh surfaces and vertex positions and connectivity information, where points on the surface are sampled to generate a shape image and surface attributes; generating a visual volumetric video-based compressed (V3C) image from the generated mesh surface; performing video-based mesh compression using the vertex positions and connectivity information; generating a V3C bitstream based on the V3C image and the video-based mesh compression; a non-transitory memory for storing an application for a processor coupled to the memory and configured to process the application; Equipped with data resulting from performing video based mesh compression using the vertex position and connectivity information is encapsulated into a vertex video component structure configured to enable hierarchical mesh encoding that encodes mesh information by separating the set of vertices into layers to generate a mesh connectivity level of detail that includes a vertex count and a face count for the mesh; generating the V3C image from the generated mesh surface includes combining mesh information in which the movements of the vertices of the patches are not tracked with mesh information in which the movements of the vertices of the patches are tracked; An apparatus comprising:

8. the vertex position and connectivity information includes triangulation information of surface patches; 8. The apparatus of claim 7.

9. If only one layer of the vertex set is implemented in the vertex video component structure, video data resulting from performing video-based mesh compression using the vertex position and connectivity information is embedded in an occupancy map.

8. The apparatus of claim 7.

10. The connectivity information is generated using a surface reconstruction algorithm including Poisson surface reconstruction or ball pivoting.

8. The apparatus of claim 7.

11. The application is further configured to apply edge erasure to the patch data while taking into account two-dimensional characteristics of edges.

8. The apparatus of claim 7.

12. the application is further configured to perform subdivision of a mesh surface on a patch-by-patch basis using the connectivity information.

8. The apparatus of claim 7.

13. one or more cameras for acquiring three-dimensional content; Perform mesh voxelization on the input mesh, performing patch generation to segment the mesh into patches containing mesh surfaces and vertex positions and connectivity information, where points on the surface are sampled to generate a shape image and surface attributes; generating a visual volumetric video-based compressed (V3C) image from the generated mesh surface; performing video-based mesh compression using the vertex positions and connectivity information; generating a V3C bitstream based on the V3C image and the video-based mesh compression; an encoder for encoding the three-dimensional content by Equipped with data resulting from performing video based mesh compression using the vertex position and connectivity information is encapsulated into a vertex video component structure configured to enable hierarchical mesh encoding that encodes mesh information by separating the set of vertices into layers to generate a mesh connectivity level of detail that includes a vertex count and a face count for the mesh; generating the V3C image from the generated mesh surface includes combining mesh information in which the movements of the vertices of the patches are not tracked with mesh information in which the movements of the vertices of the patches are tracked; A system characterized by:

14. the vertex position and connectivity information includes triangulation information of surface patches; The system of claim 13.

15. If only one layer of the vertex set is implemented in the vertex video component structure, video data resulting from performing video-based mesh compression using the vertex position and connectivity information is embedded in an occupancy map. The system of claim 13.

16. The connectivity information is generated using a surface reconstruction algorithm including Poisson surface reconstruction or ball pivoting. The system of claim 13.

17. the encoder is configured to apply edge erasure to the patch data while taking into account two-dimensional characteristics of edges; The system of claim 13.

18. the encoder is configured to perform mesh surface subdivision on a patch-by-patch basis using the connectivity information; The system of claim 13.

Citation Information

Patent Citations

  • Projection-based mesh compression

    US11373339B2

  • Robust mesh tracking and fusion by using part-based key frames and priori model

    US20190026942A1