Mesh Compression via Point Cloud Representation

By segmenting the 3D grid into surface sheets that follow connectivity and performing 2D projection sampling, the grid data is encoded using a projection-based method, which solves the problem of difficulty in effectively compressing 3D grid data in the prior art, and achieves efficient grid data compression and coding effects.

CN113939849BActive Publication Date: 2025-06-24SONY GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080042641.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-27
Filing Date
2020-12-03
Publication Date
2025-06-24
Estimated Expiration
2040-12-03

AI Technical Summary

Technical Problem

The prior art is difficult to effectively compress 3D grid data, especially in dynamic 3D scenarios, and it is impossible to effectively encode the connectivity of the grid.

Method used

Using a projection-based approach, the grid is segmented into surface patches that follow connectivity, the grid surface is sampled by a rasterization method, projected onto the 2D patch, and encoded using the same method as point cloud compression.

Benefits of technology

It realizes efficient compression and encoding of 3D mesh data, can maintain good quality in dynamic 3D scenarios, and improve rendering and point filtering algorithms through additional connectivity data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113939849B_ABST
    Figure CN113939849B_ABST
Patent Text Reader

Abstract

This document describes a method for compressing a mesh using a projection-based approach and leveraging tools and syntax already generated for projection-based point cloud compression. Similar to the V-PCC method, the mesh is segmented into surface patches, with the only difference being that these patches follow the connectivity of the mesh. Each surface patch (or 3D patch) is then projected onto a 2D patch, such that in the case of a mesh, the triangle surface sampling is similar to common rasterization methods used in computer graphics. For each patch, the positions of the projected vertices and the connectivity of these vertices are saved together in a list. The sampled surface now resembles a point cloud, and the sampled surface is encoded using the same methods used for point cloud compression. Additionally, the list of vertices and connectivity is encoded per patch, and this data is sent along with the encoded point cloud data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority under 35 U.S.C.§119(e) to U.S. Provisional Patent Application Serial No. 62 / 946,194, filed on December 10, 2019, entitled "Mesh Compression via Point Cloud Representation", the entire content of which is incorporated herein by reference for all purposes. Technical Field

[0003] The present invention relates to three - dimensional graphics. More specifically, the present invention relates to the encoding of three - dimensional graphics. Background Art

[0004] In recent years, a new method for compressing point clouds based on 3D - to - 2D projection has been under standardization. This method is also known as V - PCC (Video - based Point Cloud Compression), which maps 3D point cloud data into a number of 2D patches, then further arranges these patches into an atlas image, and then encodes the atlas image with a video encoder. The atlas image corresponds to the geometry of the points, their respective textures, and a placeholder map that indicates which positions are to be considered for point cloud reconstruction.

[0005] In 2017, MPEG issued a Call for Proposals (CfP) for point cloud compression. After the evaluation of several proposals, MPEG is currently considering two different point cloud compression techniques: 3D native encoding techniques (based on octrees and similar encoding methods), or 3D - to - 2D projection followed by traditional video encoding. In the case of dynamic 3D scenarios, MPEG is using Test Model Software 2 (TMC2) based on patch - surface modeling, projecting patches from 3D to 2D images, and encoding the 2D images with a video encoder (such as HEVC). This method has proven to be more efficient than native 3D encoding and can achieve competitive bitrates with acceptable quality.

[0006] Due to the success of the projection - based method (also known as the video - based method, or V - PCC) for encoding 3D point clouds, the standard is expected to include further 3D data (such as 3D meshes) in future versions. However, the current version of the standard is only suitable for the transmission of unconnected sets of points and thus has no mechanism for sending the connectivity of the points, although it is required in 3D mesh compression.

[0007] Methods for extending the functionality of V-PCC to meshes have also been proposed. One possible method is to use V-PCC to encode the vertices and then use a mesh compression method (such as TFAN or Edgebreaker) to encode the connectivity. The limitation of this method is that the original mesh must be dense, so the point cloud generated from the vertices is not sparse and cannot be efficiently encoded after projection. In addition, the order of the vertices affects the encoding of the connectivity, and different methods for reorganizing the mesh connectivity have been proposed. Another method for encoding sparse meshes is to use RAW patch data to encode the vertex positions in 3D. Since RAW patches directly encode (x, y, z), in this method, all vertices are encoded as RAW data, and the connectivity is encoded by a similar mesh compression method as described above. In RAW patches, the vertices can be sent in any preferred order, so the order generated from the connectivity encoding can be used. This method can encode sparse point clouds, but RAW patches are not efficient for encoding 3D data, and further data (such as the attributes of the triangular faces) may be lost from this method. Summary of the Invention

[0008] This document describes a method for compressing meshes using a projection-based approach and leveraging the tools and syntax already generated for projection-based point cloud compression. Similar to the V-PCC method, the mesh is segmented into surface patches, with the only difference being that these patches follow the connectivity of the mesh. Each surface patch (or 3D patch) is then projected onto a 2D patch, by virtue of which, in the case of a mesh, the triangular surface sampling is similar to the common rasterization methods used in computer graphics. For each patch, the positions of the projected vertices and their connectivity are saved in a list. The sampled surface now resembles a point cloud and is encoded using the same methods used for point cloud compression. In addition, the list of vertices and connectivity is encoded per patch, and this data is sent together with the encoded point cloud data.

[0009] Additional connectivity data can be interpreted as the underlying mesh generated for each patch, giving the decoder the flexibility to use or not use this additional data. This data can be used to improve rendering and point filtering algorithms. In addition, encoding the mesh using the same principles as projection-based compression results in better integration with the current V-PCC method for projection-based point cloud encoding.

[0010] In one aspect, a method of programming in a non-transitory memory of a device includes: performing mesh voxelization on an input mesh; implementing patch generation that divides the mesh into patches, the patches including a rasterized mesh surface and vertex positions and connectivity information; generating a video-based point cloud compression (V-PCC) image from the rasterized mesh surface; implementing base mesh encoding using the vertex positions and connectivity information; and generating a V-PCC bitstream based on the V-PCC image and the base mesh encoding. Mesh voxelization includes moving and / or scaling mesh values to avoid negative and non-integer values. Mesh voxelization includes finding the lowest vertex value below zero and moving the mesh values to make the lowest vertex value above zero. Implementing patch generation includes calculating the normal of each triangle. Calculating the normal of a triangle includes using the cross product between edges. The method further includes classifying the triangles based on the normal. The method further includes implementing a refinement process by analyzing adjacent triangles. Base mesh encoding includes encoding the (u, v) coordinates of the vertices. Generating the V-PCC bitstream includes base mesh signaling and using multiple layers. The first layer in the multiple layers implementation includes the original point cloud, the second layer in the multiple layers implementation includes a sparse mesh, and the third layer in the multiple layers implementation includes a dense mesh. The method further includes generating a base mesh for each patch that includes additional connectivity data, where the decoder determines whether to utilize the additional connectivity data, and where the additional connectivity data improves rendering and point filtering. Encoding the connectivity information is based on a color code. Generating the V-PCC bitstream based on the V-PCC image and the base mesh encoding utilizes the connectivity information of each patch.

[0011] In another aspect, a device includes a non-transitory memory for storing an application that is configured to: perform mesh voxelization on an input mesh; implement patch generation that divides the mesh into patches, the patches including a rasterized mesh surface and vertex positions and connectivity information; generate a video-based point cloud compression (V-PCC) image from the rasterized mesh surface; implement base mesh encoding using the vertex positions and connectivity information; and generate a V-PCC bitstream based on the V-PCC image and the base mesh encoding; and a processor coupled to the memory, the processor being configured to process the application. Mesh voxelization includes shifting and / or scaling mesh values to avoid negative and non-integer values. Mesh voxelization includes finding the lowest vertex value below zero and shifting the mesh values so that the lowest vertex value is above zero. Implementing patch generation includes calculating the normal of each triangle. Calculating the normal of a triangle includes using the cross product between edges. The application is also configured to classify triangles based on the normals. The application is also configured to implement a refinement process by analyzing adjacent triangles. Base mesh encoding includes encoding the (u, v) coordinates of vertices. Generating the V-PCC bitstream includes base mesh signaling and is implemented using multiple layers. The first layer in the multiple layer implementation includes the original point cloud, the second layer in the multiple layer implementation includes a sparse mesh, and the third layer in the multiple layer implementation includes a dense mesh. The application is also configured to generate a base mesh for each patch that includes additional connectivity data, where a decoder determines whether to utilize the additional connectivity data, and where the additional connectivity data improves rendering and point filtering. The connectivity information is encoded based on a color code. Generating the V-PCC bitstream based on the V-PCC image and the base mesh encoding utilizes the connectivity information of each patch.

[0012] On the other hand, a system includes one or more cameras for acquiring three-dimensional content and an encoder for encoding the three-dimensional content by performing the following operations: performing mesh voxelization on an input mesh of the three-dimensional content; implementing patch generation, which divides the mesh into patches, the patches including a rasterized mesh surface and vertex positions and connectivity information; generating a video-based point cloud compression (V-PCC) image from the rasterized mesh surface; implementing base mesh encoding using the vertex positions and connectivity information and generating a V-PCC bitstream based on the V-PCC image and the base mesh encoding. Mesh voxelization includes moving and / or scaling mesh values to avoid negative and non-integer values. Mesh voxelization includes finding the lowest vertex value below zero and moving the mesh values so that the lowest vertex value is above zero. Implementing patch generation includes calculating the normal of each triangle. Calculating the normal of a triangle includes using the cross product between edges. The encoder is also used to classify triangles according to the normals. The encoder is also used to implement a refinement process by analyzing adjacent triangles. Base mesh encoding includes encoding the (u, v) coordinates of vertices. Generating the V-PCC bitstream includes base mesh signaling and implementation using multiple layers. The first layer in the multiple-layer implementation includes the original point cloud, the second layer in the multiple-layer implementation includes a sparse mesh, and the third layer in the multiple-layer implementation includes a dense mesh. The encoder is also configured to generate a base mesh for each patch including additional connectivity data, where the decoder determines whether to utilize the additional connectivity data, and where the additional connectivity data improves rendering and point filtering. The connectivity information is encoded based on a color code. Generating the V-PCC bitstream based on the V-PCC image and the base mesh encoding utilizes the connectivity information of each patch. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 Shows a mesh compression method according to some embodiments.

[0014] Figure 2 Shows mesh voxelization according to some embodiments.

[0015] Figure 3 Shows an image related to patch generation according to some embodiments.

[0016] Figure 4 Shows a schematic diagram of a projected triangle for point cloud representation according to some embodiments.

[0017] Figure 5 Shows an exemplary image of vertices and triangles according to some embodiments.

[0018] Figure 6 Shows an example of encoding connectivity by indicating triangle connectivity using the color channels of a geometry image according to some embodiments.

[0019] Figure 7 Shows a network abstraction layer (NAL) unit and a multi - layer implementation for base mesh signaling according to some embodiments.

[0020] Figure 8 Shows a multi - layer implementation for base mesh signaling according to some embodiments.

[0021] Figure 9 Shows a schematic diagram of geometry refinement according to some embodiments.

[0022] Figure 10 Shows a flowchart of point cloud rendering using a mesh compression method according to some embodiments.

[0023] Figure 11 Shows a block diagram of an exemplary computing device configured to implement a mesh compression method according to some embodiments. Detailed Description

[0024] The method of compressing 3D mesh data using a point cloud representation of a mesh surface is described above. Embodiments utilize 3D surface patches to represent the point cloud and perform a temporally consistent global mapping of the 3D patch surface data to a 2D canvas image.

[0025] In 3D point cloud encoding using a video encoder, the projection from 3D to 2D is important for generating videos that will represent the point cloud. The most efficient way to generate those videos is to use 3D patches, which segment the surface of the object, and use an orthographic projection to generate a segmented depth image, which is bundled together and used as the input to the video encoder. In current point cloud standards, 3D meshes cannot be encoded because there is no defined method to encode the connectivity of the mesh. Additionally, if the vertex data is sparse, the standard performs poorly because it cannot utilize the correlations between vertices.

[0026] Methods for performing mesh encoding using video - based standards for point cloud compression are described herein. Methods for segmenting mesh surfaces and joint surface sampling and 2D patch generation are disclosed. The disclosed methods also describe each patch that is encoded for local connectivity and project the positions of the vertices onto the 2D patch. Methods for signaling connectivity and vertex positions and enabling the reconstruction of the original input mesh are also described.

[0027] Embodiments can be applied to dense time - varying meshes with mesh attributes such as texture.

[0028] Figure 1Shows a mesh compression method according to some embodiments. In step 100, mesh voxelization is performed on the input mesh. Mesh voxelization involves converting the floating-point values of the point positions of the input mesh into integers. The precision of the integers can be set by the user or automatically set. In some embodiments, mesh voxelization includes shifting the values so that there are no negative numbers. In step 102, patch creation / generation for splitting the mesh into patches is implemented. Patch generation also generates 1) the rasterized mesh surface and 2) vertex position and connectivity information. In step 104, the rasterized mesh surface is a set of points that are passed through V-PCC image generation and encoded as a V-PCC image. In step 106, vertex position and connectivity information is received for encoding the base mesh. In step 108, a V-PCC bitstream is generated based on the V-PCC image generation and the base mesh encoding. In some embodiments, fewer or additional steps are implemented. In some embodiments, the order of the steps is modified.

[0029] Mesh voxelization

[0030] Figure 2 Shows mesh voxelization according to some embodiments. As shown in FIG. 200, the original mesh is below the axis, resulting in negative numbers. Through mesh voxelization, the mesh is moved and / or scaled to avoid negative and non-integer values. In one implementation, the lowest vertex value below zero is found, and then these values can be moved so that the lowest vertex value is above zero. In some embodiments, the range of values is adapted to a specified bit range such as 11 bits (e.g., by scaling).

[0031] Image 202 shows no perceivable difference between the original mesh and the voxelized mesh.

[0032] Patch generation

[0033] The patch generation described herein is similar to the patch generation in V-PCC. However, instead of calculating the normal for each point, the normal for each triangle is calculated. The cross product between the edges is used to calculate the normal for each triangle to determine the normal vector. Then, the triangles are classified according to the normal. For example, the normals are divided into n (e.g., 6) categories (such as front, back, up, down, left, and right). Different colors are used to represent the normals to show the initial segmentation. Figure 3 Image 300 shows different colors (such as black and light gray) in grayscale because different colors represent different normals. Although it may be difficult to see in Image 300, the top surface (e.g., the top of a person's head, the top of a ball, and the top of a sports shoe) is one color (e.g., green), the first side of the person / ball is dark, representing another color (e.g., red), the bottom of the ball is another color (e.g., purple), and the front sides of the person and the ball, mostly light gray, represent another color (e.g., cyan).

[0034] The principal direction can be found by multiplying the product of the normals by the direction. The smoothing / refinement process can be achieved by looking at adjacent triangles. For example, if all adjacent triangles with a quantity higher than a threshold are blue, then this triangle is also classified as blue, even if there are anomalies initially indicating that the triangle is red. For example, a red triangle represented by reference numeral 302 can be corrected to cyan as shown by reference numeral 304.

[0035] Image 310 shows an example of a triangle with a normal vector.

[0036] Connected components of the triangles are generated to identify which triangles have the same color (e.g., triangles with the same class share at least one vertex).

[0037] Connectivity information describes how points are connected in 3D. These connections together generate triangles (more specifically, 3 different connections sharing 3 points), which thus generate a surface (described by a set of triangles). Although triangles are described herein, other geometries (e.g., rectangles) are also allowed.

[0038] By identifying triangles with different colors, the color can be used to encode the connectivity. Each triangle identified by three connections is encoded with a unique color.

[0039] By projecting the mesh onto a 2D surface, the area covered by the triangle projection is also determined by a set of pixels. If the grouped pixels are encoded with different colors, the triangles can be identified by the different colors in the image. Once the triangles are known, the connectivity can be obtained by simply identifying the three connections that form the triangle.

[0040] Each triangle is projected onto a patch. If the projected position of a vertex is already occupied, the triangle is encoded in another patch and thus goes into a list of missing triangles to be processed again later. Alternatively, a mapping graph can be used to identify overlapping vertices and still be able to represent triangles with overlapping vertices. In another alternative, points can be split into separate layers (e.g., a first set of points in one layer and a second set of points in a second layer).

[0041] The triangles are rasterized to generate points for the point cloud representation.

[0042] Figure 4A schematic diagram of a projected triangle for point cloud representation according to some embodiments is shown. Triangle 400 has been projected onto a grid 402 (e.g., a 2D projection of the triangle). Each square in grid 402 is a point in the point cloud. Some points are the original points of the vertices. As shown, when these points are projected, they are voxelized and projected into these positions. Points 404 in the 2D projection mark the vertices on the original mesh. For the area of triangle 400, points are generated by rasterizing the surface. The grid elements within the triangle become points in the point cloud, which generates the point cloud from the mesh (e.g., rasterization is performed on the projection plane).

[0043] The points added to the point cloud follow the structure of the mesh, so the point cloud geometry can be as rough as the underlying mesh. However, the geometry can be improved by sending additional positions for each rasterized pixel.

[0044] Base mesh encoding

[0045] The list of points in a face is the vertices of the triangle, and even after projection, the connectivity of the mesh is the same. Figure 5 An exemplary image of vertices and triangles according to some embodiments is shown. The vertices are black points, and the connectivity is the lines connecting the black points.

[0046] Encode the connectivity (e.g., based on a color code). In some embodiments, encode a list of integer values. Differential Pulse Code Modulation (DPCM) can be used in the list. In some embodiments, the list can be refined or intelligent mesh encoding can be implemented. In some embodiments, more complex methods are also possible (e.g., using Edgebreaker or TFAN, both are encoding algorithms).

[0047] In some embodiments, encode the (u, v) coordinates of the vertices instead of the (x, y, z). The (u, v) coordinates are positions on a 2D grid (e.g., the positions where the vertices are projected). From the projection in the geometry image, the (x, y, z) information can be determined. The DPCM method is also possible. In some embodiments, the (u, v) coordinates are stored in a list. The order can be determined by the connectivity. Based on the connectivity, it is known that certain vertices are connected, so the (u, v) values of the connected vertices should be similar, which enables predictions such as parallelogram prediction (e.g., Draco, a mesh compression algorithm).

[0048] Figure 6 An example of encoding connectivity by using the color channels of a geometry image to indicate triangle connectivity according to some embodiments is shown. For example, if certain triangles are the same, their color is yellow, and the color of different triangles can be blue, etc., so the color identifies the triangles and connectivity of the mesh.

[0049] Base mesh signaling

[0050] Send additional information per patch. In each patch information, a list of connected components (e.g., vertices) and the positions of the vertices in 2D space is sent. A more efficient representation can use a DPCM scheme for faces and vertices, as discussed herein.

[0051] Figure 7 Shows a network abstraction layer (NAL) unit and a multi-layer implementation for base mesh signaling according to some embodiments. The NAL 700 includes information such as a header, group layer, number of faces, number of vertices, number of faces, and vertex positions.

[0052] In some embodiments, a multi-layer implementation in the NAL is used to send additional layers containing connectivity information. The V-PCC unit stream 702 utilized in the multi-layer implementation is shown. The first layer (e.g., layer 0) defines the point cloud, and layer 1 defines the mesh layer. In some embodiments, these layers are related to each other. In some embodiments, additional layers are utilized.

[0053] Figure 8 Shows a multi-layer implementation for base mesh signaling according to some embodiments. In a hierarchical representation, layer_id can be used to send meshes with different resolutions. For example, layer 0 is the original point cloud, layer 1 is a sparse mesh, and layer 2 is a dense mesh. Additional layers (e.g., layer 3 is a very dense mesh) can be implemented. In some embodiments, the order of the layers is different, e.g., layer 0 is a dense mesh, layer 1 is a sparse mesh, and layer 2 is the original point cloud. In some embodiments, the additional layer only provides the difference or δ with the previous layer. For example, as Figure 8 shown, there are 3 triangles in layer 1, and layer 2 has 6 triangles, where the large triangle is divided into 4 triangles, and the division of the large triangle (e.g., the 4 triangles) is included in layer 2.

[0054] The syntax of the patch data unit can be modified to include:

[0055]

[0056]

[0057] In some embodiments, alternative encodings are implemented, such as using TFAN or Edgebreaker to encode patch connectivity, using parallelogram prediction for vertices, and / or using DPCM encoding.

[0058] Figure 9Shows a schematic diagram of geometry refinement according to some embodiments. By transmitting δ information from the base mesh surface to the actual position of the point cloud, a more accurate position of the points can be improved. Because when the mesh surface is rasterized, the generated point cloud will have a geometry similar to the mesh surface, which may be rough. The δ information can be obtained by sending δ from the mesh surface, and the normal direction of the mesh can also be considered.

[0059] As described herein, since triangles are considered flat, each triangle can send more information.

[0060] Rendering optimization and geometry filtering can also be achieved. Since the base mesh represents the surface, all the points contained in the boundaries of the triangles are logically connected. When re-projecting points, holes may appear due to geometry differences and different baseline distances. However, the renderer can use the underlying mesh information to improve the re-projection. Because it knows from the mesh that the points should be logically connected on the surface, the renderer can generate interpolated points and close the holes even without sending any additional information.

[0061] For example, in some instances, due to holes in the point cloud during projection, but since it is known that the surface is represented by triangles, all points should be filled on the surface. Therefore, even if these points are not explicitly encoded from the mesh representation, geometry filtering can be used to fill in the missing points.

[0062] As described herein, the mesh compression method uses a projection-based method, and tools and syntax already generated for projection-based point cloud compression are described herein. Similar to the V-PCC method, the mesh is divided into surface patches, and the only difference is that these patches follow the connectivity of the mesh. Then each surface patch (or 3D patch) is projected onto a 2D patch. Thus, in the case of a mesh, triangle surface sampling is similar to common rasterization methods used in computer graphics. For each patch, the positions of the projected vertices and the connectivity of these vertices are saved in a list. The sampled surface now resembles a point cloud, and the same methods used for point cloud compression are used to encode the sampled surface. In addition, the list of vertices and connectivity is encoded per patch, and this data is sent together with the encoded point cloud data.

[0063] Additional connectivity data can be interpreted as the base mesh generated for each patch, giving the decoder the flexibility to use or not use this additional data. This data can be used to improve rendering and point filtering algorithms. In addition, encoding the mesh using the same principle as projection-based compression results in better integration with the current projection-based V-PCC method of point cloud encoding.

[0064] Figure 10Shows a flowchart of point cloud rendering using a mesh compression method according to some embodiments. In step 1000, a mesh is encoded with V-PCC and / or a coded mesh is received (e.g., at a device). In step 1002, the coded mesh enters a V-PCC decoder for decoding, resulting in a point cloud 1004 and a mesh 1006. In step 1008, point cloud filtering is applied to the point cloud 1004 and the mesh 1006. In step 1010, the filtered point cloud and the mesh 1006 are used in point cloud rendering. In some embodiments, fewer or additional steps are implemented. In some embodiments, the order of the steps is modified.

[0065] Figure 11 Shows a block diagram of an exemplary computing device configured to implement a mesh compression method according to some embodiments. The computing device 1100 can be used to acquire, store, calculate, process, communicate, and / or display information, (such as images and videos including 3D content). The computing device 1100 can implement any aspect of mesh compression. Generally, the hardware structure suitable for implementing the computing device 1100 includes a network interface 1102, a memory 1104, a processor 1106, an I / O device 1108, a bus 1110, and a storage device 1112. The choice of the processor is not critical as long as a suitable processor with sufficient speed is selected. The memory 1104 can be any conventional computer memory known in the art. The storage device 1112 can include a hard disk drive, a CDROM, a CDRW, a DVD, a DVDRW, a high-definition disc / drive, an ultra-high-definition drive, a flash card, or any other storage device. The computing device 1100 can include one or more network interfaces 1102. Examples of network interfaces include network cards connected to Ethernet or other types of LANs. One or more I / O devices 1108 can include one or more of the following: a keyboard, a mouse, a monitor, a screen, a printer, a modem, a touch screen, a button interface, and other devices. One or more mesh compression applications 1130 for implementing the mesh compression method may be stored in the storage device 1112 and the memory 1104 and are processed in the manner in which applications are typically processed. Figure 11 More or fewer components as shown in can be included in the computing device 1100. In some embodiments, mesh compression hardware 1120 is included. Although Figure 11 the computing device 1100 in includes an application 1130 and hardware 1120 for the mesh compression method, the mesh compression method can be implemented on the computing device in hardware, firmware, software, or any combination thereof. For example, in some embodiments, the mesh compression application 1130 is programmed in the memory and executed using the processor. In another example, in some embodiments, the mesh compression hardware 1120 is programmed as hardware logic that includes gates specifically designed to implement the mesh compression method.

[0066] In some embodiments, one or more mesh compression applications 1130 include a number of applications and / or modules. In some embodiments, a module further includes one or more sub-modules. In some embodiments, fewer or additional modules may be included.

[0067] Examples of suitable computing devices include personal computers, laptop computers, computer workstations, servers, mainframe computers, handheld computers, personal digital assistants, cellular / mobile phones, smart devices, gaming consoles, digital cameras, digital video cameras, camera phones, smartphones, portable music players, tablet computers, mobile devices, video players, video disc burners / players (e.g., DVD burners / players, high-definition disc burners / players, ultra-high-definition disc burners / players), televisions, home entertainment systems, augmented reality devices, virtual reality devices, smart jewelry (e.g., smartwatches), transportation vehicles (e.g., autonomous transportation vehicles), or any other suitable computing device.

[0068] To utilize the mesh compression method, a device acquires or receives 3D content and processes and / or transmits the content in an optimized manner such that correct and efficient display of the 3D content can be achieved. The mesh compression method can be implemented with the help of a user or automatically without user participation.

[0069] In operation, the mesh compression method can achieve more efficient and accurate mesh compression compared to previous implementations.

[0070] In an exemplary implementation, the mesh compression described herein is implemented on top of TMC2v8.0, with only one frame and a single map. Information from the implementation includes:

[0071] Bitstream statistics:

[0072] Header: 16B 128b

[0073] vpcc unit size [VPCC_VPS]: 31B 248b

[0074] vpcc unit size [VPCC_AD]: 451967B 3615736b

[0075] vpcc unit size [VPCC_OVD]: 25655B 205240b (Ocm video = 25647B)

[0076] vpcc unit size [VPCC_GVD]: 64342B 514736b (Geo video = 64334B + 0B + 0B + 0B)

[0077] VPCC unit size [VPCC_AVD]: 72816B 582528b (Tex video = 72808B + 0B)

[0078] Total metadata: 477685B 3821480b

[0079] Total geometry: 64334B 514672b

[0080] Total texture: 72808B 582464b

[0081] Total: 614827B 4918616b

[0082] Total bitstream size 614843B

[0083] Some embodiments of mesh compression via point cloud representation

[0084] 1. A method programmed in a non - transitory memory of a device, comprising:

[0085] Performing mesh voxelization on an input mesh;

[0086] Implementing patch generation, where patch generation divides the mesh into patches, and the patches include a rasterized mesh surface and vertex position and connectivity information;

[0087] Generating a video - based point cloud compression (V - PCC) image from the rasterized mesh surface; implementing base mesh encoding with vertex position and connectivity information; and

[0088] Generating a V - PCC bitstream based on the V - PCC image and the base mesh encoding.

[0089] 2. The method according to clause 1, wherein mesh voxelization includes moving and / or scaling mesh values to avoid negative and non - integer values.

[0090] 3. The method according to clause 2, wherein mesh voxelization includes finding the lowest vertex value below zero and moving the mesh values so that the lowest vertex value is above zero.

[0091] 4. The method according to clause 1, wherein implementing patch generation includes calculating the normal of each triangle.

[0092] 5. The method according to clause 4, wherein calculating the normal of a triangle includes using the cross product between edges.

[0093] 6. The method according to clause 4, further comprising classifying the triangles according to the normal.

[0094] 7. The method according to clause 4 further includes implementing a refinement process by analyzing adjacent triangles.

[0095] 8. The method according to clause 1, wherein the base mesh encoding includes encoding the (u, v) coordinates of vertices.

[0096] 9. The method according to clause 1, wherein generating a V-PCC bitstream includes base mesh signaling and is implemented using multiple layers.

[0097] 10. The method according to clause 9, wherein the first layer in the multiple-layer implementation includes the original point cloud, the second layer in the multiple-layer implementation includes a sparse mesh, and the third layer in the multiple-layer implementation includes a dense mesh.

[0098] 11. The method according to clause 1 further includes generating a base mesh for each patch that includes additional connectivity data, wherein the decoder determines whether to utilize the additional connectivity data, and wherein the additional connectivity data improves rendering and point filtering.

[0099] 12. The method according to clause 1, wherein the connectivity information is encoded based on a color code.

[0100] 13. The method according to clause 1, wherein generating a V-PCC bitstream based on a V-PCC image and base mesh encoding utilizes the connectivity information of each patch.

[0101] 14. An apparatus, comprising:

[0102] A non-transitory memory for storing an application program, the application program being configured to:

[0103] Perform mesh voxelization on an input mesh;

[0104] Implement patch generation, which divides the mesh into patches, the patches including a rasterized mesh surface and vertex positions and connectivity information;

[0105] Generate a video-based point cloud compression (V-PCC) image from the rasterized mesh surface;

[0106] Implement base mesh encoding using the vertex positions and connectivity information; and

[0107] Generate a V-PCC bitstream based on the V-PCC image and base mesh encoding; and

[0108] A processor coupled to the memory, the processor being configured to process the application program.

[0109] 15. The apparatus according to clause 14, wherein the mesh voxelization includes moving and / or scaling mesh values to avoid negative and non-integer values.

[0110] 16. The apparatus according to clause 15, wherein the mesh voxelization includes finding the lowest vertex value below zero and shifting the mesh values to make the lowest vertex value higher than zero.

[0111] 17. The apparatus according to clause 14, wherein implementing patch generation includes calculating the normal of each triangle.

[0112] 18. The apparatus according to clause 17, wherein calculating the normal of a triangle includes using the cross product between edges.

[0113] 19. The apparatus according to clause 17, wherein the application is further configured to classify triangles according to the normal.

[0114] 20. The apparatus according to clause 17, wherein the application is further configured to implement a refinement process by analyzing adjacent triangles.

[0115] 21. The apparatus according to clause 14, wherein the base mesh encoding includes encoding the (u, v) coordinates of vertices.

[0116] 22. The apparatus according to clause 14, wherein generating a V-PCC bitstream includes base mesh signaling and is implemented using multiple layers.

[0117] 23. The apparatus according to clause 22, wherein the first layer in the multiple-layer implementation includes the original point cloud, the second layer in the multiple-layer implementation includes a sparse mesh, and the third layer in the multiple-layer implementation includes a dense mesh.

[0118] 24. The apparatus according to clause 14, wherein the application is further configured to generate a base mesh for each patch including additional connectivity data, wherein the decoder determines whether to utilize the additional connectivity data, and wherein the additional connectivity data improves rendering and point filtering.

[0119] 25. The apparatus according to clause 14, wherein the connectivity information is encoded based on a color code.

[0120] 26. The apparatus according to clause 14, wherein generating a V-PCC bitstream based on the V-PCC image and the base mesh encoding utilizes the connectivity information of each patch.

[0121] 27. A system, comprising:

[0122] One or more cameras for acquiring three-dimensional content; and

[0123] An encoder for encoding the three-dimensional content by:

[0124] Performing mesh voxelization on the input mesh of the three-dimensional content;

[0125] Implement patch generation, where patch generation divides the mesh into patches, and the patches include a rasterized mesh surface and vertex position and connectivity information;

[0126] Generate a video-based point cloud compression (V-PCC) image from the rasterized mesh surface;

[0127] Implement base mesh encoding using vertex position and connectivity information; and

[0128] Generate a V-PCC bitstream based on the V-PCC image and the base mesh encoding.

[0129] 28. The system according to clause 27, wherein mesh voxelization includes moving and / or scaling mesh values to avoid negative and non-integer values.

[0130] 29. The system according to clause 28, wherein mesh voxelization includes finding the lowest vertex value below zero and moving the mesh values so that the lowest vertex value is above zero.

[0131] 30. The system according to clause 27, wherein implementing patch generation includes calculating the normal of each triangle.

[0132] 31. The system according to clause 30, wherein calculating the normal of a triangle includes using the cross product between edges.

[0133] 32. The system according to clause 30, wherein the encoder is further configured to classify triangles according to the normal.

[0134] 33. The system according to clause 30, wherein the encoder is further configured to implement a refinement process by analyzing adjacent triangles.

[0135] 34. The system according to clause 27, wherein base mesh encoding includes encoding the (u, v) coordinates of vertices.

[0136] 35. The system according to clause 27, wherein generating the V-PCC bitstream includes base mesh signaling and is implemented using multiple layers.

[0137] 36. The system according to clause 35, wherein the first layer in the multiple-layer implementation includes the original point cloud, the second layer in the multiple-layer implementation includes a sparse mesh, and the third layer in the multiple-layer implementation includes a dense mesh.

[0138] 37. The system according to clause 27, wherein the encoder is further configured to generate a base mesh for each patch including additional connectivity data, wherein the decoder determines whether to utilize the additional connectivity data, and wherein the additional connectivity data improves rendering and point filtering.

[0139] 38. The system according to clause 27, wherein the connectivity information is encoded based on a color code.

[0140] 39. The system according to clause 27, wherein generating a V-PCC bitstream based on the V-PCC image and the base mesh encoding utilizes the connectivity information of each patch.

[0141] The present invention has been described in terms of specific embodiments incorporating details to facilitate understanding of the principles of construction and operation of the present invention. This reference to specific embodiments and their details is not intended to limit the scope of the appended claims. It will be apparent to those skilled in the art that various other modifications may be made in the embodiments chosen for illustration without departing from the spirit and scope of the invention as defined by the claims.

Claims

1. A method of programming in a non-transitory memory of a device, comprising: Performing mesh voxelization on an input mesh to generate a voxelized mesh; Implementing 3D patch generation, which divides the voxelized mesh into 3D patches, the 3D patches including a rasterized mesh surface and vertex position and connectivity information, wherein implementing 3D patch generation includes calculating the normal of each triangle, and calculating the normal of a triangle includes using the cross product between edges; Classifying triangles according to the normals; Implementing a refinement process by analyzing adjacent triangles; After performing 2D projection on the triangles, generating a video-based point cloud compression V-PCC image from the rasterized mesh surface; Implementing base mesh encoding with vertex position and connectivity information; and Generating a V-PCC bitstream based on the V-PCC image and the base mesh encoding.

2. The method according to claim 1, wherein mesh voxelization includes moving and / or scaling mesh values to avoid negative and non-integer values.

3. The method according to claim 2, wherein mesh voxelization includes finding the lowest vertex value below zero and moving the mesh values so that the lowest vertex value is above zero.

4. The method according to claim 1, wherein base mesh encoding includes encoding the (u, v) coordinates of vertices.

5. The method according to claim 1, wherein generating a V-PCC bitstream includes base mesh signaling and is implemented using multiple layers.

6. The method according to claim 5, wherein the first layer in the multiple-layer implementation includes the original point cloud, the second layer in the multiple-layer implementation includes a sparse mesh, and the third layer in the multiple-layer implementation includes a dense mesh.

7. The method according to claim 1, further comprising generating a base mesh for each patch including additional connectivity data, wherein the decoder determines whether to utilize the additional connectivity data, and wherein the additional connectivity data improves rendering and point filtering.

8. The method according to claim 1, wherein the connectivity information is encoded based on a color code.

9. The method according to claim 1, wherein when generating a V-PCC bitstream based on the V-PCC image and the base mesh encoding, the connectivity information of each patch is utilized.

10. An apparatus, comprising: A non-transitory memory for storing an application program, the application program for: Performing mesh voxelization on an input mesh to generate a voxelized mesh; Implementing 3D patch generation, which divides the voxelized mesh into 3D patches, the 3D patches including a rasterized mesh surface and vertex position and connectivity information, wherein implementing 3D patch generation includes calculating the normal of each triangle, and calculating the normal of a triangle includes using the cross product between edges; Classifying triangles according to the normals; Implementing a refinement process by analyzing adjacent triangles; After performing 2D projection on the triangles, generating a video-based point cloud compression V-PCC image from the rasterized mesh surface; Implementing base mesh encoding using vertex position and connectivity information; and Generate a V-PCC bitstream based on a V-PCC image and base mesh encoding; and a processor coupled to a memory, the processor being configured to process the application program.

11. The apparatus according to claim 10, wherein mesh voxelization includes moving and / or scaling mesh values to avoid negative and non-integer values.

12. The apparatus according to claim 11, wherein mesh voxelization includes finding the lowest vertex value below zero and moving the mesh values so that the lowest vertex value is above zero.

13. The apparatus according to claim 10, wherein base mesh encoding includes encoding the (u, v) coordinates of vertices.

14. The apparatus according to claim 10, wherein generating the V-PCC bitstream includes base mesh signaling and is implemented using multiple layers.

15. The apparatus according to claim 14, wherein the first layer in the multi-layer implementation includes the original point cloud, the second layer in the multi-layer implementation includes a sparse mesh, and the third layer in the multi-layer implementation includes a dense mesh.

16. The apparatus according to claim 10, wherein the application program is further configured to generate a base mesh including additional connectivity data for each patch, wherein the decoder determines whether to utilize the additional connectivity data, and wherein the additional connectivity data improves rendering and point filtering.

17. The apparatus according to claim 10, wherein the connectivity information is encoded based on a color code.

18. The apparatus according to claim 10, wherein the connectivity information of each patch is utilized when generating the V-PCC bitstream based on the V-PCC image and base mesh encoding.

19. A system, comprising: One or more cameras for acquiring three-dimensional content; And An encoder for encoding the three-dimensional content by: Performing mesh voxelization on an input mesh of the three-dimensional content to generate a voxelized mesh; Implement 3D patch generation, where 3D patch generation divides the voxelized mesh into 3D patches, the 3D patches including a rasterized mesh surface and vertex positions and connectivity information, wherein the implementing 3D patch generation includes calculating the normal of each triangle, and wherein calculating the normal of a triangle includes using the cross product between edges; Classifying the triangles according to the normals; Implementing a refinement process by analyzing adjacent triangles; After performing a 2D projection on the triangles, generating a video-based point cloud compression V-PCC image from the rasterized mesh surface; implementing base mesh encoding using the vertex positions and connectivity information; and Generating a V-PCC bitstream based on the V-PCC image and base mesh encoding.

20. The system according to claim 19, wherein mesh voxelization includes moving and / or scaling mesh values to avoid negative and non-integer values.

21. The system according to claim 20, wherein mesh voxelization includes finding the lowest vertex value below zero and moving the mesh values so that the lowest vertex value is above zero.

22. The system according to claim 19, wherein base mesh encoding includes encoding the (u, v) coordinates of vertices.

23. The system according to claim 19, wherein generating the V-PCC bitstream includes base mesh signaling and is implemented using multiple layers.

24. The system according to claim 23, wherein the first layer in the multiple-layer implementation includes the original point cloud, the second layer in the multiple-layer implementation includes a sparse mesh, and the third layer in the multiple-layer implementation includes a dense mesh.

25. The system according to claim 19, wherein the encoder is further configured to generate a base mesh for each patch that includes additional connectivity data, wherein the decoder determines whether to utilize the additional connectivity data, and wherein the additional connectivity data improves rendering and point filtering.

26. The system according to claim 19, wherein the connectivity information is encoded based on a color code.

27. The system according to claim 19, wherein the connectivity information of each patch is utilized when generating the V-PCC bitstream based on the V-PCC image and base mesh encoding.

Citation Information

Patent Citations

  • Scalable compression of time-consistend 3D mesh sequences

    US20120262444A1