Three-dimensional grid coding and decoding method and device
By dividing 3D meshes into intra-slice and inter-slice components for separate encoding, the method enhances compression efficiency and reduces data volume, addressing the complexity and precision challenges in 3D mesh models.
Patent Information
- Application Number
- CN202410058463.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, the data volume of the three-dimensional grid model is large, resulting in complex processing, visualization and storage, and lack of efficient and general encoding and decoding methods.
The three-dimensional grid is divided into the basic grid intra-frame film and the basic grid inter-frame film, and the intra-code and inter-frame coding modes are used respectively. The basic grid intra-frame film and inter-frame film are merged and reconstructed to generate displacement code streams, and encoding them with texture maps and auxiliary information.
It improves coding efficiency, ensures the standard consistency and integrity of the reconstruction grid, and is suitable for a variety of application scenarios.
Smart Images

Figure CN120321412A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of three-dimensional mesh encoding and decoding, and particularly relates to a three-dimensional mesh encoding and decoding method and its device. Background Art
[0002] In recent years, with the rapid development of multimedia technology, related research results have been rapidly industrialized and have become an indispensable and important part of people's lives. Three-dimensional models have become a new generation of digital media following audio, images, and videos. Three-dimensional meshes are a commonly used representation of three-dimensional models. Compared with traditional multimedia such as images and videos, three-dimensional mesh models have stronger interactivity and realism, making them more and more widely used in various fields such as commerce, manufacturing, construction, education, medicine, entertainment, art, and military.
[0003] In related technologies, with the increasing demand for the visual effects of three-dimensional mesh models, the models have become more and more complex and the accuracy of the models has also become higher, resulting in a corresponding increase in the amount of data required to represent the three-dimensional mesh; the above problems have made the processing, visualization, transmission, and storage of three-dimensional meshes become more and more complex.
[0004] Therefore, it is urgent to improve the encoding and decoding methods in the prior art and provide an efficient and general three-dimensional mesh compression algorithm. Summary of the Invention
[0005] In order to solve the above problems existing in the prior art, the present invention provides a three-dimensional mesh encoding and decoding method and its device. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0006] In the first aspect, the present invention provides a three-dimensional mesh encoding method, including:
[0007] Dividing a three-dimensional mesh into a base mesh intra-slice and a base mesh inter-slice, and performing encoding to obtain a base mesh intra-slice bitstream and a base mesh inter-slice bitstream, and at the same time obtaining a reconstructed base mesh intra-slice and a reconstructed base mesh inter-slice;
[0008] Merging the reconstructed base mesh intra-slice and the reconstructed base mesh inter-slice to obtain a reconstructed base mesh;
[0009] Performing a subdivision process on the reconstructed base mesh to obtain the displacement of the subdivided mesh, and encoding the displacement of the subdivided mesh to obtain a displacement bitstream.
[0010] In the second aspect, the present invention also provides a three-dimensional mesh decoding method, including:
[0011] Decode the in-frame slice bitstream of the base mesh frame to obtain the in-frame slice of the base mesh frame, decode the inter-frame slice of the base mesh frame to obtain the inter-frame slice of the base mesh frame, decode the in-frame slice of the base mesh frame to obtain the reconstructed in-frame slice of the base mesh frame, and decode the inter-frame slice of the base mesh frame to obtain the reconstructed inter-frame slice of the base mesh frame;
[0012] Merge the reconstructed in-frame slice of the base mesh frame and the reconstructed inter-frame slice of the base mesh frame to obtain the reconstructed base mesh;
[0013] Decode the displacement bitstream to obtain the displacement of the subdivided mesh after subdividing the reconstructed base mesh. Based on the reconstructed base mesh and the displacement of the subdivided mesh, obtain the reconstructed mesh.
[0014] In a third aspect, the present invention further provides a three-dimensional mesh encoding device, which is applied to the encoding end and includes:
[0015] A first encoding module, configured to divide a three-dimensional mesh into an in-frame slice of the base mesh frame and an inter-frame slice of the base mesh frame, and perform encoding to obtain an in-frame slice bitstream of the base mesh frame and an inter-frame slice bitstream of the base mesh frame, and at the same time obtain a reconstructed in-frame slice of the base mesh frame and a reconstructed inter-frame slice of the base mesh frame;
[0016] A processing module, configured to merge the reconstructed in-frame slice of the base mesh frame and the reconstructed inter-frame slice of the base mesh frame to obtain the reconstructed base mesh;
[0017] A second encoding module, configured to perform a subdivision process on the reconstructed base mesh to obtain the displacement of the subdivided mesh, and encode the displacement of the subdivided mesh to obtain a displacement bitstream.
[0018] In a fourth aspect, the present invention further provides a three-dimensional mesh decoding device, which is applied to the decoding end and includes:
[0019] A first decoding module, configured to decode the in-frame slice bitstream of the base mesh frame to obtain the in-frame slice of the base mesh frame, decode the inter-frame slice of the base mesh frame to obtain the inter-frame slice of the base mesh frame, decode the in-frame slice of the base mesh frame to obtain the reconstructed in-frame slice of the base mesh frame, and decode the inter-frame slice of the base mesh frame to obtain the reconstructed inter-frame slice of the base mesh frame;
[0020] A processing module, configured to merge the reconstructed in-frame slice of the base mesh frame and the reconstructed inter-frame slice of the base mesh frame to obtain the reconstructed base mesh;
[0021] A second decoding module, configured to decode the displacement bitstream to obtain the displacement of the subdivided mesh after subdividing the reconstructed base mesh. Based on the reconstructed base mesh and the displacement of the subdivided mesh, obtain the reconstructed mesh.
[0022] Advantages of the present invention:
[0023] A 3D mesh encoding and decoding method and its device provided by the present invention. Compared with the existing solution that uses an inter-frame encoding mode when the 3D mesh is similar to the reference mesh and an intra-frame encoding mode otherwise, that is, when there are some regions where the 3D mesh is not similar to the reference mesh, the inter-frame encoding mode cannot be utilized. In the present invention, the 3D mesh is divided into a base mesh intra-frame slice and a base mesh inter-frame slice, where the base mesh inter-frame slice adopts the inter-frame encoding mode and the base mesh intra-frame slice adopts the intra-frame encoding mode, which can greatly improve the encoding efficiency. In the present invention, the reconstructed base mesh intra-frame slice and the reconstructed base mesh inter-frame slice are merged to obtain the reconstructed base mesh, and displacements are generated based on the reconstructed base mesh. This framework can be consistent with the existing intra-frame encoding scheme and is also the form that should be in the standard. In addition, the reconstructed base mesh is obtained by merging the reconstructed base mesh intra-frame slice and the reconstructed base mesh inter-frame slice. This reconstructed base mesh is an entirety and has good standard consistency as the reference mesh for subsequent meshes.
[0024] The following will further elaborate on the present invention in conjunction with the accompanying drawings and embodiments. Description of the Drawings
[0025] Figure 1 is a schematic diagram of a 3D mesh encoding method provided by an embodiment of the present invention;
[0026] Figure 2 is a flowchart of generating a base mesh provided by an embodiment of the present invention;
[0027] Figure 3 is a schematic diagram of generating a base mesh provided by an embodiment of the present invention;
[0028] Figure 4 is a schematic diagram of generating a registered base mesh provided by an embodiment of the present invention;
[0029] Figure 5 is a schematic diagram of detecting mismatched regions provided by an embodiment of the present invention;
[0030] Figure 6 is a schematic diagram of boundary vertex simplification and adjustment provided by an embodiment of the present invention;
[0031] Figure 7 is a schematic diagram of mesh simplification provided by an embodiment of the present invention;
[0032] Figure 8 is a schematic diagram of base mesh encoding provided by an embodiment of the present invention;
[0033] Figure 9 is a schematic diagram of the base mesh inter-frame slice bitstream and the base mesh inter-frame slice bitstream structure provided by an embodiment of the present invention;
[0034] Figure 10It is a schematic diagram of a bitstream splicing method provided by an embodiment of the present invention;
[0035] Figure 11 It is another schematic diagram of a bitstream splicing method provided by an embodiment of the present invention;
[0036] Figure 12 It is a schematic diagram of one of the five modes of Edgebreaker provided by an embodiment of the present invention;
[0037] Figure 13 It is a schematic diagram of a texture coordinate parameterization provided by an embodiment of the present invention;
[0038] Figure 14 It is a schematic diagram of a geometric displacement vector calculation method provided by an embodiment of the present invention;
[0039] Figure 15 It is a schematic diagram of a subdivision provided by an embodiment of the present invention;
[0040] Figure 16 It is a schematic diagram of a displacement encoding provided by an embodiment of the present invention;
[0041] Figure 17 It is a schematic diagram of a deformed mesh reconstruction provided by an embodiment of the present invention;
[0042] Figure 18 It is a schematic diagram of a texture map conversion provided by an embodiment of the present invention;
[0043] Figure 19 It is a schematic diagram of a three-dimensional mesh decoding method provided by an embodiment of the present invention;
[0044] Figure 20 It is a schematic diagram of a displacement decoding provided by an embodiment of the present invention. Detailed implementation manners
[0045] The following further describes the present invention in detail with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0046] In the prior art, although there are currently many representation methods for three-dimensional meshes, the triangular mesh is still the most common representation method at present. A three-dimensional mesh can be regarded as composed of three basic elements: vertices, edges, and faces. A vertex is the most basic element in the mesh, which defines a position in a three-dimensional space. An edge is a line segment connecting two vertices in the mesh. A face can be regarded as a polygon formed by a closed path of edges. For a triangular mesh, each face is a triangle.
[0047] The information contained in a mesh is generally divided into three categories: geometric information, connectivity information, and attribute information. Geometric information refers to the positions of each vertex of the mesh in three-dimensional space. Connectivity information describes the association relationships between the elements in the mesh, that is, the connection relationships between vertices. Attribute information is optional, and it can associate attributes to the corresponding mesh elements (such as vertex colors, normal vectors, etc. can be associated with mesh vertices). The mesh can also be parameterized to map from three-dimensional space to a two-dimensional planar region. This mapping relationship is usually described by a set of parametric coordinates, called UV coordinates or texture coordinates, which are associated with mesh vertices. This two-dimensional mapping can be used to represent high-resolution attribute information, such as textures, normal vectors, etc.
[0048] In almost all application fields that use three-dimensional meshes (such as computational simulation, entertainment, medical imaging, digital cultural relics, computer design, e-commerce, etc.), with the increasing demand for better visual effects of three-dimensional mesh models, the models are becoming more and more complex and the accuracy of the models is also getting higher. Therefore, the amount of data required to represent three-dimensional meshes is correspondingly increasing. The above problems have led to the processing, visualization, transmission, and storage of three-dimensional meshes becoming increasingly complex. Three-dimensional mesh compression can be regarded as a way to solve the above problems. It reduces the size of model data and is beneficial to the processing, storage, and transmission of three-dimensional meshes. Therefore, it is necessary to propose an efficient and general three-dimensional mesh compression algorithm.
[0049] Recently, the international standardization organization MPEG in the field of audio and video coding compression has started to develop a compression standard for three-dimensional meshes, VDMC (Video-based dynamic mesh coding, video-based dynamic mesh compression). This standard is specified based on the existing V3C (Visual Volumetric Video-based Coding, visual volumetric content coding based on video) standard. The V3C standard provides a general method for compressing three-dimensional models, and the models can be presented in the form of point clouds, meshes, or panoramic videos, etc. Making the compression method of three-dimensional mesh models compatible with this standard helps the promotion and applicability of this method.
[0050] In view of this, the present invention provides a three-dimensional mesh coding method, which optimizes the three-dimensional mesh encoding and decoding method in VDMC. It is of great significance to combine the optimization method with the V3C standard; a possible optimization method is to optimize the inter-frame coding of the mesh, that is, to divide a base mesh or a sub-mesh of the base mesh into two slices, an intra-frame slice and an inter-frame slice, for separate encoding, so as to provide an efficient and general three-dimensional mesh compression algorithm.
[0051] Please refer to Figure 1 , Figure 1It is a schematic diagram of a three-dimensional mesh encoding method provided by an embodiment of the present invention. A three-dimensional mesh encoding method provided by the present invention includes:
[0052] Divide the three-dimensional mesh into base mesh intra-slice and base mesh inter-slice, and perform encoding to obtain the base mesh intra-slice bitstream and the base mesh inter-slice bitstream, and at the same time obtain the reconstructed base mesh intra-slice and the reconstructed base mesh inter-slice;
[0053] Merge the reconstructed base mesh intra-slice and the reconstructed base mesh inter-slice to obtain the reconstructed base mesh;
[0054] Perform subdivision processing on the reconstructed base mesh to obtain the displacement of the subdivided mesh, and encode the displacement of the subdivided mesh to obtain the displacement bitstream.
[0055] Specifically, please continue to refer to Figure 1 , a three-dimensional mesh encoding method provided by this embodiment, respectively obtains a base mesh bitstream, a displacement bitstream, a texture map bitstream, and an auxiliary base mesh information bitstream according to a reference base mesh, a current input mesh, an input texture map, and additional base mesh information, and mixes these bitstreams to form a mixed stream to implement the entire process of three-dimensional mesh encoding.
[0056] In this embodiment, according to the current input mesh and the reference base mesh, a base mesh is generated. The base mesh includes a base mesh intra-slice and a base mesh inter-slice, that is, the three-dimensional mesh is divided into a base mesh intra-slice and a base mesh inter-slice, and the base mesh intra-slice and the base mesh inter-slice are encoded to obtain the base mesh intra-slice bitstream and the base mesh inter-slice bitstream. The base mesh intra-slice bitstream and the base mesh inter-slice bitstream are merged to form a base mesh bitstream; at the same time, the reconstructed base mesh intra-slice and the reconstructed base mesh inter-slice are obtained, and the reconstructed base mesh intra-slice and the reconstructed base mesh inter-slice are merged to obtain the reconstructed base mesh; perform subdivision processing on the reconstructed base mesh to obtain the displacement of the subdivided mesh, encode the displacement of the subdivided mesh to obtain the displacement bitstream, and at the same time obtain the reconstructed displacement; according to the reconstructed displacement and the reconstructed base mesh, generate a reconstructed deformed mesh; according to the reconstructed deformed mesh, perform texture map conversion using the current input mesh and the input texture map to obtain a texture map, and encode the texture map to obtain a texture map bitstream; in addition, it is also necessary to encode the auxiliary base mesh information, and the auxiliary base mesh information is used to guide the decoding process at the decoding end; finally, the base mesh bitstream, the displacement bitstream, the texture map bitstream, and the auxiliary base mesh information bitstream are mixed to obtain the required bitstream.
[0057] Among them, the in-frame slice of the base mesh refers to the mesh area with a low similarity to the reference base mesh, and its connection relationship, vertex geometric coordinates, texture coordinates, etc. are directly encoded; the inter-frame slice of the base mesh refers to the mesh area with a high similarity to the reference base mesh, and its connection relationship, vertex geometric coordinates, texture coordinates, etc. are all encoded based on the reference base mesh using time-domain prediction technology, greatly improving the compression efficiency. Optionally, the reference base mesh is the base mesh corresponding to the base mesh reconstructed in the time domain; it can be understood that a mesh can include multiple in-frame slices of the base mesh and multiple inter-frame slices of the base mesh. The multiple in-frame slices of the base mesh and the multiple inter-frame slices of the base mesh are connected together, and joint processing such as texture coordinate parameterization, displacement generation, and displacement encoding can be performed, and it can be regarded as a whole reference base mesh for the subsequent time-domain mesh.
[0058] In this embodiment, compared with the existing scheme that uses the inter-frame coding mode when the three-dimensional mesh is similar to the reference mesh and the intra-frame coding mode otherwise, that is, when there are some areas of the three-dimensional mesh that are not similar to the reference mesh, the inter-frame coding mode cannot be used. In this embodiment, the three-dimensional mesh is divided into in-frame slices of the base mesh and inter-frame slices of the base mesh. Among them, the inter-frame slices of the base mesh adopt the inter-frame coding mode, and the in-frame slices of the base mesh adopt the intra-frame coding mode, which can greatly improve the coding efficiency. In this embodiment, the in-frame slices of the reconstructed base mesh and the inter-frame slices of the reconstructed base mesh are merged to obtain the reconstructed base mesh, and displacements are generated based on the reconstructed base mesh. This framework can be consistent with the existing intra-frame coding scheme and is also the standard form that should be. In addition, the in-frame slices of the reconstructed base mesh and the inter-frame slices of the reconstructed base mesh are merged to obtain the reconstructed base mesh. This reconstructed base mesh is a whole and has good standard consistency as the reference mesh for the subsequent meshes.
[0059] It should be noted that this embodiment is not only applicable to the inter-frame three-dimensional base mesh coding, but also applicable to the sub-mesh coding of three-dimensional inter-frame meshes and three-dimensional inter-frame base meshes; when the three-dimensional mesh to be encoded includes multiple sub-meshes, each sub-mesh can be encoded according to the method of the above embodiment, that is, the processing unit of the above proposed coding method can be a sub-mesh, and the sub-mesh is sometimes a Slice of the mesh; when the three-dimensional mesh to be encoded includes multiple Slices, the object of the coding method provided in the above embodiment can be a P slice.
[0060] In an optional embodiment of the present invention, please refer to Figure 2 , Figure 2 is a flowchart of generating a base mesh provided by an embodiment of the present invention. A three-dimensional mesh coding method provided by the present invention, the generation process of the base mesh includes:
[0061] Generate the in-frame slices and inter-frame slices of the current frame base grid based on the current frame input grid and the reference base grid. The reference base grid is the base grid corresponding to the already reconstructed grid, such as the base grid of the previous frame reconstructed grid adjacent in the time domain. The reference base grid can be the base grid of multiple reconstructed grids. For example, when multiple reference frames are used, the present invention describes it with a single reference base grid as an example.
[0062] First, use the inter-frame registration algorithm to deform the reference base grid so that its shape is as similar as possible to the current frame input grid, and obtain the output registered base grid and the registered subdivision grid.
[0063] Secondly, perform non-matching area detection to detect which parts of the inter-frame registered base grid are well registered and which parts are poorly registered. After non-matching area detection, the inter-frame slices of the base grid and the in-frame slices of the primary grid are obtained; among them, the inter-frame slices of the base grid are the well-registered areas taken from the registered base grid. The inter-frame slices of the base grid can carry information matching the reference base grid, such as the vertices in the reference base grid corresponding to each vertex, for use when encoding the inter-frame slices of the base grid later. The in-frame slices of the primary grid are the poorly registered areas in the input grid, and the in-frame slices of the base grid are obtained by further processing in combination with the information of the inter-frame slices of the base grid. Optionally, multiple inter-frame slices of the base grid and / or multiple in-frame slices of the base grid may be obtained simultaneously. It can be understood that some vertices in the inter-frame slices of the base grid and some vertices in the in-frame slices of the base grid have the same position, and the inter-frame slices of the base grid and the in-frame slices of the base grid are connected into an overall base grid.
[0064] Please refer to Figure 3 , Figure 3 which is a schematic diagram of generating the base grid provided by the embodiment of the present invention. Figure 3 In (a), it is the reference base grid. The gray area in the reference base grid is the area matching the inter-frame slices of the base grid, and the information of this area is used for encoding the inter-frame slices of the base grid. The blue area in the reference base grid is the non-matching area. Figure 3 In (b), it is the input grid to be encoded. The black area in the input grid is the area matching the reference base grid, and the inter-frame slices of the base grid are obtained, as shown in Figure 3 In (c). The red area in the input grid is the area not matching the reference base grid, and the in-frame slices of the primary grid are obtained, as shown in Figure 3 In (d). The in-frame slices of the primary grid are processed through the grid to obtain the in-frame slices of the base grid, as shown in Figure 3 In (e). Connect the partially overlapping vertices of the in-frame slices of the base grid and the inter-frame slices of the base grid together to form the base grid, as shown in Figure 3 In (f).
[0065] It should be noted that, in order to connect the in-frame patches and the inter-frame patches of the base mesh into a whole base mesh, it is necessary to adjust (simplify and move the positions) the vertices at the boundary of the in-frame patches of the base mesh according to the vertices at the boundary of the inter-frame patches of the base mesh. During the mesh simplification process, the boundary vertices that are the same as those of the inter-frame patches of the base mesh are kept unchanged, and finally the in-frame patches of the base mesh are obtained.
[0066] The above process of generating the base mesh is described in detail through the following embodiments, specifically:
[0067] The generation process of the registered base mesh and the registered subdivision mesh is as follows:
[0068] Use the inter-frame registration algorithm to deform the base mesh of the reference frame, that is, deform the reference base mesh so that the shape of the deformed mesh is as close as possible to the input mesh of the current frame.
[0069] Please refer to Figure 4 , Figure 4 which is a schematic diagram of generating the registered base mesh provided by the embodiment of the present invention. First, use the nearest neighbor search algorithm to deform the input mesh of the current frame onto the reference frame subdivision mesh with the reference frame subdivision mesh as the target mesh, and output the intermediate mesh; secondly, still use the nearest neighbor search algorithm to deform the generated intermediate mesh onto the reference frame subdivision mesh with the reference frame subdivision mesh as the target mesh, and the deformed mesh is the inter-frame registered subdivision mesh; then, with the generated inter-frame registered subdivision mesh as the target mesh, fit and deform the reference frame base mesh to the target mesh, and the base mesh after the fitting deformation is the inter-frame registered base mesh.
[0070] The process of detecting the mismatched region is as follows:
[0071] Please refer to Figure 5 , Figure 5 which is a schematic diagram of detecting the mismatched region provided by the embodiment of the present invention. The specific steps are as follows:
[0072] ① Traverse the patches of the registered base mesh to obtain the bounding box of each patch;
[0073] ② According to the bounding boxes of the patches of the registered base mesh, obtain the weighted average normal vectors of the input mesh and the corresponding region patches in the registered subdivision mesh;
[0074] ③ Calculate the included angle between the two weighted average normal vectors. If the included angle is greater than the set threshold, mark the corresponding patch in the registered base mesh as a mismatched patch;
[0075] ④ Return to step ① until all the patches of the registered base mesh are traversed;
[0076] ⑤Perform correction processing on the mismatched patches of the registration basic mesh, similar to dilation and erosion in image processing;
[0077] ⑥Adaptively obtain the bounding boxes of several regions in the mismatched patch set;
[0078] ⑦Delete the patches of the registration basic mesh located within the bounding box of the mismatched region to obtain the basic mesh inter-frame patches; at the same time, obtain the patches of the input mesh located within the bounding box of the mismatched region, that is, the primary mesh intra-frame patches.
[0079] The process of boundary vertex simplification and adjustment is as follows:
[0080] The purpose of boundary vertex simplification and adjustment for the primary mesh intra-frame patches: to make the number and position of vertices at the boundary between the inter-frame part of the mesh and the intra-frame part of the mesh consistent, without gaps. Please refer to Figure 6 , Figure 6 which is a schematic diagram of boundary vertex simplification and adjustment provided by an embodiment of the present invention. The specific steps are as follows:
[0081] ①Traverse the boundary points of the basic mesh inter-frame patches, and use the nearest neighbor search algorithm to find the corresponding matching points on the primary mesh intra-frame patches for each boundary point, and mark these matching points;
[0082] ②Traverse the boundary vertices of the primary mesh intra-frame patches, move the non-matching boundary points to the positions of the matching points with the closest Euclidean distance, and then remove duplicate points and degenerate faces;
[0083] ③Move the boundary matching points of the intra-frame patches processed in step ② to the positions of the corresponding boundary vertices of the basic mesh inter-frame patches.
[0084] The process of mesh simplification is as follows:
[0085] Mesh simplification is to simplify the currently input mesh to a basic mesh with relatively fewer points and faces, and to maintain the shape of the original mesh as much as possible. The key point of mesh simplification lies in the simplification operation and the corresponding error metric. Please refer to Figure 7 , Figure 7 which is a schematic diagram of mesh simplification provided by an embodiment of the present invention. Merge the vertices at both ends of the edge into one vertex and delete the connection between these two vertices. Repeat this process in the entire mesh according to certain rules to reduce the number of faces and vertices of the mesh to the target value.
[0086] During the simplification process, a certain error metric can be selected to optimize the simplification result. For example, the sum of the equation coefficients of all adjacent faces of a vertex can be selected as the error metric of the vertex, and the error metric of the corresponding edge is the sum of the error metrics of the two vertices on the edge. In short, the error generated by merging an edge is the sum of the distances from the merged vertex to all the planes adjacent to the original two vertices of the edge.
[0087] After determining the simplification operation and the corresponding error metric, start iterative mesh simplification. First, calculate the vertex error of the initial mesh to obtain the error of each edge. Then, sort each edge in ascending order of error, and select the edge with the smallest error for merging each time. At the same time, calculate the position of the merged vertex and update the errors of all edges related to the merged vertex. That is, update the order of the edge arrangement to ensure that each iteration is based on the global error metric. Simplify the faces of the mesh through iteration to the number required for lossy coding. Finally, obtain the base mesh intra-slice. The base mesh intra-slice, the base mesh intra-frame slice, and the base mesh inter-frame slice form the base mesh.
[0088] In an optional embodiment of the present invention, the base mesh intra-slice is encoded using the intra-coding mode to obtain the base mesh intra-slice bitstream; the base mesh inter-frame slice is encoded using the inter-coding mode to obtain the base mesh inter-frame slice bitstream;
[0089] According to the base mesh intra-slice bitstream, the base mesh inter-frame slice bitstream, and the additional base mesh information, obtain the base mesh bitstream; wherein, the additional base mesh information includes at least one of the number of base mesh intra-slices, the number of base mesh inter-frame slices, the length of the base mesh intra-slice bitstream, and the length of the base mesh inter-frame slice bitstream.
[0090] Specifically, please refer to Figure 8 , Figure 8 is a schematic diagram of the base mesh coding provided by the embodiment of the present invention. In this embodiment, the base mesh coding includes two parts: the base mesh inter-frame slice coding and the base mesh intra-slice coding; for the base mesh inter-frame slice, use the inter-coding mode, that is, based on the reference base mesh information, use the time-domain prediction technology to encode it to obtain the base mesh inter-frame slice bitstream; for the base mesh intra-slice, use the intra-coding mode, that is, without using the reference base mesh information, directly encode the base mesh intra-slice to obtain the base mesh intra-slice bitstream. The base mesh inter-frame slice bitstream and the base mesh intra-slice bitstream are combined to form the base mesh bitstream. Add additional base mesh information to the base mesh bitstream for parsing the base mesh inter-frame slice bitstream and the base mesh intra-slice bitstream. Optionally, the additional base mesh information can be the header information of the base mesh bitstream, including at least one of the number of base mesh inter-frame slices, the number of base mesh intra-slices, the length of the base mesh intra-slice bitstream, and the length of the base mesh inter-frame slice bitstream. The combination of the additional base mesh information, the base mesh inter-frame slice bitstream, and the base mesh intra-slice bitstream can be completed in the bitstream merging module.
[0091] Please refer to Figure 9 , Figure 9It is a schematic diagram of the sum of the base grid inter - slice bitstreams and the structure of the base grid inter - slice bitstreams provided by the embodiments of the present invention. The header information contains identification information to distinguish inter - slice data from intra - slice data; if the inter - slice data precedes the intra - slice data, the length of the inter - slice data can be identified; if the intra - slice data precedes the inter - slice data, the length of the intra - slice data can be identified; the inter - slice data and / or the intra - slice data are byte - aligned, that is, multiples of whole bytes.
[0092] When dealing with N intra - slices, for example, the header information contains identification information of the number N of intra - slices and the size of each intra - slice data. Such as encoding the numerical value N to identify the number N of intra - slices, encoding the number of bytes of each intra - slice data, or encoding the number of bytes of the first N - 1 intra - slice data.
[0093] When dealing with M inter - slices, for example, the header information contains identification information of the number M of inter - slices and the size of each inter - slice data. Such as encoding the numerical value M - 1 (or M) to identify the number M of inter - slices, encoding the number of bytes of each inter - slice data, or encoding the number of bytes of the first M - 1 inter - slice data.
[0094] In the data unit, it can be in the order of the base grid inter - slice and the base grid intra - slice bitstreams, or in the order of the base grid intra - slice and the base grid inter - slice bitstreams.
[0095] Please refer to Figure 10 and Figure 11 , Figure 10 It is a schematic diagram of a bitstream splicing method provided by the embodiments of the present invention. Figure 11 It is another schematic diagram of the bitstream splicing method provided by the embodiments of the present invention. First, write the number of base grid inter - slices, the number of base grid intra - slices, the number of bytes occupied by each base grid inter - slice bitstream segment, and the number of bytes occupied by each base grid intra - slice bitstream segment in sequence in the header information of the inter - frame base grid / sub - grid; then, store each base grid inter - slice bitstream segment and each base grid intra - slice bitstream segment in the corresponding order in the data unit.
[0096] In an optional embodiment of the present invention, the encoding of the base grid intra - slice includes at least one of encoding connection relationships, vertex geometric coordinates, and texture coordinates, and the encoding of the base grid inter - slice includes at least one of encoding connection relationships, vertex geometric coordinates, and texture coordinates; among them,
[0097] For encoding the texture coordinates of the base grid intra - slice, it includes encoding the overall offset and / or scaling of the texture coordinates of the base grid intra - slice.
[0098] For the base mesh inter - slice, encode the connection relationship, vertex geometric coordinates, and texture coordinates according to the information of the reference base mesh; wherein, the reference base mesh represents the base mesh corresponding to the reconstructed base mesh in the time domain, and the information of the reference base mesh includes the connection relationship, vertex geometric coordinates, and texture coordinates.
[0099] Specifically, in this embodiment, it is described in detail through the following process.
[0100] The base mesh inter - slice encoding process includes encoding at least one of the connection relationship information, vertex geometric information, and texture coordinate information in the base mesh inter - slice. Among them,
[0101] One way to encode the connection relationship information is to directly use the connection relationship of the reference base mesh, or a syntax element can be added to the header information to indicate whether to directly use the connection relationship of the reference base mesh.
[0102] One way to encode the texture coordinate information is to directly use the texture coordinate information of the reference base mesh, and the texture coordinates of the vertices in the base mesh inter - slice directly use the texture coordinates of the corresponding points in the reference base mesh. Or the texture coordinates of the corresponding points in the reference base mesh can be uniformly scaled or translated, and the scaling or translation information is encoded in the header information. The acquisition of the scaling or translation information utilizes the information of the base mesh inter - slice, such as according to the texture coordinates of the base mesh inter - slice.
[0103] For vertex geometric information encoding, it can be decided whether to encode the motion vector information of each vertex in the base mesh inter - slice according to the rate - distortion criterion. When encoding the motion vector, it is called P - base mesh inter - slice encoding; when not encoding the motion vector, it is called Skip - base mesh inter - slice encoding.
[0104] Whether it is P - base mesh inter - slice encoding or Skip - base mesh inter - slice encoding, it is necessary to encode the reference base mesh information, which is used to indicate which part of the reference base mesh the base mesh inter - slice refers to. One possible encoding method for the reference information is as follows:
[0105] ① Set and encode the number of reference vertices, which is used to indicate how many vertices of the reference base mesh the current base mesh inter - slice refers to;
[0106] ② Encode the indices of the corresponding reference vertices in the reference base mesh in the order of the vertex indices of the base mesh inter - slice;
[0107] ③ Set and encode the number of reference triangular patches, which is used to indicate how many triangular patches of the reference base mesh the current base mesh inter - slice refers to;
[0108] ④Encode the indices of the corresponding reference triangular patches in the reference base mesh in the order of the triangular patch indices of the base mesh frames.
[0109] For the encoding of the above reference vertex indices and reference triangular patch indices, the possible encoding methods include:
[0110] Directly binary-encode the index values.
[0111] Perform differential encoding on the index information, that is, encode the difference between the current index value and the previous index value.
[0112] For the encoding of the motion vector information of the P base mesh frame slices, a possible method is to follow the existing motion vector encoding method in V-DMC. The steps are as follows:
[0113] ①Divide every 16 vertices of the P base mesh frame slice into a group in the order of vertex indices.
[0114] ②Perform Skip mode determination at the motion vector group level according to the rate-distortion criterion to decide whether to encode the motion vectors of this group of vertices.
[0115] ③If the motion vectors of this group of vertices are not encoded, the positions of this group of vertices need to be adjusted to the positions of the corresponding vertices in the reference frame.
[0116] ④If the motion vectors of this group of vertices are encoded, it is necessary to traverse and compare the encoding bit costs of the motion vector residuals corresponding to the three motion vector prediction modes, and select the prediction mode with the smaller bit cost; when encoding, first encode the prediction mode identifier of this group of vertices, and then encode the motion vector residuals of this group of vertices.
[0117] The encoding of the base mesh intra-slice includes encoding at least one of the connection relationship information, vertex geometry information, and texture coordinate information in the base mesh intra-slice.
[0118] In one case, directly encode the connection relationship information, vertex geometry information, and texture coordinate information of the base mesh intra-slice. A possible encoding method is to use, for example, Draco to compress the static mesh, encode the connection relationship, vertex geometry coordinates, and texture coordinates of the base mesh, and finally output the base mesh intra-slice bitstream. The following introduces a usable static base mesh encoder, Draco.
[0119] The main idea of Draco to compress the static mesh is mesh compression driven by the connection relationship. It traverses all the faces of the mesh in a specific way, marks each face according to a specific rule, encodes the marks of all the faces obtained by traversal, that is, encodes the mesh connection relationship. Subsequently, encode all the vertex coordinate information in the order of traversing the connection relationship.
[0120] The main process of the Draco encoding grid includes: First, for the input grid, connection relationships are generated based on its geometric information, that is, the connection relationships of vertices in three-dimensional space. After constructing the connection relationships of faces, an initial face is selected to start traversing all the faces of the current grid, that is, traversing to generate symbols. Here, the Edgebreaker algorithm is used to traverse and generate symbols. This algorithm divides the current corner into five modes according to the state of the triangular face where it is located when traversing to the current corner. Please refer to Figure 12 , Figure 12 which is a schematic diagram of one of the five modes of Edgebreaker provided by the embodiments of the present invention.
[0121] The above five modes define the direction of traversing the next face after traversing to the current face. According to the above traversal method, corresponding symbols are generated for each face defined by the geometric information of the current grid, and then these symbols are entropy encoded to obtain the bitstream of the connection relationships defined by the geometric information of the current grid. At the same time, the order of traversing the corresponding vertices is also obtained by traversing each face. This vertex order is passed to the geometric information encoder, rearranged in the order of traversal, quantized according to the predetermined quantization parameters, and then predicted. The prediction uses the parallelogram prediction method.
[0122] The encoding method of texture coordinates is similar to that of vertex geometric information, and prediction can also be used. When the connection relationship of texture coordinates is different from that of vertex geometric information, it is necessary to encode the connection relationship of texture coordinates, or the difference from the connection relationship of vertex geometric information can also be encoded.
[0123] The texture coordinates of the base grid intra-slice can be uniformly scaled and / or translated using the texture coordinates of the base grid inter-slice, and the scaling / translation information is encoded in the header information. Among them, the adjustment schemes for scaling and / or translation include:
[0124] Adjustment scheme one: The texture coordinates of the inter-slice remain unchanged, and only the texture coordinates of the intra-slice are adjusted. Then, only the adjusted texture coordinates of the intra-slice need to be encoded when encoding the texture coordinates.
[0125] Adjustment scheme two: The texture coordinates of the inter-slice and the intra-slice are uniformly adjusted.
[0126] Another encoding method can also be used for the texture coordinates of the intra-slice, that is, the parameterized information of the texture coordinates of the intra-slice is encoded, and the decoding end decodes the texture coordinates of the intra-slice with the help of the parameterized information.
[0127] Please refer to Figure 13 , Figure 13It is a schematic diagram of texture coordinate parameterization provided by an embodiment of the present invention. First, corresponding texture coordinates are obtained from the reference base mesh for the base mesh inter-frame slices; at the same time, mesh parameterization is performed on the base mesh intra-frame slices to generate texture coordinates for the intra-frame slices; finally, the texture coordinates of the base mesh inter-frame slices and the base mesh intra-frame slices are adjusted and uniformly normalized to facilitate subsequent conversion to generate a texture map. Among them, mesh parameterization can be for the base mesh intra-frame slices or for the reconstructed base mesh intra-frame slices.
[0128] Mesh parameterization is used to generate corresponding texture coordinates for the mesh. Currently, there are many algorithms for parameterizing the mesh, such as the Isochart algorithm, the orthogonal projection algorithm, etc. In this coding framework, both of the above two schemes can be used to parameterize the reconstructed base mesh. The following briefly introduces the two algorithms:
[0129] Isochart algorithm
[0130] This algorithm uses spectral analysis to achieve stretch-driven 3D mesh parameterization, unfolds, slices, and packs the 3D mesh into a 2D texture domain. A stretch threshold is set, and its algorithm overview is as follows:
[0131] a) Calculate the surface spectral analysis to provide an initial parameterization;
[0132] b) Perform iterative stretch optimization;
[0133] c) If the stretch of this derived parameterization is less than the threshold, stop;
[0134] d) Perform surface spectral clustering to divide the surface into charts;
[0135] e) Use the graph cut algorithm to optimize the chart boundaries;
[0136] f) Iteratively segment the charts until the stretch criterion is met.
[0137] Orthogonal projection algorithm
[0138] orthoAtlas is a projection-based mesh parameterization method that generates texture coordinates for the mesh through orthogonal projection. The main process includes:
[0139] a) Calculate the mesh attributes, including the adjacent faces of each face and the area and normal vector of each face;
[0140] b) Determine the projection plane of each face according to the normal vector;
[0141] c) Start clustering all faces according to the projection plane to form a connected region, and first select the starting face for clustering;
[0142] d), starting from the starting surface, iterate to determine whether the adjacent surfaces of the surfaces added to the connected region can be added to the connected region;
[0143] e), after each connected region iteration is completed, obtain multiple connected regions;
[0144] f), determine whether to merge adjacent connected regions according to the error metric;
[0145] g), detect whether there is an overlapping region during projection, remove the overlapping surfaces and regenerate the connected regions;
[0146] h), arrange all the projected regions onto a two-dimensional image.
[0147] In an alternative embodiment of the present invention, merging the intra-frame slices of the reconstructed basic mesh and the inter-frame slices of the reconstructed basic mesh to obtain the reconstructed basic mesh includes:
[0148] Adjust the geometric information of some vertices of the intra-frame slices of the reconstructed basic mesh to make it the same as the geometric information of the corresponding part of the vertices of the inter-frame slices of the reconstructed basic mesh;
[0149] Delete the duplicate vertices and connection relationships between the intra-frame slices of the reconstructed basic mesh and the inter-frame slices of the reconstructed basic mesh.
[0150] Specifically, in this embodiment, the inter-frame slices of the reconstructed basic mesh and the intra-frame slices of the reconstructed basic mesh are merged to obtain the reconstructed basic mesh. The geometric information of the corresponding vertices of the inter-frame slices of the reconstructed basic mesh and the intra-frame slices of the reconstructed basic mesh may be different. The corresponding vertices are obtained by the nearest neighbor method, and the geometric information of the corresponding vertices of the intra-frame slices of the reconstructed basic mesh is adjusted to make the geometric information of the vertices of the intra-frame slices of the reconstructed basic mesh the same as the geometric information of the corresponding vertices of the inter-frame slices of the reconstructed basic mesh. Then, the duplicate vertices and connection relationships are deleted to obtain the reconstructed basic mesh.
[0151] In an alternative embodiment of the present invention, performing a subdivision process on the reconstructed basic mesh to obtain the displacement of the subdivided mesh, and encoding the displacement of the subdivided mesh to obtain a displacement bitstream, includes:
[0152] Using the displacement information of the reference basic mesh, predicting the prediction residual of the displacement of the inter-frame slice of the basic mesh, and encoding the prediction residual of the displacement to achieve the encoding of the displacement of the inter-frame slice of the basic mesh.
[0153] Specifically, please refer to Figure 14 , Figure 14 is a schematic diagram of a geometric displacement vector calculation method provided by an embodiment of the present invention. For the reconstructed basic mesh, perform a subdivision deformation to obtain the displacement. The basic idea of the subdivision and deformation module is as Figure 14As shown in , the same concept is applied to the input reconstructed base mesh to generate displacement vector information. Figure 14 In , the input 2D curve (represented by a 2D polyline), called the "original" curve, is first downsampled to generate a base curve / polyline, called the "simplified" curve. The subdivision scheme is then applied to the simplified polyline to generate the "subdivided" curve. The subdivided polyline is then deformed to obtain a better approximation of the original curve. That is, a geometric displacement vector ( Figure 14 The shape of the subdivided curve is made as close to the shape of the original curve as possible (as shown by the arrow in ). These geometric displacement vectors are the geometric displacement vector information output by this module. The same deformation process is also applied to the attribute information corresponding to the vertex to obtain the corresponding attribute displacement vector.
[0154] For the parameterized sub-grids, the input mesh is first subdivided. The subdivision scheme can be arbitrarily selected. One possible scheme is the midpoint subdivision scheme, that is, each triangle is subdivided into four sub-triangles in each subdivision iteration, such as Figure 15 As shown, Figure 15 A schematic diagram of subdivision provided by an embodiment of the present invention is shown in FIG. A new vertex is introduced in the middle of each edge, and the subdivision of geometric information and attribute information is performed independently because the connection relationship between geometric information and attribute information is usually different.
[0155] The method of introducing a new vertex in the middle of each edge is to calculate the newly introduced midpoint v of the edge (v1, v2) 12 The position Pos(v 12 ) is shown in formula (1):
[0156]
[0157] Among them, Pos(v1) and Pos(v2) represent the geometric coordinates of vertex v1 and vertex v2 respectively.
[0158] For the subdivided mesh, find the nearest neighbor of each point on the original input mesh (including the points on the original mesh surface), which can be accelerated by using data structures such as kdTree. The displacement vector of the geometric coordinates of each vertex of the subdivided mesh is obtained by calculating the distance between each vertex on the subdivided mesh and its nearest neighbor on the original input mesh.
[0159] The coordinate system of the calculated vertex displacement is converted from the Cartesian coordinate system to the local coordinate system. A feasible method is to convert each vertex displacement coordinate into the coordinate system composed of the corresponding vertex normal vector and two vectors tangent to the normal vector. The specific conversion process is shown in formula (2):
[0160]
[0161] Among them, represents the vertex displacement before coordinate system transformation, represents the vertex displacement after coordinate system transformation, and represent pairwise mutually orthogonal unit vectors, and is the vertex normal vector, and represent two vectors tangent to For the calculation of the vertex normal vector, a feasible method is: equal to the area-weighted sum of the normal vectors of the adjacent patches of the vertex.
[0162] For the above displacement generation scheme, in order to facilitate subsequent unified coding processing of the displacement, it is necessary to simply combine the two parts of the displacement and output the complete displacement of the current frame.
[0163] Please refer to Figure 16 , Figure 16 which is a schematic diagram of the displacement coding provided by the embodiment of the present invention. First, perform inter-frame prediction on the inter-frame slice displacement to obtain the displacement residual; secondly, perform wavelet transform, quantization, etc. on the displacement to obtain the quantized wavelet transform coefficients; finally, encode the quantized wavelet transform coefficients to obtain the displacement bitstream. Next, each step in Figure 16 will be introduced in detail:
[0164] The inter-frame slice prediction includes:
[0165] Only operate on the inter-frame slice displacement data. For the inter-frame slice displacement data, there are two possible coding methods: one is to directly code the displacement information; the other is to code the prediction residual. When coding the prediction residual of the inter-frame slice displacement, it is necessary to use this module to calculate the prediction residual of the inter-frame slice displacement data. A feasible calculation method is: calculate the difference between the current vertex displacement and the corresponding vertex displacement of the reference frame.
[0166] The wavelet transform includes:
[0167] The transform can be applied to the displacement vector to reduce the correlation between its data. An optional transform is, for example, the linear wavelet transform, and its prediction process is defined as shown in Equation (3):
[0168]
[0169] where v represents the newly inserted midpoint on the edge (v1, v2), and Signal(v), Signal(v1), and Signal(v2) respectively represent the displacement vectors corresponding to vertices v, v1, and v2. After predicting the displacement vector of vertex v, it is updated, and the update process is defined as shown in Equation (4):
[0170]
[0171] Among them, v * represents the set of all vertices adjacent to vertex v, and the transformed displacement vector is called the wavelet coefficient.
[0172] The transformed displacement vector, i.e., the wavelet coefficient, can be quantized, and there are various quantization methods. One method is shown in Eqs. (5) and (6):
[0173] disp[v].d[k] = floor(disp[v].d[k] * scale[k]) (5);
[0174]
[0175] Among them, disp[v] represents the value of the transformed displacement vector of the v-th vertex, d[k] represents the k-th value of the displacement vector, floor represents rounding down. bitDepthPosition represents the bit depth of the geometric position of the current mesh vertex, and qp[k] represents the quantization parameter of the k-th coefficient. As mentioned above, after converting the coordinate system of the displacement vector, its normal component has a more significant effect on the quality than the tangential component. Therefore, a larger quantization parameter can be used for the tangential component.
[0176] Meanwhile, according to the characteristics of wavelet transform, different quantization parameters can also be used for the newly generated vertices and the original vertices after subdivision. That is, for the vertices after subdivision, the quantization parameter is updated as shown in Eq. (7):
[0177] scale[k] = scale[k] * lodScale[k] (7);
[0178] Among them, lodScale[k] represents the coefficient of the quantization parameter at the current subdivision level.
[0179] There are various possibilities for the specific method of encoding the quantized wavelet transform coefficients. One possibility is to reuse the existing method of V-DMC: arrange the quantized wavelet transform coefficients into video frames and send them to a video encoder for encoding, or directly perform entropy encoding; another possible method is to directly perform entropy encoding on the displacement information without going through the aforementioned wavelet transform, quantization, etc.
[0180] The deformed mesh reconstruction scheme is as Figure 17 shown, Figure 17It is a schematic diagram of deformed grid reconstruction provided by an embodiment of the present invention. First, the reconstructed basic grid inter-frame slices and the reconstructed basic grid intra-frame slices are merged to restore the complete reconstructed basic grid. Then, the basic grid is subdivided. After inverse quantization and inverse wavelet transform processing of the reconstructed displacement, frame-inter slice displacement prediction can be selectively performed according to the coding settings. Finally, the displacement is obtained and superimposed on the subdivided basic grid, and the reconstructed deformed grid can be obtained finally.
[0181] In an alternative embodiment of the present invention, texture map conversion is performed according to the input original texture map, input grid, and reconstructed deformed grid. Please refer to Figure 18 , Figure 18 It is a schematic diagram of texture map conversion provided by an embodiment of the present invention, and the specific conversion steps are as follows:
[0182] ① Calculate the texture coordinates of each pixel on the texture map to be generated. For example, the texture coordinates corresponding to pixel A(i,j) are P(u,v);
[0183] ② Determine whether the texture coordinates are within a certain triangular face after parameterization of the subdivided deformed grid;
[0184] ③ If the texture coordinates do not belong to any triangular face, mark the pixel as an empty pixel, and then it can be filled with a filling algorithm;
[0185] If the texture coordinates belong to a triangular face, then,
[0186] Mark the pixel as filled;
[0187] Calculate the barycentric coordinates of the texture coordinates in the current triangular face;
[0188] According to the barycentric coordinates and the corresponding triangular face, map the two-dimensional texture coordinates to three-dimensional geometric coordinates, that is, map to the point on the subdivided deformed grid corresponding to the texture coordinates, as shown by M(x,y,z) in the figure;
[0189] Find the point on the input original grid that is closest to the three-dimensional coordinate, as shown by M'(x,y,z) in the figure;
[0190] Calculate the barycentric coordinates of the three-dimensional coordinate according to the triangular face where it is located and map it to two dimensions, and calculate its texture coordinates, that is, P'(u',v');
[0191] Sample through the texture coordinates on the input original texture map to obtain the value A'(i',j') at the corresponding pixel position;
[0192] Assign this value to the corresponding pixel A(i,j) on the texture map to be generated.
[0193] Subsequently, after obtaining the converted texture map, for the empty pixels therein, existing filling algorithms (such as the Push-Pull algorithm) can be used to fill these empty pixels. Then, existing video encoders, such as H.264 / AVC, H.265 / HEVC, H.266 / VVC, etc., can be used to encode it to obtain the bitstream of the output texture map. In addition, operations such as color space conversion and chroma subsampling can be selectively applied to make the video encoding obtain better rate-distortion performance, such as the color space conversion from RGB444 to YUV 420.
[0194] It should be noted that during the encoding process, such as the type of mesh encoder, the type of video encoder, the mesh subdivision scheme, the displacement transformation scheme, etc., different methods can be replaced according to actual needs. Therefore, the selected scheme needs to be passed to the decoding end to guide correct decoding. The auxiliary base mesh information also includes an optional reference frame list, which identifies the index list of the reference frames used by the current frame; the subdivision identifier indicates whether the base mesh of the current frame needs to be subdivided and deformed, that is, whether it contains displacement information. The auxiliary base mesh information can also include the type of static mesh encoder, the type of video encoder, the mesh subdivision scheme, the displacement transformation scheme, the coefficient arrangement scheme, and the color conversion scheme, etc.
[0195] After encoding is completed, the base mesh bitstream, texture coordinate bitstream, displacement bitstream, texture map bitstream, auxiliary information bitstream, etc. are mixed to obtain the finally output encoded bitstream.
[0196] Based on the same inventive concept, please refer to Figure 19 , Figure 19 is a schematic diagram of a three-dimensional mesh decoding method provided by an embodiment of the present invention. The decoding method is implemented corresponding to the encoding method one by one. The decoding method includes:
[0197] Decode the base mesh intra-slice bitstream to obtain the base mesh intra-slice, decode the base mesh inter-slice bitstream to obtain the base mesh inter-slice, decode the base mesh intra-slice to obtain the reconstructed base mesh intra-slice, and decode the base mesh inter-slice to obtain the reconstructed base mesh inter-slice;
[0198] Merge the reconstructed base mesh intra-slice and the reconstructed base mesh inter-slice to obtain the reconstructed base mesh;
[0199] Decode the displacement bitstream to obtain the displacement of the subdivided mesh after the reconstruction base mesh is subdivided. According to the reconstruction base mesh and the displacement of the subdivided mesh, obtain the reconstructed mesh.
[0200] Specifically, please continue to refer to Figure 19, a 3D mesh decoding method provided in this embodiment demultiplexes the bitstream obtained by the encoding end to respectively obtain a base mesh bitstream, an auxiliary base mesh information bitstream, a displacement bitstream, and a texture map bitstream. Based on these bitstreams, a 3D mesh is reconstructed to complete the entire process of 3D mesh decoding.
[0201] It should be noted that the method used in the decoding process depends on the method used in the encoding process.
[0202] In this embodiment, the base mesh bitstream is split into an intra-slice bitstream of the base mesh and an inter-slice bitstream of the base mesh. The intra-slice bitstream of the base mesh and the inter-slice bitstream of the base mesh are respectively decoded to obtain an intra-slice of the base mesh and an inter-slice of the base mesh; the intra-slice of the base mesh and the inter-slice of the base mesh are respectively decoded to obtain a reconstructed intra-slice of the base mesh and a reconstructed inter-slice of the base mesh. The reconstructed intra-slice of the base mesh and the reconstructed inter-slice of the base mesh are merged to obtain a reconstructed base mesh; after decoding the displacement bitstream, a displacement is obtained. Based on the displacement, a reconstructed deformed mesh is obtained; based on the reconstructed deformed mesh and the reconstructed base mesh, a reconstructed 3D mesh is obtained; the texture map bitstream is decoded to obtain a texture map, and further a reconstructed texture map is obtained; the vertices of the reconstructed mesh can find corresponding positions in the reconstructed texture map according to their texture coordinates and are rendered. It should be emphasized that first, the auxiliary base mesh information bitstream needs to be decoded to obtain auxiliary base mesh information, which plays a guiding role in the decoding processes of the base mesh bitstream, the displacement bitstream, and the texture map bitstream to achieve successful parsing of the above bitstreams; it can be understood that the auxiliary base mesh information includes at least one of the type of mesh encoder used, the type of video encoder used, the mesh subdivision scheme, the displacement transformation scheme, etc. In other words, according to the method used in the encoding process, the corresponding method is used for decoding in the decoding process.
[0203] In an optional embodiment of the present invention, additional base mesh information is obtained by decoding the base mesh bitstream. Based on the additional base mesh information, a decoding method corresponding to the method of encoding the intra-slice of the base mesh is determined to decode the intra-slice bitstream of the base mesh; a decoding method corresponding to the method of encoding the inter-slice of the base mesh is determined to decode the inter-slice bitstream of the base mesh; wherein, the additional base mesh information includes at least one of the number of intra-slices of the base mesh, the number of inter-slices of the base mesh, the length of the intra-slice bitstream of the base mesh, the length of the inter-slice bitstream of the base mesh, the encoding method of the intra-slice of the base mesh in the current frame, and the encoding method of the inter-slice of the base mesh in the current frame.
[0204] Specifically, in this embodiment, the decoding end first decodes the auxiliary base mesh information bitstream to obtain the auxiliary base mesh information, and determines the decoding scheme according to the auxiliary base mesh information. Among them, the auxiliary base mesh information mainly includes an intra-coding flag, which indicates whether the current frame needs to be constructed based on the base mesh of the reference frame, that is, whether it needs to be constructed based on the reference base mesh; a reference frame list, which indicates the indexes of the reference frames that the current frame needs to use, and this reference frame list is applied to the subsequent base mesh construction process; a subdivision flag, which indicates whether subsequent subdivision deformation operations need to be performed on the reconstructed base mesh; a static mesh encoder type, which guides the decoding end to use the corresponding static mesh decoder; a video encoder type, which guides the decoding end to use the corresponding video decoder; a subdivision scheme, that is, the scheme for subdividing the base mesh in the reconstructed deformed mesh, and the subdivision schemes of the encoding and decoding ends should be consistent; it also includes other optional displacement transformation schemes, coefficient arrangement schemes, etc.; it should be noted that the auxiliary base mesh information described here includes the independently transmitted auxiliary base mesh information and the header information that may be included in other bitstream parts.
[0205] In an optional embodiment of the present invention, the decoding of the intra-slice of the base mesh includes decoding at least one of the connection relationship, vertex geometric coordinates, and texture coordinates; the decoding of the inter-slice of the base mesh includes decoding at least one of the connection relationship, vertex geometric coordinates, and texture coordinates; among them,
[0206] For decoding the texture coordinates of the intra-slice of the base mesh, according to the additional base mesh information, determine the overall offset and / or scaling of the texture coordinates of the encoded intra-slice of the base mesh, and decode the intra-slice bitstream of the base mesh according to the overall offset and / or scaling.
[0207] Specifically, in this embodiment, the structure of the base mesh bitstream at the encoding end is the same as that at the decoding end. According to the structure of the base mesh bitstream, the base mesh bitstream is divided into an intra-slice bitstream of the base mesh and an inter-slice bitstream of the base mesh; if the base mesh bitstream includes multiple intra-slice bitstreams of the base mesh, it is divided into multiple intra-slice bitstreams of the base mesh; if the base mesh bitstream includes multiple inter-slice bitstreams of the base mesh, it is divided into multiple inter-slice bitstreams of the base mesh; the obtained inter-slice bitstream of the base mesh or intra-slice bitstream of the base mesh only contains one inter-slice information of the base mesh or one intra-slice information of the base mesh, which is called the base mesh slice bitstream, and each base mesh slice bitstream is decoded independently according to its type (intra or inter).
[0208] The decoding of the inter-slice of the base mesh includes:
[0209] The decoding of the base mesh inter - frame slice includes decoding at least one of the connection relationship information, vertex geometry information, and texture coordinate information in the base mesh inter - frame slice. The decoding methods of information such as connection relationship information, vertex geometry information, and texture coordinate information correspond to the encoding methods.
[0210] One way to decode the connection relationship information is to directly use the connection relationship of the reference base mesh. Whether to directly use the connection relationship of the reference base mesh can be determined according to the indication of the syntax element in the header information.
[0211] One way to decode the texture coordinate information is to directly use the texture coordinate information of the reference base mesh. The texture coordinates of the vertices in the base mesh inter - frame slice directly use the texture coordinates of the corresponding points in the reference base mesh. If the texture coordinates of the corresponding points in the reference base mesh are uniformly scaled or translated, then information on scaling or translation of the texture coordinates is required.
[0212] To decode the vertex geometry information, first, it is necessary to decode whether the type of the base mesh inter - frame slice is a P - type base mesh inter - frame slice or a Skip - type base mesh inter - frame slice. If it is a P - type base mesh inter - frame slice, then its reference information needs to be decoded first, and then the motion vector information needs to be decoded; if it is a Skip - type base mesh inter - frame slice, then only the reference information needs to be decoded.
[0213] Whether it is a P - type base mesh inter - frame slice or a Skip - type base mesh inter - frame slice, it is necessary to decode the reference information, which is used to indicate which part of the reference base mesh the base mesh inter - frame slice refers to. Corresponding to the encoding end, one possible way to decode the reference information is as follows:
[0214] ① Decode the number of reference vertices;
[0215] ② Decode the indices of the reference vertices;
[0216] ③ Decode the number of reference triangular patches;
[0217] ④ Decode the indices of the reference triangular patches.
[0218] Based on the decoded reference vertex index information and reference patch index information above, the decoding end can recover a reconstructed base mesh inter - frame slice that is exactly the same as the encoding end.
[0219] For P - type base mesh inter - frame slices, it is also necessary to further decode the motion vector information. One possible decoding method is to follow the existing method of V - DMC. The specific steps are as follows:
[0220] ① First, decode the Skip - mode flag of the motion vector group of a set of vertices;
[0221] ② If the encoding mode of this group of motion vectors is the Skip mode, there is no need to continue decoding the motion vectors, and the corresponding vertex positions in the reference frame can be directly assigned to this group of vertices;
[0222] ③ If the encoding mode of this group of motion vectors is not the Skip mode, continue to decode the motion vector prediction mode flag and all the motion vector residuals of the vertices in this group; according to the prediction mode flag, use the corresponding prediction method to calculate the motion vector prediction value of this group of vertices, and then add it to the decoded motion vector prediction residual to restore the motion vector information of this group of vertices.
[0223] For the P-based mesh inter slice, after decoding the motion vector information of each vertex, it can be superimposed on the corresponding vertex positions of the inter slice reference base mesh to decode and reconstruct the base mesh inter slice.
[0224] The decoding of the base mesh intra slice includes:
[0225] The decoding of the base mesh intra slice includes decoding at least one of the connection relationship information, vertex geometry information, and texture coordinate information in the base mesh intra slice. The decoding methods of information such as connection relationship information, vertex geometry information, and texture coordinate information correspond to the encoding methods.
[0226] If the connection relationship information, vertex geometry information, and texture coordinate information of the base mesh intra slice are directly encoded using, for example, Draco compressed static meshes, use the corresponding decoding algorithm to decode the connection relationship, vertex geometric coordinates, and texture coordinates of the base mesh intra slice, and finally output the base mesh intra slice bitstream. A available static base mesh encoder Draco is introduced below.
[0227] If the texture coordinates of the intra slice adopt the parametric information encoding mode, the decoding end decodes the intra slice texture coordinates with the help of the parametric information.
[0228] In an optional embodiment of the present invention, merging and reconstructing the base mesh intra slice and the reconstructed base mesh inter slice to obtain the reconstructed base mesh includes:
[0229] Adjust the geometric information of some vertices of the reconstructed base mesh intra slice to be the same as that of some vertices of the reconstructed base mesh inter slice;
[0230] Delete the duplicate vertices and connection relationships of the reconstructed base mesh intra slice and the reconstructed base mesh inter slice.
[0231] Specifically, in this embodiment, the process of the encoding end merging the reconstructed intra-slice of the base mesh and the reconstructed inter-slice of the base mesh is the same as that of the decoding end merging the reconstructed intra-slice of the base mesh and the reconstructed inter-slice of the base mesh. The reconstructed inter-slice of the base mesh and the reconstructed intra-slice of the base mesh are merged to obtain the reconstructed base mesh. The geometric information of the corresponding vertices of the reconstructed inter-slice of the base mesh and the reconstructed intra-slice of the base mesh may be different. The corresponding vertices are obtained by the nearest neighbor method, and the geometric information of the corresponding vertices of the reconstructed intra-slice of the base mesh is adjusted to make the geometric information of the vertices of the reconstructed intra-slice of the base mesh the same as that of the corresponding vertices of the reconstructed inter-slice of the base mesh. Then, the duplicate vertices and connection relationships are deleted to obtain the reconstructed base mesh.
[0232] In an alternative embodiment of the present invention, the displacement bitstream is decoded to obtain the displacement of the subdivided mesh after the reconstruction base mesh is subdivided. The displacement of the subdivided mesh is decoded to obtain the reconstructed deformed mesh, including:
[0233] Decoding the displacement of the subdivided mesh to obtain the prediction residual of the displacement of the inter-slice of the base mesh, and combining the displacement information of the reference base mesh to obtain the displacement of the inter-slice of the base mesh.
[0234] Specifically, please refer to Figure 20 , Figure 20 which is a schematic diagram of displacement decoding provided by the embodiment of the present invention. Displacement decoding is the inverse process of displacement encoding.
[0235] First, the displacement needs to be decoded. The decoding of the displacement needs to be consistent with the encoding end. If the encoding end encodes the displacement using video coding, the decoding end needs to use video decoding to decode; if the encoding end directly performs entropy coding on the displacement, the decoding end needs to perform entropy decoding on the displacement bitstream. After the displacement decoding is completed, the quantized wavelet transform coefficients can be obtained.
[0236] Secondly, the quantized wavelet transform coefficients are dequantized and inverse wavelet transformed. If the encoding end directly encodes the displacement of the inter-slice, the decoded displacement is output after the inverse wavelet transform; if the encoding end uses inter-frame prediction coding for the displacement of the inter-slice, the decoding end needs to perform the same prediction process to obtain the predicted value of the displacement of the inter-slice, and add it to the decoded displacement residual of the inter-slice to output the final decoded displacement.
[0237] Please continue to refer to Figure 17 , and obtain the reconstructed deformed mesh according to the displacement obtained above; obtain the reconstructed three-dimensional mesh according to the reconstructed deformed mesh and the reconstructed base mesh.
[0238] In an alternative embodiment of the present invention, under the guidance of the decoded auxiliary base mesh information, the texture map bitstream is decoded, that is, decoded using the video decoder indicated in the auxiliary base mesh information, and an optional color space conversion is performed on it to obtain an image format consistent with the input texture at the encoding end, and the finally decoded output texture map is obtained. Figure 1 The decoding end finally obtains the reconstructed deformed mesh at the decoding end and the corresponding texture map, and subsequent applications use the reconstructed deformed mesh and texture map as inputs for processing.
[0239] The decoding end finally obtains the reconstructed deformed mesh at the decoding end and the corresponding texture map, and subsequent applications use the reconstructed deformed mesh and texture map as inputs for processing.
[0240] In an alternative embodiment of the present invention, in order to ensure that multiple base mesh inter-frame slice bitstreams and intra-frame slice base mesh bitstreams can be distinguished at the decoding end, multiple identification information needs to be set in the syntax structure to distinguish multiple base mesh intra-frame slice bitstream segments and base mesh inter-frame slice bitstream segments. The syntax structure provided in this embodiment is designed based on the V-DMC syntax structure. The identification information should be placed in the header information, and both the base mesh inter-frame slice bitstream and the base mesh intra-frame slice bitstream should be placed in the data unit. The relevant syntax structure is shown in Table 1.
[0241] Table 1 Syntax Structure
[0242]
[0243] smh_inter_segment_count indicates the number of base mesh inter-frame slices.
[0244] smh_intra_segment_count indicates the number of base mesh intra-frame slices.
[0245] smh_inter_segment_byte_counts indicates the number of bytes occupied by a certain base mesh inter-frame slice bitstream in the data unit, and is used to separate multiple base mesh code inter-frame slice segments and base mesh intra-frame slice bitstream segments.
[0246] smh_intra_segment_byte_counts indicates the number of bytes occupied by a certain base mesh intra-frame slice bitstream in the data unit, and is used to separate multiple base mesh code inter-frame slice segments and base mesh intra-frame slice bitstream segments.
[0247] In the data unit syntax, the P-frame inter-base mesh type and the Skip-frame inter-mesh type are cancelled, and the non-intra base mesh is uniformly used for representation. When the base mesh / sub-mesh type is not intra-frame, its data unit contains multiple base mesh inter-frame slice bitstream segments and base mesh intra-frame slice bitstream segments. The relevant syntax structure is shown in Table 2.
[0248] Table 2 Syntax Structure
[0249]
[0250] Based on the same inventive concept, please continue to refer to Figure 1 , the present invention further provides a three-dimensional mesh encoding device, which is applied to a three-dimensional mesh encoding method provided in the above embodiments of the present invention. For the embodiments of the method, please refer to the above, and details will not be repeated here. The three-dimensional mesh encoding device includes:
[0251] A first encoding module, configured to divide a three-dimensional mesh into base mesh intra-frame slices and base mesh inter-frame slices, and perform encoding to obtain a base mesh intra-frame slice bitstream and a base mesh inter-frame slice bitstream, and simultaneously obtain a reconstructed base mesh intra-frame slice and a reconstructed base mesh inter-frame slice;
[0252] A processing module, configured to merge the reconstructed base mesh intra-frame slice and the reconstructed base mesh inter-frame slice to obtain a reconstructed base mesh;
[0253] A second encoding module, configured to perform subdivision processing on the reconstructed base mesh to obtain displacements of the subdivided meshes, and encode the displacements of the subdivided meshes to obtain a displacement bitstream.
[0254] Specifically, a three-dimensional mesh encoding device provided in this embodiment is applied to an encoding end, and includes a first encoding module, a processing module, and a second encoding module; wherein,
[0255] The first encoding module includes a base mesh generation module, a base mesh intra-frame slice encoding module, and a base mesh inter-frame slice encoding module, and is specifically configured to generate a base mesh according to a current input mesh and a reference base mesh, and encode the base mesh intra-frame slices and the base mesh inter-frame slices in the base mesh to obtain a base mesh intra-frame slice bitstream and a base mesh inter-frame slice bitstream, and a reconstructed base mesh intra-frame slice and a reconstructed base mesh inter-frame slice;
[0256] The processing module includes a mesh merging module, and is specifically configured to merge the reconstructed base mesh intra-frame slice and the reconstructed base mesh inter-frame slice to obtain a reconstructed base mesh;
[0257] The second encoding module includes a displacement generation module and a displacement encoding module, and is specifically configured to perform subdivision processing on the reconstructed base mesh to obtain displacements of the refined meshes, and encode the displacements of the refined meshes to obtain a displacement bitstream.
[0258] In addition, it further includes a deformed mesh reconstruction module, a texture map conversion module, and a texture map encoding module, which are specifically configured to obtain a reconstructed deformed mesh according to the reconstructed displacement and the reconstructed base mesh, convert a texture map according to the reconstructed deformed mesh, the input texture map, and the current input mesh, and perform texture map encoding to obtain a texture map bitstream.
[0259] Based on the same inventive concept, please continue to refer to Figure 19, the present invention also provides a three-dimensional mesh decoding device, which is applied to a three-dimensional mesh decoding method provided in the above embodiments of the present invention. For the embodiments of the method, please refer to the above, and details will not be repeated here. The three-dimensional mesh decoding device includes:
[0260] A first decoding module, configured to decode the in-slice bitstream of the base mesh frame to obtain the in-slice of the base mesh frame, decode the inter-slice of the base mesh frame to obtain the inter-slice of the base mesh frame, decode the in-slice of the base mesh frame to obtain the reconstructed in-slice of the base mesh frame, and decode the inter-slice of the base mesh frame to obtain the reconstructed inter-slice of the base mesh frame;
[0261] A processing module, configured to merge the reconstructed in-slice of the base mesh frame and the reconstructed inter-slice of the base mesh frame to obtain the reconstructed base mesh;
[0262] A second decoding module, configured to decode the displacement bitstream to obtain the displacement of the subdivided mesh after the reconstructed base mesh is subdivided, and obtain the reconstructed mesh according to the displacement of the reconstructed base mesh and the subdivided mesh.
[0263] Specifically, a three-dimensional mesh decoding device provided in this embodiment is applied to the decoding end and includes a first decoding module, a processing module, and a second decoding module; among them,
[0264] The first decoding module includes an in-slice decoding module of the base mesh frame and an inter-slice decoding module of the base mesh frame, and is specifically configured to decode the in-slice bitstream of the base mesh frame to obtain the in-slice of the base mesh frame, decode the inter-slice of the base mesh frame to obtain the inter-slice of the base mesh frame, decode the in-slice of the base mesh frame to obtain the reconstructed in-slice of the base mesh frame, and decode the inter-slice of the base mesh frame to obtain the reconstructed inter-slice of the base mesh frame;
[0265] The processing module includes a mesh merging module, and is specifically configured to merge the reconstructed in-slice of the base mesh frame and the reconstructed inter-slice of the base mesh frame to obtain the reconstructed base mesh;
[0266] The second decoding module includes a displacement encoding module and a mesh merging module, and is specifically configured to decode the displacement bitstream to obtain the displacement of the subdivided mesh after the reconstructed base mesh is subdivided, and obtain the reconstructed mesh according to the displacement of the reconstructed base mesh and the subdivided mesh.
[0267] In addition, it further includes a texture map decoding module, which is specifically configured to decode the texture map bitstream to reconstruct the texture map.
[0268] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant are intended to cover non-exclusive inclusion, so that an article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the article or device comprising said element. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The orientation or positional relationship indicated by "above", "below", "left", "right", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation of the present invention.
[0269] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0270] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A three-dimensional grid encoding method, characterized in that Including: Dividing a three-dimensional mesh into in-frame basic mesh slices and inter-frame basic mesh slices, and performing encoding to obtain an in-frame basic mesh slice bitstream and an inter-frame basic mesh slice bitstream, and simultaneously obtaining reconstructed in-frame basic mesh slices and reconstructed inter-frame basic mesh slices; Merging the reconstructed in-frame basic mesh slices and the reconstructed inter-frame basic mesh slices to obtain a reconstructed basic mesh; Performing subdivision processing on the reconstructed basic mesh to obtain displacements of the subdivided meshes, and encoding the displacements of the subdivided meshes to obtain a displacement bitstream.
2. The three-dimensional grid encoding method according to claim 1, wherein Encoding the in-frame basic mesh slices using an intra-frame encoding mode to obtain the in-frame basic mesh slice bitstream; Encoding the inter-frame basic mesh slices using an inter-frame encoding mode to obtain the inter-frame basic mesh slice bitstream; Obtaining a basic mesh bitstream according to the in-frame basic mesh slice bitstream, the inter-frame basic mesh slice bitstream, and additional basic mesh information; wherein, the additional basic mesh information includes at least one of the number of in-frame basic mesh slices, the number of inter-frame basic mesh slices, the length of the in-frame basic mesh slice bitstream, and the length of the inter-frame basic mesh slice bitstream.
3. The three-dimensional grid encoding method according to claim 1, characterized in that The encoding of the in-frame basic mesh slices includes encoding at least one of connection relationships, vertex geometric coordinates, and texture coordinates, and the encoding of the inter-frame basic mesh slices includes encoding at least one of connection relationships, vertex geometric coordinates, and texture coordinates; wherein, For encoding the texture coordinates of the in-frame basic mesh slices, it includes encoding the overall offset and / or scaling of the texture coordinates of the in-frame basic mesh slices; For the inter-frame basic mesh slices, encoding connection relationships, vertex geometric coordinates, and texture coordinates according to information of a reference basic mesh; wherein, the reference basic mesh represents the basic mesh corresponding to the reconstructed basic mesh in the time domain, and the information of the reference basic mesh includes connection relationships, vertex geometric coordinates, and texture coordinates.
4. The three-dimensional grid encoding method according to claim 1, characterized in that The merging of the reconstructed in-frame basic mesh slices and the reconstructed inter-frame basic mesh slices to obtain a reconstructed basic mesh includes: Adjusting the geometric information of some vertices of the reconstructed in-frame basic mesh slices to make it the same as the geometric information of the corresponding some vertices of the reconstructed inter-frame basic mesh slices; Deleting the duplicate vertices and connection relationships between the reconstructed in-frame basic mesh slices and the reconstructed inter-frame basic mesh slices.
5. The three-dimensional grid encoding method according to claim 1, wherein The performing of subdivision processing on the reconstructed basic mesh to obtain displacements of the subdivided meshes, and encoding the displacements of the subdivided meshes to obtain a displacement bitstream includes: Using the displacement information of a reference basic mesh to predict a prediction residual of the displacement of the inter-frame basic mesh slice, and encoding the prediction residual of the displacement to implement the encoding of the displacement of the inter-frame basic mesh slice.
6. A three-dimensional grid decoding method, characterized in that Including: Decoding the in-frame basic mesh slice bitstream to obtain in-frame basic mesh slices, decoding the inter-frame basic mesh slices to obtain inter-frame basic mesh slices, decoding the in-frame basic mesh slices to obtain reconstructed in-frame basic mesh slices, and decoding the inter-frame basic mesh slices to obtain reconstructed inter-frame basic mesh slices; Merging the reconstructed in-frame basic mesh slices and the reconstructed inter-frame basic mesh slices to obtain a reconstructed basic mesh; Decode the displacement bitstream to obtain the displacement of the subdivided mesh after reconstructing the base mesh subdivision process. Obtain the reconstructed mesh according to the reconstructed base mesh and the displacement of the subdivided mesh.
7. The three-dimensional grid decoding method according to claim 6, characterized in that Decode the base mesh bitstream to obtain additional base mesh information. Determine the decoding method corresponding to the method of encoding the in-frame slices of the base mesh according to the additional base mesh information, and decode the in-frame slice bitstream of the base mesh. Determine the decoding method corresponding to the method of encoding the inter-frame slices of the base mesh, and decode the inter-frame slice bitstream of the base mesh; wherein, the additional base mesh information includes at least one of the number of in-frame slices of the base mesh, the number of inter-frame slices of the base mesh, the length of the in-frame slice bitstream of the base mesh, the length of the inter-frame slice bitstream of the base mesh, the encoding method of the in-frame slices of the base mesh in the current frame, and the encoding method of the inter-frame slices of the base mesh in the current frame.
8. The three-dimensional grid decoding method according to claim 7, characterized in that The decoding of the in-frame slices of the base mesh includes decoding at least one of the connection relationship, vertex geometric coordinates, and texture coordinates; the decoding of the inter-frame slices of the base mesh includes decoding at least one of the connection relationship, vertex geometric coordinates, and texture coordinates; wherein, For decoding the texture coordinates of the in-frame slices of the base mesh, determine the overall offset and / or scaling for encoding the texture coordinates of the in-frame slices of the base mesh according to the additional base mesh information, and decode the in-frame slice bitstream of the base mesh according to the overall offset and / or scaling.
9. The three-dimensional grid decoding method according to claim 6, wherein The merging of the reconstructed in-frame slices of the base mesh and the reconstructed inter-frame slices of the base mesh to obtain the reconstructed base mesh includes: Adjust the geometric information of some vertices of the reconstructed in-frame slices of the base mesh to be the same as the geometric information of some vertices of the reconstructed inter-frame slices of the base mesh; Delete the duplicate vertices and connection relationships between the reconstructed in-frame slices of the base mesh and the reconstructed inter-frame slices of the base mesh.
10. The 3D mesh decoding method according to claim 6, characterized in that The decoding of the displacement bitstream to obtain the displacement of the subdivided mesh after reconstructing the base mesh subdivision process, and decoding the displacement of the subdivided mesh to obtain the reconstructed deformed mesh includes: Decode the displacement of the subdivided mesh to obtain the prediction residual of the displacement of the inter-frame slices of the base mesh, and combine the displacement information of the reference base mesh to obtain the displacement of the inter-frame slices of the base mesh.
11. A three-dimensional grid encoding device, characterized in that, Applied to the encoding end, it includes: A first encoding module for dividing a three-dimensional mesh into in-frame slices and inter-frame slices of the base mesh, and encoding to obtain an in-frame slice bitstream and an inter-frame slice bitstream of the base mesh, and at the same time obtaining reconstructed in-frame slices and reconstructed inter-frame slices of the base mesh; A processing module for merging the reconstructed in-frame slices and the reconstructed inter-frame slices of the base mesh to obtain the reconstructed base mesh; A second encoding module for performing subdivision processing on the reconstructed base mesh to obtain the displacement of the subdivided mesh, and encoding the displacement of the subdivided mesh to obtain a displacement bitstream.
12. A three-dimensional grid decoding device, characterized in that, Applied to the decoding end, it includes: A first decoding module for decoding the in-frame slice bitstream of the base mesh to obtain the in-frame slices of the base mesh, decoding the inter-frame slices of the base mesh to obtain the inter-frame slices of the base mesh, decoding the in-frame slices of the base mesh to obtain the reconstructed in-frame slices of the base mesh, and decoding the inter-frame slices of the base mesh to obtain the reconstructed inter-frame slices of the base mesh; A processing module, configured to merge the in-slice of the reconstructed base mesh frame and the inter-slice of the reconstructed base mesh frame to obtain a reconstructed base mesh; A second decoding module, configured to decode the displacement bitstream to obtain the displacement of the subdivided mesh after the reconstructed base mesh is subdivided, and obtain a reconstructed mesh according to the reconstructed base mesh and the displacement of the subdivided mesh.
Citation Information
Cited By
Three-dimensional geometric model sequence compression method, device and equipment based on binary arithmetic coding and storage medium
CN120807668A