Encoding method, decoding method, encoding apparatus, decoding apparatus, encoder, decoder, code stream, device, and storage medium
By cleaning and subdividing the three-dimensional grid at the decoding end and the encoding end, the problems of low decoding efficiency and low reconstruction quality in the existing technology are solved, and more efficient 3-dimensional grid coding and better reconstruction effects are achieved.
Patent Information
- Application Number
- PCT/CN2024/072404
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2025-07-24
AI Technical Summary
The existing three-dimensional grid encoding technology has problems with low decoding efficiency and low quality of reconstruction grids at the decoding end. Especially in the process of connecting information encoding of the three-dimensional grid, it is impossible to ensure that the reconstruction basic grid corresponds to the subdivided deformation grid in the preprocessing stage, resulting in a decrease in encoding efficiency and reconstruction quality.
The reconstructed basic mesh is cleaned at the decoding end, and the duplicate vertices and degraded surfaces are removed, and then segmented to ensure that the grid used for subdivision at the decoding end is consistent with the encoding end, thereby improving the decoding efficiency; the basic mesh is cleaned at the encoding end and then segmented to remove duplicate points to ensure coding efficiency.
Through the grid cleaning steps at the decoding end and the encoding end, the decoding efficiency and reconstruction grid quality at the decoding end are ensured, and the overall efficiency and reconstruction quality of 3D grid coding are improved.
Smart Images

Figure CN2024072404_24072025_PF_FP_ABST
Abstract
Description
Coding and decoding method and device, codec, code stream, device, storage medium Technical Field
[0001] The embodiments of the present application relate to the field of trellis coding and decoding technology, and are related to but not limited to coding and decoding methods and apparatuses, codecs, bit streams, devices, and storage media. Background Art
[0002] Video-based dynamic mesh coding (VDMC) is a standard developed by the Moving Picture Experts Group (MPEG) for compressing 3D meshes. Its main idea is to compress 3D meshes by leveraging the existing Visual Volumetric Video-based Coding (V3C) standard. However, since the connectivity information of the 3D mesh also needs to be encoded, the encoding process for 3D meshes differs slightly from that of V3C. Therefore, the syntax, semantics, and decoding operations of the V3C standard decoder need to be extended to support the decoding and reconstruction of 3D meshes.
[0003] Summary of the Invention
[0004] The encoding and decoding method and apparatus, codec, bitstream, and storage medium provided in the embodiments of the present application can improve the decoding efficiency of the decoding end. The encoding and decoding method and apparatus, codec, bitstream, and storage medium provided in the embodiments of the present application are implemented as follows:
[0005] According to a first aspect of an embodiment of the present application, a decoding method is provided, which is applied to a decoder, and the method includes: decoding a code stream to determine a reconstructed base mesh; performing mesh cleaning on the reconstructed base mesh to determine a first base mesh; subdividing the first base mesh to determine a second base mesh; and determining a reconstructed mesh based on the second base mesh.
[0006] It can be understood that in the decoding method provided in the embodiment of the present application, before subdividing the reconstructed basic grid, the reconstructed basic grid is first cleaned up, and then the cleaned basic grid (i.e., the first basic grid) is subdivided; in this way, the grid used for subdivision at the decoding end is the same as the grid used for subdivision at the encoding end, so that the decoding end can perform correct decoding. If subdivision is not performed, errors will occur, which is beneficial to improving the decoding efficiency of the decoding end.
[0007] According to a second aspect of an embodiment of the present application, an encoding method is provided, which is applied to an encoder and includes: processing an original mesh of a current image to obtain a base mesh; wherein the number of vertices of the base mesh is less than the number of vertices of the original mesh; generating a code stream based on the base mesh; performing mesh cleaning on the base mesh to determine a first base mesh; subdividing the first base mesh to determine a second base mesh; and generating a code stream based on the second base mesh.
[0008] It can be understood that in the encoding method provided in the embodiment of the present application, before the basic grid is subdivided, the basic grid is first cleaned up, and then subdivided based on the cleaned basic grid (i.e., the first basic grid); in this way, the duplicate points of the basic grid generated by the encoding process are removed through grid cleaning, thereby ensuring that the grid subdivision can proceed normally, while improving the encoding efficiency of the grid.
[0009] According to a third aspect of an embodiment of the present application, a decoding device is provided, which is applied to a decoder, and the device includes: a decoding module, configured to decode a code stream and determine a reconstructed basic grid; a grid cleaning module, configured to perform grid cleaning on the reconstructed basic grid and determine a first basic grid; a subdivision module, configured to subdivide the first basic grid and determine a second basic grid; and a reconstruction module, configured to determine a reconstructed grid based on the second basic grid.
[0010] According to a fourth aspect of an embodiment of the present application, a decoder is provided, comprising a first memory and a first processor; wherein the first memory is used to store a computer program that can be run on the first processor; and the first processor is used to execute the decoding method as described in the embodiment of the present application when running the computer program.
[0011] According to a fifth aspect of an embodiment of the present application, there is provided an encoding device, applied to an encoder, comprising: a basic mesh generation module, configured to process an original mesh of a current image to obtain a basic mesh; wherein the number of vertices of the basic mesh is less than the number of vertices of the original mesh; a first encoding module, configured to generate a code stream based on the basic mesh; a mesh cleaning module, configured to perform mesh cleaning on the basic mesh to determine a first basic mesh; a subdivision module, configured to subdivide the first basic mesh to determine a second basic mesh; and a second encoding module, configured to generate a code stream based on the second basic mesh.
[0012] According to a sixth aspect of an embodiment of the present application, an encoder is provided, comprising a second memory and a second processor; wherein the second memory is used to store a computer program that can be run on the second processor; and the second processor is used to execute the encoding method as described in the embodiment of the present application when running the computer program.
[0013] According to a seventh aspect of the embodiments of the present application, a code stream is provided, where the code stream is obtained by the encoding method described in any embodiment of the present application.
[0014] According to an eighth aspect of the embodiments of the present application, an electronic device is provided, comprising: a processor suitable for executing a computer program; a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the encoding method as described in any one of the embodiments of the present application is implemented, or when the computer program is executed by the processor, the decoding method as described in any one of the embodiments of the present application is implemented.
[0015] According to the ninth aspect of the embodiments of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed, it implements the encoding method as described in any one of the embodiments of the present application, or implements the decoding method as described in any one of the embodiments of the present application.
[0016] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings herein are incorporated into and constitute a part of this specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, serve to illustrate the technical solutions of the present application. Obviously, the drawings described below are merely some embodiments of the present application. Those skilled in the art can, without inventive effort, derive other drawings from these drawings.
[0018] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0019] FIG1 is a schematic diagram of a three-dimensional grid image provided in an embodiment of the present application;
[0020] FIG2 is a partially enlarged schematic diagram of a three-dimensional grid image provided in an embodiment of the present application;
[0021] FIG3 is a schematic diagram of a connection method of a three-dimensional grid provided in an embodiment of the present application;
[0022] FIG4 is a second schematic diagram of a three-dimensional grid image provided in an embodiment of the present application;
[0023] FIG5 is a schematic diagram of a grid data storage format provided in an embodiment of the present application;
[0024] FIG6 is a schematic diagram of properties of a three-dimensional grid image provided in an embodiment of the present application;
[0025] FIG7 is a schematic diagram of an encoding end framework of a VDMC encoding and decoding framework;
[0026] FIG8 is a schematic diagram of a decoding end framework corresponding to FIG7 ;
[0027] FIG9 is a schematic diagram of an implementation flow of a decoding method provided in an embodiment of the present application;
[0028] FIG10 is a schematic diagram of an implementation flow of the encoding method provided in an embodiment of the present application;
[0029] FIG11 is a schematic diagram of an implementation flow of determining a base grid according to an embodiment of the present application;
[0030] FIG12 is a schematic diagram of an implementation flow of a method for generating a displacement code stream provided in an embodiment of the present application;
[0031] FIG13 is a schematic diagram of a three-dimensional grid coding framework based on subdivision deformation provided in an embodiment of the present application;
[0032] FIG14 is a schematic diagram of a 3D mesh decoding framework based on subdivision deformation provided by an embodiment of the present application;
[0033] Figure 15 is a simplified grid example diagram;
[0034] Figure 16 is a schematic diagram of the Draco coding framework;
[0035] Figure 17 is a schematic diagram of the five modes of Edgebreaker;
[0036] FIG18 is a schematic diagram of various parallelogram prediction methods;
[0037] FIG19 is a schematic diagram of a displacement calculation method;
[0038] Figure 20 is a schematic diagram of subdivision;
[0039] FIG21 is a schematic diagram of texture map conversion;
[0040] FIG22 is a schematic structural diagram of a decoding device provided in an embodiment of the present application;
[0041] FIG23 is a schematic diagram of the structure of an encoding device provided in an embodiment of the present application;
[0042] FIG24 is a schematic diagram of the structure of a decoder provided in an embodiment of the present application;
[0043] Figure 25 is a schematic diagram of the structure of the encoder provided in an embodiment of the present application. DETAILED DESCRIPTION
[0044] To make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the specific technical solutions of the present application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not intended to limit the scope of the present application.
[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0046] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0047] It should be pointed out that the terms "first\second\third" etc. involved in the embodiments of the present application do not represent a specific order for the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0048] The encoder and decoder frameworks and business scenarios described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. It will be appreciated by those skilled in the art that with the evolution of encoders and decoders and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0049] It should be noted that it is possible to decode and synthesize different data format bitstreams within the same video scene. These can include at least image format, point cloud format, and mesh format. In this way, real-time immersive media interaction services can be provided for multiple data formats (e.g., meshes, point clouds, images, etc.) from different sources.
[0050] In embodiments of the present application, the data format-based approach allows for independent processing at the bitstream level of the data format. This means that, similar to tiles or slices in video coding, different data formats in this scenario can be encoded independently, enabling independent encoding and decoding based on the data format.
[0051] Generally speaking, 3D animation content uses a keyframe-based representation method, that is, each frame is a static mesh. Static meshes at different times have the same topological structure and different geometric structures. However, the amount of data of 3D dynamic meshes represented based on keyframes is extremely large, so how to effectively store, transmit and draw them has become a problem faced by the development of 3D dynamic meshes. In addition, the spatial scalability of the mesh needs to be supported for different user terminals (computers, notebooks, portable devices, mobile phones); different mesh bandwidths (broadband, narrowband, wireless) need to support the quality scalability of the mesh. Therefore, 3D dynamic mesh compression is a very critical issue.
[0052] A 3D mesh is the surface of a 3D object composed of countless polygons in space. Polygons are composed of vertices and edges. Figure 1 is an example of a 3D mesh image provided in an embodiment of the present application, and Figure 2 is a partially enlarged schematic diagram of a 3D mesh image provided in an embodiment of the present application. Figures 1 and 2 show that the mesh surface is composed of closed polygons.
[0053] Two-dimensional images have information expressed at every pixel, and their distribution is regular, so there's no need to record their position information separately. However, the distribution of vertices in a mesh in three-dimensional space is random and irregular, and the way polygons are constructed requires additional regulations. Therefore, it's necessary to record the position of each vertex in space, as well as the connection information of each polygon, to fully represent a mesh image. Figure 3 is an example diagram of the connection method for a three-dimensional mesh provided in an embodiment of the present application. As shown in Figure 3, the same number of vertices and vertex positions, due to different connection methods, result in completely different surfaces.
[0054] Texture information and material information of a 3D mesh surface are usually represented using a 2D image, where UV coordinates are used to map the 2D image to the 3D surface.
[0055] Similar to two-dimensional images, each position in the acquisition process may have corresponding attribute information, usually RGB color values, which reflect the color of the object. For three-dimensional meshes, in addition to color, the attribute information corresponding to each vertex is also commonly the reflectance value, which reflects the surface material of the object.
[0056] Therefore, three-dimensional mesh data typically includes geometric coordinate information (x, y, z), geometric connection information, UV coordinates, and an attribute map. FIG4 is a second example of a three-dimensional mesh image provided in an embodiment of the present application, and FIG5 is a schematic diagram of a mesh data storage format provided in an embodiment of the present application. As shown in FIG5 , it includes three-dimensional geometric position information (such as v 248 167 22, etc.), UV coordinates (such as vt934 1867, etc.), and connection information (such as f1 / 1, etc.). FIG6 is a schematic diagram of the attributes of a three-dimensional mesh image provided in an embodiment of the present application.
[0057] Video-based dynamic mesh coding (VDMC) is a standard developed by the Moving Picture Experts Group (MPEG) for compressing three-dimensional meshes. Its main idea is to compress three-dimensional meshes by utilizing the existing Visual Volumetric Video-based Coding (V3C) standard. Since the connection information of the three-dimensional mesh also needs to be encoded, the encoding process of the three-dimensional mesh is slightly different from that of V3C. The syntax, semantics and decoding operations of the V3C standard decoding end need to be extended to support the decoding and reconstruction of the three-dimensional mesh. Figure 7 is a schematic diagram of the encoding end framework of a VDMC encoding and decoding framework, and Figure 8 is a schematic diagram of the decoding end framework corresponding to Figure 7.
[0058] As shown in FIG7 , for an input mesh (also referred to as an original mesh or initial mesh), a base mesh, a corresponding subdivided mesh, and a deformed mesh are first obtained through a base mesh generation module 701. The input mesh is first simplified by a simplification module within the base mesh generation module 701. New texture coordinates are then generated for the simplified mesh by a mesh parameterization module within the base mesh generation module 701. The parameterized mesh is then subdivided and deformed. This involves inserting new vertices into the mesh according to a specific subdivision method to obtain a subdivided mesh. The distances between the subdivided mesh vertices and the nearest neighboring points of the input mesh, known as displacements, are then calculated to obtain a corresponding deformed mesh (also known as a subdivided deformed mesh). The vertices of the deformed mesh are the vertices in the input mesh that are the nearest neighbors of the subdivided mesh vertices. Subsequently, the vertex positions of the parameterized mesh (i.e., the mesh before the subdivided deformation) are adjusted based on the displacement information. This adjusted mesh, referred to as the base mesh, is then fed into a base mesh encoding module (base mesh compression module 702) and compressed using a mesh encoder within the base mesh compression module 702 to obtain a base mesh bitstream. In the inter-frame mode, a motion vector may be generated for each vertex of the base mesh according to the reference frame, and the base mesh compression module 702 only needs to compress the motion vector.
[0059] As shown in FIG7 , after encoding / compression, the base mesh is reconstructed. The displacement is then calculated using the reconstructed base mesh and the deformed mesh obtained during the base mesh generation phase. Subsequently, the displacement information is input to a displacement encoding module 704, which transforms and quantizes the input displacement information. The processed displacements can then be encoded using a video encoder or entropy encoder to produce a displacement bitstream (i.e., a displacement video bitstream). Furthermore, the reconstructed displacement information and the reconstructed base mesh are input to a deformed mesh reconstruction module 705 to produce a reconstructed deformed mesh. Specifically, the reconstructed displacement information is applied to the subdivided base mesh (i.e., the subdivided mesh) to produce a reconstructed deformed mesh. This mesh, along with the original input mesh and its corresponding texture map, is then input to a corresponding texture map conversion module 706 to produce a texture map corresponding to the reconstructed mesh. This texture map is then encoded using a video encoder (i.e., the texture map compression module 707) to produce a texture map bitstream (i.e., a texture map video bitstream).
[0060] The overall framework of the decoding end is shown in Figure 8. For the received code stream, the decoding end first demultiplexes the various code streams to obtain the basic grid code stream, the displacement code stream and the texture map code stream respectively. For the basic grid code stream, the basic grid is obtained by decoding using the grid decoder corresponding to the encoding end (i.e., the basic grid decoding module 801), and the basic grid is obtained by the subdivision module 802 to obtain the subdivided grid. The displacement code stream and the texture map code stream are decoded by the video decoder. For the displacement code stream, after the video is decoded, the displacement needs to be extracted from the image through the displacement decoding module 803, and the displacement reconstruction module 804 performs dequantization, inverse transformation and other steps, and then applies it to the subdivided basic grid (i.e., subdivided grid) through the deformed grid reconstruction module 805 to obtain the deformed grid reconstructed by the decoding end. The texture map code stream is decoded by the texture map decoding module 806 to obtain the texture map corresponding to the reconstructed deformed grid. The subsequent application or rendering module processes the reconstructed deformed grid and the decoded texture map as input.
[0061] In related coding schemes, the base mesh generation module 701 in the pre-processing stage obtains a base mesh and its corresponding subdivided deformed mesh (i.e., deformed mesh). This subdivided deformed mesh calculates displacements with the subdivided mesh obtained based on the reconstructed base mesh. However, due to the substitutability of the base mesh encoder, there is no guarantee that the reconstructed base mesh corresponds to the subdivided deformed mesh in the pre-processing stage, meaning that the displacement calculations may be incorrect. In addition, due to the influence of codecs, the geometric vertices of the base mesh calculated in the subdivided deformed stage and the reconstructed base mesh vary, resulting in deviations in the calculated displacements, which, to a certain extent, reduces coding efficiency and the quality of the reconstructed mesh.
[0062] Based on the above analysis, the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0063] The present application provides a decoding method, which can be applied to a decoder. FIG9 is a schematic diagram of an implementation flow of the decoding method provided in the present application. As shown in FIG9 , the method includes the following steps 901 and 904:
[0064] Step 901: Decode the code stream and determine the reconstructed base mesh.
[0065] Step 902: performing mesh cleaning on the reconstructed basic mesh to determine a first basic mesh;
[0066] Step 903: Subdivide the first basic grid to determine a second basic grid;
[0067] Step 904: Determine a reconstructed grid based on the second basic grid.
[0068] In an embodiment of the present application, before subdividing the reconstructed basic grid, the reconstructed basic grid is first cleaned, and then the cleaned basic grid (i.e., the first basic grid) is subdivided; in this way, the grid used for subdivision at the decoding end is the same as the grid used for subdivision at the encoding end, so that the decoding end can perform correct decoding. If subdivision is not performed, errors will occur, which is beneficial to improving the decoding efficiency of the decoding end.
[0069] The following describes further optional implementations and related terms of each of the above steps.
[0070] In step 901, the bitstream is decoded to determine a reconstructed base mesh.
[0071] Here, the decoding module corresponding to the encoder is used to decode the bitstream to reconstruct the base grid. If the encoder uses the base grid encoder Draco to compress the base grid, the decoder uses the inverse process of the base grid encoder Draco to process the bitstream to reconstruct the base grid.
[0072] In step 902 , mesh cleaning is performed on the reconstructed basic mesh to determine a first basic mesh; the first basic mesh can also be understood as a cleaned basic mesh.
[0073] In some embodiments, the mesh cleaning of the reconstructed base mesh includes: removing duplicate vertices and / or degenerate faces of the reconstructed base mesh; wherein the duplicate vertices refer to vertices whose distance from the current vertex is less than or equal to a first threshold; and the degenerate faces refer to mesh faces whose area is less than or equal to a second threshold.
[0074] In some embodiments, the reconstructed base mesh includes geometric coordinates of vertices of the reconstructed base mesh and connection relationships of the vertices of the reconstructed base mesh.
[0075] In some embodiments, removing duplicate vertices of the reconstructed base mesh includes: removing duplicate vertices of the reconstructed base mesh according to geometric coordinates of the vertices of the reconstructed base mesh and a connection relationship between the vertices of the reconstructed base mesh, and determining a third base mesh.
[0076] Furthermore, in some embodiments, the removing duplicate vertices of the reconstructed base mesh according to the geometric coordinates of the vertices of the reconstructed base mesh and the connection relationship of the vertices of the reconstructed base mesh to determine the third base mesh includes: removing duplicate vertices in the reconstructed base mesh according to the geometric coordinates of the vertices of the reconstructed base mesh; and updating the connection relationship of the base mesh after removing duplicate vertices according to the connection relationship of the vertices of the reconstructed base mesh to obtain the third base mesh.
[0077] In some embodiments, removing the degenerate surfaces of the reconstructed base mesh includes removing the degenerate surfaces of the third base mesh, that is, first removing duplicate vertices from the reconstructed base mesh, and then removing the degenerate surfaces based on the removal.
[0078] In the embodiment of the present application, there is no limitation on the values of the first threshold and the second threshold, which can be any predefined values. For example, in some embodiments, the first threshold is equal to 0, and the duplicate vertex refers to a vertex with the same geometric coordinates as the current vertex.
[0079] Exemplarily, in some embodiments, the second threshold is equal to 0, that is, mesh faces with an area of 0 in the reconstructed basic mesh are removed.
[0080] In one possible implementation, duplicate vertices are removed, also known as geometric vertex deduplication. The input to this module is the geometric coordinates of the reconstructed base mesh and the corresponding connectivity relationships, namely decBaseMesh.pointInput and decBaseMesh.triangleInput, including the corresponding number of geometric vertices pointCount and the number of triangles triangleCount. The output is the deduplicated geometric coordinates and the updated connectivity relationships, namely decBaseMesh.pointOutput and decBaseMesh.triangleOutput. The geometric vertex deduplication process is as follows:
[0081] In one possible implementation, the module for removing degenerate surfaces from the reconstructed base mesh takes as input the mesh decBaseMesh after vertex deduplication, and outputs the mesh cleanMesh (i.e., the first base mesh) after degenerate surfaces are removed. The process is as follows:
[0082] In step 903, the first basic grid is subdivided to determine a second basic grid.
[0083] It should be noted that the second basic mesh can also be understood as a subdivided mesh.
[0084] For step 903, in a possible implementation, the same subdivision steps as those at the encoding end may be performed according to the number of subdivisions at the encoding end, and new vertices may be inserted on the edges of the cleaned basic mesh using the indicated interpolation method to obtain a subdivided mesh.
[0085] In step 904, a reconstructed grid is determined based on the second basic grid.
[0086] In some embodiments, the decoding method further comprises: decoding the code stream to determine reconstructed displacement information. The determining the reconstructed grid based on the second basic grid comprises: determining the reconstructed grid based on the second basic grid and the reconstructed displacement information.
[0087] In a possible implementation, the geometric coordinates of the vertices of the second basic mesh (ie, the subdivided mesh) are added to the corresponding displacements in the reconstructed displacement information to obtain a reconstructed mesh.
[0088] In the embodiment of the present application, the reconstructed mesh may also be understood as a reconstructed deformed mesh.
[0089] In some embodiments, the decoding method further includes: decoding the code stream to determine a reconstructed texture map.
[0090] In one possible implementation, after obtaining the reconstructed mesh and the reconstructed texture map, the decoder can input this information into a corresponding application (such as a video conferencing application or a game application that provides immersive media interactive services), and the application renders the image content based on the reconstructed mesh and the reconstructed texture map.
[0091] The embodiment of the present application provides an encoding method, which corresponds to the processing flow of the decoding end. The method can be applied to the encoder. FIG10 is a schematic diagram of the implementation flow of the encoding method provided in the embodiment of the present application. As shown in FIG10 , the method includes the following steps 1001 and 1005:
[0092] Step 1001: Process the original mesh of the current image to obtain a base mesh; wherein the number of vertices of the base mesh is smaller than the number of vertices of the original mesh;
[0093] Step 1002: Generate a bitstream based on the basic grid;
[0094] Step 1003: performing mesh cleaning on the basic mesh to determine a first basic mesh;
[0095] Step 1004: Subdivide the first basic grid to determine a second basic grid;
[0096] Step 1005: Generate a code stream according to the second basic grid.
[0097] It can be understood that in the encoding method provided in the embodiment of the present application, before the basic grid is subdivided, the basic grid is first cleaned up, and then subdivided based on the cleaned basic grid (i.e., the first basic grid); in this way, the duplicate points of the basic grid generated by the encoding process are removed through grid cleaning, thereby ensuring that the grid subdivision can proceed normally, while improving the encoding efficiency of the grid.
[0098] The following describes further optional implementations and related terms of each of the above steps.
[0099] In step 1001, the original mesh of the current image is processed to obtain a base mesh, wherein the number of vertices of the base mesh is smaller than the number of vertices of the original mesh.
[0100] In some embodiments, the original mesh may be simplified to obtain a simplified mesh; the simplified mesh may then be mesh parameterized to obtain a base mesh. In other embodiments, the base mesh may be a mesh obtained by simplifying the original mesh, quantizing it, dequantizing it, and then mesh parameterizing it. In other words, the processing of the original mesh sequentially includes mesh simplification and mesh parameterization, or the processing of the original mesh sequentially includes mesh simplification, quantization, dequantization, and mesh parameterization.
[0101] Specifically, in some embodiments, as shown in FIG11 , processing the original grid of the current image to obtain the basic grid includes the following steps 1101 to 1104:
[0102] Step 1101: Simplify the original mesh to obtain a simplified mesh.
[0103] Mesh simplification involves reducing the original mesh to a simplified mesh with a relatively small number of points and faces, while minimizing any significant changes to the shape. The detailed process is described in the description of the mesh simplification module in the example application below and will not be repeated here.
[0104] Step 1102: quantize the simplified grid to obtain a quantized grid;
[0105] Step 1103, performing inverse quantization on the quantized grid to obtain an inverse quantized grid;
[0106] Quantizing / dequantizing the simplified mesh can be understood as quantizing / dequantizing the geometric coordinates of the vertices of the simplified mesh. The dequantized mesh is used by the subsequent mesh parameterization module, thereby ensuring that the basic mesh input by the mesh parameterization module is the same as the geometric coordinates of the basic mesh reconstructed at the decoding end, which is beneficial to improving the parameterization quality of the basic mesh reconstructed at the decoding end.
[0107] In one possible implementation, the quantization process of vertex geometric coordinates is as follows: scale = ((1 < < QP) - 1.0) / ((1 < < bitDepth) - 1.0) qBaseV[i] = Clamp (Round (baseV[i] * scale), 0.0, (1 < < bitDepth) - 1.0)
[0108] Where QP is the preset quantization parameter, bitDepth is the bit depth of the input mesh geometric coordinates, qBaseV[i] represents the geometric coordinates of the i-th vertex of the base mesh, the clamp operator is shown below, and the Round function indicates rounding.
[0109] The process of inverse quantization of vertex geometric coordinates is as follows: iscale = 1.0 / scale rBaseV[i] = qBaseV[i]*iscale
[0110] Among them, rBaseV[i] represents the geometric coordinates of the i-th vertex of the mesh after dequantization.
[0111] Step 1104: Determine the basic grid according to the inverse quantized grid.
[0112] In some embodiments, determining the basic grid according to the inverse quantized grid includes: performing grid parameterization on the inverse quantized grid to obtain the basic grid.
[0113] It will be appreciated that mesh parameterization is used to generate corresponding texture coordinates from the dequantized mesh. In implementation, algorithms such as the Isochart algorithm or the orthogonal projection algorithm can be used to parameterize the dequantized mesh. For an introduction to these two algorithms, see the description of the mesh parameterization module in the exemplary application below and will not be repeated here.
[0114] In step 1002, a code stream is generated according to the basic grid.
[0115] In the embodiments of the present application, the base grid encoder used to generate the base grid bitstream is flexible and replaceable. The encoder only needs to specify the identifier of the base grid encoder to be used, so that the decoder can use the corresponding base grid decoder for decoding. For example, the encoder uses the base grid encoder Draco to process the base grid and generate the corresponding bitstream. For details about the base grid encoder Draco, please refer to the description of the base grid compression module in the exemplary application below and will not be repeated here.
[0116] In step 1003, mesh cleaning is performed on the basic mesh to determine a first basic mesh.
[0117] In some embodiments, the code stream of the base grid is decoded to determine a reconstructed base grid; and grid cleaning is performed on the reconstructed base grid to obtain a first base grid (ie, a cleaned base grid).
[0118] It can be understood that, in conjunction with step 1004, the first base mesh is used for mesh subdivision, and the reconstructed base mesh may have duplicate vertices and / or degenerate faces. Therefore, in order to ensure the normal progress of step 1004, the base mesh is cleaned before subdividing it, including removing duplicate vertices and / or degenerate faces.
[0119] In addition, in conjunction with step 1004, it differs from the encoding framework of the related art, such as that shown in FIG. 7 above, in that the embodiment of the present application is based on the reconstructed base grid, rather than the second base grid obtained from the base grid before encoding. That is, in the embodiment of the present application, the code stream of the base grid is decoded to determine the reconstructed base grid, and the reconstructed base grid is grid-cleaned to obtain the first base grid; then, the first base grid is subdivided to obtain the second base grid; this allows the base grid encoder to encode the base grid in a more flexible manner, avoiding the impact of the encoding of the base grid on the second base grid and the reconstructed grid, while also benefiting in improving the quality of the reconstructed grid obtained based on the reconstructed base grid.
[0120] It should be noted that, in the embodiment of the present application, the first base mesh can also be understood as the cleaned base mesh, the second base mesh can also be understood as the subdivided mesh, and the final reconstructed mesh can also be understood as the reconstructed deformed mesh.
[0121] In some embodiments, the mesh cleaning of the reconstructed base mesh includes: removing duplicate vertices and / or degenerate faces of the reconstructed base mesh; wherein the duplicate vertices refer to vertices whose distance from the current vertex is less than or equal to a first threshold; and the degenerate faces refer to mesh faces whose area is less than or equal to a second threshold.
[0122] In some embodiments, the reconstructed base mesh includes geometric coordinates of vertices of the reconstructed base mesh and connection relationships of the vertices of the reconstructed base mesh.
[0123] In some embodiments, removing duplicate vertices of the reconstructed base mesh includes: removing duplicate vertices of the reconstructed base mesh according to geometric coordinates of the vertices of the reconstructed base mesh and a connection relationship between the vertices of the reconstructed base mesh, and determining a third base mesh.
[0124] Furthermore, in some embodiments, the removing duplicate vertices of the reconstructed base mesh according to the geometric coordinates of the vertices of the reconstructed base mesh and the connection relationship of the vertices of the reconstructed base mesh to determine the third base mesh includes: removing duplicate vertices in the reconstructed base mesh according to the geometric coordinates of the vertices of the reconstructed base mesh; and updating the connection relationship of the base mesh after removing duplicate vertices according to the connection relationship of the vertices of the reconstructed base mesh to obtain the third base mesh.
[0125] In some embodiments, the first threshold is equal to 0, and the repeated vertex refers to a vertex with the same geometric coordinates as the current vertex.
[0126] In some embodiments, removing the degenerate surfaces of the reconstructed base mesh includes: removing the degenerate surfaces of the third base mesh.
[0127] In some embodiments, the second threshold is equal to 0.
[0128] It should be noted that the implementation of grid cleaning at the encoding end is the same as that at the decoding end, so please refer to the above introduction to grid cleaning at the decoding end, which will not be repeated here.
[0129] In step 1004, the cleaned basic grid is subdivided to determine a second basic grid.
[0130] In the embodiment of the present application, the second basic mesh may also be understood as a subdivided mesh.
[0131] For the implementation method of grid subdivision on the encoding side, please refer to the introduction of the subdivision and deformation module on the encoding side in the exemplary application below, which will not be repeated here.
[0132] In step 1005, a code stream is generated according to the second basic grid.
[0133] In some embodiments, as shown in FIG12 , step 1005 may be implemented through steps 1201 to 1203 as follows:
[0134] Step 1201: deform the second basic mesh according to the original mesh to obtain a deformed mesh.
[0135] In one possible implementation, for the second base mesh (i.e., the subdivided mesh), the nearest neighbor points of each vertex of the second base mesh on the original mesh (including points on the mesh surface of the original mesh) are found. These nearest neighbor points constitute the vertices of the deformed mesh, and their connection relationship is the same as that of the second base mesh.
[0136] In the embodiment of the present application, there is no limitation on the method for finding the nearest neighbor of each vertex of the second basic mesh on the original mesh, and the search can be accelerated by using a data structure such as kdTree.
[0137] Step 1202: Determine the displacement information between the deformed mesh and the second basic mesh.
[0138] In a possible implementation, the distance between the geometric coordinates of each vertex of the second basic mesh and the geometric coordinates of the corresponding nearest neighbor point in the deformed mesh is calculated, so as to obtain the displacement information between the deformed mesh and the second basic mesh.
[0139] Step 1203: Generate a code stream according to the displacement information.
[0140] The encoding method of the displacement information can be understood by referring to the introduction of the displacement encoding module on the encoding end in the exemplary application below, which will not be repeated here.
[0141] In some embodiments, the encoding method further includes: decoding the code stream to determine the reconstructed displacement information; determining the reconstructed grid based on the reconstructed basic grid and the reconstructed displacement information; performing texture map conversion on the original texture map based on the original grid and the reconstructed grid to obtain a converted texture map; and generating a code stream based on the converted texture map.
[0142] The encoding process of the above texture map can be understood by referring to the relevant introduction of the texture map conversion module and the texture map compression module in the following exemplary application, and will not be repeated here.
[0143] It should also be noted that the encoding method described above corresponds to the decoding method described above. Therefore, for any technical details not disclosed in the encoding method, please refer to the corresponding explanations and descriptions of the decoding method described above for understanding. To save space and avoid repetition, some technical details of the encoding method will not be explained or described here. In addition, for any technical details not disclosed in the encoding and decoding methods described above, please refer to the relevant explanations and descriptions of the exemplary applications provided below for understanding.
[0144] The following describes exemplary applications of the embodiments of the present application.
[0145] In an embodiment of the present application, a three-dimensional mesh encoding and decoding framework based on subdivision deformation is provided. Specifically, the encoding framework divides the input mesh (i.e., the original mesh) into a base mesh and a displacement. The base mesh has fewer points or faces than the input original mesh, and the displacement is used to improve the quality of the reconstructed base mesh so that it approaches the original mesh. In an embodiment of the present application, by subdividing and deforming the reconstructed base mesh, the base mesh encoder is allowed to encode the base mesh in a more flexible manner, avoiding the influence of the base mesh encoding on the subdivided mesh and the deformed mesh, while improving the quality of the mesh (i.e., the deformed mesh) after the reconstructed base mesh is deformed by subdivision. In addition, in the base mesh generation module, by simulating the reconstruction process (i.e., quantizing and dequantizing the simplified mesh before mesh parameterization), the input of the mesh parameterization is made the same as the reconstructed base mesh, and it can also ensure that the texture coordinates of the reconstructed base mesh correspond to its geometric coordinates, thereby improving the parameterization quality of the reconstructed mesh.
[0146] The present application provides an encoding framework. FIG13 is a schematic diagram of a three-dimensional mesh encoding framework based on subdivision deformation provided by the present application. As shown in FIG13 , the input mesh passes through a base mesh generation module 1301 to obtain a base mesh with fewer points and faces. The base mesh is then encoded using a base mesh compression module 1302 (i.e., a base mesh encoder) to obtain a base mesh code stream and a reconstructed base mesh. The reconstructed base mesh is cleaned by a mesh cleaning module 1303, and the cleaned mesh is subdivided and deformed by a subdivision deformation module 1304 to obtain a displacement corresponding to the reconstructed mesh. The displacement is then processed by a displacement encoding module 1305, including transformation and quantization. The processed displacement is then encoded and reconstructed. The deformed mesh reconstruction module 1306 uses the reconstructed displacement and the reconstructed base mesh to obtain a reconstructed deformed mesh. The texture map conversion module 1307 then performs texture map conversion based on the reconstructed deformed mesh to obtain a corresponding texture map. The texture map compression module 1308 encodes the texture map to obtain a texture map code stream.
[0147] The embodiment of the present application provides a decoding framework. Figure 14 is a schematic diagram of a three-dimensional grid decoding framework based on subdivision deformation provided by the embodiment of the present application, as shown in Figure 14. The basic grid code stream is decoded by the basic grid decoder 1401 corresponding to the encoding end. The grid cleaning module 1402 performs a grid cleaning step on the decoded basic grid, that is, removes duplicate points and degenerate surfaces (surfaces with an area of 0). The displacement code stream is decoded by the displacement decoding module 1403, and then the displacement is reconstructed by the displacement reconstruction module 1404, including steps such as inverse transformation and inverse quantization. The reconstructed displacement is then applied to the subdivided basic grid (i.e., the subdivided grid, where the subdivided grid is obtained by subdividing the grid output by the subdivision module 1406) through the deformed grid reconstruction module 1405 to obtain a reconstructed deformed grid, i.e., a decoded grid or a reconstructed grid. The texture map code stream obtains a reconstructed texture map after passing through the texture map decoding module 1407, i.e., a decoded texture map. Next, the main modules of the three-dimensional grid encoding and decoding framework provided by the embodiment of the present application are specifically described.
[0148] Encoding side:
[0149] (1) Basic grid generation module 1301:
[0150] The basic grid generation module 1301 takes the original grid as input and outputs the basic grid. The module 1301 includes a grid simplification module, an optional quantization and dequantization module, and a grid parameterization module. Each module is described below.
[0151] ①Grid simplification module
[0152] The mesh simplification module is used to simplify the current input mesh (i.e., the original mesh) into a base mesh with a relatively small number of points and faces, while minimizing any significant shape changes. The mesh simplification module focuses on the simplification operation and the corresponding error energy function. A possible mesh simplification operation is shown in Figure 15. This operation merges the vertices v1 and v2 at the ends of an edge into a single vertex v and deletes the connection between them. This process is repeated throughout the mesh according to a specific rule to reduce the number of faces and vertices to the target value.
[0153] During the simplification process, you can choose an error metric to optimize the simplified results. For example, the error metric for a vertex can be the sum of the coefficients of the equations of all adjacent faces. The error metric for an edge can be the sum of the error metrics of the two vertices on the edge. In other words, the error resulting from merging an edge is the sum of the squared distances from the merged vertex to all adjacent faces of the original two vertices on the edge. Furthermore, the error metric can also take into account the displacement between the base mesh and the subdivided mesh.
[0154] After determining the simplification operation and the corresponding error metric, the mesh simplification process begins iteratively. First, the vertex errors of the initial mesh are calculated to obtain the error for each edge. Edges are then sorted from smallest to largest error, and the edge with the smallest error is merged each time. Simultaneously, the positions of the merged vertices are calculated, and the errors of all edges associated with the merged vertices are updated. This means that the order of edge arrangement is updated to ensure that each iteration is based on a global error metric. Through iteration, the mesh faces are simplified to the number required for lossy encoding.
[0155] ②Quantization and dequantization module
[0156] This step / module is an optional module, which is used to quantize and dequantize the vertex geometric coordinates of the mesh output by the simplification module for use by the subsequent mesh parameterization module, thereby ensuring that the basic mesh input by the mesh parameterization module is the same as the basic mesh reconstructed at the decoding end, thereby improving the parameterization quality of the basic mesh reconstructed at the decoding end.
[0157] The vertex quantization process is as follows: scale = ((1<<QP)-1.0) / ((1<<bitDepth)-1.0) qBaseV[i] = Clamp(Round(baseV[i]*scale),0.0,(1<<bitDepth)-1.0)
[0158] Where QP is the preset quantization parameter, bitDepth is the bit depth of the input mesh geometric coordinates, qBaseV[i] represents the geometric coordinates of the i-th vertex of the base mesh, the clamp operator is shown below, and the Round function indicates rounding.
[0159] The vertex dequantization process is as follows: iscale = 1.0 / scale rBaseV[i] = qBaseV[i]*iscale
[0160] Among them, rBaseV[i] represents the geometric coordinates of the i-th vertex of the mesh after dequantization.
[0161] ③Grid parameterization module
[0162] The mesh parameterization module generates texture coordinates for the mesh and outputs them as the base mesh. Currently, many algorithms exist for mesh parameterization, such as the Isochart algorithm and the orthogonal projection algorithm. Within this encoding framework, both of these schemes can be used to parameterize the reconstructed base mesh. The following briefly describes both algorithms.
[0163] 1) Isochart algorithm
[0164] This algorithm uses spectral analysis to implement stretch-driven 3D mesh parameterization, UV-unwrapping the 3D mesh, tiling it, and packing it into a 2D texture domain. A stretch threshold is set, and the algorithm is outlined as follows: a) to f)
[0165] a) Compute surface spectrum analysis to provide an initial parameterization;
[0166] b) performing iterations of stretch optimization;
[0167] c) if the stretch of this derived parameterization is less than a threshold, stop;
[0168] d) performing surface spectral clustering to partition the surface into charts;
[0169] e) Use graph cut algorithm to optimize chart boundaries;
[0170] f) Iteratively split the charts until the stretching criteria are met.
[0171] 2) Orthogonal projection algorithm
[0172] OrthoAtlas is a projection-based mesh parameterization method that generates texture coordinates for a mesh through orthogonal projection. Its main steps include the following a) to h):
[0173] a) Calculate mesh properties, including the neighboring faces of each face and the area and normal vector of each face;
[0174] b) Determine the projection plane of each face based on the normal vector;
[0175] c) Start clustering all faces according to the projection plane to form a connected region, first selecting the starting face of the cluster;
[0176] d) Iterating from the starting face, determining whether adjacent faces of the face added to the connected region can be added to the connected region;
[0177] e) After each connected region is iterated, multiple connected regions are obtained;
[0178] f) determining whether to merge adjacent connected regions based on the error metric;
[0179] g) Check whether there are overlapping areas during projection, remove the overlapping areas and regenerate connected areas;
[0180] h) Arrange all the projected regions into a two-dimensional image.
[0181] (2) Basic grid compression module 1302:
[0182] The input of this module 1302 is the base grid, and its output is the base grid code stream and the reconstructed base grid. In this coding framework, the base grid compression module 1302 is flexible. The base grid encoder used by this module is interchangeable. The encoder only needs to specify the identifier of the base grid encoder to enable the decoder to use the corresponding base grid decoder for decoding. The following describes one possible base grid encoder, Draco.
[0183] Draco's main approach to static mesh compression is connectivity-driven mesh compression. It traverses all faces of a mesh in a specific manner, labels each face according to specific rules, and encodes the labels of all traversed faces, effectively encoding the mesh's connectivity. It then encodes all vertex coordinates and associated vertex attribute information in the order in which the connectivity relationships were traversed. The encoding framework is shown in Figure 16.
[0184] The main process of Draco mesh encoding involves: first, for the input mesh, a connection relationship is generated based on its geometric information, namely, the connection relationship between vertices in three-dimensional space. After constructing the face connection relationship, an initial face is selected to begin traversing all faces of the current mesh, i.e., generating symbols. This traversal and symbol generation uses the Edgebreaker algorithm, which divides the traversal into five modes based on the state of the triangle face at the current corner, as shown in Figure 17.
[0185] The five modes also define the direction of traversing the next face after traversing the current face. According to the above traversal method, corresponding symbols are generated for each face defined by the current mesh geometry information, and then these symbols are entropy encoded to obtain a code stream of the connection relationship defined by the current mesh geometry information. At the same time, traversing each face also obtains the order of traversing the corresponding vertices. The vertex order is passed to the geometry information encoder, which is rearranged in the order of traversal and quantized according to the predetermined quantization parameter. It is then predicted using a parallelogram prediction method, as shown in Figure 18. The prediction methods include single parallelogram prediction (Figure 18 (a)), multi-parallelogram prediction (Figure 18 (b)) and weighted parallelogram prediction (Figure 18 (c)). Finally, the predicted geometric coordinate residual is entropy encoded.
[0186] Attribute information, or texture coordinates, also has a connection relationship. When traversing this connection relationship, the difference between the attribute information and the geometric information is compared and the difference, also known as boundary information, is encoded. Similar to the encoding of geometric information, after determining the order of attribute information traversal, the attribute information is quantized, predicted, and entropy encoded in the same manner to produce the attribute information encoded bitstream. Finally, the various bitstream components are combined to produce the final encoded bitstream.
[0187] (3) Grid cleaning module 1303:
[0188] The mesh cleaning module 1303 takes the reconstructed base mesh as input and outputs the cleaned base mesh. This module 1303 is used to remove duplicate points in the base mesh generated by encoding and other processes to ensure that the subdivision deformation module 1304 can work properly. It includes two modules: geometric vertex deduplication and degenerate surface removal. For detailed description, see the same process on the decoding end.
[0189] (4) Subdivision deformation module 1304:
[0190] The basic idea behind the subdivision deformation module 1304 is shown in FIG19 , where the same concept is applied to the input 3D mesh to generate displacements. In FIG19 , the input 2D curve (represented by a 2D polyline), referred to as the "original" curve, is first downsampled to generate a base curve / polyline, referred to as the "simplified" curve. The subdivision scheme is then applied to the simplified polyline to generate the "subdivided" curve. The subdivided polyline is then deformed to obtain a better approximation of the original curve. This means that a displacement vector is calculated for each vertex of the subdivided mesh, so that the shape of the displaced curve is as close as possible to the shape of the original curve. These displacement vectors are the displacement information output by the module.
[0191] This step takes the parameterized mesh as input and first subdivides the input mesh. Any subdivision scheme can be chosen. One possible scheme is the midpoint subdivision scheme, which subdivides each triangle into four subtriangles in each subdivision iteration, as shown in Figure 20. A new vertex is introduced in the middle of each edge. The subdivision of geometric information and attribute information is performed independently because the connection relationship between geometric information and attribute information is usually different.
[0192] For the subdivided mesh (i.e., the subdivided mesh), the nearest neighbor of each point on the original input mesh (including points on the original mesh surface) can be found. Data structures such as kdTree can be used to accelerate the search. The displacement of the geometric coordinates of each vertex on the subdivided mesh is obtained by calculating the distance between the geometric coordinates of each vertex on the subdivided mesh and its nearest neighbor on the original input mesh. This module can obtain the base mesh and the corresponding subdivided deformed mesh (i.e., the deformed mesh).
[0193] (5) Displacement encoding module 1305:
[0194] The base mesh compression module 1302 compresses and reconstructs the base mesh, subdivides the reconstructed base mesh to obtain a subdivided mesh, deforms the subdivided mesh according to the original mesh to obtain a deformed mesh, and determines the displacement of corresponding vertices between the subdivided mesh and the deformed mesh.
[0195] Displacement encoding can be implemented in a variety of ways. One approach involves transforming the coordinate system of the displacements. Specifically, the coordinate system of each vertex's displacement is converted to a coordinate system constructed using the vertex's normal vector and two components tangent to that normal vector. The displacements are then transformed using techniques such as wavelet transforms. The transformed coefficients are quantized and arranged in the image in scan order, and video encoding is applied to the image. Alternatively, entropy coding can be used to directly encode the generated or processed displacements.
[0196] (6) Deformed Mesh Reconstruction Module 1306:
[0197] Because the displacement encoding stage causes some loss in the quantization process, the displacement must be reconstructed on the encoder side to ensure consistency with the decoder side. After obtaining the reconstructed displacement, the reconstructed base mesh is subdivided to obtain a subdivided mesh. This subdivided mesh is then deformed according to the corresponding displacement to obtain a reconstructed subdivided mesh (i.e., the reconstructed subdivided mesh).
[0198] (7) Texture map conversion module 1307 and texture map compression module 1308:
[0199] The texture map conversion module 1307 first performs texture map conversion based on the input original mesh, the input original texture map, and the mesh after subdivision and deformation (ie, the reconstructed deformed mesh), as shown in FIG21 .
[0200] The steps of texture conversion are as follows a) to d):
[0201] a) Calculate the texture coordinates of each pixel on the texture map to be generated. For example, the texture coordinates corresponding to pixel A(i, j) in Figure 21 are P(u, v);
[0202] b) Determine whether the texture coordinate is within a certain triangular face after the parameterization of the subdivided deformed mesh.
[0203] c) If the texture coordinate does not belong to any triangle, mark the pixel as an empty pixel and then fill it with a filling algorithm;
[0204] d) If the texture coordinate belongs to a triangle, then
[0205] ●Mark the pixel as filled;
[0206] Calculate the center of gravity of the texture in the current triangle based on the texture coordinates;
[0207] Based on the barycentric coordinates and the corresponding triangular face, the 2D texture coordinates are mapped to 3D geometric coordinates, i.e., to the points on the subdivided deformed mesh corresponding to the texture coordinates, as shown by M(x, y, z) in FIG21 ;
[0208] Find the point on the input original grid that is closest to the 3D coordinate, as shown by M'(x,y,z) in Figure 21;
[0209] Calculate the barycentric coordinates of the 3D coordinates based on the triangle face they are on and map them to 2D to calculate their texture coordinates, i.e. P'(u',v');
[0210] ●Sampling is performed on the input original texture map using the texture coordinates to obtain the value A'(i', j') of the corresponding pixel position (see FIG21 ).
[0211] ●Assign this value to the corresponding pixel A(i,j) on the texture map to be generated (see Figure 21).
[0212] For empty pixels, existing filling algorithms (such as Push-Pull algorithm) can be used to fill these empty pixels.
[0213] Decoding end:
[0214] (1) Basic grid decoding module 1401:
[0215] The basic grid decoding module 1401 selects a basic grid decoder corresponding to the encoding end to decode the basic grid code stream to obtain a reconstructed basic grid.
[0216] (2) Grid cleaning module 1402 and subdivision module 1406:
[0217] The input to the mesh cleaning module 1402 is the decoded base mesh, and the output is the cleaned base mesh. This step is used to remove duplicate vertices and degenerate faces from the decoded base mesh. Duplicate vertices are vertices with identical geometric coordinates, and degenerate faces are triangles with zero area. This module ensures that the mesh used for subdivision on the decoding side is identical to the input mesh of the subdivision deformation module 1304 on the encoding side. The mesh cleaning step does not process texture coordinates. The mesh cleaning process is described in detail below and is mainly divided into two parts: geometric vertex deduplication and degenerate face removal.
[0218] Deduplication of geometric vertices:
[0219] The input to the geometry deduplication module is the decoded base mesh's geometric coordinates and corresponding connectivity relationships, namely decBaseMesh.pointInput and decBaseMesh.triangleInput, including the corresponding number of geometric vertices (pointCount) and triangles (triangleCount). The output is the deduplicated geometric coordinates and updated connectivity relationships, namely decBaseMesh.pointOutput and decBaseMesh.triangleOutput. The vertex deduplication process is as follows:
[0220] Degenerate surface removal:
[0221] The input of the degenerate surface removal module is the mesh decBaseMesh after vertex deduplication, and the output is the mesh after degenerate surfaces are removed.
[0222] cleanMesh. The process is as follows:
[0223] After the mesh is cleaned, the subdivision module 1406 performs the same subdivision steps as the encoder (without deformation) according to the subdivision number of the encoder, and obtains the subdivided base mesh through the indicated interpolation method. The geometric vertices of this mesh are added with the displacement obtained by subsequent decoding to obtain the deformed mesh reconstructed by the decoder.
[0224] (3) Displacement decoding and reconstruction (i.e., displacement decoding module 1403 and displacement reconstruction module 1404):
[0225] The displacement code stream is decoded by a displacement decoder (i.e., displacement decoding module 1403). If the encoder compressed the displacement using video encoding, the decoder decodes it using the corresponding video decoder and restores it from the two-dimensional image in the corresponding order according to the arrangement scheme. The displacement reconstruction module 1404 then performs inverse transformation and dequantization on the information output by the displacement decoding module 1403 to restore the displacement consistent with the encoder, i.e., the reconstructed displacement information. If entropy encoding is used, it is directly entropy decoded, followed by subsequent reconstruction steps such as inverse transformation and dequantization.
[0226] (4) Deformed Mesh Reconstruction Module 1405:
[0227] The deformed mesh reconstruction module 1405 is used to sequentially add the reconstructed displacement to each vertex of the subdivided base mesh (ie, the subdivided mesh) to obtain a reconstructed deformed mesh, which is the mesh finally output by the decoding end (ie, the reconstructed mesh).
[0228] (5) Texture map decoding module 1407:
[0229] The texture image decoding module 1407 is responsible for decoding the texture image code stream. The texture image decoding module 1407 uses a video decoder to perform decoding, and optionally performs color space conversion to obtain an image format consistent with the texture image input by the encoding end, and obtains the final decoded output texture image.
[0230] It is understood that in the embodiments of the present application:
[0231] (1) A 3D mesh encoding and decoding framework based on subdivision deformation is proposed to perform subdivision deformation on the reconstructed base mesh;
[0232] (2) A quantization / dequantization module can be applied before the grid parameterization module so that the input of the grid parameterization module is the same as the reconstructed base grid.
[0233] (3) The basic mesh on the decoding side is cleaned before subdivision, including deduplication of geometric vertices and removal of degenerate surfaces.
[0234] It should be noted that although the steps of the method of the present application are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all steps must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps; or steps in different embodiments may be combined to form a new technical solution.
[0235] Based on the above embodiments, the present invention provides a decoding device, which is applied to a decoder. FIG22 is a schematic structural diagram of the decoding device provided by the present invention. As shown in FIG22 , the decoding device 220 includes:
[0236] A decoding module 2201 is configured to decode the code stream and determine a reconstructed base grid;
[0237] A grid cleaning module 2202 is configured to perform grid cleaning on the reconstructed basic grid to determine a first basic grid;
[0238] A subdivision module 2203 is configured to subdivide the first basic grid to determine a second basic grid;
[0239] The reconstruction module 2204 is configured to determine a reconstructed grid according to the second basic grid.
[0240] In some embodiments, the decoding module 2201 is further configured to decode the code stream to determine the reconstructed displacement information.
[0241] In some embodiments, determining the reconstructed grid according to the second basic grid includes: determining the reconstructed grid according to the second basic grid and the reconstructed displacement information.
[0242] In some embodiments, the decoding module 2201 is further configured to decode the code stream to determine a reconstructed texture map.
[0243] In some embodiments, the mesh cleaning of the reconstructed base mesh includes: removing duplicate vertices and / or degenerate faces of the reconstructed base mesh; wherein the duplicate vertices refer to vertices whose distance from the current vertex is less than or equal to a first threshold; and the degenerate faces refer to mesh faces whose area is less than or equal to a second threshold.
[0244] In some embodiments, the reconstructed base mesh includes geometric coordinates of vertices of the reconstructed base mesh and connection relationships of the vertices of the reconstructed base mesh.
[0245] In some embodiments, removing duplicate vertices of the reconstructed base mesh includes: removing duplicate vertices of the reconstructed base mesh according to geometric coordinates of the vertices of the reconstructed base mesh and a connection relationship between the vertices of the reconstructed base mesh, and determining a third base mesh.
[0246] Furthermore, in some embodiments, the removing duplicate vertices of the reconstructed base mesh according to the geometric coordinates of the vertices of the reconstructed base mesh and the connection relationship of the vertices of the reconstructed base mesh to determine the third base mesh includes: removing duplicate vertices in the reconstructed base mesh according to the geometric coordinates of the vertices of the reconstructed base mesh; and updating the connection relationship of the base mesh after removing duplicate vertices according to the connection relationship of the vertices of the reconstructed base mesh to obtain the third base mesh.
[0247] In some embodiments, the first threshold is equal to 0, and the repeated vertex refers to a vertex with the same geometric coordinates as the current vertex.
[0248] In some embodiments, removing the degenerate surfaces of the reconstructed base mesh includes: removing the degenerate surfaces of the third base mesh.
[0249] In some embodiments, the second threshold is equal to 0.
[0250] The present application provides an encoding device, which is applied to an encoder. FIG23 is a schematic diagram of the structure of the encoding device provided in the present application. As shown in FIG23 , the encoding device 230 includes:
[0251] A basic mesh generation module 2301 is configured to process the original mesh of the current image to obtain a basic mesh; wherein the number of vertices of the basic mesh is smaller than the number of vertices of the original mesh;
[0252] A first encoding module 2302 is configured to generate a code stream according to the basic grid;
[0253] A grid cleaning module 2303 is configured to perform grid cleaning on the basic grid to determine a first basic grid;
[0254] A subdivision module 2304 is configured to subdivide the first basic grid to determine a second basic grid;
[0255] The second encoding module 2305 is configured to generate a code stream according to the second basic grid.
[0256] In some embodiments, generating a code stream based on the second basic grid includes: deforming the second basic grid according to the original grid to obtain a deformed mesh; determining displacement information between the deformed mesh and the second basic grid; and generating a code stream based on the displacement information.
[0257] In some embodiments, performing grid cleaning on the basic grid to determine the first basic grid includes: decoding a code stream of the basic grid to determine a reconstructed basic grid; and performing grid cleaning on the reconstructed basic grid to determine the first basic grid.
[0258] In some embodiments, the mesh cleaning of the reconstructed base mesh includes: removing duplicate vertices and / or degenerate faces of the reconstructed base mesh; wherein the duplicate vertices refer to vertices whose distance from the current vertex is less than or equal to a first threshold; and the degenerate faces refer to mesh faces whose area is less than or equal to a second threshold.
[0259] In some embodiments, the reconstructed base mesh includes geometric coordinates of vertices of the reconstructed base mesh and connection relationships of the vertices of the reconstructed base mesh.
[0260] In some embodiments, removing duplicate vertices of the reconstructed base mesh includes: removing duplicate vertices of the reconstructed base mesh according to geometric coordinates of the vertices of the reconstructed base mesh and a connection relationship between the vertices of the reconstructed base mesh, and determining a third base mesh.
[0261] Furthermore, in some embodiments, the removing duplicate vertices of the reconstructed base mesh according to the geometric coordinates of the vertices of the reconstructed base mesh and the connection relationship of the vertices of the reconstructed base mesh to determine the third base mesh includes: removing duplicate vertices in the reconstructed base mesh according to the geometric coordinates of the vertices of the reconstructed base mesh; and updating the connection relationship of the base mesh after removing duplicate vertices according to the connection relationship of the vertices of the reconstructed base mesh to obtain the third base mesh.
[0262] In some embodiments, the first threshold is equal to 0, and the repeated vertex refers to a vertex with the same geometric coordinates as the current vertex.
[0263] In some embodiments, removing the degenerate surfaces of the reconstructed base mesh includes: removing the degenerate surfaces of the third base mesh.
[0264] In some embodiments, the second threshold is equal to 0.
[0265] In some embodiments, the processing of the original grid of the current image to obtain the basic grid includes: simplifying the original grid to obtain a simplified grid; quantizing the simplified grid to obtain a quantized grid; dequantizing the quantized grid to obtain a dequantized grid; and determining the basic grid based on the dequantized grid.
[0266] In some embodiments, determining the basic grid according to the inverse quantized grid includes: performing grid parameterization on the inverse quantized grid to obtain the basic grid.
[0267] In some embodiments, the encoding device 230 also includes a deformed mesh reconstruction module, a texture map conversion module and a texture map compression module; wherein the second encoding module 2305 is further configured to decode the code stream and determine the reconstructed displacement information; the deformed mesh reconstruction module is configured to determine the reconstructed mesh based on the reconstructed basic mesh and the reconstructed displacement information; the texture map conversion module is configured to perform texture map conversion on the original texture map based on the original mesh and the reconstructed mesh to obtain a converted texture map; the texture map compression module is configured to generate a code stream based on the converted texture map.
[0268] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of this application, please refer to the description of the method embodiment of this application for understanding.
[0269] It should be noted that the division of modules in the device described in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation. In addition, the functional units in the various embodiments of the present application can be integrated into a processing unit, or they can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. It can also be implemented in the form of a combination of software and hardware.
[0270] It should be noted that, in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling an electronic device to execute all or part of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.
[0271] An embodiment of the present application provides a computer-readable storage medium storing a computer program. When the computer program is executed, it implements an encoding method on an encoder side or a decoding method on a decoder side.
[0272] An embodiment of the present application provides a decoder, as shown in FIG24 , the decoder 240 includes: a first communication interface 2401, a first memory 2402, and a first processor 2403; each component is coupled together via a first bus system 2404. It is understood that the first bus system 2404 is used to achieve connection and communication between these components. In addition to the data bus, the first bus system 2404 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in FIG24 , various buses are labeled as the first bus system 2404. Among them,
[0273] The first communication interface 2401 is used to receive and send signals when sending and receiving information with other external network elements;
[0274] A first memory 2402 is used to store computer programs that can be run on the first processor 2403;
[0275] The first processor 2403 is configured to execute the decoding method described in the embodiment of the present application when running the computer program.
[0276] It is understood that the first memory 2402 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The first memory 2402 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0277] The first processor 2403 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the first processor 2403. The above-mentioned first processor 2403 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the first memory 2402 , and the first processor 2403 reads the information in the first memory 2402 and completes the steps of the above method in combination with its hardware.
[0278] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0279] Optionally, as another embodiment, the first processor 2403 is further configured to execute any of the aforementioned decoding method embodiments when running the computer program.
[0280] The present application implements an encoder, as shown in FIG25 , the encoder 250 includes: a second communication interface 2501, a second memory 2502, and a second processor 2503; each component is coupled together via a second bus linker 2504. It is understood that the second bus linker 2504 is used to achieve connection and communication between these components. In addition to the data bus, the second bus linker 2504 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as the second bus linker 2504 in FIG25 . Among them,
[0281] The second communication interface 2501 is used to receive and send signals during the process of sending and receiving information between other external network elements;
[0282] The second memory 2502 is used to store computer programs that can be run on the second processor 2503;
[0283] The second processor 2503 is configured to, when running the computer program, execute:
[0284] Sorting the samples in the first image according to differences between the samples to obtain a first list;
[0285] A sample value of the first component having the same position coordinates as the sample in the second image is obtained according to the index number of the sample in the first list and the bit depth of the first component.
[0286] Optionally, as another embodiment, the second processor 2503 is further configured to execute the aforementioned encoding method embodiment when running the computer program.
[0287] It can be understood that the hardware functions of the second memory 2502 and the first memory 2402 are similar, and the hardware functions of the second processor 2503 and the first processor 2403 are similar; they will not be described in detail here.
[0288] An embodiment of the present application provides an electronic device, comprising: a processor adapted to execute a computer program; and a computer-readable storage medium storing the computer program, wherein the computer program, when executed by the processor, implements the encoding method and / or decoding method described in the embodiment of the present application. The electronic device can be any type of device capable of trellis encoding and / or trellis decoding, such as a mobile phone, tablet computer, laptop computer, personal computer, television, projection device, or monitoring device.
[0289] The embodiment of the present application further provides a code stream, which is generated by the encoding method described in the above embodiment.
[0290] An embodiment of the present application also provides a computer program product, including computer program instructions.
[0291] Optionally, the computer program product can be applied to the encoder or decoder in the embodiments of the present application, and the computer program instructions enable the computer to execute the corresponding processes implemented by the encoder or decoder in the various methods of the embodiments of the present application. For the sake of brevity, they are not repeated here.
[0292] The embodiment of the present application also provides a computer program.
[0293] Optionally, the computer program can be applied to the encoder or decoder in the embodiments of the present application. When the computer program runs on a computer, the computer executes the corresponding processes implemented by the encoder or decoder in the various methods of the embodiments of the present application. For the sake of brevity, they are not repeated here.
[0294] Those skilled in the art will appreciate that the units and algorithmic steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0295] It should be noted that the descriptions of the above encoder, decoder, bitstream, computer program product, computer program, storage medium, and device embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, storage medium, and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0296] It should be understood that "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments. The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other. For the sake of brevity, they will not be repeated here.
[0297] The term "and / or" in this article is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, object A and / or object B can mean: object A exists alone, object A and object B exist at the same time, and object B exists alone.
[0298] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0299] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical or other forms.
[0300] The modules described above as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules; they may be located in one place or distributed over multiple grid units; some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.
[0301] In addition, all functional modules in the embodiments of the present application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the above-mentioned integrated modules can be implemented in the form of hardware or in the form of hardware plus software functional units.
[0302] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.
[0303] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks or optical disks.
[0304] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined, if they do not conflict, to obtain new method embodiments. The features disclosed in the several product embodiments provided in this application can be arbitrarily combined, if they do not conflict, to obtain new product embodiments. The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined, if they do not conflict, to obtain new method embodiments or device embodiments.
[0305] The above is merely an embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A decoding method, which is applied to a decoder, and the method includes: Decoding a bitstream to determine a reconstructed base mesh; Performing mesh cleaning on the reconstructed base mesh to determine a first base mesh; Subdividing the first base mesh to determine a second base mesh; Determining a reconstructed mesh according to the second base mesh.
2. The method according to claim 1, wherein The method further includes: Decoding the bitstream to determine reconstructed displacement information.
3. The method according to claim 2, wherein The determining the reconstructed mesh according to the second base mesh includes: Determining the reconstructed mesh according to the second base mesh and the reconstructed displacement information.
4. The method according to claim 1, wherein The method further includes: Decoding the bitstream to determine a reconstructed texture map.
5. The method according to any one of claims 1-4, wherein, The performing mesh cleaning on the reconstructed base mesh includes: Removing duplicate vertices and / or degenerate faces of the reconstructed base mesh; wherein, the duplicate vertices refer to vertices whose distance from the current vertex is less than or equal to a first threshold; the degenerate faces refer to mesh faces whose area is less than or equal to a second threshold.
6. The method according to claim 5, wherein, The reconstructed base mesh includes the geometric coordinates of the vertices of the reconstructed base mesh and the connection relationship of the vertices of the reconstructed base mesh.
7. The method according to claim 6, wherein, Removing the duplicate vertices of the reconstructed base mesh includes: Removing the duplicate vertices of the reconstructed base mesh according to the geometric coordinates of the vertices of the reconstructed base mesh and the connection relationship of the vertices of the reconstructed base mesh, and determining a third base mesh.
8. The method according to claim 7, wherein, Removing the duplicate vertices in the reconstructed base mesh according to the geometric coordinates of the vertices of the reconstructed base mesh; Updating the connection relationship of the base mesh after removing the duplicate vertices according to the connection relationship of the vertices of the reconstructed base mesh to obtain the third base mesh.
9. The method according to any one of claims 5 - 8, wherein, The first threshold is equal to 0, and the duplicate vertices refer to vertices having the same geometric coordinates as the current vertex.
10. The method according to any one of claims 8, wherein, Removing the degenerate faces of the reconstructed base mesh includes: Removing the degenerate faces from the third base mesh.
11. The method according to any one of claims 5 - 10, wherein, The second threshold is equal to 0.
12. An encoding method, which is applied to an encoder, and the method includes: Processing the original mesh of the current image to obtain a base mesh; wherein, the number of vertices of the base mesh is less than the number of vertices of the original mesh; Generating a bitstream according to the base mesh; Performing mesh cleaning on the base mesh to determine a first base mesh; Subdividing the first base mesh to determine a second base mesh; Generating a bitstream according to the second base mesh.
13. The method according to claim 12, wherein, The generating a bitstream according to the second base mesh includes: Deforming the second base mesh according to the original mesh to obtain a deformed mesh; Determining the displacement information between the deformed mesh and the second base mesh; Generating a bitstream according to the displacement information.
14. The method according to claim 12, wherein, Performing mesh cleaning on the base mesh to determine a first base mesh includes: Decoding the bitstream of the base mesh to determine a reconstructed base mesh; Performing mesh cleaning on the reconstructed base mesh to determine a first base mesh.
15. The method according to claim 12, wherein, The mesh cleaning of the reconstructed basic mesh includes: Removing duplicate vertices and / or degenerate faces of the reconstructed basic mesh; wherein, the duplicate vertices refer to the vertices whose distance from the current vertex is less than or equal to the first threshold; the degenerate faces refer to the mesh faces whose area is less than or equal to the second threshold.
16. The method according to claim 15, wherein, The reconstructed basic mesh includes the geometric coordinates of the vertices of the reconstructed basic mesh and the connection relationship of the vertices of the reconstructed basic mesh.
17. The method according to claim 16, wherein Removing duplicate vertices of the reconstructed basic mesh includes: According to the geometric coordinates of the vertices of the reconstructed basic mesh and the connection relationship of the vertices of the reconstructed basic mesh, removing duplicate vertices of the reconstructed basic mesh to determine the third basic mesh.
18. The method according to claim 17, wherein, Removing duplicate vertices in the reconstructed basic mesh according to the geometric coordinates of the vertices of the reconstructed basic mesh; According to the connection relationship of the vertices of the reconstructed basic mesh, updating the connection relationship of the basic mesh after removing duplicate vertices to obtain the third basic mesh.
19. The method according to any one of claims 15 - 18, wherein, The first threshold is equal to 0, and the duplicate vertices refer to the vertices with the same geometric coordinates as the current vertex.
20. The method according to claim 18, wherein, Removing degenerate faces of the reconstructed basic mesh includes: Removing degenerate faces from the third basic mesh.
21. The method according to any one of claims 15-20, wherein, The second threshold is equal to 0.
22. The method according to claim 12, wherein The processing of the original mesh of the current image to obtain the basic mesh includes: Performing mesh simplification on the original mesh to obtain a simplified mesh; Quantizing the simplified mesh to obtain a quantized mesh; Performing inverse quantization on the quantized mesh to obtain an inverse quantized mesh; Determining the basic mesh according to the inverse quantized mesh.
23. The method according to claim 22, wherein The determining the basic mesh according to the inverse quantized mesh includes: Performing mesh parameterization on the inverse quantized mesh to obtain the basic mesh.
24. The method according to claim 14, wherein The method further includes: Decoding the bitstream to determine the reconstructed displacement information; Determining the reconstructed mesh according to the reconstructed basic mesh and the reconstructed displacement information; Performing texture map conversion on the original texture map according to the original mesh and the reconstructed mesh to obtain a converted texture map; Generating a bitstream according to the converted texture map.
25. A decoding device, applied to a decoder, the device includes: A decoding module configured to decode the bitstream to determine the reconstructed basic mesh; A mesh cleaning module configured to perform mesh cleaning on the reconstructed basic mesh to determine the first basic mesh; A subdivision module configured to subdivide the first basic mesh to determine the second basic mesh; A reconstruction module configured to determine the reconstructed mesh according to the second basic mesh.
26. A decoder, including a first memory and a first processor; wherein, The first memory is used to store a computer program that can run on the first processor; The first processor is configured to execute the method according to any one of claims 1 to 11 when running the computer program.
27. An encoding device, applied to an encoder, the device includes: A basic grid generation module, configured to process the original grid of the current image to obtain a basic grid; wherein, the number of vertices of the basic grid is less than the number of vertices of the original grid; A first encoding module, configured to generate a bitstream according to the basic grid; A grid cleaning module, configured to clean the basic grid to determine a first basic grid; A subdivision module, configured to subdivide the first basic grid to determine a second basic grid; A second encoding module, configured to generate a bitstream according to the second basic grid.
28. An encoder, comprising a second memory and a second processor; wherein, The second memory is used to store a computer program that can run on the second processor; The second processor is configured to execute the method according to any one of claims 12 to 24 when running the computer program.
29. A bitstream, which is obtained by the encoding method according to any one of claims 12 to 24.
30. An electronic device, comprising: A processor, adapted to execute a computer program; A computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by the processor, it implements the method according to any one of claims 1 to 11, or when the computer program is executed by the processor It implements the method according to any one of claims 12 to 24.
31. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program, and when the computer program is executed, it implements the method according to any one of claims 1 to 11, or implements the method according to any one of claims 12 to 24.
Citation Information
Patent Citations
Apparatus and method for displaced mesh compression
CN113470179A
Apparatus and method for dynamic grid coding
CN117044209A
Remeshing for efficient compression
WO2023172457A1
V-mesh bitstream structure including syntax elements and decoding process with reconstruction
WO2023172509A1
Cited By
Geological cavity model reconstruction method based on grid restoration and UV reordering
CN122134965A