Grid encoding method, grid decoding method and related equipment
By subdividing the three-dimensional grid and encoding the displacement information of some levels, the problem of low coding efficiency is solved, and the amount of data is reduced and the processing is simplified.
Patent Information
- Application Number
- CN202310743475.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-21
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2043-06-21
AI Technical Summary
In existing technologies, the encoding efficiency of three-dimensional grids is low, resulting in large data volumes and complex processing, visualization, and transmission.
By obtaining a basic grid for subdivision processing, the number of levels of displacement information to be encoded is selected to be less than or equal to the number of levels of the subdivided grid, and the information is encoded to reduce the amount of encoded data.
It improves coding efficiency, reduces data volume, and simplifies processing and transmission processes.
Smart Images

Figure CN119182908B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of computer technology, and specifically relates to a grid encoding method, a grid decoding method and related equipment. Background Art
[0002] With the rapid development of multimedia technology, 3D models have become a new generation of digital media, following audio, images, and video. Three-dimensional meshes and point clouds are two commonly used representations of 3D models. Compared to traditional multimedia such as images and videos, 3D mesh models offer greater interactivity and realism, and thus have a wide range of applications.
[0003] In related art, when encoding displacement information of subdivided grids obtained based on a three-dimensional grid at an encoding end, all displacement information of the subdivided grids is encoded, which requires encoding a large amount of data and has low encoding efficiency. Summary of the Invention
[0004] The embodiments of the present application provide a grid encoding method, a grid decoding method and related equipment, which can solve the problem of low encoding efficiency.
[0005] In a first aspect, a grid coding method is provided, comprising:
[0006] Obtaining a base grid corresponding to the three-dimensional grid, and performing subdivision processing based on the base grid to obtain a subdivided grid;
[0007] Performing a deformation operation based on the subdivided grid to obtain a deformed grid;
[0008] Selecting displacement information to be encoded from the displacement information corresponding to the subdivided grid, where the displacement information corresponding to the subdivided grid is used to represent the displacement of vertices of the subdivided grid relative to the deformed grid, and the number of levels of the displacement information to be encoded is less than or equal to the number of levels of the displacement information corresponding to the subdivided grid;
[0009] The displacement information to be encoded is encoded to obtain a first encoding result.
[0010] In a second aspect, a grid decoding method is provided, comprising:
[0011] Decoding a first encoding result in a code stream corresponding to the three-dimensional grid to obtain decoded displacement information;
[0012] Decoding the third encoding result in the code stream to obtain a basic grid, and subdividing the basic grid to obtain a subdivided grid;
[0013] Performing padding processing on the decoded displacement information to obtain padded displacement information, wherein the number of levels of the padded displacement information is greater than or equal to the number of levels of the decoded displacement information;
[0014] Reconstruction processing is performed based on the filled displacement information and the subdivided grid to obtain a reconstructed grid.
[0015] In a third aspect, a grid coding device is provided, comprising:
[0016] An acquisition module is used to acquire a basic grid corresponding to the three-dimensional grid, and perform subdivision processing based on the basic grid to obtain a subdivided grid;
[0017] A deformation module, configured to perform a deformation operation based on the subdivided grid to obtain a deformed grid;
[0018] a selection module configured to select displacement information to be encoded from the displacement information corresponding to the subdivided grid, wherein the displacement information corresponding to the subdivided grid is used to represent the displacement of vertices of the subdivided grid relative to the deformed grid, and the number of levels of the displacement information to be encoded is less than or equal to the number of levels of the displacement information corresponding to the subdivided grid;
[0019] The encoding module is used to encode the displacement information to be encoded to obtain a first encoding result.
[0020] In a fourth aspect, a grid decoding device is provided, comprising:
[0021] A first decoding module is used to decode a first encoding result in a code stream corresponding to the three-dimensional grid to obtain decoded displacement information;
[0022] a second decoding module, configured to decode the third encoding result in the code stream to obtain a basic grid, and subdivide the basic grid to obtain a subdivided grid;
[0023] a filling module, configured to perform filling processing on the decoded displacement information to obtain filled displacement information, wherein the number of levels of the filled displacement information is greater than or equal to the number of levels of the decoded displacement information;
[0024] The reconstruction module is used to perform reconstruction processing based on the filled displacement information and the subdivided grid to obtain a reconstructed grid.
[0025] In a fifth aspect, a terminal is provided, which includes a processor, a memory, and a program or instruction stored in the memory and runnable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect; or, when executed by the processor, the program or instruction, when executed by the processor, implements the steps of the method described in the second aspect.
[0026] In a sixth aspect, a terminal is provided, comprising a processor and a communication interface, wherein the processor is used to: obtain a basic grid corresponding to a three-dimensional grid, perform subdivision processing based on the basic grid to obtain a subdivided grid; perform deformation operations based on the subdivided grid to obtain a deformed grid; select displacement information to be encoded from the displacement information corresponding to the subdivided grid, the displacement information corresponding to the subdivided grid is used to characterize the displacement of the vertices of the subdivided grid relative to the deformed grid, and the number of levels of the displacement information to be encoded is less than or equal to the number of levels of the displacement information corresponding to the subdivided grid; encode the displacement information to be encoded to obtain a first encoding result.
[0027] In a seventh aspect, a terminal is provided, comprising a processor and a communication interface, wherein the processor is configured to: decode a first coding result in a code stream corresponding to a three-dimensional grid to obtain decoded displacement information; decode a third coding result in the code stream to obtain a basic grid, and subdivide the basic grid to obtain a subdivided grid; perform filling processing on the decoded displacement information to obtain filled displacement information, wherein the number of levels of the filled displacement information is greater than or equal to the number of levels of the decoded displacement information; and perform reconstruction processing based on the filled displacement information and the subdivided grid to obtain a reconstructed grid.
[0028] In an eighth aspect, a readable storage medium is provided, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the grid encoding method as described in the first aspect are implemented, or when the program or instruction is executed by a processor, the steps of the grid decoding method as described in the second aspect are implemented.
[0029] In the ninth aspect, a chip is provided, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the method described in the first aspect, or to implement the steps of the method described in the second aspect.
[0030] In the tenth aspect, a computer program / program product is provided, which is stored in a non-volatile storage medium, and the program / program product is executed by at least one processor to implement the steps of the method described in the first aspect, or to implement the steps of the method described in the second aspect.
[0031] In an embodiment of the present application, a base mesh corresponding to a three-dimensional mesh is obtained, and subdivision processing is performed based on the base mesh to obtain a subdivided mesh. A deformation operation is performed on the subdivided mesh to obtain a deformed mesh. Displacement information to be encoded is selected from the displacement information corresponding to the subdivided mesh, where the displacement information corresponding to the subdivided mesh is used to represent the displacement of vertices of the subdivided mesh relative to the deformed mesh, and the number of levels of the displacement information to be encoded is less than or equal to the number of levels of the displacement information corresponding to the subdivided mesh. The displacement information to be encoded is encoded to obtain a first encoding result. In this way, by encoding only the displacement information of some levels of the displacement information corresponding to the subdivided mesh, the amount of encoded data can be reduced, thereby improving encoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a schematic diagram of V3C encoding in related technology;
[0033] Figure 2 This is a schematic diagram of V3C decoding in the related art;
[0034] Figure 3 Schematic diagram of V-DMC coding in related art;
[0035] Figure 4 Schematic diagram of V-DMC decoding in related art;
[0036] Figure 5 This is a flowchart of a grid coding method provided by an embodiment of the present application;
[0037] Figure 6 This is one of the displacement coding schematic diagrams provided in the embodiments of the present application;
[0038] Figure 7 This is one of the displacement decoding schematic diagrams provided in the embodiment of the present application;
[0039] Figure 8 This is a flowchart of a grid decoding method provided by an embodiment of the present application;
[0040] Figure 9 This is the second displacement coding diagram provided in the embodiment of the present application;
[0041] Figure 10 This is the second displacement decoding schematic diagram provided in the embodiment of the present application;
[0042] Figure 11 This is the third schematic diagram of a displacement decoding provided in an embodiment of the present application;
[0043] Figure 12 This is the third displacement coding diagram provided in the embodiment of the present application;
[0044] Figure 13 This is the fourth displacement decoding diagram provided in the embodiment of the present application;
[0045] Figure 14 This is a schematic diagram of a grid coding process provided by an embodiment of the present application;
[0046] Figure 15 This is a simplified grid diagram provided in an embodiment of the present application;
[0047] Figure 16 This is an example diagram of displacement calculation provided by an embodiment of the present application;
[0048] Figure 17 This is a basic grid compression schematic diagram provided in an embodiment of the present application;
[0049] Figure 18 This is a schematic diagram of attribute graph conversion provided by an embodiment of the present application;
[0050] Figure 19 This is a schematic diagram of a grid decoding process provided by an embodiment of the present application;
[0051] Figure 20 This is a schematic structural diagram of a grid coding device provided in an embodiment of the present application;
[0052] Figure 21 This is a schematic structural diagram of a grid decoding device provided in an embodiment of the present application;
[0053] Figure 22 This is a schematic diagram of the structure of a communication device provided in an embodiment of the present application;
[0054] Figure 23 This is a schematic diagram of the structure of a terminal provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] The following will be combined with the accompanying drawings in the embodiments of this application to clearly describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0056] The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same type, and do not limit the number of objects, for example, the first object can be one or more. In addition, "or" in this application represents at least one of the connected objects. For example, "A or B" covers three options, namely, Option 1: including A but not including B; Option 2: including B but not including A; Option 3: including both A and B. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.
[0057] The term "indication" in this application can be either a direct indication (or explicit indication) or an indirect indication (or implicit indication). A direct indication can be understood as the sender explicitly informing the receiver of specific information, the operation to be performed, or the requested result, etc. in the instruction sent; an indirect indication can be understood as the receiver determining the corresponding information based on the instruction sent by the sender, or making a judgment and determining the operation to be performed or the requested result, etc. based on the judgment result.
[0058] The encoding and decoding end corresponding to the encoding and decoding method in the embodiment of the present application can be a terminal, which can also be called a terminal device or user equipment (UE). The terminal can be a mobile phone, a tablet computer (Tablet Personal Computer), a laptop computer (Laptop Computer) or a notebook computer, a personal digital assistant (PDA), a handheld computer, a netbook, an ultra-mobile personal computer (UMPC), a mobile Internet device (Mobile Internet Device, MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device (Wearable Device) or a vehicle-mounted device (VUE), a pedestrian terminal (Pedestrian User Equipment, PUE) and other terminal-side devices. Wearable devices include: smart watches, bracelets, headphones, glasses, etc. It should be noted that the specific type of the terminal is not limited in the embodiment of the present application.
[0059] For ease of understanding, some of the contents involved in the embodiments of this application are described below:
[0060] 1. 3D Grid
[0061] In recent years, with the rapid development of multimedia technology, related research results have been rapidly industrialized and have become an indispensable part of people's lives. Three-dimensional models have become the next generation of digital media, following audio, images, and video. Three-dimensional meshes are a commonly used representation method for 3D models. Compared to traditional multimedia such as images and videos, 3D mesh models offer greater interactivity and realism, leading to their increasing application in various fields, including commerce, manufacturing, construction, education, medicine, entertainment, art, and the military.
[0062] While there are many ways to represent 3D meshes, triangular meshes remain the most common. A 3D mesh can be considered to be composed of three basic elements: vertices, edges, and faces. Vertices are the most basic elements in a mesh, defining a position in 3D space. Edges are line segments connecting two vertices in the mesh. Faces can be considered polygons formed by closed paths of edges. For a triangular mesh, each face is a triangle.
[0063] The information contained in the mesh is usually divided into three categories: geometric information, connection information, and attribute information. Geometric information refers to the position of each vertex of the mesh in three-dimensional space. Connection information describes the association between the elements in the mesh, that is, the connection relationship between vertices. Attribute information is optional, and it can associate attributes with corresponding mesh elements (such as vertex color, normal vector, etc. can be associated with mesh vertices). Mesh parameterization can also be used to map the mesh from three-dimensional space to a two-dimensional plane area. This mapping relationship is usually described by a set of parameter coordinates, called UV coordinates or texture coordinates, which are associated with mesh vertices. This two-dimensional mapping can be used to represent high-resolution attribute information, such as textures, normal vectors, etc.
[0064] In nearly all application fields using 3D meshes (such as computational simulation, entertainment, medical imaging, digitized artifacts, computer design, and e-commerce), the demand for visually appealing 3D mesh models is increasing, leading to increasingly complex models and higher precision. Consequently, the amount of data required to represent the 3D mesh is also increasing. These issues have led to increasing complexity in the processing, visualization, transmission, and storage of 3D meshes. 3D mesh compression can be considered a solution to these problems. It reduces the size of model data and facilitates the processing, storage, and transmission of 3D meshes. Therefore, it is necessary to propose an efficient and universal 3D mesh compression algorithm.
[0065] Recently, the Moving Picture Experts Group (MPEG), an international organization for standardization specializing in audio and video coding and compression, has begun developing a compression standard for 3D meshes called Video-based Dynamic Mesh Coding (V-DMC or VDMC). This standard is based on the existing Visual Volumetric Video-based Coding (V3C) standard, which provides a general method for compressing 3D models, which can be represented by point clouds, meshes, or panoramic videos. Compatibility of 3D mesh compression methods with this standard will facilitate the method's widespread adoption and applicability. Therefore, optimizing the 3D mesh encoding and decoding methods in VDMC and integrating these optimizations with the V3C standard is of great significance. One possible optimization approach is to optimize displacement coding. In existing frameworks, displacements are calculated by calculating the distances between the vertices of the reconstructed mesh and the original mesh, aiming to improve mesh quality. In existing frameworks, displacements are encoded using a video encoder. Providing multiple options for displacement encoding can help improve coding performance.
[0066] 2. V3C standard
[0067] The V3C standard provides a method for encoding and decoding various three-dimensional media through video or image coding technology. Specifically, it converts the three-dimensional media content from a three-dimensional representation into multiple two-dimensional representations (called V3C components) through projection and other methods before encoding, and then uses existing video or image coding technology to encode the two-dimensional representation. V3C components mainly include occupancy components, geometric components, and attribute components. The occupancy component can indicate which areas in the two-dimensional representation are associated with the data of the three-dimensional representation; the geometric component represents information related to the position of the three-dimensional data in space, and the attribute component can provide attribute information corresponding to the vertex, such as material, texture, etc. In addition, the components also contain information on how to reconstruct the three-dimensional model through these components, which is called atlas information. An example diagram of the V3C standard is shown below. Figure 1 and Figure 2 shown.
[0068] Atlas information is used to link all components, and additional information for reconstructing 3D from 2D is also included in the atlas components. An atlas consists of multiple basic units, called patches. Each patch represents a region of the available 2D components and contains the information needed to project that region back into 3D space.
[0069] 3. V-DMC
[0070] V-DMC is a standard developed by MPEG for compressing 3D meshes. Its main idea is to compress 3D meshes by leveraging the existing V3C standard. Since 3D meshes contain connection information that needs to be encoded, its specific encoding process is slightly different from V3C. The syntax, semantics, and decoding operations of the V3C standard decoder need to be extended to support the decoding and reconstruction of 3D meshes. Figure 3 and Figure 4 It is the current coding and decoding framework of V-DMC.
[0071] The overall framework of the encoding end is as follows Figure 3 As shown in the figure, the input mesh is first simplified by the simplification module. Mesh parameterization is then used to generate new texture coordinates for the mesh. The parameterized mesh is then subdivided and deformed. This involves inserting new vertices according to a specific subdivision method and calculating the distances between the subdivided mesh vertices and the nearest neighbor of the input mesh, known as displacement information. The pre-subdivided mesh, known as the base mesh, is then fed into the base mesh encoding module for compression using an existing mesh encoder. In inter-frame mode, the base mesh encoding module generates motion vectors for each vertex of the base mesh based on a reference frame, and only the motion vectors are compressed.
[0072] The base mesh is reconstructed after encoding, and then the order of displacement is adjusted according to the vertex order of the reconstructed base mesh. Subsequently, the vertex displacement information after the order is adjusted is subjected to wavelet transformation, the transformed coefficients are quantized, and then the quantized coefficients are arranged into a two-dimensional image according to a specific scanning order, and the two-dimensional image is encoded using a video encoder. Then, the reconstructed displacement information is applied to the subdivided base mesh to obtain a reconstructed subdivided and deformed mesh. This mesh, the original input mesh, and its corresponding texture map are input into the corresponding texture map conversion module to obtain the texture map corresponding to the reconstructed mesh, and this texture map is also encoded using a video encoder. The parameters used in the encoding process, such as the type of video encoder used, the type of mesh encoder, the transformation parameters, the quantization parameters, etc., are passed to the decoding end through auxiliary information.
[0073] The overall framework of the decoding end is as follows Figure 4As shown, for the received code stream, the decoding end first demultiplexes each part of the code stream to obtain the basic grid code stream, the displacement video code stream, the texture map video code stream and the auxiliary information code stream respectively. For the basic grid code stream, the grid decoder indicated by the auxiliary information is used to decode the basic grid. The displacement video code stream and the texture map video code stream are decoded by the video decoder. For the displacement part, after the video is decoded, the displacement needs to be taken out from the image through the displacement decoding module, and the steps such as inverse quantization and inverse transformation are performed, and then it is applied to the subdivided basic grid to obtain the deformed grid reconstructed by the decoding end. After decoding, the texture map is the texture map corresponding to the reconstructed deformed grid. The subsequent application or rendering module processes the reconstructed deformed grid and the decoded texture map as input.
[0074] The grid encoding method, grid decoding method and related equipment provided by the embodiments of the present application are described in detail below with reference to some embodiments and their application scenarios in combination with the accompanying drawings.
[0075] See also Figure 5 , Figure 5 This is a flowchart of a grid coding method provided by an embodiment of the present application, which can be applied to a coding terminal device, such as Figure 5 As shown, the grid coding method includes the following steps:
[0076] Step 101: Obtain a basic grid corresponding to a three-dimensional grid, and perform subdivision processing based on the basic grid to obtain a subdivided grid.
[0077] Among them, the three-dimensional grid can be grid simplified to obtain the basic grid corresponding to the three-dimensional grid; or, the three-dimensional grid can be grid simplified and grid parameterized to obtain the basic grid corresponding to the three-dimensional grid; or, the basic grid corresponding to the three-dimensional grid can be obtained by other means; and so on, this embodiment does not limit this.
[0078] In addition, the base mesh may be subdivided to obtain a subdivided mesh. The subdividing process may include adding vertices to the edges of the base mesh, and the base mesh after adding the vertices becomes a subdivided mesh.
[0079] In one implementation, the grid may be represented in the form of a grid sequence, and processing the grid may be considered as processing the grid sequence.
[0080] Step 102: Perform a deformation operation based on the subdivided grid to obtain a deformed grid.
[0081] The deformation operation may be to deform the subdivided mesh into a mesh with the same shape as the three-dimensional mesh, and the mesh obtained after deformation is the deformed mesh.
[0082] Step 103: Select displacement information to be encoded from the displacement information corresponding to the subdivided grid, where the displacement information corresponding to the subdivided grid is used to represent the displacement of the vertices of the subdivided grid relative to the deformed grid, and the number of levels of the displacement information to be encoded is less than or equal to the number of levels of the displacement information corresponding to the subdivided grid.
[0083] The displacement information corresponding to the subdivided mesh can be used to represent the displacement of the vertices of the subdivided mesh relative to the vertices of the deformed mesh. The displacement information can also be described as displacement data. For example, for each vertex of the subdivided mesh, a corresponding vertex can be found on the deformed mesh, and the displacement information corresponding to the subdivided mesh includes displacement information representing the displacement of each vertex of the subdivided mesh relative to the corresponding vertex on the deformed mesh.
[0084] It should be noted that due to the subdivision process, the displacement information includes multiple levels. The more times the base grid is subdivided, the more levels of displacement information there are. The displacement information corresponding to the subdivided grid may include multiple levels of displacement information.
[0085] In which, the displacement information corresponding to the subdivided grid may include displacement components of at least one dimension, and the displacement information to be encoded corresponding to each dimension of the displacement components of the at least one dimension may be selected respectively. For the displacement components of each dimension, the number of levels of the displacement information to be encoded corresponding to each dimension selected may be the same or different.
[0086] In one embodiment, selecting the displacement information to be encoded from the displacement information corresponding to the subdivided grid may include: determining the displacement information corresponding to the subdivided grid, the displacement information corresponding to the subdivided grid includes displacement components of at least one dimension, and selecting the displacement information to be encoded corresponding to each dimension of the displacement components of the at least one dimension, wherein the number of levels of the displacement information to be encoded corresponding to each dimension is less than or equal to the number of levels of the displacement components of each dimension.
[0087] For example, the displacement information corresponding to the subdivided grid may include: displacement components of the first dimension of N levels, displacement components of the second dimension of the N levels, and displacement components of the third dimension of the N levels. The displacement information to be encoded corresponding to the first dimension may include: displacement components of the first dimension of M1 levels among the displacement components of the first dimension of the N levels; the displacement information to be encoded corresponding to the second dimension may include: displacement components of the second dimension of M2 levels among the displacement components of the second dimension of the N levels; the displacement information to be encoded corresponding to the third dimension may include: displacement components of the third dimension of M3 levels among the displacement components of the third dimension of the N levels, N is a positive integer, M1, M2 and M3 are all integers greater than or equal to 0, and M1, M2 and M3 are all less than or equal to N.
[0088] It should be noted that the first dimension can be a first coordinate axis dimension, such as an x-axis dimension; the second dimension can be a second coordinate axis dimension, such as a y-axis dimension; and the third dimension can be a third coordinate axis dimension, such as a z-axis dimension.
[0089] Step 104: Encode the displacement information to be encoded to obtain a first encoding result.
[0090] The encoding of the displacement information to be encoded to obtain the first encoding result may be encoding the selected displacement information to be encoded to obtain the first encoding result.
[0091] In addition, the encoding of the displacement information to be encoded may be video encoding of the displacement information to be encoded, or entropy encoding of the displacement information to be encoded, etc. This embodiment does not limit the specific encoding method for encoding the displacement information to be encoded.
[0092] In related technologies, the same level is used for each dimension of displacement, but the actual displacement data is not necessarily evenly distributed in each dimension. In this case, if the same level of subdivision is used for each dimension, bits for encoding displacement data will be wasted. In this embodiment, it is allowed to encode only part of the level data of each dimension of displacement, and then the displacement data recovery operation is performed on each dimension separately at the decoding end. For each dimension i of displacement, only the first n i (n i >=0) levels of data are encoded, and Nn is discarded i At the decoding end, for each dimension i of the decoded displacement data, the discarded Nn can be replaced by zero padding, for example. i The data of each level is filled in, corresponding one to one with the mesh vertices on the decoding end.
[0093] In an embodiment of the present application, a base mesh corresponding to a three-dimensional mesh is obtained, and subdivision processing is performed based on the base mesh to obtain a subdivided mesh. A deformation operation is performed on the subdivided mesh to obtain a deformed mesh. Displacement information to be encoded is selected from the displacement information corresponding to the subdivided mesh, where the displacement information corresponding to the subdivided mesh is used to represent the displacement of vertices of the subdivided mesh relative to the deformed mesh, and the number of levels of the displacement information to be encoded is less than or equal to the number of levels of the displacement information corresponding to the subdivided mesh. The displacement information to be encoded is encoded to obtain a first encoding result. In this way, by encoding only the displacement information of some levels of the displacement information corresponding to the subdivided mesh, the amount of encoded data can be reduced, thereby improving encoding efficiency.
[0094] Optionally, the displacement information corresponding to the subdivided grid includes displacement information of multiple levels, and the selecting the displacement information to be encoded from the displacement information corresponding to the subdivided grid includes:
[0095] Determining a target level among the multiple levels, wherein the target level is determined based on displacement information representing a first preset value among the displacement information of the multiple levels;
[0096] Displacement information to be encoded is selected from the displacement information corresponding to the subdivided grid based on the target level.
[0097] The first preset value may be 0.
[0098] The displacement information may include a displacement component of at least one dimension, and for the displacement component of each dimension, a target level corresponding to each dimension may be determined respectively.
[0099] In one embodiment, the displacement components of at least one dimension include at least the displacement components of the target dimension, the displacement information corresponding to the subdivided grid includes at least the displacement components of the target dimension of multiple levels, and the number of displacement components representing the displacement of the first preset value in the displacement components of the target dimension of the target level meets a preset condition. The selecting of the displacement information to be encoded from the displacement information corresponding to the subdivided grid based on the target level may be selecting the displacement information to be encoded corresponding to the target dimension from the displacement components of the target dimension of the multiple levels based on the target level. For example, the target dimension may be the first dimension, the second dimension, or the third dimension.
[0100] In addition, the target level is determined based on the displacement information of the multiple levels whose characterizing displacement as the first preset value, and may include: the ratio of the number of displacement information of the target level whose characterizing displacement as the first preset value to the total number of displacement information of the target level satisfies the preset condition; or, may include: the difference between the number of displacement information of the target level whose characterizing displacement as the first preset value and the total number of displacement information of the target level satisfies the preset condition; or, may include: the ratio of the number of displacement information of the target level whose characterizing displacement as the first preset value to the number of displacement information of the target level whose characterizing displacement is not the first preset value satisfies the preset condition; or, may include: the difference between the number of displacement information of the target level whose characterizing displacement as the first preset value and the number of displacement information of the target level whose characterizing displacement is not the first preset value satisfies the preset condition; and so on. This embodiment does not limit this.
[0101] It should be noted that, since the sum of the number of displacement information whose displacement is the first preset value and the number of displacement information whose displacement is not the first preset value is the total number of displacement information, that is, the number of displacement information whose displacement is the first preset value and the number of displacement information whose displacement is not the first preset value have a corresponding relationship, wherein the target level is determined based on the displacement information representing the displacement of the first preset value in the displacement information of the multiple levels, it can also be described as: the target level is determined based on the displacement information representing the displacement of the multiple levels that is not the first preset value. The target level is determined based on the displacement information representing the displacement of the multiple levels that is not the first preset value, which may include: the ratio of the number of displacement information representing the displacement of the target level that is not the first preset value to the total number of displacement information of the target level meets the preset condition; or the difference between the number of displacement information representing the displacement of the target level that is not the first preset value and the total number of displacement information of the target level meets the preset condition; etc. This embodiment is not limited to this.
[0102] It should be noted that the target level can be determined starting from the first level among the multiple levels. When the target level is not the first level, the displacement information to be encoded includes the displacement information from the first level to the previous level of the target level; or, the target level can be determined starting from the last level among the multiple levels, and the displacement information to be encoded includes the displacement information from the first level among the multiple levels to the target level.
[0103] In this embodiment, a target level is determined from among the multiple levels, where the target level is determined based on displacement information representing a displacement of a first preset value from among the displacement information of the multiple levels; and displacement information to be encoded is selected from the displacement information corresponding to the subdivided grids based on the target level. In this way, the displacement information to be encoded can be determined from the displacement information of the multiple levels based on the number of displacement information representing a displacement of the first preset value. This allows for the selection of displacement information from a level carrying a greater amount of information for encoding, effectively reducing the amount of encoded data and improving encoding efficiency.
[0104] Optionally, a ratio of the number of displacement information representing a displacement of a first preset value in the displacement information of the target level to the total number of displacement information of the target level meets a preset condition.
[0105] The target level corresponding to the displacement component of each dimension may be determined respectively for the displacement component of each dimension, thereby supporting the selection of displacement information of different levels for encoding for displacement components of different dimensions.
[0106] In one embodiment, for the target dimension, the ratio of the number of displacement components representing a displacement of a first preset value in the displacement components of the target dimension of the target level to the total number of displacement components of the target dimension of the target level meets a preset condition.
[0107] In this embodiment, the ratio of the number of displacement information representing the displacement of the first preset value in the displacement information of the target layer to the total number of displacement information of the target layer meets the preset conditions, thereby supporting the selection of layers with a larger proportion of non-zero displacement information for displacement encoding and discarding layers with a smaller proportion of non-zero displacement information, which can effectively reduce the amount of encoded data and improve encoding efficiency.
[0108] Optionally, determining a target level among the multiple levels includes:
[0109] Determining a target level from a first level among the multiple levels, wherein a ratio of a number of displacement information items representing a displacement of a first preset value in the displacement information items of the target level to a total number of displacement information items of the target level is greater than the first preset ratio;
[0110] In a case where the target layer is not the first layer, the displacement information to be encoded includes displacement information from the first layer to a layer previous to the target layer.
[0111] The first preset value may be 0. The first preset ratio may be 80%, 85%, 90%, 95%, etc., and this embodiment does not limit the first preset ratio. In addition, when the target layer is the first layer, the displacement information to be encoded may be empty.
[0112] In addition, a target level corresponding to each displacement component of each dimension can be determined separately, thereby supporting the selection of displacement information of different levels for encoding displacement components of different dimensions. It should be noted that in the process of determining the target level corresponding to the displacement component of each dimension, the displacement component of each dimension can be set with a corresponding first preset ratio, and the first preset ratios corresponding to the displacement components of different dimensions can be the same or different.
[0113] In one embodiment, the ratio of the number of displacement components representing a displacement of a first preset value in the displacement components of the target dimension of the target level to the total number of displacement components of the target dimension of the target level is greater than the first preset ratio; when the target level is not the first level, the displacement information to be encoded corresponding to the target dimension includes the displacement components of the target dimension from the first level to the previous level of the target level.
[0114] It should be noted that the ratio of the number of displacement information of the target level that represents displacement as the first preset value to the total number of displacement information of the target level is greater than the first preset ratio. It can also be described as: the ratio of the number of displacement information of the target level that represents displacement not as the first preset value to the total number of displacement information of the target level is less than or equal to the first preset ratio.
[0115] For example, starting from the first level of displacement information, the number k of non-zero displacement components of the target dimension in that level can be counted, and the proportion of k in that level can be calculated. If the proportion is less than or equal to a first preset ratio, the displacement components of the target dimension in that level and all levels after it are removed, and only the displacement components of the target dimension before that level are encoded. If the proportion does not meet the requirement of being less than or equal to the first preset ratio, the same statistical calculation is performed on the next level until a level that meets the conditions is found or all subdivision levels are traversed.
[0116] In this embodiment, the target level is determined starting from the first level of the multiple levels, wherein the ratio of the number of displacement information items representing displacements of a first preset value in the displacement information of the target level to the total number of displacement information items of the target level is greater than the first preset ratio; if the target level is not the first level, the displacement information to be encoded includes the displacement information from the first level to the level immediately preceding the target level. In this way, level selection can be performed starting from the first level, discarding the displacement information of the target level and subsequent levels that carry less information, which can effectively reduce the amount of encoded data and improve encoding efficiency.
[0117] Optionally, determining a target level among the multiple levels includes:
[0118] Determining a target level from the last level of the multiple levels, wherein a ratio of the number of displacement information items representing displacements of the first preset value in the displacement information items of the target level to the total number of displacement information items of the target level is less than or equal to a second preset ratio;
[0119] The displacement information to be encoded includes displacement information from a first level among the multiple levels to the target level.
[0120] The first preset value may be 0. The second preset ratio may be 5%, 8%, 10%, 15%, etc. This embodiment does not limit the second preset ratio.
[0121] In addition, a target level corresponding to each displacement component of each dimension can be determined separately for the displacement component of each dimension, thereby supporting the selection of displacement information of different levels for encoding for displacement components of different dimensions. It should be noted that in the process of determining the target level corresponding to the displacement component of each dimension, the displacement component of each dimension can be set with a corresponding second preset ratio, and the second preset ratios corresponding to the displacement components of different dimensions can be the same or different.
[0122] In one embodiment, the ratio of the number of displacement components in the target dimension of the target level that represent a displacement of a first preset value to the total number of displacement components in the target dimension of the target level is less than or equal to a second preset ratio.
[0123] It should be noted that the ratio of the number of displacement information of the target level that represents displacement as the first preset value to the total number of displacement information of the target level is less than or equal to the second preset ratio. It can also be described as: the ratio of the number of displacement information of the target level that represents displacement not as the first preset value to the total number of displacement information of the target level is greater than the second preset ratio.
[0124] For example, the displacement component of the target dimension is statistically calculated continuously from the last level forward, the number k of non-zero displacement components of the target dimension in the level is counted, and the proportion of k in the level is calculated. If the proportion of the last level is greater than the second preset ratio, the forward traversal is stopped. If the proportion of the last level meets the requirements (less than or equal to the second preset ratio), the displacement component of the target dimension of the level is removed, and the same statistical calculation and proportion judgment are continued for the previous level until a level that does not meet the requirements is encountered or all subdivided levels are traversed.
[0125] In this embodiment, the target level is determined starting from the last level of the multiple levels, wherein the ratio of the number of displacement information items representing displacements of a first preset value in the displacement information of the target level to the total number of displacement information items of the target level is less than or equal to a second preset ratio; and the displacement information to be encoded includes the displacement information from the first level of the multiple levels to the target level. In this way, level selection can be performed starting from the last level, discarding the displacement information of levels after the target level that carry less information, which can effectively reduce the amount of encoded data and improve encoding efficiency.
[0126] Optionally, the displacement information corresponding to the subdivided grid includes: displacement components of a first dimension of N levels, displacement components of a second dimension of the N levels, and displacement components of a third dimension of the N levels;
[0127] The displacement information to be encoded includes: the displacement components of the first dimension of M1 levels among the displacement components of the first dimension of the N levels, the displacement components of the second dimension of M2 levels among the displacement components of the second dimension of the N levels, and the displacement components of the third dimension of M3 levels among the displacement components of the third dimension of the N levels, N is a positive integer, M1, M2 and M3 are all integers greater than or equal to 0, and M1, M2 and M3 are all less than or equal to N.
[0128] The selected displacement information to be encoded includes: the displacement components of the first dimension of the M1 levels, the displacement components of the second dimension of the M2 levels, and the displacement components of the third dimension of the M3 levels.
[0129] In this embodiment, the displacement information to be encoded includes the displacement components of the first dimension of M1 levels among the displacement components of the first dimension of the N levels, the displacement components of the second dimension of M2 levels among the displacement components of the second dimension of the N levels, and the displacement components of the third dimension of M3 levels among the displacement components of the third dimension of the N levels, thereby supporting the selection of displacement components of different levels for displacement components of each dimensionality for displacement encoding, which can improve the flexibility of displacement encoding while improving the encoding efficiency.
[0130] Optionally, the code stream corresponding to the three-dimensional grid includes the first encoding result and a second encoding result, and the second encoding result includes the encoding result of the first flag bit;
[0131] When the first flag bit is a second preset value, the second encoding result also includes an encoding result of the first parameter; or,
[0132] When the first flag bit is a third preset value, the second encoding result also includes an encoding result of the first parameter, an encoding result of the second parameter, and an encoding result of the third parameter.
[0133] The first parameter is used to characterize the M1 levels, the second parameter is used to characterize the M2 levels, and the third parameter is used to characterize the M3 levels.
[0134] In addition, the second preset value and the third preset value are different values. For example, the second preset value may be 0, and the third preset value may be 1; or, the second preset value may be 1, and the third preset value may be 0. This embodiment does not limit the second preset value and the third preset value.
[0135] It should be noted that the displacement component of the first dimension may be a more important displacement component. In some scenarios, it may be permitted to encode only the displacement component of the first dimension. The decoding end can determine the dimension of the encoded displacement component by using the encoding result of the first flag bit carried in the bitstream. For example, the displacement component of the first dimension may be the normal component of the displacement.
[0136] In this embodiment, when the first flag is a second preset value, the second encoding result also includes the encoding result of the first parameter; when the first flag is a third preset value, the second encoding result also includes the encoding results of the first parameter, the second parameter, and the third parameter. Thus, the decoder can determine the displacement component of the encoding through the first flag; and through the parameter information included in the second encoding result, the decoder can determine the level of the displacement information selected for encoding.
[0137] Optionally, encoding the displacement information to be encoded to obtain a first encoding result includes:
[0138] Packing the displacement information to be encoded into a video frame;
[0139] Performing padding processing on the arranged video frames to obtain padded video frames;
[0140] The padded video frame is encoded to obtain a first encoding result.
[0141] The arranged video frames may be padded based on a preset video padding value, which may be a predetermined value known to the decoding end.
[0142] In this embodiment, the displacement information to be encoded is arranged into a video frame; the arranged video frame is padded to obtain a padded video frame; and the padded video frame is encoded to obtain a first encoding result. Thus, when encoding the displacement information using a video encoding method, the video padding process ensures that the resolution of each frame in the video is the same, and the pixel data can be arranged in a rectangular shape.
[0143] Optionally, encoding the displacement information to be encoded to obtain a first encoding result includes:
[0144] Entropy coding is performed on the displacement information to be encoded to obtain a first coding result.
[0145] In this implementation, entropy coding is performed on the displacement information to be encoded to obtain a first coding result, and encoding of the displacement information can be achieved through entropy coding.
[0146] As a specific embodiment, Figure 6 and Figure 7 As shown, the embodiment of the present application proposes a new encoding and decoding method for dynamic mesh displacement, which allows encoding of only part of the hierarchical data of each dimensional component of the displacement, and then performs displacement data recovery operations on each dimension separately at the decoding end. Displacement is an important part of the dynamic mesh encoding and decoding process. It stores the distance information of the vertices of the subdivided mesh obtained after the basic mesh is subdivided to the nearest neighbor point of the original input mesh along the normal direction. It is a three-dimensional data, and the x, y, and z components can be used to represent its three-dimensional data respectively. Due to the existence of the subdivision operation, the displacement has multiple subdivision levels and a base level. The base level is in front and the subdivision level is in the back. The more times the basic mesh is subdivided, the more subdivision levels of the displacement. Assuming that there are N (N>=1) levels of displacement, for each dimension i of the displacement, only the first n can be coded at the encoding end. i (n i>=0) levels of data are encoded, and Nn is discarded i At the decoding end, for each dimension i of the decoded displacement data, the discarded Nn is replaced by zero padding, for example. i The data of each level is filled in, corresponding one to one with the mesh vertices on the decoding end.
[0147] like Figure 6 As shown, at the encoding end, the displacement data is first input into the pre-processing module, which includes all operations before encoding except data selection, such as transformation, quantization, etc.; then each dimensional component of the pre-processed displacement data is input into the corresponding level selection module, and these modules will output the corresponding possible x, y, and z component label data according to different selection methods during the selection process; finally, the selected displacement level data is input into the encoding module for encoding to obtain the displacement code stream. There are many specific encoding methods for this module, such as video encoding, entropy coding, etc. Figure 7 As shown, at the decoding end, the displacement code stream is decoded using a decoding method corresponding to the encoding method. Each dimensional component of the decoded data is then fed into a corresponding data padding module. These modules pad each dimensional component of the decoded partial displacement data to a complete length based on the x, y, and z component marker data that may have been received, ultimately recovering the complete three-dimensional displacement data. Finally, the padded displacement data is fed into a post-processing module to obtain the displacement data. The operations of the post-processing module are the reverse of the operations of the pre-processing module on the encoding end, such as dequantization and inverse transformation.
[0148] See also Figure 8 , Figure 8 This is a flowchart of a grid decoding method provided by an embodiment of the present application, which can be applied to a decoding terminal device, such as Figure 8 As shown, the grid decoding method includes the following steps:
[0149] Step 201: Decode a first encoding result in a code stream corresponding to a three-dimensional grid to obtain decoded displacement information;
[0150] Step 202: Decode the third encoding result in the code stream to obtain a basic grid, and subdivide the basic grid to obtain a subdivided grid;
[0151] Step 203: performing padding processing on the decoded displacement information to obtain padded displacement information, wherein the number of levels of the padded displacement information is greater than or equal to the number of levels of the decoded displacement information;
[0152] Step 204 : Reconstruct the grid based on the filled displacement information and the subdivided grid to obtain a reconstructed grid.
[0153] Optionally, decoding the first encoding result in the code stream corresponding to the three-dimensional grid to obtain decoded displacement information includes:
[0154] Decoding a first encoding result in a code stream corresponding to the three-dimensional grid to obtain a video frame;
[0155] Decoded displacement information is obtained from the video frame, where the decoded displacement information is displacement information in the video frame excluding a preset video filling value.
[0156] Optionally, the method further includes:
[0157] Decoding the second encoding result in the code stream to obtain a target parameter, where the target parameter is used to represent the number of levels of displacement information;
[0158] The decoding of the first encoding result in the code stream corresponding to the three-dimensional grid to obtain decoded displacement information includes:
[0159] Decoding a first encoding result in a code stream corresponding to the three-dimensional grid to obtain a video frame;
[0160] Decoded displacement information is obtained from the video frame based on the target parameter.
[0161] Optionally, decoding the first encoding result in the code stream corresponding to the three-dimensional grid to obtain decoded displacement information includes:
[0162] Entropy decoding is performed on a first encoding result in a code stream corresponding to the three-dimensional grid to obtain decoded displacement information.
[0163] Optionally, performing padding processing on the decoded displacement information to obtain padded displacement information includes:
[0164] The decoded displacement information is padded based on the number of vertices of the subdivided mesh to obtain padded displacement information.
[0165] Optionally, the padded displacement information includes N levels of first-dimensional displacement components, N levels of second-dimensional displacement components, and N levels of third-dimensional displacement components;
[0166] The decoded displacement information includes the displacement components of the first dimension of M1 levels among the displacement components of the first dimension of the N levels, the displacement components of the second dimension of M2 levels among the displacement components of the second dimension of the N levels, and the displacement components of the third dimension of M3 levels among the displacement components of the third dimension of the N levels, where N is a positive integer, M1, M2 and M3 are all integers greater than or equal to 0, and M1, M2 and M3 are all less than or equal to N.
[0167] Optionally, the method further includes:
[0168] Decoding the second encoding result in the code stream to obtain a first flag bit and a target parameter;
[0169] Wherein, when the first flag bit is the second preset value, the target parameter includes the first parameter; or,
[0170] When the first flag bit is the third preset value, the target parameter includes the first parameter, the second parameter and the third parameter.
[0171] The first parameter is used to characterize the M1 levels, the second parameter is used to characterize the M2 levels, and the third parameter is used to characterize the M3 levels.
[0172] The following describes the grid encoding and decoding method through several specific embodiments:
[0173] It should be noted that for different bit-shift encoding and decoding methods, the encoding and decoding end processing methods have different implementation processes. The following will explain the bit-shift encoding and decoding process according to different encoding and decoding methods.
[0174] Example 1:
[0175] In this embodiment, the displacement information is encoded and decoded based on a video encoding and decoding method.
[0176] Encoding side:
[0177] like Figure 9 As shown, the displacement data is input into the pre-processing module, which performs a series of pre-processing operations on the original displacement data, such as transformation and quantization. Then, each dimensional component of the pre-processed displacement data is input into the corresponding level selection module. The level selection module selects the level by judging whether the proportion of the number of non-zero displacement components in the level is less than or equal to the set threshold. Among them, when different dimensional components are selected for the level, the setting of their thresholds is independent of each other. Level selection includes the following two methods:
[0178] The first method: Starting from the first level of displacement data, count the number k of non-zero displacement components in that level and calculate the proportion of k in that level. If the proportion is less than or equal to the set threshold, the displacement components of that level and all levels after it are removed, and only the displacement components before the level are encoded. If the proportion does not meet the requirement, continue to perform the same statistical calculation for the next level until a level that meets the requirements is found or all subdivision levels are traversed.
[0179] The second method is to start from the last level and continuously perform statistical calculations forward. If the proportion of the last level does not meet the requirements, the forward traversal stops. If the proportion of the last level meets the requirements, the corresponding displacement component of the level is removed, and the same statistical calculations and proportion judgments are performed on the previous level until a level that does not meet the requirements is encountered or all subdivision levels are traversed.
[0180] It should be noted that you can choose to mark the number of levels selected for each dimensional component and pass it to the decoding end; or you can choose not to mark the number of levels selected for each dimensional component, and only need to fill each dimensional component of the displacement information at the decoding end according to the number of vertices in the decoded grid. After the level selection, the selected displacement data can be arranged into the YUV image according to a predetermined rule. During the video arrangement process, in order to ensure that the resolution of each frame in the YUV video is the same and the pixel data can be arranged into a rectangle, the YUV video can be filled. The video filling value to be filled can be an agreed value known to the decoding end.
[0181] Use a video encoder to encode the YUV video to obtain a displacement code stream.
[0182] Decoding end:
[0183] Depending on whether the encoder marks the number of selected levels for each dimensional component of the displacement, there are different implementation methods for the displacement decoding method based on video decoding. Two possible implementation methods are described below.
[0184] (1) Do not mark the number of selected levels
[0185] like Figure 10 As shown, the displacement decoding process for each dimension component, without marking the number of selected levels, is as follows: For the YUV video obtained by video decoding, all image pixel values are extracted in the order in which the displacement data was arranged in the video by the encoder. The extracted image pixel values contain two parts of data: the displacement data selected by the encoder; and the video padding value of the encoder, which is a value agreed upon between the encoder and decoder. After extracting the pixel value data, the video padding portion of the pixel value data is removed according to the video padding value agreed upon with the encoder, thereby obtaining the displacement data selected by the encoder.
[0186] According to the number of vertices of the mesh obtained after the decoded mesh is subdivided, each dimensional component of the displacement data is padded with zeros until the number of data is the same as the number of vertices of the mesh obtained after the decoded mesh is subdivided. At this time, the restored complete displacement data can be obtained. It should be noted that the number of vertices of the mesh obtained after the decoded mesh is subdivided can be obtained in other modules of the dynamic mesh decoding framework. For example, the V-DMC framework can obtain it by subdividing the decoded basic mesh. The restored complete displacement data is then input into the post-processing module to obtain the original displacement data, completing the decoding of the displacement. The operation of the post-processing module is the inverse implementation of the operation of the pre-processing module on the encoding end, such as inverse quantization and inverse transformation.
[0187] (2) Mark the number of selected levels
[0188] like Figure 11 As shown, the displacement decoding process for marking the number of selected levels for each dimensional component of the displacement is as follows: Since the number of displacement levels and the number of displacement data in each level are exactly the same as the subdivided decoding basic grid, the number of displacement data selected by the encoder for each displacement dimensional component can be obtained based on the selected number of levels for each dimensional component in the subdivided decoding basic grid. Based on the number of displacement data for each dimensional component, the data for each selected displacement dimensional component is extracted from the YUV video obtained by video decoding in the order in which the displacement data was arranged into the video by the encoder. After obtaining the data for each dimensional component of the displacement, zero padding is performed at the end of the data for each dimensional component based on the number of vertices in the grid after subdividing the decoding grid until the number of data and the number of vertices are the same; the complete padded displacement data is then input into the post-processing module to obtain the original displacement data, completing the displacement decoding. The operations of the post-processing module are the inverse implementation of the operations of the pre-processing module on the encoder, such as inverse quantization and inverse transformation.
[0189] Example 2:
[0190] In this embodiment, the displacement information is encoded and decoded based on an entropy coding method.
[0191] Encoding side:
[0192] When entropy coding is selected as the displacement coding method, the displacement coding process is as follows Figure 12As shown. First, the displacement data is input into the pre-processing module, which performs a series of pre-processing operations on the original displacement data, such as transformation and quantization; then, for each dimensional component of the pre-processed displacement data, it is input into the corresponding level selection module. The level selection module selects the level by judging whether the proportion of the number of non-zero displacement components in the level is less than or equal to the set threshold. Among them, when different dimensional components are selected for level, the setting of their thresholds is independent of each other. Level selection includes the following two methods:
[0193] The first method: Starting from the first level of displacement data, count the number k of non-zero displacement components in that level and calculate the proportion of k in that level. If the proportion is less than or equal to the set threshold, the displacement components of that level and all levels after it are removed, and only the displacement components before the level are encoded. If the proportion does not meet the requirement, the same statistical calculation is performed on the next level until a level that meets the requirements is found or all subdivision levels are traversed.
[0194] The second method is to start from the last level and continuously perform statistical calculations forward. If the proportion of the last level does not meet the requirements, the forward traversal stops. If the proportion of the last level meets the requirements, the corresponding displacement component of the level is removed, and the same statistical calculations and proportion judgments are performed on the previous level until a level that does not meet the requirements is encountered or all subdivision levels are traversed.
[0195] After the hierarchical selection, the selected displacement data can be directly entropy coded. The specific entropy coding method is not limited here, and for example, context-based adaptive variable length coding (CAVLC) and context-based adaptive binary arithmetic coding (CABAC) can be used.
[0196] Decoding end
[0197] like Figure 13As shown, for the received displacement code stream, the displacement code stream is first entropy decoded to obtain each dimension component of the selected displacement, and then zero padding is performed at the end of each dimension component of the displacement according to the number of vertices of the mesh after the decoded mesh is subdivided, until the number of data is the same as the number of vertices. It should be noted that the number of vertices of the mesh after the decoded mesh is subdivided can be obtained in other modules of the dynamic mesh decoding framework, such as the V-DMC framework can be obtained by subdividing the decoded basic mesh. The restored complete displacement data is then input into the post-processing module to obtain the original displacement data, completing the decoding of the displacement. The operation of the post-processing module is the inverse implementation of the operation of the pre-processing module on the encoding end, such as inverse quantization and inverse transformation.
[0198] Example 3:
[0199] In this embodiment, the displacement information is encoded and decoded based on the V-DMC framework.
[0200] Encoding side:
[0201] The application of the embodiment of the present application at the encoding end is mainly in the displacement encoding process. Figure 14 As shown, the grid coding method performed by the encoding end includes the following processes:
[0202] (1) Mesh simplification
[0203] Mesh simplification is to simplify the current input mesh into a base mesh with relatively few points and faces, and to keep the shape of the original mesh as much as possible. The key points of mesh simplification are the simplification operation and the corresponding error metric. A feasible mesh simplification operation is Figure 15 As shown in the figure, the vertices at both ends of the edge are merged into a single vertex and the connection between the two vertices is deleted. This process is repeated throughout the mesh according to a certain rule to reduce the number of faces and vertices of the mesh to the target value.
[0204] During the simplification process, a specific error metric can be selected to optimize the simplified result. For example, the error metric for a vertex can be the sum of the coefficients of the equations of all adjacent faces. The error metric for an edge can be the sum of the error metrics of the two vertices on the edge. In other words, the error resulting from merging an edge is the sum of the distances from the merged vertex to all adjacent faces of the original two vertices on the edge.
[0205] After determining the simplification operation and the corresponding error metric, the mesh simplification process begins iteratively. First, the vertex errors of the initial mesh are calculated to obtain the error for each edge. Edges are then sorted from smallest to largest error, and the edge with the smallest error is merged each time. Simultaneously, the positions of the merged vertices are calculated, and the errors of all edges associated with the merged vertices are updated. This means that the order of edge arrangement is updated to ensure that each iteration is based on a global error metric. Through iteration, the mesh faces are simplified to the number required for lossy encoding.
[0206] (2) Mesh parameterization
[0207] Texture coordinates may be regenerated based on the reconstructed base mesh of the current frame. This step requires generating texture coordinates for each attribute map of the input mesh. If multiple attribute maps have similar characteristics, they can share the same texture coordinates. Texture coordinate generation can be done by mesh parameterization, among other methods. Numerous algorithms exist for mesh parameterization, such as the Isocharts algorithm, which uses spectral analysis to implement stretch-driven 3D mesh parameterization. The 3D mesh is then UV-unwrapped, tiled, and packed into a 2D texture domain.
[0208] (3) Subdivision deformation
[0209] The subdivision and deformation modules are optional modules that can improve the quality of the mesh reconstructed by the decoder. They can be omitted if the base mesh quality already meets the application requirements.
[0210] The basic idea of subdivision and deformation module is as follows Figure 16 As shown in , the same concept is applied to the input 3D mesh to generate displacement vector information. Figure 16 In [1], the input 2D curve (represented by a 2D polyline), called the "original" curve, is first downsampled to generate a basic curve / polyline, called the "simplified" curve. The subdivision scheme is then applied to the simplified polyline to generate the "subdivided" curve. The subdivided polyline is then deformed to obtain a better approximation of the original curve. That is, a geometric displacement vector is calculated for each vertex of the subdivided mesh so that the shape of the subdivided curve is as close as possible to the shape of the original curve. These geometric displacement vectors are the geometric displacement vector information (vertex displacement) output by this module. The same deformation process is also applied to the attribute information corresponding to the vertex to obtain the corresponding attribute displacement vector.
[0211] (4) Compression of base mesh
[0212] The basic mesh compression module compresses the basic mesh information output by the subdivision deformation module. There are two main modes of basic mesh compression: intra-frame mode and inter-frame mode. Figure 17As shown in Figure 2. In intra mode, the input to the base mesh compression module is a 3D mesh containing geometric coordinates, connectivity, and vertex attribute information. In inter mode, the input to the base mesh compression module is a motion vector and its identifier, as well as any intra-frame sub-meshes. After encoding, the encoded mesh needs to be reconstructed for processing by subsequent modules.
[0213] Among them, the grid coding in intra mode can use the existing grid encoder to encode the input basic grid, such as Draco. The grid encoder type is encoded through auxiliary information and passed to the decoding end;
[0214] In inter-frame mode, motion vectors and identifiers must be encoded. The motion vector identifier is an array of numbers greater than or equal to 0. The array size is equal to the number of vertices in the base mesh of the reference frame corresponding to the temporal motion vector identifier. This array can be directly encoded using existing entropy coding algorithms, such as CABAC.
[0215] (5) Reconstruct the base mesh
[0216] After encoding the base grid, the base grid is reconstructed. For grids encoded in intra-frame mode, the encoded grid is decoded to obtain the reconstructed grid.
[0217] For the basic grid information encoded in the inter-frame mode, it is necessary to reconstruct the inter-frame sub-grid based on the motion vector and the identifier. If there is an intra-frame sub-grid in the inter-frame mode, the reconstructed inter-frame sub-grid is merged with the intra-frame sub-grid to obtain a reconstructed basic grid.
[0218] (6) Encoding vertex displacement
[0219] To encode vertex displacements, first adjust the order of vertex displacements according to the reconstructed base mesh; then perform wavelet transform and quantization on the displacements to obtain quantized wavelet transform coefficients; then input the quantized wavelet transform coefficients into the data truncation module; then arrange the wavelet transform coefficients of the current frame into a two-dimensional image, finally forming a YUV video, which is then encoded with the help of a video encoder to obtain the displacement code stream. Encoding vertex displacements includes the following process:
[0220] a. Adjust the displacement order
[0221] Adjust the order of vertex displacements based on the reconstructed base mesh. First, subdivide the reconstructed base mesh, then adjust the order of displacements to match the subdivided base mesh vertex order. Optionally, the coordinate system of the vertex displacements can be converted from the Cartesian coordinate system to the local coordinate system.
[0222] b. Wavelet transform
[0223] A transformation can be applied to the displacement vector to reduce the correlation between its data. An optional transformation is a linear wavelet transform, and its prediction process is defined as follows:
[0224]
[0225] Where v is the newly inserted midpoint on the edge (v1, v2), Signal(v), Signal(v1), and Signal(v2) are the displacement vectors corresponding to vertices v, v1, and v2, respectively. The displacement vector of vertex v is predicted and then updated. The update process is defined as follows:
[0226]
[0227] where v * is the set of all adjacent vertices of vertex v. The transformed displacement vector is called wavelet coefficient.
[0228] c. Coefficient quantization
[0229] The transformed displacement vector, i.e., the wavelet coefficient, can be quantized. There are many ways to quantize it. One quantization formula is as follows:
[0230] disp[v].d[k]=floor(disp[v].d[k]*scale[k])
[0231]
[0232] Where disp[v] represents the transformed value of the displacement vector at the vth vertex, d[k] represents the kth value of the displacement vector, and floor indicates rounding down. scale[k] is an intermediate parameter. bitDepthPosition represents the bit depth of the current mesh vertex's geometric position, and qp[k] represents the quantization parameter for the kth coefficient. As mentioned earlier, after transforming the coordinate system of the displacement vector, its normal component has a more significant impact on quality than its tangential component, so a larger quantization parameter can be used for the tangential component.
[0233] At the same time, according to the characteristics of wavelet transform, different quantization parameters can be used for the newly generated vertices and the original vertices. That is, for the subdivided vertices, the quantization parameter update formula is as follows:
[0234] scale[k]=scale[k]*lodScale[k]
[0235] Among them, lodScale[k] represents the coefficient of the quantization parameter of the current subdivision level.
[0236] d. Level selection
[0237] For the quantized wavelet transform coefficients, before arranging them into a two-dimensional image, the x, y, and z three-dimensional components of the quantized wavelet transform coefficients are subjected to a level selection operation. The level selection operation selects the level by judging whether the proportion of the number of non-zero corresponding components in the level of the quantized wavelet transform coefficients in the level is less than or equal to the set threshold. Specifically, starting from the first level of the quantized wavelet transform coefficients, the number k of non-zero corresponding components in the level is counted, and the proportion of k in the level is calculated. If the proportion is less than or equal to the set threshold, the corresponding components of the wavelet transform coefficients of the level and all levels after the level are removed, and only the corresponding components of the wavelet transform coefficients before the level are encoded; if the proportion does not meet the requirements, the same statistical calculation is performed on the next level until a level that meets the conditions is found or all subdivided levels are traversed. Alternatively, statistical calculations are performed continuously from the last level forward. If the proportion of the last level does not meet the requirements, the forward traversal is stopped. If the proportion of the last level meets the requirements, the corresponding components of the wavelet transform coefficients of the level are removed, and the same statistical calculations and proportion judgments are performed on the previous level until a level that does not meet the requirements is encountered or all subdivided levels are traversed.
[0238] It should be noted that you can choose to mark the number of levels selected for each dimensional component and store it in the auxiliary information and transmit it to the decoding end; or you can choose not to mark the number of levels selected for each dimensional component, and only need to fill in the displacement information according to the number of decoded mesh vertices at the decoding end.
[0239] e. Two-dimensional arrangement
[0240] After the displacement data is truncated through hierarchical selection, the displacement data can be arranged on the two-dimensional image as follows:
[0241] Traverse the wavelet coefficients in order from low frequency to high frequency;
[0242] For each coefficient, determine the index of the NxM pixel block where it is located (for example, N=M=16), and store it in the raster scan order of the block;
[0243] The position of the corresponding NxM pixel block on the image is calculated according to the Morton order.
[0244] It should be noted that this is not limited to block arrangement; other arrangements, such as zigzag order and raster order, can also be used. The encoder can explicitly specify the corresponding arrangement in the bitstream. To ensure that the resolution of each frame of the generated YUV video is the same and the number of pixels can be arranged in a rectangular shape, video padding can be performed. The video padding value can be a predetermined value known by the decoder.
[0245] f. Video compression
[0246] After arranging the wavelet coefficients on a two-dimensional image, they can be directly encoded using a video encoder to obtain a shifted code stream.
[0247] (7) Reconstruction of deformed mesh
[0248] Displacement vector encoding requires obtaining the reconstructed displacement vector value. There are two methods for obtaining quantized wavelet transform coefficients from a reconstructed two-dimensional image, depending on whether the number of levels for selecting each dimensional component of the displacement is marked when selecting the level:
[0249] The first method: unmark the number of selected levels
[0250] In this way, the number of levels for selecting each dimensional component of the unmarked displacement is selected. When obtaining the quantized wavelet transform coefficients, all pixel data in the two-dimensional image are firstly taken out according to the arrangement method during encoding; the extracted pixel data contains two values, one is the quantized wavelet transform coefficient selected by the level selection module, and the other is the video filling value in the two-dimensional arrangement process. After obtaining the pixel data, the video filling part in the pixel data can be removed according to the video filling value to obtain the data of each dimensional component of the quantized wavelet transform coefficient selected by the level selection module. Then, each dimensional component of the selected quantized wavelet transform coefficient is padded with zeros according to the number of vertices after the subdivision of the reconstructed basic mesh, until the number of each dimensional component data selected is the same as the number of vertices after the subdivision of the reconstructed basic mesh, and the quantized wavelet transform coefficient can be obtained.
[0251] The second method: mark the number of selected levels
[0252] In this way, the number of levels for selecting each dimensional component of the displacement is marked. When obtaining the quantized wavelet transform coefficients, the number of levels of the quantized wavelet transform coefficients and the number of data in each level are exactly the same as those of the subdivided reconstructed basic grid. Therefore, in the subdivided reconstructed basic grid, the number of levels for selecting each dimensional component of the quantized wavelet transform coefficients selected by the level selection module can be obtained according to the number of levels for selecting each dimensional component. Then, according to the number of dimensional component data of the wavelet transform coefficients, the selected data is taken out from the two-dimensional image according to the arrangement method during encoding. Then, according to the number of vertices of the subdivided reconstructed basic grid, zero padding operations are performed at the end of each dimensional component data taken out, until the number of data is the same as the number of vertices, and the quantized wavelet transform coefficients can be obtained.
[0253] The quantized wavelet transform coefficients are dequantized and inversely transformed to obtain a displacement vector consistent with the decoding end. After obtaining the reconstructed geometric displacement vector, the subdivided reconstructed base mesh is reconstructed according to the corresponding geometric displacement vector to obtain a reconstructed subdivided deformed mesh, which is passed to the attribute graph conversion module for attribute graph conversion.
[0254] (8) Attribute Graph Conversion
[0255] The attribute map is converted based on the input original mesh, the input original attribute map and the reconstructed deformed mesh, such as Figure 18 shown.
[0256] The steps for attribute graph conversion are as follows:
[0257] a. Calculate the texture coordinates of each pixel on the attribute map to be generated. For example, the texture coordinates corresponding to pixel A(i, j) are P(u, v).
[0258] b. Determine whether the texture coordinates are within a certain triangle face after the parameterization of the subdivided deformed mesh;
[0259] c. If the texture coordinate does not belong to any triangle, mark the pixel as an empty pixel and then fill it with a filling algorithm;
[0260] d. If the texture coordinate belongs to a triangle, then:
[0261] Mark the pixel as filled;
[0262] Calculate the center of gravity coordinates of the texture in the current triangle according to the texture coordinates;
[0263] According to the barycentric coordinates and the corresponding triangular face, the two-dimensional texture coordinates are mapped to the three-dimensional geometric coordinates, that is, mapped to the points on the subdivided deformed grid corresponding to the texture coordinates, such as Figure 18 As shown in M(x,y,z);
[0264] Find the point closest to the three-dimensional coordinate on the input original grid, such as Figure 18 Medium M ′ (x,y,z)
[0265] The three-dimensional coordinates are calculated based on the center coordinates of the triangle face and mapped to two dimensions to calculate the texture coordinates, that is, P ′ (u ′ ,v ′ );
[0266] The texture coordinates are used to sample the original attribute map input to obtain the value A at the corresponding pixel position. ′ (i ′ ,j ′ );
[0267] Assign this value to the corresponding pixel A(i,j) on the attribute map to be generated.
[0268] (9) Attribute Graph Compression
[0269] After obtaining the converted attribute map, the empty pixels therein can be filled using existing filling algorithms (such as the Push-Pull algorithm). Then, existing video encoders such as H.264 / AVC, H.265 / HEVC, and H.266 / VVC can be used to encode the empty pixels to obtain the bitstream of the output attribute map. In addition, operations such as color space conversion and chroma subsampling can be selectively applied to achieve better rate-distortion performance in video encoding, such as color space conversion from RGB 444 to YUV420. In the case of multiple attribute maps, the texture coordinates corresponding to each attribute map need to be identified through auxiliary information during encoding.
[0270] (10) Auxiliary information coding
[0271] During the encoding process, various modules can choose from a variety of alternative schemes, such as the grid encoder type, video encoder type, grid subdivision scheme, and spatial displacement vector transform scheme. This encoding framework allows for the use of different schemes. Therefore, the selected scheme can be communicated to the decoder to guide correct decoding.
[0272] The auxiliary information includes the type of static grid encoder, the type of video encoder, the grid subdivision scheme, the number of grid iterations, the geometric displacement vector transformation scheme, the coefficient arrangement scheme, the number of levels selected for the displacement x, y, and z components, or the color conversion scheme, etc. For the grid part encoded using the inter-frame coding mode, it also includes the corresponding reference frame list, etc.
[0273] After all modules are encoded, the basic grid part stream, texture coordinate part stream, displacement vector video stream, attribute map video stream and auxiliary information stream are mixed to obtain the final encoded stream output by the encoding end.
[0274] Decoding end:
[0275] like Figure 19 As shown, the grid decoding method performed by the decoding end includes the following process:
[0276] (1) Auxiliary information decoding
[0277] Auxiliary information indicates the type of static mesh encoder, video encoder type, mesh subdivision scheme, mesh iteration count, geometric displacement vector transformation scheme, coefficient permutation scheme, displacement truncation flag, or color conversion scheme, among other things. For meshes encoded using inter-frame coding, this information also includes the corresponding reference frame list. Auxiliary information can be used to guide the decoder for correct decoding. After decoding, this information is represented by corresponding syntax parameters.
[0278] (2) Basic grid decoding
[0279] Base mesh decoding can be categorized into intra and inter modes. Base mesh decoding is solely responsible for decoding information representing the base mesh. For intra mode, the output of base mesh decoding is a 3D mesh containing geometric information, connectivity, and possible attribute information, including texture coordinates. For inter mode, base mesh decoding is responsible for decoding motion vectors and identifiers, as well as any intra sub-meshes. The output information is the motion vector, identifier, and intra sub-meshes.
[0280] The decoding of the basic grid information in the intra-frame mode is also the decoding of the basic grid. The basic grid can be decoded using the corresponding grid decoder according to the static grid decoder type indicated by the auxiliary information.
[0281] In inter-frame mode, the motion vector, identifier, and any intra-frame sub-grids are decoded. For intra-frame sub-grids, the grid decoder indicated by the auxiliary information is used for decoding. The motion vector identifier is obtained through entropy decoding, and the motion vector is obtained through inverse prediction and inverse quantization after entropy decoding.
[0282] (3) Segmentation
[0283] The subdivision module operates in the same way as the encoding end subdivision operation, and uses auxiliary information to indicate the number of subdivision iterations in different areas of the current basic grid.
[0284] (4) Displacement vector video decoding
[0285] The displacement vector video code stream is decoded according to the video decoder type indicated in the auxiliary information, including geometric displacement vector video and possible displacement vector videos with different attributes.
[0286] (5) Displacement vector decoding
[0287] Depending on whether the number of levels selected for each dimension of the displacement is marked, there are two ways to decode the displacement vector:
[0288] The first method: unmark the number of selected levels
[0289] In this method, the number of levels selected for each dimensional component of the displacement is not marked. When obtaining the quantized wavelet transform coefficients, all pixel data in the two-dimensional image is first extracted according to the arrangement method used during encoding. The extracted pixel data contains two values: one is the quantized wavelet transform coefficient selected by the encoder's level selection module, and the other is the video padding value during the encoder's two-dimensional arrangement process, which is a value agreed upon between the encoder and decoder. After obtaining the pixel data, the video padding portion of the pixel data can be removed based on the video padding value to obtain the data for each dimensional component of the quantized wavelet transform coefficient selected by the encoder's level selection module. Then, zero padding is performed on each dimensional component of the selected quantized wavelet transform coefficient based on the number of vertices after the decoding base grid is subdivided. Until the number of data for each dimensional component selected is the same as the number of vertices after the decoding base grid is subdivided, the quantized wavelet transform coefficient is obtained. Then, dequantization and inverse transformation operations are performed on the quantized wavelet transform coefficient to restore the corresponding displacement vector.
[0290] The second method: Mark the number of selected levels
[0291] In this method, the number of levels selected for each dimensional component of the displacement is marked. When obtaining the quantized wavelet transform coefficients, the number of levels of the quantized wavelet transform coefficients and the number of data in each level are exactly the same as the subdivided decoding base grid. Therefore, based on the number of levels selected for each dimensional component in the subdivided decoding base grid, the number of data for each dimensional component of the quantized wavelet transform coefficients selected by the encoding-end level selection module can be obtained. Then, based on the number of data for each dimensional component of the wavelet transform coefficients, the selected data is extracted from the two-dimensional image according to the arrangement method used during encoding. Then, based on the number of vertices in the subdivided decoding base grid, zero padding is performed at the end of each dimensional component data extracted until the number of data and the number of vertices are the same, thus obtaining the quantized wavelet transform coefficients. Subsequently, dequantization and inverse transformation operations are performed on the quantized wavelet transform coefficients to restore the corresponding displacement vector.
[0292] (6) Reconstruct the deformed mesh
[0293] After the base mesh and displacement vectors are decoded and reconstructed, the deformed mesh is reconstructed based on these two parts. The corresponding displacement vector is added to each vertex of the subdivided mesh. The calculation method is as follows:
[0294] deformedmesh[i].v[k]=subdivmesh[i].v[k]+displacement[k]
[0295] Among them, subdivmesh[i].v[k] is the geometric coordinates of the k-th vertex after the base mesh of the current frame (index is i) is subdivided, displacement[k] is the spatial displacement vector corresponding to the k-th vertex, and deformedmesh[i].v[k] is the geometric coordinates of the k-th vertex after subdivision and deformation in the current frame.
[0296] In addition, the attribute displacement vector is applied to the corresponding attribute value in the same way. Taking texture coordinates as an example, the calculation formula is as follows:
[0297] deformedmesh[i].vt[k]=subdivmesh[i].vt[k]+attdisplacement[k]
[0298] Among them, subdivmesh[i].vt[k] represents the k-th texture coordinate after the base mesh of the current frame is subdivided, attdisplacement[k] represents the k-th value of the attribute displacement vector, and deformedmesh[i].vt[k] represents the value of the k-th texture coordinate after the subdivision and deformation of the current frame.
[0299] (7) Attribute Graph Decoding
[0300] The attribute graph decoder is responsible for decoding the attribute graph bitstream. The attribute graph is decoded using the video decoder indicated in the auxiliary information. An optional color space conversion is performed on it to obtain the attribute graph that matches the input attribute of the encoder. Figure 1 The final decoded output attribute map is obtained by using the same image format. For multiple attribute map streams, after decoding, each attribute map is mapped to the corresponding texture coordinates according to its identifier.
[0301] After each module is processed, the decoder finally obtains the reconstructed deformed mesh and the corresponding attribute map. Subsequent applications use the reconstructed deformed mesh and attribute map as input for processing.
[0302] Example 4:
[0303] In this embodiment, the syntax structure of the code stream corresponding to the three-dimensional grid is described by taking the syntax structure designed for the V-DMC codec framework as an example:
[0304] Table 1: Atlas sequence parameter set vdmc extension RBSP syntax
[0305]
[0306]
[0307] Here, Atlas sequence parameter set vdmc extension RBSP syntax refers to the Atlas sequence parameter set vdmc extended network raw byte sequence payload (RBSP) syntax. Descriptor refers to a descriptor.
[0308] Asps_vdmc_ext_subdivision_method indicates the subdivision method of the current mesh sequence. Table 2 describes the correspondence between the subdivision method and asps_vdmc_ext_subdivision_method.
[0309] Table 2
[0310] asps_vdmc_ext_subdivision_method Name of subdivision method 0 NONE 1 MIDPOINT
[0311] Among them, Name of subdivision method refers to the name of the subdivision method.
[0312] Asps_vdmc_ext_subdivision_iteration_count indicates the number of mesh subdivision iterations. If asps_vdmc_ext_subdivision_iteration_count does not exist, it is assumed that the value is 0.
[0313] Where asps_vdmc_ext_displacement_coordinate_system indicates the coordinate system identifier of the current grid sequence. Table 3 describes the correspondence between the supported coordinate systems and asps_vdmc_ext_displacement_coordinate_system.
[0314] Table 3
[0315] asps_vdmc_ext_displacement_coordinate_system Name of displacement coordinate system 0 CANNONICAL 1 LOCAL
[0316] Name of displacement coordinate system refers to the name of the displacement coordinate system.
[0317] Wherein, asps_vdmc_ext_transform_method indicates the identifier of the wavelet transform applied to the displacement. Table 4 describes the correspondence between the supported wavelet transform methods and asps_vdmc_ext_transform_method.
[0318] Table 4
[0319]
[0320] Here, Name of transform method refers to the name of the transformation method.
[0321] Among them, asps_vdmc_ext_num_attribute_video indicates the number of attributes sent through the video sub-bitstream.
[0322] asps_vdmc_ext_attribute_type_id[i] indicates the attribute type of the attribute video data unit with index i.
[0323] asps_vdmc_ext_attribute_frame_width[i] represents the atlas frame width of the attribute video data unit with index i.
[0324] asps_vdmc_ext_attribute_frame_height[i] represents the atlas frame height of the attribute video data unit with index i.
[0325] asps_vdmc_ext_attribute_transform_method[i] is the transform identifier that applies to the attribute with index i in the attribute video data unit.
[0326] Table 5 describes the correspondence between the supported transformation methods and asps_vdmc_ext_attribute_transform_method.
[0327] Table 5
[0328] asps_vdmc_ext_attribute_transform_method Name of transform method 0 NONE 1 LINEAR_LIFTING
[0329] Here, Name of transform method refers to the name of the transformation method.
[0330] asps_vdmc_ext_packing_method equal to 0 indicates that the displacement component samples are arranged in ascending order, and asps_vdmc_ext_packing_method equal to 1 indicates that the displacement component samples are arranged in descending order.
[0331] asps_vdmc_ext_1d_displacement_flag equal to 1 indicates that only the normal component of the displacement is present in the displacement vector video. The other two components are inferred to be 0. asps_vdmc_ext_1d_displacement_flag equal to 0 indicates that all three components of the displacement are present in the displacement vector video.
[0332] Table 6: Atlas frame parameter set vdmc extension RBSP syntax
[0333]
[0334] Here, Atlas frame parameter set vdmc extension RBSP syntax refers to the Atlas frame parameter set vdmc extended RBSP syntax.
[0335] Among them, afps_vdmc_ext_overriden_flag is equal to 1, which means that the following parameters exist in the atlas frame parameter set extension:
[0336] afps_vdmc_ext_subdivision_enable_flag;
[0337] afps_vdmc_ext_displacement_coordinate_system_enable_flag;
[0338] afps_vdmc_ext_transform_method_enable_flag;
[0339] afps_vdmc_ext_transform_parameters_enable_flag;
[0340] afps_vdmc_ext_packing_block_size_enable_flag.
[0341] Among them, afps_vdmc_ext_subdivision_enable_flag equal to 1 indicates the presence of the following parameters in the atlas frame parameter set extension: afps_vdmc_ext_subdivision_method and afps_vdmc_ext_subdivision_iteration_count. When afps_vdmc_ext_subdivision_enable_flag is not present, its value is inferred to be 0.
[0342] Among them, afps_vdmc_ext_displacement_coordinate_system_enable_flag is equal to 1, indicating that afps_vdmc_ext_displacement_coordinate_system exists in the atlas frame parameter set extension.
[0343] When afps_vdmc_ext_displacement_coordinate_system_enable_flag is not present, its value is inferred to be 0.
[0344] Among them, afps_vdmc_ext_transform_method_enable_flag is equal to 1 to indicate the presence of afps_vdmc_ext_transform_method in the atlas frame parameter set extension. When afps_vdmc_ext_transform_method_enable_flag is not present, its value is inferred to be equal to 0.
[0345] afps_vdmc_ext_transform_parameters_enable_flag is equal to 1 to indicate the presence of vdmc_lifting_transform_parameters in the atlas frame parameter set extension. When afps_vdmc_ext_transform_parameters_enable_flag is not present, its value is inferred to be equal to 0.
[0346] Wherein, afps_vdmc_ext_subdivision_method indicates the subdivision method of the current frame. When afps_vdmc_ext_subdivision_method does not exist, afps_vdmc_ext_subdivision_method is equal to asps_vdmc_ext_subdivision_method.
[0347] afps_vdmc_ext_subdivision_iteration_count indicates the number of subdivision iterations. When afps_vdmc_ext_subdivision_method is equal to 0 and afps_vdmc_ext_subdivision_iteration_count is not present, afps_vdmc_ext_subdivision_iteration_count is equal to 0.
[0348] When afps_vdmc_ext_subdivision_method is not 0 and afps_vdmc_ext_subdivision_iteration_count does not exist, afps_vdmc_ext_subdivision_iteration_count is equal to asps_vdmc_ext_subdivision_iteration_count.
[0349] afps_vdmc_ext_displacement_coordinate_system indicates the coordinate system identifier of the current frame grid displacement. When afps_vdmc_ext_displacement_coordinate_system does not exist, afps_vdmc_ext_displacement_coordinate_system is equal to asps_vdmc_ext_displacement_coordinate_system.
[0350] afps_vdmc_ext_transform_method indicates an identifier applied to displacement transform. When afps_vdmc_ext_transform_method does not exist, afps_vdmc_ext_transform_method is equal to asps_vdmc_ext_transform_method.
[0351] afps_ext_disp_x_lod_enable_flag indicates the identifier of the number of layers transmitted for the current frame displacement x component. Table 7 describes the correspondence between the number of layers transmitted and afps_ext_disp_x_lod_enable_flag. The maximum value of afps_ext_disp_x_lod_enable_flag depends on afps_vdmc_ext_subdivision_iteration_count.
[0352] Table 7
[0353] afps_ext_disp_x_lod_enable_flag 0 Indicates that all levels of the x component are passed 1 Represents the first layer that passes the x component 2 Represents the first two layers that pass the x component … …
[0354] Among them, afps_ext_disp_y_lod_enable_flag indicates the identifier of the number of layers transmitted for the current frame displacement y component. Table 8 describes the correspondence between the number of layers transmitted and afps_ext_disp_y_lod_enable_flag. The maximum value of afps_ext_disp_y_lod_enable_flag depends on afps_vdmc_ext_subdivision_iteration_count.
[0355] Table 8
[0356] afps_ext_disp_y_lod_enable_flag 0 Indicates that all levels of the y component are passed 1 Represents the first layer that passes the y component 2 Represents the first two layers that pass the y component … …
[0357] Among them, afps_ext_disp_z_lod_enable_flag indicates the identifier of the number of layers of the current frame displacement z component transmitted. Table 9 describes the correspondence between the number of layers transmitted and afps_ext_disp_z_lod_enable_flag. The maximum value of afps_ext_disp_z_lod_enable_flag depends on afps_vdmc_ext_subdivision_iteration_count.
[0358] Table 9
[0359] afps_ext_disp_z_lod_enable_flag 0 Indicates that all levels of the z component are passed 1 Represents the first layer that passes the z component 2 Represents the first two layers that pass the z component … …
[0360] It should be noted that the trellis coding method provided in the embodiments of the present application can be executed by a trellis coding device, or a control module in the trellis coding device for executing the trellis coding method. The trellis coding device provided in the embodiments of the present application is described by taking the trellis coding method executed by the trellis coding device as an example.
[0361] See Figure 20 , Figure 20 This is a structural diagram of a grid coding device provided in an embodiment of the present application. Figure 20 As shown, the grid coding device 300 includes:
[0362] An acquisition module 301 is configured to acquire a base mesh corresponding to a three-dimensional mesh, and perform subdivision processing based on the base mesh to obtain a subdivided mesh;
[0363] A deformation module 302 is configured to perform a deformation operation based on the subdivided grid to obtain a deformed grid;
[0364] a selection module 303 configured to select displacement information to be encoded from the displacement information corresponding to the subdivided mesh, wherein the displacement information corresponding to the subdivided mesh is used to represent the displacement of vertices of the subdivided mesh relative to the deformed mesh, and the number of levels of the displacement information to be encoded is less than or equal to the number of levels of the displacement information corresponding to the subdivided mesh;
[0365] The encoding module 304 is configured to encode the displacement information to be encoded to obtain a first encoding result.
[0366] Optionally, the displacement information corresponding to the subdivided grid includes displacement information of multiple levels, and the selection module includes:
[0367] a determining unit, configured to determine a target level among the multiple levels, wherein the target level is determined based on displacement information representing a first preset value among the displacement information of the multiple levels;
[0368] A selection unit is configured to select displacement information to be encoded from the displacement information corresponding to the subdivided grid based on the target level.
[0369] Optionally, a ratio of the number of displacement information representing a displacement of a first preset value in the displacement information of the target level to the total number of displacement information of the target level meets a preset condition.
[0370] Optionally, the determining unit is specifically configured to:
[0371] Determining a target level from a first level among the multiple levels, wherein a ratio of a number of displacement information items representing a displacement of a first preset value in the displacement information items of the target level to a total number of displacement information items of the target level is greater than the first preset ratio;
[0372] In a case where the target layer is not the first layer, the displacement information to be encoded includes displacement information from the first layer to a layer previous to the target layer.
[0373] Optionally, the determining unit is specifically configured to:
[0374] Determining a target level from the last level of the multiple levels, wherein a ratio of the number of displacement information items representing displacements of the first preset value in the displacement information items of the target level to the total number of displacement information items of the target level is less than or equal to a second preset ratio;
[0375] The displacement information to be encoded includes displacement information from a first level among the multiple levels to the target level.
[0376] Optionally, the displacement information corresponding to the subdivided grid includes: displacement components of a first dimension of N levels, displacement components of a second dimension of the N levels, and displacement components of a third dimension of the N levels;
[0377] The displacement information to be encoded includes: the displacement components of the first dimension of M1 levels among the displacement components of the first dimension of the N levels, the displacement components of the second dimension of M2 levels among the displacement components of the second dimension of the N levels, and the displacement components of the third dimension of M3 levels among the displacement components of the third dimension of the N levels, N is a positive integer, M1, M2 and M3 are all integers greater than or equal to 0, and M1, M2 and M3 are all less than or equal to N.
[0378] Optionally, the code stream corresponding to the three-dimensional grid includes the first encoding result and a second encoding result, and the second encoding result includes the encoding result of the first flag bit;
[0379] When the first flag bit is a second preset value, the second encoding result also includes an encoding result of the first parameter; or,
[0380] When the first flag bit is a third preset value, the second encoding result also includes an encoding result of the first parameter, an encoding result of the second parameter, and an encoding result of the third parameter.
[0381] The first parameter is used to characterize the M1 levels, the second parameter is used to characterize the M2 levels, and the third parameter is used to characterize the M3 levels.
[0382] Optionally, the encoding module is specifically configured to:
[0383] Arranging the displacement information to be encoded into a video frame;
[0384] Performing padding processing on the arranged video frames to obtain padded video frames;
[0385] The padded video frame is encoded to obtain a first encoding result.
[0386] Optionally, the encoding module is specifically configured to:
[0387] Entropy coding is performed on the displacement information to be encoded to obtain a first coding result.
[0388] The trellis coding device in the embodiments of the present application can be a device, a device or electronic device with an operating system, or a component, integrated circuit, or chip in a terminal. The device or electronic device can be a mobile terminal or a non-mobile terminal. For example, the mobile terminal can include but is not limited to the types of terminals listed above, and the non-mobile terminal can be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), an ATM or a self-service machine, etc., which are not specifically limited in the embodiments of the present application.
[0389] The grid coding device provided in the embodiment of the present application can achieve Figure 5 The various processes implemented by the method embodiment achieve the same technical effect and are not described here again to avoid repetition.
[0390] It should be noted that the trellis decoding method provided in the embodiments of the present application can be executed by a trellis decoding device, or a control module within the trellis decoding device that is configured to execute the trellis decoding method. The trellis decoding device provided in the embodiments of the present application is described using the trellis decoding method executed by the trellis decoding device as an example.
[0391] See Figure 21 , Figure 21 This is a structural diagram of a grid decoding device provided in an embodiment of the present application. Figure 21 As shown, the grid decoding device 400 includes:
[0392] A first decoding module 401 is configured to decode a first encoding result in a code stream corresponding to a three-dimensional grid to obtain decoded displacement information;
[0393] A second decoding module 402 is configured to decode the third encoding result in the code stream to obtain a basic grid, and subdivide the basic grid to obtain a subdivided grid;
[0394] A filling module 403 is configured to perform filling processing on the decoded displacement information to obtain filled displacement information, wherein the number of levels of the filled displacement information is greater than or equal to the number of levels of the decoded displacement information;
[0395] The reconstruction module 404 is configured to perform reconstruction processing based on the filled displacement information and the subdivided grid to obtain a reconstructed grid.
[0396] Optionally, the first decoding module is specifically configured to:
[0397] Decoding a first encoding result in a code stream corresponding to the three-dimensional grid to obtain a video frame;
[0398] Decoded displacement information is obtained from the video frame, where the decoded displacement information is displacement information in the video frame excluding a preset video filling value.
[0399] Optionally, the device further comprises:
[0400] a third decoding module, configured to decode the second encoding result in the code stream to obtain a target parameter, wherein the target parameter is used to represent the number of levels of displacement information;
[0401] The first decoding module is specifically configured to:
[0402] Decoding a first encoding result in a code stream corresponding to the three-dimensional grid to obtain a video frame;
[0403] Decoded displacement information is obtained from the video frame based on the target parameter.
[0404] Optionally, the first decoding module is specifically configured to:
[0405] Entropy decoding is performed on a first encoding result in a code stream corresponding to the three-dimensional grid to obtain decoded displacement information.
[0406] Optionally, the filling module is specifically used to:
[0407] The decoded displacement information is padded based on the number of vertices of the subdivided mesh to obtain padded displacement information.
[0408] Optionally, the padded displacement information includes N levels of first-dimensional displacement components, N levels of second-dimensional displacement components, and N levels of third-dimensional displacement components;
[0409] The decoded displacement information includes the displacement components of the first dimension of M1 levels among the displacement components of the first dimension of the N levels, the displacement components of the second dimension of M2 levels among the displacement components of the second dimension of the N levels, and the displacement components of the third dimension of M3 levels among the displacement components of the third dimension of the N levels, where N is a positive integer, M1, M2 and M3 are all integers greater than or equal to 0, and M1, M2 and M3 are all less than or equal to N.
[0410] Optionally, the device further comprises:
[0411] a fourth decoding module, configured to decode the second encoding result in the code stream to obtain a first flag bit and a target parameter;
[0412] Wherein, when the first flag bit is the second preset value, the target parameter includes the first parameter; or,
[0413] When the first flag bit is the third preset value, the target parameter includes the first parameter, the second parameter and the third parameter.
[0414] The first parameter is used to characterize the M1 levels, the second parameter is used to characterize the M2 levels, and the third parameter is used to characterize the M3 levels.
[0415] The trellis decoding device in the embodiments of the present application can be a device, a device or electronic device with an operating system, or a component, integrated circuit, or chip in a terminal. The device or electronic device can be a mobile terminal or a non-mobile terminal. For example, the mobile terminal can include but is not limited to the types of terminals listed above, and the non-mobile terminal can be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine (ATM), or a self-service machine, etc., which are not specifically limited in the embodiments of the present application.
[0416] The grid decoding device provided in the embodiment of the present application can achieve Figure 8 The various processes implemented by the method embodiment achieve the same technical effect and are not described here again to avoid repetition.
[0417] Alternatively, as Figure 22 As shown, an embodiment of the present application further provides a communication device 500, including a processor 501 and a memory 502, wherein the memory 502 stores a program or instruction that can be run on the processor 501. For example, when the communication device 500 is an encoding end device, the program or instruction, when executed by the processor 501, implements the various steps of the above-mentioned grid encoding method embodiment and can achieve the same technical effect. When the communication device 500 is a decoding end device, the program or instruction, when executed by the processor 501, implements the various steps of the above-mentioned grid decoding method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0418] The present application also provides a terminal including a processor and a communication interface, wherein the processor is configured to: obtain a base mesh corresponding to a three-dimensional mesh, perform subdivision processing based on the base mesh to obtain a subdivided mesh; perform a deformation operation based on the subdivided mesh to obtain a deformed mesh; select displacement information to be encoded from displacement information corresponding to the subdivided mesh, wherein the displacement information corresponding to the subdivided mesh is used to represent the displacement of vertices of the subdivided mesh relative to the deformed mesh, wherein the number of levels of the displacement information to be encoded is less than or equal to the number of levels of the displacement information corresponding to the subdivided mesh; and encode the displacement information to be encoded to obtain a first encoding result. Alternatively, the processor is configured to: decode a first encoding result in a code stream corresponding to the three-dimensional mesh to obtain decoded displacement information; decode a third encoding result in the code stream to obtain a base mesh, and perform subdividing processing on the base mesh to obtain a subdivided mesh; perform padding processing on the decoded displacement information to obtain padded displacement information, wherein the number of levels of the padded displacement information is greater than or equal to the number of levels of the decoded displacement information; and perform reconstruction processing based on the padded displacement information and the subdivided mesh to obtain a reconstructed mesh.
[0419] This terminal embodiment corresponds to the above-mentioned trellis encoding method or trellis decoding method embodiment. Each implementation process and implementation method of the above-mentioned trellis encoding method or trellis decoding method embodiment can be applied to this terminal embodiment and can achieve the same technical effect. Specifically, Figure 23 A schematic diagram of the hardware structure of a terminal for implementing an embodiment of the present application.
[0420] The terminal 600 includes but is not limited to: a radio frequency unit 601, a network module 602, an audio output unit 603, an input unit 604, a sensor 605, a display unit 606, a user input unit 607, an interface unit 608, a memory 609 and at least some of the components of the processor 610.
[0421] Those skilled in the art will understand that the terminal 600 may also include a power supply (such as a battery) to power each component, and the power supply may be logically connected to the processor 610 through a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system. Figure 23 The terminal structure shown in the figure does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be repeated here.
[0422] It should be understood that in an embodiment of the present application, the input unit 604 may include a graphics processing unit (GPU) 6041 and a microphone 6042, and the GPU 6041 processes image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 606 may include a display panel 6061, and the display panel 6061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 607 includes a touch panel 6071 and at least one of other input devices 6072. The touch panel 6071 is also called a touch screen. The touch panel 6071 may include two parts: a touch detection device and a touch controller. Other input devices 6072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.
[0423] In the embodiment of the present application, after receiving downlink data from a network-side device, the radio frequency unit 601 may transmit the data to the processor 610 for processing. Furthermore, the radio frequency unit 601 may send uplink data to the network-side device. Typically, the radio frequency unit 601 includes, but is not limited to, an antenna, an amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, and the like.
[0424] The memory 609 can be used to store software programs or instructions and various data. The memory 609 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 609 may include a volatile memory or a non-volatile memory, or the memory 609 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct RAM bus random access memory (DRRAM). The memory 609 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
[0425] Processor 610 may include one or more processing units. Optionally, processor 610 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 610.
[0426] Wherein, when the terminal is an encoding terminal device:
[0427] The processor 610 is configured to:
[0428] Obtaining a base grid corresponding to the three-dimensional grid, and performing subdivision processing based on the base grid to obtain a subdivided grid;
[0429] Performing a deformation operation based on the subdivided grid to obtain a deformed grid;
[0430] Selecting displacement information to be encoded from the displacement information corresponding to the subdivided grid, where the displacement information corresponding to the subdivided grid is used to represent the displacement of vertices of the subdivided grid relative to the deformed grid, and the number of levels of the displacement information to be encoded is less than or equal to the number of levels of the displacement information corresponding to the subdivided grid;
[0431] The displacement information to be encoded is encoded to obtain a first encoding result.
[0432] Optionally, the displacement information corresponding to the subdivided grid includes displacement information of multiple levels, and the processor 610 is specifically configured to:
[0433] Determining a target level among the multiple levels, wherein the target level is determined based on displacement information representing a first preset value among the displacement information of the multiple levels;
[0434] Displacement information to be encoded is selected from the displacement information corresponding to the subdivided grid based on the target level.
[0435] Optionally, a ratio of the number of displacement information representing a displacement of a first preset value in the displacement information of the target level to the total number of displacement information of the target level meets a preset condition.
[0436] Optionally, the processor 610 is specifically configured to:
[0437] Determining a target level from a first level among the multiple levels, wherein a ratio of a number of displacement information items representing a displacement of a first preset value in the displacement information items of the target level to a total number of displacement information items of the target level is greater than the first preset ratio;
[0438] In a case where the target layer is not the first layer, the displacement information to be encoded includes displacement information from the first layer to a layer previous to the target layer.
[0439] Optionally, the processor 610 is specifically configured to:
[0440] Determining a target level from the last level of the multiple levels, wherein a ratio of the number of displacement information items representing displacements of the first preset value in the displacement information items of the target level to the total number of displacement information items of the target level is less than or equal to a second preset ratio;
[0441] The displacement information to be encoded includes displacement information from a first level among the multiple levels to the target level.
[0442] Optionally, the displacement information corresponding to the subdivided grid includes: displacement components of a first dimension of N levels, displacement components of a second dimension of the N levels, and displacement components of a third dimension of the N levels;
[0443] The displacement information to be encoded includes: the displacement components of the first dimension of M1 levels among the displacement components of the first dimension of the N levels, the displacement components of the second dimension of M2 levels among the displacement components of the second dimension of the N levels, and the displacement components of the third dimension of M3 levels among the displacement components of the third dimension of the N levels, N is a positive integer, M1, M2 and M3 are all integers greater than or equal to 0, and M1, M2 and M3 are all less than or equal to N.
[0444] Optionally, the code stream corresponding to the three-dimensional grid includes the first encoding result and a second encoding result, and the second encoding result includes the encoding result of the first flag bit;
[0445] When the first flag bit is a second preset value, the second encoding result also includes an encoding result of the first parameter;
[0446] When the first flag bit is a third preset value, the second encoding result also includes an encoding result of the first parameter, an encoding result of the second parameter, and an encoding result of the third parameter.
[0447] The first parameter is used to characterize the M1 levels, the second parameter is used to characterize the M2 levels, and the third parameter is used to characterize the M3 levels.
[0448] Optionally, the processor 610 is specifically configured to:
[0449] Arranging the displacement information to be encoded into a video frame;
[0450] Performing padding processing on the arranged video frames to obtain padded video frames;
[0451] The padded video frame is encoded to obtain a first encoding result.
[0452] Optionally, the processor 610 is specifically configured to:
[0453] Entropy coding is performed on the displacement information to be encoded to obtain a first coding result.
[0454] Wherein, when the terminal is a decoding end device:
[0455] The processor 610 is configured to:
[0456] Decoding a first encoding result in a code stream corresponding to the three-dimensional grid to obtain decoded displacement information;
[0457] Decoding the third encoding result in the code stream to obtain a basic grid, and subdividing the basic grid to obtain a subdivided grid;
[0458] Performing padding processing on the decoded displacement information to obtain padded displacement information, wherein the number of levels of the padded displacement information is greater than or equal to the number of levels of the decoded displacement information;
[0459] Reconstruction processing is performed based on the filled displacement information and the subdivided grid to obtain a reconstructed grid.
[0460] Optionally, the processor 610 is specifically configured to:
[0461] Decoding a first encoding result in a code stream corresponding to the three-dimensional grid to obtain a video frame;
[0462] Decoded displacement information is obtained from the video frame, where the decoded displacement information is displacement information in the video frame excluding a preset video filling value.
[0463] Optionally, the processor 610 is specifically configured to:
[0464] Decoding the second encoding result in the code stream to obtain a target parameter, where the target parameter is used to represent the number of levels of displacement information;
[0465] Decoding a first encoding result in a code stream corresponding to the three-dimensional grid to obtain a video frame;
[0466] Decoded displacement information is obtained from the video frame based on the target parameter.
[0467] Optionally, the processor 610 is specifically configured to:
[0468] Entropy decoding is performed on a first encoding result in a code stream corresponding to the three-dimensional grid to obtain decoded displacement information.
[0469] Optionally, performing padding processing on the decoded displacement information to obtain padded displacement information includes:
[0470] The decoded displacement information is padded based on the number of vertices of the subdivided mesh to obtain padded displacement information.
[0471] Optionally, the padded displacement information includes N levels of first-dimensional displacement components, N levels of second-dimensional displacement components, and N levels of third-dimensional displacement components;
[0472] The decoded displacement information includes the displacement components of the first dimension of M1 levels among the displacement components of the first dimension of the N levels, the displacement components of the second dimension of M2 levels among the displacement components of the second dimension of the N levels, and the displacement components of the third dimension of M3 levels among the displacement components of the third dimension of the N levels, where N is a positive integer, M1, M2 and M3 are all integers greater than or equal to 0, and M1, M2 and M3 are all less than or equal to N.
[0473] Optionally, the processor 610 is specifically configured to:
[0474] Decoding the second encoding result in the code stream to obtain a first flag bit and a target parameter;
[0475] Wherein, when the first flag bit is a second preset value, the target parameter includes the first parameter;
[0476] When the first flag bit is the third preset value, the target parameter includes the first parameter, the second parameter and the third parameter.
[0477] The first parameter is used to characterize the M1 levels, the second parameter is used to characterize the M2 levels, and the third parameter is used to characterize the M3 levels.
[0478] Specifically, the terminal of the embodiment of the present application further includes: instructions or programs stored in the memory 609 and executable on the processor 610, and the processor 610 calls the instructions or programs in the memory 609 to execute Figure 16 or Figure 17 The methods executed by the modules shown achieve the same technical effects, so they will not be described here to avoid repetition.
[0479] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned grid encoding method embodiment are implemented, or when the program or instruction is executed by a processor, the various processes of the above-mentioned grid decoding method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0480] The processor is the processor in the terminal described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.
[0481] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned grid encoding method embodiment, or to implement the various processes of the above-mentioned grid decoding method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0482] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0483] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0484] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0485] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A grid coding method, characterized in that: include: Obtaining a base grid corresponding to the three-dimensional grid, and performing subdivision processing based on the base grid to obtain a subdivided grid; Performing a deformation operation based on the subdivided grid to obtain a deformed grid; Selecting displacement information to be encoded from the displacement information corresponding to the subdivided grid, where the displacement information corresponding to the subdivided grid is used to represent the displacement of vertices of the subdivided grid relative to the deformed grid, and the number of levels of the displacement information to be encoded is less than or equal to the number of levels of the displacement information corresponding to the subdivided grid; Encoding the displacement information to be encoded to obtain a first encoding result; The displacement information corresponding to the subdivided grid includes displacement information of multiple levels, and selecting the displacement information to be encoded from the displacement information corresponding to the subdivided grid includes: Determining a target level among the multiple levels, wherein the target level is determined based on displacement information representing a first preset value among the displacement information of the multiple levels; Displacement information to be encoded is selected from the displacement information corresponding to the subdivided grid based on the target level.
2. The method according to claim 1, characterized in that The ratio of the number of displacement information representing a displacement of a first preset value in the displacement information of the target level to the total number of displacement information of the target level meets a preset condition.
3. The method according to claim 1 or 2, characterized in that Determining a target level among the multiple levels includes: Determining a target level from a first level among the multiple levels, wherein a ratio of a number of displacement information items representing a displacement of a first preset value in the displacement information items of the target level to a total number of displacement information items of the target level is greater than the first preset ratio; In a case where the target layer is not the first layer, the displacement information to be encoded includes displacement information from the first layer to a layer previous to the target layer.
4. The method according to claim 1 or 2, characterized in that Determining a target level among the multiple levels includes: Determining a target level from the last level of the multiple levels, wherein a ratio of the number of displacement information items representing displacements of the first preset value in the displacement information items of the target level to the total number of displacement information items of the target level is less than or equal to a second preset ratio; The displacement information to be encoded includes displacement information from a first level among the multiple levels to the target level.
5. The method according to any one of claims 1 to 4, characterized in that The displacement information corresponding to the subdivided grid includes: N levels of displacement components of the first dimension, N levels of displacement components of the second dimension, and N levels of displacement components of the third dimension; The displacement information to be encoded includes: the displacement components of the first dimension of M1 levels among the displacement components of the first dimension of the N levels, the displacement components of the second dimension of M2 levels among the displacement components of the second dimension of the N levels, and the displacement components of the third dimension of M3 levels among the displacement components of the third dimension of the N levels, N is a positive integer, M1, M2 and M3 are all integers greater than or equal to 0, and M1, M2 and M3 are all less than or equal to N.
6. The method according to claim 5, characterized in that The code stream corresponding to the three-dimensional grid includes the first encoding result and the second encoding result, and the second encoding result includes the encoding result of the first flag bit; When the first flag bit is a second preset value, the second encoding result also includes an encoding result of the first parameter; or, When the first flag bit is a third preset value, the second encoding result also includes an encoding result of the first parameter, an encoding result of the second parameter, and an encoding result of the third parameter. The first parameter is used to characterize the M1 levels, the second parameter is used to characterize the M2 levels, and the third parameter is used to characterize the M3 levels.
7. The method according to any one of claims 1 to 6, characterized in that The encoding of the displacement information to be encoded to obtain a first encoding result includes: Arranging the displacement information to be encoded into a video frame; Performing padding processing on the arranged video frames to obtain padded video frames; The padded video frame is encoded to obtain a first encoding result.
8. The method according to any one of claims 1 to 6, characterized in that The encoding of the displacement information to be encoded to obtain a first encoding result includes: Entropy coding is performed on the displacement information to be encoded to obtain a first coding result.
9. A grid decoding method, characterized in that: include: Decoding a first encoding result in a code stream corresponding to the three-dimensional grid to obtain decoded displacement information; Decoding the third encoding result in the code stream to obtain a basic grid, and subdividing the basic grid to obtain a subdivided grid; Performing padding processing on the decoded displacement information to obtain padded displacement information, wherein the number of levels of the padded displacement information is greater than or equal to the number of levels of the decoded displacement information; Reconstruction processing is performed based on the filled displacement information and the subdivided grid to obtain a reconstructed grid.
10. The method according to claim 9, characterized in that The decoding of the first encoding result in the code stream corresponding to the three-dimensional grid to obtain decoded displacement information includes: Decoding a first encoding result in a code stream corresponding to the three-dimensional grid to obtain a video frame; Decoded displacement information is obtained from the video frame, where the decoded displacement information is displacement information in the video frame excluding a preset video filling value.
11. The method according to claim 9, characterized in that The method further comprises: Decoding the second encoding result in the code stream to obtain a target parameter, where the target parameter is used to represent the number of levels of displacement information; The decoding of the first encoding result in the code stream corresponding to the three-dimensional grid to obtain decoded displacement information includes: Decoding a first encoding result in a code stream corresponding to the three-dimensional grid to obtain a video frame; Decoded displacement information is obtained from the video frame based on the target parameter.
12. The method according to claim 9, characterized in that The decoding of the first encoding result in the code stream corresponding to the three-dimensional grid to obtain decoded displacement information includes: Entropy decoding is performed on a first encoding result in a code stream corresponding to the three-dimensional grid to obtain decoded displacement information.
13. The method according to any one of claims 9 to 12, characterized in that The step of performing padding processing on the decoded displacement information to obtain the padded displacement information includes: The decoded displacement information is padded based on the number of vertices of the subdivided mesh to obtain padded displacement information.
14. The method according to any one of claims 9 to 13, characterized in that The padded displacement information includes N levels of first-dimensional displacement components, N levels of second-dimensional displacement components, and N levels of third-dimensional displacement components; The decoded displacement information includes the displacement components of the first dimension of M1 levels among the displacement components of the first dimension of the N levels, the displacement components of the second dimension of M2 levels among the displacement components of the second dimension of the N levels, and the displacement components of the third dimension of M3 levels among the displacement components of the third dimension of the N levels, where N is a positive integer, M1, M2 and M3 are all integers greater than or equal to 0, and M1, M2 and M3 are all less than or equal to N.
15. The method according to claim 14, characterized in that The method further comprises: Decoding the second encoding result in the code stream to obtain a first flag bit and a target parameter; Wherein, when the first flag bit is the second preset value, the target parameter includes the first parameter; or, When the first flag bit is the third preset value, the target parameter includes the first parameter, the second parameter and the third parameter. The first parameter is used to characterize the M1 levels, the second parameter is used to characterize the M2 levels, and the third parameter is used to characterize the M3 levels.
16. A grid coding device, characterized in that: The device comprises: An acquisition module is used to acquire a basic grid corresponding to the three-dimensional grid, and perform subdivision processing based on the basic grid to obtain a subdivided grid; A deformation module, configured to perform a deformation operation based on the subdivided grid to obtain a deformed grid; a selection module configured to select displacement information to be encoded from the displacement information corresponding to the subdivided grid, wherein the displacement information corresponding to the subdivided grid is used to represent the displacement of vertices of the subdivided grid relative to the deformed grid, and the number of levels of the displacement information to be encoded is less than or equal to the number of levels of the displacement information corresponding to the subdivided grid; an encoding module, configured to encode the displacement information to be encoded to obtain a first encoding result; The displacement information corresponding to the subdivided grid includes displacement information of multiple levels, and the selection module includes: a determining unit, configured to determine a target level among the multiple levels, wherein the target level is determined based on displacement information representing a first preset value among the displacement information of the multiple levels; A selection unit is configured to select displacement information to be encoded from the displacement information corresponding to the subdivided grid based on the target level.
17. A trellis decoding device, characterized in that: The device comprises: A first decoding module is used to decode a first encoding result in a code stream corresponding to the three-dimensional grid to obtain decoded displacement information; a second decoding module, configured to decode the third encoding result in the code stream to obtain a basic grid, and subdivide the basic grid to obtain a subdivided grid; a filling module, configured to perform filling processing on the decoded displacement information to obtain filled displacement information, wherein the number of levels of the filled displacement information is greater than or equal to the number of levels of the decoded displacement information; The reconstruction module is used to perform reconstruction processing based on the filled displacement information and the subdivided grid to obtain a reconstructed grid.
18. A terminal, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the grid encoding method as described in any one of claims 1 to 8; or, when executed by the processor, the program or instruction, when executed by the processor, implements the steps of the grid decoding method as described in any one of claims 9 to 15.
19. A readable storage medium, characterized in that The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the grid encoding method as described in any one of claims 1 to 8 are implemented, or when the program or instruction is executed by the processor, the steps of the grid decoding method as described in any one of claims 9 to 15 are implemented.
Citation Information
Patent Citations
Texture image compression method based on geometric information of three-dimensional model
CN105141970A
Urban three-dimensional space grid compression coding method and device and terminal equipment
CN111260784A