Encoding and decoding method, code stream, encoder, decoder and storage medium
Patent Information
- Application Number
- CN202380096894.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-30
- Publication Date
- 2025-11-14
AI Technical Summary
The existing dynamic grid coding technology increases the encoding code rate of the shift coefficient in inter-coding, resulting in a decrease in grid compression performance and unable to effectively reduce the encoding code rate of the shift coefficient.
By determining the basic grid of the current image at the encoding and decoding ends, and adaptively encoding based on the geometric position information of the initial grid and the shift coefficient in the reference image, the encoding of the shift coefficient in the current image is skipped. , the shift coefficients in the reference image are used to reconstruct the geometric position information of the grid, and the code stream of the shift coefficient is reduced.
While ensuring the quality of grid reconstruction, the encoding coding rate of the shift coefficient is reduced, the geometric coding efficiency of the grid is improved, and the encoding and decoding performance is improved.
Smart Images

Figure CN120958829A_ABST
Abstract
Description
Coding and decoding method, code stream, encoder, decoder and storage medium Technical Field
[0001] The embodiments of the present application relate to the field of dynamic grid coding and decoding technology, and in particular to a coding and decoding method, a bit stream, an encoder, a decoder, and a storage medium. Background Art
[0002] In the standard reference software for Dynamic Mesh Coding provided by the Moving Picture Experts Group (MPEG), most of them use the geometric information of the initial mesh obtained by basic mesh division and the shift coefficient to reconstruct and restore the geometric position information of the current mesh.
[0003] If the current image can be encoded using an inter-frame coding scheme, that is, if a reference image exists for the current image, the encoder can use the base grid of the reference image to determine the grid information for the current image. However, existing technical solutions are not perfect, increasing the encoding bit rate of the shift coefficients and thus reducing grid compression performance.
[0004] Summary of the Invention
[0005] The embodiments of the present application provide a coding and decoding method, a code stream, an encoder, a decoder, and a storage medium, which can reduce the coding rate of the shift coefficients while ensuring the quality of grid reconstruction, thereby improving the geometric coding efficiency of the grid.
[0006] The technical solution of the embodiment of the present application can be implemented as follows:
[0007] In a first aspect, an embodiment of the present application provides a decoding method, applied to a decoder, the method comprising:
[0008] Determine the base grid of the current image;
[0009] Subdivide the base grid and determine the geometric position information of the initial grid of the current layer in the current image;
[0010] Decoding the code stream to determine first syntax identification information;
[0011] When the first syntax identification information indicates that the shift coefficient of the current layer in the current image uses the first decoding mode, determining the shift coefficient of the current layer in the reference image;
[0012] The geometric position information of the reconstructed grid of the current layer in the current image is determined according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image.
[0013] In a second aspect, an embodiment of the present application provides an encoding method, applied to an encoder, the method comprising:
[0014] Determine the base grid of the current image according to the original grid of the current image;
[0015] Subdivide the base grid and determine the geometric position information of the initial grid of the current layer in the current image;
[0016] determining a shift coefficient of the current layer in the reference image, and determining geometric position information of a first reconstructed grid of the current layer in the current image based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image;
[0017] According to the geometric position information of the first reconstructed grid and the geometric position information of the original grid, it is determined whether to perform encoding processing on the shift coefficients of the current layer in the current image.
[0018] In a third aspect, an embodiment of the present application provides a code stream, which is generated by bit encoding based on information to be encoded; wherein the information to be encoded includes at least one of the following:
[0019] The value of the first syntax identification information, the value of the second syntax identification information, the value of the third syntax identification information, the value of the fourth syntax identification information, the reference image index of the current image, the shift coefficient of the current layer in the current image, and the mapping indication information of the current layer in the current image;
[0020] Among them, the second syntax identification information is used to indicate whether the shift coefficient of the current sequence enables the first coding mode, the third syntax identification information is used to indicate whether the shift coefficient of the current image enables the first coding mode, the first syntax identification information is used to indicate whether the shift coefficient of the current layer in the current image uses the first coding mode, and the fourth syntax identification information is used to indicate whether the basic grid of the current image uses inter-frame processing; and the current sequence includes the current image, and the LOD layer divided by the current image includes the current layer.
[0021] In a fourth aspect, an embodiment of the present application provides an encoder, comprising a first determining unit, a first subdivision unit, a first reconstruction unit, and an encoding unit, wherein:
[0022] A first determining unit is configured to determine a base grid of the current image according to the original grid of the current image;
[0023] a first subdivision unit configured to subdivide the base grid and determine geometric position information of an initial grid of a current layer in a current image;
[0024] a first reconstruction unit configured to determine a shift coefficient of the current layer in the reference image, and determine geometric position information of a first reconstructed grid of the current layer in the current image based on geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image;
[0025] The encoding unit is configured to determine whether to perform encoding processing on the shift coefficients of the current layer in the current image according to the geometric position information of the first reconstructed grid and the geometric position information of the original grid.
[0026] In a fifth aspect, an embodiment of the present application provides an encoder, comprising a first memory and a first processor; wherein,
[0027] a first memory for storing a computer program capable of running on the first processor;
[0028] The first processor is configured to execute the method according to the second aspect when running a computer program.
[0029] In a sixth aspect, an embodiment of the present application provides a decoder, comprising a second determining unit, a second subdivision unit, a decoding unit, and a second reconstruction unit, wherein:
[0030] a second determining unit configured to determine a base grid of the current image;
[0031] a second subdivision unit configured to subdivide the base grid and determine geometric position information of an initial grid of a current layer in a current image;
[0032] a decoding unit configured to decode the code stream, determine first syntax identification information; and determine a shift coefficient of the current layer in the reference image when the first syntax identification information indicates that the shift coefficient of the current layer in the current image uses the first decoding mode;
[0033] The second reconstruction unit is configured to determine the geometric position information of the reconstructed grid of the current layer in the current image according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image.
[0034] In a seventh aspect, an embodiment of the present application provides a decoder, the decoder comprising a second memory and a second processor; wherein,
[0035] a second memory for storing a computer program capable of running on the second processor;
[0036] The second processor is configured to execute the method according to the first aspect when running a computer program.
[0037] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed, implements the method described in the first aspect or the method described in the second aspect.
[0038] Embodiments of the present application provide a coding and decoding method, a bitstream, an encoder, a decoder, and a storage medium. At the encoding end, a base grid of the current image is determined based on the original grid of the current image; the base grid is subdivided to determine geometric position information of the initial grid of the current layer in the current image; a shift coefficient of the current layer in the reference image is determined, and geometric position information of a first reconstructed grid of the current layer in the current image is determined based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image; and whether to encode the shift coefficient of the current layer in the current image is determined based on the geometric position information of the first reconstructed grid and the geometric position information of the original grid. At the decoding end, a base grid of the current image is determined; the base grid is subdivided to determine geometric position information of the initial grid of the current layer in the current image; the bitstream is decoded, and first syntax identification information is determined; when the first syntax identification information indicates that the shift coefficient of the current layer in the current image uses a first decoding mode, the shift coefficient of the current layer in the reference image is determined; and the geometric position information of the reconstructed grid of the current layer in the current image is determined based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image. In this way, based on the geometric position information of the initial grid after the basic grid of the current image is subdivided and the geometric position information of the first reconstructed grid obtained by the shift coefficient in the reference image, the adaptive encoding of the shift coefficient in the current image can be determined; if the encoding of the shift coefficient in the current image is skipped, the shift coefficient in the current image does not need to be transmitted in the code stream at this time. The decoding end can determine the geometric position information of the first reconstructed grid based on the geometric position information of the initial grid and the shift coefficient in the reference image, which not only reduces the code stream of the shift coefficient, but also ensures the reconstructed geometric quality of the grid midpoint, thereby further improving the geometric information quality of the grid midpoint, and thus improving the encoding and decoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] FIG1A is a schematic diagram of a three-dimensional grid image 1;
[0040] FIG1B is a partially enlarged schematic diagram of a three-dimensional grid image;
[0041] Figure 2 is a schematic diagram of the connection method of the three-dimensional grid;
[0042] FIG3A is a second schematic diagram of a three-dimensional grid image;
[0043] FIG3B is a schematic diagram of a grid data storage format;
[0044] FIG3C is a schematic diagram of properties of a three-dimensional grid image;
[0045] FIG4 is a schematic diagram showing the composition of the overall framework of grid coding;
[0046] FIG5A is a schematic diagram of preprocessing of a two-dimensional curve;
[0047] FIG5B is a schematic diagram of generating a shift coefficient;
[0048] FIG6A is a first schematic diagram of quantization processing of grid geometric position information;
[0049] FIG6B is a second schematic diagram of quantization processing of grid geometric position information;
[0050] FIG7A is a schematic diagram of coding of the connection relationship of triangular facets;
[0051] FIG7B is a schematic diagram of encoding geometric position information;
[0052] FIG7C is a schematic diagram of texture coordinate encoding;
[0053] FIG8 is a schematic diagram showing the basic principle of the shift coefficient;
[0054] FIG9 is a schematic diagram of encoding of a shift coefficient mapped to a two-dimensional image;
[0055] FIG10 is a schematic diagram of encoding of inter-frame geometric position information;
[0056] FIG11A is a schematic diagram showing the composition of an intra-frame coding framework;
[0057] FIG11B is a schematic diagram showing the composition of an inter-frame coding framework;
[0058] FIG12A is a schematic diagram showing the composition of an intra-frame decoding framework;
[0059] FIG12B is a schematic diagram showing the composition of an inter-frame decoding framework;
[0060] FIG13A is a schematic diagram of iterative subdivision of a basic grid;
[0061] FIG13B is a schematic diagram of an LOD space structure;
[0062] FIG14 is a schematic diagram of coefficient reorganization for quantized coefficients;
[0063] FIG15 is a schematic diagram showing the basic principle of reconstructing and restoring geometric position information;
[0064] FIG16 is a schematic diagram of a mesh architecture of a codec provided in an embodiment of the present application;
[0065] FIG17 is a flowchart diagram 1 of a decoding method provided in an embodiment of the present application;
[0066] FIG18 is a second flow chart of a decoding method provided in an embodiment of the present application;
[0067] FIG19 is a flowchart diagram 1 of an encoding method provided in an embodiment of the present application;
[0068] FIG20 is a second flow chart of an encoding method provided in an embodiment of the present application;
[0069] FIG21 is a schematic diagram showing another basic principle of reconstructing and restoring geometric position information provided by an embodiment of the present application;
[0070] FIG22 is a schematic diagram of the structure of an encoder provided in an embodiment of the present application;
[0071] FIG23 is a schematic diagram of a specific hardware structure of an encoder provided in an embodiment of the present application;
[0072] FIG24 is a schematic diagram of the structure of a decoder provided in an embodiment of the present application;
[0073] FIG25 is a schematic diagram of a specific hardware structure of a decoder provided in an embodiment of the present application;
[0074] FIG26 is a schematic diagram of the composition structure of a coding and decoding system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0075] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below with reference to the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present application.
[0076] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0077] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0078] It should also be pointed out that the terms "first\second\third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0079] It should be noted that it is possible to decode and synthesize different data format bitstreams within the same video scene. These can include at least image format, point cloud format, and mesh format. In this way, real-time immersive video interaction services can be provided for multiple data formats (e.g., mesh, point cloud, image, etc.) from different sources.
[0080] In embodiments of the present application, the data format-based approach allows for independent processing at the bitstream level of the data format. This means that, similar to tiles or slices in video encoding, different data formats in this scenario can be encoded independently, enabling independent encoding and decoding based on the data format.
[0081] Generally speaking, 3D animation content uses a keyframe-based representation method, that is, each frame is a static mesh. Static meshes at different times have the same topological structure and different geometric structures. However, the amount of data of 3D dynamic meshes represented based on keyframes is extremely large, so how to effectively store, transmit and draw them has become a problem faced by the development of 3D dynamic meshes. In addition, the spatial scalability of the mesh needs to be supported for different user terminals (computers, notebooks, portable devices, mobile phones); different mesh bandwidths (broadband, narrowband, wireless) need to support the quality scalability of the mesh. Therefore, 3D dynamic mesh compression is a very critical issue.
[0082] A 3D mesh is the surface of a 3D object composed of countless polygons in space. Polygons are composed of vertices and edges. Figure 1A shows a 3D mesh image, and Figure 1B shows a partially enlarged schematic diagram of the 3D mesh image. Figures 1A and 1B show that the mesh surface is composed of closed polygons.
[0083] A two-dimensional image has information expressed at every pixel point and is distributed regularly, so there is no need to record its position information separately. However, the distribution of vertices in the mesh in three-dimensional space is random and irregular, and the way polygons are formed requires additional regulations. Therefore, it is necessary to record the position of each vertex in space and the connection information of each polygon to fully express a mesh image. As shown in Figure 2, the same number of vertices and vertex positions will form completely different surfaces due to different connection methods.
[0084] In addition to the above information, since 3D mesh images are usually encoded using existing 2D image / video encoding methods, the 3D mesh needs to be converted from 3D space to 2D images. The UV coordinates define this conversion process.
[0085] Similar to 2D images, each position in the image may have corresponding attribute information, typically RGB color values, which reflect the object's color. For 3D meshes, in addition to color, each vertex often has reflectance values, which reflect the surface material. 3D mesh attribute information is stored in 2D images, and the mapping from 2D to 3D is defined by UV coordinates.
[0086] Therefore, 3D mesh data typically includes 3D geometric position information (x, y, z), geometric connectivity, UV coordinates, and an attribute map. Figure 3A shows a 3D mesh image, Figure 3B shows the mesh data storage format, which includes 3D geometric position information, UV coordinates, and connectivity information, and Figure 3C shows the corresponding attribute diagram.
[0087] Current 3D dynamic mesh compression methods include space-time prediction methods, which improve compression efficiency by eliminating spatial and temporal correlations; principal component analysis (PCA)-based technology, which projects in the eigenvector space to concentrate energy; and wavelet-based methods, which support spatial scalability and quality scalability.
[0088] It should be noted that in the dynamic mesh coding provided by the Moving Picture Experts Group (MPEG), Figure 4 shows a schematic diagram of the overall mesh coding framework, Figure 5A shows a schematic diagram of the preprocessing of a 2D curve, and Figure 5B shows a schematic diagram of the generation of shift coefficients. The preprocessing process for a 3D mesh is similar, and on the encoding side, it is mainly divided into two parts: preprocessing and encoder. Preprocessing first generates a base mesh and shift coefficients. The preprocessing process includes: first, downsampling the original mesh to generate a simplified mesh (decimated mesh) with a significantly reduced number of vertices, or the base mesh. The base mesh is then subdivided and algorithmically generated, with newly generated vertices inserted along the edges of the base mesh to form a subdivided mesh. Finally, for each vertex in the subdivided mesh, the nearest vertex in the original mesh is found. The vector between the vertex in the subdivided mesh and the nearest vertex in the original mesh is the shift coefficient. Since the subdivision grid can be automatically generated at the codec end as long as the subdivision algorithm and the number of subdivision iterations are determined, after preprocessing, the original grid only needs to be represented as a simple basic grid and a series of shift coefficients. This can greatly reduce the amount of data without affecting the reconstruction at the decoding end.
[0089] Video Dynamic Mesh Coding (V-DMC) based on video coding can be broadly categorized into two main categories: geometric position information encoding and attribute information encoding. As shown in Figure 5, each frame of the basketball_player sequence includes two files: basketball_player_fr0001_qp12_qt12.obj and basketball_player_fr0002.png. basketball_player_fr0001_qp12_qt12.obj contains four types of information: geometric position information (x, y, z), the connectivity of geometric position triangles, texture coordinates (u, v), and the connectivity of texture coordinates. basketball_player_fr0002.png represents the texture attribute information of the current image. In current V-DMC encoders, geometric position information is jointly encoded using the Dynamic Range Arithmetic Coding (DRACO) and the video codec, while texture information is encoded directly using the video codec. Among them, Video Codec can include H.264 / Advanced Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), H.266 / Versatile Video Coding (VVC / VV-enC), etc. Therefore, the following will introduce the mesh geometric information encoding in detail.
[0090] Geometric information can be divided into the encoding of position information (geometric position information and texture position information) and the encoding of connectivity relationships (geometric position information triangle patch connectivity, texture position information connectivity). Currently, V-DMC coding is mainly divided into two coding test conditions: intra-frame coding and inter-frame coding (low latency, currently no RA test environment).
[0091] (1) Intra-frame geometric information coding (Intra coding).
[0092] 1. Mesh preprocessing.
[0093] a) As shown in Figures 5A and 5B, using a two-dimensional connection relationship as an example, the original mesh's connection relationship contains a large number of points. Before encoding the mesh's geometric information, the mesh's geometric information is first quantized or simplified, ultimately resulting in a corresponding decimated mesh as the base mesh.
[0094] b) As shown in FIG6A and FIG6B , the quantization processing of the grid is performed based on the coordinates of the triangle patch. According to the connection relationship between the quantization points, the quantization processing can be divided into the following two cases:
[0095] When two vertices share a common edge, that is, two vertices that belong to the same edge before quantization, then after quantization, all the triangles connected by the two vertices need to be connected together, which involves the vanishing of the previous triangles as shown in Figure 6A;
[0096] Otherwise, if the two vertices do not share a common edge, that is, the two vertices do not belong to the same edge, then after quantization, it is only necessary to merge the boundaries of the two vertices, as shown in FIG6B , which does not affect the number of triangles.
[0097] c) In the whole process of mesh quantization based on triangle coordinates, the core problem is how to get the best vertex based on the previous vertex coordinates. The current V-DMC will get the best quantization point in the following four modes. Assuming that the vertex distribution before quantization is V1 and V2, and the vertex coordinate after quantization is V', there are the following: V1, V2, (V1+V2) / 2 and Q -1 (V1+V2), where Q is the quantization matrix corresponding to the vertex coordinates of V1 and V2. The distortion measure D before and after basic quantization is used to select the optimal quantization point.
[0098] 2.Base mesh encoding.
[0099] a) After obtaining the base mesh, the DRACO encoder is used to encode its geometric information. This geometric information primarily includes connectivity relationships and geometric position information. The DRACO encoding process is as follows: first, the connectivity relationships are encoded. Then, the geometric position information of the points is encoded based on the connectivity relationships. Finally, the texture position information is encoded based on the connectivity relationships and geometric position information.
[0100] b) Encoding of connection relationships. DRACO uses the "Edgebreaker Coding" scheme to encode the connection relationships of the mesh. See Figure 7A for details. In Figure 7A, v represents the current vertex. Before encoding the connection relationships of the mesh, the vertices of the mesh are divided into five types: C, L, R, S, and E. The physical meaning of each symbol is as follows:
[0101] C: None of the triangles connected to the current vertex have been encoded;
[0102] L: The triangle on the left connected to the current vertex completes the encoding;
[0103] R: The triangle on the right side connected to the current vertex completes the encoding;
[0104] S: The left and right triangles connected to the current vertex have not been encoded;
[0105] E: The left and right triangles connected to the current vertex have been encoded.
[0106] Finally, the type of each vertex and the processing order of the vertices are encoded in a certain order, and the decoding end restores the geometric connection relationship of the mesh according to the processing order and type of the vertices.
[0107] c) Coding of geometric position information. After completing the coding of the vertex connection relationship, the geometric position information of each vertex is predictively coded based on the vertex connection relationship. The idea adopted by predictive coding is the "Parallelograms algorithm", as shown in Figure 7B. A simple linear fit is performed using the three vertices adjacent to the current point to be coded: the left vertex, the right vertex, and the opposite vertex: pred pos =(left+right)-opposite (1)
[0108] d) After completing the point connection relationship and geometric position information, the texture coordinates are predictively encoded based on the decoding and reconstruction of these two, as shown in Figure 7C. Similarly, assuming the current vertex is C, the left and right vertices of the current point can be obtained based on the point connection relationship. Then, the texture coordinates of the left and right vertices are used to predict the texture coordinates of the current vertex C.
[0109] 3. Displacement coefficient encoding.
[0110] a) First, after the base mesh is encoded and reconstructed, a partitioning algorithm is used to partition the base mesh to obtain the initial reconstructed mesh. Specifically, the curve corresponding to the subdivided mesh in Figure 5A is used to obtain the subdivided mesh (also called the "initial mesh") through simple linear interpolation. The coordinates of the newly inserted points are obtained by linear interpolation based on the two vertices on the current boundary:
[0111] b) Secondly, calculate the error Delta between the points in the subdivided mesh and the original mesh after division. The error Delta can be a point error in the world coordinate system. Finally, the displacement (i.e., displacement coefficient) of each point is calculated using the error Delta between each point and the normal vector Norm of each point. See Figure 8 for details. In Figure 8, the bold solid line represents the error Delta, and N and T represent the normal vector Norm. Thus, the specific calculation method is as follows: Displacement = Delta × Norm (3)
[0112] c) After calculating the displacement of each point, the spatial domain residual coefficient can be transformed into the frequency domain using the lifting transform to obtain the corresponding frequency domain residual coefficient.
[0113] d) Finally, a coefficient packing algorithm is used to map the frequency domain residual coefficients of each point into a two-dimensional image in a certain order. The current V-DMC can be arranged according to the Morton Code Order, as shown in Figure 9.
[0114] e) Finally, a traditional Video Codec is used to encode the two-dimensional image.
[0115] 4. Recoloring.
[0116] Recoloring is an algorithm on the encoder side. After the reconstruction of the encoder's geometric information is completed, the original geometric information, the original texture attribute information, and the reconstructed mesh geometric information are used to recolor the texture attribute information of the reconstructed mesh.
[0117] (2) Inter-frame geometric information coding (Inter coding).
[0118] a) Similar to the encoding above, the geometric position information includes the geometric connection relationship and the geometric position information encoding. However, it should be noted that the inter-frame geometric position information encoding only needs to encode the geometric position information (x, y, z) of the current base mesh, and does not need to encode the connection relationship and texture position information (u, v). The specific reason is as follows: if the current image can be inter-coded, then the base mesh of the reference image of the current image will be used at the encoder to obtain the mesh information of the current image. Therefore, the current image and the reference image have the same connection relationship and UV texture coordinates, only the geometric position information is different.
[0119] b) Based on a), it can be known that the only difference between the current image and the reference image is the geometric position information. Therefore, the current V-DMC performs predictive coding on the geometric position information of the current image.
[0120] Specifically as shown in Figure 10, the black dot is the point to be coded. The corresponding prediction point is obtained by using the current point in the reference image (similar to the same-position block in video coding). Then, the motion vector (MV) of the current point is predicted and coded using the neighboring points of the current point (MV of the coded vertex). The specific details are as follows. Assume that the coordinates of the current point are pos and the coordinates of the corresponding same-position point are Pred pos , then the MV of the current point is calculated as: MV = Pos-Pred pos (4)
[0121] There are two predictive coding modes in the current V-DMC:
[0122] i. Directly encode the MV of the current point;
[0123] ii. Use the neighborhood to perform predictive coding on the MV of the current point.
[0124] At the encoding end, the rate-distortion optimization algorithm is used to obtain the optimal coding mode for each coding group (CG). The current V-DMC sets the number of points for each CG to be at most 16.
[0125] Coding of texture attribute information: Current V-DMC encodes texture attribute information directly using a video codec (Video-Codec), such as AVC, HEVC, VVC, or VV-enC.
[0126] Figure 11A is a schematic diagram of the framework of an intra-frame encoder. As shown in Figure 11A, in the intra-frame encoder, a common static mesh encoder (Static Mesh Encoder) can be used to encode the simplified mesh to generate the corresponding bitstream (Compressed base mesh bitstream). Next, the reconstructed simplified mesh is used to update the displacement coefficients (Update Displacements). The updated displacement coefficients are subjected to wavelet transform (Wavelet Transform) and quantization (Quantization) to obtain the displacement coefficients. After being packaged into images and videos (Image Packing, Video Packing), they are encoded using HEVC to generate a bitstream (Compressed displacements bitstream) of the displacement coefficients. For attribute map encoding, the feature map is first transformed (Texture Transfer) according to the difference between the reconstructed geometric information and the original geometric information, and then padded (Padding) and packaged (Video Packing) and encoded using a video encoder to form an attribute bitstream (Compressed attribute bitstream).
[0127] Figure 11B is a schematic diagram of an inter-frame encoder framework. As shown in Figure 11B , the inter-frame encoder and intra-frame encoder processes are roughly the same, but the inter-frame encoder does not directly encode the simplified grid. Instead, it encodes the motion vector MV between the simplified grid of the current image and the simplified grid of the reference image, and generates the corresponding motion vector bitstream (Compressed motion bitstream).
[0128] Correspondingly, during the decoding process, the decoder can also be divided into an intra-frame decoder and an inter-frame decoder according to the type of the frame it operates on, which are used to perform intra-frame decoding and inter-frame decoding respectively.
[0129] FIG12A is a schematic diagram of intra-frame decoding. As shown in FIG12A , in the intra-frame decoder, a static mesh decoder can be used to decode the simplified mesh. A video decoder is used to decode the shift coefficient video, and the shift coefficient is obtained through video unpacking and inverse wavelet transform. The decoded simplified mesh and shift coefficient are used to obtain the decoded mesh geometry information. The attribute map is decoded directly through the video decoder.
[0130] FIG12B is a schematic diagram of inter-frame decoding. As shown in FIG12B , for an inter-frame decoder, the process is basically the same as that for an intra-frame decoder, except that the simplified grid is not directly decoded, but the motion vector is decoded and the simplified grid of the current image is calculated using the simplified grid of the previous frame image (e.g., the reference image).
[0131] In summary, in the dynamic mesh coding (Dynamic Mesh Coding) currently provided by MPEG, the dynamic mesh coding process is divided into the following steps: at the encoding end, the basic mesh generated by preprocessing is quantized and then encoded using Google's open source DRACO encoder, and the shift coefficients are encoded using HEVC after wavelet transform, quantization, and two-dimensional mapping. The two-dimensional attribute map is also directly transmitted to the HEVC encoder for encoding; at the decoding end, the basic mesh code stream is decoded by DRACO to generate a decoded basic mesh, and the shift coefficients are decoded by HEVC decoding, inverse two-dimensional mapping, inverse quantization, and inverse transformation to generate decoded shift coefficients. Then, the decoded basic mesh and the decoded shift coefficients are used together to reconstruct the three-dimensional mesh geometry, and the attribute code stream is decoded by HEVC to generate a reconstructed attribute map.
[0132] (3) General test conditions for MPEG DMC.
[0133] a. There are 2 test conditions:
[0134] Condition 1: all intra geometry is lossy and attributes are lossy;
[0135] Condition 2: Random access is lossy in geometry and attributes;
[0136] b. The general test sequence may include five categories, namely Cat1-A, Cat1-B and Cat1-C, all of which contain geometric information and color attribute information.
[0137] The following is a detailed introduction to the Displacement coefficient encoding of V-DMC.
[0138] In one specific implementation, the encoder first iteratively partitions the base mesh using a specific algorithm to obtain the corresponding mesh position information. The specific partitioning algorithm is consistent with the previous description, using linear interpolation of vertices on each boundary to obtain the corresponding geometric position information. Assuming that the entire partitioning is iterated N times, the LOD partitioning is performed based on the displacement coefficients obtained from different iterative partitioning, as shown in Figure 13A.
[0139] As shown in Figure 13A, the base mesh is linearly interpolated to obtain the corresponding mesh geometric position information. The initial geometric position information is used to calculate the error with the original mesh to obtain the displacement coefficient of each point. The LOD division can be divided into four layers: level 0, level 1, level 2, and level 3. The specific LOD spatial structure is shown in Figure 13B.
[0140] Secondly, a lifting wavelet transform is performed based on the LOD spatial structure, which can include two steps: prediction and update. The prediction algorithm is as follows:
[0141] Here, v represents the vertex to be predicted, and v1 and v2 represent the vertices at both ends of the boundary where the vertex to be predicted is located.
[0142] Again, the update steps are as follows:
[0143] Finally, the transformed coefficients are quantized and reorganized. As shown in Figure 14, this reorganization is performed on a block-by-block basis, where they are grouped into Block 0, Block 1, Block 2, and Block 3. In other words, the current V-DMC performs coefficient reorganization on a block-by-block basis, with each block being 16×16 in size. The coefficients within each block are arranged according to the Morton code to produce the corresponding 2D image. After completing this series of operations, the 2D image can be encoded using the Video Codec.
[0144] In summary, the dynamic grid encoding process can be divided into the following steps:
[0145] 1. Preprocess the original mesh by reducing the number of vertices in the mesh and simplifying the connection relationship.
[0146] 2. Subdivide the simplified mesh in step 1. For any two connected vertices in step 1, add a new point at the midpoint of the connecting line, and repeat this process twice.
[0147] 3. For each vertex in step 2, find the point in the original mesh that is closest to it and calculate the displacement coefficient of these two points.
[0148] 4. Use an encoder such as Draco to quantize the simplified grid in step 1 and then encode it.
[0149] 5. Adjust the shift coefficients in step 3 based on the reconstructed simplified grid obtained in step 4.
[0150] 6. Perform wavelet transform on the shift coefficients in step 5, and quantize the shift coefficients after wavelet transform to obtain quantized transform coefficients.
[0151] 7. Map the quantized transform coefficients from three-dimensional space to a two-dimensional image (or "image packing") to generate a two-dimensional image of shifted coefficients.
[0152] 8. Use a standard video encoder such as H.265 to encode the shift coefficient two-dimensional image in step 6.
[0153] In another specific implementation, for the decoding end, first, the Video-Codec is used to decode and reconstruct the two-dimensional image to restore the corresponding two-dimensional image. Secondly, according to the coefficient reorganization method, the lifting transform coefficient corresponding to each point can be restored. Finally, the inverse transform of the lifting wavelet transform can be used to restore the displacement coefficient of each point. After obtaining the displacement coefficient of each point, the geometric position information of the base mesh and the displacement coefficient are used to reconstruct and restore the geometric position information corresponding to the current mesh. As shown in Figure 15, the geometric position information of the level 0 layer in the base mesh and the displacement coefficient of the level 0 layer can reconstruct and restore the geometric position information of the level 0 layer in the reconstructed mesh. The geometric position information of the level 1 layer in the base mesh and the displacement coefficient of the level 1 layer can reconstruct and restore the geometric position information of the level 1 layer in the reconstruct mesh. And so on, the geometric position information of the level 3 layer in the base mesh and the displacement coefficient of the level 3 layer can reconstruct and restore the geometric position information of the level 3 layer in the reconstruct mesh.
[0154] In summary, the dynamic grid decoding process can be divided into the following steps:
[0155] 1. The basic grid code stream is decoded by a decoder such as draco to generate a decoded basic grid.
[0156] 2. The shift coefficient bit stream is decoded using a standard video encoder such as H.265 to obtain a shift coefficient two-dimensional image.
[0157] 3. Map the shift coefficient two-dimensional image from the two-dimensional image to the three-dimensional space (or "image unpacking") to obtain quantized transform coefficients.
[0158] 4. Dequantize and inverse wavelet transform the quantized transform coefficients to obtain the decoded shift coefficients.
[0159] 5. The decoded base grid and the decoded shift coefficients are combined to generate the reconstructed 3D grid geometric information.
[0160] 6. After the attribute code stream is decoded by HEVC, a reconstructed attribute graph is generated.
[0161] In the related art, existing V-DMC coding always uses the initial mesh reconstruction geometry information and displacement coefficients obtained by dividing the base mesh to reconstruct and restore the current mesh geometry information. However, for inter-frame coding of meshes, if the current image can be inter-coded, that is, there is a reference image for the current image, then when encoding or reconstructing the mesh geometry information of the current image, the reconstructed displacement coefficients of the reference image can be obtained. Based on this, an inter-frame coding scheme can be introduced, that is, the initial mesh geometry information obtained by dividing the base mesh of the current image and the reconstructed displacement coefficients of the reference image are used to obtain the corresponding predicted mesh geometry information, and the re-obtained geometry information is used together with the original mesh geometry information for predictive coding or skip coding.
[0162] Based on this foundation, an embodiment of the present application provides a coding and decoding method. For inter-frame prediction coding of mesh, the geometric information of the initial mesh obtained by dividing the Base Mesh of the current image and the reconstructed Displacement coefficient of the reference image are used to obtain the corresponding predicted mesh geometric position information. Secondly, the new predicted geometric position information is used to perform parameter fitting with the original mesh position information at the encoding end, and finally the parameter relationship obtained by fitting is used to encode the geometric position information of the current mesh. For example, the Displacement coefficients of some LOD layers of the current image can be skipped or the correlation between the Displacement coefficients between different points can be used to reduce the encoding of the Displacement coefficients of some points. This can reduce the code stream size of the Displacement coefficient encoding while ensuring the quality of the reconstructed mesh, thereby further improving the geometric coding efficiency of the mesh. It can even skip encoding the Displacement coefficients in the current image, thereby saving bit rate and improving coding and decoding performance.
[0163] The embodiment of the present application also provides a grid architecture of a codec system including a decoding method and an encoding method. Figure 16 is a schematic diagram of a grid architecture of a codec provided by the embodiment of the present application. As shown in Figure 16, the grid architecture includes one or more electronic devices 13 to 1N and a communication grid 01, wherein the electronic devices 13 to 1N can perform video interaction through the communication grid 01. During the implementation process, the electronic device can be various types of devices with codec functions. For example, the electronic device can include a mobile phone, a tablet computer, a personal computer, a personal digital assistant, a navigator, a digital phone, a video phone, a television, a sensor device, a server, etc., which is not specifically limited in the embodiment of the present application. Here, the decoder or encoder described in the embodiment of the present application can be the above-mentioned electronic device.
[0164] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0165] In one embodiment of the present application, referring to FIG17 , a flowchart of a decoding method provided by an embodiment of the present application is shown. As shown in FIG17 , the method may include:
[0166] S1701: Determine the base grid of the current image.
[0167] It should be noted that the decoding method of the embodiment of the present application may refer to an inter-frame decoding method, more specifically, an inter-frame decoding method for displacement coefficients in a dynamic grid. The decoding method can be applied to a decoder in a V-DMC, but is not limited thereto.
[0168] It should also be noted that, in the embodiments of the present application, the basic grid may also be referred to as a “simplified grid.” In some embodiments, determining the basic grid of the current image may include: decoding a bitstream to determine the basic grid of the current image.
[0169] For example, the code stream here may refer to a basic grid code stream. Then, by decoding the basic grid code stream with a dynamic grid decoder (eg, DRACO), the basic grid of the current image may be obtained.
[0170] In some embodiments, determining the base grid of the current image may include: determining the base grid of the current image based on the base grid of the reference image.
[0171] It should also be noted that, in the embodiment of the present application, the reference image is a decoded image before the current image. For example, the reference image can be the previous frame of the current image, but this is not specifically limited.
[0172] In this way, if the current image uses inter-frame decoding and the base grid of the current image is not written into the code stream (ie, decoding of the base grid of the current image is skipped), the base grid of the reference image can also be used as the base grid of the current image.
[0173] S1702: Subdivide the basic grid to determine the geometric position information of the initial grid of the current layer in the current image.
[0174] It should be noted that, in the embodiment of the present application, by dividing the current image into levels of detail (LOD), at least one layer can be determined; wherein, the at least one layer can include the current layer.
[0175] For example, as shown in Figure 13A, the geometric position information of the initial mesh obtained by subdividing the base mesh three times is considered to be the 0th iteration corresponding to the 0th layer (level 0). The newly added vertices in the first iteration constitute the 1st layer (level 1), the newly added vertices in the second iteration constitute the 2nd layer (level 2), and the newly added vertices in the third iteration constitute the 3rd layer (level 3). The specific LOD division structure is shown in Figure 13B. The top layer is the base mesh. As the iteration proceeds, the number of newly added vertices in each iteration increases successively, forming a pyramid structure. In this way, by subdividing the base mesh, the geometric position information of the initial mesh of the current layer in the current image can be determined.
[0176] In some embodiments, subdividing the base grid to determine the geometric position information of the initial grid of the current layer in the current image may include: iteratively dividing the base grid according to a grid subdivision mode to determine the geometric position information of the initial grid of the current layer in the current image.
[0177] Specifically, in the embodiments of the present application, the mesh subdivision mode can be understood as upsampling the vertices on each boundary of the base mesh, or can also be understood as interpolating the vertices on each boundary of the base mesh. Exemplarily, the mesh subdivision mode includes a subdivision algorithm and a number of subdivision iterations. In some embodiments, the subdivision algorithm can be an interpolation algorithm, for example, a linear interpolation algorithm, or a nonlinear interpolation algorithm, which is not specifically limited herein.
[0178] Here, after the base mesh is decoded and reconstructed, a subdivision algorithm is used to subdivide the base mesh to obtain the initial mesh. For example, the base mesh can be used to obtain the initial mesh (also called "subdivided mesh") through a linear interpolation algorithm. In each iteration, the coordinates of the newly inserted points are obtained by linear interpolation based on the two vertices on the current boundary:
[0179] Among them, pos1 and pos2 are the geometric position coordinates of the two end vertices on the current boundary participating in this iteration, pos new The geometric position coordinates of the new vertices added for this iteration.
[0180] S1703: Decode the code stream and determine the first syntax identification information.
[0181] It should be noted that, in the embodiment of the present application, the first syntax identification information is used to indicate whether the shift coefficient of the current layer in the current image uses the first decoding mode. In some embodiments, the method may further include:
[0182] If the value of the first syntax identification information is the first value, determining that the first syntax identification information indicates that the shift coefficient of the current layer in the current image uses the first decoding mode;
[0183] If the value of the first syntax identification information is the second value, it is determined that the first syntax identification information indicates that the shift coefficient of the current layer in the current image does not use the first decoding mode.
[0184] In the embodiment of the present application, the first value and the second value are different, and the first value and the second value can be in parameter form or in digital form. Specifically, the first syntax identification information can be a parameter written in the profile or a flag value, which is not specifically limited here.
[0185] In an embodiment of the present application, the first syntax identification information is used as a syntax element at the LOD layer level to indicate whether the shift coefficient of the LOD layer of the current image uses the first decoding mode.
[0186] In some embodiments, the method may further include:
[0187] Decoding the code stream to determine second syntax identification information;
[0188] When the second syntax identification information indicates that the shift coefficient of the current sequence enables the first decoding mode, decoding the code stream to determine the third syntax identification information;
[0189] When the third syntax identification information indicates that the shift coefficient of the current image enables the first decoding mode, the code stream is decoded to determine the first syntax identification information.
[0190] In this embodiment of the present application, the second syntax identification information is a syntax element at the Sequence Parameter Set (SPS) level, and the third syntax identification information is a syntax element at the Frame Parameter Set (FPS) level. The second syntax identification information is used to indicate whether the shift coefficients of the current sequence enable the first decoding mode, and the third syntax identification information is used to indicate whether the shift coefficients of the current image enable the first decoding mode. Here, the current sequence includes at least the current image, and the LOD layers into which the current image is divided include at least the current layer.
[0191] In some embodiments, if the value of the second grammar identification information is a first value, it is determined that the second grammar identification information indicates that the shift coefficient of the current sequence enables the first decoding mode; if the value of the second grammar identification information is a second value, it is determined that the second grammar identification information indicates that the shift coefficient of the current sequence does not enable the first decoding mode.
[0192] In some embodiments, if the value of the third grammar identification information is the first value, it is determined that the third grammar identification information indicates that the shift coefficient of the current image enables the first decoding mode; if the value of the third grammar identification information is the second value, it is determined that the third grammar identification information indicates that the shift coefficient of the current image does not enable the first decoding mode.
[0193] In the embodiment of the present application, the first value and the second value are different, and the first value and the second value can be in parameter form or in numerical form. Specifically, the second syntax identification information and the third syntax identification information can be parameters written in the profile or the value of a flag, which is not specifically limited here.
[0194] For example, for the first value and the second value, the first value can be set to 1 and the second value can be set to 0; or the first value can be set to 0 and the second value can be set to 1; or the first value can be set to true and the second value can be set to false; or the first value can be set to false and the second value can be set to true. In the embodiment of the present application, the first value is set to 1 and the second value is set to 0, but this is not a specific limitation.
[0195] That is to say, for high-level syntax elements, first determine in the sequence parameter set whether to start the decoding method of the embodiment of the present application, secondly determine in the frame parameter set whether to start the decoding method of the embodiment of the present application, and finally determine at each LOD level the decoding method of the shift coefficient of the current layer, specifically whether to use the first decoding mode or the second decoding mode to decode the shift coefficient of the current layer.
[0196] S1704: When the first syntax identification information indicates that the shift coefficient of the current layer in the current image uses the first decoding mode, determine the shift coefficient of the current layer in the reference image.
[0197] S1705: Determine the geometric position information of the reconstructed grid of the current layer in the current image according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image.
[0198] It should be noted that in the embodiment of the present application, the first decoding mode skips decoding the shift coefficients of the current layer in the current image, that is, the shift coefficients of the current layer in the current image are not decoded. In this case, the shift coefficients of the current layer in the reference image can be used. In this way, since there is no need to decode the shift coefficients of the current layer in the current image, the encoding and decoding efficiency of the geometric position information can be improved while ensuring the reconstruction quality of the geometric position information of the grid.
[0199] In some embodiments, determining the geometric position information of the reconstructed grid of the current layer in the current image based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image may include: determining the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image; determining the geometric position information of the reconstructed grid of the current layer in the current image based on the mapping relationship and the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image.
[0200] It should also be noted that, in an embodiment of the present application, the mapping relationship can be a lookup table (LUT), which can record the correspondence between the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image and the geometric position information of the reconstructed grid of the current layer in the current image; or, the mapping relationship can also be a preset function, which can characterize the correspondence between the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image and the geometric position information of the reconstructed grid of the current layer in the current image.
[0201] It should also be noted that in the embodiments of the present application, the mapping relationship may include at least one of the following: a mapping relationship based on a linear function, a mapping relationship based on a nonlinear function, and a mapping relationship based on a neural grid. For example, the fitting of such a mapping relationship may include, but is not limited to, linear fitting, curve fitting, or convolution parameter fitting, and is not specifically limited here.
[0202] It should also be noted that, in the embodiments of the present application, the mapping relationship may be established by the decoding end according to relevant parameters, or may be determined by decoding the code stream. In some embodiments, the method may further include:
[0203] Decode the code stream and determine the mapping indication information of the current layer in the current image; based on the mapping indication information, determine the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image.
[0204] That is to say, in an embodiment of the present application, the encoding end can determine the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image and the geometric position information of the reconstructed grid of the current layer in the current image, and write the mapping relationship into the code stream; in this way, the decoding end can determine the mapping relationship by decoding the code stream, and then determine the geometric position information of the reconstructed grid of the current layer in the current image.
[0205] In a specific implementation, the mapping indication information includes first indication information; wherein the first indication information is used to indicate fitting parameters of the mapping relationship. In some embodiments, determining, based on the mapping indication information, a mapping relationship between geometric position information of an initial grid of a current layer in a current image, a shift coefficient of the current layer in a reference image, and geometric position information of a reconstructed grid of the current layer in the current image may include: determining, based on the first indication information, the fitting parameters of the mapping relationship; and determining, based on the fitting parameters, a mapping relationship between geometric position information of the initial grid of the current layer in the current image, a shift coefficient of the current layer in the reference image, and geometric position information of the reconstructed grid of the current layer in the current image.
[0206] It should be noted that in an embodiment of the present application, the code stream is decoded to determine the fitting parameters of the mapping relationship; then, based on the fitting parameters, the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image can be determined.
[0207] Exemplarily, when the mapping relationship is a mapping relationship based on a linear function, the fitting parameter may include the slope and / or intercept of the linear function.
[0208] Exemplarily, when the mapping relationship is based on a nonlinear function, the fitting parameter may include at least one constant in the nonlinear function. For example, if the nonlinear function is an exponential function, the fitting parameter may include the constant a of the nonlinear function; if the nonlinear function is a polynomial function, the fitting parameter may include the coefficients a0, a1, a2, ... of the polynomial function; and if the nonlinear function is a logarithmic function, the fitting parameter may include the base a of the nonlinear function.
[0209] In another specific implementation, the mapping indication information further includes second indication information; wherein the second indication information is used to indicate the type of the mapping relationship. In some embodiments, determining the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image based on the mapping indication information may include: determining the type of the mapping relationship based on the second indication information; and determining the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image based on the type of the mapping relationship and fitting parameters.
[0210] In the embodiment of the present application, the type of mapping relationship may include a linear function type, an exponential function type, a logarithmic function type, a polynomial function type, etc., which is not specifically limited here.
[0211] It should also be noted that, in an embodiment of the present application, the code stream is decoded to determine the type of mapping relationship and fitting parameters; then, based on the type of mapping relationship and fitting parameters, the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image can be determined.
[0212] For example, the specific formula of the mapping relationship is as follows: predict mesh =f(reconMesh+refDisp,lvl) (8)
[0213] Among them, lvl represents different LOD layers, reconMesh represents the geometric position information of the initial mesh, refDisp represents the reconstruction displacement coefficient of the reference image, and predict mesh It means that the geometric position information of the reconstructed grid is obtained by using a certain functional relationship.
[0214] In a specific implementation, a simple linear function relationship is used to obtain the corresponding relationship between the geometric position information of the reconstructed grid and the geometric position information of the initial grid, as shown below:mesh =k*(reconMesh+refDisp,lvl)+b (9)
[0215] Where * and + represent vector multiplication and addition; k and b represent fitting parameters.
[0216] In another specific implementation, the mapping indication information includes third indication information; wherein the third indication information is used to indicate an index number of the mapping relationship. In some embodiments, determining, based on the mapping indication information, a mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image may include: determining, based on the third indication information, the index number of the mapping relationship; and determining, based on the index number of the mapping relationship, the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image.
[0217] In the embodiment of the present application, the encoding end and the decoding end are both pre-set with several mapping relationships, and at this time the corresponding mapping relationship can be determined according to the index sequence number.
[0218] It should also be noted that, in an embodiment of the present application, the code stream is decoded to determine the index number of the mapping relationship; then, based on the index number of the mapping relationship, the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image can be determined.
[0219] Furthermore, in some embodiments, after step S1703, referring to FIG18 , the method may further include:
[0220] S1801: When the first syntax identification information indicates that the shift coefficient of the current layer in the current image uses the second decoding mode, decode the code stream to determine the shift coefficient of the current layer in the current image.
[0221] S1802: Determine geometric position information of a reconstructed grid of the current layer in the current image according to geometric position information of the initial grid of the current layer in the current image and a shift coefficient of the current layer in the current image.
[0222] In an embodiment of the present application, the first decoding mode is different from the second decoding mode. The first decoding mode may represent skipping the decoding of the shift coefficients of the current layer in the current image, i.e., it is not necessary to decode the shift coefficients of the current layer in the current image; the second decoding mode represents decoding the shift coefficients of the current layer in the current image, i.e., it is necessary to decode the shift coefficients of the current layer in the current image.
[0223] In an embodiment of the present application, if it is necessary to decode the shift coefficient of the current layer in the current image, then in some embodiments, decoding the code stream and determining the shift coefficient of the current layer in the current image may include: decoding the code stream to determine the two-dimensional image of the current layer in the current image; performing coefficient reorganization processing on the two-dimensional image to determine the lifting transformation coefficient of the current layer; and performing inverse transformation processing on the lifting transformation coefficient of the current layer to determine the shift coefficient of the current layer.
[0224] In a specific embodiment, performing coefficient reorganization processing on a two-dimensional image to determine the lifting transform coefficients of the current layer may include: performing coefficient reorganization processing on the two-dimensional image to determine the quantization coefficients of the current layer; and performing inverse quantization processing on the quantization coefficients of the current layer to determine the lifting transform coefficients of the current layer.
[0225] Specifically, the bitstream here can refer to the shift coefficient bitstream. By decoding the shift coefficient bitstream through a video decoder (e.g., Video-Codec), the corresponding two-dimensional image can be reconstructed. Then, by using coefficient reorganization, the quantization coefficients corresponding to each point in the current layer are recovered. Dequantization is then used to recover the lifting transform coefficients corresponding to each point in the current layer. Finally, the inverse of the lifting wavelet transform is used to recover the shift coefficients corresponding to each point in the current layer.
[0226] In this way, after obtaining the shift coefficient of each point, the geometric position information of the basic grid and the decoded shift coefficient can be used to reconstruct and restore the geometric position information of the reconstructed grid.
[0227] Furthermore, in embodiments of the present application, an identification information may be set to determine the decoding method for the base grid of the current image. Specifically, in some embodiments, the method may further include: decoding the bitstream to determine fourth syntax identification information; and when the fourth syntax identification information indicates that the base grid of the current image uses inter-frame processing, performing the steps of decoding the bitstream to determine the first syntax identification information.
[0228] In some embodiments, when the fourth syntax identification information indicates that the basic grid of the current image uses inter-frame processing, the method may further include: decoding the code stream to determine the reference image index of the current image; and determining the reference image based on the reference image index of the current image.
[0229] It should be noted that during the inter-frame prediction process, an identification information can first be set to determine the decoding method of the base grid of the current image. If the base grid of the current image can be inter-frame decoded, then the reference image index of the current image needs to be transmitted. Based on such an algorithm, the embodiment of the present application first determines whether the base grid of the current image can be inter-frame decoded. If inter-frame decoding can be used, it then determines whether the shift coefficient of the current layer in the current image uses skip decoding. Otherwise, the shift coefficient of the current layer in the current image is defaulted to use the second decoding mode.
[0230] In some embodiments, the method may further include: when the fourth syntax identification information indicates that the basic grid of the current image does not use inter-frame processing, decoding the code stream and determining the shift coefficient of the current layer in the current image; determining the geometric position information of the reconstructed grid of the current layer in the current image based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the current image.
[0231] That is to say, in an embodiment of the present application, if the basic grid of the current image does not adopt inter-frame decoding, the shift coefficient of the current layer in the current image can use the second decoding mode, that is, the shift coefficient of the current layer in the current image is determined by decoding the code stream, and then the geometric position information of the reconstructed grid of the current layer in the current image is determined based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the current image.
[0232] Furthermore, during the inter-frame prediction process, a flag may be set to determine the decoding method for the base grid of the current image. If the base grid of the current image can be inter-decoded, the reference image index of the current image needs to be transmitted. In an embodiment of the present application, regardless of whether the base grid of the current image can be inter-decoded, a flag needs to be transmitted to indicate whether the shift coefficients of the current layer in the current image use skip decoding.
[0233] Furthermore, in an embodiment of the present application, the geometric position information of the initial mesh obtained by basic mesh division of the current image and the shift coefficients of the reference image are used as a method for reconstructing the geometric position information of the mesh points. This solution reduces the bitrate of the shift coefficients while reducing the bitrate of the reconstructed mesh point position information to a certain extent, thereby improving the mesh coding efficiency. In an embodiment of the present application, parameter fitting can also be used to utilize the mapping relationship between the predicted mesh point geometric position information and the original mesh point geometric position information, thereby ensuring the quality of the reconstructed mesh point geometric position information, reducing the bitrate size of the shift coefficients, and further improving the mesh coding efficiency.
[0234] This embodiment provides a decoding method that determines a base grid for a current image; subdivides the base grid to determine geometric position information of an initial grid of a current layer in the current image; decodes a bitstream and determines first syntax identification information; when the first syntax identification information indicates that the shift coefficients of the current layer in the current image use a first decoding mode, determines the shift coefficients of the current layer in a reference image; and determines the geometric position information of a reconstructed grid of the current layer in the current image based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficients of the current layer in the reference image. In this way, the first syntax identification information can be used to indicate whether to skip decoding of the shift coefficients of the current layer in the current image. When the shift coefficients of the current layer in the current image are skipped, the geometric position information of the reconstructed grid is determined based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficients of the current layer in the reference image. This not only reduces the bitstream of the shift coefficients, but also ensures the reconstructed geometric quality of the grid midpoints, thereby further improving the quality of the geometric information of the grid midpoints and thereby improving encoding and decoding efficiency.
[0235] In another embodiment of the present application, see Figure 19, which shows a schematic flow chart of an encoding method provided by an embodiment of the present application. As shown in Figure 19, the method may include:
[0236] S1901: Determine a base grid of the current image according to the original grid of the current image.
[0237] It should be noted that the encoding method in the embodiment of the present application may refer to an inter-frame encoding method, more specifically, an inter-frame encoding method for displacement coefficients in a dynamic grid. The encoding method can be applied to an encoder in a V-DMC, but is not limited thereto.
[0238] It should also be noted that in the embodiments of the present application, the base grid may also be referred to as a "simplified grid." In some embodiments, determining the base grid of the current image based on the original grid of the current image may include downsampling the original grid of the current image to determine the base grid of the current image.
[0239] For example, the original mesh of the current image may be downsampled to generate a base mesh with a significantly reduced number of vertices. Furthermore, after obtaining the base mesh, the base mesh may be encoded, and the resulting encoded bits may be written into the bitstream.
[0240] For example, the code stream here can refer to the base mesh code stream. Specifically, a dynamic mesh encoder (e.g., DRACO) can be used to encode the geometric information of the base mesh, and the resulting coded bits can be written into the base mesh code stream. The geometric information mainly includes the connection relationship and the connection relationship of the geometric position information.
[0241] Exemplarily, for the encoding of the geometric information of the basic grid, the encoding process can be: first complete the encoding of the connection relationship, then encode the geometric position information of the point based on the connection relationship of the geometric position, and finally encode the texture position information based on the connection relationship and the geometric position information.
[0242] S1902: Subdivide the basic grid and determine the geometric position information of the initial grid of the current layer in the current image.
[0243] It should be noted that, in the embodiment of the present application, at least one layer can be determined by performing LOD division on the current image; wherein, the at least one layer may include the current layer.
[0244] It should also be noted that, in an embodiment of the present application, after obtaining the basic mesh, the basic mesh can also be subdivided, and new vertices can be inserted on the boundary of the basic mesh to generate the geometric position information of the initial mesh. For example, as shown in FIG13A , the geometric position information of the initial mesh obtained by subdividing the basic mesh three times is that the basic mesh is regarded as the 0th iteration corresponding to the 0th layer (level0), the newly added vertices in the first iteration constitute the 1st layer (level1), the newly added vertices in the second iteration constitute the 2nd layer (level2), and the newly added vertices in the third iteration constitute the 3rd layer (level3). The specific LOD division structure is shown in FIG13B , where the top layer is the basic mesh. As the iteration proceeds, the number of newly added vertices in each iteration increases successively to form a pyramid structure.
[0245] In some embodiments, subdividing the base grid to determine the geometric position information of the initial grid of the current layer in the current image may include: iteratively dividing the base grid according to a grid subdivision mode to determine the geometric position information of the initial grid of the current layer in the current image.
[0246] Specifically, in the embodiments of the present application, the mesh subdivision mode can be understood as upsampling the vertices on each boundary of the base mesh, or can also be understood as interpolating the vertices on each boundary of the base mesh. Exemplarily, the mesh subdivision mode includes a subdivision algorithm and a number of subdivision iterations. In some embodiments, the subdivision algorithm can be an interpolation algorithm, for example, a linear interpolation algorithm, or a nonlinear interpolation algorithm, which is not specifically limited herein.
[0247] Here, after obtaining the base mesh of the current image, a certain subdivision algorithm will be used to subdivide the base mesh to obtain the initial mesh. For example, the base mesh can be used to obtain the initial mesh (also called "subdivided mesh") through the linear interpolation algorithm. In each iteration, the coordinates of the newly inserted points are obtained by linear interpolation based on the two vertices on the current boundary:
[0248] Among them, pos1 and pos2 are the geometric position coordinates of the two end vertices on the current boundary participating in this iteration, pos new The geometric position coordinates of the new vertices added for this iteration.
[0249] S1903: Determine the shift coefficient of the current layer in the reference image, and determine the geometric position information of the first reconstructed grid of the current layer in the current image based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image.
[0250] It should be noted that, in the embodiment of the present application, the reference image is an encoded image before the current image. For example, the reference image can be the previous frame image of the current image, but this is not specifically limited.
[0251] In some embodiments, determining the geometric position information of the first reconstructed grid of the current layer in the current image based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image may include: determining the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the first reconstructed grid of the current layer in the current image; determining the geometric position information of the first reconstructed grid of the current layer in the current image based on the mapping relationship and the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image.
[0252] It should also be noted that, in an embodiment of the present application, the mapping relationship can be a lookup table (LUT), which can record the correspondence between the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image and the geometric position information of the first reconstructed grid; or, the mapping relationship can also be a preset function, which can characterize the correspondence between the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image and the geometric position information of the first reconstructed grid.
[0253] It should also be noted that in the embodiments of the present application, the mapping relationship may include at least one of the following: a mapping relationship based on a linear function, a mapping relationship based on a nonlinear function, and a mapping relationship based on a neural grid. For example, the fitting of such a mapping relationship may include, but is not limited to, linear fitting, curve fitting, or convolution parameter fitting, and is not specifically limited here.
[0254] It should also be noted that in the embodiments of the present application, the encoding end may write the bitstream after determining the mapping relationship. In some embodiments, the method may further include: determining mapping indication information for the current layer in the current image based on the mapping relationship; encoding the mapping indication information, and writing the resulting coded bits into the bitstream.
[0255] That is to say, in an embodiment of the present application, the encoding end can determine the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image and the geometric position information of the first reconstructed grid and write the mapping relationship into the code stream; in this way, the decoding end can determine the mapping relationship by decoding the code stream, and then restore the geometric position information of the first reconstructed grid.
[0256] In a specific implementation, the mapping indication information includes first indication information; wherein the first indication information is used to indicate fitting parameters of the mapping relationship. In some embodiments, the method may further include: determining the fitting parameters of the mapping relationship; encoding the fitting parameters of the mapping relationship, and writing the obtained encoded bits into the bitstream.
[0257] In a specific implementation, the mapping indication information further includes second indication information; wherein the second indication information is used to indicate the type of the mapping relationship. In some embodiments, the method may further include: determining the type of the mapping relationship and fitting parameters; encoding the type of the mapping relationship and the fitting parameters, and writing the resulting encoded bits into the bitstream.
[0258] Exemplarily, when the mapping relationship is based on a linear function, the fitting parameter may include the slope and / or intercept of the linear function. When the mapping relationship is based on a nonlinear function, the fitting parameter may include at least one constant in the nonlinear function. For example, if the nonlinear function is an exponential function, the fitting parameter may include the constant a of the nonlinear function; if the nonlinear function is a polynomial function, the fitting parameter may include the coefficients a0, a1, a2, ... of the polynomial function; if the nonlinear function is a logarithmic function, the fitting parameter may include the base a of the nonlinear function.
[0259] Exemplarily, the types of mapping relationships may include linear function types, exponential function types, logarithmic function types, polynomial function types, etc., which are not specifically limited here.
[0260] In the embodiment of the present application, the specific formula of the mapping relationship is as follows: mesh =f(reconMesh+refDisp,lvl) (11)
[0261] Among them, lvl represents different LOD layers, reconMesh represents the geometric position information of the initial mesh, refDisp represents the reconstruction displacement coefficient of the reference image, and predict mesh It indicates that the geometric position information of the first reconstructed grid is obtained by using a certain functional relationship.
[0262] In a specific implementation, a simple linear function relationship is used to obtain the corresponding relationship between the geometric position information of the first reconstructed grid and the geometric position information of the initial grid, as shown below: mesh =k*(reconMesh+refDisp,lvl)+b (12)
[0263] Where * and + represent vector multiplication and addition; k and b represent fitting parameters.
[0264] In another specific implementation, the mapping indication information includes third indication information; wherein the third indication information is used to indicate the index number of the mapping relationship. In some embodiments, the method may further include: determining the index number of the mapping relationship; encoding the index number of the mapping relationship, and writing the obtained encoded bits into the bitstream.
[0265] In the embodiment of the present application, the encoding and decoding ends are both pre-configured with several mapping relationships. The corresponding mapping relationship can then be determined based on the index number. Thus, after determining the index number of the mapping relationship, the encoding end can write it into the bitstream. Subsequently, the decoding end can obtain the mapping relationship index number by decoding the bitstream, and then determine the corresponding mapping relationship based on the mapping relationship index number.
[0266] S1904: Determine whether to perform encoding processing on the shift coefficients of the current layer in the current image according to the geometric position information of the first reconstructed grid and the geometric position information of the original grid.
[0267] It should be noted that, in some embodiments, determining whether to encode the shift coefficient of the current layer in the current image based on the geometric position information of the first reconstructed grid and the geometric position information of the original grid may include: performing error calculation on the geometric position information of the first reconstructed grid and the geometric position information of the original grid to determine a first error result; based on the first error result, determining the encoding mode of the shift coefficient of the current layer in the current image, wherein the encoding mode is used to characterize whether to encode the shift coefficient of the current layer in the current image.
[0268] It should also be noted that, in some embodiments, the method may further include: determining the geometric position information of the second reconstructed grid of the current layer in the current image based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the current image; performing error calculation on the geometric position information of the second reconstructed grid and the geometric position information of the original grid to determine a second error result.
[0269] In an embodiment of the present application, for the shift coefficient of the current layer in the current image, the method may further include: determining the shift coefficient of the current layer in the current image based on the geometric position information of the initial grid of the current layer in the current image and the geometric position information of the original grid of the current layer in the current image.
[0270] In a specific embodiment, determining the shift coefficient of the current layer in the current image based on the geometric position information of the initial mesh of the current layer in the current image and the geometric position information of the original mesh of the current layer in the current image can include: determining the error value of the first vertex between the initial mesh and the original mesh based on the first vertex in the initial mesh, and determining the normal vector of the first vertex; calculating the shift coefficient of the first vertex based on the error value of the first vertex and the normal vector of the first vertex; wherein the first vertex is any vertex in the initial mesh.
[0271] It should be noted that in the embodiment of the present application, for the first vertex, the error value Delta between the initial mesh and the original mesh can be calculated first. The error Delta can be an error of a point in the world coordinate system. Then, the error Delta between the first vertices and the normal vector Norm of the first vertex are used to calculate the displacement coefficient of the first vertex, as shown in Figure 8. In an exemplary embodiment, a specific calculation method is as follows: Displacement = Delta × Norm (13)
[0272] Where Displacement is the displacement coefficient.
[0273] Thus, after obtaining the second error result, the coding mode of the shift coefficient of the current layer in the current image is determined according to the first error result and the second error result. Specifically, in some embodiments, the method may include:
[0274] determining an error ratio between the first error result and the second error result;
[0275] If the error ratio is less than a preset threshold, determining that the shift coefficient of the current layer in the current image uses the first coding mode;
[0276] If the error ratio is greater than a preset threshold, it is determined that the shift coefficient of the current layer in the current image uses the second encoding mode.
[0277] It should be noted that in the embodiment of the present application, the first coding mode is different from the second coding mode. The first coding mode represents skipping the encoding of the shift coefficients of the current layer in the current image; the second coding mode represents encoding the shift coefficients of the current layer in the current image.
[0278] It should also be noted that in an embodiment of the present application, when the error ratio is equal to a preset threshold, it can be determined that the shift coefficient of the current layer in the current image uses the first coding mode, or it can be determined that the shift coefficient of the current layer in the current image uses the second coding mode.
[0279] In a specific embodiment, referring to FIG20 , the method may include:
[0280] S2001: Determine geometric position information of a first reconstructed grid of a current layer in a current image according to geometric position information of an initial grid of a current layer in a current image and a shift coefficient of the current layer in a reference image.
[0281] S2002: performing error calculation on the geometric position information of the first reconstructed grid and the geometric position information of the original grid to determine a first error result.
[0282] S2003: Determine geometric position information of a second reconstructed grid of the current layer in the current image according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the current image.
[0283] S2004: performing error calculation on the geometric position information of the second reconstructed grid and the geometric position information of the original grid to determine a second error result.
[0284] S2005: Determine an error ratio between the first error result and the second error result.
[0285] S2006: If the error ratio is less than a preset threshold, skip encoding the shift coefficients of the current layer in the current image.
[0286] S2007: If the error ratio is greater than a preset threshold, encode the shift coefficient of the current layer in the current image.
[0287] In the embodiment of the present application, the first error result and the second error result can be calculated by a certain distortion. Here, the mean square error (MSE) algorithm can be used to obtain the error of the geometric position information of the mesh at different LOD layers, as shown below:
[0288] That is to say, before encoding the Displacement coefficient of the lvl layer in the current image, the error between the reconstructed mesh obtained by the original encoding method and the original mesh is first calculated, that is, the second error result, represented by Dist_org; secondly, the method of the embodiment of the present application is used: the error between the geometric position information of the initial mesh obtained by dividing the Base mesh and the Displacement coefficient in the reference image is reconstructed, that is, the first error result, represented by Dist_skip.
[0289] In this way, the error ratio can be expressed as Dist_skip / Dist_org. If Dist_skip / Dist_org is less than a preset threshold, then it can be determined that the Displacement coefficients of the lvl layer in the current image are encoded using the first encoding mode, that is, encoding is skipped, that is, the Displacement coefficients of the lvl layer in the current image are not encoded; if Dist_skip / Dist_org is greater than the preset threshold, then it can be determined that the Displacement coefficients of the lvl layer in the current image are encoded using the second encoding mode, that is, the Displacement coefficients of the lvl layer in the current image need to be encoded.
[0290] In some embodiments, the method may further include: performing rate-distortion cost calculation on the shift coefficients of the current layer in the current image according to the first coding mode to determine a first rate-distortion result; and performing rate-distortion cost calculation on the shift coefficients of the current layer in the current image according to the second coding mode to determine a second rate-distortion result;
[0291] A coding mode for a shift coefficient of a current layer in a current image is determined according to the first rate-distortion result and the second rate-distortion result.
[0292] In a specific embodiment, determining the coding mode of the shift coefficient of the current layer in the current image according to the first rate-distortion result and the second rate-distortion result may include:
[0293] If the first rate-distortion result is less than the second rate-distortion result, determining that the shift coefficient of the current layer in the current image uses the first coding mode;
[0294] If the second rate-distortion result is less than the first rate-distortion result, it is determined that the shift coefficient of the current layer in the current image uses the second coding mode.
[0295] It should be noted that in an embodiment of the present application, when the first rate-distortion result is equal to the second rate-distortion result, it can be determined that the shift coefficient of the current layer in the current image uses the first coding mode, or it can be determined that the shift coefficient of the current layer in the current image uses the second coding mode.
[0296] It should also be noted that in the embodiment of the present application, in order to more accurately measure whether the first coding mode and the second coding mode improve the coding performance, the embodiment of the present application simultaneously performs a rate-distortion trade-off on the two coding modes. Here, the rate-distortion cost method can be used to calculate the rate-distortion result after the comprehensive quality improvement and the increase in bit rate. Among them, the first rate-distortion result and the second rate-distortion result can respectively represent the rate-distortion cost of the reconstructed grid obtained by the first coding mode relative to the original grid or the reconstructed grid obtained by the second coding mode relative to the original grid, which is used to indicate whether the compression efficiency of the coding shift coefficient is skipped. For the calculation of the first rate-distortion result and the second rate-distortion result, the specific calculation formula is as follows: J = D + λ × R (15)
[0297] Here, J is the rate-distortion result, D is the distortion between the original grid and the reconstructed grid obtained by the first coding mode or the second coding mode, such as the sum of squares of corresponding point errors; λ is a quantity related to the quantization parameter QP, and R is the total geometric bitstream size divided by the number of frames.
[0298] Exemplarily, if the first rate-distortion result is smaller than the second rate-distortion result, then the first coding mode can be determined to be the optimal coding mode. At this time, the displacement coefficient of the lvl layer in the current image uses the first coding mode, that is, skip coding, that is, the Displacement coefficient of the lvl layer in the current image is not encoded; if the second rate-distortion result is smaller than the first rate-distortion result, then the second coding mode can be determined to be the optimal coding mode. At this time, the displacement coefficient of the lvl layer in the current image uses the second coding mode, that is, the Displacement coefficient of the lvl layer in the current image needs to be encoded.
[0299] Furthermore, in some embodiments, the method may further include: when the shift coefficient of the current layer in the current image uses the second coding mode, encoding the shift coefficient of the current layer in the current image, and writing the obtained coding bits into the bitstream.
[0300] In an embodiment of the present application, encoding is performed on the shift coefficients of the current layer in the current image, and the resulting encoded bits are written into the bitstream. Specifically, the encoding process may include: performing a lifting transform on the shift coefficients of the current layer in the current image to determine the lifting transform coefficients; performing quantization on the lifting transform coefficients to determine the quantization coefficients; performing coefficient reorganization on the quantization coefficients to determine a two-dimensional image; and encoding the two-dimensional image and writing the resulting encoded bits into the bitstream.
[0301] Exemplarily, the implementation steps of the shift coefficient encoding method provided in the embodiment of the present application may include:
[0302] First, the base mesh is iteratively partitioned using a specific algorithm to obtain the corresponding mesh position information. The specific partitioning algorithm is consistent with the previous content, using linear interpolation of vertices on each boundary to obtain the corresponding geometric position information. Assuming that the entire partitioning is iterated N times, the displacements obtained from different iterative partitioning are then divided into Level of Dimensions (LODs) to obtain the corresponding mesh geometric position information. The error between the initial geometric position information and the original mesh is calculated to obtain the displacement of each point.
[0303] Secondly, based on the LOD spatial structure, the shift coefficients are subjected to a lifting wavelet transform, which includes two steps: prediction and update. The prediction step is as follows:
[0304] The update steps are as follows:
[0305] Finally, the transformed coefficients are quantized and reorganized. Current V-DMC uses block-based coefficient reorganization, with each block being 16×16. The coefficients within each block are arranged using Morton codes to produce the corresponding 2D image. After completing these operations, the 2D image is encoded using a Video Codec.
[0306] Exemplarily, the shift coefficients may be directly encoded using an entropy encoder to obtain a shift coefficient code stream.
[0307] Furthermore, in some embodiments, the method may also include: determining a value of second grammar identification information, wherein the second grammar identification information is used to indicate whether the shift coefficient of the current sequence enables the first coding mode; encoding the value of the second grammar identification information, and writing the obtained coded bits into the bitstream.
[0308] Furthermore, in some embodiments, the method may also include: when the second syntax identification information indicates that the shift coefficient of the current sequence enables the first coding mode, determining the value of the third syntax identification information, wherein the third syntax identification information is used to indicate whether the shift coefficient of the current image enables the first coding mode; encoding the value of the third syntax identification information, and writing the obtained coding bits into the bitstream.
[0309] Furthermore, in some embodiments, the method may also include: when the third syntax identification information indicates that the shift coefficient of the current image enables the first coding mode, determining the value of the first syntax identification information, wherein the first syntax identification information is used to indicate whether the shift coefficient of the current layer in the current image uses the first coding mode; encoding the value of the first syntax identification information and writing the obtained coding bits into the bitstream.
[0310] In an embodiment of the present application, the second syntax identification information is a syntax element at the Sequence Parameter Set (SPS) level, the third syntax identification information is a syntax element at the frame parameter set (FPS) level, and the first syntax identification information is a syntax element at the LOD layer level. The second syntax identification information is used to indicate whether the shift coefficients of the current sequence enable the first coding mode, the third syntax identification information is used to indicate whether the shift coefficients of the current image enable the first coding mode, and the first syntax identification information is used to indicate whether the shift coefficients of the current layer use the first coding mode. Here, the current sequence includes at least the current image, and the LOD layers into which the current image is divided include at least the current layer.
[0311] In an embodiment of the present application, for the value of the first syntax identification information, if the shift coefficient of the current layer in the current image uses the first coding mode, or the encoding of the shift coefficient of the current layer in the current image is skipped, the value of the first syntax identification information is determined to be the first value; if the shift coefficient of the current layer in the current image does not use the first coding mode, or the shift coefficient of the current layer in the current image needs to be encoded, the value of the first syntax identification information is determined to be the second value.
[0312] In an embodiment of the present application, for the value of the second grammar identification information, if the shift coefficient of the current sequence enables the first coding mode, the value of the second grammar identification information is determined to be the first value; if the shift coefficient of the current sequence does not enable the first coding mode, the value of the second grammar identification information is determined to be the second value.
[0313] In an embodiment of the present application, for the value of the third grammar identification information, if the shift coefficient of the current image enables the first coding mode, the value of the third grammar identification information is determined to be the first value; if the shift coefficient of the current image does not enable the first coding mode, the value of the third grammar identification information is determined to be the second value.
[0314] It should be noted that the first value and the second value are different, and the first value and the second value can be in parameter form or in numeric form. Specifically, the first syntax identification information, the second syntax identification information, and the third syntax identification information can be parameters written into the profile or the value of a flag, which is not specifically limited here.
[0315] For example, for the first value and the second value, the first value can be set to 1 and the second value can be set to 0; or the first value can be set to 0 and the second value can be set to 1; or the first value can be set to true and the second value can be set to false; or the first value can be set to false and the second value can be set to true. In the embodiment of the present application, the first value is set to 1 and the second value is set to 0, but this is not a specific limitation.
[0316] That is to say, in an embodiment of the present application, for high-level syntax elements, it is first determined in the sequence parameter set whether to start the encoding method of the embodiment of the present application, and secondly, it is determined in the frame parameter set whether to start the encoding method of the embodiment of the present application, and finally, the encoding method of the shift coefficient of the current layer is determined at each LOD level, specifically whether to use the first encoding mode or the second encoding mode to encode the shift coefficient of the current layer.
[0317] Furthermore, in some embodiments, the method may also include: determining a value of fourth syntax identification information, wherein the fourth syntax identification information is used to indicate whether the basic grid of the current image uses an inter-frame processing method; encoding the value of the fourth syntax identification information, and writing the obtained encoding bits into the bitstream.
[0318] It should be noted that, in an embodiment of the present application, for the value of the fourth grammar identification information, if the basic grid of the current image uses the inter-frame processing method, the value of the fourth grammar identification information is determined to be the first value; if the basic grid of the current image does not use the inter-frame processing method, the value of the fourth grammar identification information is determined to be the second value.
[0319] It should also be noted that in the embodiment of the present application, the first value is different from the second value, and the first value and the second value can be in parameter form or in numerical form. Specifically, the fourth syntax identification information can be a parameter written in the profile or a flag value, which is not specifically limited here. Exemplarily, the first value is set to 1 and the second value is set to 0, but this is not specifically limited here.
[0320] In some embodiments, the method may further include: when the base grid of the current image uses an inter-frame processing method, executing the step of determining whether to perform encoding processing on the shift coefficients of the current layer in the current image.
[0321] In an embodiment of the present application, when the basic grid of the current image uses an inter-frame processing method, the method also includes: determining a reference image index of the current image based on the reference image; encoding the reference image index of the current image, and writing the obtained encoding bits into the code stream.
[0322] It should be noted that during the inter-frame prediction process, an identification information can first be set to determine the encoding method of the base grid of the current image. If the base grid of the current image can be inter-coded, then the reference image index of the current image needs to be transmitted. Based on such an algorithm, the embodiment of the present application first determines whether the base grid of the current image can be inter-coded. If inter-coding is possible, it then determines whether the shift coefficient of the current layer in the current image uses skip coding. Otherwise, the shift coefficient of the current layer in the current image is defaulted to use the second coding mode.
[0323] In some embodiments, the method may further include: determining the shift coefficient of the current layer in the current image when the fourth syntax identification information indicates that the basic grid of the current image does not use inter-frame processing; determining the geometric position information of the reconstructed grid of the current layer in the current image based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the current image.
[0324] That is to say, in an embodiment of the present application, if the basic grid of the current image does not adopt inter-frame coding, the shift coefficient of the current layer in the current image can use the second coding mode by default, that is, after determining the shift coefficient of the current layer in the current image, the geometric position information of the reconstructed grid of the current layer in the current image is determined according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the current image.
[0325] Furthermore, during the inter-frame prediction process, a flag may be first set to determine the encoding method for the base grid of the current image. If the base grid of the current image can be inter-coded, then the reference image index of the current image needs to be transmitted. In an embodiment of the present application, regardless of whether the base grid of the current image can be inter-coded, a flag needs to be transmitted to indicate whether the shift coefficients of the current layer in the current image use skip coding.
[0326] Furthermore, in an embodiment of the present application, the geometric position information of the initial mesh obtained by basic mesh division of the current image and the shift coefficients of the reference image are used as a method for reconstructing the geometric position information of the mesh points. This solution reduces the bitrate of the shift coefficients while reducing the bitrate of the reconstructed mesh point position information to a certain extent, thereby improving the mesh coding efficiency. In an embodiment of the present application, parameter fitting can also be used to utilize the mapping relationship between the predicted mesh point geometric position information and the original mesh point geometric position information, thereby ensuring the quality of the reconstructed mesh point geometric position information, reducing the bitrate size of the shift coefficients, and further improving the mesh coding efficiency.
[0327] Furthermore, an embodiment of the present application further provides a code stream, which is generated by bit encoding based on the information to be encoded; wherein the information to be encoded may include at least one of the following:
[0328] The value of the first syntax identification information, the value of the second syntax identification information, the value of the third syntax identification information, the value of the fourth syntax identification information, the reference image index of the current image, the shift coefficient of the current layer in the current image, and the mapping indication information of the current layer in the current image;
[0329] Among them, the second syntax identification information is used to indicate whether the shift coefficient of the current sequence enables the first coding mode, the third syntax identification information is used to indicate whether the shift coefficient of the current image enables the first coding mode, the first syntax identification information is used to indicate whether the shift coefficient of the current layer in the current image uses the first coding mode, and the fourth syntax identification information is used to indicate whether the basic grid of the current image uses inter-frame processing; and the current sequence includes the current image, and the LOD layer divided by the current image includes the current layer.
[0330] This embodiment provides an encoding method that determines a base grid for the current image based on an original grid of the current image; subdivides the base grid to determine geometric position information of an initial grid of a current layer in the current image; determines a shift coefficient for the current layer in a reference image, and determines geometric position information of a first reconstructed grid for the current layer in the current image based on the geometric position information of the initial grid for the current layer in the current image and the shift coefficient of the current layer in the reference image; and determines whether to encode the shift coefficient for the current layer in the current image based on the geometric position information of the first reconstructed grid and the geometric position information of the original grid. In this way, based on whether to encode the shift coefficient for the current layer in the current image, if encoding of the shift coefficient for the current layer in the current image is skipped, the geometric position information of the first reconstructed grid can be determined based on the geometric position information of the initial grid for the current layer in the current image and the shift coefficient of the current layer in the reference image. This not only reduces the bitrate of the shift coefficient, but also ensures the reconstructed geometric quality of grid midpoints, thereby further improving the quality of the geometric information of the grid midpoints and thereby improving encoding efficiency.
[0331] In another embodiment of the present application, based on the existing V-DMC coding foundation, the geometric position information of the corresponding reconstructed mesh (also referred to as "the geometric position information of the predicted mesh") is obtained by using the geometric position information after the Base Mesh division of the current image and the shift coefficients reconstructed from the reference image. By analyzing the correlation between the geometric position information of the midpoints of the predicted mesh and the geometric position information of the midpoints of the original mesh, the shift coefficients can be adaptively encoded. This can reduce the bit rate of the shift coefficient encoding while ensuring the quality of the mesh geometric position information reconstruction, thereby improving the coding efficiency of the mesh geometric position information.
[0332] For example, Figure 21 is a schematic diagram illustrating the basic principle of determining the geometric position information of a prediction grid according to an embodiment of the present application. As shown in Figure 21, the geometric position information of the prediction grid can be determined based on the shift coefficients of the base grid and the reference image. Then, based on the parameter correlation between the geometric position information of the midpoints of the prediction grid and the geometric position information of the midpoints of the original grid, it can be determined whether to encode the shift coefficients of the current image.
[0333] In a specific implementation, at the encoder end, taking the current LOD layer as an example, a certain algorithm is used to fit the relationship between the geometric position information of the predicted grid and the geometric position information of the initial grid. Assume that the relationship between the two is as follows: predict mesh =f(reconMesh+refDisp,lvl) (18)
[0334] Among them, lvl represents different LOD layers, reconMesh represents the geometric position information of the initial mesh, refDisp represents the displacement coefficient of the reference image, and predict mesh It means that the geometric position information of the corresponding prediction grid is obtained by using a certain functional relationship.
[0335] In the embodiment of the present application, considering the balance between coding complexity and coding efficiency, the relationship between the geometric position information of the predicted grid and the geometric position information of the initial grid can be obtained by using simple linear fitting, as shown below: mesh =k*(reconMesh+refDisp,lvl)+b (19)
[0336] Among them, * and + are vector multiplication and addition.
[0337] In this way, through a certain distortion, for example, the MSE algorithm can be used to obtain the error of the mesh geometric position information corresponding to different LOD layers, as shown below:
[0338] At the encoding end, before encoding the shift coefficients of the current layer in the current image, first calculate the error between the point geometric position information of the reconstructed grid obtained by the original encoding scheme and the point geometric position information of the original grid, which is Dist_org; secondly, using the method in the embodiment of the present application: the error between the point geometric position information of the predicted grid obtained by reconstructing the point geometric position information of the initial grid obtained by dividing the Base Mesh and the shift coefficients in the reference image is Dist_skip. Again, when Dist_skip / Dist_org is less than a certain threshold (Th), it means that the shift coefficients of the current layer in the current image are skipped for encoding, specifically: the shift coefficients of the current layer in the current image are not encoded, and the geometric position information of the current reconstructed grid is directly reconstructed using the geometric position information after the Base Mesh division and the shift coefficients in the reference image.
[0339] In another specific implementation, at the decoding end, first, the basic grid of the current image is obtained by analytical reconstruction, and the basic grid of the current image is divided by a partitioning algorithm to obtain the geometric position information of the initial grid. Secondly, the decoding mode of the shift coefficient of the current image is determined according to the encoding mode of the shift coefficient of the current image.
[0340] If the displacement coefficient in the current image needs to be decoded, the initial geometric position information obtained by dividing the base mesh of the current image and the displacement coefficient are used to reconstruct and restore the geometric position information of the reconstructed mesh;
[0341] If the Displacement coefficient in the current image is skipped for decoding, the initial geometric position information obtained by dividing the Base mesh of the current image and the Displacement coefficient in the reference image are directly used to reconstruct and restore the geometric position information of the reconstructed mesh.
[0342] Furthermore, in an embodiment of the present application, the displacement coefficient coding mode of each frame is modified. Specifically, based on the existing V-DMC coding scheme, in inter-frame coding, there is first an identification information (identifier) that determines the coding method of the Base Mesh of the current image. If the Base Mesh of the current image can be inter-coded, the reference image index of the current image needs to be passed. Based on such an algorithm, the embodiment of the present application first determines whether the Base Mesh of the current image can be inter-coded. If inter-coding can be used, then it will determine whether the Displacement coefficient of the current image uses skip coding. Otherwise, the Displacement coefficient of the current image is defaulted to the original coding scheme.
[0343] Furthermore, in an embodiment of the present application, the displacement coefficient coding mode for each frame is modified. Specifically, based on the existing V-DMC coding scheme, in inter-frame coding, there is first an identification information that determines the coding method of the base mesh of the current image. If the base mesh of the current image can use inter-frame coding, the reference image index of the current image needs to be transmitted. In this embodiment of the application, regardless of whether the base mesh of the current image can use inter-frame coding, it is necessary to transmit an identification information indicating whether the displacement coefficient of the current image uses skip coding.
[0344] Furthermore, in an embodiment of the present application, the displacement coefficient encoding mode for each frame is modified. Specifically, by using the geometric position information of the initial mesh obtained by dividing the Base Mesh of the current image and the Displacement coefficient of the reference image as a method for reconstructing the geometric position information of the mesh points, this encoding scheme reduces the bitrate of the encoded Displacement coefficients under the premise of reducing the geometric position information of the reconstructed mesh points to a certain extent, thereby improving the coding efficiency of the mesh. In an embodiment of the present application, the relationship between the predicted mesh point geometric position information and the original mesh geometric position information can be used by using parameter fitting to ensure the quality of the reconstructed mesh point geometric position information, thereby reducing the bitrate size of the Displacement coefficients and further improving the coding efficiency of the mesh.
[0345] That is, in an embodiment of the present application, the coding mode of the displacement coefficients in the current image is determined by utilizing the geometric position information of the points at different LOD layers after the Base Mesh is divided in the current image and the displacement coefficients reconstructed at different LOD layers in the reference image. Specifically, at the encoding end, the geometric position information of the corresponding predicted mesh is first obtained by utilizing the geometric position information of the points at different LOD layers after the Base Mesh is divided in the current image and the displacement coefficients reconstructed at different LOD layers in the reference image. Secondly, at the encoding end, the Dist_org between the geometric position information of the reconstructed mesh obtained by the original coding scheme and the original position information is calculated, and the error Dist_skip between the geometric position information of the predicted mesh and the geometric position information of the original mesh is used to adaptively determine the coding mode of the displacement coefficients in the current image. If the displacement coefficients in the current image can be skipped for coding, and the quality of the geometric position information of the reconstructed mesh can be guaranteed to be similar to that of the original coding scheme, this can further improve the efficiency of mesh geometric information coding. It should be noted here that this solution can also further improve the quality of the mesh's reconstructed geometric information by first fitting the relationship between the geometric position information of the predicted mesh and the original mesh position information, and then using the fitted position information to skip encoding. There is no restriction on the fitting method or the use of linear fitting, curve fitting or convolution parameter fitting. This solution is more about making the best use of the relationship between the position information of the points after the Base mesh division of the current image, the Displacement coefficient in the reference image, and the geometric position information of the mesh of the current image. By utilizing the correlation between the three, the encoding bitstream size of the Displacement coefficient can be reduced, and at the same time, the quality of the mesh's point reconstruction geometric information can be guaranteed, thereby further improving the quality of the mesh's point geometric information.
[0346] In the embodiment of the present application, the specific implementation of the aforementioned embodiment is described in detail through the above embodiment. It can be seen that according to the technical solution of the aforementioned embodiment, the displacement coefficient encoding mode in the current image is determined by utilizing the geometric position information of the points at different LOD layers after the base mesh is divided in the current image and the displacement coefficients reconstructed at different LOD layers in the reference image. Specifically, at the encoding end, the geometric position information of the corresponding predicted mesh is first obtained by utilizing the geometric position information of the points at different LOD layers after the base mesh is divided in the current image and the displacement coefficients reconstructed at different LOD layers in the reference image. Secondly, at the encoding end, the Dist_org between the geometric position information of the reconstructed mesh obtained by the original encoding scheme and the geometric position information of the original mesh is calculated, and the error Dist_skip between the geometric position information of the predicted mesh and the geometric position information of the original mesh is used to adaptively determine the encoding mode of the displacement coefficient in the current image. If the displacement coefficient in the current image can be skipped and the quality of the position information of the reconstructed mesh can be guaranteed to be similar to that of the original encoding scheme, the mesh geometric information encoding efficiency can be further improved. However, it should be noted that the embodiment of the present application can also further improve the quality of the reconstructed geometric information of the mesh by first fitting the relationship between the geometric position information of the predicted mesh and the geometric position information of the original mesh, and then skip encoding using the fitted geometric position information. Among them, linear fitting, curve fitting or convolution parameter fitting, etc. can be used for the fitting method, which is not specifically limited here. In addition, the embodiment of the present application is more about making use of the relationship between the geometric position information of the points after the Base mesh of the current image is divided, the Displacement coefficient in the reference image and the geometric position information of the mesh of the current image as much as possible, and by utilizing the correlation between the three, the encoding code stream size of the Displacement coefficient can be reduced, and at the same time, the quality of the point reconstruction geometric information of the mesh can be guaranteed, thereby further improving the quality of the point geometric information of the mesh.
[0347] Taking the inter-frame test environment as an example, as shown in Table 1, BD-Rate is a performance indicator for measuring lossy compression efficiency. When BD-Rate is less than 0, it means that the coding efficiency is improved compared to the existing coding scheme.
[0348] Table 1
[0349] Table 1 shows that a larger threshold Th results in a greater reduction in the displacement coefficient bitrate compared to the original encoding scheme, but this results in a reduction in reconstruction quality. Overall, when threshold Th is set to 1.06, 1.07, and 1.08, the bitrate can be reduced by approximately 12%, 30%, and 38%, respectively. However, D1 decreases by approximately -0.25dB, -0.731dB, and -0.96dB, respectively, and D2 decreases by approximately -0.279dB, -0.8dB, and -1.10dB, respectively. This, in turn, improves bitrate and enhances encoding and decoding efficiency.
[0350] In yet another embodiment of the present application, based on the same inventive concept as the aforementioned embodiment, FIG22 is a schematic diagram showing the composition structure of an encoder provided by an embodiment of the present application. As shown in FIG22 , the encoder 220 may include a first determination unit 2201, a first subdivision unit 2202, a first reconstruction unit 2203, and an encoding unit 2204, wherein:
[0351] The first determining unit 2201 is configured to determine a base grid of the current image according to the original grid of the current image;
[0352] A first subdivision unit 2202 is configured to subdivide the base grid and determine geometric position information of an initial grid of a current layer in a current image;
[0353] a first reconstruction unit 2203 configured to determine a shift coefficient of the current layer in the reference image, and determine geometric position information of a first reconstructed grid of the current layer in the current image based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image;
[0354] The encoding unit 2204 is configured to determine whether to perform encoding processing on the shift coefficients of the current layer in the current image according to the geometric position information of the first reconstructed grid and the geometric position information of the original grid.
[0355] In some embodiments, the reference picture is an encoded picture preceding the current picture.
[0356] In some embodiments, the first determining unit 2201 is further configured to perform downsampling processing on the original grid of the current image to determine the basic grid of the current image.
[0357] In some embodiments, the first determining unit 2201 is further configured to determine the shift coefficient of the current layer in the current image based on the geometric position information of the initial grid of the current layer in the current image and the geometric position information of the original grid of the current layer in the current image.
[0358] In some embodiments, the first determination unit 2201 is further configured to determine the error value of the first vertex between the initial mesh and the original mesh based on the first vertex in the initial mesh, and determine the normal vector of the first vertex; and calculate the shift coefficient of the first vertex based on the error value of the first vertex and the normal vector of the first vertex; wherein the first vertex is any vertex in the initial mesh.
[0359] In some embodiments, the first determination unit 2201 is further configured to perform error calculation on the geometric position information of the first reconstructed grid and the geometric position information of the original grid to determine a first error result; and based on the first error result, determine the encoding mode of the shift coefficient of the current layer in the current image, wherein the encoding mode is used to characterize whether the shift coefficient of the current layer in the current image is to be encoded.
[0360] In some embodiments, the first determination unit 2201 is further configured to determine the geometric position information of the second reconstructed grid of the current layer in the current image based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the current image; and perform error calculation on the geometric position information of the second reconstructed grid and the geometric position information of the original grid to determine a second error result.
[0361] In some embodiments, the first determination unit 2201 is further configured to determine the error ratio between the first error result and the second error result; if the error ratio is less than a preset threshold, the shift coefficient of the current layer in the current image is determined to use the first encoding mode; if the error ratio is greater than the preset threshold, the shift coefficient of the current layer in the current image is determined to use the second encoding mode; wherein the first encoding mode is different from the second encoding mode.
[0362] In some embodiments, the first determination unit 2201 is further configured to perform rate-distortion cost calculation on the shift coefficient of the current layer in the current image according to the first coding mode to determine a first rate-distortion result; and perform rate-distortion cost calculation on the shift coefficient of the current layer in the current image according to the second coding mode to determine a second rate-distortion result; and determine the coding mode of the shift coefficient of the current layer in the current image based on the first rate-distortion result and the second rate-distortion result.
[0363] In some embodiments, the first determination unit 2201 is further configured to determine that the shift coefficient of the current layer in the current image uses the first encoding mode if the first rate-distortion result is smaller than the second rate-distortion result; and to determine that the shift coefficient of the current layer in the current image uses the second encoding mode if the second rate-distortion result is smaller than the first rate-distortion result.
[0364] In some embodiments, the first encoding mode represents skipping encoding of the shift coefficients of the current layer in the current image; and the second encoding mode represents encoding of the shift coefficients of the current layer in the current image.
[0365] In some embodiments, the encoding unit 2204 is further configured to encode the shift coefficient of the current layer in the current image when the shift coefficient of the current layer in the current image uses the second encoding mode, and write the obtained encoding bits into the bitstream.
[0366] In some embodiments, the first determining unit 2201 is further configured to determine a value of second syntax identification information, wherein the second syntax identification information is used to indicate whether the shift coefficients of the current sequence enable the first coding mode;
[0367] The encoding unit 2204 is further configured to perform encoding processing on the value of the second syntax identification information, and write the obtained encoding bits into the bitstream.
[0368] In some embodiments, the first determining unit 2201 is further configured to determine a value of third syntax identification information when the second syntax identification information indicates that the shift coefficients of the current sequence enable the first coding mode, wherein the third syntax identification information is used to indicate whether the shift coefficients of the current image enable the first coding mode;
[0369] The encoding unit 2204 is further configured to perform encoding processing on the value of the third syntax identification information, and write the obtained encoding bits into the bitstream.
[0370] In some embodiments, the first determining unit 2201 is further configured to, when the third syntax identification information indicates that the shift coefficient of the current image uses the first coding mode, determine a value of the first syntax identification information, wherein the first syntax identification information is used to indicate whether the shift coefficient of the current layer in the current image uses the first coding mode;
[0371] The encoding unit 2204 is further configured to perform encoding processing on the value of the first syntax identification information, and write the obtained encoding bits into the bitstream.
[0372] In some embodiments, the first determining unit 2201 is further configured to determine a value of fourth syntax identification information, wherein the fourth syntax identification information is used to indicate whether the base grid of the current image uses an inter-frame processing mode;
[0373] The encoding unit 2204 is further configured to perform encoding processing on the value of the fourth syntax identification information, and write the obtained encoding bits into the bitstream.
[0374] In some embodiments, the first determining unit 2201 is further configured to determine whether to perform encoding processing on the shift coefficients of the current layer in the current image when the base grid of the current image uses an inter-frame processing method.
[0375] In some embodiments, the first determining unit 2201 is further configured to determine a reference image index of the current image based on the reference image when the base grid of the current image uses an inter-frame processing manner;
[0376] The encoding unit 2204 is further configured to perform encoding processing on the reference image index of the current image, and write the obtained encoding bits into the bitstream.
[0377] In some embodiments, the first determining unit 2201 is further configured to perform a lifting transformation on the shift coefficients of the current layer in the current image to determine lifting transformation coefficients; perform a quantization process on the lifting transformation coefficients to determine quantization coefficients; and perform a coefficient reorganization process on the quantization coefficients to determine a two-dimensional image;
[0378] The encoding unit 2204 is further configured to perform encoding processing on the two-dimensional image and write the obtained encoding bits into the bit stream.
[0379] In some embodiments, the first reconstruction unit 2203 is further configured to determine a mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the first reconstructed grid of the current layer in the current image; and determine the geometric position information of the first reconstructed grid of the current layer in the current image based on the mapping relationship and the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image.
[0380] In some embodiments, the mapping relationship includes at least one of the following: a mapping relationship based on a linear function, a mapping relationship based on a nonlinear function, and a mapping relationship based on a neural grid.
[0381] In some embodiments, the first determining unit 2201 is further configured to determine mapping indication information of a current layer in the current image based on the mapping relationship;
[0382] The encoding unit 2204 is further configured to perform encoding processing on the mapping indication information and write the obtained encoding bits into the bit stream.
[0383] In some embodiments, the first subdivision unit 2202 is further configured to iteratively divide the basic grid according to the grid subdivision mode to determine geometric position information of the initial grid of the current layer in the current image.
[0384] In some embodiments, the mesh subdivision mode includes: a subdivision algorithm and a number of subdivision iterations.
[0385] In some embodiments, the subdivision algorithm is a linear interpolation algorithm.
[0386] It is understandable that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and of course it can also be a module, or it can be non-modular. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.
[0387] If the integrated unit is implemented as a software functional module and not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or grid device, etc.) or a processor to perform all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0388] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the encoder 220. The computer-readable storage medium stores a computer program, and when the computer program is executed by the first processor, it implements the method described in any one of the aforementioned embodiments.
[0389] Based on the composition of the encoder 220 and the computer-readable storage medium, refer to Figure 23, which shows a specific hardware structure diagram of the encoder 220 provided in an embodiment of the present application. As shown in Figure 23, the encoder 220 may include: a first communication interface 2301, a first memory 2302 and a first processor 2303; each component is coupled together through a first bus system 2304. It can be understood that the first bus system 2304 is used to achieve connection and communication between these components. In addition to the data bus, the first bus system 2304 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as the first bus system 2304 in Figure 23. Among them,
[0390] The first communication interface 2301 is used to receive and send signals when sending and receiving information with other external network elements;
[0391] A first memory 2302 is used to store computer programs that can be run on the first processor 2303;
[0392] The first processor 2303 is configured to, when running the computer program, execute:
[0393] Determine the base grid of the current image according to the original grid of the current image;
[0394] Subdivide the base grid and determine the geometric position information of the initial grid of the current layer in the current image;
[0395] determining a shift coefficient of the current layer in the reference image, and determining geometric position information of a first reconstructed grid of the current layer in the current image based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image;
[0396] According to the geometric position information of the first reconstructed grid and the geometric position information of the original grid, it is determined whether to perform encoding processing on the shift coefficients of the current layer in the current image.
[0397] It is understood that the first memory 2302 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The first memory 2302 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0398] The first processor 2303 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the first processor 2303. The above-mentioned first processor 2303 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented as a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the first memory 2302 , and the first processor 2303 reads the information in the first memory 2302 and completes the steps of the above method in combination with its hardware.
[0399] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in a processor or outside a processor.
[0400] Optionally, as another embodiment, the first processor 2303 is further configured to execute the method described in any one of the aforementioned embodiments when running the computer program.
[0401] This embodiment provides an encoder in which encoding processing is performed based on whether the shift coefficient of the current layer in the current image is encoded. If encoding the shift coefficient of the current layer in the current image is skipped, the geometric position information of the first reconstructed grid can be determined based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image. This not only reduces the code rate of the shift coefficient, but also ensures the reconstructed geometric quality of the grid midpoint, thereby further improving the geometric information quality of the grid midpoint, and thus improving the encoding efficiency.
[0402] Based on the same inventive concept as the above embodiments, referring to FIG24 , a schematic diagram of the structure of a decoder provided by an embodiment of the present application is shown. As shown in FIG24 , the decoder 240 may include a second determination unit 2401, a second subdivision unit 2402, a decoding unit 2403, and a second reconstruction unit 2404, wherein:
[0403] The second determining unit 2401 is configured to determine a base grid of the current image;
[0404] The second subdivision unit 2402 is configured to subdivide the base grid and determine geometric position information of the initial grid of the current layer in the current image;
[0405] The decoding unit 2403 is configured to decode the code stream, determine the first syntax identification information; and determine the shift coefficient of the current layer in the reference image when the first syntax identification information indicates that the shift coefficient of the current layer in the current image uses the first decoding mode;
[0406] The second reconstruction unit 2404 is configured to determine the geometric position information of the reconstructed grid of the current layer in the current image according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image.
[0407] In some embodiments, the reference picture is a decoded picture preceding the current picture.
[0408] In some embodiments, the decoding unit 2403 is further configured to decode the code stream and determine the second syntax identification information; when the second syntax identification information indicates that the shift coefficient of the current sequence enables the first decoding mode, decode the code stream and determine the third syntax identification information; when the third syntax identification information indicates that the shift coefficient of the current image enables the first decoding mode, decode the code stream and determine the first syntax identification information; wherein the first syntax identification information is used to indicate whether the shift coefficient of the current layer uses the first decoding mode, and the current sequence includes the current image, and the LOD layer divided by the current image includes the current layer.
[0409] In some embodiments, the decoding unit 2403 is further configured to, when the first syntax identification information indicates that the shift coefficient of the current layer in the current image uses the second decoding mode, decode the code stream to determine the shift coefficient of the current layer in the current image;
[0410] The second determining unit 2401 is further configured to determine the geometric position information of the reconstructed grid of the current layer in the current image according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the current image.
[0411] In some embodiments, the first decoding mode is different from the second decoding mode; wherein: the first decoding mode represents skipping decoding of the shift coefficients of the current layer in the current image; the second decoding mode represents decoding of the shift coefficients of the current layer in the current image.
[0412] In some embodiments, the decoding unit 2403 is further configured to decode the code stream and determine the fourth syntax identification information; when the fourth syntax identification information indicates that the basic grid of the current image uses the inter-frame processing method, perform the step of decoding the code stream and determining the first syntax identification information.
[0413] In some embodiments, the decoding unit 2403 is further configured to decode the code stream and determine the reference image index of the current image when the fourth syntax identification information indicates that the base grid of the current image uses the inter-frame processing mode;
[0414] The second determining unit 2401 is further configured to determine a reference image according to a reference image index of the current image.
[0415] In some embodiments, the decoding unit 2403 is further configured to, when the fourth syntax identification information indicates that the base grid of the current image does not use the inter-frame processing mode, decode the code stream and determine the shift coefficient of the current layer in the current image;
[0416] The second determining unit 2401 is further configured to determine the geometric position information of the reconstructed grid of the current layer in the current image according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the current image.
[0417] In some embodiments, the decoding unit 2403 is further configured to decode the code stream to determine a two-dimensional image of a current layer in the current image;
[0418] The second determining unit 2401 is further configured to perform coefficient reorganization processing on the two-dimensional image to determine the lifting transformation coefficients of the current layer; and perform inverse transformation processing on the lifting transformation coefficients of the current layer to determine the shift coefficients of the current layer.
[0419] In some embodiments, the second determination unit 2401 is further configured to perform coefficient reorganization processing on the two-dimensional image to determine the quantization coefficients of the current layer; and perform inverse quantization processing on the quantization coefficients of the current layer to determine the lifting transformation coefficients of the current layer.
[0420] In some embodiments, the second reconstruction unit 2404 is further configured to determine the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image; and determine the geometric position information of the reconstructed grid of the current layer in the current image based on the mapping relationship and the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image.
[0421] In some embodiments, the mapping relationship includes at least one of the following: a mapping relationship based on a linear function, a mapping relationship based on a nonlinear function, and a mapping relationship based on a neural grid.
[0422] In some embodiments, the decoding unit 2403 is further configured to decode the code stream and determine mapping indication information of the current layer in the current image;
[0423] The second determination unit 2401 is further configured to determine the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image and the geometric position information of the reconstructed grid of the current layer in the current image according to the mapping indication information.
[0424] In some embodiments, the mapping indication information includes first indication information; the second determination unit 2401 is further configured to determine the fitting parameters of the mapping relationship based on the first indication information; and determine the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image based on the fitting parameters.
[0425] In some embodiments, the mapping indication information also includes second indication information; the second determination unit 2401 is further configured to determine the type of mapping relationship based on the second indication information; and determine the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image based on the type of mapping relationship and the fitting parameters.
[0426] In some embodiments, the mapping indication information includes third indication information; the second determination unit 2401 is further configured to determine the index serial number of the mapping relationship based on the third indication information; and determine the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image based on the index serial number of the mapping relationship.
[0427] In some embodiments, the decoding unit 2403 is further configured to decode the code stream to determine a base grid of the current image.
[0428] In some embodiments, the second subdivision unit 2402 is further configured to iteratively divide the basic grid according to the grid subdivision mode to determine geometric position information of the initial grid of the current layer in the current image.
[0429] In some embodiments, the mesh subdivision mode includes: a subdivision algorithm and a number of subdivision iterations.
[0430] In some embodiments, the subdivision algorithm is a linear interpolation algorithm.
[0431] It is understood that in this embodiment, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular system. Furthermore, the various components in this embodiment can be integrated into a single processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The aforementioned integrated units can be implemented in the form of hardware or software functional modules.
[0432] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, this embodiment provides a computer-readable storage medium, which is applied to the decoder 240 and stores a computer program. When the computer program is executed by the second processor, it implements any of the methods in the aforementioned embodiments.
[0433] Based on the composition of the decoder 240 and the computer-readable storage medium, refer to Figure 25, which shows a specific hardware structure diagram of the decoder 240 provided in an embodiment of the present application. As shown in Figure 25, the decoder 240 may include: a second communication interface 2501, a second memory 2502 and a second processor 2503; each component is coupled together through a second bus system 2504. It can be understood that the second bus system 2504 is used to achieve connection and communication between these components. In addition to the data bus, the second bus system 2504 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as the second bus system 2504 in Figure 25. Among them,
[0434] The second communication interface 2501 is used to receive and send signals during the process of sending and receiving information between other external network elements;
[0435] The second memory 2502 is used to store computer programs that can be run on the second processor 2503;
[0436] The second processor 2503 is configured to, when running the computer program, execute:
[0437] Determine the base grid of the current image;
[0438] Subdivide the base grid and determine the geometric position information of the initial grid of the current layer in the current image;
[0439] Decoding the code stream to determine first syntax identification information;
[0440] When the first syntax identification information indicates that the shift coefficient of the current layer in the current image uses the first decoding mode, determining the shift coefficient of the current layer in the reference image;
[0441] The geometric position information of the reconstructed grid of the current layer in the current image is determined according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image.
[0442] Optionally, as another embodiment, the second processor 2503 is further configured to execute any one of the methods described in the foregoing embodiments when running the computer program.
[0443] It can be understood that the hardware functions of the second memory 2502 are similar to those of the first memory 2302, and the hardware functions of the second processor 2503 are similar to those of the first processor 2303; they will not be described in detail here.
[0444] This embodiment provides a decoder, in which whether the shift coefficient of the current layer in the current image is skipped for decoding is indicated based on the first syntax identification information. When the shift coefficient of the current layer in the current image is skipped for decoding, the geometric position information of the reconstructed grid is determined based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image. This not only reduces the code stream of the shift coefficient, but also ensures the reconstructed geometric quality of the grid midpoint, thereby further improving the geometric information quality of the grid midpoint, and further improving the encoding and decoding efficiency.
[0445] In yet another embodiment of the present application, referring to FIG26 , a schematic diagram of the structure of a coding and decoding system provided by an embodiment of the present application is shown. As shown in FIG26 , the coding and decoding system 260 may include an encoder 2601 and a decoder 2602 .
[0446] In the embodiment of the present application, the encoder 2601 may be the encoder described in any one of the aforementioned embodiments, and the decoder 2602 may be the decoder described in any one of the aforementioned embodiments.
[0447] It should be noted that, in this application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0448] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0449] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0450] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0451] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0452] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims. Industrial Applicability
[0453] In an embodiment of the present application, at the encoding end, a base grid of the current image is determined based on the original grid of the current image; the base grid is subdivided to determine the geometric position information of the initial grid of the current layer in the current image; a shift coefficient of the current layer in the reference image is determined, and based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image, the geometric position information of a first reconstructed grid of the current layer in the current image is determined; and based on the geometric position information of the first reconstructed grid and the geometric position information of the original grid, it is determined whether to encode the shift coefficient of the current layer in the current image. At the decoding end, a base grid of the current image is determined; the base grid is subdivided to determine the geometric position information of the initial grid of the current layer in the current image; a bitstream is decoded to determine first syntax identification information; and when the first syntax identification information indicates that the shift coefficient of the current layer in the current image uses a first decoding mode, the shift coefficient of the current layer in the reference image is determined; and based on the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image, the geometric position information of the reconstructed grid of the current layer in the current image is determined. In this way, based on the geometric position information of the initial grid after the basic grid of the current image is subdivided and the geometric position information of the first reconstructed grid obtained by the shift coefficient in the reference image, the adaptive encoding of the shift coefficient in the current image can be determined; if the encoding of the shift coefficient in the current image is skipped, the shift coefficient in the current image does not need to be transmitted in the code stream at this time. The decoding end can determine the geometric position information of the first reconstructed grid based on the geometric position information of the initial grid and the shift coefficient in the reference image, which not only reduces the code stream of the shift coefficient, but also ensures the reconstructed geometric quality of the grid midpoint, thereby further improving the geometric information quality of the grid midpoint, and thus improving the encoding and decoding efficiency.
Claims
1. A decoding method, applied to a decoder, the method comprising: Determine the base grid of the current image; Subdividing the basic grid to determine geometric position information of an initial grid of a current layer in the current image; Decoding the code stream to determine first syntax identification information; When the first syntax identification information indicates that the shift coefficient of the current layer in the current image uses the first decoding mode, determining the shift coefficient of the current layer in the reference image; The geometric position information of the reconstructed grid of the current layer in the current image is determined according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image.
2. The method according to claim 1, wherein: The reference image is a decoded image before the current image.
3. The method according to claim 1, wherein: The decoding of the code stream to determine the first syntax identification information includes: Decoding the code stream to determine second syntax identification information; When the second syntax identification information indicates that the shift coefficient of the current sequence enables the first decoding mode, decoding the code stream to determine the third syntax identification information; When the third syntax identification information indicates that the shift coefficient of the current image enables the first decoding mode, decoding the bitstream to determine the first syntax identification information; The first syntax identification information is used to indicate whether the shift coefficient of the current layer uses a first decoding mode, and the current sequence includes the current image, and the LOD layer divided by the current image includes the current layer.
4. The method according to claim 1, wherein: The method further comprises: When the first syntax identification information indicates that the shift coefficient of the current layer in the current image uses the second decoding mode, decoding the code stream to determine the shift coefficient of the current layer in the current image; The geometric position information of the reconstructed grid of the current layer in the current image is determined according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the current image.
5. The method according to claim 4, wherein: The first decoding mode is different from the second decoding mode; wherein: The first decoding mode represents skipping of decoding shift coefficients of a current layer in the current image; The second decoding mode represents decoding shift coefficients of a current layer in the current image.
6. The method according to claim 1, wherein: The method further comprises: Decoding the code stream to determine fourth syntax identification information; When the fourth syntax identification information indicates that the basic grid of the current image uses an inter-frame processing mode, a step of decoding a code stream and determining the first syntax identification information is performed.
7. The method according to claim 6, wherein: When the fourth syntax identification information indicates that the base grid of the current image uses an inter-frame processing mode, the method further includes: Decoding a code stream to determine a reference image index of the current image; The reference image is determined according to the reference image index of the current image.
8. The method according to claim 6, wherein: The method further comprises: When the fourth syntax identification information indicates that the base grid of the current image does not use the inter-frame processing mode, decoding the code stream to determine the shift coefficient of the current layer in the current image; The geometric position information of the reconstructed grid of the current layer in the current image is determined according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the current image.
9. The method according to claim 4 or 8, wherein: The decoding code stream determines the shift coefficient of the current layer in the current image, including: Decoding the bitstream to determine a two-dimensional image of a current layer in the current image; Performing coefficient reorganization processing on the two-dimensional image to determine the lifting transformation coefficient of the current layer; An inverse transform process is performed on the lifting transform coefficients of the current layer to determine the shift coefficients of the current layer.
10. The method according to claim 9, wherein: The performing coefficient reorganization processing on the two-dimensional image to determine the lifting transformation coefficient of the current layer includes: Performing coefficient reorganization processing on the two-dimensional image to determine the quantization coefficient of the current layer; Dequantization is performed on the quantized coefficients of the current layer to determine the lifting transformation coefficients of the current layer.
11. The method according to claim 1, wherein: The step of determining the geometric position information of the reconstructed grid of the current layer in the current image according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image comprises: Determine a mapping relationship between geometric position information of an initial grid of a current layer in the current image, a shift coefficient of the current layer in the reference image, and geometric position information of a reconstructed grid of the current layer in the current image; The geometric position information of the reconstructed grid of the current layer in the current image is determined according to the mapping relationship, the geometric position information of the initial grid of the current layer in the current image, and the shift coefficient of the current layer in the reference image.
12. The method according to claim 11, wherein: The mapping relationship includes at least one of the following: a mapping relationship based on a linear function, a mapping relationship based on a nonlinear function, and a mapping relationship based on a neural grid.
13. The method according to claim 11, wherein: The determining of the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image includes: Decoding the code stream to determine mapping indication information of a current layer in the current image; According to the mapping indication information, a mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image and the geometric position information of the reconstructed grid of the current layer in the current image is determined.
14. The method according to claim 13, wherein: The mapping indication information includes first indication information; The determining, according to the mapping indication information, a mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image comprises: Determining fitting parameters of the mapping relationship according to the first indication information; According to the fitting parameters, a mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image and the geometric position information of the reconstructed grid of the current layer in the current image is determined.
15. The method according to claim 14, wherein: The mapping indication information further includes second indication information; The determining, according to the mapping indication information, a mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image comprises: Determine the type of the mapping relationship according to the second indication information; According to the type of the mapping relationship and the fitting parameters, the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image and the geometric position information of the reconstructed grid of the current layer in the current image is determined.
16. The method according to claim 13, wherein: The mapping indication information includes third indication information; The determining, according to the mapping indication information, a mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image, and the geometric position information of the reconstructed grid of the current layer in the current image comprises: Determine the index number of the mapping relationship according to the third indication information; According to the index number of the mapping relationship, the mapping relationship between the geometric position information of the initial grid of the current layer in the current image, the shift coefficient of the current layer in the reference image and the geometric position information of the reconstructed grid of the current layer in the current image is determined.
17. The method according to any one of claims 1 to 16, wherein: The step of determining a base grid of the current image comprises: The code stream is decoded to determine a base grid of the current image.
18. The method according to any one of claims 1 to 16, wherein: The step of subdividing the basic grid to determine geometric position information of an initial grid of a current layer in the current image includes: The basic grid is iteratively divided according to a grid subdivision mode to determine geometric position information of an initial grid of a current layer in the current image.
19. The method according to claim 18, wherein: The grid subdivision mode includes: a subdivision algorithm and a subdivision iteration number.
20. The method according to claim 19, wherein: The subdivision algorithm is a linear interpolation algorithm.
21. A coding method, applied to an encoder, the method comprising: Determine a base grid of the current image according to an original grid of the current image; Subdividing the basic grid to determine geometric position information of an initial grid of a current layer in the current image; Determine a shift coefficient of the current layer in the reference image, and determine geometric position information of a first reconstructed grid of the current layer in the current image according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image; Determine whether to perform encoding processing on the shift coefficients of the current layer in the current image according to the geometric position information of the first reconstructed grid and the geometric position information of the original grid.
22. The method according to claim 21, wherein: The reference image is an encoded image before the current image.
23. The method according to claim 21, wherein: The step of determining the base grid of the current image according to the original grid of the current image comprises: Down-sampling is performed on the original grid of the current image to determine a basic grid of the current image.
24. The method according to claim 21, wherein: The method further comprises: The shift coefficient of the current layer in the current image is determined according to the geometric position information of the initial grid of the current layer in the current image and the geometric position information of the original grid of the current layer in the current image.
25. The method according to claim 24, wherein: The determining, according to the geometric position information of the initial grid of the current layer in the current image and the geometric position information of the original grid of the current layer in the current image, the shift coefficient of the current layer in the current image comprises: Based on a first vertex in the initial mesh, determining an error value of the first vertex between the initial mesh and the original mesh, and determining a normal vector of the first vertex; A shift coefficient of the first vertex is calculated according to the error value of the first vertex and the normal vector of the first vertex; wherein the first vertex is any vertex in the initial mesh.
26. The method of claim 21, wherein: The determining whether to perform encoding processing on the shift coefficient of the current layer in the current image according to the geometric position information of the first reconstructed grid and the geometric position information of the original grid includes: Performing error calculation on the geometric position information of the first reconstructed grid and the geometric position information of the original grid to determine a first error result; Based on the first error result, a coding mode of the shift coefficient of the current layer in the current image is determined, wherein the coding mode is used to indicate whether to perform coding processing on the shift coefficient of the current layer in the current image.
27. The method according to claim 26, wherein: The method further comprises: Determining geometric position information of a second reconstructed grid of the current layer in the current image according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the current image; An error calculation is performed on the geometric position information of the second reconstructed grid and the geometric position information of the original grid to determine a second error result.
28. The method according to claim 27, wherein: The determining, based on the first error result, a coding mode of a shift coefficient of a current layer in the current image comprises: determining an error ratio between the first error result and the second error result; If the error ratio is less than a preset threshold, determining that the shift coefficient of the current layer in the current image uses a first coding mode; If the error ratio is greater than a preset threshold, determining that the shift coefficient of the current layer in the current image uses a second coding mode; The first encoding mode is different from the second encoding mode.
29. The method according to claim 27, wherein: The method further comprises: Performing rate-distortion cost calculation on the shift coefficient of the current layer in the current image according to the first coding mode to determine a first rate-distortion result; and performing rate-distortion cost calculation on the shift coefficient of the current layer in the current image according to the second coding mode to determine a second rate-distortion result; A coding mode of a shift coefficient of a current layer in the current image is determined according to the first rate-distortion result and the second rate-distortion result.
30. The method of claim 29, wherein: The determining, according to the first rate-distortion result and the second rate-distortion result, a coding mode of a shift coefficient of a current layer in the current image comprises: If the first rate-distortion result is less than the second rate-distortion result, determining that the shift coefficient of the current layer in the current image uses a first coding mode; If the second rate-distortion result is less than the first rate-distortion result, it is determined that the shift coefficient of the current layer in the current image uses a second coding mode.
31. The method according to any one of claims 28 to 30, wherein: The first coding mode represents skipping of coding shift coefficients of a current layer in the current image; The second coding mode represents encoding the shift coefficients of the current layer in the current image.
32. The method according to claim 28 or 30, wherein: The method further comprises: When the shift coefficient of the current layer in the current image uses the second coding mode, the shift coefficient of the current layer in the current image is coded, and the obtained coded bits are written into the bitstream.
33. The method of claim 21, wherein: The method further comprises: Determine a value of second grammar identification information, wherein the second grammar identification information is used to indicate whether the shift coefficient of the current sequence enables the first coding mode; The value of the second syntax identification information is encoded, and the obtained encoded bits are written into a bit stream.
34. The method of claim 33, wherein: The method further comprises: When the second grammar identification information indicates that the shift coefficient of the current sequence enables the first coding mode, determine a value of third grammar identification information, wherein the third grammar identification information is used to indicate whether the shift coefficient of the current image enables the first coding mode; The value of the third syntax identification information is encoded, and the obtained encoded bits are written into a bit stream.
35. The method of claim 34, wherein: The method further comprises: When the third syntax identification information indicates that the shift coefficient of the current image enables the first coding mode, determine a value of the first syntax identification information, wherein the first syntax identification information is used to indicate whether the shift coefficient of the current layer in the current image uses the first coding mode; The value of the first syntax identification information is encoded, and the obtained encoded bits are written into a bit stream.
36. The method of claim 21, wherein: The method further comprises: Determining a value of fourth syntax identification information, wherein the fourth syntax identification information is used to indicate whether the basic grid of the current image uses an inter-frame processing mode; The value of the fourth syntax identification information is encoded, and the obtained encoded bits are written into a bit stream.
37. The method of claim 36, wherein: The method further comprises: When the base grid of the current image uses an inter-frame processing method, a step of determining whether to perform encoding processing on the shift coefficients of the current layer in the current image is performed.
38. The method of claim 37, wherein: When the base grid of the current image uses an inter-frame processing mode, the method further includes: Determining a reference image index of the current image according to the reference image; The reference image index of the current image is encoded, and the obtained encoding bits are written into a bitstream.
39. The method of claim 32, wherein: The encoding process is performed on the shift coefficient of the current layer in the current image, and the obtained encoding bits are written into the bit stream, including: Performing a lifting transformation on the shift coefficients of the current layer in the current image to determine lifting transformation coefficients; quantizing the lifting transform coefficients to determine quantization coefficients; Performing coefficient reorganization processing on the quantized coefficients to determine a two-dimensional image; The two-dimensional image is coded and the obtained coded bits are written into a bit stream.
40. The method according to any one of claims 21 to 39, wherein: The step of determining the geometric position information of the first reconstructed grid of the current layer in the current image according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image comprises: Determine a mapping relationship between geometric position information of an initial grid of a current layer in the current image, a shift coefficient of the current layer in the reference image, and geometric position information of a first reconstructed grid of the current layer in the current image; The geometric position information of the first reconstructed grid of the current layer in the current image is determined according to the mapping relationship, the geometric position information of the initial grid of the current layer in the current image, and the shift coefficient of the current layer in the reference image.
41. The method of claim 40, wherein: The mapping relationship includes at least one of the following: a mapping relationship based on a linear function, a mapping relationship based on a nonlinear function, and a mapping relationship based on a neural grid.
42. The method of claim 40, wherein: The method further comprises: Based on the mapping relationship, determining mapping indication information of a current layer in the current image; The mapping indication information is coded and the obtained coded bits are written into a bit stream.
43. The method according to any one of claims 21 to 39, wherein: The step of subdividing the basic grid to determine geometric position information of an initial grid of a current layer in the current image includes: The basic grid is iteratively divided according to a grid subdivision mode to determine geometric position information of an initial grid of a current layer in the current image.
44. The method of claim 43, wherein: The grid subdivision mode includes: a subdivision algorithm and a subdivision iteration number.
45. The method of claim 44, wherein: The subdivision algorithm is a linear interpolation algorithm.
46. A code stream, the code stream is generated by bit encoding according to information to be encoded; wherein, The information to be encoded includes at least one of the following: The value of the first grammar identifier information, the value of the second grammar identifier information, the value of the third grammar identifier information, the value of the fourth grammar identifier information, the reference image index of the current image, the shift coefficient of the current layer in the current image, and the mapping indication information of the current layer in the current image; wherein the second grammar identifier information is used to indicate whether the shift coefficient of the current sequence enables the first coding mode, the third grammar identifier information is used to indicate whether the shift coefficient of the current image enables the first coding mode, the first grammar identifier information is used to indicate whether the shift coefficient of the current layer in the current image uses the first coding mode, and the fourth grammar identifier information is used to indicate whether the basic grid of the current image uses inter-frame processing; and the current sequence includes the current image, and the LOD layers divided by the current image include the current layer.
47. An encoder, comprising a first determining unit, a first subdividing unit, a first reconstructing unit and an encoding unit, wherein: The first determining unit is configured to determine a basic grid of the current image according to an original grid of the current image; The first subdivision unit is configured to subdivide the basic grid and determine geometric position information of an initial grid of a current layer in the current image; The first reconstruction unit is configured to determine a shift coefficient of the current layer in the reference image, and determine geometric position information of a first reconstructed grid of the current layer in the current image according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image; The encoding unit is configured to determine whether to perform encoding processing on the shift coefficients of the current layer in the current image according to the geometric position information of the first reconstructed grid and the geometric position information of the original grid.
48. An encoder, comprising a first memory and a first processor, wherein: The first memory is used to store a computer program that can be run on the first processor; The first processor is configured to execute the method according to any one of claims 21 to 45 when running the computer program.
49. A decoder, comprising a second determining unit, a second subdividing unit, a decoding unit and a second reconstructing unit, wherein: The second determining unit is configured to determine a base grid of the current image; The second subdivision unit is configured to subdivide the basic grid and determine geometric position information of an initial grid of a current layer in the current image; The decoding unit is configured to decode the code stream, determine the first syntax identification information; and determine the shift coefficient of the current layer in the reference image when the first syntax identification information indicates that the shift coefficient of the current layer in the current image uses the first decoding mode; The second reconstruction unit is configured to determine the geometric position information of the reconstructed grid of the current layer in the current image according to the geometric position information of the initial grid of the current layer in the current image and the shift coefficient of the current layer in the reference image.
50. A decoder, comprising a second memory and a second processor, wherein: The second memory is used to store a computer program that can be run on the second processor; The second processor is configured to execute the method according to any one of claims 1 to 20 when running the computer program.
51. A computer-readable storage medium, wherein: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 20 is implemented, or the method according to any one of claims 21 to 45 is implemented.