Coding and decoding method, encoder, decoder, code stream and storage medium

CN120982094APending Publication Date: 2025-11-18GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380096895.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-06-30
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

The prior art fails to effectively utilize the strong correlation between the geometric information of the grid in dynamic grid encoding, resulting in low encoding efficiency and affecting the grid compression performance.

Method used

By using the first mode on the encoding and decoding ends, the basic grid geometry information of the current image is directly determined based on the basic grid geometry information of the reference image, and the encoding and decoding processing of the motion vector is skipped, thereby reducing the encoding code stream size of the geometric information. .

Benefits of technology

It improves the coding efficiency of geometric information, improves the grid compression performance, and reduces the size of the coded code stream.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120982094A_ABST
    Figure CN120982094A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a decoding method, and the method comprises the steps: decoding a code stream at a decoding end through a decoder, and determining first identification information corresponding to a current image; under the condition that the first identification information indicates that a first basic grid corresponding to the current image uses a first mode, determining geometric information of the first basic grid according to geometric information of a second basic grid corresponding to a reference image of the current image; and determining a reconstructed original grid of the current image according to the geometric information of the first basic grid.
Need to check novelty before this filing date? Find Prior Art

Description

Coding and decoding method, encoder, decoder, code stream and storage medium Technical Field

[0001] The embodiments of the present application relate to the technical field of grid compression coding, and in particular to a coding and decoding method, an encoder, a decoder, a code stream, and a storage medium. Background Art

[0002] In the standard reference software for Dynamic Mesh Coding (DMC) provided by the Moving Picture Experts Group (MPEG), encoding and decoding the geometric information of a mesh mainly involves organizing and compressing the base mesh or motion vectors corresponding to the original mesh.

[0003] However, the current common geometric information compression schemes do not take into account the strong correlation between the geometric information of most grids, which reduces the coding efficiency of geometric information and further affects the grid compression performance.

[0004] Summary of the Invention

[0005] The embodiments of the present application provide a coding and decoding method, an encoder, a decoder, a code stream, and a storage medium, which can improve the coding efficiency of geometric information and thereby enhance the mesh compression performance.

[0006] The technical solution of the embodiment of the present application can be implemented as follows:

[0007] In a first aspect, an embodiment of the present application provides a decoding method, applied to a decoder, the method comprising:

[0008] Decoding the code stream to determine first identification information corresponding to the current image;

[0009] When the first identification information indicates that the first basic mesh corresponding to the current image uses the first mode, determining the geometric information of the first basic mesh according to geometric information of a second basic mesh corresponding to a reference image of the current image;

[0010] A reconstructed original mesh of the current image is determined according to the geometric information of the first basic mesh.

[0011] In a second aspect, an embodiment of the present application provides a decoding method applied to a decoder, wherein the decoder includes a video decoder and a grid decoder, and the grid decoder and the video decoder are used to perform the decoding method as described in the first aspect.

[0012] In a third aspect, an embodiment of the present application provides an encoding method, applied to an encoder, the method comprising:

[0013] Determining first identification information corresponding to the current image;

[0014] When the first identification information indicates that the first basic mesh corresponding to the current image uses the first mode, determining the geometric information of the first basic mesh according to geometric information of a second basic mesh corresponding to a reference image of the current image;

[0015] A shift coefficient corresponding to the current image is determined according to the geometric information of the first basic grid, and the shift coefficient is written into a bitstream.

[0016] In a fourth aspect, an embodiment of the present application provides an encoding method applied to an encoder, wherein the encoder includes a video encoder and a grid encoder, and the grid encoder and the video encoder are used to execute the encoding method as described in the third aspect.

[0017] In a fifth aspect, an embodiment of the present application provides a code stream, wherein the code stream is generated by bit encoding based on information to be encoded; wherein the information to be encoded includes at least one of the following:

[0018] First identification information, second identification information, third identification information, shift coefficient, motion vector information, fitting parameters, and geometric information of the first basic grid.

[0019] In a sixth aspect, an embodiment of the present application provides an encoder, comprising: a first determining unit and an encoding unit; wherein,

[0020] The first determining unit is configured to determine first identification information corresponding to the current image; if the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, determine geometric information of the first basic grid based on geometric information of a second basic grid corresponding to a reference image of the current image; and determine a shift coefficient corresponding to the current image based on the geometric information of the first basic grid;

[0021] The encoding unit is configured to write the shift coefficient into a bit stream.

[0022] In a seventh aspect, an embodiment of the present application provides an encoder, comprising: a first memory and a first processor; wherein,

[0023] a first memory for storing a computer program capable of running on the first processor;

[0024] The first processor is configured to execute the encoding method described in the third aspect when running a computer program.

[0025] In an eighth aspect, an embodiment of the present application provides a decoder, comprising: a decoding unit, a second determining unit; wherein,

[0026] The decoding unit is configured to decode the code stream;

[0027] The second determination unit is configured to determine first identification information corresponding to the current image; when the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, determine the geometric information of the first basic grid according to the geometric information of the second basic grid corresponding to the reference image of the current image; and determine the reconstructed original grid of the current image according to the geometric information of the first basic grid.

[0028] In a ninth aspect, an embodiment of the present application provides a decoder, comprising: a second memory and a second processor; wherein:

[0029] a second memory for storing a computer program capable of running on the second processor;

[0030] The second processor is configured to execute the method according to the first aspect when running a computer program.

[0031] In the tenth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed, it implements the method described in the first aspect or the second aspect, or implements the method described in the third aspect or the fourth aspect.

[0032] The embodiments of the present application provide a coding and decoding method, an encoder, a decoder, a bitstream, and a storage medium. At the decoding end, the decoder decodes the bitstream and determines first identification information corresponding to the current image; if the first identification information indicates that the first base grid corresponding to the current image uses the first mode, the geometric information of the first base grid is determined based on the geometric information of the second base grid corresponding to the reference image of the current image; and the reconstructed original grid of the current image is determined based on the geometric information of the first base grid. At the encoding end, the encoder determines the first identification information corresponding to the current image; if the first identification information indicates that the first base grid corresponding to the current image uses the first mode, the geometric information of the first base grid is determined based on the geometric information of the second base grid corresponding to the reference image of the current image; and the shift coefficient corresponding to the current image is determined based on the geometric information of the first base grid, and the shift coefficient is written into the bitstream. Therefore, in the embodiments of the present application, if the first identification information corresponding to the current image indicates that the corresponding first base grid uses the first mode, the geometric information of the first base grid can be directly determined based on the geometric information of the second base grid corresponding to the reference image, that is, the base grid of the reference image is directly used to complete the prediction of the base grid of the current image, and the motion vector between the reference image and the current image can no longer be encoded and decoded. That is to say, in the embodiment of the present application, the motion vector is skipped by using the first mode, thereby reducing the encoding code stream size of the geometric information of the grid, thereby improving the encoding efficiency of the geometric information and enhancing the grid compression performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] FIG1A is a schematic diagram of a three-dimensional grid image 1;

[0034] FIG1B is a partially enlarged schematic diagram of a three-dimensional grid image;

[0035] Figure 2 is a schematic diagram of the connection method of the three-dimensional grid;

[0036] FIG3A is a second schematic diagram of a three-dimensional grid image;

[0037] FIG3B is a schematic diagram of a grid data storage format;

[0038] FIG3C is a schematic diagram of properties of a three-dimensional grid image;

[0039] FIG4 is a schematic diagram showing the composition of the overall framework of grid coding;

[0040] FIG5A is a schematic diagram of preprocessing of a two-dimensional curve;

[0041] FIG5B is a schematic diagram of generating a shift coefficient;

[0042] FIG6A is a first schematic diagram of quantization processing of grid geometric position information;

[0043] FIG6B is a second schematic diagram of quantization processing of grid geometric position information;

[0044] FIG7A is a schematic diagram of coding of the connection relationship of triangular facets;

[0045] FIG7B is a schematic diagram of encoding geometric position information;

[0046] FIG7C is a schematic diagram of texture coordinate encoding;

[0047] FIG8 is a schematic diagram showing the basic principle of the shift coefficient;

[0048] FIG9 is a schematic diagram of encoding of a shift coefficient mapped to a two-dimensional image;

[0049] FIG10 is a schematic diagram of encoding of inter-frame geometric position information;

[0050] FIG11A is a schematic diagram showing the composition of an intra-frame coding framework;

[0051] FIG11B is a schematic diagram showing the composition of an inter-frame coding framework;

[0052] FIG12A is a schematic diagram showing the composition of an intra-frame decoding framework;

[0053] FIG12B is a schematic diagram showing the composition of an inter-frame decoding framework;

[0054] FIG13A is a schematic diagram of iterative subdivision of a basic grid;

[0055] FIG13B is a schematic diagram of an LOD space structure;

[0056] FIG14 is a schematic diagram of coefficient reorganization for quantized coefficients;

[0057] FIG15 is a schematic diagram showing the basic principle of reconstructing and restoring geometric position information;

[0058] FIG16 is a schematic diagram of a mesh architecture of a codec provided in an embodiment of the present application;

[0059] FIG17 is a schematic diagram of a decoding method proposed in an embodiment of the present application;

[0060] FIG18 is a schematic diagram of determining geometric information of a first basic grid in the first mode;

[0061] FIG19 is a schematic diagram of determining geometric information of the first basic grid in the second mode;

[0062] FIG20 is a schematic diagram of an encoding method proposed in an embodiment of the present application;

[0063] FIG21 is a schematic diagram of the structure of the encoder;

[0064] Figure 22 is a schematic diagram of the second structure of the encoder;

[0065] FIG23 is a schematic diagram of the structure of the decoder;

[0066] FIG24 is a second schematic diagram of the decoder structure. DETAILED DESCRIPTION

[0067] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below with reference to the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present application.

[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0069] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0070] It should also be pointed out that the terms "first\second\third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0071] It should be noted that it is possible to decode and synthesize different data format bitstreams within the same video scene. These can include at least image format, point cloud format, and mesh format. In this way, real-time immersive video interaction services can be provided for multiple data formats (e.g., mesh, point cloud, image, etc.) from different sources.

[0072] In embodiments of the present application, the data format-based approach allows for independent processing at the bitstream level of the data format. This means that, similar to tiles or slices in video encoding, different data formats in this scenario can be encoded independently, enabling independent encoding and decoding based on the data format.

[0073] Generally speaking, 3D animation content uses a keyframe-based representation method, that is, each frame is a static mesh. Static meshes at different times have the same topological structure and different geometric structures. However, the amount of data of 3D dynamic meshes represented based on keyframes is extremely large, so how to effectively store, transmit and draw them has become a problem faced by the development of 3D dynamic meshes. In addition, the spatial scalability of the mesh needs to be supported for different user terminals (computers, notebooks, portable devices, mobile phones); different network bandwidths (broadband, narrowband, wireless) need to support the quality scalability of the mesh. Therefore, 3D dynamic mesh compression is a very critical issue.

[0074] A 3D mesh is the surface of a 3D object composed of countless polygons in space. Polygons are composed of vertices and edges. Figure 1A shows a 3D mesh image, and Figure 1B shows a partially enlarged schematic diagram of the 3D mesh image. Figures 1A and 1B show that the mesh surface is composed of closed polygons.

[0075] A two-dimensional image has information expressed at every pixel point and is distributed regularly, so there is no need to record its position information separately. However, the distribution of vertices in the mesh in three-dimensional space is random and irregular, and the way polygons are formed requires additional regulations. Therefore, it is necessary to record the position of each vertex in space and the connection information of each polygon to fully express a mesh image. As shown in Figure 2, the same number of vertices and vertex positions will form completely different surfaces due to different connection methods.

[0076] In addition to the above information, since 3D mesh images are usually encoded using existing 2D image / video encoding methods, the 3D mesh needs to be converted from 3D space to 2D images. The UV coordinates define this conversion process.

[0077] Similar to 2D images, each position in the image may have corresponding attribute information, typically RGB color values, which reflect the object's color. For 3D meshes, in addition to color, each vertex often has reflectance values, which reflect the surface material. 3D mesh attribute information is stored in 2D images, and the mapping from 2D to 3D is defined by UV coordinates.

[0078] Therefore, 3D mesh data typically includes 3D geometric position information (x, y, z), geometric connectivity, UV coordinates, and an attribute map. Figure 3A shows a 3D mesh image, Figure 3B shows the mesh data storage format, which includes 3D geometric position information, UV coordinates, and connectivity information, and Figure 3C shows the corresponding attribute diagram.

[0079] Current 3D dynamic mesh compression methods include space-time prediction methods, which improve compression efficiency by eliminating spatial and temporal correlations; principal component analysis (PCA)-based technology, which projects in the eigenvector space to concentrate energy; and wavelet-based methods, which support spatial scalability and quality scalability.

[0080] It should be noted that in the dynamic mesh coding provided by the Moving Picture Experts Group (MPEG), Figure 4 shows a schematic diagram of the overall mesh coding framework, Figure 5A shows a schematic diagram of the preprocessing of a 2D curve, and Figure 5B shows a schematic diagram of the generation of shift coefficients. The preprocessing process for a 3D mesh is similar, and on the encoding side, it is mainly divided into two parts: preprocessing and encoder. Preprocessing first generates a base mesh and shift coefficients. The preprocessing process includes: first, downsampling the original mesh to generate a simplified mesh (decimated mesh) with a significantly reduced number of vertices, or the base mesh. The base mesh is then subdivided and algorithmically generated, with newly generated vertices inserted along the edges of the base mesh to form a subdivided mesh. Finally, for each vertex in the subdivided mesh, the nearest vertex in the original mesh is found. The vector between the vertex in the subdivided mesh and the nearest vertex in the original mesh is the shift coefficient. Since the subdivision grid can be automatically generated at the codec end as long as the subdivision algorithm and the number of subdivision iterations are determined, after preprocessing, the original grid only needs to be represented as a simple basic grid and a series of shift coefficients. This can greatly reduce the amount of data without affecting the reconstruction at the decoding end.

[0081] Video Dynamic Mesh Coding (V-DMC) based on video coding can be broadly categorized into two main categories: geometric position information encoding and attribute information encoding. As shown in Figure 5, each frame in the basketball_player sequence includes two files: basketball_player_fr0001_qp12_qt12.obj and basketball_player_fr0002.png. basketball_player_fr0001_qp12_qt12.obj contains four types of information: geometric position information (x, y, z), the connectivity of geometric position triangles, texture coordinates (u, v), and the connectivity of texture coordinates. basketball_player_fr0002.png represents the texture attribute information of the current frame (e.g., the current image). In current V-DMC encoders, geometric position information is jointly encoded using the Dynamic Range Arithmetic Coding (DRACO) and the video codec, while texture information is encoded directly using the video codec. Among them, Video Codec can include H.264 / Advanced Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), H.266 / Versatile Video Coding (VVC / VV-enC), etc. Therefore, the following will introduce the mesh geometric information encoding in detail.

[0082] Geometric information can be divided into the encoding of position information (geometric position information and texture position information) and the encoding of connectivity relationships (geometric position information triangle patch connectivity, texture position information connectivity). Currently, V-DMC coding is mainly divided into two coding test conditions: intra-frame coding and inter-frame coding (low latency, currently no RA test environment).

[0083] (1) Intra-frame geometric information coding (Intra coding).

[0084] 1. Mesh preprocessing.

[0085] a) As shown in Figures 5A and 5B, using a two-dimensional connection relationship as an example, the original mesh's connection relationship contains a large number of points. Before encoding the mesh's geometric information, the mesh's geometric information is first quantized or simplified, ultimately resulting in a corresponding decimated mesh as the base mesh.

[0086] b) As shown in FIG6A and FIG6B , the quantization processing of the grid is performed based on the coordinates of the triangle patch. According to the connection relationship between the quantization points, the quantization processing can be divided into the following two cases:

[0087] When two vertices share a common edge, that is, two vertices that belong to the same edge before quantization, then after quantization, all the triangles connected by the two vertices need to be connected together, which involves the vanishing of the previous triangles as shown in Figure 6A;

[0088] Otherwise, if the two vertices do not share a common edge, that is, the two vertices do not belong to the same edge, then after quantization, it is only necessary to merge the boundaries of the two vertices, as shown in FIG6B , which does not affect the number of triangles.

[0089] c) In the whole process of mesh quantization based on triangle coordinates, the core problem is how to get the best vertex based on the previous vertex coordinates. The current V-DMC will get the best quantization point in the following four modes. Assuming that the vertex distribution before quantization is V1 and V2, and the vertex coordinate after quantization is V', there are the following: V1, V2, (V1+V2) / 2 and Q -1 (V1+V2), where Q is the quantization matrix corresponding to the vertex coordinates of V1 and V2. The distortion measure D before and after basic quantization is used to select the optimal quantization point.

[0090] 2.Base mesh encoding.

[0091] a) After obtaining the base mesh, the DRACO encoder is used to encode its geometric information. This geometric information primarily includes connectivity relationships and geometric position information. The DRACO encoding process is as follows: first, the connectivity relationships are encoded. Then, the geometric position information of the points is encoded based on the connectivity relationships. Finally, the texture position information is encoded based on the connectivity relationships and geometric position information.

[0092] b) Encoding of connection relationships. DRACO uses the "Edgebreaker Coding" scheme to encode the connection relationships of the mesh. See Figure 7A for details. In Figure 7A, v represents the current vertex. Before encoding the connection relationships of the mesh, the vertices of the mesh are divided into five types: C, L, R, S, and E. The physical meaning of each symbol is as follows:

[0093] C: None of the triangles connected to the current vertex have been encoded;

[0094] L: The triangle on the left connected to the current vertex completes the encoding;

[0095] R: The triangle on the right side connected to the current vertex completes the encoding;

[0096] S: The left and right triangles connected to the current vertex have not been encoded;

[0097] E: The left and right triangles connected to the current vertex have been encoded.

[0098] Finally, the type of each vertex and the processing order of the vertices are encoded in a certain order, and the decoding end restores the geometric connection relationship of the mesh according to the processing order and type of the vertices.

[0099] c) Coding of geometric position information. After completing the coding of the vertex connection relationship, the geometric position information of each vertex is predictively coded based on the vertex connection relationship. The idea adopted by predictive coding is the "Parallelograms algorithm", as shown in Figure 7B. A simple linear fit is performed using the three vertices adjacent to the current point to be coded: the left vertex, the right vertex, and the opposite vertex: pred pos =(keft+right)-opposite (1)

[0100] d) After completing the point connection relationship and geometric position information, the texture coordinates are predictively encoded based on the decoding and reconstruction of these two, as shown in Figure 7C. Similarly, assuming the current vertex is C, the left and right vertices of the current point can be obtained based on the point connection relationship. Then, the texture coordinates of the left and right vertices are used to predict the texture coordinates of the current vertex C.

[0101] 3. Displacement coefficient encoding.

[0102] a) First, after the base mesh is encoded and reconstructed, a partitioning algorithm is used to partition the base mesh to obtain the initial reconstructed mesh. Specifically, the curve corresponding to the subdivided mesh in Figure 5A is used to obtain the subdivided mesh (also called the "initial mesh") through simple linear interpolation. The coordinates of the newly inserted points are obtained by linear interpolation based on the two vertices on the current boundary:

[0103] b) Secondly, calculate the error Delta between the points in the subdivided mesh and the original mesh after division. The error Delta can be a point error in the world coordinate system. Finally, the displacement (i.e., displacement coefficient) of each point is calculated using the error Delta between each point and the normal vector Norm of each point. See Figure 8 for details. In Figure 8, the bold solid line represents the error Delta, and N and T represent the normal vector Norm. Thus, the specific calculation method is as follows: Displacement = Delta × Norm (3)

[0104] c) After calculating the displacement of each point, the spatial domain residual coefficient can be transformed into the frequency domain using the lifting transform to obtain the corresponding frequency domain residual coefficient.

[0105] d) Finally, a coefficient packing algorithm is used to map the frequency domain residual coefficients of each point into a two-dimensional image in a certain order. The current V-DMC can be arranged according to the Morton Code Order, as shown in Figure 9.

[0106] e) Finally, a traditional Video Codec is used to encode the two-dimensional image.

[0107] 4. Recoloring.

[0108] Recoloring is an algorithm on the encoder side. After the reconstruction of the encoder's geometric information is completed, the original geometric information, the original texture attribute information, and the reconstructed mesh geometric information are used to recolor the texture attribute information of the reconstructed mesh.

[0109] (2) Inter-frame geometric information coding (Inter coding).

[0110] a) Similar to the encoding above, the geometric position information includes the geometric connection relationship and the geometric position information encoding. However, it should be noted that the geometric position information encoding between frames only needs to encode the geometric position information (x, y, z) of the current base mesh, and does not need to encode the connection relationship and texture position information (u, v). The specific reason is as follows: If the current frame can use inter-frame encoding, then the base mesh of the reference frame of the current frame (such as the reference image) will be used at the encoding end to obtain the mesh information of the current frame. Therefore, the current frame and the reference frame have the same connection relationship and UV texture coordinates, only the geometric position information is different.

[0111] b) Based on a), it can be known that there is only an error in the geometric position information between the current frame and the reference frame, so the current V-DMC performs predictive coding on the geometric position information of the current frame.

[0112] As shown in Figure 10, the black dot is the point to be coded. The current point is used in the reference frame to get the corresponding predicted point (similar to the same-position block in video coding). Then, the neighborhood point of the current point (MV of the coded vertex) is used to predict the motion vector (MV) of the current point. The specific steps are as follows: assuming that the coordinates of the current point are pos and the coordinates of the corresponding same-position point are Pred pos , then the MV of the current point is calculated as: MV = Pos-Pred pos (4)

[0113] There are two predictive coding modes in the current V-DMC:

[0114] i. Directly encode the MV of the current point;

[0115] ii. Use the neighborhood to perform predictive coding on the MV of the current point.

[0116] At the encoding end, the rate-distortion optimization algorithm is used to obtain the optimal coding mode for each coding group (CG). The current V-DMC sets the number of points for each CG to be at most 16.

[0117] Coding of texture attribute information: Current V-DMC encodes texture attribute information directly using a video codec (Video-Codec), such as AVC, HEVC, VVC, or VV-enC.

[0118] Figure 11A is a schematic diagram of the framework of an intra-frame encoder. As shown in Figure 11A, in the intra-frame encoder, a common static mesh encoder (Static Mesh Encoder) can be used to encode the simplified mesh to generate the corresponding bitstream (Compressed base mesh bitstream). Next, the reconstructed simplified mesh is used to update the displacement coefficients (Update Displacements). The updated displacement coefficients are subjected to wavelet transform (Wavelet Transform) and quantization (Quantization) to obtain the displacement coefficients. After being packaged into images and videos (Image Packing, Video Packing), they are encoded using HEVC to generate a bitstream (Compressed displacements bitstream) of the displacement coefficients. For attribute map encoding, the feature map is first transformed (Texture Transfer) according to the difference between the reconstructed geometric information and the original geometric information, and then padded (Padding) and packaged (Video Packing) and encoded using a video encoder to form an attribute bitstream (Compressed attribute bitstream).

[0119] Figure 11B is a schematic diagram of an inter-frame encoder framework. As shown in Figure 11B, the inter-frame encoder and intra-frame encoder process are roughly the same, but the inter-frame encoder does not directly encode the simplified grid. Instead, it encodes the motion vector MV between the simplified grid of the current frame and the simplified grid of the reference frame and generates the corresponding motion vector bitstream (Compressed motion bitstream).

[0120] Correspondingly, during the decoding process, the decoder can also be divided into an intra-frame decoder and an inter-frame decoder according to the type of the frame it operates on, which are used to perform intra-frame decoding and inter-frame decoding respectively.

[0121] FIG12A is a schematic diagram of intra-frame decoding. As shown in FIG12A , in the intra-frame decoder, a static mesh decoder can be used to decode the simplified mesh. A video decoder is used to decode the shift coefficient video, and the shift coefficient is obtained through video unpacking and inverse wavelet transform. The decoded simplified mesh and shift coefficient are used to obtain the decoded mesh geometry information. The attribute map is decoded directly through the video decoder.

[0122] FIG12B is a schematic diagram of inter-frame decoding. As shown in FIG12B , for an inter-frame decoder, the process is basically the same as that of an intra-frame decoder, except that the simplified grid is not directly decoded, but the motion vector is decoded and the simplified grid of the current frame is calculated using the simplified grid of the previous frame (reference frame).

[0123] In summary, in the dynamic mesh coding (Dynamic Mesh Coding) currently provided by MPEG, the dynamic mesh coding process is divided into the following steps: at the encoding end, the basic mesh generated by preprocessing is quantized and then encoded using Google's open source DRACO encoder, and the shift coefficients are encoded using HEVC after wavelet transform, quantization, and two-dimensional mapping. The two-dimensional attribute map is also directly transmitted to the HEVC encoder for encoding; at the decoding end, the basic mesh code stream is decoded by DRACO to generate a decoded basic mesh, and the shift coefficients are decoded by HEVC decoding, inverse two-dimensional mapping, inverse quantization, and inverse transformation to generate decoded shift coefficients. Then, the decoded basic mesh and the decoded shift coefficients are used together to reconstruct the three-dimensional mesh geometry, and the attribute code stream is decoded by HEVC to generate a reconstructed attribute map.

[0124] (3) General test conditions for MPEG DMC.

[0125] a. There are 2 test conditions:

[0126] Condition 1: all intra geometry is lossy and attributes are lossy;

[0127] Condition 2: Random access is lossy in geometry and attributes;

[0128] b. The general test sequence may include five categories, namely Cat1-A, Cat1-B and Cat1-C, all of which contain geometric information and color attribute information.

[0129] The following is a detailed introduction to the Displacement coefficient encoding of V-DMC.

[0130] In a specific implementation, the encoder first iterates the mesh using a certain algorithm to obtain the corresponding mesh position information. The specific division algorithm is the same as the above, and linear interpolation is performed using the vertices on each boundary to obtain the corresponding geometric position information. Assuming that the entire division is iterated N times, then the values ​​obtained by different iterative divisions are

[0131] The Displacement coefficient is used for LOD division, as shown in FIG13A .

[0132] As shown in Figure 13A, the base mesh is linearly interpolated to obtain the corresponding mesh geometric position information. The initial geometric position information is used to calculate the error with the original mesh to obtain the displacement coefficient of each point. The LOD division can be divided into four layers: level 0, level 1, level 2, and level 3. The specific LOD spatial structure is shown in Figure 13B.

[0133] Secondly, a lifting wavelet transform is performed based on the LOD spatial structure, which can include two steps: prediction and update. The prediction algorithm is as follows:

[0134] Here, v represents the vertex to be predicted, and v1 and v2 represent the vertices at both ends of the boundary where the vertex to be predicted is located.

[0135] Again, the update steps are as follows:

[0136] Finally, the transformed coefficients are quantized and reorganized. As shown in Figure 14, this reorganization is performed on a block-by-block basis, where they are grouped into Block 0, Block 1, Block 2, and Block 3. In other words, the current V-DMC performs coefficient reorganization on a block-by-block basis, with each block being 16×16 in size. The coefficients within each block are arranged according to the Morton code to produce the corresponding 2D image. After completing this series of operations, the 2D image can be encoded using the Video Codec.

[0137] In summary, the dynamic grid encoding process can be divided into the following steps:

[0138] 1. Preprocess the original mesh by reducing the number of vertices in the mesh and simplifying the connection relationship.

[0139] 2. Subdivide the simplified mesh in step 1. For any two connected vertices in step 1, add a new point at the midpoint of the connecting line, and repeat this process twice.

[0140] 3. For each vertex in step 2, find the point in the original mesh that is closest to it and calculate the displacement coefficient of these two points.

[0141] 4. Use an encoder such as Draco to quantize the simplified grid in step 1 and then encode it.

[0142] 5. Adjust the shift coefficients in step 3 based on the reconstructed simplified grid obtained in step 4.

[0143] 6. Perform wavelet transform on the shift coefficients in step 5, and quantize the shift coefficients after wavelet transform to obtain quantized transform coefficients.

[0144] 7. Map the quantized transform coefficients from three-dimensional space to a two-dimensional image (or "image packing") to generate a two-dimensional image of shifted coefficients.

[0145] 8. Use a standard video encoder such as H.265 to encode the shift coefficient two-dimensional image in step 6.

[0146] In another specific implementation, for the decoding end, first, the Video-Codec is used to decode and reconstruct the two-dimensional image to restore the corresponding two-dimensional image. Secondly, according to the coefficient reorganization method, the lifting transform coefficient corresponding to each point can be restored. Finally, the inverse transform of the lifting wavelet transform can be used to restore the Displacement coefficient of each point. After obtaining the Displacement coefficient of each point, the geometric position information of the Base mesh and the Displacement coefficient are used to reconstruct and restore the geometric position information corresponding to the current mesh. As shown in Figure 15, the geometric position information of the level0 layer in the base mesh and the Displacement coefficient of the level0 layer can reconstruct and restore the geometric position information of the level0 layer in the reconstructed mesh, and the geometric position information of the level1 layer in the Base mesh and the Displacement coefficient of the level1 layer can reconstruct and restore.

[0147] The geometric position information of the level 1 layer in the Reconstruct mesh, and so on, the geometric position information of the level 3 layer in the Base mesh and the displacement coefficient of the level 3 layer can be reconstructed to restore the geometric position information of the level 3 layer in the Reconstruct mesh.

[0148] In summary, the dynamic grid decoding process can be divided into the following steps:

[0149] 1. The basic grid code stream is decoded by a decoder such as draco to generate a decoded basic grid.

[0150] 2. The shift coefficient bit stream is decoded using a standard video encoder such as H.265 to obtain a shift coefficient two-dimensional image.

[0151] 3. Map the shift coefficient two-dimensional image from the two-dimensional image to the three-dimensional space (or "image unpacking") to obtain the quantized transform coefficients.

[0152] 4. Dequantize and inverse wavelet transform the quantized transform coefficients to obtain the decoded shift coefficients.

[0153] 5. The decoded base grid and the decoded shift coefficients are combined to generate the reconstructed 3D grid geometric information.

[0154] 6. After the attribute code stream is decoded by HEVC, a reconstructed attribute graph is generated.

[0155] In related technologies, the existing V-DMC encoding takes into account the certain differences in the geometric position information of the corresponding Base mesh points between different meshes. Therefore, when the current frame can use inter-frame prediction encoding, the Base mesh between the two will be losslessly encoded, that is, MV lossless encoding.

[0156] However, due to the dynamic multi-frame sequence corresponding to the mesh, the geometric information of most meshes is highly correlated, which is often some simple rigid translation or rigid rotation. The current common geometric information compression schemes do not take into account the strong correlation between the geometric information of most meshes, thereby reducing the encoding efficiency of the geometric information and affecting the mesh compression performance.

[0157] To address the above-mentioned issues, an embodiment of the present application provides a coding and decoding method. At the decoding end, the decoder decodes the bitstream and determines first identification information corresponding to the current image. If the first identification information indicates that the first base grid corresponding to the current image uses the first mode, the geometric information of the first base grid is determined based on the geometric information of the second base grid corresponding to the reference image of the current image. The reconstructed original grid of the current image is determined based on the geometric information of the first base grid. At the encoding end, the encoder determines the first identification information corresponding to the current image. If the first identification information indicates that the first base grid corresponding to the current image uses the first mode, the geometric information of the first base grid is determined based on the geometric information of the second base grid corresponding to the reference image of the current image. The shift coefficient corresponding to the current image is determined based on the geometric information of the first base grid, and the shift coefficient is written into the bitstream. Thus, in an embodiment of the present application, if the first identification information corresponding to the current image indicates that the corresponding first base grid uses the first mode, the geometric information of the first base grid can be directly determined based on the geometric information of the second base grid corresponding to the reference image, that is, the base grid of the reference image is directly used to complete the prediction of the base grid of the current image, and the motion vector between the reference image and the current image can no longer be coded and decoded. That is to say, in the embodiment of the present application, the motion vector is skipped by using the first mode, thereby reducing the encoding code stream size of the geometric information of the grid, thereby improving the encoding efficiency of the geometric information and enhancing the grid compression performance.

[0158] The embodiment of the present application provides a network architecture of a codec system including a decoding method and an encoding method. FIG16 is a schematic diagram of a mesh architecture of a codec provided by the embodiment of the present application. As shown in FIG16 , the mesh architecture includes one or more electronic devices 13 to 1N and a communication grid 01, wherein the electronic devices 13 to 1N can perform video interaction through the communication grid 01. During implementation, the electronic device can be various types of devices with codec functions. For example, the electronic device can include a mobile phone, a tablet computer, a personal computer, a personal digital assistant, a navigator, a digital phone, a video phone, a television, a sensor device, a server, etc., which is not specifically limited in the embodiment of the present application. Here, the decoder or encoder described in the embodiment of the present application can be the above-mentioned electronic device.

[0159] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.

[0160] An embodiment of the present application proposes a decoding method. FIG17 is a schematic diagram of the decoding method proposed in the embodiment of the present application. As shown in FIG17 , in the embodiment of the present application, the method for performing decoding processing by the decoder may include the following steps:

[0161] Step 101: Decode a code stream to determine first identification information corresponding to a current image.

[0162] In an embodiment of the present application, the decoder may first decode the code stream to determine first identification information corresponding to the current image, wherein the first identification information may be used to determine whether the current image uses the first mode.

[0163] It is understandable that in the embodiment of the present application, at the encoding end, a base mesh and a displacement coefficient (displacement value) are first generated through preprocessing. The original mesh can be downsampled to generate a base mesh with a significantly reduced number of vertices. The base mesh is then subdivided and generated through an algorithm. Newly generated vertices are inserted on the edges of the base mesh. Finally, for each vertex in the subdivided mesh, the vertex closest to it is found in the original mesh. The vector between the vertex in the subdivided mesh and the nearest vertex in the original mesh is the displacement coefficient. As long as the subdivision algorithm and the number of subdivision iterations are determined, the subdivided mesh can be automatically generated at the codec end. Therefore, after preprocessing, the original mesh only needs to be represented as a simple base mesh and a series of displacement coefficients.

[0164] Accordingly, in an embodiment of the present application, for the intra-frame mode, the code stream is decoded to determine the simplified grid corresponding to the current image, that is, the corresponding first basic grid.

[0165] It should be noted that, in an embodiment of the present application, the first mode may represent a skip coding mode of a motion vector (MV) between a current image and a reference image, that is, when prediction processing is performed through the first mode, encoding and decoding processing of the MV may be skipped.

[0166] It can be understood that, in the embodiments of the present application, the current image can be understood as the current frame, and the reference image can be understood as the reference frame.

[0167] Accordingly, in the embodiments of the present application, the first mode may be understood as using the geometric information of the second basic grid corresponding to the reference image directly as the geometric information of the second basic grid corresponding to the current image.

[0168] It can be understood that the encoding and decoding method proposed in the embodiment of the present application can be applied to the encoding and decoding processing of geometric information.

[0169] Furthermore, in an embodiment of the present application, the encoding of V-DMC can be mainly divided into two categories: geometric position information encoding and attribute information encoding. As shown in Figure 5, each frame file of the sequence basketball_player includes two files: basketball_player_fr0001_qp12_qt12.obj and basketball_player_fr0002.png. Among them, basketball_player_fr0001_qp12_qt12.obj contains four types of information: geometric position information (x, y, z), the connection relationship of geometric position information triangles, texture coordinates (u, v) and the connection relationship of texture coordinates. basketball_player_fr0002.png represents the texture attribute information of the current image. In the current V-DMC encoder, the geometric position information is jointly encoded using DRACO and Video Codec (AVC, HEVC or VVC), and the texture information encoding is directly encoded using Video Codec.

[0170] It should be noted that in the embodiments of the present application, the encoding of geometric information can be divided into: encoding of position information (geometric position information and texture position information) and encoding of connection relationships (connection relationships of triangle patches of geometric position information and connection relationships of texture position information).

[0171] For intra-frame mode, after obtaining the base mesh, the DRACO encoder is used to encode the base mesh's geometric information. The entire DRACO encoding process mainly includes: first, completing the encoding of the connection relationship, then encoding the geometric position information of the points based on the geometric connection relationship, and finally encoding the texture position information based on the connection relationship and geometric position information.

[0172] For the inter-frame mode, after obtaining the base mesh, you can choose to use the base mesh of the reference image of the current image to obtain the mesh information of the current image. Since the current image and the reference image have the same connection relationship and uv texture coordinates, only the geometric position information is different. Therefore, the geometric position information encoding between frames only needs to encode the geometric position information (x, y, z) of the current base mesh, and does not need to encode the connection relationship and texture position information (u, v).

[0173] Furthermore, in the embodiments of the present application, for inter-frame mode, the only difference between the current image and the reference image is the geometric position information. Therefore, when performing predictive coding on the geometric position information of the current image, the common V-DMC can either use the current point in the reference image to obtain the corresponding prediction point (similar to the same-position block in video coding), and then determine the MV of the current point and perform predictive coding. Alternatively, the MV of the current point can be predictively coded using the neighboring points of the current point (the MV of the vertices that have already been coded).

[0174] For example, in some embodiments, assuming that the coordinates of the current point are pos, the coordinates of the corresponding co-point are Pred pos , then the MV calculation method of the current point is as shown in formula (4).

[0175] The common V-DMC mainly includes two MV prediction coding modes, one is to directly encode the MV of the current point, and the other is to predict the MV of the current point using the neighborhood.

[0176] It can be understood that in an embodiment of the present application, if the first identification information corresponding to the current image indicates that the first basic grid corresponding to the current image uses the first mode, then it can be considered that the encoding and decoding processing can be skipped for the MV between the current image and the reference image, that is, the decoding processing of the MV between the current image and the reference image is no longer performed.

[0177] It should be noted that, in the embodiment of the present application, after the first identification information is determined by decoding the code stream, it can be determined by the value of the first identification information whether the first basic grid corresponding to the current image uses the first mode.

[0178] Exemplarily, in some embodiments, if the value of the first identification information is the first value, then it can be determined that the first basic grid corresponding to the current image uses the first mode, that is, it can be considered that the decoding processing of the MV corresponding to the current image can be skipped; if the value of the first identification information is not the first value, then it can be determined that the first basic grid corresponding to the current image does not use the first mode.

[0179] Exemplarily, in some embodiments, if the value of the first identification information is the second value, it can be determined that the first basic grid corresponding to the current image uses the second mode, where the second mode can be an inter-frame prediction mode.

[0180] Exemplarily, in some embodiments, if the value of the first identification information is a third value, it can be determined that the first basic grid corresponding to the current image uses a third mode, where the third mode can be an intra-frame prediction mode.

[0181] It should be noted that, in the embodiment of the present application, the first value, the second value, and the third value may be in parameter form or in digital form. For example, the first value may be 1, the second value may be 2, and the third value may be 3.

[0182] Furthermore, in an embodiment of the present application, before determining the first identification information, the code stream may be decoded to determine the second identification information, wherein the second identification information may be used to determine whether the current sequence allows the use of the first mode, that is, whether the prediction mode corresponding to the current sequence includes the first mode. If the second identification information corresponding to the current sequence indicates that the current sequence allows the use of the first mode, then it can be considered that the prediction mode used when performing prediction processing on the current sequence frame may include the first mode, that is, for any image frame or video frame in the current sequence, the first mode of MV skip encoding and decoding processing may be used.

[0183] It should be noted that, in the embodiment of the present application, after the second identification information is determined by decoding the code stream, whether the current sequence allows the use of the first mode can be determined by the value of the second identification information.

[0184] Exemplarily, in some embodiments, if the value of the second identification information is the fourth value, it can be determined that the current sequence allows the use of the first mode; if the value of the second identification information is the fifth value, it can be determined that the current sequence does not allow the use of the first mode.

[0185] It should be noted that, in the embodiment of the present application, the fourth value and the fifth value may be in parameter form or in digital form. For example, the fourth value may be 1, and the fifth value may be 0.

[0186] It can be understood that in an embodiment of the present application, when the second identification information indicates that the current sequence allows the use of the first mode, the code stream can continue to be decoded to determine the first identification information corresponding to the current image, that is, the determination process of the first identification information corresponding to the current image can be continued.

[0187] That is, in an embodiment of the present application, the identification information for whether the first mode can be used can be determined by parsing layer by layer, that is, the encoding end can choose to set the identification information indicating whether the first mode can be used layer by layer from the upper layer to the lower layer. For example, first, the second identification information corresponding to the VFPS (Vmesh Frames Parameters Set) indicates whether the current sequence needs to skip the MV codec scheme; secondly, after the FPS (Frame Parameters Set) inherits the VFPS, it determines whether the current image uses the MV skipping codec scheme.

[0188] Furthermore, in an embodiment of the present application, after determining the first identification information corresponding to the current image, if the first basic grid of the first identification information uses the first mode, the code stream can be further decoded to determine the third identification information; wherein the third identification information is used to determine whether the coding group (Coding Group, CG) corresponding to the first basic grid uses the first mode.

[0189] It can be understood that if the third identification information indicates that a coding group corresponding to the first basic grid uses the first mode, then it can be considered that the decoding processing of the corresponding MV is skipped for the coding group; if the third identification information indicates that a coding group corresponding to the first basic grid does not use the first mode, then it can be considered that the corresponding MV is decoded for the coding group.

[0190] It should be noted that, in the embodiment of the present application, after the third identification information is determined by decoding the code stream, whether the coding group is allowed to use the first mode can be determined by the value of the third identification information.

[0191] Exemplarily, in some embodiments, if the value of the third identification information is the sixth value, then it can be determined that the corresponding coding group uses the first mode, that is, the decoding processing of the MV corresponding to the coding group is skipped; if the value of the third identification information is not the sixth value, then it can be determined that the corresponding coding group does not use the first mode.

[0192] Exemplarily, in some embodiments, if the value of the third identification information is the seventh value, it can be determined that the corresponding coding group uses the second mode, where the second mode can be an inter-frame prediction mode.

[0193] Exemplarily, in some embodiments, if the value of the third identification information is the eighth value, it can be determined that the corresponding coding group uses the third mode, where the third mode can be an intra-frame prediction mode.

[0194] It should be noted that, in the embodiment of the present application, the sixth value, the seventh value, and the eighth value may be in parameter form or in digital form. For example, the sixth value may be 1, the seventh value may be 2, and the eighth value may be 3.

[0195] That is to say, in an embodiment of the present application, the coding group CG can be used as a unit, and for each point in the CG, the corresponding third identification information can be used to indicate whether the MV corresponding to the point in the current CG can be skipped, that is, to indicate whether the coding group uses the first mode.

[0196] It should be noted that, in the embodiments of the present application, the decoder may be a video decoder, or any decoding device including a video decoder and a trellis decoder.

[0197] Furthermore, in an embodiment of the present application, the code stream transmitted to the decoder may be a code stream including the first identification information, or may be code stream data including the code stream including the first identification information and a code stream of a simplified grid (basic grid) (or a code stream of a motion vector).

[0198] It is understandable that, in the embodiments of the present application, the current image may be a current image frame or a current video frame, which is not specifically limited in the present application.

[0199] Step 102 : When the first identification information indicates that the first basic mesh corresponding to the current image uses the first mode, determine the geometric information of the first basic mesh according to the geometric information of the second basic mesh corresponding to the reference image of the current image.

[0200] In an embodiment of the present application, after decoding the code stream and determining the first identification information corresponding to the current image, when the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, the geometric information of the first basic grid can be further determined based on the geometric information of the second basic grid corresponding to the reference image of the current image.

[0201] It can be understood that in an embodiment of the present application, if the first identification information corresponding to the current image indicates that the first basic grid corresponding to the current image uses the first mode, then it can be considered that the encoding and decoding processing can be skipped for the MV between the current image and the reference image, that is, the decoding processing of the MV between the current image and the reference image is no longer performed.

[0202] It is understandable that, in the embodiment of the present application, the first basic grid may be divided so as to determine at least one coding group corresponding to the first basic grid, that is, at least one current coding group.

[0203] Correspondingly, in an embodiment of the present application, for the reference image of the current image, after determining the second basic grid corresponding to the reference image, the second basic grid can be divided to determine at least one coding group corresponding to the second basic grid, that is, at least one reference coding group.

[0204] It should be noted that, in the embodiment of the present application, the number of points in the current coding group obtained by dividing the current image may be the same as the number of points in the reference coding group obtained by dividing the corresponding reference image.

[0205] Exemplarily, in some embodiments, the number of points in each coding group CG can be set to N, where N is an integer greater than 0, for example, the value of N is 16.

[0206] It can be understood that in an embodiment of the application, if the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, then the geometric information of the second basic grid corresponding to the reference image can be directly selected as the geometric information of the second basic grid corresponding to the current image.

[0207] In related technologies, the base mesh information of the reference image (geometric information of the second base mesh) and the MV of each point in each CG can be used to reconstruct and restore the base mesh information of the current image (geometric information of the first base mesh). In the embodiment of the present application, when it is determined that the first identification information indicates that the first base mesh corresponding to the current image uses the first mode, the base mesh information of the reference image can be directly determined as the base mesh information of the current image (geometric information of the first base mesh).

[0208] Furthermore, in an embodiment of the present application, when determining the geometric information of the first basic grid based on the geometric information of the second basic grid corresponding to the reference image of the current image, the second basic grid can be divided first to determine at least one coding group corresponding to the second basic grid; then the geometric information of the at least one coding group corresponding to the second basic grid is determined as the geometric information of the at least one coding group corresponding to the first basic grid; and then the geometric information of the first basic grid can be determined based on the geometric information of the at least one coding group corresponding to the first basic grid.

[0209] That is to say, in an embodiment of the present application, after determining that the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, no processing may be performed on the second basic grid corresponding to the reference image. Instead, the second basic grid may be directly divided to determine at least one coding group corresponding to the second basic grid. Then, the geometric information of the at least one coding group corresponding to the second basic grid may be selected and determined in turn as the geometric information of the at least one coding group corresponding to the first basic grid corresponding to the current image, and the geometric information of the first basic grid corresponding to the current image may be reconstructed.

[0210] Exemplarily, in some embodiments, Figure 18 is a schematic diagram of determining the geometric information of the first basic grid in the first mode. As shown in Figure 18, the second basic grid corresponding to the reference image can be divided into four coding groups, CG0, CG1, CG2, and CG3, and then the geometric information corresponding to each coding group can be directly determined as the geometric information of the coding group corresponding to the first basic grid, that is, the first basic grid can also be divided into four coding groups, CG0, CG1, CG2, and CG3, wherein the geometric information of each coding group is respectively determined by the geometric information of the corresponding coding group after the second basic grid corresponding to the reference image is divided.

[0211] It can be seen that in the embodiment of the present application, based on the encoding and decoding method of the first mode, you can choose to skip the encoding and decoding of the first basic grid corresponding to the current image, that is, choose to skip the encoding and decoding of the MV between the current image and the reference image, but under the premise that the Base mesh information of the default current image (the geometric information of the first basic grid) is completely consistent with the Base mesh information in the reference image (the geometric information of the second basic grid), directly use the Base mesh information in the reference image (the geometric information of the second basic grid) as the Base mesh information of the current image (the geometric information of the first basic grid), so that on the basis of ensuring the reconstruction quality of the mesh geometric position information, the bit rate of the MV encoding can be reduced, thereby improving the encoding efficiency of the mesh geometric position information.

[0212] Furthermore, in the embodiments of the present application, since the corresponding third identification information can be used to indicate whether the MV corresponding to the midpoint in the coding group corresponding to the first basic grid can be skipped, that is, to indicate whether the coding group corresponding to the first basic grid uses the first mode, for the current coding group in the at least one coding group corresponding to the first basic grid, if the third identification information indicates that the current coding group uses the first mode, a reference coding group corresponding to the current coding group can be first determined in the at least one coding group corresponding to the second basic grid; then, the geometric information of the reference coding group can be determined as the geometric information of the current coding group; and finally, the geometric information of the first basic grid can be determined based on the geometric information of the current coding group.

[0213] That is, in an embodiment of the present application, after determining that the third identification information indicates that the current coding group in the first basic grid uses the first mode, a coding group corresponding to the current coding group can be determined as a reference coding group in at least one coding group corresponding to the second basic grid, and then the geometric information of the reference coding group can be selected as the geometric information of the current image coding group, and then the geometric information of the first basic grid corresponding to the current image can be reconstructed.

[0214] Exemplarily, in some embodiments, assuming that the first basic grid can be divided into four coding groups CG0, CG1, CG2, and CG3, and the second basic grid can also be divided into four coding groups CG0, CG1, CG2, and CG3, if the third identification information corresponding to the CG1 in the first basic grid indicates that the CG1 can use the first mode, that is, the encoding and decoding processing of the MV of the point in the CG1 can be skipped, then the geometric information of the corresponding CG1 in the second basic grid can be directly determined as the geometric information of the CG1 in the first basic grid.

[0215] Furthermore, in an embodiment of the present application, when determining the geometric information of the first basic grid based on the geometric information of the second basic grid corresponding to the reference image of the current image, the fitting parameters between the first basic grid and the second basic grid can be determined first; then the geometric information of the first basic grid can be determined based on the fitting parameters and the geometric information of the second basic grid.

[0216] That is to say, in an embodiment of the present application, the geometric information between the current image and the reference image is not exactly the same, but there are certain differences. Therefore, it is also possible not to directly determine the geometric information of the second basic grid corresponding to the reference image as the geometric information of the first basic grid corresponding to the current image, but first determine the fitting parameters between the first basic grid and the second basic grid, and then use the fitting parameters to fit the second basic grid, and finally determine the geometric information of the fitted second basic grid as the geometric information of the corresponding first basic grid.

[0217] It is understandable that, in the embodiment of the present application, the fitting parameters between the first basic grid and the second basic grid may be obtained by decoding the code stream, or may be pre-set and stored at the decoding end.

[0218] It can be seen that in the embodiments of the present application, considering that the base meshes of the two adjacent frames, the current image and the reference image, will have certain translation and rotation transformations in most cases, a rotation and translation transformation matrix can be introduced, that is, parameter fitting of the Base mesh of the two adjacent frames. In this way, the Base mesh information corresponding to the reference image after fitting is used as the Base mesh information of the current image, which can ensure the quality of the mesh geometric position information.

[0219] Furthermore, in an embodiment of the present application, when the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, the motion vector of the adjacent reconstructed point corresponding to the current point in the first basic grid can be determined first; then the motion vector of the adjacent reconstructed point is determined as the motion vector of the current point; finally, the geometric information of the first basic grid can be determined based on the geometric information of the second basic grid and the motion vector of the current point.

[0220] That is to say, in an embodiment of the present application, after determining that the first basic grid corresponding to the current image uses the first mode through the first identification information, for the skipping processing of the corresponding MV, in addition to directly determining the geometric information of the second basic grid corresponding to the reference image as the geometric information of the first basic grid, it is also possible to choose to use the MV of the neighborhood point of the current point in the first basic grid as the MV of the current point to perform subsequent reconstruction processing of the geometric information.

[0221] Furthermore, in an embodiment of the present application, when the third identification information indicates that the current coding group uses the first mode, the motion vector of the adjacent reconstructed point corresponding to the current point in the current coding group can be determined first; then the motion vector of the adjacent reconstructed point is determined as the motion vector of the current point; finally, the geometric information of the first basic grid can be determined based on the geometric information of the second basic grid and the motion vector of the current point.

[0222] That is to say, in an embodiment of the present application, after determining that the current coding group uses the first mode through the third identification information, for the skipping processing of the corresponding MV, in addition to directly determining the geometric information of the reference coding group in the second basic grid corresponding to the reference image as the geometric information of the current coding group in the first basic grid, it is also possible to choose to use the MV of the neighborhood point of the current point in the current coding group as the MV of the current point to perform subsequent reconstruction processing of the geometric information.

[0223] It can be seen that in the embodiment of the present application, the skip coding and decoding method of the MV based on the first mode can be used as a lossy coding and decoding mode, that is, the predicted MV of the neighborhood is directly used as the MV information of the current point, thereby reducing the code stream size of the MV information.

[0224] Furthermore, in an embodiment of the present application, when determining the geometric information of the first basic grid based on the geometric information of the second basic grid and the motion vector of the current point, the geometric information of the reference point corresponding to the current point can be first determined based on the geometric information of the second basic grid; then the geometric information of the current point can be determined based on the geometric information of the reference point and the motion vector of the current point; finally, the geometric information of the first basic grid can be determined based on the geometric information of the current point.

[0225] Step 103: Determine a reconstructed original grid of the current image according to the geometric information of the first basic grid.

[0226] In an embodiment of the present application, after the geometric information of the first basic grid is determined, the reconstructed original grid of the current image may be further determined according to the geometric information of the first basic grid.

[0227] Furthermore, in an embodiment of the present application, the code stream may be decoded to determine the shift coefficient corresponding to the current image.

[0228] Correspondingly, in an embodiment of the present application, when determining the reconstructed original grid of the current image based on the geometric information of the first basic grid, the corresponding shift vector can be first determined based on the shift coefficient, and then the reconstructed original grid of the current image can be determined based on the shift vector and the geometric information of the first basic grid.

[0229] It should be noted that in an embodiment of the present application, after determining the shift coefficient corresponding to the current image, the corresponding shift vector can be further determined based on the shift coefficient; then, the geometric information can be reconstructed based on the shift vector and the geometric information of the corresponding first basic grid, thereby determining the reconstructed original grid of the current image.

[0230] It should be noted that, in the embodiment of the present application, since the shift coefficients are generated by wavelet transforming the shift vectors, it is possible to select to perform inverse wavelet transform on the shift coefficients respectively, so as to determine the corresponding shift vectors.

[0231] It should be noted that in an embodiment of the present application, when reconstructing geometric information based on the shift vector and the geometric information of the first basic grid of the current image, the first basic grid can be subdivided first to determine the subdivided grid of the current image; then the corresponding reconstructed original grid can be determined according to the shift vector and the subdivided grid.

[0232] It can be understood that in the embodiment of the present application, the shift vector is obtained through the original grid and subdivided grid of the current image. Therefore, after determining the shift vector and the subdivided grid, the geometric information can be further reconstructed based on the shift vector and the subdivided grid to obtain the reconstructed original grid corresponding to the current image.

[0233] Furthermore, in an embodiment of the present application, after decoding the bitstream and determining the first identification information corresponding to the current image, that is, after step 101, and before determining the reconstructed original mesh of the current image based on the geometric information of the first basic mesh, that is, before step 103, the method for decoding by the decoder may further include the following steps:

[0234] Step 104 : When the first identification information indicates that the first basic grid uses the second mode, determine motion vector information corresponding to the first basic grid.

[0235] In an embodiment of the present application, after decoding the code stream and determining the first identification information corresponding to the current image, when the first identification information indicates that the first basic grid corresponding to the current image uses the second mode, the information of the motion vector corresponding to the first basic grid can be determined first.

[0236] It should be noted that, in the embodiment of the present application, the second mode may be an inter-frame prediction mode.

[0237] It is understood that in the embodiments of the present application, based on the second mode, i.e., the inter-frame prediction mode, when reconstructing the geometric information of the first base grid, the motion vector information corresponding to the first base grid can be first determined. The motion vector information corresponding to the first base grid can include the motion vector of each point in the second base grid.

[0238] Furthermore, in an embodiment of the present application, when determining the motion vector information corresponding to the first basic grid, for the current coding group in at least one coding group corresponding to the first basic grid, the motion vector corresponding to the current point in the current coding group can be determined by decoding the code stream, and then the motion vector information corresponding to the first basic grid can be determined.

[0239] That is to say, in an embodiment of the present application, for the inter-frame mode, at the encoding end, you can choose to use the current point in the reference image to obtain the corresponding prediction point (similar to the same-position block in video encoding), and then determine the MV of the current point, and then write the MV of the current point into the code stream and transmit it to the decoding end. At the decoding end, the MV of the current point can be directly determined by decoding the code stream.

[0240] Furthermore, in an embodiment of the present application, when determining the motion vector information corresponding to the first basic grid, for the current coding group in at least one coding group corresponding to the first basic grid, the motion vector of the current point can be determined based on the motion vector of the adjacent reconstructed point corresponding to the current point in the current coding group, and then the motion vector information corresponding to the first basic grid can be determined.

[0241] That is to say, in an embodiment of the present application, for the inter-frame mode, at the decoding end, you can also choose to use the neighborhood points of the current point (MVs of vertices that have completed encoding) to predict the MV of the current point, and then determine the MV of the current point.

[0242] Step 105 : Determine the geometric information of the first basic grid based on the geometric information of the second basic grid and the motion vector information corresponding to the first basic grid.

[0243] In an embodiment of the present application, if the first identification information indicates that the first basic grid uses the second mode, after determining the motion vector information corresponding to the first basic grid, the geometric information of the first basic grid can be further determined based on the geometric information of the second basic grid and the motion vector information corresponding to the first basic grid.

[0244] Furthermore, in an embodiment of the present application, when determining the geometric information of the first basic grid based on the geometric information of the second basic grid and the motion vector information corresponding to the first basic grid, the reference point corresponding to the current point can be first determined in at least one coding group corresponding to the second basic grid and the reference coding group corresponding to the current coding group; then, the geometric information of the current point can be determined based on the geometric information of the reference point and the motion vector of the current point; finally, the geometric information of the first basic grid can be determined based on the geometric information of the current point.

[0245] That is to say, in an embodiment of the present application, after determining that the first identification information indicates that the first basic grid corresponding to the current image uses the second mode, the motion vector of the point in the current coding group in the first basic grid can be determined first, and then the geometric information of the current coding group of the first basic grid corresponding to the current image can be determined in combination with the geometric information of the reference coding group in the second basic grid, and then the geometric information of the first basic grid corresponding to the current image can be reconstructed.

[0246] Exemplarily, in some embodiments, Figure 19 is a schematic diagram of determining the geometric information of the first basic grid in the second mode. As shown in Figure 19, the second basic grid corresponding to the reference image can be divided into four coding groups CG0, CG1, CG2, and CG3. Correspondingly, the first basic grid of the current image can also be divided into four coding groups CG0, CG1, CG2, and CG3. Then, the geometric information of the first basic grid of the current image can be reconstructed and restored based on the motion vector MV of the point in the coding group corresponding to the first basic grid and the geometric information of the corresponding coding group after the second basic grid is divided.

[0247] Furthermore, in an embodiment of the present application, after decoding the bitstream and determining the first identification information corresponding to the current image, that is, after step 101, and before determining the reconstructed original mesh of the current image based on the geometric information of the first basic mesh, that is, before step 103, the method for decoding by the decoder may further include the following steps:

[0248] Step 106: When the first identification information indicates that the first basic grid uses the third mode, decode the code stream to determine geometric information of the first basic grid.

[0249] In an embodiment of the present application, after decoding the code stream and determining the first identification information corresponding to the current image, if the first identification information indicates that the first basic grid corresponding to the current image uses the third mode, the code stream can be decoded to determine the geometric information of the first basic grid.

[0250] It should be noted that, in the embodiment of the present application, the third mode may be an intra-frame prediction mode.

[0251] It can be understood that in an embodiment of the present application, based on the third mode, i.e., the intra-frame prediction mode, when reconstructing the geometric information of the first basic grid, the geometric information of the first basic grid corresponding to the current image can be directly determined by decoding the code stream.

[0252] That is to say, in an embodiment of the present application, for the intra-frame prediction mode, the first basic grid of the current image can be obtained through a grid decoder, wherein the grid decoder decodes the code stream, thereby directly or indirectly determining the first basic grid corresponding to the current image.

[0253] It is understandable that, in the embodiments of the present application, the coding and decoding method proposed in the present application can be applied to both intra-frame coding and decoding and inter-frame coding and decoding, which is not specifically limited in the present application.

[0254] It should be noted that, in the embodiments of the present application, the decoder for performing decoding processing may include a video decoder and a grid decoder.

[0255] Exemplarily, in some embodiments, the code stream may be decoded by a grid decoder to determine the geometric information of the first basic grid.

[0256] It should be noted that, in the embodiment of the present application, for intra-frame coding and decoding, the grid decoder can decode the code stream to obtain the first basic grid of the corresponding current image.

[0257] Accordingly, in an embodiment of the present application, for intra-frame encoding and decoding, the grid decoder can receive the code stream of the first basic grid transmitted by the encoding end, and determine the first basic grid of the current image by decoding the code stream of the first basic grid.

[0258] Exemplarily, in some embodiments, the bitstream may be decoded by a grid decoder to determine the motion vector information corresponding to the first basic grid.

[0259] It should be noted that, in the embodiment of the present application, for inter-frame coding and decoding, the grid decoder can decode the code stream to obtain the corresponding motion vector of the current image, and then can further determine the first basic grid of the current image based on the motion vector.

[0260] Accordingly, in an embodiment of the present application, for inter-frame coding and decoding, the grid decoder can receive the code stream of the motion vector transmitted by the encoding end, and determine the motion vector of the current image by decoding the code stream of the motion vector, and then use the motion vector of the current image and the second basic grid of the decoded previous frame (reference image) to further determine the first basic grid of the current image.

[0261] For example, in some embodiments, a bitstream may be decoded by a video decoder to determine a shift coefficient corresponding to the current image.

[0262] In summary, the decoding method proposed in steps 101 to 106 above provides a base mesh information (geometric information of the base mesh) encoding and decoding scheme, namely, a skip encoding and decoding scheme. Based on the proposed first skip encoding and decoding mode, the geometric information of the second base mesh of the reference image can be directly used without any processing of the second base mesh. The second base mesh is directly divided and processed to obtain the corresponding mesh reconstruction geometric information. Ultimately, this reconstructed geometric information is used to reconstruct and restore the geometric information of the first base mesh of the current image.

[0263] That is to say, the encoding and decoding method proposed in the embodiment of the present application can consider that the Base mesh information of the current image (the geometric information of the first basic mesh) is completely consistent with the Base mesh information in the reference image (the geometric information of the second basic mesh). Therefore, the Base mesh information in the reference image can be directly used as the Base mesh information of the current image, thereby reducing the bit rate of the MV encoding while ensuring the reconstruction quality of the mesh geometric position information, thereby improving the encoding efficiency of the mesh geometric position information.

[0264] It can be seen that the coding and decoding method proposed in the embodiment of the present application is a Base Mesh information coding scheme, that is, a skip coding scheme. First, in the pre-processing link of the encoding end, if the current image can be inter-frame coded, then the MSE distortion measure will be used in the pre-processing link to calculate the distortions corresponding to the three coding modes: intra-frame coding algorithm, inter-frame coding algorithm and skip coding algorithm of the Base mesh. By comparing the distortion of the reconstructed mesh and the original mesh geometric information of the three, when the error between the distortion corresponding to the inter-frame coding and the skip coding of the Base mesh is within a certain range compared to the best distortion of the first two, it is considered that the current image can skip coding the MV when encoding the Base mesh (first mode), otherwise the existing coding algorithm (second mode or third mode) is used to encode the Base mesh information. If the Base mesh information in the current image can skip coding the MV, and it can be guaranteed that the quality of the position information of the mesh after reconstruction is not much different from the original coding scheme, this can further improve the mesh's geometric information coding efficiency.

[0265] It should be noted that in the embodiment of the present application, based on the first mode, in addition to controlling the encoding of the MV components corresponding to the Base mesh of the entire frame sequence through the preprocessing link, it is also possible to adaptively decide whether to skip encoding in units of CG, thereby reducing the encoding code stream size of the mesh geometric information while ensuring the reconstruction quality of the current image mesh geometric information, thereby further improving the mesh geometric information encoding efficiency.

[0266] The embodiment of the present application proposes a decoding method, in which, at the decoding end, the decoder decodes the code stream and determines the first identification information corresponding to the current image; when the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, the geometric information of the first basic grid is determined according to the geometric information of the second basic grid corresponding to the reference image of the current image; and the reconstructed original grid of the current image is determined according to the geometric information of the first basic grid. It can be seen that in the embodiment of the present application, if the first identification information corresponding to the current image indicates that the corresponding first basic grid uses the first mode, then the geometric information of the first basic grid can be directly determined according to the geometric information of the second basic grid corresponding to the reference image, that is, the basic grid of the reference image is directly used to complete the prediction of the basic grid of the current image, and the motion vector between the reference image and the current image can no longer be encoded and decoded. That is to say, in the embodiment of the present application, the first mode is used to implement skipping of motion vectors, thereby reducing the size of the encoding code stream of the geometric information of the grid, thereby improving the encoding efficiency of the geometric information and improving the grid compression performance.

[0267] An embodiment of the present application proposes an encoding method. FIG20 is a schematic diagram of the encoding method proposed in the embodiment of the present application. As shown in FIG20 , in the embodiment of the present application, the encoding method performed by the encoder may include the following steps:

[0268] Step 201: Determine first identification information corresponding to the current image.

[0269] In an embodiment of the present application, the encoder may first determine first identification information corresponding to the current image, wherein the first identification information may be used to determine whether the current image uses the first mode.

[0270] It is understandable that in the embodiment of the present application, at the encoding end, a base mesh and a displacement coefficient (displacement value) are first generated through preprocessing. The original mesh can be downsampled to generate a base mesh with a significantly reduced number of vertices. The base mesh is then subdivided and generated through an algorithm. Newly generated vertices are inserted on the edges of the base mesh. Finally, for each vertex in the subdivided mesh, the vertex closest to it is found in the original mesh. The vector between the vertex in the subdivided mesh and the nearest vertex in the original mesh is the displacement coefficient. As long as the subdivision algorithm and the number of subdivision iterations are determined, the subdivided mesh can be automatically generated at the codec end. Therefore, after preprocessing, the original mesh only needs to be represented as a simple base mesh and a series of displacement coefficients.

[0271] Accordingly, in an embodiment of the present application, for the intra-frame mode, the simplified grid corresponding to the current image, that is, the corresponding first basic grid, can be directly determined based on the original grid corresponding to the current image.

[0272] It should be noted that, in an embodiment of the present application, the first mode may represent a skip coding mode of a motion vector (MV) between a current image and a reference image, that is, when prediction processing is performed through the first mode, encoding and decoding processing of the MV may be skipped.

[0273] It can be understood that, in the embodiments of the present application, the current image can be understood as the current frame, and the reference image can be understood as the reference frame.

[0274] Accordingly, in the embodiments of the present application, the first mode may be understood as using the geometric information of the second basic grid corresponding to the reference image directly as the geometric information of the second basic grid corresponding to the current image.

[0275] It can be understood that the encoding and decoding method proposed in the embodiment of the present application can be applied to the encoding and decoding processing of geometric information.

[0276] Furthermore, in an embodiment of the present application, the encoding of V-DMC can be mainly divided into two categories: geometric position information encoding and attribute information encoding. As shown in Figure 5, each frame file of the sequence basketball_player includes two files: basketball_player_fr0001_qp12_qt12.obj and basketball_player_fr0002.png. Among them, basketball_player_fr0001_qp12_qt12.obj contains four types of information: geometric position information (x, y, z), the connection relationship of geometric position information triangles, texture coordinates (u, v) and the connection relationship of texture coordinates. basketball_player_fr0002.png represents the texture attribute information of the current image. In the current V-DMC encoder, the geometric position information is jointly encoded using DRACO and Video Codec (AVC, HEVC or VVC), and the texture information encoding is directly encoded using Video Codec.

[0277] It should be noted that in the embodiments of the present application, the encoding of geometric information can be divided into: encoding of position information (geometric position information and texture position information) and encoding of connection relationships (connection relationships of triangle patches of geometric position information and connection relationships of texture position information).

[0278] For intra-frame mode, after obtaining the base mesh, the DRACO encoder is used to encode the base mesh's geometric information. The entire DRACO encoding process mainly includes: first, completing the encoding of the connection relationship, then encoding the geometric position information of the points based on the geometric connection relationship, and finally encoding the texture position information based on the connection relationship and geometric position information.

[0279] For the inter-frame mode, after obtaining the base mesh, you can choose to use the base mesh of the reference image of the current image to obtain the mesh information of the current image. Since the current image and the reference image have the same connection relationship and uv texture coordinates, only the geometric position information is different. Therefore, the geometric position information encoding between frames only needs to encode the geometric position information (x, y, z) of the current base mesh, and does not need to encode the connection relationship and texture position information (u, v).

[0280] Furthermore, in the embodiments of the present application, for inter-frame mode, the only difference between the current image and the reference image is the geometric position information. Therefore, when performing predictive coding on the geometric position information of the current image, the common V-DMC can either use the current point in the reference image to obtain the corresponding prediction point (similar to the same-position block in video coding), and then determine the MV of the current point and perform predictive coding. Alternatively, the MV of the current point can be predictively coded using the neighboring points of the current point (the MV of the vertices that have already been coded).

[0281] For example, in some embodiments, assuming that the coordinates of the current point are pos, the coordinates of the corresponding co-point are Pred pos , then the MV calculation method of the current point is as shown in formula (4).

[0282] The common V-DMC mainly includes two MV prediction coding modes, one is to directly encode the MV of the current point, and the other is to predict the MV of the current point using the neighborhood.

[0283] It can be understood that in an embodiment of the present application, if the first identification information corresponding to the current image indicates that the first basic grid corresponding to the current image uses the first mode, then it can be considered that the encoding and decoding processing can be skipped for the MV between the current image and the reference image, that is, the decoding processing of the MV between the current image and the reference image is no longer performed.

[0284] It should be noted that, in the embodiment of the present application, after determining whether to use the first mode, the value of the first identification information can be set to indicate whether the first basic grid corresponding to the current image uses the first mode.

[0285] Exemplarily, in some embodiments, if the value of the first identification information is the first value, then it can be determined that the first basic grid corresponding to the current image uses the first mode, that is, it can be considered that the decoding processing of the MV corresponding to the current image can be skipped; if the value of the first identification information is not the first value, then it can be determined that the first basic grid corresponding to the current image does not use the first mode.

[0286] Exemplarily, in some embodiments, if the value of the first identification information is the second value, it can be determined that the first basic grid corresponding to the current image uses the second mode, where the second mode can be an inter-frame prediction mode.

[0287] Exemplarily, in some embodiments, if the value of the first identification information is a third value, it can be determined that the first basic grid corresponding to the current image uses a third mode, where the third mode can be an intra-frame prediction mode.

[0288] It should be noted that, in the embodiment of the present application, the first value, the second value, and the third value may be in parameter form or in digital form. For example, the first value may be 1, the second value may be 2, and the third value may be 3.

[0289] Furthermore, in an embodiment of the present application, before determining the first identification information, the second identification information may be determined first, wherein the second identification information may be used to determine whether the current sequence allows the use of the first mode, that is, whether the prediction mode corresponding to the current sequence includes the first mode. If the second identification information corresponding to the current sequence indicates that the current sequence allows the use of the first mode, then it can be considered that the prediction mode used when performing prediction processing on the current sequence frame may include the first mode, that is, for any image frame or video frame in the current sequence, the first mode of MV skip encoding and decoding processing may be used.

[0290] It should be noted that, in the embodiment of the present application, after determining whether the first mode is allowed to be used, it is possible to determine whether the current sequence is allowed to use the first mode by setting the value of the second identification information.

[0291] Exemplarily, in some embodiments, if the value of the second identification information is the fourth value, it can be determined that the current sequence allows the use of the first mode; if the value of the second identification information is the fifth value, it can be determined that the current sequence does not allow the use of the first mode.

[0292] It should be noted that, in the embodiment of the present application, the fourth value and the fifth value may be in parameter form or in digital form. For example, the fourth value may be 1, and the fifth value may be 0.

[0293] It can be understood that in an embodiment of the present application, when the second identification information indicates that the current sequence allows the use of the first mode, the first identification information corresponding to the current image can be continued to be determined, that is, the determination process of the first identification information corresponding to the current image can be continued.

[0294] That is, in an embodiment of the present application, the identification information for whether the first mode can be used can be determined by parsing layer by layer, that is, the encoding end can choose to set the identification information indicating whether the first mode can be used layer by layer from the upper layer to the lower layer. For example, first, the second identification information corresponding to the VFPS (Vmesh Frames Parameters Set) indicates whether the current sequence needs to skip the MV codec scheme; secondly, after the FPS (Frame Parameters Set) inherits the VFPS, it determines whether the current image uses the MV skipping codec scheme.

[0295] Furthermore, in an embodiment of the present application, after determining the first identification information corresponding to the current image, if the first basic grid of the first identification information uses the first mode, the third identification information can be further determined; wherein the third identification information is used to determine whether the coding group (Coding Group, CG) corresponding to the first basic grid uses the first mode.

[0296] It can be understood that if the third identification information indicates that a coding group corresponding to the first basic grid uses the first mode, then it can be considered that the decoding processing of the corresponding MV is skipped for the coding group; if the third identification information indicates that a coding group corresponding to the first basic grid does not use the first mode, then it can be considered that the corresponding MV is decoded for the coding group.

[0297] It should be noted that, in the embodiment of the present application, after determining whether the coding group corresponding to the first basic grid uses the first mode, the value of the third identification information can be set to indicate whether the coding group is allowed to use the first mode.

[0298] Exemplarily, in some embodiments, if the value of the third identification information is the sixth value, then it can be determined that the corresponding coding group uses the first mode, that is, the decoding processing of the MV corresponding to the coding group is skipped; if the value of the third identification information is not the sixth value, then it can be determined that the corresponding coding group does not use the first mode.

[0299] Exemplarily, in some embodiments, if the value of the third identification information is the seventh value, it can be determined that the corresponding coding group uses the second mode, where the second mode can be an inter-frame prediction mode.

[0300] Exemplarily, in some embodiments, if the value of the third identification information is the eighth value, it can be determined that the corresponding coding group uses the third mode, where the third mode can be an intra-frame prediction mode.

[0301] It should be noted that, in the embodiment of the present application, the sixth value, the seventh value, and the eighth value may be in parameter form or in digital form. For example, the sixth value may be 1, the seventh value may be 2, and the eighth value may be 3.

[0302] That is to say, in an embodiment of the present application, the coding group CG can be used as a unit, and for each point in the CG, the corresponding third identification information can be used to indicate whether the MV corresponding to the point in the current CG can be skipped, that is, to indicate whether the coding group uses the first mode.

[0303] Furthermore, in an embodiment of the present application, when determining the first identification information corresponding to the current image, for the first basic grid corresponding to the current image, the first distortion corresponding to the first mode, the second distortion corresponding to the second mode, and the third distortion corresponding to the third mode can be calculated respectively according to the rate-distortion algorithm; and then the first identification information corresponding to the current image is determined based on the first distortion, the first distortion, and the third distortion.

[0304] It should be noted that in an embodiment of the present application, the MSE algorithm can be used in the preprocessing stage to calculate the distortion corresponding to the skip coding mode (first mode). When the distortion is within a certain range compared to the original optimal distortion, it means that the current image can use MV skip coding, and finally the first mode is selected.

[0305] For example, in some embodiments, in the pre-processing phase of the encoding end, the distortion D1 of intra-frame prediction coding (third mode), the distortion D2 of inter-frame prediction coding (second mode), and the distortion D3 of MV skip coding (first mode) are calculated respectively. Then, the errors between the three distortions D1, D2, and D3 can be compared. Among them, since inter-frame prediction coding will reduce the encoded bitstream, skip coding will further reduce the encoded bitstream size. Therefore, when the distortion D3 of skip coding is within a certain error range compared to the distortion D1 of the intra-frame, it is considered that the Base mesh of the current image can use skip coding, that is, the MV component is not encoded.

[0306] It can be understood that in the embodiments of the present application, the method for determining the second identification information and the third identification information can also refer to the above-mentioned first identification information, and use the MSE algorithm to calculate the distortion corresponding to the skip coding mode (first mode) in the preprocessing stage, and determine whether to use the first mode by comparing the distortion with the original optimal distortion.

[0307] Furthermore, in an embodiment of the present application, after the first identification information is determined, the first identification information may be written into a code stream and transmitted to a decoding end.

[0308] Furthermore, in an embodiment of the present application, the code stream transmitted to the decoder may be a code stream including the first identification information, or may be code stream data including the code stream including the first identification information and a code stream of a simplified grid (basic grid) (or a code stream of a motion vector).

[0309] It is understandable that, in the embodiments of the present application, the current image may be a current image frame or a current video frame, which is not specifically limited in the present application.

[0310] Step 202: When the first identification information indicates that the first basic mesh corresponding to the current image uses the first mode, determine the geometric information of the first basic mesh according to the geometric information of the second basic mesh corresponding to the reference image of the current image.

[0311] In an embodiment of the present application, after determining the first identification information corresponding to the current image, when the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, the geometric information of the first basic grid can be further determined based on the geometric information of the second basic grid corresponding to the reference image of the current image.

[0312] It can be understood that in an embodiment of the present application, if the first identification information corresponding to the current image indicates that the first basic grid corresponding to the current image uses the first mode, then it can be considered that the encoding and decoding processing can be skipped for the MV between the current image and the reference image, that is, the decoding processing of the MV between the current image and the reference image is no longer performed.

[0313] It is understandable that, in the embodiment of the present application, the first basic grid may be divided so as to determine at least one coding group corresponding to the first basic grid, that is, at least one current coding group.

[0314] Correspondingly, in an embodiment of the present application, for the reference image of the current image, after determining the second basic grid corresponding to the reference image, the second basic grid can be divided to determine at least one coding group corresponding to the second basic grid, that is, at least one reference coding group.

[0315] It should be noted that, in the embodiment of the present application, the number of points in the current coding group obtained by dividing the current image may be the same as the number of points in the reference coding group obtained by dividing the corresponding reference image.

[0316] Exemplarily, in some embodiments, the number of points in each coding group CG can be set to N, where N is an integer greater than 0, for example, the value of N is 16.

[0317] It can be understood that in an embodiment of the application, if the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, then the geometric information of the second basic grid corresponding to the reference image can be directly selected as the geometric information of the second basic grid corresponding to the current image.

[0318] In related technologies, the base mesh information of the reference image (geometric information of the second base mesh) and the MV of each point in each CG can be used to reconstruct and restore the base mesh information of the current image (geometric information of the first base mesh). In the embodiment of the present application, when it is determined that the first identification information indicates that the first base mesh corresponding to the current image uses the first mode, the base mesh information of the reference image can be directly determined as the base mesh information of the current image (geometric information of the first base mesh).

[0319] Furthermore, in an embodiment of the present application, when determining the geometric information of the first basic grid based on the geometric information of the second basic grid corresponding to the reference image of the current image, the second basic grid can be divided first to determine at least one coding group corresponding to the second basic grid; then the geometric information of the at least one coding group corresponding to the second basic grid is determined as the geometric information of the at least one coding group corresponding to the first basic grid; and then the geometric information of the first basic grid can be determined based on the geometric information of the at least one coding group corresponding to the first basic grid.

[0320] That is to say, in an embodiment of the present application, after determining that the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, no processing may be performed on the second basic grid corresponding to the reference image. Instead, the second basic grid may be directly divided to determine at least one coding group corresponding to the second basic grid. Then, the geometric information of the at least one coding group corresponding to the second basic grid may be selected and determined in turn as the geometric information of the at least one coding group corresponding to the first basic grid corresponding to the current image, and the geometric information of the first basic grid corresponding to the current image may be reconstructed.

[0321] Exemplarily, in some embodiments, as shown in FIG18 , the second basic grid corresponding to the reference image can be divided into four coding groups, CG0, CG1, CG2, and CG3, and then the geometric information corresponding to each coding group can be directly determined as the geometric information of the coding group corresponding to the first basic grid, that is, the first basic grid can also be divided into four coding groups, CG0, CG1, CG2, and CG3, wherein the geometric information of each coding group is respectively determined by the geometric information of the corresponding coding group after the second basic grid corresponding to the reference image is divided.

[0322] It can be seen that in the embodiment of the present application, based on the encoding and decoding method of the first mode, you can choose to skip the encoding and decoding of the first basic grid corresponding to the current image, that is, choose to skip the encoding and decoding of the MV between the current image and the reference image, but under the premise that the Base mesh information of the default current image (the geometric information of the first basic grid) is completely consistent with the Base mesh information in the reference image (the geometric information of the second basic grid), directly use the Base mesh information in the reference image (the geometric information of the second basic grid) as the Base mesh information of the current image (the geometric information of the first basic grid), so that on the basis of ensuring the reconstruction quality of the mesh geometric position information, the bit rate of the MV encoding can be reduced, thereby improving the encoding efficiency of the mesh geometric position information.

[0323] Furthermore, in the embodiments of the present application, since the corresponding third identification information can be used to indicate whether the MV corresponding to the midpoint in the coding group corresponding to the first basic grid can be skipped, that is, to indicate whether the coding group corresponding to the first basic grid uses the first mode, for the current coding group in the at least one coding group corresponding to the first basic grid, if the third identification information indicates that the current coding group uses the first mode, a reference coding group corresponding to the current coding group can be first determined in the at least one coding group corresponding to the second basic grid; then, the geometric information of the reference coding group can be determined as the geometric information of the current coding group; and finally, the geometric information of the first basic grid can be determined based on the geometric information of the current coding group.

[0324] That is, in an embodiment of the present application, after determining that the third identification information indicates that the current coding group in the first basic grid uses the first mode, a coding group corresponding to the current coding group can be determined as a reference coding group in at least one coding group corresponding to the second basic grid, and then the geometric information of the reference coding group can be selected as the geometric information of the current image coding group, and then the geometric information of the first basic grid corresponding to the current image can be reconstructed.

[0325] Exemplarily, in some embodiments, assuming that the first basic grid can be divided into four coding groups CG0, CG1, CG2, and CG3, and the second basic grid can also be divided into four coding groups CG0, CG1, CG2, and CG3, if the third identification information corresponding to the CG1 in the first basic grid indicates that the CG1 can use the first mode, that is, the encoding and decoding processing of the MV of the point in the CG1 can be skipped, then the geometric information of the corresponding CG1 in the second basic grid can be directly determined as the geometric information of the CG1 in the first basic grid.

[0326] Furthermore, in an embodiment of the present application, when determining the geometric information of the first basic grid based on the geometric information of the second basic grid corresponding to the reference image of the current image, the fitting parameters between the first basic grid and the second basic grid can be determined first; then the geometric information of the first basic grid can be determined based on the fitting parameters and the geometric information of the second basic grid.

[0327] That is to say, in an embodiment of the present application, the geometric information between the current image and the reference image is not exactly the same, but there are certain differences. Therefore, it is also possible not to directly determine the geometric information of the second basic grid corresponding to the reference image as the geometric information of the first basic grid corresponding to the current image, but first determine the fitting parameters between the first basic grid and the second basic grid, and then use the fitting parameters to fit the second basic grid, and finally determine the geometric information of the fitted second basic grid as the geometric information of the corresponding first basic grid.

[0328] It is understandable that, in the embodiment of the present application, the fitting parameters between the first basic grid and the second basic grid may be transmitted to the decoding end through a code stream, or may be pre-set and stored in the decoding end.

[0329] It can be seen that in the embodiments of the present application, considering that the base meshes of the two adjacent frames, the current image and the reference image, will have certain translation and rotation transformations in most cases, a rotation and translation transformation matrix can be introduced, that is, parameter fitting of the Base mesh of the two adjacent frames. In this way, the Base mesh information corresponding to the reference image after fitting is used as the Base mesh information of the current image, which can ensure the quality of the mesh geometric position information.

[0330] Furthermore, in an embodiment of the present application, when the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, the motion vector of the adjacent reconstructed point corresponding to the current point in the first basic grid can be determined first; then the motion vector of the adjacent reconstructed point is determined as the motion vector of the current point; finally, the geometric information of the first basic grid can be determined based on the geometric information of the second basic grid and the motion vector of the current point.

[0331] That is to say, in an embodiment of the present application, after determining that the first basic grid corresponding to the current image uses the first mode through the first identification information, for the skipping processing of the corresponding MV, in addition to directly determining the geometric information of the second basic grid corresponding to the reference image as the geometric information of the first basic grid, it is also possible to choose to use the MV of the neighborhood point of the current point in the first basic grid as the MV of the current point to perform subsequent reconstruction processing of the geometric information.

[0332] Furthermore, in an embodiment of the present application, when the third identification information indicates that the current coding group uses the first mode, the motion vector of the adjacent reconstructed point corresponding to the current point in the current coding group can be determined first; then the motion vector of the adjacent reconstructed point is determined as the motion vector of the current point; finally, the geometric information of the first basic grid can be determined based on the geometric information of the second basic grid and the motion vector of the current point.

[0333] That is to say, in an embodiment of the present application, after determining that the current coding group uses the first mode through the third identification information, for the skipping processing of the corresponding MV, in addition to directly determining the geometric information of the reference coding group in the second basic grid corresponding to the reference image as the geometric information of the current coding group in the first basic grid, it is also possible to choose to use the MV of the neighborhood point of the current point in the current coding group as the MV of the current point to perform subsequent reconstruction processing of the geometric information.

[0334] It can be seen that in the embodiment of the present application, the skip coding and decoding method of the MV based on the first mode can be used as a lossy coding and decoding mode, that is, the predicted MV of the neighborhood is directly used as the MV information of the current point, thereby reducing the code stream size of the MV information.

[0335] Furthermore, in an embodiment of the present application, when determining the geometric information of the first basic grid based on the geometric information of the second basic grid and the motion vector of the current point, the geometric information of the reference point corresponding to the current point can be first determined based on the geometric information of the second basic grid; then the geometric information of the current point can be determined based on the geometric information of the reference point and the motion vector of the current point; finally, the geometric information of the first basic grid can be determined based on the geometric information of the current point.

[0336] Step 203: Determine a shift coefficient corresponding to the current image according to the geometric information of the first basic grid, and write the shift coefficient into the bitstream.

[0337] In an embodiment of the present application, after the geometric information of the first basic grid is determined, a shift coefficient corresponding to the current image may be further determined based on the geometric information of the first basic grid, and the shift coefficient may be written into the bitstream.

[0338] Furthermore, in an embodiment of the present application, when determining the shift coefficient corresponding to the current image based on the geometric information of the first basic grid, the shift vector corresponding to the current image can be updated based on the geometric information of the first basic grid to determine the shift coefficient corresponding to the current image.

[0339] Furthermore, in an embodiment of the present application, when using a shift vector to determine a corresponding shift coefficient, the encoder may choose to perform wavelet transform processing on the shift vectors respectively, so as to determine the shift coefficient.

[0340] It should be noted that in an embodiment of the present application, the shift vector can be determined by the refined grid of the current image and the first original grid of the current image; wherein, the original grid is simplified to determine the corresponding first original grid, that is, the simplified grid, and the simplified grid is refined to determine the corresponding refined grid.

[0341] It is understood that in the embodiments of the present application, during the preprocessing process, the original mesh of the current image can be simplified to obtain a simplified mesh (decimated mesh), or also called a base mesh. The simplified mesh can then be subdivided to obtain a subdivided mesh. Finally, for each vertex in the subdivided mesh, the point in the original mesh closest to it is found, and the displacement vector of the two points is calculated.

[0342] Furthermore, in an embodiment of the present application, after determining the shift coefficient corresponding to the current image according to the geometric information of the first basic grid, the shift coefficient can be written into the code stream and transmitted to the decoding end.

[0343] Furthermore, in an embodiment of the present application, after determining the first identification information corresponding to the current image, that is, after step 201, and before determining the shift coefficient corresponding to the current image based on the geometric information of the first basic grid and writing the shift coefficient into the bitstream, that is, before step 203, the method for the encoder to perform encoding processing may further include the following steps:

[0344] Step 204: When the first identification information indicates that the first basic grid uses the second mode, determine motion vector information corresponding to the first basic grid.

[0345] In an embodiment of the present application, after determining the first identification information corresponding to the current image, when the first identification information indicates that the first basic grid corresponding to the current image uses the second mode, the information of the motion vector corresponding to the first basic grid can be determined first.

[0346] It should be noted that, in the embodiment of the present application, the second mode may be an inter-frame prediction mode.

[0347] It is understood that in the embodiments of the present application, based on the second mode, i.e., the inter-frame prediction mode, when reconstructing the geometric information of the first base grid, the motion vector information corresponding to the first base grid can be first determined. The motion vector information corresponding to the first base grid can include the motion vector of each point in the second base grid.

[0348] Furthermore, in an embodiment of the present application, when determining the motion vector information corresponding to the first basic grid, for the current coding group in at least one coding group corresponding to the first basic grid, a reference point corresponding to the current point is determined in the reference image and in the reference coding group corresponding to the current coding group; then, based on the motion vector of the reference point, the motion vector of the current point is determined, and the motion vector of the current point is written into the code stream.

[0349] That is to say, in an embodiment of the present application, for the inter-frame mode, at the encoding end, you can choose to use the current point in the reference image to obtain the corresponding prediction point (similar to the same-position block in video encoding), and then determine the MV of the current point, and then write the MV of the current point into the code stream and transmit it to the decoding end. At the decoding end, the MV of the current point can be directly determined by decoding the code stream.

[0350] Furthermore, in an embodiment of the present application, when determining the motion vector information corresponding to the first basic grid, for the current coding group in at least one coding group corresponding to the first basic grid, the motion vector of the current point can be determined based on the motion vector of the adjacent reconstructed point corresponding to the current point in the current coding group, and then the motion vector information corresponding to the first basic grid can be determined.

[0351] That is to say, in an embodiment of the present application, for the inter-frame mode, at the decoding end, you can also choose to use the neighborhood points of the current point (MVs of vertices that have completed encoding) to predict the MV of the current point, and then determine the MV of the current point.

[0352] Step 205: Determine the geometric information of the first basic grid based on the geometric information of the second basic grid and the motion vector information corresponding to the first basic grid.

[0353] In an embodiment of the present application, if the first identification information indicates that the first basic grid uses the second mode, after determining the motion vector information corresponding to the first basic grid, the geometric information of the first basic grid can be further determined based on the geometric information of the second basic grid and the motion vector information corresponding to the first basic grid.

[0354] Furthermore, in an embodiment of the present application, when determining the geometric information of the first basic grid based on the geometric information of the second basic grid and the motion vector information corresponding to the first basic grid, the reference point corresponding to the current point can be first determined in at least one coding group corresponding to the second basic grid and the reference coding group corresponding to the current coding group; then, the geometric information of the current point can be determined based on the geometric information of the reference point and the motion vector of the current point; finally, the geometric information of the first basic grid can be determined based on the geometric information of the current point.

[0355] That is to say, in an embodiment of the present application, after determining that the first identification information indicates that the first basic grid corresponding to the current image uses the second mode, the motion vector of the point in the current coding group in the first basic grid can be determined first, and then the geometric information of the current coding group of the first basic grid corresponding to the current image can be determined in combination with the geometric information of the reference coding group in the second basic grid, and then the geometric information of the first basic grid corresponding to the current image can be reconstructed.

[0356] Exemplarily, in some embodiments, as shown in FIG19 , the second basic grid corresponding to the reference image can be divided into four coding groups, CG0, CG1, CG2, and CG3. Correspondingly, the first basic grid of the current image can also be divided into four coding groups, CG0, CG1, CG2, and CG3. Then, the geometric information of the first basic grid of the current image can be reconstructed and restored based on the motion vectors of the points in the coding group corresponding to the first basic grid and the geometric information of the corresponding coding group after the second basic grid is divided.

[0357] Furthermore, in an embodiment of the present application, after determining the first identification information corresponding to the current image, that is, after step 201, and before determining the shift coefficient corresponding to the current image based on the geometric information of the first basic grid and writing the shift coefficient into the bitstream, that is, before step 203, the method for the encoder to perform encoding processing may further include the following steps:

[0358] Step 206: When the first identification information indicates that the first basic grid uses the third mode, decode the code stream to determine geometric information of the first basic grid.

[0359] In an embodiment of the present application, after determining the first identification information corresponding to the current image, when the first identification information indicates that the first basic grid corresponding to the current image uses the third mode, the geometric information of the first basic grid can be determined according to the original grid of the current image.

[0360] It should be noted that, in the embodiment of the present application, the third mode may be an intra-frame prediction mode.

[0361] It can be understood that in an embodiment of the present application, based on the third mode, i.e., the intra-frame prediction mode, when reconstructing the geometric information of the first basic grid, the geometric information of the first basic grid corresponding to the current image can be directly determined through the original grid.

[0362] It is understandable that, in the embodiments of the present application, the coding and decoding method proposed in the present application can be applied to both intra-frame coding and decoding and inter-frame coding and decoding, which is not specifically limited in the present application.

[0363] It should be noted that, in the embodiments of the present application, the encoder used to perform encoding processing may include a video encoder and a grid encoder.

[0364] Exemplarily, in some embodiments, the geometric information of the first basic grid may be written into the bitstream by a grid encoder.

[0365] It should be noted that, in the embodiment of the present application, for intra-frame encoding and decoding, the grid encoder can write the geometric information of the first basic grid into the code stream and transmit it to the decoding end.

[0366] Accordingly, in an embodiment of the present application, for intra-frame encoding and decoding, the grid decoder can receive the code stream of the first basic grid transmitted by the encoding end, and determine the first basic grid of the current image by decoding the code stream of the first basic grid.

[0367] Exemplarily, in some embodiments, the motion vector information corresponding to the first basic grid may be written into the bitstream through a grid encoder.

[0368] It should be noted that, in the embodiment of the present application, for inter-frame coding and decoding, the grid encoder may write the motion vector information corresponding to the first basic grid into the bitstream.

[0369] Accordingly, in an embodiment of the present application, for inter-frame coding and decoding, the grid decoder can receive the code stream of the motion vector transmitted by the encoding end, and determine the motion vector of the current image by decoding the code stream of the motion vector, and then use the motion vector of the current image and the second basic grid of the decoded previous frame (reference image) to further determine the first basic grid of the current image.

[0370] Exemplarily, in some embodiments, the shift coefficient corresponding to the current image may be written into the bitstream by a video encoder.

[0371] In summary, the decoding method proposed in steps 201 to 206 above provides a base mesh information (geometric information of the base mesh) encoding and decoding scheme, namely, a skip encoding and decoding scheme. Based on the proposed first skip encoding and decoding mode, the geometric information of the second base mesh of the reference image can be directly used without any processing of the second base mesh. The second base mesh is directly divided and processed to obtain the corresponding mesh reconstruction geometric information. Ultimately, this reconstructed geometric information is used to reconstruct and restore the geometric information of the first base mesh of the current image.

[0372] That is to say, the encoding and decoding method proposed in the embodiment of the present application can consider that the Base mesh information of the current image (the geometric information of the first basic mesh) is completely consistent with the Base mesh information in the reference image (the geometric information of the second basic mesh). Therefore, the Base mesh information in the reference image can be directly used as the Base mesh information of the current image, thereby reducing the bit rate of the MV encoding while ensuring the reconstruction quality of the mesh geometric position information, thereby improving the encoding efficiency of the mesh geometric position information.

[0373] It can be seen that the coding and decoding method proposed in the embodiment of the present application is a Base Mesh information coding scheme, that is, a skip coding scheme. First, in the pre-processing link of the encoding end, if the current image can be inter-frame coded, then the MSE distortion measure will be used in the pre-processing link to calculate the distortions corresponding to the three coding modes: intra-frame coding algorithm, inter-frame coding algorithm and skip coding algorithm of the Base mesh. By comparing the distortion of the reconstructed mesh and the original mesh geometric information of the three, when the error between the distortion corresponding to the inter-frame coding and the skip coding of the Base mesh is within a certain range compared to the best distortion of the first two, it is considered that the current image can skip coding the MV when encoding the Base mesh (first mode), otherwise the existing coding algorithm (second mode or third mode) is used to encode the Base mesh information. If the Base mesh information in the current image can skip coding the MV, and it can be guaranteed that the quality of the position information of the mesh after reconstruction is not much different from the original coding scheme, this can further improve the mesh's geometric information coding efficiency.

[0374] It should be noted that in the embodiment of the present application, based on the first mode, in addition to controlling the encoding of the MV components corresponding to the Base mesh of the entire frame sequence through the preprocessing link, it is also possible to adaptively decide whether to skip encoding in units of CG, thereby reducing the encoding code stream size of the mesh geometric information while ensuring the reconstruction quality of the current image mesh geometric information, thereby further improving the mesh geometric information encoding efficiency.

[0375] The embodiment of the present application proposes a coding method, in which, at the coding end, the encoder determines the first identification information corresponding to the current image; when the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, the geometric information of the first basic grid is determined according to the geometric information of the second basic grid corresponding to the reference image of the current image; the shift coefficient corresponding to the current image is determined according to the geometric information of the first basic grid, and the shift coefficient is written into the code stream. It can be seen that in the embodiment of the present application, if the first identification information corresponding to the current image indicates that the corresponding first basic grid uses the first mode, then the geometric information of the first basic grid can be directly determined according to the geometric information of the second basic grid corresponding to the reference image, that is, the basic grid of the reference image is directly used to complete the prediction of the basic grid of the current image, and the motion vector between the reference image and the current image can no longer be encoded and decoded. That is to say, in the embodiment of the present application, the first mode is used to implement skipping of motion vectors, thereby reducing the size of the coding code stream of the geometric information of the grid, thereby improving the coding efficiency of the geometric information and improving the grid compression performance.

[0376] Based on the above embodiment, another embodiment of the present application proposes a coding and decoding method, which is applied to an encoder and a decoder, wherein the encoder includes a video encoder, a trellis encoder, and a preprocessor. The decoder includes a video decoder and a trellis decoder.

[0377] It should be noted that, in the embodiments of the present application, the coding and decoding method can be used for intra-frame coding and decoding, and can also be used for inter-frame coding and decoding, which is not specifically limited in the present application.

[0378] It should be noted that, in the embodiment of the present application, the preprocessor can be used to generate a base grid and a shift vector according to the original grid of the current image.

[0379] It is understood that in the embodiments of the present application, during the preprocessing process, the original mesh of the current image can be simplified to obtain a simplified mesh (decimated mesh), or a base mesh. The base mesh can then be subdivided to obtain a subdivided mesh. Finally, for each vertex in the subdivided mesh, the point in the original mesh closest to it is found, and the displacement vector of the two points is calculated.

[0380] Furthermore, in an embodiment of the present application, after the preprocessor generates a corresponding basic grid based on the original grid, the grid encoder can be used to encode the basic grid and then generate a code stream of the basic grid.

[0381] It should be noted that, in the embodiment of the present application, for intra-frame coding, the grid encoder can encode the basic grid of the current image to obtain a code stream of the basic grid.

[0382] Accordingly, in an embodiment of the present application, for intra-frame coding, the video encoder may first determine the reconstructed basic network of the current image using the bitstream of the generated basic grid; and then may update the shift vector by reconstructing the basic grid.

[0383] That is to say, in an embodiment of the present application, after the grid encoder completes encoding of the basic grid of the current image and generates a code stream of the basic grid, the video encoder can use the code stream of the basic grid to complete the reconstruction of the basic grid, and use the reconstructed basic grid to complete the update of the shift vector.

[0384] It should be noted that, in the embodiment of the present application, for inter-frame coding, the grid encoder can encode the motion vector of the current image to obtain a code stream of the motion vector.

[0385] Accordingly, in an embodiment of the present application, for inter-frame coding, the video encoder can first use the code stream of the generated motion vector to determine the motion vector of the current image, and then further determine the reconstructed basic network of the current image based on the motion vector; finally, the shift vector can be updated by reconstructing the basic grid.

[0386] That is to say, in an embodiment of the present application, after the grid encoder completes encoding of the motion vector of the current image and generates a code stream of the motion vector, the video encoder can use the code stream of the motion vector to complete the reconstruction of the basic grid, and use the reconstructed basic grid to complete the update of the shift vector.

[0387] It can be understood that, in the embodiment of the present application, after the grid encoder generates the code stream of the basic grid or the code stream of the motion vector, it can transmit the code stream of the basic grid or the code stream of the motion vector to the decoding end.

[0388] In common V-DMC coding, when inter-frame prediction coding is available for the current image, the base mesh corresponding to the current image and the base mesh of the reference image are used for prediction coding. Specifically, the base mesh in the current image is first divided into different coding groups (CGs), each of which can have a fixed number of points N (e.g., 16). Secondly, a rate-distortion optimization algorithm is used to select the optimal prediction coding mode for each CG: neighborhood prediction or direct encoding.

[0389] Because the geometric position information of the corresponding base mesh points between different meshes may vary, when the current image can be coded using inter-frame predictive coding, the base meshes between the two will be losslessly coded, i.e., MV lossless coding. However, due to the dynamic multi-frame sequence corresponding to the meshes, the geometric information of most meshes is highly correlated, often consisting of simple rigid translations or rotations. Therefore, a coding method can be considered to skip encoding the base mesh of the current image.

[0390] Common V-DMC coding, when determining whether the current image can use inter-frame prediction, first uses the Mean Square Error (MSE) algorithm in the pre-processing stage on the encoding end to adaptively select intra-frame prediction or inter-frame prediction. The base mesh of the intra-frame prediction algorithm is obtained by performing a simple quantization process on the original mesh geometric information. There is no reference information, and DRACO is currently used for encoding. Inter-frame prediction coding uses the base mesh of the reference image for processing. In the pre-processing stage, the base mesh (the connection relationship of the triangular facets) in the base reference image is corrected. In other words, only the geometric information of the points in the reference image needs to be modified. Therefore, it is necessary to encode the geometric position information error between the reference image and each point in the current image.

[0391] Furthermore, the encoding and decoding method proposed in the embodiment of the present application can establish a new mode in the preprocessing link of the encoding end, such as the first mode. The first mode directly uses the base mesh of the reference image, and does not perform any processing on the base mesh. The base mesh is directly divided and processed to obtain the corresponding mesh reconstruction geometric information, and finally the reconstructed geometric information is used to reconstruct and restore the geometric information of the current mesh.

[0392] It is understood that in the embodiment of the present application, based on the encoding and decoding scheme of the first mode, it can be assumed that the base mesh geometric information of the current image is completely consistent with the base mesh geometric information in the reference image, so the MV in the current image can be skipped. That is, the first mode can be a skip coding mode for the MV.

[0393] Correspondingly, in an embodiment of the present application, it is necessary to use the MSE algorithm in the preprocessing stage to calculate the distortion corresponding to the skip coding mode. When the distortion is within a certain range compared to the optimal distortion, it means that the current image can use MV skip coding, and finally the first mode is selected, and the base mesh of the reference image is directly used as the base mesh information of the current image.

[0394] In contrast, based on the common inter-frame prediction mode (second mode), the Base mesh information of the current image is reconstructed and restored by using the Base mesh information of the reference image and the MV of each point in each CG. The first mode proposed in the embodiment of the present application, i.e., skip encoding of the Base mesh, can assume that the Base mesh information of the current image is completely consistent with the Base mesh information in the reference image, and directly use the Base mesh information in the reference image as the Base mesh information of the current image. This can reduce the bit rate of the MV encoding while ensuring the quality of the reconstruction of the mesh geometric position information, thereby improving the encoding efficiency of the mesh geometric position information.

[0395] Furthermore, in an embodiment of the present application, for the skip coding mode of MV, i.e., the first mode, in the pre-processing link of the encoding end, the distortion D1 of the intra-frame prediction coding (third mode), the distortion D2 of the inter-frame prediction coding (second mode), and the distortion D3 of the MV skip coding (first mode) are calculated respectively. Then, the errors between the three distortions D1, D2, and D3 can be compared. Among them, since the inter-frame prediction coding will reduce the encoded bit stream, the skip coding will further reduce the size of the encoded bit stream. Therefore, when the distortion D3 of the skip coding is within a certain error range compared to the intra-frame distortion D1, it is considered that the Base mesh of the current image can adopt skip coding, that is, the MV component is not encoded.

[0396] Finally, through a certain amount of distortion, the MSE algorithm is used to obtain the error of the geometric position information of the mesh at different LOD layers, as shown in the following formula:

[0397] In other words, the encoding and decoding method proposed in the embodiments of this application establishes a new model in the preprocessing phase on the encoder side: the base mesh of the reference image is directly used without any processing. Instead, the base mesh is directly divided and processed to obtain the corresponding mesh reconstruction geometry information. Finally, this reconstructed geometry information is used to reconstruct and restore the geometry information of the current mesh. Based on this scheme, the base mesh geometry information of the current image is completely consistent with the base mesh geometry information of the reference image, so the MV in the current image can be skipped.

[0398] Furthermore, in an embodiment of the present application, for the skip coding mode of MV, that is, the first mode, at the decoding end, the header information corresponding to the current image is decoded, indicating that the current image is intra-frame coded (third mode), inter-frame coded (second mode) or MV skip coding (first mode). Then, the decoding mode of the current Base mesh can be determined according to the coding mode of the Base mesh of the current image. Among them, if the Base mesh of the current image needs to be decoded, the Base mesh information of the reference image and the MV information of each point obtained by parsing are used to reconstruct and restore the geometric position information of each point of the Base mesh of the current image; if the Base mesh of the current image skips coding, the Base mesh information in the reference image is directly used as the Base mesh information of the current image.

[0399] It can be understood that in the embodiments of the present application, the encoding mode of the Base mesh in the current image mesh can be determined by utilizing the distortion corresponding to the skip encoding of the MV corresponding to the Base mesh at the encoding end. Similarly, it is also possible to determine whether the MV corresponding to the point in the current CG can be skip encoded for each point in the CG based on the CG, thereby ensuring the overall mesh reconstruction quality, reducing the size of the base mesh encoding code stream, and improving the encoding efficiency of the mesh geometric information.

[0400] It is understood that in the embodiments of this application, common V-DMC encoding of MVs has two coding modes: neighborhood prediction coding and direct coding. These two modes are lossless coding, and can further skip MV coding, i.e., the first mode is a new lossy coding mode. In the first mode, the predicted MV can be directly used as the MV information of the current point, thereby reducing the bitstream size of the MV information, ensuring the quality of mesh reconstruction, and further improving the quality of the mesh's reconstructed geometric information.

[0401] It is understandable that in the embodiments of the present application, based on the first mode, the base mesh of the reference image is directly used as the base mesh information of the current image. This is achieved on the assumption that the base meshes of two adjacent frames are consistent. If there is translation and rotation transformation between two adjacent frames, then a rotation and translation transformation matrix can be introduced, that is, parameter fitting of the base meshes of the two adjacent frames can be introduced. In this way, the base mesh corresponding to the reference image after fitting is used as the base mesh information of the current image. This can ensure the quality of the mesh geometric position information and reduce the encoding bitstream size of the MV, thereby further improving the encoding efficiency of the mesh geometric information.

[0402] It can be seen that the coding and decoding method proposed in the embodiment of the present application is a Base Mesh information coding scheme, that is, a skip coding scheme. First, in the pre-processing link of the encoding end, if the current image can be inter-frame coded, then the MSE distortion measure will be used in the pre-processing link to calculate the distortions corresponding to the three coding modes: intra-frame coding algorithm, inter-frame coding algorithm and skip coding algorithm of the Base mesh. By comparing the distortion of the reconstructed mesh and the original mesh geometric information of the three, when the error between the distortion corresponding to the inter-frame coding and the skip coding of the Base mesh is within a certain range compared to the best distortion of the first two, it is considered that the current image can skip coding the MV when encoding the Base mesh (first mode), otherwise the existing coding algorithm (second mode or third mode) is used to encode the Base mesh information. If the Base mesh information in the current image can skip coding the MV, and it can be guaranteed that the quality of the position information of the mesh after reconstruction is not much different from the original coding scheme, this can further improve the mesh's geometric information coding efficiency.

[0403] It should be noted that in the embodiment of the present application, based on the first mode, in addition to controlling the encoding of the MV components corresponding to the Base mesh of the entire frame sequence through the preprocessing link, it is also possible to adaptively decide whether to skip encoding in units of CG, thereby reducing the encoding code stream size of the mesh geometric information while ensuring the reconstruction quality of the current image mesh geometric information, thereby further improving the mesh geometric information encoding efficiency.

[0404] The embodiment of the present application proposes a coding and decoding method. At the decoding end, the decoder decodes the code stream and determines the first identification information corresponding to the current image. When the first identification information indicates that the first base grid corresponding to the current image uses the first mode, the geometric information of the first base grid is determined based on the geometric information of the second base grid corresponding to the reference image of the current image. The reconstructed original grid of the current image is determined based on the geometric information of the first base grid. At the encoding end, the encoder determines the first identification information corresponding to the current image. When the first identification information indicates that the first base grid corresponding to the current image uses the first mode, the geometric information of the first base grid is determined based on the geometric information of the second base grid corresponding to the reference image of the current image. The shift coefficient corresponding to the current image is determined based on the geometric information of the first base grid, and the shift coefficient is written into the code stream. Therefore, in the embodiment of the present application, if the first identification information corresponding to the current image indicates that the corresponding first base grid uses the first mode, the geometric information of the first base grid can be directly determined based on the geometric information of the second base grid corresponding to the reference image, that is, the base grid of the reference image is directly used to complete the prediction of the base grid of the current image, and the motion vector between the reference image and the current image can no longer be coded and decoded. That is to say, in the embodiment of the present application, the motion vector is skipped by using the first mode, thereby reducing the encoding code stream size of the geometric information of the grid, thereby improving the encoding efficiency of the geometric information and enhancing the grid compression performance.

[0405] Based on the above embodiment, in another embodiment of the present application, based on the same inventive concept as the above embodiment, FIG21 is a schematic diagram of the composition structure of an encoder. As shown in FIG21 , the encoder 100 may include: a first determining unit 111, an encoding unit 112, wherein:

[0406] The first determining unit 111 is configured to determine first identification information corresponding to the current image; if the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, determine geometric information of the first basic grid based on geometric information of a second basic grid corresponding to a reference image of the current image; and determine a shift coefficient corresponding to the current image based on the geometric information of the first basic grid;

[0407] The encoding unit 112 is configured to write the shift coefficient into a bit stream.

[0408] It is understood that in this embodiment, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular system. Furthermore, the various components in this embodiment can be integrated into a single processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The aforementioned integrated units can be implemented in the form of hardware or software functional modules.

[0409] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0410] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the encoder 100. The computer-readable storage medium stores a computer program, and when the computer program is executed by the first processor, it implements the method described in any one of the aforementioned embodiments.

[0411] Based on the composition of the above-mentioned encoder 100 and the computer-readable storage medium, Figure 22 is a second schematic diagram of the composition structure of the encoder. As shown in Figure 22, the encoder 100 may include: a first memory 121 and a first processor 122, a first communication interface 123 and a first bus system 124. The first memory 121, the first processor 122, and the first communication interface 123 are coupled together through the first bus system 124. It can be understood that the first bus system 124 is used to realize the connection and communication between these components. In addition to the data bus, the first bus system 124 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as the first bus system 124 in Figure 10. Among them,

[0412] The first communication interface 123 is used to receive and send signals during the process of sending and receiving information between other external network elements;

[0413] The first memory 121 is used to store computer programs that can be run on the first processor;

[0414] The first processor 122 is configured to, when running the computer program, determine first identification information corresponding to a current image; when the first identification information indicates that a first basic grid corresponding to the current image uses a first mode, determine geometric information of the first basic grid based on geometric information of a second basic grid corresponding to a reference image of the current image; determine a shift coefficient corresponding to the current image based on the geometric information of the first basic grid, and write the shift coefficient into a bitstream.

[0415] It is understood that the first memory 121 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The first memory 121 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0416] The first processor 122 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the first processor 122. The above-mentioned first processor 122 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the first memory 121 , and the first processor 122 reads the information in the first memory 121 and completes the steps of the above method in combination with its hardware.

[0417] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0418] Optionally, as another embodiment, the first processor 122 is further configured to execute the method described in any one of the aforementioned embodiments when running the computer program.

[0419] FIG23 is a schematic diagram of the first structure of a decoder. As shown in FIG23 , the decoder 200 may include: a decoding unit 211 and a second determining unit 212; wherein,

[0420] The decoding unit 211 is configured to decode the code stream;

[0421] The second determination unit 212 is configured to determine the first identification information corresponding to the current image; when the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, determine the geometric information of the first basic grid according to the geometric information of the second basic grid corresponding to the reference image of the current image; and determine the reconstructed original grid of the current image according to the geometric information of the first basic grid.

[0422] It is understood that in this embodiment, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular system. Furthermore, the various components in this embodiment can be integrated into a single processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The aforementioned integrated units can be implemented in the form of hardware or software functional modules.

[0423] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0424] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the decoder 200. The computer-readable storage medium stores a computer program, and when the computer program is executed by the first processor, it implements the method described in any one of the aforementioned embodiments.

[0425] Based on the composition of the above-mentioned decoder 200 and the computer-readable storage medium, Figure 24 is a second schematic diagram of the composition structure of the decoder. As shown in Figure 24, the decoder 200 may include: a second memory 221 and a second processor 222, a second communication interface 223 and a second bus system 224. The second memory 221 and the second processor 222, and the second communication interface 223 are coupled together through the second bus system 224. It can be understood that the second bus system 224 is used to realize the connection and communication between these components. In addition to the data bus, the second bus system 224 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are marked as the second bus system 224 in Figure 12. Among them,

[0426] The second communication interface 223 is used to receive and send signals during the process of sending and receiving information with other external network elements;

[0427] The second memory 221 is used to store computer programs that can be run on the second processor;

[0428] The second processor 222 is used to decode the code stream and determine the first identification information corresponding to the current image when running the computer program; when the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, determine the geometric information of the first basic grid based on the geometric information of the second basic grid corresponding to the reference image of the current image; and determine the reconstructed original grid of the current image based on the geometric information of the first basic grid.

[0429] It is understood that the second memory 221 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The second memory 221 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0430] The second processor 222 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the second processor 222. The above-mentioned second processor 222 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the second memory 221 , and the second processor 222 reads the information in the second memory 221 and completes the steps of the above method in combination with its hardware.

[0431] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0432] Optionally, as another embodiment, the second processor 222 is further configured to execute the method described in any one of the aforementioned embodiments when running the computer program.

[0433] An embodiment of the present application provides a codec. At the decoding end, the decoder decodes a bitstream and determines first identification information corresponding to a current image. If the first identification information indicates that the first base grid corresponding to the current image uses a first mode, the decoder determines the geometric information of the first base grid based on the geometric information of the second base grid corresponding to the reference image of the current image. The reconstructed original grid of the current image is determined based on the geometric information of the first base grid. At the encoding end, the encoder determines the first identification information corresponding to the current image. If the first identification information indicates that the first base grid corresponding to the current image uses a first mode, the decoder determines the geometric information of the first base grid based on the geometric information of the second base grid corresponding to the reference image of the current image. The encoder determines the shift coefficient corresponding to the current image based on the geometric information of the first base grid and writes the shift coefficient into the bitstream. Thus, in an embodiment of the present application, if the first identification information corresponding to the current image indicates that the corresponding first base grid uses a first mode, the geometric information of the first base grid can be directly determined based on the geometric information of the second base grid corresponding to the reference image, that is, the base grid of the reference image is directly used to complete the prediction of the base grid of the current image, and the motion vector between the reference image and the current image can no longer be encoded and decoded. That is to say, in the embodiment of the present application, the motion vector is skipped by using the first mode, thereby reducing the encoding code stream size of the geometric information of the grid, thereby improving the encoding efficiency of the geometric information and enhancing the grid compression performance.

[0434] Furthermore, an embodiment of the present application provides a code stream, which is generated by bit-coding information to be coded; wherein the information to be coded includes at least one of the following:

[0435] First identification information, second identification information, third identification information, shift coefficient, motion vector information, fitting parameters, and geometric information of the first basic grid.

[0436] It should be noted that, in the embodiments of the present application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0437] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0438] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0439] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0440] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0441] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims. Industrial Applicability

[0442] The embodiments of the present application provide a coding and decoding method, an encoder, a decoder, a bitstream, and a storage medium. At the decoding end, the decoder decodes the bitstream and determines first identification information corresponding to the current image; if the first identification information indicates that the first base grid corresponding to the current image uses the first mode, the geometric information of the first base grid is determined based on the geometric information of the second base grid corresponding to the reference image of the current image; and the reconstructed original grid of the current image is determined based on the geometric information of the first base grid. At the encoding end, the encoder determines the first identification information corresponding to the current image; if the first identification information indicates that the first base grid corresponding to the current image uses the first mode, the geometric information of the first base grid is determined based on the geometric information of the second base grid corresponding to the reference image of the current image; and the shift coefficient corresponding to the current image is determined based on the geometric information of the first base grid, and the shift coefficient is written into the bitstream. Therefore, in the embodiments of the present application, if the first identification information corresponding to the current image indicates that the corresponding first base grid uses the first mode, the geometric information of the first base grid can be directly determined based on the geometric information of the second base grid corresponding to the reference image, that is, the base grid of the reference image is directly used to complete the prediction of the base grid of the current image, and the motion vector between the reference image and the current image can no longer be encoded and decoded. That is to say, in the embodiment of the present application, the motion vector is skipped by using the first mode, thereby reducing the encoding code stream size of the geometric information of the grid, thereby improving the encoding efficiency of the geometric information and enhancing the grid compression performance.

Claims

1. A decoding method, applied to a decoder, the method comprising: Decoding the code stream to determine first identification information corresponding to the current image; When the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, determining the geometric information of the first basic grid according to the geometric information of the second basic grid corresponding to the reference image of the current image; A reconstructed original grid of the current image is determined according to the geometric information of the first basic grid.

2. The method according to claim 1, wherein: The method further comprises: Decoding the code stream to determine the second identification information; When the second identification information indicates that the current sequence allows the use of the first mode, a process for determining the first identification information corresponding to the current image is executed.

3. The method according to claim 1 or 2, wherein: The method further comprises: When the first identification information indicates that the first basic grid uses the first mode, decode the code stream and determine third identification information; wherein the third identification information is used to determine whether the coding group corresponding to the first basic grid uses the first mode.

4. The method according to claim 3, wherein: The method further comprises: For a current coding group in the at least one coding group corresponding to the first basic grid, when the third identification information indicates that the current coding group uses the first mode, determining a reference coding group corresponding to the current coding group in the at least one coding group corresponding to the second basic grid; Determining the geometric information of the reference coding group as the geometric information of the current coding group; The geometric information of the first basic grid is determined according to the geometric information of the current encoding group.

5. The method according to any one of claims 1 to 3, wherein: The determining the geometric information of the first basic grid according to the geometric information of the second basic grid corresponding to the reference image of the current image includes: Dividing the second basic grid to determine at least one coding group corresponding to the second basic grid; Correspondingly determining the geometric information of at least one coding group corresponding to the second basic grid as the geometric information of at least one coding group corresponding to the first basic grid; The geometric information of the first basic grid is determined according to the geometric information of at least one encoding group corresponding to the first basic grid.

6. The method according to claim 1, wherein: The method further comprises: When the first identification information indicates that a first basic grid corresponding to the current image uses the first mode, determining a motion vector of an adjacent reconstruction point corresponding to a current point in the first basic grid; Determining the motion vector of the adjacent reconstructed point as the motion vector of the current point; Based on the geometric information of the second basic mesh and the motion vector of the current point, the geometric information of the first basic mesh is determined.

7. The method according to claim 3, wherein: The method further comprises: When the third identification information indicates that the current coding group uses the first mode, determining a motion vector of an adjacent reconstructed point corresponding to a current point in the current coding group; Determining the motion vector of the adjacent reconstructed point as the motion vector of the current point; Based on the geometric information of the second basic mesh and the motion vector of the current point, the geometric information of the first basic mesh is determined.

8. The method according to claim 6 or 7, wherein: The determining, based on the geometric information of the second basic grid and the motion vector of the current point, the geometric information of the first basic grid comprises: Determine geometric information of a reference point corresponding to the current point based on geometric information of the second basic grid; Determine the geometric information of the current point according to the geometric information of the reference point and the motion vector of the current point; The geometric information of the first basic grid is determined according to the geometric information of the current point.

9. The method according to claim 1, wherein: The determining the geometric information of the first basic grid according to the geometric information of the second basic grid corresponding to the reference image of the current image includes: determining a fitting parameter between the first base grid and the second base grid; The geometric information of the first basic mesh is determined according to the fitting parameters and the geometric information of the second basic mesh.

10. The method according to claim 2, wherein: After the decoded code stream determines the first identification information corresponding to the current image and before determining the reconstructed original grid of the current image according to the geometric information of the first basic grid, the method further includes: When the first identification information indicates that the first basic grid uses the second mode, determining motion vector information corresponding to the first basic grid; Based on the geometric information of the second basic grid and the motion vector information corresponding to the first basic grid, the geometric information of the first basic grid is determined.

11. The method according to claim 10, wherein: The determining the motion vector information corresponding to the first basic grid includes: For a current coding group in at least one coding group corresponding to the first basic grid, a code stream is decoded to determine a motion vector corresponding to a current point in the current coding group.

12. The method according to claim 10, wherein: The determining the motion vector information corresponding to the first basic grid includes: For a current coding group in at least one coding group corresponding to the first basic grid, a motion vector of the current point is determined according to motion vectors of adjacent reconstructed points corresponding to the current point in the current coding group.

13. The method according to claim 11 or 12, wherein: The determining, based on the geometric information of the second basic grid and the motion vector information corresponding to the first basic grid, the geometric information of the first basic grid includes: Determine a reference point corresponding to the current point in at least one coding group corresponding to the second basic grid and in a reference coding group corresponding to the current coding group; Determine the geometric information of the current point according to the geometric information of the reference point and the motion vector of the current point; The geometric information of the first basic grid is determined according to the geometric information of the current point.

14. The method according to claim 2, wherein: After the decoded code stream determines the first identification information corresponding to the current image and before determining the reconstructed original grid of the current image according to the geometric information of the first basic grid, the method further includes: When the first identification information indicates that the first basic grid uses the third mode, the code stream is decoded to determine the geometric information of the first basic grid.

15. The method according to any one of claims 1, 10, and 14, wherein: The method further comprises: The code stream is decoded to determine a shift coefficient corresponding to the current image.

16. The method according to claim 15, wherein: The step of determining the reconstructed original grid of the current image according to the geometric information of the first basic grid comprises: Determine a corresponding shift vector according to the shift coefficient; A reconstructed original mesh of the current image is determined based on the shift vector and the geometric information of the first basic mesh.

17. The method according to claim 14, wherein: The method further comprises: The bit stream is decoded by a grid decoder to determine the geometric information of the first basic grid.

18. The method according to claim 11, wherein: The method further comprises: The bit stream is decoded by a grid decoder to determine the motion vector information corresponding to the first basic grid.

19. The method according to claim 16, wherein: The method further comprises: The bit stream is decoded by a video decoder to determine the shift coefficient corresponding to the current image.

20. A decoding method, applied to a decoder, wherein: The decoder comprises a video decoder and a trellis decoder, The grid decoder and the video decoder are used to perform the decoding method as described in any one of claims 1-19.

21. A coding method, applied to an encoder, the method comprising: Determine first identification information corresponding to the current image; When the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, determining the geometric information of the first basic grid according to the geometric information of the second basic grid corresponding to the reference image of the current image; A shift coefficient corresponding to the current image is determined according to the geometric information of the first basic grid, and the shift coefficient is written into a bitstream.

22. The method according to claim 21, wherein: The method further comprises: Determining second identification information; When the second identification information indicates that the current sequence allows the use of the first mode, a process for determining the first identification information corresponding to the current image is executed.

23. The method according to claim 21 or 22, wherein: The method further comprises: When the first identification information indicates that the first basic grid uses the first mode, third identification information is determined; wherein the third identification information is used to determine whether the coding group corresponding to the first basic grid uses the first mode.

24. The method according to claim 23, wherein: The method further comprises: For a current coding group in the at least one coding group corresponding to the first basic grid, when the third identification information indicates that the current coding group uses the first mode, determining a reference coding group corresponding to the current coding group in the at least one coding group corresponding to the second basic grid; Determining the geometric information of the reference coding group as the geometric information of the current coding group; The geometric information of the first basic grid is determined according to the geometric information of the current encoding group.

25. The method according to any one of claims 21 to 23, wherein: The determining the geometric information of the first basic grid according to the geometric information of the second basic grid corresponding to the reference image of the current image includes: Dividing the second basic grid to determine at least one coding group corresponding to the second basic grid; Correspondingly determining the geometric information of at least one coding group corresponding to the second basic grid as the geometric information of at least one coding group corresponding to the first basic grid; The geometric information of the first basic grid is determined according to the geometric information of at least one encoding group corresponding to the first basic grid.

26. The method according to claim 21, wherein: The method further comprises: When the first identification information indicates that a first basic grid corresponding to the current image uses the first mode, determining a motion vector of an adjacent reconstruction point corresponding to a current point in the first basic grid; Determining the motion vector of the adjacent reconstructed point as the motion vector of the current point; Based on the geometric information of the second basic mesh and the motion vector of the current point, the geometric information of the first basic mesh is determined.

27. The method according to claim 23, wherein: The method further comprises: When the third identification information indicates that the current coding group uses the first mode, determining a motion vector of an adjacent reconstructed point corresponding to a current point in the current coding group; Determining the motion vector of the adjacent reconstructed point as the motion vector of the current point; Based on the geometric information of the second basic mesh and the motion vector of the current point, the geometric information of the first basic mesh is determined.

28. The method according to claim 26 or 27, wherein: The determining, based on the geometric information of the second basic grid and the motion vector of the current point, the geometric information of the first basic grid comprises: Determine geometric information of a reference point corresponding to the current point based on geometric information of the second basic grid; Determine the geometric information of the current point according to the geometric information of the reference point and the motion vector of the current point; The geometric information of the first basic grid is determined according to the geometric information of the current point.

29. The method according to claim 21, wherein: The determining the geometric information of the first basic grid according to the geometric information of the second basic grid corresponding to the reference image of the current image includes: determining a fitting parameter between the first base grid and the second base grid; The geometric information of the first basic mesh is determined according to the fitting parameters and the geometric information of the second basic mesh.

30. The method of claim 22, wherein: After determining the first identification information corresponding to the current image, and determining the shift coefficient corresponding to the current image according to the geometric information of the first basic grid, and before writing the shift coefficient into the bitstream, the method further includes: When the first identification information indicates that the first basic grid uses the second mode, determining motion vector information corresponding to the first basic grid; Based on the geometric information of the second basic grid and the motion vector information corresponding to the first basic grid, the geometric information of the first basic grid is determined.

31. The method according to claim 30, wherein: The determining the motion vector information corresponding to the first basic grid includes: For a current coding group in at least one coding group corresponding to the first basic grid, determining, in the reference image, a reference point corresponding to the current point in a reference coding group corresponding to the current coding group; The motion vector of the current point is determined according to the motion vector of the reference point, and the motion vector of the current point is written into the code stream.

32. The method of claim 30, wherein: The determining the motion vector information corresponding to the first basic grid includes: For a current coding group in at least one coding group corresponding to the first basic grid, a motion vector of the current point is determined according to motion vectors of adjacent reconstructed points corresponding to the current point in the current coding group.

33. The method according to claim 31 or 32, wherein: The determining, based on the geometric information of the second basic grid and the motion vector information corresponding to the first basic grid, the geometric information of the first basic grid includes: Determine a reference point corresponding to the current point in at least one coding group corresponding to the second basic grid and in a reference coding group corresponding to the current coding group; Determine the geometric information of the current point according to the geometric information of the reference point and the motion vector of the current point; The geometric information of the first basic grid is determined according to the geometric information of the current point.

34. The method of claim 22, wherein: After determining the first identification information corresponding to the current image, and determining the shift coefficient corresponding to the current image according to the geometric information of the first basic grid, and before writing the shift coefficient into the bitstream, the method further includes: When the first identification information indicates that the first basic grid uses the third mode, geometric information of the first basic grid is determined according to an original grid of the current image.

35. The method according to any one of claims 21, 30, and 34, wherein: The determining, according to the geometric information of the first basic grid, a shift coefficient corresponding to the current image includes: The shift vector corresponding to the current image is updated according to the geometric information of the first basic grid to determine the shift coefficient corresponding to the current image.

36. The method according to any one of claims 21 to 23, wherein: The determining the first identification information corresponding to the current image includes: For a first basic grid corresponding to the current image, respectively calculating a first distortion corresponding to the first mode, a second distortion corresponding to the second mode, and a third distortion corresponding to the third mode according to a rate-distortion algorithm; First identification information corresponding to the current image is determined according to the first distortion, the second distortion, and the third distortion.

37. The method of claim 36, wherein: The method further comprises: The first identification information is written into the code stream.

38. The method of claim 34, wherein: The method further comprises: The geometric information of the first basic grid is written into the bitstream through the grid encoder.

39. The method of claim 31, wherein: The method further comprises: The motion vector information corresponding to the first basic grid is written into the bit stream through the grid encoder.

40. The method according to any one of claims 1, 30, and 34, wherein: The method further comprises: The shift coefficient corresponding to the current image is written into a bit stream through a video encoder.

41. A coding method, applied to an encoder, wherein: The encoder includes a video encoder and a trellis encoder, The grid encoder and the video encoder are used to perform the encoding method as described in any one of claims 21-40.

42. A code stream, the code stream is generated by bit encoding according to information to be encoded; wherein, The information to be encoded includes at least one of the following: First identification information, second identification information, third identification information, shift coefficient, motion vector information, fitting parameters, and geometric information of the first basic grid.

43. An encoder, comprising: The first determining unit is a coding unit; wherein, The first determining unit is configured to determine first identification information corresponding to the current image; when the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, determine the geometric information of the first basic grid according to the geometric information of the second basic grid corresponding to the reference image of the current image; and determine the shift coefficient corresponding to the current image according to the geometric information of the first basic grid; The encoding unit is configured to write the shift coefficient into a bit stream.

44. An encoder, comprising: a first memory and a first processor; wherein, A first memory, for storing a computer program that can be run on the first processor; The first processor is configured to execute the method according to any one of claims 21 to 40 when running a computer program.

45. A decoder, the decoder comprising: A decoding unit, a second determining unit; wherein, The decoding unit is configured to decode the code stream; The second determination unit is configured to determine first identification information corresponding to the current image; when the first identification information indicates that the first basic grid corresponding to the current image uses the first mode, determine the geometric information of the first basic grid according to the geometric information of the second basic grid corresponding to the reference image of the current image; and determine the reconstructed original grid of the current image according to the geometric information of the first basic grid.

46. ​​A decoder, the decoder comprising: A second memory and a second processor; wherein, A second memory for storing a computer program that can be run on a second processor; The second processor is configured to execute the method according to any one of claims 1 to 20 when running a computer program.

47. A computer-readable storage medium storing a computer program, wherein the computer program, when executed, implements the method according to any one of claims 1 to 20, or implements the method according to any one of claims 21 to 40.