Coding method, decoding method, code stream, encoder, decoder, and storage medium

By determining the optimal encoding mode on each coordinate dimension of the vertices of the basic grid at the encoding end, the problem of low motion vector encoding and decoding efficiency in the prior art is solved, and the dynamic grid encoding and decoding performance is improved.

WO2025148072A1PCT designated stage expired Publication Date: 2025-07-17GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/072194
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-12
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

In the existing dynamic mesh encoding methods, the motion vector encoding and decoding mode of the basic mesh is not perfect enough, resulting in low encoding and decoding efficiency and reducing the dynamic mesh encoding and decoding performance.

Method used

At the encoding end, the motion vector corresponding to the vertices in the basic grid in each coordinate dimension is determined, and the optimal encoding mode is determined by skipping the encoding cost estimation of the encoding and prediction modes, and an independent flag bit indication decoder is used for decoding.

Benefits of technology

Improve the encoding and decoding efficiency of basic grids and improve the encoding and decoding performance of dynamic grids.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024072194_17072025_PF_FP_ABST
    Figure CN2024072194_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a coding method, a decoding method, a code stream, an encoder, a decoder, and a storage medium, which can improve coding and decoding efficiency of a basic mesh in dynamic mesh coding and decoding. The decoding method comprises: analyzing a code stream, and determining first syntactic information corresponding to a vertex in a basic mesh of a current image, the first syntactic information representing whether or not to skip decoding of a motion vector, and the first syntactic information acting on the vertex or acting on each coordinate dimension corresponding to the vertex; and, based on the first syntactic information, determining motion vector decoding information corresponding to the vertex.
Need to check novelty before this filing date? Find Prior Art

Description

Coding and decoding method, code stream, encoder, decoder and storage medium Technical Field

[0001] The embodiments of the present application relate to the field of dynamic grid coding and decoding technology, and in particular to a coding and decoding method, a bit stream, an encoder, a decoder, and a storage medium. Background Art

[0002] In the standard reference software for Dynamic Mesh Coding (DMC) provided by the Moving Picture Experts Group (MPEG), inter-frame encoding of the base mesh in the current image typically involves encoding the motion vectors of the vertices within the base mesh. On the decoder side, the base mesh corresponding to the current image is reconstructed by decoding the motion vectors of the vertices within the base mesh and combining them with the connectivity information from the base mesh of the reference image corresponding to the current image.

[0003] However, the current method for determining the encoding and decoding mode of motion vectors is not perfect enough, resulting in a high encoding and decoding cost of motion vectors, thereby reducing the encoding and decoding efficiency of the basic grid and further reducing the encoding and decoding performance of DMC.

[0004] Summary of the Invention

[0005] The embodiments of the present application provide a coding and decoding method, a code stream, an encoder, a decoder, and a storage medium, which can improve the coding and decoding efficiency of the basic grid, thereby improving the coding and decoding performance of the DMC.

[0006] The technical solution of the embodiment of the present application can be implemented as follows:

[0007] In a first aspect, an embodiment of the present application provides a decoding method, applied to a decoder, the method comprising:

[0008] A decoding method, applied to a decoder, comprising:

[0009] Parsing the bitstream to determine first syntax information corresponding to a vertex in a base mesh of the current image; the first syntax information indicating whether to skip decoding of a motion vector; the first syntax information being applied to the vertex or to at least one coordinate dimension corresponding to the vertex;

[0010] Determine the motion vector decoding information corresponding to the vertex based on the first syntax information.

[0011] In a second aspect, an embodiment of the present application provides an encoding method, applied to an encoder, the method comprising:

[0012] Determine a base grid of the current image and motion vectors corresponding to vertices in the base grid in each coordinate dimension of at least one coordinate dimension;

[0013] When performing inter-frame coding on the basic grid, skip coding and coding cost estimation of at least one prediction mode are performed on the motion vector corresponding to each coordinate dimension of the vertex to determine a first coding cost and at least one second coding cost corresponding to each coordinate dimension; the first coding cost corresponds to the skip coding mode; and the at least one second coding cost corresponds to the at least one prediction mode;

[0014] First syntax information is determined according to the first coding cost and at least one second coding cost corresponding to each coordinate dimension; the first syntax information indicates whether to skip encoding of the motion vector.

[0015] In a third aspect, an embodiment of the present application provides a code stream, wherein the code stream is generated by bit encoding based on information to be encoded; wherein the information to be encoded includes at least first syntax information; the first syntax information indicates whether to skip encoding of a motion vector; and the first syntax information is determined by the following method:

[0016] When performing inter-frame coding on a vertex in a base mesh, performing skip coding and coding cost estimation of at least one prediction mode on a motion vector corresponding to each coordinate dimension of the vertex, and determining a first coding cost and at least one second coding cost corresponding to each coordinate dimension; the first coding cost corresponds to the skip coding mode; and the at least one second coding cost corresponds to the at least one prediction mode;

[0017] The first grammatical information is determined according to the first coding cost and the at least one second coding cost.

[0018] In a fourth aspect, an embodiment of the present application provides a decoder, including:

[0019] a parsing portion configured to parse the bitstream and determine first syntax information corresponding to a vertex in a base mesh of a current image; the first syntax information indicating whether to skip decoding of a motion vector; the first syntax information being applied to the vertex or to each coordinate dimension corresponding to the vertex;

[0020] The decoding part is configured to determine the motion vector decoding information corresponding to the vertex based on the first syntax information.

[0021] In a fifth aspect, an embodiment of the present application provides a decoder, the decoder comprising a first memory and a first processor; wherein,

[0022] a first memory for storing a computer program capable of running on the first processor;

[0023] The first processor is configured to execute the decoding method as described in the first aspect when running a computer program.

[0024] In a sixth aspect, an embodiment of the present application provides an encoder, including:

[0025] A motion vector determining part, configured to determine a base grid of a current image and a motion vector corresponding to a vertex in the base grid in each coordinate dimension;

[0026] a coding cost determining unit configured to, when performing inter-frame coding on the basic grid, perform skip coding and coding cost estimation of at least one prediction mode on the motion vector corresponding to each coordinate dimension of the vertex, and determine a first coding cost and at least one second coding cost corresponding to each coordinate dimension; the first coding cost corresponds to the skip coding mode; and the at least one second coding cost corresponds to the at least one prediction mode;

[0027] The syntax information determination part is configured to determine first syntax information based on the first coding cost corresponding to each coordinate dimension and at least one second coding cost; the first syntax information represents whether to skip the encoding of the motion vector.

[0028] In a seventh aspect, an embodiment of the present application provides an encoder, comprising a second memory and a second processor; wherein,

[0029] a second memory for storing a computer program capable of running on the second processor;

[0030] The second processor is used to execute the encoding method as described in the second aspect when running the computer program.

[0031] In an eighth aspect, an embodiment of the present application provides a storage medium storing a computer program, which, when executed, implements the decoding method as described in the first aspect, or implements the encoding method as described in the second aspect.

[0032] The embodiments of the present application provide a coding and decoding method, a code stream, an encoder, a decoder, and a storage medium. At the encoding end, the basic grid of the current image and the motion vectors corresponding to the vertices in the basic grid in each coordinate dimension are determined; when the basic grid is inter-coded, the motion vectors corresponding to the vertices in each coordinate dimension are skipped and the coding cost of at least one prediction mode is estimated to determine the first coding cost and at least one second coding cost corresponding to each coordinate dimension; the first coding cost corresponds to the skip coding mode; the at least one second coding cost corresponds to at least one prediction mode; based on the first coding cost and at least one second coding cost corresponding to each coordinate dimension, first syntax information is determined; the first syntax information indicates whether to skip the coding of the motion vector. At the decoding end, the code stream is parsed to determine the first syntax information corresponding to the vertices in the basic grid of the current image; the first syntax information indicates whether to skip the decoding of the motion vector; the first syntax information acts on the vertex, or acts on each coordinate dimension corresponding to the vertex separately; based on the first syntax information, the decoding information of the motion vector corresponding to the vertex is determined. In this way, by comparing the coding costs of the motion vectors of the vertices in the basic grid in each coordinate dimension under skip coding and at least one prediction mode, the optimal processing method corresponding to the vertex in the coordinate dimension can be determined, such as skip coding or the optimal prediction mode, and the first syntax information is used to indicate the optimal processing method to the decoder, thereby improving the coding and decoding efficiency and thus improving the coding and decoding performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] FIG1A is a schematic diagram of a three-dimensional grid image 1;

[0034] FIG1B is a partially enlarged schematic diagram of a three-dimensional grid image;

[0035] Figure 2 is a schematic diagram of the connection method of the three-dimensional grid;

[0036] FIG3A is a second schematic diagram of a three-dimensional grid image;

[0037] FIG3B is a schematic diagram of a grid data storage format;

[0038] FIG3C is a schematic diagram of properties of a three-dimensional grid image;

[0039] FIG4 is a schematic diagram showing the composition of the overall framework of grid coding;

[0040] FIG5A is a schematic diagram of preprocessing of a two-dimensional curve;

[0041] FIG5B is a schematic diagram of generating a shift coefficient;

[0042] FIG6A is a first schematic diagram of quantization processing of grid geometric position information;

[0043] FIG6B is a second schematic diagram of quantization processing of grid geometric position information;

[0044] FIG7A is a schematic diagram of coding of the connection relationship of triangular facets;

[0045] FIG7B is a schematic diagram of encoding geometric position information;

[0046] FIG7C is a schematic diagram of texture coordinate encoding;

[0047] FIG8 is a schematic diagram showing the basic principle of the shift coefficient;

[0048] FIG9 is a schematic diagram of encoding of a shift coefficient mapped to a two-dimensional image;

[0049] FIG10 is a schematic diagram of encoding of inter-frame geometric position information;

[0050] FIG11A is a schematic diagram showing the composition of an intra-frame coding framework;

[0051] FIG11B is a schematic diagram showing the composition of an inter-frame coding framework;

[0052] FIG12A is a schematic diagram showing the composition of an intra-frame decoding framework;

[0053] FIG12B is a schematic diagram showing the composition of an inter-frame decoding framework;

[0054] FIG13 is a schematic diagram of a mesh architecture of a codec provided in an embodiment of the present application;

[0055] FIG14 is a schematic diagram of a flow chart of an encoding method provided in an embodiment of the present application;

[0056] FIG15 is a schematic diagram of a flowchart of a decoding method provided in an embodiment of the present application;

[0057] FIG16 is a schematic diagram of the structure of an encoder provided in an embodiment of the present application;

[0058] FIG17 is a schematic diagram of a specific hardware structure of an encoder provided in an embodiment of the present application;

[0059] FIG18 is a schematic diagram of the structure of a decoder provided in an embodiment of the present application;

[0060] FIG19 is a schematic diagram of a specific hardware structure of a decoder provided in an embodiment of the present application;

[0061] FIG20 is a schematic diagram of the composition structure of a coding and decoding system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0062] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below with reference to the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present application.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0064] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0065] It should also be pointed out that the terms "first\second\third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0066] It should be noted that it is possible to decode and synthesize different data format bitstreams within the same video scene. These can include at least image format, point cloud format, and mesh format. In this way, real-time immersive video interaction services can be provided for multiple data formats (e.g., mesh, point cloud, image, etc.) from different sources.

[0067] In embodiments of the present application, the data format-based approach allows for independent processing at the bitstream level of the data format. This means that, similar to tiles or slices in video encoding, different data formats in this scenario can be encoded independently, enabling independent encoding and decoding based on the data format.

[0068] Generally speaking, 3D animation content uses a keyframe-based representation method, that is, each frame is a static mesh. Static meshes at different times have the same topological structure and different geometric structures. However, the amount of data of 3D dynamic meshes represented based on keyframes is extremely large, so how to effectively store, transmit and draw them has become a problem faced by the development of 3D dynamic meshes. In addition, the spatial scalability of the mesh needs to be supported for different user terminals (computers, notebooks, portable devices, mobile phones); different mesh bandwidths (broadband, narrowband, wireless) need to support the quality scalability of the mesh. Therefore, 3D dynamic mesh compression is a very critical issue.

[0069] A 3D mesh is the surface of a 3D object composed of countless polygons in space. Polygons are composed of vertices and edges. Figure 1A shows a 3D mesh image, and Figure 1B shows a partially enlarged schematic diagram of the 3D mesh image. Figures 1A and 1B show that the mesh surface is composed of closed polygons.

[0070] A two-dimensional image has information expressed at every pixel point and is distributed regularly, so there is no need to record its position information separately. However, the distribution of vertices in the mesh in three-dimensional space is random and irregular, and the way polygons are formed requires additional regulations. Therefore, it is necessary to record the position of each vertex in space and the connection information of each polygon to fully express a mesh image. As shown in Figure 2, the same number of vertices and vertex positions will form completely different surfaces due to different connection methods.

[0071] In addition to the above information, since 3D mesh images are usually encoded using existing 2D image / video encoding methods, the 3D mesh needs to be converted from 3D space to 2D images. The UV coordinates define this conversion process.

[0072] Similar to 2D images, each position in the image may have corresponding attribute information, typically RGB color values, which reflect the object's color. For 3D meshes, in addition to color, each vertex often has reflectance values, which reflect the surface material. 3D mesh attribute information is stored in 2D images, and the mapping from 2D to 3D is defined by UV coordinates.

[0073] Therefore, 3D mesh data typically includes 3D geometric position information (x, y, z), geometric connectivity, UV coordinates, and an attribute map. Figure 3A shows a 3D mesh image, Figure 3B shows the mesh data storage format, which includes 3D geometric position information, UV coordinates, and connectivity information, and Figure 3C shows the corresponding attribute diagram.

[0074] Current 3D dynamic mesh compression methods include space-time prediction methods, which improve compression efficiency by eliminating spatial and temporal correlations; principal component analysis (PCA)-based technology, which projects in the eigenvector space to concentrate energy; and wavelet-based methods, which support spatial scalability and quality scalability.

[0075] It should be noted that in the dynamic mesh coding provided by the Moving Picture Experts Group (MPEG), Figure 4 shows a schematic diagram of the overall mesh coding framework, Figure 5A shows a schematic diagram of the preprocessing of a 2D curve, and Figure 5B shows a schematic diagram of the generation of shift coefficients. The preprocessing process for a 3D mesh is similar, and on the encoding side, it is mainly divided into two parts: preprocessing and encoder. Preprocessing first generates a base mesh and shift coefficients. The preprocessing process includes: first, downsampling the original mesh to generate a simplified mesh (decimated mesh) with a significantly reduced number of vertices, or the base mesh. The base mesh is then subdivided and algorithmically generated, with newly generated vertices inserted along the edges of the base mesh to form a subdivided mesh. Finally, for each vertex in the subdivided mesh, the nearest vertex in the original mesh is found. The vector between the vertex in the subdivided mesh and the nearest vertex in the original mesh is the shift coefficient. Since the subdivision grid can be automatically generated at the codec end as long as the subdivision algorithm and the number of subdivision iterations are determined, after preprocessing, the original grid only needs to be represented as a simple basic grid and a series of shift coefficients. This can greatly reduce the amount of data without affecting the reconstruction at the decoding end.

[0076] Video Dynamic Mesh Coding (V-DMC) based on video coding can be broadly categorized into two main categories: geometric position information encoding and attribute information encoding. For example, each frame in the basketball_player sequence includes two files: basketball_player_fr0001_qp12_qt12.obj and basketball_player_fr0002.png. basketball_player_fr0001_qp12_qt12.obj contains four types of information: geometric position information (x, y, z), the connectivity of geometric position triangles, texture coordinates (u, v), and the connectivity of texture coordinates. basketball_player_fr0002.png represents the texture attribute information of the current image. In current V-DMC encoders, geometric position information is jointly encoded using the Dynamic Range Arithmetic Coding (DRACO) and the video codec, while texture information encoding is performed directly using the video codec. Among them, Video Codec can include H.264 / Advanced Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), H.266 / Versatile Video Coding (VVC / VV-enC), etc. Therefore, the following will introduce the mesh geometric information encoding in detail.

[0077] Geometric information can be divided into the encoding of position information (geometric position information and texture position information) and the encoding of connectivity relationships (geometric position information triangle patch connectivity, texture position information connectivity). Currently, V-DMC coding is mainly divided into two coding test conditions: intra-frame coding and inter-frame coding (low latency, currently no RA test environment).

[0078] (1) Intra-frame geometric information coding (Intra coding).

[0079] 1. Mesh preprocessing.

[0080] a) As shown in Figures 5A and 5B, using a two-dimensional connection relationship as an example, the original mesh's connection relationship contains a large number of points. Before encoding the mesh's geometric information, the mesh's geometric information is first quantized or simplified, ultimately resulting in a corresponding decimated mesh as the base mesh.

[0081] b) As shown in FIG6A and FIG6B , the quantization processing of the grid is performed based on the coordinates of the triangle patch. According to the connection relationship between the quantization points, the quantization processing can be divided into the following two cases:

[0082] When two vertices share a common edge, that is, two vertices that belong to the same edge before quantization, then after quantization, all the triangles connected by the two vertices need to be connected together, which involves the vanishing of the previous triangles as shown in Figure 6A;

[0083] Otherwise, if the two vertices do not share a common edge, that is, the two vertices do not belong to the same edge, then after quantization, it is only necessary to merge the boundaries of the two vertices, as shown in FIG6B , which does not affect the number of triangles.

[0084] c) In the whole process of mesh quantization based on triangle coordinates, the core problem is how to get the best vertex based on the previous vertex coordinates. The current V-DMC will get the best quantization point in the following four modes. Assuming that the vertex distribution before quantization is V1 and V2, and the vertex coordinate after quantization is V', there are the following: V1, V2, (V1+V2) / 2 and Q -1 (V1+V2), where Q is the quantization matrix corresponding to the vertex coordinates of V1 and V2. The distortion measure D before and after basic quantization is used to select the optimal quantization point.

[0085] 2.Base mesh encoding.

[0086] a) After obtaining the base mesh, the DRACO encoder is used to encode its geometric information. This geometric information primarily includes connectivity relationships and geometric position information. The DRACO encoding process is as follows: first, the connectivity relationships are encoded. Then, the geometric position information of the points is encoded based on the connectivity relationships. Finally, the texture position information is encoded based on the connectivity relationships and geometric position information.

[0087] b) Encoding of connection relationships. DRACO uses the "Edgebreaker Coding" scheme to encode the connection relationships of the mesh. See Figure 7A for details. In Figure 7A, v represents the current vertex. Before encoding the connection relationships of the mesh, the vertices of the mesh are divided into five types: C, L, R, S, and E. The physical meaning of each symbol is as follows:

[0088] C: None of the triangles connected to the current vertex have been encoded;

[0089] L: The triangle on the left connected to the current vertex completes the encoding;

[0090] R: The triangle on the right side connected to the current vertex completes the encoding;

[0091] S: The left and right triangles connected to the current vertex have not been encoded;

[0092] E: The left and right triangles connected to the current vertex have been encoded.

[0093] Finally, the type of each vertex and the processing order of the vertices are encoded in a certain order, and the decoding end restores the geometric connection relationship of the mesh according to the processing order and type of the vertices.

[0094] c) Coding of geometric position information. After completing the coding of the vertex connection relationship, the geometric position information of each vertex is predictively coded based on the vertex connection relationship. The idea adopted by predictive coding is the "Parallelograms algorithm", as shown in Figure 7B. A simple linear fit is performed using the three vertices adjacent to the current point to be coded: the left vertex, the right vertex, and the opposite vertex: pred pos =(left+right)-opposite (1)

[0095] d) After completing the point connection relationship and geometric position information, the texture coordinates are predictively encoded based on the decoding and reconstruction of these two, as shown in Figure 7C. Similarly, assuming the current vertex is C, the left and right vertices of the current point can be obtained based on the point connection relationship. Then, the texture coordinates of the left and right vertices are used to predict the texture coordinates of the current vertex C.

[0096] 3. Displacement coefficient encoding.

[0097] a) First, after the base mesh is encoded and reconstructed, a partitioning algorithm is used to partition the base mesh to obtain the initial reconstructed mesh. Specifically, the curve corresponding to the subdivided mesh in Figure 5A is used to obtain the subdivided mesh (also called the "initial mesh") through simple linear interpolation. The coordinates of the newly inserted points are obtained by linear interpolation based on the two vertices on the current boundary:

[0098] b) Secondly, calculate the error Delta between the points in the subdivided mesh and the original mesh after division. The error Delta can be a point error in the world coordinate system. Finally, the displacement (i.e., displacement coefficient) of each point is calculated using the error Delta between each point and the normal vector Norm of each point. See Figure 8 for details. In Figure 8, the bold solid line represents the error Delta, and N and T represent the normal vector Norm. Thus, the specific calculation method is as follows: Displacement = Delta × Norm (3)

[0099] c) After calculating the displacement of each point, the spatial domain residual coefficient can be transformed into the frequency domain using the lifting transform to obtain the corresponding frequency domain residual coefficient.

[0100] d) Finally, a coefficient packing algorithm is used to map the frequency domain residual coefficients of each point into a two-dimensional image in a certain order. The current V-DMC can be arranged according to the Morton Code Order, as shown in Figure 9.

[0101] e) Finally, a traditional Video Codec is used to encode the two-dimensional image.

[0102] 4. Recoloring.

[0103] Recoloring is an algorithm on the encoder side. After the reconstruction of the encoder's geometric information is completed, the original geometric information, the original texture attribute information, and the reconstructed mesh geometric information are used to recolor the texture attribute information of the reconstructed mesh.

[0104] (2) Inter-frame geometric information coding (Inter coding).

[0105] a) Similar to the encoding above, the geometric position information includes the geometric connection relationship and the geometric position information encoding. However, it should be noted that the inter-frame geometric position information encoding only needs to encode the geometric position information (x, y, z) of the current base mesh, and does not need to encode the connection relationship and texture position information (u, v). The specific reason is as follows: if the current image can be inter-coded, then the base mesh of the reference image of the current image will be used at the encoder to obtain the mesh information of the current image. Therefore, the current image and the reference image have the same connection relationship and UV texture coordinates, only the geometric position information is different.

[0106] b) Based on a), it can be known that the only difference between the current image and the reference image is the geometric position information. Therefore, the current V-DMC performs predictive coding on the geometric position information of the current image.

[0107] Specifically as shown in Figure 10, the black dot is the point to be coded. The corresponding prediction point is obtained by using the current point in the reference image (similar to the same-position block in video coding). Then, the motion vector (MV) of the current point is predicted and coded using the neighboring points of the current point (MV of the coded vertex). The specific details are as follows. Assume that the coordinates of the current point are pos and the coordinates of the corresponding same-position point are Pred pos , then the MV of the current point is calculated as: MV = Pos-Pred pos (4)

[0108] There are two predictive coding modes in the current V-DMC:

[0109] i. Directly encode the MV of the current point;

[0110] ii. Use the neighborhood to perform predictive coding on the MV of the current point.

[0111] At the encoding end, the rate-distortion optimization algorithm is used to obtain the optimal coding mode for each coding group (CG). The current V-DMC sets the number of points for each CG to be at most 16.

[0112] Coding of texture attribute information: Current V-DMC encodes texture attribute information directly using a video codec (Video-Codec), such as AVC, HEVC, VVC, or VV-enC.

[0113] Figure 11A is a schematic diagram of the framework of an intra-frame encoder. As shown in Figure 11A, in the intra-frame encoder, a common static mesh encoder (Static Mesh Encoder) can be used to encode the base mesh to generate the corresponding bitstream (Compressed base mesh bitstream). Next, the reconstructed base mesh is used to update the displacement coefficients (Update Displacements). The updated displacement coefficients are subjected to wavelet transform (Wavelet Transform) and quantization (Quantization) to obtain the displacement coefficients. After being packaged into images and videos (Image Packing, Video Packing), they are encoded using HEVC to generate a bitstream (Compressed displacements bitstream) of the displacement coefficients. For attribute map encoding, the feature map is first transformed (Texture Transfer) according to the difference between the reconstructed geometric information and the original geometric information, and then padded (Padding) and packed (Video Packing) and encoded using a video encoder to form an attribute bitstream (Compressed attribute bitstream).

[0114] Figure 11B is a schematic diagram of an inter-frame encoder framework. As shown in Figure 11B, the inter-frame encoder and intra-frame encoder process are roughly the same, but the inter-frame encoder does not directly encode the base grid. Instead, it encodes the motion vector MV between the base grid of the current image and the base grid of the reference image and generates the corresponding motion vector bitstream (Compressed motion bitstream).

[0115] Correspondingly, during the decoding process, the decoder can also be divided into an intra-frame decoder and an inter-frame decoder according to the type of the frame it operates on, which are used to perform intra-frame decoding and inter-frame decoding respectively.

[0116] FIG12A is a schematic diagram of intra-frame decoding. As shown in FIG12A , in the intra-frame decoder, a static mesh decoder can be used to decode the base mesh. A video decoder is used to decode the shift coefficient video, and the shift coefficient is obtained through video unpacking and inverse wavelet transform. The decoded base mesh and shift coefficient are used to obtain the decoded mesh geometry information. The attribute map is decoded directly through the video decoder.

[0117] Figure 12B is a schematic diagram of inter-frame decoding. As shown in Figure 12B, the process for an inter-frame decoder is basically the same as that for an intra-frame decoder. However, instead of directly decoding the base grid, the motion vector is decoded and the base grid of the current image is calculated using the base grid of the previous frame (e.g., the reference image).

[0118] In summary, in the dynamic mesh coding (Dynamic Mesh Coding) currently provided by MPEG, the dynamic mesh coding process is divided into the following steps: at the encoding end, the basic mesh generated by preprocessing is quantized and then encoded using Google's open source DRACO encoder, and the shift coefficients are encoded using HEVC after wavelet transform, quantization, and two-dimensional mapping. The two-dimensional attribute map is also directly transmitted to the HEVC encoder for encoding; at the decoding end, the basic mesh code stream is decoded by DRACO to generate a decoded basic mesh, and the shift coefficients are decoded by HEVC decoding, inverse two-dimensional mapping, inverse quantization, and inverse transformation to generate decoded shift coefficients. Then, the decoded basic mesh and the decoded shift coefficients are used together to reconstruct the three-dimensional mesh geometry, and the attribute code stream is decoded by HEVC to generate a reconstructed attribute map.

[0119] (3) General test conditions for MPEG DMC.

[0120] a. There are 2 test conditions:

[0121] Condition 1: all intra geometry is lossy and attributes are lossy;

[0122] Condition 2: Random access is lossy in geometry and attributes;

[0123] b. The general test sequence may include five categories, namely Cat1-A, Cat1-B and Cat1-C, all of which contain geometric information and color attribute information.

[0124] In the related art, when performing inter-frame coding on the basic grid, a unified coding mode is usually used to encode the motion vectors of the vertices in the basic grid in the three coordinate dimensions (x, y, z), and a unified coding mode flag is used to indicate its coding and decoding mode. It can be seen that the inter-frame coding mode of the basic grid in the related art is not perfect enough, which reduces the coding and decoding efficiency of the basic grid. If the motion vector of the vertex in each coordinate dimension of the basic grid can be determined to encode the optimal coding mode of the coordinate dimension, and an independent flag is used to indicate the coding mode corresponding to each coordinate dimension, so that the decoder adopts the corresponding decoding method for decoding, the coding and decoding efficiency of the basic grid can be greatly improved, thereby improving the coding and decoding performance of DMC.

[0125] Based on this foundation, an embodiment of the present application provides a coding and decoding method, which determines the basic grid of the current image at the coding end; when performing inter-frame coding on the basic grid, determines the motion vector corresponding to the vertex of the basic grid in each coordinate dimension; based on the motion vector corresponding to the vertex in each coordinate dimension, encodes in at least one coding mode and determines at least one coding cost in each coordinate dimension; based on at least one coding cost, determines the dimension coding flag corresponding to the vertex in each coordinate dimension and the motion vector coding information corresponding to the vertex; the dimension coding flag corresponding to each coordinate dimension represents the motion vector prediction mode corresponding to the coordinate dimension. In this way, for the motion vector of the basic grid in each coordinate dimension, the optimal coding mode corresponding to the coordinate dimension can be determined by comparing at least one coding cost, and the dimension coding flag corresponding to each coordinate dimension is used to indicate the encoder's decoding method for the coordinate dimension, thereby improving coding and decoding efficiency and thus improving coding and decoding performance.

[0126] The embodiment of the present application also provides a grid architecture of a codec system including a decoding method and an encoding method. Figure 13 is a schematic diagram of a grid architecture of a codec provided by the embodiment of the present application. As shown in Figure 13, the grid architecture includes one or more electronic devices 13 to 1N and a communication grid 01, wherein the electronic devices 13 to 1N can perform video interaction through the communication grid 01. During the implementation process, the electronic device can be various types of devices with codec functions. For example, the electronic device can include a mobile phone, a tablet computer, a personal computer, a personal digital assistant, a navigator, a digital phone, a video phone, a television, a sensor device, a server, etc., which is not specifically limited in the embodiment of the present application. Here, the decoder or encoder described in the embodiment of the present application can be the above-mentioned electronic device.

[0127] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0128] In another embodiment of the present application, referring to FIG14 , a schematic flow chart of an encoding method provided by an embodiment of the present application is shown. As shown in FIG14 , the method may include:

[0129] S1401: Determine a base grid of the current image and motion vectors corresponding to vertices in the base grid in each coordinate dimension.

[0130] It should be noted that the encoding method in the embodiment of the present application is applied to an encoder. The encoding method in the embodiment of the present application may refer to an inter-frame encoding method, more specifically, an inter-frame encoding method for a base grid in a dynamic grid. The encoding method can be applied to an encoder in a V-DMC, but is not limited thereto.

[0131] It should also be noted that in the embodiments of the present application, the base grid may also be referred to as a "simplified grid." In some embodiments, determining the base grid of the current image based on the original grid of the current image may include quantizing or downsampling the original grid of the current image (i.e., the current frame) to determine the base grid of the current image.

[0132] For example, first, the original mesh of the current image may be downsampled to generate a basic mesh with a significantly reduced number of vertices.

[0133] In an embodiment of the present application, a base mesh includes multiple vertices. In some embodiments, the vertices in the base mesh may include every vertex in the base mesh. When encoding the base mesh in inter-frame coding mode, the encoder determines a motion vector for the vertex in the base mesh. The motion vector for the vertex is the motion vector between a first geometric coordinate of the vertex in the base mesh corresponding to the current image and a second geometric coordinate of a reference vertex in a reference base mesh corresponding to a reference image of the current image.

[0134] In an embodiment of the present application, the motion vector of a vertex corresponds to multiple coordinate dimensions. For example, the motion vector of a vertex in a three-dimensional grid corresponds to three coordinate dimensions of x, y, and z. When the encoder performs inter-frame encoding of the motion vectors of the vertices in the basic grid, it first determines the motion vector corresponding to the vertex in each coordinate dimension, and adopts at least one encoding mode for the motion vector corresponding to each coordinate dimension to determine the optimal encoding mode for the vertex in the coordinate dimension. It should be noted that the first coordinate dimension, the second coordinate dimension, or the third coordinate dimension in the embodiment of the present application can be any coordinate dimension of x, y, or z, and the embodiment of the present application does not limit this.

[0135] In some embodiments, the vertices in the base mesh include: each vertex in the current group of at least one group of the base mesh, or each vertex in the base mesh. For example, in the case of non-group coding, the vertices include each vertex in the base mesh; in the case of group coding, the vertices include each vertex in the current group of at least one group of the base mesh. The specific selection depends on the actual situation and is not limited in the embodiments of this application. They will be explained in the following embodiments.

[0136] S1402: When performing inter-frame coding on the basic grid, skip coding and coding cost estimation of at least one prediction mode are performed on the motion vector corresponding to each coordinate dimension of the vertex in at least one coordinate dimension to determine a first coding cost and at least one second coding cost corresponding to each coordinate dimension.

[0137] In an embodiment of the present application, the encoder may perform intra-frame encoding and inter-frame encoding on the basic grid, respectively, and determine the image-level encoding method corresponding to the basic grid of the current image by comparing the encoding cost of intra-frame encoding and the encoding cost of inter-frame encoding. It should be noted that, in an embodiment of the present application, the image level is equivalent to the frame level, and the image-level syntax information is equivalent to the frame-level syntax information. Among them, when the basic grid is inter-frame encoded, the encoder estimates the encoding cost of at least one prediction mode for the motion vector corresponding to the vertex in each coordinate dimension, so as to determine the optimal encoding mode and motion vector encoding value corresponding to each coordinate dimension. In some embodiments, the encoding value includes encoding bits.

[0138] In some embodiments, the at least one coordinate dimension may include one or more of the x, y, and z dimensions.

[0139] In some embodiments, the encoder determines second syntax information; the second syntax information represents whether the current image is intra-frame encoded or inter-frame encoded; the second syntax information is entropy encoded, and the obtained encoding bits are written into the code stream. The second syntax information can act on the sequence level and / or the image level, and the specific selection is made according to the actual situation, which is not limited by the embodiments of the present application. For example, when the second syntax information is 0, it indicates that the encoding mode of the basic grid of the current image is the inter-frame encoding mode; when the second syntax information is 1, it indicates that the encoding mode of the basic grid of the current image is the intra-frame encoding mode; when the second syntax information is 2, it indicates that the encoding mode of the basic grid of the current image is the skip encoding mode; 3 is a reserved bit. The specific selection is made according to the actual situation, which is not limited by the embodiments of the present application.

[0140] In this embodiment of the present application, the first coding cost corresponds to a skip coding mode; and the at least one second coding cost corresponds to at least one prediction mode. For each coordinate dimension, the encoder performs skip coding and estimates the coding cost of at least one prediction mode on the motion vector corresponding to the vertex in that coordinate dimension, determining the first coding cost and at least one second coding cost for that coordinate dimension, thereby determining the first coding cost and at least one second coding cost for each coordinate dimension.

[0141] In an embodiment of the present application, the coding cost can be evaluated and determined by bit rate and / or distortion cost. In some embodiments, since the motion vector encoding of the vertex in the basic mesh is usually lossless coding, the bit rate can be used to estimate the coding cost in each coordinate dimension. For each coordinate dimension, the encoder can determine the first bit rate estimate and at least one second bit rate estimate corresponding to the coordinate dimension by skipping coding and bit rate estimation of at least one prediction mode for the motion vector corresponding to the vertex in the coordinate dimension; the first bit rate estimate is used as the first coding cost, and the at least one second bit rate estimate is used as the at least one second coding cost. In some embodiments, quality distortion can also be introduced as a coding cost. For each coordinate dimension, the encoder can determine the first distortion cost and at least one second distortion cost corresponding to the vertex in the coordinate dimension by skipping coding and distortion cost estimation of at least one prediction mode for the motion vector corresponding to the vertex in the coordinate dimension; the first distortion cost is used as the first coding cost, and the at least one second distortion cost is used as the at least one second coding cost.

[0142] In some embodiments, for the vertices in the basic grid, it includes: for each vertex of the current group in at least one group of the basic grid, for each coordinate dimension, skip coding estimation and coding estimation of at least one prediction mode are performed on the motion vector corresponding to the coordinate dimension of each vertex of the current group, and the first coding cost and at least one second coding cost corresponding to the coordinate dimension of each vertex of the current group are determined; based on the first coding cost and at least one second coding cost corresponding to the coordinate dimension of each vertex of the current group, the first coding cost and at least one second coding cost corresponding to the coordinate dimension of the current group are determined, thereby determining the first coding cost and at least one second coding cost corresponding to each coordinate dimension.

[0143] Among them, for each coordinate dimension, the sum or weighted average of the first coding costs corresponding to each vertex in the current group in the coordinate dimension is counted to determine the first coding cost corresponding to the current group in the coordinate dimension; the sum or weighted average of the second coding costs corresponding to the same prediction mode corresponding to each vertex in the current group in the coordinate dimension is counted to determine at least one second coding cost corresponding to the current group in the coordinate dimension; thereby determining the first coding cost and at least one second coding cost corresponding to the current group in each coordinate dimension.

[0144] For example, the coding cost can be determined by the bit rate estimate value, which can be determined by the bit rate estimate value bits. ijRepresents the sum of the estimated bitrates for the motion vector of the i-th coordinate dimension (i can be 0, 1, or 2) of each vertex in the current group, encoded using the j-th prediction mode. Assuming there are three coordinate dimensions and two prediction modes, using skip mode and two prediction modes for coding cost estimation, 3*3=9 estimated bitrates can be calculated.

[0145] It should be noted that each group in the at least one group is processed in the same manner.

[0146] In some embodiments, the vertices in the base mesh include: each vertex in the base mesh. For each coordinate dimension, skip coding and coding cost estimation of at least one prediction mode are performed on the motion vector corresponding to each vertex in the base mesh in the coordinate dimension to determine a first coding cost and at least one second coding cost corresponding to each vertex in the base mesh in the coordinate dimension; based on the first coding cost and at least one second coding cost corresponding to each vertex in the base mesh in the coordinate dimension, the first coding cost and at least one second coding cost corresponding to the base mesh in the coordinate dimension are determined, thereby determining the first coding cost and at least one second coding cost corresponding to each coordinate dimension.

[0147] Among them, for each coordinate dimension, the sum or weighted average of the first coding costs corresponding to each vertex in the basic grid in the coordinate dimension is counted to determine the first coding cost corresponding to the basic grid in the coordinate dimension; the sum or weighted average of the second coding costs corresponding to the same prediction mode corresponding to each vertex in the basic grid in the coordinate dimension is counted to determine at least one second coding cost corresponding to the basic grid in the coordinate dimension; thereby determining the first coding cost and at least one second coding cost corresponding to the basic grid in each coordinate dimension.

[0148] It should be noted that, for the weighted average, the weight corresponding to each vertex can be the same or different.

[0149] In the embodiment of the present application, each vertex corresponds to an index. For example, if the total number of vertices in the base mesh is N, the index of the vertices in the base mesh can be from 0 to N-1; wherein N is an integer greater than or equal to 3. In this way, the encoder can encode the vertices in the base mesh according to the index corresponding to each vertex according to a preset encoding order, such as in ascending order. It should be noted that the preset encoding order can also be other orders, which are specifically selected according to actual conditions and are not limited in the embodiment of the present application.

[0150] In an embodiment of the present application, the encoder may divide each vertex in the base mesh into at least one group according to the index, and implement encoding of the base mesh in groups. For example, the size of the group may be 16, indicating that the vertices in the base mesh are grouped in units of 16 vertices, and the number of vertices in the last group is 1 to 16. The size of the group may also take other values, which are specifically selected according to the actual situation and are not limited in the embodiment of the present application. Taking the total number of vertices of the base mesh as 40 as an example, the group size is 16, then the indices of the vertices contained in the first group are 0 to 15, the indices of the vertices contained in the second group are 16 to 31, and the indices of the vertices contained in the third group are 32 to 39.

[0151] In some embodiments, the encoder may encode the base grid in group coding as a default mode, or may determine whether to group code the base grid based on fourth syntax information and / or fifth syntax information. The fourth syntax information indicates whether to group code the base grids of all images in the current sequence to which the current image belongs, and the fifth syntax information indicates whether to group code the base grid of the current image.

[0152] In some embodiments, fourth grammatical information and / or fifth grammatical information are determined; when the fourth grammatical information and / or the fifth grammatical information represent grouped encoding of the basic mesh, grouping is performed based on the index corresponding to each vertex in the basic mesh and the number of vertices, and at least one group is determined.

[0153] Exemplarily, if the value of the fourth grammatical information and / or the fifth grammatical information is the first identification value, it is determined that the basic grid is to be grouped and encoded; if the value of the fourth grammatical information and / or the fifth grammatical information is the second identification value, it is determined that the basic grid is not to be grouped and encoded. In an embodiment of the present application, the first identification value and the second identification value are different. Here, the first identification value and the second identification value can be in parameter form, such as the first identification value can be TRUE or true, and the second identification value can be FALSE or false, or can be in digital form, such as the first identification value can be 1 and the second identification value can be 0. Conversely, the first identification value can also be FALSE or false, the second identification value can be TRUE or true, or the first identification value can be 0, the second identification value can be 1, and so on. The specific selection is made according to the actual situation and is not limited in the embodiment of the present application. In some embodiments, the fourth grammatical information and / or the fifth grammatical information here can be a parameter written in the profile, or can be the value of a flag, which is not specifically limited here.

[0154] In some embodiments, in the case of group coding, the encoder may also determine the sixth grammatical information and / or the seventh grammatical information based on the number of vertices in the group. The sixth grammatical information and / or the seventh grammatical information represent the number of vertices in the group. Exemplarily, the sixth grammatical information may represent the number of vertices of all images in the current sequence; the seventh grammatical information may represent the number of vertices of the current image. In other words, the sixth grammatical information and / or the seventh grammatical information represent the size of the current group. The encoder performs entropy coding on the sixth grammatical information and / or the seventh grammatical information, and writes the obtained coding bits into the bitstream. Exemplarily, the sixth grammatical information and / or the seventh grammatical information may be 16 or other values, which may be selected based on actual conditions and are not limited in the embodiments of the present application.

[0155] It should be noted that when the total number of vertices of the base mesh is not an integer multiple of the number of vertices in the group, the number of vertices in the last group is less than the number of vertices in the group.

[0156] In some embodiments, whether to group encode the base grid may also be determined based on the sixth grammatical information and / or the seventh grammatical information. When the sixth grammatical information and / or the seventh grammatical information is a preset identification value, the encoder encodes the base grid without grouping. For example, the preset identification value may be 0 or a character-based identification value; the preset identification value may be preset in the form of a reserved bit, and the specific selection is based on actual circumstances and is not limited in the present embodiment.

[0157] In some embodiments, the at least one second coding cost corresponds to at least one of the at least one prediction mode and the no-prediction mode. The process of determining the at least one coding cost in S1402 may be as follows, where the vertex in the following process includes each vertex in the base mesh, or each vertex in the current group of the at least one group of the base mesh:

[0158] For each coordinate dimension, determining a first motion vector prediction value corresponding to the vertex in the coordinate dimension by using a no-prediction mode; the no-prediction mode indicates that the motion vector prediction value corresponding to the coordinate dimension is set to a preset vector prediction value;

[0159] Estimating a coding cost based on the first motion vector prediction value to determine a second coding cost corresponding to the no-prediction mode;

[0160] and / or,

[0161] Determine a second motion vector prediction value according to a first weighted average of motion vector reconstruction values ​​corresponding to a first number of neighboring points of the vertex in the coordinate dimension;

[0162] Estimating a coding cost based on the second motion vector prediction value to determine a second coding cost corresponding to a first prediction mode in at least one prediction mode;

[0163] and / or,

[0164] Determining a first offset value according to first weighted values ​​of motion vector reconstruction values ​​corresponding to a first number of neighboring points of the vertex in the coordinate dimension;

[0165] Determining a third motion vector prediction value according to the first weighted average value and the first offset value;

[0166] performing coding cost estimation based on the third motion vector prediction value to determine a second coding cost corresponding to a second prediction mode in at least one prediction mode;

[0167] and / or,

[0168] Determine a fourth motion vector prediction value according to a second weighted average of motion vector reconstruction values ​​corresponding to a second number of neighboring points of the vertex in the coordinate dimension;

[0169] performing coding cost estimation based on the fourth motion vector prediction value to determine a second coding cost corresponding to a third prediction mode in at least one prediction mode;

[0170] and / or,

[0171] Determine a second offset value according to a second weighted average of motion vector reconstruction values ​​corresponding to a second number of neighboring points of the vertex in the coordinate dimension;

[0172] determining a fifth motion vector prediction value according to the second weighted average value and the second offset value;

[0173] A coding cost is estimated according to the fifth motion vector prediction value to determine a second coding cost corresponding to a fourth prediction mode in at least one prediction mode.

[0174] Thus, at least one of the second coding cost corresponding to the no prediction mode, the second coding cost corresponding to the first prediction mode, the second coding cost corresponding to the second prediction mode, the second coding cost corresponding to the third prediction mode and the second coding cost corresponding to the fourth prediction mode is determined as at least one second coding cost.

[0175] In some embodiments, the first number of neighbor points may include all neighbor points of the vertex, and the second number of neighbor points may include some neighbor points of the vertex. The first number of neighbor points or the second number of neighbor points may include one or more of the following:

[0176] In the current image, the vertices that are connected to the vertex and whose encoding order is before the vertex;

[0177] and / or, a vertex with the same index as the vertex in the reference image corresponding to the current image;

[0178] and / or, all vertices that are connected to the vertex's co-location in the reference image;

[0179] And / or, in the reference image, the vertices with the same index as their neighbor points in the current image.

[0180] In some embodiments, the first weighted average of the motion vector reconstruction values ​​corresponding to the first number of neighbor points in the coordinate dimension can be calculated by the formula OK. Where predCount is the first number, is the weighted value of the motion vector reconstruction value corresponding to the neighbor point in the coordinate dimension. The weights of different neighbor points can be the same or different.

[0181] In some embodiments, the third motion vector prediction value can be obtained by the formula Determine. Where bias is the first bias value. When the value is greater than or equal to the preset threshold, the first number predCount is right-shifted by the preset number of bits to determine the first bias value bias. When bias=predCount>>1.

[0182] In some embodiments, the second weighted average of the motion vector reconstruction values ​​corresponding to the second number of neighboring points in the coordinate dimension can be calculated by the formula OK. Where predCount' is the second number, is the weighted value of the motion vector reconstruction value corresponding to the neighbor point in the coordinate dimension. The weights of different neighbor points can be the same or different.

[0183] In some embodiments, the fifth motion vector prediction value can be obtained by the formula Determine. Where bias' is the second bias value. When the value is greater than or equal to the preset threshold, the second number predCount' is right-shifted by the preset number of bits to determine the second bias value bias'. When bias'=predCount'>>1.

[0184] S1403: Determine first syntax information according to a first coding cost and at least one second coding cost corresponding to each coordinate dimension.

[0185] In embodiments of the present application, the first syntax information indicates whether to skip encoding of motion vectors. In some embodiments, the first syntax information may apply to all dimensions of a vertex, or may apply to each coordinate dimension corresponding to the vertex. In other words, the first syntax information may be a single piece of syntax information corresponding to all dimensions of a vertex, or a single piece of syntax information corresponding to each coordinate dimension of a vertex.

[0186] In the embodiment of the present application, the first coding cost and the at least one second coding cost represent the coding performance of the skip coding mode and each prediction mode in each coordinate dimension. Thus, by comparing the coding performance of the skip coding mode and the at least one prediction mode in each coordinate dimension, the optimal coding mode corresponding to the vertex can be determined, and the first syntax information can be determined based on the determined optimal coding mode.

[0187] In some embodiments, the encoder encodes the first syntax information corresponding to the vertex and writes the obtained coded bits into the bitstream.

[0188] It can be understood that at the encoding end, the motion vectors corresponding to the base grid of the current image and the vertices in the base grid in each coordinate dimension are determined; when performing inter-frame encoding on the base grid, the motion vectors corresponding to the vertices in each coordinate dimension are estimated for skip coding and at least one prediction mode, and a first coding cost and at least one second coding cost corresponding to each coordinate dimension are determined; the first coding cost corresponds to the skip coding mode; the at least one second coding cost corresponds to at least one prediction mode; first syntax information is determined based on the first coding cost and at least one second coding cost corresponding to each coordinate dimension; the first syntax information indicates whether to skip the encoding of the motion vector. By comparing the coding costs of the motion vectors of the vertices in the base grid in each coordinate dimension under skip coding and at least one prediction mode, the optimal processing method for the vertex in that coordinate dimension can be determined, such as skip coding or the optimal prediction mode, and the first syntax information is used to indicate the optimal processing method to the decoder, thereby improving coding efficiency and thus improving coding performance.

[0189] In some embodiments of the present application, the first syntax information is applied to each coordinate dimension corresponding to the vertex. The first syntax information is determined based on the first coding cost and at least one second coding cost corresponding to each coordinate dimension, including:

[0190] For each coordinate dimension, when the first coding cost is less than or equal to each second coding cost of at least one second coding cost, the first syntax information corresponding to the vertex in the coordinate dimension is determined to be a first value; the first value represents skipping the encoding of the motion vector of the vertex in the coordinate dimension.

[0191] For each coordinate dimension, when the first coding cost is greater than or equal to any second coding cost of at least one second coding cost, the first grammatical information corresponding to the vertex in the coordinate dimension is determined to be a second value; the second value indicates that the encoding of the motion vector of the vertex in the coordinate dimension is not skipped.

[0192] For each coordinate dimension, when the first coding cost is greater than or equal to any second coding cost of at least one second coding cost, the target prediction mode corresponding to the vertex in each coordinate dimension is determined according to the at least one second coding cost; and the third grammatical information corresponding to the vertex in each coordinate dimension is determined according to the target prediction mode corresponding to the vertex in each coordinate dimension.

[0193] In this embodiment, for each coordinate dimension, when the first coding cost is less than or equal to each second coding cost of at least one second coding cost, the motion vector encoding corresponding to the vertex in the coordinate dimension is set to a preset value.

[0194] The first value and the second value are different. The first value can be TRUE or true, and the second value can be FALSE or false. They can also be in digital form, such as the first value can be 1 and the second value can be 0. Conversely, the first value can also be FALSE or false, and the second value can be TRUE or true, or the first value can be 0 and the second value can be 1, etc. The specific selection depends on the actual situation and is not limited in the embodiments of the present application.

[0195] That is, for each coordinate dimension, if the first coding cost corresponding to the skip coding mode is the minimum of the first coding cost and at least one second coding cost, or is equal to the second coding cost corresponding to each prediction mode, then it is determined that coding is skipped for that coordinate dimension, and the motion vector corresponding to the vertex in that coordinate dimension is set to a preset value. Exemplarily, the preset value can be 0, or it can be set to other values ​​depending on the actual situation. The specific selection is based on the actual situation and is not limited in the embodiments of this application.

[0196] For each coordinate dimension, if the first coding cost corresponding to the skip coding mode is not the minimum value of the first coding cost and at least one second coding cost, or is equal to the second coding cost corresponding to any prediction mode, it is determined that the coordinate dimension will not be skipped, and the optimal prediction mode, that is, the target prediction mode, is determined based on the at least one second coding cost, and the third grammatical information corresponding to the vertex in the coordinate dimension is determined based on the target prediction mode.

[0197] It should be noted that, in the embodiment of the present application, the third syntax information corresponding to the vertices in different coordinate dimensions may be the same or different.

[0198] In this embodiment of the present application, at least one preset identification value may be preconfigured to correspond to at least one prediction mode, with each of the at least one preset identification value corresponding to one of the at least one prediction mode. In this way, the third syntax information corresponding to each vertex in each coordinate dimension may be determined based on the preset identification value corresponding to the optimal prediction mode determined for each vertex in each coordinate dimension.

[0199] In some embodiments, for each coordinate dimension, the encoder may determine the third syntax information corresponding to the vertex in that coordinate dimension based on the minimum second encoding cost among at least one second encoding cost, thereby determining the third syntax information corresponding to the vertex in each coordinate dimension. The encoder performs the same processing for each coordinate dimension to determine the third syntax information corresponding to the vertex in each coordinate dimension.

[0200] In some embodiments, the first syntax information corresponding to the vertex and the third syntax information corresponding to the vertex in each coordinate dimension are encoded, and the obtained encoding bits are written into the bitstream.

[0201] In some embodiments, motion vector prediction is performed based on the target prediction mode corresponding to the vertex in each coordinate dimension to determine the target motion vector prediction value corresponding to the vertex in each coordinate dimension; based on the motion vector corresponding to the vertex in each coordinate dimension and the target motion vector prediction value, the motion vector residual corresponding to the vertex in each coordinate dimension is determined; the motion vector residual corresponding to the vertex in each coordinate dimension is encoded to determine the motion vector encoding corresponding to the vertex in each coordinate dimension.

[0202] In some embodiments, the motion vector corresponding to each coordinate dimension of the vertex is encoded and written into the bitstream.

[0203] In some embodiments, for each coordinate dimension, when the absolute value of the motion vector residual corresponding to the vertex in the coordinate dimension is greater than a first preset residual threshold and less than a second preset residual threshold, the eighth grammatical information corresponding to the vertex in the coordinate dimension is determined as a third value, and the ninth grammatical information corresponding to the vertex in the coordinate dimension is determined as a fourth value; the second preset residual threshold is greater than the first preset residual threshold; the motion vector residual corresponding to the vertex in the coordinate dimension is encoded, and the motion vector code corresponding to the vertex in the coordinate dimension is determined, thereby determining the motion vector code corresponding to the vertex in each coordinate dimension. The motion vector residual can be a positive value or a negative value.

[0204] In some embodiments, for each coordinate dimension, when the absolute value of the motion vector residual corresponding to the vertex in the coordinate dimension is greater than or equal to the second preset residual threshold, the eighth grammatical information corresponding to the vertex in the coordinate dimension is determined as the third value, and the ninth grammatical information corresponding to the vertex in the coordinate dimension is determined as the third value; the remainder of the motion vector residual corresponding to the vertex in the coordinate dimension is encoded to determine the motion vector code corresponding to the vertex in the coordinate dimension, thereby determining the motion vector code corresponding to the vertex in each coordinate dimension.

[0205] Exemplarily, the first preset residual threshold may be 0, and the second preset residual threshold may be 1. The specific selection is made according to actual conditions and is not limited in the embodiments of the present application.

[0206] In some embodiments, entropy coding is performed on the eighth syntax information and the ninth syntax information corresponding to the vertex in each coordinate dimension, and the obtained coding bits and the motion vector coding corresponding to the vertex in each coordinate dimension are written into the bitstream.

[0207] In some embodiments, for the case of non-group coding, the vertex in the above process refers to each vertex in the base mesh, and the first syntax information is applied to the base mesh, specifically, to each coordinate dimension corresponding to the base mesh, to indicate whether to skip coding the coordinate dimension corresponding to each vertex in the base mesh. For the case of group coding, the vertex in the above process refers to each vertex in the current group in at least one group in the base mesh, and the first syntax information is applied to the current group, specifically, to each coordinate dimension corresponding to the current group, to indicate whether to skip coding the coordinate dimension corresponding to each vertex in the current group.

[0208] In some embodiments, motion vector prediction is performed based on the target prediction mode corresponding to the vertex in each coordinate dimension to determine the target motion vector prediction value corresponding to the vertex in each coordinate dimension; based on the motion vector corresponding to the vertex in each coordinate dimension and the target motion vector prediction value, the motion vector residual corresponding to the vertex in each coordinate dimension is determined; the motion vector residual corresponding to the vertex in each coordinate dimension is encoded to determine the motion vector encoding corresponding to the vertex in each coordinate dimension.

[0209] It can be understood that by comparing the first coding cost with at least one second coding cost of the motion vector of the vertex in each coordinate dimension, the optimal coding mode corresponding to the vertex in the coordinate dimension can be determined, and the first syntax information corresponding to each coordinate dimension is used in combination with the third syntax information to indicate the encoder how to decode the coordinate dimension, thereby improving the coding efficiency and further improving the coding performance.

[0210] In some embodiments of the present application, the first syntax information is applied to the vertex, that is, to the three coordinate dimensions corresponding to the vertex. The first syntax information is determined based on the first coding cost and at least one second coding cost corresponding to each coordinate dimension, including:

[0211] When the sum of the first coding costs of the vertex in three coordinate dimensions is less than or equal to the sum of the three coordinate dimensions of each second coding cost in at least one second coding cost, the first syntax information is determined to be the fifth value; the fifth value represents skipping the encoding of the motion vector of the vertex.

[0212] When the sum of the first coding costs of the vertex in three coordinate dimensions is greater than or equal to the sum of the three coordinate dimensions of each second coding cost in at least one second coding cost, the first syntax information is determined to be the sixth value; the sixth value indicates that the encoding of the motion vector of the vertex is not skipped.

[0213] The fifth value and the first value may be the same or different, the sixth value and the second value may be the same or different, and the fifth value and the sixth value may be different. The fifth value may be TRUE or true, and the sixth value may be FALSE or false. They may also be in numerical form, such as the fifth value may be 1 and the sixth value may be 0. Conversely, the fifth value may be FALSE or false, and the sixth value may be TRUE or true, or the fifth value may be 0 and the sixth value may be 1, etc. The specific selection depends on the actual situation and is not limited in the embodiments of the present application.

[0214] When the sum of the first coding costs of the vertex in three coordinate dimensions is less than or equal to the sum of the three coordinate dimensions of each second coding cost in at least one second coding cost, the motion vector corresponding to the vertex is set to a preset value.

[0215] That is, for the three coordinate dimensions, the sum of the first coding costs corresponding to the skip coding mode in the three coordinate dimensions is determined; the sum of each second coding cost in the at least one second coding cost in the three coordinate dimensions is determined, and then at least one second coding cost sum value is determined. If the sum of the first coding costs is less than or equal to each second coding cost sum value in the at least one second coding cost sum value, then it is determined that all three coordinate dimensions of the vertex are skipped, the first syntax information is determined to be the fifth value, and the motion vector corresponding to the vertex is set to a preset value. In other words, the motion vector corresponding to the vertex in the three coordinate dimensions is set to the preset value.

[0216] In this embodiment, when the sum of the first coding costs of a vertex in three coordinate dimensions is greater than or equal to the sum of the three coordinate dimensions of each second coding cost in at least one second coding cost, the target prediction mode corresponding to the vertex in each coordinate dimension is determined based on the at least one second coding cost corresponding to the vertex in each coordinate dimension; and the third grammatical information corresponding to the vertex in each coordinate dimension is determined based on the target prediction mode corresponding to the vertex in each coordinate dimension.

[0217] That is to say, for the three coordinate dimensions, if the sum of the first coding costs is greater than or equal to each second coding cost and value in at least one second coding cost and value, it is determined not to skip encoding of the three coordinate dimensions of the vertex, the first syntax information is determined to be the sixth value, and the target prediction mode is determined based on the comparison of at least one second coding cost of the vertex in each coordinate dimension, and then the third syntax information is determined based on the target prediction mode.

[0218] In some embodiments, the first syntax information corresponding to the vertex and the third syntax information corresponding to the vertex in each coordinate dimension are encoded, and the obtained encoding bits are written into the bitstream.

[0219] In this embodiment, motion vector prediction is performed based on the target prediction mode corresponding to the vertex in each coordinate dimension to determine the target motion vector prediction value corresponding to the vertex in each coordinate dimension; based on the motion vector corresponding to the vertex in each coordinate dimension and the target motion vector prediction value, the motion vector residual corresponding to the vertex in each coordinate dimension is determined; the motion vector residual corresponding to the vertex in each coordinate dimension is encoded to determine the motion vector encoding corresponding to the vertex in each coordinate dimension.

[0220] In some embodiments, the motion vector corresponding to each coordinate dimension of the vertex is encoded and written into the bitstream.

[0221] In some embodiments, for the case of non-group coding, the vertex in the above process refers to each vertex in the base mesh, and the first syntax information is applied to the base mesh to indicate whether to skip coding for each vertex in the base mesh, that is, whether to skip coding for the three coordinate dimensions corresponding to each vertex in the base mesh. For the case of group coding, the vertex in the above process refers to each vertex in the current group in at least one group in the base mesh, and the first syntax information is applied to the current group to indicate whether to skip coding for each vertex in the current group, that is, whether to skip coding for the three coordinate dimensions corresponding to each vertex in the current group.

[0222] It can be understood that the optimal encoding mode corresponding to the vertex can be determined by comparing the first encoding cost with at least one second encoding cost of the motion vector of the vertex in each coordinate dimension, and the first grammatical information combined with the third grammatical information can be used to instruct the encoder on how to decode the coordinate dimension, thereby improving the encoding efficiency and further improving the encoding performance.

[0223] In some embodiments of the present application, the first syntax information is applied to each coordinate dimension corresponding to the vertex, and the tenth syntax information is applied to the vertex, that is, to the three coordinate dimensions corresponding to the vertex. The first syntax information is determined based on the first coding cost and at least one second coding cost corresponding to each coordinate dimension, including:

[0224] The tenth grammatical information is determined based on the first coding cost and at least one second coding cost corresponding to each coordinate dimension; the tenth grammatical information represents whether to skip encoding of the motion vector of the vertex; when the tenth grammatical information represents that encoding of the motion vector of the vertex is not skipped, the first grammatical information corresponding to the vertex in each coordinate dimension is determined based on the first coding cost and at least one second coding cost corresponding to each coordinate dimension; the first grammatical information represents whether to skip encoding of the motion vector corresponding to the vertex in the coordinate dimension.

[0225] When the sum of the first coding costs of the vertex in three coordinate dimensions is less than or equal to the sum of the three coordinate dimensions of each second coding cost in at least one second coding cost, the tenth grammatical information is determined to be the seventh value; the seventh value represents skipping the encoding of the motion vector of the vertex.

[0226] When the sum of the first coding costs of the vertex in three coordinate dimensions is greater than or equal to the sum of the three coordinate dimensions of each second coding cost in at least one second coding cost, the tenth grammatical information is determined to be the eighth value; the eighth value indicates that the encoding of the motion vector of the vertex is not skipped.

[0227] The seventh value and the first value may be the same or different, the eighth value and the second value may be the same or different, and the seventh value and the eighth value may be different. The seventh value may be TRUE or true, and the eighth value may be FALSE or false. They may also be in numerical form, such as the seventh value may be 1 and the eighth value may be 0. Conversely, the seventh value may be FALSE or false, and the eighth value may be TRUE or true, or the seventh value may be 0 and the eighth value may be 1, etc. The specific selection is based on actual conditions and is not limited in the embodiments of the present application.

[0228] In this embodiment, when the sum of the first coding costs of a vertex in three coordinate dimensions is less than or equal to the sum of the three coordinate dimensions of each second coding cost in at least one second coding cost, the motion vector encoding corresponding to the vertex is set to a preset value.

[0229] That is, for the three coordinate dimensions, the sum of the first coding costs corresponding to the skip coding mode in the three coordinate dimensions is determined; the sum of each second coding cost in the at least one second coding cost in the three coordinate dimensions is determined, and then at least one second coding cost sum value is determined. If the sum of the first coding costs is less than or equal to each second coding cost sum value in the at least one second coding cost sum value, then it is determined that all three coordinate dimensions of the vertex are skipped, the tenth syntax information is determined to be the seventh value, and the motion vector corresponding to the vertex is set to a preset value. In other words, the motion vectors corresponding to the vertex in the three coordinate dimensions are all set to preset values.

[0230] For each of the three coordinate dimensions, if the sum of the first coding costs is greater than or equal to each of the second coding costs and values ​​in at least one second coding cost and value, it is determined not to skip coding for all three coordinate dimensions of the vertex, the tenth grammatical information is determined to be the eighth value, and for each coordinate dimension, the coding cost of the skip coding mode is compared with that of at least one prediction mode to determine whether to skip coding for the coordinate dimension, that is, according to the first coding cost corresponding to each coordinate dimension and at least one second coding cost, the first grammatical information corresponding to the vertex in each coordinate dimension is determined.

[0231] Similarly, in this embodiment, determining the first syntax information corresponding to the vertex in each coordinate dimension according to the first encoding cost and at least one second encoding cost corresponding to each coordinate dimension may include:

[0232] For each coordinate dimension, when the first coding cost is less than or equal to each second coding cost of at least one second coding cost, the first syntax information corresponding to the vertex in the coordinate dimension is determined to be a first value; the first value represents skipping the encoding of the motion vector of the vertex in the coordinate dimension.

[0233] For each coordinate dimension, when the first coding cost is greater than or equal to any second coding cost of at least one second coding cost, the first grammatical information corresponding to the vertex in the coordinate dimension is determined to be a second value; the second value indicates that the encoding of the motion vector of the vertex in the coordinate dimension is not skipped.

[0234] When the sum of the first coding costs of the vertex in the three coordinate dimensions is greater than or equal to the sum of the three coordinate dimensions of each second coding cost in at least one second coding cost, that is, the tenth syntax information is the eighth value, and the encoding of the motion vector of the vertex is not skipped, for each coordinate dimension, when the first coding cost of the vertex in the coordinate dimension is less than or equal to each second coding cost in at least one second coding cost, the motion vector corresponding to the vertex in the coordinate dimension is set to a preset value. For each coordinate dimension, when the first coding cost of the vertex in the coordinate dimension is greater than or equal to any second coding cost in at least one second coding cost, the target prediction mode corresponding to the vertex in the coordinate dimension is determined based on the at least one second coding cost; and based on the target prediction mode corresponding to the vertex in the coordinate dimension, the third syntax information corresponding to the vertex in the coordinate dimension is determined.

[0235] In this embodiment, motion vector prediction is performed based on the target prediction mode corresponding to the vertex in the coordinate dimension to determine the target motion vector prediction value corresponding to the vertex in the coordinate dimension; the motion vector residual corresponding to the vertex in the coordinate dimension is determined based on the motion vector corresponding to the vertex in the coordinate dimension and the target motion vector prediction value; the motion vector residual corresponding to the vertex in the coordinate dimension is encoded to determine the motion vector encoding corresponding to the vertex in the coordinate dimension. Thus, encoding is achieved for each coordinate dimension.

[0236] In this embodiment, for the case of non-grouped coding, the vertex in the above process refers to each vertex in the base mesh. In this case, the tenth syntax information is applied to the base mesh and is used to indicate whether to skip coding for each vertex in the base mesh, that is, whether to skip coding for the three coordinate dimensions corresponding to each vertex in the base mesh. The first syntax information is applied to the base mesh, specifically, each coordinate dimension corresponding to the base mesh, and is used to indicate whether to skip coding for that coordinate dimension corresponding to each vertex in the base mesh.

[0237] In this embodiment, for group coding, the vertex in the above process refers to each vertex in the current group of at least one group in the base mesh. The first syntax information applies to the current group and indicates whether to skip coding for each vertex in the current group, that is, whether to skip coding for the three coordinate dimensions corresponding to each vertex in the current group. The first syntax information applies to the current group, specifically, each coordinate dimension corresponding to the current group, and indicates whether to skip coding for that coordinate dimension corresponding to each vertex in the current group.

[0238] It can be understood that the optimal coding mode corresponding to the vertex can be determined by comparing the first coding cost with at least one second coding cost of the motion vector of the vertex in each coordinate dimension, and the tenth grammatical information can be combined with the first grammatical information, or combined with the third grammatical information to instruct the encoder on how to decode the coordinate dimension, thereby improving the coding efficiency and further improving the coding performance.

[0239] Referring to Figure 15, it shows a schematic flow chart of a decoding method provided by an embodiment of the present application. As shown in Figure 15, the method may include:

[0240] S1501. Parse the code stream to determine first syntax information corresponding to vertices in the base mesh of the current image; the first syntax information indicates whether to skip decoding of motion vectors; the first syntax information acts on the vertices, or acts on at least one coordinate dimension corresponding to the vertices.

[0241] S1502: Determine motion vector decoding information corresponding to the vertex based on the first syntax information.

[0242] It should be noted that the decoding method of the embodiment of the present application is applied to a decoder. The decoding method of the embodiment of the present application may refer to an inter-frame decoding method, and more specifically, may be an inter-frame decoding method for a base grid in a dynamic grid. The decoding method can be applied to a decoder in a V-DMC, but is not limited thereto.

[0243] In this embodiment of the present application, the decoder parses the bitstream to determine first syntax information. The first syntax information applies to vertices and, for example, indicates whether to skip decoding the motion vector for the vertex in three coordinate dimensions. Alternatively, the first syntax information applies to each of at least one coordinate dimension corresponding to the vertex, indicating whether to skip decoding the motion vector for the vertex in that coordinate dimension.

[0244] In some embodiments, the decoder can directly parse the first syntax information; the decoder can also parse the code stream to determine the second syntax information corresponding to the base grid; when the second syntax information represents inter-frame decoding, the decoder continues to parse the code stream to determine the first syntax information corresponding to the vertices in the base grid of the current image. Exemplarily, when the second syntax information is 0, it indicates that the encoding mode of the base grid of the current image is inter-frame coding mode; when the second syntax information is 1, it indicates that the encoding mode of the base grid of the current image is intra-frame coding mode; when the second syntax information is 2, it indicates that the encoding mode of the base grid of the current image is skip coding mode; 3 is a reserved bit. The specific selection is based on actual conditions and is not limited in the embodiments of this application.

[0245] The following three cases are explained:

[0246] Case 1: In some embodiments, the decoder parses the bitstream to determine the first syntax information corresponding to each coordinate dimension of at least one coordinate dimension for a vertex in the base mesh; the first syntax information indicates whether to skip decoding of the motion vector corresponding to that coordinate dimension. That is, the first syntax information applies to each coordinate dimension of the at least one coordinate dimension, and each coordinate dimension corresponds to a piece of first syntax information indicating whether to skip decoding of the motion vector corresponding to that coordinate dimension.

[0247] For each coordinate dimension of at least one coordinate dimension, when the first syntax information corresponding to the coordinate dimension is a first value, the motion vector decoded value corresponding to the vertex in the coordinate dimension is set to a preset value; thereby determining the motion vector decoded value corresponding to the vertex in each coordinate dimension; and determining the motion vector decoded information corresponding to the vertex based on the motion vector decoded value corresponding to the vertex in each coordinate dimension. The first value indicates that motion vector decoding for the coordinate dimension is skipped.

[0248] For each coordinate dimension in at least one coordinate dimension, when the first syntax information corresponding to the vertex in the coordinate dimension is the second value, continue to parse the code stream to determine the third syntax information corresponding to the vertex in the coordinate dimension; the second value indicates that the motion vector decoding of the coordinate dimension is not skipped; determine the target prediction mode corresponding to the vertex in the coordinate dimension according to the third syntax information, and continue to parse the code stream to determine the motion vector residual decoding value corresponding to the vertex in the coordinate dimension; perform motion vector prediction according to the target prediction mode corresponding to the vertex in the coordinate dimension, and determine the target motion vector prediction value corresponding to the vertex in the coordinate dimension; determine the motion vector decoding value corresponding to the vertex in the coordinate dimension according to the target motion vector prediction value and the motion vector residual decoding value corresponding to the vertex in the coordinate dimension, thereby determining the motion vector decoding value corresponding to the vertex in each coordinate dimension; determine the motion vector decoding information corresponding to the vertex according to the motion vector decoding value corresponding to the vertex in each coordinate dimension.

[0249] Case 2: In some embodiments, the decoder parses the bitstream to determine the first syntax information corresponding to a vertex in the base mesh. This first syntax information indicates whether to skip decoding the motion vector for the vertex. Specifically, the first syntax information applies to three coordinate dimensions and indicates whether to skip decoding the motion vector for the vertex in those three coordinate dimensions.

[0250] In some embodiments, when the first syntax information is a fifth value, the motion vector decoding information of the vertex is determined to be a preset value.

[0251] In some embodiments, when the first syntax information is the sixth value, for each coordinate dimension in at least one coordinate dimension, continue to parse the code stream to determine the third syntax information corresponding to the vertex in the coordinate dimension; determine the target prediction mode corresponding to the vertex in the coordinate dimension based on the third syntax information corresponding to the vertex in the coordinate dimension, and continue to parse the code stream to determine the motion vector residual decoding value corresponding to the vertex in the coordinate dimension; perform motion vector prediction based on the target prediction mode corresponding to the vertex in the coordinate dimension to determine the target motion vector prediction value corresponding to the vertex in the coordinate dimension; determine the motion vector decoding value corresponding to the vertex in the coordinate dimension based on the target motion vector prediction value and the motion vector residual decoding value corresponding to the vertex in the coordinate dimension, thereby determining the motion vector decoding value corresponding to the vertex in each coordinate dimension; determine the motion vector decoding information corresponding to the vertex based on the motion vector decoding value corresponding to the vertex in each coordinate dimension.

[0252] Case 3: In some embodiments, the decoder parses the bitstream to determine the tenth syntax information corresponding to the vertex in the base mesh; the tenth syntax information indicates whether to skip decoding of the motion vector of the vertex;

[0253] When the tenth grammatical information indicates that decoding of the motion vector of the vertex is not skipped, the code stream continues to be parsed to determine the first grammatical information corresponding to the vertex in the basic mesh in each coordinate dimension; the first grammatical information indicates whether to skip decoding of the motion vector corresponding to the vertex in the basic mesh in the coordinate dimension.

[0254] That is, the tenth syntax information acts on the three coordinate dimensions corresponding to the vertex, indicating whether to skip decoding the motion vector for the three coordinate dimensions of the vertex. The first syntax information acts on each coordinate dimension corresponding to the vertex, indicating whether to skip decoding the motion vector corresponding to that coordinate dimension.

[0255] In some embodiments, when the tenth syntax information indicates that decoding of a motion vector of a vertex is skipped, a decoded motion vector value corresponding to the vertex is set to a preset value.

[0256] In some embodiments, for a third coordinate dimension in at least one coordinate dimension, when the tenth syntax information indicates that decoding of the motion vector of the vertex is not skipped, and the first syntax information corresponding to the vertex in the first coordinate dimension and the second coordinate dimension are both the first value, the first syntax information corresponding to the vertex in the third coordinate dimension is determined to be the second value. Exemplarily, the three coordinate dimensions correspond to dimension numbers (0, 1, 2), where 0 corresponds to the first coordinate dimension, 1 corresponds to the second coordinate dimension, and 2 corresponds to the third coordinate dimension. The decoder parses in the order of the dimension numbers.

[0257] That is to say, when the tenth grammatical information indicates that the decoding of the motion vector of the vertex is not skipped, and the first grammatical information of the two coordinate dimensions 0 and 1 parsed previously are both the first values, it is not necessary to continue parsing the first grammatical information of coordinate dimension 2, and the first grammatical information of coordinate dimension 2 can be directly determined as the second value.

[0258] In the above three cases, the process of the decoder performing motion vector prediction according to the target prediction mode corresponding to the vertex in the coordinate dimension and determining the target motion vector prediction value corresponding to the vertex in the coordinate dimension may include:

[0259] If the target prediction mode corresponding to the vertex in the coordinate dimension indicates no prediction mode, the target motion vector prediction value corresponding to the vertex in the coordinate dimension is set to a preset prediction value. The preset prediction value is the same as the preset prediction value in the encoder, and can be 0 or other values, depending on the actual situation and is not limited in this embodiment of the application.

[0260] Alternatively, when the target prediction mode corresponding to the vertex in the coordinate dimension represents no prediction mode, the encoder can also directly determine the motion vector residual decoding value corresponding to the vertex in the coordinate dimension as the motion vector decoding value corresponding to the vertex in the coordinate dimension.

[0261] When the target prediction mode corresponding to the vertex in the coordinate dimension represents the first prediction mode, the target motion vector prediction value corresponding to the vertex in the coordinate dimension is determined based on the first weighted average of the motion vector reconstruction values ​​corresponding to the first number of neighboring points of the vertex in the coordinate dimension. Here, the process of determining the target motion vector prediction value corresponding to the first prediction mode by the decoder is consistent with the same process in the encoder and is not repeated here.

[0262] In the case where the target prediction mode corresponding to the vertex in the coordinate dimension represents the second prediction mode, the first bias value is determined according to the first weighted value of the motion vector reconstruction value corresponding to the first number of neighboring points of the vertex in the coordinate dimension; based on the first weighted value and the first bias value, the target motion vector prediction value corresponding to the vertex in the coordinate dimension is determined. Wherein, when the first weighted value is greater than or equal to the preset threshold, the first number is shifted right by a preset number of bits to determine the first bias value. When the first weighted value is less than the preset threshold, the first number is shifted right by a preset number of bits and inverted to determine the first bias value. Here, the process of the decoder determining the target motion vector prediction value corresponding to the second prediction mode is consistent with the same process in the encoder, and will not be repeated here.

[0263] When the target prediction mode corresponding to the vertex in the coordinate dimension represents the third prediction mode, the target motion vector prediction value corresponding to the vertex in the coordinate dimension is determined based on the second weighted average of the motion vector reconstruction values ​​corresponding to the second number of neighboring points of the vertex in the coordinate dimension. Here, the process of determining the target motion vector prediction value corresponding to the third prediction mode by the decoder is consistent with the same process in the encoder and is not repeated here.

[0264] In the case where the target prediction mode corresponding to the vertex in the coordinate dimension represents the fourth prediction mode, the second bias value is determined according to the second weighted value of the motion vector reconstruction value corresponding to the second number of neighboring points of the vertex in the coordinate dimension; the motion vector prediction value corresponding to the vertex in the coordinate dimension is determined according to the second weighted value and the second bias value. Wherein, when the second weighted value is greater than or equal to the preset threshold, the second number is shifted right by a preset number of bits to determine the second bias value. When the second weighted value is less than the preset threshold, the second number is shifted right by a preset number of bits and inverted to determine the second bias value. Here, the process of the decoder determining the target motion vector prediction value corresponding to the fourth prediction mode is consistent with the same process in the encoder, and will not be repeated here.

[0265] In the above process, the first number of neighbor points or the second number of neighbor points includes one or more of the following:

[0266] In the current image, the vertices that are connected to the vertex and whose encoding order is before the vertex;

[0267] and / or, a vertex with the same index as the vertex in the reference image corresponding to the current image;

[0268] and / or, all vertices that are connected to the vertex's co-location in the reference image;

[0269] And / or, in the reference image, the vertices with the same index as their neighbor points in the current image.

[0270] In the above process, the process of continuing to parse the code stream and determining the motion vector residual decoding value corresponding to the coordinate dimension includes:

[0271] By continuing to parse the bitstream, determining the eighth syntax information and the ninth syntax information corresponding to the coordinate dimension;

[0272] When the eighth syntax information is the third value and the ninth syntax information is the fourth value, the code stream continues to be parsed to determine the motion vector residual decoding value corresponding to the coordinate dimension.

[0273] When the eighth syntax information is the third value and the ninth syntax information is the fourth value, continue to parse the code stream to determine the remainder of the motion vector residual corresponding to the coordinate dimension; determine the motion vector residual decoding value based on the remainder of the motion vector residual.

[0274] The motion vector residual decoding value can be a positive number or a negative number.

[0275] In some embodiments, the decoder may perform decoding in a grouped decoding mode or in a non-grouped decoding mode. The decoder may use grouped decoding as a default decoding mode or use non-grouped decoding as a default decoding mode. The decoder may also determine whether to perform grouped decoding by parsing the fourth syntax information and / or the fifth syntax information.

[0276] In some embodiments, the fourth syntax information indicates whether the base meshes of all images in the sequence are to be grouped and encoded; and the fifth syntax information indicates whether the base mesh of the current image is to be grouped and encoded. When the fourth syntax information and / or the fifth syntax information indicates that the base mesh is to be grouped and decoded, the grouping is performed based on the index corresponding to each vertex in the base mesh, and at least one group is determined.

[0277] In some embodiments, when the fourth syntax information and / or the fifth syntax information indicate group decoding, the decoder may further determine sixth syntax information and / or seventh syntax information by parsing the bitstream; the sixth syntax information and / or the seventh syntax information represent the number of vertices in a group. The decoder groups the vertices in the base mesh based on the index corresponding to each vertex in the base mesh and the number of vertices in the group, and determines at least one group.

[0278] In some embodiments, when the fourth syntax information and / or the fifth syntax information indicates that the basic grid is not to be grouped and decoded, the first syntax information corresponding to each coordinate dimension of the basic grid in at least one coordinate dimension is determined by parsing the code stream.

[0279] For the grouped decoding method, the vertices in the base mesh include: each vertex in the current group of at least one group of the base mesh. For the non-grouped decoding method, the vertices in the base mesh include: each vertex in the base mesh.

[0280] In some embodiments, for the first case above, for the group decoding mode, determining the first syntax information corresponding to each coordinate dimension of the vertices in the base mesh in at least one coordinate dimension by parsing the bitstream includes:

[0281] Determine, by parsing the bitstream, first syntax information corresponding to each coordinate dimension of the current group in at least one coordinate dimension;

[0282] For each coordinate dimension, the first grammatical information corresponding to the coordinate dimension of the current group is determined as the first grammatical information corresponding to the coordinate dimension of each vertex of the current group, thereby determining the first grammatical information corresponding to each coordinate dimension of each vertex of the current group in at least one coordinate dimension.

[0283] The above process continues to parse the code stream to determine the third syntax information corresponding to the vertex in the coordinate dimension, including:

[0284] By continuing to parse the code stream, determining the third syntax information corresponding to the coordinate dimension of the current group;

[0285] The third syntax information corresponding to the coordinate dimension of the current group is determined as the third syntax information corresponding to the coordinate dimension of each vertex of the current group.

[0286] For example, the group coding syntax of the first case above can be expressed as:

[0287] Or, expressed as:

[0288] For the first case above, for the non-grouped decoding method, the first syntax information corresponding to the vertices in the basic mesh in each coordinate dimension is determined by parsing the bitstream, including:

[0289] Determine, by parsing the bitstream, first syntax information corresponding to each coordinate dimension of the basic grid in at least one coordinate dimension;

[0290] For each coordinate dimension, the first grammatical information corresponding to the basic grid in the coordinate dimension is determined as the first grammatical information corresponding to the coordinate dimension of each vertex in the basic grid, thereby determining the first grammatical information corresponding to each coordinate dimension of each vertex in the basic grid in at least one coordinate dimension.

[0291] The above process continues to parse the code stream to determine the third syntax information corresponding to the vertex in the coordinate dimension, including:

[0292] By continuing to parse the code stream, determining the third syntax information corresponding to the basic grid in the coordinate dimension;

[0293] The third syntax information corresponding to the coordinate dimension of the basic mesh is determined as the third syntax information corresponding to the coordinate dimension of each vertex in the basic mesh.

[0294] In some embodiments, for the second case above, for the group decoding method, the above-mentioned determination of the first syntax information corresponding to each coordinate dimension of the vertex in the basic mesh by parsing the bitstream includes:

[0295] Determine the first syntax information corresponding to the current group by parsing the code stream;

[0296] The first syntax information corresponding to the current group is determined as the first syntax information corresponding to each vertex of the current group. That is, whether to skip decoding of the three coordinate dimensions corresponding to each vertex of the current group is determined based on the first syntax information corresponding to the current group.

[0297] The above process continues to parse the code stream to determine the third syntax information corresponding to the vertex in the coordinate dimension, including:

[0298] By continuing to parse the code stream, determining the third syntax information corresponding to the coordinate dimension of the current group;

[0299] The third syntax information corresponding to the coordinate dimension of the current group is determined as the third syntax information corresponding to the coordinate dimension of each vertex of the current group.

[0300] For example, the group coding syntax of the second case above can be expressed as:

[0301] Or, expressed as:

[0302] For the second case above, for the non-group decoding mode, the first syntax information corresponding to the vertices in the basic mesh in each coordinate dimension is determined by parsing the bitstream, including:

[0303] Determine the first syntax information corresponding to the basic grid by parsing the code stream;

[0304] The first syntax information corresponding to the base mesh is determined as the first syntax information corresponding to each vertex in the base mesh, that is, whether to skip decoding of the three coordinate dimensions corresponding to each vertex of the base mesh is determined based on the first syntax information corresponding to the base mesh.

[0305] The above process continues to parse the code stream to determine the third syntax information corresponding to the vertex in the coordinate dimension, including:

[0306] By continuing to parse the code stream, determining the third syntax information corresponding to the basic grid in the coordinate dimension;

[0307] The third syntax information corresponding to the coordinate dimension of the basic mesh is determined as the third syntax information corresponding to the coordinate dimension of each vertex in the basic mesh.

[0308] In some embodiments, for the third case, in a group decoding mode, parsing the code stream to determine the tenth syntax information corresponding to the vertices in the base mesh includes:

[0309] Determining tenth grammatical information corresponding to the current group by parsing the bitstream;

[0310] The tenth syntax information corresponding to the current group is determined as the tenth syntax information corresponding to each vertex of the current group.

[0311] The above-mentioned method of determining the first syntax information corresponding to each coordinate dimension of the vertices in the basic mesh by parsing the bitstream includes:

[0312] Determine, by parsing the bitstream, first syntax information corresponding to each coordinate dimension of the current group in at least one coordinate dimension;

[0313] For each coordinate dimension, the first grammatical information corresponding to the coordinate dimension of the current group is determined as the first grammatical information corresponding to the coordinate dimension of each vertex of the current group, thereby determining the first grammatical information corresponding to each coordinate dimension of each vertex of the current group in at least one coordinate dimension.

[0314] The above process continues to parse the code stream to determine the third syntax information corresponding to the vertex in the coordinate dimension, including:

[0315] By continuing to parse the code stream, determining the third syntax information corresponding to the coordinate dimension of the current group;

[0316] The third syntax information corresponding to the coordinate dimension of the current group is determined as the third syntax information corresponding to the coordinate dimension of each vertex of the current group.

[0317] For example, the group coding syntax of the third case above can be expressed as:

[0318] Or, expressed as:

[0319] For the third case above, for the non-grouped decoding mode, the above parsing of the code stream to determine the tenth syntax information corresponding to the vertices in the basic mesh includes:

[0320] Determine tenth grammatical information corresponding to the basic grid by parsing the bitstream;

[0321] The tenth syntax information corresponding to the base mesh is determined as the tenth syntax information corresponding to each vertex in the base mesh.

[0322] The above-mentioned method of determining the first syntax information corresponding to each coordinate dimension of the vertices in the basic mesh by parsing the bitstream includes:

[0323] Determine, by parsing the bitstream, first syntax information corresponding to each coordinate dimension of the basic grid in at least one coordinate dimension;

[0324] For each coordinate dimension, the first grammatical information corresponding to the basic grid in the coordinate dimension is determined as the first grammatical information corresponding to the coordinate dimension of each vertex in the basic grid, thereby determining the first grammatical information corresponding to each coordinate dimension of each vertex in the basic grid in at least one coordinate dimension.

[0325] The above process continues to parse the code stream to determine the third syntax information corresponding to the vertex in the coordinate dimension, including:

[0326] By continuing to parse the code stream, determining the third syntax information corresponding to the basic grid in the coordinate dimension;

[0327] The third syntax information corresponding to the coordinate dimension of the basic mesh is determined as the third syntax information corresponding to the coordinate dimension of each vertex in the basic mesh.

[0328] It can be understood that at the decoder, the bitstream is parsed to determine the first syntax information corresponding to the vertices in the base mesh of the current image. This first syntax information indicates whether to skip decoding of motion vectors. The first syntax information is applied to the vertex, or to each coordinate dimension corresponding to the vertex. Based on the first syntax information, the motion vector decoding information corresponding to the vertex is determined. In this way, the optimal processing method for the vertex can be determined based on the first syntax information, thereby improving decoding efficiency and, in turn, decoding performance.

[0329] It should be noted that for the first and third cases above, the order of parsing the content from the code stream must meet the following principles. In addition, there is no limit on the parsed content to be added before, between, or after each content to be parsed:

[0330] a. The prediction mode syntax element of the first coordinate dimension must be parsed after the first syntax information of the first coordinate dimension;

[0331] b. The prediction mode syntax element of the second coordinate dimension must be parsed after the first syntax information of the second coordinate dimension;

[0332] c. The prediction mode syntax element of the third coordinate dimension must be parsed after the first syntax information of the third coordinate dimension;

[0333] d. The decoded motion vector residual values ​​of the first coordinate dimension of all vertices in the group must be parsed after the prediction mode syntax element of the first coordinate dimension;

[0334] e. The decoded motion vector residual values ​​of the second coordinate dimension of all vertices in the group must be parsed after the prediction mode syntax element of the second coordinate dimension;

[0335] f. The decoded motion vector residual values ​​of the third coordinate dimension of all vertices in the group must be parsed after the prediction mode syntax element of the third coordinate dimension;

[0336] An example of an implementation method that meets the above-mentioned limiting conditions is as follows:

[0337] a. Parse the first syntax information of the first coordinate dimension and execute step b;

[0338] b. If the first syntax information of the first coordinate dimension indicates not to skip, first parse the prediction mode syntax element of the first coordinate dimension, then parse the first syntax information of the second coordinate dimension, and execute step c;

[0339] If the first syntax information of the first coordinate dimension indicates skipping, the first syntax information of the second coordinate dimension is parsed and step c is executed;

[0340] c. If the first syntax information of the second coordinate dimension indicates not to skip, first parse the prediction mode syntax element of the second coordinate dimension, then parse the first syntax information of the third coordinate dimension, and execute step d;

[0341] If the first syntax information of the second coordinate dimension indicates skipping, the first syntax information of the third coordinate dimension is parsed and step d is executed;

[0342] d. If the first syntax information of the third coordinate dimension indicates not to skip, parse the prediction mode syntax element of the third coordinate dimension and execute step e;

[0343] If the first syntax information of the third coordinate dimension indicates skipping, executing step e;

[0344] e. If the first syntax information of the first coordinate dimension indicates not to skip, decode the residual decoded value of the motion vector of the first coordinate dimension of vertex 0 in the decoding group; if the first syntax information of the second coordinate dimension indicates not to skip, decode the residual decoded value of the motion vector of the second coordinate dimension of vertex 0 in the decoding group; if the first syntax information of the third coordinate dimension indicates not to skip, decode the residual decoded value of the motion vector of the third coordinate dimension of vertex 0 in the decoding group; if the first syntax information of the first coordinate dimension indicates not to skip, decode the residual decoded value of the motion vector of the first coordinate dimension of vertex 1 in the decoding group; if the first syntax information of the second coordinate dimension indicates not to skip, decode the residual decoded value of the motion vector of the second coordinate dimension of vertex 1 in the decoding group; if the first syntax information of the third coordinate dimension indicates not to skip, decode the residual decoded value of the motion vector of the third coordinate dimension of vertex 1 in the decoding group; ..., if the first The first syntax information of the coordinate dimension indicates not skipping, and the decoded value of the motion vector residual of the first coordinate dimension of vertex i-2 in the decoding group is decoded. If the first syntax information of the second coordinate dimension indicates not skipping, the decoded value of the motion vector residual of the second coordinate dimension of vertex i-2 in the decoding group is decoded. If the first syntax information of the third coordinate dimension indicates not skipping, the decoded value of the motion vector residual of the third coordinate dimension of vertex i-2 in the decoding group is decoded. If the first syntax information of the first coordinate dimension indicates not skipping, the decoded value of the motion vector residual of the first coordinate dimension of vertex i-1 in the decoding group is decoded. If the first syntax information of the second coordinate dimension indicates not skipping, the decoded value of the motion vector residual of the second coordinate dimension of vertex i-1 in the decoding group is decoded. If the first syntax information of the third coordinate dimension indicates not skipping, the decoded value of the motion vector residual of the third coordinate dimension of vertex i-1 in the decoding group is decoded.

[0345] Another implementation example is as follows:

[0346] a. Parse the first syntax information of the first coordinate dimension, parse the first syntax information of the second coordinate dimension, parse the first syntax information of the third coordinate dimension, and execute step b;

[0347] b. If the first syntax information of the first coordinate dimension indicates not to skip, then parse the prediction mode syntax element of the first coordinate dimension; if the first syntax information of the second coordinate dimension indicates not to skip, then parse the prediction mode syntax element of the second coordinate dimension; if the first syntax information of the third coordinate dimension indicates not to skip, then parse the prediction mode syntax element of the third coordinate dimension, and execute step c;

[0348] c. If the first syntax information of the first coordinate dimension indicates not to skip, decode the residual decoded value of the motion vector of the first coordinate dimension of vertex 0 in the decoding group; if the first syntax information of the second coordinate dimension indicates not to skip, decode the residual decoded value of the motion vector of the second coordinate dimension of vertex 0 in the decoding group; if the first syntax information of the third coordinate dimension indicates not to skip, decode the residual decoded value of the motion vector of the third coordinate dimension of vertex 0 in the decoding group; if the first syntax information of the first coordinate dimension indicates not to skip, decode the residual decoded value of the motion vector of the first coordinate dimension of vertex 1 in the decoding group; if the first syntax information of the second coordinate dimension indicates not to skip, decode the residual decoded value of the motion vector of the second coordinate dimension of vertex 1 in the decoding group; if the first syntax information of the third coordinate dimension indicates not to skip, decode the residual decoded value of the motion vector of the third coordinate dimension of vertex 1 in the decoding group; ..., if the first The first syntax information of the coordinate dimension indicates not skipping, and the decoded value of the motion vector residual of the first coordinate dimension of vertex i-2 in the decoding group is decoded. If the first syntax information of the second coordinate dimension indicates not skipping, the decoded value of the motion vector residual of the second coordinate dimension of vertex i-2 in the decoding group is decoded. If the first syntax information of the third coordinate dimension indicates not skipping, the decoded value of the motion vector residual of the third coordinate dimension of vertex i-2 in the decoding group is decoded. If the first syntax information of the first coordinate dimension indicates not skipping, the decoded value of the motion vector residual of the first coordinate dimension of vertex i-1 in the decoding group is decoded. If the first syntax information of the second coordinate dimension indicates not skipping, the decoded value of the motion vector residual of the second coordinate dimension of vertex i-1 in the decoding group is decoded. If the first syntax information of the third coordinate dimension indicates not skipping, the decoded value of the motion vector residual of the third coordinate dimension of vertex i-1 in the decoding group is decoded.

[0349] For the first case above, the bit rate saving effect on the inter-frame basic grid can be shown in Table 1, as follows:

[0350] Table 1

[0351] For the second case above, the bit rate saving effect on the inter-frame basic grid can be shown in Table 2, as follows:

[0352] Table 2

[0353] For the second case above, the bit rate saving effect on the inter-frame basic grid can be shown in Table 3, as follows:

[0354] Table 3

[0355] It can be seen that the embodiments of the present application improve the encoding and decoding efficiency, thereby improving the encoding and decoding performance.

[0356] In yet another embodiment of the present application, based on the same inventive concept as the aforementioned embodiment, FIG16 is a schematic diagram illustrating the structure of an encoder provided by an embodiment of the present application. As shown in FIG16 , the encoder 190 may include a motion vector determination unit 1901, a coding cost determination unit 1902, and a syntax information determination unit 1903, wherein:

[0357] The motion vector determining section 1901 is configured to determine a base mesh of a current image and a motion vector corresponding to each coordinate dimension of at least one coordinate dimension of a vertex in the base mesh;

[0358] The coding cost determining portion 1902 is configured to, when performing inter-frame coding on the base mesh, estimate coding costs of skip coding and at least one prediction mode for the motion vector corresponding to each coordinate dimension of the vertex, and determine a first coding cost and at least one second coding cost corresponding to each coordinate dimension; the first coding cost corresponds to the skip coding mode; and the at least one second coding cost corresponds to the at least one prediction mode.

[0359] The syntax information determination part 1903 is configured to determine first syntax information according to the first coding cost corresponding to each coordinate dimension and at least one second coding cost; the first syntax information represents whether to skip the encoding of the motion vector.

[0360] In some embodiments, the coding cost determination part 1902 is further configured to, for each coordinate dimension, determine that the first grammatical information corresponding to the vertex in the coordinate dimension is a first value when the first coding cost is less than or equal to each second coding cost of the at least one second coding cost; the first value represents skipping the encoding of the motion vector of the vertex in the coordinate dimension.

[0361] In some embodiments, the grammatical information determination part 1903 is further configured to, for each coordinate dimension, determine that the first grammatical information corresponding to the vertex in the coordinate dimension is a second value when the first coding cost is greater than or equal to any second coding cost of the at least one second coding cost; the second value indicates that the encoding of the motion vector of the vertex in the coordinate dimension is not skipped.

[0362] In some embodiments, the encoder 190 also includes an encoding part; the encoding part is configured to set the motion vector corresponding to the vertex in the coordinate dimension to a preset value for each coordinate dimension when the first encoding cost is less than or equal to each second encoding cost of the at least one second encoding cost.

[0363] In some embodiments, the grammatical information determination part 1903 is further configured to, for each coordinate dimension, determine the target prediction mode corresponding to the vertex in each coordinate dimension according to the at least one second coding cost when the first coding cost is greater than or equal to any second coding cost of the at least one second coding cost; and determine the third grammatical information corresponding to the vertex in each coordinate dimension according to the target prediction mode corresponding to the vertex in each coordinate dimension.

[0364] In some embodiments, the encoding part is further configured to perform motion vector prediction based on the target prediction mode corresponding to the vertex in each coordinate dimension, and determine the target motion vector prediction value corresponding to the vertex in each coordinate dimension; determine the motion vector residual corresponding to the vertex in each coordinate dimension based on the motion vector corresponding to the vertex in each coordinate dimension and the target motion vector prediction value; encode the motion vector residual corresponding to the vertex in each coordinate dimension, and determine the motion vector encoding corresponding to the vertex in each coordinate dimension.

[0365] In some embodiments, the syntax information determination part 1903 is further configured to determine that the first syntax information is a fifth value when the sum of the first coding costs of the vertex in the three coordinate dimensions is less than or equal to the sum of the three coordinate dimensions of each second coding cost in the at least one second coding cost; the fifth value represents skipping the encoding of the motion vector of the vertex.

[0366] In some embodiments, the encoding part is further configured to set the motion vector corresponding to the vertex to a preset value when the sum of the first encoding costs of the vertex in the three coordinate dimensions is less than or equal to the sum of the three coordinate dimensions of each second encoding cost in the at least one second encoding cost.

[0367] In some embodiments, the syntax information determination part 1903 is further configured to determine that the first syntax information is a sixth value when the sum of the first coding costs of the vertex in the three coordinate dimensions is greater than or equal to the sum of the three coordinate dimensions of each second coding cost in the at least one second coding cost; the sixth value indicates that the encoding of the motion vector of the vertex is not skipped.

[0368] In some embodiments, the grammatical information determination part 1903 is further configured to determine the target prediction mode corresponding to the vertex in each coordinate dimension according to at least one second coding cost corresponding to the vertex in each coordinate dimension when the sum of the first coding costs of the vertex in the three coordinate dimensions is greater than or equal to the sum of the three coordinate dimensions of each second coding cost of the at least one second coding cost; and determine the third grammatical information corresponding to the vertex in each coordinate dimension according to the target prediction mode corresponding to the vertex in each coordinate dimension.

[0369] In some embodiments, the grammatical information determination part 1903 is further configured to determine the tenth grammatical information based on the first coding cost and at least one second coding cost corresponding to each coordinate dimension; the tenth grammatical information represents whether to skip encoding the motion vector of the vertex; in the case where the tenth grammatical information represents not skipping encoding the motion vector of the vertex, the first grammatical information corresponding to the vertex in each coordinate dimension is determined based on the first coding cost and at least one second coding cost corresponding to each coordinate dimension; the first grammatical information represents whether to skip encoding the motion vector corresponding to the vertex in the coordinate dimension.

[0370] In some embodiments, the grammatical information determination part 1903 is further configured to determine that the tenth grammatical information is the seventh value when the sum of the first coding costs of the vertex in the three coordinate dimensions is less than or equal to the sum of the three coordinate dimensions of each second coding cost in the at least one second coding cost; the seventh value represents skipping the encoding of the motion vector of the vertex.

[0371] In some embodiments, the encoding part is further configured to set the motion vector encoding corresponding to the vertex to a preset value when the sum of the first encoding costs of the vertex in the three coordinate dimensions is less than or equal to the sum of the three coordinate dimensions of each second encoding cost in the at least one second encoding cost.

[0372] In some embodiments, the grammatical information determination part 1903 is further configured to determine that the tenth grammatical information is an eighth value when the sum of the first coding costs of the vertex in the three coordinate dimensions is greater than or equal to the sum of the three coordinate dimensions of each second coding cost in the at least one second coding cost; the eighth value indicates that the encoding of the motion vector of the vertex is not skipped.

[0373] In some embodiments, the at least one second coding cost corresponds to at least one prediction mode and at least one of the no-prediction mode; the coding cost determination part 1902 is further configured to determine, for each coordinate dimension, the first motion vector prediction value corresponding to the vertex in the coordinate dimension through the no-prediction mode; the no-prediction mode characterization sets the motion vector prediction value corresponding to the coordinate dimension to a preset vector prediction value; based on the first motion vector prediction value, the coding cost is estimated to determine the second coding cost corresponding to the no-prediction mode; and / or, based on the first weighted average of the motion vector reconstruction values ​​corresponding to the coordinate dimension of the first number of neighboring points of the vertex, the second motion vector prediction value is determined; based on the second motion vector prediction value, the coding cost is estimated to determine the second coding cost corresponding to the first prediction mode in the at least one prediction mode; and / or, based on the first weighted average of the motion vector reconstruction values ​​corresponding to the coordinate dimension of the first number of neighboring points of the vertex, the second motion vector prediction value is determined. value, determine a first bias value; determine a third motion vector prediction value based on the first weighted value and the first bias value; perform coding cost estimation based on the third motion vector prediction value, and determine the second coding cost corresponding to the second prediction mode in the at least one prediction mode; and / or, determine a fourth motion vector prediction value based on the second weighted average value of the motion vector reconstruction values ​​corresponding to the second number of neighboring points of the vertex in the coordinate dimension; perform coding cost estimation based on the fourth motion vector prediction value, and determine the second coding cost corresponding to the third prediction mode in the at least one prediction mode; and / or, determine a second bias value based on the second weighted value of the motion vector reconstruction values ​​corresponding to the second number of neighboring points of the vertex in the coordinate dimension; determine a fifth motion vector prediction value based on the second weighted value and the second bias value; perform coding cost estimation based on the fifth motion vector prediction value, and determine the second coding cost corresponding to the fourth prediction mode in the at least one prediction mode.

[0374] In some embodiments, the coding cost determination portion 1902 is further configured to, when the first weighted average value is greater than or equal to a preset threshold, right-shift the first number by a preset number of bits to determine the first offset value.

[0375] In some embodiments, the coding cost determining portion 1902 is further configured to, when the first weighted value is less than a preset threshold, right-shift the first number by a preset number of bits and invert it to determine the first bias value.

[0376] In some embodiments, the coding cost determining portion 1902 is further configured to, when the second weighted value is greater than or equal to a preset threshold, right-shift the second number by a preset number of bits to determine the second offset value.

[0377] In some embodiments, the coding cost determination part 1902 is further configured to, when the second weighted average value is less than a preset threshold, right-shift the second number by a preset number of bits and invert it to determine the second bias value.

[0378] In some embodiments, the encoding part is further configured to encode the first syntax information corresponding to the vertex and write the obtained encoding bits into the bitstream; or, encode the first syntax information corresponding to the vertex and the third syntax information corresponding to the vertex in each coordinate dimension, and write the obtained encoding bits into the bitstream.

[0379] In some embodiments, the encoding part is further configured to, for each coordinate dimension, determine the eighth grammatical information corresponding to the vertex in the coordinate dimension as a third value, and determine the ninth grammatical information corresponding to the vertex in the coordinate dimension as a fourth value, when the absolute value of the motion vector residual corresponding to the vertex in the coordinate dimension is greater than the first preset residual threshold and less than the second preset residual threshold; the second preset residual threshold is greater than the first preset residual threshold; encode the motion vector residual corresponding to the vertex in the coordinate dimension, determine the motion vector code corresponding to the vertex in the coordinate dimension, and thus determine the motion vector code corresponding to the vertex in each coordinate dimension.

[0380] In some embodiments, the encoding part is further configured to, for each coordinate dimension, determine the eighth grammatical information corresponding to the vertex in the coordinate dimension as a third value, and determine the ninth grammatical information corresponding to the vertex in the coordinate dimension as the third value when the absolute value of the motion vector residual corresponding to the vertex in the coordinate dimension is greater than or equal to the second preset residual threshold; encode the remainder of the motion vector residual corresponding to the vertex in the coordinate dimension, determine the motion vector code corresponding to the vertex in the coordinate dimension, and thus determine the motion vector code corresponding to the vertex in each coordinate dimension.

[0381] In some embodiments, the encoding part is further configured to perform entropy encoding on the eighth grammatical information and the ninth grammatical information corresponding to each coordinate dimension of the vertex, and write the obtained encoding bits and the motion vector encoding corresponding to each coordinate dimension of the vertex into the bitstream.

[0382] In some embodiments, the encoding part is further configured to determine second syntax information; the second syntax information represents intra-frame encoding or inter-frame encoding of the current image; entropy encoding is performed on the second syntax information, and the obtained encoding bits are written into the bitstream.

[0383] In some embodiments, the vertices in the basic grid include: each vertex in the current group in at least one group of the basic grid; the coding cost determination part 1902 is also configured to perform skip coding coding estimation and at least one prediction mode coding estimation on the motion vector corresponding to each vertex of the current group in the coordinate dimension for each coordinate dimension, and determine the first coding cost and at least one second coding cost corresponding to each vertex of the current group in the coordinate dimension; determine the first coding cost and at least one second coding cost corresponding to the coordinate dimension of each vertex of the current group based on the first coding cost and at least one second coding cost corresponding to the coordinate dimension of each vertex of the current group, thereby determining the first coding cost and at least one second coding cost corresponding to the current group in each coordinate dimension.

[0384] In some embodiments, the vertices in the basic mesh include: each vertex in the basic mesh; the coding cost determination part 1902 is further configured to perform skip coding and coding cost estimation of at least one prediction mode for the motion vector corresponding to each vertex in the basic mesh in the coordinate dimension for each coordinate dimension, and determine the first coding cost and at least one second coding cost corresponding to each vertex in the basic mesh in the coordinate dimension; determine the first coding cost and at least one second coding cost corresponding to the basic mesh in the coordinate dimension based on the first coding cost and at least one second coding cost corresponding to each vertex in the basic mesh in the coordinate dimension, thereby determining the first coding cost and at least one second coding cost corresponding to the basic mesh in each coordinate dimension.

[0385] In some embodiments, the encoder also includes a first grouping part; the first grouping part is configured to determine fourth grammatical information and / or fifth grammatical information; when the fourth grammatical information and / or the fifth grammatical information represent group encoding of the basic mesh, grouping is performed based on the index corresponding to each vertex in the basic mesh and the number of vertices in the group to determine the at least one group.

[0386] In some embodiments, the encoding part is further configured to determine sixth grammatical information and / or seventh grammatical information based on the number of vertices in the group; perform entropy encoding on the sixth grammatical information and / or the seventh grammatical information, and write the obtained coded bits into the bitstream.

[0387] In some embodiments, the first number of neighbor points or the second number of neighbor points include one or more of the following:

[0388] In the current image, the vertices that are connected to the vertex and whose coding order is before the vertex;

[0389] and / or, a vertex with the same index as the vertex in a reference image corresponding to the current image;

[0390] and / or, in the reference image, all vertices that are connected to the same point as the vertex;

[0391] and / or, in the reference image, a vertex having the same neighbor point index as the vertex in the current image.

[0392] It should be noted that the description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of the present invention, please refer to the description of the method embodiment of the present invention for understanding.

[0393] It is understood that in the embodiments of the present application, a "portion" may be a circuit portion, a processor portion, a program or software portion, and so forth. It may also be a module or non-modular. Furthermore, the various components in this embodiment may be integrated into a single processing unit, each unit may exist physically as a separate unit, or two or more units may be integrated into a single unit. These integrated units may be implemented in either hardware or software functional modules.

[0394] If the integrated unit is implemented as a software functional module and not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or grid device, etc.) or a processor to perform all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0395] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the encoder 190. The computer-readable storage medium stores a computer program, and when the computer program is executed by the first processor, it implements the method described in any one of the aforementioned embodiments.

[0396] Based on the composition of the encoder 190 and the computer-readable storage medium, refer to Figure 17, which shows a specific hardware structure diagram of the encoder 190 provided in an embodiment of the present application. As shown in Figure 20, the encoder 190 may include: a first communication interface 2301, a first memory 2302 and a first processor 2303; each component is coupled together through a first bus system 2304. It can be understood that the first bus system 2304 is used to achieve connection and communication between these components. In addition to the data bus, the first bus system 2304 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as the first bus system 2304 in Figure 20. Among them,

[0397] The first communication interface 2301 is used to receive and send signals when sending and receiving information with other external network elements;

[0398] A first memory 2302 is used to store computer programs that can be run on the first processor 2303;

[0399] The first processor 2303 is configured to execute the encoding method applied to the encoder in the embodiment of the present application when running the computer program.

[0400] It is understood that the first memory 2302 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The first memory 2302 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0401] The first processor 2303 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the first processor 2303. The above-mentioned first processor 2303 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented as a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the first memory 2302 , and the first processor 2303 reads the information in the first memory 2302 and completes the steps of the above method in combination with its hardware.

[0402] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0403] Optionally, as another embodiment, the first processor 2303 is further configured to execute the encoding method applied to an encoder described in any one of the aforementioned embodiments when running the computer program.

[0404] Based on the same inventive concept as the above embodiment, referring to FIG18 , a schematic diagram of the structure of a decoder provided by an embodiment of the present application is shown. As shown in FIG18 , the decoder 240 may include a parsing part 2401 and a decoding part 2402 , wherein:

[0405] The parsing portion 2401 is configured to parse the bitstream and determine first syntax information corresponding to a vertex in a base mesh of a current image; the first syntax information indicates whether to skip decoding of a motion vector; the first syntax information is applied to the vertex or to at least one coordinate dimension corresponding to the vertex;

[0406] The decoding part 2402 is configured to determine the motion vector decoding information corresponding to the vertex based on the first syntax information.

[0407] In some embodiments, parsing the code stream to determine first syntax information corresponding to vertices in a base mesh of the current image includes:

[0408] Parsing the code stream to determine second syntax information corresponding to the basic grid;

[0409] In some embodiments, the parsing part 2401 is further configured to, when the second syntax information indicates inter-frame decoding, continue parsing the code stream to determine the first syntax information corresponding to the vertices in the base mesh of the current image.

[0410] In some embodiments, the parsing part 2401 is further configured to determine the first syntax information corresponding to each coordinate dimension of the vertices in the basic mesh in at least one coordinate dimension by parsing the code stream; the first syntax information represents whether to skip decoding of the motion vector corresponding to the coordinate dimension.

[0411] In some embodiments, the parsing part 2401 is further configured to determine first syntax information corresponding to the vertex in the basic mesh by parsing the code stream; the first syntax information indicates whether to skip decoding of the motion vector of the vertex.

[0412] In some embodiments, the parsing part 2401 is further configured to parse the code stream to determine the tenth syntax information corresponding to the vertex in the basic mesh; the tenth syntax information indicates whether to skip decoding the motion vector of the vertex; when the tenth syntax information indicates that the decoding of the motion vector of the vertex is not skipped, continue to parse the code stream to determine the first syntax information corresponding to each coordinate dimension of the vertex in the basic mesh in at least one coordinate dimension; the first syntax information indicates whether to skip decoding the motion vector corresponding to the vertex in the basic mesh in the coordinate dimension.

[0413] In some embodiments, the decoding part 2402 is further configured to set the decoded value of the motion vector corresponding to the vertex to a preset value when the tenth syntax information indicates that decoding of the motion vector of the vertex is skipped.

[0414] In some embodiments, the decoding part 2402 is further configured to set the motion vector decoding value corresponding to the vertex in each coordinate dimension to a preset value when the first syntax information corresponding to the coordinate dimension is a first value for each coordinate dimension; thereby determining the motion vector decoding value corresponding to the vertex in each coordinate dimension; and determining the motion vector decoding information corresponding to the vertex based on the motion vector decoding value corresponding to the vertex in each coordinate dimension.

[0415] In some embodiments, the decoding part 2402 is further configured to, for each coordinate dimension, continue to parse the code stream when the first syntax information corresponding to the vertex in the coordinate dimension is the second value, and determine the third syntax information corresponding to the vertex in the coordinate dimension; determine the target prediction mode corresponding to the vertex in the coordinate dimension according to the third syntax information, and continue to parse the code stream to determine the motion vector residual decoding value corresponding to the vertex in the coordinate dimension; perform motion vector prediction according to the target prediction mode corresponding to the vertex in the coordinate dimension, and determine the target motion vector prediction value corresponding to the vertex in the coordinate dimension; determine the motion vector decoding value corresponding to the vertex in the coordinate dimension according to the target motion vector prediction value and the motion vector residual decoding value corresponding to the vertex in the coordinate dimension, thereby determining the motion vector decoding value corresponding to the vertex in each coordinate dimension; determine the motion vector decoding information corresponding to the vertex according to the motion vector decoding value corresponding to the vertex in each coordinate dimension.

[0416] In some embodiments, the parsing part 2402 is further configured to determine that, for the third coordinate dimension in each coordinate dimension, the first grammatical information corresponding to the vertex in the third coordinate dimension is the second value when the tenth grammatical information represents that decoding of the motion vector of the vertex is not skipped and the first grammatical information corresponding to the first coordinate dimension and the second coordinate dimension of the vertex are both the first value.

[0417] In some embodiments, the decoding part 2402 is further configured to determine the motion vector decoding information of the vertex as a preset decoding value when the first syntax information is a fifth value.

[0418] In some embodiments, the decoding part 2402 is further configured to, when the first syntax information is the sixth value, continue to parse the code stream for each coordinate dimension to determine the third syntax information corresponding to the vertex in the coordinate dimension; determine the target prediction mode corresponding to the vertex in the coordinate dimension based on the third syntax information corresponding to the vertex in the coordinate dimension, and continue to parse the code stream to determine the motion vector residual decoding value corresponding to the vertex in the coordinate dimension; perform motion vector prediction based on the target prediction mode corresponding to the vertex in the coordinate dimension to determine the target motion vector prediction value corresponding to the vertex in the coordinate dimension; determine the motion vector decoding value corresponding to the vertex in the coordinate dimension based on the target motion vector prediction value and the motion vector residual decoding value corresponding to the vertex in the coordinate dimension, thereby determining the motion vector decoding value corresponding to the vertex in each coordinate dimension; determine the motion vector decoding information corresponding to the vertex based on the motion vector decoding value corresponding to the vertex in each coordinate dimension.

[0419] In some embodiments, the decoding part 2402 is further configured to set the target motion vector prediction value corresponding to the vertex in the coordinate dimension to a preset prediction value when the target prediction mode corresponding to the vertex in the coordinate dimension represents no prediction mode.

[0420] In some embodiments, the decoding part 2402 is further configured to determine the motion vector residual decoding value corresponding to the vertex in the coordinate dimension as the motion vector decoding value corresponding to the vertex in the coordinate dimension when the target prediction mode corresponding to the vertex in the coordinate dimension represents no prediction mode.

[0421] In some embodiments, the decoding part 2402 is further configured to determine the target motion vector prediction value corresponding to the vertex in the coordinate dimension based on the first weighted average of the motion vector reconstruction values ​​corresponding to the coordinate dimension of a first number of neighboring points of the vertex, when the target prediction mode corresponding to the vertex in the coordinate dimension represents a first prediction mode.

[0422] In some embodiments, the decoding part 2402 is further configured to determine a first bias value based on a first weighted value of a motion vector reconstruction value corresponding to a first number of neighboring points of the vertex in the coordinate dimension when the target prediction mode corresponding to the vertex in the coordinate dimension represents a second prediction mode; and determine the target motion vector prediction value corresponding to the vertex in the coordinate dimension based on the first weighted value and the first bias value.

[0423] In some embodiments, the decoding part 2402 is further configured to determine the target motion vector prediction value corresponding to the vertex in the coordinate dimension based on the second weighted average of the motion vector reconstruction values ​​corresponding to the coordinate dimension of a second number of neighboring points of the vertex, when the target prediction mode corresponding to the vertex in the coordinate dimension represents a third prediction mode.

[0424] In some embodiments, the decoding part 2402 is further configured to determine a second bias value based on a second weighted value of a motion vector reconstruction value corresponding to a second number of neighboring points of the vertex in the coordinate dimension when the target prediction mode corresponding to the vertex in the coordinate dimension represents a fourth prediction mode; and determine the motion vector prediction value corresponding to the vertex in the coordinate dimension based on the second weighted value and the second bias value.

[0425] In some embodiments, the decoding part 2402 is further configured to, when the first weighted value is greater than or equal to a preset threshold, right-shift the first number by a preset number of bits to determine the first offset value.

[0426] In some embodiments, the decoding part 2402 is further configured to, when the first weighted value is less than a preset threshold, right-shift the first number by a preset number of bits and invert it to determine the first bias value.

[0427] In some embodiments, the decoding part 2402 is further configured to, when the second weighted value is greater than or equal to a preset threshold, right-shift the second number by a preset number of bits to determine the second offset value.

[0428] In some embodiments, the decoding part 2402 is further configured to, when the second weighted value is less than a preset threshold, right-shift the second number by a preset number of bits and invert it to determine the second bias value.

[0429] In some embodiments, the first number of neighbor points or the second number of neighbor points includes one or more of the following:

[0430] In the current image, the vertices that are connected to the vertex and whose coding order is before the vertex;

[0431] and / or, a vertex with the same index as the vertex in a reference image corresponding to the current image;

[0432] and / or, in the reference image, all vertices that are connected to the same point as the vertex;

[0433] and / or, in the reference image, a vertex having the same neighbor point index as the vertex in the current image.

[0434] In some embodiments, the vertices in the basic mesh include: each vertex of the current group in at least one group of the basic mesh; the parsing part 2401 is further configured to determine the first syntax information corresponding to each coordinate dimension of the current group in at least one coordinate dimension by parsing the code stream; for each coordinate dimension, the first syntax information corresponding to the current group in the coordinate dimension is determined as the first syntax information corresponding to each vertex of the current group in the coordinate dimension, thereby determining the first syntax information corresponding to each vertex of the current group in each coordinate dimension.

[0435] In some embodiments, the vertices in the basic mesh include: each vertex in the basic mesh; the parsing part 2401 is further configured to determine the first syntax information corresponding to each coordinate dimension of the basic mesh in at least one coordinate dimension by parsing the code stream; for each coordinate dimension, the first syntax information corresponding to the basic mesh in the coordinate dimension is determined as the first syntax information corresponding to each vertex in the basic mesh in the coordinate dimension, thereby determining the first syntax information corresponding to each vertex in the basic mesh in each coordinate dimension.

[0436] In some embodiments, the vertices in the basic mesh include: each vertex of the current group in at least one group of the basic mesh; the parsing part 2401 is further configured to determine the first syntax information corresponding to the current group by parsing the code stream; and determine the first syntax information corresponding to the current group as the first syntax information corresponding to each vertex of the current group.

[0437] In some embodiments, the vertices in the basic mesh include: each vertex in the basic mesh; the parsing part 2401 is further configured to determine the first syntax information corresponding to the basic mesh by parsing the code stream; and determine the first syntax information corresponding to the basic mesh as the first syntax information corresponding to each vertex in the basic mesh.

[0438] In some embodiments, the vertices in the basic mesh include: each vertex of the current group in at least one group of the basic mesh; the parsing part 2401 is further configured to determine the tenth grammatical information corresponding to the current group by parsing the code stream; and determine the tenth grammatical information corresponding to the current group as the tenth grammatical information corresponding to each vertex of the current group.

[0439] In some embodiments, the vertices in the basic mesh include: each vertex in the basic mesh; the parsing part 2401 is further configured to determine the tenth grammatical information corresponding to the basic mesh by parsing the code stream; and determine the tenth grammatical information corresponding to the basic mesh as the tenth grammatical information corresponding to each vertex in the basic mesh.

[0440] In some embodiments, the vertices in the basic grid include: each vertex of the current group in at least one group of the basic grid; the parsing part 2401 is further configured to determine the third syntax information corresponding to the current group in the coordinate dimension by continuing to parse the code stream; and determine the third syntax information corresponding to the current group in the coordinate dimension as the third syntax information corresponding to each vertex of the current group in the coordinate dimension.

[0441] In some embodiments, the vertices in the basic mesh include: each vertex in the basic mesh; the parsing part 2401 is further configured to determine the third syntax information corresponding to the basic mesh in the coordinate dimension by continuing to parse the code stream; and determine the third syntax information corresponding to the basic mesh in the coordinate dimension as the third syntax information corresponding to each vertex in the basic mesh in the coordinate dimension.

[0442] In some embodiments, the parsing part 2401 is further configured to determine the fourth syntax information and / or the fifth syntax information by parsing the bitstream;

[0443] The decoder 240 also includes a second grouping part; the second grouping part is configured to group based on the index corresponding to each vertex in the basic mesh and determine the at least one group when the fourth grammatical information and / or the fifth grammatical information represent group decoding of the basic mesh.

[0444] In some embodiments, the parsing part 2401 is further configured to determine the first syntax information corresponding to the basic grid in each coordinate dimension by parsing the code stream when the fourth syntax information and / or the fifth syntax information indicates that the basic grid is not to be grouped and decoded.

[0445] In some embodiments, the second grouping portion is further configured to determine sixth syntax information and / or seventh syntax information by parsing the bitstream; the sixth syntax information and / or the seventh syntax information represent the number of vertices in the group;

[0446] The vertices in the basic mesh are grouped according to the index corresponding to each vertex in the basic mesh and the number of vertices in the group to determine the at least one group.

[0447] In some embodiments, the parsing part 2401 is further configured to determine the eighth grammatical information and the ninth grammatical information corresponding to the coordinate dimension by continuing to parse the code stream; when the eighth grammatical information is the third value and the ninth grammatical information is the fourth value, continue to parse the code stream to determine the motion vector residual decoding value corresponding to the coordinate dimension.

[0448] In some embodiments, the decoding part 2402 is further configured to continue parsing the code stream when the eighth grammatical information is the third value and the ninth grammatical information is the fourth value, and determine the remainder of the motion vector residual corresponding to the coordinate dimension; and determine the motion vector residual decoding value based on the remainder of the motion vector residual.

[0449] It is understood that in this embodiment, a "portion" may be a circuit portion, a processor portion, a program portion, or software portion, and may also be a module or non-modular. Furthermore, the various components in this embodiment may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional modules.

[0450] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, this embodiment provides a computer-readable storage medium, which is applied to the decoder 240 and stores a computer program. When the computer program is executed by the second processor, it implements any of the methods in the aforementioned embodiments.

[0451] It should be noted that the description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of the present invention, please refer to the description of the method embodiment of the present invention for understanding.

[0452] Based on the composition of the decoder 240 and the computer-readable storage medium, refer to Figure 19, which shows a specific hardware structure diagram of the decoder 240 provided in an embodiment of the present application. As shown in Figure 19, the decoder 240 may include: a second communication interface 2501, a second memory 2502 and a second processor 2503; each component is coupled together through a second bus system 2504. It can be understood that the second bus system 2504 is used to achieve connection and communication between these components. In addition to the data bus, the second bus system 2504 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as the second bus system 2504 in Figure 19. Among them,

[0453] The second communication interface 2501 is used to receive and send signals during the process of sending and receiving information between other external network elements;

[0454] The second memory 2502 is used to store computer programs that can be run on the second processor 2503;

[0455] The second processor 2503 is configured to execute the decoding method applied to the decoder provided in the embodiment of the present application when running the computer program.

[0456] Optionally, as another embodiment, the second processor 2503 is further configured to execute any one of the methods described in the foregoing embodiments when running the computer program.

[0457] It can be understood that the hardware functions of the second memory 2502 are similar to those of the first memory 2302, and the hardware functions of the second processor 2503 are similar to those of the first processor 2303; they will not be described in detail here.

[0458] In yet another embodiment of the present application, referring to FIG20 , a schematic diagram of the structure of a coding and decoding system provided by an embodiment of the present application is shown. As shown in FIG20 , the coding and decoding system 260 may include an encoder 2601 and a decoder 2602 .

[0459] In the embodiment of the present application, the encoder 2601 may be the encoder described in any one of the aforementioned embodiments, and the decoder 2602 may be the decoder described in any one of the aforementioned embodiments.

[0460] It should be noted that, in this application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0461] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0462] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0463] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0464] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0465] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims. Industrial Applicability

[0466] In an embodiment of the present application, at the encoding end, the motion vectors corresponding to the base grid of the current image and the vertices in the base grid in each coordinate dimension are determined; when performing inter-frame encoding on the base grid, skip encoding and coding cost estimation of at least one prediction mode are performed on the motion vectors corresponding to the vertices in each coordinate dimension, and a first coding cost and at least one second coding cost corresponding to each coordinate dimension are determined; the first coding cost corresponds to the skip encoding mode; at least one second coding cost corresponds to at least one prediction mode; based on the first coding cost and at least one second coding cost corresponding to each coordinate dimension, first syntax information is determined; the first syntax information indicates whether to skip encoding of the motion vector. At the decoding end, the bitstream is parsed to determine the first syntax information corresponding to the vertices in the base grid of the current image; the first syntax information indicates whether to skip decoding of the motion vector; the first syntax information acts on the vertex, or acts on each coordinate dimension corresponding to the vertex separately; based on the first syntax information, the decoding information of the motion vector corresponding to the vertex is determined. In this way, by comparing the coding costs of the motion vectors of the vertices in the basic grid in each coordinate dimension under skip coding and at least one prediction mode, the optimal processing method corresponding to the vertex in the coordinate dimension can be determined, such as skip coding or the optimal prediction mode, and the first syntax information is used to indicate the optimal processing method to the decoder, thereby improving the coding and decoding efficiency and thus improving the coding and decoding performance.

Claims

1. A decoding method, applied to a decoder, the method comprising: Analyzing a bitstream to determine first syntax information corresponding to vertices in a base grid of a current image; The first syntax information indicating whether to skip decoding of a motion vector; The first syntax information acting on the vertices, or acting on at least one coordinate dimension corresponding to the vertices respectively; Based on the first syntax information, determining motion vector decoding information corresponding to the vertices.

2. The method according to claim 1, wherein The analyzing a bitstream to determine first syntax information corresponding to vertices in a base grid of a current image includes: Analyzing the bitstream to determine second syntax information corresponding to the base grid; In a case where the second syntax information indicates inter-frame decoding, continuing to analyze the bitstream to determine first syntax information corresponding to vertices in the base grid of the current image.

3. The method according to claim 1 or 2, wherein The analyzing a bitstream to determine first syntax information corresponding to vertices in a base grid of a current image includes: By analyzing the bitstream, determining first syntax information corresponding to each coordinate dimension in the at least one coordinate dimension of a vertex in the base grid; the first syntax information indicating whether to skip decoding of a motion vector corresponding to this coordinate dimension.

4. The method according to claim 1 or 2, wherein The analyzing a bitstream to determine first syntax information corresponding to vertices in a base grid of a current image includes: By analyzing the bitstream, determining first syntax information corresponding to a vertex in the base grid; the first syntax information indicating whether to skip decoding of a motion vector of the vertex.

5. The method according to claim 1 or 2, wherein The analyzing a bitstream to determine first syntax information corresponding to vertices in a base grid of a current image includes: Analyzing the bitstream to determine tenth syntax information corresponding to a vertex in the base grid; the tenth syntax information indicating whether to skip decoding of a motion vector of the vertex; In a case where the tenth syntax information does not indicate skipping decoding of a motion vector of the vertex, continuing to analyze the bitstream to determine first syntax information corresponding to each coordinate dimension in at least one coordinate dimension of the vertex in the base grid; the first syntax information indicating whether to skip decoding of a motion vector corresponding to the vertex in this coordinate dimension in the base grid.

6. The method according to claim 5, wherein, The method further comprises: In a case where the tenth syntax information indicates skipping decoding of a motion vector of the vertex, setting a motion vector decoding value corresponding to the vertex to a preset value.

7. The method according to claim 3 or 5, wherein The based on the first syntax information, determining motion vector decoding information corresponding to the vertices includes: For each coordinate dimension in the at least one coordinate dimension, in a case where the first syntax information corresponding to this coordinate dimension is a first value, setting a motion vector decoding value corresponding to the vertex in this coordinate dimension to a preset value; thereby determining motion vector decoding values corresponding to the vertex in each coordinate dimension; According to the motion vector decoding values corresponding to the vertex in each coordinate dimension, determining motion vector decoding information corresponding to the vertex.

8. The method according to claim 3 or 5, wherein The based on the first syntax information, determining motion vector decoding information corresponding to the vertices includes: For each of the at least one coordinate dimension, when the first syntax information corresponding to the vertex in this coordinate dimension is a second value, continue to parse the bitstream to determine the third syntax information corresponding to the vertex in this coordinate dimension; Determine the target prediction mode corresponding to the vertex in this coordinate dimension according to the third syntax information, and continue to parse the bitstream to determine the decoded value of the motion vector residual corresponding to the vertex in this coordinate dimension; Perform motion vector prediction according to the target prediction mode corresponding to the vertex in this coordinate dimension to determine the target motion vector prediction value corresponding to the vertex in this coordinate dimension; Determine the decoded value of the motion vector corresponding to the vertex in this coordinate dimension according to the target motion vector prediction value and the decoded value of the motion vector residual corresponding to the vertex in this coordinate dimension, so as to determine the decoded value of the motion vector corresponding to the vertex in each coordinate dimension; Determine the decoded motion vector information corresponding to the vertex according to the decoded value of the motion vector corresponding to the vertex in each coordinate dimension.

9. The method according to claim 5, wherein The method further includes: For the third coordinate dimension among the at least one coordinate dimension, when the tenth syntax information indicates not to skip the decoding of the motion vector of the vertex, and the first syntax information corresponding to the vertex in the first coordinate dimension and the second coordinate dimension are both first values, determine that the first syntax information corresponding to the vertex in the third coordinate dimension is a second value.

10. The method according to claim 4, wherein, The determining the decoded motion vector information corresponding to the vertex based on the first syntax information includes: When the first syntax information is a fifth value, determine the decoded motion vector information of the vertex as a preset decoded value.

11. The method according to claim 4, wherein The determining the decoded motion vector information corresponding to the vertex based on the first syntax information includes: When the first syntax information is a sixth value, for each of the at least one coordinate dimension, continue to parse the bitstream to determine the third syntax information corresponding to the vertex in this coordinate dimension; Determine the target prediction mode corresponding to the vertex in this coordinate dimension according to the third syntax information corresponding to the vertex in this coordinate dimension, and continue to parse the bitstream to determine the decoded value of the motion vector residual corresponding to the vertex in this coordinate dimension; Perform motion vector prediction according to the target prediction mode corresponding to the vertex in this coordinate dimension to determine the target motion vector prediction value corresponding to the vertex in this coordinate dimension; Determine the decoded value of the motion vector corresponding to the vertex in this coordinate dimension according to the target motion vector prediction value and the decoded value of the motion vector residual corresponding to the vertex in this coordinate dimension, so as to determine the decoded value of the motion vector corresponding to the vertex in each coordinate dimension; Determine the decoded motion vector information corresponding to the vertex according to the decoded value of the motion vector corresponding to the vertex in each coordinate dimension.

12. The method according to claim 8 or 11, wherein The performing motion vector prediction according to the target prediction mode corresponding to the vertex in this coordinate dimension to determine the target motion vector prediction value corresponding to the vertex in this coordinate dimension includes: When the target prediction mode corresponding to the vertex in this coordinate dimension represents no prediction mode, set the target motion vector prediction value corresponding to the vertex in this coordinate dimension to a preset prediction value.

13. The method according to claim 8 or 11, wherein, The method further includes: When the target prediction mode corresponding to the vertex in this coordinate dimension represents no prediction mode, determine the motion vector residual decoding value corresponding to the vertex in this coordinate dimension as the motion vector decoding value corresponding to the vertex in this coordinate dimension.

14. The method according to claim 8 or 11, wherein The determining the target motion vector prediction value corresponding to the vertex in this coordinate dimension according to the target prediction mode corresponding to the vertex in this coordinate dimension includes: When the target prediction mode corresponding to the vertex in this coordinate dimension represents the first prediction mode, determine the target motion vector prediction value corresponding to the vertex in this coordinate dimension according to the first weighted average of the motion vector reconstruction values corresponding to the first number of neighbor points of the vertex in this coordinate dimension.

15. The method according to claim 8 or 11, wherein The determining the target motion vector prediction value corresponding to the vertex in this coordinate dimension according to the target prediction mode corresponding to the vertex in this coordinate dimension includes: When the target prediction mode corresponding to the vertex in this coordinate dimension represents the second prediction mode, determine a first offset value according to the first weighted value of the motion vector reconstruction values corresponding to the first number of neighbor points of the vertex in this coordinate dimension; Determine the target motion vector prediction value corresponding to the vertex in this coordinate dimension according to the first weighted value and the first offset value.

16. The method according to claim 8 or 11, wherein, The determining the target motion vector prediction value corresponding to the vertex in this coordinate dimension according to the target prediction mode corresponding to the vertex in this coordinate dimension includes: When the target prediction mode corresponding to the vertex in this coordinate dimension represents the third prediction mode, determine the target motion vector prediction value corresponding to the vertex in this coordinate dimension according to the second weighted average of the motion vector reconstruction values corresponding to the second number of neighbor points of the vertex in this coordinate dimension.

17. The method according to claim 8 or 11, wherein The determining the target motion vector prediction value corresponding to the vertex in this coordinate dimension according to the target prediction mode corresponding to the vertex in this coordinate dimension includes: When the target prediction mode corresponding to the vertex in this coordinate dimension represents the fourth prediction mode, determine a second offset value according to the second weighted value of the motion vector reconstruction values corresponding to the second number of neighbor points of the vertex in this coordinate dimension; Determine the motion vector prediction value corresponding to the vertex in this coordinate dimension according to the second weighted value and the second offset value.

18. The method according to claim 15, wherein, The determining the first offset value according to the first weighted value of the motion vector reconstruction values corresponding to the first number of neighbor points of the vertex includes: When the first weighted value is greater than or equal to a preset threshold, shift the first number to the right by a preset number of bits to determine the first offset value.

19. The method according to claim 15, wherein, The determining the first offset value according to the first weighted value of the motion vector reconstruction values corresponding to the first number of neighbor points of the vertex includes: When the first weighted value is less than a preset threshold, the first number is right-shifted by a preset number of bits and inverted to determine the first offset value.

20. The method according to claim 17, wherein The determining of the second offset value according to the second weighted value of the motion vector reconstruction value corresponding to the second number of neighboring points of the vertex in the coordinate dimension includes: When the second weighted value is greater than or equal to a preset threshold, the second number is right-shifted by a preset number of bits to determine the second bias value. Set value.

21. The method according to claim 17, wherein, The determining of the second offset value according to the second weighted value of the motion vector reconstruction value corresponding to the second number of neighboring points of the vertex in the coordinate dimension includes: When the second weighted value is less than a preset threshold, the second number is right-shifted by a preset number of bits and inverted to determine the second offset value.

22. The method according to any one of claims 14-21, wherein, The first number of neighbor points or the second number of neighbor points includes one or more of the following: In the current image, a vertex that is connected to the vertex and whose encoding order is before the vertex; and / or, a vertex having the same index as the vertex in a reference image corresponding to the current image; and / or, in the reference image, all vertices that are connected to the same point of the vertex; And / or, in the reference image, a vertex having the same neighbor point index as the vertex in the current image.

23. The method according to claim 3, or any one of claims 5 to 9, or any one of claims 12 to 22, wherein, The vertices in the basic mesh include: each vertex in a current group in at least one group of the basic mesh; and determining, by parsing the bitstream, first syntax information corresponding to each coordinate dimension of the vertices in the basic mesh in the at least one coordinate dimension includes: Determine, by parsing the bitstream, first syntax information corresponding to each coordinate dimension of the current group in the at least one coordinate dimension; For each coordinate dimension in the at least one coordinate dimension, the first grammatical information corresponding to the current group in the coordinate dimension is determined as the first grammatical information corresponding to each vertex of the current group in the coordinate dimension, thereby determining the first grammatical information corresponding to each vertex of the current group in each coordinate dimension in the at least one coordinate dimension.

24. The method according to claim 3, or any one of claims 5-9, or any one of claims 12-22, wherein, The vertices in the basic mesh include: each vertex in the basic mesh; and determining, by parsing the bitstream, the first syntax information corresponding to each coordinate dimension of the vertices in the basic mesh in the at least one coordinate dimension includes: Determine, by parsing the bitstream, first syntax information corresponding to each coordinate dimension of the basic grid in the at least one coordinate dimension; For each coordinate dimension, the first grammatical information corresponding to the basic grid in the coordinate dimension is determined as the first grammatical information corresponding to each vertex in the basic grid in the coordinate dimension, thereby determining the first grammatical information corresponding to each vertex in the basic grid in each coordinate dimension in at least one coordinate dimension.

25. The method according to claim 4 or 10, or any one of claims 11-22, wherein, The vertices in the basic mesh include: each vertex of a current group in at least one group of the basic mesh; and determining the first syntax information corresponding to the vertices in the basic mesh by parsing the bitstream includes: Determine the first syntax information corresponding to the current group by parsing the bitstream; Determine the first syntax information corresponding to the current group as the first syntax information corresponding to each vertex of the current group.

26. The method according to claim 4 or 10, or any one of claims 11-22, wherein The vertices in the base mesh include: each vertex in the base mesh; Determining the first syntax information corresponding to the vertices in the base mesh by parsing the bitstream includes: Determine the first syntax information corresponding to the base mesh by parsing the bitstream; Determine the first syntax information corresponding to the base mesh as the first syntax information corresponding to each vertex in the base mesh.

27. The method according to any one of claims 5-9, or any one of claims 12-23, wherein The vertices in the base mesh include: each vertex of the current group in at least one group of the base mesh; Parsing the bitstream to determine the tenth syntax information corresponding to the vertices in the base mesh includes: Determine the tenth syntax information corresponding to the current group by parsing the bitstream; Determine the tenth syntax information corresponding to the current group as the tenth syntax information corresponding to each vertex of the current group.

28. The method according to any one of claims 5-9 or any one of claims 12-23, wherein, The vertices in the base mesh include: each vertex in the base mesh; Parsing the bitstream to determine the tenth syntax information corresponding to the vertices in the base mesh includes: Determine the tenth syntax information corresponding to the base mesh by parsing the bitstream; Determine the tenth syntax information corresponding to the base mesh as the tenth syntax information corresponding to each vertex in the base mesh.

29. The method according to claim 8 or 11, or any one of claims 12 - 22, wherein, The vertices in the base mesh include: each vertex of the current group in at least one group of the base mesh; Continuing to parse the bitstream to determine the third syntax information corresponding to the vertex in this coordinate dimension includes: Determine the third syntax information corresponding to the current group in this coordinate dimension by continuing to parse the bitstream; Determine the third syntax information corresponding to the current group in this coordinate dimension as the third syntax information corresponding to each vertex of the current group in this coordinate dimension.

30. The method according to claim 8 or 11, or any one of claims 12 - 22, wherein, The vertices in the base mesh include: each vertex in the base mesh; Continuing to parse the bitstream to determine the third syntax information corresponding to the vertex in this coordinate dimension includes: Determine the third syntax information corresponding to the base mesh in this coordinate dimension by continuing to parse the bitstream; Determine the third syntax information corresponding to the base mesh in this coordinate dimension as the third syntax information corresponding to each vertex in the base mesh in this coordinate dimension.

31. The method according to any one of claims 23, 25, or 29, wherein The method further includes: Determine the fourth syntax information and / or the fifth syntax information by parsing the bitstream; When the fourth syntax information and / or the fifth syntax information indicates grouped decoding of the base mesh, group based on the index corresponding to each vertex in the base mesh to determine the at least one group.

32. The method according to any one of claims 24, 26, or 30, wherein The determining the first syntax information corresponding to the base mesh in each coordinate dimension by parsing the bitstream includes: When the fourth syntax information and / or the fifth syntax information does not indicate grouped decoding of the base mesh, determine the first syntax information corresponding to the base mesh in each coordinate dimension by parsing the bitstream.

33. The method according to claim 31, wherein The grouping based on the index corresponding to each vertex in the base mesh to determine the at least one group includes: Determine the sixth syntax information and / or the seventh syntax information by parsing the bitstream; the sixth syntax information and / or the seventh syntax information represent the number of intra-group vertices; Group the vertices in the base mesh according to the index corresponding to each vertex in the base mesh and the number of intra-group vertices to determine the at least one group.

34. The method according to claim 8 or 11, wherein The continued parsing of the bitstream to determine the decoded value of the motion vector residual corresponding to the coordinate dimension includes: Determine the eighth syntax information and the ninth syntax information corresponding to the coordinate dimension by continuing to parse the bitstream; When the eighth syntax information is the third value and the ninth syntax information is the fourth value, continue to parse the bitstream to determine the decoded value of the motion vector residual corresponding to the coordinate dimension.

35. The method according to claim 34, wherein, The continued parsing of the bitstream to determine the decoded value of the motion vector residual corresponding to the coordinate dimension includes: When the eighth syntax information is the third value and the ninth syntax information is the fourth value, continue to parse the bitstream to determine the remainder of the motion vector residual corresponding to the coordinate dimension; Determine the decoded value of the motion vector residual according to the remainder of the motion vector residual.

36. An encoding method applied to an encoder, the method includes: Determine the base mesh of the current image and the motion vectors corresponding to each vertex in the base mesh in each coordinate dimension of at least one coordinate dimension; When performing inter-frame encoding on the base mesh, estimate the encoding cost of skip encoding and at least one prediction mode for the motion vectors corresponding to the vertices in each coordinate dimension, and determine the first encoding cost corresponding to each coordinate dimension and at least one second encoding cost; the first encoding cost corresponds to the skip encoding mode; The at least one second encoding cost corresponds to the at least one prediction mode; Determine the first syntax information according to the first encoding cost corresponding to each coordinate dimension and the at least one second encoding cost; The first syntax information represents whether to skip the encoding of the motion vector.

37. The method according to claim 36, wherein, The determining the first syntax information according to the first encoding cost corresponding to each coordinate dimension and the at least one second encoding cost includes: For each coordinate dimension, when the first encoding cost is less than or equal to each second encoding cost among the at least one second encoding cost, determine that the first syntax information corresponding to the vertex in this coordinate dimension is the first value; the first value represents skipping the encoding of the motion vector of the vertex in this coordinate dimension.

38. The method according to claim 36, wherein The determining the first syntax information according to the first encoding cost corresponding to each coordinate dimension and the at least one second encoding cost includes: For each coordinate dimension, when the first encoding cost is greater than or equal to any second encoding cost among the at least one second encoding cost, determine that the first syntax information corresponding to the vertex in this coordinate dimension is the second value; the second value represents not skipping the encoding of the motion vector of the vertex in this coordinate dimension.

39. The method according to claim 37, wherein, The method further includes: For each of the coordinate dimensions, when the first coding cost is less than or equal to each of the at least one second coding cost, set the motion vector corresponding to the vertex in this coordinate dimension to a preset value.

40. The method according to claim 38, wherein, The method further includes: For each of the coordinate dimensions, when the first coding cost is greater than or equal to any one of the at least one second coding cost, determine the target prediction mode corresponding to the vertex in each coordinate dimension according to the at least one second coding cost; Determine the third syntax information corresponding to the vertex in each coordinate dimension according to the target prediction mode corresponding to the vertex in each coordinate dimension.

41. The method according to claim 40, wherein, The method further includes: Perform motion vector prediction according to the target prediction mode corresponding to the vertex in each coordinate dimension, and determine the target motion vector prediction value corresponding to the vertex in each coordinate dimension; Determine the motion vector residual corresponding to the vertex in each coordinate dimension according to the motion vector and the target motion vector prediction value corresponding to the vertex in each coordinate dimension; Encode the motion vector residual corresponding to the vertex in each coordinate dimension to determine the motion vector encoding corresponding to the vertex in each coordinate dimension.

42. The method according to claim 36, wherein, The determining of the first syntax information according to the first coding cost corresponding to each coordinate dimension and the at least one second coding cost includes: When the sum of the first coding costs of the vertex in the three coordinate dimensions is less than or equal to the sum of the three coordinate dimensions of each of the at least one second coding cost, determine the first syntax information as a fifth value; the fifth value represents skipping the encoding of the motion vector of the vertex.

43. The method according to claim 42, wherein The method further includes: When the sum of the first coding costs of the vertex in the three coordinate dimensions is less than or equal to the sum of the three coordinate dimensions of each of the at least one second coding cost, set the motion vector corresponding to the vertex to a preset value.

44. The method according to claim 36, wherein The determining of the first syntax information according to the first coding cost corresponding to each coordinate dimension and the at least one second coding cost includes: When the sum of the first coding costs of the vertex in the three coordinate dimensions is greater than or equal to the sum of the three coordinate dimensions of each of the at least one second coding cost, determine the first syntax information as a sixth value; the sixth value represents not skipping the encoding of the motion vector of the vertex.

45. The method according to claim 44, wherein, The method further includes: When the sum of the first coding costs of the vertex in the three coordinate dimensions is greater than or equal to the sum of the three coordinate dimensions of each of the at least one second coding cost, determine the target prediction mode corresponding to the vertex in each coordinate dimension according to the at least one second coding cost corresponding to the vertex in each coordinate dimension; Determine the third syntax information corresponding to the vertex in each coordinate dimension according to the target prediction mode corresponding to the vertex in each coordinate dimension.

46. The method according to claim 36, wherein, The determining of the first syntax information according to the first coding cost corresponding to each coordinate dimension and the at least one second coding cost includes: Determine tenth syntax information according to the first coding cost corresponding to each coordinate dimension and at least one second coding cost; the tenth syntax information indicates whether to skip coding of the motion vector of the vertex. In the case where the tenth syntax information indicates not to skip coding of the motion vector of the vertex, determine the first syntax information corresponding to the vertex in each coordinate dimension according to the first coding cost corresponding to each coordinate dimension and at least one second coding cost; the first syntax information indicates whether to skip coding of the motion vector corresponding to the vertex in this coordinate dimension.

47. The method according to claim 46, wherein The determining the tenth syntax information according to the first coding cost corresponding to each coordinate dimension and at least one second coding cost includes: In the case where the sum of the first coding costs of the vertex in the three coordinate dimensions is less than or equal to the sum of the three coordinate dimensions of each second coding cost among the at least one second coding cost, determine the tenth syntax information as a seventh value; the seventh value indicates skipping coding of the motion vector of the vertex.

48. The method according to claim 47, wherein, The method further includes: In the case where the sum of the first coding costs of the vertex in the three coordinate dimensions is less than or equal to the sum of the three coordinate dimensions of each second coding cost among the at least one second coding cost, set the motion vector coding corresponding to the vertex to a preset value.

49. The method according to claim 46, wherein The determining the tenth syntax information corresponding to the vertex according to the first coding cost corresponding to each coordinate dimension and at least one second coding cost includes: In the case where the sum of the first coding costs of the vertex in the three coordinate dimensions is greater than or equal to the sum of the three coordinate dimensions of each second coding cost among the at least one second coding cost, determine the tenth syntax information as an eighth value; the eighth value indicates not to skip coding of the motion vector of the vertex.

50. The method according to any one of claims 36 - 49, wherein, The at least one second coding cost corresponds to at least one of at least one prediction mode and a non-prediction mode; estimating the coding cost of at least one prediction mode for the motion vector corresponding to the vertex in each coordinate dimension, and determining the at least one second coding cost corresponding to each coordinate dimension includes: For each coordinate dimension, determine a first motion vector prediction value corresponding to the vertex in this coordinate dimension through a non-prediction mode; the non-prediction mode indicates setting the motion vector prediction value corresponding to this coordinate dimension to a preset vector prediction value. Based on the first motion vector prediction value, estimate the coding cost and determine the second coding cost corresponding to the non-prediction mode. And / or, Determine a second motion vector prediction value according to the first weighted average of the motion vector reconstruction values corresponding to the first number of neighbor points of the vertex in this coordinate dimension. Based on the second motion vector prediction value, estimate the coding cost and determine the second coding cost corresponding to the first prediction mode in the at least one prediction mode. And / or, Determine a first offset value according to the first weighted value of the motion vector reconstruction values corresponding to the first number of neighbor points of the vertex in this coordinate dimension. Determine a third motion vector prediction value according to the first weighted value and the first offset value. Estimate the coding cost according to the third motion vector prediction value, and determine the second coding cost corresponding to the second prediction mode in the at least one prediction mode; And / or, Determine a fourth motion vector prediction value according to a second weighted average of motion vector reconstruction values corresponding to a second number of neighbor points of the vertex in this coordinate dimension; Estimate the coding cost according to the fourth motion vector prediction value, and determine the second coding cost corresponding to the third prediction mode in the at least one prediction mode; And / or, Determine a second bias value according to a second weighted value of motion vector reconstruction values corresponding to a second number of neighbor points of the vertex in this coordinate dimension; Determine a fifth motion vector prediction value according to the second weighted value and the second bias value; Estimate the coding cost according to the fifth motion vector prediction value, and determine the second coding cost corresponding to the fourth prediction mode in the at least one prediction mode.

51. The method according to claim 50, wherein The determining the first bias value according to the first weighted value of motion vector reconstruction values corresponding to a first number of neighbor points of the vertex in this coordinate dimension includes: When the first weighted average is greater than or equal to a preset threshold, shift the first number to the right by a preset number of bits to determine the first bias value.

52. The method according to claim 50, wherein, The determining the first bias value according to the first weighted value of motion vector reconstruction values corresponding to a first number of neighbor points of the vertex in this coordinate dimension includes: When the first weighted value is less than the preset threshold, shift the first number to the right by a preset number of bits and take the inverse to determine the first bias value.

53. The method according to claim 50, wherein The determining the second bias value according to the second weighted value of motion vector reconstruction values corresponding to a second number of neighbor points of the vertex in this coordinate dimension includes: When the second weighted value is greater than or equal to a preset threshold, shift the second number to the right by a preset number of bits to determine the second bias value.

54. The method according to claim 50, wherein, The determining the second bias value according to the second weighted value of motion vector reconstruction values corresponding to a second number of neighbor points of the vertex in this coordinate dimension includes: When the second weighted average is less than the preset threshold, shift the second number to the right by a preset number of bits and take the inverse to determine the second bias value.

55. The method according to any one of claims 36 - 54, wherein, The method further includes: Encode the first syntax information corresponding to the vertex, and write the obtained encoded bits into the bitstream; Or, Encode the first syntax information corresponding to the vertex and the third syntax information corresponding to the vertex in each coordinate dimension, and write the obtained encoded bits into the bitstream.

56. The method according to claim 41, wherein, The encoding the motion vector residual corresponding to the vertex in each coordinate dimension to determine the motion vector encoding corresponding to the vertex in each coordinate dimension includes: For each coordinate dimension, when the absolute value of the motion vector residual corresponding to the vertex in this coordinate dimension is greater than a first preset residual threshold and less than a second preset residual threshold, determine the eighth syntax information corresponding to the vertex in this coordinate dimension as a third value, And determine the ninth syntax information corresponding to the vertex in this coordinate dimension as a fourth value; the second preset residual threshold is greater than the first preset residual threshold; Encode the motion vector residual corresponding to the vertex in this coordinate dimension to determine the motion vector encoding corresponding to the vertex in this coordinate dimension, so as to determine the motion vector encoding corresponding to the vertex in each coordinate dimension.

57. The method according to claim 41, wherein, The encoding of the motion vector residual corresponding to the vertex in each coordinate dimension to determine the motion vector encoding corresponding to the vertex in each coordinate dimension includes: For each of the coordinate dimensions, when the absolute value of the motion vector residual corresponding to the vertex in this coordinate dimension is greater than or equal to a second preset residual threshold, determine the eighth syntax information corresponding to the vertex in this coordinate dimension as a third value, and determine the ninth syntax information corresponding to the vertex in this coordinate dimension as the third value; Encode the remainder of the motion vector residual corresponding to the vertex in this coordinate dimension to determine the motion vector encoding corresponding to the vertex in this coordinate dimension, so as to determine the motion vector encoding corresponding to the vertex in each coordinate dimension.

58. The method according to claim 56 or 57, wherein, The method further includes: Perform entropy encoding on the eighth syntax information and the ninth syntax information corresponding to the vertex in each coordinate dimension, and write the obtained encoded bits and the motion vector encoding corresponding to the vertex in each coordinate dimension into the bitstream.

59. The method according to any one of claims 36 - 58, wherein, The method further includes: Determine second syntax information; the second syntax information represents intra-frame encoding or inter-frame encoding of the current image; Perform entropy encoding on the second syntax information, and write the obtained encoded bits into the bitstream.

60. The method according to any one of claims 36 - 59, wherein, The vertices in the base grid include: each vertex of the current group in at least one group of the base grid; the encoding cost estimation of skipping encoding and at least one prediction mode for the motion vector corresponding to the vertex in each coordinate dimension to determine the first encoding cost corresponding to each coordinate dimension and at least one second encoding cost includes: For each of the coordinate dimensions, perform encoding estimation of skipping encoding and encoding estimation of at least one prediction mode for the motion vector corresponding to each vertex of the current group in this coordinate dimension, and determine the first encoding cost and at least one second encoding cost corresponding to each vertex of the current group in this coordinate dimension; According to the first encoding cost and at least one second encoding cost corresponding to each vertex of the current group in this coordinate dimension, determine the first encoding cost and at least one second encoding cost corresponding to the current group in this coordinate dimension, so as to determine the first encoding cost and at least one second encoding cost corresponding to the current group in each coordinate dimension.

61. The method according to any one of claims 36 - 59, wherein The vertices in the base grid include: each vertex in the base grid; the encoding cost estimation of skipping encoding and at least one prediction mode for the motion vector corresponding to the vertex in each coordinate dimension to determine the first encoding cost corresponding to each coordinate dimension and at least one second encoding cost includes: For each of the coordinate dimensions, perform encoding cost estimation of skipping encoding and at least one prediction mode for the motion vector corresponding to each vertex in the base grid in this coordinate dimension, and determine the first encoding cost and at least one second encoding cost corresponding to each vertex in the base grid in this coordinate dimension; According to the first coding cost and at least one second coding cost corresponding to each vertex in the basic mesh in the coordinate dimension, the first coding cost and at least one second coding cost corresponding to the basic mesh in the coordinate dimension are determined, thereby determining the first coding cost and at least one second coding cost corresponding to the basic mesh in each coordinate dimension.

62. The method according to claim 60, wherein, The method further comprises: determining fourth grammatical information and / or fifth grammatical information; In a case where the fourth grammatical information and / or the fifth grammatical information represents group encoding of the basic mesh, grouping is performed based on an index corresponding to each vertex in the basic mesh and the number of vertices in a group to determine the at least one group.

63. The method according to claim 62, wherein, The method further comprises: Determine sixth grammatical information and / or seventh grammatical information according to the number of vertices in the group; Perform entropy coding on the sixth syntax information and / or the seventh syntax information, and write the obtained coding bits into a bitstream.

64. The method according to claim 50, wherein, The first number of neighbor points or the second number of neighbor points includes one or more of the following: In the current image, a vertex that is connected to the vertex and whose encoding order is before the vertex; and / or, a vertex having the same index as the vertex in a reference image corresponding to the current image; and / or, in the reference image, all vertices that are connected to the same point of the vertex; And / or, in the reference image, a vertex having the same neighbor point index as the vertex in the current image.

65. A decoder comprising: A parsing part, configured to parse the code stream to determine first syntax information corresponding to vertices in a base mesh of a current image; The first syntax information indicates whether to skip decoding of the motion vector; The first syntax information acts on the vertex, or acts on the vertex corresponding to at least one coordinate dimension of ; The decoding part is configured to determine the motion vector decoding information corresponding to the vertex based on the first syntax information.

66. An encoder comprising: A motion vector determination part, configured to determine a basic grid of a current image and a motion vector corresponding to each coordinate dimension of a vertex in the basic grid in at least one coordinate dimension; The coding cost determination part is configured to perform skip coding and coding cost estimation of at least one prediction mode on the motion vector corresponding to each coordinate dimension of the vertex when performing inter-frame coding on the basic grid, and determine a first coding cost and at least one second coding cost corresponding to each coordinate dimension; the first coding cost corresponds to the skip coding mode; The at least one second encoding cost corresponds to the at least one prediction mode; The grammatical information determining part is configured to determine the first grammatical information according to the first coding cost corresponding to each coordinate dimension and at least one second coding cost; The first syntax information indicates whether to skip encoding of the motion vector.

67. A decoder comprising: a first memory and a first processor; The first memory stores a computer program executable on the first processor, and the first processor implements the method according to any one of claims 1 to 35 when executing the program.

68. An encoder, comprising: a second memory and a second processor; the second memory stores a computer program that can run on the second processor, and when the second processor executes the program, the method described in any one of claims 36-64 is implemented.

69. A bitstream, which is generated by performing bit encoding on information to be encoded; wherein, The information to be encoded at least includes first syntax information; the first syntax information represents whether to skip encoding of motion vectors; the first syntax information is determined by the following method: When performing inter-frame encoding on the vertices in the base grid, estimate the encoding cost of skip encoding and at least one prediction mode for the motion vectors corresponding to each coordinate dimension of the vertices, and determine the first encoding cost corresponding to each coordinate dimension and at least one second encoding cost; the first encoding cost corresponds to the skip encoding mode; the at least one second encoding cost corresponds to the at least one prediction mode; Determine the first syntax information according to the first encoding cost and the at least one second encoding cost.

70. A storage medium, on which a computer program is stored, when the computer program is executed by a first processor, the method described in any one of claims 1-35 is implemented; or, when the computer program is executed by a second processor, the method described in any one of claims 36-64 is implemented.

Citation Information

Patent Citations

  • Mesh patch syntax

    US20230306644A1

  • Curvature-Guided Inter-Patch 3D Inpainting for Dynamic Mesh Coding

    US20230316647A1

  • Mesh vertex displacements coding

    US20230412849A1

  • 2D mesh geometry and motion vector compression

    US6047088A

  • Encoding device, decoding device, encoding method, and decoding method

    WO2023238867A1