Video coding and decoding method, device, equipment, system and storage medium

CN121925678APending Publication Date: 2026-04-24GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
Filing Date
2023-09-13
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In the prior art, the displacement coefficient prediction of three-dimensional grid videos is inaccurate, resulting in poor encoding and decoding effects.

Method used

By organizing the detailed layer of the refined grid vertices of the three-dimensional grid image, determining the M adjacent points of the vertex, and predicting them based on these displacement coefficients, thereby obtaining the reconstruction displacement coefficient or displacement coefficient residual value.

Benefits of technology

It improves the prediction accuracy of the vertex displacement coefficient in three-dimensional grid videos and improves the encoding and decoding effect of three-dimensional grid videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121925678A_ABST
    Figure CN121925678A_ABST
Patent Text Reader

Abstract

The invention provides a video encoding and decoding method, device, equipment, system and storage medium, when geometric information of a current three-dimensional grid image is encoded and decoded, detail layer organization is performed on vertexes of a refined grid of the current three-dimensional grid image to obtain a detail layer structure of the vertexes of the refined grid, and the detail layer structure comprises N detail layers. And for the jth vertex of the ith detail layer in the N detail layers, determining M adjacent points of the jth vertex in the displacement coefficient coded and decoded vertexes of the refined grid. And determining a predicted displacement coefficient of the jth vertex based on the displacement coefficients of the M adjacent points. And obtaining a reconstruction displacement coefficient of the jth vertex based on the predicted displacement coefficient of the jth vertex. Namely, according to the embodiment of the invention, the displacement coefficient of the current vertex is predicted based on the displacement coefficient of the adjacent point of the current vertex, so that the prediction accuracy of the displacement coefficient of the vertex can be improved, and the coding and decoding effects of the three-dimensional grid video are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding and decoding method, device, equipment, system, and storage medium Technical Field

[0001] The present application relates to the field of video coding and decoding technology, and in particular to a video coding and decoding method, apparatus, device, system, and storage medium. Background Art

[0002] A 3D mesh is a three-dimensional object surface composed of countless polygons in space. In the 3D mesh video compression process, the 3D mesh image is first preprocessed to generate a basic mesh and displacement coefficients, which are then encoded.

[0003] Currently, displacement coefficients are encoded by performing wavelet transform, quantization, and two-dimensional mapping before being encoded using a video encoder. During the wavelet transform process, the encoder first sorts the vertices of the 3D mesh image by detail layer. For each vertex in the current detail layer, the displacement coefficient is predicted. However, current displacement coefficient prediction methods are inaccurate, resulting in poor encoding and decoding performance for the derived 3D mesh video.

[0004] Summary of the Invention

[0005] The embodiments of the present application provide a video encoding and decoding method, apparatus, device, system, and storage medium, which can improve the prediction accuracy of the displacement coefficients of vertices in a three-dimensional mesh video, thereby improving the encoding and decoding effect of the three-dimensional mesh video.

[0006] In a first aspect, the present application provides a video decoding method, applied to a decoder, comprising:

[0007] performing detail layer organization on vertices of a refined mesh of a current three-dimensional mesh image to obtain a detail layer structure of the vertices of the refined mesh, wherein the refined mesh is obtained by subdividing a base mesh of the current three-dimensional mesh image, and the detail layer structure includes N detail layers, where N is a positive integer;

[0008] For a j-th vertex of an i-th detail layer among the N detail layers, determining M neighboring points of the j-th vertex from decoded vertices of the refinement grid, wherein the decoded vertices are vertices whose displacement coefficients have been decoded, i is a non-negative integer less than N, and j is a non-negative integer;

[0009] Determining a predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points;

[0010] Based on the predicted displacement coefficient of the j-th vertex, a reconstructed displacement coefficient of the j-th vertex is obtained.

[0011] In a second aspect, an embodiment of the present application provides a video encoding method, applied to an encoder, comprising:

[0012] performing detail layer organization on vertices of a refined mesh of a current three-dimensional mesh image to obtain a detail layer structure of the vertices of the refined mesh, wherein the refined mesh is obtained by subdividing a base mesh of the current three-dimensional mesh image, and the detail layer structure includes N detail layers, where N is a positive integer;

[0013] For a j-th vertex of an i-th detail layer among the N detail layers, determine M neighboring points of the j-th vertex from among the encoded vertices of the refined mesh, where the encoded vertices are vertices whose displacement coefficients have been encoded, i is a non-negative integer less than N, and j is a non-negative integer;

[0014] Determining a predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points;

[0015] Based on the predicted displacement coefficient of the j-th vertex, a displacement coefficient residual value of the j-th vertex is obtained.

[0016] In a third aspect, the present application provides a video decoding device for executing the method of the first aspect or its respective implementations. Specifically, the device includes a functional unit for executing the method of the first aspect or its respective implementations.

[0017] In a fourth aspect, the present application provides a video encoding device for executing the method of the second aspect or its respective implementations. Specifically, the device includes a functional unit for executing the method of the second aspect or its respective implementations.

[0018] In a fifth aspect, the present application provides a video decoder comprising a processor and a memory, wherein the memory is configured to store a computer program, and the processor is configured to call and execute the computer program stored in the memory to perform the method of the first aspect or its respective implementations.

[0019] In a sixth aspect, the present application provides a video encoder comprising a processor and a memory, wherein the memory is configured to store a computer program, and the processor is configured to call and execute the computer program stored in the memory to perform the method of the second aspect or its respective implementations.

[0020] In a seventh aspect, the present application provides a video encoding and decoding system, comprising a video encoder and a video decoder. The video decoder is configured to execute the method of the first aspect or its respective implementations, and the video encoder is configured to execute the method of the second aspect or its respective implementations.

[0021] In an eighth aspect, the present application provides a chip for implementing the method of any one of the first and second aspects above, or their respective implementations. Specifically, the chip includes: a processor for calling and executing a computer program from a memory, so that a device equipped with the chip executes the method of any one of the first and second aspects above, or their respective implementations.

[0022] In a ninth aspect, the present application provides a computer-readable storage medium for storing a computer program, which enables a computer to execute the method of any one of the first to second aspects above or any of their implementations.

[0023] In a tenth aspect, the present application provides a computer program product, comprising computer program instructions, which enable a computer to execute the method of any one of the above-mentioned first to second aspects or their respective implementations.

[0024] In an eleventh aspect, the present application provides a computer program which, when executed on a computer, enables the computer to execute the method in any one of the first to second aspects or their respective implementations.

[0025] Based on the above technical solution, when encoding and decoding the geometric information of the current three-dimensional grid image, the vertices of the refined grid of the current three-dimensional grid image are first organized into detail layers to obtain the detail layer structure of the vertices of the refined grid, and the detail layer structure includes N detail layers. For the j-th vertex of the i-th detail layer in the N detail layers, the M neighboring points of the j-th vertex are determined among the vertices whose displacement coefficients of the refined grid have been encoded and decoded. Then, based on the displacement coefficients of the M neighboring points, the predicted displacement coefficient of the j-th vertex is determined. Finally, based on the predicted displacement coefficient of the j-th vertex, the reconstructed displacement coefficient or the displacement coefficient residual value of the j-th vertex is obtained. That is to say, when predicting the vertices in the three-dimensional grid, the embodiment of the present application predicts the displacement coefficient of the current vertex based on the displacement coefficients of the neighboring points of the current vertex, which can improve the prediction accuracy of the displacement coefficient of the vertex, thereby improving the encoding and decoding effect of the three-dimensional grid video. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] FIG1 is a schematic block diagram of a video encoding and decoding system according to an embodiment of the present application;

[0027] Figure 2 is a schematic diagram of a three-dimensional grid image;

[0028] Figure 3 is a schematic diagram of three-dimensional grid connection;

[0029] FIG4 is a schematic diagram of data storage of a three-dimensional grid;

[0030] FIG5 is a schematic diagram of preprocessing of a two-dimensional curve provided in an embodiment of the present application;

[0031] FIG6 is a schematic diagram of generating displacement coefficients;

[0032] FIG7 is a schematic diagram of an intra-frame encoder;

[0033] FIG8 is a schematic diagram of an intraframe decoder;

[0034] FIG9 is a schematic diagram of an inter-frame encoder;

[0035] FIG10 is a schematic diagram of an inter-frame decoder;

[0036] Figure 11 is a schematic diagram of LOD division;

[0037] FIG12 is a flow chart of a video decoding method according to an embodiment of the present application;

[0038] FIG13 is a flow chart of a video encoding method according to an embodiment of the present application;

[0039] FIG14 is a schematic block diagram of a video decoding device provided in an embodiment of the present application;

[0040] FIG15 is a schematic block diagram of a video encoding apparatus according to an embodiment of the present application;

[0041] FIG16 is a schematic block diagram of an electronic device provided in an embodiment of the present application;

[0042] FIG17 is a schematic block diagram of a video encoding and decoding system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0043] The present application can be applied to the field of image coding and decoding, the field of video coding and decoding, the field of hardware video coding and decoding, the field of dedicated circuit video coding and decoding, the field of real-time video coding and decoding, etc. For example, the solution of the present application can be combined with an audio and video coding standard (AVS), such as the H.264 / audio and video coding (AVC) standard, the H.265 / high efficiency video coding (HEVC) standard, and the H.266 / versatile video coding (VVC) standard. Alternatively, the solution of the present application can be combined with other proprietary or industry standards and operated, and the standards include ITU-TH.261, ISO / IEC MPEG-1 Visual, ITU-TH.262 or ISO / IEC MPEG-2 Visual, ITU-TH.263, ISO / IEC MPEG-4 Visual, ITU-TH.264 (also known as ISO / IEC MPEG-4 AVC), including scalable video coding (SVC) and multi-view video coding (MVC) extensions. It should be understood that the technology of this application is not limited to any specific coding standard or technology.

[0044] The high-degree-of-freedom immersive coding system can be roughly divided into the following links according to the task line: data acquisition, data organization and expression, data encoding and compression, data decoding and reconstruction, data synthesis and rendering, and finally presenting the target data to the user.

[0045] The encoding involved in the embodiment of the present application is mainly video encoding and decoding. For ease of understanding, the video encoding and decoding system involved in the embodiment of the present application is first introduced in conjunction with Figure 1.

[0046] FIG1 is a schematic block diagram of a video encoding and decoding system involved in an embodiment of the present application. It should be noted that FIG1 is only an example, and the video encoding and decoding system of the embodiment of the present application includes but is not limited to that shown in FIG1. ​​As shown in FIG1, the video encoding and decoding system 100 includes an encoding device 110 and a decoding device 120. The encoding device is used to encode (which can be understood as compressing) the video data to generate a code stream, and transmit the code stream to the decoding device. The decoding device decodes the code stream generated by the encoding device to obtain decoded video data.

[0047] The encoding device 110 of the embodiment of the present application can be understood as a device with a video encoding function, and the decoding device 120 can be understood as a device with a video decoding function, that is, the embodiment of the present application includes a wider range of devices for the encoding device 110 and the decoding device 120, such as smartphones, desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, car computers, etc.

[0048] In some embodiments, the encoding device 110 may transmit the encoded video data (eg, a code stream) to the decoding device 120 via a channel 130. The channel 130 may include one or more media and / or devices capable of transmitting the encoded video data from the encoding device 110 to the decoding device 120.

[0049] In one example, the channel 130 includes one or more communication media that enable the encoding device 110 to transmit the encoded video data directly to the decoding device 120 in real time. In this example, the encoding device 110 can modulate the encoded video data according to a communication standard and transmit the modulated video data to the decoding device 120. The communication media includes wireless communication media, such as radio frequency spectrum. Optionally, the communication media may also include wired communication media, such as one or more physical transmission lines.

[0050] In another example, channel 130 includes a storage medium that can store the video data encoded by encoding device 110. The storage medium includes various locally accessible data storage media, such as optical disks, DVDs, and flash memories. In this example, decoding device 120 can retrieve the encoded video data from the storage medium.

[0051] In another example, the channel 130 may include a storage server that can store the video data encoded by the encoding device 110. In this example, the decoding device 120 can download the stored encoded video data from the storage server. Alternatively, the storage server can store the encoded video data and transmit the encoded video data to the decoding device 120, such as a web server (e.g., for a website), a file transfer protocol (FTP) server, etc.

[0052] In some embodiments, the encoding device 110 includes a video encoder 112 and an output interface 113. The output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.

[0053] In some embodiments, the encoding device 110 may further include a video source 111 in addition to the video encoder 112 and the input interface 113 .

[0054] The video source 111 may include at least one of a video acquisition device (eg, a video camera), a video archive, a video input interface, and a computer graphics system, wherein the video input interface is used to receive video data from a video content provider, and the computer graphics system is used to generate video data.

[0055] The video encoder 112 encodes the video data from the video source 111 to generate a bitstream. The video data may include one or more pictures or a sequence of pictures. The bitstream contains the coding information of the picture or picture sequence in the form of a bitstream. The coding information may include the coded picture data and associated data. The associated data may include a sequence parameter set (SPS), a picture parameter set (PPS), and other syntax structures. The SPS may contain parameters that apply to one or more sequences. The PPS may contain parameters that apply to one or more pictures. The syntax structure refers to a set of zero or more syntax elements arranged in a specified order in the bitstream.

[0056] The video encoder 112 transmits the encoded video data directly to the decoding device 120 via the output interface 113. The encoded video data may also be stored in a storage medium or a storage server for subsequent reading by the decoding device 120.

[0057] In some embodiments, the decoding device 120 includes an input interface 121 and a video decoder 122 .

[0058] In some embodiments, the decoding device 120 may further include a display device 123 in addition to the input interface 121 and the video decoder 122 .

[0059] The input interface 121 includes a receiver and / or a modem and can receive the encoded video data via the channel 130 .

[0060] The video decoder 122 is configured to decode the encoded video data to obtain decoded video data, and transmit the decoded video data to the display device 123 .

[0061] The decoded video data is displayed on the display device 123. The display device 123 may be integrated with the decoding device 120 or external to the decoding device 120. The display device 123 may include various display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.

[0062] In addition, Figure 1 is only an example, and the technical solution of the embodiment of the present application is not limited to Figure 1. For example, the technology of the present application can also be applied to unilateral video encoding or unilateral video decoding.

[0063] The embodiments of the present application relate to encoding and decoding of three-dimensional network videos. The following introduces relevant technologies of three-dimensional network videos involved in the embodiments of the present application.

[0064] (1) Three-dimensional grid data format.

[0065] A three-dimensional grid is the surface of a three-dimensional object composed of countless polygons in space. Polygons are composed of vertices and edges. Figure 2 (a) shows a three-dimensional grid image, and Figure 2 (b) shows a partial enlarged view of the three-dimensional grid image. It can be seen that the grid surface is composed of closed polygons.

[0066] A two-dimensional image has information expressed at each pixel point and is distributed regularly, so there is no need to record its position information separately. However, the distribution of vertices in the mesh in three-dimensional space is random and irregular, and the way polygons are formed requires additional regulations. Therefore, it is necessary to record the position of each vertex in space and the connection information of each polygon to fully express a mesh image. As shown in Figure 3, the same number of vertices and vertex positions will form completely different surfaces due to different connection methods.

[0067] In addition to the above information, since 3D meshes are typically encoded using existing video encoding methods for 2D images, it is necessary to convert the 3D mesh from 3D space to a 2D image. UV coordinates define this conversion process. UV coordinates refer to a 2D plane. U represents the horizontal direction, and V represents the vertical direction.

[0068] Similar to 2D images, each position in the image may have corresponding attribute information, typically RGB color values, which reflect the object's color. For 3D meshes, in addition to color, each vertex often has reflectance values, which reflect the surface material. 3D mesh attribute information is stored in 2D images, and the mapping from 2D to 3D is defined by UV coordinates.

[0069] Therefore, 3D mesh data typically includes geometric coordinate information (x, y, z), geometric connectivity information, UV coordinates, and an attribute map. For example, for a mesh image as shown in Figure 4 (a), its corresponding data storage format can be as shown in Figure 4 (b), which stores 3D geometric coordinates, UV coordinates, and connectivity information, and Figure 4 (c) shows the corresponding attribute map.

[0070] (2) Compression of three-dimensional grid.

[0071] The three-dimensional mesh can be encoded and decoded using the Dynamic Mesh Coding (DMC) codec framework of the Moving Picture Experts Group (MPEG).

[0072] Specifically, a base grid and displacement coefficients may be generated through preprocessing, and then the base grid and displacement coefficients may be encoded and decoded using a DMC encoding and decoding framework.

[0073] The preprocessing process of three-dimensional grid can be compared with the same method.

[0074] FIG5 is a schematic diagram of preprocessing of a two-dimensional curve provided in an embodiment of the present application.

[0075] As shown in Figure 5, the original mesh is first downsampled to generate a base mesh with a significantly reduced number of vertices. This base mesh is then subdivided and algorithmically generated, with newly generated vertices inserted along the edges of the base mesh to create a subdivided mesh. Then, as shown in Figure 6, for each vertex in the subdivided mesh, the nearest vertex in the original mesh is found. The vector between the vertex in the subdivided mesh and the nearest vertex in the original mesh is the displaced coefficient. Since the subdivided mesh can be automatically generated at the encoder and decoder end once the subdivision algorithm and number of iterations are determined, after preprocessing, the original mesh only needs to be represented as a simple base mesh and a series of displacement coefficients. This significantly reduces the amount of data and does not affect reconstruction at the decoder end.

[0076] The DMC codec framework can be divided into the DMC intra-frame codec framework and the DMC inter-frame codec framework.

[0077] For DMC intra-frame codec framework:

[0078] At the encoding end, as shown in Figure 7, the basic grid generated by preprocessing is quantized and then encoded using an open source static grid encoder to obtain the basic network code stream.

[0079] The static grid decoder decodes the base grid bitstream to obtain two reconstructed base grids. The displacement coefficients are then updated based on the reconstructed quantized base network. These updated displacement coefficients are then subjected to wavelet transform to obtain the transform coefficients. The transform coefficients are then quantized, two-dimensionally mapped, and encoded using High Efficiency Video Coding (HEVC) to obtain the displacement coefficient bitstream.

[0080] The two-dimensional attribute map is also directly transmitted to the HEVC encoder for encoding. For example, the above-mentioned displacement coefficient code stream is unpacked to obtain the reconstructed quantized transform coefficients. The reconstructed quantized transform coefficients are dequantized to obtain the reconstructed transform coefficients. Next, the reconstructed transform coefficients are subjected to an inverse wavelet transform to obtain the reconstructed displacement coefficients. At the same time, the basic grid code stream is decoded and dequantized to obtain the reconstructed basic grid. Based on the reconstructed basic grid and the reconstructed displacement coefficients, the reconstructed deformed grid is obtained. Based on the reconstructed deformed grid, the two-dimensional attribute map is attribute-converted to obtain a further attribute map. The further attribute map is filled, converted to a color space, and encoded to obtain the attribute code stream.

[0081] The basic network code stream, displacement coefficient code stream and attribute code stream generated above are processed by the composite module and then output as a code stream.

[0082] On the decoding side, as shown in Figure 8, the decoding composite module decodes the code stream output by the encoder to obtain the base grid code stream, the displacement coefficient code stream, and the attribute code stream. The base grid code stream is decoded by the static grid decoder to obtain the reconstructed quantized base grid. This reconstructed quantized base grid is then dequantized to obtain the reconstructed base grid.

[0083] The displacement code stream is decoded, the image is unpacked, inversely quantized, and inversely wavelet transformed to obtain the decoded displacement coefficients. Then, based on the reconstructed base grid and the decoded displacement coefficients, a 3D grid is reconstructed to obtain the decoded 3D grid.

[0084] After video decoding of the attribute code stream, color space conversion is performed to obtain a decoded attribute map.

[0085] For DMC inter-frame codec framework:

[0086] On the encoding side, as shown in Figure 9, due to the use of inter-frame mode, the base mesh does not need to encode its connection information. Only the motion vectors between the vertex geometric coordinates of the current frame and the vertex geometric coordinates of the reference frame need to be encoded. The remaining modules are consistent with the intra-frame encoding process. On the decoding side, as shown in Figure 10, the motion vectors are decoded from the bitstream and combined with the connection information of the reference frame to obtain the base mesh. The remaining modules are consistent with the intra-frame decoding process.

[0087] (3) General test conditions for MPEG DMC.

[0088] 1) There are 2 test conditions:

[0089] Condition 1: all intra, geometry lossy, and attribute lossy.

[0090] Condition 2: random access, lossy geometry, and lossy attributes.

[0091] 2) Common test sequences include Cat1-A, Cat1-B and Cat1-C, all of which contain geometric and color attribute information.

[0092] The current encoding and decoding process of the displacement coefficient is:

[0093] As can be seen from the above, the current MPEG V-DMC obtains displacement coefficients by calculating the vector between the vertex in the subdivided mesh and the nearest vertex in the original mesh. The displacement coefficients are then subjected to wavelet transform and finally mapped to two dimensions before being encoded using existing video coding methods.

[0094] The vertices of the subdivided mesh are organized according to the Level of Details (LOD). Assuming the number of subdivision iterations is 2, as shown in Figure 11, the vertices (circles) in the base mesh are first defined as belonging to LOD0. After the first iteration, a new vertex is inserted into each edge of the base mesh, and the two endpoints of the edge where each new vertex is located are recorded (for example, points a and b will be recorded as the two endpoints of point d). All these new vertices (square points) are defined as LOD1. Then, a second iteration is performed on the new subdivided mesh, and a new vertex is inserted into each edge. The two endpoints of the edge where each new vertex is located are recorded (for example, points a and d will be recorded as the two endpoints of point g; points d and e will be recorded as the two endpoints of point i). All these new vertices (cross points) are defined as LOD2, and so on.

[0095] Wavelet transform is performed based on the vertex order defined by the detail layer. For example, there are three detail layers.

[0096] On the encoding side, perform the following steps a) and b) for each LOD in the order of LOD2 and LOD1:

[0097] a) Prediction: For each vertex of the current LOD, calculate the weighted average of the displacement coefficients of its two endpoints as its predicted value dp i , then calculate the displacement coefficient d of the current vertex i The residual delta between the predicted value and the i For example, as shown in formula (1) and formula (2): dp i =predWeight1*d i1 +predWeight2*d i2 (1) delta i =d i -dp i(2)

[0098] b) Update: For each vertex of the current LOD, use the displacement coefficient residual delta of the current vertex i Update the displacement coefficient values ​​of its two endpoints. For example, as shown in formula (3) and formula (4): i1 =d i1 +updateWeight1*delta i (3) d i2 =d i2 +updateWeight2*delta i (4)

[0099] On the decoding side, perform the following steps a) and b) for each LOD in the order of LOD1 and LOD2:

[0100] a) Update: For each vertex of the current LOD, use the displacement coefficient residual decoding value of the current vertex Update the decoded values ​​of the displacement coefficients at its two endpoints. For example, as shown in formula (5) and formula (6):

[0101] b) Prediction: For each vertex of the current LOD, calculate the weighted average of the displacement coefficient decoding values ​​of its two endpoints as its predicted value Then the residual decoding value of the displacement coefficient of the current vertex is Reconstruct the displacement coefficient of the current vertex with the predicted value For example, as shown in formula (7) and formula (8):

[0102] As can be seen from the above, during the current prediction process, vertices in LOD2 can only be predicted through vertices in LOD1 and / or LOD0. For example, point i is predicted by the displacement coefficients of endpoints d and e, which makes the prediction inaccurate and results in poor encoding and decoding effects of 3D mesh videos.

[0103] In order to solve the above technical problems, the embodiment of the present application predicts the displacement coefficient of the current vertex based on the displacement coefficient of the neighboring points of the current vertex when predicting the vertices in the three-dimensional mesh, which can improve the prediction accuracy of the vertex displacement coefficient and thus improve the encoding and decoding effect of the three-dimensional mesh video.

[0104] 12 , the video decoding method provided in the embodiment of the present application is introduced by taking the decoding end as an example.

[0105] FIG12 is a schematic flow chart of a video decoding method according to an embodiment of the present application. It should be understood that the decoding method can be performed by a decoder. For example, the decoding method can be applied to the intra-frame decoding framework shown in FIG8 or the inter-frame decoding framework shown in FIG10. For ease of description, the following description uses a decoder as an example.

[0106] As shown in FIG12 , the decoding method may include:

[0107] S101 , performing detail layer organization on the vertices of the refined mesh of the current three-dimensional mesh image to obtain a detail layer structure of the vertices of the refined mesh.

[0108] The refined grid is obtained by subdividing the basic grid of the current three-dimensional grid image, and the detail layer structure includes N detail layers, where N is a positive integer.

[0109] As can be seen from the above, the embodiment of the present application decodes the three-dimensional grid video.

[0110] The above-mentioned current three-dimensional grid image can be understood as one or one frame of three-dimensional grid image in the three-dimensional grid video to be decoded.

[0111] In some embodiments, the 3D mesh video may also be referred to as a frame sequence, a current mesh video, or a current 3D mesh video.

[0112] In some embodiments, the current three-dimensional grid image may also be referred to as a three-dimensional grid image currently to be decoded, a three-dimensional grid image to be decoded, a grid image to be decoded, etc.

[0113] In the embodiment of the present application, there is no restriction on the specific decoding mode of the current 3D grid image. In other words, the current 3D grid image can be decoded using the intra-frame decoding mode or the inter-frame decoding mode, and the embodiment of the present application does not impose any restriction on this.

[0114] In one example, when the embodiments of the present application utilize intra-frame prediction mode for decoding, as shown in Figure 8 , the decoder decodes the bitstream of the current 3D mesh image using the decomposition module to obtain a base mesh bitstream and a displacement coefficient bitstream. The decoder then decodes the base network bitstream using a static network decoder to obtain a quantized base network. The quantized base mesh is then dequantized to obtain the base mesh of the current 3D mesh image.

[0115] In one example, when the embodiments of the present application utilize intra-frame prediction mode for decoding, as shown in Figure 9, the decoder decodes the bitstream of the current 3D mesh image using the decomposition module to obtain a bitstream of motion vectors and displacement coefficients for the base mesh. Subsequently, the base mesh of the current 3D mesh image is obtained based on the reference base mesh of the current 3D mesh image and the motion vector of the base mesh of the current 3D mesh image.

[0116] After obtaining the base mesh of the current 3D mesh image based on the above steps, the decoder refines (or subdivides) the base mesh to obtain a refined mesh for the current 3D mesh image. Exemplarily, the decoder determines the number of subdivision iterations of the current 3D mesh network and, based on the number of subdivision iterations, refines the base mesh of the current 3D mesh image to obtain a refined mesh for the current 3D mesh image.

[0117] Next, the decoding end organizes the vertices of the refined mesh of the current three-dimensional mesh image into detail layers to obtain a detail layer structure of the vertices of the refined mesh. For example, as shown in FIG11 , the decoding end first divides the vertices of the base mesh of the current three-dimensional mesh image into LOD0, divides the vertices obtained by the first subdivision iteration of the base mesh into LOD1, and divides the vertices obtained by the second subdivision iteration of the base mesh into LOD2, and so on. The vertices of the refined mesh can be divided into different detail layers LOD to obtain a detail layer structure of the vertices of the refined mesh. The detail layer structure can include one LOD or multiple LODs. For ease of description, the number of detail layers included in the detail layer structure is recorded as N.

[0118] After the decoding end obtains the detail layer structure of the refined grid of the current three-dimensional grid image based on the above steps, it executes the following step S102.

[0119] S102 . For the j-th vertex of the i-th detail layer among the N detail layers, determine M neighboring points of the j-th vertex among the decoded vertices of the refined mesh.

[0120] Wherein, i is a non-negative integer less than N, and j is a non-negative integer.

[0121] In the embodiment of the present application, the decoding end decodes the N detail layers in the opposite order to the encoding end. For example, the encoding end encodes the displacement coefficients of the vertices in each detail layer in the order of LOD2, LOD1, and LOD0. Correspondingly, the decoding end decodes the displacement coefficients of the vertices in each detail layer in the order of LOD0, LOD1, and LOD2.

[0122] In the embodiment of the present application, the decoded detail layer can be understood as a detail layer whose displacement coefficients have been decoded, and the displacement coefficients of each vertex in the detail layer whose displacement coefficients have been decoded have been decoded. The decoded vertex can be understood as a vertex whose displacement coefficients have been decoded.

[0123] The decoding process of the displacement coefficient of each vertex included in each of the N detail layers included in the above-mentioned detail layer structure is basically the same at the decoding end. For the sake of convenience of description, the decoding of the displacement coefficient of the j-th vertex in the i-th detail layer among the N detail layers is taken as an example to illustrate.

[0124] When decoding the displacement coefficient of the j-th vertex, the decoding end first determines the predicted displacement coefficient of the j-th vertex. In some embodiments, the predicted displacement coefficient of the j-th vertex is also called the predicted value of the displacement coefficient of the j-th vertex.

[0125] Currently, when determining the predicted displacement coefficient for the jth vertex, the two endpoints of the edge on which the jth vertex lies in the refined mesh are used as the two predicted points for the jth vertex. Since the displacement coefficients of the two predicted points at these two endpoints have already been decoded, the displacement coefficient of the jth vertex is determined based on the displacement coefficients of these two predicted points. However, in a 3D mesh, the displacement coefficients of adjacent or neighboring vertices are highly correlated. Current prediction methods do not consider the displacement coefficients of the jth vertex's neighbors, resulting in inaccurate prediction of the jth vertex's displacement coefficient and unsatisfactory decoding of 3D mesh videos.

[0126] In order to solve this technical problem, the embodiment of the present application takes into account the displacement coefficients of the neighboring points of the j-th vertex when predicting the displacement coefficient of the j-th vertex, thereby improving the prediction accuracy of the displacement coefficient of the j-th vertex and thus improving the decoding effect of the three-dimensional mesh video.

[0127] In some embodiments, the number of neighboring points of each vertex in the refined mesh is the same, for example, 10.

[0128] In some embodiments, different vertices in the refined mesh may have different numbers of neighboring points. For example, vertex 1 may have 10 neighboring points, and vertex 2 may have 13 neighboring points.

[0129] In the embodiment of the present application, the decoding end determines the M neighboring points of the j-th vertex among the decoded vertices of the refined grid in the following specific ways, but not limited to:

[0130] In a first approach, the decoding end determines M neighboring points of the j-th vertex from decoded vertices in the i-th detail layer and at least one detail layer whose displacement coefficients have been decoded among the N detail layers.

[0131] For example, assuming that the i-th detail layer is LOD2 and the j-th vertex is the third vertex in LOD2, the displacement coefficients of each vertex in detail layers LOD1 and LOD0 have been decoded, as well as the displacement coefficients of the first and second vertices in LOD2.

[0132] In an example, the vertices included in the detail layer LOD1 and the first vertex and the second vertex in LDO2 may be determined as M neighboring points of the j-th vertex.

[0133] In an example, the vertices included in the detail layer LOD0 and the first vertex and the second vertex in LDO2 may be determined as M neighboring points of the j-th vertex.

[0134] In an example, the vertices included in the detail layers LOD0 and LOD1, and the first vertex and the second vertex in LDO2 may be determined as M neighboring points of the j-th vertex.

[0135] In another example, during mesh refinement based on equilateral triangles, the decoder may search for the M vertices with the shortest distances to the j-th vertex among all vertices whose displacement coefficients have been decoded in the N detail layers, and use them as the M neighboring points of the current j-th vertex. For example, in detail layers before the i-th detail layer, such as the i-1-th detail layer and the i-2-th detail layer, and among vertices whose displacement coefficients have been decoded in the i-th detail layer, the M vertices with the shortest distances to the j-th vertex may be determined as the M neighboring points of the j-th vertex.

[0136] For example, assume that the i-th detail layer mentioned above is LOD2, and the j-th vertex is the third vertex in LOD2. At this time, the displacement coefficients of each vertex in the detail layers LOD1 and LOD0 have been decoded, and the displacement coefficients of the first vertex and the second vertex in LDO2 have been decoded. Based on this, the decoding end can determine the distance between each vertex included in the detail layers LOD1 and LOD0 and the j-th vertex, and determine the distance between the first vertex and the second vertex in LDO2 and the j-th vertex respectively. Based on the distance, the M vertices with the smallest distance to the j-th vertex are selected from the vertices included in LOD1 and LOD0, as well as the first vertex and the second vertex in LDO2, as the M neighboring points of the j-th vertex.

[0137] In the second method, the decoding end determines the decoded vertices of the refined mesh that have an edge connection with the j-th vertex and an edge connection step length of 1 as the neighboring points of the j-th vertex.

[0138] In the second implementation, the decoding end searches for vertices included in the thinning network that are directly connected to the j-th vertex and whose displacement coefficients have been decoded, and determines these vertices as neighboring points of the j-th vertex.

[0139] In the embodiment of the present application, in the refined mesh, if there is an edge connection between vertices and the step length of the edge connection is 1, it can be understood that the two vertices are directly connected and there are no other vertices in between. In other words, the step length of the edge connection is 1, which can be understood as a 1-hop connection.

[0140] For example, as shown in Figure 11, assuming that the jth vertex is vertex i, and that it has an edge connection with vertex i with a step length of 1 (where an edge connection step length of 1 can be understood as being directly connected to the i-th vertex with no other vertices in between), the vertices include vertex g, vertex h, vertex d, vertex e, vertex k, and vertex m. Among these six vertices, the vertices whose displacement coefficients have been decoded are determined as the neighboring points of the jth vertex. Assuming that the jth vertex is vertex d, the vertices in the refined mesh that have an edge connection with vertex d with a step length of 1 include vertex g, vertex j, vertex i, and vertex k.

[0141] The M neighboring points of the j-th vertex determined by the second method include the two endpoints corresponding to the j-th vertex. The two endpoints corresponding to the j-th vertex are the two endpoints of the edge where the j-th vertex is located when the j-th vertex is subdivided and inserted. For example, in Figure 11, vertex d is inserted into edge ab during subdivision and insertion, so the two endpoints corresponding to vertex d are the two endpoints of edge ab, namely vertex a and vertex b. For another example, vertex g is inserted into edge ad during subdivision and insertion, so the two endpoints corresponding to vertex g are the two endpoints of edge ad, namely vertex a and vertex d. For another example, vertex i is inserted into edge ed during subdivision and insertion, so the two endpoints corresponding to vertex i are the two endpoints of edge ed, namely vertex e and vertex d.

[0142] In a possible implementation of the second method, the decoding end determines the decoded vertices of the displacement coefficients of the vertices of the refined grid, which have an edge connection with the j-th vertex and an edge connection step of 1, as the neighboring points of the j-th vertex. The specific process may be that the decoding end first determines the vertices of the displacement coefficients of the vertices included in the refined grid, and then searches for the vertices of the displacement coefficients of the vertices that have an edge connection with the j-th vertex and an edge connection step of 1 among the vertices, and then determines these vertices as the M neighboring points of the j-th vertex.

[0143] In another possible implementation method of the second method, the decoding end determines the decoded vertices of the displacement coefficients that are connected to the j-th vertex with an edge and an edge connection step of 1 among the vertices of the refined grid as the neighboring points of the j-th vertex. The specific process can be that the decoding end first determines all vertices that are connected to the j-th vertex with an edge and an edge connection step of 1 among the vertices included in the refined grid, and then searches for the decoded vertices of the displacement coefficients among all vertices that are connected to the j-th vertex with an edge connection step of 1, and determines them as the M neighboring points of the j-th vertex.

[0144] In some embodiments, the decoding end may also use other methods to determine the M neighboring points of the j-th vertex, which is not limited in this embodiment of the present application.

[0145] Based on the above steps, the decoding end determines the M neighboring points of the j-th vertex and then executes the following step S103.

[0146] S103. Determine the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points.

[0147] In an embodiment of the present application, the decoding end determines the M neighboring points of the j-th vertex based on the above steps, and then determines the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of these M neighboring points. Due to the embodiment of the present application, when determining the predicted displacement coefficient of the j-th vertex, the displacement coefficients of the neighboring points of the j-th vertex are taken into account, thereby improving the prediction accuracy of the displacement coefficient.

[0148] The embodiment of the present application does not limit the specific manner in which the decoding end determines the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of M neighboring points.

[0149] In a possible implementation, the average value of the displacement coefficients of the M neighboring points is determined as the predicted displacement coefficient of the j-th vertex.

[0150] In another possible implementation, the weighted average of the displacement coefficients of the M neighboring points is determined as the predicted displacement coefficient of the j-th vertex. In this case, the above S103 includes the following steps:

[0151] S103-A1, determining the weights corresponding to M neighboring points;

[0152] S103-A2: Based on the weights corresponding to the M neighboring points, determine a weighted average of the displacement coefficients of the M neighboring points as the predicted displacement coefficient of the j-th vertex.

[0153] In this implementation, the decoding end first determines the weight corresponding to each of the M neighboring points.

[0154] The embodiment of the present application does not limit the specific manner in which the decoding end determines the weight corresponding to each of the M neighboring points.

[0155] In one example, the weight corresponding to each of the M neighboring points is a default value. That is, the encoder and decoder determine one or more default values ​​as the weight corresponding to each of the M neighboring points.

[0156] In one example, the encoder writes the weight corresponding to each of the M neighboring points into the bitstream, so that the decoder can obtain the weight corresponding to each of the M neighboring points by decoding the bitstream.

[0157] The embodiment of the present application does not limit the specific value of the weight corresponding to each of the M neighboring points.

[0158] In one example, the weight corresponding to each of the M neighboring points is equal. For example, the weight corresponding to each of the M neighboring points is a.

[0159] In one example, weights corresponding to at least two neighboring points among the M neighboring points are unequal.

[0160] For example, among the M neighboring points, the weights corresponding to some neighboring points are unequal, and the weights corresponding to some neighboring points are equal.

[0161] For another example, the weight corresponding to each of the M neighboring points is not equal.

[0162] In one example, the reciprocal of the distance between each of the M neighboring points and the j-th vertex can be used as the weight corresponding to the neighboring point. In this case, the closer the neighboring point is to the j-th vertex, the larger the weight corresponding to the neighboring point, and the farther the neighboring point is from the j-th vertex, the smaller the weight corresponding to the neighboring point.

[0163] Exemplarily, the weights corresponding to the M neighboring points are added together to equal 1.

[0164] Based on the above steps, the decoding end determines the weight corresponding to each of the M neighboring points, and then executes the above S103-A2 step to determine the weighted average of the displacement coefficients of these M neighboring points based on the weight of each of the M neighboring points, and determines the weighted average as the predicted displacement coefficient of the j-th vertex.

[0165] Exemplarily, the decoding end determines the predicted displacement coefficient of the j-th point by the following formula (9):

[0166] in, is the predicted displacement coefficient of the j-th vertex, predWeight m The weight corresponding to the neighboring point with index m among the M neighboring points of the j-th vertex, is the displacement coefficient of the neighboring point with index m among the M neighboring points of the j-th vertex. “*” is the multiplication motion operator.

[0167] In some embodiments, before the above S103, that is, before the decoding end determines the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points, the method of the embodiment of the present application further includes the following steps 1 and 2:

[0168] Step 1: Decode the code stream to obtain the displacement coefficient residual values ​​of all vertices in the i-th detail layer;

[0169] Step 2: Based on the displacement coefficient residual values ​​of all vertices in the i-th detail layer, the displacement coefficients of some or all decoded vertices in the refined mesh are updated.

[0170] In an embodiment of the present application, when decoding the current three-dimensional grid image, the decoding end decodes each of the N refinement layers corresponding to the refined grid of the current three-dimensional grid image one by one. For example, the vertices included in LOD0 are decoded first. Since the vertices in LDO0 are vertices of the basic grid and there are no predicted vertices, the encoding end can encode the geometric information of each vertex in LOD0 (including position coordinates and connection relationships) into the bitstream. In this way, the decoding end can obtain the position coordinates and connection relationships of each vertex in LOD0 by directly decoding the bitstream, and then construct the basic grid of the current three-dimensional grid image. At the same time, the decoding end decodes the displacement coefficient bitstream to obtain the displacement coefficient of each vertex in LOD0. Then, the decoding end decodes each vertex in LDO1 as a displacement coefficient. Specifically, the decoding end first decodes the displacement coefficient bitstream to obtain the residual value of the displacement coefficient of each vertex in LDO1. Next, based on the residual value of the displacement coefficient of each vertex in LDO1, the displacement coefficient of each vertex in LOD0 and the displacement coefficient of the vertex whose displacement coefficient in LDO1 has been decoded are partially or entirely updated. Based on the updated displacement coefficients in LOD0 and LDO1, the displacement coefficients of the remaining vertices in LOD1 are decoded. Then, based on the residual value of the displacement coefficient of each vertex in LDO2, the displacement coefficient of each vertex in LOD0 and LOD1 and the displacement coefficient of the vertex whose displacement coefficient in LDO2 has been decoded are partially or entirely updated. Based on the updated displacement coefficient of each vertex in LOD0, LDO1 and LDO2, the displacement coefficient of the remaining vertices in LOD2 are decoded. By analogy, the displacement coefficients of each vertex in each detail layer of N detail layers can be decoded, thereby obtaining a reconstructed three-dimensional mesh image of the current three-dimensional mesh image.

[0171] The embodiment of the present application does not limit the specific method of updating the displacement coefficients of some or all decoded vertices in the refined grid based on the displacement coefficient residual values ​​of all vertices in the i-th detail layer in the above step 2.

[0172] In some embodiments, the decoding end updates the displacement coefficients of the two endpoints of each vertex in the refined mesh based on the displacement coefficient residual value of each vertex in the i-th detail layer. Taking the j-th vertex in the i-th detail layer as an example, based on the displacement coefficient residual value of the j-th vertex, the displacement coefficients of the two endpoints corresponding to the j-th vertex are updated, wherein the two end vertices corresponding to the j-th vertex are the two endpoints of the edge where the j-th vertex is when the j-th vertex is subdivided and inserted. For example, in Figure 11, vertex d is inserted into edge ab when it is subdivided and inserted, so the two endpoints corresponding to vertex d are the two endpoints of edge ab, namely vertex a and vertex b. For another example, vertex g is inserted into edge ad when it is subdivided and inserted, so the two endpoints corresponding to vertex g are the two endpoints of edge ad, namely vertex a and vertex d.

[0173] In some embodiments, the decoding end updates the displacement coefficients of the M neighboring points of each vertex based on the displacement coefficient residual value of each vertex in the i-th detail layer. Taking the j-th vertex in the i-th detail layer as an example, the displacement coefficients of the M neighboring points of the j-th vertex are updated based on the displacement coefficient residual value of the j-th vertex.

[0174] In an embodiment of the present application, the decoding end updates the displacement coefficients of the two endpoints corresponding to the j-th vertex based on the residual value of the displacement coefficient of the j-th vertex, and updates the displacement coefficients of M neighboring points based on the residual value of the displacement coefficient of the j-th vertex. The specific methods are basically the same. For the sake of convenience of description, the third vertex is used to replace one of the two endpoints corresponding to the j-th vertex or one of the M neighboring points, and the process of updating the displacement coefficient of the third vertex based on the residual value of the displacement coefficient of the j-th vertex is introduced.

[0175] First, the decoding end determines the update weight of the third vertex, then determines the product of the update weight and the residual value of the displacement coefficient of the j-th vertex, and determines the difference between the displacement coefficient of the third vertex and the product as the updated displacement coefficient of the third vertex.

[0176] For example, the decoding end determines the update weight of the first vertex, determines the product of the update weight and the residual value of the displacement coefficient of the j-th vertex, and determines the difference between the displacement coefficient of the first vertex and the product as the updated displacement coefficient of the first vertex.

[0177] For another example, the decoding end determines the update weight of the second vertex, determines the product of the update weight and the residual value of the displacement coefficient of the j-th vertex, and determines the difference between the displacement coefficient of the second vertex and the product as the updated displacement coefficient of the second vertex.

[0178] For another example, for a certain neighboring point among the M neighboring points, the decoding end determines the update weight of the neighboring point, determines the product of the updated weight and the residual value of the displacement coefficient of the j-th vertex, and determines the difference between the displacement coefficient of the neighboring point and the product as the updated displacement coefficient of the neighboring point.

[0179] Based on the above steps, the decoder can update the displacement coefficients of some or all decoded vertices in the refined mesh based on the displacement coefficient residual values ​​of each vertex in the i-th detail layer. Then, based on the updated displacement coefficients, the displacement coefficients of the vertices in the i-th detail layer are introduced.

[0180] For example, taking the j-th vertex in the i-th detail layer as an example, the predicted displacement coefficient of the j-th vertex is determined based on the updated displacement coefficients of the M neighboring points of the j-th vertex. For example, the evaluation value or weighted average of the updated displacement coefficients of the M neighboring points of the j-th vertex is determined as the predicted displacement coefficient of the j-th vertex.

[0181] In some embodiments, the decoding end determines the predicted displacement coefficient of the jth vertex based on the displacement coefficients of the neighboring points of the jth vertex only when the number of neighboring points of the jth vertex is greater than or equal to a preset value.

[0182] For example, the decoder determines the predicted displacement coefficient of the jth vertex based on the displacement coefficients of the jth vertex's neighbors only when the number of neighbors of the jth vertex equals a preset value. In other words, if M equals a preset value, the decoder determines the predicted displacement coefficient of the jth vertex based on the displacement coefficients of M neighboring points.

[0183] The embodiment of the present application does not limit the specific value of the preset value, for example, the preset value is 6. That is, when the number of neighboring points of the j-th vertex is 6, the decoding end determines the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of M neighboring points.

[0184] In some embodiments, if M is not equal to 6, for example, when the number of neighboring points of the j-th vertex is less than 6 or greater than 6, the decoding end determines the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the two endpoints corresponding to the j-th vertex. The specific method for determining the two endpoints corresponding to the j-th vertex can be referred to the description of the above embodiment and is not repeated here.

[0185] After the decoding end determines the predicted displacement coefficient of the j-th vertex based on the above steps, it executes the following step S104.

[0186] S104 : Obtain a reconstructed displacement coefficient of the j-th vertex based on the predicted displacement coefficient of the j-th vertex.

[0187] In the embodiment of the present application, there is no limitation on the specific method of obtaining the reconstructed displacement coefficient of the j-th vertex based on the predicted displacement coefficient of the j-th vertex.

[0188] In some embodiments, the predicted displacement coefficient of the jth vertex can be determined as the reconstructed displacement coefficient of the jth vertex, or the predicted displacement coefficient of the jth vertex can be adjusted based on a preset adjustment method to obtain the reconstructed displacement coefficient of the jth vertex.

[0189] In some embodiments, the decoding end decodes the code stream to obtain the residual value of the displacement coefficient of the j-th vertex, and determines the reconstructed displacement coefficient of the j-th vertex based on the residual value of the displacement coefficient of the j-th vertex and the predicted displacement coefficient of the j-th vertex.

[0190] For example, the decoding end performs video decoding, image unpacking, and inverse quantization on the displacement coefficient code stream of the current three-dimensional grid image to obtain the residual value of the displacement coefficient of the j-th vertex.

[0191] Next, based on the residual value of the displacement coefficient of the j-th vertex and the predicted displacement coefficient of the j-th vertex, the reconstructed displacement coefficient of the j-th vertex is determined.

[0192] For example, the sum of the residual value of the displacement coefficient of the j-th vertex and the predicted displacement coefficient of the j-th vertex is determined as the reconstructed displacement coefficient of the j-th vertex.

[0193] Exemplarily, the decoding end determines the reconstruction displacement coefficient of the j-th vertex by the following formula (10):

[0194] in, is the reconstruction displacement coefficient of the j-th vertex, is the residual value of the displacement coefficient of the j-th vertex, is the predicted displacement coefficient of the j-th vertex.

[0195] The above embodiment describes the process by which a decoder determines the reconstructed value of the displacement coefficient for the jth vertex in the i-th detail layer among N detail layers. Referring to the above steps, the decoder can determine the reconstructed value of the displacement coefficient for each vertex in the i-th detail layer. Furthermore, referring to the above steps, the displacement coefficients for the vertices in each of the N detail layers can be determined, thereby obtaining the displacement coefficients for each vertex in the refined mesh of the current 3D mesh image. Subsequently, based on the displacement coefficients for each vertex, each vertex in the refined mesh is offset to obtain a reconstructed 3D mesh for the current 3D mesh image.

[0196] In the video decoding method provided by the embodiment of the present application, when decoding the geometric information of the current three-dimensional grid image, the decoding end first determines the refined grid of the current three-dimensional grid image, organizes the vertices of the refined grid of the current three-dimensional grid image into detail layers, and obtains the detail layer structure of the vertices of the refined grid, wherein the detail layer structure includes N detail layers. For the j-th vertex of the i-th detail layer in the N detail layers, the M neighboring points of the j-th vertex are determined among the decoded vertices of the refined grid. Then, based on the displacement coefficients of the M neighboring points, the predicted displacement coefficient of the j-th vertex is determined. Finally, based on the predicted displacement coefficient of the j-th vertex, the reconstructed displacement coefficient of the j-th vertex is obtained. In other words, when predicting vertices in the three-dimensional grid, the embodiment of the present application predicts the displacement coefficient of the current vertex based on the displacement coefficients of the neighboring points of the current vertex, which can improve the prediction accuracy of the vertex displacement coefficient and thus improve the encoding and decoding effect of the three-dimensional grid video.

[0197] The above describes the multi-view video decoding method of the present application by taking the decoding end as an example, and the following describes it by taking the encoding end as an example.

[0198] FIG13 is a schematic flow chart of a video encoding method according to an embodiment of the present application. It should be understood that the decoding method can be performed by a decoder. For example, the decoding method can be applied to the intra-frame decoding framework shown in FIG7 or the inter-frame decoding framework shown in FIG9. For ease of description, the following description uses a decoder as an example.

[0199] As shown in FIG13 , the encoding method may include:

[0200] S201 : Perform detail layer organization on the vertices of the refined mesh of the current three-dimensional mesh image to obtain a detail layer structure of the vertices of the refined mesh.

[0201] The detail layer structure includes N detail layers, where N is a positive integer.

[0202] As can be seen from the above, the embodiment of the present application encodes a three-dimensional grid video.

[0203] The above-mentioned current 3D grid image can be understood as one or one frame of 3D grid image in the 3D grid video to be encoded.

[0204] In some embodiments, the 3D mesh video may also be referred to as a frame sequence, a current mesh video, or a current 3D mesh video.

[0205] In some embodiments, the current three-dimensional grid image may also be referred to as a three-dimensional grid image currently to be encoded, a three-dimensional grid image to be encoded, a grid image to be encoded, etc.

[0206] In the embodiment of the present application, there is no restriction on the specific encoding mode of the current 3D grid image. In other words, the current 3D grid image can be encoded using an intra-frame encoding mode or an inter-frame encoding mode, and the embodiment of the present application does not impose any restriction on this.

[0207] In one example, when encoding in intra-frame prediction mode is used in an embodiment of the present application, as shown in FIG7 , the encoder quantizes the base mesh of the current 3D mesh image and then encodes it to obtain a base mesh bitstream. The encoder then decodes the base mesh bitstream using a static network decoder to obtain a quantized base mesh. The quantized base mesh is then dequantized to obtain a reconstructed base mesh for the current 3D mesh image.

[0208] In one example, if the embodiments of the present application utilize intra-frame prediction mode for encoding, as shown in Figure 9, the encoder quantizes the base grid of the current 3D mesh image and then processes it against the reference base grid of the current 3D mesh image to obtain a motion vector for the base grid of the current 3D mesh image. The motion vector is then encoded to obtain a motion bitstream. The decoder then processes the motion bitstream using a motion decoder and the reconstructed base grid to obtain the reconstructed base grid of the current 3D mesh image.

[0209] After obtaining the base mesh of the current 3D mesh image based on the above steps, the encoder refines the base mesh to obtain a refined mesh for the current 3D mesh image. Exemplarily, the encoder determines a number of subdivision iterations for the current 3D mesh network and, based on the number of subdivision iterations, refines the base mesh of the current 3D mesh image to obtain a refined mesh for the current 3D mesh image.

[0210] Next, the encoder organizes the vertices of the refined mesh of the current three-dimensional mesh image into detail layers to obtain a detail layer structure of the vertices of the refined mesh. For example, as shown in FIG11 , the encoder first divides the vertices of the base mesh of the current three-dimensional mesh image into LOD0, divides the vertices obtained by the first subdivision iteration of the base mesh into LOD1, and divides the vertices obtained by the second subdivision iteration of the base mesh into LOD2, and so on. The vertices of the refined mesh can be divided into different detail layers LOD to obtain a detail layer structure of the vertices of the refined mesh. The detail layer structure can include one LOD or multiple LODs. For ease of description, the number of detail layers included in the detail layer structure is recorded as N.

[0211] After the encoder obtains the detail layer structure of the refined grid of the current three-dimensional grid image based on the above steps, it executes the following step S202.

[0212] S202 . For the j-th vertex of the i-th detail layer among the N detail layers, determine M neighboring points of the j-th vertex among the encoded vertices of the refined grid.

[0213] Wherein, i is a non-negative integer less than N, and j is a non-negative integer.

[0214] In the embodiment of the present application, the encoding order of the N detail layers at the encoder is opposite to the encoding order of the N detail layers at the decoder. For example, during encoding, the encoder encodes the displacement coefficients of the vertices in each detail layer in the order of LOD2, LOD1, and LOD0. Correspondingly, the decoder decodes the displacement coefficients of the vertices in each detail layer in the order of LOD0, LOD1, and LOD2.

[0215] The encoding process of the displacement coefficient of each vertex included in each of the N detail layers included in the above-mentioned detail layer structure is basically the same at the encoding end. For the sake of convenience of description, the encoding of the displacement coefficient of the j-th vertex in the i-th detail layer among the N detail layers is taken as an example to illustrate.

[0216] When encoding the displacement coefficient of the j-th vertex, the encoding end first determines the predicted displacement coefficient of the j-th vertex. In some embodiments, the predicted displacement coefficient of the j-th vertex is also called the predicted value of the displacement coefficient of the j-th vertex.

[0217] Currently, when determining the predicted displacement coefficient for the jth vertex, the two endpoints of the edge on which the jth vertex lies in the refined mesh are used as the two predicted points for the jth vertex. Since the displacement coefficients of the two predicted points at these two endpoints are already encoded, the displacement coefficient of the jth vertex is determined based on the displacement coefficients of these two predicted points. However, in a 3D mesh, the displacement coefficients of adjacent or neighboring vertices are highly correlated. Current prediction methods do not consider the displacement coefficients of the jth vertex's neighbors, resulting in inaccurate prediction of the jth vertex's displacement coefficient and unsatisfactory encoding of 3D mesh videos.

[0218] In order to solve this technical problem, the embodiment of the present application takes into account the displacement coefficients of the neighboring points of the j-th vertex when predicting the displacement coefficient of the j-th vertex, thereby improving the prediction accuracy of the displacement coefficient of the j-th vertex and thus improving the encoding effect of the three-dimensional mesh video.

[0219] In this embodiment of the present application, the encoder determines the M neighboring points of the j-th vertex among the encoded vertices of the refined grid in the following specific ways, but not limited to:

[0220] In a first approach, the encoder determines M neighboring points of the jth vertex from among the encoded vertices in the i-th detail layer and at least one detail layer whose displacement coefficients are encoded among the N detail layers.

[0221] For example, assuming that the i-th detail layer is LOD2 and the j-th vertex is the third vertex in LOD2, the displacement coefficients of each vertex in detail layers LOD1 and LOD0 have been encoded, as well as the displacement coefficients of the first and second vertices in LOD2.

[0222] In an example, the vertices included in the detail layer LOD1 and the first vertex and the second vertex in LDO2 may be determined as M neighboring points of the j-th vertex.

[0223] In an example, the vertices included in the detail layer LOD0 and the first vertex and the second vertex in LDO2 may be determined as M neighboring points of the j-th vertex.

[0224] In an example, the vertices included in the detail layers LOD0 and LOD1, and the first vertex and the second vertex in LDO2 may be determined as M neighboring points of the j-th vertex.

[0225] In another example, during mesh refinement based on equilateral triangles, the encoder may search for M vertices with the shortest distance to the j-th vertex among all vertices whose displacement coefficients have been encoded in the current N detail layers, and use them as the M neighboring points of the current j-th vertex. For example, in detail layers before the i-th detail layer, such as the i-1-th detail layer and the i-2-th detail layer, and among vertices whose displacement coefficients have been encoded in the i-th detail layer, the encoder may determine the M vertices with the shortest distance to the j-th vertex as the M neighboring points of the j-th vertex.

[0226] For example, assume that the i-th detail layer is LOD1, and the j-th vertex is the third vertex in LOD1. At this time, the displacement coefficients of each vertex in the detail layer LOD2 have been encoded, and the displacement coefficients of the first vertex and the second vertex in LDO1 have been encoded. Based on this, the encoding end can determine the distance between each vertex included in the detail layer LOD2 and the j-th vertex, and determine the distance between the first vertex and the second vertex in LDO1 and the j-th vertex respectively. Based on the distance, the M vertices with the smallest distance to the j-th vertex are selected from the vertices included in LOD2 and the first vertex and the second vertex in LDO1 as the M neighboring points of the j-th vertex.

[0227] In the second method, the encoder determines the vertices of the refined mesh that are connected to the j-th vertex by an edge and whose displacement coefficients have been encoded with an edge connection step size of 1 as the neighboring points of the j-th vertex.

[0228] In the second implementation, the encoder searches for vertices included in the refinement network whose displacement coefficients are directly connected to the j-th vertex and have been encoded, and determines these vertices as neighboring points of the j-th vertex.

[0229] In the embodiment of the present application, in the refined mesh, if there is an edge connection between vertices and the step length of the edge connection is 1, it can be understood that the two vertices are directly connected and there are no other vertices in between. In other words, the step length of the edge connection is 1, which can be understood as a 1-hop connection.

[0230] For example, as shown in Figure 11, assuming that the j-th vertex is vertex i, and that it is connected to vertex i by an edge with a step length of 1 (where the step length of 1 can be understood as being directly connected to the i-th vertex with no other vertices in between), the vertices include vertex g, vertex h, vertex d, vertex e, vertex k, and vertex m. Among these six vertices, the vertices with encoded displacement coefficients are determined as the neighboring points of the j-th vertex. Assuming that the j-th vertex is vertex d, the vertices in the refined mesh that are connected to vertex d by an edge with a step length of 1 include vertex g, vertex j, vertex i, and vertex k.

[0231] The M neighboring points of the j-th vertex determined by the second method include the two endpoints corresponding to the j-th vertex. The two endpoints corresponding to the j-th vertex are the two endpoints of the edge where the j-th vertex is located when the j-th vertex is subdivided and inserted. For example, in Figure 11, vertex d is inserted into edge ab during subdivision and insertion, so the two endpoints corresponding to vertex d are the two endpoints of edge ab, namely vertex a and vertex b. For another example, vertex g is inserted into edge ad during subdivision and insertion, so the two endpoints corresponding to vertex g are the two endpoints of edge ad, namely vertex a and vertex d. For another example, vertex i is inserted into edge ed during subdivision and insertion, so the two endpoints corresponding to vertex i are the two endpoints of edge ed, namely vertex e and vertex d.

[0232] In a possible implementation of the second method, the encoding end determines the displacement coefficient-encoded vertices in the refined grid that have an edge connection with the j-th vertex and an edge connection step of 1 as the neighboring points of the j-th vertex. The specific process may be that the encoding end first determines the vertices with encoded displacement coefficients among the vertices included in the refined grid, and then searches for the vertices with encoded displacement coefficients that have an edge connection with the j-th vertex and an edge connection step of 1 among the vertices, and then determines these vertices as the M neighboring points of the j-th vertex.

[0233] In another possible implementation method of the second method, the encoding end determines the displacement coefficient encoded vertices that have an edge connection with the j-th vertex and an edge connection step length of 1 among the vertices of the refined grid as the neighboring points of the j-th vertex. The specific process can be that the encoding end first determines all vertices that have an edge connection with the j-th vertex and an edge connection step length of 1 among the vertices included in the refined grid, and then searches for the displacement coefficient encoded vertices among all vertices that have an edge connection with the j-th vertex and an edge connection step length of 1, and determines them as the M neighboring points of the j-th vertex.

[0234] In some embodiments, the encoding end may also use other methods to determine the M neighboring points of the j-th vertex, which is not limited in this embodiment of the present application.

[0235] After the encoder determines the M neighboring points of the j-th vertex based on the above steps, it executes the following step S203.

[0236] S203 : Determine the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points.

[0237] In an embodiment of the present application, the encoding end determines the M neighboring points of the j-th vertex based on the above steps, and then determines the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of these M neighboring points. Due to the embodiment of the present application, when determining the predicted displacement coefficient of the j-th vertex, the displacement coefficients of the neighboring points of the j-th vertex are taken into account, thereby improving the prediction accuracy of the displacement coefficient.

[0238] The embodiment of the present application does not limit the specific method for the encoder to determine the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of M neighboring points.

[0239] In a possible implementation, the average value of the displacement coefficients of the M neighboring points is determined as the predicted displacement coefficient of the j-th vertex.

[0240] In another possible implementation, the weighted average of the displacement coefficients of the M neighboring points is determined as the predicted displacement coefficient of the j-th vertex. In this case, the above S203 includes the following steps:

[0241] S203-A1, determining the weights corresponding to the M neighboring points;

[0242] S203-A2: Based on the weights corresponding to the M neighboring points, determine a weighted average of the displacement coefficients of the M neighboring points as the predicted displacement coefficient of the j-th vertex.

[0243] In this implementation, the encoder first determines the weight corresponding to each of the M neighboring points.

[0244] The embodiment of the present application does not limit the specific method by which the encoder determines the weight corresponding to each of the M neighboring points.

[0245] In one example, the weight corresponding to each of the M neighboring points is a default value. That is, the encoding end and the encoding end determine one or more default values ​​as the weight corresponding to each of the M neighboring points.

[0246] In one example, the encoder writes the weight corresponding to each of the M neighboring points into the bitstream. In this way, the encoder can obtain the weight corresponding to each of the M neighboring points by encoding the bitstream.

[0247] The embodiment of the present application does not limit the specific value of the weight corresponding to each of the M neighboring points.

[0248] In one example, the weight corresponding to each of the M neighboring points is equal. For example, the weight corresponding to each of the M neighboring points is a.

[0249] In one example, weights corresponding to at least two neighboring points among the M neighboring points are unequal.

[0250] For example, among the M neighboring points, the weights corresponding to some neighboring points are unequal, and the weights corresponding to some neighboring points are equal.

[0251] For another example, the weight corresponding to each of the M neighboring points is not equal.

[0252] In one example, the reciprocal of the distance between each of the M neighboring points and the j-th vertex can be used as the weight corresponding to the neighboring point. In this case, the closer the neighboring point is to the j-th vertex, the larger the weight corresponding to the neighboring point, and the farther the neighboring point is from the j-th vertex, the smaller the weight corresponding to the neighboring point.

[0253] Exemplarily, the weights corresponding to the M neighboring points are added together to equal 1.

[0254] Based on the above steps, the encoding end determines the weight corresponding to each of the M neighboring points, and then executes the above S203-A2 step to determine the weighted average of the displacement coefficients of these M neighboring points based on the weight of each of the M neighboring points, and determines the weighted average as the predicted displacement coefficient of the j-th vertex.

[0255] Exemplarily, the encoding end determines the predicted displacement coefficient of the j-th point through the following formula (11).

[0256] Among them, dp j is the predicted displacement coefficient of the j-th vertex, predWeightm The weight corresponding to the neighboring point with index m among the M neighboring points of the j-th vertex, *d jm is the displacement coefficient of the neighboring point with index m among the M neighboring points of the j-th vertex. “*” is the multiplication motion operator.

[0257] In some embodiments, the encoder determines the predicted displacement coefficient of the jth vertex based on the displacement coefficients of the neighboring points of the jth vertex only when the number of neighboring points of the jth vertex is greater than or equal to a preset value.

[0258] For example, the encoder determines the predicted displacement coefficient of the jth vertex based on the displacement coefficients of the jth vertex's neighbors only when the number of neighbors of the jth vertex equals a preset value. In other words, if M equals a preset value, the encoder determines the predicted displacement coefficient of the jth vertex based on the displacement coefficients of M neighboring points.

[0259] The embodiment of the present application does not limit the specific value of the preset value, for example, the preset value is 6. That is, when the number of neighboring points of the j-th vertex is 6, the encoder determines the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points.

[0260] In some embodiments, if M is not equal to 6, for example, when the number of neighboring points of the j-th vertex is less than 6 or greater than 6, the encoder determines the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the two endpoints corresponding to the j-th vertex. The specific method for determining the two endpoints corresponding to the j-th vertex can be found in the description of the above embodiment and is not repeated here.

[0261] After the encoder determines the predicted displacement coefficient of the j-th vertex based on the above steps, it executes the following step S204.

[0262] S204 : Obtain a displacement coefficient residual value of the j-th vertex based on the predicted displacement coefficient of the j-th vertex.

[0263] In the embodiment of the present application, there is no limitation on the specific method of obtaining the displacement coefficient residual value of the j-th vertex based on the predicted displacement coefficient of the j-th vertex.

[0264] In some embodiments, the encoding end obtains a displacement coefficient residual value of the j-th vertex based on the displacement coefficient of the j-th vertex and the predicted displacement coefficient of the j-th vertex.

[0265] For example, the difference between the displacement coefficient of the j-th vertex and the predicted displacement coefficient of the j-th vertex is determined as the displacement coefficient residual value of the j-th vertex.

[0266] For example, the encoder determines the displacement coefficient residual value of the j-th vertex by the following formula (12): j=d j -dp j (12)

[0267] Among them, d j is the displacement coefficient of the j-th vertex, delta j is the displacement coefficient residual value of the j-th vertex, dp j is the predicted displacement coefficient of the j-th vertex.

[0268] In some embodiments, after the above S204, that is, after the encoder obtains the displacement coefficient residual value of the j-th vertex based on the predicted displacement coefficient of the j-th vertex, the method of the embodiment of the present application further includes the following step 3:

[0269] Step 3: Based on the displacement coefficient residual values ​​of all vertices in the i-th detail layer, the displacement coefficients of some or all encoded vertices in the refined mesh are updated.

[0270] The above embodiment describes the process of determining the displacement coefficient residual value of the jth vertex in the i-th detail layer among N detail layers by the encoder. Referring to the above steps, the encoder can determine the displacement coefficient residual value of each vertex in the i-th detail layer.

[0271] In an embodiment of the present application, when encoding a current three-dimensional mesh image, the encoder encodes each of the N refinement layers corresponding to the refined mesh of the current three-dimensional mesh image one by one. For example, the encoder encodes each vertex in LDO2 as a displacement coefficient. Specifically, based on the above steps, the encoder determines the residual value of the displacement coefficient of each vertex in LDO2. Then, based on the residual value of the displacement coefficient of each vertex in LDO2, the encoder updates some or all of the displacement coefficients of each vertex in LOD1 and LOD0, as well as the displacement coefficients of the vertices in LDO2 whose displacement coefficients have been encoded. Based on the updated displacement coefficients in LOD1 and LDO0, the encoder encodes the displacement coefficients of the remaining vertices in LOD2. Then, the encoder determines the residual value of the displacement coefficient of each vertex in LDO1. Then, based on the residual value of the displacement coefficient of each vertex in LDO1, the encoder updates some or all of the displacement coefficients of each vertex in LOD0, as well as the displacement coefficients of the vertices in LDO1 whose displacement coefficients have been encoded. Based on the updated displacement coefficients of each vertex in LOD0 and LDO1, the displacement coefficients of the remaining vertices in LOD1 are encoded. Similarly, the displacement coefficients of each vertex in each detail layer of N detail layers can be encoded to obtain the displacement coefficient code stream of the current 3D mesh image.

[0272] The embodiment of the present application does not limit the specific method of updating the displacement coefficients of some or all encoded vertices in the refined grid based on the displacement coefficient residual values ​​of all vertices in the i-th detail layer in the above step 3.

[0273] In some embodiments, the encoding end updates the displacement coefficients of the two endpoints of each vertex in the refined mesh based on the displacement coefficient residual value of each vertex in the i-th detail layer. Taking the j-th vertex in the i-th detail layer as an example, based on the displacement coefficient residual value of the j-th vertex, the displacement coefficients of the two endpoints corresponding to the j-th vertex are updated, wherein the two end vertices corresponding to the j-th vertex are the two endpoints of the edge where the j-th vertex is when the j-th vertex is subdivided and inserted. For example, in Figure 11, vertex d is inserted into edge ab when it is subdivided and inserted, so the two endpoints corresponding to vertex d are the two endpoints of edge ab, namely vertex a and vertex b. For another example, vertex g is inserted into edge ad when it is subdivided and inserted, so the two endpoints corresponding to vertex g are the two endpoints of edge ad, namely vertex a and vertex d.

[0274] In some embodiments, the encoder updates the displacement coefficients of the M neighboring points of each vertex based on the displacement coefficient residual value of each vertex in the i-th detail layer. Taking the j-th vertex in the i-th detail layer as an example, the displacement coefficients of the M neighboring points of the j-th vertex are updated based on the displacement coefficient residual value of the j-th vertex.

[0275] In an embodiment of the present application, the encoding end updates the displacement coefficients of the two endpoints corresponding to the j-th vertex based on the residual value of the displacement coefficient of the j-th vertex, and updates the displacement coefficients of M neighboring points based on the residual value of the displacement coefficient of the j-th vertex. The specific methods are basically the same. For the sake of convenience of description, the third vertex is used to replace one of the two endpoints corresponding to the j-th vertex or one of the M neighboring points, and the process of updating the displacement coefficient of the third vertex based on the residual value of the displacement coefficient of the j-th vertex is introduced.

[0276] First, the encoding end determines the update weight of the third vertex, then determines the product of the update weight and the residual value of the displacement coefficient of the j-th vertex, and determines the difference between the displacement coefficient of the third vertex and the product as the updated displacement coefficient of the third vertex.

[0277] For example, the encoding end determines the update weight of the first vertex, determines the product of the update weight and the residual value of the displacement coefficient of the j-th vertex, and determines the difference between the displacement coefficient of the first vertex and the product as the updated displacement coefficient of the first vertex.

[0278] For another example, the encoding end determines the update weight of the second vertex, determines the product of the update weight and the residual value of the displacement coefficient of the j-th vertex, and determines the difference between the displacement coefficient of the second vertex and the product as the updated displacement coefficient of the second vertex.

[0279] For another example, for a certain neighboring point among the M neighboring points, the encoding end determines the update weight of the neighboring point, determines the product of the updated weight and the residual value of the displacement coefficient of the j-th vertex, and determines the difference between the displacement coefficient of the neighboring point and the product as the updated displacement coefficient of the neighboring point.

[0280] Based on the above steps, the encoder can update the displacement coefficients of some or all of the vertices whose displacement coefficients have been encoded in the refined mesh based on the displacement coefficient residual value of each vertex in the i-th detail layer. Then, based on the updated displacement coefficients, the displacement coefficients of the vertices in the i-th detail layer are introduced.

[0281] For example, taking the j-th vertex in the i-th detail layer as an example, the predicted displacement coefficient of the j-th vertex is determined based on the updated displacement coefficients of the M neighboring points of the j-th vertex. For example, the evaluation value or weighted average of the updated displacement coefficients of the M neighboring points of the j-th vertex is determined as the predicted displacement coefficient of the j-th vertex.

[0282] The above embodiment describes the process by which the encoder determines the displacement coefficient residual value for the jth vertex in the i-th detail layer among N detail layers. Referring to the above steps, the encoder can determine the displacement coefficient residual value for each vertex in the N detail layers. Subsequently, the encoder quantizes, image-packs, and video-encodes the displacement coefficient residual values ​​for each vertex in the N detail layers to obtain a displacement coefficient bitstream for the current 3D mesh image.

[0283] The video encoding method provided by the embodiment of the present application is that when the encoding end encodes the current three-dimensional grid image, it first determines the refined grid of the current three-dimensional grid image, organizes the vertices of the refined grid of the current three-dimensional grid image into detail layers, and obtains the detail layer structure of the vertices of the refined grid, and the detail layer structure includes N detail layers. For the j-th vertex of the i-th detail layer in the N detail layers, the M neighboring points of the j-th vertex are determined among the encoded vertices of the refined grid. Then, based on the displacement coefficients of the M neighboring points, the predicted displacement coefficient of the j-th vertex is determined. Finally, based on the predicted displacement coefficient of the j-th vertex, the displacement coefficient residual value of the j-th vertex is obtained. That is to say, when the embodiment of the present application predicts the vertices in the three-dimensional grid, the displacement coefficient of the current vertex is predicted based on the displacement coefficients of the neighboring points of the current vertex, which can improve the prediction accuracy of the vertex displacement coefficient, thereby improving the encoding effect of the three-dimensional grid video.

[0284] It should be understood that Figures 12 to 13 are merely examples of the present application and should not be understood as limiting the present application.

[0285] The preferred embodiments of the present application are described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the specific details in the above embodiments. Within the technical concept of the present application, a variety of simple modifications can be made to the technical solution of the present application, and these simple modifications all fall within the scope of protection of the present application. For example, the various specific technical features described in the above specific embodiments can be combined in any suitable manner unless there is any contradiction. In order to avoid unnecessary repetition, the present application will not further explain various possible combinations. For another example, the various different embodiments of the present application can also be arbitrarily combined, and as long as they do not violate the ideas of the present application, they should also be regarded as the contents disclosed in the present application.

[0286] It should also be understood that in the various method embodiments of the present application, the size of the sequence numbers of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. In addition, in the embodiments of the present application, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships can exist. Specifically, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this application generally indicates that the related objects before and after are in an "or" relationship.

[0287] The above text describes in detail a method embodiment of the present application in conjunction with Figures 12 to 13 , and the following text describes in detail a device embodiment of the present application in conjunction with Figures 14 to 15 .

[0288] FIG14 is a schematic block diagram of a video decoding device provided in an embodiment of the present application. The video decoding device 10 is applied to the above-mentioned video decoder.

[0289] As shown in FIG14 , the video decoding apparatus 10 includes:

[0290] a refinement unit 11 configured to organize vertices of a refined mesh of a current three-dimensional mesh image into detail layers to obtain a detail layer structure of the vertices of the refined mesh, wherein the refined mesh is obtained by subdividing a base mesh of the current three-dimensional mesh image, and the detail layer structure includes N detail layers, where N is a positive integer;

[0291] a neighboring point determination unit 12 configured to determine, for a j-th vertex of an i-th detail layer among the N detail layers, M neighboring points of the j-th vertex from decoded vertices of the refined mesh, where the decoded vertices are vertices whose displacement coefficients have been decoded, i is a non-negative integer less than N, j is a non-negative integer, and M is a positive integer;

[0292] A prediction unit 13 is configured to determine a predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points;

[0293] The reconstruction unit 14 is configured to obtain a reconstruction displacement coefficient of the j-th vertex based on the predicted displacement coefficient of the j-th vertex.

[0294] In some embodiments, the neighboring point determination unit 12 is specifically configured to determine M neighboring points of the j-th vertex from among the decoded vertices in the i-th detail layer and at least one detail layer whose displacement coefficients have been decoded in the N detail layers.

[0295] In some embodiments, the neighboring point determination unit 12 is specifically configured to determine, among the vertices of the refined mesh, the decoded vertices of the displacement coefficients that are connected to the j-th vertex with an edge and an edge connection step of 1 as neighboring points of the j-th vertex.

[0296] In some embodiments, the prediction unit 13 is specifically configured to determine the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points if M is equal to a preset value.

[0297] In some embodiments, the prediction unit 13 is configured to determine an average value of the displacement coefficients of the M neighboring points as the predicted displacement coefficient of the j-th vertex.

[0298] In some embodiments, the prediction unit 13 is used to determine the weights corresponding to the M neighboring points; based on the weights corresponding to the M neighboring points, determine the weighted average of the displacement coefficients of the M neighboring points as the predicted displacement coefficient of the j-th vertex.

[0299] In some embodiments, the weight corresponding to each of the M neighboring points is equal.

[0300] In some embodiments, the weights corresponding to at least two neighboring points among the M neighboring points are unequal.

[0301] In some embodiments, the reconstruction unit 14 is specifically used to decode the code stream to obtain the residual value of the displacement coefficient of the j-th vertex; based on the residual value of the displacement coefficient of the j-th vertex and the predicted displacement coefficient of the j-th vertex, obtain the reconstructed displacement coefficient of the j-th vertex.

[0302] In some embodiments, the reconstruction unit 14 is specifically configured to determine the sum of the displacement coefficient residual value of the j-th vertex and the predicted displacement coefficient of the j-th vertex as the reconstructed displacement coefficient of the j-th vertex.

[0303] In some embodiments, the prediction unit 13 is further used to decode the code stream to obtain the displacement coefficient residual values ​​of all vertices in the i-th detail layer before determining the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points; update the displacement coefficients of some or all decoded vertices in the refined grid based on the displacement coefficient residual values ​​of all vertices in the i-th detail layer; and determine the predicted displacement coefficient of the j-th vertex based on the updated displacement coefficients of the M neighboring points.

[0304] In some embodiments, the prediction unit 13 is specifically used to update the displacement coefficients of the two endpoints corresponding to the j-th vertex based on the residual value of the displacement coefficient of the j-th vertex, and the two end vertices corresponding to the j-th vertex are the two endpoints of the edge where the j-th vertex is located when the j-th vertex is subdivided and inserted.

[0305] In some embodiments, the prediction unit 13 is specifically configured to update the displacement coefficients of the M neighboring points for the j-th vertex based on the displacement coefficient residual value of the j-th vertex.

[0306] In some embodiments, the prediction unit 13 is specifically used to determine the update weight of the third vertex; determine the product of the update weight and the residual value of the displacement coefficient of the j-th vertex; and determine the difference between the displacement coefficient of the third vertex and the product as the updated displacement coefficient of the third vertex; wherein the third vertex is any one of the two endpoints corresponding to the j vertices, or any one of the M neighboring points.

[0307] It should be understood that the device embodiments and method embodiments may correspond to each other, and similar descriptions may refer to the method embodiments. To avoid repetition, no further description is given here. Specifically, the device 10 shown in FIG14 can execute the decoding method of the decoding end of the embodiment of the present application, and the aforementioned and other operations and / or functions of the various units in the device 10 are respectively for implementing the corresponding processes in each method such as the decoding method of the decoding end. For the sake of brevity, no further description is given here.

[0308] FIG15 is a schematic block diagram of a video encoding device provided in an embodiment of the present application, which is applied to the above-mentioned encoder.

[0309] As shown in FIG15 , the video encoding apparatus 20 may include:

[0310] a refinement unit 21 configured to organize vertices of a refined mesh of a current 3D mesh image into detail layers to obtain a detail layer structure of the vertices of the refined mesh, wherein the refined mesh is obtained by subdividing a base mesh of the current 3D mesh image, and the detail layer structure includes N detail layers, where N is a positive integer;

[0311] a neighboring point determination unit 22 configured to determine, for a j-th vertex of an i-th detail layer among the N detail layers, M neighboring points of the j-th vertex from among the encoded vertices of the refined mesh, where the encoded vertices are vertices whose displacement coefficients have been encoded, i is a non-negative integer less than N, j is a non-negative integer, and M is a positive integer;

[0312] A prediction unit 23 is configured to determine a predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points;

[0313] The residual unit 24 is configured to obtain a residual value of the displacement coefficient of the j-th vertex based on the predicted displacement coefficient of the j-th vertex.

[0314] In some embodiments, the neighboring point determination unit 22 is specifically configured to determine M neighboring points of the j-th vertex from among the decoded vertices in the i-th detail layer and at least one detail layer in which the displacement coefficients of the N detail layers have been encoded.

[0315] In some embodiments, the neighboring point determination unit 22 is specifically configured to determine, among the vertices of the refined grid, displacement coefficient-encoded vertices that have an edge connection with the j-th vertex and an edge connection step length of 1 as neighboring points of the j-th vertex.

[0316] In some embodiments, the prediction unit 23 is specifically configured to determine a predicted displacement coefficient point of the j-th vertex based on the displacement coefficients of the M neighboring points if the M is equal to a preset value.

[0317] In some embodiments, the prediction unit 23 is specifically configured to determine an average value of the displacement coefficients of the M neighboring points as the predicted displacement coefficient of the j-th vertex.

[0318] In some embodiments, the prediction unit 23 is specifically used to determine the weights corresponding to the M neighboring points; based on the weights corresponding to the M neighboring points, determine the weighted average of the displacement coefficients of the M neighboring points as the predicted displacement coefficient of the j-th vertex.

[0319] In some embodiments, the weight corresponding to each of the M neighboring points is equal.

[0320] In some embodiments, the weights corresponding to at least two neighboring points among the M neighboring points are unequal.

[0321] In some embodiments, the residual unit 24 is specifically configured to obtain a residual value of the displacement coefficient of the j-th vertex based on the displacement coefficient of the j-th vertex and the predicted displacement coefficient of the j-th vertex.

[0322] In some embodiments, the residual unit 24 is specifically configured to determine a difference between the displacement coefficient of the j-th vertex and the predicted displacement coefficient of the j-th vertex as a residual value of the displacement coefficient of the j-th vertex.

[0323] In some embodiments, the residual unit 24, after determining the displacement coefficient residual value of each vertex in the i-th detail layer, is also used to update the displacement coefficients of some or all encoded vertices in the refined grid based on the displacement coefficient residual values ​​of all vertices in the i-th detail layer.

[0324] In some embodiments, the residual unit 24 is specifically used to update the displacement coefficients of the two endpoints corresponding to the j-th vertex based on the residual value of the displacement coefficient of the j-th vertex. The two end vertices corresponding to the j-th vertex are the two endpoints of the edge where the j-th vertex is located when the j-th vertex is subdivided and inserted.

[0325] In some embodiments, the residual unit 24 is specifically configured to update the displacement coefficients of the M neighboring points for the j-th vertex based on the residual value of the displacement coefficient of the j-th vertex.

[0326] In some embodiments, the residual unit 24 is specifically used to determine the update weight of the neighboring point; determine the product of the updated weight and the residual value of the displacement coefficient of the j-th vertex; and determine the sum of the displacement coefficient of the neighboring point and the product as the updated displacement coefficient of the neighboring point; wherein the neighboring point is any one of the two endpoints corresponding to the j vertices, or any one of the M neighboring points.

[0327] It should be understood that the device embodiments and the method embodiments may correspond to each other, and similar descriptions may refer to the method embodiments. To avoid repetition, they will not be described here. Specifically, the device 20 shown in Figure 15 may correspond to the corresponding subject in the encoding method of the encoding end of the embodiment of the present application, and the aforementioned and other operations and / or functions of each unit in the device 20 are respectively for implementing the corresponding processes in each method such as the encoding method of the encoding end. For the sake of brevity, they will not be described here.

[0328] The above describes the apparatus and system of the embodiment of the present application from the perspective of functional units in conjunction with the accompanying drawings. It should be understood that the functional unit can be implemented in the form of hardware, can be implemented by instructions in the form of software, or can be implemented by a combination of hardware and software units. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software units in the decoding processor. Optionally, the software unit can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps in the above method embodiment in conjunction with its hardware.

[0329] FIG16 is a schematic block diagram of an electronic device provided in an embodiment of the present application.

[0330] As shown in FIG16 , the electronic device 30 may be a video encoder or a video decoder as described in an embodiment of the present application. The electronic device 30 may include:

[0331] The memory 33 and the processor 32 are configured to store a computer program 34 and transmit the program code 34 to the processor 32. In other words, the processor 32 can call and run the computer program 34 from the memory 33 to implement the method in the embodiment of the present application.

[0332] For example, the processor 32 may be configured to execute the steps of the above method according to the instructions in the computer program 34 .

[0333] In some embodiments of the present application, the processor 32 may include but is not limited to:

[0334] General-purpose processor, Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.

[0335] In some embodiments of the present application, the memory 33 includes but is not limited to:

[0336] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).

[0337] In some embodiments of the present application, the computer program 34 may be divided into one or more units, which are stored in the memory 33 and executed by the processor 32 to implement the method provided by the present application. The one or more units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program 34 in the electronic device 30.

[0338] As shown in FIG16 , the electronic device 30 may further include:

[0339] The transceiver 33 may be connected to the processor 32 or the memory 33 .

[0340] The processor 32 can control the transceiver 33 to communicate with other devices. Specifically, it can send information or data to other devices or receive information or data sent by other devices. The transceiver 33 may include a transmitter and a receiver. The transceiver 33 may further include an antenna, and the number of antennas may be one or more.

[0341] It should be understood that the various components in the electronic device 30 are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.

[0342] FIG17 is a schematic block diagram of a video encoding and decoding system provided in an embodiment of the present application.

[0343] As shown in Figure 17, the video encoding and decoding system 40 may include: a video encoder 41 and a video decoder 42, wherein the video encoder 41 is used to execute the video encoding method involved in the embodiment of the present application, and the video decoder 42 is used to execute the video decoding method involved in the embodiment of the present application.

[0344] The present application also provides a computer storage medium having a computer program stored thereon, which, when executed by a computer, enables the computer to perform the method of the above-mentioned method embodiment. In other words, the present application also provides a computer program product containing instructions, which, when executed by a computer, enables the computer to perform the method of the above-mentioned method embodiment.

[0345] The present application also provides a code stream, which is generated according to the above encoding method.

[0346] When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid state drive (SSD)).

[0347] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0348] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the unit is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0349] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, the functional units in the various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0350] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A video decoding method, characterized in that: include: Organizing the vertices of the refined mesh of the current three-dimensional mesh image in detail layers to obtain a detail layer structure of the vertices of the refined mesh, wherein the refined mesh is obtained by subdividing the base mesh of the current three-dimensional mesh image, and the detail layer structure includes N detail layers, where N is a positive integer; For a j-th vertex of an i-th detail layer among the N detail layers, determine M neighboring points of the j-th vertex from decoded vertices of the refined mesh, wherein the decoded vertices are vertices whose displacement coefficients have been decoded, i is a non-negative integer less than N, j is a non-negative integer, and M is a positive integer; Determine a predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points; Based on the predicted displacement coefficient of the j-th vertex, a reconstructed displacement coefficient of the j-th vertex is obtained.

2. The method according to claim 1, characterized in that The step of determining M neighboring points of the j-th vertex among the decoded vertices of the refined grid comprises: Among the decoded vertices in the i-th detail layer and in at least one detail layer whose displacement coefficients have been decoded among the N detail layers, M neighboring points of the j-th vertex are determined.

3. The method according to claim 1, characterized in that The step of determining M neighboring points of the j-th vertex among the decoded vertices of the refined grid comprises: Among the vertices of the refined mesh, the vertices with decoded displacement coefficients that are edge-connected to the j-th vertex and whose edge connection step length is 1 are determined as neighboring points of the j-th vertex.

4. The method according to claim 1, characterized in that: The step of determining the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points includes: If M is equal to a preset value, the predicted displacement coefficient of the j-th vertex is determined based on the displacement coefficients of the M neighboring points.

5. The method according to any one of claims 1 to 4, characterized in that: The step of determining the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points includes: The average value of the displacement coefficients of the M neighboring points is determined as the predicted displacement coefficient of the j-th vertex.

6. The method according to any one of claims 1 to 4, characterized in that: The step of determining the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points includes: Determine the weights corresponding to the M neighboring points; Based on the weights corresponding to the M neighboring points, a weighted average of the displacement coefficients of the M neighboring points is determined as the predicted displacement coefficient of the j-th vertex.

7. The method according to claim 6, characterized in that The weight corresponding to each of the M neighboring points is equal.

8. The method according to claim 6, characterized in that The weights corresponding to at least two neighboring points among the M neighboring points are not equal.

9. The method according to any one of claims 1 to 4, characterized in that: The step of obtaining the reconstruction displacement coefficient of the j-th vertex based on the predicted displacement coefficient of the j-th vertex includes: Decoding the bitstream to obtain the residual value of the displacement coefficient of the j-th vertex; Based on the displacement coefficient residual value of the j-th vertex and the predicted displacement coefficient of the j-th vertex, a reconstructed displacement coefficient of the j-th vertex is obtained.

10. The method according to claim 9, characterized in that The step of obtaining the reconstruction displacement coefficient of the j-th vertex based on the displacement coefficient residual value of the j-th vertex and the predicted displacement coefficient of the j-th vertex comprises: The sum of the displacement coefficient residual value of the j-th vertex and the predicted displacement coefficient of the j-th vertex is determined as the reconstructed displacement coefficient of the j-th vertex.

11. The method according to claim 9, characterized in that Before determining the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points, the method further includes: Decoding the bitstream to obtain displacement coefficient residual values ​​of all vertices in the i-th detail layer; Based on the displacement coefficient residual values ​​of all vertices in the i-th detail layer, updating the displacement coefficients of some or all decoded vertices in the refined grid; The step of determining the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points includes: Based on the updated displacement coefficients of the M neighboring points, the predicted displacement coefficient of the j-th vertex is determined.

12. The method according to claim 11, characterized in that The updating of the displacement coefficients of some or all decoded vertices of the refined mesh based on the displacement coefficient residual values ​​of all vertices in the i-th detail layer comprises: For the j-th vertex, the displacement coefficients of the two endpoints corresponding to the j-th vertex are updated based on the residual value of the displacement coefficient of the j-th vertex. The two end vertices corresponding to the j-th vertex are the two endpoints of the edge where the j-th vertex is located when the j-th vertex is subdivided and inserted.

13. The method according to claim 11, characterized in that The updating of the displacement coefficients of some or all decoded vertices in the refined mesh based on the displacement coefficient residual values ​​of all vertices in the i-th detail layer includes: For the j-th vertex, the displacement coefficients of the M neighboring points are updated based on the displacement coefficient residual value of the j-th vertex.

14. The method according to claim 12 or 13, characterized in that Based on the displacement coefficient residual value of the j-th vertex, the displacement coefficient of the third vertex is updated, including: Determining an update weight of the third vertex; Determine the product of the update weight and the displacement coefficient residual value of the j-th vertex; Determine the difference between the displacement coefficient of the third vertex and the product as the updated displacement coefficient of the third vertex; The third vertex is any one of the two endpoints corresponding to the j vertices, or any one of the M adjacent points.

15. A video encoding method, characterized in that: include: Organizing the vertices of the refined mesh of the current three-dimensional mesh image in detail layers to obtain a detail layer structure of the vertices of the refined mesh, wherein the refined mesh is obtained by subdividing the base mesh of the current three-dimensional mesh image, and the detail layer structure includes N detail layers, where N is a positive integer; For a j-th vertex of an i-th detail layer among the N detail layers, determine M neighboring points of the j-th vertex among the encoded vertices of the refined grid, wherein the encoded vertices are vertices whose displacement coefficients have been encoded, i is a non-negative integer less than N, j is a non-negative integer, and M is a positive integer; Determine a predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points; Based on the predicted displacement coefficient of the j-th vertex, a displacement coefficient residual value of the j-th vertex is obtained.

16. The method according to claim 15, characterized in that The step of determining M neighboring points of the j-th vertex among the encoded vertices of the refined grid comprises: Among the decoded vertices in the i-th detail layer and at least one detail layer in which the displacement coefficients have been encoded among the N detail layers, M neighboring points of the j-th vertex are determined.

17. The method according to claim 16, characterized in that The step of determining M neighboring points of the j-th vertex among the encoded vertices of the refined grid comprises: Among the vertices of the refined mesh, the displacement coefficient encoded vertices which are connected to the j-th vertex by an edge and whose edge connection step length is 1 are determined as neighboring points of the j-th vertex.

18. The method according to claim 15, characterized in that The step of determining the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points includes: If M is equal to a preset value, the predicted displacement coefficient of the j-th vertex is determined based on the displacement coefficients of the M neighboring points.

19. The method according to any one of claims 15 to 18, characterized in that: The step of determining the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points includes: The average value of the displacement coefficients of the M neighboring points is determined as the predicted displacement coefficient of the j-th vertex.

20. The method according to any one of claims 15 to 18, characterized in that: The step of determining the predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points includes: Determine the weights corresponding to the M neighboring points; Based on the weights corresponding to the M neighboring points, a weighted average of the displacement coefficients of the M neighboring points is determined as the predicted displacement coefficient of the j-th vertex.

21. The method according to claim 20, characterized in that The weight corresponding to each of the M neighboring points is equal.

22. The method according to claim 20, characterized in that The weights corresponding to at least two neighboring points among the M neighboring points are not equal.

23. The method according to any one of claims 15 to 18, characterized in that: The step of obtaining the displacement coefficient residual value of the j-th vertex based on the predicted displacement coefficient of the j-th vertex includes: Based on the displacement coefficient of the j-th vertex and the predicted displacement coefficient of the j-th vertex, a residual value of the displacement coefficient of the j-th vertex is obtained.

24. The method according to claim 23, characterized in that The step of obtaining the displacement coefficient residual value of the j-th vertex based on the displacement system of the j-th vertex and the predicted displacement coefficient of the j-th vertex comprises: The difference between the displacement coefficient of the j-th vertex and the predicted displacement coefficient of the j-th vertex is determined as the displacement coefficient residual value of the j-th vertex.

25. The method according to claim 15, characterized in that After determining the displacement coefficient residual value of each vertex in the i-th detail layer, the method further includes: Based on the displacement coefficient residual values ​​of all vertices in the i-th detail layer, the displacement coefficients of some or all encoded vertices in the refined mesh are updated.

26. The method according to claim 25, characterized in that The updating of the displacement coefficients of some or all of the encoded vertices in the refined mesh based on the displacement coefficient residual values ​​of all vertices in the i-th detail layer includes: For the j-th vertex, the displacement coefficients of the two endpoints corresponding to the j-th vertex are updated based on the residual value of the displacement coefficient of the j-th vertex. The two end vertices corresponding to the j-th vertex are the two endpoints of the edge where the j-th vertex is located when the j-th vertex is subdivided and inserted.

27. The method according to claim 25, characterized in that The updating of the displacement coefficients of some or all of the encoded vertices in the refined mesh based on the displacement coefficient residual values ​​of all vertices in the i-th detail layer includes: For the j-th vertex, the displacement coefficients of the M neighboring points are updated based on the displacement coefficient residual value of the j-th vertex.

28. The method according to claim 26 or 27, characterized in that Based on the displacement coefficient residual value of the j-th vertex, the displacement coefficient of the third vertex is updated, including: Determining an update weight of the third vertex; Determine the product of the update weight and the displacement coefficient residual value of the j-th vertex; Determine the sum of the displacement coefficient of the third vertex and the product as the updated displacement coefficient of the third vertex; The third vertex is any one of the two endpoints corresponding to the j vertices, or any one of the M adjacent points.

29. A video decoding device, characterized in that: include: A refinement unit, configured to organize the vertices of the refined mesh of the current three-dimensional mesh image into detail layers to obtain a detail layer structure of the vertices of the refined mesh, wherein the refined mesh is obtained by subdividing the base mesh of the current three-dimensional mesh image, and the detail layer structure includes N detail layers, where N is a positive integer; a neighboring point determination unit, configured to determine, for a j-th vertex of an i-th detail layer among the N detail layers, M neighboring points of the j-th vertex from decoded vertices of the refined grid, wherein the decoded vertices are vertices whose displacement coefficients have been decoded, i is a non-negative integer less than N, j is a non-negative integer, and M is a positive integer; A prediction unit, configured to determine a predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points; The reconstruction unit is used to obtain the reconstruction displacement coefficient of the j-th vertex based on the predicted displacement coefficient of the j-th vertex.

30. A video encoding device, characterized in that: include: A refinement unit, configured to organize the vertices of the refined mesh of the current three-dimensional mesh image into detail layers to obtain a detail layer structure of the vertices of the refined mesh, wherein the refined mesh is obtained by subdividing the base mesh of the current three-dimensional mesh image, and the detail layer structure includes N detail layers, where N is a positive integer; a neighboring point determination unit, configured to determine, for a j-th vertex of an i-th detail layer among the N detail layers, M neighboring points of the j-th vertex among the encoded vertices of the refined grid, wherein the encoded vertices are vertices whose displacement coefficients have been encoded, wherein i is a non-negative integer less than N, j is a non-negative integer, and M is a positive integer; A prediction unit, configured to determine a predicted displacement coefficient of the j-th vertex based on the displacement coefficients of the M neighboring points; The residual unit is used to obtain a residual value of the displacement coefficient of the j-th vertex based on the predicted displacement coefficient of the j-th vertex.

31. An electronic device, characterized in that: including a processor and a memory; The memory shown is used to store computer programs; The processor is used to call and run the computer program stored in the memory to implement the method described in any one of claims 1 to 14 or 15 to 28 above.

32. A video encoding and decoding system, characterized in that: include: Video encoders and video decoders; The video decoder is used to implement the method described in any one of claims 1 to 14 above; The video encoder is used to implement the method described in any one of claims 15 to 28.

33. A computer-readable storage medium, characterized in that: For storing computer programs; The computer program enables a computer to execute the method according to any one of claims 1 to 14 or 15 to 28 above.