Coding method, decoding method, system, and related device

By directly encoding the actual motion trend information of moving vertices in a 3D mesh sequence, the problem of low compression rate and high bit rate in existing 3D mesh encoding and decoding schemes is solved, achieving more efficient data compression and transmission.

WO2025260841A9PCT designated stage Publication Date: 2026-04-16HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2026-04-16

AI Technical Summary

Technical Problem

Existing 3D mesh encoding and decoding schemes have low compression rates when dealing with vertices with non-uniform motion, resulting in large data volume, high bit rate, and high encoding complexity, making it difficult to effectively compress and transmit 3D mesh data.

Method used

By acquiring the actual motion trend information of the moving vertices in the 3D mesh sequence and directly encoding this trend information instead of calculating the prediction residual, the encoding is performed using the actual motion trend information, which reduces the range of quantized values, reduces the number of bits, improves the compression rate, and reduces the bit rate.

Benefits of technology

It improves the compression rate of 3D meshes, shortens the bitstream length, reduces encoding complexity and bit rate, enhances transmission and storage efficiency, and reduces screen stuttering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025081405_16042026_PF_FP_ABST
    Figure CN2025081405_16042026_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a coding method, a decoding method, a system, and a related device. In the method, when a three-dimensional mesh is coded, actual motion trend information of a three-dimensional mesh to be coded in a three-dimensional mesh sequence may be acquired, the actual motion trend information indicating an actual motion trend of a motion vertex that performs non-uniform motion in the three-dimensional mesh sequence in a slot corresponding to said three-dimensional mesh; and coded data of said three-dimensional mesh is acquired on the basis of the actual motion trend information, so as to code the three-dimensional mesh. The method can achieve a better coding and decoding effect.
Need to check novelty before this filing date? Find Prior Art

Description

Encoding and decoding method, system and related device

[0001] The present application claims priority to the Chinese patent application No. 202410799179.5, filed on June 19, 2024, and entitled "Encoding and decoding method, system and related device", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present application relates to the technical field of three-dimensional mesh, in particular to an encoding and decoding method, system and related device. BACKGROUND

[0003] Three-dimensional mesh (for example, dynamic mesh) can be used to express volumetric video, digital human, computer graphics (CG) content, etc. The three-dimensional mesh is a three-dimensional structure composed of a series of vertices and corresponding edges, for example, the three-dimensional structure can include a plurality of planar patches, each patch being composed of a group of vertices and corresponding edges. The data of the three-dimensional mesh includes the coordinates of the vertices and the connection relationship between the vertices. The number of vertices of the three-dimensional mesh is large, so that the data of the three-dimensional mesh includes a lot of coordinate information of the vertices, resulting in a large amount of data of the three-dimensional mesh. Therefore, before transmitting the three-dimensional mesh, it is generally necessary to compress and encode the three-dimensional mesh.

[0004] However, for three-dimensional meshes with non-uniform motion vertices, the current encoding and decoding scheme needs to be improved. SUMMARY

[0005] The present application provides an encoding and decoding method, system and related device to improve the current encoding and decoding scheme of three-dimensional mesh.

[0006] In a first aspect, the present application provides an encoding method, comprising: obtaining actual motion trend information of a first three-dimensional mesh to be encoded in a three-dimensional mesh sequence, the actual motion trend information indicating an actual motion trend of a motion vertex moving at a non-uniform speed in the first three-dimensional mesh corresponding time slot in the three-dimensional mesh sequence; obtaining encoding data of the first three-dimensional mesh based on the actual motion trend information; and encoding the encoding data into a code stream corresponding to the three-dimensional mesh sequence.

[0007] The motion trend information can be motion acceleration.

[0008] In the embodiments of the present application, when encoding the current frame in the three-dimensional mesh sequence, the encoding end can obtain the encoding data of the current frame based on actual motion trend information of the current frame, wherein the actual motion trend information indicates actual motion trend of a motion vertex in the three-dimensional mesh sequence in a corresponding time slot of the current frame. In the case that the motion vertex in the three-dimensional mesh sequence moves at a non-uniform speed (for example, moves at a variable speed or rotates by itself), although the vertex position of the motion vertex in the three-dimensional mesh sequence can change, causing the three-dimensional mesh to deform, the actual motion trend of the motion vertex in the corresponding time slot of each frame in the three-dimensional mesh sequence is less changed. By encoding the current frame in the three-dimensional mesh sequence based on the actual motion trend information of the current frame, better encoding effect can be obtained.

[0009] In a possible implementation manner, the obtaining the encoding data of the first three-dimensional mesh based on the actual motion trend information comprises: obtaining prediction motion trend information of the first three-dimensional mesh, the prediction motion trend information indicating predicted motion trend of the motion vertex in a corresponding time slot of the first three-dimensional mesh; and obtaining the encoding data based on the actual motion trend information and the prediction motion trend information.

[0010] In the embodiments of the present application, when the motion vertex in the three-dimensional mesh to be encoded moves at a non-uniform speed, although the difference between the uniform speeds of the same motion vertex in the corresponding time slots of different frames is large, the actual motion trend information (for example, offset acceleration) of the same motion vertex in the corresponding time slots of different frames is close, so that the prediction motion information (including the prediction motion trend information) predicted by the present application for the motion vertex is closer to the actual motion information (including the actual motion trend information) of the motion vertex, thereby improving the accuracy of the prediction information (here, the prediction motion trend) of the motion vertex of the current frame.

[0011] In a possible implementation manner, the encoding data comprises a residual error between the actual motion trend information and the prediction motion trend information.

[0012] In the embodiment of the present application, in the case that the current frame includes a scene in which the motion vertex performs non-uniform motion, the encoding end of the present application can encode the residual error between the actual motion trend information and the predicted motion trend information of the current frame to realize the encoding of the current frame. In the case that the three-dimensional mesh performs self-deformation or deformation (for example, cloth deformation) due to non-uniform motion, the vertex position of the current frame predicted based on the historical frame in the related art is not accurate. However, the actual motion trend information of the same motion vertex in different frames in the three-dimensional mesh sequence is almost unchanged. Therefore, the predicted motion trend information of the current frame obtained by the present application is closer to the actual motion trend information of the current frame, thereby reducing the residual error to be encoded, reducing the data amount after compression, and improving the data compression rate. Since the prediction residual error of each motion vertex of the three-dimensional mesh is small, the interval formed by the quantization result of the residual error between the predicted motion trend information and the actual motion trend information of each motion vertex of the three-dimensional mesh is small, so that the number of binary bits occupied by the encoding data of the residual error is smaller, thereby shortening the length of the encoding data of the three-dimensional mesh, and further shortening the length of the code stream corresponding to the three-dimensional mesh sequence and reducing the code rate.

[0013] In a possible implementation manner, the encoding data includes the actual motion trend information.

[0014] In the related art, when the three-dimensional mesh is encoded, it is assumed that the three-dimensional mesh performs uniform motion, and the motion speed of the vertex of the three-dimensional mesh is unchanged between adjacent frames. Based on the assumption, the vertex position of the current frame is predicted by using the vertex position of the historical frame, and then the residual error (also referred to as position residual error) between the predicted position and the actual position of the vertex of the current frame is encoded to realize the encoding of the current frame. However, when the three-dimensional mesh is a model of a flexible body or an elastic body or other easily deformable three-dimensional object, in the case that the vertex in the three-dimensional mesh performs non-uniform motion (for example, the vertex performs self-rotation motion, or the vertex performs variable speed motion, or the three-dimensional mesh performs self-deformation in the three-dimensional mesh sequence), the predicted vertex position of the current frame based on the assumption that the vertex performs uniform motion is not accurate, so that the predicted position of the vertex deviates from the actual position of the vertex too much. In the case that the vertex performs non-uniform motion in the three-dimensional mesh sequence, the predicted positions of a large number of vertices are all based on the assumption that the vertex performs uniform motion, resulting in a large difference in the sizes of the prediction residual errors of a large number of vertices in the three-dimensional mesh, and a low compression rate of the three-dimensional mesh.

[0015] When encoding the data, the data can be quantized, and then the quantized values are converted into binary sequences for encoding. The higher the frequency of the quantized values, the smaller the interval formed by the quantized values (for example, [-1, 1]). Moreover, the higher the frequency of the quantized values, the fewer the number of bits of the binary converted from the quantized values. Therefore, the encoding scheme of the three-dimensional mesh in the related technology causes the interval formed by the quantized values of a large number of vertices in the three-dimensional mesh to be larger (for example, [-10, 10]), thereby causing the number of binary bits occupied by the encoded data of the current frame to be larger, and further causing the length of the bitstream of the three-dimensional mesh sequence to be longer and the code rate to be higher.

[0016] In the embodiment of the present application, when encoding the current frame in the three-dimensional mesh sequence, the encoding end does not need to calculate the prediction residual of the vertex (for example, the residual between the predicted position and the actual position of the vertex) or encode the residual, but directly encodes the actual motion trend information of the current frame to obtain the encoded data of the current frame, and encodes the encoded data into the bitstream corresponding to the three-dimensional mesh sequence. Then in the case where the motion vertices move at a non-uniform speed in the three-dimensional mesh sequence, a large number of motion vertices in the current frame can have similar actual motion trend (for example, acceleration) information. When encoding the actual motion trend information of a large number of motion vertices in the current frame, the interval formed by the quantized values of the actual motion trend information can be smaller, the number of binary bits occupied by the encoded data of the current frame can be smaller, thereby improving the compression rate, shortening the length of the bitstream corresponding to the three-dimensional mesh sequence, and reducing the code rate. Moreover, compared with the scheme of encoding the residual between the actual motion trend information and the predicted motion trend information of the current frame, the embodiment of the present application only needs to calculate and encode the actual motion trend information of the current frame, and the encoding complexity is lower, thereby shortening the encoding delay.

[0017] In a possible implementation, the bitstream includes a first syntax structure corresponding to the first three-dimensional mesh, and the first syntax structure includes an identifier of the first three-dimensional mesh.

[0018] In a possible implementation, the first syntax structure further includes information corresponding to the motion vertex, and the information of the motion vertex includes: a first residual of the motion vertex on a first component; a second residual of the motion vertex on a second component; and a third residual of the motion vertex on a third component.

[0019] The actual motion information of the motion vertex can include actual component information of the motion vertex on three components, respectively, actual component information on a first component, actual component information on a second component, and actual component information on a third component.

[0020] The predicted motion information of the motion vertex can include predicted component information of the motion vertex on three components, respectively, predicted component information on a first component, predicted component information on a second component, and predicted component information on a third component.

[0021] The three components are three components in a three-dimensional coordinate system, which is a Cartesian coordinate system or a cylindrical coordinate system.

[0022] When the three-dimensional coordinate system is a Cartesian coordinate system, the three components can be an x component, a y component, and a z component, respectively.

[0023] When the three-dimensional coordinate system is a cylindrical coordinate system, the three components can be a radius component, a height component, and an angle (for example, a deflection angle) component, respectively.

[0024] The first residual is a residual between actual component information (for example, actual motion trend information (for example, actual motion acceleration)) of the motion vertex on the first component and predicted component information (for example, predicted motion trend information (for example, predicted motion acceleration)) on the first component.

[0025] The second residual is a residual between actual component information and predicted component information on the second component. For example, the actual component information is actual motion velocity, and the predicted component information is predicted motion velocity; or the actual component information is actual motion acceleration, and the predicted component information is predicted motion acceleration.

[0026] The third residual is a residual between actual component information and predicted component information on the third component. For example, the actual component information is actual motion velocity, and the predicted component information is predicted motion velocity; or the actual component information is actual motion acceleration, and the predicted component information is predicted motion acceleration.

[0027] At least one of the first residual, the second residual, and the third residual is a residual between predicted motion acceleration and actual motion acceleration on the corresponding component.

[0028] In a possible implementation, the first syntax structure further includes: a vertex identifier of the motion vertex.

[0029] When a part of vertices in the first three-dimensional mesh is encoded, the first syntax structure can include a vertex identifier of the part of vertices, so as to facilitate the decoding end to locate the three types of residuals of which motion vertex.

[0030] In a possible implementation, the obtaining the prediction motion trend information of the first three-dimensional mesh comprises: obtaining N sets of reconstruction data corresponding to N second three-dimensional meshes that have been encoded in the three-dimensional mesh sequence, N being an integer greater than or equal to a set value; and determining the prediction motion trend information based on the N sets of reconstruction data.

[0031] The N second three-dimensional meshes and the N sets of reconstruction data are in a one-to-one correspondence.

[0032] In the embodiments of the present application, because the decoding end cannot obtain the original data of each frame in the three-dimensional mesh sequence, only the reconstruction data of each frame can be obtained by decoding. In order to ensure that the prediction motion trend information calculated by the encoding end and the decoding end for the same frame is the same, the encoding end can obtain the reconstruction data of at least three historical frames that have been encoded in the three-dimensional mesh sequence when obtaining the prediction motion trend information of the current frame, and determine the prediction motion trend information of the current frame based on the reconstruction data (rather than the original data of the at least three historical frames), so as to ensure that the prediction motion trend information of the current frame determined by the encoding end is consistent with the prediction motion trend information of the current frame obtained by the decoding end, thereby ensuring the accuracy of the reconstruction data of the three-dimensional mesh sequence, and ensuring that the prediction motion information calculated by the encoding end and the decoding end for the current frame is consistent, so as to facilitate the accurate decoding of the current frame by the decoding end, improve the decoding accuracy, ensure that the position difference of the motion vertex between the reconstructed three-dimensional mesh and the original three-dimensional mesh is small, and improve the quality of the reconstructed three-dimensional mesh.

[0033] In a possible implementation, the obtaining the actual motion trend information of the first three-dimensional mesh to be encoded in the three-dimensional mesh sequence comprises: determining the actual motion information of the first three-dimensional mesh based on the original data of the first three-dimensional mesh and the original data of a plurality of third three-dimensional meshes that have been encoded in the three-dimensional mesh sequence.

[0034] In the present application, when the encoding end obtains the actual motion trend information of the current frame, in order to ensure the accuracy of the actual motion trend information, the actual motion trend information of the current frame can be determined based on the original data of the current frame and the original data of at least two historical frames that have been encoded, thereby ensuring the accuracy of the actual motion information calculated by the encoding end.

[0035] In a possible implementation, the obtaining the N sets of reconstruction data corresponding to the N second three-dimensional meshes in the three-dimensional mesh sequence comprises: obtaining N sets of encoded data corresponding to the N second three-dimensional meshes; obtaining N sets of predicted motion trend information of the N second three-dimensional meshes, the N sets of predicted motion trend information being N predicted motion trends of the motion vertices in N time slots corresponding to the N second three-dimensional meshes; and decoding the N sets of encoded data based on the N sets of predicted motion trend information to obtain the N sets of reconstruction data. The N second three-dimensional meshes and the N sets of predicted motion information are in a one-to-one correspondence.

[0036] In this application, the process of obtaining, at the encoding end, reconstruction data of at least three historical frames that have been encoded before a current frame is the same as the process of obtaining, at the decoding end, reconstruction data of a frame that has been encoded in the three-dimensional mesh sequence, so as to ensure that the reconstruction data of the N second three-dimensional meshes obtained at the encoding and decoding sides is consistent.

[0037] In a possible implementation, the decoding the N sets of encoded data based on the N sets of predicted motion trend information to obtain the N sets of reconstruction data of the N second three-dimensional meshes comprises: decoding the N sets of encoded data to obtain N sets of residuals between N sets of actual motion trend information of the N second three-dimensional meshes and the N sets of predicted motion trend information, the N sets of actual motion trend information comprising N actual motion trends of the motion vertices in N time slots corresponding to the N sets of second three-dimensional meshes; determining, based on the N sets of predicted motion trend information and the N sets of residuals, the N sets of actual motion trend information (substantially the N sets of actual motion trend information obtained by reconstruction, also referred to as N sets of reconstruction motion information) of the N second three-dimensional meshes, the N sets of actual motion trend information being N reconstruction information of the N actual motion trends of the motion vertices in N time slots corresponding to the N second three-dimensional meshes; and determining the N sets of reconstruction data based on the N sets of actual motion trend information and reconstruction data of S fourth three-dimensional meshes corresponding to each of the N second three-dimensional meshes, the S fourth three-dimensional meshes being three-dimensional meshes that have been decoded in the three-dimensional mesh sequence, and S being an integer greater than or equal to 2.

[0038] The N second three-dimensional meshes and the N sets of reconstruction motion information are in a one-to-one correspondence.

[0039] The N second three-dimensional meshes and the N sets of actual motion information are also in a one-to-one correspondence.

[0040] In the embodiment of the present application, the encoder can use the N residual errors of the N historical frames decoded from the encoding data of the N historical frames and the N sets of predicted motion trend information of the N historical frames to reconstruct N sets of actual motion trend information (expressed as reconstructed motion information) of the N historical frames; then, the encoder can determine the reconstructed data of a historical frame in the N historical frames based on the reconstructed data of at least two frames that have been encoded before the historical frame and the reconstructed actual motion trend information of the historical frame. The reconstructed data can include the reconstructed positions of the motion vertices of the current frame and optionally the connection relationship (for example, the topology of the current frame) between the motion vertices, so that the reconstruction of each historical frame in the N historical frames can be realized. In the case that the three-dimensional mesh sequence can be deformed, the offset accelerations of the same motion vertices (referring to the same vertex identification or the same vertex ordering position) of different frames are close, so that in the deformation scenario of the three-dimensional mesh sequence, the correct decoding of the three-dimensional mesh can be realized at a reduced code rate, the decoding data amount is reduced due to the small residual error, and the transmission efficiency and storage efficiency are improved.

[0041] In a possible implementation, at least one of the first residual error, the second residual error and the third residual error is a residual error of motion acceleration.

[0042] In a possible implementation, the first component is an X component, the second component is a Y component, and the third component is a Z component.

[0043] In a possible implementation, the first component is an angle component, the second component is a height component, and the third component is a radius component.

[0044] In a possible implementation, the first three-dimensional mesh, the N second three-dimensional meshes and the plurality of third three-dimensional meshes have the same topology.

[0045] The topology of each three-dimensional mesh represents the connection relationship of the vertices in the three-dimensional mesh, the connection relationship of the vertices of each three-dimensional mesh in the three-dimensional mesh sequence remains unchanged, and the number and identification of the vertices of each three-dimensional mesh are the same, so that the change of the motion trend information of the same vertex in the three-dimensional mesh sequence can be tracked, and the predicted motion trend information and the actual motion trend information of the motion vertex can be determined.

[0046] In a possible implementation, the display order of the N second three-dimensional meshes in the three-dimensional mesh sequence is before the display order of the first three-dimensional mesh in the three-dimensional mesh sequence.

[0047] For example, the N second three-dimensional meshes can be decoded and then rendered and displayed, and the display order can be before the rendering image of the first three-dimensional mesh.

[0048] In a possible implementation, the display order of the plurality of third three-dimensional meshes in the sequence of three-dimensional meshes is before the display order of the first three-dimensional mesh in the sequence of three-dimensional meshes.

[0049] For example, the plurality of third three-dimensional meshes can be decoded and rendered to be displayed, and the display order of the plurality of third three-dimensional meshes can be before the rendering image of the first three-dimensional mesh.

[0050] The N second three-dimensional meshes and the plurality of third three-dimensional meshes can be the same, the number of the third three-dimensional meshes is also N, and N is greater than or equal to 3. Or the N second three-dimensional meshes and the plurality of third three-dimensional meshes are completely different three-dimensional meshes, or are not completely the same three-dimensional meshes. The number of the plurality of third three-dimensional meshes is at least two.

[0051] In a second aspect, the present application provides a decoding method, comprising: obtaining a code stream corresponding to a sequence of three-dimensional meshes; obtaining, from the code stream, encoding data of a first three-dimensional mesh to be decoded in the sequence of three-dimensional meshes; obtaining actual motion trend information of the first three-dimensional mesh based on the encoding data, the actual motion trend information indicating an actual motion trend of a motion vertex in the sequence of three-dimensional meshes that moves at a non-uniform speed in a time slot corresponding to the first three-dimensional mesh; and obtaining reconstruction data of the first three-dimensional mesh based on the actual motion trend information.

[0052] In the embodiments of the present application, when a current frame in the sequence of three-dimensional meshes is decoded, a decoding end can obtain encoding data of the current frame from a code stream, and obtain actual motion trend information of the current frame based on the encoding data, the actual motion trend information indicating an actual motion trend of a motion vertex in the sequence of three-dimensional meshes that moves at a non-uniform speed in a time slot corresponding to the current frame. In the case that the motion vertex moves at a non-uniform speed (for example, variable speed motion or self-rotation motion) in the sequence of three-dimensional meshes, although the vertex position of the motion vertex in the sequence of three-dimensional meshes can change, causing the three-dimensional mesh to deform, the actual motion trend of the motion vertex in the time slot corresponding to each frame in the sequence of three-dimensional meshes is relatively small. By decoding the current frame in the sequence of three-dimensional meshes based on the actual motion trend information of the current frame, a better encoding effect can be obtained.

[0053] In a possible implementation, the actual motion trend information of the first three-dimensional mesh is obtained based on the encoding data, comprising: obtaining predicted motion trend information of the first three-dimensional mesh, the predicted motion trend information indicating a predicted motion trend of the motion vertex in the time slot corresponding to the first three-dimensional mesh; and obtaining the actual motion trend information based on the encoding data and the predicted motion trend information.

[0054] In the embodiments of the present application, when decoding the current frame, the decoding end can obtain the prediction motion trend information of the current frame, which is the same as the prediction motion trend information of the current frame obtained by the encoding end when encoding the current frame. Thus, the decoding end can use the prediction motion information to decode the encoding data of the current frame, to ensure the accuracy of the reconstructed data of the decoded current frame, for example, the closeness of the reconstructed position of the motion vertex in the current frame to the original position. In addition, the decoding end can also obtain the encoding data of the current frame from the code stream, which is obtained based on the actual motion information and the prediction motion information. When the motion vertex in the three-dimensional mesh sequence moves at a non-uniform speed, although the difference between the uniform speed rates of the same motion vertex in the time slots corresponding to different frames is large, the actual motion trend information (for example, the offset acceleration) of the same motion vertex in the time slots corresponding to different frames is close, so that the difference between the prediction motion information (including the predicted motion trend information) predicted by the present application and the actual motion information (including the actual motion trend information) of the motion vertex is extremely small. Thus, the length of the code stream corresponding to the three-dimensional mesh sequence is shortened, the code rate is lower, the transmission bandwidth can be occupied, and the transmission efficiency and storage efficiency of the code stream (the same storage space can store more encoding data of frames) can be improved. In addition, when decoding the current frame in the three-dimensional mesh sequence, the encoding end can use the above prediction motion information to decode the encoding data of the current frame to obtain the reconstructed data of the current frame, which can reduce the decoding amount of the code stream of the three-dimensional mesh sequence and reduce the picture freezing of the three-dimensional mesh sequence.

[0055] In a possible implementation, the encoding data includes a residual error between the actual motion trend information and the prediction motion trend information.

[0056] In the embodiments of the present application, the encoding data includes a residual error between the actual motion information and the prediction motion information of the current frame. Thus, when the motion vertex in the three-dimensional mesh sequence moves at a non-uniform speed, although the difference between the uniform speed rates of the same motion vertex in the time slots corresponding to different frames is large, the actual motion trend information (for example, the offset acceleration) of the same motion vertex in the time slots corresponding to different frames is close, so that the residual error of the decoding end for the current frame is extremely small, thereby reducing the decoding data amount, reducing the occupation of the transmission bandwidth, and improving the transmission efficiency.

[0057] In a possible implementation, the encoding data includes the actual motion trend information.

[0058] In the embodiments of the present application, when decoding the current frame in the three-dimensional mesh sequence, the decoding end can directly decode the encoding data of the current frame in the code stream to obtain the actual motion trend information of the current frame. In the case that the motion vertices move at a non-uniform speed in the three-dimensional mesh sequence, there can be a large number of motion vertices in the current frame that have similar actual motion trend (for example, acceleration) information, and the encoding data of the current frame includes the actual motion trend information of the motion vertices, so that the actual motion trend information of a large number of motion vertices in the encoding data is similar or identical. In this way, the number of binary bits occupied by the encoding data is smaller, the length of the code stream is shorter, and the code rate is lower. Then, by decoding the encoding data, the decompression rate can be improved. In addition, compared with the scheme in which the encoding data includes the residual between the actual motion trend information and the predicted motion trend information of the current frame, the embodiments of the present application only need to decode the encoding data of the current frame, without calculating the predicted motion trend information and without decoding the residual, so that the decoding complexity of the present application is lower, thereby shortening the decoding delay.

[0059] In a possible implementation, the encoding data has a first syntax structure, and the first syntax structure includes an identifier of the first three-dimensional mesh.

[0060] In a possible implementation, the first syntax structure further includes information corresponding to the motion vertex, and the information corresponding to the motion vertex includes: a first residual of the motion vertex on a first component; a second residual of the motion vertex on a second component; and a third residual of the motion vertex on a third component.

[0061] In a possible implementation, the first syntax structure further includes a vertex identifier of the motion vertex.

[0062] In a possible implementation, the obtaining of the predicted motion trend information of the first three-dimensional mesh includes: based on the code stream, obtaining N sets of reconstruction data corresponding to N second three-dimensional meshes that have been decoded in the three-dimensional mesh sequence, N being an integer greater than or equal to a set value; and based on the N sets of reconstruction data, determining the predicted motion trend information.

[0063] In the embodiments of the present application, when decoding the current frame in the three-dimensional mesh sequence, the decoding end can directly decode the encoding data of the current frame in the code stream to obtain the actual motion trend information of the current frame. In the case that the motion vertices move at a non-uniform speed in the three-dimensional mesh sequence, there can be a large number of motion vertices in the current frame that have similar actual motion trend (for example, acceleration) information, and the encoding data of the current frame includes the actual motion trend information of the motion vertices, so that the actual motion trend information of a large number of motion vertices in the encoding data is similar or identical. In this way, the number of binary bits occupied by the encoding data is smaller, the length of the code stream is shorter, and the code rate is lower. Then, by decoding the encoding data, the decompression rate can be improved. In addition, compared with the scheme in which the encoding data includes the residual between the actual motion trend information and the predicted motion trend information of the current frame, the embodiments of the present application only need to decode the encoding data of the current frame, without calculating the predicted motion trend information and without decoding the residual, so that the decoding complexity of the present application is lower, thereby shortening the decoding delay.

[0064] In the embodiments of the present application, in order to ensure that the prediction motion trend information calculated by the encoding end and the decoding end for the same frame is the same, the decoding end can obtain the reconstruction data of at least three historical frames in the three-dimensional mesh sequence when obtaining the prediction motion trend information of the current frame, and determine the prediction motion trend information of the current frame based on the reconstruction data (rather than the original data of the at least three historical frames), so as to ensure that the prediction motion trend information determined by the decoding end for the current frame is consistent with the prediction motion trend information obtained by the encoding end for the current frame, thereby ensuring accurate decoding of the three-dimensional mesh sequence by the decoding end, ensuring that the vertex position difference between the reconstructed three-dimensional mesh and the original three-dimensional mesh is small, and improving the quality of the reconstructed three-dimensional mesh.

[0065] In a possible implementation, the reconstruction data of the first three-dimensional mesh is obtained based on the actual motion trend information, including: based on the code stream, reconstruction data corresponding to a plurality of third three-dimensional meshes in the three-dimensional mesh sequence that have been decoded is obtained; and based on the actual motion trend information and the reconstruction data of the plurality of third three-dimensional meshes, the reconstruction data of the first three-dimensional mesh is obtained.

[0066] In the embodiments of the present application, the decoding end can obtain the reconstruction data of at least two frames that have been decoded before the current frame based on the code stream, and then obtain the reconstruction data of the current frame by using the actual motion trend information of the current frame and the reconstruction data, thereby ensuring accurate reconstruction of the three-dimensional mesh.

[0067] In a possible implementation, at least one of the first residual, the second residual, and the third residual is a residual of motion acceleration.

[0068] In a possible implementation, the first component is an X component, the second component is a Y component, and the third component is a Z component.

[0069] In a possible implementation, the first component is an angle component, the second component is a height component, and the third component is a radius component.

[0070] In a possible implementation, the first three-dimensional mesh, the N second three-dimensional meshes, and the plurality of third three-dimensional meshes have the same topological structure.

[0071] The effects of the second aspect and possible implementation manners are similar to those of the first aspect and possible implementation manners, and will not be repeated here.

[0072] In a third aspect, the present application provides an encoding device, comprising: an obtaining module configured to obtain actual motion trend information of a first three-dimensional mesh to be encoded in a three-dimensional mesh sequence, the actual motion trend information indicating actual motion trend of a motion vertex in the three-dimensional mesh sequence that moves at a non-uniform speed in a time slot corresponding to the first three-dimensional mesh; an encoding module configured to obtain encoding data of the first three-dimensional mesh based on the actual motion trend information, and encode the encoding data into a bitstream corresponding to the three-dimensional mesh sequence.

[0073] In a fourth aspect, the present application provides a decoding device, comprising: an obtaining module configured to obtain a bitstream corresponding to a three-dimensional mesh sequence; a decoding module configured to obtain encoding data of a first three-dimensional mesh to be decoded in the three-dimensional mesh sequence from the bitstream, and obtain actual motion trend information of the first three-dimensional mesh based on the encoding data, the actual motion trend information indicating actual motion trend of a motion vertex in the three-dimensional mesh sequence that moves at a non-uniform speed in a time slot corresponding to the first three-dimensional mesh; and obtain reconstruction data of the first three-dimensional mesh based on the actual motion trend information.

[0074] In a fifth aspect, the present application provides an encoding device, comprising: one or more processors; and a memory configured to store one or more programs, wherein the one or more programs, when executed by the one or more processors, enable the one or more processors to implement the method according to any one of the first aspect.

[0075] In a sixth aspect, the present application provides a decoding device, comprising: one or more processors; and a memory configured to store one or more programs, wherein the one or more programs, when executed by the one or more processors, enable the one or more processors to implement the method according to any one of the second aspect.

[0076] In a seventh aspect, the present application provides a computer-readable storage medium, which stores program instructions, wherein the program instructions, when executed by a processor, enable the processor to perform the method according to any one of the first to second aspects.

[0077] In an eighth aspect, the present application provides a computer program product, which comprises computer program code, wherein the computer program code, when executed on a processor, enables the processor to perform the method according to any one of the first to second aspects.

[0078] In a ninth aspect, the present application provides a chip, comprising a processor and a memory, wherein the memory is configured to store a computer program, and the processor is configured to invoke and run the computer program stored in the memory to perform the method according to any one of the first to second aspects.

[0079] In a tenth aspect, the present application provides a bitstream, the bitstream corresponding to a sequence of three-dimensional meshes, the bitstream comprising encoded data of a first three-dimensional mesh, the encoded data being obtained based on actual motion trend information of the first three-dimensional mesh, the actual motion trend information indicating actual motion trend of a motion vertex in non-uniform motion in a time slot corresponding to the first three-dimensional mesh in the sequence of three-dimensional meshes.

[0080] In a possible implementation, the encoded data comprises a residual between the actual motion trend information and predicted motion trend information of the motion vertex in the time slot corresponding to the first three-dimensional mesh, the predicted motion trend information indicating predicted motion trend of the motion vertex in the time slot corresponding to the first three-dimensional mesh.

[0081] In a possible implementation, the encoded data comprises the actual motion trend information.

[0082] In a possible implementation, the bitstream comprises a first syntax structure corresponding to the first three-dimensional mesh, the first syntax structure comprising an identifier of the first three-dimensional mesh.

[0083] In a possible implementation, the first syntax structure further comprises information of the motion vertex, the information of the motion vertex comprising: a first residual of the motion vertex in a first component; a second residual of the motion vertex in a second component; a third residual of the motion vertex in a third component.

[0084] In a possible implementation, the information of the motion vertex further comprises: a vertex identifier of the motion vertex.

[0085] In a possible implementation, at least one of the first residual, the second residual and the third residual is a residual of motion acceleration.

[0086] In a possible implementation, the first component is an X component, the second component is a Y component, and the third component is a Z component.

[0087] In a possible implementation, the first component is an angle component, the second component is a height component, and the third component is a radius component.

[0088] In an eleventh aspect, the present application provides a computer readable storage medium, the computer readable storage medium having stored thereon the bitstream according to any one of the above tenth aspect.

[0089] In a twelfth aspect, the present application provides a method for transmitting an encoded bitstream of video data, the method comprising: obtaining the bitstream from a storage medium, the bitstream being the bitstream of the tenth aspect or any possible implementation of the tenth aspect combined with the tenth aspect, and stored in the storage medium; and transmitting the bitstream.

[0090] In a thirteenth aspect, the present application provides a system for transmitting an encoded bitstream of video data, the system comprising: an obtaining unit, configured to obtain the bitstream from a storage medium, the bitstream being the bitstream of the tenth aspect or any possible implementation of the tenth aspect combined with the tenth aspect, and stored in the storage medium; and a transmitting unit, configured to transmit the bitstream.

[0091] In a fourteenth aspect, the present application provides a method for storing an encoded bitstream of video data, the method comprising: receiving the bitstream of the tenth aspect or any possible implementation of the tenth aspect combined with the tenth aspect; and storing the bitstream into a storage medium.

[0092] In a fifteenth aspect, the present application provides a system for storing an encoded bitstream of video data, the system comprising: a receiving unit, configured to receive the bitstream of the tenth aspect or any possible implementation of the tenth aspect combined with the tenth aspect; and a storing unit, configured to store the bitstream. BRIEF DESCRIPTION OF DRAWINGS

[0093] FIG. 1A is an architecture diagram of a coding system according to an embodiment of the present application;

[0094] FIG. 1B is a flowchart of an encoding method according to an embodiment of the present application;

[0095] FIG. 1C is a flowchart of an encoding method according to an embodiment of the present application;

[0096] FIG. 2A is a flowchart of an encoding method according to an embodiment of the present application;

[0097] FIG. 2B is a flowchart of an encoding method according to an embodiment of the present application;

[0098] FIG. 2C is a flowchart of an encoding method according to an embodiment of the present application;

[0099] FIG. 2D is a flowchart of an encoding method according to an embodiment of the present application;

[0100] FIG. 3A is a flowchart of an encoding method according to an embodiment of the present application;

[0101] FIG. 3B is a flowchart of an encoding method according to an embodiment of the present application;

[0102] FIG. 3C is a flowchart of an encoding method according to an embodiment of the present application;

[0103] FIG. 3D is a flowchart of an encoding method according to an embodiment of the present application;

[0104] FIG. 3E is a flowchart of an encoding method according to an embodiment of the present application;

[0105] FIG. 3F is a flowchart of a decoding method according to an embodiment of the present application;

[0106] FIG. 3G is a flowchart of a decoding method according to an embodiment of the present application;

[0107] FIG. 4A is a flowchart of a decoding method according to an embodiment of the present application;

[0108] FIG. 4B is a flowchart of a decoding method according to an embodiment of the present application;

[0109] FIG. 4C is a flowchart of a decoding method according to an embodiment of the present application;

[0110] FIG. 5A is a flowchart of a decoding method according to an embodiment of the present application;

[0111] FIG. 5B is a flowchart of a decoding method according to an embodiment of the present application;

[0112] FIG. 5C is a flowchart of a decoding method according to an embodiment of the present application;

[0113] FIG. 5D is a flowchart of a decoding method according to an embodiment of the present application;

[0114] FIG. 5E is a flowchart of a decoding method according to an embodiment of the present application;

[0115] FIG. 6 is a block diagram of an encoding apparatus according to an embodiment of the present application;

[0116] FIG. 7 is a block diagram of a decoding apparatus according to an embodiment of the present application;

[0117] FIG. 8 is a structural schematic diagram of an apparatus according to an embodiment of the present application. DETAILED DESCRIPTION

[0118] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the embodiments of the present application, any other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0119] The terms "first", "second", etc. in the description embodiments of the present application and claims and drawings are only used for distinguishing purposes of description and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying sequence. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, comprising a series of steps or units. The method, system, product or device is not necessarily limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0120] It should be understood that in the embodiments of the present application, "at least one" refers to one or more, and "multiple" refers to two or more. "And / or" is used to describe the relationship between the associated objects, which means that there can be three relationships, for example, "A and / or B" can represent three cases: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can mean a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0121] The following first introduces the terms related to the embodiments of the present application:

[0122] 1) Three-dimensional mesh (e.g., dynamic mesh) (3-Dimension Mesh, 3D Mesh): also known as three-dimensional model, is a three-dimensional structure composed of a series of vertices and corresponding edges. Illustratively, the three-dimensional structure can include planar patches (also known as polygons, usually quadrilaterals or triangles). Each planar patch can be composed of a set of vertices and corresponding edges formed by connecting the vertices. The three-dimensional mesh can be used to express volumetric video, digital human, computer graphics (CG) content, etc. The data of the three-dimensional mesh usually includes: vertex coordinates, connection relationship, texture coordinates and texture map, etc. The vertex coordinates are used to indicate the position information of the vertex in the three-dimensional (3D) space, and the connection relationship is used to indicate the vertex composition of the patch (e.g., triangular patch) in the three-dimensional mesh.

[0123] 2) Topology of three-dimensional mesh: represents the connection relationship between multiple vertices within all patches in the three-dimensional mesh.

[0124] 3) Current frame: at the encoding end, "current frame" represents the three-dimensional mesh to be encoded; at the decoding end, "current frame" represents the three-dimensional mesh to be decoded.

[0125] 4) History frame: at the encoding end, the "history frame" means the three-dimensional mesh that has been encoded before the current frame (the three-dimensional mesh to be encoded); at the decoding end, the "history frame" means the three-dimensional mesh that has been decoded before the current frame (the three-dimensional mesh to be decoded).

[0126] 5) Frame sequence: also referred to as three-dimensional mesh sequence, is a plurality of three-dimensional meshes with the same topology arranged in sequence, which can be the display sequence or the encoding sequence of the three-dimensional meshes, which is not limited herein. Among them, the plurality of three-dimensional meshes in the three-dimensional mesh sequence have the same vertices, and the "same vertices" indicates the vertices with the same vertex identification or vertex sequence number.

[0127] For example, the three-dimensional mesh sequence is composed of three-dimensional mesh 1 to three-dimensional mesh n, where n is an integer greater than 1, and each of the three-dimensional mesh 1 to three-dimensional mesh n has vertices 1 to vertices m, m is an integer greater than or equal to 3. And the vertex connection relationship of vertices 1 to vertices m in three-dimensional mesh 1 is the same as the vertex connection relationship of vertices 1 to vertices m in each of three-dimensional mesh 2 to three-dimensional mesh n. However, the vertex positions of the same vertices (such as vertex 1) in different three-dimensional meshes in the three-dimensional mesh sequence can be different, and the present application does not limit whether the vertex positions of the "same vertices" (also referred to as "the same vertex") in different three-dimensional meshes in the three-dimensional mesh sequence are the same.

[0128] 6) Motion vertex: due to the same vertices in the three-dimensional mesh sequence, and the vertex position of the same vertex in the three-dimensional mesh sequence can change, the motion information of any vertex in the three-dimensional mesh sequence can be tracked. Among them, the motion vertex means the vertex with motion information (or the vertex in motion) in the three-dimensional mesh sequence. The motion vertex can move at a constant speed or non-constant speed in the three-dimensional mesh sequence, which is not limited herein.

[0129] 7) Time slot corresponding to the three-dimensional mesh: can be the starting presentation time corresponding to the three-dimensional mesh, can be the end presentation time corresponding to the three-dimensional mesh, can be the middle time of the presentation period corresponding to the three-dimensional mesh, and can be the entire presentation period (time period) corresponding to the three-dimensional mesh, which is not limited by the embodiments of the present application; for example, at a frame rate of 25 frames per second, the presentation time of each three-dimensional mesh is 40 milliseconds, then the time slot corresponding to the first three-dimensional mesh can be 0 milliseconds, can be 39 milliseconds, can be 19 milliseconds (middle time), and can be the first 40 milliseconds time period, the time slot corresponding to the second three-dimensional mesh can be 40 milliseconds, can be 79 milliseconds, can be 59 milliseconds, and can be the second 40 milliseconds time period, and so on.

[0130] 8) motion tendency: the motion tendency of a motion vertex, refers to the tendency of the motion vertex in motion, including the acceleration of the motion vertex, optionally, also including the speed of the motion vertex; the motion tendency can be used to calculate the position or coordinates of the motion vertex; in a given coordinate system, the motion tendency includes all projection components of the acceleration / speed of the motion vertex in the coordinate system; taking the Cartesian coordinate system as an example, the motion tendency includes the projection component of the acceleration / speed of the motion vertex in the x-axis, the projection component in the y-axis and the projection component in the z-axis; the motion tendency of the motion vertex in the time slot corresponding to the three-dimensional grid can be the average acceleration / average speed of the motion vertex in the time slot corresponding to the three-dimensional grid, or the instantaneous acceleration / instantaneous speed of the motion vertex in the time slot corresponding to the three-dimensional grid, which is not limited by the embodiments of the application.

[0131] 9) motion tendency information: information used to describe or indicate the motion tendency.

[0132] 10) actual motion tendency: the actual motion tendency of a motion vertex, refers to the motion tendency calculated according to the position information or coordinate information of the vertex in the current frame and several frames before the current frame.

[0133] 11) actual motion tendency information: information used to describe or indicate the actual motion tendency;

[0134] 12) predicted motion tendency: the predicted motion tendency of a motion vertex, refers to the motion tendency calculated according to the position information or coordinate information of the vertex in several frames (not including the current frame) before the current frame.

[0135] 13) predicted motion tendency information: information used to describe or indicate the predicted motion tendency.

[0136] In each of the following embodiments, the current frame of the encoding end corresponds to the first three-dimensional grid of the encoding side (including the encoding method and the encoding device) in the summary, the N historical frames of the encoding end correspond to the N second three-dimensional grids of the encoding side (including the encoding method and the encoding device) in the summary, the M historical frames of the encoding end correspond to the plurality of third three-dimensional grids of the encoding side (including the encoding method and the encoding device) in the summary. The S historical frames of the encoding end correspond to the plurality of fourth three-dimensional grids of the encoding side (including the encoding method and the encoding device) in the summary.

[0137] In each of the following embodiments, the current frame of the decoding side corresponds to the first three-dimensional mesh of the decoding side (including the decoding method and the decoding device) in the summary of the invention, the N historical frames of the decoding side correspond to the N second three-dimensional meshes of the decoding side (including the decoding method and the decoding device) in the summary of the invention, and the M historical frames of the decoding side correspond to the plurality of third three-dimensional meshes of the decoding side (including the decoding method and the decoding device) in the summary of the invention.

[0138] FIG. 1A is a schematic diagram of a coding system architecture of a three-dimensional mesh according to an embodiment of the present application. As shown in FIG. 1A, the system can include an encoding end 100 and a decoding end 200.

[0139] As shown by the dashed arrows in FIG. 1A, the encoding end 100 and the decoding end 200 can be communicatively connected (for example, directly connected through a network or indirectly connected through a server, a communication device, etc., which is not limited herein), or the encoding end 100 and the decoding end 200 can not be connected.

[0140] The encoding end 100 can encode the obtained three-dimensional mesh data and output a bitstream.

[0141] For example, the encoding end 100 can output the bitstream to a local storage device, a storage medium, or to a remote device or a server (such as a video website), or output the bitstream to the decoding end 200 through the above-mentioned communication connection, which is not limited herein.

[0142] The decoding end 200 can decode the input bitstream to obtain reconstructed three-dimensional mesh data. For example, the decoding end can read the bitstream from a local storage device or a storage medium, or receive the bitstream from a remote device or a server (such as a video website), or receive the input bitstream through the above-mentioned communication connection, and then decode the bitstream.

[0143] Although FIG. 1A shows the encoding end 100 and the decoding end 200 as independent devices, device embodiments can also include both the encoding end 100 and the decoding end 200, or both the corresponding encoding function of the encoding end 100 and the corresponding decoding function of the decoding end 200. In these embodiments, the encoding end 100 or the corresponding encoding function of the encoding end 100, the decoding end 200 or the corresponding function of the decoding end 200, can be implemented by hardware, software, or a combination of hardware and software, including one or more signal processing and / or application specific integrated circuits.

[0144] The encoding end 100 and the decoding end 200 can include any of various devices, including any type of handheld or fixed device or wearable device, such as a notebook computer or laptop computer, a smartphone, a tablet or tablet computer, a desktop computer, a smart television, a vehicle display screen, a smart bracelet, a smart watch, an augmented reality (AR) device, a virtual reality (VR) device, and the like, and can not use or use any type of operating system. In some cases, the encoding end 100 and the decoding end 200 can be wireless communication devices equipped with wireless communication components.

[0145] The application scenario of the system shown in FIG. 1A can be various business scenarios involving the encoding and decoding of three-dimensional meshes. For example, a video conference scenario, a video call scenario, an online education scenario, a remote tutoring scenario, a live streaming scenario, a cloud gaming scenario, and the like, and the present application does not limit this.

[0146] For example, the system is applied to a live streaming scenario, the encoding end 100 can be a server of a live streaming application (App), and the decoding end 200 can be a terminal device such as a mobile phone running the live streaming application. The three-dimensional mesh can be a three-dimensional mesh of clothes of a digital person in live streaming, the encoding end 100 encodes the three-dimensional mesh of the clothes of the digital person, the decoding end 200 decodes the three-dimensional mesh, and displays an image of the clothes of the digital person rendered by the decoded three-dimensional mesh.

[0147] For another example, the system can be applied to a cloud gaming scenario, the encoding end 100 can be a cloud end providing a game application, and the decoding end 200 can be a terminal device such as a mobile phone running the game application. The three-dimensional mesh can be a three-dimensional mesh of virtual clothes of a virtual character in the game, the encoding end 100 encodes the three-dimensional mesh of the virtual clothes of the virtual character, the decoding end 200 decodes the three-dimensional mesh, and displays an image of the virtual clothes of the virtual character rendered by the decoded three-dimensional mesh.

[0148] The encoding and decoding processes of the three-dimensional mesh by the encoding end 100 and the decoding end 200 will be described below in conjunction with FIG. 1A.

[0149] Among the three-dimensional mesh sequence to be encoded / decoded, the three-dimensional mesh includes a motion vertex, and the encoding and decoding processes of the present application are described by taking the motion vertex in the three-dimensional mesh sequence as an example of a vertex moving at a non-uniform speed. The encoding and decoding processes of the encoding end 100 and the decoding end 200 of the present application for the motion vertex moving at a uniform speed in the three-dimensional mesh sequence are the same as the encoding and decoding processes for the motion vertex moving at a non-uniform speed.

[0150] The non-uniform motion can be a motion in which at least one of the speed and direction of motion changes. For example, the non-uniform motion can be a rotational motion (either uniform or non-uniform), or it can be a variable-speed motion (e.g., interval running).

[0151] For example, the 3D object corresponding to this 3D mesh is cloth.

[0152] In a cloud gaming scenario, the three-dimensional object can be virtual clothing worn by a virtual character. When the virtual character undergoes a self-rotation (e.g., dancing ballet), it can cause the virtual clothing to rotate as well. This allows the corresponding three-dimensional mesh sequence to include non-uniform vertices undergoing self-rotation, causing the virtual clothing to deform.

[0153] For example, in a cloud gaming scenario, a virtual avatar wearing virtual clothes runs at high speed, causing the virtual clothes to deform. This means that the corresponding 3D mesh sequence of the virtual clothes can include moving vertices undergoing variable speed motion, resulting in deformation of the virtual clothes.

[0154] Of course, the three-dimensional object corresponding to the three-dimensional mesh in this application can also be a flexible object other than cloth, such as clay, ball, jelly, etc. When there are vertices in the three-dimensional mesh of these objects that are undergoing non-uniform motion, the encoding and decoding method and system of this application are also applicable.

[0155] In other embodiments, the moving vertex can also be a vertex in a three-dimensional mesh sequence that moves at a constant speed. In other words, the encoding and decoding method and system of this application can be applied not only to scenarios where vertices in a three-dimensional mesh sequence move at a non-uniform speed, but also to scenarios where vertices in a three-dimensional mesh sequence move at a constant speed.

[0156] In some scenarios, such as when the 3D mesh is a deformable 3D object like a flexible body (or other scenarios, without limitation), the vertices in this 3D mesh do not move at a constant velocity, but rather at a non-constant velocity. Related technologies for encoding 3D meshes assume that all vertices in the mesh move at a constant velocity. Specifically, this technology predicts the predicted positions of vertices in the current frame (i.e., the 3D mesh to be encoded) based on the assumption of constant velocity motion, calculates the residual between the predicted and actual positions, and encodes this residual into the bitstream to encode the current frame. However, because the vertices do not move at a constant velocity, the predicted positions are inaccurate, and the encoding and decoding methods for 3D meshes in these technologies need improvement.

[0157] The following example uses a 3D mesh as a 3D model of the virtual clothing worn by a virtual character in a cloud gaming scene to illustrate the processing procedure of the system in this application.

[0158] As shown in FIG. 1A, the process can include the following steps:

[0159] S102: The encoding end 100 can obtain the encoding data of the current frame in the three-dimensional mesh sequence, and encode the encoding data into the code stream corresponding to the three-dimensional mesh sequence.

[0160] The encoding end 100 can obtain the actual motion trend information of the current frame in the three-dimensional mesh sequence, and based on the actual motion trend information, obtain the encoding data of the current frame, and then encode the encoding data into the code stream corresponding to the three-dimensional mesh sequence.

[0161] For example, the encoding data can include residual or actual motion trend information, etc.

[0162] The encoding data is data before being converted into a bit stream, so the encoding data can also be referred to as to-be-encoded data.

[0163] For example, the encoding end 100 can perform entropy encoding or compression encoding on the encoding data, etc. to convert the encoding data into a bit stream and write it into the code stream.

[0164] The actual motion trend information indicates the actual motion trend of the vertex (also referred to as the motion vertex) in the three-dimensional mesh sequence that moves at a non-uniform speed in the time slot corresponding to the current frame.

[0165] In some embodiments, each frame in the three-dimensional mesh sequence can have original data. The original data of the three-dimensional mesh includes the original position information of each vertex in the three-dimensional mesh, and optionally, the connection relationship between each vertex in the three-dimensional mesh (also referred to as the topology of the three-dimensional mesh).

[0166] In order to encode the original position information of the vertex of the current frame, the actual motion trend information of the current frame can be calculated to obtain the encoding data of the current frame.

[0167] The three-dimensional mesh sequence of the present application can be generated by the encoding end 100, or can be generated by other devices or software connected to the encoding end 100, which is not limited here.

[0168] The three-dimensional mesh sequence (also referred to as the frame sequence) can be generated by a physical simulator or by a physical simulation algorithm. The topology of each frame in the frame sequence is the same, and the vertex identification of each frame is the same. The topology of the three-dimensional mesh can include the connection relationship between each vertex in all vertices in the three-dimensional mesh. Therefore, the vertex position of the same vertex in the frame sequence in different frames can change, but the connection relationship between the vertex and other vertices remains unchanged.

[0169] The topology of different frames in the three-dimensional mesh sequence (also referred to as frame sequence) generated by the physical simulator or the physical simulation algorithm is kept unchanged, but the vertex position of the same vertex in the three-dimensional mesh in different frames can change (for example, non-uniform motion).

[0170] For example, the encoding end 100 can obtain the original data of the current frame and the historical frame from the physical simulator, which can include the original position information of the vertices of the current frame and the historical frame.

[0171] It should be understood that the three-dimensional mesh sequence of the present application is not limited to being generated by a physical simulator or by a physical simulation algorithm, but can also be generated in other ways, as long as the topology of different frames in the three-dimensional mesh sequence is kept unchanged and has the same vertices (referring to the same vertex identification or vertex sequence number).

[0172] For example, the encoding end 100 can collect the original position (for example, three-dimensional coordinates) information of each vertex in the current frame (specifically, the three-dimensional mesh to be encoded) and the original position information of each vertex in the historical frame (specifically, the three-dimensional mesh that has been encoded before the three-dimensional mesh to be encoded).

[0173] Alternatively, when only part of the vertices in the three-dimensional mesh need to be encoded and decoded, the original position information of the vertices of the current frame and the historical frame collected by the encoding end 100 can also be the original position of part of the vertices in the three-dimensional mesh, not the original position information of each vertex.

[0174] For example, in the cloud game scenario, the cloud can generate the three-dimensional mesh of the virtual clothes worn by the virtual character in the game through the physical simulation algorithm, and collect the original position information of each vertex in the three-dimensional mesh of the clothes.

[0175] In the three-dimensional mesh sequence generated by the physical simulation algorithm or by the physical simulator, the topologies of different frames are the same, but the original position information of the vertices under the topology changes between different frames, therefore, the method of the present application can calculate the actual motion trend information of the vertices in the three-dimensional mesh based on the changed original position of the vertices while keeping the topology of the three-dimensional mesh unchanged, so as to realize the encoding and decoding of the three-dimensional mesh.

[0176] S103: The encoding end 100 outputs the code stream.

[0177] For example, the encoding end 100 can output the code stream to a local storage device or a storage medium, or to a server, etc., which is not limited here.

[0178] S201: The decoding end 200 inputs the code stream.

[0179] For example, the decoding end 200 can read the code stream from a local storage device or storage medium, or receive the code stream from a remote device or server (such as a video website) and the like to obtain the input code stream.

[0180] S202: The decoding end 200 decodes the input code stream to obtain reconstructed three-dimensional mesh data.

[0181] For example, the decoding end 200 can decode the code stream to obtain the encoding data of the current frame to be decoded in the three-dimensional mesh sequence from the code stream, and obtain the actual motion trend information of the current frame based on the encoding data, so as to reconstruct the vertex position of the current frame by using the actual motion trend information, thereby obtaining the reconstructed data (such as the reconstructed position information of the vertex and the above-mentioned topological structure) of the current frame to reconstruct the current frame.

[0182] The current frame is a three-dimensional mesh to be decoded.

[0183] Optionally, the system can further include a rendering end (not shown) which can be in communication connection with the decoding end 200. The rendering end can be software, hardware, chips and the like having a function of rendering the three-dimensional mesh into an image, which is not limited here.

[0184] The rendering end can render the three-dimensional mesh data obtained by the decoding end 200 through S202 to obtain an image or an image sequence (such as a video) for on-screen display.

[0185] For example, in a cloud game scenario, the rendering end can render the three-dimensional mesh of the virtual clothes reconstructed by the decoding end 200 to obtain a rendering image of the virtual clothes for on-screen display.

[0186] The encoding process of the three-dimensional mesh (such as the process performed by the encoding end 100 in FIG. 1A) and the decoding process of the three-dimensional mesh (such as the process performed by the decoding end 200 in FIG. 1A) of the present application will be described below in combination with different embodiments.

[0187] FIG. 1B is a method flowchart for encoding a current frame (a three-dimensional mesh to be encoded) according to an embodiment of the present application. The method flowchart can be implemented based on the architecture shown in FIG. 1A, but is not limited to the architecture shown in FIG. 1A. The method flowchart can include the following steps:

[0188] S330: The encoding end obtains the actual motion trend information of the current frame to be encoded in the three-dimensional mesh sequence.

[0189] The actual motion trend information indicates the actual motion trend of the motion vertex in the non-uniform motion in the current frame corresponding time slot in the three-dimensional mesh sequence.

[0190] The actual motion trend can be a motion acceleration. Since the current frame is a three-dimensional structure, the actual motion trend can be a motion acceleration in at least one of the three components in a three-dimensional coordinate system.

[0191] In some embodiments, when the actual motion trend is a motion acceleration in one component (e.g., the X component), the encoding end can further obtain actual motion rate information of the current frame, the actual motion rate information being actual motion rates of the motion apex in the other two components (e.g., the Y component and the X component) in the corresponding time slot of the current frame.

[0192] In some embodiments, when the actual motion trend is a motion acceleration in two components (e.g., the X component and the Y component), the encoding end can further obtain actual motion rate information of the current frame, the actual motion rate information being an actual motion rate of the motion apex in the other component (e.g., the Z component) in the corresponding time slot of the current frame.

[0193] The data of the sequence of three-dimensional meshes can include original data of each three-dimensional mesh to be encoded, the original data can include a topology of the three-dimensional mesh and original vertex positions of all or part of the vertices in the three-dimensional mesh, and optionally further include a frame rate of the sequence of three-dimensional meshes, and optionally further include an instantaneous frame rate of each three-dimensional mesh in the sequence of three-dimensional meshes.

[0194] The sequence of three-dimensional meshes has a frame rate, and an inverse of the frame rate is an overall motion time of the motion apex in the sequence of three-dimensional meshes.

[0195] Any frame in the sequence of three-dimensional meshes can have its own instantaneous frame rate, for example, an average of the instantaneous frame rates of all frames in the sequence of three-dimensional meshes is the frame rate of the sequence of three-dimensional meshes. The instantaneous frame rates of different frames in the sequence of three-dimensional meshes can be the same or different, which is not limited here. Based on this, the time slot corresponding to a frame in the sequence of three-dimensional meshes can be an inverse of the instantaneous frame rate of the frame, and the sum of the time slots corresponding to each frame in the sequence of three-dimensional meshes can be the overall motion time.

[0196] In the method flow shown in FIG. 2B below, the data of the sequence of three-dimensional meshes can be divided into original data of the current frame and original data of the historical frame.

[0197] In order to encode the current frame (e.g., the original vertex position of the motion apex), the embodiments of the present application can realize encoding of the current frame based on the actual motion trend information of the motion apex of the current frame to obtain the encoding data of the current frame.

[0198] S331: The encoding end obtains the encoding data of the current frame based on the actual motion trend information of the current frame.

[0199] The actual motion trend information can be directly encoded by the encoding end, or the current frame can be encoded based on the actual motion trend information and other information of the current frame to obtain the encoding data of the vertex position of each motion vertex of the current frame.

[0200] The encoding data is data obtained by encoding the vertex position of the motion vertex of the current frame, but the encoding data is not a 0, 1 bit sequence.

[0201] The encoding end encodes the encoding data into a code stream corresponding to the three-dimensional mesh sequence.

[0202] The encoding end can perform entropy encoding or compression encoding on the encoding data to convert the encoding data into a 0, 1 bit sequence, and write the 0, 1 bit sequence into the code stream.

[0203] In the embodiments of the present application, when encoding the current frame in the three-dimensional mesh sequence, the encoding end can obtain the encoding data of the current frame based on the actual motion trend information of the current frame, wherein the actual motion trend information indicates the actual motion trend of a motion vertex in the three-dimensional mesh sequence that moves at a non-uniform speed in a time slot corresponding to the current frame. In the case where the motion vertex in the three-dimensional mesh sequence moves at a non-uniform speed (for example, variable speed motion or self-rotation motion), although the vertex position of the motion vertex in the three-dimensional mesh sequence can change, causing the three-dimensional mesh to deform, the actual motion trend of the motion vertex in the time slot corresponding to each frame in the three-dimensional mesh sequence is relatively small. By encoding the current frame in the three-dimensional mesh sequence based on the actual motion trend information of the current frame, better encoding effect can be obtained.

[0204] FIG. 1C is a flowchart of a method for encoding a current frame (a three-dimensional mesh to be encoded) according to an embodiment of the present application. The method can be implemented based on the architecture shown in FIG. 1A, but is not limited to the architecture shown in FIG. 1A. The method can be combined with the flow shown in FIG. 1B, but is not limited to the flow shown in FIG. 1B. As shown in FIG. 1C, the method can include the following steps:

[0205] S3301: The encoding end obtains actual motion trend information of a current frame to be encoded in a three-dimensional mesh sequence.

[0206] The execution principle of S3301 is the same as that of S330 shown in FIG. 1B, which will not be described here.

[0207] S3312: The encoding end encodes the actual motion trend information to obtain encoding data of the current frame.

[0208] For example, the encoding data includes the actual motion trend information.

[0209] For example, the encoding end can quantize, compress, or at least one encoding method to the actual motion trend information to obtain the encoding data of the current frame. The application does not limit the encoding method of the actual motion trend information.

[0210] S332: The encoding end encodes the encoding data into the code stream corresponding to the three-dimensional mesh sequence.

[0211] In the related art, when encoding the three-dimensional mesh, it is assumed that the three-dimensional mesh moves at a constant speed, and the motion speed of the vertex of the three-dimensional mesh between adjacent frames is constant. Based on this assumption, the vertex position of the current frame is predicted using the vertex position of the historical frame, and then the residual (also referred to as position residual) between the predicted position and the actual position of the vertex of the current frame is encoded to realize the encoding of the current frame. However, when the three-dimensional mesh is a model of a flexible or elastic three-dimensional object that is easily deformed, the vertex in the three-dimensional mesh moves at a non-constant speed (for example, the vertex rotates itself or the vertex moves at a variable speed, or the three-dimensional mesh deforms itself in the three-dimensional mesh sequence), and the predicted vertex position of the current frame based on the assumption that the vertex moves at a constant speed is not accurate, resulting in a large deviation between the predicted position of the vertex and the actual position of the vertex. In addition, in the scenario where the vertex moves at a non-constant speed in the three-dimensional mesh sequence, the predicted positions of a large number of vertices are all based on the assumption that the vertex moves at a constant speed, resulting in a large difference in the predicted residual of a large number of vertices in the three-dimensional mesh, and a low compression rate of the three-dimensional mesh.

[0212] When encoding the data, the data can be quantized, and then the quantized value is converted into a binary sequence for encoding. The higher the frequency of the quantized value, the smaller the interval formed by the quantized value (for example, [-1, 1]). In addition, the higher the frequency of the quantized value, the fewer the number of bits of the binary converted from the quantized value. Therefore, the encoding scheme of the three-dimensional mesh in the related art will cause the interval formed by the quantized value of the residual of a large number of vertices in the three-dimensional mesh to be larger (for example, [-10, 10]), resulting in a larger number of binary bits occupied by the encoding data of the current frame, and further causing the length of the code stream of the three-dimensional mesh sequence to be longer and the code rate to be higher.

[0213] In the embodiments of the present application, when encoding the current frame in the sequence of three-dimensional meshes, the encoding end does not need to calculate the prediction residual of the vertex (for example, the residual between the predicted position and the actual position of the vertex) and encode the residual, but directly encodes the actual motion trend information of the current frame to obtain the encoding data of the current frame, and encodes the encoding data into the code stream corresponding to the sequence of three-dimensional meshes. Then in the case where the motion vertex moves at a non-uniform speed in the sequence of three-dimensional meshes, there can be a large number of motion vertices in the current frame that have similar actual motion trend (for example, acceleration) information. Then when encoding the actual motion trend information of a large number of motion vertices in the current frame, the interval formed by the quantized values of the actual motion trend information can be smaller, and the number of binary bits occupied by the encoding data of the current frame in the code stream can be smaller, thereby improving the compression rate and shortening the length of the code stream corresponding to the sequence of three-dimensional meshes and reducing the code rate. In addition, compared with the scheme of encoding the residual between the actual motion trend information and the predicted motion trend information of the current frame, the encoding complexity of the embodiments of the present application is lower because only the actual motion trend information of the current frame needs to be calculated and encoded, thereby shortening the encoding delay.

[0214] For example, in a specific application scenario, a stationary curtain is pushed open by a person, and then the three-dimensional mesh of the curtain can have a large number of groups of vertices, and the actual motion trend (for example, motion acceleration) of each group of vertices is basically the same. For the reader's understanding, the reader can compare the three-dimensional mesh to an image to be encoded, and compare the actual motion trend information of the vertex in the three-dimensional mesh to the pixel value of the image. If a large block of pixels in the image has similar pixel values, the compression rate of the image will be higher. Similarly, if the actual motion trend information (for example, acceleration) of a large number of vertices in the three-dimensional mesh to be encoded is similar or the same, the compression rate of the three-dimensional mesh can be improved by directly encoding the actual motion trend information of the motion vertex of the three-dimensional mesh to realize the encoding of the three-dimensional mesh.

[0215] FIG. 2A is a flowchart of a method for encoding a current frame (a three-dimensional mesh to be encoded) provided by the embodiments of the present application. The method flowchart can be implemented based on the architecture shown in FIG. 1A, but is not limited to the architecture shown in FIG. 1A. The method can be combined with FIG. 1B or FIG. 1C (but is not limited to being combined with FIG. 1B or FIG. 1C). The method flowchart can include the following steps:

[0216] S3100a: The encoding end obtains the predicted motion trend information of the current frame to be encoded in the sequence of three-dimensional meshes.

[0217] The predicted motion trend information indicates the predicted motion trend (for example, predicted motion acceleration) of the motion vertex in the time slot corresponding to the current frame;

[0218] S3100b: The encoding end obtains the actual motion trend information of the current frame.

[0219] The implementation principle and extended embodiments can refer to S330 shown in FIG. 1B, which will not be repeated here.

[0220] Here, the present application does not limit the execution sequence between S3100a and S3100b, and the two can be executed in series or in parallel.

[0221] S3102: The encoding end obtains the encoding data of the current frame based on the predicted motion trend information of the current frame and the actual motion trend information of the current frame.

[0222] Here, the vertex position of the motion vertex of the current frame can be encoded based on the predicted motion trend information of the current frame and the actual motion trend information of the current frame in any manner to obtain the encoding data of the current frame, which is not limited here.

[0223] S332: The encoding end encodes the encoding data into the code stream corresponding to the three-dimensional mesh sequence.

[0224] In a possible implementation, the encoding data includes a residual error between the actual motion trend information and the predicted motion trend information. The encoding end can determine the residual error between the actual motion trend information and the predicted motion trend information, and then encode the residual error into the code stream corresponding to the three-dimensional mesh sequence.

[0225] For example, the encoding end can quantize and entropy encode the residual error to convert the residual error into a bit sequence, and write the bit sequence into the code stream.

[0226] The encoding manner of encoding the residual error is not limited here.

[0227] In the related art, when encoding the three-dimensional mesh, it is assumed that the three-dimensional mesh moves at a constant speed to predict the vertex position of the current frame, so that the predicted position of the vertex deviates greatly from the actual position of the vertex, and further causes the residual error of the vertex encoding of the three-dimensional mesh to be too large, the data amount after compression is large, and the compression rate is low.

[0228] In encoding the prediction residual of each vertex of the three-dimensional mesh, the prediction residual of each vertex needs to be quantized, and the quantized value is converted into a binary sequence for encoding. The higher the frequency of the quantized value, the smaller the interval formed by the quantized value of the prediction residual of each vertex of the three-dimensional mesh (for example, [-1, 1]). In addition, the higher the frequency of the quantized value, the fewer the number of bits of the binary converted from the quantized value. Therefore, when the residual of the vertex of the three-dimensional mesh is larger, the interval formed by the quantized value of the residual of each vertex is larger (for example, [-10, 10]), which leads to a larger number of binary bits occupied by the encoding result (bit sequence) of the residual, and further leads to a longer length of the code stream of the three-dimensional mesh sequence and a higher code rate.

[0229] However, in the embodiment of the present application, in the case that the current frame includes a motion vertex performing non-uniform motion, the encoding end of the present application encodes the residual between the actual motion trend information and the predicted motion trend information of the current frame to realize the encoding of the current frame. In the case that the three-dimensional mesh is self-deformed or deformed by non-uniform motion (for example, cloth deformation), the vertex position predicted by the related art based on the historical frame is not accurate. However, the actual motion trend information of the same motion vertex in different frames in the three-dimensional mesh sequence is almost unchanged. Therefore, the predicted motion trend information of the current frame obtained by the present application is closer to the actual motion trend information of the current frame, thereby reducing the residual to be encoded, and reducing the amount of compressed data and improving the data compression rate. Since the prediction residual of each motion vertex of the three-dimensional mesh is small, the interval formed by the quantized result of the residual between the predicted motion trend information and the actual motion trend information of each motion vertex of the three-dimensional mesh is small, which leads to a smaller number of binary bits occupied by the encoding result (bit sequence) of the residual, and further shortens the length of the code stream corresponding to the three-dimensional mesh sequence and reduces the code rate.

[0230] In other embodiments, the encoding end can determine the difference between the predicted motion trend information and the actual motion trend information of the current frame in a manner other than residual, and encode the current frame into the three-dimensional mesh sequence based on the difference, which is not limited here.

[0231] When the motion vertex in the three-dimensional mesh to be encoded performs non-uniform motion, the difference between the uniform rates of the same motion vertex in different frames is large. In the related art, when the three-dimensional mesh is encoded, whether the vertex performs uniform motion or variable motion, the predicted position of the vertex in the current frame is predicted based on the actual position of the vertex in the previous frame under the assumption that the vertex performs uniform motion, which leads to an inaccurate predicted position of the vertex in the current frame.

[0232] In the method flow corresponding to FIG. 2A, when the motion vertex in the three-dimensional mesh to be encoded moves at a non-uniform speed, although the difference between the uniform speeds of the same motion vertex in the time slots corresponding to different frames is large, the actual motion trend information (for example, the offset acceleration) of the same motion vertex in the time slots corresponding to different frames is close, so that the predicted motion information (including the predicted motion trend information) predicted by the present application for the motion vertex is closer to the actual motion information (including the actual motion trend information) of the motion vertex, thereby improving the accuracy of the predicted information (here, the predicted motion trend) of the motion vertex of the current frame, and the actual motion trend information of the current frame is also accurate. Then, when encoding the current frame in the sequence of three-dimensional meshes, the encoding terminal obtains the encoding data of the current frame based on the accurate actual motion trend information and the accurate predicted motion trend information of the current frame, thereby facilitating the encoding and decoding of the three-dimensional mesh.

[0233] FIG. 2B is a method flow diagram for encoding a current frame (a three-dimensional mesh to be encoded) according to an embodiment of the present application. The method flow can be implemented based on the architecture shown in FIG. 1A, but is not limited to the architecture shown in FIG. 1A. In addition, the method flow shown in FIG. 2B can be combined with FIG. 2A, FIG. 1B, FIG. 1C, and any possible implementation thereof. As shown in FIG. 2B, the method flow can include the following steps:

[0234] S300a: The encoding terminal obtains N sets of reconstruction data of N historical frames in the sequence of three-dimensional meshes.

[0235] The N historical frames and the N sets of reconstruction data correspond to each other in a one-to-one manner, and one historical frame corresponds to one set of reconstruction data.

[0236] The historical frame is a three-dimensional mesh that has been encoded in the sequence of three-dimensional meshes.

[0237] For example, the historical frame is a three-dimensional mesh that has been encoded before the current frame.

[0238] In some embodiments, the number of historical frames in S300a can be N, and N is an integer greater than or equal to a preset threshold (for example, 3).

[0239] Since the historical frame has been encoded, the historical frame can be encoded by the encoding method of the present application, or by other known or future encoding methods, which are not limited here.

[0240] The reconstruction data of one historical frame can be referred to as one set of reconstruction data, and here N sets of reconstruction data corresponding to N historical frames can be obtained.

[0241] The encoding end of the present application can obtain reconstructed data of the historical frames obtained by decoding the encoded data of the historical frames. The decoding process can be performed by the encoding end of the present application or by another decoding device, which is not limited herein.

[0242] The reconstructed data of each of the N historical frames can include reconstructed positions of the motion vertices in the historical frame (i.e., reconstructed data of the vertex positions), and optionally, a topology corresponding to the historical frame. Since the topology of each frame in the three-dimensional mesh sequence is the same, the topology can not need to be reconstructed. Optionally, the reconstructed data of the historical frame can further include a frame rate of the three-dimensional mesh sequence, and optionally, a temporal frame rate of each three-dimensional mesh in the three-dimensional mesh sequence.

[0243] In some embodiments, the frame rate of the three-dimensional mesh sequence or the temporal frame rate of each three-dimensional mesh in the three-dimensional mesh sequence can not be encoded into the code stream. The frame rate and the temporal frame rate can be agreed upon by the encoding end and the decoding end in advance.

[0244] S301a: The encoding end determines the predicted motion information of the current frame based on the reconstructed data of the N historical frames.

[0245] For example, the encoding end can determine the predicted motion information of the motion vertices in the current frame based on the reconstructed positions of the motion vertices in each of the N historical frames.

[0246] The predicted motion information is predicted motion information of the motion vertices in the three-dimensional mesh sequence that move at a non-uniform speed in the time slot corresponding to the current frame. The predicted motion information can include the predicted motion trend information and optionally, predicted motion rate information.

[0247] The predicted motion rate information is a predicted motion rate of the motion vertices in the time slot corresponding to the current frame.

[0248] In this step, the encoding end can use the historical frames that have been encoded before the current frame to predict the motion information (represented by the predicted motion information) of the current frame.

[0249] S301b: The encoding end determines the actual motion information of the current frame based on the original data of the M historical frames and the original data of the current frame.

[0250] The actual motion information is actual motion information of the motion vertices in the three-dimensional mesh sequence that move at a non-uniform speed in the time slot corresponding to the current frame. The actual motion information can include the actual motion trend information and optionally, actual motion rate information.

[0251] The actual motion rate information is an actual motion rate of the motion vertices in the time slot corresponding to the current frame.

[0252] M is an integer greater than or equal to 2.

[0253] The data of the sequence of three-dimensional meshes in the method flow shown in FIG. 2A can include original data of each three-dimensional mesh to be encoded, which can include the topology of the three-dimensional mesh and the original positions of all or part of the vertices in the three-dimensional mesh.

[0254] In the method flow, the data of the sequence of three-dimensional meshes can include original data of the current frame and original data of the historical frames.

[0255] For example, in this embodiment, the original data of the M historical frames and the original data of the current frame are the data obtained through S101 in FIG. 1A. The original data of the historical frames can include the original positions of the motion vertices in the historical frames. The original data of the current frame can include the original positions of the motion vertices in the current frame.

[0256] For example, the original data in S301b can be obtained from a physical simulator.

[0257] S301a and S301b can be executed in series or in parallel, and the present application does not limit this.

[0258] The historical frames in S301b can be the same as the historical frames in S301a, so that M=N, and both M and N are greater than or equal to 3. Alternatively, the historical frames in S301a and S301b can also not be completely the same (i.e., there are some historical frames that are the same between the two steps), or the historical frames in S301a and S301b can be completely different.

[0259] S302: The encoding end encodes the current frame into a code stream corresponding to the sequence of three-dimensional meshes based on the predicted motion information of the current frame and the actual motion information of the current frame.

[0260] The implementation principle of S302 is the same as that of S3102 in the method flow shown in FIG. 2A, and will not be described here.

[0261] In the corresponding method flow of FIG. 2B, because the decoding end side cannot obtain the original data of each frame in the three-dimensional mesh sequence, only the reconstructed data of each frame can be obtained through decoding. In order to ensure that the prediction motion information calculated by the encoding end and the decoding end side for the same frame is the same, when obtaining the prediction motion information of the current frame, the encoding end can obtain the reconstructed data of at least three encoded historical frames of the three-dimensional mesh sequence, and determine the prediction motion information of the current frame based on the reconstructed data (rather than the original data of the at least three historical frames), so as to ensure that the prediction motion information of the current frame determined by the encoding end is consistent with the prediction motion information obtained by the decoding end side, thereby ensuring the accuracy of the reconstructed data of the three-dimensional mesh sequence. When obtaining the actual motion information of the current frame, the encoding end can determine the actual motion information of the current frame based on the original data of the current frame and the original data of at least two historical frames, so as to ensure the accuracy of the actual motion information. Finally, the current frame is encoded into a code stream based on the actual motion information and the prediction motion information of the current frame. In this way, on the one hand, the residual of the encoding is reduced, and on the other hand, the accuracy of the actual motion information calculated by the encoding end is ensured, and the prediction motion information calculated by the encoding end and the decoding end for the current frame remains consistent, so as to facilitate the accurate decoding of the current frame by the decoding end, improve the decoding accuracy, ensure that the vertex position difference between the reconstructed three-dimensional mesh and the original three-dimensional mesh is small, and improve the quality of the reconstructed three-dimensional mesh.

[0262] FIGS. 3A-3E are various method flow diagrams for encoding a three-dimensional mesh according to embodiments of the present application; these method flows can be implemented based on the architecture shown in FIG. 1A, the flows shown in FIGS. 1B, 1C, 2A-2D, but are not limited to the above-mentioned architecture and flows.

[0263] The method flow shown in FIG. 3A can be a specific example of the method flow shown in FIG. 2B to implement the encoding of the current frame, which can mainly include the following steps:

[0264] S3001a: The encoding end determines the actual motion information of the motion vertex of the i-1th frame based on the reconstructed position of the motion vertex of each of the i-3th, i-2th, and i-1th frames and the instantaneous frame rate of each of the i-2th and i-1th frames, and uses the actual motion information of the motion vertex of the i-1th frame as the prediction motion information of the motion vertex of the i-th frame.

[0265] Optionally, the instantaneous frame rate of the i-3th frame can also be combined.

[0266] The i-3th, i-2th, and i-1th frames are an example of the N historical frames in the method flow shown in FIG. 2B.

[0267] The reconstructed data of the N historical frames obtained in S300a in the method flow shown in FIG. 2B can include the vertex position of the motion vertex in each of the i-3th frame, the i-2th frame and the i-1th frame, and optionally, can further include the instantaneous frame rate of each of the i-2th frame and the i-1th frame. The explanation and description of the instantaneous frame rate can refer to the related explanation and description of the embodiment of FIG. 2A, which will not be repeated here. The instantaneous frame rate in S3001a can be the instantaneous frame rate of each frame in the three-dimensional mesh sequence agreed by the encoding end. For example, the three-dimensional mesh data obtained by S101 at the encoding end can include the instantaneous frame rate of each frame.

[0268] The frame of the present application is a three-dimensional mesh, which can include position information of each vertex. The position information is a three-dimensional coordinate. For example, the coordinate of a vertex c can be a three-dimensional Cartesian coordinate c(x, y, z), and for another example, the coordinate of the vertex c can be a cylindrical coordinate c(h, r, θ), where h represents height, r represents radius, and θ represents angle (for example, deflection angle).

[0269] The different frames in the frame sequence of the present application have the same topology of the three-dimensional mesh (for example, three-dimensional model), for example, can share the same three-dimensional coordinate system, but due to the change of the vertex position of the three-dimensional mesh between different frames, the vertex position of the same vertex in different frames can be different, in other words, in the case of deformation of the three-dimensional mesh, the position of the vertex in the three-dimensional mesh in different frames can change.

[0270] The current frame is the i-th frame, i is an integer greater than or equal to 4, and the three frames (specifically, the i-3th frame, the i-2th frame and the i-1th frame) before the i-th frame that have been completed encoding can be used as the N historical frames. Since in the scene where the motion vertex moves at a non-uniform speed, the motion information (for example, motion acceleration) of the motion vertex between adjacent frames is relatively close, in order to improve the accuracy of the predicted motion information of the motion vertex of the current frame, the actual motion information of the motion vertex of the previous frame (i-1th frame) before the current frame (which belongs to the actual motion information reconstructed for the i-1th frame) can be used as the predicted motion information of the motion vertex of the current frame, to improve the difference between the predicted motion information of the current frame and the actual motion information of the current frame, thereby reducing the residual between the actual motion information and the predicted motion information of the current frame.

[0271] In the present embodiment, when determining the actual motion information of the motion vertex of the i-1th frame (the actual motion information reconstructed for the i-1th frame), the actual motion information of the motion vertex of the i-1th frame (the actual motion information reconstructed for the i-1th frame) can be calculated based on the reconstructed positions of the motion vertices of the i-3th, i-2th and i-1th frames, and the actual motion information of the vertex of the i-1th frame is taken as the predicted motion information of the motion vertex in the current frame.

[0272] In other embodiments, when determining the actual motion information of the motion vertex of the i-1th frame (the actual motion information reconstructed for the i-1th frame), the actual motion information of the motion vertex of the i-1th frame (the actual motion information reconstructed for the i-1th frame) can also be calculated based on the reconstructed positions of the motion vertices of any two frames before the i-1th frame (not limited to the i-3th and i-2th frames) and the i-1th frame, and taken as the predicted motion information of the motion vertex in the current frame.

[0273] The first three frames (e.g. the 1st, 2nd and 3rd frames) in the frame sequence can be encoded by using the currently known single-frame compression method, which is not limited herein.

[0274] In other embodiments, when determining the predicted motion information of the motion vertex of the current frame, the actual motion information of the previous frame (e.g. the i-1th frame) of the current frame (the actual motion information reconstructed for the i-1th frame) is not necessarily taken as the predicted motion information of the current frame, but the actual motion information of any one of the frames (e.g. the i-2th, i-3th or i-4th frame) that has been encoded before the current frame (also the actual motion information reconstructed) can be taken as the predicted motion information of the current frame. Therefore, when determining the predicted motion information of the motion vertex of the current frame, the N historical frames referred to are not necessarily the first three frames that have been encoded before the i-1th frame and adjacent to the current frame in the display order. In other embodiments, the N historical frames shown in FIG. 2B can also be the i-4th, i-2th and i-1th frames, or the i-5th, i-4th and i-3th frames (the reconstructed actual motion information of the i-3th frame is calculated as the predicted motion information of the i-1th frame), or the i-5th, i-3th and i-1th frames, etc. In other words, the N historical frames can be any at least three frames that have been encoded before the i-1th frame, which is not limited herein.

[0275] The predicted motion information includes predicted motion trend information of the motion vertex in the time slot corresponding to the i-1th frame, which in the present embodiment can be specifically the reconstructed result of the actual motion trend information of the motion vertex in the time slot corresponding to the i-1th frame.

[0276] Then, when determining the reconstruction result of the actual motion trend information, not only the reconstruction position of the motion apex of each of the N historical frames is referred to, but also the time slot corresponding to each of the N historical frames is referred to, the time slot being the reciprocal of the instantaneous frame rate of the corresponding frame. For example, the time slot corresponding to the i-1th frame is the time interval between the i-2th frame and the i-1th frame, which is the reciprocal of the instantaneous frame rate of the i-1th frame.

[0277] S3001b: The encoding end determines the actual motion information of the motion apex of the i th frame based on the original position of the motion apex of each of the i-2th frame, the i-1th frame and the i th frame and the instantaneous frame rate of each frame (for example, the instantaneous frame rate of the i-1th frame and the i th frame, and optionally, the instantaneous frame rate of the i-2th frame).

[0278] Wherein, S3001a and S3001b can be executed in series or in parallel, and the present application does not limit this.

[0279] Wherein, the original position of the motion apex of each of the i-2th frame, the i-1th frame and the i th frame in S3001b can be realized by obtaining the original data of the M historical frames shown in FIG. 2B and the original data of the current frame, where M = 2, and the M historical frames are the i-2th frame and the i-1th frame. The way of obtaining the instantaneous frame rate of the M historical frames is the same as the way of obtaining the instantaneous frame rate of each frame in S3001a, which will not be described here.

[0280] In the present embodiment, in order to improve the calculation accuracy of the actual motion information of the current frame (i th frame), the original position of the motion apex in the current frame and the original position of the motion apex in each of the M historical frames and the instantaneous frame rate of each frame can be used to calculate the actual motion information of the current frame, and in the present example, the M historical frames are the two frames (specifically, the i-2th frame and the i-1th frame) that have been encoded before the current frame (for example, the i th frame).

[0281] In other embodiments, when calculating the actual motion information of a frame (for example, the i-1th frame or the i th frame, which is not limited here) in the frame sequence, the M historical frames referred to are not limited to two frames that have been encoded before the frame, but can also be a larger number of frames (for example, 3 frames or 4 frames), that is, M ≥ 2, M being an integer, and the M historical frames are not limited to frames adjacent to the frame in display order, but can also be non-adjacent frames.

[0282] When the actual motion information of a frame is calculated, if the M historical frames referred to are three or more frames that have been encoded before the frame, the actual motion information of the frame can be calculated by combining the implementation principle of the case where the M historical frames referred to are two frames and using an algorithm such as weighted average or average to calculate the obtained multiple actual motion information, without limitation.

[0283] For example, when the actual motion information of the i-th frame is calculated, the M historical frames referred to include the (i-3)-th frame, the (i-2)-th frame and the (i-1)-th frame. When the actual motion information of the motion vertex of the i-th frame is determined based on the original positions of the motion vertices of the (i-3)-th frame, the (i-2)-th frame, the (i-1)-th frame and the i-th frame, the actual motion information of the motion vertex of the (i-1)-th frame can be calculated based on the original positions of the motion vertices of the (i-3)-th frame, the (i-2)-th frame and the (i-1)-th frame, the actual motion information of the motion vertex of the i-th frame can be calculated based on the original positions of the motion vertices of the (i-3)-th frame, the (i-2)-th frame and the i-th frame, and then the actual motion information of the motion vertex of the (i-1)-th frame and the actual motion information of the motion vertex of the i-th frame are averaged (or weighted summation or other operations are performed on the actual motion information of each frame) to obtain the final actual motion information of the motion vertex of the i-th frame.

[0284] S3001a details how to determine the actual motion information (essentially the reconstructed actual motion information) of the motion vertex of the (i-1)-th frame by using the vertex positions (specifically the reconstructed positions) of the (i-3)-th frame, the (i-2)-th frame and the (i-1)-th frame and the corresponding instantaneous frame rates. Similarly, in S3001b, a similar method can be used to determine the actual motion information (not the reconstructed actual motion information) of the motion vertex of the i-th frame by using the vertex positions (specifically the original positions) of the (i-2)-th frame, the (i-1)-th frame and the i-th frame and the instantaneous frame rates of these frames.

[0285] The actual motion information includes actual motion trend information of the motion vertex in a time slot corresponding to the i-th frame.

[0286] When the actual motion trend information is determined, not only the original positions of the motion vertices of the M historical frames are referred to, but also the time slots corresponding to the M historical frames are referred to. The time slot corresponding to each frame is the reciprocal of the instantaneous frame rate of the corresponding frame. For example, the time slot corresponding to the i-th frame is the time interval between the (i-1)-th frame and the i-th frame, which is the reciprocal of the instantaneous frame rate of the i-th frame.

[0287] S3002: The encoder can calculate the residual error between the predicted motion information of the motion vertex of the i-th frame and the actual motion information of the motion vertex of the i-th frame, and encode the residual error to obtain the bitstream.

[0288] wherein each motion vertex has a predicted motion information and an actual motion information, the encoder can calculate the difference between the actual motion information and the predicted motion information of each motion vertex to obtain the residual of each motion vertex, and then entropy encode the residual to realize the encoding of the original position of each motion vertex in the current frame.

[0289] wherein the entropy encoding can include but not limited to: Shannon encoding, Huffman encoding and arithmetic coding, etc., which are not limited here.

[0290] In addition, whether the actual motion information or the predicted motion information of the current frame, can include the offset acceleration of the motion vertex in at least one offset direction.

[0291] For example, when calculating the residual of a certain motion vertex in the i-th frame, the actual offset acceleration and the predicted offset acceleration of the certain vertex in the i-th frame in the X-axis direction of the three-dimensional Cartesian coordinate system can be calculated, and the actual offset rate and the predicted offset rate of the certain vertex in the i-th frame in the Y-axis direction of the three-dimensional Cartesian coordinate system can be calculated, and the actual offset acceleration and the predicted offset acceleration (or offset rate) of the certain vertex in the i-th frame in the Z-axis direction of the three-dimensional Cartesian coordinate system can be calculated. And encode the three residuals to encode the i-th frame into a code stream.

[0292] In the corresponding embodiment of FIG. 3A, the encoder can use the original positions of the motion vertices of at least two adjacent historical frames before the current frame and the instantaneous frame rates of the historical frames to calculate the actual motion information of the motion vertex of the current frame, so as to improve the calculation accuracy of the actual motion information. And the encoder can use the reconstructed positions of the motion vertices of at least three adjacent historical frames before the current frame and the instantaneous frame rates of the historical frames to calculate the actual motion information (which is the reconstructed actual motion information) of the motion vertex of the previous frame of the current frame, and use the actual motion information of the motion vertex of the previous frame of the current frame as the predicted motion information of the motion vertex of the current frame. Since the actual motion trend information of the motion vertex of the adjacent frame is very close in the scene where the motion vertex does not move, using the actual motion information of the motion vertex of the previous frame of the current frame as the predicted motion information of the motion vertex of the current frame can improve the closeness between the predicted motion information of the current frame and the actual motion information of the current frame, thereby reducing the encoding residual, reducing the encoding amount, improving the compression ratio, shortening the code stream length and reducing the code rate, improving the transmission efficiency and storage efficiency.

[0293] FIG. 3B is a flowchart of a method for encoding a current frame, which illustrates the principle of the encoding method of the present application by taking the encoding of a motion vertex (e.g., vertex P) in the current frame as an example. FIG. 3B can be implemented based on the architecture shown in FIG. 1A, the flows shown in FIG. 1B, FIG. 1C, FIG. 2A to FIG. 2D, FIG. 3A and FIG. 3B, but is not limited to the above-mentioned architecture and flows.

[0294] As shown in FIG. 3B, the process of encoding each motion vertex in the current frame is introduced by taking the encoding of vertex P in the current frame (i-th frame) as an example.

[0295] It should be understood that even if the vertex in the current frame moves at a constant speed in the sequence of three-dimensional meshes, the process of encoding and decoding the motion vertex of the present application is also applicable to the vertex moving at a constant speed.

[0296] FIG. 3B shows the three-dimensional mesh diagrams of the i-th frame to be encoded, the i-3-th frame, the i-2-th frame and the i-1-th frame which have been encoded.

[0297] Referring to FIG. 3B, the i-3-th frame, the i-2-th frame and the i-1-th frame can share the same three-dimensional coordinate system, but the position of vertex P in different frames has changed, so that the original positions of vertex P in the i-3-th frame, the i-2-th frame, the i-1-th frame and the i-th frame in the same three-dimensional coordinate system are P i-3 , P i-2 , P i-1 , P i , respectively. In the case that the position of vertex P moves at a non-constant speed in the sequence of three-dimensional meshes, the position of vertex P is shifted from position P i-3 to position P i-2 , position P i-1 and position P i , respectively. For example, the instantaneous frame rate of the i-th frame in the sequence of frames is fi, and the time interval (i.e., the time slot corresponding to the i-th frame) Δti between the i-1-th frame and the i-th frame is 1 / fi.

[0298] For the convenience of description, the encoding process of the present application is illustrated by taking the original positions of vertex P in the i-3-th frame, the i-2-th frame and the i-1-th frame as the same as the reconstructed positions of vertex P in the i-3-th frame, the i-2-th frame and the i-1-th frame.

[0299] Specifically, when the encoding end calculates the actual motion information (i.e., the reconstructed actual motion information, such as the acceleration in at least one shift direction) of vertex P in the i-1-th frame based on the reconstructed positions of vertex P in the i-3-th frame, the i-2-th frame and the i-1-th frame as the prediction motion information of vertex P in the i-th frame, the following process can be used to achieve this:

[0300] The encoding end can be based on the position P of vertex P. i-3 and position P i-2 To calculate the position of vertex P from the position P in the (i-3)th frame. i-3 Offset to the position P of vertex P in frame i-2. i-2 Offset displacement Δ in at least one offset direction s2 (e.g., the difference between the x components); and based on this offset displacement Δ s2 Given the time interval Δt2, calculate the offset rate v of vertex P from frame i-3 to frame i-2 in at least one of the aforementioned offset directions. i-2 For example, v i-2 =Δ s2 / Δt2;

[0301] Similarly, the encoding end can be based on the position P of vertex P. i-2 and position P i-1 To calculate the offset displacement Δ of vertex P from frame i-2 to frame i-1 in at least one of the aforementioned offset directions. s1 (e.g., the difference between the x components); and based on this offset displacement Δ s1 Given the time interval Δt1, calculate the offset rate v of P from frame i-2 to frame i-1 in at least one of the aforementioned offset directions. i-1 For example, v i-1 =Δ s1 / Δt1;

[0302] Then, the encoder can base its work on the aforementioned offset rate v of vertex P in the (i-2)th frame. i-2 And the aforementioned offset rate v of vertex P in the (i-1)th frame i-1 The position P of vertex P in the (i-2)th frame is calculated using the time interval Δt(i-1) between the (i-2)th and (i-1)th frames. i-2 Offset to position P in frame i-1 i-1 The offset acceleration a in at least one of the aforementioned offset directions i-1 For example, a i-1 =(v i-1 -v i-2 ) / Δt(i-1);

[0303] In this way, the encoder can calculate the reconstruction information of the actual motion information of vertex P in the (i-3), (i-2), and (i-1)th frames, as well as the corresponding instantaneous frame rate, based on the reconstruction position of vertex P in each of the (i-3), (i-2), and (i-1)th frames, and the corresponding instantaneous frame rate (e.g., the offset acceleration a in at least one offset direction mentioned above). i-1 ), to serve as the predicted motion information for vertex P of the i-th frame.

[0304] Similarly, please refer to Figure 3B to calculate the position P of vertex P from the (i-2)th frame. i-2 Offset to position P in frame i-1 i-1 The offset acceleration a in at least one of the aforementioned offset directions i-1 The principle is similar; the encoding end can also base its work on the original positions of vertex P in the (i-2)th, (i-1)th, and ith frames (e.g., position P). i-2 Location P i-1 Location P i The position P of vertex P in the (i-1)th frame is calculated using the corresponding instantaneous frame rate. i-1 Offset to position P in frame i i The offset acceleration a in at least one of the aforementioned offset directions i To obtain the actual motion information of vertex P in the i-th frame (e.g., the offset acceleration a in at least one offset direction mentioned above). i ).

[0305] Finally, please refer to Figure 3B. The encoder can control the offset acceleration a. i With offset acceleration a i-1 The difference is encoded to encode the residual between the actual motion information and the predicted motion information of vertex P in the i-th frame, thereby completing the encoding of vertex P in the i-th frame. The encoding process for other vertices in the i-th frame is similar to that for vertex P, and will not be described in detail here.

[0306] Figure 3C is a flowchart of a method for encoding the current frame (a three-dimensional mesh to be encoded) provided in an embodiment of this application. The method can be implemented based on the architecture shown in Figure 1A, the process shown in Figure 2B, Figure 3A, and Figure 3B, but is not limited to combining the architecture and process shown in Figure 1A, Figure 1B, Figure 1C, Figure 2B, Figure 3A, and Figure 3B.

[0307] In the method flow shown in Figure 3C, the encoding end can convert the vertex positions of each moving vertex in the historical frame and the current frame from the three-dimensional Cartesian coordinate system to the cylindrical coordinate system, so as to calculate the actual motion information and predicted motion information of the vertex in the three components of height, radius and angle, so as to realize the encoding of the vertex of the current frame.

[0308] In the method flow shown in Figure 3C, regardless of whether it is actual motion information or predicted motion information, the motion information may include the offset motion information of the current frame's moving vertex in the three offset directions of the cylindrical coordinate system. The offset motion information includes the offset rate of the moving vertex in the height direction (also expressed as the height component), the offset rate of the moving vertex in the radius direction (also expressed as the radius component), and the offset acceleration of the vertex in the angular direction (also expressed as the angular component).

[0309] wherein the offset rate of the motion vertex on the height component is also referred to as the height change rate of the vertex, the offset rate of the motion vertex on the radius component is also referred to as the radius change rate of the vertex, and the offset acceleration of the motion vertex on the angle component is also referred to as the angular acceleration of the vertex.

[0310] As shown in FIG. 3C, the method can mainly include the following steps:

[0311] S501: The encoding end converts the reconstructed positions of the motion vertex in each of the i-3th frame, the i-2th frame and the i-1th frame from the three-dimensional Cartesian coordinate system to the cylindrical coordinate system to obtain the height, radius and angle (specifically, the reconstructed height, reconstructed radius and reconstructed angle) of the motion vertex in each of the i-3th frame, the i-2th frame and the i-1th frame; and the encoding end converts the original positions of the motion vertex in each of the i-2th frame, the i-1th frame and the i-th frame from the three-dimensional Cartesian coordinate system to the cylindrical coordinate system to obtain the height, radius and angle (specifically, the original height, original radius and original angle) of the motion vertex in each of the i-2th frame, the i-1th frame and the i-th frame.

[0312] The reconstructed positions of the motion vertex in each of the i-3th frame, the i-2th frame and the i-1th frame can be obtained by S300a shown in FIG. 2B.

[0313] The original positions of the motion vertex in each of the i-2th frame, the i-1th frame and the i-th frame can be obtained by S101 shown in FIG. 1A.

[0314] Specifically, the encoding end can convert the coordinates (x, y, z) of the reconstructed position of the motion vertex of the i-3th frame obtained by the three-dimensional Cartesian coordinate system to the cylindrical coordinate system with the barycenter of the i-3th frame as the coordinate system origin to obtain the component values (h, r, θ) of the height component, the radius component and the angle component of the motion vertex of the i-3th frame, wherein h represents the height, r represents the radius, and θ represents the angle. The i-3th frame is a three-dimensional grid, the three-dimensional grid includes a plurality of vertices, and the barycenter of the i-3th frame can be the barycenter of all vertices in the three-dimensional grid or other determined barycenter, which is not limited here.

[0315] Similarly, the encoding end can convert the reconstructed positions of the motion vertex in each of the i-2th frame and the i-1th frame obtained by the three-dimensional Cartesian coordinate system to the cylindrical coordinate system with the barycenter of the i-3th frame as the coordinate system origin, so that different frames share the same cylindrical coordinate system, so as to facilitate accurate calculation of the motion information of the vertex, to obtain the component values of the height component, the radius component and the angle component of the motion vertex in each of the i-2th frame and the i-1th frame.

[0316] Of course, the origin of the same coordinate system shared by different frames is not limited to the barycenter of the i-3th frame, but can also be the barycenter of other frames, for example, the barycenter of the first frame in the frame sequence, which is not limited here, as long as the vertices of the frame sequence share the same coordinate system.

[0317] Similarly, the encoding end can convert the original positions of the motion vertices in each of the i-2th, i-1th and i th frames from the three-dimensional Cartesian coordinate system to the cylindrical coordinate system to obtain the height, radius and angle (specifically, the original height, original radius and original angle) of the motion vertices in each of the i-2th, i-1th and i th frames.

[0318] Taking the vertex P shown in FIG. 3B as an example, the encoding end can obtain the vertex coordinates P i-3 , P i-2 , P i-1 , P i of the vertex P in the i-3th, i-2th, i-1th and i th frames respectively through coordinate system conversion in this step.

[0319] S502a: Based on the reconstructed height, reconstructed radius and reconstructed angle of the motion vertices of the i-3th, i-2th and i-1th frames, and the instantaneous frame rate of each frame (for example, the i-2th and i-1th frames), the encoding end calculates the reconstructed offset speed of the vertex in the i-1th frame in the height component, the reconstructed offset speed of the vertex in the i-1th frame in the radius component, and the reconstructed offset acceleration of the vertex in the i-1th frame in the angle component to obtain the reconstructed actual motion information of the motion vertex in the i-1th frame (which can be used as the predicted motion information of the motion vertex in the i th frame).

[0320] As described in the above embodiments, the actual motion information can include the offset acceleration of the motion vertex in the i-1th frame in at least one of the three offset directions, and in the embodiment of FIG. 3C, the calculated offset acceleration corresponds to the offset direction of the angle component. In other embodiments, the calculated offset acceleration can also correspond to at least one of the height component and the radius component, which is not limited here.

[0321] For ease of illustration, the instantaneous frame rate of each frame in the three-dimensional grid sequence is taken as an example to illustrate the method of the present application, and the time slot of each frame is the time interval Δt, where Δt = 1 / f, and f is the frame rate of the three-dimensional grid sequence.

[0322] In S502a, taking the vertex P as an example in combination with FIG. 3B, when calculating the reconstructed offset speed of the motion vertex in the i-1th frame in the height direction (the actual offset speed is calculated, and the same reasoning applies), the encoding end can first calculate the reconstructed offset speed of the vertex P in the i-1th frame in the height direction based on the reconstructed height of the vertex P in the i-1th frame and the instantaneous frame rate of the i-1th frame, that is, the reconstructed offset speed of the vertex P in the i-1th frame in the height direction is calculated as follows: i-2offset displacement Δh i-1 offset displacement Δh i-1 wherein the offset displacement Δh i-1 is the difference between the height component of the position P i-1 and the height component of the position P i-2 ; and then, based on the offset displacement Δh i-1 and the time interval Δt, the position P i-2 of the vertex P in the i-1th frame is calculated from the position P i-1 offset rate in the height h direction For example The time interval Δt is the time interval between two adjacent frames.

[0323] In S502a, taking the vertex P as an example in combination with FIG. 3B, when calculating the reconstructed offset rate of the vertex P in the i-1th frame in the radius direction (and the actual offset rate is calculated in the same way), the encoding end can first calculate the offset displacement Δr i-2 of the vertex P from the position P i-1 in the radius r component i-1 wherein the offset displacement Δr i-1 is the difference between the radius component of the position P i-1 and the radius component of the position P i-2 ; and then, based on the offset displacement Δr i-1 and the time interval Δt, the position P i-2 of the vertex P in the i-1th frame is calculated from the position P i-1 offset rate in the radius r direction For example

[0324] In S502a, taking the vertex P as an example in combination with FIG. 3B, when calculating the reconstructed offset acceleration of the vertex P in the i-1th frame in the angle component (and the actual reconstructed offset acceleration is calculated in the same way), the encoding end can first calculate the offset rate of the vertex P in the angle component in the i-2th frame and the offset rate of the vertex P in the angle component in the i-2th frame Then, based on the two offset rates and the time interval Δt, the offset acceleration of the vertex P in the i-1th frame in the angle component is calculated For example

[0325] wherein, when calculating the above offset rate and the above offset rate , the implementation process is similar to the implementation principle of calculating the offset rate of the vertex P in the i-1th frame in the radius direction. Specifically, when calculating the offset rate When the time interval Δt is known, the encoding end can calculate the displacement of vertex P from position P i-3 to position P i-2 the displacement angle Δθ on the angle component i-2 wherein the displacement angle Δθ i-2 is the difference between the angle component of position P i-3 and the angle component of position P i-2 ; then, based on the displacement angle Δθ i-2 and the time interval Δt, the encoding end can calculate the displacement of vertex P from position P i-3 to position P i-2 the displacement rate on the angle direction (also expressed as angular velocity), for example

[0326] Similarly, the encoding end can use the angle component of vertex P at position P i-2 in the i-2 frame, the angle component of vertex P at position P i-1 in the i-1 frame, and the time interval Δt to calculate the displacement rate of vertex P on the angle component in the i-1 frame (also expressed as angular velocity) as the predicted displacement rate of vertex P on the angle component in the i frame.

[0327] S502b: The encoding end calculates the actual displacement rate of the motion vertex in the i frame on the height direction, the actual displacement rate of vertex P in the i frame on the radius component, and the actual displacement acceleration of vertex P in the i frame on the angle component based on the original height, the original radius, and the original angle of the motion vertex in the i-2 frame, the i-1 frame, and the i frame, respectively, and the instantaneous frame rate of each frame (e.g., the i-1 frame and the i frame), to obtain the actual motion information of the motion vertex in the i frame.

[0328] S502a and S502b can be executed in series or in parallel, and the present application does not limit this.

[0329] The execution principle of S502b is similar to that of S502a, and the only difference is that when calculating the actual displacement rate or the actual displacement acceleration, the components of the cylindrical coordinate system used are actual components (e.g., original height, original radius, and original angle), which will not be described here.

[0330] In combination with the embodiment of FIG. 3B, taking vertex P as an example, the actual displacement rate of vertex P in the i frame on the height component the actual displacement rate of vertex P in the i frame on the radius component the actual displacement acceleration of vertex P in the i frame on the angle component

[0331] S503: The encoding end encodes the residual of the actual motion information of the vertex of the i-th frame and the predicted motion information of the vertex of the i-th frame to obtain the code stream of the i-th frame.

[0332] Specifically, taking the vertex P as an example, the encoding end can calculate the difference between the actual offset rate of the vertex P of the i-th frame in the height component and the predicted offset rate in the height component (for example, ) to obtain the residual of the vertex P of the i-th frame in the height component; and the encoding end can calculate the difference between the actual offset rate of the vertex P of the i-th frame in the radius component and the predicted offset rate in the radius component (for example, ) to obtain the residual of the vertex P of the i-th frame in the radius component; and the encoding end can calculate the difference between the actual offset acceleration of the vertex P of the i-th frame in the angle component and the predicted offset acceleration in the angle component (for example, ) to obtain the residual of the vertex P of the i-th frame in the angle component.

[0333] In the method flow shown in FIG. 3C, the encoding end encodes not the residual of the coordinate value between the predicted position and the accurate position of the vertex of the current frame, but the residual of the motion information between the predicted motion information and the accurate motion information of the vertex, in the case of non-uniform motion of the vertex in the three-dimensional mesh, the offset acceleration of the vertex of the adjacent frame is close, by mining the offset acceleration of the deformable three-dimensional mesh (such as cloth) in the offset direction, the residual of the actual offset acceleration and the predicted offset acceleration of the vertex of the current frame is encoded, which can reduce the data amount of the encoded residual, thereby improving the compression rate of the deformable three-dimensional mesh sequence of the flexible body such as cloth or the elastic body, and further reducing the code rate.

[0334] And in the method flow shown in FIG. 3C, the encoding end converts the vertex positions of the historical frames and the current frame into the cylindrical coordinate system to encode the current frame, so that when the moving vertex performs rotational motion in the three-dimensional mesh sequence (for example, a person wearing a virtual clothes performs ballet dance, driving the virtual clothes to perform rotational motion), the angular acceleration of the moving vertex in the angle component changes very little between frames, and then the three-dimensional mesh is encoded by calculating the residual of the actual offset acceleration and the predicted offset acceleration in the angle component and encoding the residual, so that the encoding end of the present application can match the motion mode of the three-dimensional mesh by calculating the components (height, radius, angle) of the motion information, and further make the predicted motion information and the actual motion information more close, thereby improving the data compression rate.

[0335] FIG. 3D is another method flowchart for encoding the current frame (the three-dimensional mesh to be encoded) according to an embodiment of the present application, which can be implemented based on the architecture shown in FIG. 1A, the flowcharts shown in FIGS. 1B, 1C, 2A-2D, 3A, and 3B, but is not limited to the above-mentioned architecture and flowcharts.

[0336] The method flowchart shown in FIG. 3D is identical to most of the method flowchart shown in FIG. 3C, with the difference being that in the method flowchart shown in FIG. 3D, the motion information (including the predicted motion information and the actual motion information) calculated by the encoding end for the current frame is three offset accelerations on three offset components, not only the offset acceleration of the motion vertex on the angle component as in the method flowchart shown in FIG. 3C, but also the offset acceleration of the motion vertex on the radius component and the offset acceleration of the motion vertex on the height component.

[0337] Specifically, taking the vertex P shown in FIG. 3B as an example, the actual motion information can include the actual offset acceleration of the vertex P of the i-th frame on the height component, the actual offset acceleration of the vertex P of the i-th frame on the radius component, and the actual offset acceleration of the vertex P of the i-th frame on the angle component. The predicted motion information can include the reconstructed actual offset acceleration of the vertex P of the i-1-th frame on the height component, the reconstructed actual offset acceleration of the vertex P of the i-1-th frame on the radius component, and the reconstructed actual offset acceleration of the vertex P of the i-1-th frame on the angle component.

[0338] As shown in FIG. 3D, the method can mainly include the following steps:

[0339] S601: The encoding end converts the reconstructed positions of the motion vertices of each of the i-3-th, i-2-th, and i-1-th frames from the three-dimensional Cartesian coordinate system to the cylindrical coordinate system to obtain the height, radius, and angle (specifically, the reconstructed height, reconstructed radius, and reconstructed angle) of the motion vertices of each of the i-3-th, i-2-th, and i-1-th frames, and converts the original positions of the motion vertices of each of the i-2-th, i-1-th, and i-th frames from the three-dimensional Cartesian coordinate system to the cylindrical coordinate system to obtain the height, radius, and angle (specifically, the original height, original radius, and original angle) of the motion vertices of each of the i-2-th, i-1-th, and i-th frames.

[0340] The execution principle of S601 is identical to that of S501 in the method flowchart shown in FIG. 3C, which will not be described here again.

[0341] S602a: The encoding end calculates the reconstructed offset rate of the motion vertex of the i-2th frame in the height component, the reconstructed offset rate of the motion vertex of the i-1th frame in the height component, the reconstructed offset rate of the motion vertex of the i-2th frame in the radius component, the reconstructed offset rate of the motion vertex of the i-1th frame in the radius component, the reconstructed offset rate of the motion vertex of the i-2th frame in the angle component, and the reconstructed offset rate of the motion vertex of the i-1th frame in the angle component based on the reconstructed height, the reconstructed radius, and the reconstructed angle of the vertex of the i-3th frame, the i-2th frame, and the i-1th frame, and the instantaneous frame rate of each frame (e.g., the i-2th frame and the i-1th frame).

[0342] After S601, the encoding end can perform S602a.

[0343] Specifically, taking the vertex P above as an example, the encoding end can calculate the offset rate of the vertex P of the i-1th frame in the height component based on the height component value of the vertex P of the i-2th frame and the i-1th frame respectively, and the instantaneous frame rate of the i-1th frame (here, the reconstructed offset rate), and the specific implementation process can refer to the related introduction of the embodiment of FIG. 3C, which is not described herein again. Similarly, the encoding end can calculate the offset rate of the vertex P of the i-2th frame in the height component based on the height component value of the vertex P of the i-3th frame and the i-2th frame respectively, and the instantaneous frame rate of the i-2th frame (here, the reconstructed offset rate).

[0344] Similarly, the encoding end can calculate the offset rate of the vertex P of the i-1th frame in the radius component based on the radius component value of the vertex P of the i-2th frame and the i-1th frame respectively, and the instantaneous frame rate of the i-1th frame (here, the reconstructed offset rate), and the specific implementation process can refer to the related introduction of the embodiment of FIG. 3C, which is not described herein again. Similarly, the encoding end can calculate the offset rate of the vertex P of the i-2th frame in the radius component based on the radius component value of the vertex P of the i-3th frame and the i-2th frame respectively, and the instantaneous frame rate of the i-2th frame (here, the reconstructed offset rate).

[0345] Similarly, the encoding end can calculate the offset rate of the vertex P of the i-1th frame in the angle component based on the angle component value of the vertex P of the i-2th frame and the i-1th frame respectively, and the instantaneous frame rate of the i-1th frame (here, the reconstructed offset rate), and the specific implementation process can refer to the related introduction of the embodiment of FIG. 3C, which is not described herein again. Similarly, the encoding end can calculate the offset rate of the vertex P of the i-2th frame in the angle component based on the angle component value of the vertex P of the i-3th frame and the i-2th frame respectively, and the instantaneous frame rate of the i-2th frame reconstructed displacement rate of the motion vertex on the height component of the i-2th frame, the reconstructed displacement rate of the motion vertex on the height component of the i-1th frame, and the instantaneous frame rate of the i-1th frame, to obtain the reconstructed actual motion information of the motion vertex of the i-1th frame (which can be used as the predicted motion information of the motion vertex of the i-1th frame).

[0346] S603a: the encoding end calculates the reconstructed displacement acceleration of the motion vertex on the height component of the i-1th frame based on the reconstructed displacement rate of the motion vertex on the height component of the i-2th frame, the reconstructed displacement rate of the motion vertex on the height component of the i-1th frame, and the instantaneous frame rate of the i-1th frame, and calculates the reconstructed displacement acceleration of the motion vertex on the radius component of the i-1th frame based on the reconstructed displacement rate of the motion vertex on the radius component of the i-2th frame, the reconstructed displacement rate of the motion vertex on the radius component of the i-1th frame, and the instantaneous frame rate of the i-1th frame, and calculates the reconstructed displacement acceleration of the motion vertex on the angle component of the i-1th frame based on the reconstructed displacement rate of the motion vertex on the angle component of the i-2th frame, the reconstructed displacement rate of the motion vertex on the angle component of the i-1th frame, and the instantaneous frame rate of the i-1th frame, to obtain the reconstructed actual motion information of the motion vertex of the i-1th frame (which can be used as the predicted motion information of the motion vertex of the i-1th frame).

[0347] After S602a, the encoding end can perform S603a, and taking the vertex P as an example, the encoding end can calculate the reconstructed displacement acceleration of the vertex P on the angle component of the i-1th frame based on the displacement rate of the vertex P on the angle component of the i-2th frame, the displacement rate of the vertex P on the angle component of the i-1th frame, and the instantaneous frame rate of the i-1th frame. The specific implementation process can refer to the related description of the embodiment of FIG. 3C, which will not be described here. Similarly, the encoding end can calculate the reconstructed displacement acceleration of the vertex P on the height component of the i-1th frame based on the reconstructed displacement rate of the vertex P on the height component of the i-2th frame, the reconstructed displacement rate of the vertex P on the height component of the i-1th frame, and the instantaneous frame rate of the i-1th frame.

[0348]

[0349] Similarly, the encoding end can calculate the reconstructed displacement acceleration of the vertex P on the radius component of the i-1th frame based on the reconstructed displacement rate of the vertex P on the radius component of the i-2th frame, the reconstructed displacement rate of the vertex P on the radius component of the i-1th frame, and the instantaneous frame rate of the i-1th frame.

[0350] ​​​S602b: The encoding end calculates the actual displacement rate of the motion vertex of the i-1th frame in the height component, the actual displacement rate of the motion vertex of the ith frame in the height component, the actual displacement rate of the motion vertex of the i-1th frame in the radius component, the actual displacement rate of the motion vertex of the ith frame in the radius component, the actual displacement rate of the motion vertex of the i-1th frame in the angle component, and the actual displacement rate of the motion vertex of the ith frame in the angle component based on the original height, the original radius, and the original angle of the motion vertex of the i-2th frame, the i-1th frame, and the ith frame, and the instantaneous frame rate of each frame (the i-1th frame and the ith frame).

[0351] After S601, the encoding end can also perform S602b.

[0352] The present application does not limit the execution order of S602a and S602b, which can be executed in series or in parallel.

[0353] The implementation principle of S602b is similar to that of S602a, which will not be described here.

[0354] S603b: The encoding end calculates the actual displacement acceleration of the motion vertex of the ith frame in the height component based on the actual displacement rate of the motion vertex of the i-1th frame in the height component, the actual displacement rate of the motion vertex of the ith frame in the height component, and the instantaneous frame rate of the ith frame; calculates the actual displacement acceleration of the motion vertex of the ith frame in the radius component based on the actual displacement rate of the motion vertex of the i-1th frame in the radius component, the actual displacement rate of the motion vertex of the ith frame in the radius component, and the instantaneous frame rate of the ith frame; and calculates the actual displacement acceleration of the motion vertex of the ith frame in the angle component based on the actual displacement rate of the motion vertex of the i-1th frame in the angle component, the actual displacement rate of the motion vertex of the ith frame in the angle component, and the instantaneous frame rate of the ith frame, to obtain the actual motion information of the motion vertex of the ith frame.

[0355] After S602b, the encoding end can perform S603b, and taking the vertex P as an example, the encoding end can calculate the actual displacement acceleration of the vertex P of the ith frame in the angle component based on the actual displacement rate of the vertex P of the i-1th frame and the ith frame in the angle component and the instantaneous frame rate of the ith frame The specific implementation process can refer to the related description of the embodiment of FIG. 3C, which will not be described here.

[0356] Similarly, the encoding end can calculate the actual displacement acceleration of the vertex P of the ith frame in the height component based on the actual displacement rate of the vertex P of the i-1th frame and the ith frame in the height component and the instantaneous frame rate of the ith frame

[0357] Similarly, the encoding end can calculate the actual offset acceleration of the vertex P of the i-th frame in the radius component based on the actual offset rate of the vertex P of the i-1-th frame and the i-th frame in the radius component and the instantaneous frame rate of the i-th frame

[0358] S604: The encoding end can calculate the residual of the offset acceleration in the three offset directions based on the actual offset acceleration of the vertex of the i-th frame in the three offset directions (take the vertex P as an example, specifically, the actual offset acceleration of the vertex P in the height component the actual offset acceleration of the vertex P in the radius component the actual offset acceleration of the vertex P in the angle component ) and the predicted offset acceleration of the vertex of the i-th frame in the three offset directions (take the vertex P as an example, specifically, the reconstructed offset acceleration of the vertex P in the height component the reconstructed offset acceleration of the vertex P in the radius component the reconstructed offset acceleration of the vertex P in the angle component ), and encode the residual to obtain the code stream.

[0359] Specifically, take the vertex P as an example, the encoding end can calculate the residual of the actual offset acceleration of the vertex P of the i-th frame in the height component and the offset acceleration of the vertex P of the i-th frame in the height component

[0360] In addition, the encoding end can calculate the residual of the actual offset acceleration of the vertex P of the i-th frame in the radius component and the offset acceleration of the vertex P of the i-th frame in the radius component

[0361] In addition, the encoding end can calculate the residual of the actual offset acceleration of the vertex P of the i-th frame in the angle component and the offset acceleration of the vertex P of the i-th frame in the radius component

[0362] Then, the above three residuals of the vertex P are encoded to realize the encoding of the vertex P of the i-th frame, and the encoding process of other vertices of the i-th frame is the same, which will not be described here, so as to realize the encoding of the i-th frame and obtain the code stream.

[0363] ​​​Compared with the method flow shown in FIG. 3C, in the method flow shown in FIG. 3D, the encoding end can encode the residual of the offset acceleration of the vertex of the current frame in three offset directions, without including the step of encoding the residual of the offset rate in the offset direction, compared with encoding the residual of the offset rate of the predicted vertex of the current frame in one offset direction (which is greatly different from the actual offset rate of the vertex in the offset direction, resulting in a large residual), and encoding the residual of the offset rate. In the embodiments of the present application, the predicted offset acceleration of the current frame in the offset direction can be closer to the actual acceleration, thereby further reducing the residual size of the encoding, thereby improving the compression rate of the deformable three-dimensional mesh sequence of the flexible body or the elastic body, and further reducing the code rate.

[0364] FIG. 3E is another method flow chart for encoding the current frame (the three-dimensional mesh to be encoded) provided by the embodiments of the present application, which can be implemented based on the architecture shown in FIG. 1A, the flows shown in FIGS. 1B, 1C, 2A-2D, 3A, and 3B, but is not limited to the above-mentioned architecture and flow.

[0365] The implementation principle of the method flow shown in FIG. 3E is the same as that of the method flow shown in FIG. 3D, and the difference is that, in the method flow shown in FIG. 3D, the motion information (including the predicted motion information and the actual motion information) calculated by the encoding end for the current frame is three offset accelerations in three offset components (h, r, θ) of the cylindrical coordinate system, while in the method flow shown in FIG. 3E, the encoding end does not need to perform coordinate system conversion on the vertex coordinates of the current frame and the historical frame, but directly calculates the motion information (including the predicted motion information and the actual motion information) of the vertex in the Cartesian coordinate system, which is three offset accelerations in three offset components (x, y, z) of the Cartesian coordinate system.

[0366] Specifically, taking the vertex P shown in FIG. 3B as an example, the coordinates of the vertex P are the coordinates in the three-dimensional Cartesian coordinate system (including x, y, and z directions). The actual motion information of the vertex P of the i-th frame can include the actual offset acceleration of the vertex P of the i-th frame in the x direction, the actual offset acceleration of the vertex P of the i-th frame in the y direction, and the actual offset acceleration of the vertex P of the i-th frame in the z direction. The predicted motion information of the vertex P of the i-th frame can include the reconstructed actual offset acceleration of the vertex P of the i-1-th frame in the x direction, the reconstructed actual offset acceleration of the vertex P of the i-1-th frame in the y direction, and the reconstructed actual offset acceleration of the vertex P of the i-1-th frame in the z direction.

[0367] As shown in FIG. 3E, the method can mainly include the following steps:

[0368] S702a: The encoding end calculates the reconstructed displacement rate of the motion vertex of the i-2th frame in each of the x, y, and z directions based on the reconstructed position of the motion vertex of each of the i-3th, i-2th, and i-1th frames and the instantaneous frame rate of the i-2th and i-1th frames, and calculates the reconstructed displacement rate of the motion vertex of the i-1th frame in each of the x, y, and z directions.

[0369] wherein the reconstructed position of the motion vertex of each of the i-3th, i-2th, and i-1th frames is a coordinate in the Cartesian coordinate system, and the reconstructed position includes an x component value, a y component value, and a z component value.

[0370] Specifically, continuing to use the vertex P as an example, the encoding end can calculate the reconstructed displacement rate of the vertex P of the i-2th frame in each of the x, y, and z directions (respectively denoted by ) based on the vertex position of the vertex P of each of the i-3th and i-2th frames and the instantaneous frame rate of the i-2th frame, and calculate the reconstructed displacement rate of the vertex P of the i-1th frame in each of the x, y, and z directions (respectively denoted by ) based on the vertex position of the vertex P of each of the i-2th and i-1th frames and the instantaneous frame rate of the i-1th frame.

[0371] The execution principle of S702a is similar to that of S602a in the method flow shown in FIG. 3D, which will not be described herein again.

[0372] S703a: The encoding end calculates the reconstructed displacement acceleration of the motion vertex of the i-1th frame in each of the x, y, and z directions based on the reconstructed displacement rate of the motion vertex of the i-2th frame in each of the x, y, and z directions, the reconstructed displacement rate of the motion vertex of the i-1th frame in each of the x, y, and z directions, and the instantaneous frame rate of the i-1th frame, to obtain the reconstructed actual motion information of the motion vertex of the i-1th frame (which can be used as the predicted motion information of the motion vertex of the i-1th frame).

[0373] After S702a, the encoding end can perform S703a. Continuing to use the vertex P as an example, the encoding end can calculate the reconstructed displacement acceleration of the vertex P of the i-1th frame in each of the x, y, and z directions (respectively denoted by ) based on the reconstructed displacement rate of the vertex P of the i-2th frame in each of the x, y, and z directions (respectively denoted by ), the reconstructed displacement rate of the vertex P of the i-1th frame in each of the x, y, and z directions (respectively denoted by ), and the instantaneous frame rate of the i-1th frame. The specific implementation principle is similar to that of S603a in the embodiment of FIG. 3D, which will not be described herein again.

[0374] S702b: The encoding end calculates the actual displacement rate of the motion vertex of the i-1th frame in each of the x, y, and z directions based on the original position of the motion vertex of each of the i-2th, i-1th, and ith frames and the instantaneous frame rate of the i-1th and ith frames, and calculates the actual displacement rate of the motion vertex of the ith frame in each of the x, y, and z directions.

[0375] Continuing to take the vertex P as an example, the encoding end can calculate the actual displacement rate of the vertex P of the i-1th frame in each of the x, y, and z directions (respectively denoted as vx, vy, and vz) based on the original position of the vertex P of each of the i-2th and i-1th frames and the instantaneous frame rate of the i-1th frame. The encoding end can calculate the actual displacement rate of the vertex P of the ith frame in each of the x, y, and z directions (respectively denoted as vx, vy, and vz) based on the original position of the vertex P of each of the i-1th and ith frames and the instantaneous frame rate of the ith frame.

[0376] The implementation principle of S702b is similar to that of S702a, and thus will not be described herein.

[0377] S703b: The encoding end calculates the actual displacement acceleration of the motion vertex of the ith frame in each of the x, y, and z directions based on the actual displacement rate of the motion vertex of the i-1th frame in each of the x, y, and z directions, the actual displacement rate of the motion vertex of the ith frame in each of the x, y, and z directions, and the instantaneous frame rate of the ith frame, to obtain the actual motion information of the motion vertex of the ith frame.

[0378] After S702b, the encoding end can perform S703b. Continuing to take the vertex P as an example, the encoding end can calculate the actual displacement acceleration of the vertex P of the ith frame in each of the x, y, and z directions (respectively denoted as ax, ay, and az) based on the actual displacement rate of the vertex P of the i-1th frame in each of the x, y, and z directions (respectively denoted as vx, vy, and vz), the actual displacement rate of the vertex P of the ith frame in each of the x, y, and z directions (respectively denoted as vx, vy, and vz), and the instantaneous frame rate of the ith frame. The implementation principle of S703b is similar to that of S703a, and thus will not be described herein.

[0379] The implementation principle of S703b is similar to that of S703a, and thus will not be described herein.

[0380] ​​​S704: The encoding end can calculate the residual of the offset acceleration in the three offset directions based on the actual offset acceleration of the vertex of the i-th frame in the x, y, z three offset directions (for example, the vertex P, specifically, the three offset accelerations in the x, y, z directions ), and the offset acceleration of the vertex of the i-th frame in the x, y, z three offset directions (for example, the vertex P, specifically, the three reconstructed offset accelerations in the x, y, z directions ), and encode the residual to obtain the code stream.

[0381] After S703a and S703b, the encoding end can perform S704.

[0382] Specifically, taking the vertex P as an example, the encoding end can calculate the residual of the actual offset acceleration of the vertex P of the i-th frame in the x direction and the predicted offset acceleration of the vertex P of the i-th frame in the x direction

[0383] In addition, the encoding end can calculate the residual of the actual offset acceleration of the vertex P of the i-th frame in the y direction and the predicted offset acceleration of the vertex P of the i-th frame in the y direction

[0384] In addition, the encoding end can calculate the residual of the actual offset acceleration of the vertex P of the i-th frame in the z direction and the predicted offset acceleration of the vertex P of the i-th frame in the z direction

[0385] Then, the above three residuals of the vertex P are encoded to realize the encoding of the vertex P of the i-th frame, and the encoding process of other vertices of the i-th frame is the same, which will not be described here, so as to realize the encoding of the i-th frame to obtain the code stream.

[0386] ​​​Differently from the method flow shown in FIG. 3C and FIG. 3D, in the method flow shown in FIG. 3E, the encoding end does not need to make coordinate system conversion on the position of the motion vertex of each frame, and can directly encode the vertex of the three-dimensional grid established on the three-dimensional Cartesian coordinate system to encode the residual of the offset acceleration of the vertex in the three offset directions of the three-dimensional Cartesian coordinate system. In this way, when the vertex in the three-dimensional grid makes non-uniform motion (for example, a virtual person running at an accelerated speed drives the virtual clothes to make non-uniform motion) with a small change in the motion direction (for example, straight line motion) in the sequence of three-dimensional grids, the acceleration of the motion vertex of the three-dimensional grid under the Cartesian coordinate system can change obviously. Then, the encoding end calculates the offset motion information (including offset acceleration) of the motion vertex in each offset direction of the Cartesian coordinate system, so that the offset direction of the offset acceleration can match the motion direction and motion mode of the non-uniform motion, thereby improving the closeness between the predicted motion information and the actual motion information of the present application, reducing the residual, and improving the compression rate.

[0387] It should be understood that the above-mentioned FIG. 3C, FIG. 3D and FIG. 3E only exemplarily show three method flows of the encoding end of the present application for encoding the current frame. In other embodiments, the present application can also provide more encoding methods, and the object encoded by the encoding method is still the residual of the actual motion information and the predicted motion information of the vertex of the current frame, and the motion information (actual motion information and predicted motion information) is the respective offset motion information in the three offset directions of the three-dimensional coordinate system of the current frame, wherein the offset acceleration in at least one of the respective offset motion information in the three offset directions exists, and the offset motion information in the remaining offset directions can be offset speed or offset acceleration.

[0388] Please return to S300a shown in FIG. 2B. In order to obtain the reconstructed data of the N historical frames, the method flow shown in FIG. 2C or FIG. 2D can be used.

[0389] FIG. 2C is a method flow chart for obtaining the reconstructed data of N historical frames of the current frame (the three-dimensional grid to be encoded) provided by an embodiment of the present application. The method flow can be implemented based on the architecture shown in FIG. 1A, and can be implemented based on FIG. 2A, FIG. 2B and any possible implementation. As shown in FIG. 2C, the method flow includes the following steps:

[0390] S3201a: The encoding end obtains N sets of encoding data corresponding to the N historical frames (corresponding to the N second three-dimensional grids in the invention) of the sequence of three-dimensional grids which have been encoded.

[0391] Wherein, the N historical frames and the N sets of encoding data are one-to-one corresponding.

[0392] The N historical frames are N three-dimensional meshes that have been encoded before the current frame, where N is an integer greater than or equal to 3.

[0393] The encoding end can then obtain N sets of encoding data of the N three-dimensional meshes that have been encoded.

[0394] In some embodiments, when the N historical frames are three-dimensional meshes other than the first three frames (i.e., the first frame, the second frame, and the third frame) in the sequence of three-dimensional meshes, each of the N historical frames is encoded by the encoding method of the present application, and the decoding of the encoding data can be performed using the process shown in FIG. 2C to obtain the reconstruction data of the N historical frames.

[0395] In other embodiments, when the N historical frames include any one of the first frame, the second frame, and the third frame, any one of the first frame, the second frame, and the third frame can be decoded by using any decoding method in the prior art or to be developed in the future to obtain the reconstruction data of the frame, and the reconstruction data of the frames other than the first frame, the second frame, and the third frame in the N historical frames can be obtained using the process shown in FIG. 2C.

[0396] In some embodiments, the encoding data of each of the N historical frames can include the residual between the predicted motion information of the historical frame and the actual motion information of the historical frame.

[0397] S3201b: The encoding end obtains N sets of predicted motion information of the N historical frames.

[0398] The N sets of predicted motion information are N sets of predicted motion information of the motion vertex that moves at a non-uniform speed in the corresponding time slots of the N historical frames in the sequence of three-dimensional meshes, and the N sets of predicted motion information can include N sets of predicted motion trend information and, optionally, N sets of predicted motion rate information.

[0399] The N sets of predicted motion trend information are the predicted motion trends of the motion vertex in the corresponding time slots of the N historical frames.

[0400] The N sets of predicted motion rate information are the predicted motion rates of the motion vertex in the corresponding time slots of the N historical frames.

[0401] The principle of obtaining the predicted motion information (e.g., the predicted motion trend information of the current frame) of each of the N historical frames and the principle of obtaining the predicted motion information (e.g., the predicted motion trend information of the current frame) of the current frame in the various embodiments of the encoding method described above are the same, and thus will not be described herein.

[0402] S3202: The encoding end decodes the N sets of encoding data based on the N sets of predicted motion information to obtain N sets of reconstruction data of the N historical frames.

[0403] In the method flow shown in FIG. 3E, the implementation principle of the encoding end obtaining the reconstruction data of each historical frame in the N historical frames is the same as the implementation principle of the decoding end obtaining the reconstruction data of the current frame through FIG. 4A. In the embodiment of the present application, the process of the encoding end obtaining the reconstruction data of at least three historical frames of the current frame before encoding is the same as the process of the decoding end obtaining the reconstruction data of a frame of the three-dimensional mesh sequence that has been encoded, which can ensure the correct decoding of the decoding end to the current frame and improve the decoding accuracy.

[0404] FIG. 2D is a method flow chart for obtaining the reconstruction data of N historical frames of a current frame (a three-dimensional mesh to be encoded) according to an embodiment of the present application. The method flow can be implemented based on the architecture shown in FIG. 1A, and can be implemented based on FIG. 1B, FIG. 1C, FIG. 2A, FIG. 2B, FIG. 2C and any possible implementation related thereto. As shown in FIG. 2D, the method flow includes the following steps:

[0405] S3201a: The encoding end obtains N sets of encoding data of N historical frames of the three-dimensional mesh sequence that have been encoded.

[0406] The execution of S3201a is the same as the execution principle of S3201a shown in FIG. 2C, which will not be repeated here.

[0407] S3201c: The encoding end decodes the N sets of encoding data to obtain N residuals.

[0408] The N residuals are the residuals between N sets of actual motion information (for example, actual motion trend information) of the N historical frames and the N sets of predicted motion information (for example, predicted motion trend information) described above.

[0409] The N sets of actual motion information include N actual motion trend information of the motion vertex in N time slots corresponding to the N sets of second three-dimensional meshes.

[0410] S3201c is executed after S3201a.

[0411] S3201b: The encoding end obtains N sets of predicted motion information of the N historical frames.

[0412] The N sets of predicted motion information include N predicted motion trend information of the motion vertex in N time slots corresponding to the N sets of historical frames.

[0413] The present application does not limit the execution order between S3201a and S3201b.

[0414] The present application does not limit the execution order between S3201c and S3201b.

[0415] S3203: The encoding end determines N sets of reconstructed motion information of the N historical frames based on the N sets of predicted motion information and the N residuals.

[0416] The N sets of reconstructed motion information include N sets of reconstructed information of actual N motion trend information of the motion vertex in the N time slots of the N historical frames.

[0417] N is an integer greater than or equal to 3.

[0418] S3204: The encoding end determines N sets of reconstructed data of the N historical frames based on the N sets of reconstructed motion information and the reconstructed data of S historical frames corresponding to each of the N historical frames.

[0419] The encoding end can determine a set of reconstructed data of one of the N historical frames by using a set of reconstructed motion information of the historical frame and S sets of reconstructed data of S historical frames corresponding to the historical frame. The determination of the reconstructed data of the other remaining historical frames in the N historical frames is similar. The S historical frames are S decoded historical frames in the three-dimensional mesh sequence, and thus the encoding end can obtain the reconstructed data of the S historical frames, which can include the reconstructed positions of the motion vertices.

[0420] S is an integer greater than or equal to 2, and the S historical frames corresponding to one of the N historical frames are at least two three-dimensional meshes in the three-dimensional mesh sequence that have been encoded before the corresponding historical frame.

[0421] For example, the current frame is the 7th frame, and the N historical frames are the 4th frame to the 6th frame. When determining the reconstructed data of the 4th frame, the encoding end can determine the reconstructed data of the 4th frame by using the reconstructed motion information of the 4th frame and at least two sets of reconstructed data of at least two frames (for example, the 2nd frame and the 3rd frame) that have been decoded before the 4th frame.

[0422] Similarly, when determining the reconstructed data of the 5th frame, the encoding end can determine the reconstructed data of the 5th frame by using the reconstructed motion information of the 5th frame and at least two sets of reconstructed data of at least two frames (for example, the 2nd frame and the 4th frame, which are not limited here) that have been decoded before the 5th frame.

[0423] Similarly, when determining the reconstructed data of the 6th frame, the encoding end can determine the reconstructed data of the 6th frame by using the reconstructed motion information of the 6th frame and at least two sets of reconstructed data of at least two frames (for example, the 3rd frame and the 4th frame, which are not limited here) that have been decoded before the 6th frame.

[0424] It should be understood that the S historical frames corresponding to each of the N historical frames are completely independent, and the S historical frames corresponding to different historical frames in the N historical frames are not associated.

[0425] The implementation principle of the encoding end obtaining the reconstruction data of each of the N historical frames is the same as the implementation principle of the decoding end obtaining the reconstruction data of the current frame (the three-dimensional mesh to be decoded, for example, the i-th frame) in the embodiments of FIG. 4C and FIG. 5A to FIG. 5E. For details, refer to the decoding process of the decoding end on the three-dimensional mesh to be decoded, which will not be described here.

[0426] In the method flow corresponding to FIG. 4C, the encoding end can obtain N sets of actual motion information (expressed as reconstructed motion information) of the N historical frames by using N residuals of the N historical frames obtained by decoding the encoding data of the N historical frames and N sets of predicted motion information of the N historical frames. Then, the encoding end can determine the reconstruction data of one of the N historical frames based on the reconstruction data of at least two frames that have been encoded before the historical frame and the reconstructed actual motion information of the historical frame. The reconstruction data can include the reconstructed positions of the motion vertices of the current frame, and optionally, the connection relationship between the motion vertices (for example, the topology of the current frame), so as to realize the reconstruction of each of the N historical frames. In the case that the three-dimensional mesh sequence can be deformed, the offset accelerations of the same motion vertices (referring to the same vertex identifier or the same vertex ordering position) of different frames are close, and both the predicted motion information and the actual motion information can include motion trend information (for example, offset acceleration). Therefore, in the case that the three-dimensional mesh sequence is deformed, the correct decoding of the three-dimensional mesh can be realized with a smaller code rate, the decoding data amount can be reduced, and the transmission efficiency and the storage efficiency can be improved.

[0427] The following exemplary introduces the structures of several code streams of the present application:

[0428] The following code stream structure is the data structure of the encoding data of the frames after the third frame in the three-dimensional mesh sequence. The data structure of the first frame in the code stream of the three-dimensional mesh sequence can include the topology of any frame in the three-dimensional mesh sequence, and the topologies of the frames are the same.

[0429] Embodiment one

[0430] In the data structure of the code stream, the three-dimensional mesh identifier can be marked first to indicate that the data structure is the data structure of the frame. If all the vertices in the frame are processed according to the coding and decoding method of the present application, the data structure of a frame can not carry the identifier of the vertex, and the syntax structures of all the vertices can be arranged in sequence in the data structure of the frame.

[0431] The actual motion information of the motion vertex in each of the above embodiments can include actual component information of the motion vertex on three components, respectively actual component information on a first component, actual component information on a second component, and actual component information on a third component.

[0432] The predicted motion information of the motion vertex can include predicted component information of the motion vertex on three components, respectively predicted component information on a first component, predicted component information on a second component, and predicted component information on a third component.

[0433] The three components are three components in a three-dimensional coordinate system, which is a Cartesian coordinate system or a cylindrical coordinate system.

[0434] When the three-dimensional coordinate system is a Cartesian coordinate system, the three components can be an x component, a y component, and a z component, respectively.

[0435] When the three-dimensional coordinate system is a cylindrical coordinate system, the three components can be a radius component, a height component, and an angle (for example, a deflection angle) component, respectively.

[0436] The first residual is a residual between actual component information (for example, actual motion acceleration) of the motion vertex on a first component and predicted component information (for example, predicted motion acceleration) on the first component;

[0437] The second residual is a residual between actual component information of the motion vertex on a second component and predicted component information on the second component. For example, the actual component information is actual motion velocity, and the predicted component information is predicted motion velocity; or the actual component information is actual motion acceleration, and the predicted component information is predicted motion acceleration.

[0438] The third residual is a residual between actual component information of the motion vertex on a third component and predicted component information on the third component. For example, the actual component information is actual motion velocity, and the predicted component information is predicted motion velocity; or the actual component information is actual motion acceleration, and the predicted component information is predicted motion acceleration.

[0439] At least one of the first residual, the second residual, and the third residual is a residual between predicted motion acceleration and actual motion acceleration on the corresponding component.

[0440] The code stream data structure is a data structure other than the first three frames of the three-dimensional mesh sequence. The first three frames are not transmitted by encoding compression, but are encoded static models and vertex connection relationships. Among them, the vertex connection relationship of each frame in the three-dimensional mesh sequence remains unchanged and can be reused, combined with the vertex position obtained by decoding reconstruction, to recover the three-dimensional mesh.

[0441] Embodiment Two

[0442] Embodiment three

[0443] Embodiment four

[0444] Embodiment five

[0445] For example, when the vertex positions of some vertices in the three-dimensional mesh are encoded, the first data structure of a frame of the three-dimensional mesh can include not only the three-dimensional mesh identifier but also the vertex identifiers of the moving vertices and the information of the moving vertices, and the information of the moving vertices can include three residuals of the moving vertices on the first component, the second component and the third component, which is the same as Embodiment one, and will not be described herein again.

[0446] It should be noted that the above example shows the structures of five code streams, but this does not limit the code stream structures, and the present application does not make a specific limitation on this.

[0447] The decoding process of the code stream corresponding to the three-dimensional mesh sequence at the decoding end of the present application will be introduced below. The code stream can be the code stream of the three-dimensional mesh sequence obtained by encoding in any one of the above embodiments. The data structure of the code stream can refer to the above description.

[0448] FIG. 3F is a flowchart of a method for decoding a current frame (a three-dimensional mesh to be decoded) according to an embodiment of the present application. The method flowchart can be implemented based on the architecture shown in FIG. 1A, and the code stream can be the code stream obtained by the encoding method at the encoding end according to any one of the above embodiments. As shown in FIG. 3F, the method flowchart includes the following steps:

[0449] S440: The decoding end obtains the code stream corresponding to the three-dimensional mesh sequence.

[0450] It should be understood that the vertices moving at a uniform speed in the three-dimensional mesh sequence can also be decoded by using the decoding method of each embodiment of the present application.

[0451] In a possible implementation, each three-dimensional mesh in the three-dimensional mesh sequence has the same topological structure. For example, the number of vertices, the identifiers of the vertices and the connection relationship between the vertices of each frame in the three-dimensional mesh sequence are the same. In this way, it is convenient to track and determine the motion information of a moving vertex moving at a non-uniform speed in the three-dimensional mesh sequence.

[0452] S441: The decoding end obtains the encoding data of the current frame from the code stream.

[0453] The encoding data is not the original bit sequence in the code stream, but the data obtained after decoding the original bit sequence in the code stream.

[0454] The encoding data can also be referred to as decoded data at the decoding end.

[0455] For example, the decoding end can decode (for example, entropy decode) the code stream to read the encoding data of the current frame from the code stream, or read the encoding data of the current frame from other storage media or devices, which is not limited here.

[0456] S442: The decoding end obtains actual motion trend information of the current frame based on the encoding data.

[0457] For example, the decoding end can obtain the actual motion trend information from the encoding data, or the decoding end obtains the actual motion trend information of the current frame by using the encoding data in combination with other data, which is not limited here.

[0458] The actual motion trend information indicates an actual motion trend of a motion vertex in the three-dimensional mesh sequence that moves at a non-uniform speed in a time slot corresponding to the current frame.

[0459] The actual motion trend can be motion acceleration. Since the current frame is a three-dimensional structure, the actual motion trend can be motion acceleration in at least one of three components in a three-dimensional coordinate system.

[0460] In some embodiments, when the actual motion trend is motion acceleration in one component (for example, the X component), the encoding end can also obtain actual motion rate information of the current frame based on the encoding data, the actual motion rate information being actual motion rates of the motion vertex in the other two components (for example, the Y component and the X component) in the time slot corresponding to the current frame.

[0461] In some embodiments, when the actual motion trend is motion acceleration in two components (for example, the X component and the Y component), the encoding end can also obtain actual motion rate information of the current frame based on the encoding data, the actual motion rate information being an actual motion rate of the motion vertex in the other component (for example, the Z component) in the time slot corresponding to the current frame.

[0462] Wherein, the three-dimensional mesh sequence has a frame rate, and the reciprocal of the frame rate is the overall motion time of the motion vertex in the three-dimensional mesh sequence.

[0463] And any frame in the three-dimensional mesh sequence can have its own instantaneous frame rate. For example, the average of the instantaneous frame rates of all frames in the three-dimensional mesh sequence is the frame rate of the three-dimensional mesh sequence. The instantaneous frame rates of different frames in the three-dimensional mesh sequence can be the same or different, which is not limited here. Based on this, the time slot corresponding to a frame in the three-dimensional mesh sequence can be the reciprocal of the instantaneous frame rate of the frame, and the sum of the time slots corresponding to each frame in the three-dimensional mesh sequence can be the overall motion time.

[0464] S443: The decoding end decodes the actual motion trend information to obtain the reconstructed data of the current frame.

[0465] The reconstructed data of the current frame can include the reconstructed positions of the motion vertices in the current frame, and optionally can also include the connection relationship between the motion vertices. Considering that the topological structure of each frame in the three-dimensional mesh sequence is the same, the connection relationship can not be decoded, because the decoding end can obtain the topological structure when decoding the first frame or the frame before the current frame of the three-dimensional mesh sequence.

[0466] In the embodiments of the present application, when the decoding end decodes the current frame in the three-dimensional mesh sequence, the decoding end can obtain the encoding data of the current frame from the code stream, and obtain the actual motion trend information of the current frame based on the encoding data, the actual motion trend information indicating the actual motion trend of the motion vertex in the three-dimensional mesh sequence that moves at a non-uniform speed in the time slot corresponding to the current frame. In the scenario where the motion vertex in the three-dimensional mesh sequence moves at a non-uniform speed (for example, variable speed motion or self-rotation motion), although the vertex position of the motion vertex in the three-dimensional mesh sequence can change, causing the three-dimensional mesh to deform, the actual motion trend of the motion vertex in the time slot corresponding to each frame in the three-dimensional mesh sequence is relatively small. By decoding the current frame in the three-dimensional mesh sequence based on the actual motion trend information of the current frame, better decoding effect can be obtained.

[0467] FIG. 3G is a flowchart of a method for encoding a current frame (a three-dimensional mesh to be encoded) provided by the embodiments of the present application, the method flow can be implemented based on the architecture shown in FIG. 1A, but is not limited to the architecture shown in FIG. 1A, and the method can be combined with the flow shown in FIG. 3F, but is not limited to the flow shown in FIG. 3F, as shown in FIG. 3G, the method flow can include the following steps:

[0468] S440: The decoding end obtains a code stream corresponding to the three-dimensional mesh sequence.

[0469] The implementation principle of this step is the same as that of S440 shown in FIG. 3F, which will not be described here.

[0470] S441: The decoding end obtains the encoding data of the current frame from the code stream.

[0471] The implementation principle of this step is the same as that of S441 shown in FIG. 3F, which will not be described here.

[0472] S442a: The decoding end decodes the encoding data to obtain the actual motion trend information of the current frame.

[0473] For example, the encoding data includes the actual motion trend information.

[0474] The decoding end can perform at least one decoding method such as decompression on the encoded data to obtain actual motion trend information of the current frame.

[0475] S443: The decoding end decodes the actual motion trend information to obtain reconstructed data of the current frame.

[0476] The implementation principle of this step is the same as S443 shown in FIG. 3F, and will not be described here again.

[0477] In the embodiments of the present application, when the decoding end decodes the current frame in the three-dimensional mesh sequence, the actual motion trend information of the current frame can be directly decoded from the encoded data of the current frame in the code stream. In the case where the motion vertices move at a non-uniform speed in the three-dimensional mesh sequence, a large number of motion vertices in the current frame can have similar actual motion trend (for example, acceleration) information, and the encoded data of the current frame includes the actual motion trend information of the motion vertices, so that the actual motion trend information of a large number of motion vertices in the encoded data is similar or identical. In this way, the number of binary bits occupied by the encoded data encoded into the code stream is smaller, the length of the code stream is shorter, and the code rate is lower. Then, the decompression rate can be improved by decoding the encoded data.

[0478] FIG. 4A is a flowchart of a method for decoding a current frame (a three-dimensional mesh to be decoded) provided in the embodiments of the present application. The method can be implemented based on the architecture shown in FIG. 1A, but is not limited to the architecture shown in FIG. 1A. The method can be combined with the flows shown in FIG. 3F and FIG. 3G, but is not limited to the combination with the flows shown in FIG. 3F and FIG. 3G. The code stream obtained by the decoding end can be the code stream obtained by the encoding end in any one of the above embodiments.

[0479] As shown in FIG. 4A, the method flow includes the following steps:

[0480] S400: The decoding end obtains a code stream corresponding to a three-dimensional mesh sequence.

[0481] For details, refer to the description of S440 of FIG. 3F, which will not be described here again.

[0482] S401a: The decoding end obtains predicted motion trend information of the current frame.

[0483] The predicted motion trend information indicates a predicted motion trend (for example, a predicted motion acceleration) of the motion vertex in a time slot corresponding to the current frame;

[0484] As shown by the dotted arrow in FIG. 4A, the decoding end can acquire the prediction motion trend information based on the code stream. Alternatively, the decoding end can decode the code stream in advance to obtain the prediction motion trend information of the current frame and cache locally, and then when decoding the current frame, the decoding end can acquire the cached prediction motion trend information locally.

[0485] The present application does not limit the execution order of S401a and S401b, which can be executed in series or in parallel. For example, S401b can be executed prior to S401a, which is not limited herein.

[0486] S401b: The decoding end acquires the encoding data of the current frame from the code stream.

[0487] The encoding data includes data encoded based on the actual motion trend information and the prediction motion trend information, the actual motion trend information being the actual motion trend of the motion vertex in the time slot corresponding to the current frame.

[0488] The implementation principle of S401b is the same as that of S441 shown in FIG. 3F, which is not repeated here.

[0489] S403: The decoding end obtains the actual motion trend information of the current frame based on the prediction motion trend information and the encoding data.

[0490] For example, the encoding data includes data encoded based on the actual motion trend information and the prediction motion trend information, and then the decoding end can obtain the actual motion trend information of the current frame based on the prediction motion trend information and the encoding data.

[0491] S443: The decoding end obtains the reconstruction data of the current frame based on the actual motion trend information.

[0492] The implementation principle of S443 is the same as that of S443 shown in FIG. 3F, which is not repeated here.

[0493] In a possible implementation, the encoding data includes the residual error between the actual motion trend information and the prediction motion trend information.

[0494] In the embodiment of the present application, the encoding data includes the residual error between the actual motion information and the prediction motion information of the current frame, so that when the motion vertex in the three-dimensional mesh to be decoded moves at a non-uniform speed, although the uniform speed rates of the same motion vertex in the time slots corresponding to different frames are quite different, the actual motion trend information (for example, the offset acceleration) of the same motion vertex in the time slots corresponding to different frames is close, so that the residual error of the decoding end for the current frame is extremely small, thereby reducing the amount of decoding data, reducing the occupation of transmission bandwidth, and improving the transmission efficiency.

[0495] In the method flow corresponding to FIG. 4A, when decoding the current frame, the decoding end can obtain the prediction motion trend information of the current frame, which is the same as the prediction motion trend information of the current frame obtained by the encoding end when encoding the current frame. In this way, the decoding end can use the prediction motion information to decode the encoding data of the current frame, to ensure the accuracy of the reconstructed data of the decoded current frame, for example, the closeness between the reconstructed position of the motion vertex in the current frame and the original position. In addition, the decoding end can also obtain the encoding data of the current frame from the code stream, which is obtained based on the actual motion information and the prediction motion information. When the motion vertex in the three-dimensional mesh sequence moves at a non-uniform speed, although the difference between the uniform speed rates of the same motion vertex in the time slots corresponding to different frames is large, the actual motion trend information (such as the offset acceleration) of the same motion vertex in the time slots corresponding to different frames is close, so that the difference between the predicted motion information (including the predicted motion trend information) predicted by the present application and the actual motion information (including the actual motion trend information) of the motion vertex is extremely small. Then the length of the code stream of the three-dimensional mesh sequence is shortened, the code rate is lower, the transmission bandwidth can be occupied, and the transmission efficiency and storage efficiency of the code stream can be improved (the same storage space can store more encoding data of frames). In addition, when decoding the current frame in the three-dimensional mesh sequence, the encoding end can use the above-mentioned prediction motion information to decode the encoding data of the current frame to obtain the reconstructed data of the current frame, which can reduce the decoding amount of the code stream of the three-dimensional mesh sequence and reduce the situation of picture freezing corresponding to the three-dimensional mesh sequence.

[0496] FIG. 4B is a method flowchart for decoding a current frame (a three-dimensional mesh to be decoded) provided by an embodiment of the present application. The method flowchart can be implemented based on the architecture shown in FIG. 1A, but is not limited to the architecture shown in FIG. 1A. In addition, the method flowchart shown in FIG. 4B can be combined with FIG. 3F, FIG. 3G, FIG. 4A and any possible implementation thereof. As shown in FIG. 4B, the method flowchart can include the following steps:

[0497] S400: The decoding end obtains a code stream corresponding to a three-dimensional mesh sequence.

[0498] The execution principle of S400 is the same as that of S400 in FIG. 4A, which will not be described here.

[0499] S4101a: The decoding end obtains the reconstructed data of N decoded historical frames in the three-dimensional mesh sequence based on the code stream.

[0500] Each of the N history frames has a set of reconstruction data, and the reconstruction data of the N history frames in S4101a is N sets of reconstruction data corresponding to the N history frames respectively, where N is an integer greater than or equal to a preset threshold (for example, 3).

[0501] For example, the decoding end can obtain the reconstruction data of the N history frames by decoding the N sets of encoded data of the N history frames in the code stream.

[0502] The reconstruction data of a history frame can include the reconstruction position of the motion vertex, optionally include the connection relationship between the motion vertices, and optionally further include the instantaneous frame rate corresponding to the history frame.

[0503] Of course, the instantaneous frame rate corresponding to each frame in the three-dimensional mesh sequence can also not be encoded into the code stream, but be pre-agreed by the encoding end and the decoding end, so that the decoding end can determine the instantaneous frame rate of the current frame according to the frame identifier of the current frame.

[0504] In some embodiments, when the decoding end obtains the N sets of reconstruction data of the N history frames, the decoding end can also read the N sets of reconstruction data of the N history frames, where the decoding end can decode the N sets of encoded data of the N history frames in the code stream in advance before decoding the current frame, and store the N sets of reconstruction data of the N history frames obtained by decoding.

[0505] The N sets of reconstruction data corresponding to each of the N history frames can include the reconstruction position of the motion vertex in the history frame (i.e., the reconstruction data of the vertex position), and optionally further include the topology structure corresponding to the history frame. Since the topology structure of each frame in the three-dimensional mesh sequence is the same, the topology structure does not need to be reconstructed. Optionally, the set of reconstruction data of the history frame can further include the frame rate of the three-dimensional mesh sequence, and optionally further include the instantaneous frame rate of each three-dimensional mesh in the three-dimensional mesh sequence.

[0506] S4102a: The decoding end determines the prediction motion information of the current frame based on the N sets of reconstruction data.

[0507] The execution principle of S4102a is the same as that of S301a shown in FIG. 2B of the encoding end, and specific implementation processes can be referred to the specific implementation process of S301a, which will not be described here.

[0508] The prediction motion information is the predicted motion information of the motion vertex in the three-dimensional mesh sequence that moves at a non-uniform speed in the time slot corresponding to the current frame, and the prediction motion information can include the prediction motion trend information and optionally include the prediction motion rate information.

[0509] The prediction motion rate information is the predicted motion rate of the motion vertex in the time slot corresponding to the current frame.

[0510] S401b: The decoding end obtains the encoding data of the current frame from the code stream.

[0511] The execution principle of S401b is the same as that of S401b in the embodiment of FIG. 4A, and thus is not described here again.

[0512] S403: The decoding end determines the reconstruction data of the current frame based on the prediction motion information and the encoding data of the current frame.

[0513] The execution principle of S403 is the same as that of S403 in the embodiment of FIG. 4A, and thus is not described here again.

[0514] In addition, S4101a is executed before S4102a, and the present application does not limit the execution order between S401b and S4101a. After S4102a and S401b, S403 can be executed.

[0515] In the method flow corresponding to FIG. 4B, in order to ensure that the prediction motion information calculated by the encoding end and the decoding end for the same frame is the same, when the decoding end obtains the prediction motion information of the current frame, the decoding end can obtain the reconstruction data of at least three historical frames of the three-dimensional mesh sequence that have been encoded, and determine the prediction motion information of the current frame based on the reconstruction data (rather than the original data of the at least three historical frames), so as to ensure that the prediction motion information determined by the decoding end for the current frame is consistent with the prediction motion information obtained by the encoding end for the current frame, thereby ensuring accurate decoding of the three-dimensional mesh sequence by the decoding end, ensuring that the vertex position difference between the reconstructed three-dimensional mesh and the original three-dimensional mesh is small, and improving the quality of the reconstructed three-dimensional mesh.

[0516] FIG. 4C is a method flow diagram provided by an embodiment of the present application for decoding a current frame (a three-dimensional mesh to be decoded). The method flow can be implemented based on the architecture shown in FIG. 1A, but is not limited to the architecture shown in FIG. 1A. In addition, the method flow shown in FIG. 4C can be combined with FIG. 3F, FIG. 3G, FIG. 4A, FIG. 4B, and any possible implementation thereof, and can also be combined with FIG. 3F, FIG. 3G, FIG. 4A, FIG. 4B, and any possible implementation thereof. As shown in FIG. 4C, the method flow can include the following steps:

[0517] S4101a: The decoding end obtains N sets of reconstruction data of N historical frames of the three-dimensional mesh sequence that have been decoded based on the code stream.

[0518] The execution principle of S4101a in FIG. 4C is the same as that of S4101a shown in FIG. 4B, and thus is not described here again.

[0519] S4102a: The decoding end determines the prediction motion information of the current frame based on the N sets of reconstruction data.

[0520] The execution principle of S4102a is the same as that of S4102a shown in FIG. 4B, and details are not repeated here.

[0521] S4101b: Decoding the encoding data of the current frame to obtain a residual.

[0522] The encoding data of the current frame can be obtained by S401b shown in FIG. 4B, and details are not repeated here.

[0523] The residual is the residual between the actual motion information of the current frame and the predicted motion information.

[0524] S4103: The decoding end determines the reconstructed motion information of the current frame based on the predicted motion information and the residual.

[0525] The reconstructed motion information is the reconstructed information of the actual motion information (such as the actual motion trend information) of the current frame, and the reconstructed motion information is also referred to as the reconstructed actual motion information.

[0526] For example, the decoding end can add the predicted motion information and the residual to obtain the reconstructed motion information of the current frame.

[0527] S4104: The decoding end determines the reconstructed data of the current frame based on the reconstructed motion information and the reconstructed data of the M historical frames in the three-dimensional mesh sequence.

[0528] The decoding end can obtain the reconstructed data of the M historical frames by decoding the encoding data of the M historical frames in the code stream. Alternatively, the decoding end can read the decoded data of the M historical frames from other devices, and the other devices can decode the encoding data of the M historical frames in the code stream to obtain the decoded data of the M historical frames.

[0529] M is an integer greater than or equal to 2.

[0530] In the method flow corresponding to FIG. 4C, the decoder can use the residual between the actual motion information of the current frame obtained by decoding the code stream and the predicted motion information, and the determined predicted motion information of the current frame, to reconstruct the actual motion information (expressed as reconstructed motion information) of the current frame; then, the decoder can determine the reconstructed data of the current frame based on the reconstructed data of at least two frames decoded before the current frame and the actual motion information of the current frame. The reconstructed data can include the reconstructed positions of the motion vertices of the current frame, and optionally, the connection relationship (for example, the topology of the current frame) between the motion vertices, so as to realize the reconstruction of the current frame. In the case that the three-dimensional mesh sequence can be deformed, the offset accelerations of the same motion vertices (referring to the same vertex identifier or the same vertex ordering position) of different frames are close, and both the predicted motion information and the actual motion information can include motion trend information (for example, offset acceleration). Therefore, in the deformation scene of the three-dimensional mesh sequence, correct decoding of the three-dimensional mesh can be realized with a smaller code rate, and the data amount of decompression can be reduced due to the small residual.

[0531] FIGS. 5A-5E are respectively various method flow diagrams for decoding a three-dimensional mesh provided by embodiments of the present application; the method flows can be realized based on the architecture shown in FIG. 1A, the flows shown in FIGS. 3F, 4A-4C, and various possible implementations.

[0532] The method flow shown in FIG. 5A can be a specific example of the method flow shown in FIG. 4C, to realize the decoding of the current frame.

[0533] In the method flow shown in FIG. 5A, the decoder can use the reconstructed positions of the motion vertices in at least three historical frames adjacent to the current frame and the instantaneous frame rates of at least two of the at least three historical frames, to determine the predicted motion information of the motion vertices in the current frame; and based on the residual between the actual motion information of the current frame obtained by decoding the code stream and the predicted motion information, and the predicted motion information, to reconstruct the actual motion information of the motion vertices in the current frame; finally, based on the reconstructed positions of the motion vertices of the previous two frames (two decoded and reconstructed frames before and adjacent to the current frame) of the current frame and the actual motion information of the current frame, to reconstruct the reconstructed positions of the motion vertices of the current frame, to realize the decoding of the current frame. The method can mainly include the following steps:

[0534] S4001a: the decoder determines the actual motion information (substantially the reconstructed actual motion information) of the motion vertex of the i-1th frame based on the reconstructed positions of the motion vertices of the i-3th, i-2th and i-1th frames and the instantaneous frame rates of the i-2th and i-1th frames, and uses the actual motion information of the i-1th frame as the predicted motion information of the motion vertex of the i th frame.

[0535] The reconstructed position of the motion vertex of each frame can be obtained from the reconstructed data of the N historical frames obtained in S4101a in the embodiment of FIG. 4B or FIG. 4C.

[0536] The instantaneous frame rate of each frame can also be obtained from the reconstructed data of the N historical frames, or the instantaneous frame rate of each frame in the three-dimensional mesh sequence can be agreed in advance by the encoding end and the decoding end, so that the decoding end can determine the instantaneous frame rate of the corresponding frame according to the frame identifier of each frame.

[0537] The execution principle of S4001a is the same as that of S3001a shown in FIG. 3A, and the specific implementation process and other possible implementation manners can be referred to S3001a and the possible other implementation manners extended from S3001a.

[0538] The three historical frames used in S4001a can be examples of the N historical frames mentioned in the embodiments of the encoding end and the decoding end.

[0539] S4001b: The decoding end decodes the received code stream to obtain the residual between the actual motion information and the predicted motion information of the motion vertex of the i-th frame.

[0540] For example, the decoding end can entropy-decode the code stream to obtain the residual between the actual motion information and the predicted motion information of the vertex of the i-th frame.

[0541] S4001a and S4001b can be executed in series or in parallel, and the present application does not limit this.

[0542] S4002: The decoding end reconstructs the actual motion information of the motion vertex of the i-th frame based on the residual and the predicted motion information of the motion vertex of the i-th frame obtained in S4001a, to obtain the reconstructed actual motion information (also referred to as reconstructed motion information) of the motion vertex of the i-th frame.

[0543] For example, the residual is the difference between the actual motion information and the predicted motion information of the motion vertex of the i-th frame, and then in this step, the decoding end can add the residual to the predicted motion information of the motion vertex of the i-th frame obtained in S4001a, to reconstruct the actual motion information of the motion vertex of the i-th frame.

[0544] The actual motion information of the motion vertex of the i-th frame corresponds to the predicted motion information of the motion vertex of the i-th frame.

[0545] The actual motion information of the motion vertex of the i-th frame may include respective offset motion information of the motion vertex in three offset directions of a three-dimensional coordinate system, wherein the offset motion information of the motion vertex in at least one of the three offset directions is offset acceleration, and the offset motion information of the motion vertex in the remaining offset directions may be offset velocity or offset acceleration, which is not limited herein. However, the at least one offset direction is the same as the offset direction corresponding to the offset acceleration in the predicted motion information.

[0546] Regardless of the actual motion information or the predicted motion information of the current frame, the offset acceleration of the motion vertex in at least one offset direction is included, and the residual error includes the residual error between the actual offset motion information and the predicted offset motion information of the same offset direction.

[0547] S4003: The decoding end determines the reconstructed position of the motion vertex of the i-th frame based on the reconstructed position of the motion vertex of the i-1-th frame, the reconstructed position of the motion vertex of the i-2-th frame, the instantaneous frame rate of each frame (for example, the instantaneous frame rate of the i-1-th frame, the instantaneous frame rate of the i-th frame, and optionally the instantaneous frame rate of the i-2-th frame), and the reconstructed information (also referred to as reconstructed motion information) of the actual motion information of the vertex of the i-th frame.

[0548] In this step, the reconstructed position of the motion vertex of the i-1-th frame and the reconstructed position of the motion vertex of the i-2-th frame may be obtained in the same manner as the reconstructed position of the motion vertex of the i-1-th frame and the reconstructed position of the motion vertex of the i-2-th frame used in S4001a, which is not repeated herein.

[0549] In S4003, the decoding end may reconstruct the vertex position of the motion vertex of the current frame (here, the i-th frame) based on the reconstructed positions of the motion vertices of at least two historical frames decoded before the current frame, the instantaneous frame rate, and the reconstructed information of the actual motion information of the motion vertex of the current frame reconstructed in S4002, to obtain the reconstructed positions of the motion vertices in the current frame, thereby completing the reconstruction of the current frame.

[0550] In S4003, the at least two historical frames used may be examples of the M historical frames mentioned in the embodiments of the encoding end and the decoding end.

[0551] It should be understood that the frame identifiers of the N historical frames used by the encoding end and the N historical frames used by the decoding end are the same.

[0552] Similarly, it should be understood that the frame identifiers of the M historical frames used by the encoding end and the M historical frames used by the decoding end are the same.

[0553] In other embodiments, when reconstructing the reconstructed position of the motion vertex of a frame (e.g. the i-1th frame or the ith frame, without limitation), the M history frames referred to are not limited to two frames that have been coded before the frame, but can also be a larger number of frames (e.g. 3 frames or 4 frames), i.e. M≥2, M being an integer, and the M history frames are not limited to frames adjacent to the frame in display order, but can also be non-adjacent frames.

[0554] Then when reconstructing the reconstructed position of the motion vertex of a frame, when the M history frames referred to are three or more frames that have been decoded before the frame, the implementation principle when the M history frames referred to are two frames can be combined, and an algorithm (without limitation) such as weighted average or averaging is used to operate the multiple reconstructed positions of the motion vertex obtained to obtain the reconstructed position of the motion vertex of the frame.

[0555] For example, the M history frames (here the i-1th frame and the i-2th frame) referred to in S4003 can also be the i-3th frame, the i-2th frame and the i-1th frame. Then when determining the reconstructed position of the motion vertex of the ith frame based on the reconstructed position of the motion vertex of each of the i-3th frame, the i-2th frame and the i-1th frame, and the reconstructed information of the actual motion information of the motion vertex of the ith frame, the candidate reconstructed positions of the motion vertex of the ith frame can be calculated based on the reconstructed position of the motion vertex of each of the i-3th frame, the i-2th frame and the i-1th frame, and the reconstructed information of the actual motion information of the motion vertex of the ith frame; and the candidate reconstructed positions of the motion vertex of the ith frame can be calculated based on the reconstructed position of the motion vertex of each of the i-2th frame, the i-1th frame and the ith frame, and the reconstructed information of the actual motion information of the motion vertex of the ith frame; then the two candidate reconstructed positions of the motion vertex of the ith frame are averaged (or weighted summation, etc.) to obtain the reconstructed position of the motion vertex of the ith frame in S4003.

[0556] Optionally, after S4003, the method can further comprise S4004.

[0557] S4004: The decoding end performs smoothing processing on the reconstructed position of the motion vertex of the ith frame to obtain the reconstructed ith frame.

[0558] In consideration of the vertex positions of the motion vertices in the current frame obtained by the decoder through S4003, the three-dimensional mesh formed by connecting the motion vertices according to the topology of the three-dimensional mesh sequence (for example, the reconstructed current frame) can have a wrinkled mesh with an insufficiently smooth surface, and then the decoder can perform smoothing processing on the reconstructed positions of the motion vertices of the i-th frame obtained by S4003 to correct the reconstructed positions of some motion vertices, so that the surface of the three-dimensional mesh formed by the vertex positions of the motion vertices of the i-th frame after the smoothing processing (the i-th frame) is smoother, and the reconstructed i-th frame is obtained.

[0559] In the corresponding embodiment of FIG. 5A, the decoder can use the reconstructed positions of the motion vertices of the previous three decoded frames before the current frame and the corresponding instantaneous frame rates to calculate the reconstructed information of the actual motion information of the motion vertices of the previous frame of the current frame as the predicted motion information of the motion vertices of the current frame. The reconstructed information of the actual motion information or the predicted motion information can include the acceleration of the motion vertices of the current frame in at least one offset direction. Since the acceleration of the motion vertices of adjacent frames is close when the three-dimensional mesh is moving at a non-uniform speed, using the reconstructed information of the actual motion information of the previous frame as the predicted motion information of the current frame can improve the accuracy of the predicted motion information of the current frame, so that it is closer to the actual motion information of the motion vertices of the current frame. Then, the actual motion information of the current frame is reconstructed using the residual between the actual motion information and the predicted motion information of the current frame obtained by decoding the code stream and the predicted motion information of the current frame obtained by decoding, wherein the predicted motion information is close to the actual motion information, so that the reconstructed actual motion information is also closer to the actual motion information of the current frame. Finally, based on the reconstructed positions of the motion vertices of the two historical frames decoded before the current frame and adjacent to the current frame in display order and the reconstructed information of the actual motion information of the motion vertices of the current frame and the corresponding instantaneous frame rates, the reconstructed positions of the motion vertices of the current frame are reconstructed to ensure correct decoding of the three-dimensional mesh by the decoder in the case of small residual, small bandwidth occupation and low code rate of the code stream.

[0560] In another embodiment, the application also provides a decoding method which can be implemented based on the flow shown in FIG. 3G.

[0561] The flow can include obtaining the reconstructed information of the actual motion information of the motion vertices of the i-th frame by decoding the encoding data of the current frame as shown in FIG. 5A; and S4003 as shown in FIG. 5A to obtain the reconstructed positions of the motion vertices of the i-th frame.

[0562] In the method flow, the prediction motion information of the motion vertex of the i-th frame is not calculated, but the actual motion information of the current frame decoded from the encoded data is directly used to reconstruct the vertex position of the current frame, which has lower encoding complexity.

[0563] FIG. 5B is a flowchart of a method for decoding a current frame, taking a vertex P of the current frame as an example, which can be implemented based on the architecture shown in FIG. 1A, and can be implemented based on the flows shown in FIG. 4C and FIG. 5A, but is not limited to the embodiments in combination with FIG. 1A, FIG. 4C and FIG. 5A.

[0564] As shown in FIG. 5B, taking decoding the vertex P in the current frame (i-th frame) as an example, the process of decoding each vertex of the current frame is introduced.

[0565] It should be understood that even if the vertex in the current frame moves at a constant speed in the sequence of three-dimensional meshes, the encoding and decoding process of the motion vertex of the present application is also applicable to the vertex moving at a constant speed.

[0566] FIG. 5B shows the three-dimensional mesh diagrams of the decoded i-3-th frame, i-2-th frame and i-1-th frame.

[0567] Referring to FIG. 5B, the reconstructed positions of the vertex P in the i-3-th frame, i-2-th frame and i-1-th frame in the three-dimensional mesh are P i-3 , P i-2 and P i-1 respectively, and the actual position (also referred to as the original position), the reconstructed position and the target position after smoothing of the vertex P in the i-th frame are P i , P i ' and P i " respectively. In the case where the vertex P moves at a non-constant speed in the sequence of three-dimensional meshes, the position of the vertex P is shifted from P i-3 to P i-2 and P i-1 in turn. For example, the instantaneous frame rate of the i-th frame in the sequence of frames is fi, and the time interval (i.e. the time slot corresponding to the i-th frame) Δti = 1 / fi between the i-1-th frame and the i-th frame.

[0568] For ease of illustration, the decoding process of the present application is described taking the original position of the vertex P in the i-3-th frame, i-2-th frame and i-1-th frame as the same as the reconstructed position of the vertex P in the i-3-th frame, i-2-th frame and i-1-th frame as an example.

[0569] Specifically:

[0570] Firstly:

[0571] When the decoding end calculates the actual motion information of the vertex P in the i-1th frame based on the respective reconstructed positions of the vertex P in the i-3th, i-2th and i-1th frames, and uses the actual motion information of the vertex P in the i-1th frame as the predicted motion information of the vertex P in the i th frame, the specific implementation principle is completely the same as the process of determining the predicted motion information of the vertex P in the i th frame mentioned in FIG. 3B, and thus will not be described herein.

[0572] In this way, the decoding end can calculate the reconstructed information of the actual motion information of the vertex P in the i-1th frame (for example, the acceleration a i-1 in the at least one displacement direction) based on the respective reconstructed positions of the vertex P in the i-3th, i-2th and i-1th frames and the corresponding frame rates, and use the reconstructed information of the actual motion information of the vertex P in the i-1th frame as the predicted motion information of the vertex P in the i th frame.

[0573] Secondly,

[0574] The decoding end can decode the code stream to obtain the residual of the actual motion information and the predicted motion information of the vertex P in the i th frame, for example, the residual of the acceleration a i -a i-1 .

[0575] Then,

[0576] The decoding end can add the predicted motion information of the vertex P in the i th frame and the decoded residual to obtain the reconstructed information of the actual motion information of the vertex P in the i th frame (for example, the acceleration a i =a i-1 +a i -a i-1 in the at least one displacement direction and the displacement rate in the remaining displacement direction).

[0577] In addition,

[0578] The decoding end can also obtain the reconstructed position of the vertex P in the i-1th frame (for example, the position P i-1 ), and determine the actual motion information of the vertex P in the i-1th frame (for example, the displacement rate v i-1 of the vertex P in the at least one displacement direction, but not the acceleration) based on the reconstructed position of the vertex P in the i-1th frame and the reconstructed position of the vertex P in the i-2th frame and the respective instantaneous frame rates (at least one instantaneous frame rate) of the two frames.

[0579] Then,

[0580] The decoding end can use the reconstructed position of the vertex P in the i-1th frame (for example, the position P i-1), and the actual motion information of the reconstructed vertex P of the (i-1)th frame (e.g., the offset rate v in at least one of the offset directions mentioned above). i-1 And the actual motion information of vertex P in the reconstructed i-th frame (e.g., the offset acceleration a in at least one offset direction mentioned above). i and the offset rate in the remaining offset directions. (e.g., the y-direction) (e.g., the z-direction) to calculate the reconstructed position P of vertex P in the i-th frame. i ′.

[0581] For example, the offset acceleration a in the actual motion information of vertex P in the i-th frame above. i The corresponding offset direction and the motion information reconstructed from vertex P in the (i-1)th frame (e.g., the offset rate v in at least one of the offset directions mentioned above). i-1 The corresponding offset direction is the X direction in the three-dimensional Cartesian coordinate system. For example, the position of vertex P in frame i-1 is P. i-1 The coordinates are (x1, y1, z1), and the reconstructed position P of the vertex P to be solved in the i-th frame is P1. i The coordinates of ′ are (x2, y2, z2).

[0582] For example, when calculating x2, it can be achieved using the following formula 1.1:

[0583] For example, when calculating y2, it can be achieved using the following formula 2.1:

[0584] For example, when calculating z2, it can be achieved using the following formula 3.1:

[0585] In other embodiments, the formula 2.1 above... Can be replaced with Or replace with and Any value between.

[0586] Similarly, in formula 3.1 above... Can be replaced with Or replace with and Any value between.

[0587] in, This represents the offset rate of the motion vertex in the i-1th frame in the Z component within the corresponding time slot of the i-1th frame;

[0588] represents the offset rate of the motion vertex of the i-1th frame in the Y component in the corresponding time slot of the i-1th frame.

[0589] In this way, the decoding end can decode the reconstructed position P i ′(x2, y2, z2) of the vertex P of the i th frame, and similarly, the reconstructed positions of other vertices of the i th frame can be reconstructed. Although the reconstructed position P i ′(x2, y2, z2) of the vertex P of the i th frame, but the error is within an acceptable range, so that the above-mentioned error in the image sequence obtained by rendering the decoded three-dimensional mesh is difficult to be observed by the naked eye. i

[0590] Optionally, the decoding end can perform smoothing processing on the reconstructed position of the motion vertex of the i th frame, for example, as shown in FIG. 5B, the position of the vertex P in the i th frame can be smoothed from the reconstructed position P i ′ to the target position P i ″, so that the image displayed after the smoothed three-dimensional mesh is rendered has higher quality, smaller noise, and higher smoothness.

[0591] FIG. 5C is a method flowchart provided by an embodiment of the present application for decoding a current frame (a three-dimensional mesh to be decoded), which can be implemented based on the architecture shown in FIG. 1A, the flows shown in FIG. 4C, FIG. 5A, and FIG. 5B, but is not limited to the architecture and flow shown in FIG. 1A, FIG. 4C, FIG. 5A, and FIG. 5B.

[0592] The code stream in the flow shown in FIG. 5C can be the code stream obtained by the flow shown in FIG. 3C.

[0593] In the method flow shown in FIG. 5C, the decoding end can convert the reconstructed positions of the motion vertices in the decoded and reconstructed historical frames from the three-dimensional Cartesian coordinate system to the cylindrical coordinate system, to calculate the predicted motion information of the motion vertices of the current frame on the three components of the height, radius, and angle of the motion vertices, and then complete the reconstruction of the actual motion information of the vertices of the current frame, so as to reconstruct the current frame by using the actual motion information of the historical frames and the current frame.

[0594] In the method flow shown in FIG. 5C, whether the actual motion information or the predicted motion information, the motion information can include the offset motion information of the vertices of the current frame in the three offset directions of the cylindrical coordinate system, and the offset motion information includes the offset rate of the vertices in the height direction (also referred to as the height component), the offset rate of the vertices in the radius direction (also referred to as the radius component), and the offset acceleration of the vertices in the angle direction (also referred to as the angle component).

[0595] ​The offset rate of the vertex on the height component is also referred to as the height change rate of the vertex, the offset rate of the vertex on the radius component is also referred to as the radius change rate of the vertex, and the offset acceleration of the vertex on the angle component is also referred to as the angular acceleration.

[0596] As shown in FIG. 5C, the method can mainly include the following steps:

[0597] S801a: The decoding end decodes the code stream to obtain a residual between actual motion information and predicted motion information of the i-th frame.

[0598] The residual includes a residual of the offset rate of the motion vertex on the height component, a residual of the offset rate of the motion vertex on the radius component, and a residual of the acceleration of the motion vertex on the angle component between the i-th frame and the (i-1)-th frame.

[0599] Specifically, referring to the data input in S503 in FIG. 3C, the residual includes a residual between the actual offset rate of the motion vertex on the height component of the i-th frame and the reconstructed offset rate of the motion vertex on the height component of the (i-1)-th frame, a residual between the actual offset rate of the motion vertex on the radius component of the i-th frame and the reconstructed offset rate of the motion vertex on the radius component of the (i-1)-th frame, and a residual between the actual offset acceleration of the motion vertex on the angle component of the i-th frame and the reconstructed offset acceleration of the motion vertex on the angle component of the (i-1)-th frame.

[0600] S801b: The decoding end converts the reconstructed position of the motion vertex of each of the (i-3)-th frame, the (i-2)-th frame and the (i-1)-th frame from the three-dimensional Cartesian coordinate system to the cylindrical coordinate system to obtain the reconstructed height, the reconstructed radius and the reconstructed angle of the vertex of each of the (i-3)-th frame, the (i-2)-th frame and the (i-1)-th frame.

[0601] S802b and S802c can be executed after S801b, and the present application does not limit the execution order between S802b and S802c, which can be executed in series or in parallel.

[0602] In addition, the present application also does not limit the execution order between S801a and S801b, which can be executed in series or in parallel.

[0603] The execution principle of S801b is similar to that of S501 in the embodiment of FIG. 3C, and the specific implementation process is described above, which will not be described here.

[0604] S802b: The decoding end calculates the reconstructed offset velocity of the vertex in the height component of the i-1th frame, the reconstructed offset velocity of the vertex in the radius component of the i-1th frame, and the reconstructed offset acceleration of the vertex in the angle component of the i-1th frame based on the reconstructed height, the reconstructed radius, and the reconstructed angle of the motion vertex of the i-3th frame, the i-2th frame, and the i-1th frame, and the instantaneous frame rate of each frame (e.g., the i-2th frame and the i-1th frame), to obtain the reconstructed actual motion information of the motion vertex of the i-1th frame (which can be used as the predicted motion information of the motion vertex of the i-1th frame).

[0605] The execution principle of S802b is exactly the same as that of S502a in the embodiment of FIG. 5C, and thus will not be described here.

[0606] For example, in combination with FIG. 5B, taking the vertex P as an example, the reconstructed offset velocity of the vertex P in the height component of the i-1th frame can be obtained through S802b the reconstructed offset velocity of the vertex P in the radius component of the i-1th frame the reconstructed offset acceleration of the vertex in the angle component of the i-1th frame

[0607] After S801a and S802b, S803a can be executed.

[0608] S803a: The decoding end determines the reconstructed information of the actual motion information of the motion vertex of the i-1th frame based on the residual obtained in S801a and the predicted motion information of the motion vertex of the i-1th frame obtained in S802b.

[0609] Specifically, continuing to take the vertex P as an example, the decoding end can add the residual of the offset velocity in the height component between the same vertex of the i-1th frame and the i-1th frame obtained in S801a (e.g. ) to the reconstructed offset velocity of the vertex P in the height component of the i-1th frame obtained in S802b to obtain the reconstructed offset velocity of the vertex P in the height component of the i-1th frame (specifically, the reconstructed information of the actual offset velocity).

[0610] Similarly, the decoding end can add the residual of the offset velocity in the radius component between the same vertex of the i-1th frame and the i-1th frame obtained in S801a (e.g. ) to the reconstructed offset velocity of the vertex P in the radius component of the i-1th frame obtained in S802b to obtain the reconstructed offset velocity of the vertex P in the radius component of the i-1th frame

[0611] Similarly, the decoding end can obtain the residual of the offset acceleration in the angular component between the same vertex of the i-th frame and the (i-1)-th frame obtained by S801a (e.g., ), and the reconstructed offset acceleration of vertex P in the angular component of the (i-1)th frame obtained by S802b. The values ​​are added together to obtain the reconstructed offset acceleration of vertex P in the angular component of the i-th frame.

[0612] For example, the reconstruction information of the actual motion information of vertex P in the i-th frame may include the aforementioned reconstruction offset rate. Reconstruction offset rate and reconstruction offset acceleration

[0613] It should be understood that in the method flow shown in Figure 5C, three graphic shapes are used to represent the data that needs to be processed. These graphics are not intended to limit the technical solution of this application, but are intended to help readers understand how the decoding end of this application uses the obtained data to decode and obtain the reconstructed position of the motion vertex.

[0614] S802c: The decoding end can calculate the reconstruction offset rate of the motion vertex in the angular component of the i-1 frame based on the reconstruction angle of the motion vertex in each frame of the i-2 and i-1 frames obtained through S801b and the instantaneous frame rate of the i-1 frame.

[0615] For ease of explanation, the method of Figures 5C to 5E of this application is illustrated by taking the example that the instantaneous frame rate of each frame in the three-dimensional mesh sequence is the same. The time slot of each frame is the time interval Δt, where Δt = 1 / f, and f is the frame rate of the three-dimensional mesh sequence.

[0616] Continuing with vertex P in Figure 5B as an example, the decoding end can utilize the position P of vertex P in the (i-2)th frame. i-2 The angle components, and the position P of vertex P in the (i-1)th frame. i-1 The angle component and the time slot Δt of the (i-1)th frame are used to calculate the reconstruction offset rate of vertex P in the (i-1)th frame on the angle component. (Also expressed as the reconstructed angular velocity), for example Wherein, the offset angle Δθ i-1 For position P i-1 Angular components and position P i-2 The difference between the angular components;

[0617] S804 can be executed after S801a, S803a, and optionally after S802c. In some embodiments, the reconstruction of the height and radius of the vertices in S804 can also be performed before S802c. S802c is mainly used to obtain the data needed to reconstruct the angle of the vertices.

[0618] S804: Based on the reconstructed actual motion information of the motion apex of the i-th frame obtained in S803a, the reconstructed position converted reconstructed height, reconstructed radius, and reconstructed angle of the motion apex of the (i-1)-th frame, and the reconstructed displacement rate of the motion apex of the (i-1)-th frame in the angle component, and the instantaneous frame rate of the i-th frame, the decoding end determines the reconstructed height, reconstructed radius, and reconstructed angle of the motion apex of the i-th frame.

[0619] Specifically, the apex P shown in FIG. 5B is taken as an example to illustrate the determination of the reconstructed height, reconstructed radius, and reconstructed angle of the apex P of the i-th frame.

[0620] The actual motion information of the apex P of the i-th frame reconstructed by the decoding end includes the reconstructed displacement rate of the apex P in the height component the reconstructed displacement rate of the apex P in the radius component and the reconstructed displacement acceleration of the apex P in the angle component

[0621] The reconstructed height, reconstructed radius, and reconstructed angle corresponding to the reconstructed position of the apex P in the (i-1)-th frame are h i-1 , r i-1 , and θ i-1 , respectively.

[0622] The reconstructed displacement rate of the apex P in the angle component of the (i-1)-th frame is

[0623] The time slot of the i-th frame is Δt.

[0624] The decoding end can determine the reconstructed height h i , reconstructed radius r i , and reconstructed angle θ i of the apex P of the i-th frame as follows.

[0625] Based on the reconstructed displacement rate of the apex P in the height component within the time slot from the (i-1)-th frame to the i-th frame (for example, the displacement rate ) and the time slot Δt of the i-th frame, the decoding end determines the displacement of the apex P in the height component from the (i-1)-th frame to the i-th frame; and based on the displacement and the height h i-1 corresponding to the position of the apex P in the (i-1)-th frame, the decoding end determines the height h i of the apex P of the i-th frame.

[0626] The displacement rate of the apex P in the height component within the time slot from the (i-1)-th frame to the i-th frame can be the reconstructed displacement rate , the reconstructed displacement rate , or the average of and , etc. and Any value within the range. The decoding end can be based on the position P of vertex P in the (i-2)th frame. i-2 The height component, and the position P of vertex P in the (i-1)th frame. i-1 The height component and the time slot Δt of the (i-1)th frame are used to calculate the reconstruction offset rate of vertex P in the (i-1)th frame on the height component. For example Wherein, the offset height Δh i-1 For position P i-1 Height components and position P i-2 The difference in height components.

[0627] For example, the decoding end can obtain the reconstructed height h of vertex P in the i-th frame using formula 1.2, formula 1.3, or formula 1.4. i .

[0628] Similarly, the decoding end can base its work on the offset rate of vertex P in the radial component within the time slot of the i-th frame (e.g., offset rate). The offset displacement of vertex P from frame (i-1) to frame (i-1) is determined based on the time slot Δt of frame i. Then, based on this offset displacement and the radius component r corresponding to the position of vertex P in frame (i-1), the offset displacement of vertex P in the radius component r is determined. i-1 To determine the radius r of vertex P in the i-th frame. i .

[0629] Wherein, the offset rate of vertex P from frame (i-1) to frame i in the radius component can be the offset rate described above. It could also be the offset rate. It can also be and The average value is equal to and Any value between [values]. The decoding end can base its value on the position P of vertex P in the (i-2)th frame. i-2 The radius component, and the position P of vertex P in the (i-1)th frame. i-1 The radius component and the time interval Δt are used to calculate the reconstruction offset rate of vertex P in the (i-1)th frame on the radius component. For example Wherein, the offset radius Δr i-1 For position P i-1 The radius component and position P i-2 The difference between the radius components.

[0630] For example, the decoding end can obtain the reconstruction radius r of vertex P in the i-th frame using formula 2.2, formula 2.3, or formula 2.4.i .

[0631] The decoding end can determine the offset rate (e.g., offset rate) of vertex P on the intrinsic angular component of the time slot of frame i-1 (the time interval from frame i-2 to frame i-1). The reconstructed offset acceleration of vertex P in the angular component of the i-th frame. And the time slot Δt of the i-th frame, to determine the time slot of vertex P in the i-th frame (from position P in the (i-1)-th frame). i-1 Reconstruction position P offset to the i-th frame i The offset angle is calculated based on the intrinsic angular component; then, based on this offset angle and the angular component θ corresponding to the position of vertex P in the (i-1)th frame,... i-1 To determine the reconstruction angle θ corresponding to the reconstruction position of vertex P in the i-th frame. i .

[0632] Wherein, the offset rate of vertex P from frame i-2 to frame i-1 in terms of the angular component can be the offset rate described above. It could also be the offset rate. It can also be and The average value is equal to and Any value between [values]. The decoding end can base its value on the position P of vertex P in the (i-2)th frame. i-2 The angle components, and the position P of vertex P in the (i-1)th frame. i-1 The angle component and the time slot Δt of the (i-1)th frame are used to calculate the reconstruction offset rate of vertex P in the (i-1)th frame on the angle component. For example Wherein, the offset angle Δθ i-1 For position P i-1 Angular components and position P i-2 The difference between the angular components. Offset rate. The calculation process is as follows:

[0633] For example, the decoding end can reconstruct the angle θ of vertex P in the i-th frame using formula 3.2, 3.3, or 3.4. i .

[0634] S805: The decoding end converts the reconstructed height, reconstructed radius, and reconstructed angle of the motion vertex of the i-th frame into a Cartesian coordinate system to obtain the reconstructed position of the motion vertex of the i-th frame.

[0635] In the method flow shown in FIG. 5C, the decoding end decodes not the residual of the coordinate value between the predicted position and the accurate position of the motion vertex of the current frame, but the residual of the predicted motion information and the actual motion information of the motion vertex, and by decoding the residual of the displacement acceleration of the vertex in the deformable three-dimensional mesh (such as cloth) in any displacement direction in the sequence of three-dimensional meshes, the data amount of the decoded residual can be reduced, thereby reducing the code stream length of the sequence of deformable three-dimensional meshes such as flexible body or elastic body, and reducing the code rate.

[0636] In the method flow shown in FIG. 5C, the encoding end converts the reconstructed positions of the motion of the N historical frames and the M historical frames into the cylindrical coordinate system to decode the current frame, so that when the motion vertex in the sequence of frames performs a self-rotating motion (such as a person wearing a virtual clothes performing ballet dance, driving the virtual clothes to perform a self-rotating motion), the angular acceleration of the motion vertex in the sequence of three-dimensional meshes about the cylindrical coordinate system is almost unchanged, and then the decoding end calculates the predicted displacement motion information (including displacement acceleration) of the motion vertex in each displacement direction in the cylindrical coordinate system, so that the calculated predicted motion information can match the motion mode of the three-dimensional mesh, and then the predicted motion information is closer to the actual motion information, thereby improving the accuracy of the reconstructed actual motion information of the current frame, and ensuring the accurate decoding of the current frame.

[0637] Different from the embodiment of FIG. 5C, the application also provides a decoding method, which can be implemented based on the flow shown in FIG. 3G.

[0638] The method flow of the embodiment and the embodiment of FIG. 5C are mostly the same, and the difference is that the reconstructed actual motion information of the current frame obtained by S803a shown in FIG. 5C is not implemented according to the process shown in FIG. 5C, but the actual motion information (such as actual motion trend information) of the current frame is directly decoded from the encoding data of the current frame to obtain the reconstructed actual motion information of the current frame. The other processes are the same as the flow of FIG. 5C, which will not be described here.

[0639] In this way, when decoding the current frame, the residual does not need to be decoded, the predicted motion information does not need to be calculated, and the reconstructed result of the actual motion information does not need to be superimposed based on the predicted motion information and the residual, but the reconstructed result of the actual motion information can be directly decoded from the encoding data, and the decoding complexity is lower, thereby reducing the time delay of encoding and decoding.

[0640] FIG. 5D is another method flowchart for decoding a current frame (a three-dimensional mesh to be decoded) according to an embodiment of the present application. The method flowchart can be implemented based on the architecture shown in FIG. 1A, the flowcharts shown in FIG. 4C, FIG. 5A, and FIG. 5B, but is not limited to the architecture and flowcharts shown in FIG. 1A, FIG. 4C, FIG. 5A, and FIG. 5B.

[0641] The method flowchart shown in FIG. 5D is the same as the method flowchart shown in FIG. 5C in most parts, except that in the method flowchart shown in FIG. 5D, the actual motion information reconstructed by the decoding end for the current frame is three offset accelerations in three offset components, not only the offset acceleration of the motion vertex in the angle component as in the method flowchart shown in FIG. 5C, but also the offset acceleration of the vertex in the radius component and the offset acceleration of the vertex in the height component.

[0642] Specifically, taking the vertex P shown in FIG. 5B as an example, the actual motion information reconstructed by the vertex P in the i-th frame of the decoding end can include the reconstructed offset acceleration of the vertex P in the i-th frame in the height component, the reconstructed offset acceleration of the vertex P in the i-th frame in the radius component, and the reconstructed offset acceleration of the vertex P in the i-th frame in the angle component. The predicted motion information of the vertex P in the i-th frame determined by the decoding end can include the reconstructed offset acceleration of the vertex P in the i-1-th frame in the height component, the reconstructed offset acceleration of the vertex P in the i-1-th frame in the radius component, and the reconstructed offset acceleration of the vertex P in the i-1-th frame in the angle component.

[0643] The code stream received in the method flowchart shown in FIG. 5D can be the code stream encoded by the method flowchart shown in FIG. 3D.

[0644] As shown in FIG. 5D, the method can mainly include the following steps:

[0645] S901a: Based on the reconstructed positions of the motion vertices in the i-3-th frame, the i-2-th frame, and the i-1-th frame, and the instantaneous frame rates of the i-2-th frame and the i-1-th frame, the decoding end calculates the reconstructed offset acceleration of the motion vertex in the i-1-th frame in the height component, the reconstructed offset acceleration of the motion vertex in the i-1-th frame in the radius component, and the reconstructed offset acceleration of the motion vertex in the i-1-th frame in the angle component, to obtain the reconstructed actual motion information of the motion vertex in the i-1-th frame (which can be used as the predicted motion information of the motion vertex in the i-th frame).

[0646] The principle of the implementation process of S901a is the same as that of S601, S602a, and S603 in the embodiment of FIG. 3D, which will not be repeated here.

[0647] S901b: The decoding end decodes the received code stream to obtain the residual between the actual motion information and the predicted motion information of the i-th frame.

[0648] The residual includes a residual of a velocity of the motion vertex between the i-th frame and the i-1-th frame on the height component, a residual of a velocity of the motion vertex on the radius component, and a residual of an acceleration of the motion vertex on the angle component.

[0649] Specifically, referring to the input data in S604 in FIG. 3D, the residual includes a residual between an actual acceleration of the motion vertex of the i-th frame on the height component and a reconstructed acceleration of the motion vertex of the i-1-th frame on the height component, a residual between an actual acceleration of the motion vertex of the i-th frame on the radius component and a reconstructed acceleration of the motion vertex of the i-1-th frame on the radius component, and a residual between an actual acceleration of the motion vertex of the i-th frame on the angle component and a reconstructed acceleration of the motion vertex of the i-1-th frame on the angle component.

[0650] After S901a and S901b, S902a can be performed.

[0651] S902a: The decoding end reconstructs the actual motion information of the motion vertex of the i-th frame based on the residual obtained in S901b and the predicted motion information of the motion vertex of the i-th frame obtained in S901a.

[0652] The decoding end can add the residual of the acceleration of the motion vertex P between the i-th frame and the i-1-th frame on the height component (e.g. ) obtained in S901b to the reconstructed acceleration of the motion vertex P on the height component of the i-1-th frame obtained in S901a to obtain the reconstructed acceleration of the motion vertex P on the height component of the i-th frame

[0653] The decoding end can add the residual of the acceleration of the motion vertex P between the i-th frame and the i-1-th frame on the radius component (e.g. ) obtained in S901b to the reconstructed acceleration of the motion vertex P on the radius component of the i-1-th frame obtained in S901a to obtain the reconstructed acceleration of the motion vertex P on the radius component of the i-th frame

[0654] The decoding end can add the residual of the acceleration of the motion vertex P between the i-th frame and the i-1-th frame on the angle component (e.g. ) obtained in S901b to the reconstructed acceleration of the motion vertex P on the angle component of the i-1-th frame obtained in S802b to obtain the reconstructed acceleration of the motion vertex P on the angle component of the i-th frame

[0655] For example, the actual motion information of the vertex P of the i-th frame includes the reconstructed offset acceleration in the height component the reconstructed offset acceleration in the height component the reconstructed offset acceleration in the radius component

[0656] S901c: the decoding end calculates the reconstructed offset velocity of the motion vertex of the i-1-th frame in the height component, the reconstructed offset velocity of the motion vertex of the i-1-th frame in the radius component, and the reconstructed offset velocity of the motion vertex of the i-1-th frame in the angle component based on the reconstructed position of each motion vertex of the i-2-th frame and the i-1-th frame and the instantaneous frame rate of the i-1-th frame, and the decoding end can obtain the reconstructed height, the reconstructed radius, and the reconstructed angle of the motion vertex of the i-1-th frame through coordinate system conversion.

[0657] The specific implementation principle of the decoding end performing S901c is similar to the principle of calculating the reconstructed offset velocity of the motion vertex of the i-1-th frame in the height component and the reconstructed offset velocity of the vertex of the i-1-th frame in the radius component through S501 and S502a in the embodiment of FIG. 3C, so as to realize the respective reconstructed offset velocities of the motion vertex of the i-1-th frame in the height component, the radius component, and the angle component. Here, no longer be repeated.

[0658] After S901c and S902a, the decoding end can perform S903.

[0659] S903: the decoding end can determine the reconstructed height, the reconstructed radius, and the reconstructed angle of the motion vertex of the i-th frame based on the reconstructed actual motion information of the motion vertex of the i-th frame obtained by the decoding end in S902a, the reconstructed height, the reconstructed radius, and the reconstructed angle of the motion vertex of the i-1-th frame after position conversion obtained by the decoding end in S901c, the reconstructed offset velocity of the motion vertex of the i-1-th frame in the height component, the reconstructed offset velocity of the motion vertex of the i-1-th frame in the radius component, and the reconstructed offset velocity of the motion vertex of the i-1-th frame in the angle component, and the instantaneous frame rate of at least one of the i-1-th frame and the i-th frame.

[0660] Specifically, the vertex P shown in FIG. 5B is taken as an example to illustrate:

[0661] The actual motion information of the vertex P of the i-th frame reconstructed by the decoding end includes the reconstructed offset acceleration in the height component the reconstructed offset acceleration in the radius component and the reconstructed offset acceleration in the angle component

[0662] The reconstructed height, the reconstructed radius, and the reconstructed angle corresponding to the position of the vertex P in the i-1-th frame are hi-1 r i-1 ,θ i-1 ;

[0663] The reconstruction offset rate of vertex P in the height component of the (i-1)th frame is The reconstruction offset rate of vertex P in the (i-1)th frame on the radius component is The reconstruction offset rate of vertex P in the (i-1)th frame on the angular component is

[0664] The time slot for each frame is Δt.

[0665] The following decoder can reconstruct the reconstructed height h of vertex P in the i-th frame. i Reconstruction radius r i Reconstructing angle θ i :

[0666] The decoding end can reconstruct the offset rate (e.g., offset rate) of vertex P in the height component based on the offset from the (i-2)th frame to the (i-1)th frame. The reconstructed offset acceleration of vertex P in the height component of the i-th frame. And the time slot of the i-th frame, to determine the position P of vertex P from the (i-1)-th frame. i-1 Reconstruction position P in frame i i The reconstructed offset height is calculated based on the height component; then, based on this reconstructed offset height and the height component h corresponding to the position of vertex P in the (i-1)th frame,... i-1 To determine the reconstructed height h corresponding to the reconstructed position of vertex P in the i-th frame. i .

[0667] Wherein, the offset rate of vertex P from frame i-2 to frame i-1 in the height component can be the offset rate described above. It could also be the offset rate. It can also be and The average value is equal to and Any value between [a_i] and [a_i]. The decoding end can base its position on vertex P in the (i-2)th frame. i-2 The height component, and the position P of vertex P in the (i-1)th frame. i-1 The height component and the time slot Δt of the (i-1)th frame are used to calculate the reconstruction offset rate of vertex P in the (i-1)th frame on the height component. For example Wherein, the offset height Δh i-1 For position P i-1 Height component and position P i-2 The difference in height components. Offset rate. The calculation process of the reconstructed offset velocity of the vertex P in the radius component from the i-2th frame to the i-1th frame is as follows:

[0668] For example, the decoding end can reconstruct the height h of the vertex P of the ith frame by formula 4.1 or formula 4.2 or formula 4.3. i .

[0669] The decoding end can determine the offset radius of the reconstructed position P of the vertex P of the ith frame in the radius component from the position P i-1 of the vertex P of the i-1th frame to the reconstructed position P i of the vertex P of the ith frame based on the reconstructed offset velocity of the vertex P in the radius component from the i-2th frame to the i-1th frame (for example, the reconstructed offset velocity ), the reconstructed offset acceleration of the vertex P in the radius component of the ith frame , and the time slot of the ith frame, and determine the reconstructed radius r i-1 corresponding to the reconstructed position of the vertex P of the ith frame based on the offset radius and the radius component r i corresponding to the position of the vertex P of the i-1th frame.

[0670] The reconstructed offset velocity of the vertex P in the radius component from the i-2th frame to the i-1th frame can be the reconstructed offset velocity , can be the reconstructed offset velocity , can be the average value of and , or any value between and . The decoding end can calculate the reconstructed offset velocity of the vertex P in the radius component of the i-1th frame based on the radius component of the position P i-2 of the vertex P of the i-2th frame, the radius component of the position P i-1 of the vertex P of the i-1th frame, and the time slot Δt of the i-1th frame. For example , the offset radius Δr i-1 is the difference between the radius component of the position P i-1 and the radius component of the position P i-2 . The calculation process of the reconstructed offset velocity is as follows:

[0671] For example, the decoding end can reconstruct the reconstructed radius r i of the vertex P of the ith frame by formula 5.1 or formula 5.2 or formula 5.3.

[0672] The decoding end can reconstruct the offset rate (e.g., reconstruction offset rate) based on the angular component of vertex P offset from frame i-2 to frame i-1. The offset acceleration in the angular component reconstructed from vertex P of the i-th frame. And the i-th frame time slot Δt, to determine the position P of vertex P from the (i-1)-th frame. i-1 Reconstruction position P in frame i i The offset angle in the angular component; then based on this offset angle and the angular component θ corresponding to the position of vertex P in the (i-1)th frame. i-1 To determine the reconstruction angle θ corresponding to the reconstruction position of vertex P in the i-th frame. i .

[0673] Wherein, the offset rate of vertex P from frame i-2 to frame i-1 in terms of the angular component can be the reconstructed offset rate described above. It could also be the reconstruction offset rate. It can also be and The average value is equal to and Any value between [a_i] and [a_i]. The decoding end can base its position on vertex P in the (i-2)th frame. i-2 The angle components, and the position P of vertex P in the (i-1)th frame. i-1 The angle component and the time slot Δt of the (i-1)th frame are used to calculate the reconstruction offset rate of vertex P in the (i-1)th frame on the angle component. For example Wherein, the offset angle Δθ i-1 For position P i-1 Angular components and position P i-2 The difference between the angular components. Offset rate. The calculation process:

[0674] For example, the decoding end can reconstruct the reconstructed angle θ of vertex P in the i-th frame using formula 6.1, 6.2, or 6.3. i .

[0675] S904: The decoding end converts the reconstructed height, reconstructed radius, and reconstructed angle of the motion vertex of the i-th frame into a Cartesian coordinate system to obtain the reconstructed position of the motion vertex of the i-th frame.

[0676] Differently from the method flow shown in FIG. 5C, in the method flow shown in FIG. 5D, the decoding end can decode the residual of the offset acceleration of the vertex of the current frame in three offset directions to reduce the decoded residual and improve the decoding rate. When decoding the current frame, the decoding end can determine the offset rate of the motion vertex of the previous frame in three offset directions by using the reconstructed positions of the vertices of the previous two frames (two frames that have been decoded before the current frame and are adjacent to the current frame), and then decode the current frame by using the offset rate of the motion vertex of the previous frame in three offset directions, and the reconstructed position of the vertex of the previous frame, and the actual motion information of the current frame that has been reconstructed, to ensure the correct decoding of the three-dimensional mesh in the case of small residual and low code rate. The code rate of the sequence of deformable three-dimensional mesh of flexible bodies or elastic bodies can be reduced.

[0677] Differently from the embodiment of FIG. 5D, the application further provides a decoding method that can be implemented based on the flow shown in FIG. 3G.

[0678] The method flow of the embodiment and the embodiment of FIG. 5D are mostly the same, and the difference is that the reconstructed actual motion information of the current frame obtained by S902a shown in FIG. 5D is not implemented according to the process shown in FIG. 5D, but is directly decoded from the encoding data of the current frame (including the actual motion information of the current frame, such as the actual motion trend information) to obtain the reconstructed actual motion information of the current frame. The other processes are the same as the flow of FIG. 5D, which will not be described here.

[0679] In this way, when decoding the current frame, the residual does not need to be decoded, the predicted motion information does not need to be calculated, and the reconstructed result of the actual motion information does not need to be obtained by superimposing the predicted motion information and the residual, but the reconstructed result of the actual motion information can be directly decoded from the encoding data, which has lower decoding complexity, thereby reducing the time delay of encoding and decoding.

[0680] FIG. 5E is another method flow chart for decoding the current frame (the three-dimensional mesh to be decoded) provided by the embodiment of the application, which can be implemented based on the architecture shown in FIG. 1A, the flows shown in FIG. 4C, FIG. 5A and FIG. 5B, but is not limited to the architecture and flow shown in FIG. 1A, FIG. 4C, FIG. 5A and FIG. 5B.

[0681] The implementation principle of the method flow shown in FIG. 5E is the same as that of the method flow shown in FIG. 5D, and the difference is that, in the method flow shown in FIG. 5D, the residual of the motion information of the current frame and the actual motion information of the reconstructed current frame are all three offset accelerations in the three offset components (h, r, θ) of the cylindrical coordinate system, while in the method flow shown in FIG. 5E, the decoding end does not need to perform coordinate system conversion on the vertex coordinates of the current frame and the historical frame, but directly performs reconstruction of the actual motion information of the vertex of the current frame in the Cartesian coordinate system, and performs reconstruction of the offset rate of the previous frame in the Cartesian coordinate system.

[0682] Specifically, taking the vertex P shown in FIG. 5B as an example, the coordinates of the vertex P are coordinates in a three-dimensional Cartesian coordinate system (including x, y, and z directions). The reconstructed actual motion information of the vertex P of the i-th frame can include a reconstructed offset acceleration of the vertex P of the i-th frame in the x direction, a reconstructed offset acceleration of the vertex P of the i-th frame in the y direction, and a reconstructed offset acceleration of the vertex P of the i-th frame in the z direction. The predicted motion information of the vertex P of the i-th frame can include a reconstructed offset acceleration of the vertex P of the i-1-th frame in the x direction, a reconstructed offset acceleration of the vertex P of the i-1-th frame in the y direction, and a reconstructed offset acceleration of the vertex P of the i-1-th frame in the z direction.

[0683] The code stream received in the method flow shown in FIG. 5E can be the code stream encoded by the method flow shown in FIG. 3E.

[0684] As shown in FIG. 5E, the method can mainly include the following steps:

[0685] S1101a: Based on the reconstructed positions of the motion vertices of the i-3-th frame, the i-2-th frame, and the i-1-th frame, and the instantaneous frame rates of the i-2-th frame and the i-1-th frame, the decoding end calculates the reconstructed offset accelerations of the motion vertices of the i-1-th frame in each of the x, y, and z directions to obtain the reconstructed actual motion information of the motion vertices of the i-1-th frame (which can be used as the predicted motion information of the motion vertices of the i-th frame).

[0686] The reconstructed positions of the motion vertices of the i-3-th frame, the i-2-th frame, and the i-1-th frame are coordinates in the Cartesian coordinate system, and the reconstructed positions include x component values, y component values, and z component values.

[0687] In this step, the decoding end can first determine the reconstructed displacement rate of the motion vertex in the i-2 frame in each of the x, y, z directions, and the reconstructed displacement rate of the motion vertex in the i-1 frame in each of the x, y, z directions, using the reconstructed positions of the motion vertex of the i-3 frame, the i-2 frame, and the i-1 frame, and the instantaneous frame rates of the i-2 frame and the i-1 frame; and then calculate the reconstructed acceleration of the motion vertex in the i-1 frame in each of the x, y, z directions, using the reconstructed displacement rate of the motion vertex in the i-2 frame in each of the x, y, z directions, the reconstructed displacement rate of the motion vertex in the i-1 frame in each of the x, y, z directions, and the instantaneous frame rate of the i-1 frame.

[0688] Taking the vertex P shown in FIG. 5B as an example, the decoding end can calculate the reconstructed acceleration of the vertex P in the i-1 frame in each of the x, y, z directions (respectively denoted as ), using the reconstructed displacement rate of the vertex P in the i-2 frame in each of the x, y, z directions (respectively denoted as ), the reconstructed displacement rate of the vertex P in the i-1 frame in each of the x, y, z directions (respectively denoted as ), and the instantaneous frame rate of the i-1 frame, to obtain the reconstructed actual motion information of the vertex P in the i-1 frame as the predicted motion information of the vertex P in the i frame.

[0689] The implementation principle of S1101a is similar to that of S702a and S703a in the embodiment of FIG. 3E, which will not be described herein again.

[0690] S1001b: The decoding end decodes the received code stream to obtain the residual between the actual motion information and the predicted motion information of the motion vertex in the i frame.

[0691] The residual includes the respective acceleration residuals of the motion vertex in the x, y, and z directions between the i frame and the i-1 frame.

[0692] Specifically, referring to the data input in S704 in FIG. 3E, the residual includes the residual between the actual acceleration of the motion vertex in the x direction in the i frame and the reconstructed acceleration of the motion vertex in the x direction in the i-1 frame, the residual between the actual acceleration of the motion vertex in the y direction in the i frame and the reconstructed acceleration of the motion vertex in the y direction in the i-1 frame, and the residual between the actual acceleration of the motion vertex in the z direction in the i frame and the reconstructed acceleration of the motion vertex in the z direction in the i-1 frame.

[0693] After S1101a and S1101b, S1102a can be performed.

[0694] S1102a: The decoding end reconstructs the actual motion information of the vertex of the i-th frame based on the residual obtained in S1101b and the predicted motion information of the vertex of the i-th frame obtained in S1101a.

[0695] Continuing with the vertex P shown in FIG. 5B as an example, the actual motion information reconstructed for the vertex P of the i-th frame can include the reconstructed offset acceleration of the vertex P in each of the x, y, and z directions (denoted as

[0696] The principle of calculating the reconstructed offset acceleration of the vertex P of the i-th frame in each component in S1102a is similar to that of calculating the reconstructed offset acceleration of the vertex P of the i-th frame in the angular direction in the embodiment of FIG. 5D, and thus is not described herein again.

[0697] S1101c: The decoding end calculates the reconstructed offset velocity of the motion vertex of the i-1-th frame in the x direction, the reconstructed offset velocity of the motion vertex of the i-1-th frame in the y direction, and the reconstructed offset velocity of the motion vertex of the i-1-th frame in the z direction based on the reconstructed positions of the motion vertex of each of the i-2-th frame and the i-1-th frame and the instantaneous frame rate of each of the i-1-th frame and the i-th frame.

[0698] The principle of implementing S1101c by the decoding end is similar to that of implementing S901c in the embodiment of FIG. 5D, and thus is not described herein again.

[0699] After S1101c and S1102a, the decoding end can perform S1103.

[0700] S1103: The decoding end can reconstruct the position of the vertex of the i-th frame based on the reconstructed actual motion information of the vertex of the i-th frame obtained in S1102a, the reconstructed offset velocities of the motion vertex of the i-1-th frame in each of the x, y, and z directions obtained in S1101c, and the instantaneous frame rate of the i-th frame and the reconstructed position of the motion vertex of the i-1-th frame.

[0701] The principle of implementing S1103 by the decoding end is similar to that of implementing S903 in the embodiment of FIG. 5D, and thus is not described herein again.

[0702] ​Differently from the method flow shown in FIG. 5C and FIG. 5D, in the method flow shown in FIG. 5E, the decoding end does not need to make coordinate system conversion on the position of the vertex of each frame, and can directly decode the vertex of the three-dimensional grid established on the three-dimensional Cartesian coordinate system to decode the residual of the offset acceleration of the vertex in three offset directions on the three-dimensional Cartesian coordinate system, so as to realize decoding of the current frame, and improve the decoding rate. In this way, when the vertex in the three-dimensional grid makes non-uniform motion with small motion direction change in the three-dimensional grid sequence (for example, a virtual person running at an accelerated speed drives the virtual clothes to make non-uniform motion), the acceleration of the vertex of the three-dimensional grid under the Cartesian coordinate system can be small, and then the decoding end calculates the predicted motion information of the vertex in each offset direction on the Cartesian coordinate system and the reconstructed actual offset motion information (including offset acceleration), so that the offset direction of the offset acceleration can match the motion direction and motion mode of the non-uniform motion, thereby improving the closeness of the predicted motion information and the actual motion information of the present application, reducing the residual, reducing the decoding data amount, improving the transmission efficiency and storage efficiency.

[0703] Differently from the embodiment of FIG. 5E, the application further provides a decoding method, which can be implemented based on the flow shown in FIG. 3G.

[0704] The method flow of the embodiment and the embodiment of FIG. 5E are mostly the same, and the difference is only that: as shown in FIG. 5E, the reconstructed actual motion information of the current frame obtained by S1102a is not implemented according to the process shown in FIG. 5E, but the encoding data of the current frame (including the actual motion information of the current frame, such as the actual motion trend information) is directly decoded to obtain the reconstructed actual motion information of the current frame. Other processes are the same as the flow of FIG. 5E, which will not be repeated here.

[0705] In this way, when decoding the current frame, there is no need to decode the residual, no need to calculate the predicted motion information, and no need to superimpose the reconstructed result of the actual motion information based on the predicted motion information and the residual, but the encoding data can be directly decoded to obtain the reconstructed result of the actual motion information, and the decoding complexity is lower, thereby reducing the coding delay.

[0706] It should be understood that the above FIG. 5C, FIG. 5D, and FIG. 5E only exemplarily show three method flows of decoding the current frame by the decoding end of the present application, and in other embodiments, the present application can also provide more decoding methods, and the code stream decoded by the decoding method is still the residual of the actual motion information and the predicted motion information of the vertex of the current frame, and the motion information (the actual motion information and the predicted motion information) is the respective offset motion information in the three offset directions of the three-dimensional coordinate system of the current frame, wherein there is offset acceleration in at least one of the respective offset motion information in the three offset directions, and the offset motion information in the remaining offset directions can be offset velocity or offset acceleration.

[0707] FIG. 6 is a structural schematic diagram of an encoding device 1000 of the present application. As shown in FIG. 6, the encoding device 1000 of the present embodiment can be applied to the above-mentioned encoding end. The encoding device 1000 can include an obtaining module 1001 and an encoding module 1002. The obtaining module 1001 is configured to obtain actual motion trend information of a first three-dimensional mesh to be encoded in a three-dimensional mesh sequence, the actual motion trend information indicating an actual motion trend of a motion vertex moving at a non-uniform speed in a time slot corresponding to the first three-dimensional mesh in the three-dimensional mesh sequence. The encoding module 1002 is configured to obtain encoding data of the first three-dimensional mesh based on the actual motion trend information, and encode the encoding data into a code stream corresponding to the three-dimensional mesh sequence.

[0708] In a possible implementation, the obtaining module 1001 is further configured to obtain predicted motion trend information of the first three-dimensional mesh, the predicted motion trend information indicating a predicted motion trend of the motion vertex in the time slot corresponding to the first three-dimensional mesh. The encoding module 1002 is further configured to obtain the encoding data based on the actual motion trend information and the predicted motion trend information.

[0709] In a possible implementation, the encoding data includes a residual between the actual motion trend information and the predicted motion trend information.

[0710] In a possible implementation, the encoding data includes the actual motion trend information.

[0711] In a possible implementation, the code stream includes a first syntax structure corresponding to the first three-dimensional mesh, and the first syntax structure includes an identifier of the first three-dimensional mesh.

[0712] In a possible implementation, the first syntax structure further includes information corresponding to the motion vertex, and the information corresponding to the motion vertex includes a first residual of the motion vertex in a first component, a second residual of the motion vertex in a second component, and a third residual of the motion vertex in a third component.

[0713] In a possible implementation, the first syntax structure further includes: a vertex identifier of the motion vertex.

[0714] In a possible implementation, the obtaining module 1001 is specifically configured to obtain N sets of reconstruction data corresponding to the N second three-dimensional meshes that have been encoded in the sequence of three-dimensional meshes, N being an integer greater than or equal to a set value; and determine the prediction motion trend information based on the N sets of reconstruction data.

[0715] In a possible implementation, the obtaining module 1001 is specifically configured to determine actual motion information of the first three-dimensional mesh based on original data of the first three-dimensional mesh and original data of a plurality of third three-dimensional meshes that have been encoded in the sequence of three-dimensional meshes.

[0716] In a possible implementation, the obtaining module 1001 is specifically configured to obtain N sets of encoding data corresponding to the N second three-dimensional meshes; obtain N sets of prediction motion trend information of the N second three-dimensional meshes, the N sets of prediction motion trend information being N prediction motion trends of the motion vertex in N time slots corresponding to the N second three-dimensional meshes; and decode the N sets of encoding data based on the N sets of prediction motion trend information to obtain the N sets of reconstruction data. The N second three-dimensional meshes and the N sets of prediction motion information are in a one-to-one correspondence.

[0717] In a possible implementation, the obtaining module 1001 is specifically configured to decode the N sets of encoding data to obtain N sets of residuals between N sets of actual motion trend information of the N second three-dimensional meshes and the N sets of prediction motion trend information, the N sets of actual motion trend information including N actual motion trends of the motion vertex in N time slots corresponding to the N sets of second three-dimensional meshes; determine N sets of actual motion trend information of the N second three-dimensional meshes based on the N sets of prediction motion trend information and the N sets of residuals, the N sets of actual motion trend information being N reconstruction information of the N actual motion trends of the motion vertex in N time slots corresponding to the N second three-dimensional meshes; and determine the N sets of reconstruction data based on the N sets of actual motion trend information and reconstruction data of S fourth three-dimensional meshes corresponding to each of the N second three-dimensional meshes, the S fourth three-dimensional meshes being three-dimensional meshes that have been decoded in the sequence of three-dimensional meshes, S being an integer greater than or equal to 2.

[0718] In a possible implementation, at least one of the first residual, the second residual, and the third residual is a residual of motion acceleration.

[0719] In a possible implementation, the first component is an X component, the second component is a Y component, and the third component is a Z component.

[0720] In a possible implementation, the first component is an angle component, the second component is a height component, and the third component is a radius component.

[0721] In a possible implementation, the first three-dimensional mesh, the N second three-dimensional meshes, and the plurality of third three-dimensional meshes have the same topology.

[0722] In a possible implementation, the N second three-dimensional meshes are displayed in the three-dimensional mesh sequence in an order earlier than the first three-dimensional mesh in the three-dimensional mesh sequence.

[0723] In a possible implementation, the plurality of third three-dimensional meshes are displayed in the three-dimensional mesh sequence in an order earlier than the first three-dimensional mesh in the three-dimensional mesh sequence.

[0724] The encoding apparatus 1000 in this embodiment can be used to implement the technical solutions of the above-described various encoding method embodiments, and has similar implementation principles and technical effects, which will not be described here again.

[0725] FIG. 7 is a structural schematic diagram of a decoding apparatus 2000 in the application. As shown in FIG. 7, the decoding apparatus 2000 in this embodiment can be applied to the decoding end described above. The decoding apparatus 2000 can include an acquisition module 2001 and a decoding module 2002. The acquisition module 2001 is configured to acquire a code stream corresponding to a three-dimensional mesh sequence. The decoding module 2002 is configured to acquire, from the code stream, encoding data of a first three-dimensional mesh to be decoded in the three-dimensional mesh sequence, acquire actual motion trend information of the first three-dimensional mesh based on the encoding data, where the actual motion trend information indicates an actual motion trend of a motion vertex that moves at a non-uniform speed in the three-dimensional mesh sequence in a time slot corresponding to the first three-dimensional mesh, and obtain reconstruction data of the first three-dimensional mesh based on the actual motion trend information.

[0726] In a possible implementation, the acquisition module 2001 is specifically configured to acquire predicted motion trend information of the first three-dimensional mesh, where the predicted motion trend information indicates a predicted motion trend of the motion vertex in the time slot corresponding to the first three-dimensional mesh, and obtain the actual motion trend information based on the encoding data and the predicted motion trend information.

[0727] In a possible implementation, the encoding data includes a residual error between the actual motion trend information and the predicted motion trend information.

[0728] In a possible implementation, the encoded data comprises the actual motion trend information.

[0729] In a possible implementation, the encoded data has a first syntax structure, and the first syntax structure comprises an identification of the first three-dimensional mesh.

[0730] In a possible implementation, the first syntax structure further comprises information corresponding to the motion vertex, and the information corresponding to the motion vertex comprises: a first residual of the motion vertex on a first component; a second residual of the motion vertex on a second component; and a third residual of the motion vertex on a third component.

[0731] In a possible implementation, the first syntax structure further comprises: a vertex identification of the motion vertex.

[0732] In a possible implementation, the decoding module 2002 is further configured to: acquire, based on the code stream, N sets of reconstruction data corresponding to N second three-dimensional meshes that have been decoded in the sequence of three-dimensional meshes, N being an integer greater than or equal to a set value; and determine the predicted motion trend information based on the N sets of reconstruction data.

[0733] In a possible implementation, the decoding module 2002 is specifically configured to: acquire, based on the code stream, reconstruction data corresponding to a plurality of third three-dimensional meshes that have been decoded in the sequence of three-dimensional meshes; and obtain the reconstruction data of the first three-dimensional mesh based on the actual motion trend information and the reconstruction data of the plurality of third three-dimensional meshes.

[0734] In a possible implementation, at least one of the first residual, the second residual, and the third residual is a residual of motion acceleration.

[0735] In a possible implementation, the first component is an X component, the second component is a Y component, and the third component is a Z component.

[0736] In a possible implementation, the first component is an angle component, the second component is a height component, and the third component is a radius component.

[0737] In a possible implementation, the first three-dimensional mesh, the N second three-dimensional meshes, and the plurality of third three-dimensional meshes have the same topological structure.

[0738] The decoding apparatus 2000 of this embodiment can be used to execute the technical solutions of each of the decoding method embodiments described above, and has similar implementation principles and technical effects, which are not described herein again.

[0739] The application further provides a coding system, which can include the coding device 1000 and the decoding device 2000.

[0740] FIG. 8 is a schematic structural diagram of the device 1200 provided by the application. The device 1200 can include a processor 1201 and optionally a transceiver circuit 1202. Optionally, the device 1200 further includes a memory 1203.

[0741] The various components of the device 1200 are coupled together by a bus 1204, which can include a data bus, a power bus, a control bus, and a state signal bus, etc. However, for the sake of clarity, the various buses are referred to as the bus 1204 in the figure.

[0742] Optionally, the memory 1203 can be used to store the instructions in the above method embodiments.

[0743] The processor 1201 can be used to execute the instructions in the memory 1203 and control the transceiver circuit 1202 to receive signals and control the transceiver circuit 1202 to send signals.

[0744] The device 1200 can be an electronic device or a chip in an electronic device or a server, etc. at the coding / decoding end in the above method embodiments.

[0745] In the implementation process, the steps of the above method embodiments can be completed by integrated logic circuits of hardware in the processor 1201 or instructions in the form of software. The processor 1201 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiments of the application can be directly embodied in a hardware coding processor for execution, or a combination of hardware and software modules in the coding processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage medium is located in the memory, and the processor reads the information in the memory and combines it with the hardware to complete the steps of the above method.

[0746] The memory mentioned in each of the above embodiments can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example and not limitation, many forms of RAM can be used, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory of the system and method described herein is intended to include, but not be limited to, these and any other suitable types of memory.

[0747] In a possib...

Claims

1. An encoding method characterized by comprising: The method comprises: obtaining actual motion trend information of a first three-dimensional mesh to be encoded in a three-dimensional mesh sequence, the actual motion trend information indicating actual motion trend of a motion vertex in non-uniform motion in a time slot corresponding to the first three-dimensional mesh in the three-dimensional mesh sequence; obtaining encoding data of the first three-dimensional mesh based on the actual motion trend information; encoding the encoding data into a code stream corresponding to the three-dimensional mesh sequence.

2. The method of claim 1, wherein, The obtaining of the encoding data of the first three-dimensional mesh based on the actual motion trend information comprises: obtaining predicted motion trend information of the first three-dimensional mesh, the predicted motion trend information indicating predicted motion trend of the motion vertex in the time slot corresponding to the first three-dimensional mesh; obtaining the encoding data based on the actual motion trend information and the predicted motion trend information.

3. The method of claim 2, wherein, The encoding data comprises a residual error between the actual motion trend information and the predicted motion trend information.

4. The method of claim 1, wherein, The encoding data comprises the actual motion trend information.

5. The method according to any one of claims 1 to 4, characterized in that, The code stream comprises a first syntax structure corresponding to the first three-dimensional mesh, and the first syntax structure comprises an identifier of the first three-dimensional mesh.

6. The method of claim 5, wherein, The first syntax structure further comprises information of the motion vertex, and the information of the motion vertex comprises: a first residual error of the motion vertex in a first component; a second residual error of the motion vertex in a second component; a third residual error of the motion vertex in a third component.

7. The method of claim 6, wherein, The information of the motion vertex further comprises: a vertex identifier of the motion vertex.

8. The method of claim 2, wherein, The obtaining of the predicted motion trend information of the first three-dimensional mesh comprises: obtaining N sets of reconstruction data corresponding to N second three-dimensional meshes that have been encoded in the three-dimensional mesh sequence, N being an integer greater than or equal to a set value; determining the predicted motion trend information based on the N sets of reconstruction data.

9. The method according to any one of claims 1 to 8, characterized in that, The obtaining of the actual motion trend information of the first three-dimensional mesh in the three-dimensional mesh sequence comprises: determining actual motion information of the first three-dimensional mesh based on original data of the first three-dimensional mesh and original data of a plurality of third three-dimensional meshes that have been encoded in the three-dimensional mesh sequence.

10. The method according to claim 8 or 9, characterized in that, The obtaining of the N sets of reconstruction data corresponding to the N second three-dimensional meshes that have been encoded in the three-dimensional mesh sequence comprises: obtaining N sets of encoding data corresponding to the N second three-dimensional meshes; obtaining N sets of predicted motion trend information of the N second three-dimensional meshes, the N sets of predicted motion trend information being N predicted motion trends of the motion vertex in N time slots corresponding to the N second three-dimensional meshes; decoding the N sets of encoding data based on the N predicted motion trends to obtain the N sets of reconstruction data.

11. The method of claim 6, wherein, At least one of the first residual error, the second residual error and the third residual error is a residual error of motion acceleration.

12. A decoding method, comprising: The method comprises: obtaining a code stream corresponding to a three-dimensional mesh sequence; obtaining encoding data of a first three-dimensional mesh to be decoded in the three-dimensional mesh sequence from the code stream; obtaining actual motion trend information of the first three-dimensional mesh based on the encoded data, the actual motion trend information indicating actual motion trend of a motion vertex in the three-dimensional mesh sequence that is in non-uniform motion in a time slot corresponding to the first three-dimensional mesh; obtaining reconstructed data of the first three-dimensional mesh based on the actual motion trend information.

13. The method of claim 12, wherein, The obtaining actual motion trend information of the first three-dimensional mesh based on the encoded data comprises: obtaining predicted motion trend information of the first three-dimensional mesh, the predicted motion trend information indicating predicted motion trend of the motion vertex in the time slot corresponding to the first three-dimensional mesh; obtaining the actual motion trend information based on the encoded data and the predicted motion trend information.

14. The method of claim 13, wherein, The encoded data comprises a residual error between the actual motion trend information and the predicted motion trend information.

15. The method of claim 12, wherein, The encoded data comprises the actual motion trend information.

16. The method according to any one of claims 12 to 15, characterized in that, The bitstream comprises a first syntax structure corresponding to the first three-dimensional mesh, and the first syntax structure comprises an identifier of the first three-dimensional mesh.

17. The method of claim 16, wherein, The first syntax structure further comprises information of the motion vertex, and the information of the motion vertex comprises: a first residual error of the motion vertex in a first component; a second residual error of the motion vertex in a second component; a third residual error of the motion vertex in a third component.

18. The method of claim 17, wherein, The information of the motion vertex further comprises: a vertex identifier of the motion vertex.

19. The method of claim 13, wherein, The obtaining predicted motion trend information of the first three-dimensional mesh comprises: obtaining, based on the bitstream, N sets of reconstructed data corresponding to N second three-dimensional meshes in the three-dimensional mesh sequence that have been decoded, N being an integer greater than or equal to a set value; determining the predicted motion trend information based on the N sets of reconstructed data.

20. The method of any one of claims 12 to 19, wherein, The obtaining reconstructed data of the first three-dimensional mesh based on the actual motion trend information comprises: obtaining, based on the bitstream, reconstructed data corresponding to a plurality of third three-dimensional meshes in the three-dimensional mesh sequence that have been decoded; obtaining the reconstructed data of the first three-dimensional mesh based on the actual motion trend information and the reconstructed data of the plurality of third three-dimensional meshes.

21. The method of claim 17, wherein, At least one of the first residual error, the second residual error and the third residual error is a residual error of motion acceleration.

22. A coding system characterized by comprise: an encoding device and a decoding device; the encoding device is configured to implement the method in any one of claims 1 to 11; the decoding device is configured to implement the method in any one of claims 12 to 21.

23. An encoding apparatus, comprising: comprise: one or more processors; a memory configured to store one or more programs; when the one or more programs are executed by the one or more processors, the encoding device is caused to implement the method in any one of claims 1 to 11.

24. A decoding apparatus, comprising: comprise: one or more processors; a memory configured to store one or more programs; when the one or more programs are executed by the one or more processors, the decoding device is caused to implement the method in any one of claims 12 to 21.

25. A computer-readable storage medium, characterized in that, A computer program comprising computer program code which, when executed on a computer, causes the computer to perform the method of any one of claims 1 to 11, or the method of any one of claims 12 to 21.

26. A computer program product, characterised in that, The computer program product comprises computer program code which, when the computer program code is run on a computer, causes the computer to perform the method of any one of claims 1 to 11, or the method of any one of claims 12 to 21.