Motion coding of dynamic grids using intra and inter graphic Fourier transforms
By using graphical Fourier transform (GFT) domain to represent motion data in dynamic grid encoding, the problems of low compression efficiency and high computational complexity in the prior art are solved, and the bit rate reduction and gradual reconstruction are achieved.
Patent Information
- Application Number
- CN202380072238.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-14
- Filing Date
- 2023-10-04
- Publication Date
- 2025-05-16
AI Technical Summary
When encoding motion data of dynamic grids, it is difficult to effectively utilize spatiotemporal correlations, resulting in low compression efficiency and high computational complexity.
The graphical Fourier transform (GFT) domain is used to represent motion data, and the GFT coefficients are derived through intra- or inter-frame grid connectivity and encoded into bitstreams. This approach allows discovery and exploitation of more signal correlations, lower bitrates, and supports progressive reconstruction.
Through the representation of the GFT domain, the bit rate of motion encoding is reduced, and the progressive reconstruction of grid geometry is supported, and the calculation complexity is low, which is suitable for real-time execution.
Smart Images

Figure CN120019412A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of European application number 22306565.7 filed on October 14, 2022, which is incorporated herein by reference in its entirety. Background Art
[0002] Computer generated or camera captured objects are often modeled by dynamic meshes. High quality representation and rendering of content containing dynamic meshes requires a large amount of data. In addition, efficient compression techniques help to deliver such content to consumers and store it. Typically, the geometry (vertex positions) of the mesh can be encoded directly or encoded relative to the geometry of a reference mesh. In the latter, a motion field representing the spatial relationship between the mesh and the reference mesh is encoded. The motion data of the encoded motion field usually contains spatial and temporal correlations. When designing encoding techniques, exploiting the spatiotemporal correlations present in the motion data can lead to a computationally efficient compression process. Summary of the invention
[0003] Disclosed herein are apparatus and methods for encoding motion data, which is a component in a dynamic mesh encoding process. As disclosed herein, motion vectors representing displacements between corresponding vertices of corresponding meshes in a sequence are represented in a graph Fourier transform (GFT) domain. GFT can be derived based on intra-frame mesh connectivity or based on inter-frame mesh connectivity. In the former, explicit motion data is used for representation, while in the latter, implicit motion data is used for representation, as disclosed herein. Compared with using the spatial domain to represent motion data, using the GFT domain to represent motion data allows more signal correlations to be discovered and utilized, which in turn leads to a reduced bit rate for motion coding. In addition, representing motion data by spectral coefficients (i.e., GFT coefficients) allows progressive reconstruction of mesh geometry, that is, progressively increasing the accuracy of the reconstructed vertex positions of the dynamic mesh. The techniques disclosed herein for encoding motion data have low computational complexity and can therefore be performed in real time.
[0004] Aspects disclosed in the present disclosure describe methods for encoding mesh data. These methods include receiving a mesh sequence, including geometric data of vertices of the meshes in the sequence, and then encoding motion data into a bitstream of encoded mesh data. The motion data represents the spatial displacement between corresponding vertices of corresponding meshes from the mesh sequence. Encoding the motion data includes transforming the geometric data based on the GFT to obtain GFT coefficients representing the motion data, and then encoding the GFT coefficients into a bitstream. Aspects disclosed herein also describe methods for decoding mesh data. These methods include receiving a bitstream of encoded mesh data including encoded motion data, and decoding the motion data from the bitstream. Decoding the motion data includes decoding the GFT coefficients representing the motion data, and then inversely transforming the decoded GFT coefficients based on the GFT to obtain decoded geometric data of the vertices of the meshes in the sequence.
[0005] Aspects disclosed in the present disclosure describe a device for encoding mesh data. The device includes at least one processor and a memory storing instructions. When executed by the at least one processor, the instructions cause the device to receive a mesh sequence, including geometric data of vertices of the meshes in the sequence, and encode motion data into a bit stream encoding the mesh data. The motion data represents the spatial displacement between corresponding vertices from the corresponding meshes in the sequence. Motion data encoding includes transforming the geometric data based on the GFT to obtain GFT coefficients representing the motion data, and then encoding the GFT coefficients into a bit stream. Aspects disclosed in the present disclosure also describe a device for decoding mesh data. The device includes at least one processor and a memory storing instructions. When executed by the at least one processor, the instructions cause the device to receive a bit stream of encoded mesh data including encoded motion data, and decode the motion data from the bit stream. Motion data decoding includes decoding the GFT coefficients representing the motion data, and then inversely transforming the decoded GFT coefficients based on the GFT to obtain decoded geometric data of the vertices of the meshes in the sequence.
[0006] Aspects disclosed in the present disclosure describe a non-transitory computer-readable medium including instructions executable by at least one processor to implement a method for encoding mesh data. These methods include receiving a mesh sequence, including geometric data of vertices of the meshes in the sequence, and encoding motion data into a bitstream of encoded mesh data. The motion data represents the spatial displacement between corresponding vertices of corresponding meshes from the mesh sequence. The motion data encoding includes transforming the geometric data based on the GFT to obtain GFT coefficients representing the motion data, and then encoding the GFT coefficients into a bitstream. Aspects disclosed in the present disclosure also describe a non-transitory computer-readable medium including instructions executable by at least one processor to implement a method for decoding mesh data. These methods include receiving a bitstream of encoded mesh data including encoded motion data, and decoding the motion data from the bitstream. The decoding of the motion data includes decoding the GFT coefficients representing the motion data, and then inversely transforming the decoded GFT coefficients based on the GFT to obtain decoded geometric data of the vertices of the meshes in the sequence.
[0007] This summary is provided to introduce in a simplified form a selection of concepts that will be further described in the following detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to limitations that address any or all disadvantages noted in any part of this disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 Surface refinement using an iterative subdivision process is shown in accordance with aspects of the present disclosure.
[0009] Figure 2 is a functional block diagram of an example system for dynamic grid encoding according to aspects of the present disclosure.
[0010] Figure 3 is a functional block diagram of an example system for dynamic grid decoding according to aspects of the present disclosure.
[0011] Figure 4 is a functional block diagram of an example basic trellis encoder according to aspects of the present disclosure.
[0012] Figure 5 is a functional block diagram of an example basic trellis decoder according to aspects of the present disclosure.
[0013] Figure 6 Different types of diagrams are shown in accordance with aspects of the present disclosure.
[0014] Figure 7 An inter-graph construction according to aspects of the present disclosure is shown.
[0015] Figure 8 is a flow chart of an example method for encoding mesh data according to aspects of the present disclosure.
[0016] Fig. 9 is a flow chart of an example method for decoding mesh data according to aspects of the present disclosure. DETAILED DESCRIPTION
[0017] The present disclosure applies to the field of motion data computation and coding in the context of dynamic mesh compression. Recently, the MPEG 3D Graphics Coding (MPEG-3DGC) group solicited proposals (CfP) for codec technologies related to compression of time-varying volumetric meshes (V-Mesh). See, CfP for Dynamic Mesh Coding, ISO / IEC JTC1 / SC 29 / WG 7, 2021. In response, the solution proposed by Mammou et al. was selected to become the MPEG V-Mesh test model, which will be used as the basis for future development of the standard. See, K. Mammou, J. Kim, A. Tourapis, D. Podborski, and K. Kolarov, “MPEG input document m59281-v4-[V-CG] Apple's DynamicMesh Coding CfP Response,” ISO / IEC JTC 1 / SC 29 / WG 7, 2022 (“Mammou”).
[0018] As further described herein, the dynamic mesh encoding described in Mammou suggests first decomposing a given mesh to be encoded into a base mesh and a displacement vector representing the spatial difference between the given mesh and the base mesh. Then, the base mesh and the displacement vector are encoded separately. The encoding of the base mesh can be performed by any static mesh encoding technique or with reference to a previously encoded base mesh (i.e., a reference base mesh). In the latter, motion vectors representing the displacements between corresponding vertices of the base mesh and the reference base mesh are encoded. Aspects of the present disclosure describe alternative techniques for calculating and encoding these motion vectors. Although these aspects are disclosed herein in the context of encoding a base mesh, as described in Mammou and applied to the V-MESH (V-DMC) encoding standard, these aspects can be applied to encoding any dynamic mesh that maintains the same connectivity and the same number of vertices.
[0019] In Mammou, the encoding of mesh geometry is based on a surface subdivision scheme that starts with a simple three-dimensional (3D) mesh called a base mesh. The base mesh contains a relatively small number of vertices and faces that are iteratively refined in a predictable way. To this end, a subdivision process is used to add new vertices and faces to the base mesh by iteratively subdividing existing faces into smaller sub-faces. The new vertices are then displaced to new positions according to predefined rules to gradually refine the mesh shape, thereby obtaining increasingly smooth and / or increasingly complex surfaces, such as Figure 1 shown.
[0020] Figure 1 Surface refinement using an iterative subdivision process 100 is shown. Figure 1 In the example of , the octahedral model 110, i.e. the basic mesh, is refined. Increasingly finer meshes (i.e. mesh subdivisions) 120-160 are generated, each of which is an iterative result of a subdivision of the previous mesh. Figure 1 , the finest mesh 160 is shown (the other 120-150 mesh subdivisions are shown in their faceted form) after a rendering operation (eg, using an interpolated shading rendering method) has been applied to demonstrate the smoothness of the resulting mesh subdivision.
[0021] Different surface subdivision schemes can be applied to the base mesh (e.g., 110). See, for example, A. Benton, "Advanced Graphics-Subdivision Surfaces", Cambridge University. In Mammou, a simple midpoint subdivision scheme is used, as described further below. Since the connectivity of the base mesh can be refined in a predictable manner using a set of subdivision rules known to both the encoder and the decoder, the only connectivity information that needs to be encoded and provided to the decoder is the connectivity of the base mesh. In addition to the base mesh connectivity, the base mesh geometry and the displacement vectors must also be encoded and provided to the decoder, as described below with reference to Figure 2-5 Further description of dynamic mesh encoding proposed in Mammou.
[0022] Typically, a mesh is a representation of a surface, including vertices associated with 3D positions on the surface; these vertices are connected by edges to form planes that approximate the surface (such as triangles). Other information can be associated with each vertex of the mesh, namely vertex attributes (e.g., normal vectors and color values). In addition, the surface can be further represented by various attributes, such as texture. Typically, the texture of the surface is described by a two-dimensional (2D) image, namely a texture map. In order to associate the faces of the mesh (e.g., triangles) with the corresponding texture data, the faces of the mesh are mapped into a 2D space (e.g., UV parameter space) associated with the texture map. Similarly, the surface can be associated with other data types, which are provided by other attribute maps, characterizing other physical properties of the surface (e.g., surface reflectivity and transparency), which may be required for realistic rendering of the surface. Therefore, the surface representation of mesh data includes topological data and attribute data - the topology of the surface is represented by the mesh M (including geometry and connectivity information, and possibly vertex attributes), and the attributes of the surface are represented by the attribute map A (including the attribute map and corresponding mapping information). Aspects described here with respect to texture data (represented by texture maps) are applicable to other types of data (typically represented by attribute maps).
[0023] Figure 2 2 is a functional block diagram of an example system 200 for dynamic mesh coding. The system 200 illustrates coding of a sequence of frames F(i), wherein data associated with frame i includes a mesh M(i) 205 and corresponding attribute maps A(i) 210. The system 200 includes a mesh decomposer 220 (e.g., part of a pre-processing unit) and an encoder 230. The mesh decomposer 220 is configured to decompose a received mesh M(i) 205 into a base mesh m(i) 222 and corresponding displacement vectors d(i) 224. The generated base mesh m(i) 222 and displacement vectors d(i) 224 are fed into the encoder 230 along with the corresponding attribute maps A(i) 210. The encoder 230 encodes the obtained data - m(i), d(i), and A(i), thereby generating corresponding bitstreams, including a base mesh bitstream 270, a mesh displacement bitstream 275, and an attribute map bitstream 280. The operation of mesh decomposer 220 and the operation of encoder 230 will be further described below.
[0024] The decomposer 220 is configured to decompose the mesh M(i) 205 into a base mesh m(i) 222 and a corresponding displacement vector d(i) 224. To generate the base mesh m(i), the decomposer 220 extracts the mesh M(i) by subsampling the vertices of the mesh (e.g., generating Figure 1Then, a mesh subdivision (e.g., 120) is generated by subdividing the basic mesh m(i), that is, each surface of the basic mesh is subdivided into a plurality of sub-surfaces, introducing additional new vertices. Figure 1 As shown, any subdivision scheme may be optionally applied iteratively. For example, each triangle of the basic mesh surface may be split into four sub-triangles by introducing three new vertices in the middle of the edge of the triangle and connecting these three vertices.
[0025] Next, the decomposer 220 determines displacement vectors d(i) 224 for the corresponding vertices of the subdivided base mesh, so that when applied to these vertices, a deformed mesh is generated that spatially fits the given mesh M(i) 205 to be encoded. Decomposing the given mesh M(i) in this way - allowing the base mesh m(i) and its corresponding displacement vectors d(i) to be encoded, rather than directly encoding the mesh M(i) - improves compression efficiency. This is because the base mesh m(i) has fewer vertices relative to the mesh M(i) and can therefore be encoded with a relatively small number of bits. In addition, the displacement vectors d(i) can be efficiently encoded using, for example, a wavelet transform implemented by the subdivision structure. Subsequently, the subdivision structure used does not need to be explicitly encoded as it can be determined by the decoder. For example, the decoder can subdivide the decoded base mesh based on the subdivision scheme type and the subdivision iteration count that can be signaled in the bitstream.
[0026] like Figure 2 As shown, the encoder 230 includes a basic grid encoder 235, a basic grid decoder 240, a grid displacement encoder 245, a grid displacement decoder 250, a grid reconstructor 255 and an attribute map encoder 260. The basic grid encoder 235 is configured to encode the basic grid m(i) into a coded basic grid cm(i) and generate a basic grid bit stream 270 therefrom. The basic grid decoder 240 is configured to reconstruct (decode) the basic grid from the coded basic grid cm(i) to generate a reconstructed quantized basic grid m'(i) and a reconstructed basic grid m"(i). The basic grid encoder 235 and the decoder 240 refer to Figure 4 and Figure 5 Further description. The grid displacement encoder 245 receives the basic grid m ( i ) and the reconstructed quantized basic grid m'( i ) as input, based on which it is configured to encode the received displacement vector d(i) into a coded displacement vector cd( i), and generates a grid displacement bitstream 275 therefrom. The grid displacement decoder 250 is configured to reconstruct (decode) a displacement vector from the encoded displacement vector cd(i), thereby generating a reconstructed displacement vector d″(i). Based on the reconstructed basic grid m″(i) and the reconstructed displacement vector d″(i), the grid reconstructor 255 is configured to reconstruct (decode) the grid into a reconstructed grid DM(i). To this end, the reconstructed basic grid m ” (i) is subdivided (according to the subdivision scheme used), and then the reconstructed displacement vector d''(i) is applied to the subdivided base mesh, in effect deforming the subdivided base mesh to obtain DM(i). Based on the mesh M(i) and the reconstructed mesh DM( i ), the property graph encoder 260 is configured to encode (multiple) property graphs A(i) 210 into (multiple) encoded property graphs and generate a property graph bitstream 280 therefrom.
[0027] Specifically, the grid displacement encoder 245 is a displacement vector d ( i ) is encoded, as described above, the displacement vectors are associated with the corresponding vertices of the subdivided basic mesh. To this end, the displacement vectors are first updated based on the reconstructed quantized basic mesh m'(i). Then, according to the subdivision scheme used, a wavelet transform is applied to represent the updated displacement vectors d'(i) - that is, the wavelet coefficients are extracted according to the subdivision process that has been used to subdivide the basic mesh. These wavelet coefficients are then quantized, packed into a 2D image, and compressed by a video encoder. The grid displacement decoder 250 generally operates inversely to the grid displacement encoder 245. Therefore, the grid displacement decoder 250 employs a video decoder to decode the packed 2D image compressed by the video encoder of the grid displacement encoder 245 (if the video encoder is lossy). The grid displacement decoder 250 then unpacks the 2D image to obtain the quantized wavelet coefficients, and applies inverse quantization, followed by an inverse wavelet transform, thereby generating the reconstructed displacement vectors d"(i).
[0028] Note that a video encoder is applied to the tasks of compressing the packed wavelet coefficients (via the trellis displacement encoder 245) and compressing the attribute map(s) (via the attribute map encoder 260). Any video encoding method (lossless or lossy) may be used for these tasks, depending on the requirements of a particular application.
[0029] Figure 32 is a functional block diagram of an example system 300 for dynamic grid decoding. The system 300 is configured to generally operate inversely to the system 200, and includes a decoder 330 and a grid reconstructor 360. The decoder 330 includes a base grid decoder 335, a grid displacement decoder 340, and a property map decoder 350. The base grid decoder 335 decodes the base grid m" (i) from the base grid bitstream 310, 270, as shown in FIG. Figure 5 As further described. The grid displacement decoder 340 decodes the reconstructed displacement vector d" (i) from the grid displacement bit stream 315, 275, and performs reference Figure 2 The steps described for the grid displacement decoder 250 are as follows. The attribute map decoder 350 decodes the attribute map from the attribute map bitstream 320, 280, inversely to the operation of the attribute map encoder 260, to generate a reconstructed attribute map DA(i) 375. The output of the decoder 330 (reconstructed basic grid m" (i) and reconstructed displacement vectors d" (i)) are used by the grid reconstructor 360 to reconstruct the decoded grid DM(i) 370.
[0030] Figure 4 4 is a functional block diagram of an example basic grid encoder 400. The basic grid encoder 400 includes a quantizer 420, a static grid encoder 440, a motion encoder 450, and a selector 460. Figure 2 As described above with respect to the basic grid encoder 235, the basic grid encoder 400 is configured to encode the basic grid m(i) into a basic grid bitstream 480. To this end, two encoders 440, 450 may be employed. Thus, after quantization 420, the static grid encoder 440 encodes the quantized basic grid qm(i) according to any static grid encoding method. Additionally, after quantization 420, the motion encoder 450 encodes the quantized basic grid qm(i) relative to a reference basic grid, i.e., a reconstructed quantized basic grid, denoted as m'(j). For example, the reference basic grid m'(j) may be associated with a previously reconstructed quantized basic grid m'(i-1) of the frame sequence F(i). Thus, the motion encoder 450 encodes a motion field f(i) that describes the motion that the vertices of m'(j) must undergo in order to reach the corresponding positions of the corresponding vertices of qm(i) (or vice versa), as further described below.
[0031] Therefore, when the motion encoder 450 is used, it is assumed that the base mesh and the reference base mesh share the same number of vertices and the same vertex connectivity, that is, only the positions of the corresponding vertices change over time. In order to maintain the same number of vertices and the same vertex connectivity in the base meshes of the frame sequence, the encoder 400 can, for example, track the transformation of the geometry applied to the previous base mesh and apply it to the current base mesh. Under such conditions, the motion encoder 450 can be configured to first calculate the motion field f(i) and then encode the calculated motion field f(i) into the base mesh bitstream 480. The motion field f(i) contains the motion vectors of the corresponding vertices in the quantized base mesh qm(i) and the reference reconstructed quantized base mesh m'(j), as shown below: f(i)=P1(i)-P2(j), (1) Wherein P1(i) is a vector containing geometric data (vertex positions) of the quantized base mesh qm(i), and wherein P2(j) is a vector containing geometric data (corresponding vertex positions) of the reference reconstructed quantized base mesh m'(j). In one aspect, for example, the motion encoder 350 can further adjust the motion vector (e.g., based on neighboring motion vectors) and then encode the adjusted motion vector using an entropy encoder.
[0032] The selector 460 can perform a selection of whether to use the output of the static mesh encoder 440 or the output of the motion encoder 450. In Mammou, it is recommended to select the bitstream of the encoder (440 or 450) that results in the smallest geometric distortion. The preferred method is to consider the total rate-distortion cost introduced by the dynamic mesh encoding (via the encoder 230) when selecting between the output of the static mesh encoder 440 and the output of the motion encoder 450. Therefore, a rate-distortion optimization that takes into account topological and photometric distortions and bit rate levels can be performed. This rate-distortion optimization can result in the selection of an encoder (440 or 450) that will provide a more efficient encoding, corresponding to the best rate-distortion cost, as disclosed in application number EP22306231.6 entitled "Rate Distortion Optimization for TimeVaryingTextured Mesh Compression", the disclosure of which is incorporated herein by reference in its entirety.
[0033] Figure 5 5 is a functional block diagram of an example basic grid decoder 500. The basic grid decoder 500 generally operates inversely to the basic grid encoder 400. It 500 includes a static grid decoder 540, a motion decoder 550, and an inverse quantizer 560. As described above with reference to Figure 3As described above with respect to the basic grid decoder 335, the basic grid decoder 500 is configured to decode the reconstructed basic grid m″(i) from the basic grid bitstream 520, 480. To this end, the basic grid decoder 500 directs the incoming basic grid stream 520 (representing the encoded basic grid cm(i)) to the static grid decoder 540 or the motion decoder 550. This directing can be performed based on signaling in the bitstream 520 indicating whether the encoded basic grid cm(i) is encoded by the static grid encoder 440 or the motion encoder 450. If the bitstream 520 is If the bitstream 520 is directed to a static grid decoder 540, the decoder decodes the basic grid from the bitstream 520 to produce a reconstructed quantized basic grid m'(i). Otherwise, if the bitstream 520 is directed to a motion decoder 550, the decoder decodes the motion field from the bitstream 520 and adds the reconstructed (decoded) motion field to the reference reconstructed quantized basic grid m'(j), thereby producing a reconstructed quantized basic grid m'(i). The obtained m'(i) is then provided to the inverse quantizer 560, which thereby generates a reconstructed basic grid m"(i). As described above, the basic grid decoder 500 is also used in the encoder 230, wherein it 240 provides the reconstructed quantized basic grid m'(i) and the reconstructed basic grid m"(i) to the grid displacement encoder 245 and the grid reconstructor 255, respectively.
[0034] Various aspects of the present disclosure describe the application of a Graphics Fourier Transform (GFT) to compute and encode a motion field f( i ) (i.e., motion vector). See A. Ortega, P. Frossard, J. J. M. F. Moura and P. Vandergheynst, “Graph signal processing: Overview, challenges, and applications”, Transactions of the IEEE, Vol. 106, No. 5, pp. 808-828, 2018. The GFT is an extension of the classical Fourier transform to a more general domain: data residing on irregular graphs. A 3D mesh model is an example of such data. “Irregular” in this context means that each vertex in the mesh can be connected to a variable number of other vertices, so that the network of vertex connections across the mesh is irregular. Such a network can be described by a planar graph, denoted G = (V, E), where V represents the set of mesh vertices (graph nodes) and E represents the set of mesh edges (connections between vertices).
[0035] Figure 6Different types of graphs are shown. In practice, aspects disclosed herein are typically applied to graphs 610 with simple connectivity ("simple" graphs). A simple graph is a graph in which: 1) links between different nodes are undirected (i.e., edges have no direction); 2) there are no multiple links between any pair of nodes (i.e., there cannot be more than one edge connecting any pair of vertices, as shown in graph 620); 3) there are no cycles around any nodes (i.e., each edge connects two different vertices, rather than any one vertex to itself, as shown in graph 630); and 4) the graph links are unweighted (i.e., edges have no weights associated with them, as they are all considered equally important, which is equivalent to giving each edge a weight of 1).
[0036] Karni et al. showed how to use the GFT to obtain "spectral compression" of 3D mesh geometry. See Z. Karni and C. Gotsman, "Spectral compression of mesh geometry", SIGGRAPH'00, New Orleans, Louisiana, USA, 2000 ("Karni"). In Karni, it is assumed that the vertex position vectors of a mesh (considering the x, y, and z coordinates separately) can be represented as linear combinations of a small number of orthogonal basis vectors. Such orthogonal basis vectors can be obtained from the combined mesh Laplacian matrix. This is similar in principle to the transform coding technique used in the JPEG image compression standard, which is based on using discrete cosine transform (DCT) basis vectors to obtain a more compact representation of image pixel data.
[0037] The computation of the combined mesh Laplacian matrix L depends only on the mesh (graph) connectivity. For a mesh with n vertices, L is a square n×n matrix and is computed as follows: L=DA, (2) where A is a symmetric adjacency matrix of dimension n×n. If vertex i (i.e., v i ) is connected to vertex j (i.e., v j ), then the matrix elements of A at positions (i, j) and (j, i) have the value "1", and have the value "0" otherwise. D is a "degree matrix" of dimension n×n, which contains the sum of the adjacency matrix values on the main diagonal of the corresponding rows (or columns), and is zero at all other positions. The value of the diagonal element i in D (i.e., the element of D(i, i)) is considered to be the degree or valence of vertex i, denoted by deg(v i ) represents the number of edges connected to that vertex. The formal mathematical definition of L can be written as:
[0038] To obtain the basis vectors, the eigenvectors (n×1 column vectors) and eigenvalues (n scalars) of the matrix L are computed. The eigenvalues are then sorted in ascending order of their magnitudes, and their corresponding eigenvectors are sorted accordingly. The normalized versions of the ordered eigenvectors of the Laplacian matrix L, the Laplacian eigenvectors, form the orthogonal basis vectors and are denoted here as L eigenvectors .
[0039] Taubin showed that when the connectivity information of the grid is used for computation, the Laplacian eigenvectors form the vector space An orthogonal basis of (where n is the number of vertices in the mesh) is provided, and such an orthogonal basis can therefore be used to represent mesh geometric data. See G Taubin, "A Signal Processing Approach to Fair Surface Design", SIGGRAPH'95, Los Angeles, CA, USA, 1995. The representation of mesh geometric data by Laplace eigenvectors can be analogized to the representation provided by Fourier basis vectors, where the corresponding eigenvalues can be analogized to the corresponding frequencies associated with the Fourier basis vectors. Thus, the arrangement of the eigenvalues from lowest to highest magnitude, and the arrangement of their corresponding eigenvectors in the same order, effectively places all of the "lowest frequency" basis vectors first, followed by increasingly "higher frequency" basis vectors. Thus, the eigenvector corresponding to an eigenvalue of zero can be considered to be the "DC" component (using the above analogy).
[0040] As Karni said, each dimension of the mesh's geometric data (i.e., each vertex position vector X = {x1, x2, ..., x n}, Y = {y1, y2, ..., y n}, Z = {z1, z2, ..., z n}) can be projected onto the same set of Laplace eigenvectors (basis vectors) by matrix multiplication to obtain 3 sets of spectral coefficients (i.e., GFT coefficients), each of which is a vector of size 1 × n. For example, for X and the corresponding set of spectral coefficients, each coefficient in the set indicates "how many" corresponding basis vectors (eigenvectors) are needed to represent X as a linear combination of all eigenvectors.
[0041] When encoding the GFT coefficients, since the coefficients are usually quantized before entropy coding, there will be some irreversible losses, resulting in lossy reconstruction (decoding) of the mesh geometry data. However, the key advantage of this transform coding method is that for a relatively smooth mesh, the resulting coefficients will have large amplitudes only for those coefficients corresponding to the lower frequency basis vectors, while the other coefficients will have zero or close to zero values. Therefore, a good approximation of the original mesh can be obtained by encoding only a part of the coefficients (those corresponding to the lower frequency basis vectors). In addition, the coefficients (part or all) can be encoded and transmitted so that the decoder can gradually improve the reconstructed mesh based on the coefficients received so far. Therefore, an elegant progressive reconstruction of the mesh geometry data (shape) can be achieved with different quality levels (i.e., different accuracy levels of the reconstruction of the mesh vertex position vectors X, Y and Z). Since the mesh connectivity (based on its derived Laplacian eigenvectors) remains unchanged, the only information that must be encoded over the frame sequence is the changes in the basic mesh geometry (i.e., the changes in the vertex position vectors X, Y and Z).
[0042] Therefore, since the Laplace eigenvector (i.e., the matrix L eigenvectots ) is independent of the mesh geometry, so these feature vectors can be independently calculated at the decoder end based on the connectivity data of the mesh. The connectivity data of the mesh can in turn be provided to the decoder in the same bitstream (e.g., bitstream 480) that represents the geometry of the basic meshes in the sequence F(i), or independently provided to the decoder from another source. It is not necessary to provide indices to the decoder to sort the feature vectors, as the decoder can sort the eigenvalues in the same manner as the encoder and sort the feature vectors accordingly. Note also that since the Laplace eigenvectors are orthogonal and contain real values (no complex numbers), the inversion of L can be obtained by simply transposing eigenvectors , that is, L eigenvectors -1 =L eigenvectors T .
[0043] The main known limitation of applying GFT to represent mesh geometry (or other data associated with mesh vertices) is that it requires computing the eigenvectors of the Laplacian matrix at both the encoder and decoder sides. Performing such computations for very large meshes (e.g., more than a few thousand vertices) is both time-consuming and susceptible to numerical instabilities, which can lead to undesirable results. However, as described herein, when applying GFT to small meshes such as base meshes, no such limitation exists.
[0044] In the first method disclosed herein, the GFT is derived based on intra-frame mesh connectivity. In this method, the motion vector is explicitly represented in the GFT domain. Therefore, the motion vector associated with a vertex from a mesh (e.g., a base mesh) can be represented by the difference between the GFT coefficient of the vertex and the GFT coefficient of the corresponding vertex from another mesh (e.g., a reference base mesh). In the second method disclosed herein, the GFT is derived based on inter-frame mesh connectivity. In this method, the motion vector is implicitly represented by an inter-frame graph. That is, a linear graph (corresponding vertices on each connected mesh sequence) can be constructed, and then each of these linear graphs can be transformed into the GFT domain. Therefore, the motion vector is implicitly represented by a representation of the changing geometry of the mesh sequence in the GFT domain. In both methods, depending on the desired mesh reconstruction quality level, a subset of the calculated GFT coefficients can be selected for encoding (e.g., by motion encoder 450) and / or decoding (e.g., by motion decoder 550) to facilitate progressive representation.
[0045] In the literature by Thanou et al., GFT is used to encode motion vectors of dynamic point clouds. See, D. Thanou, PAChou and P. Frossard, "Graph-based compression of dynamic 3D point clouds sequences", IEEE Transactions on Image Processing, Vol. 25, No. 4, pp. 1765-1778, 2016. Therein, GFT is applied to motion vectors that are first explicitly calculated between different frames of the sequence. However, the aspects disclosed herein apply GFT directly to the vertex x, y, z positions of a mesh across different frames, and therefore do not require direct calculation of motion vectors because they are represented in the GFT domain. In addition, in Thanou, since GFT is applied to dynamic point clouds (not dynamic meshes), an octree data structure must be constructed to spatially organize the input point cloud data and calculate graphs on these data. No additional data structure is required here to apply GFT to the geometry of the mesh. In addition, the motion estimation problem in Thanou is formulated as a feature matching problem on a dynamic graph, where the features are wavelet coefficients computed from the spectral graph wavelet (SGW) at each node of the graph at different scales. See DK Hammond, P. Vandergheynst and R. Gribonval, Wavelets on graphs via spectralgraph theory, Applied and Computational Harmonic Analysis, Vol. 30, No. 2, pp. 129-150, 2011. These feature descriptors are then used to compute point-to-point correspondences between graphs of different frames. In contrast, the aspects described herein do not require point-by-point feature matching.
[0046] Next, various aspects of the above-mentioned first method (in which the GFT is derived based on the intra-frame mesh connectivity) are described with reference to encoding steps C1.1-C1.10 and decoding steps D1.1-D1.6. These steps are described with reference to the following: 1) the base mesh of the corresponding frame in the sequence; 2) the reconstructed quantized reference base mesh m′(j) (which can be the first base mesh in the frame sequence or the first base mesh in the frame group (GOF)); and 3) the base mesh m(i) (which can be the next base mesh in the sequence or the next base mesh in the GOF). Note that these steps are generally applicable to meshes in the mesh sequence that maintain the same connectivity and the same number of mesh vertices.
[0047] Therefore, steps C1.1-C1.10 may be performed to encode the motion data by the motion encoder 450, as described below.
[0048] Step C1.1: Calculate the Laplacian matrix L based on the intra-frame mesh connectivity. The mesh connectivity can be derived from the first mesh in the sequence (e.g., m′(j)). Then, calculate the eigenvectors and eigenvalues of L, sort the eigenvectors according to the order of the sorted eigenvalues, and normalize the sorted eigenvectors to obtain the orthogonal basis vector L eigenvectors .
[0049] Step C1.2: Change the vertex (i.e. X m′(j) ={x1, x2, ..., x n}、Y m′(j) ={y1, y2, ..., y n} and Z m′(j) ={z1, z2, ..., z n}) The associated geometric data is projected onto L eigenvectors to obtain the GFT coefficient (CX m′(j) , CY m′(j) , CZ m′(j) ), as follows: The operator × represents matrix multiplication, L eigenvectors is an n×n matrix, where each column represents an eigenvector (basis vector), and CX m′(j) CY m′(j) and CZ m′(j) The GFT coefficients in are all vectors of size 1×n.
[0050] Step C1.3: Repeat step C1.2 for the geometric data associated with the vertices of the base mesh m(i) whose motion vectors are calculated relative to m′(j) to obtain the GFT coefficients CX m(i) CY m(i) and CZ m(i) .
[0051] Step C1.4: Select a subset of the GFT coefficients to be encoded, thus available to the decoder for signal reconstruction. The most important (largest magnitude) coefficients will generally be at the front of the 1×n vector of coefficients, since they correspond to the lowest frequencies. The more coefficients that are retained, the more accurate the final reconstruction of the base grid vertex positions will be. For the x, y and z coefficient vectors, the same coefficients must be kept (i.e., coefficients corresponding to the same eigenvector index). In the simplest case, the same cut-off can be used for all coefficient vectors, for all base grids - for example, 50% of the lowest frequency coefficients can be retained and the rest are set to 0. If coefficients are discarded in a linear manner (i.e., sequentially rather than by non-linear indexing into the coefficient array), then there is no need to provide the decoder with information about which eigenvector indices of the coefficients that have been retained. Alternatively, instead of selecting a subset of the spectral coefficients, the encoder can progressively provide all coefficients to the decoder, and the decoder can decide when to stop the decoding process, for example, when the reconstructed signal has reached a sufficient quality level, or when a bit rate limit has been reached at the decoder system.
[0052] Step C1.5: Quantize the (selected) GFT coefficients of m′(j) (according to the chosen quantization method). Then, the quantized coefficients (denoted as CX′ m′(j ) 、 CY′ m′(j) and CZ′ m′(j) ) is entropy encoded and transmitted to the decoder.
[0053] Step C1.6: Dequantize the (selected) GFT coefficients of m′(j). The dequantized GFT coefficients of m′(j) are expressed as and Steps C1.5 and C1.6 are performed to avoid accumulation of quantization errors between frames when calculating consecutive motion vectors.
[0054] Step C1.7: Calculate the differences in the (selected) GFT coefficients for the x, y and z components as follows: The matrix MV m′(j),m(i) represents the motion vector between the basic grids m′(j) and m(i) in the GFT domain. The motion matrix MV m′(j),m(i) is a matrix of size 3×n.
[0055] Step C1.8: Quantize the values in the motion matrix (according to the selected quantization method). The quantized motion matrix is denoted as MV′ m′(j),m(i) .
[0056] Step C1.9: Entropy encode the quantized motion matrix and transmit it to the decoder. In one aspect, any entropy encoding method can be applied to different frequency bands (of corresponding eigenvalues) independently or collectively.
[0057] Step C1.10: Steps C1.2-C1.9 are repeated for each subsequent basic grid in the sequence (or in the same GOF) (except steps C1.5 and C1.6, since only the coefficients of the first reference basic grid are quantized and sent to the decoder), wherein the reference basic grid for each iteration is, for example, the reconstructed basic grid m′(i) from the previous iteration. Note that the choice of reference basic grid in each iteration is not restricted, although it may be beneficial to update the reference grid in each iteration in order to improve the chances of having smaller motion vectors compared to measuring motion relative to a reference grid located several frames before the frame of the basic grid.
[0058] The encoded motion data may be decoded by motion decoder 550, generally by reversing the motion encoding steps C1.1-C1.10 above, as described below in steps D1.1-D1.6.
[0059] Step D1.1: Calculate the Laplacian matrix L based on the intra-frame mesh connectivity. The mesh connectivity can be derived from the decoded version of m′(j). Then, calculate the eigenvectors and eigenvalues of L, sort the eigenvectors according to the order of the sorted eigenvalues, and normalize the sorted eigenvectors to obtain the orthogonal basis vectors L eigenvectots .
[0060] Step D1.2: Entropy decoding and dequantization of the GFT coefficients of m'(j) (received from step C1.5). These dequantized coefficients are represented by and
[0061] Step D1.3: Entropy decoding and dequantization of motion vector MV m′(j),m(i) (received by step C1.9). These dequantized motion vectors are represented as
[0062] Step D1.4: Dequantize the motion vector and the dequantized coefficients and Add to reconstruct the GFT coefficients of m(i) (the reverse of the encoder operation in step C1.7):
[0063] Step D1.5: Reconstruct the coefficients by and With L eigenvectors TThe corresponding Laplace eigenvector linear combination is used to reconstruct the (x, y, z) vertex position value of m(i), as shown below: Where the operator × represents matrix multiplication, and the operator T represents matrix transposition. Similarly, referring to the (x, y, z) vertex position values of the base mesh m′(j) (i.e., and ) can be based on and Rebuild.
[0064] Step D1.6: Repeat steps D1.3-D1.5 for each successive base mesh in the sequence (or in the same GOF), iteratively updating the reference base mesh as needed (this depends on how the reference base mesh is selected in the encoder). The motion vectors can be obtained from the geometry data of the corresponding vertices of the base mesh and the reference base mesh. In this case, for example, and The corresponding vertex in gets the motion vector.
[0065] Aspects of the above second method (where the GFT is derived based on inter-frame grid connectivity) are described below with reference to encoding steps C2.1-C2.8 and decoding steps D2.1-2.5. In this case, inter-frame pictures can be used to implicitly represent motion vectors. Figure 7 An example for constructing an inter-frame graph is shown in Figure 7 In the example of , a dynamic mesh sequence 710, 720, 730 is shown, where each basic mesh includes n = 3 vertices. The basic meshes across the frame sequence maintain the same connectivity because only their vertex positions (indexed by j = {1, 2, 3}) change across frames. The mesh sequence 710, 720, 730 spans M frames, indexed by f = {1, 2, ..., M}. Three inter-frame graphs are shown - G1, G2 and G3 - each of which connects corresponding vertices across M frames. Therefore, for each vertex j, an inter-frame graph G is constructed across the mesh sequence 710, 720, 730. j , which has M nodes (one per frame), such as Figure 7 Note that the M frames may be frames from the entire frame sequence, or may be frames of a GOF (in the case where each GOF is processed separately).
[0066] Because the inter-frame graph G jis linear (i.e., vertices of the first base mesh in the sequence are connected only to corresponding vertices of the second base mesh in the sequence, which in turn are connected only to corresponding vertices of the third base mesh in the sequence, etc.), the only information needed to build the graph (on both the encoder and decoder sides) is the number of vertices in the base meshes and the number of frames in the sequence (or GOF) that need to be connected together. Also note that because the inter-frame graph construction is independent of the input mesh geometry or connectivity, L eigenvectors The feature vectors and feature values in can be pre-computed by both the encoder and the decoder and reused for a mesh sequence (or GOF) with the same number of frames in the sequence (or GOF) and the same number of mesh vertices (e.g., the same number of vertices in each basic mesh).
[0067] The following steps C2.1-C2.8 may be performed to encode the motion data using the implicit motion data representation by the motion encoder 450. Steps C2.1-C2.8 are applicable to a sequence of meshes that maintain the same connectivity and the same number of mesh vertices.
[0068] Step C2.1: Connect corresponding vertices v across M frames j , construct an inter-frame graph G across M frames of the grid sequence j , j = {1, 2, ..., n}, as referenced Figure 7 Then, for each graph G j , j = {1, 2, ..., n} perform the following steps.
[0069] Step C2.2: As explained with reference to equations (2) and (3), based on G j Connectivity, calculate the M×M adjacency matrix A j , M×M degree matrix D j and the M×M combined Laplacian matrix L j =D j -A j .
[0070] Step C2.3: Laplacian matrix L j The Laplace eigenvectors and eigenvalues are calculated and sorted, and then the sorted Laplace eigenvectors are normalized to obtain the orthogonal basis vectors An M×M matrix.
[0071] Step C2.4: Construct G representing the span of frames f = {1, 2, ..., M} j A 1×M vector of vertex geometry data, i.e. X j ={x j,f=1 , x j,f=2 , ..., x j,f=M},Y j ={yj,f=1 ,y j,f=2 , ..., y j,f=M}, Z j ={z j,f=1 , z j,f=2 , ..., z j,f=M}.
[0072] Step C2.5: Project the geometric vector from step C2.4 onto To obtain the GFT coefficient CX j , CY j , CZ j The 3×M matrix is as follows: The × operator represents matrix multiplication.
[0073] Step C2.6: Quantize the GFT coefficients (according to the selected quantization method). The quantized coefficients are represented by CX′ j CY′ j and CZ′ j .
[0074] Step C2.7: Select the GFT coefficients to be encoded, from which signal reconstruction can be performed. The method of selection can be different, but the same coefficients should be selected for the x, y and z coefficient vectors (i.e., the coefficients corresponding to the same eigenvector index). The same coefficients should also be selected across the GFT vectors. j , j = {1, 2, ..., n} are selected. For example, if the 50% lowest frequency coefficients are selected for graph G1, then the lowest frequency coefficients should be selected for all other graphs (G2, ..., G n ) select the same coefficients to avoid introducing distortion between the reconstructed base mesh vertices. Furthermore, to avoid encoding the feature vector indices, the unselected coefficients should be discarded in a linear manner (i.e., sequentially rather than by non-linear indexing into the coefficient array). Alternatively, all GFT coefficients can be encoded and transmitted progressively to the decoder so that the latter determines when to stop decoding. For example, the decoder can stop decoding the received coefficients when the quality of the reconstructed dynamic base mesh is satisfactory, or when the decoder system has reached its bit rate limit.
[0075] Step C2.8: Entropy encode the (selected) quantized GFT coefficients and transmit them to the decoder. In one aspect, any entropy encoding method can be applied independently or collectively to different graphs G j quantization coefficient of .
[0076] The encoded motion data may be decoded by motion decoder 550, generally the reverse of the encoding steps C2.1-C2.8 above, as described below in steps D2.1-D2.6.
[0077] Step D2.1: Construct inter-frame graph G j , j = {1, 2, ..., n}, as in step C2.1. Then, for each graph G j , j = {1, 2, ..., n} perform the following steps.
[0078] Step D2.2: Get the Laplacian matrix L j , as in step C2.2.
[0079] Step D2.3: Obtain orthonormal basis vectors As in step C2.3.
[0080] Step D2.4: Decoding and dequantization corresponding to G j The GFT coefficients of these dequantized coefficients are expressed as and
[0081] Step D2.5: Reconstruct the GFT coefficients by and With The corresponding Laplace eigenvector linear combination in the reconstruction representation G j A vector of geometric data for the vertices, as follows: Wherein the operator × represents matrix multiplication, and the operator T represents matrix transposition. The motion data may be obtained from the geometric data of consecutive corresponding vertices across the inter-frame graph.
[0082] Figure 8 8 is a flow chart of an example method 800 for encoding mesh data according to aspects of the present disclosure. Method 800 begins at step 810 by receiving a mesh sequence comprising geometric data for vertices of the meshes in the sequence. Method 800 performs encoding of motion data into a bitstream (e.g., bitstream 480) encoding the mesh data. The motion data represents spatial displacements between corresponding vertices from corresponding meshes in the sequence. Encoding of the motion data includes, in step 820, transforming the geometric data of the mesh sequence into GFT coefficients representing the motion data based on the GFT. Then, in step 830, encoding the GFT coefficients into a bitstream. In one aspect, method 800 also includes selecting a subset of the GFT coefficients, wherein only the selected subset of the GFT coefficients is encoded into the bitstream (e.g., as explained in steps C1.4 and C2.7 above).
[0083] In one aspect, the motion data is encoded using an explicit motion data representation, wherein the GFT is derived based on the intra-frame mesh connectivity of the meshes in the sequence (see step C1.1). In this aspect, the method 800 also includes: 1) transforming the geometric data associated with the vertices of the first mesh of the sequence based on the GFT to obtain a first set of GFT coefficients (see step C1.2); 2) transforming the geometric data associated with the vertices of the second mesh of the sequence based on the GFT to obtain a second set of GFT coefficients (see step c1.3); 3) encoding the first set of GFT coefficients (see step c 1.5); and 4) encoding the spectral difference between the corresponding GFT coefficients of the first set and the second set, wherein the spectral difference represents the motion vector associated with the vertex of the second mesh (see steps C1.7-C1.9). The encoded first set of GFT coefficients and the encoded spectral difference are then added to a bitstream (e.g., bitstream 480) containing the encoded motion data.
[0084] In another aspect, motion data is encoded using an implicit motion data representation. In this aspect, method 800 begins by constructing an inter-graph including corresponding vertices of meshes across a sequence (see step C2.1), and then deriving a GFT based on the inter-mesh connectivity of the inter-graph (see steps C2.2-C2.3). Method 800 also includes 1) transforming geometric data of vertices across the inter-graph based on the GFT to obtain GFT coefficients (see step C2.5), which represent motion vectors associated with corresponding vertices across the inter-graph; and 2) encoding the GFT coefficients (see step C2.8). The encoded GFT coefficients are then added to a bitstream (e.g., bitstream 480) containing the encoded motion data.
[0085] Fig. 9 9 is a flow chart of an example method 900 for decoding mesh data according to aspects of the present disclosure. Method 900 begins at step 910 by receiving a bitstream of encoded mesh data, which includes encoded motion data representing spatial displacements between corresponding vertices of corresponding meshes in a mesh sequence. Method 900 includes decoding of motion data in steps 920-930. In step 920, GFT coefficients representing the motion data are decoded. Then, in step 930, the decoded GFT coefficients are inversely transformed based on the GFT to obtain decoded geometric data for the vertices of the meshes in the sequence. The motion data can be obtained from the decoded geometric data. In one aspect, the motion data can be decoded progressively. In this case, only a subset of the GFT coefficients are decoded (and used for reconstruction by the decoder) from the GFT coefficients encoded into the bitstream (e.g., as explained in steps C1.4 and C2.7 above).
[0086] In the case where motion data is encoded using an explicit motion data representation, wherein the GFT is derived based on mesh connectivity of meshes in a sequence (see step D1.1), method 900 further includes: 1) decoding a first set of GFT coefficients from a bitstream, which is calculated by the encoder based on the GFT using geometric data associated with vertices of a first mesh of the sequence (see step D1.2); 2) decoding spectral differences between corresponding GFT coefficients of the first set and a second set of GFT coefficients from the bitstream, the second set being calculated by the encoder based on the GFT using geometric data associated with vertices of a second mesh of the sequence, the decoded spectral differences representing motion vectors associated with the vertices of the second mesh (see step D1.3); 3) inversely transforming the first set of GFT coefficients based on the GFT to obtain geometric data of the vertices of the first mesh (see step D1.5); 4) obtaining GFT coefficients of the second set by adding corresponding GFT coefficients of the first set and the spectral differences (see step D1.4); and 5) inversely transforming the second set of GFT coefficients based on the GFT to obtain geometric data of the vertices of the second mesh (see step D1.5). In this case, the motion vectors may be obtained from the geometric data of corresponding vertices of the first mesh and the second mesh.
[0087] In the case of encoding motion data using implicit motion data representation, method 900 begins by constructing an inter-graph including corresponding vertices of meshes across a sequence (see step D2.1), and then, based on the inter-graph mesh connectivity of corresponding vertices across the inter-graph, deriving a GFT (see steps D2.2-D2.3). Method 900 also includes: 1) decoding GFT coefficients from the bitstream, which are calculated by the encoder based on the GFT using geometric data of corresponding vertices across the inter-graph (see step D2.4), wherein the decoded GFT coefficients represent motion vectors associated with corresponding vertices across the inter-graph; and 2) inversely transforming the decoded GFT coefficients based on the GFT to obtain geometric data of vertices across the inter-graph (see step D2.5). In this case, the motion vectors can be obtained from the geometric data of consecutive corresponding vertices across the inter-graph.
[0088] The illustrations of the aspects described herein are intended to provide an overall understanding of the structure, function and operation of each aspect. The illustrations are not intended to be a complete description of all elements and features of the devices and systems utilizing the structures or methods described herein. Many other aspects may be apparent to those skilled in the art after reading this disclosure. Other aspects may be utilized and derived from this disclosure so that structural and logical replacements and changes may be made without departing from the scope of this disclosure. Therefore, this disclosure and the accompanying drawings should be considered illustrative rather than restrictive.
[0089] The description of these aspects is provided to enable making or using these aspects. Various modifications to these aspects will be apparent, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but to conform to the widest possible scope consistent with the principles and novel features defined by the appended claims.
Claims
1. A method for encoding grid data, comprising: Receive a sequence of meshes, including geometric data of vertices of the meshes in the sequence; as well as Encoding motion data into a bitstream encoding mesh data, the motion data representing spatial displacements between corresponding vertices from corresponding meshes in the sequence, the motion data encoding comprising: transforming the geometric data based on a Graphic Fourier Transform (GFT) to obtain GFT coefficients representing the motion data, and Encode the GFT coefficients into a bitstream.
2. The method according to claim 1, wherein encoding of the motion data further comprises: A subset of the GFT coefficients is selected, wherein encoding of the GFT coefficients comprises encoding the selected subset of the GFT coefficients.
3. The method according to claim 1 or 2, wherein the encoding of the motion data further comprises: Based on the intra-mesh connectivity of the meshes in the sequence, the GFT is derived.
4. The method according to claim 3, wherein encoding of the motion data further comprises: Transforming geometric data associated with vertices of a first mesh of the sequence based on the GFT to obtain a first set of GFT coefficients; transforming the geometric data associated with the vertices of the second mesh of the sequence based on the GFT to obtain a second set of GFT coefficients; encoding a first set of GFT coefficients; as well as encoding a spectral difference between corresponding GFT coefficients of the first set and the second set, wherein the spectral difference represents a motion vector associated with a vertex of the second mesh, The encoded motion data includes the encoded first set and the encoded spectral difference.
5. The method according to claim 1 or 2, wherein encoding of motion data further comprises: constructing an inter-frame graph including corresponding vertices of meshes across the sequence; as well as Based on the inter-frame mesh connectivity of the inter-frame graph, the GFT is derived.
6. The method according to claim 5, wherein encoding of the motion data further comprises: transforming geometric data of vertices across the inter-frame image based on a GFT to obtain GFT coefficients, wherein the GFT coefficients represent motion vectors associated with corresponding vertices across the inter-frame image; as well as Encode the GFT coefficients, The encoded motion data includes encoded GFT coefficients.
7. A method for decoding grid data, comprising: receiving a bitstream of encoded mesh data, the encoded mesh data including encoded motion data representing spatial displacements between corresponding vertices from respective meshes in a sequence of meshes; as well as Decode motion data from the bitstream. Motion data decoding includes: Decoding the GFT coefficients representing the motion data, and The decoded GFT coefficients are inversely transformed based on the GFT to obtain decoded geometric data of the vertices of the mesh in the sequence.
8. The method of claim 7, wherein the decoding of the motion data further comprises: The motion data is progressively decoded, wherein decoding of the GFT coefficients comprises decoding a subset of the GFT coefficients encoded into the bitstream.
9. The method according to claim 7 or 8, wherein the decoding of the motion data further comprises: Based on the intra-mesh connectivity of the meshes in the sequence, the GFT is derived.
10. The method of claim 9, wherein the decoding of the motion data further comprises: decoding from the bitstream a first set of GFT coefficients computed by the encoder based on the GFT using geometry data associated with vertices of a first mesh of the sequence; decoding a spectral difference from the bitstream, the spectral difference between corresponding GFT coefficients of a first set of GFT coefficients and a second set, the second set being calculated by the encoder based on the GFT using geometry data associated with vertices of a second mesh of the sequence, wherein the spectral difference represents motion vectors associated with the vertices of the second mesh; Performing an inverse transformation on the GFT coefficients of the first set based on the GFT to obtain geometric data of vertices of the first mesh; Obtaining a second set of GFT coefficients by adding corresponding GFT coefficients of the first set and the spectral difference; as well as Based on the GFT, the GFT coefficients of the second set are inversely transformed to obtain geometric data of the vertices of the second mesh.
11. The method according to claim 10, further comprising: The motion vectors are obtained from the geometric data of corresponding vertices of the first mesh and the second mesh.
12. The method according to claim 7 or 8, wherein the decoding of the motion data further comprises: constructing an inter-frame graph including corresponding vertices of meshes across the sequence; as well as Based on the inter-frame mesh connectivity of the inter-frame graph, the GFT is derived.
13. The method of claim 12, wherein the decoding of the motion data further comprises: Decoding GFT coefficients from the bitstream, which are calculated by the encoder based on the GFT using geometry data of corresponding vertices of the inter-frame graph, wherein the GFT coefficients represent motion vectors associated with corresponding vertices of the inter-frame graph; as well as The decoded GFT coefficients are inversely transformed based on GFT to obtain the geometric data of the vertices of the inter-frame graph.
14. The method according to claim 13, further comprising: Motion vectors are obtained from the geometry data of consecutive vertices across the inter-frame graph.
15. An apparatus for encoding grid data, comprising: at least one processor; as well as a memory storing instructions that, when executed by the at least one processor, cause the apparatus to: Receives a sequence of meshes, including the geometry data for the vertices of the meshes in the sequence, and Encoding motion data into a bit stream encoding mesh data, the motion data representing spatial displacements between corresponding vertices from corresponding meshes in the sequence, the motion data encoding comprising: transforming the geometric data based on the GFT to obtain GFT coefficients representing the motion data, and Encode the GFT coefficients into a bitstream.
16. The device according to claim 15, wherein: The instructions also cause the device to: A subset of the GFT coefficients is selected, wherein encoding of the GFT coefficients comprises encoding the selected subset of the GFT coefficients.
17. The device according to claim 15 or 16, wherein: The instructions also cause the device to: Based on the intra-mesh connectivity of the meshes in the sequence, the GFT is derived.
18. The device according to claim 15 or 16, wherein: The instructions also cause the device to: constructing an inter-frame graph comprising corresponding vertices of meshes across the sequence; and Based on the inter-frame mesh connectivity of the inter-frame graph, the GFT is derived.
19. An apparatus for decoding grid data, comprising: at least one processor; as well as a memory storing instructions that, when executed by the at least one processor, cause the apparatus to: receiving a bitstream of encoded mesh data, the encoded mesh data comprising encoded motion data representing spatial displacements between corresponding vertices of respective meshes from a sequence of meshes, and Decode motion data from the bitstream. Motion data decoding includes: Decoding the GFT coefficients representing the motion data, and The decoded GFT coefficients are inversely transformed based on the GFT to obtain decoded geometric data for the vertices of the mesh in the sequence.
20. The device according to claim 19, wherein The instructions also cause the device to: The motion data is progressively decoded, wherein decoding of the GFT coefficients comprises decoding a subset of the GFT coefficients encoded into the bitstream.
21. The device according to claim 19 or 20, wherein: The instructions also cause the device to: Based on the intra-mesh connectivity of the meshes in the sequence, the GFT is derived.
22. The device according to claim 19 or 20, wherein: The instructions also cause the device to: constructing an inter-frame graph comprising corresponding vertices of meshes across the sequence; and Based on the inter-frame mesh connectivity of the inter-frame graph, the GFT is derived.
23. A non-transitory computer readable medium comprising instructions executable by at least one processor to perform a method for encoding mesh data, the method comprising: receiving a mesh sequence comprising geometric data of vertices of the meshes in the sequence; as well as Encoding motion data into a bitstream encoding mesh data, the motion data representing spatial displacements between corresponding vertices from corresponding meshes in the sequence, the motion data encoding comprising: transforming the geometric data based on the GFT to obtain GFT coefficients representing the motion data, and Encode the GFT coefficients into a bitstream.
24. A non-transitory computer readable medium comprising instructions executable by at least one processor to perform a method for decoding mesh data, the method comprising: receiving a bitstream of encoded mesh data, the encoded mesh data including encoded motion data representing spatial displacements between corresponding vertices from respective meshes in a sequence of meshes; as well as Decode motion data from the bitstream. Motion data decoding includes: Decoding the GFT coefficients representing the motion data, and The decoded GFT coefficients are inversely transformed based on the GFT to obtain decoded geometric data for the vertices of the mesh in the sequence.