Encoding method, decoding method, encoders, decoders and storage medium
By introducing the weights corresponding to points in three-dimensional grid coding, the prediction, update and quantization process is optimized, and the problem of unconsidered differences in importance of different points is solved, improving coding efficiency and reconstruction quality.
Patent Information
- Application Number
- PCT/CN2024/070330
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-03
- Publication Date
- 2025-07-10
AI Technical Summary
The prior art does not fully consider the degree of importance of different points in three-dimensional grid coding, resulting in insufficient coding efficiency of geometric information.
By introducing the weights corresponding to the points, adaptive prediction and update based on the spatial connection relationship of the points, optimize the quantization process, and improve the transformation coding efficiency.
It improves the coding efficiency of three-dimensional grid geometric information, reduces the coding rate, and ensures the reconstruction quality.
Smart Images

Figure CN2024070330_10072025_PF_FP_ABST
Abstract
Description
Coding and decoding method, codec and storage medium Technical Field
[0001] The present application relates to the field of video coding and decoding technology, and in particular to a coding and decoding method, a codec, and a storage medium. Background Art
[0002] Video dynamic mesh coding (V-DMC) requires encoding the geometric information of points in a mesh. Improving the coding efficiency of geometric information is a problem that needs to be solved.
[0003] Summary of the Invention
[0004] The embodiments of the present application provide a coding method, a codec, and a storage medium to improve the coding efficiency of geometric information of points in a grid. The following introduces various aspects of the present application.
[0005] In a first aspect, a decoding method is provided, which is applied to a decoder, comprising: parsing a bitstream to determine quantization coefficients of points in a three-dimensional grid; inversely quantizing the quantization coefficients of the points in the three-dimensional grid to determine the inverse quantization coefficients of the points in the three-dimensional grid; inversely transforming the inverse quantization coefficients of the points in a first layer to determine the shift coefficients of the points in the first layer; and predicting the shift coefficients of the points in a second layer based on the shift coefficients of the points in the first layer; wherein the points in the three-dimensional grid belong to multiple levels of detail (LOD) layers, the first layer and the second layer are two adjacent layers in the multiple LOD layers, and the first layer is a layer above the second layer.
[0006] In a second aspect, a coding method is provided, which is applied to an encoder, including: performing detail level LOD division on points in a three-dimensional grid to determine multiple LOD layers, wherein the multiple LOD layers include adjacent first and second layers, and the first layer is the upper layer of the second layer; predicting the shift coefficients of the points in the second layer to determine the residual coefficients of the points in the second layer; and transforming the shift coefficients of the points in the first layer according to the residual coefficients of the points in the second layer.
[0007] According to a third aspect, a decoder is provided, comprising: a first determination unit configured to parse a code stream and determine quantization coefficients of points in a three-dimensional grid; a second determination unit configured to dequantize the quantization coefficients of the points in the three-dimensional grid and determine the dequantization coefficients of the points in the three-dimensional grid; a third determination unit configured to inversely transform the dequantization coefficients of the points in the first layer and determine the shift coefficients of the points in the first layer; and a prediction unit configured to predict the shift coefficients of the points in the second layer based on the shift coefficients of the points in the first layer; wherein the points in the three-dimensional grid belong to multiple LOD layers, the first layer and the second layer are two adjacent layers among the multiple LOD layers, and the first layer is the upper layer of the second layer.
[0008] In a fourth aspect, a decoder is provided, comprising: a memory for storing a computer program; and a processor for executing the method described in the first aspect when running the computer program.
[0009] In a fifth aspect, an encoder is provided, comprising: a first determination unit, configured to perform LOD division on points in a three-dimensional grid, and determine multiple LOD layers, wherein the multiple LOD layers include adjacent first and second layers, and the first layer is the upper layer of the second layer; a second determination unit, configured to predict the shift coefficients of the points in the second layer, and determine the residual coefficients of the points in the second layer; and a transformation unit, configured to transform the shift coefficients of the points in the first layer according to the residual coefficients of the points in the second layer.
[0010] In a sixth aspect, an encoder is provided, comprising: a memory for storing a computer program; and a processor for executing the method described in the second aspect when running the computer program.
[0011] In a seventh aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed, the method as described in the first aspect or the method as described in the second aspect is implemented.
[0012] In an eighth aspect, a non-volatile computer-readable storage medium for storing a bit stream is provided, wherein the bit stream is generated by an encoding method using an encoder, or the bit stream is decoded by a decoding method using a decoder, wherein the decoding method is the method described in the first aspect, and the encoding method is the method described in the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] FIG1A is a schematic diagram of a three-dimensional grid image.
[0014] FIG1B is a partially enlarged view of the three-dimensional grid image.
[0015] FIG2 is a schematic diagram of the connection method of the three-dimensional grid.
[0016] FIG3A is a schematic diagram of a three-dimensional grid image.
[0017] FIG3B is a schematic diagram of a grid data storage format.
[0018] FIG3C is a property diagram of a three-dimensional grid image.
[0019] Figure 4 is the overall framework diagram of grid coding.
[0020] FIG5A is a schematic diagram of a grid preprocessing process.
[0021] FIG5B is a schematic diagram illustrating a method for generating shift coefficients.
[0022] FIG6A is a schematic diagram of a quantization processing method for mesh geometric information.
[0023] FIG. 6B is another schematic diagram of a quantization processing method for mesh geometric information.
[0024] FIG. 7A is a schematic diagram showing an encoding method for the connection relationship of triangular facets.
[0025] FIG7B is a schematic diagram of an encoding method for geometric information.
[0026] FIG. 7C is a schematic diagram of a texture coordinate encoding method.
[0027] FIG8A is a schematic diagram showing the basic principle of vertex shift coefficients.
[0028] FIG8B is a schematic diagram showing a mapping method of mapping the shift coefficient to a two-dimensional image.
[0029] FIG9 is a schematic diagram of an encoding method for inter-frame geometric information.
[0030] FIG10A is a schematic diagram of an intra-frame coding method.
[0031] FIG10B is a schematic diagram of an inter-frame coding method.
[0032] FIG11 is a schematic diagram of a basic grid division method.
[0033] FIG12 is a schematic diagram of the hierarchical structure of the basic grid.
[0034] FIG13 is a schematic diagram of a coefficient reorganization method of shifted coefficients.
[0035] FIG14 is a schematic diagram of a coding block structure in a two-dimensional image.
[0036] FIG15 is a schematic diagram showing the relationship among the basic grid, the shift coefficient, and the reconstructed grid.
[0037] FIG16 is a diagram showing an example of the connection relationship of points in a three-dimensional grid.
[0038] FIG. 17 is a diagram illustrating an example of a method for determining a quantization parameter.
[0039] FIG. 18 is another diagram illustrating another example of the connection relationship between points in a three-dimensional grid.
[0040] FIG19 is a flow chart of the decoding method provided in an embodiment of the present application.
[0041] Figure 20 is a flow chart of the encoding method provided in an embodiment of the present application.
[0042] FIG21 is a diagram illustrating an example of the prediction and transformation process of the shift coefficients.
[0043] FIG22 is a diagram showing an example of the structure of the LOD layer.
[0044] FIG23 is a schematic diagram of the structure of a decoder provided in an embodiment of the present application.
[0045] FIG24 is a schematic diagram of the structure of a decoder provided in another embodiment of the present application.
[0046] FIG25 is a schematic diagram of the structure of an encoder provided in one embodiment of the present application.
[0047] FIG26 is a schematic diagram of the structure of an encoder provided in another embodiment of the present application. DETAILED DESCRIPTION
[0048] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below with reference to the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present application.
[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0050] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0051] It should also be pointed out that the terms "first\second\third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0052] Generally speaking, 3D animation content uses a keyframe-based representation method, that is, each frame is a static mesh. Static meshes at different times have the same topological structure and different geometric structures. However, the amount of data of 3D dynamic meshes represented based on keyframes is extremely large, so how to effectively store, transmit and draw them has become a problem faced by the development of 3D dynamic meshes. In addition, the spatial scalability of the mesh needs to be supported for different user terminals (computers, notebooks, portable devices, mobile phones); different network bandwidths (broadband, narrowband, wireless) need to support the quality scalability of the mesh. Therefore, 3D dynamic mesh compression is a very critical issue.
[0053] A 3D mesh is the surface of a three-dimensional object composed of multiple polygons in space. Polygons can be composed of vertices and edges. Figure 1A shows a 3D mesh image, and Figure 1B shows a magnified portion of the 3D mesh image. As can be seen from Figures 1A and 1B, a mesh surface is typically composed of multiple closed polygons.
[0054] The pixel distribution of a two-dimensional image is regular, so there's no need to record its geometric information (or position information). However, the random and irregular distribution of mesh vertices in three-dimensional space, as well as the way polygons are constructed, require additional recording. Therefore, for a three-dimensional mesh, it's necessary to record not only the spatial positions of the vertices but also the connectivity information of the polygons within the mesh to fully represent the mesh image. As shown in Figure 2, the same number and positions of vertices can produce completely different surfaces due to different connectivity methods.
[0055] In addition to the above information, since 3D mesh images are usually encoded using existing 2D image / video encoding methods, it is necessary to convert the 3D mesh image from 3D space to 2D space. 3D mesh encoding usually uses UV coordinates to define this conversion process.
[0056] Similar to 2D images, each vertex may have corresponding attribute information. This attribute information is typically an RGB color value, reflecting the object's color. For 3D mesh images, in addition to color, each vertex's attribute information often includes reflectance values, which reflect the object's surface material. The attribute information of a 3D mesh image can be stored in a 2D image, with the mapping from 2D to 3D being specified by UV coordinates.
[0057] Therefore, 3D mesh data typically includes 3D geometric position information (x, y, z), the connectivity of triangular facets within that geometric position information, texture coordinates (u, v), the connectivity of those texture coordinates, and an attribute map. Figure 3A shows a 3D mesh image, and Figure 3B shows the mesh data storage format, which includes 3D geometric position information, texture coordinates, and connectivity information. Figure 3C shows the corresponding attribute map.
[0058] Current 3D dynamic mesh compression methods include space-time prediction methods, which improve compression efficiency by eliminating spatial and temporal correlations; principal component analysis (PCA)-based techniques, which project in the eigenvector space to concentrate energy; and wavelet-based methods, which support spatial and quality scalability.
[0059] Figure 4 is a diagram of the overall framework of grid coding. Figure 5A is a schematic diagram of the two-dimensional curve preprocessing process. Figure 5B is a schematic diagram of the generation of shift coefficients. The preprocessing process of a three-dimensional grid is similar to the preprocessing process of a two-dimensional curve. At the encoding end, it is mainly divided into two parts: preprocessing and encoding. First, the base grid and shift coefficients can be generated through preprocessing. The preprocessing process includes: first, downsampling the original grid (original mesh) to generate a simplified grid (decimated mesh) with a significantly reduced number of vertices, or a base grid / basic grid (base mesh). Then, the simplified grid is subdivided, and the newly generated vertices are inserted on the edges of the simplified grid to obtain a subdivided grid (subdivided mesh) or an initial grid. Finally, for each vertex in the subdivided grid, the point closest to it in the original grid is found, and the displacement value of these two points is calculated. After preprocessing, the simplified grid and shift coefficient are input into the encoder to generate a bitstream.
[0060] V-DMC encoding can be broadly categorized into two main categories: geometric position information encoding and attribute information encoding. Each frame in the basketball_player sequence contains two files: basketball_player_fr0001_qp12_qt12.obj and basketball_player_fr0002.png. basketball_player_fr0001_qp12_qt12.obj contains four types of information: geometric position information (x, y, z), the connectivity of geometric position triangles, texture coordinates (u, v), and the connectivity of texture coordinates. basketball_player_fr0002.png represents the attribute information of the current frame. In current V-DMC encoders, geometric position information is jointly encoded using DRACO and a video codec (AVC, HEVC, or VVC), while texture information is encoded directly using the video codec. Therefore, the following section provides a detailed description of mesh geometric information encoding.
[0061] Geometric information can be divided into the encoding of position information (geometric position information and texture position information) and the encoding of connectivity relationships (geometric position information triangle patch connectivity, texture position information connectivity). Currently, V-DMC coding is mainly divided into two coding test conditions: intra-frame coding and inter-frame coding (low latency, currently no RA test environment).
[0062] Intra-frame geometry coding
[0063] 1. Mesh preprocessing
[0064] a) As shown in Figures 5A and 5B, the 2D curve preprocessing process is used as an example. The original mesh contains a large number of points in its connectivity. Before encoding the mesh geometry information, the mesh geometry information is first quantized or simplified, ultimately obtaining the corresponding simplified mesh as the base mesh.
[0065] b) As shown in FIG6A and FIG6B , the quantization processing of the grid is performed based on the coordinates of the triangle patch. According to the connection relationship between the quantization points, the quantization processing is divided into the following two cases:
[0066] If two vertices belong to the same edge before quantization, then after quantization, all the triangles connected to the two vertices need to be connected together. As shown in Figure 6A, this quantization process involves the disappearance of the previous triangles. If the two vertices do not belong to the same edge, then after quantization, only the boundaries of the two vertices need to be merged. As shown in Figure 6B, this quantization process has no effect on the number of triangles.
[0067] c) In the whole process of mesh quantization based on triangle patch coordinates, the core problem is how to get the best vertex based on the previous vertex coordinates. The current V-DMC will select the best quantization point in the following four modes. Assuming that the vertex distribution before quantization is V1 and V2, and the vertex coordinate after quantization is V', there are the following quantization points: V1, V2, (V1+V2) / 2 and Q - 1 (V1+V2), where Q is the quantization matrix corresponding to the vertex coordinates of V1 and V2. Finally, the optimal quantization point is selected based on the distortion measure D before and after quantization.
[0068] 2. Basic grid encoding
[0069] a) After obtaining the base mesh, the DRACO encoder is used to encode its geometric information. This geometric information primarily includes geometric position information and its connection relationships. The DRACO encoding process is as follows: first, the connection relationships are encoded. Then, the geometric position information of the points is encoded based on the connection relationships. Finally, the texture position information is encoded based on the connection relationships and geometric position information.
[0070] b) Coding of connection relationships. DRACO uses the "Edgebreaker Coding" scheme to encode the connection relationships of the grid. Specifically, as shown in Figure 7A:
[0071] Before encoding the connectivity of the mesh, the vertices of the mesh are divided into five types, CLRSE, where each symbol represents the following meaning:
[0072] iC: None of the triangles connected to the current vertex have completed encoding;
[0073] ii.L: The left triangle connected to the current vertex is encoded;
[0074] iii.R: The right triangle connected to the current vertex completes the encoding;
[0075] iv.S: The triangles on the left and right sides of the current vertex have not been encoded yet;
[0076] vE: The left and right triangles connected to the current vertex have been encoded.
[0077] Finally, the type of each vertex and the processing order of the vertices are encoded in a certain order, and the decoding end restores the geometric connection relationship of the mesh according to the processing order and type of the vertices.
[0078] c) Geometric Position Information Encoding. After completing the encoding of the vertex connectivity, the geometric position information of each vertex is predictively encoded based on the vertex connectivity. The idea behind predictive encoding is the "parallelogram algorithm," as shown in Figure 7B.
[0079] Use the three vertices adjacent to the current point to be encoded: the left vertex, the right vertex, and the opposite vertex for a simple linear fit: pred pos =(left+right)-opposite
[0080] d) After completing the point connection relationship and geometric position information, the texture coordinates are predictively encoded based on the decoding and reconstruction of these two, as shown in Figure 7C. Similarly, assuming the current vertex is C, the left and right vertices of the current point can be obtained based on the point connection relationship. Then, the texture coordinates of the left and right vertices are used to predict the texture coordinates of the current vertex C.
[0081] 3. Displacement Coding
[0082] a) First, after the encoding and reconstruction of the base mesh is completed, a certain partitioning algorithm will be used to partition the base mesh to obtain the initial reconstructed mesh. For details, please refer to the curve corresponding to the subdivided mesh in Figure 5A. The curve corresponding to the simplified mesh can be obtained by simple linear interpolation. The coordinates of the newly inserted point are obtained by linear interpolation based on the two vertices on the current boundary:
[0083] b) Next, the point-to-point error Delta between the subdivided mesh and the original mesh is calculated. This error Delta can be expressed as a point-to-point error in the world coordinate system. Finally, the Displacement value for each point is calculated using the point-to-point error Delta and the point's normal vector Norm, as shown in Figure 8A.
[0084] The specific calculation method is as follows: Displacement = Delta × Norm
[0085] c) After calculating the displacement of each point, the spatial domain residual coefficient can be transformed into the frequency domain using the lifting wavelet transform to obtain the corresponding frequency domain residual coefficient.
[0086] d) Finally, the Packing algorithm is used to map the frequency domain residual coefficients of each point into a two-dimensional image in a certain order. The current V-DMC is arranged in the order of the Morton code, as shown in Figure 8B.
[0087] e) Finally, the traditional Video Codec can be used to encode the two-dimensional image.
[0088] 4. Recoloring
[0089] Recoloring is an algorithm on the encoder side. After the reconstruction of the encoder's geometric information is completed, the original geometric information, the original texture attribute information, and the reconstructed mesh geometric information are used to recolor the texture attribute information of the reconstructed mesh.
[0090] Inter-frame geometry information coding
[0091] a) Similar to intra-frame geometric information encoding, including geometric connection relationship and geometric position information encoding. However, it should be noted that inter-frame geometric position information encoding only needs to encode the geometric position information (x, y, z) of the current base mesh, and does not need to encode the connection relationship and texture position information (u, v). The specific reason is as follows: If the current frame can use inter-frame encoding, then the base mesh of the reference frame of the current frame will be used at the encoder to obtain the mesh information of the current frame. Therefore, the current frame and the reference frame have the same connection relationship and UV texture coordinates, only the geometric position information is different.
[0092] b) Based on a), we know that there is only error in the geometric position information between the current frame and the reference frame, so the current V-DMC performs predictive coding on the geometric position information of the current frame.
[0093] As shown in Figure 9, the black vertex in the middle is the vertex to be encoded / decoded. The current point is used in the reference frame to obtain the corresponding prediction point (similar to the same-position block in video encoding). Then, the neighborhood point of the current point (the MV of the coded vertex) is used to predict the MV of the current point. As shown below, assuming that the coordinates of the current point are pos and the coordinates of the corresponding same-position point are Pred_pos, the MV of the current point is calculated as: MV = Pos-Pred pos
[0094] There are two predictive coding modes in the current V-DMC:
[0095] i. Directly encode the MV of the current point
[0096] ii. Use the neighborhood to predict the MV of the current point
[0097] At the encoding end, a rate-distortion optimization algorithm is used to obtain the optimal coding mode for each coding group (CG). The current V-DMC sets the maximum number of points for each CG to 16.
[0098] Coding of texture attribute information: The current V-DMC encodes texture attribute information directly using Video-Codec, such as AVC, HEVC, VVC or VV-Enc.
[0099] Figure 10A is a schematic diagram of intra-frame coding. As shown in Figure 10A, in the intra-frame encoder, a common static mesh encoder can be used to encode the simplified mesh to generate a corresponding bitstream (compressed base mesh bitstream). Next, the reconstructed simplified mesh is used to update the displacement coefficient. The updated displacement coefficient is subjected to wavelet transform and quantization to obtain the displacement coefficient. After image packing, high efficiency video coding (HEVC) is used for encoding to generate a bitstream of displacement coefficients (compressed displacements bitstream). For attribute map encoding, the feature map is first transformed (texture transfer) according to the difference between the reconstructed geometric information and the original geometric information, and then padded and color space converted and encoded using a video encoder (video coding) to form a compressed attribute bitstream.
[0100] Figure 10B is a schematic diagram of inter-frame coding. As shown in Figure 10B, the inter-frame encoder and the intra-frame encoder process are roughly the same, but the inter-frame encoder does not directly encode the simplified grid. Instead, it encodes the motion vector between the simplified grid of the current frame and the simplified grid of the reference frame (motion encoder) and generates a corresponding motion vector bitstream (compressed motion bitstream).
[0101] Common test conditions for MPEG DMC
[0102] 1) There are two test conditions for MPEG DMC:
[0103] Condition 1: all intra geometry is lossy and attributes are lossy;
[0104] Condition 2: Random access is lossy in geometry and attributes;
[0105] 2) Common test sequences include Cat1-A, Cat1-B and Cat1-C, a total of five categories, all of which contain geometric and color attribute information.
[0106] Next, we will introduce the Displacement coding of V-DMC in more detail.
[0107] Encoding algorithm:
[0108] On the encoder side, the base mesh is first iteratively partitioned according to a specific algorithm to obtain the geometric position information of the points in the subdivided mesh. The specific partitioning algorithm is consistent with the background description, that is, linear interpolation is performed using the vertices on each boundary to obtain the geometric position information of the points in the subdivided mesh. Assuming that the entire partitioning process is iterated N times, these N iterative partitioning processes can be used to perform LOD partitioning, resulting in multiple levels, as shown in Figure 11.
[0109] After obtaining the geometric position information of the subdivided grid midpoint, the geometric position information of the subdivided grid midpoint and the geometric position information of the original grid midpoint can be used to calculate the error to obtain the shift coefficient of the subdivided grid midpoint. The LOD hierarchy of the shift coefficient is shown in Figure 12.
[0110] Secondly, a lifting transform is performed based on the LOD spatial structure. The lifting transform includes two steps: prediction and update. The prediction algorithm is as follows:
[0111] The update algorithm is as follows:
[0112] Finally, the transformed coefficients are quantized and the quantized coefficients are reorganized, as shown in FIG13 .
[0113] As shown in Figure 14, V-DMC can reorganize coefficients on a block-by-block basis, where each block can be, for example, 16×16 in size. The coefficients within each block can be arranged using a Morton code to produce the corresponding two-dimensional image. After completing these operations, the two-dimensional image can be encoded using a Video Codec.
[0114] Decoding algorithm:
[0115] First, the Video Codec is used to decode and reconstruct the 2D image to obtain the corresponding 2D image. Second, the transform coefficients corresponding to each point are recovered by coefficient reorganization. Finally, the inverse of the lifting transform is used to recover the shift coefficients of each point. After obtaining the shift coefficients for each point, as shown in Figure 15, the geometric position information of the base grid and the shift coefficients can be used to reconstruct the geometric position information corresponding to the current grid.
[0116] The foregoing article describes in detail the encoding and decoding process of the geometric information of the points in the three-dimensional grid. How to improve the encoding and decoding efficiency of geometric information is a technical problem that this application needs to solve. After research, it was found that the importance of different points in the three-dimensional grid may be different. The difference in importance between different points can be reflected in the different connection relationships between different points, or the different number of neighborhood points of different points. When encoding and decoding the shift coefficients of points in the three-dimensional grid, the related technology did not fully consider the differences in importance of different points, thereby limiting the encoding and decoding efficiency of geometric information.
[0117] For example, in the V-DMC scheme provided by related art, the base grid is first divided to determine the geometric information of the subdivided grid. Next, the difference between the original geometric information of the 3D grid and the geometric information of the subdivided grid is calculated to determine the shift coefficients of the points in the 3D grid. Then, using the lifting transform, the shift coefficients are transformed from the spatial domain to the frequency domain, thereby completing the encoding of the geometric information of the V-DMC. The lifting transform scheme provided by related art includes two processes: prediction and transformation (or update). The prediction scheme uses the mean prediction of two neighboring points that are collinear with the current point, while the transformation scheme uses the residual coefficients of the current point's neighboring point set (the points predicted based on the current point) to update the shift coefficient of the current point. Based on this encoding scheme, the prediction scheme removes spatial redundancy between adjacent points, and the transformation scheme further removes frequency redundancy between adjacent points, which can improve the encoding efficiency of the geometric information of the 3D grid to a certain extent. However, in the prediction scheme provided by related art, the prediction weight is fixed at 0.5, and the update weight is usually fixed at 0.125 (indicated by the syntax element in the adaptation parameter set (APS)). Using fixed weights for prediction and / or transformation does not fully consider the differences in importance between different points, thereby limiting the coding efficiency of the geometric information of V-DMC.
[0118] The connection relationship between different points in the grid and the surrounding points may be different, so the number of neighborhood points of the same point may also be different. If the number of neighborhood points of two points is different, the number of points predicted based on the two points may be different, so the importance of the two points in the grid may be different. Taking Figure 16 as an example, the number of neighborhood points collinear with point 1 is 6, which means that the shift coefficients of the 6 points are all predicted based on the shift coefficient of point 1, so the residual coefficients of the 6 points can be used to update the shift coefficient of point 1. Similarly, the number of neighborhood points collinear with point 2 is 5, which means that the shift coefficients of the 5 points are all predicted based on the shift coefficient of point 2, so the residual coefficients of the 5 points can be used to update the shift coefficient of point 1. By comparing point 1 and point 2, it can be seen that the importance of the two in the grid is different. If a fixed prediction weight and / or update weight is used, it is not reasonable. Therefore, the embodiment of the present application introduces the weights corresponding to the points, and performs prediction and / or updating based on the weights corresponding to the points (the weights corresponding to the points can be introduced in the prediction process, the weights corresponding to the points can be introduced in the update process, or the weights corresponding to the points can be introduced in both the prediction and update processes). Prediction and / or updating based on the weights corresponding to the points helps to improve the coding efficiency of the geometric information of V-DMC.
[0119] The above describes the problems existing in the prediction and / or transformation process of the shift coefficients and possible improvements. The following describes the problems existing in the quantization process of the shift coefficients and possible improvements.
[0120] The V-DMC coding scheme provided by the related art needs to quantize the transformation coefficients after transforming the shift coefficients of each point. V-DMC performs layer-by-layer quantization based on the spatial structure of LOD. As shown in Figure 17, the quantization parameter QP of each layer Lvl It can be calculated as follows: QP Lvl =liftingQP×Scale Lvl-1
[0121] Among them, Lvl represents the index of each layer LOD, liftingQP is the initial quantization parameter, and Scale is the adjustment parameter of each layer LOD based on liftingQP.
[0122] After obtaining the quantization parameters of each layer, the following method can be used to adaptively quantize the shift coefficient of each point:
[0123] Among them, bitDepth represents the effective bit depth of the shift coefficient in the current geometric coding, Res represents the transform coefficient of the current point to be coded, and QP Lvl Represents the quantization parameter of the current LOD layer.
[0124] At the decoding end, the quantization coefficient of each point is recovered by entropy decoding analysis, and the quantization parameter QP of each LOD layer is obtained based on the same method as the encoding end. Lvl The shift coefficient of the current point is dequantized based on the quantization parameter. The specific calculation method is as follows:
[0125] In summary, the initial quantization parameter liftingQP and the quantization adjustment parameter Scale of different LOD layers are used to obtain the quantization parameter QP of different LOD layers. Lvl , to a certain extent, the coding efficiency of the geometric information of the three-dimensional grid can be improved. This is because the shift coefficient is adaptively transformed and coded using a lifting transformation based on the LOD spatial structure, resulting in the shift coefficient of the LOD high layer being a low-frequency coefficient, while the shift coefficient of the LOD low layer is a high-frequency coefficient. Therefore, based on the above quantization scheme, adaptive quantization coding can be implemented for different LOD layers. However, for points on the same layer, since the spatial connection relationship of each point is inconsistent, the importance of each point in the lifting transformation process will be inconsistent. However, the above quantization coding scheme does not fully take this into account, which limits the coding and decoding efficiency of the geometric information of the three-dimensional grid to a certain extent.
[0126] As shown in Figure 18, points P1, P2, and P3 are at the same level of depth (LOD). P1's neighborhood point set (i.e., the points predicted based on P1) includes six neighbors, P2's neighborhood point set (i.e., the points predicted based on P2) includes six neighbors, and P3's neighborhood point set (i.e., the points predicted based on P3) includes only five neighbors. For these three points, the importance is ranked as follows: P1 = P2 > P3. The higher the importance, the greater the impact of the point's shift coefficient on the reconstruction quality of the entire mesh. Therefore, when quantizing highly important shift coefficients, the quantization level of these shift coefficients can be minimized (i.e., the quantization parameter QP is reduced); while when quantizing less important shift coefficients, the quantization level can be appropriately increased (i.e., the quantization parameter QP is increased). This ensures the final reconstruction quality and improves the encoding bitrate of the geometric information. This quantization coding scheme fully considers the importance of different points, thereby improving the encoding and decoding efficiency of the geometric information in V-DMC.
[0127] 19 , the following will first provide a detailed example of how the embodiment of the present application is implemented in the decoding process from the perspective of the decoding end.
[0128] Figure 19 is a flowchart of a decoding method provided in an embodiment of the present application. The method of Figure 19 can be performed by a decoder. The decoder can be a decoder that supports V-DMC.
[0129] 19 , in step S1910 , the code stream is parsed to determine the quantization coefficients of the points in the three-dimensional grid. It should be understood that the quantization coefficients of the points mentioned in the embodiments of the present application refer to the quantization coefficients corresponding to the shift coefficients of the points.
[0130] In step S1920 , the quantized coefficients of the points in the three-dimensional grid are dequantized to determine the dequantized coefficients of the points in the three-dimensional grid.
[0131] In step S1930, the inverse quantization coefficients of the points in the first layer are inversely transformed to determine the shift coefficients of the points in the first layer. The points in the three-dimensional grid belong to multiple LOD layers, and the first layer mentioned in step S1930 can be any LOD layer among the multiple LOD layers.
[0132] In step S1940, the shift coefficients of the points in the second layer are predicted based on the shift coefficients of the points in the first layer. The second layer is an LOD layer adjacent to the first layer among the multiple LOD layers, and the second layer is a layer below the first layer. Steps S1930 and S1940 may be two sub-processes of the inverse transformation of the lifting transformation.
[0133] FIG19 illustrates the decoding process of the first and second layers as an example. Each LOD layer in the multiple LOD layers can be processed in a similar manner to obtain the shift coefficients of the points in the three-dimensional grid. Then, based on the shift coefficients of the points in the three-dimensional grid and the geometric information of the points in the subdivided grid (or the initial reconstructed grid) (the geometric information of the subdivided grid can also be called the initial geometric information of the three-dimensional grid), the reconstructed geometric information of the points in the three-dimensional grid can be determined. For example, the geometric information of the corresponding points in the subdivided grid can be shifted according to the shift coefficients of the points in the three-dimensional grid to determine the reconstructed geometric information of the points in the three-dimensional grid. The geometric information of the points in the subdivided grid can be obtained based on the division of the base grid. For example, the code stream can be parsed first to determine the geometric information of the points in the base grid; then, the base grid can be divided according to the geometric information of the points in the base grid to determine the geometric information of the points in the subdivided grid.
[0134] According to Figure 19, the complete decoding operation includes inverse quantization operation, inverse transformation operation, prediction operation, etc. In order to improve the decoding efficiency of the geometric information of the three-dimensional grid, one or more of the inverse quantization operation, inverse transformation operation, and prediction operation can be optimized based on the weights corresponding to the points in the three-dimensional grid (or the spatial connection relationship of the points). The following examples illustrate the implementation methods provided by the embodiments of the present application from the three perspectives of inverse quantization operation, inverse transformation operation, and prediction operation. It should be understood that the embodiments of the present application can optimize only one of inverse quantization, inverse transformation, and prediction; or, the embodiments of the present application can optimize any two of inverse quantization, inverse transformation, and prediction, thereby further improving the decoding efficiency of the geometric information of the three-dimensional grid; or, the embodiments of the present application can optimize inverse quantization, inverse transformation, and prediction at the same time, thereby further improving the decoding efficiency of the geometric information of the three-dimensional grid.
[0135] For ease of understanding, the following describes the inverse transform operation for points in the three-dimensional grid from the perspective of the first point to be decoded. The first point to be decoded can be any point to be decoded in the first layer. The neighboring points of the first point to be decoded are called first neighboring points. The first neighboring points can be points in the lower level of detail (LOD) layer of the first point to be decoded. As an example, the first neighboring points are points predicted based on the first point to be decoded.
[0136] Step S1930 in Figure 19, i.e., inversely transforming the inverse quantization coefficients of the points in the first layer, may include: inversely transforming the inverse quantization coefficients of the first point to be decoded in the first layer according to the inverse quantization coefficients of the first neighborhood point and the updated weights. In the related art, the updated weight adopts a fixed value (i.e., 0.125). Unlike the related art, in the embodiment of the present application, the updated weight is determined based on the weight corresponding to the first neighborhood point. The updated weight is determined based on the weight corresponding to the neighborhood point, which can make the points with higher importance have a greater impact on the inverse transformation results of the points to be decoded, which will make the updated results more accurate and help improve the decoding efficiency of the geometric information of the points in the three-dimensional grid.
[0137] The weight corresponding to the first neighborhood point can be determined based on the spatial connectivity of the first neighborhood point in the three-dimensional grid. For example, the weight corresponding to the first neighborhood point can be related to the number of neighboring points of the first neighborhood point (hereinafter referred to as second neighborhood points). Determining the weight corresponding to a point based on the number of neighboring points fully utilizes the connectivity information of the points in the three-dimensional grid, thereby making the determined point weight more reasonable.
[0138] The embodiment of the present application does not specifically limit the definition of the second neighborhood point. For example, the second neighborhood point can be a point in the same LOD layer as the first neighborhood point. Alternatively, the second neighborhood point can be a point in the next lower LOD layer than the first neighborhood point. Alternatively, the second neighborhood point can include points in both of the aforementioned LOD layers. As an example, the second neighborhood point is a point predicted based on the first neighborhood point.
[0139] In some implementations, the weight corresponding to the first neighborhood point can be equal to the number of the second neighborhood points. For example, if the first neighborhood points include point 1 and point 2, point 1 includes 5 neighborhood points, and point 2 includes 6 neighborhood points, then the weight of point 1 can be 5, and the weight of point 2 can be 6.
[0140] In some implementations, the weight corresponding to the first neighborhood point can be determined based on the product of the number of second neighborhood points and the first predicted weight. For example, the weight corresponding to the first neighborhood point can be equal to the product of the number of second neighborhood points and the first predicted weight. The first predicted weight mentioned here can be a predefined fixed value. For example, the value of the first predicted weight can be 0.5. When determining the weight corresponding to the point, taking the predicted weight into consideration can make the determined weight more reasonable.
[0141] For example, the weight corresponding to the first neighborhood point may be determined based on the following formula:
[0142] In the above formula, weight represents the weight corresponding to any point in the first neighborhood, and predWeight is the first predicted weight (which can be 0.5). P represents the second neighborhood, and weight[i] represents the initial weight corresponding to the i-th point in the second neighborhood. The initial weight corresponding to the i-th point can be set to 1. For example, if the first neighborhood includes points 1 and 2, point 1 includes 5 neighboring points, and point 2 includes 6 neighboring points, then based on the above formula, the weight of point 1 can be 2.5, and the weight of point 2 can be 3 (the value of the first predicted weight is 0.5).
[0143] In some implementations, the weight corresponding to the first neighborhood point can be determined based on the weight corresponding to the LOD layer where the first neighborhood point is located. For example, the weight corresponding to the first neighborhood point can be equal to the weight corresponding to the LOD layer where the first neighborhood point is located. The weight corresponding to the LOD layer can be understood as a level weight. The level weight of an LOD layer can be determined, for example, based on the ratio of the number of points in the LOD layer to the total number of points to be decoded.
[0144] As mentioned above, the update weight can be determined based on the weight corresponding to the first neighborhood point. As a possible implementation method, the update weight can be equal to the weight corresponding to the first neighborhood point. In another possible implementation method, the update weight can be determined based on the weight corresponding to the first neighborhood point and the first prediction weight (the first prediction weight can be a predefined fixed value, for example, the value of the first prediction weight can be 0.5). For example, the update weight can be determined based on the product of the weight corresponding to the first neighborhood point and the first prediction weight, or the update weight can be equal to the product of the weight corresponding to the first neighborhood point and the first prediction weight. Taking the prediction weight into consideration when determining the update weight can make the weight distribution of the prediction and update process more reasonable.
[0145] Exemplarily, the update weight may be determined based on the following formula: updateWeight i =predWeight×weight i
[0146] In the above formula, updateWeight i Indicates the updated weight corresponding to the i-th point in the first neighborhood point. predWeight indicates the first prediction weight, and the value of the first prediction weight can be 0.5. weight i The weight corresponding to the first neighborhood point can be determined using the implementation method described above.
[0147] The above describes in detail the weight corresponding to the first neighborhood point and the method for determining the updated weight. Based on the determination of the weight corresponding to the first neighborhood point and the updated weight, the inverse transformation of the inverse quantization coefficient of the first point to be decoded can be performed according to the following formula:
[0148] In the above formula, Signal represents the shift coefficient of the first point to be decoded. updateWeight i Indicates the update weight corresponding to the i-th point in the first neighborhood. updateWeight i The calculation method of can be found in the previous text. P represents the point set formed by the first neighborhood point, and i represents the i-th point in the first neighborhood point. i Represents the inverse quantized coefficient of the first neighborhood point.
[0149] In addition to the solution of performing inverse transformation based on the weight corresponding to the point, the embodiment of the present application can also apply other inverse transformation solutions. For example, the inverse transformation can be performed based on the adjustment parameter (Scale) of the updated weight, as shown in the following formula:
[0150] In the above formula, updateWeight represents the update weight (can be a fixed value, such as 0.125). n-i-1 is the adjustment parameter for updating the weight, n represents the number of LOD layers, i represents the index of the LOD layer. P represents the point set formed by the first neighborhood point, and i represents the i-th point in the first neighborhood point. i Represents the inverse quantized coefficient of the first neighborhood point.
[0151] The above describes in detail the inverse transformation operation of the points in the three-dimensional grid. The following describes the prediction operation of the points in the three-dimensional grid from the perspective of the second point to be decoded. The second point to be decoded can be any point to be decoded in the second layer. The neighborhood point of the second point to be decoded is called the third neighborhood point. The embodiment of the present application does not specifically limit the definition of the third neighborhood point. For example, the third neighborhood point can be a point in the same LOD layer as the second point to be decoded. Alternatively, the third neighborhood point can be a point in the next layer of the LOD layer where the second point to be decoded is located. Alternatively, the third neighborhood point can include points in the above two LOD layers at the same time. As an example, the third neighborhood point can be a point in the first layer (the layer above the LOD layer where the second point to be decoded is located) that is collinear with the second point to be decoded.
[0152] Step S1940 in Figure 19, i.e. predicting the shift coefficients of the points in the second layer based on the shift coefficients of the points in the first layer, may include: predicting the second point to be decoded based on the shift coefficients of the third neighborhood points and the second prediction weight. In the related art, the prediction weight adopts a fixed value (i.e. 0.5). Different from the related art, in the embodiment of the present application, the second prediction weight is determined based on the weight corresponding to the third neighborhood point. The greater the weight corresponding to a point in the third neighborhood point, the greater the influence of the point on the prediction result of the second point to be decoded. The prediction weight is determined based on the weight corresponding to the neighborhood point, which can make the points with higher importance have a greater influence on the prediction result of the point to be decoded, which will make the prediction process more reasonable and help improve the decoding efficiency of the geometric information of the points in the three-dimensional grid.
[0153] The weight corresponding to the third neighboring point can be determined based on the spatial connectivity of the third neighboring point in the three-dimensional grid. For example, the weight corresponding to the third neighboring point can be related to the number of neighboring points of the third neighboring point (hereinafter referred to as the fourth neighboring points). Determining the weight corresponding to a point based on the number of neighboring points fully utilizes the connectivity information of the points in the three-dimensional grid, which can make the determined point weight more reasonable.
[0154] The embodiment of the present application does not specifically limit the definition of the fourth neighborhood point. For example, the fourth neighborhood point can be a point in the same LOD layer as the third neighborhood point. Alternatively, the fourth neighborhood point can be a point in the LOD layer below the LOD layer of the third neighborhood point. Alternatively, the fourth neighborhood point can include points in both of the aforementioned LOD layers. As an example, the fourth neighborhood point is a point predicted based on the third neighborhood point.
[0155] In some implementations, the weight corresponding to the third neighborhood point can be equal to the number of the fourth neighborhood points. For example, if the third neighborhood points include point 1 and point 2, point 1 includes 5 neighborhood points, and point 2 includes 6 neighborhood points, then the weight of point 1 can be 5, and the weight of point 2 can be 6.
[0156] In some implementations, the weight corresponding to the third neighborhood point can be determined based on the product of the number of fourth neighborhood points and the first predicted weight. For example, the weight corresponding to the third neighborhood point can be equal to the product of the number of fourth neighborhood points and the first predicted weight. The first predicted weight mentioned here can be a predefined fixed value. For example, the value of the first predicted weight can be 0.5. Considering the predicted weight when determining the weight corresponding to the point can make the determined weight more reasonable.
[0157] For example, the weight corresponding to the third neighborhood point can be determined based on the following formula:
[0158] In the above formula, weight represents the weight corresponding to any point in the third neighborhood, and predWeight is the first predicted weight (which can be 0.5). P represents the fourth neighborhood, and weight[i] represents the initial weight corresponding to the i-th point in the fourth neighborhood. The initial weight corresponding to the i-th point can be set to 1. For example, if the third neighborhood includes points 1 and 2, point 1 includes 5 neighboring points, and point 2 includes 6 neighboring points, then based on the above formula, the weight of point 1 can be 2.5, and the weight of point 2 can be 3 (the first predicted weight is 0.5).
[0159] In some implementations, the weight corresponding to the third neighborhood point can be determined based on the weight corresponding to the LOD layer where the third neighborhood point is located. For example, the weight corresponding to the third neighborhood point can be equal to the weight corresponding to the LOD layer where the third neighborhood point is located. The weight corresponding to the LOD layer can be understood as a level weight. The level weight of an LOD layer can be determined, for example, based on the ratio of the number of points in the LOD layer to the total number of points to be decoded.
[0160] As previously mentioned, the second prediction weight can be determined based on the weight corresponding to the third neighborhood point. For example, the second prediction weight can be equal to the weight corresponding to the third neighborhood point. Alternatively, the second prediction weight can be based on the product of the weight corresponding to the third neighborhood point and a parameter.
[0161] As mentioned above, the second point to be decoded can be predicted based on the shift coefficient of the third neighborhood point and the second prediction weight. A possible prediction method for the second point to be decoded is given below.
[0162] Exemplarily, the second point to be decoded may be predicted based on the following formula:
[0163] In the above formula, Signal represents the prediction result of the second point to be decoded, Neigh1 and Neigh2 represent the shift coefficients of the two neighboring points collinear with the second point to be decoded. Neigh1 and Neigh2 are located in the LOD layer above the second point to be decoded. predWeight1 and predWeight2 represent the prediction weights corresponding to Neigh1 and Neigh2, respectively (i.e., the second prediction weights mentioned above).
[0164] predWeight1 and predWeight2 can be calculated using the following formula: predWeight1 = weight1 predWeight2 = weight2
[0165] sumWeight can be calculated based on the following formula: sumWeight=predWeight1+predWeight2
[0166] The above describes in detail the inverse transformation and prediction operations for points in a 3D grid. The following describes the inverse quantization operations for points in a 3D grid from the perspective of a third point to be decoded. The third point to be decoded can be any point to be decoded in the 3D grid.
[0167] Step S1920 in Figure 19, i.e., dequantizing the quantization coefficients of the points in the three-dimensional grid, may include: dequantizing the quantization coefficients of the third point to be decoded according to the weight corresponding to the third point to be decoded. In the embodiment of the present application, when dequantizing the points to be decoded in the three-dimensional grid, the weight corresponding to the point to be decoded is taken into account. For example, a suitable quantization parameter may be selected according to the weight corresponding to the point to be decoded; or, the dequantization result may be processed differently according to the weight corresponding to the point to be decoded. Considering the weight of the point to be decoded in the dequantization process can make the dequantization result more reasonable, which helps to improve the decoding efficiency of the geometric information of the points in the three-dimensional grid.
[0168] The weight corresponding to the third point to be decoded can be determined based on the spatial connectivity of the third point to be decoded in the three-dimensional grid. For example, the weight corresponding to the third point to be decoded can be related to the number of neighboring points (hereinafter referred to as fifth neighboring points) of the third point to be decoded. Determining the weight corresponding to a point based on the number of neighboring points fully utilizes the connectivity information of the points in the three-dimensional grid, making the determined point weight more reasonable.
[0169] The present embodiment does not specifically limit the definition of the fifth neighboring point. For example, the fifth neighboring point can be a point in the same LOD layer as the third point to be decoded. Alternatively, the fifth neighboring point can be a point in the LOD layer below the LOD layer of the third point to be decoded. Alternatively, the fifth neighboring point can include points in both of the aforementioned LOD layers. As an example, the fifth neighboring point is a point predicted based on the third point to be decoded.
[0170] In some implementations, the weight corresponding to the third to-be-decoded point may be equal to the number of the fifth neighboring points. For example, if the third to-be-decoded point includes point 1 and point 2, point 1 includes 5 neighboring points, and point 2 includes 6 neighboring points, then the weight of point 1 may be 5, and the weight of point 2 may be 6.
[0171] In some implementations, the weight corresponding to the third to-be-decoded point can be determined based on the product of the number of fifth neighboring points and the first predicted weight. For example, the weight corresponding to the third to-be-decoded point can be equal to the product of the number of fifth neighboring points and the first predicted weight. The first predicted weight mentioned herein can be a predefined fixed value. For example, the value of the first predicted weight can be 0.5. Considering the predicted weight when determining the weight corresponding to a point can make the determined weight more reasonable.
[0172] Exemplarily, the weight corresponding to the third to-be-decoded point may be determined based on the following formula:
[0173] In the above formula, weight represents the weight corresponding to any point in the third to-be-decoded point, and predWeight is the first predicted weight (which can be 0.5). P represents the fifth neighborhood point, and weight[i] represents the initial weight corresponding to the i-th point in the fifth neighborhood point. The initial weight corresponding to the i-th point can be set to 1. For example, if the third to-be-decoded point includes point 1 and point 2, point 1 includes 5 neighborhood points, and point 2 includes 6 neighborhood points, then based on the above formula, the weight of point 1 can be 2.5, and the weight of point 2 can be 3 (the first predicted weight is 0.5).
[0174] As mentioned above, the quantization coefficient of the third point to be decoded can be dequantized according to the weight corresponding to the third point to be decoded. A possible implementation method is given below.
[0175] For example, in some implementations, the quantization coefficient of the third point to be decoded can be inversely quantized according to the first quantization parameter. The first quantization parameter can be determined based on the weight corresponding to the third point to be decoded. That is, the quantization parameter can be adaptively set for the point to be decoded according to the weight corresponding to the point to be decoded. For example, if the weight corresponding to the third point to be decoded is large, it means that the importance of the third point to be decoded is high. For points to be decoded with a high degree of importance, the value of the first quantization parameter can be small (indicating that the degree of quantization of the point to be decoded is low). For another example, if the weight corresponding to the third point to be decoded is small, it means that the importance of the third point to be decoded is low. For points to be decoded with a low degree of importance, the value of the first quantization parameter can be large (indicating that the degree of quantization of the point to be decoded is high). According to the design method of the above-mentioned quantization parameters, the reconstruction quality can be guaranteed, thereby improving the decoding efficiency of the geometric information of the three-dimensional grid.
[0176] The first quantization parameter can be directly determined based on the weight corresponding to the third point to be decoded. For example, a mapping relationship between the weight corresponding to the point to be decoded and the quantization parameter can be pre-established, and the first quantization parameter can be determined based on the mapping relationship. Alternatively, the first quantization parameter can also be determined based on the weight corresponding to the third point to be decoded and the second quantization parameter. The second quantization parameter mentioned here can be the quantization parameter corresponding to the target LOD layer (i.e., the LOD layer to which the third point to be decoded belongs). The second quantization parameter can be determined based on the initial quantization parameter and the quantization adjustment parameter.
[0177] Exemplarily, the second quantization parameter may be determined based on the following formula: QP Lvl =liftingQP×Scale Lvl
[0178] In the above formula, Lvl represents the index of the target LOD layer, QP Lvl is the second quantization parameter, liftingQP is the initial quantization parameter, and Scale is the quantization adjustment parameter.
[0179] A more specific example of inverse quantization of the third point to be decoded is given below.
[0180] For example, the third point to be decoded may be inversely quantized based on the following formula:
[0181] In the above formula, InvRes i Represents the inverse quantization result of the third point to be decoded, Res i Indicates the quantized coefficient of the third point to be decoded, bitDepth represents the effective bit depth of the shift coefficient, QP Lvl Indicates the quantization parameter corresponding to the target LOD layer (the LOD layer where the third point to be decoded is located), weight irepresents the weight corresponding to the third to-be-decoded point (the method for determining the weight corresponding to the third to-be-decoded point can be found in the above text and will not be described in detail here).
[0182] The decoding method provided by the embodiment of the present application is described in detail above in conjunction with Figure 19. The encoding method provided by the embodiment of the present application is described in detail below in conjunction with Figure 20.
[0183] Figure 20 is a flow chart of an encoding method provided in an embodiment of the present application. The method of Figure 20 may be performed by an encoder. The encoder may be an encoder that supports V-DMC.
[0184] 20 , in step S2010 , LOD division is performed on points in a three-dimensional grid to determine multiple LOD layers, including a first layer and a second layer adjacent to each other, where the first layer is the upper layer of the second layer.
[0185] In step S2020, the shift coefficients of the points in the second layer are predicted to determine the residual coefficients of the points in the second layer. It should be understood that the residual coefficients of the points mentioned in the embodiments of the present application refer to the residual coefficients corresponding to the shift coefficients of the points, or the predicted residuals of the shift coefficients.
[0186] In step S2030, the shift coefficients of the points in the first layer are transformed according to the residual coefficients of the points in the second layer. Steps S2020 and S2030 may belong to two sub-processes of the lifting transformation.
[0187] It should be understood that the shift coefficients of points in the 3D mesh can be determined based on the difference between the geometric information of the points in the subdivided mesh and the original geometric information of the 3D mesh. For example, the base mesh can be divided based on the geometric information of the points in the base mesh, and the geometric information of the points in the subdivided mesh can be determined. Then, the shift coefficients of the points in the 3D mesh can be determined based on the geometric information of the points in the subdivided mesh and the original geometric information of the 3D mesh.
[0188] FIG20 illustrates the encoding process of the first and second layers as an example. Each of the multiple LOD layers can be processed in a similar manner to obtain transform coefficients for the points in the three-dimensional grid. Furthermore, in some implementations, FIG20 may further include step S2040, quantizing the transform coefficients for the points in the three-dimensional grid. Quantization of the transform coefficients for the points in the three-dimensional grid may also be performed layer by layer.
[0189] As can be seen from Figure 20, the complete encoding operation includes prediction operation, transformation operation and quantization operation. In order to improve the encoding efficiency of the geometric information of the three-dimensional grid, one or more of the prediction operation, transformation operation and quantization operation can be optimized based on the weights of the points in the three-dimensional grid (or the spatial connection relationship of the points). The following examples illustrate the optimization scheme provided by the embodiment of the present application from the three perspectives of prediction operation, transformation operation and quantization operation. It should be understood that the embodiment of the present application can optimize only one of prediction, transformation and quantization; or, the embodiment of the present application can optimize any two of prediction, transformation and quantization, thereby further improving the encoding efficiency of the geometric information of the three-dimensional grid; or, the embodiment of the present application can optimize prediction, transformation and quantization at the same time, thereby further improving the encoding efficiency of the geometric information of the three-dimensional grid.
[0190] The following describes the transformation operations for points in a three-dimensional grid from the perspective of a first point to be encoded. The first point to be encoded can be any point to be encoded in the first layer. The neighboring points of the first point to be encoded are referred to as first neighboring points. The first neighboring points can be points in the lower level of detail (LOD) layer of the first point to be encoded. As an example, the first neighboring points are points predicted based on the first point to be encoded.
[0191] Step S2030 in Figure 20, i.e., transforming the shift coefficients of the points in the first layer according to the residual coefficients of the points in the second layer, may include: transforming the shift coefficients of the first point to be encoded according to the residual coefficients of the first neighborhood points and the updated weights. In the related art, the updated weights use a fixed value (i.e., 0.125). Unlike the related art, in the embodiment of the present application, the updated weights are determined based on the weights corresponding to the first neighborhood points. The updated weights are determined based on the weights corresponding to the neighborhood points, which can make the transformation results more reasonable and help improve the coding efficiency of the geometric information of the points in the three-dimensional grid.
[0192] The weight corresponding to the first neighborhood point can be determined based on the spatial connectivity of the first neighborhood point in the three-dimensional grid. For example, the weight corresponding to the first neighborhood point can be related to the number of neighboring points of the first neighborhood point (hereinafter referred to as second neighborhood points). Determining the weight corresponding to a point based on the number of neighboring points fully utilizes the connectivity information of the points in the three-dimensional grid, which can make the determined point weight more reasonable.
[0193] The embodiment of the present application does not specifically limit the definition of the second neighborhood point. For example, the second neighborhood point can be a point in the same LOD layer as the first neighborhood point. Alternatively, the second neighborhood point can be a point in the next lower LOD layer than the first neighborhood point. Alternatively, the second neighborhood point can include points in both of the aforementioned LOD layers. As an example, the second neighborhood point is a point predicted based on the first neighborhood point.
[0194] In some implementations, the weight corresponding to the first neighborhood point can be equal to the number of the second neighborhood points. For example, if the first neighborhood points include point 1 and point 2, point 1 includes 5 neighborhood points, and point 2 includes 6 neighborhood points, then the weight of point 1 can be 5, and the weight of point 2 can be 6.
[0195] In some implementations, the weight corresponding to the first neighborhood point can be determined based on the product of the number of second neighborhood points and the first predicted weight. For example, the weight corresponding to the first neighborhood point can be equal to the product of the number of second neighborhood points and the first predicted weight. The first predicted weight mentioned here can be a predefined fixed value. For example, the value of the first predicted weight can be 0.5. When determining the weight corresponding to the point, considering the predicted weight can make the determined weight more reasonable.
[0196] For example, the weight corresponding to the first neighborhood point may be determined based on the following formula:
[0197] In the above formula, weight represents the weight corresponding to any point in the first neighborhood, and predWeight is the first predicted weight (which can be 0.5). P represents the second neighborhood, and weight[i] represents the initial weight corresponding to the i-th point in the second neighborhood. The initial weight corresponding to the i-th point can be set to 1. For example, if the first neighborhood includes points 1 and 2, point 1 includes 5 neighboring points, and point 2 includes 6 neighboring points, then based on the above formula, the weight of point 1 can be 2.5, and the weight of point 2 can be 3 (the first predicted weight is 0.5).
[0198] In some implementations, the weight corresponding to the first neighborhood point can be determined based on the weight corresponding to the LOD layer where the first neighborhood point is located. For example, the weight corresponding to the first neighborhood point can be equal to the weight corresponding to the LOD layer where the first neighborhood point is located. The weight corresponding to the LOD layer can be understood as a level weight. The level weight of an LOD layer can be determined, for example, based on the ratio of the number of points in the LOD layer to the total number of points to be decoded.
[0199] As mentioned above, the update weight can be determined based on the weight corresponding to the first neighborhood point. As a possible implementation method, the update weight can be equal to the weight corresponding to the first neighborhood point. In another possible implementation method, the update weight can be determined based on the weight corresponding to the first neighborhood point and the first prediction weight (the first prediction weight can be a predefined fixed value, for example, the value of the first prediction weight can be 0.5). For example, the update weight can be determined based on the product of the weight corresponding to the first neighborhood point and the first prediction weight, or the update weight can be equal to the product of the weight corresponding to the first neighborhood point and the first prediction weight. Taking the prediction weight into consideration when determining the update weight can make the weight distribution of the prediction and update process more reasonable.
[0200] Exemplarily, the update weight may be determined based on the following formula: updateWeight i =predWeight×weight i
[0201] In the above formula, updateWeight i Indicates the updated weight corresponding to the i-th point in the first neighborhood point. predWeight indicates the first prediction weight, and the value of the first prediction weight can be 0.5. weight i The weight corresponding to the first neighborhood point can be determined using any of the above implementation methods.
[0202] The above describes in detail the weight corresponding to the first neighborhood point and the method for determining the updated weight. Based on the determination of the weight corresponding to the first neighborhood point and the updated weight, the shift coefficient of the first point to be encoded can be transformed according to the following formula:
[0203] In the above formula, Signal represents the shift coefficient of the first point to be encoded. updateWeight i Indicates the update weight corresponding to the i-th point in the first neighborhood. updateWeight i The calculation method of can be found in the previous text. P represents the point set formed by the first neighborhood point, and i represents the i-th point in the first neighborhood point. i Represents the residual coefficient (shifted residual coefficient) of the first neighborhood point.
[0204] In addition to the scheme of transforming based on the weight corresponding to the point, the embodiment of the present application does not exclude other inverse transformation schemes. For example, the transformation can be performed based on the adjustment parameter (Scale) of the updated weight, as shown in the following formula:
[0205] In the above formula, updateWeight represents the update weight (can be a fixed value, such as 0.125). n-i-1 is the adjustment parameter for updating the weight, n represents the number of LOD layers, i represents the index of the LOD layer. P represents the point set formed by the first neighborhood point, and i represents the i-th point in the first neighborhood point. i Represents the residual coefficient (shifted residual coefficient) of the first neighborhood point.
[0206] The above describes in detail the transformation operation of the points in the three-dimensional grid. The following describes the prediction operation of the points in the three-dimensional grid from the perspective of the second point to be encoded. The second point to be encoded can be any point to be encoded in the second layer. The neighborhood point of the second point to be encoded is called the third neighborhood point. The embodiment of the present application does not specifically limit the definition of the third neighborhood point. For example, the third neighborhood point can be a point in the same LOD layer as the second point to be encoded. Alternatively, the third neighborhood point can be a point in the next layer of the LOD layer where the second point to be encoded is located. Alternatively, the third neighborhood point can include points in the above two LOD layers at the same time. As an example, the third neighborhood point can be a point in the first layer (the layer above the LOD layer where the second point to be encoded is located) that is collinear with the second point to be encoded.
[0207] Step S2020 in Figure 20, i.e., predicting the shift coefficients of the points in the second layer, may include: predicting the second point to be encoded based on the shift coefficients of the third neighborhood points and the second prediction weight. In the related art, the prediction weight adopts a fixed value (i.e., 0.5). Different from the related art, in the embodiment of the present application, the second prediction weight is determined based on the weight corresponding to the third neighborhood point. For example, the greater the weight corresponding to a point in the third neighborhood point, the greater the influence of the point on the prediction result of the second point to be encoded. The prediction weight is determined based on the weight corresponding to the neighborhood point, which can make the points with higher importance have a greater influence on the prediction result of the point to be encoded, which will make the prediction process more reasonable and help to improve the coding efficiency of the geometric information of the points in the three-dimensional grid.
[0208] The weight corresponding to the third neighboring point can be determined based on the spatial connectivity of the third neighboring point in the three-dimensional grid. For example, the weight corresponding to the third neighboring point can be related to the number of neighboring points of the third neighboring point (hereinafter referred to as the fourth neighboring points). Determining the weight corresponding to a point based on the number of neighboring points fully utilizes the connectivity information of the points in the three-dimensional grid, which can make the determined point weight more reasonable.
[0209] The embodiment of the present application does not specifically limit the definition of the fourth neighborhood point. For example, the fourth neighborhood point can be a point in the same LOD layer as the third neighborhood point. Alternatively, the fourth neighborhood point can be a point in the LOD layer below the LOD layer of the third neighborhood point. Alternatively, the fourth neighborhood point can include points in both of the aforementioned LOD layers. As an example, the fourth neighborhood point is a point predicted based on the third neighborhood point.
[0210] In some implementations, the weight corresponding to the third neighborhood point can be equal to the number of the fourth neighborhood points. For example, if the third neighborhood points include point 1 and point 2, point 1 includes 5 neighborhood points, and point 2 includes 6 neighborhood points, then the weight of point 1 can be 5, and the weight of point 2 can be 6.
[0211] In some implementations, the weight corresponding to the third neighborhood point can be determined based on the product of the number of fourth neighborhood points and the first predicted weight. For example, the weight corresponding to the third neighborhood point can be equal to the product of the number of fourth neighborhood points and the first predicted weight. The first predicted weight mentioned here can be a predefined fixed value. For example, the value of the first predicted weight can be 0.5. Considering the predicted weight when determining the weight corresponding to the point can make the determined weight more reasonable.
[0212] For example, the weight corresponding to the third neighborhood point can be determined based on the following formula:
[0213] In the above formula, weight represents the weight corresponding to any point in the third neighborhood, and predWeight is the first predicted weight (which can be 0.5). P represents the fourth neighborhood, and weight[i] represents the initial weight corresponding to the i-th point in the fourth neighborhood. The initial weight corresponding to the i-th point can be set to 1. For example, if the third neighborhood includes points 1 and 2, point 1 includes 5 neighboring points, and point 2 includes 6 neighboring points, then based on the above formula, the weight of point 1 can be 2.5, and the weight of point 2 can be 3 (the first predicted weight is 0.5).
[0214] In some implementations, the weight corresponding to the third neighboring point can be determined based on the weight corresponding to the LOD layer where the first neighboring point is located. For example, the weight corresponding to the third neighboring point can be equal to the weight corresponding to the LOD layer where the third neighboring point is located. The weight corresponding to the LOD layer can be understood as a level weight. The level weight of an LOD layer can be determined, for example, based on the ratio of the number of points in the LOD layer to the total number of points to be decoded.
[0215] As previously mentioned, the second prediction weight can be determined based on the weight corresponding to the third neighborhood point. For example, the second prediction weight can be equal to the weight corresponding to the third neighborhood point. Alternatively, the second prediction weight can be based on the product of the weight corresponding to the third neighborhood point and a parameter.
[0216] As mentioned above, the second point to be coded can be predicted based on the shift coefficient of the third neighborhood point and the second prediction weight. A possible prediction method for the second point to be coded is given below.
[0217] Exemplarily, the second point to be encoded may be predicted based on the following formula:
[0218] In the above formula, Signal represents the prediction result of the second point to be coded, Neigh1 and Neigh2 represent the shift coefficients of the two neighboring points collinear with the second point to be coded. Neigh1 and Neigh2 are located in the LOD layer above the LOD layer of the second point to be coded. predWeight1 and predWeight2 represent the prediction weights corresponding to Neigh1 and Neigh2, respectively (i.e., the second prediction weights mentioned above).
[0219] predWeight1 and predWeight2 can be calculated using the following formula: predWeight1 = weight1 predWeight2 = weight2
[0220] sumWeight can be calculated based on the following formula: sumWeight=predWeight1+predWeight2
[0221] The above describes in detail the transformation and prediction operations for points in a 3D grid. The following describes the quantization operations for points in a 3D grid from the perspective of a third point to be coded. The third point to be coded can be any point to be coded in the 3D grid.
[0222] Step S2040 in Figure 20, i.e., quantizing the transformation coefficients of the points in the three-dimensional grid, may include: quantizing the transformation coefficients of the third point to be encoded according to the weight corresponding to the third point to be encoded. In the embodiment of the present application, when quantizing the points to be encoded in the three-dimensional grid, the weight corresponding to the point to be encoded is taken into consideration. For example, a suitable quantization parameter may be selected based on the weight corresponding to the point to be encoded; or, the quantization result may be processed differently based on the weight corresponding to the point to be encoded. Considering the weight of the point to be encoded in the quantization process can make the quantization result more reasonable, which helps to improve the encoding efficiency of the geometric information of the points in the three-dimensional grid.
[0223] The weight corresponding to the third point to be encoded can be determined based on the spatial connectivity of the third point to be encoded in the three-dimensional grid. For example, the weight corresponding to the third point to be encoded can be related to the number of neighboring points (hereinafter referred to as fifth neighboring points) of the third point to be encoded. Determining the weight corresponding to a point based on the number of neighboring points fully utilizes the connectivity information of the points in the three-dimensional grid, which can make the determined point weight more reasonable.
[0224] The embodiment of the present application does not specifically limit the definition of the fifth neighboring point. For example, the fifth neighboring point can be a point in the same LOD layer as the third point to be encoded. Alternatively, the fifth neighboring point can be a point in the LOD layer below the LOD layer in which the third point to be encoded resides. Alternatively, the fifth neighboring point can include points in both of the aforementioned LOD layers. As an example, the fifth neighboring point is a point predicted based on the third point to be encoded.
[0225] In some implementations, the weight corresponding to the third to-be-encoded point may be equal to the number of the fifth neighboring points. For example, if the third to-be-encoded point includes point 1 and point 2, point 1 includes 5 neighboring points, and point 2 includes 6 neighboring points, then the weight of point 1 may be 5, and the weight of point 2 may be 6.
[0226] In some implementations, the weight corresponding to the third point to be encoded can be determined based on the product of the number of fifth neighboring points and the first predicted weight. For example, the weight corresponding to the third point to be encoded can be equal to the product of the number of fifth neighboring points and the first predicted weight. The first predicted weight mentioned here can be a predefined fixed value. For example, the value of the first predicted weight can be 0.5. Considering the predicted weight when determining the weight corresponding to the point can make the determined weight more reasonable.
[0227] Exemplarily, the weight corresponding to the third to-be-encoded point may be determined based on the following formula:
[0228] In the above formula, weight represents the weight corresponding to any point in the third to-be-encoded point, and predWeight is the first predicted weight (which can be 0.5). P represents the fifth neighborhood point, and weight[i] represents the initial weight corresponding to the i-th point in the fifth neighborhood point. The initial weight corresponding to the i-th point can be set to 1. For example, if the third to-be-encoded point includes point 1 and point 2, point 1 includes 5 neighborhood points, and point 2 includes 6 neighborhood points, then based on the above formula, the weight of point 1 can be 2.5, and the weight of point 2 can be 3 (the first predicted weight is 0.5).
[0229] As mentioned above, the transform coefficient of the third point to be coded can be quantized according to the weight corresponding to the third point to be coded. A possible implementation is given below.
[0230] For example, in some implementations, the transform coefficient of the third point to be encoded can be quantized according to the first quantization parameter. The first quantization parameter can be determined based on the weight corresponding to the third point to be encoded. That is, the quantization parameter can be adaptively set for the point to be encoded according to the weight corresponding to the point to be encoded. For example, if the weight corresponding to the third point to be encoded is large, it means that the importance of the third point to be encoded is high. For points to be encoded with a high degree of importance, the value of the first quantization parameter can be small (indicating that the degree of quantization of the point to be encoded is low). For another example, if the weight corresponding to the third point to be encoded is small, it means that the importance of the third point to be encoded is low. For points to be encoded with a low degree of importance, the value of the first quantization parameter can be large (indicating that the degree of quantization of the point to be encoded is high). According to the design method of the above-mentioned quantization parameters, the reconstruction quality can be guaranteed, thereby improving the geometric coding efficiency of points in the three-dimensional grid.
[0231] The first quantization parameter can be directly determined based on the weight corresponding to the third point to be encoded. For example, a mapping relationship between the weight corresponding to the point to be encoded and the quantization parameter can be established in advance, and the first quantization parameter can be determined based on the mapping relationship. Alternatively, the first quantization parameter can also be determined based on the weight corresponding to the third point to be encoded and the second quantization parameter. The second quantization parameter mentioned here can be the quantization parameter corresponding to the target LOD layer (i.e., the LOD layer to which the third point to be encoded belongs). The second quantization parameter can be determined based on the initial quantization parameter and the quantization adjustment parameter.
[0232] Exemplarily, the second quantization parameter may be determined based on the following formula: QP Lvl =liftingQP×Scale Lvl
[0233] In the above formula, Lvl represents the index of the target LOD layer, QP Lvl is the second quantization parameter, liftingQP is the initial quantization parameter, and Scale is the quantization adjustment parameter.
[0234] A more specific example of quantization of the third point to be encoded is given below.
[0235] For example, the third to-be-coded point may be quantized based on the following formula:
[0236] In the above formula, QRes i Represents the quantization result of the third point to be coded, Rea i Indicates the transform coefficient of the third point to be coded, bitDepth represents the effective bit depth of the shift coefficient, QP Lvl Indicates the quantization parameter corresponding to the target LOD layer (the LOD layer where the third point to be encoded is located), weight irepresents the weight corresponding to the third point to be encoded (the method for determining the weight corresponding to the third point to be encoded can be found in the previous text and will not be described in detail here).
[0237] The encoding and decoding method provided in the embodiment of the present application is described in detail above in conjunction with Figures 19 and 20. It should be noted that, in the absence of conflict, the terms "coefficient" and "information" or "value" mentioned in any of the above embodiments can be used interchangeably. For example, the shift coefficient mentioned above can be replaced by a shift value or shift information, and the residual coefficient can be replaced by a residual value or residual information.
[0238] It should also be noted that in the above examples, the corresponding prediction operation on the decoding side can also be called an inverse prediction operation. The transform operation on the encoding side can also be called an update operation. The inverse transform operation on the decoding side can also be called an update operation or an inverse update operation.
[0239] The following describes the embodiments of the present application in more detail with reference to specific examples. It should be noted that the examples below are only intended to help those skilled in the art understand the embodiments of the present application, and are not intended to limit the embodiments of the present application to the specific numerical values or specific scenarios illustrated. Those skilled in the art can obviously make various equivalent modifications or changes based on the examples given, and such modifications or changes also fall within the scope of the embodiments of the present application.
[0240] Example 1:
[0241] Example 1: Obtain the weight of each point based on the spatial connection relationship between adjacent points in a three-dimensional grid i , where i = 0, 1, 2…pointCount, pointCount represents the number of points in the three-dimensional grid. Then, the shift coefficient of each point is transformed with the help of the LOD spatial structure. Among them, the prediction scheme still adopts the prediction coding scheme provided by the relevant technology, that is, the shift coefficients of two collinear adjacent points are used to predict the shift coefficient of the middle position point, and the prediction weight is fixed at 0.5. When updating the shift coefficient of each point, the update weight will be based on the weight weight of each point. i And the prediction weights are adaptively obtained.
[0242] 1.1. Overall coding process
[0243] As shown in Figure 21, based on the LOD spatial structure, the shift coefficients of the entire set of points to be encoded are divided into low-frequency coefficients L(N) and high-frequency coefficients H(N). The low-frequency coefficients L(N) are then used to predict the high-frequency coefficients H(N). The predicted high-frequency coefficients H(N) are then used to adaptively update the low-frequency coefficients L(N).
[0244] Figure 22 shows a three-layer LOD spatial structure. The encoding order is from the lowest LOD layer to the highest LOD layer, i.e., Lv12->Lv11->Lv10. First, the entire shift coefficient to be encoded is divided into high-frequency coefficients Lv12 and low-frequency coefficients Lv10+Lv11. After encoding the shift coefficients at Lv12, the remaining shift coefficients are further divided into high-frequency coefficients Lv11 and low-frequency coefficients Lv10.
[0245] 1.2 Encoding Algorithm
[0246] Assign an initial weight to all points to be coded i If set to 1, all the code points are considered to be equally important.
[0247] The initial weights are updated according to the connection relationship between the points in the three-dimensional grid. The specific update formula is as follows:
[0248] Where predWeight is the prediction weight. The current V-DMC sets the prediction weight to a fixed value of 0.5. The P set represents the set of neighboring points predicted based on the current point.
[0249] Prediction is performed based on the LOD spatial structure. The specific prediction algorithm is shown in the following formula: signal = predWeight × (Neigh1 + Neigh2)
[0250] Among them, Neigh1 and Neigh2 represent the shift coefficients of two neighboring points that are collinear with the current point, and these two neighboring points are located in the upper LOD of the current layer.
[0251] After completing the prediction of the shift coefficients of all points in the current layer, the predicted residual of the shift coefficient of each point is used to update the shift coefficient of the point in the upper LOD. The specific update formula is as follows:
[0252] Among them, updateWeight i The specific calculation method of updateWeight is as follows: i =predWeight×weight i
[0253] predWeight is the prediction weight, weight i is the weight of each point.
[0254] sumWeight is calculated as follows:
[0255] signal iRepresents the shifted residual coefficients of each point predicted based on the current point.
[0256] The above prediction and update process is repeated continuously, and the transformation is performed from the lower layer to the upper layer according to the spatial structure of the LOD. Finally, the shift coefficients of the entire grid are transformed from the spatial domain to the frequency domain, thereby completing the geometric encoding of the entire grid.
[0257] 1.3 Decoding Algorithm
[0258] The decoding algorithm is a completely opposite process to the encoding algorithm. After completing the entropy decoding of the shift coefficients, the decoder continuously performs inverse transformations from the higher levels of the LOD to the lowest level of the LOD based on the LOD spatial structure, thereby reconstructing the geometric information of the entire mesh. The specific decoding algorithm is described below.
[0259] First, all points to be decoded are assigned an initial weight weight i If set to 1, all points to be decoded are considered equally important.
[0260] The weights are updated according to the connection relationship between the points in the three-dimensional grid. The specific update formula is as follows:
[0261] Where predWeight is the prediction weight value. The current V-DMC sets the prediction weight to a fixed value of 0.5. The P set is the neighborhood point set for prediction based on the current point.
[0262] According to the LOD spatial structure, the inverse transformation is performed, and the spatial connection relationship between points is used to obtain the neighborhood point set based on the current point for prediction. The specific inverse transformation formula is as follows:
[0263] Among them, updateWeight i The specific calculation method of updateWeight is as follows: i =predWeight×weight i
[0264] predWeight is the prediction weight, weight i is the weight of each point.
[0265] sumWeight is calculated as follows:
[0266] signal i Represents the shifted residual coefficients of the point predicted based on the current point.
[0267] According to the spatial structure of LOD, reverse prediction is performed. The specific prediction algorithm is shown in the following formula: signal+=predWeight×(Neigh1+Neigh2)
[0268] Among them, Neigh1 and Neigh2 represent two neighboring points on the same line of the current point, and these two neighboring points are located in the upper LOD.
[0269] The inverse transformation and inverse prediction process described above is repeated continuously, and the inverse transformation is performed from the upper layer to the lower layer according to the LOD spatial structure, and finally the shift coefficients of the entire grid are transformed from the frequency domain to the spatial domain, thereby completing the geometric encoding of the entire grid.
[0270] After completing the inverse transformation of the shift coefficient of each point in the three-dimensional grid, the initial geometric position information (geometric information of the subdivided grid) and the reconstructed shift coefficient of each point can be added to determine the reconstructed geometric information of the three-dimensional grid.
[0271] Example 1 proposes a shift coding scheme that performs lifting transformation based on the spatial connection relationship of the grid. Specifically, the points after the basic grid division are divided into LOD space, and the shift of the points in the entire grid is lifted and transformed (including prediction and transformation) based on the LOD space structure. The prediction process uses two colinear neighboring points of the current point for prediction coding, and the transformation process adaptively updates the shift coefficient of the current point according to the shift coefficient, weight and prediction weight of each point in the set of neighboring points predicted based on the current point. That is, Example 1 adjusts the fixed update weight used in the update algorithm to: update according to the number of neighboring points of each point to be encoded, the weight of each neighboring point and the shift residual coefficient of the neighboring point. This coding scheme can further improve the geometric coding efficiency of the grid by effectively utilizing the spatial relationship between points to transform the shift coefficient of each point.
[0272] The technical solution provided in Example 1 was tested using test conditions C1 and C2, based on V-DMC-4.0. Taking the lossy test environment with geometric lossy attributes as an example, BD-Rate is a performance indicator for measuring compression efficiency. When BD-Rate is less than 0, it indicates that the encoding efficiency has improved compared to traditional encoding schemes. As can be seen in the table below, compared to V-DMC-4.0, under the condition of C1, D1 and D2 increased by approximately 2.4% and 2.4%, respectively, and the Luma component increased by approximately 0.4%. Under the condition of C2, D1 and D2 increased by 2.1% and 2.1%, respectively, and the Luma component increased by approximately 0.4%.
[0273] Example 2:
[0274] Example 2: Get the weight of each point based on the spatial connection relationship between adjacent points in the grid i , where i = 0, 1, 2…pointCount, where pointCount represents the number of points in the three-dimensional grid. During the encoding process, a lifting transform is first performed on the shift coefficients of the current grid, transforming the shift coefficients of each point from the spatial domain to the frequency domain. Then, after the transformation is completed, the shift coefficients are quantized. During quantization, the quantization parameter of the current point to be encoded is adaptively determined based on the weight of the current point.
[0275] 2.1 The entire encoding process
[0276] As shown in Figure 22, based on the LOD spatial structure, the points of the entire grid to be encoded are divided into three LOD layers. The encoding order is from the lowest LOD layer to the highest LOD layer, that is, Lv12->Lv11->Lv10. First, the entire shift coefficient to be encoded is divided into high-frequency coefficients Lv12 and low-frequency coefficients Lv10+Lv11. After completing the encoding of the shift coefficients of the Lv12 layer, the remaining shift coefficients are further divided into high-frequency coefficients Lv11 and low-frequency coefficients Lv10.
[0277] 2.2 Encoding Algorithm
[0278] Assign an initial weight to all points to be coded i If set to 1, all the code points are considered to be equally important.
[0279] The initial weights are updated according to the connection relationship between the points in the three-dimensional grid. The specific update formula is as follows:
[0280] Where predWeight is the prediction weight. The current V-DMC sets the prediction weight to a fixed value of 0.5. The P set represents the set of neighboring points predicted based on the current point.
[0281] Prediction is performed based on the LOD spatial structure. The specific prediction algorithm is shown in the following formula: Signal = predWeight × (Neigh1 + Neigh2)
[0282] Among them, Neigh1 and Neigh2 represent the shift coefficients of two neighboring points that are collinear with the current point, and these two neighboring points are located in the upper LOD of the current layer.
[0283] After completing the prediction of the displacement coefficients of all points in the current layer, the displacement coefficients of the midpoints in the upper LOD are updated using the prediction residual of each point's displacement coefficient. The specific update formula is as follows:
[0284] Among them, updateWeight is the update weight, and P set is the point set predicted based on the current point.
[0285] The shift coefficients of the current layer are adaptively quantized. The specific quantization method is as follows:
[0286] Among them, Res i Indicates the transform coefficient of the current point's shift coefficient, bitDepth represents the effective bit depth of the shift coefficient in the current geometric coding, QP Lvl Represents the quantization parameter of the current LOD layer, weight i is the weight corresponding to the current point.
[0287] According to the spatial structure of LOD, the prediction, update and quantization process described above is repeated continuously to complete the geometric encoding of the entire grid.
[0288] 2.3 Decoding Algorithm
[0289] The decoding algorithm is a completely opposite process to the encoding algorithm. When encoding the point shift coefficients, the encoder continuously performs inverse transformations from the higher level of the LOD to the lower level of the LOD based on the spatial structure of the LOD. After completing the entropy decoding of the shift coefficients, the decoder continuously performs inverse transformations from the higher level of the LOD to the lower level of the LOD based on the spatial structure of the LOD, thereby reconstructing the geometric information of the entire mesh.
[0290] Assign an initial weight to all points to be decoded i If set to 1, all points to be decoded are considered equally important.
[0291] The initial weights are updated according to the connection relationship between the points in the three-dimensional grid. The specific update formula is as follows:
[0292] Where predWeight is the prediction weight. The current V-DMC sets the prediction weight to a fixed value of 0.5. The P set represents the set of neighboring points predicted based on the current point.
[0293] The shift coefficients of the current layer are adaptively dequantized. The specific dequantization method is as follows:
[0294] Among them, Res i Indicates the quantization coefficient of the shift coefficient of the current point, bitDepth represents the effective bit depth of the shift coefficient in the current geometric coding, QP Lvl Represents the quantization parameter of the current LOD layer, weight iis the weight of winning at the current point, InvRes i Indicates the inverse quantization result of the current point.
[0295] Perform inverse transformation according to the spatial structure of LOD. The specific inverse transformation formula is as follows:
[0296] Among them, updateWeight is the update weight, and P set is the set of neighborhood points predicted based on the current point.
[0297] Perform reverse prediction based on the spatial structure of LOD. For specific prediction, refer to the following formula: Signal = predWeight × (Neigh1 + Neigh2)
[0298] Among them, Neigh1 and Neigh2 represent the shift coefficients of two neighboring points that are collinear with the current point, and these two neighboring points are located in the upper LOD layer of the current point.
[0299] The above inverse quantization, inverse transformation and inverse prediction processes are repeated continuously to complete the geometric decoding of the entire grid.
[0300] After completing the inverse transformation of the shift coefficient of each point in the three-dimensional grid, the initial geometric position information (geometric information of the subdivided grid) and the reconstructed shift coefficient of each point can be added to determine the reconstructed geometric information of the three-dimensional grid.
[0301] This solution proposes a shift coding scheme with adaptive quantization based on the spatial connectivity of the grid. Specifically, the points after the basic grid is divided are subjected to LOD spatial partitioning; based on the spatial structure of the LOD, the shifts of the points in the grid are subjected to a lifting transformation (including prediction and transformation). During the prediction process, two colinear neighboring points of the current point are used for predictive coding. The transformation process adaptively updates the shift coefficient of the current point based on the shift coefficient of each point in the domain point set predicted based on the current point and a fixed update weight. After the shift coefficient transformation is completed, the shift coefficient is quantized and encoded. During quantization coding, the shift coefficients of different points in the current grid are adaptively quantized and encoded based on their weights. The weights of the points can be obtained by utilizing the spatial connectivity between the points in the current grid. A smaller quantization level can be used for points with larger weights and higher importance, while a higher quantization level can be used for points with smaller weights and lower importance. Based on this scheme, the bit rate of the grid's geometric information can be reduced while ensuring the reconstruction quality of the entire grid, thereby further improving the coding efficiency of the grid's geometric information.
[0302] The technical solution provided in Example 1 was tested based on V-DMC-4.0 using test conditions C1 and C2. Taking the lossy test environment with geometric lossy attributes as an example, BD-Rate is a performance indicator for measuring compression efficiency. When BD-Rate is less than 0, it indicates that the encoding efficiency has improved compared to traditional encoding schemes. As can be seen from the table below, compared to V-DMC-4.0, under the condition of C1, D1 and D2 increased by approximately 0.2% and 0.2%, respectively, and the Luma component increased by approximately 0.0% to 0.1%; under the condition of C2, D1 and D2 increased by 0.2% and 0.2%, respectively, and the Luma component increased by approximately 0.1%.
[0303] Example 3:
[0304] Based on Example 1 or Example 2, after obtaining the weight of the current point, the prediction weight of each point can be adaptively updated according to the weight of the current point. That is, the prediction algorithm of the current point is modified as follows:
[0305] Among them, predWeight1 can be obtained by using the weight of the first neighborhood point, that is, predWeight1=weight1 predWeight2=weight2
[0306] Then, the calculation method of sumWeight is as follows: sumWeight=ptedWeight1+predWeight2
[0307] Based on such a coding scheme, when predictive coding is performed on the shift coefficient of the current point, the spatial connection relationship between the grid points is also utilized, thereby further improving the coding efficiency of the geometric information of the grid.
[0308] The method embodiment of the present application is described in detail above in conjunction with Figures 1 to 22. The device embodiment of the present application is described in detail below in conjunction with Figures 23 to 26. It should be understood that the description of the method embodiment corresponds to the description of the device embodiment. Therefore, for parts not described in detail, reference can be made to the above method embodiment.
[0309] Figure 23 is a schematic diagram of the structure of a decoder provided by an embodiment of the present application. The decoder 2300 shown in Figure 23 includes a first determination unit 2310, a second determination unit 2320, a third determination unit 2330, and a prediction unit 2340. The first determination unit 2310 is configured to parse the code stream and determine the quantization coefficients of the points in the three-dimensional grid. The second determination unit 2320 is configured to dequantize the quantization coefficients of the points in the three-dimensional grid to determine the dequantization coefficients of the points in the three-dimensional grid. The third determination unit 2330 is configured to perform an inverse transform on the dequantization coefficients of the points in the first layer to determine the shift coefficients of the points in the first layer. The prediction unit 2340 is configured to predict the shift coefficients of the points in the second layer based on the shift coefficients of the points in the first layer. The points in the three-dimensional grid belong to multiple LOD layers, where the first layer and the second layer are adjacent layers in the multiple LOD layers, and the first layer is the layer above the second layer.
[0310] In some implementations, the third determination unit 2330 is configured to: perform an inverse transformation on the inverse quantization coefficient of the first point to be decoded in the first layer based on the inverse quantization coefficient of the first neighborhood point and the updated weight, wherein the first neighborhood point is a neighborhood point of the first point to be decoded, and the updated weight is determined based on the weight corresponding to the first neighborhood point.
[0311] In some implementations, the updated weight is determined based on the weight corresponding to the first neighborhood point and the first prediction weight.
[0312] In some implementations, the updated weight is determined based on the product of the weight corresponding to the first neighborhood point and the first prediction weight.
[0313] In some implementations, the weight corresponding to the first neighborhood point is related to the number of second neighborhood points, and the second neighborhood points are neighborhood points of the first neighborhood point; or, the weight corresponding to the first neighborhood point is determined based on the weight corresponding to the LOD layer where the first neighborhood point is located.
[0314] In some implementations, the weight corresponding to the first neighborhood point is equal to the number of the second neighborhood points; or, the weight corresponding to the first neighborhood point is equal to the product of the number of the second neighborhood points and the first prediction weight.
[0315] In some implementations, the second neighborhood point is a point predicted based on the first neighborhood point.
[0316] In some implementations, the first neighborhood point is a point predicted based on the first point to be decoded.
[0317] In some implementations, the prediction unit 2340 is configured to predict the second point to be decoded based on the shift coefficient of the third neighborhood point and the second prediction weight, where the second point to be decoded is any point in the second layer, the third neighborhood point is the neighborhood point of the second point to be decoded in the first layer, and the second prediction weight is determined based on the weight corresponding to the third neighborhood point.
[0318] In some implementations, the second prediction weight is equal to the weight corresponding to the third neighborhood point.
[0319] In some implementations, the weight corresponding to the third neighborhood point is associated with the number of fourth neighborhood points, and the fourth neighborhood point is a neighborhood point of the third neighborhood point; or, the weight corresponding to the third neighborhood point is determined based on the weight corresponding to the LOD layer where the third neighborhood point is located.
[0320] In some implementations, the weight corresponding to the third neighborhood point is equal to the number of the fourth neighborhood points; or, the weight corresponding to the third neighborhood point is equal to the product of the number of the fourth neighborhood points and the first prediction weight.
[0321] In some implementations, the fourth neighborhood point is a point predicted based on the third neighborhood point.
[0322] In some implementations, the third neighborhood point is a point in the first layer that is collinear with the second point to be decoded.
[0323] In some implementations, the second determining unit 2320 is configured to: inversely quantize the quantization coefficient of the third point to be decoded according to a weight corresponding to the third point to be decoded, wherein the third point to be decoded is any point in the three-dimensional network.
[0324] In some implementations, the second determining unit 2320 is configured to: inversely quantize the quantized coefficient of the third point to be decoded according to a first quantization parameter, where the first quantization parameter is determined based on a weight corresponding to the third point to be decoded.
[0325] In some implementations, the first quantization parameter is determined based on the weight corresponding to the third point to be decoded and the second quantization parameter, wherein the third point to be decoded is located in the target LOD layer among the multiple LOD layers, and the second quantization parameter is the quantization parameter corresponding to the target LOD layer.
[0326] In some implementations, the second quantization parameter is determined based on an initial quantization parameter and a quantization adjustment parameter.
[0327] In some implementations, the greater the weight corresponding to the third point to be decoded, the smaller the value of the first quantization parameter.
[0328] In some implementations, the weight corresponding to the third point to be decoded is related to the number of fifth neighboring points, and the fifth neighboring points are neighboring points of the third point to be decoded.
[0329] In some implementations, the weight corresponding to the third point to be decoded is equal to the number of the fifth neighborhood points; or, the weight corresponding to the third point to be decoded is equal to the product of the number of the fifth neighborhood points and the first prediction weight.
[0330] In some implementations, the fifth neighborhood point is a point predicted based on the third point to be decoded.
[0331] In some implementations, the value of the first prediction weight is a predefined fixed value.
[0332] In some implementations, the value of the first prediction weight is 0.5.
[0333] In some implementations, the decoder 2300 further includes: a fourth determination unit configured to: parse the code stream to determine the geometric information of the points in the basic grid; divide the basic grid according to the geometric information of the points in the basic grid to determine the geometric information of the points in the subdivided grid; and determine the reconstructed geometric information of the points in the three-dimensional grid based on the shift coefficients of the points in the three-dimensional grid and the geometric information of the points in the subdivided grid.
[0334] It is understandable that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and of course it can also be a module, or it can be non-modular. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.
[0335] If the integrated unit is implemented in the form of a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0336] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the decoder 2300. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the decoding method described in any one of the aforementioned embodiments.
[0337] Based on the composition of the above-mentioned decoder and the computer-readable storage medium, refer to Figure 24, which shows a schematic diagram of the specific hardware structure of the decoder provided by an embodiment of the present application. As shown in Figure 24, the decoder 2400 may include: a communication interface 2410, a memory 2420 and a processor 2430; each component is coupled together through a bus system 2440. It can be understood that the bus system 2440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 2440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 2440 in Figure 24. Among them,
[0338] Communication interface 2410, used for sending and receiving signals when sending and receiving information with other external network elements;
[0339] Memory 2420, for storing computer programs;
[0340] The processor 2430 is configured to, when running the computer program, execute:
[0341] Parse the bitstream and determine the quantization coefficients of the points in the three-dimensional grid;
[0342] Dequantizing the quantized coefficients of the points in the three-dimensional grid to determine the dequantized coefficients of the points in the three-dimensional grid;
[0343] Performing an inverse transformation on the inverse quantized coefficients of the points in the first layer to determine the shift coefficients of the points in the first layer;
[0344] Predicting the shift coefficients of the points in the second layer according to the shift coefficients of the points in the first layer;
[0345] The points in the three-dimensional grid belong to multiple LOD layers, the first layer and the second layer are two adjacent layers in the multiple LOD layers, and the first layer is the upper layer of the second layer.
[0346] It is understood that the memory 2420 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DRRAM). The memory 2420 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0347] Processor 2430 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method may be completed by hardware integrated logic circuits or software instructions in processor 2430. The above-mentioned processor 2430 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 2420, and the processor 2430 reads the information in the memory 2420 and completes the steps of the above method in combination with its hardware.
[0348] It is understood that the embodiments described herein can be implemented with hardware, software, firmware, middleware, microcode or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein or a combination thereof. For software implementation, the technology described herein can be implemented by a module (such as a process, a function, etc.) that performs the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in a processor or outside a processor.
[0349] Optionally, as another embodiment, the processor 2430 is further configured to execute the decoding method described in any one of the aforementioned embodiments when running the computer program.
[0350] Figure 25 is a schematic diagram of the structure of an encoder provided by an embodiment of the present application. The encoder 2500 of Figure 25 includes a first determination unit 2510, a second determination unit 2520, and a transformation unit 2530. The first determination unit 2510 is configured to perform detail level LOD division on the points in the three-dimensional grid and determine multiple LOD layers, wherein the multiple LOD layers include adjacent first and second layers, and the first layer is the upper layer of the second layer. The second determination unit 2520 is configured to predict the shift coefficients of the points in the second layer and determine the residual coefficients of the points in the second layer. The transformation unit 2530 is configured to transform the shift coefficients of the points in the first layer according to the residual coefficients of the points in the second layer.
[0351] In some implementations, the transformation unit 2530 is configured to transform the shift coefficient of the first point to be encoded in the first layer according to the residual coefficient of the first neighborhood point and the updated weight, wherein the first neighborhood point is the neighborhood point of the first point to be encoded in the second layer, and the updated weight is determined based on the weight corresponding to the first neighborhood point.
[0352] In some implementations, the updated weight is determined based on the weight corresponding to the first neighborhood point and the first prediction weight.
[0353] In some implementations, the updated weight is determined based on the product of the weight corresponding to the first neighborhood point and the first prediction weight.
[0354] In some implementations, the weight corresponding to the first neighborhood point is related to the number of second neighborhood points, and the second neighborhood points are neighborhood points of the first neighborhood point; or, the weight corresponding to the first neighborhood point is determined based on the weight corresponding to the LOD layer where the first neighborhood point is located.
[0355] In some implementations, the weight corresponding to the first neighborhood point is equal to the number of the second neighborhood points; or, the weight corresponding to the first neighborhood point is equal to the product of the number of the second neighborhood points and the first prediction weight.
[0356] In some implementations, the second neighborhood point is a point predicted based on the first neighborhood point.
[0357] In some implementations, the first neighborhood point is a point predicted based on the first point to be encoded.
[0358] In some implementations, the second determination unit 2520 is configured to predict the second point to be encoded based on the shift coefficient of the third neighborhood point and the second prediction weight, where the second point to be encoded is any point in the second layer, the third neighborhood point is a neighborhood point of the second point to be encoded, and the second prediction weight is determined based on the weight corresponding to the third neighborhood point.
[0359] In some implementations, the second prediction weight is equal to the weight corresponding to the third neighborhood point.
[0360] In some implementations, the weight corresponding to the third neighborhood point is associated with the number of fourth neighborhood points, and the fourth neighborhood point is a neighborhood point of the third neighborhood point; or, the weight corresponding to the third neighborhood point is determined based on the weight corresponding to the LOD layer where the third neighborhood point is located.
[0361] In some implementations, the weight corresponding to the third neighborhood point is equal to the number of the fourth neighborhood points; or, the weight corresponding to the third neighborhood point is equal to the product of the number of the fourth neighborhood points and the first prediction weight.
[0362] In some implementations, the fourth neighborhood point is a point predicted based on the third neighborhood point.
[0363] In some implementations, the third neighborhood point is a point in the first layer that is collinear with the second point to be encoded.
[0364] In some implementations, the encoder 2500 further includes a quantization unit configured to quantize the transform coefficient of the third point to be encoded according to a weight corresponding to the third point to be encoded, wherein the third point to be encoded is any point in the three-dimensional network.
[0365] In some implementations, the quantization unit is configured to: quantize the transform coefficient of the third point to be encoded according to a first quantization parameter, where the first quantization parameter is determined based on a weight corresponding to the third point to be encoded.
[0366] In some implementations, the first quantization parameter is determined based on the weight corresponding to the third point to be encoded and the second quantization parameter, wherein the third point to be encoded is located in the target LOD layer among the multiple LOD layers, and the second quantization parameter is the quantization parameter corresponding to the target LOD layer.
[0367] In some implementations, the second quantization parameter is determined based on an initial quantization parameter and a quantization adjustment parameter.
[0368] In some implementations, the greater the weight corresponding to the third to-be-encoded point, the smaller the value of the first quantization parameter.
[0369] In some implementations, the weight corresponding to the third point to be encoded is related to the number of fifth neighborhood points, and the fifth neighborhood points are neighborhood points of the third point to be encoded.
[0370] In some implementations, the weight corresponding to the third point to be encoded is equal to the number of the fifth neighborhood points; or, the weight corresponding to the third point to be encoded is equal to the product of the number of the fifth neighborhood points and the first prediction weight.
[0371] In some implementations, the fifth neighborhood point is a point predicted based on the third point to be encoded.
[0372] In some implementations, the value of the first prediction weight is a predefined fixed value.
[0373] In some implementations, the value of the first prediction weight is 0.5.
[0374] In some implementations, the encoder 2500 also includes a third determination unit, configured to divide the basic grid according to the geometric information of the points in the basic grid, and determine the geometric information of the points in the subdivided grid; and determine the shift coefficient of the points in the three-dimensional grid according to the geometric information of the points in the subdivided grid and the original geometric information of the three-dimensional grid.
[0375] It is understandable that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and of course it can also be a module, or it can be non-modular. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.
[0376] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, ROM, RAM, a magnetic disk, or an optical disk.
[0377] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the encoder 2500. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the decoding method described in any one of the aforementioned embodiments.
[0378] Based on the composition of the above-mentioned encoder and the computer-readable storage medium, refer to Figure 26, which shows a specific hardware structure diagram of the encoder 2600 provided in an embodiment of the present application. As shown in Figure 26, the encoder 2600 may include: a communication interface 2610, a memory 2620 and a processor 2630; each component is coupled together through a bus system 2640. It can be understood that the bus system 2640 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 2640 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 2640 in Figure 26. Among them,
[0379] Communication interface 2610, used for sending and receiving signals when sending and receiving information with other external network elements;
[0380] Memory 2620, for storing computer programs;
[0381] The processor 2630 is configured to, when running the computer program, execute:
[0382] Performing LOD division on points in the three-dimensional grid to determine a plurality of LOD layers, wherein the plurality of LOD layers include a first layer and a second layer adjacent to each other, and the first layer is a layer above the second layer;
[0383] Predicting the shift coefficients of the points in the second layer to determine the residual coefficients of the points in the second layer;
[0384] The shift coefficients of the points in the first layer are transformed according to the residual coefficients of the points in the second layer.
[0385] It will be appreciated that the memory 2620 in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a ROM, PROM, EPROM, EEPROM, or flash memory. The volatile memory may be a RAM, which serves as an external cache. By way of example and not limitation, many forms of RAM are available, such as SRAM, DRAM, SDRAM, DDRSDRAM, ESDRAM, SLDRAM, and DRRAM. The memory 2620 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0386] Processor 2630 may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be performed by hardware integrated logic circuits or software instructions within processor 2630. Processor 2630 may be a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules within the decoding processor. The software modules can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 2620. Processor 2630 reads information from memory 2620 and, in conjunction with its hardware, completes the steps of the above method.
[0387] It is understood that the embodiments described herein can be implemented with hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more ASICs, DSPs, DSPDs, PLDs, FPGAs, general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or a combination thereof. For software implementation, the technology described herein can be implemented by modules (e.g., processes, functions, etc.) that perform the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0388] Optionally, as another embodiment, the processor 2630 is further configured to execute the encoding method described in any one of the aforementioned embodiments when running the computer program.
[0389] An embodiment of the present application also provides a computer-readable storage medium, which is a non-volatile computer-readable storage medium for storing a bit stream. The bit stream can be generated by an encoding method of an encoder, or the bit stream can be decoded by a decoding method of a decoder, wherein the decoding method can be the decoding method described in any of the foregoing embodiments, and the encoding method can be the encoding method described in any of the foregoing embodiments.
[0390] It should be noted that, in this application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0391] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0392] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0393] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0394] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0395] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A decoding method, applied to a decoder, comprising: Analyzing a bitstream to determine quantization coefficients of points in a three-dimensional grid; Performing inverse quantization on the quantization coefficients of the points in the three-dimensional grid to determine inverse quantization coefficients of the points in the three-dimensional grid; Performing inverse transformation on the inverse quantization coefficients of the points in the first layer to determine shift coefficients of the points in the first layer; Predicting shift coefficients of points in the second layer according to the shift coefficients of the points in the first layer; Wherein, the points in the three-dimensional grid belong to multiple levels of detail (LOD) layers, the first layer and the second layer are two adjacent layers among the multiple LOD layers, and the first layer is the upper layer of the second layer.
2. The method according to claim 1, wherein The performing inverse transformation on the inverse quantization coefficients of the points in the first layer includes: Performing inverse transformation on the inverse quantization coefficient of a first point to be decoded in the first layer according to the inverse quantization coefficients of first neighboring points and an update weight, where the first neighboring points are neighboring points of the first point to be decoded, and the update weight is determined based on the weights corresponding to the first neighboring points.
3. The method according to claim 2, wherein The update weight is determined based on the weights corresponding to the first neighboring points and a first prediction weight.
4. The method according to claim 3, wherein, The update weight is determined based on the product of the weights corresponding to the first neighboring points and the first prediction weight.
5. The method according to any one of claims 2 to 4, wherein The weight corresponding to the first neighboring points is related to the number of second neighboring points, and the second neighboring points are neighboring points of the first neighboring points; or, the weight corresponding to the first neighboring points is determined based on the weight corresponding to the LOD layer where the first neighboring points are located.
6. The method according to claim 5, wherein: The weight corresponding to the first neighboring points is equal to the number of the second neighboring points; or, The weight corresponding to the first neighboring points is equal to the product of the number of the second neighboring points and the first prediction weight.
7. The method according to claim 5 or 6, wherein The second neighboring points are points predicted based on the first neighboring points.
8. The method according to any one of claims 2 to 7, wherein, The first neighboring points are points predicted based on the first point to be decoded.
9. The method according to any one of claims 1 to 8, wherein The predicting shift coefficients of points in the second layer according to the shift coefficients of the points in the first layer includes: Predicting a second point to be decoded according to the shift coefficients of third neighboring points and a second prediction weight, where the second point to be decoded is any point in the second layer, the third neighboring points are neighboring points of the second point to be decoded in the first layer, and the second prediction weight is determined based on the weights corresponding to the third neighboring points.
10. The method according to claim 9, wherein, The second prediction weight is equal to the weight corresponding to the third neighboring points.
11. The method according to claim 9 or 10, wherein, The weight corresponding to the third neighboring points is associated with the number of fourth neighboring points, and the fourth neighboring points are neighboring points of the third neighboring points; or, the weight corresponding to the third neighboring points is determined based on the weight corresponding to the LOD layer where the third neighboring points are located.
12. The method according to claim 11, wherein: The weight corresponding to the third neighboring points is equal to the number of the fourth neighboring points; or, The weight corresponding to the third neighboring points is equal to the product of the number of the fourth neighboring points and the first prediction weight.
13. The method according to claim 11 or 12, wherein, The fourth neighboring points are points predicted based on the third neighboring points.
14. The method according to any one of claims 9 to 13, wherein, The third neighboring points are points collinear with the second point to be decoded in the first layer.
15. The method according to any one of claims 1 to 14, wherein The inverse quantization of the quantization coefficients of the points in the three-dimensional grid includes: Inverse quantizing the quantization coefficients of the third point to be decoded according to the weight corresponding to the third point to be decoded, where the third point to be decoded is any point in the three-dimensional network.
16. The method according to claim 15, wherein, The inverse quantization of the quantization coefficients of the third point to be decoded according to the weight corresponding to the third point to be decoded includes: Inverse quantizing the quantization coefficients of the third point to be decoded according to a first quantization parameter, where the first quantization parameter is determined based on the weight corresponding to the third point to be decoded.
17. The method according to claim 16, wherein, The first quantization parameter is determined based on the weight corresponding to the third point to be decoded and a second quantization parameter, where the third point to be decoded is located in a target LOD layer among the multiple LOD layers, and the second quantization parameter is the quantization parameter corresponding to the target LOD layer.
18. The method according to claim 17, wherein, The second quantization parameter is determined based on an initial quantization parameter and a quantization adjustment parameter.
19. The method according to any one of claims 16 to 18, wherein The greater the weight corresponding to the third point to be decoded, the smaller the value of the first quantization parameter.
20. The method according to any one of claims 15 to 19, wherein, The weight corresponding to the third point to be decoded is related to the number of fifth neighborhood points, where the fifth neighborhood points are the neighborhood points of the third point to be decoded.
21. The method according to claim 20, wherein: The weight corresponding to the third point to be decoded is equal to the number of the fifth neighborhood points; or, The weight corresponding to the third point to be decoded is equal to the product of the number of the fifth neighborhood points and a first prediction weight.
22. The method according to claim 20 or 21, wherein The fifth neighborhood points are the points predicted based on the third point to be decoded.
23. The method according to claim 3, 4, 6, 12 or 21, wherein, The value of the first prediction weight is a predefined fixed value.
24. The method according to claim 23, wherein, The value of the first prediction weight is 0.
5.
25. The method according to any one of claims 1 to 24, wherein, The method further includes: Analyzing the code stream to determine the geometric information of the points in the base grid; Dividing the base grid according to the geometric information of the points in the base grid to determine the geometric information of the points in the subdivided grid; Determining the reconstructed geometric information of the points in the three-dimensional grid according to the shift coefficients of the points in the three-dimensional grid and the geometric information of the points in the subdivided grid.
26. An encoding method applied to an encoder, including: Performing a level of detail (LOD) division on the points in a three-dimensional grid to determine multiple LOD layers, where the multiple LOD layers include adjacent first and second layers, and the first layer is the upper layer of the second layer; Predicting the shift coefficients of the points in the second layer to determine the residual coefficients of the points in the second layer; Transforming the shift coefficients of the points in the first layer according to the residual coefficients of the points in the second layer.
27. The method according to claim 26, wherein The transforming the shift coefficients of the points in the first layer according to the residual coefficients of the points in the second layer includes: Transforming the shift coefficients of the first point to be encoded in the first layer according to the residual coefficients of the first neighborhood points and an updated weight, where the first neighborhood points are the neighborhood points of the first point to be encoded in the second layer, and the updated weight is determined based on the weight corresponding to the first neighborhood points.
28. The method according to claim 27, wherein, The updated weight is determined based on the weight corresponding to the first neighborhood points and a first prediction weight.
29. The method according to claim 28, wherein The updated weight is determined based on the product of the weight corresponding to the first neighborhood points and the first prediction weight.
30. The method according to any one of claims 27 to 29, wherein The weight corresponding to the first neighborhood point is related to the number of second neighborhood points, where the second neighborhood points are the neighborhood points of the first neighborhood point; or, the weight corresponding to the first neighborhood point is determined based on the weight corresponding to the LOD layer where the first neighborhood point is located.
31. The method according to claim 30, wherein: The weight corresponding to the first neighborhood point is equal to the number of second neighborhood points; or, The weight corresponding to the first neighborhood point is equal to the product of the number of second neighborhood points and the first prediction weight.
32. The method according to claim 30 or 31, wherein, The second neighborhood points are the points predicted based on the first neighborhood point.
33. The method according to any one of claims 27 to 32, wherein The first neighborhood point is the point predicted based on the first point to be encoded.
34. The method according to any one of claims 26 to 33, wherein, The predicting the shift coefficient of the points in the second layer includes: Predicting a second point to be encoded according to the shift coefficient of the third neighborhood points and the second prediction weight, where the second point to be encoded is any point in the second layer, the third neighborhood points are the neighborhood points of the second point to be encoded, and the second prediction weight is determined based on the weight corresponding to the third neighborhood points.
35. The method according to claim 34, wherein The second prediction weight is equal to the weight corresponding to the third neighborhood points.
36. The method according to claim 34 or 35, wherein, The weight corresponding to the third neighborhood points is associated with the number of fourth neighborhood points, where the fourth neighborhood points are the neighborhood points of the third neighborhood points; or, the weight corresponding to the third neighborhood points is determined based on the weight corresponding to the LOD layer where the third neighborhood points are located.
37. The method according to claim 36, wherein: The weight corresponding to the third neighborhood points is equal to the number of fourth neighborhood points; or, The weight corresponding to the third neighborhood points is equal to the product of the number of fourth neighborhood points and the first prediction weight.
38. The method according to claim 36 or 37, wherein, The fourth neighborhood points are the points predicted based on the third neighborhood point.
39. The method according to any one of claims 34 to 38, wherein, The third neighborhood points are the points in the first layer that are collinear with the second point to be encoded.
40. The method according to any one of claims 26 to 39, wherein, The method further includes: Quantizing the transform coefficient of the third point to be encoded according to the weight corresponding to the third point to be encoded, where the third point to be encoded is any point in the three-dimensional network.
41. The method according to claim 40, wherein, The quantizing the transform coefficient of the third point to be encoded according to the weight corresponding to the third point to be encoded includes: Quantizing the transform coefficient of the third point to be encoded according to a first quantization parameter, where the first quantization parameter is determined based on the weight corresponding to the third point to be encoded.
42. The method according to claim 41, wherein, The first quantization parameter is determined based on the weight corresponding to the third point to be encoded and a second quantization parameter, where the third point to be encoded is located in a target LOD layer among the multiple LOD layers, and the second quantization parameter is the quantization parameter corresponding to the target LOD layer.
43. The method according to claim 42, wherein The second quantization parameter is determined based on an initial quantization parameter and a quantization adjustment parameter.
44. The method according to any one of claims 41 to 43, wherein The greater the weight corresponding to the third point to be encoded, the smaller the value of the first quantization parameter.
45. The method according to any one of claims 40 to 44, wherein, The weight corresponding to the third point to be encoded is related to the number of fifth neighborhood points, where the fifth neighborhood points are the neighborhood points of the third point to be encoded.
46. The method according to claim 45, wherein: The weight corresponding to the third point to be encoded is equal to the number of fifth neighborhood points; or, The weight corresponding to the third point to be coded is equal to the product of the number of the fifth neighborhood points and the first prediction weight.
47. The method according to claim 45 or 46, wherein, The fifth neighborhood points are the points predicted based on the third point to be coded.
48. The method according to claim 28, 29, 31, 37 or 46, wherein, The value of the first prediction weight is a predefined fixed value.
49. The method according to claim 48, wherein, The value of the first prediction weight is 0.
5.
50. The method according to any one of claims 26 to 49, wherein The method further includes: Dividing the basic grid according to the geometric information of the points in the basic grid to determine the geometric information of the points in the subdivided grid; Determining the shift coefficient of the points in the three-dimensional grid according to the geometric information of the points in the subdivided grid and the original geometric information of the three-dimensional grid.
51. A decoder, comprising: A first determination unit configured to parse a bitstream to determine the quantization coefficient of the points in the three-dimensional grid; A second determination unit configured to perform inverse quantization on the quantization coefficient of the points in the three-dimensional grid to determine the inverse quantization coefficient of the points in the three-dimensional grid; A third determination unit configured to perform inverse transformation on the inverse quantization coefficient of the points in the first layer to determine the shift coefficient of the points in the first layer; A prediction unit configured to predict the shift coefficient of the points in the second layer according to the shift coefficient of the points in the first layer; Wherein, the points in the three-dimensional grid belong to multiple levels of detail LOD layers, the first layer and the second layer are two adjacent layers among the multiple LOD layers, and the first layer is the upper layer of the second layer.
52. A decoder, the decoder comprising: A memory for storing a computer program; A processor for, when running the computer program, executing the method according to any one of claims 1 to 25.
53. An encoder, comprising: A first determination unit configured to perform levels of detail LOD division on the points in the three-dimensional grid to determine multiple LOD layers, the multiple LOD layers including adjacent first and second layers, and the first layer is the upper layer of the second layer; A second determination unit configured to predict the shift coefficient of the points in the second layer to determine the residual coefficient of the points in the second layer; Coefficient; A transformation unit configured to transform the shift coefficient of the points in the first layer according to the residual coefficient of the points in the second layer.
54. An encoder, the encoder comprising: A memory for storing a computer program; A processor for, when running the computer program, executing the method according to any one of claims 16 to 50.
55. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program, and when the computer program is executed, it implements the method according to any one of claims 1 to 25, or the method according to any one of claims 26 to 50.
56. A non-volatile computer-readable storage medium storing a bitstream, the bitstream being generated by using an encoding method of an encoder, or the bitstream being decoded by using a decoding method of a decoder, wherein, The decoding method is the method according to any one of claims 1 to 25, and the encoding method is the method according to any one of claims 26 to 50.
Citation Information
Patent Citations
Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
US20210409769A1
Transform method, inverse transform method, encoder, decoder and storage medium
US20220207781A1
Point cloud encoding and decoding method and system, and point cloud encoder and point cloud decoder
US20230232004A1
Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method
WO2021045603A1