Encoding method, decoding method, encoding device, and decoding device

By generating predicted values ​​through inter-frame prediction and encoding the prediction residuals, the problem of low encoding and decoding efficiency of 3D mesh data is solved, achieving more efficient encoding and decoding and optimizing the processing of displacement vectors.

CN121942017APending Publication Date: 2026-04-28PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
Filing Date
2024-10-08
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, there is room for improvement in the encoding and decoding of 3D mesh data, especially in the efficiency of encoding and decoding displacement vectors.

Method used

Inter-frame prediction is used to generate predicted values, and the prediction residual is encoded by an encoding device and decoded by a decoding device. The predicted value of the first displacement data is generated using inter-frame prediction, and it is determined whether to use inter-frame prediction to optimize the encoding and decoding process.

Benefits of technology

It improves the efficiency of encoding and decoding three-dimensional mesh data, reduces the amount of code, appropriately processes the encoded and decoded data, and improves the encoding and decoding processing of displacement vectors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121942017A_ABST
    Figure CN121942017A_ABST
Patent Text Reader

Abstract

In a method for encoding displacement data of a three-dimensional point, a predicted value of first displacement data, which is the displacement data to be encoded, is generated by performing inter-frame prediction using second displacement data at a different time from the first displacement data (step S9101), a predicted residual is generated using the first displacement data and the generated predicted value (step S9102), and the predicted residual is encoded by performing inter-frame prediction using the second displacement data. And encoding the generated prediction residual (step S9103).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to encoding methods, etc. Background Technology

[0002] Patent document 1 proposes a method and apparatus for encoding and decoding three-dimensional mesh data.

[0003] Existing technical documents Patent documents Patent document 1: Japanese Patent Application Publication No. 2006-187015. Summary of the Invention

[0004] The problem that the invention aims to solve Further improvements are desired in the processing of encoding or decoding related to displacement vectors. The purpose of this disclosure is to improve the processing of encoding or decoding related to displacement vectors.

[0005] Methods for solving problems One aspect of the encoding method of the present invention is a method for encoding displacement data of three-dimensional points. It involves performing inter-frame prediction using second displacement data to generate a predicted value of first displacement data. The first displacement data is the displacement data to be encoded, and the second displacement data is data from a different time than the first displacement data. A prediction residual is generated using the first displacement data and the generated prediction value, and the generated prediction residual is encoded.

[0006] Furthermore, these general or specific methods can be implemented through systems, devices, integrated circuits, computer programs, or recording media such as computer-readable CD-ROMs, or through any combination of systems, devices, integrated circuits, computer programs, and recording media.

[0007] Invention Effects This disclosure helps to improve coding processes related to displacement vectors, etc. Attached Figure Description

[0008] Figure 1 This is a conceptual diagram representing a three-dimensional mesh used in the implementation.

[0009] Figure 2 This is a conceptual diagram representing the basic elements of a three-dimensional mesh used in the implementation method.

[0010] Figure 3 This is a conceptual diagram representing the mapping of implementation methods.

[0011] Figure 4 This is a block diagram illustrating a structural example of an encoding / decoding system according to an implementation method.

[0012] Figure 5This is a block diagram illustrating a structural example of an encoding device according to an implementation method.

[0013] Figure 6 This is a block diagram illustrating another structural example of the encoding device in an embodiment.

[0014] Figure 7 This is a block diagram illustrating a structural example of a decoding device according to an implementation method.

[0015] Figure 8 This is a block diagram illustrating another structural example of the decoding device in an embodiment.

[0016] Figure 9 This is a conceptual diagram representing a structural example of a bitstream implementation.

[0017] Figure 10 This is a conceptual diagram representing another structural example of a bitstream implementation.

[0018] Figure 11 This is a conceptual diagram representing another structural example of a bitstream implementation.

[0019] Figure 12 This is a block diagram illustrating a specific example of an encoding / decoding system implemented in this way.

[0020] Figure 13 This is a conceptual diagram representing a structural example of point group data in an implementation method.

[0021] Figure 14 This is a conceptual diagram representing a data file example of point group data in an implementation method.

[0022] Figure 15 This is a conceptual diagram representing a structural example of grid data in an implementation method.

[0023] Figure 16 This is a conceptual diagram representing a data file example of grid data in an implementation method.

[0024] Figure 17 This is a conceptual diagram representing the types of three-dimensional data used in the implementation.

[0025] Figure 18 This is a block diagram illustrating a structural example of a three-dimensional data encoder according to an implementation method.

[0026] Figure 19 This is a block diagram illustrating a structural example of a three-dimensional data decoder implementation.

[0027] Figure 20 This is a block diagram illustrating another structural example of a three-dimensional data encoder according to an implementation method.

[0028] Figure 21 This is a block diagram illustrating another structural example of a three-dimensional data decoder implementation.

[0029] Figure 22 This is a conceptual diagram representing a specific example of the encoding process in the implementation method.

[0030] Figure 23 This is a conceptual diagram representing a specific example of the decoding process in the implementation method.

[0031] Figure 24 This is a block diagram illustrating an implementation example of the encoding device in the embodiment.

[0032] Figure 25 This is a block diagram illustrating an implementation example of the decoding device in the embodiment.

[0033] Figure 26 This is a block diagram illustrating another structural example of an encoding / decoding system according to an implementation method.

[0034] Figure 27 This is a block diagram illustrating another structural example of the encoding device in an embodiment.

[0035] Figure 28 This is a block diagram illustrating another structural example of the decoding device in an embodiment.

[0036] Figure 29 This is a block diagram illustrating another structural example of the encoding device in an embodiment.

[0037] Figure 30 This is a block diagram illustrating another structural example of the decoding device in an embodiment.

[0038] Figure 31 This is a flowchart illustrating the processing of the encoding device in the implementation method.

[0039] Figure 32 This is an explanatory diagram that conceptually represents the encoding of a grid frame in an implementation method.

[0040] Figure 33 This is a flowchart illustrating the processing of the decoding device in the implementation method.

[0041] Figure 34 This is an explanatory diagram conceptually representing the decoding of a grid frame in an implementation method.

[0042] Figure 35 This is a block diagram illustrating a structural example of a decoding device according to an implementation method.

[0043] Figure 36 This is a block diagram illustrating a structural example of a decoding device according to an implementation method.

[0044] Figure 37 This is an explanatory diagram showing an example of sub-segmentation in an implementation method.

[0045] Figure 38This is an explanatory diagram illustrating an example of the displacement of vertices after subdivision and displacement in an implementation method.

[0046] Figure 39 This is an explanatory diagram showing an example of the vertices of the original mesh in the implementation method.

[0047] Figure 40 This is an explanatory diagram showing an example of a grid used in an implementation method.

[0048] Figure 41 This is an explanatory diagram illustrating an example of grid-to-subgrid division in an implementation method.

[0049] Figure 42 This is a first explanatory diagram illustrating an example of packing displacement information into an image frame in an embodiment.

[0050] Figure 43 This is a second explanatory diagram illustrating an example of packing displacement information into an image frame in an embodiment.

[0051] Figure 44 This is a third explanatory diagram illustrating an example of packing displacement information into an image frame in an embodiment.

[0052] Figure 45 This is a block diagram illustrating a detailed structural example of the decoding device in an embodiment.

[0053] Figure 46 This is an explanatory diagram showing the coordinates of vertices in a three-dimensional mesh of the implementation method.

[0054] Figure 47 This is an explanatory diagram showing the prediction information of the implementation method.

[0055] Figure 48 This is a block diagram illustrating a structural example of an encoding device according to an implementation method.

[0056] Figure 49 This is a flowchart illustrating a specific example of the encoding process in the implementation method.

[0057] Figure 50 This is a block diagram illustrating a structural example of a decoding device according to an implementation method.

[0058] Figure 51 This is a flowchart illustrating a specific example of the decoding process in the implementation method.

[0059] Figure 52 This is a flowchart illustrating a specific example of the encoding process in the implementation method.

[0060] Figure 53 This is a flowchart illustrating a specific example of the decoding process in the implementation method.

[0061] Figure 54This is an explanatory diagram illustrating an example of the syntax for implementing the method.

[0062] Figure 55 This is an explanatory diagram illustrating an example of the syntax for implementing the method.

[0063] Figure 56 This is an explanatory diagram illustrating an example of the syntax for implementing the method.

[0064] Figure 57 This is a block diagram illustrating a structural example of an encoding device according to an implementation method.

[0065] Figure 58 This is a block diagram illustrating a structural example of a decoding device according to an implementation method.

[0066] Figure 59 This is an explanatory diagram showing the positional relationship of three-dimensional points in the implementation method.

[0067] Figure 60 This is an explanatory diagram illustrating the method for generating LoD in an implementation scheme.

[0068] Figure 61 This is an explanatory diagram illustrating the method for generating LoD in an implementation scheme.

[0069] Figure 62 This is an explanatory diagram illustrating a method for generating predicted values ​​of displacement vectors in an implementation scheme.

[0070] Figure 63 This is an explanatory diagram illustrating an example of calculating the predicted value in an implementation method.

[0071] Figure 64 This is an explanatory diagram illustrating an example of calculating the transformation coefficients in an implementation method.

[0072] Figure 65 This is an explanatory diagram illustrating an example of inter-frame prediction of the transform coefficients in an implementation method.

[0073] Figure 66 This is an explanatory diagram illustrating an example of the syntax for implementing the method.

[0074] Figure 67 This is an explanatory diagram illustrating an example of the syntax for implementing the method.

[0075] Figure 68 This is an explanatory diagram illustrating an example of the calculation of the predicted residual in an implementation method.

[0076] Figure 69 This is an explanatory diagram illustrating an example of the calculation of the predicted residual in an implementation method.

[0077] Figure 70 This is an explanatory diagram illustrating an example of predicted value information for the displacement vector in an implementation method.

[0078] Figure 71 This is an explanatory diagram illustrating a method for generating predicted values ​​of displacement vectors in an implementation scheme.

[0079] Figure 72 This is an explanatory diagram illustrating an example of predicted value information for the displacement vector in an implementation method.

[0080] Figure 73 This is an explanatory diagram illustrating an example of predicted value information for the displacement vector in an implementation method.

[0081] Figure 74 This is an explanatory diagram illustrating an example of predicted value information for the displacement vector in an implementation method.

[0082] Figure 75 This is an explanatory diagram illustrating a method for generating predicted values ​​of displacement vectors in an implementation scheme.

[0083] Figure 76 This is an explanatory diagram illustrating an example of predicted value information for the displacement vector in an implementation method.

[0084] Figure 77 This is an explanatory diagram illustrating an example of predicted value information for motion vectors in an implementation method.

[0085] Figure 78 This is an explanatory diagram illustrating an example of predicted value information for motion vectors in an implementation method.

[0086] Figure 79 This is an explanatory diagram illustrating an example of predicted value information for motion vectors in an implementation method.

[0087] Figure 80 This is an explanatory diagram showing an example of an encoded object point in an implementation method.

[0088] Figure 81 This is an explanatory diagram illustrating an example of predicted value information for the displacement vector in an implementation method.

[0089] Figure 82 This is an explanatory diagram showing an example of a temporal dv implementation.

[0090] Figure 83 This is an explanatory diagram illustrating an example of predicted value information for the displacement vector in an implementation method.

[0091] Figure 84 This is an explanatory diagram showing an example of a reference target for a DVG implementation.

[0092] Figure 85 This is an explanatory diagram illustrating an example of the syntax for implementing the method.

[0093] Figure 86This is an explanatory diagram showing an example of a reference target for a DVG implementation.

[0094] Figure 87 This is an explanatory diagram showing an example of a reference target for a DVG implementation.

[0095] Figure 88 This is an explanatory diagram showing an example of a reference target for a DVG implementation.

[0096] Figure 89 This is an explanatory diagram illustrating an example of the syntax for implementing the method.

[0097] Figure 90 This is an explanatory diagram illustrating an example of the syntax for implementing the method.

[0098] Figure 91 This is a flowchart illustrating the encoding process of the implementation method.

[0099] Figure 92 This is a flowchart illustrating the decoding process of the implementation method. Detailed Implementation

[0100] <Summary of this disclosure> For example, three-dimensional (3D) meshes are used in computer graphics. For instance, a computer graphics image might consist of multiple frames that are distinct in time, each represented by a 3D mesh.

[0101] Furthermore, a 3D mesh consists of vertex information representing the positions of multiple vertices in 3D space, connection information representing the connections between these vertices, and attribute information representing the properties of each vertex or face. Each face is constructed based on the connection relationships between multiple vertices. Such a 3D mesh can be used to represent various computer graphics and images.

[0102] Furthermore, efficient encoding and decoding of 3D meshes are desired for transmission and storage. Arithmetic encoding and decoding can be used for efficient encoding and decoding of 3D meshes.

[0103] Further improvements are desired in the processing of encoding or decoding related to three-dimensional data. The purpose of this disclosure is to improve the processing of encoding or decoding related to three-dimensional data.

[0104] The inventions obtained from the disclosure of this specification will be illustrated below, and the effects obtained by the inventions will be explained.

[0105] (1) An encoding method is a method for encoding displacement data of three-dimensional points, which performs inter-frame prediction using second displacement data to generate a prediction value of first displacement data, wherein the first displacement data is the displacement data that is the object of encoding, the second displacement data is data at a different time than the first displacement data, a prediction residual is generated using the first displacement data and the generated prediction value, and the generated prediction residual is encoded.

[0106] According to the above technical solution, the encoding device encodes the prediction residual generated by inter-frame prediction, thereby enabling appropriate encoding of the displacement data of three-dimensional points. By using inter-frame prediction, when the difference between the first displacement data (the object of encoding) and the second displacement data at different times is relatively small, it is possible to reduce the amount of code, thereby potentially improving the encoding process. Thus, the above encoding method can help improve encoding processes related to displacement vectors.

[0107] (2) According to the encoding method described in (1), when generating the predicted value of the first displacement data, it is determined whether to perform the inter-frame prediction when generating the predicted value of the first displacement data. If it is determined that the inter-frame prediction is to be performed, the predicted value of the first displacement data is generated by performing the inter-frame prediction. If it is determined that the inter-frame prediction is not to be performed, the inter-frame prediction is not performed, and the predicted value of the first displacement data is generated.

[0108] According to the above technical solution, when encoding the first displacement data as the encoding object, the encoding device can pre-determine whether to use inter-frame prediction in the generation of the predicted value of the displacement data, and switch whether to use inter-frame prediction in the generation of the predicted value of the displacement data based on the determination result. Thus, for example, if using inter-frame prediction can improve the encoding process, it can be used in the encoding process; if using inter-frame prediction cannot improve (or degrades) the encoding process, it can be not used in the encoding process. Therefore, it is possible to further reduce the amount of code. In this way, the above encoding method can help to further improve encoding processes related to displacement vectors, etc.

[0109] (3) According to the encoding method described in (2), when determining whether to perform the inter-frame prediction, a first sum and a second sum are calculated. The first sum is the sum of the transformation coefficients of the first displacement data, and the second sum is the sum of the prediction residuals when the inter-frame prediction is applied to the transformation coefficients. If the second sum is determined to be less than the first sum, the inter-frame prediction is determined to be performed. If the second sum is determined to be not less than the first sum, the inter-frame prediction is determined not to be performed.

[0110] According to the above technical solution, the encoding device can switch whether to use inter-frame prediction in the generation of the predicted value of the displacement data by comparing the sum of the transform coefficients when inter-frame prediction is applied to the transform coefficients of the first displacement data to be encoded with the sum of the transform coefficients when inter-frame prediction is not used (in other words, the sum of the transform coefficients when inter-frame prediction is not used). Specifically, if the sum of the transform coefficients when it is determined that inter-frame prediction is applied is small, it can be determined that inter-frame prediction is used. Therefore, by making the determination easier, it is possible to reduce the amount of code. In this way, the above encoding method can help improve the encoding processing related to displacement vectors, etc.

[0111] (4) According to the encoding method described in (1), information indicating whether the inter-frame prediction was used in the generation of the prediction residual of the first displacement data is also sent.

[0112] According to the above technical solution, the encoding device transmits information indicating whether inter-frame prediction was used during encoding. This information, in turn, enables the decoding device, which receives and decodes the encoded displacement data, to communicate whether inter-frame prediction was used during encoding. This facilitates proper decoding of the encoded data. Specifically, it helps to decode data encoded using inter-frame prediction and to decode data encoded without inter-frame prediction. Thus, the above encoding method helps improve encoding processes related to displacement vectors.

[0113] (5) According to the encoding method described in (1), the three-dimensional point includes multiple three-dimensional points, and the multiple three-dimensional points belong to one of multiple layers. When generating the predicted value of the first displacement data, for each of the multiple three-dimensional points to which one or more layers belong, it is determined whether to perform the inter-frame prediction when generating the predicted value of the first displacement data for the three-dimensional points belonging to that layer among the multiple three-dimensional points. For the layer that is determined to perform the inter-frame prediction among the multiple layers, the predicted value of the first displacement data is generated by performing the inter-frame prediction. For the layer that is determined not to perform the inter-frame prediction among the multiple layers, the inter-frame prediction is not performed, and the predicted value of the first displacement data is generated.

[0114] According to the above technical solution, when encoding the first displacement data, which is the object of encoding, the encoding device can pre-determine whether to use inter-frame prediction in the generation of the predicted value of the displacement data according to each level to which the three-dimensional point belongs. Based on the determination result, it switches whether to use inter-frame prediction in the generation of the predicted value of the displacement data according to each level to which the three-dimensional point belongs. Therefore, for example, it is possible to use inter-frame prediction in the encoding process of levels where using inter-frame prediction can improve the encoding process, and not to use inter-frame prediction in the encoding process of levels where using inter-frame prediction cannot improve (or degrades) the encoding process. As a result, it is possible to further reduce the amount of code. In this way, the above encoding method can help to further improve the encoding process related to displacement vectors, etc.

[0115] (6) According to the encoding method described in (5), information indicating the level of the inter-frame prediction used in the generation of the prediction residual of the first displacement data is also transmitted.

[0116] According to the above technical solution, the encoding device transmits information indicating whether inter-frame prediction was used at each level corresponding to a 3D point during encoding. For the decoding device that receives and decodes the encoded displacement data, this information can be transmitted to facilitate proper decoding of the encoded data. Specifically, it helps to decode data from levels encoded using inter-frame prediction and to decode data from levels encoded without inter-frame prediction. Thus, the above encoding method helps improve encoding processing related to displacement vectors.

[0117] (7) A decoding method is a method for decoding displacement data of three-dimensional points, which obtains a prediction residual by decoding encoded data, generates a prediction value of first displacement data by performing inter-frame prediction using second displacement data, wherein the first displacement data is the displacement data that is the object of decoding, the second displacement data is data at a different time than the first displacement data, and the first displacement data is generated using the prediction residual and the generated prediction value.

[0118] According to the above technical solution, the decoding device decodes the prediction residual generated by inter-frame prediction, thereby enabling appropriate decoding of the displacement data of three-dimensional points. By using inter-frame prediction, when the difference between the first displacement data (the object of decoding) and the second displacement data at different times is relatively small, it is possible to reduce the amount of code, thereby potentially improving the decoding process. Thus, the above decoding method can help improve decoding processes related to displacement vectors.

[0119] (8) According to the decoding method described in (7), when generating the predicted value of the first displacement data, it is determined whether to perform the inter-frame prediction when generating the predicted value of the first displacement data. If it is determined that the inter-frame prediction is to be performed, the predicted value of the first displacement data is generated by performing the inter-frame prediction. If it is determined that the inter-frame prediction is not to be performed, the inter-frame prediction is not performed, and the predicted value of the first displacement data is generated.

[0120] According to the above technical solution, when decoding the first displacement data, which is the object of decoding, the decoding device can pre-determine whether to use inter-frame prediction in the generation of the predicted value of the displacement data, and switch whether to use inter-frame prediction in the generation of the predicted value of the displacement data based on the determination result. This may further reduce the amount of code. Thus, the above decoding method can help to further improve decoding processing related to displacement vectors.

[0121] (9) According to the decoding method described in (7), information indicating whether the inter-frame prediction was used in the generation of the prediction residual of the first displacement data is also received.

[0122] According to the above technical solution, the decoding device can determine whether the encoding device used inter-frame prediction during encoding by receiving information indicating whether inter-frame prediction was used during encoding. Furthermore, if the encoding device used inter-frame prediction during encoding, the decoding device can use inter-frame prediction to decode the displacement vector; if the encoding device did not use inter-frame prediction during encoding, the decoding device can decode the displacement vector without using inter-frame prediction. Therefore, the decoding device can appropriately decode the data encoded by the encoding device.

[0123] (10) According to the decoding method described in (7), the three-dimensional point includes multiple three-dimensional points, and the multiple three-dimensional points belong to one of multiple layers. When generating the predicted value of the first displacement data, for each of the multiple three-dimensional points to which one or more layers belong, it is determined whether to perform the inter-frame prediction when generating the predicted value of the first displacement data for the three-dimensional points belonging to that layer among the multiple three-dimensional points. For the layer that is determined to perform the inter-frame prediction among the multiple layers, the predicted value of the first displacement data is generated by performing the inter-frame prediction. For the layer that is determined not to perform the inter-frame prediction among the multiple layers, the inter-frame prediction is not performed, and the predicted value of the first displacement data is generated.

[0124] According to the above technical solution, when decoding the first displacement data, which is the object of encoding, the decoding device can pre-determine whether to use inter-frame prediction in the generation of the predicted value of the displacement data according to each level to which the three-dimensional point belongs. Based on the determination result, it switches whether to use inter-frame prediction in the generation of the predicted value of the displacement data according to each level to which the three-dimensional point belongs. This may further reduce the amount of code. Thus, the above decoding method can help to further improve the encoding processing related to displacement vectors.

[0125] (11) According to the decoding method described in (10), information indicating the level of the inter-frame prediction used in the generation of the prediction residual of the first displacement data is also received.

[0126] According to the above technical solution, the decoding device can determine whether the encoding device used inter-frame prediction for each level of the 3D points during encoding by receiving information indicating whether inter-frame prediction was used during encoding. Furthermore, for data from levels where inter-frame prediction was used during encoding, the decoding device can use inter-frame prediction to decode the displacement vectors; for data from levels where inter-frame prediction was not used during encoding, the decoding device can decode the displacement vectors without using inter-frame prediction. Therefore, the decoding device can appropriately decode the data encoded by the encoding device.

[0127] (12) An encoding apparatus for encoding displacement data of three-dimensional points, comprising a memory and a circuit capable of accessing the memory, wherein the circuit generates a prediction value of first displacement data by performing inter-frame prediction using second displacement data during operation, the first displacement data being the displacement data to be encoded, the second displacement data being data at a different time from the first displacement data, generating a prediction residual using the first displacement data and the generated prediction value, and encoding the generated prediction residual.

[0128] The above technical solution achieves the same effect as the above encoding method.

[0129] (13) A decoding apparatus for decoding displacement data of three-dimensional points, comprising a memory and a circuit capable of accessing the memory, wherein the circuit obtains a prediction residual by decoding encoded data during operation, generates a prediction value of first displacement data by performing inter-frame prediction using second displacement data, the first displacement data being the displacement data to be decoded, the second displacement data being data at a different time from the first displacement data, and the first displacement data being generated using the prediction residual and the generated prediction value.

[0130] The above technical solution achieves the same effect as the above decoding method.

[0131] Furthermore, these general or specific methods can be implemented by systems, devices, integrated circuits, computer programs, or recording media such as computer-readable CD-ROMs, or by any combination of systems, devices, integrated circuits, computer programs, or recording media.

[0132] The embodiments will now be described in detail with reference to the accompanying drawings.

[0133] Furthermore, the embodiments described below are general or specific examples. The numerical values, shapes, materials, constituent elements, arrangement positions of constituent elements, connection methods, steps, and order of steps shown in the following embodiments are examples and are not intended to limit the present invention. In addition, any constituent elements in the following embodiments that are not described in the independent claims representing the highest-level concept are described as arbitrary constituent elements.

[0134] (Implementation Method) In this embodiment, the encoding method and the decoding method are described.

[0135] <Performance and Terminology> Here, the following expressions and terms are used.

[0136] (1) Three-dimensional mesh A 3D mesh is a collection of faces, representing, for example, a 3D object. Furthermore, a 3D mesh primarily consists of vertex information, connectivity information, and attribute information. 3D meshes are sometimes represented as polygon meshes or grids. Additionally, 3D meshes can vary over time. A 3D mesh can contain metadata related to vertex, connectivity, and attribute information, as well as other additional information.

[0137] (2) Vertex information Vertex information represents information about vertices. For example, vertex information indicates the position of a vertex in three-dimensional space. Additionally, vertices correspond to the vertices of the faces that make up a three-dimensional mesh. Vertex information is sometimes represented as "geometric information." It can also sometimes be represented as positional information.

[0138] (3) Connection information Connection information represents the connections between vertices. For example, connection information represents the connections between faces or edges used to form a 3D mesh. Connection information is sometimes expressed as "Connectivity". Additionally, connection information can sometimes be expressed as face information.

[0139] (4) Attribute information Attribute information represents the properties of a vertex or face. For example, attribute information represents the color, image, and normal vector of a vertex or face. Attribute information is sometimes represented as "Texture".

[0140] (5) face A face is an element that constitutes a three-dimensional mesh. Specifically, a face is a polygon on a plane in three-dimensional space. For example, a face can be defined as a triangle in three-dimensional space.

[0141] (6) Plane A plane is a two-dimensional plane in three-dimensional space. For example, a polygon can be formed on a plane, and multiple polygons can be formed on multiple planes.

[0142] (7) Bit stream A bitstream corresponds to encoded information. Bitstreams can also be represented as streams, encoded bitstreams, compressed bitstreams, or encoded signals.

[0143] (8) Encoding and Decoding Encoding can be replaced by other representations such as saving, containing, writing, recording, signaling, sending, notifying, storing, or compressing, and these representations can also be interchanged. For example, encoding information can mean including the information in a bitstream. Furthermore, encoding information into a bitstream can mean encoding the information to generate a bitstream containing the encoded information.

[0144] Furthermore, the concept of decoding can be replaced by reading, interpreting, reading, inputting, exporting, obtaining, receiving, extracting, restoring, reconstructing, compressing / decompressing, etc., and these concepts can also be interchanged. For example, decoding information can mean obtaining information from a bitstream. Additionally, decoding information from a bitstream can mean decoding the bitstream itself to obtain the information contained within it.

[0145] (9) Ordinal number In the description, first and second ordinal numbers are sometimes attached to constituent elements, etc. These ordinal numbers can be changed as appropriate. In addition, new ordinal numbers can be assigned to constituent elements, etc., or ordinal numbers can be removed. Furthermore, regarding these ordinal numbers, they are sometimes assigned to elements for the purpose of identifying elements, and sometimes they do not correspond to a meaningful order.

[0146] 3D Mesh Figure 1 This is a conceptual diagram representing a three-dimensional mesh according to this embodiment. The three-dimensional mesh is composed of multiple faces. For example, each face is a triangle. The vertices of these triangles are defined in three-dimensional space. Furthermore, the three-dimensional mesh represents a three-dimensional object. Each face may have color or an image.

[0147] Figure 2This is a conceptual diagram representing the basic elements of the three-dimensional mesh in this embodiment. The three-dimensional mesh consists of vertex information, connection information, and attribute information. Vertex information represents the position of the vertices of a face in three-dimensional space. Connection information represents the connections between vertices. A face can be determined by vertex information and connection information. That is, through vertex information and connection information, a colorless three-dimensional object is formed in three-dimensional space.

[0148] Attribute information can be mapped to vertices or faces. Attribute information corresponding to vertices is sometimes expressed as "Attribute Per Point". Attribute information corresponding to vertices can represent the attributes of the vertex itself or the attributes of the faces connected to the vertex.

[0149] For example, color can be used as attribute information to establish a correspondence with vertices. The color corresponding to a vertex can be the color of the vertex itself, or the color of the face connected to that vertex. The color of a face can be the average of multiple colors corresponding to multiple vertices of the face. Furthermore, normal vectors can be used as attribute information to establish a correspondence with vertices or faces. Such normal vectors can represent the front and back faces of a face.

[0150] Alternatively, a 2D image can be used as attribute information to establish a correspondence with a face. The 2D image corresponding to a face is also represented as a texture image or an "attribute map". Alternatively, information representing the mapping between a face and a 2D image can also be used as attribute information to establish a correspondence with the face. Information representing such a mapping is sometimes represented as mapping information, vertex information of the texture image, texture coordinates, or "attribute UV coordinates".

[0151] Furthermore, information such as color, image, and motion graphics used as attribute information is sometimes represented as "Parametric Space".

[0152] This attribute information allows textures to be reflected in three-dimensional objects. That is, by using vertex information, connectivity information, and attribute information, colored three-dimensional objects can be formed in three-dimensional space.

[0153] In addition, as described above, attribute information corresponds to vertices or faces, but it can also correspond to edges.

[0154] Figure 3 This is a conceptual diagram illustrating the mapping in this embodiment. For example, a region of a two-dimensional image in a two-dimensional plane can be mapped to a surface of a three-dimensional mesh in three-dimensional space. Specifically, the coordinate information of a region in the two-dimensional image is mapped to a surface of the three-dimensional mesh. Thus, an image reflecting the mapped region of the two-dimensional image is generated on the surface of the three-dimensional mesh.

[0155] By using mapping, it is possible to separate the two-dimensional image used as attribute information from the three-dimensional mesh. For example, in the encoding of the three-dimensional mesh, the two-dimensional image can be encoded using image encoding or video encoding methods.

[0156] <System Structure> Figure 4 This is a block diagram illustrating a structural example of the encoding / decoding system according to this embodiment. Figure 4 The encoding and decoding system includes an encoding device 100 and a decoding device 200.

[0157] For example, encoding device 100 acquires a three-dimensional mesh and encodes it into a bitstream. Then, encoding device 100 outputs the bitstream to network 300. For example, the bitstream contains the encoded three-dimensional mesh and control information for decoding the encoded three-dimensional mesh. By encoding the three-dimensional mesh, its information is compressed.

[0158] Network 300 transmits the bit stream from encoding device 100 to decoding device 200. Network 300 can be the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. Network 300 is not necessarily limited to two-way communication; it can also be a one-way communication network used for terrestrial digital broadcasting or satellite broadcasting, etc.

[0159] Alternatively, the Network 300 can be replaced by recording media such as DVD (Digital Versatile Disc) or BD (Blu-Ray Disc (registered trademark)).

[0160] Decoding device 200 acquires the bitstream and decodes the 3D mesh from the bitstream. Through decoding the 3D mesh, its information is decompressed. For example, decoding device 200 decodes the 3D mesh using a decoding method corresponding to the encoding method used by encoding device 100 to encode the 3D mesh. That is, encoding device 100 and decoding device 200 perform encoding and decoding according to their respective encoding and decoding methods.

[0161] Furthermore, the 3D mesh before encoding can also be represented as the original 3D mesh. Additionally, the 3D mesh after decoding can also be represented as a reconstructed 3D mesh.

[0162] <Encoding device> Figure 5 This is a block diagram illustrating a structural example of the encoding apparatus 100 according to this embodiment. For example, the encoding apparatus 100 includes a vertex information encoder 101, a connection information encoder 102, and an attribute information encoder 103.

[0163] Vertex information encoder 101 is a circuit that encodes vertex information. For example, vertex information encoder 101 encodes vertex information into a bit stream according to a specified format for vertex information.

[0164] The connection information encoder 102 is a circuit that encodes connection information. For example, the connection information encoder 102 encodes the connection information into a bit stream according to a specified format.

[0165] The attribute information encoder 103 is a circuit that encodes attribute information. For example, the attribute information encoder 103 encodes the attribute information into a bit stream according to a specified format.

[0166] Vertex information, connectivity information, and attribute information can be encoded using either variable-length encoding or fixed-length encoding. Variable-length encoding can correspond to Huffman coding or context-adaptive binary arithmetic coding (CABAC), etc.

[0167] The vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 can be integrated into one unit. Alternatively, the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 can be further subdivided into multiple constituent elements.

[0168] Figure 6 This is a block diagram illustrating another structural example of the encoding device 100 according to this embodiment. For example, the encoding device 100, in addition to... Figure 5 In addition to the structure shown, it also includes a preprocessor 104 and a postprocessor 105.

[0169] The preprocessor 104 is a circuit that processes vertex information, connectivity information, and attribute information before encoding. For example, the preprocessor 104 may perform transformation, separation, or reuse processing on the 3D mesh before encoding. More specifically, for example, the preprocessor 104 may separate vertex information, connectivity information, and attribute information from the 3D mesh before encoding.

[0170] The post-processor 105 is a circuit that processes the encoded vertex information, connectivity information, and attribute information. For example, the post-processor 105 may perform conversion, separation, or multiplexing processing on the encoded vertex information, connectivity information, and attribute information. More specifically, for example, the post-processor 105 may multiplex the encoded vertex information, connectivity information, and attribute information in a bitstream. Furthermore, for example, the post-processor 105 may further perform variable-length encoding on the encoded vertex information, connectivity information, and attribute information.

[0171] <Decoding device> Figure 7 This is a block diagram illustrating a structural example of the decoding apparatus 200 according to this embodiment. For example, the decoding apparatus 200 includes a vertex information decoder 201, a connection information decoder 202, and an attribute information decoder 203.

[0172] Vertex information decoder 201 is a circuit that decodes vertex information. For example, vertex information decoder 201 decodes vertex information from a bitstream according to a specified format for vertex information.

[0173] The connection information decoder 202 is a circuit that decodes connection information. For example, the connection information decoder 202 decodes the connection information from the bitstream according to a specified format.

[0174] The attribute information decoder 203 is a circuit that decodes attribute information. For example, the attribute information decoder 203 decodes the attribute information from the bitstream according to the format specified for the attribute information.

[0175] Decoding vertex information, connectivity information, and attribute information can use either variable-length decoding or fixed-length decoding. Variable-length decoding can correspond to Huffman coding or context-adaptive binary arithmetic coding (CABAC), etc.

[0176] The vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 can be integrated into one unit. Alternatively, the vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 can be further subdivided into multiple constituent elements.

[0177] Figure 8 This is a block diagram illustrating another structural example of the decoding apparatus 200 according to this embodiment. For example, the decoding apparatus 200, in addition to... Figure 7 In addition to the structure shown, it also includes a preprocessor 204 and a postprocessor 205.

[0178] The preprocessor 204 is a circuit that processes vertex information, connectivity information, and attribute information before decoding. For example, the preprocessor 204 may perform conversion, separation, or multiplexing processing on the bitstream before decoding the vertex information, connectivity information, and attribute information.

[0179] More specifically, for example, the preprocessor 204 may separate the sub-bitstream corresponding to vertex information, the sub-bitstream corresponding to connection information, and the sub-bitstream corresponding to attribute information from the bitstream. Furthermore, for example, the preprocessor 204 may pre-decode the bitstream using variable-length decoding before decoding the vertex information, connection information, and attribute information.

[0180] The post-processor 205 is a circuit that processes the vertex information, connectivity information, and attribute information after decoding. For example, the post-processor 205 may perform conversion, separation, or multiplexing processing on the decoded vertex information, connectivity information, and attribute information. More specifically, for example, the post-processor 205 may reuse the decoded vertex information, connectivity information, and attribute information in a 3D mesh.

[0181] <Bitstream> Vertex information, connectivity information, and attribute information are encoded and stored in a bitstream. The relationship between this information and the bitstream is shown below.

[0182] Figure 9 This is a conceptual diagram illustrating a structural example of the bitstream in this embodiment. In this example, connection information, vertex information, and attribute information are integrated within the bitstream. For example, the connection information, vertex information, and attribute information may be contained in a single file.

[0183] Alternatively, multiple parts of this information can be stored sequentially, such as the first part of connection information, the first part of vertex information, the first part of attribute information, the second part of connection information, the second part of vertex information, the second part of attribute information, and so on. These multiple parts can correspond to multiple parts that are different in time, multiple parts that are different in space, or multiple faces.

[0184] Furthermore, the order in which connection information, vertex information, and attribute information are saved is not limited to the examples above; a different saving order can also be used.

[0185] Figure 10 This is a conceptual diagram illustrating another structural example of the bitstream in this embodiment. In this example, multiple files are contained within the bitstream, with connection information, vertex information, and attribute information stored in different files. Here, a file containing connection information, a file containing vertex information, and a file containing attribute information are shown, but the storage format is not limited to this example. For instance, two types of information—connection information, vertex information, and attribute information—could be contained in one file, and the remaining type of information could be contained in other files.

[0186] Alternatively, this information can be divided into more files for storage. For example, multiple parts of connection information, multiple parts of vertex information, and multiple parts of attribute information can be stored in multiple files. These multiple parts can correspond to different parts in time, different parts in space, or different faces.

[0187] Furthermore, the order in which connection information, vertex information, and attribute information are saved is not limited to the examples above; a different saving order can also be used.

[0188] Figure 11 This is a conceptual diagram illustrating another structural example of the bitstream in this embodiment. In this example, the bitstream is composed of multiple separable sub-bitstreams, with connection information, vertex information, and attribute information stored in different sub-bitstreams.

[0189] Here, sub-bitstreams containing connection information, sub-bitstreams containing vertex information, and sub-bitstreams containing attribute information are shown, but the storage format is not limited to these examples.

[0190] For example, two types of information—connection information, vertex information, and attribute information—can be contained in one sub-bitstream, while the remaining type is contained in other sub-bitstreams. Specifically, attribute information, such as 2D image data, can be stored separately from the sub-bitstreams containing connection information and vertex information, in a sub-bitstream that follows an image encoding scheme.

[0191] Furthermore, each sub-bitstream can contain multiple files. Also, multiple parts of connection information, multiple parts of vertex information, and multiple parts of attribute information can be stored in multiple files.

[0192] Furthermore, the order in which connection information, vertex information, and attribute information are saved is not limited to... Figure 9 , Figure 10 and Figure 11 The example can also use a different saving order than the one described above. For instance, they can be saved in the bitstream in the order of vertex information, connectivity information, and attribute information. Alternatively, they can be saved in the bitstream in any other order, such as connectivity information, attribute information, and vertex information; vertex information, attribute information, and connectivity information; attribute information, connectivity information, and vertex information; or attribute information, vertex information, and connectivity information.

[0193] Alternatively, connection information, vertex information, and attribute information can be divided into multiple data segments, and these multiple data segments can be stored in a bitstream in a periodic or random order.

[0194] <Specific example> Figure 12 This is a block diagram illustrating a specific example of the encoding / decoding system of this embodiment. Figure 12 The encoding and decoding system includes a three-dimensional data encoding system 110, a three-dimensional data decoding system 210, and an external connector 310.

[0195] The 3D data encoding system 110 includes a controller 111, an input / output processor 112, a 3D data encoder 113, a 3D data generator 115, and a system multiplexer 114. The 3D data decoding system 210 includes a controller 211, an input / output processor 212, a 3D data decoder 213, a system demultiplexer 214, a prompter 215, and a user interface 216.

[0196] In the 3D data encoding system 110, sensor data is input from the sensor terminal to the 3D data generator 115. The 3D data generator 115 generates 3D data, such as point group data or mesh data, based on the sensor data and inputs it to the 3D data encoder 113.

[0197] For example, the 3D data generator 115 generates vertex information and corresponding connection and attribute information. The 3D data generator 115 can process the vertex information while generating the connection and attribute information. For example, the 3D data generator 115 can reduce the amount of data by deleting duplicate vertices, or perform vertex information transformations (position offset, rotation, or normalization, etc.). Additionally, the 3D data generator 115 can also render the attribute information.

[0198] In addition, the 3D data generator 115 in Figure 12 The middle part is a component of the three-dimensional data encoding system 110, but it can also be configured externally independently of the three-dimensional data encoding system 110.

[0199] Sensor terminals that provide sensor data for generating 3D data can be, for example, mobile objects such as cars, flying objects such as airplanes, portable terminals, or cameras. Alternatively, distance sensors such as LiDAR, millimeter-wave radar, infrared sensors or rangefinders, stereo cameras, or combinations of multiple monocular cameras can also be used as sensor terminals.

[0200] Sensor data can include the distance (position) of an object, images from a single-lens reflex camera, images from a stereo camera, color, reflectivity, sensor pose, orientation, gyroscope, sensing position (GPS information or altitude), speed, acceleration, sensing time, temperature, air pressure, humidity, or magnetism, etc.

[0201] 3D data encoder 113 corresponds to Figure 5 The encoding device 100 is shown in the figure. For example, the three-dimensional data encoder 113 encodes three-dimensional data to generate encoded data. In addition, the three-dimensional data encoder 113 generates control information in the encoding of the three-dimensional data. Furthermore, the three-dimensional data encoder 113 inputs the encoded data and the control information together into the system multiplexer 114.

[0202] 3D data can be encoded using either geometric methods or video codecs. Geometric encoding can be further categorized as geometry-based encoding, while video codec encoding can be categorized as video-based encoding.

[0203] System multiplexer 114 multiplexes the encoded data and control information input from 3D data encoder 113, generating multiplexed data using a prescribed multiplexing method. Alternatively, system multiplexer 114 may multiplex other media such as images, audio, subtitles, application data, or document files, or reference timing information, together with the encoded data and control information of the 3D data. Furthermore, system multiplexer 114 can multiplex attribute information associated with sensor data or 3D data.

[0204] For example, multiplexed data can be in file form for storage or packet form for transmission. These methods can utilize ISOBMFF or ISOBMFF-based methods. Additionally, MPEG-DASH, MMT, MPEG-2 TS Systems, or RTP can also be used.

[0205] Furthermore, the multiplexed data is output as a transmission signal by the input / output processor 112 to the external connector 310. The multiplexed data can be transmitted as a transmission signal via wired or wireless means. Alternatively, the multiplexed data can be stored in internal memory or a storage device. The multiplexed data can be transmitted to a cloud server via the Internet or stored in an external storage device.

[0206] For example, the transmission or storage of multiplexed data is carried out through methods such as broadcasting or communication, corresponding to the medium used for transmission or storage. As communication protocols, HTTP, FTP, TCP, UDP, IP, or combinations thereof can be used. Furthermore, both pull-type and push-type communication methods can be used.

[0207] Wired transmission can utilize Ethernet (registered trademark), USB, RS-232C, HDMI (registered trademark), or coaxial cable, among others. Wireless transmission can utilize 3GPP (registered trademark), IEEE-standard 3G / 4G / 5G, wireless LAN, Wi-Fi, Bluetooth, or millimeter wave. Furthermore, broadcast methods can include DVB-T2, DVB-S2, DVB-C2, ATSC 3.0, and ISDB-S3.

[0208] In addition, sensor data can be input to the 3D data generator 115 or the system multiplexer 114. Furthermore, 3D data or encoded data can be directly output as a transmission signal to the external connector 310 via the input / output processor 112. The transmission signal output from the 3D data encoding system 110 is input to the 3D data decoding system 210 via the external connector 310.

[0209] In addition, the various actions of the three-dimensional data encoding system 110 can be controlled by the controller 111 that executes the application program.

[0210] In the 3D data decoding system 210, the transmission signal is input to the input / output processor 212. The input / output processor 212 decodes the multiplexed data, which is in file or packet form, from the transmission signal and inputs the multiplexed data to the system demultiplexer 214. The system demultiplexer 214 obtains the encoded data and control information from the multiplexed data and inputs it to the 3D data decoder 213. The system demultiplexer 214 can also extract other media or reference time information from the multiplexed data.

[0211] 3D data decoder 213 corresponds to Figure 7 The decoding device 200 is shown in the figure. For example, the 3D data decoder 213 decodes the 3D data from the encoded data based on a predefined encoding method. Then, the 3D data is presented to the user by the prompter 215.

[0212] Additionally, sensor data and other supplementary information can be input to the prompter 215. The prompter 215 can then provide 3D data based on this supplementary information. Furthermore, user instructions can be input from the user terminal to the user interface 216. The prompter 215 can then provide 3D data based on these input instructions.

[0213] In addition, the input / output processor 212 can obtain three-dimensional data and encoded data from the external connector 310.

[0214] In addition, the various actions of the 3D data decoding system 210 can be controlled by the controller 211 that executes the application program.

[0215] Figure 13 This is a conceptual diagram illustrating a structural example of point group data in this embodiment. Point group data is data representing a group of points of a three-dimensional object.

[0216] Specifically, a point cloud consists of multiple points, possessing positional information representing the three-dimensional coordinates of each point, as well as attribute information representing the attributes of each point. The positional information is also represented by geometric information.

[0217] The categories of attribute information can be, for example, color or reflectivity. A point can be mapped to attribute information of one category, to attribute information of multiple different categories, or to attribute information that has multiple values ​​for the same category.

[0218] Figure 14 This is a conceptual diagram representing an example of a data file for point group data in this embodiment. In this example, a one-to-one correspondence is shown between location information items and attribute information items, illustrating the location and attribute information of N points constituting the point group data. In this example, the location information is information representing three-dimensional coordinate positions using the x, y, and z axes, and the attribute information is information representing colors using RGB. A PLY file or similar can be used as a representative data file for point group data.

[0219] Figure 15 This is a conceptual diagram illustrating a structural example of the mesh data in this embodiment. Mesh data is data used in CG (Computer Graphics) and the like, and is three-dimensional mesh data that represents the three-dimensional shape of an object using multiple faces. Each face is also represented as a polygon, having a polygonal shape such as a triangle or quadrilateral.

[0220] Specifically, in addition to the multiple points that make up the point group, the 3D mesh also consists of multiple edges and multiple faces. Each point is also represented as a vertex or position. Each edge corresponds to a line segment connecting two vertices. Each face corresponds to a region enclosed by three or more edges.

[0221] In addition, 3D meshes possess positional information representing the 3D coordinates of vertices. This positional information is also represented as vertex information or geometric information. Furthermore, 3D meshes possess connectivity information representing the relationships between multiple vertices that constitute an edge or face. This connectivity information is also represented as connectivity. Additionally, 3D meshes possess attribute information representing properties specific to vertices, edges, or faces. Attribute information in 3D meshes is also represented as texture.

[0222] For example, attribute information can represent the color, reflectivity, or normal vector for vertices, edges, or faces. The direction of the normal vector can represent the front and back faces of a face.

[0223] As a data file format for grid data, an object file can also be used.

[0224] Figure 16This is a conceptual diagram representing an example of a data file for mesh data in this embodiment. In this example, the data file contains position information G(1) to G(N) of N vertices constituting a 3D mesh, and attribute information A1(1) to A1(N) of the N vertices. Additionally, in this example, it contains M attribute information A2(1) to A2(M). The items in the attribute information may not correspond one-to-one with vertices or one-to-one with faces. Furthermore, the attribute information may not exist.

[0225] Connection information is represented by a combination of vertex indices. n[1, 3, 4] represents the face of the triangle formed by vertices n=1, n=3, and n=4. Additionally, m[2, 4, 6] indicates that the attribute information for m=2, m=4, and m=6 corresponds to the three vertices respectively.

[0226] Furthermore, the actual content of the attribute information can be recorded in other files. And, a pointer to this content can be associated with vertices or faces. For example, the attribute information representing the image of a face can be stored in a two-dimensional attribute map file. Furthermore, the filename of the attribute map and the two-dimensional coordinate values ​​in the attribute map can be recorded in attribute information A2(1) to A2(M). The method of specifying attribute information for a face is not limited to these methods; any method can be used.

[0227] Figure 17 This is a conceptual diagram representing the types of three-dimensional data in this embodiment. Point cluster data and mesh data can represent static objects or dynamic objects. Static objects are objects that do not change over time, while dynamic objects are objects that change over time. Static objects can correspond to three-dimensional data for any point in time.

[0228] For example, point group data at any given time point is sometimes represented as a PCC frame. Similarly, grid data at any given time point is sometimes represented as a grid frame. Furthermore, both PCC frames and grid frames can sometimes be represented simply as frames.

[0229] Furthermore, the area of ​​the object can be restricted to a certain range, as with typical image data, or it can be unrestricted, as with map data. Additionally, the density of points or areas can be determined in various ways. Sparse point clusters or sparse grid data, or dense point clusters or dense grid data, can be used.

[0230] Next, the encoding and decoding of point groups or 3D meshes will be described. The apparatus, processing, or syntax for encoding and decoding vertex information of 3D meshes in this disclosure can be applied to the encoding and decoding of point groups.

[0231] Furthermore, the apparatus, processing, or syntax for encoding and decoding attribute information of point groups in this disclosure can be applied to the encoding and decoding of connection information or attribute information of three-dimensional meshes.

[0232] Furthermore, at least a portion of the processing can be shared in the encoding and decoding of point group data and grid data. This reduces the size of the circuitry and software programs.

[0233] Figure 18 This is a block diagram illustrating a structural example of the three-dimensional data encoder 113 according to this embodiment. In this example, the three-dimensional data encoder 113 includes a vertex information encoder 121, an attribute information encoder 122, a metadata encoder 123, and a multiplexer 124. The vertex information encoder 121, the attribute information encoder 122, and the multiplexer 124 can correspond to... Figure 6 Vertex information encoder 101, attribute information encoder 103, and post-processor 105, etc.

[0234] Furthermore, in this example, the 3D data encoder 113 encodes the 3D data using a geometry-based encoding method. In this geometry-based encoding method, the 3D structure is taken into account. Moreover, in this geometry-based encoding method, the structural information obtained from the encoding of vertex information is used to encode the attribute information.

[0235] Specifically, firstly, the vertex information, attribute information, and metadata contained in the 3D data generated from the sensor data are input to the vertex information encoder 121, attribute information encoder 122, and metadata encoder 123, respectively. Here, the connectivity information contained in the 3D data can be processed in the same way as the attribute information. Furthermore, in the case of point group data, the position information can be processed as vertex information.

[0236] Vertex information encoder 121 encodes vertex information into compressed vertex information and outputs the compressed vertex information as encoded data to multiplexer 124. In addition, vertex information encoder 121 generates metadata for the compressed vertex information and outputs it to multiplexer 124. Furthermore, vertex information encoder 121 generates structural information and outputs it to attribute information encoder 122.

[0237] The attribute information encoder 122 uses the structural information generated by the vertex information encoder 121 to encode the attribute information into compressed attribute information, and outputs the compressed attribute information as encoded data to the multiplexer 124. In addition, the attribute information encoder 122 generates metadata of the compressed attribute information and outputs it to the multiplexer 124.

[0238] Metadata encoder 123 encodes compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to multiplexer 124. The metadata encoded by metadata encoder 123 can be used for encoding vertex information and attribute information.

[0239] Multiplexer 124 multiplexes compressed vertex information, compressed vertex information metadata, compressed attribute information, compressed attribute information metadata, and compressed metadata into a bitstream. Furthermore, multiplexer 124 inputs the bitstream to the system layer.

[0240] Figure 19 This is a block diagram illustrating a structural example of the 3D data decoder 213 according to this embodiment. In this example, the 3D data decoder 213 includes a vertex information decoder 221, an attribute information decoder 222, a metadata decoder 223, and a demultiplexer 224. The vertex information decoder 221, the attribute information decoder 222, and the demultiplexer 224 can correspond to... Figure 8 Vertex information decoder 201, attribute information decoder 203, and preprocessor 204, etc.

[0241] In this example, the 3D data decoder 213 decodes the 3D data using a geometry-based encoding method. In this geometry-based decoding, the 3D structure is considered. Furthermore, in this geometry-based decoding, the structural information obtained from decoding the vertex information is used to decode the attribute information.

[0242] Specifically, first, the bitstream is input from the system layer to demultiplexer 224. Demultiplexer 224 separates compressed vertex information, compressed vertex information metadata, compressed attribute information, compressed attribute information metadata, and compressed metadata from the bitstream. The compressed vertex information and compressed vertex information metadata are input to vertex information decoder 221. The compressed attribute information and compressed attribute information metadata are input to attribute information decoder 222. The metadata is input to metadata decoder 223.

[0243] Vertex information decoder 221 uses the metadata of compressed vertex information to decode vertex information from compressed vertex information. Furthermore, vertex information decoder 221 generates structure information and outputs it to attribute information decoder 222. Attribute information decoder 222 uses the structure information generated by vertex information decoder 221 and the metadata of compressed attribute information to decode attribute information from compressed attribute information. Metadata decoder 223 decodes metadata from compressed metadata. The metadata decoded by metadata decoder 223 can be used for decoding vertex information and attribute information.

[0244] Subsequently, vertex information, attribute information, and metadata are output as 3D data from the 3D data decoder 213. Additionally, this metadata, for example, is metadata of vertex information and attribute information, which can be used in the application.

[0245] Figure 20 This is a block diagram illustrating another structural example of the 3D data encoder 113 according to this embodiment. In this example, the 3D data encoder 113 includes a vertex image generator 131, an attribute image generator 132, a metadata generator 133, an image encoder 134, a metadata encoder 123, and a multiplexer 124. The vertex image generator 131, the attribute image generator 132, and the image encoder 134 can correspond to... Figure 6 Vertex information encoder 101 and attribute information encoder 103, etc.

[0246] Furthermore, in this example, the 3D data encoder 113 encodes the 3D data using a video-based coding method. In this video-based coding method, multiple 2D images are generated from the 3D data, and these images are then encoded using an image coding method. Here, the image coding method can be HEVC (High Efficiency Video Coding) or VVC (Versatile Video Coding), etc.

[0247] Specifically, firstly, the vertex information and attribute information contained in the 3D data generated from the sensor data are input to the metadata generator 133. Furthermore, the vertex information and attribute information are input to the vertex image generator 131 and the attribute image generator 132, respectively. Additionally, the metadata contained in the 3D data is input to the metadata encoder 123. Here, the connectivity information contained in the 3D data can be processed in the same way as the attribute information. Furthermore, in the case of point group data, the position information can be processed as vertex information.

[0248] Metadata generator 133 generates mapping information for multiple two-dimensional images based on vertex information and attribute information. Furthermore, metadata generator 133 inputs the mapping information into vertex image generator 131, attribute image generator 132, and metadata encoder 123.

[0249] Vertex image generator 131 generates vertex images based on vertex information and mapping information and inputs them to image encoder 134. Attribute image generator 132 generates attribute images based on attribute information and mapping information and inputs them to image encoder 134.

[0250] The image encoder 134 encodes the vertex image and attribute image into compressed vertex information and compressed attribute information respectively according to the image encoding method, and outputs the compressed vertex information and compressed attribute information as encoded data to the multiplexer 124. In addition, the image encoder 134 generates metadata of compressed vertex information and metadata of compressed attribute information and outputs them to the multiplexer 124.

[0251] Metadata encoder 123 encodes compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to multiplexer 124. The compressible metadata contains mapping information. Furthermore, the metadata encoded by metadata encoder 123 can be used for encoding vertex information and attribute information.

[0252] Multiplexer 124 multiplexes compressed vertex information, compressed vertex information metadata, compressed attribute information, compressed attribute information metadata, and compressed metadata in the bitstream. Furthermore, multiplexer 124 inputs the bitstream to the system layer.

[0253] Figure 21 This is a block diagram illustrating another structural example of the 3D data decoder 213 according to this embodiment. In this example, the 3D data decoder 213 includes a vertex information generator 231, an attribute information generator 232, an image decoder 234, a metadata decoder 223, and a demultiplexer 224. The vertex information generator 231, the attribute information generator 232, and the image decoder 234 can be coupled with... Figure 8 The vertex information decoder 201 and the attribute information decoder 203 correspond to each other.

[0254] In this example, the 3D data decoder 213 decodes the 3D data using a video-based encoding method. In this video-based decoding, multiple 2D images are decoded according to an image encoding method, and 3D data is generated from these images. Here, the image encoding method can be HEVC (High Efficiency Video Coding) or VVC (Versatile Video Coding), etc.

[0255] Specifically, first, the bitstream is input from the system layer to demultiplexer 224. Demultiplexer 224 separates compressed vertex information, compressed vertex information metadata, compressed attribute information, compressed attribute information metadata, and compressed metadata from the bitstream. The compressed vertex information, compressed vertex information metadata, compressed attribute information, and compressed attribute information metadata are input to image decoder 234. The compressed metadata is input to metadata decoder 223.

[0256] The image decoder 234 decodes the vertex image according to the image encoding method. At this time, the image decoder 234 uses the metadata of the compressed vertex information to decode the vertex image from the compressed vertex information. Furthermore, the image decoder 234 inputs the vertex image to the vertex information generator 231. Additionally, the image decoder 234 decodes the attribute image according to the image encoding method. At this time, the image decoder 234 uses the metadata of the compressed attribute information to decode the attribute image from the compressed attribute information. Furthermore, the image decoder 234 inputs the attribute image to the attribute information generator 232.

[0257] Metadata decoder 223 decodes metadata from compressed metadata. The metadata decoded by metadata decoder 223 contains mapping information used in the generation of vertex information and attribute information. Furthermore, the metadata decoded by metadata decoder 223 can be used for decoding vertex images and attribute images.

[0258] Vertex information generator 231 reconstructs vertex information from the vertex image according to the mapping information contained in the metadata decoded by metadata decoder 223. Attribute information generator 232 reconstructs attribute information from the attribute image according to the mapping information contained in the metadata decoded by metadata decoder 223.

[0259] Subsequently, vertex information, attribute information, and metadata are output as 3D data from the 3D data decoder 213. Additionally, this metadata, for example, is metadata of vertex information and attribute information that can be used in applications.

[0260] Figure 22 This is a conceptual diagram illustrating a specific example of the encoding process in this embodiment. Figure 22 A 3D data encoder 113 and a description encoder 148 are shown. In this example, the 3D data encoder 113 includes a 2D data encoder 141 and a mesh data encoder 142. The 2D data encoder 141 includes a texture encoder 143. The mesh data encoder 142 includes a vertex information encoder 144 and a connectivity information encoder 145.

[0261] Vertex information encoder 144, connection information encoder 145, and texture encoder 143 can correspond to Figure 6 Vertex information encoder 101, connection information encoder 102, and attribute information encoder 103, etc.

[0262] For example, the two-dimensional data encoder 141 acts as a texture encoder 143, encoding the texture corresponding to the attribute information as two-dimensional data according to the image encoding method or video encoding method, thereby generating a texture file.

[0263] Furthermore, the mesh data encoder 142 operates as both the vertex information encoder 144 and the connectivity information encoder 145, generating a mesh file by encoding vertex and connectivity information. The mesh data encoder 142 can further encode texture mapping information. The encoded mapping information can then be included within the mesh file.

[0264] Furthermore, the description encoder 148 generates a description file by encoding descriptions corresponding to metadata, such as text data. The description encoder 148 can encode the description at the system layer. For example, the description encoder 148 can be included in... Figure 12 In the system multiplexer 114.

[0265] The above steps generate a bitstream containing texture files, mesh files, and description files. These files can be multiplexed in the bitstream as glTF (Graphics Language Transmission Format) or USD (Universal Scene Description) files.

[0266] Furthermore, the 3D data encoder 113 may have two mesh data encoders as mesh data encoder 142. For example, one mesh data encoder encodes the vertex and connection information of a static 3D mesh, while the other mesh data encoder encodes the vertex and connection information of a dynamic 3D mesh.

[0267] Furthermore, the bitstream can contain two mesh files. For example, one mesh file corresponds to a static 3D mesh, and the other mesh file corresponds to a dynamic 3D mesh.

[0268] Furthermore, static 3D meshes can be intraframe 3D meshes encoded using intraframe prediction, while dynamic 3D meshes can be interframe 3D meshes encoded using interframe prediction. Additionally, information for dynamic 3D meshes can be derived from the difference between the vertex or connectivity information of the intraframe 3D mesh and the vertex or connectivity information of the interframe 3D mesh.

[0269] Figure 23 This is a conceptual diagram illustrating a specific example of the decoding process in this embodiment. Figure 23The image shows a 3D data decoder 213, a description decoder 248, and a prompter 247. In this example, the 3D data decoder 213 includes a 2D data decoder 241, a mesh data decoder 242, and a mesh reconstructor 246. The 2D data decoder 241 includes a texture decoder 243. The mesh data decoder 242 includes a vertex information decoder 244 and a connectivity information decoder 245.

[0270] Vertex information decoder 244, connection information decoder 245, texture decoder 243, and mesh reconstructor 246 can be used with Figure 8 The vertex information decoder 201, connection information decoder 202, attribute information decoder 203, and post-processor 205 correspond to these. The prompter 247 can be associated with... Figure 12 The corresponding prompts are 215, etc.

[0271] For example, the two-dimensional data decoder 241 acts as a texture decoder 243, and decodes the texture corresponding to the attribute information as two-dimensional data according to the texture file, based on the image encoding method or video encoding method.

[0272] Furthermore, the mesh data decoder 242 acts as both the vertex information decoder 244 and the connectivity information decoder 245, decoding vertex and connectivity information from the mesh file. The mesh data decoder 242 can further decode texture-specific mapping information from the mesh file.

[0273] Furthermore, the description decoder 248 decodes the description corresponding to metadata such as text data from the description file. The description decoder 248 can decode the description at the system layer. For example, the description decoder 248 can be included in... Figure 12 In the system demultiplexer 214.

[0274] Mesh Reconstructor 246, as described, reconstructs the 3D mesh based on vertex information, connectivity information, and texture. Hint 247, as described, renders and outputs the 3D mesh.

[0275] Through the above actions, a 3D mesh is reconstructed and output from a bitstream including texture files, mesh files, and description files.

[0276] Furthermore, the 3D data decoder 213 may have two mesh data decoders as mesh data decoder 242. For example, one mesh data decoder decodes the vertex information and connectivity information of a static 3D mesh, while the other mesh data decoder decodes the vertex information and connectivity information of a dynamic 3D mesh.

[0277] Furthermore, the bitstream can contain two mesh files. For example, one mesh file corresponds to a static 3D mesh, and the other mesh file corresponds to a dynamic 3D mesh.

[0278] Furthermore, static 3D meshes can be intra-frame 3D meshes encoded using intra-frame prediction, while dynamic 3D meshes can be inter-frame 3D meshes encoded using inter-frame prediction. Additionally, as information for dynamic 3D meshes, the difference between the vertex or connectivity information of intra-frame 3D meshes and the vertex or connectivity information of inter-frame 3D meshes can be used.

[0279] The coding method for dynamic 3D meshes is sometimes referred to as DMC (Dynamic Mesh Coding). Additionally, video-based coding of dynamic 3D meshes is sometimes referred to as V-DMC (Video-based Dynamic Mesh Coding).

[0280] Point cloud compression is sometimes referred to as PCC. Additionally, video-based point cloud compression is sometimes referred to as V-PCC. Furthermore, geometry-based point cloud compression is sometimes referred to as G-PCC.

[0281] <Implementation Example> Figure 24 This is a block diagram illustrating an implementation example of the encoding device 100 according to this embodiment. The encoding device 100 includes circuitry 151 and a memory 152. For example, Figure 5 The multiple components of the encoding device 100 shown are transmitted through... Figure 24 The circuit 151 and memory 152 shown are used to implement this.

[0282] Circuit 151 is a circuit that performs information processing and can access memory 152. For example, circuit 151 is a dedicated or general-purpose circuit for encoding a three-dimensional mesh. Circuit 151 can be a processor like a CPU. Alternatively, circuit 151 can be an assembly of multiple circuits.

[0283] Memory 152 is a dedicated or general-purpose memory that stores the information used by circuit 151 to encode the three-dimensional mesh. Memory 152 can be a circuit and can be connected to circuit 151. Alternatively, memory 152 can be contained within circuit 151. Alternatively, memory 152 can be a collection of multiple circuits. Furthermore, memory 152 can be a disk or optical disk, or it can be a storage device or recording medium. Memory 152 can be non-volatile memory or volatile memory.

[0284] For example, memory 152 can store a three-dimensional mesh or a bit stream. Additionally, memory 152 can store the program used by circuit 151 to encode the three-dimensional mesh.

[0285] Furthermore, in the encoding device 100, it is not necessary to implement Figure 5 All of the multiple constituent elements shown here may be exempt from the multiple processes shown here. Figure 5 A portion of the multiple constituent elements shown may be included in other devices, and a portion of the multiple processes shown herein may be executed by other devices. Furthermore, in the encoding device 100, the multiple constituent elements of this disclosure may be implemented in any combination, and the multiple processes of this disclosure may be performed in any combination.

[0286] Figure 25 This is a block diagram illustrating an implementation example of the decoding device 200 according to this embodiment. The decoding device 200 includes circuitry 251 and a memory 252. For example, Figure 7 The multiple components of the decoding device 200 shown are transmitted through... Figure 25 The circuit 251 and memory 252 shown are used to implement this.

[0287] Circuit 251 is a circuit that performs information processing and can access memory 252. For example, circuit 251 is a dedicated or general-purpose circuit for decoding a three-dimensional mesh. Circuit 251 can be a processor like a CPU. Alternatively, circuit 251 can be an assembly of multiple circuits.

[0288] Memory 252 is a dedicated or general-purpose memory that stores the information used by circuit 251 to decode the three-dimensional mesh. Memory 252 can be a circuit and can be connected to circuit 251. Alternatively, memory 252 can be contained within circuit 251. Alternatively, memory 252 can be an assembly of multiple circuits. Alternatively, memory 252 can be a disk or optical disk, or it can take the form of a storage device or recording medium. Memory 252 can be non-volatile memory or volatile memory.

[0289] For example, memory 252 can store a three-dimensional mesh or a bit stream. Additionally, memory 252 can also store the program used by circuit 251 to decode the three-dimensional mesh.

[0290] Furthermore, in the decoding device 200, it is not necessary to implement... Figure 7 All of the multiple constituent elements shown here may be exempt from the multiple processes shown here. Figure 7 A portion of the multiple constituent elements shown may be included in other devices, and a portion of the multiple processes shown herein may be executed by other devices. Furthermore, in the decoding device 200, the multiple constituent elements of this disclosure may be arbitrarily combined to implement the present disclosure, and the multiple processes of this disclosure may be arbitrarily combined to perform the present disclosure.

[0291] The encoding and decoding methods, which include steps performed by the constituent elements of the encoding apparatus 100 and decoding apparatus 200 of this disclosure, can be executed by any device or system. For example, part or all of the encoding and decoding methods can be executed by a computer equipped with a processor, memory, and input / output circuits. In this case, the encoding and decoding methods can be executed by the computer executing a program for causing the computer to perform the encoding and decoding methods.

[0292] In addition, programs and bit streams can be recorded on non-transitory computer-readable recording media such as CD-ROMs.

[0293] An example of a program could be a bitstream. For instance, a bitstream containing an encoded 3D mesh includes syntactic elements for the decoding device 200 to decode the 3D mesh. Furthermore, the bitstream causes the decoding device 200 to decode the 3D mesh according to the syntactic elements contained within it. Therefore, a bitstream can function similarly to a program.

[0294] The aforementioned bitstream can be an encoded bitstream containing the encoded three-dimensional grid, or a multiplexed bitstream containing the encoded three-dimensional grid and other information.

[0295] Furthermore, the components of the encoding device 100 and the decoding device 200 can be constructed from dedicated hardware, general-purpose hardware that executes the aforementioned programs, or a combination thereof. Additionally, the general-purpose hardware can consist of a memory storing programs and a general-purpose processor that reads and executes the programs from the memory. Here, the memory can be a semiconductor memory or a hard disk, and the general-purpose processor can be a CPU.

[0296] In addition, dedicated hardware can consist of memory and a dedicated processor. For example, a dedicated processor can refer to the memory used to record data to execute encoding and decoding methods.

[0297] Furthermore, as described above, each component of the encoding device 100 and the decoding device 200 can be a circuit. These circuits can be configured as a single circuit or they can be separate circuits. Additionally, these circuits can correspond to dedicated hardware or general-purpose hardware that executes the aforementioned programs. Moreover, the encoding device 100 and the decoding device 200 can be implemented as integrated circuits.

[0298] Alternatively, the encoding device 100 can be a transmitting device for transmitting a three-dimensional mesh. The decoding device 200 can be a receiving device for receiving a three-dimensional mesh.

[0299] <Displacement Encoding and Decoding> Here, as an example, the following terminology will be used.

[0300] (1) Image An image is a unit of data consisting of a collection of pixels, containing either a picture or a block smaller than a picture. Images include both moving images and still images.

[0301] (2) Pictures An image is a unit of image processing consisting of a collection of pixels, also known as a frame or field.

[0302] (3) blocks A block is a processing unit consisting of a specific number of pixels. Blocks are also referred to using the terms shown in the examples below. The shape of a block is not particularly limited. A block can be, for example, an M×N pixel rectangle or an M×M pixel square. Blocks can also be triangular, circular, or other shapes. Examples of blocks are described below.

[0303] • Slices, tiles, or bricks • CTU, superblock, or basic partitioning unit • VPDU, Hardware Processing Division Unit • CU, Processing Block Unit, Prediction Block Unit (PU), or Orthogonal Transform Block Unit (TU) ·Sub-block (4) Pixels or samples A pixel or sample is the smallest point in an image; in other words, it is the smallest unit. A pixel or sample includes not only pixels at integer positions but also pixels at sub-pixel positions generated based on those integer positions.

[0304] (5) Pixel value or sample value Pixel values ​​or sample values ​​are the inherent values ​​of a pixel. Pixel values ​​or sample values ​​include luminance values, chrominance values, RGB grayscale levels, and may also include depth values ​​or binary values ​​of 0 or 1.

[0305] (6) Sign A flag represents more than one bit. For example, a flag may be a parameter or index represented by two or more bits. Flags not only represent values ​​expressed in binary numbers, but sometimes also values ​​expressed in numeric values ​​other than binary numbers.

[0306] (7) Signal Signals are symbolized or encoded to convey information. Signals include discrete digital signals or continuous analog signals.

[0307] (8) Stream or bitstream A stream or bitstream is a sequence of digital data that represents the flow of digital data. A stream or bitstream can be a single stream or can be structured as multiple streams with multiple levels. Streams or bitstreams can be transmitted using a single transmission path via serial communication or using multiple transmission paths via packet communication.

[0308] (9) Difference In the case of scalars, differences can include simple differences (xy) and difference calculations. Differences can include the absolute value of the difference (|xy|), the square of the difference (x^2-y^2), the square root of the difference (√(xy)), the weighted difference (ax-by, where a and b are constants), or the offset difference (x-y+a, where a is the offset), etc.

[0309] (10) and In the case of scalars, sums can include simple sums (x+y) and addition calculations. Sums can include the absolute value of the sum (|x+y|), the sum of squares (x^2+y^2), the square root of the sum (√(x+y)), the weighted sum (ax+by, where a and b are constants), or the offset sum (x+y+a, where a is the offset), etc.

[0310] (11) "based on" The phrase "based on" implies that objects other than "that" can also be considered. Additionally, "based on" is sometimes used when a direct result is obtained, or when the result is obtained through intermediate steps.

[0311] (12) "Being used" or "Using" Expressions like "what was used" or "what was used" imply that things other than "what" can also be considered. Furthermore, expressions like "was used" or "was used" are sometimes used when a direct result is obtained, or when a result is obtained through intermediate steps.

[0312] (13) Prohibition "Prohibition" can be translated as "not permitted". In addition, "not prohibited / prohibited" or "permitted / permitted" does not necessarily mean "obligation".

[0313] (14) “Restriction” or “Limitation” "Restriction" or "limitation" can be translated as "not permitted / not allowed" or "not allowed / permitted." Furthermore, "prohibited / not prohibited" or "not permitted / permitted" does not necessarily imply an "obligation." Additionally, what is prohibited in quantity / quality can be partial or complete.

[0314] (15) Chroma The term chroma, represented by the symbols Cb or Cr, is an adjective that indicates one of two color difference signals associated with a primary color in a sample arrangement or a single sample. The term chroma is sometimes used in place of chrominance.

[0315] (16) Luminance The term "luma" is an adjective denoted by the symbol or subscript Y or L, representing the monochromatic signal of a sample arrangement or a single sample that is related to the primary color. The term "luma" is sometimes used in place of "luminance".

[0316] The encoding and decoding system of this embodiment will be described below.

[0317] A typical 3D model (also known as a 3D model) represents an object digitally, allowing users to search the model in 3D using scaling, translation, and rotation while rendering over time. One method of constructing such a representation is to use triangles to build a 3D mesh. In the model described above, the positions of the triangle vertices, the connectivity between the triangle vertices, and the properties associated with them (normals or UV mapping, etc.) are stored.

[0318] To store all this information in uncompressed form, a very large storage area is required, thus requiring a very large bandwidth for transmission. The triangles forming the mesh, especially in temporal and spatial proximity, mostly exhibit repeating patterns and similar properties. Utilizing these repetitions, efficient encoding and decoding methods can be developed for storage and transmission. One such encoding and decoding method is Video-based Dynamic Mesh Coding (V-DMC).

[0319] Figure 26 This is a block diagram illustrating another configuration example of the encoding / decoding system according to this embodiment. For example... Figure 26 As shown, the encoding and decoding system includes an encoding device 100 and a decoding device 200.

[0320] The encoding / decoding system accepts a 3D mesh (also known as a 3D mesh) as input in the form of vertices' 3D coordinates (vertices information), connectivity (connectivity information), and associated attributes (attribute information). Furthermore, a 3D mesh can include not only geometry but also texture mapping.

[0321] The encoding device 100 takes the input 3D mesh (also called the input 3D mesh or input grid) in the form of the three-dimensional coordinates of the vertices, connectivity, and associated attributes. The encoding device 100 encodes all associated information into a stream. The stream can consist of a single bit stream or multiple bit streams.

[0322] Network 300 transmits the stream generated by the encoding device to the decoding device 200. Network 300 can be the Internet, WAN (Wide Area Network), LAN (Local Area Network), or any combination thereof. Furthermore, network 300 is not necessarily limited to a two-way communication network; it can also be a one-way communication network transmitting broadcast waves such as terrestrial digital broadcasts or satellite broadcasts. Alternatively, recording media such as DVDs (Digital Versatile Discs) or BDs (Blu-ray Discs) can be used instead of network 300.

[0323] The bitstream is sent to the decoding device 200 via network 300. The decoding device 200 decodes the bitstream and uses the 3D coordinates, connectivity, and association properties of the decoded vertices to generate a 3D mesh. The decoding device 200 outputs the generated 3D mesh (also known as the output 3D mesh or output grid).

[0324] Figure 27 This is a diagram showing another configuration example of the encoding device 100.

[0325] like Figure 27 As shown, the encoding device 100 includes a preprocessor 1103 and a compressor 1106.

[0326] Encoding device 100 reads in input mesh 1101 and attribute map 1102 and passes them to preprocessor 1103. Preprocessor 1103 processes the input mesh to extract base mesh 1104 and displacement data 1105. Attribute map 1102, along with the extracted base mesh 1104 and displacement data 1105, is then transmitted to compressor 1106.

[0327] Furthermore, compressor 1106 compresses the base mesh 1104, displacement data 1105, and attribute map 1102 to generate bitstream 1107. By also including metadata 1108 in the bitstream 1107, compressor 1106 can send additional information to decoding device 200.

[0328] Figure 28 This is a diagram showing another configuration example of the decoding device 200.

[0329] like Figure 28 As shown, the decoding device 200 includes a decompressor 2102 and a post-processor 2106.

[0330] Decoding device 200 reads bitstream 2101 and passes it to decompressor 2102. Decompressor 2102 decompresses the base mesh 2103, displacement data 2104, and attribute map 2108 from bitstream 2101 and passes them to postprocessor 2106. An example of displacement data 2104 is a displacement vector.

[0331] Furthermore, the post-processor 2106 generates the output mesh 2107 by processing the base mesh 2103 according to the displacement data 2104 and the attribute mapping 2108. The post-processor 2106 can also use information from the metadata 2105 to generate the output mesh 2107.

[0332] Figure 29 This is a block diagram illustrating yet another structural example of the encoding device 100 of this embodiment.

[0333] In this example, the encoding device 100 includes a volumetric capturer 511, a projector 512, a basic mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and one or more other types of encoders 516 as options.

[0334] Volume capturer 511 captures content and outputs the captured content to projector 512.

[0335] Projector 512 projects content onto a 3D mesh frame containing vertex geometric coordinates (vertex coordinates representing the position of vertices), texture coordinates, and connectivity data (connection information). This data is output to a base mesh encoder 513, a translation encoder 514, an attribute encoder 515, and one or more other types of encoders 516 as options. Each encoder compresses the data into a bitstream.

[0336] Figure 30 This is a block diagram illustrating another structural example of the decoding device 200 of this embodiment.

[0337] In this example, the decoding device 200 includes a basic mesh decoder 613, a displacement decoder 614, an attribute decoder 615, one or more other types of decoders 616, and a 3D reconstructor 617.

[0338] The bitstream is sent to a base mesh decoder 613, a displacement decoder 614, an attribute decoder 615, and one or more other types of decoders 616 as options. These decoders decode the bitstream to generate decoded data containing vertex geometric coordinates, texture coordinates, and connectivity data. The decoded data is then sent to a 3D reconstructor 617 to reconstruct a 3D mesh frame.

[0339] The encoding process performed by the encoding device 100 will be described in detail below.

[0340] Figure 31 This is a flowchart illustrating the processing of the encoding device 100. Figure 32 This is a conceptual illustration of the encoding of a grid frame. (See reference...) Figure 31 and Figure 32 Explain the processing of the encoding device 100.

[0341] In step S101, the encoding device 100 reads a 3D mesh frame and its attributes as an input mesh frame. The input mesh frame is the mesh frame input to the encoding device 100. An example of a 3D mesh frame as an input mesh frame is shown as mesh frame 1301 (see reference). Figure 32 ).

[0342] In step S102, the encoding device 100 performs decimation processing on the input mesh frame read in step S101 to generate a base mesh frame with fewer vertices than the input mesh frame. The base mesh frame generated by decimating mesh frame 1301 is represented as base mesh frame 1302 (see reference). Figure 32 ).

[0343] In step S103, the encoding device 100 calculates the displacement information used by the decoding device 200 to reconstruct the mesh frame. The displacement information is equivalent to the displacement vector from the vertices of the base mesh frame generated in step S102 towards the vertices of the input mesh frame. One method for calculating the displacement information is to subtract the coordinates of the vertices of the base mesh frame from the coordinates of the vertices of the input mesh frame. The displacement information calculated based on mesh frame 1301 and base mesh frame 1302 is represented as displacement information 1303 (see reference). Figure 32 The displacement information 1303 is in vector form; in other words, it is represented as a displacement vector.

[0344] In step S104, the encoding device 100 encodes the base grid frame generated in step S102, the displacement information generated in step S103, and the attributes of the input grid frame into a bitstream (equivalent to a compressed bitstream). An example of this bitstream is shown as bitstream 1304 (see reference). Figure 32 ).

[0345] Specifically, bitstream 1304 contains vertex coordinates and connectivity information of vertices A, C, E, and F, displacement information, a video bitstream containing texture data, and compressed attribute mapping (see reference). Figure 32 The displacement information contains displacement information used to displace vertices based on vertex coordinates obtained from the base mesh frame after sub-segmentation. The compressed attribute map is used to apply texture data to the texture coordinates of the mesh frame reconstructed using the base mesh frame and displacement information.

[0346] The decoding process performed by the decoding device 200 will be described in detail below.

[0347] Figure 33 This is a flowchart illustrating the processing of the decoding device 200. Figure 34 This is a conceptual illustration of the decoding of a 3D mesh. (See reference...) Figure 33 and Figure 34 Explain the processing of the decoding device 200.

[0348] In step S201, the decoding device 200 decodes the base grid frame and attributes from the bitstream (equivalent to a compressed bitstream). An example of a decoded base grid frame (equivalent to a decoded base grid frame) is shown as decoded base grid frame 2301 (see reference). Figure 34 ).

[0349] In step S202, the decoding device 200 generates sub-segmented vertices by sub-segmenting the base mesh frame decoded in step S201. An example of a base mesh frame including sub-segmented vertices is represented as base mesh frame 2302 (see reference). Figure 34 ).

[0350] In step S203, the decoding device 200 decodes the displacement information from the bitstream (equivalent to a compressed bitstream). An example of the decoded displacement information is represented as displacement information 2303 (refer to...). Figure 34 The displacement information 2303 is in vector form; in other words, it is represented as a displacement vector.

[0351] In step S204, the decoding device 200 reconstructs the shape of the mesh frame by moving the vertices of the base mesh frame, including the sub-segmented vertices, to new positions using displacement information, and then restores the mesh frame by applying attribute information. One example of an attribute is texture. An example of the reconstructed mesh frame is represented as mesh frame 2304 (see reference). Figure 34 ).

[0352] Figure 35 This is a block diagram illustrating a structural example of the decoding device according to this embodiment.

[0353] Figure 35 This is an example of a block diagram representing a typical intra-frame decoding.

[0354] Figure 35 The decoding device shown includes a demultiplexer 1231, a switch 1232, a static grid decoder 1233, a grid buffer 1234, a motion decoder 1235, a basic grid reconstructor 1236, an inverse quantizer 1237, a video decoder 1238, an image unpacker 1239, an inverse quantizer 1240, an inverse wavelet transform 1241, a reconstructor 1242, a video decoder 1243, and a color converter 1244.

[0355] Demultiplexer 1231 obtains the compressed bitstream and separates the compressed data associated with the base grid, the image containing displacement data (also known as the displacement bitstream), and the image containing attribute data (also known as the attribute bitstream). The compressed data associated with the base grid is passed to switch 1232. Switch 1232 determines whether to perform intra-frame decoding or inter-frame decoding based on the parameters within the bitstream.

[0356] When intra-frame decoding is selected, the bitstream is passed to a static mesh decoder 1233 that generates the quantized base mesh. The static mesh decoder 1233 is, for example, a decoder that uses an edgebreaker algorithm in the decoding of 3D mesh data. The static mesh decoder 1233 generates the quantized base mesh based on the bitstream. The quantized base mesh generated by the static mesh decoder 1233 is stored in a mesh buffer 1234 for reference when inter-frame decoding is selected.

[0357] When inter-frame decoding is selected, switch 1232 passes compressed data associated with the base mesh to motion decoder 1235. Motion decoder 1235 receives the previously decoded quantized base mesh and decodes motion data representing the difference in vertex coordinates between the quantized base mesh stored in mesh buffer 1234 and the current quantized base mesh. Base mesh reconstructor 1236 uses the motion data and the quantized base mesh stored in mesh buffer 1234 to reconstruct the current quantized base mesh. The quantized base mesh obtained from inter-frame or intra-frame decoding is passed to inverse quantizer 1237 to obtain the decoded base mesh.

[0358] The image containing displacement data is passed to the video decoder 1238 because its bitstream contains displacement data in image format with two chromaticity information points and one luminance information point. The video decoder 1238 decodes the data using a video frame decompression method. In another example, an arithmetic decoder is used to decode the displacement data. This decompressed data is passed to the image unpacker 1239, which extracts wavelet coefficients associated with each vertex from the decompressed data in image form. The inverse quantizer 1240 inverse-quantizes the quantized wavelet coefficients using the three components associated with each vertex. The inverse wavelet transform 1241 performs an inverse transform on the result, ultimately obtaining the decoded displacement data. The decoded displacement data and the decoded base mesh are passed to the reconstructor 1242. The reconstructor 1242 subdivides the edges of the decoded base mesh and uses the decoded displacement data to displace the vertices, obtaining the decoded mesh.

[0359] To obtain the decoded attribute bitstream, the image containing the attribute data is passed to another video decoder 1243. For color space and color format conversion, the decoded attribute bitstream is further processed in the color converter 1244 to obtain the decoded attribute map.

[0360] Figure 36 This is a block diagram illustrating a structural example of the decoding device according to this embodiment.

[0361] Figure 36 This represents an example of a reconstructor that obtains a decoded 3D mesh 1256 from the decoded base mesh 1251 and the decoded displacement data 1254.

[0362] The decoded base grid 1251 is passed to the sub-segmenter 1252.

[0363] Sub-segmenter 1252 subdivides any two connected vertices of the overall 3D mesh by appending a new vertex between them. To generate a predefined number of vertices, this process can be repeated multiple times, including vertices created through previous sub-segmentation steps. The repeated sub-segments of the overall 3D mesh generate new levels of detail (LoD). The sub-segmented mesh 1253 and the decoded displacement data 1254 are passed to displacementr 1255. Displacer 1255 moves each vertex to a new position according to the corresponding displacement data, thereby generating the decoded 3D mesh 1256.

[0364] The following describes sub-segmentation. Sub-segmentation is performed by a sub-segmenter (specifically, sub-segmenter 1206 or sub-segmenter 2204).

[0365] Figure 37 This is an illustrative diagram representing an example of sub-segmentation.

[0366] Figure 37 The basic mesh shown in (a) contains vertices A, B and C, as well as connection information representing their connectivity.

[0367] exist Figure 37 (b) shows the mesh generated by the first sub-segmentation, in other words, the mesh after the first sub-segmentation. The sub-segmenter generates vertices D, E, and F, as well as connection information representing their connectivity, in the first sub-segmentation. The mesh generated by the sub-segmenter is also referred to as LoD1 or the first LoD.

[0368] Vertex D of the mesh after the first sub-segmentation is generated by sub-segmenting vertices A and B. Similarly, vertex E is generated by sub-segmenting vertices B and C. Vertex F is generated by sub-segmenting vertices A and C.

[0369] Furthermore, as an example, vertex D can be the midpoint of line segment AB (in other words, edge AB) connecting vertices A and B, which will form the basis of this generation. Similarly, vertex E can be the midpoint of line segment AC. Vertex F can be the midpoint of line segment BC.

[0370] exist Figure 37 Figure (c) shows the mesh generated by the second sub-segmentation, in other words, the mesh after the second sub-segmentation. In the second sub-segmentation, the sub-segmenter generates vertices G, H, I, J, K, L, M, N, and O, as well as connection information representing their connectivity. The mesh generated by the sub-segmenter is also referred to as LoD2 or the second LoD.

[0371] Vertex G of the mesh after the second sub-segmentation is generated by sub-segmenting vertices A and D. Similarly, vertex H is generated by sub-segmenting vertices A and E. Vertex I is generated by sub-segmenting vertices B and D. Vertex J is generated by sub-segmenting vertices D and F. Vertex K is generated by sub-segmenting vertices E and F. Vertex L is generated by sub-segmenting vertices C and E. Vertex M is generated by sub-segmenting vertices B and F. Vertex N is generated by sub-segmenting vertices C and F. Vertex O is generated by sub-segmenting vertices D and E.

[0372] Furthermore, as an example, vertex G can be the midpoint of line segment AD (in other words, edge AD) connecting vertices A and D, which will serve as the basis for this generation. Similarly, vertex H can be the midpoint of line segment AE. Vertex I can be the midpoint of line segment BD. Vertex J can be the midpoint of line segment DF. Vertex K can be the midpoint of line segment EF. Vertex L can be the midpoint of line segment CE. Vertex M can be the midpoint of line segment BF. Vertex N can be the midpoint of line segment CF. Vertex O can be the midpoint of line segment DE.

[0373] The following is for reference Figure 38 and Figure 39 The displacement of the vertices is described. The displacement of the vertices is performed by the reconstructor 2209.

[0374] Figure 38 This is an explanatory diagram showing an example of the displacement of vertices after sub-segmentation and displacement. Figure 39 This is an illustrative diagram representing an example of the vertices of the original mesh.

[0375] Figure 38 The basic mesh shown in (a) includes vertices A, B, C and Z, as well as connection information representing their connectivity.

[0376] exist Figure 38 (b) shows the mesh generated by the first sub-segmentation, in other words, the mesh after the first sub-segmentation (i.e., the first LoD). The sub-segmenter generates vertices S, T, U, X, or Y, and connection information representing their connectivity, in the first sub-segmentation. Regarding vertices S, T, U, X, or Y, and... Figure 37 The vertices D, E, and F shown in (b) are the same.

[0377] exist Figure 38 (c) shows the mesh generated by the second sub-segmentation; in other words, it shows the mesh after the second sub-segmentation (i.e., the second LoD). The sub-segmenter generates vertices D, E, F, G, and H in the second sub-segmentation, along with connection information representing their connectivity. Regarding vertices D, E, F, G, and H, and... Figure 37 The vertices G, H, I, J, K, L, M, N, or O shown in (c) are the same.

[0378] Figure 38 (d) represents the mesh containing the vertices after displacement following subdivision. Figure 38 The vertices A, B, C, D, E, F, G, H, S, T, U, X, Y, and Z shown in (d) are located using displacement information from Figure 38 The position of the vertex shown in (c) after displacement.

[0379] Figure 39The original grid shown is an example of the grid input to the encoding device 100, i.e., the grid before encoding.

[0380] Figure 38 The grid shown has the same Figure 39 The shape shown is close to that of the original mesh. The displacement information is generated by the displacement vector calculator 1207 of the encoding device 100 as information representing the displacement from the vertices of the base mesh to the vertices of the original mesh. Therefore, by using the displacement information generated in this way, the mesh is reconstructed to generate a mesh with a shape close to that of the original mesh.

[0381] Decoding device 200 can output Figure 38 The grid shown in (d) is a grid.

[0382] Next, refer to Figure 40 and Figure 41 This describes the division of a grid into subgrids.

[0383] The mesh is divided into multiple smaller parts, each of which can be encoded. During mesh division, the vertices of the mesh are segmented in a way that allows the coordinates and connectivity of the vertices within each segment to be encoded independently.

[0384] Figure 40 This is an illustrative diagram representing an example of a grid. Figure 41 This is an illustrative diagram showing an example of dividing a grid into subgrids.

[0385] Figure 40 The grid shown is the original grid, sometimes referred to as the full grid in contrast to the subgrid.

[0386] Figure 41 Indicates will Figure 40 The diagram shows the case where the entire mesh is divided into two sub-mesh segments. Regarding vertices A, B, and C of the entire mesh (refer to...). Figure 40 Vertex A is copied to vertices A1 and A2, vertex B is copied to vertices B1 and B2, and vertex C is copied to vertices C1 and C2, thus creating two sub-meshes (i.e., sub-mesh 1 and sub-mesh 2) from the full mesh. Sub-mesh 1 and sub-mesh 2 are meshes that can be decoded independently.

[0387] The following is for reference Figure 42 , Figure 43 and Figure 44 This indicates that displacement information is packaged into the image frame.

[0388] Figure 42 , Figure 43 and Figure 44 This is an illustrative diagram showing an example of packing displacement information into an image frame. Additionally, an image frame can also be referred to as a video frame.

[0389] Vertex displacement data can be encoded into image frame data, for example, by mapping to each component of a YUV format image frame (i.e., each of the Y component (Y plane), U component (U plane), and V component (V plane). This will be explained below as an example. Alternatively, vertex displacement data can also be encoded into image frame data by mapping to each component of an RGB format image frame (each of the R component, G component, and B component).

[0390] The decoding device 200 can use an image encoding module to extract displacement data. The displacement data can be the X, Y, or Z components in a global coordinate system (e.g., a Cartesian coordinate system), or the shape of normal, tangent, or two tangent components in a local coordinate system. Methods for mapping displacement data to image frames include the following.

[0391] For example, in method 1, displacement data is configured in the image frame according to the scan order. An example of displacement data packing in this case is as follows: Figure 42 As shown, displacement data is directly mapped to image frames according to a predefined scanning order.

[0392] Additionally, since the height and width of an image frame are fixed, sometimes the displacement data cannot be fully contained within the frame. In such cases, the remaining portion of the image frame is filled with padding data (also known as padded data) (see [reference]). Figure 42 ).

[0393] For example, in method 2, the displacement data is separated into multiple LoDs and mapped to the Y, U, and V components of the image frame. An example of this type of displacement data packing is shown below. Figure 43 As shown. Here, the displacement data of the next LoD image frame begins immediately after the end of the displacement data of the previous LoD. Similar to method 1, if the displacement data does not fit perfectly into the image frame, the end of the image frame is padded (see [reference]). Figure 43 ).

[0394] For example, in method 3, the displacement data corresponding to the LoD is mapped to the Y, U, and V components of the image frame in a manner different from that in method 2. An example of the packing of displacement data in this case is as follows... Figure 44 As shown. Therefore, each LoD can be decoded independently. In method 3, the displacement data of each LoD is padded in the middle, and CTU alignment is performed together with the padding at the end of the video frame (see [reference]). Figure 44 ).

[0395] Figure 45 This is a block diagram illustrating a detailed structural example of the decoding apparatus 200 according to this embodiment. Specifically, Figure 45 This illustrates an example of the structure of the geometric coordinate decoder provided by the decoding device 200.

[0396] In this example, the decoding device 200 includes a frame header decoder 631, a vertex geometry predictor 632, a vertex geometry difference decoder 633, and a reconstructor 634.

[0397] The frame header decoder 631 reads the bitstream, decodes the frame headers in the bitstream, and determines whether to perform intra-frame decoding (intra-frame prediction) or inter-frame decoding (inter-frame prediction) on the frame data.

[0398] When inter-frame decoding is selected, the frame data contained in the bitstream is output to the vertex geometry predictor 632.

[0399] Vertex geometry predictor 632 outputs the predicted information to reconstructor 634. An example of the predicted information is a motion vector.

[0400] Reconstructor 634 uses the prediction information and the vertex coordinates relative to previously decoded frames to output the three-dimensional coordinates (vertex geometric coordinates) of the vertices.

[0401] On the other hand, when intra-frame decoding is selected, the frame data contained in the bitstream is output to the vertex geometry differential decoder 633.

[0402] The vertex geometric coordinate difference decoder 633 decodes frame data encoded as differences between the coordinates of vertices contained in the frame in order to generate vertex coordinates. Only one of the vertex geometric coordinates from the vertex geometric coordinate difference decoder 633 and the reconstructor 634 is used to generate the decoded 3D mesh frame.

[0403] Figure 46 This is an explanatory diagram showing the coordinates of vertices in the three-dimensional mesh of this embodiment. Specifically, Figure 46 This represents an example of decoding the entire frame of a 3D mesh using the actual vertex coordinates (positions) contained in the bitstream.

[0404] like Figure 46 As shown in (a), using an orthogonal coordinate system (x, y, z), the coordinates of vertex A contained in the 3D mesh frame at time (t) are decoded as (6, 8, 9). Similarly, the coordinates of vertex B are decoded as (10, 6, 7), and the coordinates of vertex C are decoded as (14, 8, 9). The same applies to vertices D through G.

[0405] Figure 47 This is an explanatory diagram showing the prediction information of this embodiment. Specifically, Figure 47This represents another example of decoding the entire frame of the three-dimensional grid frame at time (t) using the frame at time (t-1) (past frames) and the prediction information contained in the bitstream.

[0406] The coordinates (6, 8, 9) of vertex A in the frame being decoded (the current frame) are decoded by adding the coordinates (4, 7, 8) of vertex A in past frames to the value (2, 1, 1) associated with vertex A as indicated by the prediction information. Similarly, the coordinates (10, 6, 7) of vertex B in the current frame are decoded by adding the coordinates (8, 6, 7) of vertex B in past frames to the value (2, 0, 0) associated with vertex B as indicated by the prediction information.

[0407] Hereinafter, an example of the structure of the encoding device of this embodiment will be described.

[0408] Figure 48 This is a block diagram illustrating a structural example of the encoding device according to this embodiment.

[0409] Figure 48 The encoding device shown includes a decimator 4801, a sub-segmenter 4802, a displacement vector calculator 4803, a wavelet transformer 4804, an inter-frame predictor 4805, a quantizer 4806, an image packer 4807, a video encoder 4808, an inverse quantizer 4811, a reconstructor 4812, and a reference buffer 4813.

[0410] Extractor 4801 acquires a mesh frame (equivalent to the original 3D mesh frame, also called the original mesh frame or original mesh) input to the encoding device, and performs decimation processing (in other words, interval culling) on ​​the acquired mesh frame to generate a base mesh frame (also called the base mesh). Decimation processing is the process of deleting (in other words, interval culling) a portion of the vertices contained in the original mesh. Decimation processing may include processes that change the position of at least a portion of the vertices contained in the original mesh, or processes that change the connectivity of at least a portion of the vertices contained in the original mesh. Decimation processing is also simply referred to as decimation.

[0411] The base mesh generated by the extraction process has fewer vertices than the original mesh. The vertices of the base mesh can be located at different positions than those of the original mesh. Furthermore, the connectivity of the vertices in the base mesh can differ from that of the original mesh. The extractor 4801 provides the generated base mesh frame to the sub-segmenter 4802.

[0412] Sub-segmenter 4802 performs sub-segmentation processing on the base mesh frame generated by extractor 4801. Sub-segmentation processing can be a process of subdividing the base mesh frame by sub-segmentation. Sub-segmenter 4802 provides the sub-segmented base mesh frame to displacement vector calculator 4803.

[0413] Specifically, the sub-segmenter 4802 can sub-segment the mesh frame by generating new vertices between two interconnected vertices contained in the mesh frame. By repeatedly generating new vertices, the number of vertices contained in the mesh frame can be made to a predetermined number. Through the repeated sub-segmentation in the mesh frame as a whole (in other words, the multiple executions of sub-segmentation), multiple LoD (Level of Detail) levels are generated.

[0414] The displacement vector calculator 4803 obtains the original grid frame acquired by the encoding device and the sub-segmented base grid frame from the sub-segmenter 4802. Based on the vertices of the base grid frame and the vertices generated through the sub-segmentation of the base grid frame, the displacement vector calculator 4803 calculates the vectors pointing towards the corresponding vertices of the original grid frame as displacement vectors. The displacement vector calculator 4803 provides the calculated displacement vectors to the wavelet transform 4804.

[0415] Wavelet transform 4804 performs wavelet transform processing on the displacement vector calculated by displacement vector calculator 4803 to obtain transform coefficients (also called wavelet coefficients). Wavelet transform 4804 provides the obtained wavelet coefficients to inter-frame predictor 4805. In the wavelet transform, wavelet transform 4804 assigns vertices to multiple LoD levels, and by applying, for example, lifting transform to the displacement vector of the vertex, it can calculate wavelet coefficients representing various components from low-frequency components to high-frequency components.

[0416] Inter-frame predictor 4805 uses inter-frame prediction to calculate the prediction residual of the wavelet coefficients of the displacement vector of the target frame. Specifically, inter-frame predictor 4805 uses the wavelet coefficients of the displacement vector of the encoded frame (also called the reference frame) stored in the reference buffer 4813 to perform inter-frame prediction on the wavelet coefficients of the displacement vector of the target frame, thereby calculating the prediction residual of the wavelet coefficients of the displacement vector of the target frame.

[0417] The quantizer 4806 quantizes the prediction residuals of the wavelet coefficients calculated by the inter-frame predictor 4805. The quantizer 4806 can quantize the prediction residuals of the wavelet coefficients according to each LoD level. The quantizer 4806 provides the quantized prediction residuals to the image packer 4807 and the inverse quantizer 4811.

[0418] Image packer 4807 generates an image containing the prediction residual quantized by quantizer 4806. Image packer 4807 generates the aforementioned image by mapping the prediction residual quantized by quantizer 4806 to pixels in a two-dimensional image format. Image packer 4807 provides the generated image to video encoder 4808. In the process of mapping the quantized prediction residual to pixels in a two-dimensional image format, mapping information representing the allocation of the quantized prediction residual to pixels in the two-dimensional image format can be used.

[0419] The video encoder 4808 encodes the image generated by the image packer 4807 into a bitstream (also called a shifted bitstream) (in other words, generates a shifted bitstream). The video encoder 4808 outputs the shifted bitstream. The shifted bitstream can be a bitstream containing shift information in the form of an image. The image form can be, for example, a form containing two chroma information values ​​and one luminance information value. Furthermore, the video encoder 4808 can utilize a general-purpose module with the function of converting images into bitstreams. By using a highly reliable general-purpose module as the video encoder 4808, the above functions can be performed more reliably.

[0420] Inverse quantizer 4811 inverse-quantizes the prediction residual quantized by quantizer 4806, thereby generating the prediction residual of wavelet coefficients. Specifically, inverse quantizer 4811 can inverse-quantize the prediction residual quantized by quantizer 4806 according to each LoD level, thereby generating the prediction residual. Inverse quantizer 4811 provides the generated prediction residual of wavelet coefficients to reconstructor 4812.

[0421] The reconstructor 4812 restores (also called reconstructs) the wavelet coefficients based on the prediction residuals of the wavelet coefficients provided by the inverse quantizer 4811 and the reference frames stored in the reference buffer 4813. The reconstructed wavelet coefficients are then stored in the reference buffer 4813.

[0422] The reference buffer 4813 is a storage device, for example, storing wavelet coefficients of a reference frame. The wavelet coefficients stored in the reference buffer 4813 can be used for inter-frame prediction performed by the inter-frame predictor 4805.

[0423] Figure 49 This is a flowchart illustrating a specific example of the encoding process in this embodiment.

[0424] In step S4901, the inter-frame predictor 4805 calculates the intra-frame sum of the transform coefficients of the displacement vector (also known as sum_nointer).

[0425] In step S4902, the inter-frame predictor 4805 calculates the intra-frame sum of the prediction residuals (also known as sum_inter) when inter-frame prediction is applied to the transformation coefficients of the displacement vector.

[0426] In addition, the prediction residual of inter-frame prediction can be calculated by subtracting the transformation coefficient of the displacement vector of the three-dimensional point (also known as the reference point) of the reference frame from the transformation coefficient of the displacement vector of the three-dimensional point of the coded object in the coded object frame.

[0427] In step S4903, the inter-frame predictor 4805 determines whether the sum_inter calculated in step S4902 is less than the sum_nointer calculated in step S4901. If the sum_inter is less than the sum_nointer (yes in step S4903), the process proceeds to step S4904; otherwise (no in step S4903), the process proceeds to step S4911.

[0428] In step S4904, the inter-frame predictor 4805 decides to use inter-frame prediction to encode the transformation coefficients of the displacement vector within the frame and outputs the prediction residual to the quantizer 4806.

[0429] In step S4905, information indicating that the transform coefficients of the intra-frame displacement vectors are encoded using inter-frame prediction is appended to the header (e.g., the stream header). For example, by setting the information in the header indicating that the transform coefficients of the displacement vectors are encoded using inter-frame prediction to 1, it is possible to indicate that the transform coefficients of the intra-frame displacement vectors are encoded using inter-frame prediction.

[0430] In step S4911, it is decided not to use inter-frame prediction to encode the transform coefficients of the intra-frame displacement vector, and the transform coefficients are output to the quantizer 4806.

[0431] In step S4912, information indicating that the transform coefficients of the intra-frame displacement vectors are encoded without using inter-frame prediction is appended to the header. For example, by setting the information in the header that indicates the transform coefficients of the displacement vectors are encoded using inter-frame prediction to 0, it is possible to indicate that the transform coefficients of the intra-frame displacement vectors are encoded without using inter-frame prediction.

[0432] Figure 50 This is a block diagram illustrating a structural example of the decoding device according to this embodiment.

[0433] The decoding device includes a video decoder 5001, an image unpacker 5002, an inverse quantizer 5003, a reconstructor 5004, an inverse wavelet transform 5005, a reconstructor 5006, and a reference buffer 5011.

[0434] The video decoder 5001 acquires the shifted bitstream and decodes it into an image. This image can be an image contained by mapping quantized wavelet coefficients to pixels in a two-dimensional image format. The video decoder 5001 provides this image to the image unpacker 5002. Furthermore, the video decoder 5001 can utilize a general-purpose module with the function of converting bitstreams into images. By utilizing a highly reliable general-purpose module as the video decoder 5001, the above functions can be performed more reliably.

[0435] Image unpacker 5002 extracts quantized wavelet coefficients from the image provided by video decoder 5001. In the process of extracting the quantized wavelet coefficients from the image, a mapping representing the allocation of pixels in a two-dimensional image format can be used. Image unpacker 5002 provides the quantized wavelet coefficients extracted from the image to inverse quantizer 5003.

[0436] The inverse quantizer 5003 inverse-quantizes the prediction residuals of the quantized wavelet coefficients provided by the image unpacker 5002, thereby generating prediction residuals of the wavelet coefficients. Specifically, the inverse quantizer 5003 can generate prediction residuals of the wavelet coefficients by inverse-quantizing the prediction residuals of the quantized wavelet coefficients according to each LoD level.

[0437] Reconstructor 5004 restores (also called reconstructs) the transform coefficients of the decoded target frame based on the prediction residuals of the transform coefficients of the target frame and the transform coefficients of the reference frame. Reconstructor 5004 provides the restored transform coefficients to inverse wavelet transformor 5005 and reference buffer 5011.

[0438] The inverse wavelet transform 5005 generates a displacement vector (equivalent to a decoded displacement vector) by performing an inverse wavelet transform on the wavelet coefficients provided from the reconstructor 5004. The inverse wavelet transform is equivalent to the inverse of the wavelet transform performed by the wavelet transform 4804. Specifically, in the inverse wavelet transform, the inverse wavelet transform 5005 can calculate the displacement vector of the vertices by applying an inverse lifting transform to the wavelet coefficients. The inverse wavelet transform 5005 provides the generated decoded displacement vector to the reconstructor 5006.

[0439] Reconstructor 5006 uses the decoded displacement vector and the decoded base mesh provided by inverse wavelet transform 5005 to reconstruct the mesh (equivalent to the decoded mesh). Reconstructor 5006 outputs the reconstructed decoded mesh.

[0440] The reference buffer 5011 is a storage device, for example, storing wavelet coefficients of a reference frame. The wavelet coefficients stored in the reference buffer 5011 can be used to restore the transform coefficients of the decoded object frame performed by the reconstructor 5004.

[0441] Figure 51 This is a flowchart illustrating a specific example of the decoding process in this embodiment.

[0442] In step S5101, the reconstructor 5004 determines whether information indicating that the transform coefficients of the intra-frame displacement vectors have been encoded using inter-frame prediction is appended to the header. This information, for example, is information indicating that disp_frame_inter_mode is 1. If it is determined that the above information is appended to the header (yes in step S5101), the process proceeds to step S5102; otherwise (no in step S5101), the process proceeds to step S5111.

[0443] In step S5102, the reconstructor 5004 determines that the transform coefficients of the displacement vector within the frame have been encoded using inter-frame prediction, adds the transform coefficients of the reference frame to the prediction residual provided by the inverse quantizer 5003, decodes the transform coefficients and outputs them.

[0444] In step S5111, the reconstructor 5004 determines that the transform coefficients of the intra-frame displacement vector have been encoded without using inter-frame prediction, and outputs the transform coefficients provided by the inverse quantizer 5003.

[0445] Figure 52 This is a flowchart illustrating a specific example of the encoding process in this embodiment.

[0446] exist Figure 52 The coding process illustrated shows an example where the inter-frame prediction unit for the displacement vector is determined for each LoD, and information indicating whether inter-frame prediction is applied for each LoD is appended to the header: `disp_lod_inter_mode`. This allows switching the inter-frame prediction on or off for each LoD, improving coding efficiency. For example, there are cases where subtle motion is easily treated with inter-frame prediction while coarse motion is difficult to treat. In such cases, enabling inter-frame prediction at the lower level (high-frequency component set) and disabling it at the higher level (low-frequency component set) can help improve coding efficiency.

[0447] In step S5201, the inter-frame predictor 4805 begins loop A, which involves repeatedly executing steps S5202-S5206 and steps S5211-S5212 (described later). In loop A, one or more LoDs of interest are focused on, and processing is performed on these LoDs, ultimately controlling the processing to cover all LoDs. The LoDs of interest are also referred to as the LoDs of interest. Loop A can also be called the LoD loop.

[0448] In step S5202, the inter-frame predictor 4805 calculates the sum of the transformation coefficients of the displacement vector within the LoD of interest (also known as sum_lod_nointer).

[0449] In step S5203, the inter-frame predictor 4805 calculates the sum of the prediction residuals within the LoD of interest (also known as sum_lod_inter) when the inter-frame prediction is applied to the transform coefficients of the displacement vector.

[0450] In step S5204, the inter-frame predictor 4805 determines whether the sum_lod_inter calculated in step S5203 is less than the sum_lod_nointer calculated in step S5202. If the sum_lod_inter is less than the sum_lod_nointer (yes in step S5204), the process proceeds to step S5205; otherwise (no in step S5204), the process proceeds to step S5211.

[0451] In step S5205, the inter-frame predictor 4805 decides to use inter-frame prediction to encode the transformation coefficients of the displacement vector within the LoD of interest, and outputs the prediction residual to the quantizer 4806.

[0452] In step S5206, the inter-frame predictor 4805 appends information indicating that the transform coefficients of the displacement vectors within the LoD of interest have been encoded using inter-frame prediction to the header (e.g., the stream header). For example, by setting the information disp_lod_inter_mode(i) contained in the header, which indicates that the transform coefficients of the displacement vectors have been encoded using inter-frame prediction, to 1, it is possible to indicate that the transform coefficients of the displacement vectors within the LoD of interest have been encoded using inter-frame prediction. Furthermore, i is the ordinal number of the LoD of interest, indicating which LoD the interest is. The same applies thereafter.

[0453] In step S5211, the inter-frame predictor 4805 decides not to use inter-frame prediction to encode the transform coefficients of the displacement vector within the LoD of interest, and outputs the transform coefficients to the quantizer 4806.

[0454] In step S5212, the inter-frame predictor 4805 appends information indicating that the transform coefficients of the displacement vector within the LoD of interest are encoded without using inter-frame prediction to the header. For example, by setting the information disp_lod_inter_mode(i) contained in the header, which indicates that the transform coefficients of the displacement vector are encoded by inter-frame prediction, to 0, it is possible to indicate that the transform coefficients of the displacement vector within the LoD of interest are encoded without using inter-frame prediction.

[0455] In step S5207, the inter-frame predictor 4805 performs the end processing of loop A. Specifically, the inter-frame predictor 4805 determines whether the processing of steps S5202~S5206 and steps S5211~S5212 (where the processing of steps S5205 and S5206 and the processing of steps S5211 and S5212 are only one of the two corresponding to the determination result of step S5204) has been executed on all LoDs. If not, control is performed to focus on the LoDs that have not yet been executed and execute the processing.

[0456] Figure 53 This is a flowchart illustrating a specific example of the decoding process in this embodiment.

[0457] exist Figure 53 The decoding process shown illustrates the following example: The bitstream encoded by appending information indicating whether inter-frame prediction was applied per LoD to the header is decoded, with the inter-frame prediction unit determined for each LoD displacement vector. This allows for switching the inter-frame prediction on or off per LoD, facilitating proper decoding of the bitstream, which improves coding efficiency.

[0458] In step S5301, the refactoring unit 5004 begins the processing of loop A, which involves repeatedly executing steps S5302-S5203 and step S5311 (described later). In loop A, one or more LoDs are considered, and processing is performed on these LoDs, ultimately controlling the processing to cover all LoDs. The LoDs considered are also referred to as the LoDs of interest. Loop A can also be called the LoD loop.

[0459] In step S5302, the reconstructor 5004 determines whether the header contains information indicating that the transform coefficients of the displacement vectors within the LoD of interest have been encoded using inter-frame prediction. This information, for example, indicates that disp_lod_inter_mode(i) is 1. If the header contains this information (yes in step S5302), the process proceeds to step S5303; otherwise (no in step S5302), the process proceeds to step S5311.

[0460] In step S5303, the reconstructor 5004 determines that the transform coefficients of the displacement vector within the LoD of interest have been encoded using inter-frame prediction, adds the transform coefficients of the reference frame to the prediction residual provided by the inverse quantizer 5003, decodes the transform coefficients and outputs them.

[0461] In step S5311, the reconstructor 5004 determines that the transform coefficients of the displacement vector within the LoD of interest have been encoded without using inter-frame prediction, and outputs the transform coefficients provided by the inverse quantizer.

[0462] In step S5304, the reconstructor 5004 performs the end processing of loop A. Specifically, the reconstructor 5004 determines whether the processing of steps S5302~S5203 and step S5311 has been performed on all LoDs (wherein, the processing of step S5203 and the processing of step S5311 are only one of the two corresponding to the determination result of step S5302). If not, control is performed to focus on the LoDs that have not yet been processed and perform the processing.

[0463] Figure 54 This is an explanatory diagram illustrating an example of the syntax of this implementation method.

[0464] Figure 54 The examples of syntax shown represent examples of the structure of information contained in the bitstream generated by the encoding device.

[0465] Figure 54 The syntax shown includes a displacement vector_header. The displacement vector_header contains the disp_frame_inter_mode.

[0466] `disp_frame_inter_mode` indicates whether inter-frame prediction was used to encode the displacement vectors within the frame. For example, a value of 1 indicates that inter-frame prediction was used to encode the displacement vectors within the frame, while a value of 0 indicates that inter-frame prediction was not used to encode the displacement vectors within the frame. Therefore, the decoding device can determine whether inter-frame prediction was used to encode the displacement vectors within the frame and can decode the bitstream appropriately.

[0467] Furthermore, the unit for attaching `disp_frame_inter_mode` is not limited to the frame unit. For example, `disp_frame_inter_mode` can be attached at the sub-grid level. Thus, by switching the use of inter-frame prediction at the sub-grid level, coding efficiency can be improved. For instance, by using inter-frame prediction for sub-grids with low motion, such as background objects, and not using it for sub-grids with high motion, such as foreground objects, coding efficiency can be improved.

[0468] Furthermore, `disp_frame_inter_mode` can be appended to the header on a sequence-by-sequence basis. This reduces the header's bit size. For example, when encoding sequences with low overall motion, inter-frame prediction is used within the sequence as a whole; when encoding sequences with high overall motion, inter-frame prediction is not used within the sequence as a whole, thereby reducing the header's bit size and improving coding efficiency. Additionally, when inter-frame prediction is used within the sequence as a whole, `disp_frame_inter_mode` and `disp_lod_inter_mode` do not need to be appended to the header. This further reduces the header's bit size.

[0469] Figure 55 This is an explanatory diagram illustrating an example of the syntax of this implementation method.

[0470] Figure 55 The syntax shown includes a displacement vector_header. The displacement vector_header contains disp_lod_inter_mode[i].

[0471] `disp_lod_inter_mode[i]` indicates whether inter-frame prediction was used to encode the displacement vectors belonging to the i-th LoD when generating the LoD and encoding the displacement vectors. For example, a value of 1 indicates that inter-frame prediction was used to encode the displacement vectors of the 3D points belonging to the i-th LoD, and a value of 0 indicates that inter-frame prediction was not used to encode the displacement vectors of the 3D points belonging to the i-th LoD. Therefore, the decoding device can determine whether the displacement vectors of the 3D points belonging to the i-th LoD were encoded using inter-frame prediction, and can appropriately decode the bitstream.

[0472] Figure 56 This is an explanatory diagram illustrating an example of the syntax of this implementation method.

[0473] Figure 56 The syntax shown includes displacement vector_header. Displacement vector_header contains disp_frame_inter_mode, disp_lod_inter_mode_present, and disp_lod_inter_mode[i].

[0474] disp_frame_inter_mode and Figure 54 The disp_frame_inter_mode shown is the same.

[0475] `disp_lod_inter_mode_present` indicates whether `disp_lod_inter_mode` is included in the header. For example, a value of 1 indicates that `disp_lod_inter_mode` is included in the header, and a value of 0 indicates that `disp_lod_inter_mode` is not included in the header. Furthermore, when `disp_frame_inter_mode` = 1, in order to perform inter-frame prediction on a frame-by-frame basis, the value of `disp_lod_inter_mode_present` can be assumed to be 0, and `disp_lod_inter_mode` is not appended to the header. This reduces the amount of code in the header.

[0476] disp_lod_inter_mode[i] and Figure 55 The disp_lod_inter_mode[i] shown is the same.

[0477] Alternatively, when disp_frame_inter_mode = 0, the value of disp_lod_inter_mode_present can be omitted, and disp_lod_inter_mode[i] can be included in the header regardless of the value of disp_lod_inter_mode_present.

[0478] In addition, `disp_frame_inter_mode`, `disp_lod_inter_mode`, or `disp_lod_inter_mode_present` can also be entropy-encoded and appended to the header. For example, the values ​​can be binarized and arithmetic-coded. Alternatively, to reduce processing overhead, fixed-length encoding can be used.

[0479] Figure 57 This is a block diagram illustrating a structural example of the encoding device according to this embodiment.

[0480] Figure 57 The encoding device shown includes an extractor 5701, a sub-segmenter 5702, a displacement vector calculator 5703, a wavelet transformer 5704, a LoD-based inter-frame predictor 5705, a quantizer 5706, a switch 5707, an image packer 5708, a video encoder 5709, an arithmetic encoder 5710, an inverse quantizer 5711, a reconstructor 5712, and a reference buffer 5713.

[0481] Extractor 5701, sub-segmenter 5702, and displacement vector calculator 5703 are respectively connected to... Figure 48 The extractor 4801, sub-segmenter 4802 and displacement vector calculator 4803 shown are the same.

[0482] Wavelet transform 5704 performs wavelet transform processing on the displacement vector calculated by displacement vector calculator 5703 to obtain transform coefficients (also called wavelet coefficients). Wavelet transform 5704 provides the obtained wavelet coefficients to the LoD-based inter-frame predictor 5705. In the wavelet transform, wavelet transform 5704 assigns vertices to multiple LoD levels, and by applying, for example, lifting transform to the displacement vector of that vertex, it can calculate wavelet coefficients representing various components from low-frequency components to high-frequency components.

[0483] The LoD-based inter-frame predictor 5705 outputs the prediction residuals of the wavelet coefficients of the displacement vector of the target frame for each LoD using inter-frame prediction. Specifically, the LoD-based inter-frame predictor 5705 performs inter-frame prediction on the wavelet coefficients of the displacement vector of the target frame for each LoD using the wavelet coefficients of the displacement vector of the encoded frame (also called the reference frame) stored in the reference buffer 5713, thereby outputting the prediction residuals of the wavelet coefficients of the displacement vector of the target frame. The LoD-based inter-frame predictor 5705 can determine whether to use inter-frame prediction to encode the transform coefficients of the displacement vector for each LoD and switch accordingly.

[0484] The quantizer 5706 quantizes the prediction residuals of the wavelet coefficients calculated by the LoD-based inter-frame predictor 5705. The quantizer 5706 quantizes the prediction residuals of the wavelet coefficients at each LoD level. The quantizer 5706 provides the quantized prediction residuals to the image packer 5708 or the arithmetic encoder 5710 via the switch 5707, and also to the inverse quantizer 5711.

[0485] Switch 5707 is a switch that toggles whether the prediction residual quantized by quantizer 5706 is provided to image packer 5708 or arithmetic encoder 5710.

[0486] Image packer 5708 and video encoder 5709 respectively with Figure 48 The image packer 4807 and video encoder 4808 shown are the same.

[0487] The arithmetic encoder 5710 encodes the prediction residual quantized by the quantizer 5706 into a bit stream (also known as a shifted bit stream) through arithmetic encoding (in other words, generates a shifted bit stream). The arithmetic encoder 5710 outputs the shifted bit stream.

[0488] Inverse quantizer 5711 inverse-quantizes the prediction residuals quantized by quantizer 5706, thereby generating prediction residuals for wavelet coefficients. Specifically, inverse quantizer 5711 inverse-quantizes the prediction residuals quantized by quantizer 5706 at each LoD level to generate prediction residuals. Inverse quantizer 5711 provides the generated prediction residuals for wavelet coefficients to reconstructor 5712.

[0489] Reconstructor 5712 restores (also called reconstructs) the wavelet coefficients based on the prediction residuals of the wavelet coefficients provided by inverse quantizer 5711 and the reference frames stored in reference buffer 5713. Reconstructor 5712 stores the restored wavelet coefficients in reference buffer 5713.

[0490] The reference buffer 5713 is a storage device, for example, storing wavelet coefficients of a reference frame. The wavelet coefficients stored in the reference buffer 5713 can be used for inter-frame prediction in the LoD-based inter-frame predictor 5705.

[0491] Figure 57 The encoding device shown can switch between using the video encoder 5709 and the arithmetic encoder 5710 to encode the transformation coefficients of the displacement vector using arithmetic encoding. Therefore, the encoding device can efficiently encode the displacement vector using arithmetic encoding even when the video encoder 5709 is unavailable.

[0492] Alternatively, when using arithmetic coding, the transformation coefficients of the displacement vector can be encoded by switching between using inter-frame prediction for each LoD. Generally, with arithmetic coding, the less variation in the input values, the more efficient the compression. Therefore, by varying the inter-frame prediction suppression value for each LoD, the coding efficiency can be improved for the prediction residuals of the displacement vector's transformation coefficients.

[0493] Alternatively, information indicating whether the transformation coefficients of the displacement vector were encoded using the video encoder 5709 or encoded using arithmetic encoder 5710 can be appended to the header. This allows the decoding device to appropriately switch the decoding method by referring to the header information.

[0494] In addition, Figure 57 The example described uses the case where the transformation coefficients of the displacement vector are encoded using video encoder 5709 or arithmetic encoder 5710 via arithmetic coding, but a structure that always uses arithmetic coding is also possible. According to this structure, coding efficiency may be improved compared to using inter-frame prediction at all LoD levels when arithmetic coding of the transformation coefficients of the displacement vector is performed.

[0495] Figure 58 The decoding device shown includes a video decoder 5801, an image unpacker 5802, an arithmetic decoder 5803, a switch 5804, an inverse quantizer 5805, a reconstructor 5806, an inverse wavelet transform 5807, a reconstructor 5808, and a reference buffer 5811.

[0496] The video decoder 5801 and the image unpacker 5802 are respectively connected to... Figure 50 The video decoder 5001 and image unpacker 5002 shown are the same.

[0497] The arithmetic decoder 5803 acquires the shifted bitstream and performs arithmetic decoding on the prediction residuals contained in the acquired shifted bitstream. Additionally, the arithmetic decoder 5803 can also decode various header information.

[0498] Switch 5804 is a switch that toggles whether the prediction residual provided by the image unpacker 5802 is provided to the inverse quantizer 5805 or the prediction residual provided by the arithmetic decoder 5803 is provided to the inverse quantizer 5805.

[0499] The inverse quantizer 5805 inverse-quantizes the prediction residuals of the quantized wavelet coefficients provided by the image unpacker 5802 or the arithmetic decoder 5803 via the switch 5804, thereby generating the prediction residuals of the wavelet coefficients. Specifically, the inverse quantizer 5805 inverse-quantizes the prediction residuals of the quantized wavelet coefficients at each LoD level, thereby generating the prediction residuals of the wavelet coefficients.

[0500] Reconstructor 5806 restores (also known as reconstructs) the transform coefficients of the decoded target frame based on the prediction residuals of the transform coefficients of the target frame and the transform coefficients of the reference frame. Reconstructor 5806 provides the restored transform coefficients to inverse wavelet transformor 5807 and reference buffer 5811. Reconstructor 5806 can also determine whether inter-frame prediction has been applied per LoD based on header information and switch the restoration method accordingly.

[0501] The inverse wavelet transform 5807 generates a displacement vector (equivalent to decoding the displacement vector) by performing an inverse wavelet transform on the wavelet coefficients provided from the reconstructor 5806. The inverse wavelet transform is equivalent to the inverse of the wavelet transform performed by wavelet transform 5704. Specifically, in the inverse wavelet transform, inverse wavelet transform 5807 can calculate the displacement vector of the vertices by applying an inverse lifting transform to the wavelet coefficients.

[0502] The inverse wavelet transform 5807 provides the generated decoded displacement vector to the reconstructor 5808.

[0503] Reconstructor 5808 uses the decoded displacement vector and the decoded base mesh provided by inverse wavelet transform 5807 to reconstruct the mesh (equivalent to decoding the mesh). Reconstructor 5808 outputs the reconstructed decoded mesh.

[0504] The reference buffer 5811 is a storage device, for example, storing wavelet coefficients of a reference frame. The wavelet coefficients stored in the reference buffer 5811 can be used to restore the transform coefficients of the decoded object frame by the reconstructor 5806.

[0505] Figure 58 The decoding device shown can also decode header information, determine whether the encoding device used a video encoder (e.g., video encoder 5709) to encode the transform coefficients of the displacement vector, or used an arithmetic encoder (e.g., arithmetic encoder 5710) to encode the vector using arithmetic encoding, and switch the decoding method accordingly. Thus, even when a video encoder (e.g., video encoder 5709) cannot be used, the decoding device can still appropriately decode the bitstream obtained by efficiently encoding the displacement vector using arithmetic encoding.

[0506] Figure 59 This is an explanatory diagram showing the positional relationship of three-dimensional points in this embodiment.

[0507] As a method for encoding the displacement vector of a 3D point, consider the following approach: calculate the predicted value of the displacement vector of a 3D point, and encode the difference (prediction residual) between the original displacement vector value and the predicted value. For example, if the displacement vector of a 3D point p has a value of Ap and a predicted value of Pp, the encoding device 100 encodes the absolute value of the difference, Diffp = |Ap-Pp|, and the information indicating whether (Ap-Pp) is positive or negative. In this case, if the predicted value Pp can be generated with high precision, the value of the absolute value of the difference, Diffp, becomes smaller. Therefore, for example, the encoding device 100 can reduce the amount of code by using entropy encoding, where a smaller value results in a smaller number of bits generated.

[0508] As a method for generating predicted values ​​of displacement vectors by the encoding device 100, the displacement vectors of other three-dimensional points surrounding the three-dimensional point of the object to be encoded are considered. Here, the three-dimensional points surrounding the three-dimensional point refer to other three-dimensional points within a specified distance (a specified range) from the three-dimensional point. For example, in the case where there are three-dimensional points of the object to be encoded, namely three-dimensional point p = (x1, y1, z1) and three-dimensional point q = (x2, y2, z2), the Euclidean distance d(p, q) between three-dimensional point p and three-dimensional point q is √((x1-y1)). 2 +(x2-y2) 2 +(x³-y³) 2If the position of the three-dimensional point q is less than a certain threshold THd, the encoding device 100 determines that the position of the three-dimensional point q is close to the position of the three-dimensional point p, and determines that the value of the displacement vector of the three-dimensional point q is used in the generation of the predicted value of the displacement vector of the three-dimensional point p.

[0509] Alternatively, other methods can be used to calculate the distance, such as Mahalanobis distance.

[0510] Alternatively, for example, the encoding device 100 may determine that a 3D point that is farther away from the object to be encoded than a specified distance (outside the specified range) will not be used in the prediction. For example, if a 3D point r exists and the distance d(p, r) between the 3D point p and the 3D point r is greater than or equal to a threshold THd, then the encoding device 100 determines that the 3D point r will not be used in the prediction. Furthermore, the specified distance can be arbitrarily determined and is not particularly limited.

[0511] Additionally, the encoding device 100 can append the value of the threshold THd to the header of the bitstream.

[0512] For example, when encoding the displacement vector of a three-dimensional point of the object to be encoded using a predicted value, the encoding device 100 may utilize the already encoded displacement vector or the already decoded displacement vector when using the displacement vector of the surrounding three-dimensional point used in the generation of the predicted value.

[0513] Furthermore, when the decoding device 200 decodes the displacement vector of the three-dimensional point of the decoding object using the predicted value, it utilizes the already decoded displacement vector when using the displacement vector of the surrounding three-dimensional point used in the generation of the predicted value.

[0514] Therefore, the same prediction value is generated during encoding and decoding. As a result, the decoding device 200 is able to correctly decode the bitstream of three-dimensional points generated by the encoding device 100.

[0515] Furthermore, it is stated that points surrounding a 3D point represent other 3D points within a specified range from that point, but are not necessarily limited to this. For example, in Figure 47 In the case of the shown 3D point D (i.e., vertex D), there exist 3D points A, B, C, E, F, and G as surrounding 3D points. However, surrounding 3D points (in other words, neighboring points) can also be selected based on any one or more of the following conditions A and B. That is, neighboring points are points selected according to certain conditions and are referenced for predicting the information of the 3D points of the encoded object. Neighboring points can also be referred to as reference 3D points, reference points, or reference vertices, etc.

[0516] Condition A: A 3D point that is connected to a 3D point of the object. • Condition B: 3D points that have been encoded or decoded prior to the object's 3D points. For example, when selecting 3D points that satisfy conditions A and B as adjacent points, and encoding or decoding 3D point D and its adjacent points in the order of 3D point A, 3D point C, 3D point E, 3D point F, 3D point D, 3D point B, and 3D point G, 3D points A, C, E, and F can be selected as adjacent points of 3D point D. Since 3D points A, C, E, and F are connected to 3D point D, their displacement vector values ​​are likely to be similar. Furthermore, since 3D points A, C, E, and F have been encoded or decoded before 3D point D, their displacement vectors can be used to calculate the predicted value of the displacement vector of 3D point D. Therefore, the accuracy of the predicted value of the displacement vector of 3D point D can be improved, thus increasing encoding efficiency.

[0517] Furthermore, in addition to conditions A and B mentioned above, the number of neighboring points of a 3D point can be limited to a specified value (NumNeiCnt) or less. For example, by setting NumNeiCnt = 3, the number of neighboring points of a 3D point can be limited to three or less. This reduces the memory capacity required to store information about the neighboring points of a 3D point and also reduces the processing load when calculating the predicted displacement vector. Moreover, the specified value can be arbitrarily determined and is not particularly limited.

[0518] Alternatively, for example, the encoding device 100 may append the aforementioned specified value, in other words, NumNeiCnt, representing the maximum number of adjacent points, to the header of the data unit for encoding and thus append it to the bit stream.

[0519] Thus, by decoding the header of the bitstream, the decoding device 200 is able to appropriately decode a bitstream that limits the number of the largest adjacent points to below NumNeiCnt.

[0520] Furthermore, when there are more 3D points satisfying conditions A and B than NumNeiCnt as adjacent points, adjacent points can be selected in order of increasing distance from the 3D point being encoded or decoded. For example, when NumNeiCnt = 3, there are four 3D points A, C, E, and F that satisfy conditions A and B as adjacent points of 3D point D. Also, if 3D points A, C, E, and F are closest to 3D point D in that order, then 3D points A, C, and E can be selected as adjacent points of 3D point D. Since 3D points A, C, and E are connected to 3D point D and are close to it, their displacement vector values ​​are likely to be close to the displacement vector value of 3D point D. Additionally, 3D points A, C, and E may have been encoded or decoded before 3D point D. Therefore, the displacement vectors of three-dimensional points A, C, and E can be used to calculate the predicted value of the displacement vector of three-dimensional point D.

[0521] This improves the accuracy of the predicted displacement vector of a 3D point D. Furthermore, by limiting the number of neighboring points, the memory capacity for storing information about the neighboring points of a 3D point can be reduced, as can the processing power required to calculate the predicted displacement vector.

[0522] Furthermore, the method for selecting neighboring points of a 3D point is not limited to the above. For example, if the 3D point Z to be encoded is generated from 3D points X and Y constituting the base grid through sub-segmentation, then the 3D points X and Y of the base grid can also be considered as neighboring points of 3D point Z. Essentially, when generating 3D point Z from 3D points X and Y through sub-segmentation, since 3D point Z exists on the straight line connecting 3D points X and Y, 3D points X and Y become neighboring points of 3D point Z. Additionally, the displacement vector value of 3D point Z is highly likely to be close to the displacement vectors of 3D points X and Y. By using 3D points X and Y as neighboring points and calculating the predicted value of the displacement vector of 3D point Z using their displacement vectors, encoding efficiency can be improved.

[0523] The following explains the method for generating LoD.

[0524] Figure 60 and Figure 61 This is an explanatory diagram illustrating the method for generating LoD in this embodiment.

[0525] When encoding the displacement vectors of 3D points, the encoding device sometimes uses the positional information of the 3D points to classify them into more than one level before encoding. Here, each level used in the classification is called a LoD (Level of Detail). Each LoD is assigned a unique identifier (e.g., a number). For example, the 0th LoD is also called LoD0, the first LoD is also called LoD1, the nth LoD is also called LoDn, and the (n-1)th LoD is also called LoD(n-1).

[0526] use Figure 60 and Figure 61 This section explains the method for generating LoD (LoD). Furthermore, if the encoding or decoding device cannot calculate the position or distance information of 3D points within the frame of the object to be encoded or decoded, the position or distance information of the 3D points corresponding to those points within the already encoded or decoded frame can be used. Therefore, it is possible to efficiently encode 3D points of the object to be encoded or decoded by classifying them into more than one level.

[0527] Figure 60 Points a0, a1, a2, b0, b1, b2, c0, c1, and c2 are shown as the 3D points to be encoded. Furthermore, d(x, y) represents the distance between point x and point y.

[0528] By setting the threshold values ​​for each layer of LoD to be larger in higher layers (those closer to LoD0), higher layers become more clusters of points with greater distances between them (also known as sparse clusters), while lower layers become more clusters of points with closer distances between them (also known as dense clusters). Here, LoD0 is the highest layer (refer to...). Figure 61 ).

[0529] If point y is at a distance d(x, y) from point x that is greater than the threshold of the LoD to which point x belongs but less than the threshold of the LoD above that LoD, then point y belongs to the same LoD as point x. Furthermore, if point x belongs to LoD0, which is the highest-level LoD, then point y belongs to the same LoD as point x if its distance d(x, y) from point x is greater than the threshold of the LoD to which point x belongs.

[0530] First, the encoding device selects point a0 as the initial point and assigns it to LoD0. Next, the encoding device extracts point a1 whose distance from point a0 is greater than the LoD0 threshold Thres_Lod[0] and assigns it to LoD0. Then, the encoding device extracts point a2 whose distance from point a1 is greater than the LoD0 threshold Thres_Lod[0] and assigns it to LoD0. In this way, the encoding device constructs LoD0 such that the distance between all points within LoD0 is greater than the threshold Thres_Lod[0].

[0531] Next, the encoding device selects point b0, which has not yet been assigned a LoD, and assigns it to LoD1. Then, the encoding device selects point b1, which is at a distance greater than the LoD1 threshold Thres_Lod[1] and has not yet been assigned a LoD, and assigns it to LoD1. Next, the encoding device selects point b2, which is at a distance greater than the LoD1 threshold Thres_Lod[1] and has not yet been assigned a LoD, and assigns it to LoD1. Thus, the encoding device constructs LoD1 such that the distance between each point within LoD1 is greater than the threshold Thres_Lod[1].

[0532] Next, the encoding device selects point c0, which has not yet been assigned a LoD, and assigns it to LoD2. Next, the encoding device selects point c1, which is at a distance greater than the LoD2 threshold Thres_Lod[2] and has not yet been assigned a LoD, and assigns it to LoD2. Next, the encoding device selects point c2, which is at a distance greater than the LoD2 threshold Thres_Lod[2] and has not yet been assigned a LoD, and assigns it to LoD2. In this way, LoD2 is constructed such that the distance between each point in LoD2 is greater than the threshold Thres_Lod[2].

[0533] The threshold for each LoD can be appended to the header of the bitstream. For example, in Figure 60 In the case of Thres_Lod[0], Thres_Lod[1] and Thres_Lod[2], the thresholds Thres_Lod[0], Thres_Lod[1] and Thres_Lod[2] can be appended to the header of the bitstream.

[0534] Additionally, for the lowest level of LoD, all 3D points that haven't yet been allocated LoD can be assigned. In this case, by not attaching a threshold to the lowest level of LoD to the header, the amount of code in the header can be reduced. For example, this could be achieved in... Figure 60 In the case of Thres_Lod[0] and Thres_Lod[1] being appended to the header by the encoding device, and Thres_Lod[2] not being appended to the header, the decoding device infers that Thres_Lod[2] is a value of 0.

[0535] Alternatively, the level number of the LoD can be appended to the header. This allows the decoding device to determine whether the LoD is the lowest level.

[0536] Furthermore, when the LoD level is single-level, that is, when the displacement vectors of 3D points are encoded without generating LoDs, the encoding device can omit the LoD generation process described in the above example. Alternatively, the encoding device can set the LoD level to 1 and apply the LoD generation method described in the above example. In this case, the encoding device can perform the LoD generation process by setting all 3D points to belong to the same LoD. Thus, the encoding device can reduce the processing time for LoD generation.

[0537] Furthermore, the encoding or decoding of displacement vectors described in this embodiment can also be applied to methods other than the LoD generation method described above. For example, if the LoD level to which a 3D point belongs is predetermined, the encoding efficiency can be improved by applying the displacement vector encoding or decoding method described in this embodiment.

[0538] Furthermore, the methods for generating LoD are not limited to those mentioned above; for example, such as... Figure 35 and Figure 36 As shown, the LoD can also be determined based on the number of sub-segments performed from the base mesh. For example, if the base mesh is sub-segmented twice, the 3D points contained in the base mesh can be assigned to LoD0, the points generated from the base mesh's 3D points through one sub-segment can be assigned to LoD1, and the points generated through two sub-segments can be assigned to LoD2. This reduces the processing time for LoD generation.

[0539] The method for selecting the initial 3D point when constructing each LoD can also depend on the encoding order during displacement vector encoding. For example, the encoding device selects the 3D point that was first encoded during displacement vector encoding as the initial point a0 of LoD0, and uses point a0 as the base point to select points a1 and a2 to construct LoD0. Furthermore, the encoding device can also select the 3D point whose displacement vector was first encoded from among the 3D points that do not belong to LoD0 at the current moment as the initial point b0 of LoD1. That is, the encoding device can also select the 3D point whose displacement vector was first encoded from among the 3D points of LoDs that do not belong to a level below LoD(n-1) as the initial point n0 of LoDn. Therefore, during decoding, the same initial point selection method can be used (specifically, the method of selecting the 3D point whose displacement vector was first decoded from among the 3D points of LoDs that do not belong to a level below LoD(n-1) as the initial point n0 of LoDn) to construct the same LoD as during encoding, and the bitstream can be decoded appropriately.

[0540] Figure 62This is an explanatory diagram illustrating the method for generating the predicted value of the displacement vector in this embodiment.

[0541] The encoding device can use the information from the LoD (LoD) to generate predicted values ​​of the displacement vectors of three-dimensional points.

[0542] For example, the encoding device could generate LoD1 using the encoded and decoded displacement vectors contained in LoD0 and LoD1, while sequentially encoding the 3D points contained in LoD0. In this way, the encoding device can use the encoded and decoded displacement vectors contained in LoDn' (where n'≤n) to generate predicted values ​​of the displacement vectors of the 3D points contained in LoDn.

[0543] Furthermore, the predicted value of the displacement vector of a 3D point can be generated by calculating the average of the displacement vectors of a certain number of 3D points that are the neighboring 3D points of the 3D point of the encoded object, which are also the neighboring 3D points of the encoded and decoded object. This certain number is, for example, the number of neighboring 3D points of the 3D point of the encoded object (e.g., N). In this case, the value N is appended to the header of the bitstream, etc.

[0544] Furthermore, the value N, representing the number of neighboring points (i.e., N) used in the calculation of the predicted value, can be appended to each 3D point used to generate the predicted value. Thus, the encoding device can select an appropriate N neighboring points for each 3D point that is the object of the predicted value generation, thereby improving the accuracy of the predicted value and reducing the prediction residual. Additionally, the encoding device can append the value N to the beginning of the bitstream and fix it within the bitstream (in other words, the value N can be used as a fixed value in the encoding of the 3D points contained in the bitstream). Therefore, the encoding device does not need to encode or decode the value N for each 3D point, thus reducing the processing load. Furthermore, the encoding device can also encode the value N separately for each LoD. Therefore, the encoding device has the potential to improve encoding efficiency by selecting an appropriate value N for each LoD.

[0545] Furthermore, the predicted displacement vector of a 3D point can also be calculated based on the weighted average of N encoded and decoded neighboring points. For example, the encoding device can consider using the distance information between the 3D point of the encoded object and its N neighboring points to perform a weighted average. In this regard, using... Figure 62 Please provide an explanation.

[0546] When the encoding device uses a different value N for each LoD, for example, the value of N can be set larger for higher-level LoDs and smaller for lower-level LoDs. In higher-level LoDs, the distances between the 3D points belonging to that LoD are relatively large; therefore, by setting a larger value for N, it is possible to select more surrounding 3D points and perform averaging, thereby improving prediction accuracy. Conversely, in lower-level LoDs, the distances between the 3D points belonging to that LoD are relatively small; therefore, by setting a smaller value for N, it is possible to suppress the averaging processing load and perform more efficient prediction.

[0547] The predicted value of a point P belonging to LoDN is generated based on the reconstructed point P' belonging to LoDN' (where N'≤N). Here, point P' is selected as a neighboring point based on connectivity and distance.

[0548] Furthermore, the predicted values ​​of the displacement vectors can also be calculated using an unweighted average. This reduces the processing load.

[0549] like Figure 62 As shown, point a2 is predicted based on points a0 and a1. Additionally, point b2 is predicted based on points a0, a1, a2, b0, and b1. Alternatively, the selection of neighboring points can vary depending on the number N used in the prediction. For example, when N=5, points a0, a1, a2, b0, and b1 are selected as neighbors of point b2; when N=4, points a0, a1, a2, and b1 are selected based on distance information.

[0550] Additionally, as an example, 3D points contained in the base mesh are included in LoD0, 3D points generated from the base mesh through a 1st sub-segment are included in LoD1, and 3D points generated from the base mesh through a 2nd sub-segment are included in LoD2.

[0551] For example, when using a weighted average of adjacent points in the prediction, the predicted value a2p of point a2 is calculated by the weighted average of points a0 and a1 (refer to Equations (1) and (2)). Here, A i It is the value of the displacement vector of point ai.

[0552] [Number 1] in, [Number 2] Furthermore, the predicted value b2p for point b2 is calculated using a weighted average of points a0, a1, a2, b0, and b1 (refer to Equations (3), (4), and (5)). Here, B i It is the value of the displacement vector of point bi.

[0553] [Number 3] in, [Number 4] [Number 5] Alternatively, references at the same level can be omitted when generating predicted displacement vector values. This reduces the processing load. Furthermore, when generating 3D points at the midpoint between two 3D points through sub-segmentation, the weight w can be adjusted. i The value is fixed at 0.5. This reduces the processing volume.

[0554] Alternatively, when encoding the displacement vector value of a three-dimensional point, the encoding device can calculate the difference between the predicted value generated based on the neighboring points of the three-dimensional point and the three-dimensional point itself (also called the transformation coefficient, see equations (6) and (7) below), and encode the value by quantizing the calculated transformation coefficient. Here, transformation coefficient a2c is the transformation coefficient of point a2, and transformation coefficient b2c is the transformation coefficient of point b2.

[0555] [Number 6] [Number 7] For example, an encoding device can perform quantization by dividing the transform coefficients by the quantization scale. In this case, the smaller the quantization scale, the smaller the error that may be generated by quantization (quantization error), and conversely, the larger the quantization scale, the larger the quantization error.

[0556] The quantized value of transform coefficient a2c is set as quantization value a2q, and the quantized value of transform coefficient b2r is set as quantization value b2q (refer to Equations 8 and 9 below). QS_LoD0 is the quantization ratio of LoD0, and QS_LoD1 is the quantization ratio of LoD1.

[0557] [Number 8] [Number 9] Furthermore, the encoding device can change the quantization ratio for each LoD. For example, the higher the LoD, the smaller the quantization ratio, and the lower the LoD, the larger the quantization ratio. The displacement vector value of a 3D point belonging to a higher layer might be used as the prediction value of the displacement vector of a 3D point belonging to a lower layer. Therefore, by reducing the quantization ratio of the higher layer, quantization errors that may occur at the higher layer can be suppressed, the accuracy of the prediction value can be improved, and thus the encoding efficiency can be improved. Additionally, the encoding device can append the quantization ratio to the header, etc., for each LoD. Thus, the encoding device can help the decoding device correctly decode the quantization ratio and appropriately decode the bitstream.

[0558] Additionally, the encoding device can convert the quantized transform coefficients from signed integer values ​​to unsigned integer values. For example, the encoding device can convert the quantized value a2q, which is a signed integer value, to the quantized value a2u, which is an unsigned integer value, as described below.

[0559] When the quantized value a²q is less than 0: a²u = -1 - (2 × a²q) For cases other than those mentioned above: a2u=2×a2q (Equation 10) Furthermore, for example, the encoding device can convert the quantized value b2q, which is a signed integer value, into the quantized value b2u, which is an unsigned integer value, as described below.

[0560] When the quantized value b2q is less than 0: b2u = -1 - (2 × b2q) For cases other than those mentioned above: b2u = 2 × b2q (Equation 11) Therefore, the encoding device has the advantage of not needing to consider the generation of negative integers when entropy encoding the transform coefficients.

[0561] In addition, the encoding device does not necessarily need to convert signed integer values ​​to unsigned integer values; for example, the signed bits can be entropy encoded separately.

[0562] Furthermore, the encoding method for transform coefficients is not limited to this. For example, the encoding device can also perform arithmetic encoding on a per-bit basis using context, such as the sign bit representing the positive or negative of the transform coefficient and the binary data representing the absolute value of the transform coefficient. Therefore, the encoding device has the potential to improve the encoding efficiency of the transform coefficients of the displacement vector.

[0563] Furthermore, if the quantization of the transform coefficients of the displacement vector is not required, this process can be skipped, and the transform coefficients can be directly arithmetically encoded. This reduces processing time.

[0564] Figure 63 This is an explanatory diagram illustrating an example of the calculation of the predicted value in this embodiment.

[0565] Reference Figure 63 This example illustrates how to generate the LoD and calculate the predicted values ​​of the displacement vectors for each 3D point.

[0566] exist Figure 63 In the diagram, points a0, a1, and a2 are 3D points contained in the base mesh and belong to LoD0. Points b0 and b1 are 3D points generated from the 3D points contained in the base mesh through a first-order sub-segmentation and belong to LoD1. Points c0, c1, c2, and c3 are 3D points generated from the 3D points contained in the base mesh through a second-order sub-segmentation and belong to LoD2.

[0567] When point b0 is a 3D point generated by subsegmenting points a0 and a1, the predicted value of the displacement vector of point b0 can be calculated using points a0 and a1.

[0568] Furthermore, when point c2 is a 3D point generated by subsegmenting points a1 and b1, the predicted value of the displacement vector of point c2 can be calculated using points a1 and b1.

[0569] Figure 64 This is an explanatory diagram illustrating an example of the calculation of the transformation coefficients in this embodiment.

[0570] Reference Figure 64 This illustrates an example of calculating the transformation coefficients by subtracting the predicted values ​​from the displacement vectors of each three-dimensional point.

[0571] exist Figure 64 In this diagram, transformation coefficients a0c, a1c, and a2c are transformation coefficients contained in the base mesh, belonging to LoD0. Transformation coefficients b0c and b1c are 3D points generated from the 3D points contained in the base mesh through a first-order sub-segmentation, belonging to LoD1. Transformation coefficients c0c, c1c, c2c, and c3c are 3D points generated from the 3D points contained in the base mesh through a second-order sub-segmentation, belonging to LoD2.

[0572] The transformation coefficient b0c of point b0 is obtained by subtracting the predicted value b0p of point b0 from the value of point b0.

[0573] b0c=b0-b0p The predicted value b0p can be the average of the values ​​at point a0 and point a1.

[0574] b0p=(a0+a1) / 2 In addition, the transformation coefficient c2c of point c2 is obtained by subtracting the predicted value c2p of point c2 from the value of point c2.

[0575] c2c = c2 - c2p The predicted value c2p can be the average of the values ​​at point a1 and point b1.

[0576] c2p=(a1+b1) / 2 Alternatively, a lifting transform can be applied: the transform coefficients of the displacement vectors of the 3D points contained in the lower layer of the LoD are calculated, and these transform coefficients are fed back to the upper layer for encoding. By applying the lifting transform, the transform coefficients of the low-frequency components of the displacement vectors can be collected in the upper layer, and the transform coefficients of the high-frequency components of the displacement vectors can be collected in the lower layer. Thus, for example, by quantizing the transform coefficients of the high-frequency components in the lower layer, the amount of information can be reduced, thereby improving encoding efficiency.

[0577] Figure 65 This is an explanatory diagram illustrating an example of inter-frame prediction of transform coefficients in this embodiment.

[0578] The transform coefficients of the displacement vectors of three-dimensional points can also be used for inter-frame prediction using the transform coefficients of the displacement vectors of frames that are temporally different from the frame being encoded. (See reference...) Figure 65 This needs to be explained.

[0579] Transformation coefficients of displacement vectors of 3D points, for example, can be used for inter-frame prediction by using the transformation coefficients of displacement vectors of frames that have just been encoded or decoded.

[0580] For example, if Frame(t) at time t is the frame to be encoded, inter-frame prediction can be performed using the transformation coefficients of the displacement vector of the frame that was just encoded or decoded, i.e., Frame(t-1) at time t-1.

[0581] Specifically, we consider using the transform coefficient a0c' of the displacement vector of the 3D point a0 corresponding to a0 in Frame(t-1), and use inter-frame prediction to encode the transform coefficient a0c of the displacement vector of the 3D point a0 in Frame(t). More specifically, we consider encoding the value after subtracting a0c' from a0c. When the 3D point a0 in Frame(t) and the 3D point a0' in Frame(t-1) correspond between frames, the values ​​of their displacement vectors are likely to be close. Therefore, by subtracting a0c' from a0c, we can further reduce the transform coefficient and improve the coding efficiency of entropy coding.

[0582] Furthermore, inter-frame prediction is not limited to referencing the frame just encoded; it can refer to any frame. In this case, information from the referenced frame can be appended to the header. This allows the decoding device to appropriately decode the bitstream by referring to the same frame referenced by the encoding device. Additionally, inter-frame prediction can be performed by referring to multiple frames. For example, by using double prediction with two reference frames, coding efficiency can be improved. Furthermore, when using double prediction, the average of the transform coefficients of each displacement vector in the two reference frames can be used as the inter-frame prediction value. This allows for the generation of prediction values ​​with high accuracy, further improving coding efficiency.

[0583] Alternatively, the information `disp_inter_mode`, indicating whether inter-frame prediction is applied to the transform coefficients of the displacement vector, can be appended to the header. Thus, for example, when the inter-frame motion changes significantly and the correspondence between 3D points in the frames cannot be obtained, inter-frame prediction for the transform coefficients of the displacement vector can be turned off (e.g., `disp_inter_mode = 0`), and when the inter-frame motion changes are small and the correspondence between 3D points in the frames can be obtained, inter-frame prediction can be turned on (e.g., `disp_inter_mode = 1`). By adaptively controlling the on / off state of inter-frame prediction in this way, coding efficiency can be improved.

[0584] Furthermore, if the adaptive switching of inter-frame prediction is performed on a frame-by-frame basis, `disp_frame_inter_mode` can be appended to the header storing frame information; if it is performed on a sequence-by-sequence basis, `disp_seq_inter_mode` can be appended to the header storing sequence information. This allows for the control of enabling or disabling inter-frame prediction of the shift vector transform coefficients on a frame-by-frame or sequence-by-sequence basis, thereby improving coding efficiency.

[0585] Alternatively, inter-frame prediction information (disp_lod_inter_mode) indicating whether the transform coefficients of the displacement vector are applied can be prepared for each LoD, and the application of inter-frame prediction can be switched on a per LoD basis. For example, consider the following scenario: the coding unit compares the generated code amounts for the case where inter-frame prediction is applied to the transform coefficients of the displacement vector with and without application for each LoD, selects the one with the smaller generated code amount, and appends its information to the header. The decoding unit decodes the transform coefficients of the displacement vector based on the information appended to the header. Thus, by switching the application of inter-frame prediction on a per LoD basis, coding efficiency can be improved.

[0586] The encoding device can decode the quantized transform coefficients through inverse quantization and reconstruction, and use them for subsequent prediction of the 3D points of the encoded object. Specifically, the encoding device can calculate the inverse quantization value by multiplying the quantized transform coefficients by the quantization ratio, and add the inverse quantization value to the predicted value to obtain the decoded value. For example, the encoding device can calculate the inverse quantization value a2iq based on the quantized value a2q as follows, and can calculate the inverse quantization value b2iq based on the quantized value b2q as follows.

[0587] a2iq=a2q×QS_LoD0 b2iq=b2q×QS_LoD1 (Equation 12) Furthermore, the encoding device can calculate the reconstructed value a2rec based on the inverse quantization value a2iq as follows, and can also calculate the reconstructed value b2rec based on the inverse quantization value b2iq as follows.

[0588] a2rec=a2iq+a2p b2rec=b2iq+b2p (Equation 13) Furthermore, this embodiment describes a method for generating predicted values ​​of displacement vectors of three-dimensional points by constructing one or more LoDs using an encoding device, but it is not limited to this. For example, it can also be applied to the case of generating predicted values ​​of displacement vectors of three-dimensional points by constructing a single LoD, or to the case of generating predicted values ​​of displacement vectors of three-dimensional points without generating LoDs.

[0589] In this case, all 3D points belong to the same LoD (e.g., LoD0). Therefore, when the encoding device encodes or decodes the 3D points contained in LoD0 sequentially, it can also use the already encoded and decoded displacement vectors contained in LoD0 to generate predicted values ​​for the 3D points belonging to LoD0. In this way, the encoding device may be able to reduce processing time by encoding without generating multiple levels of LoDs.

[0590] Furthermore, when the quantization of the transform coefficients of the displacement vector is not required, the encoding device can skip the quantization and inverse quantization processes and directly add the arithmetically decoded transform coefficients to the predicted value to obtain the decoded value. This reduces processing time.

[0591] Figure 66 This is an explanatory diagram illustrating an example of the syntax of this implementation method.

[0592] Figure 66 The example syntax shown illustrates an example of the structure of information contained in the bitstream generated by the encoding device.

[0593] Figure 66The syntax shown includes a displacement vector_header. The displacement vector_header includes NumLoD, NumOfPoint[i], Thres_Lod[i], NumNeiCnt[i], THd[i], and QS[i].

[0594] NumLoD represents the number of levels in a LoD.

[0595] NumOfPoint[i] represents the number of 3D points belonging to level i. Alternatively, if the encoding device appends the total number of 3D points, AllNumOfPoint, to other headers, NumOfPoint[NumLoD-1] (i.e., the number of 3D points belonging to the lowest level) may not be appended to the header. In this case, NumOfPoint[NumLoD-1] can be calculated using the following (Equation 14).

[0596] [Number 10] Thres_Lod[i] represents the threshold of the LoD at level i. The encoding device constructs the LoDi such that the distance between points within the LoDi is greater than the threshold Thres_Lod[i]. Furthermore, the value of Thres_Lod[NumLoD-1] (i.e., the threshold of the lowest-level LoD) may not be appended to the header. In this case, Thres_Lod[NumLoD-1] can be estimated to be 0. This allows for a reduction in the amount of code in the header.

[0597] NumNeiCnt[i] represents the upper limit of the number of neighboring points used to generate the predicted value of a 3D point belonging to level i. When the number of neighboring points M is less than NumNeiCnt[i] (i.e., M < NumNeiCnt[i]), the encoding device can also use M neighboring points to calculate the predicted value. Additionally, if it is not necessary to make the value of NumNeiCnt[i] different for each LoD, the encoding device can also append one NumNeiCnt to the header.

[0598] THd[i] represents the upper limit of the distance of 3D points used in the prediction of 3D points of the object being encoded or decoded in level i. The encoding device may also exclude 3D points whose distance from the 3D points of the object being encoded or decoded is greater than THd[i] from the prediction. In addition, if it is not necessary to make the value of THd[i] different in each LoD, one THd may be appended to the header.

[0599] QS[i] represents the quantization ratio of level i.

[0600] Alternatively, the encoding device can entropy-encode NumLoD, Thres_Lod[i], NumNeiCnt[i], THd[i], or QS[i] and append them to the header. For example, consider the encoding device binarizing each value and performing arithmetic encoding. Furthermore, the encoding device can also encode with a fixed length to reduce processing overhead.

[0601] Furthermore, the encoding device does not necessarily need to append NumLoD, Thres_Lod[i], NumNeiCnt[i], THd[i], or QS[i] to the header; for example, these can be specified by the profile or level of the standard. This allows for a reduction in the number of bits in the header.

[0602] Figure 67 This is an explanatory diagram illustrating an example of the syntax of this implementation method.

[0603] Figure 67 The example syntax shown illustrates an example of the structure of information contained in the bitstream generated by the encoding device.

[0604] Figure 67 The syntax shown includes displacement vector_data. Displacement vector_data may contain dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k] related to each level from the 0th to the NumLoDth of LoD (also known as the jth level).

[0605] `dispd_is_zero[k]` indicates whether the absolute value of the transformation coefficient of the k-th component of the displacement vector of the i-th 3D point (vertex[i]) contained in the j-th level of LoD is 0. A value of 1 indicates that the absolute value of the transformation coefficient of the k-th component is 0, and a value of 0 indicates that the absolute value of the transformation coefficient of the k-th component is 1 or higher.

[0606] `dispd_is_one[k]` indicates whether the absolute value of the transformation coefficient of the k-th component of the displacement vector of the i-th 3D point (vertex[i]) contained in the j-th level of LoD is 1. A value of 1 indicates that the absolute value of the prediction residual of the k-th component is 1, and a value of 0 indicates that the absolute value of the prediction residual of the k-th component is 2 or more.

[0607] Furthermore, if dispd_is_one[k] is not included in the bitstream, the decoding device can also presume this value to be 0. This prevents setting an indeterminate value for dispd_is_one[k] during decoding, allowing for appropriate decoding processing.

[0608] dispd_minus2[k] represents the value obtained by subtracting the value 2 from the absolute value of the transformation coefficient of the k-th component of the displacement vector of the i-th 3D point (vertex[i]) contained in the j-th level of LoD.

[0609] Furthermore, if dispd_minus2[k] is not included in the bitstream, the decoding device can presume this value to be 0. This prevents setting an indeterminate value for dispd_minus2[k] during decoding, allowing for appropriate decoding processing.

[0610] dispd_sign[k] represents the sign bit of the displacement vector of the k-th component of the ith 3D point (vertex[i]) contained in the j-th level of the LoD. A value of 1 indicates that the transformation coefficient of the k-th component is negative, and a value of 0 indicates that the transformation coefficient of the k-th component is positive.

[0611] Furthermore, regarding the k-th component, when the displacement vector is represented by an orthogonal coordinate system, the first component can represent the x-component, the second component the y-component, and the third component the z-component. Alternatively, when the displacement vector is represented by a local coordinate system, the first component can represent the normal component, the second component the tangent component, and the third component the binormal component. Thus, regardless of whether the displacement vector is represented by an orthogonal coordinate system or a local coordinate system, a common syntactic structure can be utilized.

[0612] Furthermore, the transformation coefficients dispd[k] of the k-th component of the displacement vector of the i-th 3D point (i.e., vertex[i]) can be obtained using the information described above. Figure 68 The calculation is performed using the operations shown.

[0613] Import Figure 67 The syntax structure shown allows the encoding device to reduce the frequency of encoding and appending dispd_is_one[k], dispd_minus2[k], and dispd_sign[k] to the bitstream when encoding transform coefficients that are prone to becoming dispd[k] = 0, thereby potentially improving encoding efficiency. Furthermore, for example, when encoding prediction residuals that are prone to becoming dispd[k] = 1 or 0, the encoding device can reduce the frequency of encoding and appending dispd_minus2[k] to the bitstream, thereby potentially improving encoding efficiency.

[0614] Furthermore, this embodiment shows an example where the transform coefficients are assumed to be easily 0 or 1 (dispd[k] = 0 or 1), but it is not limited to this; the same processing can be applied to any dispd[k]. For example, when encoding transform coefficients that are easily 2 (dispd[k] = 2), dispd_is_two[k] and dispd_minus3[k] can be newly introduced. Therefore, when encoding transform coefficients that are easily 2 (dispd[k] = 2), the frequency of encoding and appending dispd_minus3[k] to the bitstream can be reduced, resulting in the possibility of improving encoding efficiency. Additionally, in this case, it is possible to... Figure 69 The computation process shown is used to calculate dispd[k].

[0615] Alternatively, the encoding device can binarize at least one of dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k], and apply context-based arithmetic coding. For example, since dispd_is_zero[k], dispd_is_one[k], and dispd_sign[k] are each 1 bit, the encoding device can allocate one context for each of them, update the probability of occurrence based on the frequency of occurrence of 0 and 1, and then encode them. Therefore, it is possible to improve the encoding efficiency. Alternatively, the encoding device can binarize dispd_minus2[k] using Exponential Golomb, allocate context for each bit, update the probability of occurrence based on the frequency of occurrence of 0 and 1, and then encode them. Therefore, it is possible to improve the encoding efficiency.

[0616] Furthermore, the encoding device can also allocate different contexts for each component of dispd as the context assigned to dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k]. This could potentially improve encoding efficiency when the value of dispd differs for each component. Alternatively, the encoding device can also allocate the same context for each component of dispd as the context assigned to dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k]. This could potentially improve encoding efficiency when the values ​​of the components of dispd are similar.

[0617] The decoding device can transform the decoded quantized transform coefficients from unsigned integer values ​​to signed integer values ​​using the opposite method to the encoding device. Therefore, by entropy encoding the transform coefficients, it is possible to appropriately decode bitstreams generated without considering the generation of negative integers.

[0618] Furthermore, it is not always necessary to transform unsigned integer values ​​into signed integer values. For example, when decoding a bitstream generated by entropy encoding the sign bit, the decoding device can also decode the sign bit. Moreover, the decoding method for the transform coefficients is not limited to this; for example, arithmetic decoding can be performed on the binary data representing the sign of the transform coefficient and the absolute value of the transform coefficient, using context for each bit. Thus, the decoding device can appropriately decode bitstreams that improve the encoding efficiency of the transform coefficients of the shift vector.

[0619] The decoding device decodes the quantized transform coefficients, which have been transformed into signed integer values, through inverse quantization and reconstruction, and uses this information for subsequent prediction of the 3D points of the decoded object. Specifically, the decoding device calculates the inverse quantization value by multiplying the quantized transform coefficients by the quantization ratio, and adds the inverse quantization value to the predicted value to obtain the decoded value.

[0620] For example, the decoded unsigned quantized value a2u is transformed into the signed value a2q as shown below. Additionally, ">>" indicates a shift operation.

[0621] When the least significant bit (LSB) of a2u is 1; a2q = -((a2u+1) >> 1) In cases other than those mentioned above: a2q = (a2u >> 1) (Equation 15) Alternatively, for example, the decoded unsigned quantized value b2u is transformed into a signed value b2q as described below.

[0622] The case where b2u's LSB is 1; b2q = -((b2u+1)>>1) In cases other than those mentioned above: b2q = (b2u >> 1) (Equation 16) The decoding device calculates the reconstructed value after inverse quantization. The reconstructed value can be used for subsequent predictions of the 3D points of the decoded object.

[0623] For example, the decoding device can calculate the inverse quantization value a2iq based on the quantization value a2q as follows, and can calculate the inverse quantization value b2iq based on the quantization value b2q as follows.

[0624] a2iq=a2q×QS_LoD0 b2iq=b2q×QS_LoD1 (Equation 17) Furthermore, the decoding device can calculate the reconstructed value a2rec based on the inverse quantization value a2iq as follows, and can calculate the reconstructed value b2rec based on the inverse quantization value b2iq as follows.

[0625] a2rec=a2iq+a2p b2rec=b2iq+b2p (Equation 18) The above example illustrates how the encoding device calculates and generates the average of the displacement vectors of a certain number of neighboring 3D points of the encoded and decoded 3D point of the coded object as the predicted value of the displacement vector of the 3D point. However, it is not limited to this example and the predicted value can be generated by other methods.

[0626] For example, the encoding device can directly use the displacement vector of the nearest 3D point among the encoded and decoded neighboring 3D points of the 3D point of the object to be encoded as the prediction value. Furthermore, the encoding device can append a prediction mode value (PredMode) to each 3D point, allowing selection of prediction values. For example, the encoding device can set the total number of prediction modes M, assign an average value to prediction mode 0, assign the displacement vector of 3D point A to prediction mode 1, ..., assign the displacement vector of 3D point Z to prediction mode M-1, and append the prediction mode used for prediction to the bitstream for each 3D point. The 3D points A to Z for which displacement vectors are assigned to prediction modes 1 to prediction mode M-1 can be used sequentially, starting with the closest 3D point among the encoded and decoded neighboring 3D points of the object to be encoded.

[0627] Figure 70 This is an explanatory diagram illustrating an example of the predicted value information of the displacement vector in this embodiment. Figure 71 This is an explanatory diagram illustrating the method for generating the predicted value of the displacement vector in this embodiment.

[0628] Figure 70 An example of prediction information for point b2 is shown when the number of adjacent 3D points N used in the prediction is 4 and the number of prediction patterns M is 5. The prediction information includes information indicating the prediction values ​​used in each prediction pattern for more than one prediction pattern. Figure 70 The example of the predicted value information shown is a table representing the predicted values ​​used in each of the more than one prediction models.

[0629] exist Figure 70 In the example shown, the predicted values ​​used in the prediction of point b2 are, for example, points a0, a1, a2, and b1, which are adjacent 3D points (see reference). Figure 71Accordingly, the "average of points a0, a1, a2 and b1" is assigned as the predicted value for prediction mode 0.

[0630] In addition, Figure 70 In the above, "point b1" is assigned as a predicted value in prediction mode 1. "point b2" is assigned as a predicted value in prediction mode 2. "point a1" is assigned as a predicted value in prediction mode 3. "point a0" is assigned as a predicted value in prediction mode 4.

[0631] Additionally, the numerical value that uniquely represents a prediction pattern is also called the prediction pattern value. Here, we will assume that the prediction pattern value of prediction pattern m is m. As an example, prediction pattern values ​​are used sequentially starting from the smallest integer value.

[0632] The allocation of prediction mode values ​​can be determined based on the order of their distances from the 3D points of the object being encoded. For example, the closer a 3D point is to the 3D point of the object being encoded, the smaller the prediction mode value the encoding device assigns it. In the example above, the 3D point with the smallest distance to the 3D point b2 (i.e., closest to 3D point b2) is point b1, the second smallest distance to 3D point b2 is point a2, the third smallest distance to 3D point b2 is point a1, and the fourth smallest distance to 3D point b2 is point a0.

[0633] Therefore, due to the small distance, the difference between the displacement vector and the predicted value is relatively small. This allows for the allocation of smaller prediction mode values ​​to points with a higher probability of being selected as predicted values, thus reducing the number of bits used to encode the prediction mode values. Additionally, smaller prediction mode values ​​can be preferentially assigned to 3D points belonging to the same LoD as the 3D points of the coded object.

[0634] in addition, Figure 72 An example of the predicted value information for point a2 is shown when the number of neighboring 3D points N is 2 and the number of prediction patterns M is 5.

[0635] exist Figure 72 In the example of predicted value information shown, the prediction of point a2 uses predicted values, for example, points a0 and a1, which are adjacent 3D points. Accordingly, in Figure 72 In the above, the "average of points a0 and a1" is assigned as the predicted value for prediction mode 0.

[0636] In addition, "point a1" is assigned as a predicted value in prediction mode 1. "point a0" is assigned as a predicted value in prediction mode 2.

[0637] In addition, when there are fewer than 4 adjacent points, for prediction modes with unassigned prediction values, information indicating that they are not being used can be set (referred to as "not available" in the figure).

[0638] In addition, Figure 73 The image shows an example of predicted value information when the displacement vector is represented by an orthogonal coordinate system (XYZ coordinate system).

[0639] exist Figure 73 In the example shown, the values ​​used in the prediction of point b2 are, for example, points a0, a1, a2, and b1, which are adjacent 3D points (see reference). Figure 71 Correspondingly, in Figure 73 In the prediction model 0, the predicted values ​​are assigned coordinates (Xave, Yave, Zave) representing the average of points a0, a1, a2, and b1. Here, Xave can be calculated using the average or weighted average of Xb1, Xa2, Xa1, and Xa0. Yave can be calculated using the average or weighted average of Yb1, Yb2, Ya1, and Ya0. Zave can be calculated using the average or weighted average of Zb1, Zb2, Za1, and Za0.

[0640] Furthermore, the coordinates of "point b1" (Xb1, Yb1, Zb1) were assigned as the predicted value for prediction mode 1. The coordinates of "point b2" (Xa2, Ya2, Za2) were assigned as the predicted value for prediction mode 2. The coordinates of "point a1" (Xa1, Ya1, Za1) were assigned as the predicted value for prediction mode 3. The coordinates of "point a0" (Xa0, Ya0, Za0) were assigned as the predicted value for prediction mode 4.

[0641] For example, the encoding device may select prediction mode 2 (i.e., prediction mode value 2) and encode the XYZ components of the displacement vector of the three-dimensional point of the object to be encoded using prediction values ​​Xa2, Ya2, and Za2, respectively. In this case, the encoding device appends prediction mode value 2 to the bitstream.

[0642] Furthermore, the above example illustrates the case where the displacement vector is in an orthogonal coordinate system, but it is not limited to this. For example, it can also be applied to displacement vectors represented by local coordinate systems.

[0643] Alternatively, the prediction mode number M can be appended to the bitstream. Alternatively, the prediction mode number M can be left unattached to the bitstream, its value specified by a standard profile or level. Furthermore, the prediction mode number M can also be set to a value calculated based on the number of three-dimensional points N used in the prediction (e.g., M = N + 1).

[0644] When the encoding device generates a predicted value for the displacement vector of a three-dimensional point by attaching a prediction mode value (PredMode) to each three-dimensional point, as an example of the method of assigning prediction values ​​to each prediction mode, an example is shown of using the distance information of the three-dimensional point of the object being encoded to assign the displacement vector of the adjacent point as the prediction value. However, it is not limited to this, and the method of assigning prediction values ​​to the prediction mode can also be changed by some means.

[0645] For example, the encoding device can calculate a median value based on the predicted values ​​assigned to each prediction mode and assign the calculated median value to prediction mode 0. In this way, the encoding device can assign the median value as a predicted value to the prediction mode with the smaller predicted value. Therefore, the encoding device can generate candidate predicted values ​​that prioritize the median value of the displacement vectors of adjacent points, thus improving encoding efficiency.

[0646] Reference Figures 74-76 The changes to the allocation of forecast values ​​that used the median value are explained.

[0647] Figure 74 This is an explanatory diagram illustrating an example of predicted value information for the displacement vector in an implementation method. Figure 75 This is an explanatory diagram illustrating the method for generating the predicted value of the displacement vector in this embodiment. Figure 76 This is an explanatory diagram illustrating an example of the predicted value information of the displacement vector in this embodiment.

[0648] exist Figure 75 In the example shown, the number of 3D points used in the prediction is N = 4, and the number of prediction patterns is M = 4. Point a2 is predicted based on points a0 and a1. Additionally, point b2 is predicted based on points a0, a1, a2, b0, and b1.

[0649] Here, as an example, the following situation is illustrated: points b1, a2, a1, and a0 are arranged in order of their closest distance to the 3D points of the encoded object. The displacement vectors of the closest 3D points are assigned to the prediction modes with smaller prediction mode values. The magnitudes of each prediction value are set as b1 > a1 > a0 > a2.

[0650] The encoding device calculates the median value of the predicted values ​​for the prediction pattern. For example, the encoding device can sort the n predicted values ​​assigned to the prediction pattern in ascending or descending order and use the (n / 2)th value as the median value. Alternatively, the method for calculating the median value can be switched between cases where n is odd and cases where n is even.

[0651] For example, when n is odd, the encoding device can set the (n / 2)th predicted value (with decimal places discarded) from the sorted 0th to (n-1)th predicted values ​​as the center value. Furthermore, when n is even, the encoding device can use the (n / 2-1)th and (n / 2)th predicted values ​​from the sorted 0th to (n-1)th predicted values ​​as candidates A and B for the center value, and adopt one of A and B as the center value using a certain method. For example, the one of A and B that is closer to the 3D point of the object to be encoded can be set as the center value.

[0652] exist Figure 75 In the example shown, since n = 4, the median value can be calculated using the same method as when n is even. For example, if b1, a1, a0, and a2 are sorted in ascending order, they become a2, a0, a1, and b1. In this case, the (n / 2-1)th predicted value is a0, and the n / 2th is a1, which are chosen as candidates A and B for the median value. Furthermore, a1 is closer to the 3D point of the encoded object than a0, so a1 is selected as the median value.

[0653] In this case, such as Figure 76 As shown, the encoding device assigns the predicted value a1, selected as the median, to prediction mode 0, and assigns the predicted value b1, originally assigned to prediction mode 0, to prediction mode 2, which has been assigned the predicted value a1. In other words, the encoding device swaps the predicted values ​​of prediction mode 0 and prediction mode 2. Thus, the encoding device can generate candidate predicted values ​​that prioritize the median value of the displacement vectors of adjacent points, thereby improving encoding efficiency.

[0654] Furthermore, as a method for assigning predicted values ​​to prediction modes, an example using a median value is shown, but it is not limited to this. For example, the encoding device can calculate an average value based on the predicted values ​​assigned to each prediction mode and assign predicted values ​​close to the average value to prediction mode 0. Thus, it is possible to generate candidate predicted values ​​that prioritize displacement vectors that are close to the average of the displacement vectors of neighboring points, thereby improving encoding efficiency.

[0655] Alternatively, the encoding device may initially calculate a central value based on the displacement vectors of adjacent points and assign it to prediction mode 0, and then assign the displacement vectors of the surrounding three-dimensional points other than the central value to prediction modes 1 and later using the distance information of those three-dimensional points.

[0656] Furthermore, the encoding device can append information indicating whether the center value is prioritized (also called center value priority information) to the header, etc. Alternatively, if the center value priority information indicates that the center value is prioritized, the encoding device can assign the center value to prediction mode 0 using the method described above; otherwise, it can assign a prediction value to the prediction mode regardless of the center value. Thus, by adaptively switching between the desired center value priority and other conditions while encoding, the encoding device may be able to improve encoding efficiency. Furthermore, the decoding device can appropriately decode the bitstream based on the center value priority information appended to the header, etc.

[0657] Additionally, as an example of a change in the assignment of forecast values ​​that prioritizes the median, an example is shown where the median is assigned to forecast pattern 0 and the forecast value originally assigned to forecast pattern 0 is replaced with the forecast pattern assigned the median, but this is not a limitation.

[0658] Figure 77 This is an explanatory diagram illustrating an example of the predicted value information of the displacement vector in this embodiment.

[0659] For example, it could be, such as Figure 77 As shown, the encoding device assigns a central value to prediction mode 0, assigns the predicted value originally assigned to prediction mode 0 to prediction mode 1, assigns the predicted value originally assigned to prediction mode 1 to prediction mode 2, and so on, shifting the predicted values ​​assigned to each prediction mode until a value is reassigned to the prediction mode that originally had a central value assigned to it. This allows the generation of predicted value information that prioritizes the central value of the displacement vectors of adjacent points and prioritizes candidate predicted values ​​that are close to each other, thus improving encoding efficiency.

[0660] Additionally, examples of forecast assignments prioritizing the median or average are shown, but this is not a limitation.

[0661] Figure 78 and Figure 79 This is an explanatory diagram illustrating an example of the predicted value information of the displacement vector in this embodiment.

[0662] For example, such as Figure 78 As shown, the encoding device calculates statistical information about the predicted values ​​of the prediction pattern. This statistical information can include, for example, the median, mean, variance, or standard deviation of neighboring points. Furthermore, the encoding device can adjust the allocation of predicted values ​​based on the calculated statistical information (see reference). Figure 79 ).

[0663] Reference Figures 80-83 This section describes a variation of setting the prediction value for a 3D point 'a' of the encoded object in the encoded object frame.

[0664] Figure 80 This is an explanatory diagram showing an example of a point representing the encoding object in this embodiment. Figure 81 This is an explanatory diagram illustrating an example of the predicted value information of the displacement vector in this embodiment. Figure 82 This is an explanatory diagram showing an example of a temporal dv in this embodiment. Figure 83 This is an explanatory diagram illustrating an example of the predicted value information of the displacement vector in this embodiment.

[0665] Encoding devices can be like Figure 81 As shown in the predicted value information, the three-dimensional point a of the encoded object in the encoded object frame is set (refer to...). Figure 80 The encoding device can set 0 (no prediction) as the prediction value for prediction mode 0, and set the average value of the displacement vectors (dv0, dv1, and dv2, respectively) of adjacent points a, b, and c as the prediction value for prediction mode 1. Furthermore, the encoding device can set the displacement vectors (dv0, dv1, and dv2, respectively) of adjacent points a, b, and c as the prediction values ​​for prediction modes 2, 3, and 4, respectively.

[0666] Furthermore, the predicted values ​​assigned to each prediction model are not limited to these; other predicted values ​​can also be assigned.

[0667] Furthermore, the encoding device can, for example, use a reference frame (reference frame) that is different from the frame to be encoded. Figure 82 The displacement vector within the reference frame is assigned to the prediction value. Specifically, as the prediction value for the 3D point a of the encoded object, the encoding device can use the displacement vector (hereinafter referred to as temporal dv) of the corresponding point a' of the 3D point a of the encoded object within the reference frame. When the object moves with a certain motion, the value of the displacement vector of the 3D point of the encoded object is likely to be a value that is close to the value of the displacement vector of the corresponding point of the 3D point of the encoded object within the reference frame. Therefore, by adding the temporal dv as the prediction value to the prediction candidate, it is possible to improve the encoding efficiency.

[0668] Figure 83 An example of prediction information is shown when the prediction value of prediction mode 5 is appended as a temporary dv. Alternatively, temporary dv can also be appended as the prediction value of other prediction modes (i.e., any one of prediction modes 0-4). Furthermore, it is also possible to... Figure 81 The predicted value of any prediction mode shown in the prediction information is changed to temporary dv.

[0669] Furthermore, the encoding device can calculate the temporal dv according to the prediction unit (DVG) of the displacement vector within the reference frame and store it in memory as the temporal dv of the corresponding point a', using the temporal dv of the DVG to which the corresponding point a' belongs. This reduces the amount of storage required.

[0670] Furthermore, the encoding device can calculate the temporal dv of the DVG based on the displacement vectors of the 3D points belonging to the DVG. For example, the average value of the displacement vectors of the 3D points belonging to the DVG can also be used as the temporal dv of the DVG. Thus, while reducing the storage required to maintain the temporal dv, it is possible to improve encoding efficiency by appending the temporal dv to the prediction candidates.

[0671] Alternatively, for example, the encoding device can calculate the global displacement vector (hereinafter, globaldv) of the target frame and append the globaldv to the prediction candidate as a prediction value. The encoding device can, for example, calculate the globaldv based on the average displacement vector within the target frame or a reference frame. Furthermore, the encoding device can append the calculated globaldv to the bitstream. Thus, the decoding device can decode the globaldv appended by the encoding device as a prediction candidate from the bitstream, and can append the same globaldv as the encoding device to the prediction candidate.

[0672] Furthermore, for example, the encoding device can select at least two displacement vectors from the displacement vectors appended to the prediction candidates as new prediction values, and append the average of the selected two or more displacement vectors as a new prediction value to the prediction candidates. This could potentially improve encoding efficiency.

[0673] Furthermore, for example, the encoding device can store one or more previously used displacement vectors as new prediction values ​​in the memory, and append at least one of these displacement vectors as new prediction values ​​to the prediction candidates. This could potentially improve encoding efficiency. Alternatively, the encoding device can periodically or irregularly save displacement vectors used in encoding or decoding in the memory (i.e., the memory described above that stores one or more previously used displacement vectors), and delete old displacement vectors that have been stored for a certain period of time from the time they were saved. In this way, by updating the displacement vectors stored in the memory, the encoding device can assign new displacement vectors to the prediction candidates, potentially improving encoding efficiency.

[0674] When encoding the displacement vectors of 3D points using an encoding device, a DVG (Digital Void Generation) can be set as the prediction unit according to the encoding or decoding order, and encoding or decoding can be performed on a per-DVG basis. For example, consider defining the number of 3D points contained in a DVG (DVGSize), and dividing the 3D points into multiple DVGs for encoding or decoding according to the encoding or decoding order. Furthermore, the encoding or decoding order of the 3D point displacement vectors can be arbitrary. For example, a LoD (Location of Detail) can be generated and encoded or decoded sequentially for each LoD level. Alternatively, a LoD can be omitted, and the displacement vectors can be encoded or decoded according to the encoding or decoding order of the 3D point position information (vertex). Additionally, Morton code can be generated using the 3D point position information, and encoding or decoding can be performed according to the Morton code order.

[0675] The following is an explanation of the definition of DVG.

[0676] Figure 84 This is an explanatory diagram showing an example of a reference target for DVG in this embodiment.

[0677] exist Figure 84 In the example of reference targets in the DVG shown (also known as the first example), 3D points within the same DVG are defined as non-referenceable. For example, 3D points within the same DVG could be set to not be added as adjacent points.

[0678] Additionally, 3D points within different encoded or decoded DVGs are defined as referential. For example, 3D points within different encoded or decoded DVGs can be set to not be appended as adjacent points.

[0679] Additionally, the size of the DVG can be recorded in the header, etc. (see reference) Figure 85 For example, with a DVG size (DVGSize) of 16, information such as "DVGSize = 16" can be appended to the header. Alternatively, DVGSize can be set to 2^n, and the value of n can be appended to the header.

[0680] In addition, it can encode or decode 3D points within the same DVG in parallel.

[0681] Figure 85 This is an explanatory diagram illustrating an example of the syntax of this implementation method.

[0682] Figure 85 The example syntax shown illustrates an example of the structure of information contained in the bitstream generated by the encoding device.

[0683] Figure 85The syntax shown includes a displacement vector_header. The displacement vector_header contains DVGSize.

[0684] DVGSize represents the number of three-dimensional points contained in the DVG.

[0685] Figure 86 This is an explanatory diagram showing an example of a reference target for DVG in this embodiment.

[0686] exist Figure 86 In the example of reference targets in the DVG shown (also known as the second example), encoded or decoded 3D points within the same DVG are defined as referable. Uncoded or undecoded 3D points are defined as non-referable. For example, encoded or decoded 3D points within the same DVG can be set as addable as adjacent points, while uncoded / decoded 3D points can be set as non-addable.

[0687] Furthermore, 3D points within different encoded or decoded DVGs are defined as referential. For example, 3D points within different encoded or decoded DVGs can be set to be appended as adjacent points.

[0688] Additionally, the size of the DVG can also be recorded in the header, etc. (see reference) Figure 85 For example, if the DVG size (DVGSize) is 16, information such as "DVGSize = 16" can be appended to the header. Alternatively, DVGSize can be set to 2^n, and the value of n can be appended to the header.

[0689] Therefore, even for 3D points within the same DVG, encoded / decoded 3D points can be set as references to improve prediction accuracy and coding efficiency.

[0690] Figure 87 This is an explanatory diagram showing an example of a reference target for DVG in this embodiment.

[0691] exist Figure 87 In the example of reference targets in the DVG shown (also known as the third example), encoded or decoded 3D points within the same DVG are defined as referable. Uncoded or undecoded 3D points are defined as non-referable. For example, encoded or decoded 3D points within the same DVG can be set as addable as adjacent points, while uncoded or undecoded 3D points can be set as non-addable as adjacent points.

[0692] In addition, 3D points within different DVGs are defined as non-referenceable. For example, 3D points within different DVGs can be set to not be added as adjacent points.

[0693] Additionally, the size of the DVG can also be recorded in the header, etc. (see reference) Figure 85 For example, with a DVG size (DVGSize) of 16, information such as "DVGSize = 16" can be appended to the header. Alternatively, DVGSize can be set to 2^n, and the value of n can be appended to the header.

[0694] In this way, by prohibiting references between DVGs and eliminating dependencies between DVGs, multiple DVGs can be encoded or decoded in parallel.

[0695] Furthermore, by enabling reference to encoded or decoded 3D points within the same DVG, prediction accuracy can be improved, potentially leading to increased coding efficiency.

[0696] In reference Figure 84 The description illustrates the following example: when encoding the displacement vectors of 3D points, the DVG is set according to the encoding or decoding order, and encoding or decoding is performed on a per-DVG basis. For example, the following example is shown: defining the number of 3D points contained in a DVG (DVGSize), and dividing the 3D points into multiple DVGs for encoding or decoding according to the encoding or decoding order. Here, the prediction mode PredMode used for encoding the displacement vectors, or the information disp_dvg_inter_mode regarding whether to apply inter-frame prediction, can be set per DVG. In this case, the 3D points contained in the same DVG share PredMode or disp_dvg_inter_mode, and the same value can be set. Therefore, by reducing the amount of code in PredMode or disp_dvg_inter_mode, encoding efficiency can be improved. Alternatively, it is not limited to each DVG; PredMode or disp_dvg_inter_mode can also be set according to other sets of 3D points.

[0697] Figure 88 This is an explanatory diagram showing an example of a reference target for DVG in this embodiment.

[0698] exist Figure 88 In the example of reference targets in the DVG shown (also known as the fourth example), encoded or decoded 3D points within the same DVG are defined as referable. Uncoded or undecoded 3D points are defined as non-referable. For example, encoded or decoded 3D points within the same DVG can be set as addable as adjacent points, while uncoded or undecoded 3D points can be set as non-addable.

[0699] Furthermore, 3D points within different encoded or decoded DVGs are defined as referential. For example, 3D points within different encoded or decoded DVGs can be set to be appended as adjacent points.

[0700] In addition, the encoding device can attach a PredMode or disp_dvg_inter_mode to each DVG and use the same PredMode or dips_dvg_inter_mode to predictively encode the 3D points within the DVG.

[0701] Furthermore, the encoding device can determine whether to attach PredMode or disp_dvg_inter_mode to each DVG. For example, the encoding device can use the variance of the displacement vectors of decoded 3D points within different DVGs to calculate the PredMode or disp_dvg_inter_mode of the DVG to which the 3D points of the encoded object belong. Furthermore, if the calculated variance value is above a threshold, PredMode or disp_dvg_inter_mode can be attached to the DVG; otherwise, it can be left unattached. If PredMode or disp_dvg_inter_mode is not attached, it can be assumed that PredMode = 0 or disp_dvg_inter_mode = 0.

[0702] Additionally, the size of the DVG can also be recorded in the header, etc. (see reference) Figure 85 For example, with a DVG size (DVGSize) of 16, information such as "DVGSize = 16" can be appended to the header. Alternatively, DVGSize can be set to 2^n, and the value of n can be appended to the header.

[0703] In this way, even 3D points within the same DVG can be configured to reference already encoded or decoded 3D points, thereby improving prediction accuracy and potentially increasing coding efficiency.

[0704] Furthermore, by attaching PredMode or dips_dvg_inter_mode to each DVG, compared to attaching PredMode or dips_dvg_inter_mode to each 3D point, overhead can be reduced, potentially improving coding efficiency.

[0705] Figure 89 This is an explanatory diagram illustrating an example of the syntax of this implementation method.

[0706] Figure 89The example syntax shown illustrates an example of the structure of information contained in the bitstream generated by the encoding device.

[0707] Figure 89 The syntax shown includes a displacement vector_header. The displacement vector_header contains DVGSize.

[0708] DVGSize represents the unit of the displacement vector of the predicted 3D point. For each DVGSize 3D point, a PredMode or disp_dvg_inter_mode is appended. The same PredMode or disp_dvg_inter_mode is used to encode and decode the 3D points within the same DVG.

[0709] Alternatively, when encoding the displacement vector into LoD levels, a different DVGSize can be set for each LoD level. In this case, the DVGSize for each LoD level can be appended to the header. Thus, the decoding device can correctly decode the bitstream generated by setting the DVGSize for each LoD level.

[0710] For example, when encoding displacement vectors using LoD levels, increasing the transform coefficients of the lower-level LoD set tends to decrease their values. Therefore, the lower the LoD level, the easier it is to improve inter-frame prediction accuracy. Thus, by increasing the DVGSize value in the lower-level LoD levels, and sharing dips_dvg_inter_mode across multiple 3D points, the amount of code encoded for disp_dvg_inter_mode can be reduced, potentially improving coding efficiency. On the other hand, in the aforementioned LoD levels, increasing the transform coefficients of the low-frequency components of the transform set tends to increase their values. Therefore, the lower the LoD level, the easier it is to decrease inter-frame prediction accuracy. Thus, by decreasing the DVGSize value in the lower-level LoD levels, it is possible to finely configure whether to perform inter-frame prediction, potentially improving coding efficiency.

[0711] Figure 90 This is an explanatory diagram illustrating an example of the syntax of this implementation method.

[0712] Figure 90 The example syntax shown represents an example of the structure of information contained in the bitstream generated by the encoding device.

[0713] Figure 90The syntax shown includes displacement_vector_data. Displacement_vector_data may contain PredMode, disp_dvg_inter_mode, dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k] related to each of the 0th to NumLoD levels of LoD (also known as the jth level).

[0714] PredMode represents information about the prediction mode used to encode or decode the displacement vector of the i-th 3D point at the j-th level. PredMode takes values ​​from 0 to M-1 (where M is the total number of prediction modes). If PredMode is not included in the bitstream (i.e., the condition of the if statement "maxdiff >= Thfix[i] && NumPredMode[i] > 1" is not met), the value of PredMode can be presumed to be 0. Alternatively, it can be any value from 0 to M-1, not limited to 0. Furthermore, the presumed value for PredMode when it is not included in the bitstream can be appended to the header, etc. Additionally, PredMode can be binarized using truncated unary code and arithmetic encoded with the number of prediction modes assigned prediction values.

[0715] `disp_dvg_inter_mode` indicates whether inter-frame prediction is used to encode or decode the i-th displacement vector of the j-th LoD level. A value of 1 indicates that inter-frame prediction is applied, and a value of 0 indicates that inter-frame prediction is not applied.

[0716] dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], dispd_sign[k] and Figure 67 The information in each part is the same.

[0717] The following describes an example of the encoding process in this embodiment.

[0718] Figure 91 This is a flowchart illustrating the encoding process in this embodiment.

[0719] Figure 91 The encoding process shown is a method for encoding the displacement data of three-dimensional points performed by the encoding device.

[0720] In step S9101, the encoding device generates a predicted value for the first displacement data by performing inter-frame prediction using second displacement data at a different time than the first displacement data, which is the displacement data to be encoded.

[0721] In step S9102, the encoding device uses the first displacement data and the generated prediction value to generate a prediction residual.

[0722] In step S9103, the encoding device encodes the generated prediction residuals.

[0723] Therefore, the encoding device encodes the prediction residuals generated through inter-frame prediction, thereby enabling appropriate encoding of the displacement data of three-dimensional points. By using inter-frame prediction, it is possible to reduce the amount of code when the difference between the first displacement data to be encoded and the second displacement data at different times is relatively small, thus potentially improving the encoding process. In this way, the above encoding method can help improve encoding processes related to displacement vectors.

[0724] For example, when generating the predicted value of the first displacement data, it can be determined whether to perform inter-frame prediction. If it is determined that inter-frame prediction should be performed, the predicted value of the first displacement data is generated by performing inter-frame prediction. If it is determined that inter-frame prediction should not be performed, the predicted value of the first displacement data is generated without performing inter-frame prediction.

[0725] Therefore, when encoding the first displacement data to be encoded, the encoding device can pre-determine whether to use inter-frame prediction in the generation of the predicted values ​​of the displacement data, and switch between using inter-frame prediction in the generation of the predicted values ​​of the displacement data based on the determination result. Thus, for example, inter-frame prediction can be used in the encoding process if it can improve the encoding process, and it can be omitted in the encoding process if it cannot improve (or degrades) the encoding process. This makes it possible to further reduce the amount of code. In this way, the above encoding method can help to further improve encoding processes related to displacement vectors, etc.

[0726] For example, when determining whether to perform inter-frame prediction, the sum of the transformation coefficients of the first displacement data (i.e., the first sum) and the sum of the prediction residuals when inter-frame prediction is applied to the transformation coefficients (i.e., the second sum) are calculated. If the second sum is less than the first sum, inter-frame prediction is performed. If the second sum is not less than the first sum, inter-frame prediction is not performed.

[0727] Therefore, the encoding device can compare the sum of the transform coefficients when inter-frame prediction is applied to the transform coefficients of the first displacement data to be encoded with the sum of the transform coefficients (in other words, the sum of the transform coefficients when inter-frame prediction is not used), thereby switching whether to use inter-frame prediction in the generation of the predicted value of the displacement data. Specifically, when it is determined that the sum of the transform coefficients when inter-frame prediction is applied is small, it can be determined that inter-frame prediction is used. Therefore, by making the determination easier, it is possible to reduce the amount of code. In this way, the above encoding method can help improve the encoding processing related to displacement vectors, etc.

[0728] For example, information indicating whether inter-frame prediction was used in the generation of the prediction residual of the first displacement data can also be sent.

[0729] Therefore, by sending information indicating whether inter-frame prediction was used during encoding, the encoding device can communicate this information to the decoding device, which receives and decodes the encoded displacement data. This facilitates proper decoding of the encoded data. Specifically, it enables decoding of data encoded using inter-frame prediction and decoding of data encoded without inter-frame prediction. Thus, the above encoding method helps improve encoding processes related to displacement vectors.

[0730] For example, the three-dimensional points may include multiple three-dimensional points, each belonging to one of multiple layers. When generating the predicted value of the first displacement data, for each of the multiple three-dimensional points belonging to one or more layers, it is determined whether to perform inter-frame prediction when generating the predicted value of the first displacement data for the three-dimensional points belonging to that layer. For the layers determined to perform inter-frame prediction, the predicted value of the first displacement data is generated by performing inter-frame prediction. For the layers determined not to perform inter-frame prediction, the predicted value of the first displacement data is generated without performing inter-frame prediction.

[0731] Therefore, when encoding the first displacement data to be encoded, the encoding device can pre-determine whether to use inter-frame prediction in the generation of the predicted value of the displacement data for each level to which the three-dimensional point belongs. Based on the determination result, it switches whether to use inter-frame prediction in the generation of the predicted value of the displacement data for each level to which the three-dimensional point belongs. Thus, for example, it is possible to use inter-frame prediction in the encoding process of levels where using inter-frame prediction can improve the encoding process, and not to use inter-frame prediction in the encoding process of levels where using inter-frame prediction cannot improve (or degrades) the encoding process. As a result, it is possible to further reduce the amount of code. In this way, the above encoding method can help to further improve the encoding process related to displacement vectors, etc.

[0732] For example, information indicating the level of inter-frame prediction used in the generation of the prediction residual of the first displacement data can also be sent.

[0733] Therefore, by sending information indicating whether inter-frame prediction was used at each level corresponding to the 3D point during encoding, the encoding device can communicate to the decoding device that receives and decodes the encoded displacement data whether inter-frame prediction was used at each level corresponding to the 3D point during encoding. This facilitates proper decoding of the encoded data. Specifically, it facilitates decoding data of levels encoded using inter-frame prediction and decoding data of levels encoded without inter-frame prediction. Thus, the above encoding method helps improve encoding processing related to displacement vectors.

[0734] The following describes an example of the decoding process in this embodiment.

[0735] Figure 92 This is a flowchart illustrating the decoding process in this embodiment.

[0736] Figure 92 The decoding process shown is a method for decoding the displacement data of three-dimensional points performed by the decoding device.

[0737] In step S9201, the decoding device obtains the prediction residual by decoding the encoded data.

[0738] In step S9202, the decoding device generates a predicted value for the first displacement data by performing inter-frame prediction using second displacement data at a different time than the first displacement data, which is the displacement data to be decoded.

[0739] In step S9203, the decoding device uses the prediction residual and the generated prediction value to generate the first displacement data.

[0740] Therefore, the decoding device decodes the prediction residuals generated through inter-frame prediction, thereby enabling appropriate decoding of the displacement data of three-dimensional points. By using inter-frame prediction, when the difference between the first displacement data to be decoded and the second displacement data at different times is relatively small, it is possible to reduce the amount of code, thereby potentially improving the decoding process. Thus, the above decoding method can help improve decoding processes related to displacement vectors.

[0741] For example, when generating the predicted value of the first displacement data, it can be determined whether to perform inter-frame prediction. If it is determined that inter-frame prediction is to be performed, the predicted value of the first displacement data is generated by performing inter-frame prediction. If it is determined that inter-frame prediction is not to be performed, the predicted value of the first displacement data is generated without performing inter-frame prediction.

[0742] Therefore, when decoding the first displacement data, which is the object of decoding, the decoding device can pre-determine whether to use inter-frame prediction in the generation of the predicted value of the displacement data, and switch whether to use inter-frame prediction in the generation of the predicted value of the displacement data based on the determination result. This may further reduce the amount of code. Thus, the above decoding method can help to further improve decoding processing related to displacement vectors.

[0743] For example, information indicating whether inter-frame prediction was used in the generation of the prediction residual of the first displacement data can be further received.

[0744] Therefore, by receiving information indicating whether inter-frame prediction was used during encoding, the decoding device can determine whether the encoding device used inter-frame prediction during encoding. Furthermore, the decoding device can decode the displacement vector using inter-frame prediction if it did, and can decode the displacement vector without using inter-frame prediction if it did not. Thus, the decoding device can appropriately decode the data encoded by the encoding device.

[0745] For example, the three-dimensional points may include multiple three-dimensional points, each belonging to one of multiple layers. When generating the predicted value of the first displacement data, for each of the multiple three-dimensional points belonging to one or more layers, it is determined whether to perform inter-frame prediction when generating the predicted value of the first displacement data for the three-dimensional points belonging to that layer. For the layers that are determined to perform inter-frame prediction, the predicted value of the first displacement data is generated by performing inter-frame prediction. For the layers that are determined not to perform inter-frame prediction, the predicted value of the first displacement data is generated without performing inter-frame prediction.

[0746] Therefore, when decoding the first displacement data, which is the object of encoding, the decoding device can pre-determine whether to use inter-frame prediction in the generation of the predicted value of the displacement data for each level to which the three-dimensional point belongs. Based on the determination result, it switches whether to use inter-frame prediction in the generation of the predicted value of the displacement data for each level to which the three-dimensional point belongs. This may further reduce the amount of code. Thus, the above decoding method can help to further improve the encoding processing related to displacement vectors.

[0747] For example, information indicating the level of inter-frame prediction used in the generation of the prediction residual of the first displacement data can be further received.

[0748] Therefore, by receiving information indicating whether inter-frame prediction was used at each level corresponding to a 3D point during encoding, the decoding device can determine whether the encoding device used inter-frame prediction at each level. Furthermore, for data from levels where inter-frame prediction was used during encoding, the decoding device can decode the displacement vector using inter-frame prediction; for data from levels where inter-frame prediction was not used, the decoding device can decode the displacement vector without using inter-frame prediction. Thus, the decoding device can appropriately decode the data encoded by the encoding device.

[0749] <Other examples> The above describes the encoding and decoding devices according to the embodiments, but the methods of encoding and decoding devices are not limited to the embodiments described above. Modifications that can be conceived by those skilled in the art can be implemented in the embodiments, and the multiple constituent elements in the embodiments can be arbitrarily combined.

[0750] For example, the processes performed by specific components in the implementation can be performed by other components instead of specific components. Furthermore, the order of multiple processes can be changed, or multiple processes can be executed in parallel.

[0751] Furthermore, as described above, at least a portion of the various structures of this disclosure can be implemented as an integrated circuit. At least a portion of the various processes of this disclosure can be utilized as an encoding method or a decoding method. A program for causing a computer to execute the encoding method or the decoding method can also be used. Additionally, a non-transitory computer-readable recording medium on which the program is recorded can also be used. Furthermore, a bitstream for causing a decoding device to perform decoding processing can also be used.

[0752] Furthermore, at least a portion of the various structures and processes disclosed herein can also be used as a transmitting device, a receiving device, a transmitting method, and a receiving method. A program for causing a computer to execute the transmitting method or the receiving method can also be used. Additionally, a non-transitory computer-readable recording medium containing the program can also be used.

[0753] Industrial utilization potential This disclosure is useful, for example, for encoding devices, decoding devices, transmitting devices, and receiving devices related to three-dimensional meshes, and can be applied to computer graphics systems and three-dimensional data display systems.

[0754] Explanation of reference numerals in the attached figures 100 encoding device 101, 121, 144 vertex information encoder 102, 145 connect information encoder 103, 122 Attribute Information Encoder 104, 204, 1103 preprocessors 105, 205, 2106 post-processors 110 Three-Dimensional Data Encoding System 111, 211 controllers 112, 212 Input / Output Processors 113 3D Data Encoder 114 System Multiplexer 115 3D Data Generator 123 metadata encoder 124 multiplexer 131 Vertex Image Generator 132-attribute image generator 133 Metadata Generator 134 Image Encoder 141 Two-Dimensional Data Encoder 142 Grid Data Encoder 143 Texture Encoder 148 describes the encoder 151, 251 circuits 152, 252 memory 200 decoding device 201, 221, 244 vertex information decoder 202, 245 Connection Information Decoder 203, 222 Attribute Information Decoder 210 3D Data Decoding System 213 3D Data Decoder 214 System Demultiplexer 215, 247 prompts 216 User Interface 223 metadata decoder 224 demultiplexer 231 Vertex Information Generator 232 Attribute Information Generator 234 Image Decoder 241 Two-Dimensional Data Decoder 242 Grid Data Decoder 243 Texture Decoder 246 Mesh Reconstructor 248 Description Decoder 300 Network 310 External Connector 613 Basic Mesh Decoder 614 displacement decoder 615 Attribute Decoder 616 Other Types of Decoders 617 3D Reconstructor 631 frame header decoder 632 Vertex Geometry Coordinate Predictor 633 Vertex Geometric Coordinate Differential Decoder 634 Reconfigurator 1231 Demultiplexer 1232 switch 1233 Static Mesh Decoder 1234 Grid Buffer 1235 Motion Decoder 1236 Basic Mesh Reconstructor 1237 Inverse Quantizer 1238, 1243 video decoders 1239 Image Unpacker 1240 Inverse Quantizer 1241 Inverse Wavelet Transform 1242 Reconstructor 1244 Color Converter The base grid after 1251 decoding 1252 sub-segmenter Mesh after 1253 subdivisions Displacement data after 1254 decoding 1255 displacement device 3D mesh after 1256 decoding 4801, 5701 extractors 4802 and 5702 sub-dividers 4803, 5703 Displacement Vector Calculator 4804 and 5704 wavelet transforms 4805 inter-frame predictor 4806 and 5706 quantizers 4807, 5708 image packers 4808 and 5709 video encoders 4811, 5003 Inverse Quantizer 4812, 5004, 5006, 5806, 5808 Reconfigurators 4813, 5811 reference buffers 5001 and 5801 video decoders 5002, 5802 Image Unpacker 5005 and 5807 inverse wavelet transform 5705 LoD-based inter-frame predictor 5707 and 5804 switches 5710 Arithmetic Encoder 5803 Arithmetic Decoder 5805 Inverse Quantizer

Claims

1. An encoding method for encoding displacement data of three-dimensional points, characterized in that, Inter-frame prediction using the second displacement data is performed to generate a predicted value for the first displacement data, where the first displacement data is the displacement data used for encoding, and the second displacement data is data from a different time point than the first displacement data. The prediction residual is generated using the first displacement data and the generated prediction value. The generated prediction residuals are encoded.

2. The encoding method according to claim 1, characterized in that, When generating the predicted value of the first displacement data Determine whether to perform the inter-frame prediction when generating the predicted value of the first displacement data. If it is determined that inter-frame prediction should be performed, the predicted value of the first displacement data is generated by performing the inter-frame prediction. If it is determined that inter-frame prediction should not be performed, then inter-frame prediction is not performed, and the predicted value of the first displacement data is generated.

3. The encoding method according to claim 2, characterized in that, When determining whether to perform the inter-frame prediction Calculate a first sum and a second sum. The first sum is the sum of the transform coefficients for the first displacement data, and the second sum is the sum of the prediction residuals when the inter-frame prediction is applied to the transform coefficients. If the second sum is determined to be less than the first sum, then the inter-frame prediction is performed. If it is determined that the second sum is not less than the first sum, then it is determined that the inter-frame prediction will not be performed.

4. The encoding method according to claim 1, characterized in that, It also sends information indicating whether the inter-frame prediction was used in the generation of the prediction residual of the first displacement data.

5. The encoding method according to claim 1, characterized in that, The three-dimensional points include multiple three-dimensional points. The multiple three-dimensional points each belong to one of multiple levels. When generating the predicted value of the first displacement data For each of the plurality of 3D points belonging to one or more layers, it is determined whether to perform the inter-frame prediction when generating the predicted value of the first displacement data for the 3D points belonging to that layer among the plurality of 3D points. For the level among the one or more levels that is determined to perform the inter-frame prediction, the predicted value of the first displacement data is generated by performing the inter-frame prediction. For the layers that are determined not to perform inter-frame prediction among the one or more layers, inter-frame prediction is not performed, and the predicted value of the first displacement data is generated.

6. The encoding method according to claim 5, characterized in that, It also sends information indicating that the level of the inter-frame prediction was used in the generation of the prediction residual of the first displacement data.

7. A decoding method, which is a method for decoding displacement data of three-dimensional points, characterized in that, The prediction residual is obtained by decoding the encoded data. A predicted value for the first displacement data is generated by performing inter-frame prediction using the second displacement data. The first displacement data is the displacement data to be decoded, and the second displacement data is data from a different time point than the first displacement data. The first displacement data is generated using the predicted residual and the generated predicted value.

8. The decoding method according to claim 7, characterized in that, When generating the predicted value of the first displacement data Determine whether to perform the inter-frame prediction when generating the predicted value of the first displacement data. If it is determined that inter-frame prediction should be performed, the predicted value of the first displacement data is generated by performing the inter-frame prediction. If it is determined that inter-frame prediction should not be performed, then inter-frame prediction is not performed, and the predicted value of the first displacement data is generated.

9. The decoding method according to claim 7, characterized in that, It also receives information indicating whether the inter-frame prediction was used in the generation of the prediction residual of the first displacement data.

10. The decoding method according to claim 7, characterized in that, The three-dimensional points include multiple three-dimensional points. The multiple three-dimensional points each belong to one of multiple levels. When generating the predicted value of the first displacement data For each of the plurality of 3D points belonging to one or more layers, it is determined whether to perform the inter-frame prediction when generating the predicted value of the first displacement data for the 3D points belonging to that layer among the plurality of 3D points. For the level among the one or more levels that is determined to perform the inter-frame prediction, the predicted value of the first displacement data is generated by performing the inter-frame prediction. For the layers that are determined not to perform inter-frame prediction among the one or more layers, inter-frame prediction is not performed, and the predicted value of the first displacement data is generated.

11. The decoding method according to claim 10, characterized in that, It also receives information indicating the level of the inter-frame prediction used in the generation of the prediction residual in the first displacement data.

12. An encoding device for encoding displacement data of three-dimensional points, characterized in that, have: Memory; and The circuit is capable of accessing the memory. During operation, the circuit generates a predicted value for the first displacement data by performing inter-frame prediction using the second displacement data. The first displacement data is the displacement data used for encoding, and the second displacement data is data from a different time point than the first displacement data. The circuit uses the first displacement data and the generated predicted value to generate a prediction residual during operation. The circuit encodes the generated prediction residual during operation.

13. A decoding device for decoding displacement data of three-dimensional points, characterized in that, have: Memory; and The circuit is capable of accessing the memory. The circuit obtains the prediction residual by decoding the encoded data during operation. During operation, the circuit generates a predicted value for the first displacement data by performing inter-frame prediction using the second displacement data. The first displacement data is the displacement data to be decoded, and the second displacement data is data from a different time point than the first displacement data. The circuit uses the predicted residual and the generated predicted value to generate the first displacement data during operation.

Citation Information

Patent Citations

  • Progressive three-dimensional mesh information coding / decoding method, and apparatus therefor

    JP2006187015A