Encoding method, decoding method, encoding device, and decoding device
The method improves three-dimensional mesh data encoding and decoding by adapting processes based on layer comparisons and parameter referencing, enhancing efficiency and error prevention in parameter handling.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
- Filing Date
- 2025-10-27
- Publication Date
- 2026-05-07
AI Technical Summary
Existing encoding and decoding methods for three-dimensional mesh data are inefficient and require improvements, particularly in handling parameters used in the encoding processing.
A method for symbolizing parameters in three-dimensional point encoding that involves obtaining parameter sets with different numbers of layers and executing specific processes based on the comparison of these layers, including referencing higher or same hierarchy parameters and calculating difference information.
This approach enhances the encoding and decoding processes by allowing for adaptive processing based on layer comparisons, preventing errors and improving efficiency in handling three-dimensional point parameters.
Smart Images

Figure JP2025037621_07052026_PF_FP_ABST
Abstract
Description
Symbolization method, decoding method, symbolization device, and decoding device
[0001] This disclosure relates to a symbolization method and the like.
[0002] In Patent Document 1, methods and apparatuses for encoding and decoding three-dimensional mesh data have been proposed.
[0003] Japanese Patent Application Laid-Open No. 2006-187015
[0004] Further improvement is desired for encoding or decoding processing related to parameters used in the encoding processing of three-dimensional points. The purpose of this disclosure is to improve the encoding or decoding processing related to parameters used in the encoding processing of three-dimensional points.
[0005] The symbolization method according to one aspect of the present invention is a method for symbolizing parameters used in the encoding processing of three-dimensional points, which includes obtaining a first parameter set having a first number of layers and a second parameter set having a second number of layers, which is referred to in the symbolization of the first parameter set, and when the first number of layers is greater than the second number of layers, executing a first process, and when the first number of layers is not greater than the second number of layers, executing a second process different from the first process.
[0006] These general or specific aspects may be implemented by a system, an apparatus, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, an apparatus, an integrated circuit, a computer program, and a recording medium.
[0007] This disclosure can contribute to the improvement of encoding processing and the like related to parameters used in the encoding processing of three-dimensional points.
[0008] This is a conceptual diagram showing a three-dimensional mesh according to the embodiment. This is a conceptual diagram showing the basic elements of a three-dimensional mesh according to the embodiment. This is a conceptual diagram showing mapping according to the embodiment. This is a block diagram showing an example configuration of an encoding / decoding system according to the embodiment. This is a block diagram showing an example configuration of an encoding device according to the embodiment. This is a block diagram showing another example configuration of an encoding device according to the embodiment. This is a block diagram showing an example configuration of a decoding device according to the embodiment. This is a block diagram showing another example configuration of a decoding device according to the embodiment. This is a conceptual diagram showing an example configuration of a bitstream according to the embodiment. This is a conceptual diagram showing yet another example configuration of a bitstream according to the embodiment. This is a block diagram showing a concrete example of an encoding / decoding system according to the embodiment. This is a conceptual diagram showing an example configuration of point cloud data according to the embodiment. This is a conceptual diagram showing an example data file for point cloud data according to the embodiment. This is a conceptual diagram showing an example configuration of mesh data according to the embodiment. This is a conceptual diagram showing an example data file for mesh data according to the embodiment. This is a conceptual diagram showing the types of three-dimensional data according to the embodiment. This is a block diagram showing an example configuration of a three-dimensional data encoder according to the embodiment. This is a block diagram showing an example configuration of a three-dimensional data decoder according to the embodiment. This is a block diagram showing another example configuration of a three-dimensional data encoder according to the embodiment. This is a block diagram showing another example configuration of a three-dimensional data decoder according to the embodiment. This is a conceptual diagram showing a specific example of the encoding process according to the embodiment. This is a conceptual diagram showing a specific example of the decoding process according to the embodiment. This is a block diagram showing an implementation example of the encoding device according to the embodiment. This is a block diagram showing an implementation example of the decoding device according to the embodiment. This is a block diagram showing another configuration example of the encoding / decoding system according to the embodiment. This is a block diagram showing another configuration example of the encoding device according to the embodiment. This is a block diagram showing another configuration example of the decoding device according to the embodiment. This is a flowchart showing the processing of the encoding device according to the embodiment. This is an explanatory diagram conceptually showing the encoding of a mesh frame according to the embodiment. This is a flowchart showing the processing of the decoding device according to the embodiment.This is an explanatory diagram conceptually showing the decoding of a mesh frame according to the embodiment. This is a block diagram showing an example of the configuration of a decoding device according to the embodiment. This is a block diagram showing an example of the configuration of a decoding device according to the embodiment. This is an explanatory diagram showing an example of subdivision according to the embodiment. This is an explanatory diagram showing an example of vertex displacement after subdivision according to the embodiment. This is an explanatory diagram showing an example of vertices of the original mesh according to the embodiment. This is an explanatory diagram showing an example of a mesh according to the embodiment. This is an explanatory diagram showing an example of subdivision of a mesh according to the embodiment. This is a first explanatory diagram showing an example of packing displacement information into an image frame according to the embodiment. This is a second explanatory diagram showing an example of packing displacement information into an image frame according to the embodiment. This is a third explanatory diagram showing an example of packing displacement information into an image frame according to the embodiment. This is a block diagram showing a detailed example of the configuration of a decoding device according to the embodiment. This is an explanatory diagram showing the coordinates of vertices in a three-dimensional mesh according to the embodiment. This is an explanatory diagram showing prediction information according to the embodiment. This is a block diagram showing an example of the configuration of an encoding device according to the embodiment. This is a flowchart showing a specific example of the encoding process according to the embodiment. This is a block diagram showing an example of the configuration of a decoding device according to the embodiment. This is a flowchart showing a specific example of the decoding process according to the embodiment. This is a flowchart showing a specific example of the encoding process according to the embodiment. This is a flowchart showing a specific example of the decoding process according to the embodiment. This is an explanatory diagram showing an example of syntax according to the embodiment. This is an explanatory diagram showing an example of syntax according to the embodiment. This is an explanatory diagram showing an example of syntax according to the embodiment. This is a block diagram showing an example of the configuration of an encoding device according to the embodiment. This is a block diagram showing an example of the configuration of a decoding device according to the embodiment. This is an explanatory diagram showing the positional relationship of three-dimensional points according to the embodiment. This is an explanatory diagram showing the method of generating LoD in predicted values of displacement vectors in the embodiment. This is an explanatory diagram showing an example of calculating predicted values in the embodiment. This is an explanatory diagram showing an example of calculating conversion coefficients in the embodiment. This is an explanatory diagram showing an example of interpretation of conversion coefficients in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment.This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of calculation of predicted residuals in the embodiment. This is an explanatory diagram showing an example of calculation of predicted residuals in the embodiment. This is an explanatory diagram showing an example of predicted value information for displacement vectors in the embodiment. This is an explanatory diagram showing a method for generating predicted values for displacement vectors in the embodiment. This is an explanatory diagram showing an example of predicted value information for displacement vectors in the embodiment. This is an explanatory diagram showing an example of predicted value information for displacement vectors in the embodiment. This is an explanatory diagram showing an example of predicted value information for displacement vectors in the embodiment. This is an explanatory diagram showing an example of predicted value information for displacement vectors in the embodiment. This is an explanatory diagram showing an example of predicted value information for motion vectors in the embodiment. This is an explanatory diagram showing an example of predicted value information for motion vectors in the embodiment. This is an explanatory diagram showing an example of points to be coded in the embodiment. This is an explanatory diagram showing an example of predicted value information for displacement vectors in the embodiment. This is an explanatory diagram showing an example of a temporal dv in the embodiment. This is an explanatory diagram showing an example of predicted value information for displacement vectors in the embodiment. This is an explanatory diagram showing an example of a DVG reference in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of a DVG reference in the embodiment. This is an explanatory diagram showing an example of a DVG reference in the embodiment. This is an explanatory diagram showing an example of a DVG reference in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is a flowchart showing the encoding process in the embodiment. This is a flowchart showing the decoding process in the embodiment. This is an explanatory diagram showing an example of a method for calculating predicted values of displacement vectors in the embodiment. This is an explanatory diagram showing an example of a displacement vector in the embodiment. This is an explanatory diagram showing an example of a method for calculating conversion coefficients in the embodiment. This is a flowchart showing an example of an encoding process in the embodiment. This is a flowchart showing an example of a decoding process in the embodiment. This is an explanatory diagram showing an example of a definition of a method for predicting displacement vectors in the embodiment.This is an explanatory diagram showing an example of a header in the embodiment. This is an explanatory diagram showing an example of a header in the embodiment. This is a flowchart showing the encoding process in the embodiment. This is a flowchart showing the decoding process in the embodiment. This is an explanatory diagram showing an example of the lifting transformation process in the embodiment. This is an explanatory diagram showing an example of the lifting transformation process in the embodiment. This is an explanatory diagram showing an example of the weight determination process in the embodiment. This is an explanatory diagram showing an example of the lifting transformation process in the embodiment. This is an explanatory diagram showing an example of the input coefficients, prediction coefficients, and prediction residuals before and after modification in the embodiment. This is an explanatory diagram showing an example of the weight determination process in the embodiment. This is an explanatory diagram showing an example of the inverse lifting transformation process in the embodiment. This is an explanatory diagram showing an example of the inverse lifting transformation process in the embodiment. This is a flowchart showing the encoding process in the embodiment. This is a flowchart showing the decoding process in the embodiment. This is an explanatory diagram showing an example of the syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is a flowchart showing the encoding process in the embodiment. This is a flowchart showing the decoding process in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment.This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is a flowchart showing the estimation process of fps_transform_method in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is a flowchart showing the estimation process of mdu_transform_method in the embodiment. This is a flowchart showing the encoding process in the embodiment. This is a flowchart showing the decoding process in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is a flowchart showing an example of syntax in the embodiment. This is a flowchart showing an example of syntax in the embodiment. This is a flowchart showing an example of syntax in the embodiment. This is a flowchart showing an example of syntax in the embodiment. This is a flowchart showing an example of syntax in the embodiment. This is a flowchart showing an example of syntax in the embodiment. This is a flowchart showing an example of syntax in the embodiment. This is a flowchart showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of a header when the header structure of the upper header and the current header are different in the embodiment. This is an explanatory diagram showing an example of a reference in the embodiment. This is an explanatory diagram showing an example of a reference in the embodiment. This is an explanatory diagram showing an example of a reference in the embodiment. This is an explanatory diagram showing an example of a reference in the embodiment. This is an explanatory diagram showing an example of a reference in the embodiment. This is an explanatory diagram showing an example of a reference in the embodiment. This is an explanatory diagram showing an example of a reference in the embodiment. This is an explanatory diagram showing an example of a reference in the embodiment. This is an explanatory diagram showing an example of a reference in the embodiment. This is an explanatory diagram showing an example of a syntax in the embodiment. This is an explanatory diagram showing an example of a syntax in the embodiment. This is an explanatory diagram showing an example of a syntax in the embodiment. This is an explanatory diagram showing an example of a syntax in the embodiment.This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of a header when the header structure of the upper header and the current header are different in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of syntax in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment. This is an explanatory diagram showing an example of reference in the embodiment.This is an explanatory diagram showing an example of syntax in the embodiment. This is a flowchart showing the encoding process in the embodiment. This is a flowchart showing the decoding process in the embodiment.
[0009] <Summary of this disclosure> For example, a three-dimensional (3D) mesh is used in computer graphics images. For example, a computer graphics image may consist of multiple frames that are different in time, and each frame may be represented by a three-dimensional mesh.
[0010] Furthermore, a three-dimensional mesh consists of vertex information indicating the position of each of several vertices in three-dimensional space, connection information indicating the connections between the multiple vertices, and attribute information indicating the attributes of each vertex or face. Each face is constructed according to the connection relationships of the multiple vertices. Various computer graphics images can be represented using such a three-dimensional mesh.
[0011] Furthermore, efficient encoding and decoding of three-dimensional meshes are expected for the transmission and storage of these meshes. Arithmetic coding and decoding may be used for efficient encoding and decoding of three-dimensional meshes.
[0012] Further improvements are desired in the encoding or decoding of three-dimensional data. This disclosure aims to improve the encoding or decoding of three-dimensional data.
[0013] The following describes examples of inventions that can be obtained from the disclosures in this specification, and explains the effects and other benefits that can be obtained from such inventions.
[0014] (1) A method for encoding parameters used in encoding processing of three-dimensional points, comprising: obtaining a first parameter set having a first number of layers and a second parameter set having a second number of layers that is referenced in encoding the first parameter set; executing a first process when the first number of layers is greater than the second number of layers; and executing a second process different from the first process when the first number of layers is not greater than the second number of layers.
[0015] According to the above embodiment, the encoding device that performs the encoding method can perform different processing depending on whether the number of levels in the first parameter set (i.e., the first level) is greater than or less than the number of levels in the second parameter set (i.e., the second level). This allows the encoding device to switch the processing to be performed depending on whether the number of levels in the first parameter set (i.e., the first level) is greater than or less than the number of levels in the second parameter set (i.e., the second level). In this way, the above encoding method can contribute to improving the encoding processing related to parameters used in the encoding processing of three-dimensional points.
[0016] (2) The encoding method according to (1), wherein the first process includes a process of referring to a second parameter of a higher hierarchy than the first hierarchy included in the second parameter set for a first parameter of a hierarchy included in the first parameter set, and the second process includes a process of referring to a second parameter of the same hierarchy as the first parameter of each hierarchy included in the first parameter set for a first parameter of that hierarchy included in the second parameter set.
[0017] According to the above embodiment, the encoding device that performs the encoding method can avoid, in the first processing, referencing a parameter in the second parameter set that is at the same level as the first parameter in the first level when processing a reference for the first parameter of the first level. As a result, the encoding device can prevent the first processing from being properly executed due to the lack of a reference target if, for example, there is no parameter in the second parameter set that is at the same level as the first level. Thus, the above encoding method can contribute to improving the encoding process for parameters used in the encoding process of three-dimensional points.
[0018] (3) The encoding method according to (2), wherein the first process includes a process that references the second parameter of the highest level included in the second parameter set for each first parameter of the level included in the first parameter set.
[0019] According to the above embodiment, the encoding device that performs the encoding method can avoid referencing a parameter in the second parameter set that is at the same level as the first parameter in the first level in the first processing step when processing a reference for the first parameter of the first level. Furthermore, assuming that the second parameter of the highest level in the second parameter set always exists, the encoding device references the second parameter of the highest level included in the second parameter set, thereby more effectively avoiding the first processing not being executed properly due to the absence of a reference target. In this way, the above encoding method can contribute to improving the encoding process for parameters used in the encoding process of three-dimensional points.
[0020] (4) The encoding method according to (2) or (3), wherein the first process further includes a process of calculating difference information showing the difference between the first parameter of one hierarchy and the referenced second parameter, and the second process further includes a process of calculating difference information showing the difference between the first parameter of one hierarchy and the referenced second parameter.
[0021] According to the above embodiment, the encoding device that performs the encoding method can perform subsequent difference calculation processing using a second parameter of a reference that has been appropriately set. Specifically, in the first process, the encoding device can calculate the difference between the first parameter of the first hierarchy and the second parameter of the reference that has been appropriately set. Furthermore, in the second process, the encoding device can calculate the difference between the first parameter of the first hierarchy and the second parameter of the reference that has been appropriately set. In this way, the above encoding method can contribute to improving the encoding process for parameters used in the encoding process of three-dimensional points.
[0022] (5) The encoding method according to (1), wherein the first parameter set is included in the frame parameter set of the frame to which the three-dimensional point belongs, and the second parameter set is included in the sequence parameter set of the sequence to which the three-dimensional point belongs.
[0023] According to the above embodiment, the encoding device that performs the encoding method can perform different processing depending on whether the number of layers in the frame parameter set of the frame to which the three-dimensional point to be encoded belongs is greater than or less than the number of layers in the sequence parameter set of the sequence to which the three-dimensional point belongs, thereby allowing it to switch the processing to be performed. In this way, the above encoding method can contribute to improving the encoding processing related to the parameters used in the encoding processing of three-dimensional points.
[0024] (6) The encoding method according to (1), wherein the first parameter set is included in the mesh patch data unit of the mesh patch to which the three-dimensional point belongs, and the second parameter set is included in the frame parameter set of the frame to which the three-dimensional point belongs.
[0025] According to the above embodiment, the encoding device that performs the encoding method can perform different processing depending on whether the number of layers of the mesh patch data unit of the mesh patch to which the three-dimensional point to be encoded belongs is greater than or less than the number of layers of the frame parameter set of the frame to which the three-dimensional point belongs, thereby allowing the device to switch the processing to be performed. In this way, the above encoding method can contribute to improving the encoding processing related to the parameters used in the encoding processing of three-dimensional points.
[0026] (7) The encoding method according to any one of (1) to (6), wherein the first parameter set includes parameters for quantization processing, and the second parameter set includes parameters for quantization processing.
[0027] According to the above embodiment, the encoding device that performs the encoding method can switch the processing to be performed depending on whether the number of levels in the first parameter set (i.e., the first level) is greater than the number of levels in the second parameter set (i.e., the second level), when the first parameter set and the second parameter set are quantization parameters. The above encoding method can contribute to improving the encoding processing related to parameters used in the encoding processing of three-dimensional points.
[0028] (8) A decoding method for parameters used in decoding a three-dimensional point, comprising: obtaining a first parameter set having a first number of levels and a second parameter set having a second number of levels that is referenced in decoding the first parameter set; executing a first process when the first number of levels is greater than the second number of levels; and executing a second process different from the first process when the first number of levels is not greater than the second number of levels.
[0029] According to the above embodiment, the decoding device that performs the decoding method can perform different processing depending on whether the number of levels in the first parameter set (i.e., the first level) is greater than or less than the number of levels in the second parameter set (i.e., the second level). This allows the decoding device to switch the processing to be performed depending on whether the number of levels in the first parameter set (i.e., the first level) is greater than or less than the number of levels in the second parameter set (i.e., the second level). In this way, the above decoding method can contribute to improving the decoding process for parameters used in the decoding process of three-dimensional points.
[0030] (9) The decoding method according to (8), wherein the first process includes a process of referring to a second parameter of a higher hierarchy than the first hierarchy included in the second parameter set for a first parameter of a hierarchy included in the first parameter set, and the second process includes a process of referring to a second parameter of the same hierarchy as the first parameter of each hierarchy included in the first parameter set for a first parameter of that hierarchy included in the second parameter set.
[0031] According to the above embodiment, the decoding device executing the decoding method can avoid, in the first process, referencing a parameter in the second parameter set that is at the same level as the first parameter in the first level when processing a reference for the first parameter of the first level. This prevents the decoding device from failing to execute the first process properly if, for example, there is no parameter in the second parameter set that is at the same level as the first level, due to the lack of a reference target. Thus, the above decoding method can contribute to improving the decoding process for parameters used in the decoding process of three-dimensional points.
[0032] (10) The decoding method according to (9), wherein the first process includes a process that references the second parameter of the highest level included in the second parameter set for each first parameter of the level included in the first parameter set.
[0033] According to the above embodiment, the decoding device executing the decoding method can avoid referencing a parameter in the second parameter set that is at the same level as the first parameter in the first level in the first processing step when processing a reference for the first parameter of the first level. Furthermore, assuming that the second parameter of the highest level in the second parameter set always exists, the decoding device references the second parameter of the highest level included in the second parameter set, thereby more effectively avoiding the first processing not being executed properly due to the absence of a reference target. In this way, the above decoding method can contribute to improving the decoding process for parameters used in the decoding process of three-dimensional points.
[0034] (11) The decoding method according to (9) or (10), wherein the first process further includes decoding difference information from a bitstream showing the difference between the first parameter of one hierarchy and the referenced second parameter, and the second process further includes decoding difference information from a bitstream showing the difference between the first parameter of one hierarchy and the referenced second parameter, and adding the difference shown in the decoded difference information to the second parameter.
[0035] According to the above embodiment, the decoding device that performs the decoding method can perform subsequent addition processing using a second parameter of a reference that has been appropriately set. Specifically, in the first process, the decoding device can add to the second parameter the difference between the first parameter of the first hierarchy and the second parameter of the reference that has been appropriately set. Furthermore, in the second process, the decoding device can add to the second parameter the difference between the first parameter of the first hierarchy and the second parameter of the reference that has been appropriately set. In this way, the above decoding method can contribute to improving the decoding process for parameters used in the decoding process of three-dimensional points.
[0036] (12) The decoding method according to (8), wherein the first parameter set is included in the frame parameter set of the frame to which the three-dimensional point belongs, and the second parameter set is included in the sequence parameter set of the sequence to which the three-dimensional point belongs.
[0037] According to the above embodiment, the decoding device that performs the decoding method can perform different processing depending on whether the number of levels in the frame parameter set of the frame to which the three-dimensional point to be decoded belongs is greater than or less than the number of levels in the sequence parameter set of the sequence to which the three-dimensional point belongs, thereby allowing it to switch the processing to be performed. In this way, the above decoding method can contribute to improving the decoding process related to the parameters used in the decoding process of three-dimensional points.
[0038] (13) The decoding method according to (8), wherein the first parameter set is included in the mesh patch data unit of the mesh patch to which the three-dimensional point belongs, and the second parameter set is included in the frame parameter set of the frame to which the three-dimensional point belongs.
[0039] According to the above aspect, a decoding apparatus that executes a decoding method can execute different processes depending on whether the number of hierarchical levels of a mesh patch data unit of a mesh patch to which a three-dimensional point to be decoded belongs is greater than or not greater than the number of hierarchical levels of a frame parameter set of a frame to which the three-dimensional point belongs. Thereby, the process to be executed can be switched. Thus, the above decoding method can contribute to the improvement of the decoding process regarding parameters used for the decoding process of three-dimensional points.
[0040] (14) The decoding method according to any one of (8) to (13), wherein the first parameter set includes parameters for quantization processing, and the second parameter set includes parameters for quantization processing.
[0041] According to the above aspect, a decoding apparatus that executes a decoding method can switch the process to be executed according to whether the number of hierarchical levels of the first parameter set (i.e., the first number of hierarchical levels) is greater than the number of hierarchical levels of the second parameter set (i.e., the second number of hierarchical levels) when the first parameter set and the second parameter set are quantization parameters. The above decoding method can contribute to the improvement of the decoding process regarding parameters used for the decoding process of three-dimensional points.
[0042] (15) An encoding apparatus comprising a memory and a circuit accessible to the memory, wherein the circuit executes, in operation, an encoding method for parameters used for the encoding process of three-dimensional points. In the encoding method, a first parameter set having a first number of hierarchical levels and a second parameter set having a second number of hierarchical levels, which is referred to in the encoding of the first parameter set, are obtained. When the first number of hierarchical levels is greater than the second number of hierarchical levels, a first process is executed. When the first number of hierarchical levels is not greater than the second number of hierarchical levels, a second process different from the first process is executed.
[0043] According to the above aspect, the encoding apparatus has the same effect as the above encoding method.
[0044] A decoding device includes a memory and a circuit accessible to the memory. In operation, the circuit executes a decoding method for parameters used in decoding three-dimensional points. In the decoding method, a first parameter set having a first number of layers and a second parameter set having a second number of layers, which is referred to in decoding the first parameter set, are obtained. When the first number of layers is greater than the second number of layers, a first process is executed. When the first number of layers is not greater than the second number of layers, a second process different from the first process is executed.
[0045] According to the above aspect, the decoding device has the same effect as the above decoding method.
[0046] These general or specific aspects may be implemented in a system, an apparatus, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented in any combination of a system, an apparatus, an integrated circuit, a computer program, or a recording medium.
[0047] Hereinafter, embodiments will be specifically described with reference to the drawings.
[0048] Note that each of the embodiments described below shows general or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present invention. In addition, among the components in the following embodiments, components not described in the independent claims indicating the most general concept are described as optional components.
[0049] (Embodiment) In this embodiment, an encoding method, a decoding method, and the like will be described.
[0050] <Expression and Terms> Here, the following expressions and terms are used.
[0051] (1) Three-dimensional mesh A three-dimensional mesh is a collection of multiple faces, for example, representing a three-dimensional object. A three-dimensional mesh mainly consists of vertex information, connection information, and attribute information. A three-dimensional mesh may be expressed as a polygon mesh or a mesh. A three-dimensional mesh may also have temporal changes. A three-dimensional mesh may contain metadata related to vertex information, connection information, and attribute information, as well as other additional information.
[0052] (2) Vertex Information Vertex information is information that indicates a vertex. For example, vertex information indicates the position of a vertex in three-dimensional space. Also, a vertex corresponds to the vertices of the faces that make up a three-dimensional mesh. Vertex information is sometimes expressed as "Geometry". Vertex information is also sometimes expressed as position information.
[0053] (3) Connection Information Connection information is information that indicates the connections between vertices. For example, connection information indicates the connections that make up the faces or edges of a three-dimensional mesh. Connection information is sometimes expressed as "Connectivity". Connection information is also sometimes expressed as face information.
[0054] (4) Attribute Information Attribute information is information that indicates the attributes of a vertex or face. For example, attribute information indicates attributes such as color, image, and normal vector associated with a vertex or face. Attribute information is sometimes expressed as "Texture".
[0055] (5) A face is an element that makes up a three-dimensional mesh. Specifically, a face is a polygon on a plane in three-dimensional space. For example, a face can be defined as a triangle in three-dimensional space.
[0056] (6) A plane is a two-dimensional plane in three-dimensional space. For example, a polygon is formed on a plane, and multiple polygons are formed on multiple planes.
[0057] (7) Bitstream A bitstream corresponds to encoded information. A bitstream can also be expressed as a stream, encoded bitstream, compressed bitstream, or encoded signal.
[0058] (8) The expressions "encode" and "decode" may be replaced with expressions such as "store," "include," "write," "describe," "signalize," "transmit," "notify," "save," or "compress," and these expressions may be interchangeable. For example, encoding information may mean including information in a bitstream. Also, encoding information into a bitstream may mean encoding information and generating a bitstream that contains the encoded information.
[0059] Furthermore, the expression "decode" may be replaced with expressions such as "read out," "decipher," "read," "load," "derive," "obtain," "receive," "extract," "restore," "reconstruct," "decompress," or "expand," and these expressions may be interchangeable. For example, decoding information may mean obtaining information from a bitstream. Also, decoding information from a bitstream may mean decoding the bitstream and obtaining the information contained in the bitstream.
[0060] (9) In the explanation of ordinal numbers, first and second ordinal numbers may be assigned to components, etc. These ordinal numbers may be rearranged as appropriate. Ordinal numbers may also be newly assigned to components, etc., or removed. Furthermore, these ordinal numbers may be assigned to elements in order to identify them, and may not correspond to a meaningful order.
[0061] <Three-Dimensional Mesh> Figure 1 is a conceptual diagram showing a three-dimensional mesh according to this embodiment. The three-dimensional mesh is composed of multiple faces. For example, each face is a triangle. The vertices of these triangles are defined in three-dimensional space. The three-dimensional mesh represents a three-dimensional object. Each face may have a color or image.
[0062] Figure 2 is a conceptual diagram showing the basic elements of a three-dimensional mesh according to this embodiment. The three-dimensional mesh consists of vertex information, connection information, and attribute information. Vertex information indicates the position of the vertices of a face in three-dimensional space. Connection information indicates the connections between vertices. A face can be identified by the vertex information and connection information. In other words, a colorless three-dimensional object is formed in three-dimensional space by the vertex information and connection information.
[0063] Attribute information may be associated with vertices or faces. Attribute information associated with vertices may be expressed as "Attribute Per Point". Attribute information associated with vertices may indicate the attributes of the vertex itself or the attributes of the faces connected to the vertex.
[0064] For example, a color may be associated with a vertex as attribute information. The color associated with a vertex may be the color of the vertex itself, or the color of the face connected to the vertex. The color of a face may be the average of multiple colors associated with multiple vertices of the face. In addition, a normal vector may be associated with a vertex or face as attribute information. Such a normal vector can represent the front and back of a face.
[0065] Furthermore, a two-dimensional image may be associated with a surface as attribute information. The two-dimensional image associated with a surface is also referred to as a texture image or "Attribute Map". Additionally, information indicating the mapping between the surface and the two-dimensional image may be associated with the surface as attribute information. Such information indicating the mapping may be referred to as mapping information, vertex information of the texture image, texture coordinates, or "Attribute UV Coordinate".
[0066] Furthermore, information such as color, images, and moving images used as attribute information may be expressed as "Parametric Space."
[0067] Such attribute information can be used to reflect textures onto three-dimensional objects. In other words, vertex information, connection information, and attribute information allow a three-dimensional object with color to be formed in three-dimensional space.
[0068] In the above, attribute information is associated with vertices or faces, but it may also be associated with edges.
[0069] Figure 3 is a conceptual diagram illustrating the mapping according to this embodiment. For example, a region of a two-dimensional image in a two-dimensional plane can be mapped to a surface of a three-dimensional mesh in three-dimensional space. Specifically, coordinate information of a region in a two-dimensional image is associated with a surface of a three-dimensional mesh. As a result, the image of the region mapped in the two-dimensional image is reflected on the surface of the three-dimensional mesh.
[0070] By using mapping, a two-dimensional image used as attribute information can be separated from a three-dimensional mesh. For example, in encoding a three-dimensional mesh, the two-dimensional image may be encoded using an image encoding scheme or a video encoding scheme.
[0071] <System Configuration> Figure 4 is a block diagram showing an example configuration of the encoding and decoding system according to this embodiment. In Figure 4, the encoding and decoding system comprises an encoding device 100 and a decoding device 200.
[0072] For example, the encoding device 100 acquires a three-dimensional mesh and encodes the three-dimensional mesh into a bitstream. The encoding device 100 then outputs the bitstream to the network 300. For example, the bitstream includes the encoded three-dimensional mesh and control information for decoding the encoded three-dimensional mesh. By encoding the three-dimensional mesh, the information of the three-dimensional mesh is compressed.
[0073] Network 300 transmits the bitstream from the encoding device 100 to the decoding device 200. Network 300 may be the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. Network 300 is not necessarily limited to bidirectional communication; it may also be a one-way communication network for terrestrial digital broadcasting or satellite broadcasting, etc.
[0074] Furthermore, the network 300 may be replaced by recording media such as DVD (Digital Versatile Disc) or BD (Blu-Ray Disc®).
[0075] The decoding device 200 acquires the bitstream and decodes the three-dimensional mesh from the bitstream. The decoding of the three-dimensional mesh expands the information of the three-dimensional mesh. For example, the decoding device 200 decodes the three-dimensional mesh according to a decoding method that corresponds to the encoding method used by the encoding device 100 to encode the three-dimensional mesh. That is, the encoding device 100 and the decoding device 200 perform encoding and decoding according to their respective encoding and decoding methods.
[0076] The three-dimensional mesh before encoding can also be referred to as the original three-dimensional mesh. Similarly, the three-dimensional mesh after decoding can be referred to as the reconstructed three-dimensional mesh.
[0077] <Encoding Device> Figure 5 is a block diagram showing an example configuration of the encoding device 100 according to this embodiment. For example, the encoding device 100 includes a vertex information encoder 101, a connection information encoder 102, and an attribute information encoder 103.
[0078] The vertex information encoder 101 is an electrical circuit that encodes vertex information. For example, the vertex information encoder 101 encodes vertex information into a bitstream according to a defined format for vertex information.
[0079] The connection information encoder 102 is an electrical circuit that encodes connection information. For example, the connection information encoder 102 encodes connection information into a bitstream according to a specified format for connection information.
[0080] The attribute information encoder 103 is an electrical circuit that encodes attribute information. For example, the attribute information encoder 103 encodes attribute information into a bitstream according to a format defined for the attribute information.
[0081] Variable-length coding or fixed-length coding may be used to encode vertex information, connection information, and attribute information. Variable-length coding may correspond to Huffman coding or context-adaptive binary arithmetic coding (CABAC), etc.
[0082] The vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be integrated into a single unit. Alternatively, each of the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be further subdivided into multiple components.
[0083] Figure 6 is a block diagram showing another configuration example of the encoding device 100 according to this embodiment. For example, the encoding device 100 includes a preprocessor 104 and a postprocessor 105 in addition to the configuration shown in Figure 5.
[0084] The preprocessor 104 is an electrical circuit that performs processing before encoding vertex information, connection information, and attribute information. For example, the preprocessor 104 may perform transformation processing, separation processing, or multiplexing processing on the three-dimensional mesh before encoding. More specifically, for example, the preprocessor 104 may separate vertex information, connection information, and attribute information from the three-dimensional mesh before encoding.
[0085] The post-processor 105 is an electrical circuit that performs processing after encoding the vertex information, connection information, and attribute information. For example, the post-processor 105 may perform conversion processing, separation processing, or multiplexing processing on the encoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 105 may multiplex the encoded vertex information, connection information, and attribute information into a bitstream. Alternatively, for example, the post-processor 105 may further perform variable-length encoding on the encoded vertex information, connection information, and attribute information.
[0086] <Decoding Device> Figure 7 is a block diagram showing an example configuration of the decoding device 200 according to this embodiment. For example, the decoding device 200 includes a vertex information decoder 201, a connection information decoder 202, and an attribute information decoder 203.
[0087] The vertex information decoder 201 is an electrical circuit that decodes vertex information. For example, the vertex information decoder 201 decodes vertex information from a bitstream according to a format defined for vertex information.
[0088] The connection information decoder 202 is an electrical circuit that decodes connection information. For example, the connection information decoder 202 decodes connection information from a bitstream according to a format defined for connection information.
[0089] The attribute information decoder 203 is an electrical circuit that decodes attribute information. For example, the attribute information decoder 203 decodes attribute information from a bitstream according to a format defined for attribute information.
[0090] Variable-length decoding or fixed-length decoding may be used for decoding vertex information, connection information, and attribute information. Variable-length decoding may correspond to Huffman coding or context-adaptive binary arithmetic coding (CABAC), etc.
[0091] The vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be integrated. Alternatively, each of the vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be subdivided into multiple components.
[0092] Figure 8 is a block diagram showing another configuration example of the decoding device 200 according to this embodiment. For example, the decoding device 200 includes a pre-processor 204 and a post-processor 205 in addition to the configuration shown in Figure 7.
[0093] The preprocessor 204 is an electrical circuit that performs processing before decoding the vertex information, connection information, and attribute information. For example, the preprocessor 204 may perform conversion processing, separation processing, or multiplexing processing on the bitstream before decoding the vertex information, connection information, and attribute information.
[0094] More specifically, for example, the preprocessor 204 may separate the bitstream into sub-bitstreams corresponding to vertex information, connection information, and attribute information. Alternatively, for example, the preprocessor 204 may perform variable-length decoding on the bitstream before decoding the vertex information, connection information, and attribute information.
[0095] The post-processor 205 is an electrical circuit that performs processing after decoding the vertex information, connection information, and attribute information. For example, the post-processor 205 may perform conversion processing, separation processing, or multiplexing processing on the decoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 205 may multiplex the decoded vertex information, connection information, and attribute information into a three-dimensional mesh.
[0096] <Bitstream> Vertex information, connection information, and attribute information are encoded and stored in the bitstream. The relationship between this information and the bitstream is shown below.
[0097] Figure 9 is a conceptual diagram showing an example of the bitstream configuration according to this embodiment. In this example, connection information, vertex information, and attribute information are integrated in the bitstream. For example, connection information, vertex information, and attribute information may be included in a single file.
[0098] Furthermore, multiple parts of this information may be stored sequentially, such as the first part of connection information, the first part of vertex information, the first part of attribute information, the second part of connection information, the second part of vertex information, the second part of attribute information, and so on. These multiple parts may correspond to multiple parts that are different in time, multiple parts that are different in space, or multiple faces that are different.
[0099] Furthermore, the storage order of connection information, vertex information, and attribute information is not limited to the example above, and a different storage order may be used.
[0100] Figure 10 is a conceptual diagram showing another example of the bitstream configuration according to this embodiment. In this example, multiple files are included in the bitstream, and connection information, vertex information, and attribute information are stored in different files. Here, a file containing connection information, a file containing vertex information, and a file containing attribute information are shown, but the storage format is not limited to this example. For example, two types of information from connection information, vertex information, and attribute information may be included in one file, and the remaining type of information may be included in another file.
[0101] Alternatively, this information may be divided and stored in more files. For example, multiple parts of connection information may be stored in multiple files, multiple parts of vertex information may be stored in multiple files, and multiple parts of attribute information may be stored in multiple files. These multiple parts may correspond to multiple parts that are different in time, multiple parts that are different in space, or multiple faces that are different.
[0102] Furthermore, the storage order of connection information, vertex information, and attribute information is not limited to the example above, and a different storage order may be used.
[0103] Figure 11 is a conceptual diagram showing another example of the bitstream configuration according to this embodiment. In this example, the bitstream is composed of multiple separable sub-bitstreams, and connection information, vertex information, and attribute information are stored in different sub-bitstreams.
[0104] Here, sub-bitstreams containing connection information, sub-bitstreams containing vertex information, and sub-bitstreams containing attribute information are shown, but the storage format is not limited to these examples.
[0105] For example, two types of information from connection information, vertex information, and attribute information may be included in one sub-bitstream, and the remaining type of information may be included in another sub-bitstream. Specifically, attribute information of a two-dimensional image, etc., may be stored in a sub-bitstream compliant with an image encoding scheme, separate from the sub-bitstreams of connection information and vertex information.
[0106] Furthermore, each sub-bitstream may contain multiple files. Multiple parts of connection information may be stored in multiple files, multiple parts of vertex information may be stored in multiple files, and multiple parts of attribute information may be stored in multiple files.
[0107] Furthermore, the storage order of connection information, vertex information, and attribute information is not limited to the examples in Figures 9, 10, and 11, and a different storage order may be used. For example, they may be stored in the bitstream in the order of vertex information, connection information, and attribute information. Alternatively, they may be stored in the bitstream in any of the following orders: connection information, attribute information, and vertex information; vertex information, attribute information, and connection information; attribute information, connection information, and vertex information; or attribute information, vertex information, and connection information.
[0108] Furthermore, connection information, vertex information, and attribute information may each be divided into multiple data points, and these multiple data points may be stored in a bitstream in a periodic or random order.
[0109] <Specific Example> Figure 12 is a block diagram showing a specific example of the encoding and decoding system according to this embodiment. In Figure 12, the encoding and decoding system comprises a three-dimensional data encoding system 110, a three-dimensional data decoding system 210, and an external connector 310.
[0110] The three-dimensional data encoding system 110 comprises a controller 111, an input / output processor 112, a three-dimensional data encoder 113, a three-dimensional data generator 115, and a system multiplexer 114. The three-dimensional data decoding system 210 comprises a controller 211, an input / output processor 212, a three-dimensional data decoder 213, a system demultiplexer 214, a presenter 215, and a user interface 216.
[0111] In the three-dimensional data encoding system 110, sensor data is input from the sensor terminal to the three-dimensional data generator 115. The three-dimensional data generator 115 generates three-dimensional data, such as point cloud data or mesh data, from the sensor data and inputs it to the three-dimensional data encoder 113.
[0112] For example, the 3D data generator 115 generates vertex information and connection information and attribute information corresponding to the vertex information. The 3D data generator 115 may process the vertex information when generating the connection information and attribute information. For example, the 3D data generator 115 may reduce the amount of data by deleting duplicate vertices or perform transformations on the vertex information (such as position shifting, rotation, or normalization). The 3D data generator 115 may also render the attribute information.
[0113] Furthermore, although the three-dimensional data generator 115 is a component of the three-dimensional data encoding system 110 in Figure 12, it may also be located externally and independently of the three-dimensional data encoding system 110.
[0114] The sensor terminal that provides sensor data for generating three-dimensional data may be, for example, a moving object such as an automobile, a flying object such as an airplane, a mobile terminal, or a camera. Alternatively, distance sensors such as LIDAR, millimeter-wave radar, infrared sensors, or rangefinders, stereo cameras, or combinations of multiple monocular cameras may be used as sensor terminals.
[0115] The sensor data may include the distance (position) of the object, monocular camera images, stereo camera images, color, reflectivity, sensor attitude, orientation, gyroscope, sensing position (GPS information or altitude), speed, acceleration, sensing time, temperature, atmospheric pressure, humidity, or magnetism.
[0116] The three-dimensional data encoder 113 corresponds to the encoding device 100 shown in Figure 5, etc. For example, the three-dimensional data encoder 113 encodes three-dimensional data and generates encoded data. The three-dimensional data encoder 113 also generates control information during the encoding of three-dimensional data. The three-dimensional data encoder 113 then inputs the encoded data along with the control information to the system multiplexer 114.
[0117] The encoding method for three-dimensional data may be a geometry-based encoding method or a video codec-based encoding method. Here, the geometry-based encoding method can also be referred to as a geometry-based encoding method. The video codec-based encoding method can also be referred to as a video-based encoding method.
[0118] The system multiplexer 114 multiplexes the encoded data and control information input from the three-dimensional data encoder 113 and generates multiplexed data using a predetermined multiplexing scheme. The system multiplexer 114 may also multiplex other media such as video, audio, subtitles, application data, or document files, or reference time information, along with the encoded data and control information of the three-dimensional data. Furthermore, the system multiplexer 114 may also multiplex sensor data or attribute information related to the three-dimensional data.
[0119] For example, the multiplexed data may have a file format for storage or a packet format for transmission. ISOBMFF or an ISOBMFF-based format may be used as these formats. Alternatively, MPEG-DASH, MMT, MPEG-2 TS Systems, or RTP may be used.
[0120] The multiplexed data is then output as a transmission signal to the external connector 310 by the input / output processor 112. The multiplexed data may be transmitted as a transmission signal by wire or by wireless. Alternatively, the multiplexed data may be stored in internal memory or storage device. The multiplexed data may be transmitted to a cloud server via the internet or stored in an external storage device.
[0121] For example, the transmission or storage of multiplexed data is carried out in a manner appropriate to the medium for transmission or storage, such as broadcasting or telecommunications. Communication protocols such as HTTP, FTP, TCP, UDP, IP, or combinations thereof may be used. Furthermore, either a pull-type or push-type communication method may be used.
[0122] For wired transmission, Ethernet®, USB, RS-232C, HDMI®, or coaxial cable may be used. For wireless transmission, 3GPP®, IEEE 3G / 4G / 5G, wireless LAN, Wi-Fi, Bluetooth, or millimeter wave may be used. Furthermore, as a broadcasting method, for example, DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 may be used.
[0123] The sensor data may also be input to the three-dimensional data generator 115 or the system multiplexer 114. Alternatively, the three-dimensional data or encoded data may be output directly as a transmission signal to the external connector 310 via the input / output processor 112. The transmission signal output from the three-dimensional data encoding system 110 is input to the three-dimensional data decoding system 210 via the external connector 310.
[0124] Furthermore, each operation of the three-dimensional data encoding system 110 may be controlled by a controller 111 that executes an application program.
[0125] In the three-dimensional data decoding system 210, the transmission signal is input to the input / output processor 212. The input / output processor 212 decodes the transmission signal into multiplexed data in file format or packet format and inputs the multiplexed data to the system demultiplexer 214. The system demultiplexer 214 obtains encoded data and control information from the multiplexed data and inputs them to the three-dimensional data decoder 213. The system demultiplexer 214 may also extract other media or reference time information from the multiplexed data.
[0126] The three-dimensional data decoder 213 corresponds to the decoding device 200 shown in Figure 7, etc. For example, the three-dimensional data decoder 213 decodes three-dimensional data from encoded data based on a predetermined encoding scheme. The three-dimensional data is then presented to the user by the presenter 215.
[0127] In addition, additional information such as sensor data may be input to the display device 215. The display device 215 may present three-dimensional data based on the additional information. Furthermore, user instructions may be input from the user terminal to the user interface 216. The display device 215 may then present three-dimensional data based on the input instructions.
[0128] The input / output processor 212 may also acquire three-dimensional data and encoded data from the external connector 310.
[0129] Furthermore, each operation of the three-dimensional data decoding system 210 may be controlled by a controller 211 that executes an application program.
[0130] Figure 13 is a conceptual diagram showing an example of the configuration of point cloud data according to this embodiment. Point cloud data is data representing a three-dimensional object.
[0131] Specifically, a point cloud consists of multiple points and contains positional information indicating the three-dimensional coordinate position of each point, as well as attribute information indicating the attributes of each point. Positional information is also expressed as geometry.
[0132] The type of attribute information may be, for example, color or reflectance. A single point may be associated with attribute information relating to one type, a single point may be associated with attribute information relating to multiple different types, or a single point may be associated with attribute information having multiple values for the same type.
[0133] Figure 14 is a conceptual diagram showing an example of a data file for point cloud data according to this embodiment. In this example, there is a one-to-one correspondence between location information items and attribute information items, and it shows the location information and attribute information of N points that constitute the point cloud data. In this example, the location information is information that indicates the three-dimensional coordinate position on the three axes x, y, and z, and the attribute information is information that indicates the color in RGB. A PLY file or the like can be used as a typical data file for point cloud data.
[0134] Figure 15 is a conceptual diagram showing an example of the configuration of mesh data according to this embodiment. Mesh data is data used in CG (Computer Graphics), etc., and is three-dimensional mesh data that shows the three-dimensional shape of an object with multiple faces. Each face is also represented as a polygon and has the shape of a polygon such as a triangle or quadrilateral.
[0135] Specifically, a three-dimensional mesh consists of multiple points that make up a point cloud, as well as multiple edges and multiple faces. Each point can also be expressed as a vertex or position. Each edge corresponds to a line segment connected by two vertices. Each face corresponds to a region enclosed by three or more edges.
[0136] Furthermore, a three-dimensional mesh contains positional information indicating the three-dimensional coordinate positions of its vertices. This positional information is also referred to as vertex information or geometry. A three-dimensional mesh also contains connection information indicating the relationships between multiple vertices that constitute an edge or face. This connection information is also referred to as connectivity. Finally, a three-dimensional mesh contains attribute information indicating the attributes of vertices, edges, or faces. This attribute information in a three-dimensional mesh is also referred to as texture.
[0137] For example, attribute information may indicate the color, reflectance, or normal vector for a vertex, edge, or face. The direction of the normal vector can represent the front and back of the face.
[0138] Object files or similar formats may be used as the data file format for mesh data.
[0139] Figure 16 is a conceptual diagram showing an example of a data file for mesh data according to this embodiment. In this example, the data file includes position information G(1) to G(N) for N vertices constituting the three-dimensional mesh, and attribute information A1(1) to A1(N) for the N vertices. In addition, this example includes M pieces of attribute information A2(1) to A2(M). The items of the attribute information do not have to correspond one-to-one with vertices, nor do they have to correspond one-to-one with faces. Furthermore, attribute information may not exist at all.
[0140] Connection information is indicated by a combination of vertex indices. n[1, 3, 4] represents a triangular face composed of three vertices n=1, n=3, and n=4. Also, m[2, 4, 6] indicates that attribute information m=2, m=4, and m=6 correspond to the three vertices, respectively.
[0141] Furthermore, the actual content of the attribute information may be described in a separate file. A pointer to that content may be associated with a vertex or face, etc. For example, attribute information indicating an image for a face may be stored in a two-dimensional attribute map file. The file name of the attribute map and the two-dimensional coordinate values in the attribute map may be described in attribute information A2(1) to A2(M). The method of specifying attribute information for a face is not limited to these methods, and any method may be used.
[0142] Figure 17 is a conceptual diagram showing the types of three-dimensional data according to this embodiment. Point cloud data and mesh data may represent static objects or dynamic objects. Static objects are objects that do not change over time, while dynamic objects are objects that change over time. Static objects may correspond to three-dimensional data for any given point in time.
[0143] For example, point cloud data for any given time point may be referred to as a PCC frame. Similarly, mesh data for any given time point may be referred to as a mesh frame. Furthermore, both PCC frames and mesh frames may simply be referred to as frames.
[0144] Furthermore, the object's area may be limited to a certain range, like in regular video data, or it may not be limited, like in map data. Also, the density of points or surfaces can be determined in various ways. Sparse point cloud data or sparse mesh data may be used, or dense point cloud data or dense mesh data may be used.
[0145] Next, the encoding and decoding of point clouds or three-dimensional meshes will be described. The apparatus, processing, or syntax for encoding and decoding vertex information of a three-dimensional mesh in this disclosure may be applied to the encoding and decoding of point clouds. The apparatus, processing, or syntax for encoding and decoding of point clouds in this disclosure may be applied to the encoding and decoding of vertex information of a three-dimensional mesh.
[0146] Furthermore, the apparatus, processing, or syntax for encoding and decoding point cloud attribute information in this disclosure may also be applied to the encoding and decoding of connection information or attribute information of a three-dimensional mesh.
[0147] Furthermore, at least some of the processing may be shared between the encoding and decoding of point cloud data and the encoding and decoding of mesh data. This can reduce the size of the circuit and software programs.
[0148] Figure 18 is a block diagram showing an example configuration of a three-dimensional data encoder 113 according to this embodiment. In this example, the three-dimensional data encoder 113 comprises a vertex information encoder 121, an attribute information encoder 122, a metadata encoder 123, and a multiplexer 124. The vertex information encoder 121, the attribute information encoder 122, and the multiplexer 124 may correspond to the vertex information encoder 101, the attribute information encoder 103, and the post-processor 105 in Figure 6, etc.
[0149] In this example, the three-dimensional data encoder 113 encodes the three-dimensional data according to a geometry-based encoding scheme. In encoding according to a geometry-based encoding scheme, the three-dimensional structure is taken into consideration. In addition, in encoding according to a geometry-based encoding scheme, attribute information is encoded using the configuration information obtained in encoding the vertex information.
[0150] Specifically, first, vertex information, attribute information, and metadata contained in the three-dimensional data generated from sensor data are input to the vertex information encoder 121, attribute information encoder 122, and metadata encoder 123, respectively. Here, connection information contained in the three-dimensional data may be treated in the same way as attribute information. In the case of point cloud data, position information may be treated as vertex information.
[0151] The vertex information encoder 121 encodes the vertex information into compressed vertex information and outputs the compressed vertex information as encoded data to the multiplexer 124. The vertex information encoder 121 also generates metadata for the compressed vertex information and outputs it to the multiplexer 124. Furthermore, the vertex information encoder 121 generates configuration information and outputs it to the attribute information encoder 122.
[0152] The attribute information encoder 122 uses the configuration information generated by the vertex information encoder 121 to encode the attribute information into compressed attribute information and outputs the compressed attribute information as encoded data to the multiplexer 124. The attribute information encoder 122 also generates metadata for the compressed attribute information and outputs it to the multiplexer 124.
[0153] The metadata encoder 123 encodes compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The metadata encoded by the metadata encoder 123 may be used for encoding vertex information and attribute information.
[0154] The multiplexer 124 multiplexes the compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, the metadata of the compressed attribute information, and the compressed metadata into a bitstream. The multiplexer 124 then inputs the bitstream to the system layer.
[0155] Figure 19 is a block diagram showing an example configuration of a three-dimensional data decoder 213 according to this embodiment. In this example, the three-dimensional data decoder 213 includes a vertex information decoder 221, an attribute information decoder 222, a metadata decoder 223, and a demultiplexer 224. The vertex information decoder 221, attribute information decoder 222, and demultiplexer 224 may correspond to the vertex information decoder 201, attribute information decoder 203, and preprocessor 204 in Figure 8, etc.
[0156] In this example, the three-dimensional data decoder 213 decodes the three-dimensional data according to a geometry-based coding scheme. In decoding according to a geometry-based coding scheme, the three-dimensional structure is taken into consideration. In addition, in decoding according to a geometry-based coding scheme, attribute information is decoded using the configuration information obtained in the decoding of vertex information.
[0157] Specifically, first, the bitstream is input from the system layer to the demultiplexer 224. The demultiplexer 224 separates the compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, the metadata of the compressed attribute information, and the compressed metadata from the bitstream. The compressed vertex information and the metadata of the compressed vertex information are input to the vertex information decoder 221. The compressed attribute information and the metadata of the compressed attribute information are input to the attribute information decoder 222. The metadata is input to the metadata decoder 223.
[0158] The vertex information decoder 221 decodes vertex information from compressed vertex information using the metadata of the compressed vertex information. The vertex information decoder 221 also generates configuration information and outputs it to the attribute information decoder 222. The attribute information decoder 222 decodes attribute information from compressed attribute information using the configuration information generated by the vertex information decoder 221 and the metadata of the compressed attribute information. The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 may be used for decoding vertex information and attribute information.
[0159] Subsequently, vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. For example, this metadata is metadata for vertex information and attribute information, and can be used in application programs.
[0160] Figure 20 is a block diagram showing another configuration example of the three-dimensional data encoder 113 according to this embodiment. In this example, the three-dimensional data encoder 113 includes a vertex image generator 131, an attribute image generator 132, a metadata generator 133, a video encoder 134, a metadata encoder 123, and a multiplexer 124. The vertex image generator 131, the attribute image generator 132, and the video encoder 134 may correspond to the vertex information encoder 101 and the attribute information encoder 103 in Figure 6, etc.
[0161] In this example, the three-dimensional data encoder 113 encodes the three-dimensional data according to a video-based encoding scheme. In encoding according to a video-based encoding scheme, multiple two-dimensional images are generated from the three-dimensional data, and these multiple two-dimensional images are encoded according to a video encoding scheme. Here, the video encoding scheme may be HEVC (High Efficiency Video Coding) or VVC (Versatile Video Coding), etc.
[0162] Specifically, first, vertex information and attribute information contained in the three-dimensional data generated from sensor data are input to the metadata generator 133. The vertex information and attribute information are then input to the vertex image generator 131 and the attribute image generator 132, respectively. Furthermore, metadata contained in the three-dimensional data is input to the metadata encoder 123. Here, connection information contained in the three-dimensional data may be treated similarly to attribute information. In the case of point cloud data, position information may be treated as vertex information.
[0163] The metadata generator 133 generates map information for multiple two-dimensional images from vertex information and attribute information. The metadata generator 133 then inputs the map information to the vertex image generator 131, the attribute image generator 132, and the metadata encoder 123.
[0164] The vertex image generator 131 generates a vertex image based on the vertex information and map information and inputs it to the video encoder 134. The attribute image generator 132 generates an attribute image based on the attribute information and map information and inputs it to the video encoder 134.
[0165] The video encoder 134 encodes the vertex image and attribute image into compressed vertex information and compressed attribute information, respectively, according to the video encoding scheme, and outputs the compressed vertex information and compressed attribute information as encoded data to the multiplexer 124. The video encoder 134 also generates metadata for the compressed vertex information and metadata for the compressed attribute information and outputs them to the multiplexer 124.
[0166] The metadata encoder 123 encodes compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The compressible metadata includes map information. The metadata encoded by the metadata encoder 123 may also be used for encoding vertex information and attribute information.
[0167] The multiplexer 124 multiplexes the compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, the metadata of the compressed attribute information, and the compressed metadata into a bitstream. The multiplexer 124 then inputs the bitstream to the system layer.
[0168] Figure 21 is a block diagram showing another configuration example of the three-dimensional data decoder 213 according to this embodiment. In this example, the three-dimensional data decoder 213 includes a vertex information generator 231, an attribute information generator 232, a video decoder 234, a metadata decoder 223, and a demultiplexer 224. The vertex information generator 231, the attribute information generator 232, and the video decoder 234 may correspond to the vertex information decoder 201 and the attribute information decoder 203 in Figure 8, etc.
[0169] In this example, the three-dimensional data decoder 213 decodes the three-dimensional data according to a video-based coding scheme. In decoding according to a video-based coding scheme, multiple two-dimensional images are decoded according to the video coding scheme, and three-dimensional data is generated from the multiple two-dimensional images. Here, the video coding scheme may be HEVC (High Efficiency Video Coding) or VVC (Versatile Video Coding), etc.
[0170] Specifically, first, the bitstream is input from the system layer to the demultiplexer 224. The demultiplexer 224 separates the compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, the metadata of the compressed attribute information, and the compressed metadata from the bitstream. The compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, and the metadata of the compressed attribute information are input to the video decoder 234. The compressed metadata is input to the metadata decoder 223.
[0171] The video decoder 234 decodes the vertex image according to the video encoding scheme. In doing so, the video decoder 234 decodes the vertex image from the compressed vertex information using the metadata of the compressed vertex information. The video decoder 234 then inputs the vertex image to the vertex information generator 231. The video decoder 234 also decodes the attribute image according to the video encoding scheme. In doing so, the video decoder 234 decodes the attribute image from the compressed attribute information using the metadata of the compressed attribute information. The video decoder 234 then inputs the attribute image to the attribute information generator 232.
[0172] The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 includes map information used for generating vertex information and attribute information. The metadata decoded by the metadata decoder 223 may also be used for decoding vertex images and attribute images.
[0173] The vertex information generator 231 reconstructs vertex information from the vertex image according to the map information contained in the metadata decoded by the metadata decoder 223. The attribute information generator 232 reconstructs attribute information from the attribute image according to the map information contained in the metadata decoded by the metadata decoder 223.
[0174] Subsequently, vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. For example, this metadata is metadata for vertex information and attribute information, and can be used in application programs.
[0175] Figure 22 is a conceptual diagram showing a specific example of the encoding process according to this embodiment. Figure 22 shows a three-dimensional data encoder 113 and a description encoder 148. In this example, the three-dimensional data encoder 113 comprises a two-dimensional data encoder 141 and a mesh data encoder 142. The two-dimensional data encoder 141 comprises a texture encoder 143. The mesh data encoder 142 comprises a vertex information encoder 144 and a connectivity information encoder 145.
[0176] The vertex information encoder 144, the connection information encoder 145, and the texture encoder 143 may correspond to the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 in Figure 6, etc.
[0177] For example, the two-dimensional data encoder 141 operates as a texture encoder 143 and generates a texture file by encoding the texture corresponding to the attribute information as two-dimensional data according to an image encoding scheme or video encoding scheme.
[0178] Furthermore, the mesh data encoder 142 operates as a vertex information encoder 144 and a connection information encoder 145, and generates a mesh file by encoding vertex information and connection information. The mesh data encoder 142 may also encode mapping information for textures. The encoded mapping information may then be included in the mesh file.
[0179] Furthermore, the description encoder 148 generates a description file by encoding the description corresponding to metadata such as text data. The description encoder 148 may encode the description at the system layer. For example, the description encoder 148 may be included in the system multiplexer 114 in Figure 12.
[0180] The above operation generates a bitstream containing texture files, mesh files, and description files. These files may be multiplexed into the bitstream in a file format such as glTF (Graphics Language Transmission Format) or USD (Universal Scene Description).
[0181] The three-dimensional data encoder 113 may also include two mesh data encoders, namely the mesh data encoder 142. For example, one mesh data encoder encodes the vertex and connection information of a static three-dimensional mesh, while the other mesh data encoder encodes the vertex and connection information of a dynamic three-dimensional mesh.
[0182] Correspondingly, two mesh files may be included in the bitstream. For example, one mesh file may correspond to a static 3D mesh, and the other mesh file may correspond to a dynamic 3D mesh.
[0183] Furthermore, a static three-dimensional mesh may be a three-dimensional mesh of an intraframe encoded using intraprediction, and a dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe encoded using interprediction. In addition, as information for the dynamic three-dimensional mesh, the difference information between the vertex information or connection information of the intraframe three-dimensional mesh and the vertex information or connection information of the interframe three-dimensional mesh may be used.
[0184] Figure 23 is a conceptual diagram showing a specific example of the decoding process according to this embodiment. Figure 23 shows a three-dimensional data decoder 213, a description decoder 248, and an presenter 247. In this example, the three-dimensional data decoder 213 comprises a two-dimensional data decoder 241, a mesh data decoder 242, and a mesh reconstructor 246. The two-dimensional data decoder 241 comprises a texture decoder 243. The mesh data decoder 242 comprises a vertex information decoder 244 and a connection information decoder 245.
[0185] The vertex information decoder 244, connection information decoder 245, texture decoder 243, and mesh reconstructor 246 may correspond to the vertex information decoder 201, connection information decoder 202, attribute information decoder 203, and post-processor 205, etc., shown in Figure 8. The presenter 247 may correspond to the presenter 215, etc., shown in Figure 12.
[0186] For example, the two-dimensional data decoder 241 operates as a texture decoder 243 and decodes the texture corresponding to the attribute information from the texture file as two-dimensional data according to the image encoding scheme or video encoding scheme.
[0187] Furthermore, the mesh data decoder 242 operates as a vertex information decoder 244 and a connection information decoder 245, decoding vertex information and connection information from the mesh file. The mesh data decoder 242 may also decode mapping information to textures from the mesh file.
[0188] Furthermore, the description decoder 248 decodes the description corresponding to metadata such as text data from the description file. The description decoder 248 may decode the description at the system layer. For example, the description decoder 248 may be included in the system demultiplexer 214 in Figure 12.
[0189] The mesh reconstructor 246 reconstructs a three-dimensional mesh from vertex information, connection information, and textures according to the description. The presenter 247 renders and outputs the three-dimensional mesh according to the description.
[0190] The above process reconstructs and outputs a 3D mesh from a bitstream containing texture files, mesh files, and description files.
[0191] The three-dimensional data decoder 213 may also include two mesh data decoders, which are mesh data decoders 242. For example, one mesh data decoder decodes the vertex and connection information of a static three-dimensional mesh, while the other mesh data decoder decodes the vertex and connection information of a dynamic three-dimensional mesh.
[0192] Correspondingly, two mesh files may be included in the bitstream. For example, one mesh file may correspond to a static 3D mesh, and the other mesh file may correspond to a dynamic 3D mesh.
[0193] Furthermore, a static three-dimensional mesh may be a three-dimensional mesh of an intraframe encoded using intraprediction, and a dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe encoded using interprediction. In addition, as information for the dynamic three-dimensional mesh, the difference information between the vertex information or connection information of the intraframe three-dimensional mesh and the vertex information or connection information of the interframe three-dimensional mesh may be used.
[0194] A coding scheme for dynamic three-dimensional meshes is sometimes called DMC (Dynamic Mesh Coding). Similarly, a video-based coding scheme for dynamic three-dimensional meshes is sometimes called V-DMC (Video-based Dynamic Mesh Coding).
[0195] The encoding method for point clouds is sometimes called PCC (Point Cloud Compression). The video-based encoding method for point clouds is sometimes called V-PCC (Video-based Point Cloud Compression). The geometry-based encoding method for point clouds is sometimes called G-PCC (Geometry-based Point Cloud Compression).
[0196] <Implementation Example> Figure 24 is a block diagram showing an implementation example of the encoding device 100 according to this embodiment. The encoding device 100 includes a circuit 151 and a memory 152. For example, the multiple components of the encoding device 100 shown in Figure 5, etc., are implemented by the circuit 151 and memory 152 shown in Figure 24.
[0197] Circuit 151 is an information processing circuit and is a circuit that can access memory 152. For example, circuit 151 is a dedicated or general-purpose electrical circuit for encoding a three-dimensional mesh. Circuit 151 may also be a processor such as a CPU. Alternatively, circuit 151 may be a collection of multiple electrical circuits.
[0198] Memory 152 is a dedicated or general-purpose memory in which information for the circuit 151 to encode a three-dimensional mesh is stored. Memory 152 may be an electrical circuit, or it may be connected to circuit 151. Memory 152 may also be included in circuit 151. Memory 152 may also be a collection of multiple electrical circuits. Memory 152 may also be a magnetic disk or an optical disk, or it may be described as storage or a recording medium. Memory 152 may also be a non-volatile memory or a volatile memory.
[0199] For example, memory 152 may store a three-dimensional mesh or a bitstream. Alternatively, memory 152 may store a program for circuit 151 to encode the three-dimensional mesh.
[0200] Furthermore, not all of the multiple components shown in Figure 5, etc., are to be implemented in the encoding device 100, nor are all of the multiple processes shown herein to be performed. Some of the multiple components shown in Figure 5, etc., may be included in other devices, and some of the multiple processes shown herein may be executed by other devices. In addition, the multiple components of this disclosure may be implemented in any combination in the encoding device 100, and the multiple processes of this disclosure may be performed in any combination.
[0201] Figure 25 is a block diagram showing an example of the implementation of the decoding device 200 according to this embodiment. The decoding device 200 includes a circuit 251 and a memory 252. For example, the multiple components of the decoding device 200 shown in Figure 7, etc., are implemented by the circuit 251 and memory 252 shown in Figure 25.
[0202] Circuit 251 is an information processing circuit and is a circuit that can access memory 252. For example, circuit 251 is a dedicated or general-purpose electrical circuit for decoding a three-dimensional mesh. Circuit 251 may also be a processor such as a CPU. Alternatively, circuit 251 may be a collection of multiple electrical circuits.
[0203] Memory 252 is a dedicated or general-purpose memory that stores information for circuit 251 to decode the three-dimensional mesh. Memory 252 may be an electrical circuit and may be connected to circuit 251. Memory 252 may also be included in circuit 251. Memory 252 may also be a collection of multiple electrical circuits. Memory 252 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory 252 may also be a non-volatile memory or a volatile memory.
[0204] For example, memory 252 may store a three-dimensional mesh or a bitstream. Alternatively, memory 252 may store a program for circuit 251 to decode the three-dimensional mesh.
[0205] Furthermore, the decoding device 200 does not need to implement all of the components shown in Figure 7, etc., nor does it need to perform all of the processes shown herein. Some of the components shown in Figure 7, etc., may be included in other devices, and some of the processes shown herein may be performed by other devices. In addition, the decoding device 200 may implement the components of this disclosure in any combination, and the processes of this disclosure may be performed in any combination.
[0206] The encoding and decoding methods, including the steps performed by each component of the encoding device 100 and decoding device 200 of this disclosure, may be performed by any device or system. For example, part or all of the encoding and decoding methods may be performed by a computer equipped with a processor, memory, and input / output circuits, etc. In this case, the encoding and decoding methods may be performed by the computer executing a program that causes the computer to perform the encoding and decoding methods.
[0207] Furthermore, a non-temporary computer-readable recording medium such as a CD-ROM may contain either a program or a bitstream.
[0208] An example of a program may be a bitstream. For instance, a bitstream containing an encoded three-dimensional mesh may include syntax elements to cause the decoding device 200 to decode the three-dimensional mesh. The bitstream then causes the decoding device 200 to decode the three-dimensional mesh according to the syntax elements contained within the bitstream. Therefore, a bitstream can perform a similar role to a program.
[0209] The above bitstream may be an encoded bitstream containing an encoded three-dimensional mesh, or a multiplexed bitstream containing an encoded three-dimensional mesh and other information.
[0210] Furthermore, each component of the encoding device 100 and the decoding device 200 may be made of dedicated hardware, general-purpose hardware that executes the above-mentioned program, or a combination thereof. The general-purpose hardware may also consist of a memory on which the program is stored, and a general-purpose processor that reads the program from the memory and executes it. Here, the memory may be semiconductor memory or a hard disk, and the general-purpose processor may be a CPU.
[0211] Furthermore, dedicated hardware may consist of memory and a dedicated processor, etc. For example, a dedicated processor may refer to memory for recording data and execute an encoding method and a decoding method.
[0212] Furthermore, each component of the encoding device 100 and the decoding device 200 may be an electrical circuit, as described above. These electrical circuits may form a single electrical circuit as a whole, or they may be separate electrical circuits. These electrical circuits may correspond to dedicated hardware, or they may correspond to general-purpose hardware that executes the above-mentioned program, etc. Furthermore, the encoding device 100 and the decoding device 200 may be implemented as integrated circuits.
[0213] Furthermore, the encoding device 100 may be a transmitting device that transmits a three-dimensional mesh. The decoding device 200 may be a receiving device that receives a three-dimensional mesh.
[0214] <Encoding and Decoding of Displacements> Here, the following terms are used as examples.
[0215] (1) Image An image is a data unit composed of a collection of pixels, and includes a picture or a block smaller than a picture. Images include both moving images and still images.
[0216] (2) A picture is an image processing unit composed of a collection of pixels, and is also called a frame or field.
[0217] (3) A block is a processing unit consisting of a specific number of pixels. The term block is also used in the examples shown below. The shape of a block is not particularly limited. A block may be, for example, a rectangular shape of M x N pixels, or a square shape of M x M pixels. A block may also be triangular, circular, or have other shapes. Examples of blocks are as follows.
[0218] - Slices, tiles, or bricks - CTU, superblock, or basic partitioning unit - VPDU, hardware processing partitioning unit - CU, processing block unit, prediction block unit (PU), or orthogonal transformation block unit (TU) - subblocks
[0219] (4) A pixel or sample is the smallest point, or in other words, the smallest unit, of an image. A pixel or sample includes not only pixels at integer positions but also pixels at sub-pixel positions generated from pixels at integer positions.
[0220] (5) Pixel values or sample values Pixel values or sample values are eigenvalues of a pixel. Pixel values or sample values may include luma values, chroma values, or RGB tonal levels, and may also include depth values or binary values of 0 or 1.
[0221] (6) Flags A flag represents one or more bits. A flag is, for example, a parameter or index represented by two or more bits. A flag may also represent a value that is not represented by a binary number, but by a non-binary number.
[0222] (7) A signal is a symbol or encoded information used to transmit information. Signals include discrete digital signals or continuous analog signals.
[0223] (8) Stream or Bitstream A stream or bitstream is a sequence of digital data that represents a flow of digital data. A stream or bitstream may be a single stream or may consist of multiple streams having multiple layers. A stream or bitstream may be transmitted by serial communication using a single transmission path or by packet communication using multiple transmission paths.
[0224] (9) Differences: For scalar quantities, differences can include simple differences (x - y) and difference calculations. Differences can include the absolute value of the difference (|x - y|), the square of the difference (x^2 - y^2), the square root of the difference (√(x - y)), weighted differences (ax - by, where a and b are constants), or offset differences (x - y + a, where a is the offset).
[0225] (10) For scalar quantities, sums can include simple sums (x + y) and addition operations. Sums can include the absolute value of the sum (|x + y|), the sum of squares (x^2 + y^2), the square root of the sum (√(x + y)), a weighted sum (ax + by, where a and b are constants), or an offset sum (x + y + a, where a is the offset).
[0226] (11) "Based on" The expression "based on something" means that other things besides that "something" may also be considered. Also, "based on" can be used when a direct result is obtained, or when a result is obtained after intermediate results.
[0227] (12) "Used" or "Used" The expression "something was used" or "something was used" means that something other than that "something" may also be considered. Also, the expression "used" or "used" can be used when a direct result is obtained, or when a result is obtained after an intermediate result.
[0228] (13) Prohibition "To prohibit" can be rephrased as "not to permit." Also, "not prohibited / prohibited" or "permitted / permitted" does not necessarily mean "obligation."
[0229] (14) "Restriction" or "Limitation" "Restriction" or "Limitation" can be rephrased as "not permitted / not allowed" or "not permitted / permitted." Also, "prohibited / not prohibited" or "not permitted / permitted" does not necessarily mean "obligation." Furthermore, the prohibition, quantitatively or qualitatively, may be partial or entire.
[0230] (15) Chroma The term chroma is an adjective represented by the symbols Cb or Cr, indicating that a sample sequence or a single sample represents one of two color difference signals associated with a primary color. The term chroma is sometimes used instead of the term chrominance.
[0231] (16) Luma The term luma is an adjective represented by the symbol or subscript Y or L, indicating that a sample sequence or single sample represents a monochrome signal relating to a primary color. The term luma is sometimes used as an alternative to the term luminance.
[0232] The coding and decoding system of this embodiment will be described below.
[0233] A typical three-dimensional model (also called a 3D model) digitally represents an object so that the user can explore the model using zoom, pan, and rotate in all three dimensions while rendering it over time. One way to construct such a representation is to build a 3D mesh using triangles. In the above model, the positions of the triangle vertices, the connectivity of the triangle vertices to each other, and their associated attributes (such as normals or UV patches) are stored.
[0234] Storing all this information in an uncompressed format requires a very large amount of memory, and therefore a very large bandwidth for transmission. The triangles that form a mesh often have attributes similar to repeating patterns, especially in temporal and spatial neighborhoods. These repetitions can be used to develop efficient encoding and decoding methods for storage and transmission. One such encoding and decoding method is Video-based Dynamic Mesh Coding (V-DMC).
[0235] Figure 26 is a block diagram showing different configuration examples of the coding and decoding system according to this embodiment. As shown in Figure 26, the coding and decoding system includes an coding device 100 and a decoding device 200.
[0236] The encoding / decoding system accepts a three-dimensional mesh (also called a 3D mesh) as input in the form of three-dimensional coordinates of vertices (vertex information), connectivity (connection information), and associated attributes (attribute information). Note that the 3D mesh can include not only geometry but also texture maps.
[0237] The encoding device 100 captures the input 3D mesh (also called the input 3D mesh or input mesh) in the form of the three-dimensional coordinates of the vertices, connectivity, and associated attributes. The encoding device 100 encodes all the associated information into a stream. The stream may consist of a single bitstream or multiple bitstreams.
[0238] Network 300 transmits the stream generated by the encoding device to the decoding device 200. Network 300 may be the Internet, a WAN (Wide Area Network), a LAN (Local Area Network), or any combination thereof. Furthermore, network 300 is not necessarily limited to a bidirectional communication network, but may also be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Alternatively, a recording medium such as a DVD (Digital Versatile Disc) or a BD (Blue-Ray Disc) on which the stream is recorded may be used instead of network 300.
[0239] The stream is transmitted to the decoder 200 via the network 300. The decoder 200 decodes the bitstream and generates a three-dimensional mesh using the three-dimensional coordinates, connectivity, and associated attributes of the decoded vertices. The decoder 200 outputs the generated three-dimensional mesh (also called the output 3D mesh or output mesh).
[0240] Figure 27 shows another example of the encoding device 100 configuration.
[0241] As shown in Figure 27, the encoding device 100 includes a preprocessor 1103 and a compressor 1106.
[0242] The encoding device 100 reads the input mesh 1101 and attribute map 1102 and passes them to the preprocessor 1103. The preprocessor 1103 processes the input mesh and extracts the base mesh 1104 and displacement data 1105. The attribute map 1102, along with the extracted base mesh 1104 and displacement data 1105, is passed to the compressor 1106.
[0243] Furthermore, the compressor 1106 compresses the base mesh 1104, displacement data 1105, and attribute map 1102 to generate a bitstream 1107. The compressor 1106 can transmit additional information to the decoder 200 by further including metadata 1108 in the bitstream 1107.
[0244] Figure 28 shows another example of the configuration of the decoding device 200.
[0245] As shown in Figure 28, the decoding device 200 includes an expander 2102 and a post-processing unit 2106.
[0246] The decoding device 200 reads the bitstream 2101 and passes it to the decompressor 2102. The decompressor 2102 decompresses the base mesh 2103, displacement data 2104, and attribute map 2108 from the bitstream 2101 and passes them to the post-processor 2106. An example of displacement data 2104 is a displacement vector.
[0247] Furthermore, the post-processor 2106 generates the output mesh 2107 by processing the base mesh 2103 according to the displacement data 2104 and attribute map 2108. The post-processor 2106 may also use information from metadata 2105 to generate the output mesh 2107.
[0248] Figure 29 is a block diagram showing yet another configuration example of the encoding device 100 according to this embodiment.
[0249] In this example, the encoding device 100 includes a volumetric capturer 511, a projector 512, a base mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and optionally one or more other type encoders 516.
[0250] The volumetric capture device 511 captures the content and outputs the captured content to the projector 512.
[0251] The projector 512 projects content onto a three-dimensional mesh frame containing vertex geometry coordinates (vertex coordinates indicating the position of vertices), texture coordinates, and connectivity data (connection information). This data is output to the base mesh encoder 513, the displacement encoder 514, the attribute encoder 515, and optionally one or more other type encoders 516. Each encoder compresses the data into a bitstream.
[0252] Figure 30 is a block diagram showing yet another configuration example of the decoding device 200 according to this embodiment.
[0253] In this example, the decoding device 200 comprises a base mesh decoder 613, a displacement decoder 614, an attribute decoder 615, one or more other type decoders 616, and a three-dimensional reconstructor 617.
[0254] The bitstream is sent to the base mesh decoder 613, the displacement decoder 614, the attribute decoder 615, and optionally one or more other type decoders 616. These decoders decode the bitstream to generate decoded data containing vertex geometry coordinates, texture coordinates, and connectivity data. The decoded data is then sent to the 3D reconstructor 617, where the 3D mesh frame is reconstructed.
[0255] The encoding process performed by the encoding device 100 will be described in detail below.
[0256] Figure 31 is a flowchart showing the processing of the encoding device 100. Figure 32 is an explanatory diagram conceptually illustrating the encoding of a mesh frame. The processing of the encoding device 100 will be explained with reference to Figures 31 and 32.
[0257] In step S101, the encoding device 100 reads the input mesh frame, which is a 3D mesh frame, and its attributes. The input mesh frame is the mesh frame input to the encoding device 100. An example of the input mesh frame, a 3D mesh frame, is shown as mesh frame 1301 (see Figure 32).
[0258] In step S102, the encoding device 100 generates a base mesh frame with fewer vertices than the input mesh frame by performing a decimation process on the input mesh frame read in step S101. The base mesh frame generated by decimating the mesh frame 1301 is shown as the base mesh frame 1302 (see Figure 32).
[0259] In step S103, the encoding device 100 calculates displacement information that the decoding device 200 uses to reconstruct the mesh frame. The displacement information corresponds to a displacement vector from the vertices of the base mesh frame generated in step S102 to the vertices of the input mesh frame. One method for calculating the displacement information is to subtract the coordinates of the vertices of the base mesh frame from the coordinates of the vertices of the input mesh frame. The displacement information calculated from the mesh frame 1301 and the base mesh frame 1302 is shown as displacement information 1303 (see Figure 32). The displacement information 1303 is in vector form, or in other words, it is expressed as a displacement vector.
[0260] In step S104, the encoding device 100 encodes the base mesh frame generated in step S102, the displacement information generated in step S103, and the attributes of the input mesh frame into a bitstream (corresponding to a compressed bitstream). An example of a bitstream is shown as bitstream 1304 (see Figure 32).
[0261] Specifically, bitstream 1304 includes a video bitstream containing vertex coordinates and connection information for vertices A, C, E, and F, displacement information, and texture data, as well as a compressed attribute map (see Figure 32). The displacement information includes displacement information for displacing vertices based on vertex coordinates obtained from the subdivided base mesh frame. The compressed attribute map is texture coordinates for applying texture data to the mesh frame reconstructed using the base mesh frame and displacement information.
[0262] The decoding process performed by the decoding device 200 will be described in detail below.
[0263] Figure 33 is a flowchart showing the processing of the decoding device 200. Figure 34 is an explanatory diagram conceptually illustrating the decoding of a 3D mesh. The processing of the decoding device 200 will be explained with reference to Figures 33 and 34.
[0264] In step S201, the decoding device 200 decodes the base mesh frame and attributes from the bitstream (corresponding to the compressed bitstream). An example of the decoded base mesh frame (corresponding to the decoded base mesh frame) is shown as the decoded base mesh frame 2301 (see Figure 34).
[0265] In step S202, the decoding device 200 generates subdivided vertices by performing a subdivision process on the base mesh frame decoded in step S201. An example of a base mesh frame containing subdivided vertices is shown as base mesh frame 2302 (see Figure 34).
[0266] In step S203, the decoding device 200 decodes displacement information from the bitstream (corresponding to a compressed bitstream). An example of the decoded displacement information is shown as displacement information 2303 (see Figure 34). Displacement information 2303 is in vector form, or in other words, it is represented as a displacement vector.
[0267] In step S204, the decoding device 200 reconstructs the shape of the mesh frame by moving the vertices of the base mesh frame, including the sub-divided vertices, to new positions using displacement information, and then restores the mesh frame by applying attribute information. An example of an attribute is a texture. An example of the reconstructed mesh frame is shown as mesh frame 2304 (see Figure 34).
[0268] Figure 35 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.
[0269] Figure 35 shows an example of a typical intra-decryption block diagram.
[0270] The decoding device shown in Figure 35 comprises an inverse multiplexer 1231, a switch 1232, a static mesh decoder 1233, a mesh buffer 1234, a motion decoder 1235, a base mesh reconstructor 1236, an inverse quantizer 1237, a video decoder 1238, an image umpacker 1239, an inverse quantizer 1240, an inverse wavelet converter 1241, a reconstructor 1242, a video decoder 1243, and a color converter 1244.
[0271] The demultiplexer 1231 acquires the compressed bitstream and separates the compressed data for the base mesh, the video containing displacement data (also called the displacement bitstream), and the video containing attribute data (also called the attribute bitstream). The compressed data for the base mesh is passed to the switch 1232. The switch 1232 decides whether to perform intra-decoding or inter-decoding based on the parameters in the bitstream.
[0272] If intra-decoding is selected, the bitstream is passed to a static mesh decoder 1233 that generates a quantized base mesh. The static mesh decoder 1233 is, for example, a decoder that uses an edge breaker algorithm to decode 3D mesh data. The static mesh decoder 1233 generates a quantized base mesh from the bitstream. The quantized base mesh generated by the static mesh decoder 1233 is stored in the mesh buffer 1234 for reference when inter-decoding is selected.
[0273] Switch 1232 passes compressed data about the base mesh to the motion decoder 1235 if inter-decoding is selected. The motion decoder 1235 receives the previously decoded, quantized base mesh and decodes motion data representing the difference in vertex coordinates between the quantized base mesh stored in the mesh buffer 1234 and the current quantized base mesh. The motion data and the quantized base mesh stored in the mesh buffer 1234 are used by the base mesh reconstructor 1236 to reconstruct the current quantized base mesh. The quantized base mesh obtained from inter-decoding or intra-decoding is passed to the inverse quantizer 1237 to obtain the decoded base mesh.
[0274] The video containing displacement data is passed to the video decoder 1238 because the bitstream contains displacement data in an image format having two chroma information and one luma information. The video decoder 1238 decodes the data using the video frame decompression method. In another example, the displacement data is decoded using an arithmetic decoder. This decompressed data is passed to the image umpacker 1239, which extracts wavelet coefficients associated with each vertex from the decompressed image data. The inverse quantizer 1240 inverse quantizes the quantized wavelet coefficients with three components associated with each vertex. The inverse wavelet transformer 1241 inverse transforms the result to obtain the finally decoded displacement data. The decoded displacement data and the decoded base mesh are passed to the reconstructor 1242. The reconstructor 1242 subdivides the edges of the decoded base mesh, displaces the vertices using the decoded displacement data, and obtains the decoded mesh.
[0275] The video containing attribute data is passed to another video decoder 1243 to obtain a decoded attribute bitstream. The decoded attribute bitstream is further processed by a color converter 1244 for color space and color format conversion to obtain a decoded attribute map.
[0276] Figure 36 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.
[0277] Figure 36 shows an example of a reconstructor that obtains a decoded 3D mesh 1256 from a decoded base mesh 1251 and decoded displacement data 1254.
[0278] The decoded base mesh 1251 is passed to the sub-divider 1252.
[0279] The sub-decomposer 1252 subdivides any two connected vertices in the entire 3D mesh by adding a new vertex between them. This process can be repeated several times, including the vertices created in the previous sub-decomposition step, to generate a predetermined number of vertices. Each iteration of sub-decomposition across the entire 3D mesh generates a new level of detail (LoD). The subdivided mesh 1253 and the decoded displacement data 1254 are passed to the displacementr 1255. The displacementr 1255 generates the decoded 3D mesh 1256 by moving each vertex to a new position according to the corresponding displacement data.
[0280] Subpartitioning will be explained below. Subpartitioning is performed by a subpartitioner (specifically, subpartitioner 1206 or subpartitioner 2204).
[0281] Figure 37 is an explanatory diagram showing an example of subdivision.
[0282] The base mesh shown in Figure 37(a) includes vertices A, B, and C, and connectivity information indicating their connectivity.
[0283] Figure 37(b) shows the mesh generated by the first subdivision, in other words, the mesh after the first subdivision. In the first subdivision, the subdivision generator generates vertices D, E, and F, and connection information indicating their connectivity. The mesh generated by the subdivision generator is also called LoD1 or the first LoD.
[0284] Vertex D of the mesh after the first subdivision is a vertex generated by the subdivision based on vertices A and B. Similarly, vertex E is a vertex generated by the subdivision based on vertices B and C. Vertex F is a vertex generated by the subdivision based on vertices A and C.
[0285] For example, vertex D could be the midpoint of the line segment AB (or side AB) connecting vertices A and B, which were the source of its generation. Similarly, vertex E could be the midpoint of line segment AC, and vertex F could be the midpoint of line segment BC.
[0286] Figure 37(c) shows the mesh generated by the second subdivision, in other words, the mesh after the second subdivision. In the second subdivision, the subdivision generator generates vertices G, H, I, J, K, L, M, N, and O, and connection information indicating their connectivity. The mesh generated by the subdivision generator is also called LoD2 or the second LoD.
[0287] Vertex G of the mesh after the second subdivision is a vertex generated by the subdivision based on vertices A and D. Similarly, vertex H is a vertex generated by the subdivision based on vertices A and E. Vertex I is a vertex generated by the subdivision based on vertices B and D. Vertex J is a vertex generated by the subdivision based on vertices D and F. Vertex K is a vertex generated by the subdivision based on vertices E and F. Vertex L is a vertex generated by the subdivision based on vertices C and E. Vertex M is a vertex generated by the subdivision based on vertices B and F. Vertex N is a vertex generated by the subdivision based on vertices C and F. Vertex O is a vertex generated by the subdivision based on vertices D and E.
[0288] For example, vertex G could be the midpoint of the line segment AD (or edge AD) connecting vertices A and D, which were the source of its generation. Similarly, vertex H could be the midpoint of line segment AE. Vertex I could be the midpoint of line segment BD. Vertex J could be the midpoint of line segment DF. Vertex K could be the midpoint of line segment EF. Vertex L could be the midpoint of line segment CE. Vertex M could be the midpoint of line segment BF. Vertex N could be the midpoint of line segment CF. Vertex O could be the midpoint of line segment DE.
[0289] In the following sections, the displacement of the vertices will be explained with reference to Figures 38 and 39. The displacement of the vertices is performed by the reconstructor 2209.
[0290] Figure 38 is an explanatory diagram showing an example of vertex displacement after subdivision. Figure 39 is an explanatory diagram showing an example of vertices in the original mesh.
[0291] The base mesh shown in Figure 38(a) includes vertices A, B, C, and Z, and connectivity information indicating their connectivity.
[0292] Figure 38(b) shows the mesh generated by the first subdivision, in other words, the mesh after the first subdivision (i.e., the first LoD). In the first subdivision, the subdivision generator generates vertices S, T, U, X, or Y and connection information indicating their connectivity. Vertices S, T, U, X, or Y are the same as vertices D, E, and F shown in Figure 37(b).
[0293] Figure 38(c) shows the mesh generated by the second subdivision, in other words, the mesh after the second subdivision (i.e., the second LoD). In the second subdivision, the subdivision tool generates vertices D, E, F, G, and H, and connection information indicating their connectivity. Vertices D, E, F, G, and H are the same as vertices G, H, I, J, K, L, M, N, or O shown in Figure 37(c).
[0294] Figure 38(d) shows the mesh including the vertices after displacement following subdivision. The vertices A, B, C, D, E, F, G, H, S, T, U, X, Y, and Z shown in Figure 38(d) are located at positions that have been displaced using displacement information from the positions of those vertices shown in Figure 38(c).
[0295] The original mesh shown in Figure 39 is an example of the mesh input to the encoding device 100, that is, the mesh before encoding.
[0296] The mesh shown in Figure 38 has a shape similar to the original mesh shown in Figure 39. The displacement information is generated by the displacement vector calculator 1207 of the encoding device 100 as information indicating the displacement from the vertices of the base mesh to the vertices of the original mesh. By reconstructing the mesh using the displacement information generated in this way, a mesh with a shape similar to the original mesh is generated.
[0297] The decoding device 200 can output the mesh shown in Figure 38(d).
[0298] Next, we will explain the division of the mesh into submeshes with reference to Figures 40 and 41.
[0299] A mesh can be divided into multiple smaller parts, and each divided part can be encoded. When dividing a mesh, the vertices of the mesh are divided in such a way that the coordinates and connectivity of the vertices included in each part can be encoded independently.
[0300] Figure 40 is an explanatory diagram showing an example of a mesh. Figure 41 is an explanatory diagram showing an example of dividing a mesh into submeshes.
[0301] The mesh shown in Figure 40 is the original mesh, and is sometimes called the full mesh in contrast to the submesh.
[0302] Figure 41 shows how the full mesh shown in Figure 40 is divided into two submeshes. For vertices A, B, and C of the full mesh (see Figure 40), vertex A is duplicated into vertices A1 and A2, vertex B is duplicated into vertices B1 and B2, and vertex C is duplicated into vertices C1 and C2, thereby creating two submeshes (i.e., the first submesh and the second submesh) from the full mesh. The first submesh and the second submesh are meshes that can be decoded independently.
[0303] In the following sections, the packing of displacement information into image frames will be explained with reference to Figures 42, 43, and 44.
[0304] Figures 42, 43, and 44 are explanatory diagrams illustrating examples of packing displacement information into image frames. Note that image frames can also be referred to as video frames.
[0305] Vertex displacement data is encoded as image frame data by mapping it to each component of an image frame in YUV format (i.e., the Y component (Y Plane), U component (U Plane), and V component (V Plane) respectively). This case is explained below as an example. Alternatively, vertex displacement data may be encoded as image frame data by mapping it to each component of an image frame in RGB format (the R component, G component, and B component, respectively).
[0306] The decoding device 200 can use an image coding module to extract displacement data. The displacement data may be in the form of X, Y, or Z components in a global coordinate system (e.g., a Cartesian coordinate system), or normal, tangent, or both tangent components in a local coordinate system. Methods for mapping displacement data to an image frame include the following:
[0307] For example, in the first method, displacement data is arranged in scan order within the image frame. An example of packing displacement data in this case is shown in Figure 42. The displacement data is directly mapped to the image frame according to a predefined scan order.
[0308] Note that since the height and width of the image frame are fixed, the displacement data may not fit perfectly within the frame. In such cases, the remaining portion of the image frame is padded with padding data (also called padded data) (see Figure 42).
[0309] For example, in the second method, displacement data is separated into multiple Lines of Data (LDs) and mapped to the Y, U, and V components of the image frame. An example of packing the displacement data in this case is shown in Figure 43. Here, the displacement data for the image frame of the next LD starts immediately after the displacement data for the previous LD ends. Similar to the first method, if the displacement data does not fit perfectly into the image frame, padding is applied to the end of the image frame (see Figure 43).
[0310] For example, in the third method, the displacement data corresponding to the LoD is mapped to the Y, U, and V components of the image frame in a different manner than in the second method. An example of the packing of the displacement data in this case is shown in Figure 44. In this way, each LoD can be decoded independently. In the third method, intermediate padding is performed on the displacement data of each LoD, and CTU alignment is performed together with padding at the end of the video frame (see Figure 44).
[0311] Figure 45 is a block diagram showing a detailed configuration example of the decoding device 200 according to this embodiment. Specifically, Figure 45 shows an example of the configuration of the geometry coordinate decoder included in the decoding device 200.
[0312] In this example, the decoding device 200 includes a frame header decoder 631, a vertex geometry coordinate predictor 632, a vertex geometry coordinate difference decoder 633, and a reconstructor 634.
[0313] The frame header decoder 631 reads the bitstream, decodes the frame header in the bitstream, and decides whether to intra-decode (intra-predict) or inter-decode (inter-predict) the frame data.
[0314] If inter-decoding is selected, the frame data contained in the bitstream is output to the vertex geometry coordinate predictor 632.
[0315] The vertex geometry coordinate predictor 632 outputs prediction information to the reconstructor 634. An example of prediction information is a motion vector.
[0316] The reconstructor 634 outputs the three-dimensional coordinates of the vertices (vertex geometry coordinates) using prediction information along with the vertex coordinates from previously decoded frames.
[0317] On the other hand, if intra-decoding is selected, the frame data contained in the bitstream is output to the vertex geometry coordinate difference decoder 633.
[0318] The vertex geometry coordinate difference decoder 633 decodes the frame data, which has been encoded as the difference between the coordinates of the vertices contained in the frame, in order to generate vertex coordinates. Only one of the vertex geometry coordinates from the vertex geometry coordinate difference decoder 633 or the reconstructor 634 is used to generate the decoded three-dimensional mesh frame.
[0319] Figure 46 is an explanatory diagram showing the coordinates of vertices in a three-dimensional mesh according to this embodiment. Specifically, Figure 46 shows an example in which the entire frame of a three-dimensional mesh frame is decoded using the actual vertex coordinates (positions) contained in the bitstream.
[0320] The coordinates of vertex A in the three-dimensional mesh frame at time (t) are decoded as (6, 8, 9) using the Cartesian coordinate system (x, y, z), as shown in Figure 46(a). Similarly, the coordinates of vertex B are decoded as (10, 6, 7), and the coordinates of vertex C are decoded as (14, 8, 9). The same applies to vertices D through G.
[0321] Figure 47 is an explanatory diagram showing the prediction information according to this embodiment. Specifically, Figure 47 shows another example in which the entire frame of the three-dimensional mesh frame at time (t) is decoded using the frame at time (t-1) (past frame) and the prediction information contained in the bitstream.
[0322] The coordinates of vertex A in the frame to be decoded (the current frame), which are (6, 8, 9), are decoded by adding the coordinates of vertex A in the past frame, which are (4, 7, 8), and the value for vertex A indicated by the prediction information, which is (2, 1, 1). Similarly, the coordinates of vertex B in the current frame, which are (10, 6, 7), are decoded by adding the coordinates of vertex B in the past frame, which are (8, 6, 7), and the value for vertex B indicated by the prediction information, which is (2, 0, 0).
[0323] Hereafter, an example of the configuration of the encoding device according to this embodiment will be described.
[0324] Figure 48 is a block diagram showing an example configuration of the encoding device according to this embodiment.
[0325] The encoding device shown in Figure 48 comprises a decimator 4801, a subdivider 4802, a displacement vector calculator 4803, a wavelet converter 4804, an interpreter 4805, a quantizer 4806, an image packer 4807, a video encoder 4808, an inverse quantizer 4811, a reconstructor 4812, and a reference buffer 4813.
[0326] The decimator 4801 acquires the mesh frame (corresponding to the original 3D mesh frame, also called the original mesh frame or original mesh) input to the encoding device, and generates a base mesh frame (also called the base mesh) by performing a decimation process (in other words, a thinning process) on the acquired mesh frame. Decimation is a process that deletes (in other words, thins out) some of the vertices included in the original mesh. Decimation may include a process that changes the position of at least some of the vertices included in the original mesh, or a process that changes the connectivity of at least some of the vertices included in the original mesh. Decimation is also simply called decimation.
[0327] The base mesh generated by the decimation process has fewer vertices than the original mesh. The vertices of the base mesh may be located in different positions than the vertices of the original mesh. Also, the connectivity of the vertices of the base mesh may differ from that of the vertices of the original mesh. The decimator 4801 provides the generated base mesh frame to the sub-divider 4802.
[0328] The sub-divider 4802 performs sub-division processing on the base mesh frame generated by the decimator 4801. Sub-division processing can be a process of subdividing the base mesh frame into sub-subdivisions. The sub-divider 4802 provides the subdivided base mesh frame to the displacement vector calculator 4803.
[0329] Specifically, the sub-divider 4802 can subdivide a mesh frame by generating a new vertex between two connected vertices within the mesh frame. By repeating the generation of new vertices, the number of vertices in the mesh frame can be set to a predetermined number. Through repeated subdivision across the entire mesh frame (in other words, multiple executions of subdivision), multiple Level of Detail (LoD) hierarchies are generated.
[0330] The displacement vector calculator 4803 obtains the original mesh frame acquired by the encoding device and also obtains the sub-subdivided base mesh frame from the sub-subdivider 4802. The displacement vector calculator 4803 calculates a displacement vector from the vertices of the base mesh frame and the vertices generated by the sub-subdivision of the base mesh frame, pointing to the corresponding vertices of the original mesh frame. The displacement vector calculator 4803 provides the calculated displacement vector to the wavelet converter 4804.
[0331] The wavelet transformer 4804 obtains transformation coefficients (also called wavelet coefficients) by applying a wavelet transform to the displacement vector calculated by the displacement vector calculator 4803. The wavelet transformer 4804 provides the obtained wavelet coefficients to the interpreter 4805. In the wavelet transform, the wavelet transformer 4804 can calculate wavelet coefficients representing various components from low-frequency to high-frequency components by assigning vertices to multiple LoD hierarchies and applying, for example, a lifting transform to the displacement vectors of those vertices.
[0332] The interpreter 4805 calculates the predicted residual of the wavelet coefficients of the displacement vector of the frame to be encoded using interpretation. Specifically, the interpreter 4805 calculates the predicted residual of the wavelet coefficients of the displacement vector of the frame to be encoded by interpreting the wavelet coefficients of the displacement vector of the frame to be encoded using the wavelet coefficients of the displacement vector of the encoded frame (also called the reference frame) stored in the reference buffer 4813.
[0333] The quantizer 4806 quantizes the prediction residuals of the wavelet coefficients calculated by the interpreter 4805. The quantizer 4806 can quantize the prediction residuals of the wavelet coefficients for each LoD level. The quantizer 4806 provides the quantized prediction residuals to the image packer 4807 and the inverse quantizer 4811.
[0334] The image packer 4807 generates an image containing the predicted residuals quantized by the quantizer 4806. The image packer 4807 can generate the above image by mapping the predicted residuals quantized by the quantizer 4806 to pixels of a two-dimensional image format. The image packer 4807 provides the generated image to the video encoder 4808. In the process of mapping the quantized predicted residuals to pixels of a two-dimensional image format, mapping information representing the assignment of the quantized predicted residuals to pixels of a two-dimensional image format may be used.
[0335] The video encoder 4808 encodes the image generated by the image packer 4807 into a bitstream (also called a displacement bitstream) (in other words, it generates a displacement bitstream). The video encoder 4808 outputs the displacement bitstream. The displacement bitstream may be a bitstream that contains displacement information in image format. The image format may be, for example, a format that contains two chroma information and one luma information. The video encoder 4808 can utilize a general-purpose module that has the function of converting an image to a bitstream. By utilizing a highly reliable general-purpose module as the video encoder 4808, the above function can be executed more reliably.
[0336] The inverse quantizer 4811 generates predicted residuals of wavelet coefficients by inverse quantizing the predicted residuals quantized by the quantizer 4806. Specifically, the inverse quantizer 4811 can inverse quantize the predicted residuals quantized by the quantizer 4806 for each LoD level, thereby generating predicted residuals. The inverse quantizer 4811 provides the generated predicted residuals of wavelet coefficients to the reconstructor 4812.
[0337] The reconstructor 4812 reconstructs (also called reconstructing) the wavelet coefficients from the predicted residuals of the wavelet coefficients provided by the inverse quantizer 4811 and the reference frame stored in the reference buffer 4813. The reconstructor 4812 stores the reconstructed wavelet coefficients in the reference buffer 4813.
[0338] The reference buffer 4813 is a memory device that stores, for example, the wavelet coefficients of a reference frame. The wavelet coefficients stored in the reference buffer 4813 can be used for interprediction by the interpreter 4805.
[0339] Figure 49 is a flowchart showing a specific example of the encoding process according to this embodiment.
[0340] In step S4901, the inter predictor 4805 calculates the sum of the transformation coefficients of the displacement vector within the frame (also called sum_nointer).
[0341] In step S4902, the interpreter 4805 calculates the sum of prediction residuals within the frame (also called sum_inter) when the interpretation is applied to the transformation coefficients of the displacement vector.
[0342] The prediction residual for interpretation can be calculated by subtracting the transformation coefficient of the displacement vector of the three-dimensional point (also called the reference point) corresponding to the three-dimensional point to be encoded in the reference frame from the transformation coefficient of the displacement vector of the three-dimensional point to be encoded in the frame to be encoded.
[0343] In step S4903, the inter predictor 4805 determines whether sum_inter calculated in step S4902 is smaller than sum_nointer calculated in step S4901. If it determines that sum_inter is smaller than sum_nointer (Yes in step S4903), the process proceeds to step S4904; otherwise (No in step S4903), the process proceeds to step S4911.
[0344] In step S4904, the interpreter 4805 decides to encode the transformation coefficients of the displacement vectors in the frame using interpretation, and outputs the prediction residual to the quantizer 4806.
[0345] In step S4905, information indicating that the transformation coefficients of the displacement vectors within the frame have been encoded using interprediction is added to the header (for example, the stream header). For example, by setting the information disp_frame_inter_mode, which is included in the header and indicates that the transformation coefficients of the displacement vectors have been encoded using interprediction, to 1, it can be shown that the transformation coefficients of the displacement vectors within the frame have been encoded using interprediction.
[0346] In step S4911, it is decided to encode the transformation coefficients of the displacement vectors within the frame without using interprediction, and the transformation coefficients are output to the quantizer 4806.
[0347] In step S4912, information is added to the header indicating that the transformation coefficients of the displacement vectors within the frame were encoded without inter-prediction. For example, by setting the information disp_frame_inter_mode, which is included in the header and indicates that the transformation coefficients of the displacement vectors were encoded using inter-prediction, to 0, it can be shown that the transformation coefficients of the displacement vectors within the frame were encoded without inter-prediction.
[0348] Figure 50 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.
[0349] The decoding device comprises a video decoder 5001, an image umpacker 5002, an inverse quantizer 5003, a reconstructor 5004, an inverse wavelet converter 5005, a reconstructor 5006, and a reference buffer 5011.
[0350] The video decoder 5001 acquires a displacement bitstream and decodes the acquired displacement bitstream into an image. The image may be an image contained in which quantized wavelet coefficients are mapped to pixels of a two-dimensional image format. The video decoder 5001 provides the image to the image unpacker 5002. The video decoder 5001 can utilize a general-purpose module that has the function of converting a bitstream to an image. By utilizing a highly reliable general-purpose module as the video decoder 5001, the above function can be performed more reliably.
[0351] The image umpacker 5002 extracts quantized wavelet coefficients from the image provided by the video decoder 5001. A mapping that represents the assignment of the quantized wavelet coefficients to pixels in a two-dimensional image format may be used for the process of extracting the quantized wavelet coefficients from the image. The image umpacker 5002 provides the quantized wavelet coefficients extracted from the image to the inverse quantizer 5003.
[0352] The inverse quantizer 5003 generates predicted residuals of wavelet coefficients by inverse quantizing the predicted residuals of the quantized wavelet coefficients provided by the image umpacker 5002. Specifically, the inverse quantizer 5003 can generate predicted residuals of wavelet coefficients by inverse quantizing the predicted residuals of the quantized wavelet coefficients for each LoD hierarchy.
[0353] The reconstructor 5004 reconstructs (also called reconstructing) the transformation coefficients of the frame to be decoded from the predicted residual of the transformation coefficients of the frame to be decoded and the transformation coefficients of the reference frame. The reconstructor 5004 provides the reconstructed transformation coefficients to the inverse wavelet converter 5005 and the reference buffer 5011.
[0354] The inverse wavelet transformer 5005 generates a displacement vector (corresponding to the decoded displacement vector) by applying an inverse wavelet transform to the wavelet coefficients provided by the reconstructor 5004. The inverse wavelet transform is equivalent to the inverse transform of the wavelet transform performed by the wavelet transformer 4804. Specifically, the inverse wavelet transformer 5005 can calculate the vertex displacement vector by applying an inverse lifting transform to the wavelet coefficients in the inverse wavelet transform. The inverse wavelet transformer 5005 provides the generated decoded displacement vector to the reconstructor 5006.
[0355] The reconstructor 5006 reconstructs the mesh (corresponding to the decoded mesh) using the decoded displacement vectors and the decoded base mesh provided by the inverse wavelet transformer 5005. The reconstructor 5006 outputs the reconstructed decoded mesh.
[0356] The reference buffer 5011 is a memory device that stores, for example, the wavelet coefficients of a reference frame. The wavelet coefficients stored in the reference buffer 5011 can be used by the reconstructor 5004 to reconstruct the conversion coefficients of the frame to be decoded.
[0357] Figure 51 is a flowchart showing a specific example of the decoding process according to this embodiment.
[0358] In step S5101, the reconstructor 5004 determines whether information indicating that the transformation coefficients of the displacement vectors in the frame have been encoded using interprediction is attached to the header. This information is, for example, information indicating that disp_frame_inter_mode is 1. If it is determined that the above information is attached to the header (Yes in step S5101), the process proceeds to step S5102; otherwise (No in step S5101), the process proceeds to step S5111.
[0359] In step S5102, the reconstructor 5004 determines that the transformation coefficients of the displacement vectors in the frame have been encoded using interprediction, adds the transformation coefficients of the reference frame to the prediction residuals provided by the inverse quantizer 5003, decodes the transformation coefficients, and outputs them.
[0360] In step S5111, the reconstructor 5004 determines that the transformation coefficients of the displacement vectors in the frame have been encoded without interprediction, and outputs the transformation coefficients provided by the inverse quantizer 5003.
[0361] Figure 52 is a flowchart showing a specific example of the encoding process according to this embodiment.
[0362] The encoding process shown in Figure 52 involves determining the unit of inter-prediction of the displacement vector for each LoD, and adding information disp_lod_inter_mode to the header indicating whether inter-prediction was applied for each LoD. This allows for switching inter-prediction on or off for each LoD, improving encoding efficiency. For example, inter-prediction is often accurate for fine movements and less accurate for large movements. In such cases, enabling inter-prediction in lower LoD layers where high-frequency components are concentrated and disabling it in higher LoD layers where low-frequency components are concentrated can contribute to improving encoding efficiency.
[0363] In step S5201, the inter predictor 4805 initiates loop A, which repeatedly executes the processes described in steps S5202 to S5206 and steps S5211 to S5212. Loop A focuses on each of the one or more LoDs, executes processing for the focused LoD, and ultimately controls the system to process all LoDs. The focused LoD is also called the "focused LoD." Loop A may also be called the LoD loop.
[0364] In step S5202, the interpreter 4805 calculates the sum of the transformation coefficients of the displacement vector within the focus LoD (also called sum_lod_nointer).
[0365] In step S5203, the interpreter 4805 calculates the sum of prediction residuals within the focus LoD (also called sum_lod_inter) when the interpretation is applied to the transformation coefficients of the displacement vector.
[0366] In step S5204, the interface predictor 4805 determines whether sum_lod_inter calculated in step S5203 is smaller than sum_lod_nointer calculated in step S5202. If it determines that sum_lod_inter is smaller than sum_lod_nointer (Yes in step S5204), the process proceeds to step S5205; otherwise (No in step S5204), the process proceeds to step S5211.
[0367] In step S5205, the interpreter 4805 decides to encode the transformation coefficients of the displacement vector in the LoD of interest using interpretation, and outputs the prediction residual to the quantizer 4806.
[0368] In step S5206, the inter predictor 4805 adds information to the header (e.g., the stream header) indicating that the transformation coefficients of the displacement vectors in the LoD of interest have been encoded using inter prediction. For example, by setting the information disp_lod_inter_mode(i) included in the header, which indicates that the transformation coefficients of the displacement vectors have been encoded using inter prediction, to 1, it can be shown that the transformation coefficients of the displacement vectors in the LoD of interest have been encoded using inter prediction. Note that i is the ordinal number of the LoD of interest, indicating which LoD it is. The same applies hereafter.
[0369] In step S5211, the interpreter 4805 decides to encode the transformation coefficients of the displacement vector in the LoD of interest without using interpretation, and outputs the transformation coefficients to the quantizer 4806.
[0370] In step S5212, the inter predictor 4805 adds information to the header indicating that the transformation coefficients of the displacement vectors in the LoD of interest were encoded without using inter prediction. For example, by setting the information disp_lod_inter_mode(i) included in the header, which indicates that the transformation coefficients of the displacement vectors were encoded using inter prediction, to 0, it can be indicated that the transformation coefficients of the displacement vectors in the LoD of interest were encoded without using inter prediction.
[0371] In step S5207, the inter predictor 4805 performs the termination process for loop A. Specifically, the inter predictor 4805 determines whether the processes in steps S5202 to S5206 and steps S5211 to S5212 (however, only one of the processes in steps S5205 and S5206, and the processes in steps S5211 and S5212, depending on the result of the determination in step S5204) have been executed for all LoDs. If they have not been executed, the predictor 4805 controls the process to focus on the LoDs that have not yet been executed.
[0372] Figure 53 is a flowchart showing a specific example of the decoding process according to this embodiment.
[0373] The decoding process shown in Figure 53 involves determining the unit of inter prediction of the displacement vector for each LoD, and adding information disp_lod_inter_mode to the header indicating whether inter prediction was applied for each LoD, before decoding the encoded bitstream. This allows switching inter prediction on or off for each LoD, which can contribute to properly decoding a bitstream with improved encoding efficiency.
[0374] In step S5301, the reconfigurator 5004 initiates loop A, which repeatedly executes the processes described in steps S5302 to S5203 and step S5311. Loop A focuses on each of the one or more LoDs, executes processing on the focused LoD, and ultimately controls the system to ensure that processing is performed on all LoDs. The focused LoD is also called the "focused LoD." Loop A may also be called the LoD loop.
[0375] In step S5302, the reconstructor 5004 determines whether information indicating that the transformation coefficients of the displacement vectors in the LoD of interest have been encoded using inter-prediction is added to the header. This information is, for example, information indicating that disp_lod_inter_mode(i) is 1. If it is determined that this information is added to the header (Yes in step S5302), the process proceeds to step S5303; otherwise (No in step S5302), the process proceeds to step S5311.
[0376] In step S5303, the reconstructor 5004 determines that the transformation coefficients of the displacement vector in the LoD of interest have been encoded using interprediction, and decodes the transformation coefficients by adding the transformation coefficients of the reference frame to the prediction residual provided by the inverse quantizer 5003, and outputs them.
[0377] In step S5311, the reconstructor 5004 determines that the transformation coefficients of the displacement vector in the LoD of interest have been encoded without interprediction, and outputs the transformation coefficients provided by the inverse quantizer.
[0378] In step S5304, the reconfigurator 5004 performs the termination process for loop A. Specifically, the reconfigurator 5004 determines whether the processes in steps S5302 to S5203 and step S5311 (however, only one of the processes in step S5203 and step S5311, depending on the result of the determination in step S5302) have been executed for all LoDs. If they have not been executed, the reconfigurator controls the process to focus on the LoDs that have not yet been executed.
[0379] Figure 54 is an explanatory diagram showing an example of syntax according to this embodiment.
[0380] The syntax example shown in Figure 54 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.
[0381] The syntax shown in Figure 54 includes `displacement vector_header`, which includes `disp_frame_inter_mode`.
[0382] `disp_frame_inter_mode` indicates whether or not the displacement vectors within the frame were encoded using inter-prediction. For example, a value of 1 may indicate that the displacement vectors within the frame were encoded using inter-prediction, while a value of 0 may indicate that the displacement vectors within the frame were encoded without inter-prediction. This allows the decoder to determine whether or not the displacement vectors within the frame were encoded using inter-prediction and to decode the bitstream appropriately.
[0383] Note that the unit to which `disp_frame_inter_mode` is added is not limited to the frame unit. For example, `disp_frame_inter_mode` may be added at the submesh unit level. This allows for improved coding efficiency by switching the use of inter-prediction on a submesh basis. For example, coding efficiency can be improved by using inter-prediction for submeshes with small movement, such as background objects, and not using inter-prediction for submeshes with large movement, such as foreground objects.
[0384] Furthermore, disp_frame_inter_mode may be added on a sequence-by-sequence basis. This reduces the amount of code in the header. For example, when encoding a sequence with small overall motion, inter prediction can be used for the entire sequence, while when encoding a sequence with large overall motion, inter prediction can not be used for the entire sequence, thereby reducing the amount of code in the header and improving encoding efficiency. Note that when inter prediction is used for the entire sequence, disp_frame_inter_mode and disp_lod_inter_mode do not need to be added to the header. This also reduces the amount of code in the header.
[0385] Figure 55 is an explanatory diagram showing an example of syntax according to this embodiment.
[0386] The syntax shown in Figure 55 includes `displacement vector_header`, which includes `disp_lod_inter_mode[i]`.
[0387] disp_lod_inter_mode[i] is information indicating whether or not the displacement vector belonging to the i-th LoD was encoded using inter-prediction when generating an LoD and encoding the displacement vector. For example, a value of 1 may indicate that the displacement vector of the three-dimensional point belonging to the i-th LoD was encoded using inter-prediction, and a value of 0 may indicate that the displacement vector of the three-dimensional point belonging to the i-th LoD was encoded without inter-prediction. This allows the decoder to determine whether or not the displacement vector of the three-dimensional point belonging to the i-th LoD was encoded using inter-prediction, and to appropriately decode the bitstream.
[0388] Figure 56 is an explanatory diagram showing an example of syntax according to this embodiment.
[0389] The syntax shown in Figure 56 includes a displacement vector_header. The displacement vector_header includes disp_frame_inter_mode, disp_lod_inter_mode_present, and disp_lod_inter_mode[i].
[0390] disp_frame_inter_mode is the same as disp_frame_inter_mode shown in Figure 54.
[0391] `disp_lod_inter_mode_present` indicates whether `disp_lod_inter_mode` is included in the header. For example, a value of 1 indicates that `disp_lod_inter_mode` is included in the header, and a value of 0 indicates that `disp_lod_inter_mode` is not included in the header. Note that if `disp_frame_inter_mode=1`, inter-frame prediction is performed on a frame-by-frame basis, so the value of `disp_lod_inter_mode_present` is estimated to be 0, and `disp_lod_inter_mode` does not need to be added to the header. This reduces the amount of code in the header.
[0392] disp_lod_inter_mode[i] is the same as disp_lod_inter_mode[i] shown in Figure 55.
[0393] In addition, if disp_frame_inter_mode=0, the value of disp_lod_inter_mode_present may be omitted, and disp_lod_inter_mode[i] may be included in the header regardless of the value of disp_lod_inter_mode_present.
[0394] Furthermore, disp_frame_inter_mode, disp_lod_inter_mode, or disp_lod_inter_mode_present may be entropy encoded and appended to the header. For example, each value may be binarized and then arithmetic encoded. Alternatively, to reduce processing load, it may be encoded using a fixed length.
[0395] Figure 57 is a block diagram showing an example of the configuration of the encoding device according to this embodiment.
[0396] The encoding device shown in Figure 57 comprises a decimator 5701, a subdivider 5702, a displacement vector calculator 5703, a wavelet converter 5704, an LoD-based inter predictor 5705, a quantizer 5706, a switch 5707, an image packer 5708, a video encoder 5709, an arithmetic encoder 5710, an inverse quantizer 5711, a reconstructor 5712, and a reference buffer 5713.
[0397] The decimator 5701, the sub-divider 5702, and the displacement vector calculator 5703 are the same as the decimator 4801, sub-divider 4802, and displacement vector calculator 4803 shown in Figure 48, respectively.
[0398] The wavelet transformer 5704 obtains transformation coefficients (also called wavelet coefficients) by applying a wavelet transform to the displacement vector calculated by the displacement vector calculator 5703. The wavelet transformer 5704 provides the obtained wavelet coefficients to the LoD-based interpreter 5705. In the wavelet transform, the wavelet transformer 5704 can calculate wavelet coefficients representing various components from low-frequency to high-frequency components by assigning vertices to multiple LoD hierarchy levels and applying, for example, a lifting transform to the displacement vectors of those vertices.
[0399] The LoD-based interpreter 5705 outputs the predicted residual of the wavelet coefficients of the displacement vector of the frame to be encoded using interpretation for each LoD. Specifically, for each LoD, the LoD-based interpreter 5705 outputs the predicted residual of the wavelet coefficients of the displacement vector of the frame to be encoded by interpreting the wavelet coefficients of the displacement vector of the frame to be encoded using the wavelet coefficients of the displacement vector of the encoded frame (also called a reference frame) stored in the reference buffer 5713. The LoD-based interpreter 5705 can determine and switch for each LoD whether or not to encode the transformation coefficients of the displacement vector using interpretation.
[0400] The quantizer 5706 quantizes the predicted residuals of the wavelet coefficients calculated by the LoD-based interpreter 5705. The quantizer 5706 quantizes the predicted residuals of the wavelet coefficients for each LoD level. The quantizer 5706 provides the quantized predicted residuals to the image packer 5708 or the arithmetic encoder 5710 via the switch 5707, and also to the inverse quantizer 5711.
[0401] Switch 5707 is a switch that determines whether the predicted residuals quantized by the quantizer 5706 are provided to the image packer 5708 or to the arithmetic encoder 5710.
[0402] The image packer 5708 and the video encoder 5709 are the same as the image packer 4807 and video encoder 4808 shown in Figure 48, respectively.
[0403] The arithmetic encoder 5710 encodes the prediction residuals quantized by the quantizer 5706 into a bitstream (also called a displacement bitstream) using arithmetic coding (in other words, it generates a displacement bitstream). The arithmetic encoder 5710 outputs the displacement bitstream.
[0404] The inverse quantizer 5711 generates predicted residuals of wavelet coefficients by inverse quantizing the predicted residuals quantized by the quantizer 5706. Specifically, the inverse quantizer 5711 generates predicted residuals by inverse quantizing the predicted residuals quantized for each LoD level by the quantizer 5706, for each LoD level. The inverse quantizer 5711 provides the generated predicted residuals of wavelet coefficients to the reconstructor 5712.
[0405] The reconstructor 5712 reconstructs (also called reconstructing) the wavelet coefficients from the predicted residuals of the wavelet coefficients provided by the inverse quantizer 5711 and the reference frame stored in the reference buffer 5713. The reconstructor 5712 stores the reconstructed wavelet coefficients in the reference buffer 5713.
[0406] The reference buffer 5713 is a memory device that stores, for example, the wavelet coefficients of a reference frame. The wavelet coefficients stored in the reference buffer 5713 can be used for interpretation by the LoD-based interpreter 5705.
[0407] The encoding device shown in Figure 57 can switch between encoding the transformation coefficients of the displacement vector using the video encoder 5709 or encoding them using arithmetic encoding with the arithmetic encoder 5710. This allows the encoding device to efficiently encode the displacement vector using arithmetic encoding even when the video encoder 5709 cannot be used.
[0408] Even when using arithmetic coding, it is possible to encode the displacement vector transformation coefficients by switching whether or not to use interpretation for each LoD. Generally, arithmetic coding becomes more efficient as the change in the input value decreases, so coding efficiency can be improved by suppressing the change in the predicted residual value of the displacement vector transformation coefficients by interpretation for each LoD.
[0409] Furthermore, information indicating whether the transformation coefficients of the displacement vector were encoded using the video encoder 5709 or by arithmetic encoding using the arithmetic encoder 5710 may be added to the header. This allows the decoding device to appropriately switch the decoding method by referring to the header information.
[0410] In Figure 57, an example is shown where the transformation coefficients of the displacement vector are either encoded using the video encoder 5709 or encoded using arithmetic encoding with the arithmetic encoder 5710. However, a configuration that always uses arithmetic encoding is also possible. With this configuration, it is possible to improve the encoding efficiency when arithmetic encoding the transformation coefficients of the displacement vector compared to the case where interpretation is used at all LoD levels.
[0411] The decoding device shown in Figure 58 comprises a video decoder 5801, an image umpacker 5802, an arithmetic decoder 5803, a switch 5804, an inverse quantizer 5805, a reconstructor 5806, an inverse wavelet converter 5807, a reconstructor 5808, and a reference buffer 5811.
[0412] The video decoder 5801 and the image ampucker 5802 are the same as the video decoder 5001 and the image ampucker 5002 shown in Figure 50, respectively.
[0413] The arithmetic decoder 5803 acquires the displacement bitstream and arithmetically decodes the predicted residuals contained in the acquired displacement bitstream. The arithmetic decoder 5803 may also decode various header information.
[0414] Switch 5804 is a switch that toggles whether to provide the inverse quantizer 5805 with the predicted residuals provided by the image umpacker 5802, or with the inverse quantizer 5805 with the predicted residuals provided by the arithmetic decoder 5803.
[0415] The inverse quantizer 5805 generates predicted residuals of wavelet coefficients by inverse quantizing the predicted residuals of quantized wavelet coefficients provided from the image umpacker 5802 or arithmetic decoder 5803 via switch 5804. Specifically, the inverse quantizer 5805 generates predicted residuals of wavelet coefficients by inverse quantizing the predicted residuals of quantized wavelet coefficients for each LoD hierarchy.
[0416] The reconstructor 5806 reconstructs (also called reconstructs) the transformation coefficients of the frame to be decoded from the predicted residual of the transformation coefficients of the frame to be decoded and the transformation coefficients of the reference frame. The reconstructor 5806 provides the reconstructed transformation coefficients to the inverse wavelet converter 5807 and the reference buffer 5811. The reconstructor 5806 may determine from the header information whether interpretation was applied for each Line of Data (LoD) and switch the reconstruction method.
[0417] The inverse wavelet transformer 5807 generates a displacement vector (corresponding to the decoded displacement vector) by applying an inverse wavelet transform to the wavelet coefficients provided by the reconstructor 5806. The inverse wavelet transform is equivalent to the inverse transform of the wavelet transform performed by the wavelet transformer 5704. Specifically, the inverse wavelet transformer 5807 can calculate the vertex displacement vector by applying an inverse lifting transform to the wavelet coefficients in the inverse wavelet transform.
[0418] The inverse wavelet converter 5807 provides the generated decoded displacement vector to the reconstructor 5808.
[0419] The reconstructor 5808 reconstructs the mesh (corresponding to the decoded mesh) using the decoded displacement vectors and the decoded base mesh provided by the inverse wavelet transformer 5807. The reconstructor 5808 outputs the reconstructed decoded mesh.
[0420] The reference buffer 5811 is a memory device that stores, for example, the wavelet coefficients of a reference frame. The wavelet coefficients stored in the reference buffer 5811 can be used by the reconstructor 5806 to reconstruct the conversion coefficients of the frame to be decoded.
[0421] The decoding device shown in Figure 58 decodes the header information and may determine whether the encoding device encoded the transformation coefficients of the displacement vector using a video encoder (e.g., video encoder 5709) or arithmetic encoding using an arithmetic encoder (e.g., arithmetic encoder 5710), and switch the decoding method. This allows the decoding device to appropriately decode a bitstream in which the displacement vector has been efficiently encoded using arithmetic encoding even when a video encoder (e.g., video encoder 5709) cannot be used.
[0422] Figure 59 is an explanatory diagram showing the positional relationship of three-dimensional points according to this embodiment.
[0423] One possible method for encoding the displacement vector of a three-dimensional point is to calculate a predicted value for the displacement vector of a given three-dimensional point and encode the difference (prediction residual) between the original displacement vector value and the predicted value. For example, if the value of the displacement vector of a given three-dimensional point p is Ap and the predicted value is Pp, the encoding device 100 encodes the absolute difference value Diffp = |Ap - Pp|, which indicates the absolute value of the difference, and information indicating the sign of (Ap - Pp). In this case, if the predicted value Pp can be generated with high accuracy, the value of the absolute difference value Diffp will become smaller. Therefore, for example, the encoding device 100 can reduce the amount of encoding by performing entropy encoding using an encoding table (or context) where the number of generated bits decreases as the value decreases.
[0424] One possible method for the encoding device 100 to generate predicted displacement vectors is to use the displacement vectors of other three-dimensional points located around the three-dimensional point to be encoded. Here, a three-dimensional point located around a three-dimensional point refers to another three-dimensional point located within a predetermined distance (within a predetermined range) from the three-dimensional point. For example, if there is a three-dimensional point p = (x1, y1, z1) which is the three-dimensional point to be encoded, and a three-dimensional point q = (x2, y2, z2), then the Euclidean distance between three-dimensional point p and three-dimensional point q is d(p, q) = √((x1 - y1)). 2 + (x² - y²) 2 + (x³ - y³) 2 If the value is smaller than a certain threshold THd, the encoding device 100 determines that the position of the three-dimensional point q is close to the position of the three-dimensional point p, and decides to use the value of the displacement vector of the three-dimensional point q to generate a predicted value of the displacement vector of the three-dimensional point p.
[0425] Note that the distance calculation method may be different; for example, the Mahalanobis distance may be used.
[0426] Furthermore, for example, the encoding device 100 may determine that a three-dimensional point that is more than a predetermined distance from the three-dimensional point to be encoded (outside a predetermined range) will not be used for prediction. For example, if a three-dimensional point r exists and the distance d(p,r) between three-dimensional point p and three-dimensional point r is greater than or equal to a threshold THd, the encoding device 100 may determine that the three-dimensional point r will not be used for prediction. The predetermined distance can be arbitrarily determined and is not particularly limited.
[0427] The encoding device 100 may also add the threshold value THd to the bitstream header.
[0428] For example, when the encoding device 100 encodes the displacement vector of a three-dimensional point to be encoded using predicted values, it uses the displacement vectors of surrounding three-dimensional points used to generate the predicted values, or it uses a displacement vector that has already been encoded or a displacement vector that has already been decoded.
[0429] Furthermore, when the decoding device 200 decodes the displacement vector of a three-dimensional point to be decoded using predicted values, it uses the displacement vectors of surrounding three-dimensional points used to generate the predicted values, and also uses the displacement vectors that have already been decoded.
[0430] As a result, the same predicted values are generated during encoding and decoding. Therefore, the decoding device 200 can correctly decode the bitstream of three-dimensional points generated by the encoding device 100.
[0431] It has been explained that points surrounding a three-dimensional point refer to other three-dimensional points within a predetermined range from that three-dimensional point, but this is not necessarily the only definition. For example, in the case of three-dimensional point D (i.e., vertex D) shown in Figure 47, there are three-dimensional points A, B, C, E, F, and G as surrounding three-dimensional points, but surrounding three-dimensional points (in other words, adjacent points) may be selected based on one or more of the following conditions A and B. In other words, adjacent points are points selected according to certain conditions and are referenced to predict the information of the three-dimensional point to be encoded. Adjacent points may also be referred to as reference three-dimensional points, reference points, or reference vertices.
[0432] Condition A: A 3D point that is connected to the target 3D point. Condition B: A 3D point that has been encoded or decoded before the target 3D point.
[0433] For example, when three-dimensional points satisfying conditions A and B above are selected as adjacent points, and three-dimensional point D and its adjacent points are encoded or decoded in the order of three-dimensional point A, three-dimensional point C, three-dimensional point E, three-dimensional point F, three-dimensional point D, three-dimensional point B, and three-dimensional point G, then three-dimensional points A, C, E, and F may be selected as adjacent points to three-dimensional point D. Since three-dimensional points A, C, E, and F are connected to three-dimensional point D, their displacement vectors are likely to be close in value. Furthermore, since three-dimensional points A, C, E, and F are encoded or decoded before three-dimensional point D, the displacement vectors of three-dimensional points A, C, E, and F can be used to calculate the predicted value of the displacement vector of three-dimensional point D. This improves the accuracy of the predicted value of the displacement vector of three-dimensional point D, thereby improving encoding efficiency.
[0434] In addition to conditions A and B above, the number of adjacent points to a three-dimensional point may be limited to a predetermined value (NumNeiCnt) or less. For example, by setting NumNeiCnt = 3, the number of adjacent points to a three-dimensional point may be limited to three or less. This reduces the memory capacity required to store information on adjacent points to a three-dimensional point, and also reduces the processing load when calculating the predicted displacement vector. The predetermined value can be arbitrarily determined and is not particularly limited.
[0435] Furthermore, for example, the encoding device 100 may add the above-mentioned predetermined value, in other words, NumNeiCnt, which indicates the maximum number of adjacent points, to the header of the data unit and encode it, thereby adding it to the bitstream.
[0436] As a result, the decoding device 200 can properly decode a bitstream in which the maximum number of adjacent points is limited to NumNeiCnt or less by decoding the bitstream header.
[0437] Furthermore, if there are more three-dimensional points satisfying the above conditions A and B than NumNeiCnt, adjacent points may be selected in order of proximity to the three-dimensional point to be encoded or decoded. For example, if NumNeiCnt = 3, and there are four adjacent three-dimensional points to three-dimensional point D that satisfy the above conditions A and B, namely three-dimensional points A, C, E, and F, and the distance to three-dimensional point D is in the order of three-dimensional points A, C, E, and F, then three-dimensional points A, C, E, and E may be selected as adjacent points to three-dimensional point D. Three-dimensional points A, C, and E are connected to three-dimensional point D and are close to three-dimensional point D, so the values of their displacement vectors are likely to be close to the values of the displacement vector of three-dimensional point D. Furthermore, 3D points A, C, and E are encoded or decoded before 3D point D. Therefore, the displacement vectors of 3D points A, C, and E can be used to calculate the predicted value of the displacement vector of 3D point D.
[0438] This improves the accuracy of the predicted displacement vector for a three-dimensional point D. Furthermore, by limiting the number of adjacent points, it reduces the memory capacity required to store information about adjacent points of a three-dimensional point, and also reduces the processing load when calculating the predicted displacement vector.
[0439] The method for selecting adjacent points to a three-dimensional point is not limited to the above. For example, if the three-dimensional point to be encoded is a three-dimensional point Z generated by subdivision from three-dimensional points X and Y that constitute the base mesh, then three-dimensional points X and Y of the base mesh may be used as adjacent points to three-dimensional point Z. Basically, when three-dimensional point Z is generated by subdivision from three-dimensional points X and Y, three-dimensional point Z lies on the straight line connecting three-dimensional points X and three-dimensional points Y, so three-dimensional points X and Y are adjacent points to three-dimensional point Z. In addition, the displacement vector of three-dimensional point Z is likely to be close in value to the displacement vectors of three-dimensional points X and Y, and the encoding efficiency can be improved by using three-dimensional points X and Y as adjacent points and calculating the predicted value of the displacement vector of three-dimensional point Z using their displacement vectors.
[0440] The method for generating the Line of Deposition (LoD) will be explained below.
[0441] Figures 60 and 61 are explanatory diagrams showing the method for generating LoD in this embodiment.
[0442] When encoding displacement vectors of three-dimensional points, encoding devices may classify each three-dimensional point into one or more levels using the positional information of the three-dimensional points before encoding. Here, each level used for classification is called a Level of Detail (LoD). Each LoD is assigned a unique identifier (e.g., a number). For example, the 0th LoD is also called LoD0, the 1st LoD is also called LoD1, the nth LoD is also called LoDn, and the (n-1)th LoD is also called LoD(n-1).
[0443] The method for generating the Line of Data (LoD) will be explained using Figures 60 and 61. If the encoding or decoding device cannot calculate the position or distance information of a three-dimensional point within the frame to be encoded or decoded, the position or distance information of a corresponding three-dimensional point within an already encoded or decoded frame may be used. This may allow for efficient encoding by classifying the three-dimensional points to be encoded or decoded into one or more layers.
[0444] Figure 60 shows the three-dimensional points to be encoded: points a0, a1, a2, b0, b1, b2, c0, c1, and c2. Note that d(x, y) represents the distance between point x and point y.
[0445] By setting the threshold values for each layer of the Line of D to be larger for higher layers (layers closer to Line of D0), higher layers become point groups where the distance between three-dimensional points is greater (also called sparse point groups), and lower layers become point groups where the distance between three-dimensional points is smaller (also called dense point groups). Here, Line of D0 is the top layer (see Figure 61).
[0446] Point y belongs to the same LoD as point x when the distance d(x, y) from point x is greater than the threshold of the LoD to which point x belongs and less than or equal to the threshold of the higher LoD. When point x belongs to LoD0, which is the topmost layer, point y belongs to the same LoD as point x when the distance d(x, y) from point x is greater than the threshold of the LoD to which point x belongs.
[0447] First, the encoding device selects point a0 as the initial point and assigns it to LoD0. Next, the encoding device extracts point a1 whose distance from point a0 is greater than the threshold Thres_Lod[0] of LoD0 and assigns it to LoD0. Next, the encoding device extracts point a2 whose distance from point a1 is greater than the threshold Thres_Lod[0] of LoD0 and assigns it to LoD0. In this way, the encoding device constructs LoD0 such that the distance between each point within LoD0 is greater than the threshold Thres_Lod[0].
[0448] Next, the encoding device selects point b0 to which no LoD has been assigned and assigns it to LoD1. Next, the encoding device selects point b1 whose distance from point b0 is greater than the threshold Thres_Lod[1] of LoD1 and to which no LoD has been assigned, and assigns it to LoD1. Next, the encoding device selects point b2 whose distance from point b1 is greater than the threshold Thres_Lod[1] of LoD1 and to which no LoD has been assigned, and assigns it to LoD1. In this way, the encoding device constructs LoD1 such that the distance between each point within LoD1 is greater than the threshold Thres_Lod[1].
[0449] Next, the encoding device selects point c0 to which no LoD has been assigned and assigns it to LoD2. Next, the encoding device selects point c1 whose distance from point c0 is greater than the threshold Thres_Lod[2] of LoD2 and to which no LoD has been assigned, and assigns it to LoD2. Next, the encoding device selects point c2 whose distance from point c1 is greater than the threshold Thres_Lod[2] of LoD2 and to which no LoD has been assigned, and assigns it to LoD2. In this way, the encoding device constructs LoD2 such that the distance between each point within LoD2 is greater than the threshold Thres_Lod[2].
[0450] The threshold value of each LoD may be added to the header of the bitstream. For example, in the case of FIG. 60, the threshold values Thres_Lod[0], Thres_Lod[1], and Thres_Lod[2] may be added to the header of the bitstream.
[0451] Also, for the lowest layer of LoD, all the three-dimensional points for which LoD has not yet been assigned may be assigned. In this case, since the threshold value of the lowest layer of LoD is not added to the header, there is an effect of reducing the amount of code in the header. For example, in the case of FIG. 60, the threshold values Thres_Lod[0] and Thres_Lod[1] may be added to the header by the encoding device, and Thres_Lod[2] may not be added to the header, and the decoding device may estimate Thres_Lod[2] to have a value of 0.
[0452] Also, the number of layers of LoD may be added to the header. Thereby, it becomes possible for the decoding device to determine whether LoD is the lowest layer.
[0453] Note that when the number of layers of LoD is one layer, that is, when encoding the displacement vector of the three-dimensional point without generating LoD, the encoding device may omit the LoD generation process described in the above example. Alternatively, the encoding device may apply the LoD generation method described in the above example with the LoD layer set to 1. In this case, the encoding device may execute the LoD generation process assuming that all three-dimensional points belong to the same LoD. Thereby, the encoding device can reduce the processing time for generating LoD.
[0454] Note that the encoding or decoding of the displacement vector described in this embodiment may be applied in addition to the LoD generation method described above. For example, even when the LoD layer to which the three-dimensional point belongs is determined in advance, the encoding efficiency may be improved by applying the displacement vector encoding method or decoding method described in this embodiment.
[0455] The method for generating the Level of Data (LoD) is not limited to the method described above. For example, as shown in Figures 35 and 36, the LoD to which a point belongs may be determined according to the number of subdivisions performed from the base mesh. For example, if the base mesh is subdivided twice, the three-dimensional points included in the base mesh may first be assigned to LoD0, the points generated by the first subdivision from the three-dimensional points of the base mesh may be assigned to LoD1, and the points generated by the second subdivision may be assigned to LoD2. This reduces the processing time for generating the LoD.
[0456] The method for selecting the initial three-dimensional points when constructing each LoD may depend on the coding order during displacement vector coding. For example, the coding device selects the first three-dimensional point coded during displacement vector coding as the initial point a0 of LoD0, and then selects points a1 and a2 from point a0 to construct LoD0. Then, the coding device may select the three-dimensional point that was coded earliest among the three-dimensional points that do not currently belong to LoD0 as the initial point b0 of LoD1. In other words, the coding device may select the three-dimensional point that was coded earliest among the three-dimensional points that do not belong to the LoDs of the hierarchy below LoD(n-1) as the initial point n0 of LoDn. This allows the same LoD as during coding to be constructed during decoding using the same initial point selection method (specifically, selecting the three-dimensional point that was coded earliest among the three-dimensional points that do not belong to the LoDs of the hierarchy below LoD(n-1) as the initial point n0 of LoDn), and the bitstream to be decoded appropriately.
[0457] Figure 62 is an explanatory diagram showing the method for generating predicted values of the displacement vector in this embodiment.
[0458] The encoding device can generate predicted values for the displacement vectors of three-dimensional points using information from the Line of Deposition (LoD).
[0459] For example, if the encoding device encodes the three-dimensional points in LoD0 sequentially, it may generate LoD1 using the encoded and decoded displacement vectors contained in LoD0 and LoD1. In this way, the encoding device can generate predicted values of the displacement vectors of the three-dimensional points contained in LoDn using the encoded and decoded displacement vectors contained in LoDn' (where n'≦n).
[0460] Furthermore, the predicted displacement vector of a three-dimensional point can be generated by calculating the average of the displacement vectors of a certain number of three-dimensional points that are adjacent to the three-dimensional point to be encoded and have been encoded and decoded. This certain number is, for example, the number of adjacent points to the three-dimensional point to be encoded (e.g., N points). In this case, the value N is added to the bitstream header, etc.
[0461] The value N, which indicates the number of adjacent points (i.e., N points) used to calculate the predicted value, may be added to each three-dimensional point from which a predicted value is generated. This allows the encoding device to select an appropriate N adjacent points for each three-dimensional point from which a predicted value is generated, thereby improving the accuracy of the predicted value and reducing the prediction residual. Alternatively, the encoding device may add the value N to the bitstream header and fix it within the bitstream (in other words, the value N may be used as a common fixed value in encoding the three-dimensional points included in the bitstream). This eliminates the need for the encoding device to encode or decode the value N for each three-dimensional point, thus reducing the processing load. Furthermore, the encoding device may encode the value N separately for each Line of D (LoD). This may allow the encoding device to improve encoding efficiency by selecting an appropriate value N for each LoD.
[0462] Furthermore, the predicted displacement vector of a three-dimensional point may be calculated from the weighted average of N neighboring points that have been encoded and decoded. The encoding device can, for example, perform a weighted average using the distance information between the three-dimensional point to be encoded and each of the N neighboring points. This will be explained using Figure 62.
[0463] When an encoding device encodes using a different value N for each Line of Data (LoD), it may set the value of N to be larger for higher layers of the LoD and smaller for lower layers. In the higher layers of the LoD, the distance between three-dimensional points belonging to that LoD is relatively large, so by setting a large value of N, it may be possible to improve prediction accuracy by selecting and averaging a relatively large number of surrounding three-dimensional points. In the lower layers of the LoD, the distance between three-dimensional points belonging to that LoD is relatively small, so by setting a small value of N, it may be possible to perform efficient prediction while reducing the processing load of averaging.
[0464] The predicted value of point P belonging to LoDN is generated from the reconstructed point P' belonging to LoDN' (where N'≦N). Here, it is assumed that adjacent points are selected from point P' based on connectivity and distance.
[0465] Furthermore, the predicted displacement vector can be calculated from an unweighted average value. This reduces the amount of processing required.
[0466] As shown in Figure 62, point a2 is predicted from points a0 and a1. Similarly, point b2 is predicted from points a0, a1, a2, b0, and b1. The points selected as adjacent points used for prediction may change depending on the number of adjacent points N used for prediction. For example, when N=5, points a0, a1, a2, b0, and b1 may be selected as adjacent points to point b2, and when N=4, points a0, a1, a2, and b1 may be selected based on distance information.
[0467] For example, the three-dimensional points included in the base mesh are included in LoD0, the three-dimensional points generated from the three-dimensional points of the base mesh by one subdivision are included in LoD1, and the three-dimensional points generated from the three-dimensional points of the base mesh by two subdivisions are included in LoD2.
[0468] For example, if the weighted mean of adjacent points is used for prediction, the predicted value a2p for point a2 is calculated by the weighted mean of points a0 and a1 (see Equations 1 and 2). Here, A iis the value of the displacement vector of point ai.
[0469]
[0470] However,
[0471]
[0472] Also, the predicted value b2p of point b2 is calculated by the weighted average of point a0, point a1, point a2, point b0, and point b1 (see (Equation 3), (Equation 4), and (Equation 5)). Here, B i is the value of the displacement vector of point bi.
[0473]
[0474] However,
[0475]
[0476]
[0477] Note that when generating the predicted value of the displacement vector, references in the same hierarchy may not be made. This can reduce the processing amount. Also, when generating a three-dimensional point by subdivision at the intermediate position between two three-dimensional points, the weight w i may be fixed at 0.5. This can reduce the processing amount.
[0478] When the encoding device encodes the value of the displacement vector of a three-dimensional point, it may calculate the difference value (also referred to as the conversion coefficient, see (Equation 6) and (Equation 7) below) between the predicted value generated from the adjacent points of the three-dimensional point and the three-dimensional point, and perform encoding using quantization of the calculated conversion coefficient. Here, the conversion coefficient a2c is the conversion coefficient of point a2, and the conversion coefficient b2c is the conversion coefficient of point b2.
[0479]
[0480]
[0481] For example, the encoding device can perform quantization by dividing the conversion coefficient by the quantization scale. In this case, the smaller the quantization scale, the smaller the error (quantization error) that can occur due to quantization, and conversely, the larger the quantization scale, the larger the quantization error.
[0482] The quantized value a2q is obtained by quantizing the conversion coefficient a2c, and the quantized value b2q is obtained by quantizing the conversion coefficient b2r (see Equations 8 and 9 below). QS_LoD0 is the quantization scale of LoD0, and QS_LoD1 is the quantization scale of LoD1.
[0483]
[0484]
[0485] Furthermore, the encoding device may change the quantization scale value for each Level of Derivation (LoD). For example, the quantization scale can be made smaller for higher-level LoDs and larger for lower-level LoDs. Since the displacement vector values of three-dimensional points belonging to higher-level layers may be used as predicted values for the displacement vectors of three-dimensional points belonging to lower-level layers, encoding efficiency can be improved by reducing the quantization scale of higher-level layers to suppress quantization errors that may occur in higher-level layers and increase the accuracy of the predicted values. The encoding device may also add the quantization scale to the header or other elements for each LoD. This allows the encoding device to contribute to the decoding device correctly decoding the quantization scale and appropriately decoding the bitstream.
[0486] The encoding device may convert the converted coefficients after quantization from signed integer values to unsigned integer values. For example, the encoding device may convert the quantized value a2q, which is a signed integer, to the quantized value a2u, which is an unsigned integer, as shown below.
[0487] If the quantized value a²q is less than 0: a²u = -1 - (2 × a²q) Otherwise: a²u = 2 × a²q (Equation 10)
[0488] Alternatively, for example, the encoding device may convert the quantized value b2q, which is a signed integer value, into the quantized value b2u, which is an unsigned integer value, as shown below.
[0489] If the quantized value b²q is less than 0: b²u = -1 - (2 × b²q) Otherwise: b²u = 2 × b²q (Equation 11)
[0490] This has the advantage that the encoding device does not need to consider the occurrence of negative integers when entropy encoding the conversion coefficients.
[0491] Furthermore, the encoding device does not necessarily need to convert signed integer values to unsigned integer values; for example, the sign bit can be separately entropy encoded.
[0492] Furthermore, the encoding method for the conversion coefficients is not limited to this. For example, the encoding device may arithmetically encode the sign bit representing the sign of the conversion coefficient and the binarized data of the absolute value of the conversion coefficient bit by bit using context. This may allow the encoding device to improve the encoding efficiency of the conversion coefficients of the displacement vector.
[0493] If quantization of the transformation coefficients of the displacement vector is not required, this process can be skipped, and the transformation coefficients can be arithmetically coded directly. This can reduce processing time.
[0494] Figure 63 is an explanatory diagram showing an example of how to calculate predicted values in this embodiment.
[0495] Referring to Figure 63, we will explain an example of generating LoD and calculating the predicted displacement vectors for each three-dimensional point.
[0496] In Figure 63, points a0, a1, and a2 are three-dimensional points included in the base mesh and belong to LoD0. Points b0 and b1 are three-dimensional points generated by one subdivision from the three-dimensional points included in the base mesh and belong to LoD1. Points c0, c1, c2, and c3 are three-dimensional points generated by two subdivisions from the three-dimensional points included in the base mesh and belong to LoD2.
[0497] If point b0 is a three-dimensional point generated by subdividing from points a0 and a1, the predicted displacement vector of point b0 can be calculated using points a0 and a1.
[0498] Furthermore, if point c2 is a three-dimensional point generated by subdividing from points a1 and b1, the predicted value of the displacement vector of point c2 can be calculated using points a1 and b1.
[0499] Figure 64 is an explanatory diagram showing an example of the calculation of the conversion coefficient in this embodiment.
[0500] Referring to Figure 64, we will explain an example of calculating the transformation coefficient by subtracting the predicted values from the displacement vectors of each three-dimensional point.
[0501] In Figure 64, the transformation coefficients a0c, a1c, and a2c are transformation coefficients included in the base mesh and belong to LoD0. The transformation coefficients b0c and b1c are three-dimensional points generated from three-dimensional points included in the base mesh by one subdivision and belong to LoD1. The transformation coefficients c0c, c1c, c2c, and c3c are three-dimensional points generated from three-dimensional points included in the base mesh by two subdivisions and belong to LoD2.
[0502] The transformation coefficient b0c for point b0 is obtained by subtracting the predicted value b0p for point b0 from the value of point b0.
[0503] b0c = b0 - b0p The predicted value b0p may be the average of the values at point a0 and point a1.
[0504] b0p = (a0 + a1) / 2 Also, the transformation coefficient c2c for point c2 is obtained by subtracting the predicted value c2p for point c2 from the value of point c2.
[0505] c2c = c2 - c2p. The predicted value c2p may be the average of the values at point a1 and point b1.
[0506] c2p = (a1 + b1) / 2 Alternatively, an encoding method (lifting transform) may be applied that calculates the transformation coefficients of the displacement vectors of three-dimensional points contained in the lower layers of the LoD and feeds these transformation coefficients back to the upper layers for encoding. By applying the lifting transform, the transformation coefficients of the low-frequency components of the displacement vectors can be collected in the upper layers, and the transformation coefficients of the high-frequency components of the displacement vectors can be collected in the lower layers. This allows for improved encoding efficiency, for example, by reducing the amount of information in the transformation coefficients of the high-frequency components in the lower layers through quantization.
[0507] Figure 65 is an explanatory diagram showing an example of interpretation of conversion coefficients in this embodiment.
[0508] The transformation coefficients for the displacement vectors of three-dimensional points may be interpreted using the transformation coefficients of the displacement vectors of frames that are temporally different from the frame being encoded. This will be explained with reference to Figure 65.
[0509] The transformation coefficients for the displacement vectors of three-dimensional points can be predicted using, for example, the transformation coefficients of the displacement vectors of the frame that was encoded or decoded immediately beforehand.
[0510] For example, if Frame(t), which is the frame at time t, is the frame to be encoded, the interpretation can be predicted using the transformation coefficients of the displacement vector of the frame that was encoded or decoded immediately before, i.e., Frame(t-1), which is the frame at time t-1.
[0511] Specifically, it is conceivable to encode a0c, which is the transformation coefficient of the displacement vector of a 3D point a0 in Frame(t), using interpretation with a0c', which is the transformation coefficient of the displacement vector of a0', the corresponding 3D point a0 in Frame(t-1). More specifically, it is conceivable to encode the value obtained by subtracting a0c' from a0c. If there is a correspondence between the 3D point a0 in Frame(t) and the 3D point a0' in Frame(t-1), the displacement vector values are likely to be close, so by subtracting a0c' from a0c, the transformation coefficient can be further reduced, thereby improving the encoding efficiency of entropy coding.
[0512] Interpretation is not limited to referencing the immediately preceding encoded frame; it may reference any frame. In this case, information about the referenced frame may be added to the header. This allows the decoder to correctly decode the bitstream by referencing the same frame referenced by the encoder. Interpretation may also refer to multiple frames. For example, using dual prediction with two reference frames can improve encoding efficiency. When using dual prediction, the average of the transformation coefficients of the displacement vectors of the two reference frames may be used as the interpretation value. This allows for the generation of prediction values with high accuracy and improves encoding efficiency.
[0513] Furthermore, information called `disp_inter_mode` may be added to the header to indicate whether or not to apply inter-prediction to the transformation coefficients of the displacement vector. This allows for adaptive control of whether inter-prediction is on or off, thereby improving encoding efficiency. For example, if the change in motion between frames is large and a correspondence between three-dimensional points cannot be established between frames, inter-prediction of the transformation coefficients of the displacement vector is turned off (e.g., `disp_inter_mode=0`), and if the change in motion between frames is small and a correspondence between three-dimensional points can be established between frames, inter-prediction is turned on (e.g., `disp_inter_mode=1`).
[0514] Furthermore, if the adaptive switching of interpretation is done on a frame-by-frame basis, disp_frame_inter_mode may be added to the header that stores the frame information, and if it is done on a sequence-by-sequence basis, disp_seq_inter_mode may be added to the header that stores the sequence information. This allows for control over the on / off switching of interpretation of the transformation coefficients of the displacement vector on a frame-by-frame or sequence-by-sequence basis, thereby improving coding efficiency.
[0515] Alternatively, information called `disp_lod_inter_mode` indicating whether or not to apply inter-prediction to the transformation coefficients of the displacement vector may be provided for each LoD, and the application of inter-prediction may be switched for each LoD. For example, the encoding device could compare the generated code amount with and without applying inter-prediction to the transformation coefficients of the displacement vector for each LoD, select the one with the smaller generated code amount, and add that information to the header. The decoding device would then decode the transformation coefficients of the displacement vector according to the information added to the header. This allows for improved encoding efficiency by switching whether or not to apply inter-prediction for each LoD.
[0516] The encoding device can decode the quantized transformation coefficients through inverse quantization and reconstruction, and use them to predict the three-dimensional points and beyond of the target to be encoded. Specifically, the encoding device can calculate the inverse quantized value by multiplying the quantized transformation coefficient by the quantization scale, and obtain the decoded value by adding the inverse quantized value and the predicted value. For example, the encoding device can calculate the inverse quantized value a2iq from the quantized value a2q as shown below, and can also calculate the inverse quantized value b2iq from the quantized value b2q as shown below.
[0517] a2iq=a2q×QS_LoD0 b2iq=b2q×QS_LoD1 (Formula 12)
[0518] Furthermore, the encoding device can calculate the reconstructed value a2rec from the inverse quantized value a2iq as follows, and can also calculate the reconstructed value b2rec from the inverse quantized value b2iq as follows.
[0519] a2rec=a2iq+a2p b2rec=b2iq+b2p (Formula 13)
[0520] In this embodiment, the encoding device is shown to generate predicted values of the displacement vectors of three-dimensional points by configuring one or more Lines of Data (LoDs). However, this is not necessarily the only method, and it may also be applied to cases where a single-layer LoD is configured to generate predicted values of the displacement vectors of three-dimensional points, or to cases where no LoD is generated to generate predicted values of the displacement vectors of three-dimensional points.
[0521] In this case, all three-dimensional points belong to the same LoD (e.g., LoD0). Therefore, when the encoding device encodes or decodes the three-dimensional points in LoD0 in order, it may generate the predicted values of the three-dimensional points belonging to LoD0 using the encoded and decoded displacement vectors included in LoD0. In this way, the encoding device may be able to reduce processing time by encoding without generating multiple layers of LoD.
[0522] Furthermore, if quantization of the transformation coefficients of the displacement vector is unnecessary, the encoding device may skip the quantization and dequantization processes and simply add the arithmetically decoded transformation coefficients to the predicted value to obtain the decoded value. This reduces processing time.
[0523] Figure 66 is an explanatory diagram showing an example of syntax in this embodiment.
[0524] The syntax example shown in Figure 66 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.
[0525] The syntax shown in Figure 66 includes a distribute vector_header, which includes NumLoD, NumOfPoint[i], Thres_Lod[i], NumNeiCnt[i], THd[i], and QS[i].
[0526] NumLoD indicates the number of levels in the Level of Development (LoD).
[0527] NumOfPoint[i] indicates the number of three-dimensional points belonging to hierarchical level i. Note that if the encoding device adds the total number of three-dimensional points, AllNumOfPoint, to a separate header, NumOfPoint[NumLoD-1] (i.e., the number of three-dimensional points belonging to the lowest level) may not be added to the header. In this case, NumOfPoint[NumLoD-1] can be calculated by the following (Equation 14).
[0528]
[0529] Thres_Lod[i] indicates the threshold for the LoD of layer i. The encoding device constructs the LoD such that the distance between each point in the LoD is greater than the threshold Thres_Lod[i]. Note that the value of Thres_Lod[NumLoD-1] (i.e., the threshold for the lowest layer's LoD) may not be added to the header. In this case, Thres_Lod[NumLoD-1] can be estimated to be 0. This reduces the amount of code in the header.
[0530] NumNeiCnt[i] indicates the upper limit of the number of neighboring points used to generate predicted values for three-dimensional points belonging to hierarchy i. If the number of neighboring points M is less than NumNeiCnt[i] (i.e., M < NumNeiCnt[i]), the encoding device may calculate the predicted value using M neighboring points. Also, if it is not necessary to have different values for NumNeiCnt[i] for each LoD, the encoding device may add one NumNeiCnt to the header.
[0531] THd[i] indicates the upper limit of the distance of the three-dimensional point used to predict the three-dimensional point to be encoded or decoded at level i. The encoding device may choose not to use three-dimensional points whose distance from the three-dimensional point to be encoded or decoded is greater than THd[i] for prediction. If it is not necessary to have different values for THd[i] in each LoD, a single THd may be added to the header.
[0532] QS[i] indicates the quantization scale of hierarchy i.
[0533] The encoding device may entropy encode NumLoD, Thres_Lod[i], NumNeiCnt[i], THd[i], or QS[i] and add them to the header. For example, the encoding device could binarize each value and perform arithmetic encoding. Alternatively, the encoding device may encode using a fixed length to reduce processing load.
[0534] Furthermore, the encoding device does not necessarily need to add NumLoD, Thres_Lod[i], NumNeiCnt[i], THd[i], or QS[i] to the header; these can be specified, for example, in the profile or level of the standard. This can reduce the number of bits in the header.
[0535] Figure 67 is an explanatory diagram showing an example of syntax in this embodiment.
[0536] The syntax example shown in Figure 67 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.
[0537] The syntax shown in Figure 67 includes a displacement vector_data. The displacement vector_data may include dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k] for each of the 0th to NumLoD levels (also known as the jth level) of LoD.
[0538] dispd_is_zero[k] is information indicating whether the absolute value of the k-th component of the transformation coefficient of the displacement vector of the i-th three-dimensional point (i.e., vertex[i]) included in the j-th hierarchy of LoD is 0. A value of 1 indicates that the absolute value of the k-th component's transformation coefficient is 0, and a value of 0 may indicate that the absolute value of the k-th component's transformation coefficient is 1 or greater.
[0539] dispd_is_one[k] is information indicating whether the absolute value of the transformation coefficient of the kth component of the displacement vector of the i-th three-dimensional point (i.e., vertex[i]) included in the j-th hierarchy of LoD is 1. A value of 1 indicates that the absolute value of the predicted residual of the kth component is 1, and a value of 0 may indicate that the absolute value of the predicted residual of the kth component is 2 or greater.
[0540] Furthermore, if dispd_is_one[k] is not included in the bitstream, the decryption device may estimate its value to be 0. This prevents an undefined value from being set for dispd_is_one[k] during decryption, allowing the decryption process to be performed appropriately.
[0541] dispd_minus2[k] is information that indicates the value obtained by subtracting 2 from the absolute value of the transformation coefficient of the kth component of the displacement vector of the i-th three-dimensional point (vertex[i]) included in the j-th hierarchy of LoD.
[0542] If dispd_minus2[k] is not included in the bitstream, the decoder may estimate its value to be 0. This prevents an undefined value from being set for dispd_minus2[k] during decoding, allowing the decoding process to be performed appropriately.
[0543] dispd_sign[k] represents the sign bit of the displacement vector of the kth component of the i-th three-dimensional point (vertex[i]) included in the j-th hierarchy of LoD. A value of 1 may indicate that the transformation coefficient of the kth component is negative, and a value of 0 may indicate that the transformation coefficient of the kth component is positive.
[0544] Furthermore, for the k-th component, if the displacement vector is expressed in a Cartesian coordinate system, the first component may represent the x-component, the second component the y-component, and the third component the z-component. Also, if the displacement vector is expressed in a local coordinate system, the first component may represent the normal component, the second component the tangent component, and the third component the binormal component. This allows for the use of a common syntax structure whether the displacement vector is expressed in a Cartesian or local coordinate system.
[0545] The transformation coefficient dispd[k] of the k-th component of the displacement vector of the i-th three-dimensional point (i.e., vertex[i]) may be calculated using the above information by the calculation process shown in Figure 68.
[0546] By introducing the syntax configuration shown in Figure 67, the encoding device can reduce the frequency with which it encodes and adds dispd_is_one[k], dispd_minus2[k], and dispd_sign[k] to the bitstream when encoding conversion coefficients that tend to result in dispd[k] = 0, thereby potentially improving encoding efficiency. Furthermore, when encoding prediction residuals that tend to result in dispd[k] = 1 or 0, the encoding device can reduce the frequency with which it encodes and adds dispd_minus2[k] to the bitstream, thereby potentially improving encoding efficiency.
[0547] In this embodiment, we have shown an example assuming a case where the conversion coefficient dispd[k] is likely to be 0 or 1, but the process is not necessarily limited to this, and the same processing may be applied to any dispd[k]. For example, when encoding a conversion coefficient that is likely to be dispd[k] = 2, dispd_is_two[k] and dispd_minus3[k] may be newly introduced. This reduces the frequency with which dispd_minus3[k] is encoded and added to the bitstream when encoding a conversion coefficient that is likely to be dispd[k] = 2, and as a result, the encoding efficiency may be improved. In this case, dispd[k] may be calculated by the calculation process shown in Figure 69.
[0548] The encoding device may also apply context-based arithmetic coding by binarizing at least one of dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k]. For example, since each of dispd_is_zero[k], dispd_is_one[k], and dispd_sign[k] is 1 bit, the encoding device may assign one context to each of them and encode while updating the probability of occurrence based on the frequency of occurrence of 0 and 1. This may improve encoding efficiency. Alternatively, the encoding device may binarize dispd_minus2[k] using Exponential Golomb, assign a context to each bit, and encode while updating the probability of occurrence based on the frequency of occurrence of 0 and 1. This may also improve encoding efficiency.
[0549] The encoding device may assign a different context to each component of dispd as the context to be assigned to dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k]. This may improve encoding efficiency when the values of dispd differ for each component. Alternatively, the encoding device may assign the same context to each component of dispd as the context to be assigned to dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k]. This may improve encoding efficiency when the values of each component of dispd are similar.
[0550] The decoding device may convert the decoded quantized transformation coefficients from unsigned integer values to signed integer values in the reverse manner of the encoding device. This allows for proper decoding of the generated bitstream without considering the occurrence of negative integers when entropy coding the transformation coefficients.
[0551] It is not always necessary to convert from an unsigned integer value to a signed integer value. For example, if the decoder decodes a bitstream generated by separately entropy encoding the sign bit, it may decode the sign bit. Furthermore, the decoding method of the conversion coefficients by the decoder is not limited to this; for example, the sign bit representing the sign of the conversion coefficient and the binarized data of the absolute value of the conversion coefficient may be arithmetically decoded bit by bit using context. This allows the decoder to appropriately decode a bitstream with improved encoding efficiency of the conversion coefficients of the displacement vector.
[0552] The decoding device decodes the quantized conversion coefficients, which have been converted to signed integer values, through inverse quantization and reconstruction, and uses them to predict the three-dimensional point and beyond of the target to be decoded. Specifically, the decoding device calculates the inverse quantized value by multiplying the quantized conversion coefficient by the decoded quantization scale, and obtains the decoded value by adding the inverse quantized value and the predicted value.
[0553] For example, the decoded unsigned quantized value a2u is converted to a signed value a2q as shown below. Note that ">>" indicates a bit shift operation.
[0554] If the least significant bit (LSB) of a2u is 1, then a2q = -((a2u+1)>>1). Otherwise, a2q = (a2u>>1). (Equation 15)
[0555] Furthermore, for example, the decoded unsigned quantized value b2u is converted to a signed value b2q as follows.
[0556] If the LSB of b2u is 1, then b2q = -((b2u+1)>>1). Otherwise, b2q = (b2u>>1). (Equation 16)
[0557] The decoding device calculates a reconstructed value after inverse quantization. This reconstructed value can be used to predict the three-dimensional point and beyond that of the target being decoded.
[0558] For example, the decoding device can calculate the inverse quantization value a2iq from the quantization value a2q as shown below, and can also calculate the inverse quantization value b2iq from the quantization value b2q as shown below.
[0559] a2iq=a2q×QS_LoD0 b2iq=b2q×QS_LoD1 (Formula 17)
[0560] Furthermore, the decoding device can calculate the reconstructed value a2rec from the inverse quantization value a2iq as follows, and can also calculate the reconstructed value b2rec from the inverse quantization value b2iq as follows.
[0561] a2rec=a2iq+a2p b2rec=b2iq+b2p (Formula 18)
[0562] In the above example, we showed that the encoding device generates a predicted displacement vector of a three-dimensional point by calculating the average of the displacement vectors of three-dimensional points up to a certain number of neighboring points that have been encoded and decoded for the three-dimensional point to be encoded. However, this is not necessarily the only method, and the predicted value can be generated by other methods.
[0563] For example, the encoding device may use the displacement vector of the closest 3D point among the encoded and decoded adjacent 3D points of the 3D point to be encoded as the predicted value. Alternatively, the encoding device may add a prediction mode value (PredMode) to each 3D point, allowing the user to select a predicted value. For example, the encoding device may provide a total of M prediction modes, assign the average value to prediction mode 0, assign the displacement vector of 3D point A to prediction mode 1, ..., assign the displacement vector of 3D point Z to prediction mode M-1, and add the prediction mode used for prediction to the bitstream for each 3D point. The 3D points A to Z to which the displacement vectors are assigned from prediction mode 1 to prediction mode M-1 may be used in order from the closest adjacent 3D points of the 3D point to be encoded, among the encoded and decoded adjacent 3D points of the 3D point to be encoded.
[0564] Figure 70 is an explanatory diagram showing an example of predicted displacement vector information in this embodiment. Figure 71 is an explanatory diagram showing a method for generating predicted displacement vector values in this embodiment.
[0565] Figure 70 shows an example of prediction value information used to predict point b2 when the number of adjacent three-dimensional points N used for prediction is 4 and the number of prediction modes M is 5. The prediction value information includes information indicating the prediction value used for each of the one or more prediction modes. The example of prediction value information shown in Figure 70 is a table showing the prediction value used for each of the one or more prediction modes.
[0566] In the example shown in Figure 70, the predicted values used to predict point b2 are, for example, the adjacent three-dimensional points a0, a1, a2, and b1 (see Figure 71). Correspondingly, the "average value of points a0, a1, a2, and b1" is assigned as the predicted value for prediction mode 0.
[0567] Furthermore, in Figure 70, "point b1" is assigned as the predicted value for prediction mode 1. "point b2" is assigned as the predicted value for prediction mode 2. "point a1" is assigned as the predicted value for prediction mode 3. "point a0" is assigned as the predicted value for prediction mode 4.
[0568] The numerical value that uniquely identifies a prediction mode is also called the prediction mode value. Here, we will explain assuming that the prediction mode value for prediction mode m is m. As an example, prediction mode values are used in order from smallest integer values.
[0569] The assignment of prediction mode values may be determined in order of distance from the three-dimensional point to be encoded. For example, an encoding device can assign relatively small prediction mode values to three-dimensional points that are closer to the three-dimensional point to be encoded. In the above example, the three-dimensional point that is closest to the three-dimensional point b2 is point b1, the next closest is point a2, the next closest is point a1, and the next closest is point a0.
[0570] This allows for assigning smaller prediction mode values to points where the difference between the displacement vector and the predicted value is relatively small due to the small distance, making them more likely to be selected as prediction values. This reduces the number of bits required to encode the prediction mode values. Alternatively, smaller prediction mode values may be preferentially assigned to three-dimensional points belonging to the same Line of Disorder (LoD) as the three-dimensional point being encoded.
[0571] Figure 72 shows an example of the predicted value information used to predict point a2 when the number of adjacent three-dimensional points N used for prediction is 2 and the number of prediction modes M is 5.
[0572] In the example of prediction information shown in Figure 72, the prediction values used to predict point a2 are, for example, adjacent three-dimensional points a0 and a1. Correspondingly, in Figure 72, the "average value of points a0 and a1" is assigned as the prediction value for prediction mode 0.
[0573] Additionally, "point a1" is assigned as the predicted value in prediction mode 1. "point a0" is assigned as the predicted value in prediction mode 2.
[0574] Furthermore, if the number of adjacent points is less than four, information indicating that the prediction mode will not be used (indicated as "not available" in the diagram) may be set for prediction modes where no prediction value has been assigned.
[0575] Figure 73 shows an example of predicted value information when the displacement vector is expressed in a Cartesian coordinate system (XYZ coordinate system).
[0576] In the example shown in Figure 73, the values used to predict point b2 are, for example, the adjacent three-dimensional points a0, a1, a2, and b1 (see Figure 71). Correspondingly, in Figure 73, the coordinates (Xave, Yave, Zave) of the "average value of points a0, a1, a2, and b1" are assigned as the predicted values for prediction mode 0. Here, Xave can be calculated as the average of Xb1, Xa2, Xa1, and Xa0, or a weighted average. Yave can be calculated as the average of Yb1, Yb2, Ya1, and Ya0, or a weighted average. Zave can be calculated as the average of Zb1, Zb2, Za1, and Za0, or a weighted average.
[0577] Additionally, the coordinates (Xb1, Yb1, Zb1) of "point b1" are assigned as the predicted value for prediction mode 1. The coordinates (Xa2, Ya2, Za2) of "point b2" are assigned as the predicted value for prediction mode 2. The coordinates (Xa1, Ya1, Za1) of "point a1" are assigned as the predicted value for prediction mode 3. The coordinates (Xa0, Ya0, Za0) of "point a0" are assigned as the predicted value for prediction mode 4.
[0578] For example, the encoding device may select prediction mode 2 (i.e., prediction mode value 2) and encode the XYZ components of the displacement vector of the three-dimensional point to be encoded using the predicted values Xa2, Ya2, and Za2, respectively. In this case, the encoding device adds prediction mode value 2 to the bitstream.
[0579] In the above example, the case where the displacement vector is in a Cartesian coordinate system was used as an example, but it is not necessarily limited to this, and can also be applied to displacement vectors expressed in a local coordinate system, for example.
[0580] The number of prediction modes M may be appended to the bitstream. Alternatively, the number of prediction modes M may not be appended to the bitstream, but its value may be defined in the standard's profile or level, etc. Furthermore, the number of prediction modes M may be a value calculated from the number of three-dimensional points N used for prediction (for example, M = N + 1).
[0581] When the encoding device generates predicted displacement vectors for three-dimensional points by adding a PredMode value to each three-dimensional point, an example of how to assign predicted values to each PredMode is shown, where the displacement vectors of adjacent points are assigned as predicted values using distance information from the three-dimensional point to be encoded. However, this is not the only method, and the method of assigning predicted values to PredModes can be changed in any way.
[0582] For example, the encoding device may calculate a median from the predicted values assigned to each prediction mode and assign the calculated median to prediction mode 0. In this way, the encoding device may assign the median as the predicted value to the prediction mode with the smallest predicted mode value. This allows the encoding device to generate candidate predicted values that prioritize the median of the displacement vectors of adjacent points, thereby improving encoding efficiency.
[0583] The changes to the assignment of predicted values using the median will be explained with reference to Figures 74 to 76.
[0584] Figure 74 is an explanatory diagram showing an example of predicted displacement vector information in this embodiment. Figure 75 is an explanatory diagram showing a method for generating predicted displacement vector values in this embodiment. Figure 76 is an explanatory diagram showing an example of predicted displacement vector information in this embodiment.
[0585] In the example shown in Figure 75, the number of three-dimensional points used for prediction is N = 4, and the number of prediction modes is M = 4. Point a2 is predicted from points a0 and a1. Point b2 is predicted from points a0, a1, a2, b0, and b1.
[0586] Here, we show an example where points b1, a2, a1, and a0 are assumed to be in order of proximity to the three-dimensional point to be encoded, and the displacement vector of the closest three-dimensional point is assigned to the prediction mode with the smallest prediction mode value. The magnitude of each prediction value is assumed to be b1 > a1 > a0 > a2.
[0587] The encoding device calculates the median of the predicted values in prediction mode. For example, the encoding device can sort the n predicted values assigned to prediction mode in ascending or descending order and use the (n / 2)th value as the median. The method of calculating the median can be switched depending on whether n is odd or even.
[0588] For example, if n is odd, the encoding device can use the (n / 2)th predicted value (with decimal places truncated) from the 0th to the (n-1)th predicted values after sorting as the median. If n is even, the encoding device can use the (n / 2-1)th predicted value and the n / 2nd predicted value from the 0th to the (n-1)th predicted values after sorting as candidate median A and B, and may adopt either A or B as the median by some method. For example, the median can be the one between A and B that is closer in distance to the three-dimensional point to be encoded.
[0589] In the example shown in Figure 75, n = 4, so the median can be calculated using the method for calculating the median when n is even. For example, sorting b1, a1, a0, and a2 in ascending order results in a2, a0, a1, and b1. In this case, the (n / 2-1)th predicted value is a0, and the n / 2nd is a1, which are candidates A and B for the median. Since a1 is closer to the three-dimensional point to be encoded than a0, a1 is selected as the median.
[0590] In this case, as shown in Figure 76, the encoding device assigns the selected median value a1 to prediction mode 0, and assigns the originally assigned prediction value b1 to prediction mode 2, which was assigned to prediction value a1. In other words, the encoding device swaps the prediction values of prediction mode 0 and prediction mode 2. This allows the encoding device to generate candidate prediction values that prioritize the median of the displacement vectors of adjacent points, thereby improving encoding efficiency.
[0591] While the example shown uses the median to assign predicted values to prediction modes, this is not the only method. For example, the encoding device may calculate the average value from the predicted values assigned to each prediction mode and assign a predicted value close to the average value to prediction mode 0. This allows for the generation of candidate predicted values that prioritize displacement vectors close to the average of the displacement vectors of adjacent points, thereby improving encoding efficiency.
[0592] Alternatively, the encoding device may first calculate the median value from the displacement vectors of adjacent points and assign it to prediction mode 0, and then assign the displacement vectors of surrounding three-dimensional points other than the median value to prediction modes 1 and above using the distance information of those three-dimensional points.
[0593] Furthermore, the encoding device may add information to the header, etc., indicating whether to prioritize the median (also called median-prioritizing information). If the median-prioritizing information indicates that the median should be prioritized, the encoding device may assign the median to prediction mode 0 using the method described above, and if not, it may assign a predicted value to the prediction mode regardless of the median. This allows the encoding device to potentially improve encoding efficiency by adaptively switching between cases where median priority is desired and cases where it is not. In addition, the decoding device can appropriately decode the bitstream based on the median-prioritizing information added to the header, etc.
[0594] As an example of changing the prediction value assignment to prioritize the median, we showed an example where the median is assigned to prediction mode 0, and the prediction value originally assigned to prediction mode 0 is swapped with the prediction mode to which the median was assigned. However, this is not necessarily the only example.
[0595] Figure 77 is an explanatory diagram showing an example of predicted displacement vector information in this embodiment.
[0596] For example, as shown in Figure 77, the encoding device may shift the predicted values assigned to each prediction mode until the value is reassigned to the prediction mode to which the median was originally assigned. This can be done by assigning the median to prediction mode 0, assigning the predicted values originally assigned to prediction mode 0 to prediction mode 1, assigning the predicted values originally assigned to prediction mode 1 to prediction mode 2, and so on. This allows for the generation of predicted value information that prioritizes the median of the displacement vectors of adjacent points while prioritizing candidates for predicted values that are close in distance, thereby improving encoding efficiency.
[0597] Furthermore, while examples of assigning predicted values prioritizing the median or mean have been shown, this is not necessarily the only approach.
[0598] Figures 78 and 79 are explanatory diagrams showing examples of predicted displacement vector information in this embodiment.
[0599] For example, the encoding device calculates statistical information about the predicted values in prediction mode, as shown in Figure 78. This statistical information may include, for example, the median, mean, variance, or standard deviation of adjacent points. The encoding device can then change the assignment of predicted values based on the calculated statistical information (see Figure 79).
[0600] A modified example of setting the predicted value of the three-dimensional point a to be encoded in the frame to be encoded will be explained with reference to Figures 80 to 83.
[0601] Figure 80 is an explanatory diagram showing an example of a point to be encoded in this embodiment. Figure 81 is an explanatory diagram showing an example of predicted displacement vector information in this embodiment. Figure 82 is an explanatory diagram showing an example of a temporal dv in this embodiment. Figure 83 is an explanatory diagram showing an example of predicted displacement vector information in this embodiment.
[0602] The encoding device may set the predicted value of the three-dimensional point a (see Figure 80) to be encoded in the frame to be encoded as shown in the predicted value information in Figure 81. Specifically, the encoding device may set 0 (no prediction) as the predicted value for prediction mode 0, and may set the average value of the displacement vectors of adjacent points a, b, and c (dv0, dv1, and dv2, respectively) as the predicted value for prediction mode 1. The encoding device may also set the displacement vectors of adjacent points a, b, and c (dv0, dv1, and dv2, respectively) as the predicted values for prediction modes 2, 3, and 4, respectively.
[0603] Note that the predicted values assigned to each prediction mode are not limited to these; other predicted values may also be assigned.
[0604] Furthermore, the encoding device may, for example, assign a displacement vector from a reference frame different from the frame to be encoded (see Figure 82) as a predicted value. Specifically, the encoding device can use the displacement vector (hereinafter referred to as "temporal dv") of the corresponding point a' of the three-dimensional point a to be encoded in an already encoded or decoded reference frame as a predicted value for the three-dimensional point a to be encoded. When the target object is moving at a constant rate, the value of the displacement vector of the three-dimensional point to be encoded tends to be relatively close to the value of the displacement vector of the corresponding point of the three-dimensional point to be encoded in the reference frame. Therefore, adding the temporal dv as a predicted value to the prediction candidates may improve encoding efficiency.
[0605] Figure 83 shows an example of prediction value information when the prediction value for prediction mode 5 is added as a temporary DV. Note that a temporary DV may also be added as a prediction value for other prediction modes (i.e., any of prediction modes 0 to 4). In addition, the prediction value of any of the prediction modes in the prediction value information shown in Figure 81 may be changed to a temporary DV.
[0606] The encoding device may calculate the temporal dv for each prediction unit (Displacement Vector Group, DVG) of the displacement vector in the reference frame, store it in memory, and use the temporal dv of the DVG to which the corresponding point a' belongs as the temporal dv of the corresponding point a'. This can reduce the amount of memory used.
[0607] The encoding device may calculate the temporary DV of the DVG from the displacement vectors of the three-dimensional points belonging to the DVG. For example, the average value of the displacement vectors of the three-dimensional points belonging to the DVG may be used as the temporary DV of the DVG. This may reduce the amount of memory required to store the temporary DV while potentially improving encoding efficiency by adding the temporary DV to the prediction candidates.
[0608] Furthermore, for example, the encoding device may calculate the global displacement vector (hereinafter referred to as global dv) of the frame to be encoded and add the global dv to the prediction candidates as the predicted value. The encoding device can calculate the global dv from, for example, the average value of the displacement vectors in the frame to be encoded or the reference frame. The encoding device may also add the calculated global dv to the bitstream. As a result, the decoding device can decode the global dv added as a prediction candidate by the encoding device from the bitstream and add the same global dv as the encoding device to the prediction candidates.
[0609] Furthermore, the encoding device may, for example, select at least two displacement vectors from the displacement vectors added to the prediction candidates as new predicted values, and add the average value of the two or more selected displacement vectors as a new predicted value to the prediction candidates. This may improve encoding efficiency.
[0610] Furthermore, the encoding device may, for example, store one or more previously used displacement vectors in memory as new predicted values, and add at least one of these displacement vectors as a new predicted value to the prediction candidates. This may improve encoding efficiency. The encoding device may also periodically or irregularly store the displacement vectors used for encoding or decoding in memory (i.e., the memory in which the one or more previously used displacement vectors mentioned above are stored), and delete old displacement vectors from memory after a certain period of time has elapsed since their storage. In this way, the encoding device can update the displacement vectors stored in memory to assign new displacement vectors to the prediction candidates, potentially improving encoding efficiency.
[0611] When an encoding device encodes the displacement vectors of three-dimensional points, it may provide DVGs (Digital Vector Generators) as prediction units according to the encoding or decoding order, and encode or decode each DVG. For example, it is conceivable to define the number of three-dimensional points included in a DVG (DVGSize) and divide the three-dimensional points into multiple DVGs according to the encoding or decoding order for encoding or decoding. The encoding or decoding order of the displacement vectors of three-dimensional points does not matter. For example, it is possible to generate an Order of Data (LD) and encode or decode sequentially for each LD hierarchy. Alternatively, without generating an LD, the displacement vectors may be encoded or decoded in the same order as the encoding or decoding order of the three-dimensional point position information (vertex). Furthermore, a Morton code may be generated using the three-dimensional point position information, and the encoding or decoding may be performed in the order of the Morton code.
[0612] Examples of DVG definitions will be explained below.
[0613] Figure 84 is an explanatory diagram showing an example of a reference to the DVG in this embodiment.
[0614] In the example of a DVG reference shown in Figure 84 (also referred to as the first example), three-dimensional points within the same DVG are defined as not being able to be referenced. For example, three-dimensional points within the same DVG may be defined as not being able to be added as adjacent points.
[0615] Furthermore, three-dimensional points in different encoded or decoded DVGs are defined as referable. For example, three-dimensional points in different encoded or decoded DVGs may be made inaccessible as adjacent points.
[0616] Furthermore, the size of the DVG may be included in the header or elsewhere (see Figure 85). For example, if the size of the DVG (DVGSize) is 16, the information "DVGSize = 16" may be added to the header. Alternatively, DVGSize may be 2^n, and the value of n may be added to the header.
[0617] Furthermore, three-dimensional points within the same DVG can be encoded or decoded in parallel.
[0618] Figure 85 is an explanatory diagram showing an example of syntax in this embodiment.
[0619] The syntax example shown in Figure 85 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.
[0620] The syntax shown in Figure 85 includes `displacement vector_header`, which includes `DVGSize`.
[0621] DVGSize indicates the number of three-dimensional points included in the DVG.
[0622] Figure 86 is an explanatory diagram showing an example of a reference to the DVG in this embodiment.
[0623] In the example of DVG references shown in Figure 86 (also known as the second example), encoded or decoded 3D points within the same DVG are defined as referable. Unencoded or undecoded 3D points are defined as not referable. For example, encoded or decoded 3D points within the same DVG may be allowed to be added as adjacent points, while unencoded or undecoded 3D points may not be allowed to be added.
[0624] Furthermore, three-dimensional points in different encoded or decoded DVGs are defined as referable. For example, three-dimensional points in different encoded or decoded DVGs may be adjacent.
[0625] Furthermore, the size of the DVG may be included in the header or elsewhere (see Figure 85). For example, if the size of the DVG (DVGSize) is 16, the information "DVGSize = 16" may be added to the header. Alternatively, DVGSize may be 2^n, and the value of n may be added to the header.
[0626] By doing this, even within the same DVG, encoded and decoded 3D points can be referenced, thereby improving prediction accuracy and encoding efficiency.
[0627] Figure 87 is an explanatory diagram showing an example of a reference to the DVG in this embodiment.
[0628] In the example of a DVG reference shown in Figure 87 (also known as the third example), encoded or decoded three-dimensional points within the same DVG are defined as referable. Unencoded or undecoded three-dimensional points are defined as not referable. For example, encoded or decoded three-dimensional points within the same DVG may be allowed to be added as adjacent points, while unencoded or undecoded three-dimensional points may not be allowed to be added as adjacent points.
[0629] Furthermore, three-dimensional points within different DVGs are defined as not being able to reference each other. For example, three-dimensional points within different DVGs may be defined as not being able to be added as adjacent points.
[0630] Furthermore, the size of the DVG may be included in the header or elsewhere (see Figure 85). For example, if the size of the DVG (DVGSize) is 16, the information "DVGSize = 16" may be added to the header. Alternatively, DVGSize may be 2^n, and the value of n may be added to the header.
[0631] By prohibiting references between DVGs in this way, dependencies between DVGs are eliminated, making it possible for multiple DVGs to be encoded or decoded in parallel.
[0632] Furthermore, by allowing the referencing of encoded or decoded 3D points within the same DVG, it may be possible to improve prediction accuracy and encoding efficiency.
[0633] In the explanation referring to Figure 84, an example was shown in which, when encoding the displacement vector of a three-dimensional point, DVGs are provided according to the encoding or decoding order, and encoding or decoding is performed for each DVG. For example, an example was shown in which the number of three-dimensional points included in a DVG (DVGSize) is defined, and the three-dimensional points are divided into multiple DVGs according to the encoding or decoding order and encoded or decoded. Here, the prediction mode PredMode for encoding the displacement vector, or the information disp_dvg_inter_mode on whether to apply interpretation, may be set for each DVG. In this case, three-dimensional points included in the same DVG will share PredMode or disp_dvg_inter_mode, and the same value may be set. This can improve encoding efficiency by reducing the amount of code in PredMode or disp_dvg_inter_mode. Furthermore, it is not necessarily required to set PredMode or disp_dvg_inter_mode for each DVG; it is also acceptable to allow setting PredMode or disp_dvg_inter_mode for each other set of three-dimensional points.
[0634] Figure 88 is an explanatory diagram showing an example of a reference to the DVG in this embodiment.
[0635] In the example of a DVG reference shown in Figure 88 (also known as the fourth example), encoded or decoded three-dimensional points within the same DVG are defined as referable. Unencoded or undecoded three-dimensional points are defined as not referable. For example, encoded or decoded three-dimensional points within the same DVG may be allowed to be added to adjacent points, while unencoded or undecoded three-dimensional points may not be allowed to be added.
[0636] Furthermore, three-dimensional points in different encoded or decoded DVGs are defined as referential. For example, three-dimensional points in different encoded or decoded DVGs may be adjacent to adjacent points.
[0637] Alternatively, the encoding device may add a PredMode or dips_dvg_inter_mode to each DVG and predictively encode the three-dimensional points within the DVG using the same PredMode or dips_dvg_inter_mode.
[0638] Furthermore, the encoding device may determine whether or not to add PredMode or disp_dvg_inter_mode to each DVG. For example, the encoding device may calculate the PredMode or disp_dvg_inter_mode of the DVG to which the three-dimensional point to be encoded belongs using the variance value of the displacement vector of the decoded three-dimensional point in a different DVG. If the calculated variance value is greater than or equal to a threshold, PredMode or disp_dvg_inter_mode may be added to the DVG; otherwise, PredMode or disp_dvg_inter_mode may not be added. If PredMode or disp_dvg_inter_mode is not added, it may be assumed that PredMode = 0 or disp_dvg_inter_mode = 0.
[0639] Furthermore, the size of the DVG may be included in the header or elsewhere (see Figure 85). For example, if the size of the DVG (DVGSize) is 16, the information "DVGSize = 16" may be added to the header. Alternatively, DVGSize may be 2^n, and the value of n may be added to the header.
[0640] Thus, by allowing the referencing of encoded or decoded three-dimensional points within the same DVG, it may be possible to improve prediction accuracy and encoding efficiency.
[0641] Furthermore, by adding PredMode or dips_dvg_inter_mode to each DVG, it may be possible to reduce overhead and improve encoding efficiency compared to adding PredMode or dips_dvg_inter_mode to each three-dimensional point.
[0642] Figure 89 is an explanatory diagram showing an example of syntax in this embodiment.
[0643] The syntax example shown in Figure 89 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.
[0644] The syntax shown in Figure 89 includes a `displacement vector_header`, which includes a `DVGSize`.
[0645] DVGSize represents the unit for predicting the displacement vector of a three-dimensional point. For each DVGSize of three-dimensional points, PredMode or disp_dvg_inter_mode is assigned, and three-dimensional points within the same DVG are encoded and decoded using the same PredMode or disp_dvg_inter_mode.
[0646] Furthermore, when encoding the displacement vector by dividing it into LoD hierarchies, it may be possible to set a different DVGSize for each LoD hierarchy. In that case, the DVGSize for each LoD hierarchy may be added to the header. This allows the decoding device to correctly decode the bitstream generated by setting the DVGSize for each LoD hierarchy.
[0647] For example, when encoding displacement vectors using a LoD hierarchy, the lifting transform tends to concentrate high-frequency components in lower LoD levels, resulting in smaller transform coefficients. Therefore, the accuracy of interpretation tends to improve with lower LoD levels. Consequently, increasing the DVGSize value in lower LoD levels and sharing dips_dvg_inter_mode across many three-dimensional points can reduce the amount of code required to encode disp_dvg_inter_mode, potentially improving encoding efficiency. On the other hand, the lifting transform tends to concentrate low-frequency components in the above LoD hierarchy, resulting in larger transform coefficients. Consequently, the accuracy of interpretation tends to decrease with lower LoD levels. Therefore, decreasing the DVGSize value in lower LoD levels and allowing for finer control over whether or not to perform interpretation can potentially improve encoding efficiency.
[0648] Figure 90 is an explanatory diagram showing an example of syntax in this embodiment.
[0649] The syntax example shown in Figure 90 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.
[0650] The syntax shown in Figure 90 includes displacement_vector_data. displacement_vector_data may include PredMode, disp_dvg_inter_mode, dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k] for each of the 0th to NumLoD levels (also known as the jth level) of LoD.
[0651] PredMode is information indicating the prediction mode for encoding or decoding the displacement vector of the i-th 3D point in the j-th hierarchy. PredMode takes a value from 0 to M-1 (where M is the total number of prediction modes). If PredMode is not included in the bitstream (i.e., the condition in the if statement "maxdiff >= Thfix[i] && NumPredMode[i] > 1" is not met), PredMode may be estimated to be 0. Note that the estimated value is not limited to 0; any value from 0 to M-1 may be used. Alternatively, the estimated value for when PredMode is not included in the bitstream may be added separately in a header or similar. Furthermore, PredMode may be arithmetic encoded by binarizing it using a truncated unary code with the number of prediction modes to which the predicted value is assigned.
[0652] `disp_dvg_inter_mode` indicates whether the i-th displacement vector of the j-th LoD hierarchy is encoded or decoded using interpretation. A value of 1 indicates that interpretation is applied, and a value of 0 indicates that interpretation is not applied.
[0653] The information for dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k] is the same as the information in Figure 67.
[0654] Hereafter, an example of the encoding process in this embodiment will be described.
[0655] Figure 91 is a flowchart showing the encoding process in this embodiment.
[0656] The encoding process shown in Figure 91 is a method for encoding three-dimensional point displacement data performed by an encoding device.
[0657] In step S9101, the encoding device generates a predicted value for the first displacement data by performing interpretation using the first displacement data, which is the displacement data to be encoded, and the second displacement data from a different time.
[0658] In step S9102, the encoding device generates a predicted residual using the first displacement data and the generated predicted value.
[0659] In step S9103, the encoding device encodes the generated prediction residual.
[0660] As a result, the encoding device can appropriately encode the displacement data of three-dimensional points by encoding the prediction residual generated by interpretation. By using interpretation, it may be possible to reduce the amount of encoding when the difference between the first displacement data to be encoded and the second displacement data at a different time is relatively small, thereby potentially improving the encoding process. Thus, the above encoding method can contribute to improvements in encoding processing related to displacement vectors.
[0661] For example, when generating predicted values for the first displacement data, it is possible to determine whether or not to perform interpretation when generating the predicted values for the first displacement data. If it is determined that interpretation should be performed, the predicted values for the first displacement data are generated by performing interpretation. If it is determined that interpretation should not be performed, the predicted values for the first displacement data may be generated without performing interpretation.
[0662] This allows the encoding device to determine in advance whether or not to use interpretation to generate predicted values for the displacement data when encoding the first displacement data to be encoded, and to switch whether or not to use interpretation to generate predicted values for the displacement data according to the result of that determination. This allows, for example, the encoding device to use interpretation when it can improve the encoding process, and to not use interpretation when it cannot improve (or degrades) the encoding process. This has the potential to further reduce the amount of code. In this way, the above encoding method can contribute to further improvements in encoding processes related to displacement vectors.
[0663] For example, when deciding whether or not to perform interpretation, one could calculate a first sum, which is the sum of the transformation coefficients for the first displacement data, and a second sum, which is the sum of the prediction residuals when interpretation is applied to the transformation coefficients. If the second sum is determined to be smaller than the first sum, one can decide to perform interpretation; if the second sum is not smaller than the first sum, one can decide not to perform interpretation.
[0664] This allows the encoding device to switch whether or not to use interpretation to generate predicted values for displacement data by comparing the sum of the transformation coefficients when interpretation is applied to the transformation coefficients of the first displacement data to be encoded with the above transformation coefficients (in other words, the sum of the transformation coefficients when interpretation is not used). Specifically, it can decide to use interpretation if it determines that the sum of the transformation coefficients when interpretation is applied is small. This makes it possible to reduce the amount of coding with a simpler decision. In this way, the above encoding method can contribute to improvements in encoding processing related to displacement vectors.
[0665] For example, information indicating whether or not interpretation was used to generate the predicted residuals of the first displacement data may also be transmitted.
[0666] This allows the encoding device to transmit information indicating whether or not interpretation was used during encoding, thereby informing the decoding device, which receives and decodes the encoded displacement data, whether or not interpretation was used during encoding. This can contribute to the proper decoding of the encoded data. Specifically, it can contribute to ensuring that data encoded using interpretation is decoded using interpretation, and data encoded without interpretation is decoded without interpretation. In this way, the above encoding method can contribute to improvements in encoding processing related to displacement vectors.
[0667] For example, a three-dimensional point may include multiple three-dimensional points, each of which belongs to one of several hierarchies. When generating predicted values for the first displacement data, for each of the one or more hierarchies to which the multiple three-dimensional points belong, it is determined whether or not to perform interpretation when generating predicted values for the first displacement data for the three-dimensional points belonging to that hierarchy. For the hierarchies to which it is determined that interpretation should be performed, the predicted values for the first displacement data are generated by performing interpretation. For the hierarchies to which it is determined that interpretation should not be performed, the predicted values for the first displacement data are generated without performing interpretation.
[0668] As a result, when encoding the first displacement data to be encoded, the encoding device can pre-determine whether or not to use interpretation to generate predicted values of the displacement data for each hierarchy to which the three-dimensional point belongs, and switch whether or not to use interpretation to generate predicted values of the displacement data for each hierarchy to which the three-dimensional point belongs, according to the result of that determination. This makes it possible, for example, to use interpretation in the encoding process of hierarchies where using interpretation can improve the encoding process, and to avoid using interpretation in the encoding process of hierarchies where using interpretation does not improve (or degrades) the encoding process. This has the potential to further reduce the amount of code. In this way, the above encoding method can contribute to further improvements in the encoding process related to displacement vectors.
[0669] For example, information indicating a hierarchy using interpretation to generate the predicted residuals of the first displacement data may also be transmitted.
[0670] This allows the encoding device to transmit information indicating whether interpretation was used for each hierarchical level to which the three-dimensional points belong during encoding. This informs the decoding device, which receives and decodes the encoded displacement data, whether interpretation was used for each hierarchical level to which the three-dimensional points belong during encoding. This can contribute to the proper decoding of the encoded data. Specifically, it can contribute to ensuring that data from hierarchical levels encoded using interpretation is decoded using interpretation, and data from hierarchical levels encoded without interpretation is decoded without interpretation. Thus, the above encoding method can contribute to improvements in encoding processing related to displacement vectors.
[0671] Hereafter, an example of the decoding process in this embodiment will be described.
[0672] Figure 92 is a flowchart showing the decoding process in this embodiment.
[0673] The decoding process shown in Figure 92 is a method for decoding three-dimensional point displacement data performed by a decoding device.
[0674] In step S9201, the decoding device obtains the predicted residual by decoding the encoded data.
[0675] In step S9202, the decoding device generates a predicted value for the first displacement data by performing an interpretation using the first displacement data, which is the displacement data to be decoded, and the second displacement data from a different time period.
[0676] In step S9203, the decoding device generates first displacement data using the predicted residual and the generated predicted value.
[0677] As a result, the decoding device can appropriately decode the displacement data of three-dimensional points by decoding the prediction residual generated by interpretation. By using interpretation, it may be possible to reduce the amount of code when the difference between the first displacement data to be decoded and the second displacement data at a different time is relatively small, thereby potentially improving the decoding process. Thus, the above decoding method can contribute to improvements in the decoding process related to displacement vectors.
[0678] For example, when generating predicted values for the first displacement data, it is possible to determine whether or not to perform interpretation when generating the predicted values for the first displacement data. If it is determined that interpretation should be performed, the predicted values for the first displacement data are generated by performing interpretation. If it is determined that interpretation should not be performed, the predicted values for the first displacement data may be generated without performing interpretation.
[0679] This allows the decoding device to determine in advance whether or not to use interpretation to generate predicted values for the displacement data when decoding the first displacement data to be decoded, and to switch whether or not to use interpretation to generate predicted values for the displacement data according to the result of that determination. This may further reduce the amount of coding required. In this way, the above decoding method can contribute to further improvements in decoding processing related to displacement vectors.
[0680] For example, information may also be received indicating whether or not interpretation was used to generate the predicted residuals of the first displacement data.
[0681] This allows the decoding device to know whether the encoding device used interpretation during encoding by receiving information indicating whether or not interpretation was used during encoding. If the encoding device used interpretation during encoding, the decoding device can decode the displacement vector using interpretation; if the encoding device did not use interpretation during encoding, the decoding device can decode the displacement vector without interpretation. This enables the decoding device to properly decode the data encoded by the encoding device.
[0682] For example, a three-dimensional point may include multiple three-dimensional points, each of which belongs to one of several hierarchies. When generating predicted values for the first displacement data, for each of the one or more hierarchies to which the multiple three-dimensional points belong, it is determined whether or not to perform interpretation when generating predicted values for the first displacement data for the three-dimensional points belonging to that hierarchy. For the hierarchies to which it is determined that interpretation should be performed, the predicted values for the first displacement data are generated by performing interpretation. For the hierarchies to which it is determined that interpretation should not be performed, the predicted values for the first displacement data are generated without performing interpretation.
[0683] As a result, when decoding the first displacement data to be encoded, the decoding device can, in advance, determine whether or not to use interpretation to generate predicted values of the displacement data for each hierarchy to which the three-dimensional point belongs, and then switch whether or not to use interpretation to generate predicted values of the displacement data for each hierarchy to which the three-dimensional point belongs, according to the result of that determination. This has the potential to further reduce the amount of coding. Thus, the above decoding method can contribute to further improvements in coding processing related to displacement vectors.
[0684] For example, information indicating a hierarchy using interpretation to generate the predicted residuals of the first displacement data may also be received.
[0685] This allows the decoding device to receive information indicating whether interpretation was used for each hierarchy to which the three-dimensional points belong during encoding, thereby knowing whether the encoding device used interpretation for each hierarchy to which the three-dimensional points belong during encoding. The decoding device can then decode the displacement vector using interpretation for data in the hierarchy to which the encoding device used interpretation during encoding, and decode the displacement vector without interpretation for data in the hierarchy to which the encoding device did not use interpretation during encoding. As a result, the decoding device can appropriately decode the data encoded by the encoding device.
[0686] In the following sections, we will describe an example in which an encoding device or a decoding device calculates a predicted value of the displacement vector using the vectors of the three-dimensional points.
[0687] The explanation in Figure 62 shows that the predicted displacement vector of a three-dimensional point is generated by calculating the average value of the displacement vectors of adjacent three-dimensional points, or a weighted average value using distance information between three-dimensional points. Here, we will explain an example in which an encoding device or decoding device calculates the predicted displacement vector using a vector possessed by a three-dimensional point. One example of a vector possessed by a three-dimensional point is the normal vector of the three-dimensional point. The normal vector of a three-dimensional point may, for example, be the normal vector of a triangle included in the three-dimensional mesh, with the three-dimensional point as one of its vertices, but is not limited to this.
[0688] Figure 93 is an explanatory diagram showing an example of a method for calculating the predicted value of the displacement vector in this embodiment.
[0689] Figure 93 shows three-dimensional points a0, a1, a2, b0, b1, c0, c1, c2, and c3, corresponding to the hierarchy to which they belong (e.g., LoD).
[0690] In Figure 93, for example, if a three-dimensional point (also simply called a point) b0 is a point generated by subdivision from points a0 and a1, then point b0 may be a point predicted from points a0 and a1.
[0691] Furthermore, if point c2 is a point generated by subdivision from points a1 and b1, then point c2 may be a point predicted from points a1 and b1.
[0692] When an encoding or decoding device generates a predicted value of the displacement vector of point b0 to be encoded or decoded using the displacement vectors of points a0 and a1, it may also generate a predicted value of the displacement vector of point b0 using the normal vectors of each point.
[0693] For example, an encoding or decoding device may generate a predicted value for point b0 by preferentially using the displacement vectors of adjacent points that have normal vectors similar to the normal vector of point b0 to be encoded or decoded.
[0694] The similarity of normal vectors can be calculated, for example, using the dot product of the normal vectors. Alternatively, the similarity of normal vectors can be calculated, for example, using the cosine value of the angle Θ between the normal vectors.
[0695] For example, the similarity of the normal vectors between three-dimensional point b0 and three-dimensional point a0, cos_b0_v0, and the similarity of the normal vectors between three-dimensional point b0 and three-dimensional point a1, cos_b0_v1, can be expressed by the following equations (Equation 19).
[0696] cos_b0_v0 = normal(b0)・normal(a0) / (|normal(b0)|・|normal(a0)|) cos_b0_v1 = normal(b0)・normal(a1) / (|normal(b0)|・|normal(a1)|) (Equation 19)
[0697] Here, "•" indicates the dot product of vectors, and "||" indicates the magnitude of vectors.
[0698] In this case, if each normal vector is a unit vector, then |normal(b0)|=|normal(a0)|=|normal(a1)|=1, and the similarity is calculated using the dot product of the normal vectors as shown below (Equation 20).
[0699] cos_b0_v0 = normal(b0)・normal(a0) cos_b0_v1 = normal(b0)・normal(a1) (Equation 20)
[0700] Note that the similarity score can take values between -1 and 1.
[0701] To make the similarity score a positive value greater than or equal to 0, you may add +1. In other words, by adding +1 as shown below, the similarity score can take values between 0 and 2 (see (Equation 21) below).
[0702] cos_b0_v0 = normal(b0)・normal(a0) + 1 cos_b0_v1 = normal(b0)・normal(a1) + 1 (Equation 21)
[0703] As a method for calculating the predicted displacement vector of a three-dimensional point b0 using the above similarity, one can, for example, use a weighted average based on similarity (see (Equations 22) and (23) below). Specifically, the predicted displacement vector b0p of a three-dimensional point b0 can be calculated by a weighted average of the displacement vectors of three-dimensional points a0 and a1, based on their normal vectors.
[0704]
[0705]
[0706] Here, the predicted value b0p is the predicted value of the displacement vector of the three-dimensional point b0. D0 is the value of the displacement vector of the three-dimensional point a0, and D1 is the value of the displacement vector of the three-dimensional point a1.
[0707] Figure 94 is an explanatory diagram showing an example of a normal vector in this embodiment. Figure 94 shows examples of normal vectors for three-dimensional points a0, a1, and b0 (normal(a0), normal(a1), and normal(b0), respectively). In Figure 94, "normal(x)" represents the normal vector of three-dimensional point x.
[0708] In the example shown in Figure 94, since the similarity between normal(b0) and normal(a1) is relatively high, the encoding or decoding device may prioritize using the displacement vector value of the three-dimensional point a1 to generate a predicted value for the displacement vector of the three-dimensional point b0. This may allow the encoding or decoding device to improve its encoding efficiency.
[0709] As described above, the encoding or decoding device can calculate the predicted value of the displacement vector of a three-dimensional point b0 by calculating the predicted value of the displacement vector of the three-dimensional point b0 using a weighted average based on similarity using the normal vector, thereby prioritizing the displacement vector values of adjacent points that have a normal vector similar to that of the three-dimensional point b0. The fact that the normal vectors of two or more three-dimensional points are similar to each other indicates that there is a high probability that those two or more three-dimensional points lie on the same plane, and in that case, the displacement vector values are also likely to be similar. This is because displacement vectors tend to have a direction close to the direction of the normal vector. Therefore, as in this embodiment, by generating a predicted value of the displacement vector of the three-dimensional point to be encoded or decoded by prioritizing the displacement vector values of adjacent points that have a normal vector similar to that of the three-dimensional point to be encoded or decoded, it is possible to improve encoding efficiency.
[0710] Another example of a vector possessed by a three-dimensional point is its displacement vector. In other words, the encoding or decoding device may use the displacement vector of a three-dimensional point as the vector possessed by the three-dimensional point, instead of the normal vector of the three-dimensional point as described herein, to calculate the predicted value of the displacement vector of the three-dimensional point. In this case, the encoding or decoding device can calculate the predicted value of the displacement vector of the three-dimensional point to be encoded by a weighted average using the similarity of the displacement vectors. This may improve the encoding efficiency of the encoding or decoding device by prioritizing the displacement vector values of adjacent points whose displacement vectors are similar to the three-dimensional point to be encoded or decoded when generating the predicted value of the displacement vector of the three-dimensional point to be encoded or decoded. When the decoding device calculates the predicted value of the displacement vector of the three-dimensional point to be decoded, it may use the estimated value of the displacement vector of the three-dimensional point to calculate the similarity with other three-dimensional points, and then use a weighted average using that similarity.
[0711] Figure 95 is an explanatory diagram showing an example of a displacement vector in this embodiment.
[0712] In Figure 95, for example, if point c2 is a point generated by subdivision from points a1 and ab, then point c2 may be a point predicted from points a1 and b1. Figure 95 shows examples of normal vectors for three-dimensional points a1, b1 and c2. In Figure 95, "normal(x)" represents the normal vector of three-dimensional point x.
[0713] When an encoding or decoding device generates a predicted value of the displacement vector of point c2 to be encoded or decoded using the displacement vector of point a1 and point b1, it may also generate a predicted value of the displacement vector of point c2 using the normal vector of each point.
[0714] For example, an encoding or decoding device may preferentially use the displacement vectors of adjacent points that have normal vectors similar to the normal vector of point c2 to be encoded or decoded, in order to generate a predicted value for point c2.
[0715] For example, the similarity of the normal vectors between point c2 and point a1, cos_c2_v0, and the similarity of the normal vectors between point c2 and point b1, cos_c2_v1, can be expressed as follows, similar to the explanation in Figure 94.
[0716] cos_c2_v0 = normal(c2)・normal(a1) + 1 cos_c2_v1 = normal(c2)・normal(b1) + 1 (Equation 24)
[0717] As a method for calculating the predicted displacement vector of point c2 using the above similarity, for example, a weighted average based on similarity can be used, as shown in the following equation (see (Equations 25) and (26) below). Specifically, the predicted displacement vector c2p of point c2 can be calculated by a weighted average of the displacement vectors of points a1 and b1, based on their normal vectors.
[0718]
[0719]
[0720] Here, the predicted value c2p is the predicted value of the displacement vector at point c2. D0 is the value of the displacement vector at point a1, and D1 is the value of the displacement vector at point b1.
[0721] In the example shown in Figure 95, since the similarity between normal(c2) and normal(a1) is relatively high, the encoding or decoding device may prioritize using the displacement vector value of point a1 to generate a predicted value for the displacement vector of point c2. This may allow the encoding or decoding device to improve its encoding efficiency.
[0722] Furthermore, the method for calculating the similarity cos_v_v0 between the three-dimensional point v to be encoded or decoded and its adjacent point v0, and the similarity cos_v_v1 between the three-dimensional point v and its adjacent point v1, may be as shown below (Equation 27).
[0723] cos_v_v0 = max(0.0, normal(v)・normal(v0)) cos_v_v1 = max(0.0, normal(v)・normal(v1)) (Equation 27)
[0724] Here, max(A, B) indicates that it returns the larger of A and B. That is, max(0.0, B) returns the value of B if B is positive, and returns 0.0 if B is negative, so the range of the similarity value can be made to be between 0 and 1 using the above (Equation 27). When the similarity is a negative value, that is, between -1 and less than 0, it indicates that the angle between the normal vectors is greater than 90 degrees. In this case, the three-dimensional points may belong to different planes, and the displacement vectors may have different values. Therefore, the encoding or decoding device may be able to improve encoding efficiency by setting the similarity to 0 when the similarity is less than 0, so that it is not used in calculating the predicted value.
[0725] For example, when normal(v)・normal(v0) = -0.5, it is indicated that the angle between the normal vector of the three-dimensional point v to be encoded or decoded and the normal vector of its adjacent point v0 is 90 degrees or greater. In this case, the similarity cos_v_v0 = 0 is calculated using (Equation 27) above. As a result, the value of w in (Equation 26) above becomes 0, and the displacement vector of the adjacent point v0 is no longer used in calculating the predicted value of the three-dimensional point v. This improves the accuracy of the predicted value of v, which may allow the encoding or decoding device to improve its encoding efficiency.
[0726] The similarity cos_v_v0 between the three-dimensional point v to be encoded or decoded and its adjacent point v0, and the similarity cos_v_v1 between the three-dimensional point v and its adjacent point v1 may be calculated as shown below (Equation 28).
[0727] cos_v_v0 = max(0.0, normal(v)・normal(v0) + offset) cos_v_v1 = max(0.0, normal(v)・normal(v1) + offset) (Equation 28)
[0728] Thus, the encoding or decoding device may control the range of possible similarity values by adding an offset to the dot product of the normal vector normal(v) and the normal vector normal(v0), and to the dot product of the normal vector normal(v) and the normal vector normal(v1).
[0729] For example, in the above equation (28), when offset=0, the range of similarity values is 0 or more and 1 or less, but when offset=-0.5, the range of similarity values becomes 0 or more and 0.5 or less. When offset=0, adjacent points where normal(v)・normal(v0) is 0.5 (that is, adjacent points where the angle between the normal vector of point v and the normal vector of the adjacent point is 45 degrees) are used in calculating the predicted value. On the other hand, when offset=-0.5, the above adjacent points where normal(v)・normal(v0) is 0.5 have a similarity of 0, so they can be excluded from calculating the predicted value.
[0730] Furthermore, when offset=0, adjacent points where normal(v) and normal(v0) are -0.5 (i.e., adjacent points where the angle between the normal vector of point v and the normal vector of the adjacent point is 135 degrees) are not used in calculating the predicted value. On the other hand, when offset=1.0, adjacent points where normal(v) and normal(v0) are -0.5 have a similarity of 0.5, so they can be used in calculating the predicted value.
[0731] In this way, by switching the offset value, it is possible to switch the displacement vectors of adjacent points used to calculate the predicted values, thereby improving coding efficiency.
[0732] The offset value may be switched at the sequence level, frame level, submesh level, or per LoD. For example, when switching the offset per LoD, the offset value may be made smaller as the LoD hierarchy deepens. For example, as the LoD hierarchy deepens, the shape changes become larger, and the change in the normal vector for each 3D point may become larger. Therefore, the encoding or decoding device may be able to improve prediction accuracy by preferentially selecting points with normal vectors that have a relatively high similarity to the normal vector of point v, which is the target of encoding or decoding, and calculating the predicted value.
[0733] Specifically, the encoding or decoding device can, for example, set offset=0 for LoD1 and offset=-0.5 for LoD2. By doing so, the encoding or decoding device can potentially improve encoding efficiency by using adjacent points where the angle between the normal vector of point v to be encoded or decoded and the normal vector of the adjacent point is less than 90 degrees for calculating the predicted value in LoD1, which has a relatively shallow hierarchy, and by using adjacent points where the angle between the normal vector of point v and the normal vector of the adjacent point is less than 45 degrees for calculating the predicted value in LoD2, which has a relatively deep hierarchy. In this way, the encoding or decoding device can potentially improve encoding efficiency by switching the offset value for each LoD hierarchy.
[0734] The offset value may be added to the header information at the sequence level, frame level, submesh level, or LoD level. For example, the encoding device may add an offset for each LoD to the header. The decoding device can then correctly decode the bitstream by decoding the LoD-specific offset added to the header and using the decoded offset for each LoD to calculate the predicted value. If prediction is not used when encoding or decoding three-dimensional points belonging to LoD0, the offset for LoD0 does not need to be added to the bitstream. This reduces the amount of code in the header.
[0735] If both cos_v_v0 and cos_v_v1 are 0, the predicted value may be set to 0. In this case, the encoding or decoding device will be performing the encoding or decoding of the displacement vector of the three-dimensional point v without prediction. If the similarity between the adjacent points at both ends (in other words, adjacent points v0 and v1) is 0, the displacement vectors of both points may be different, so it may be possible to improve encoding efficiency by not making a prediction.
[0736] If both cos_v_v0 and cos_v_v1 are 0, the encoding or decoding device may use the average value of the displacement vectors of points v0 and v1 to generate a predicted value for the three-dimensional point v to be encoded or decoded. If the similarity between the adjacent points at both ends is 0, it is difficult to determine which displacement vector value to use, so it may be decided to use both equally and use their average value (i.e., the average value of the displacement vectors of points v0 and v1). This may improve encoding efficiency compared to the case without prediction. In this case, setting the offset to -1 will always make both cos_v_v0 and cos_v_v1 0, and the device can switch to calculating the predicted value using the average value. By switching the offset value in this way, the encoding or decoding device can switch whether to calculate the predicted value of the displacement vector of the three-dimensional point v to be encoded or decoded using the average value of the displacement vectors of adjacent points, or based on the normal vector.
[0737] Figure 96 is an explanatory diagram showing an example of a method for calculating the conversion coefficient in this embodiment.
[0738] Figure 96 shows an example of calculating the transformation coefficient by subtracting the predicted value from the displacement vector of each three-dimensional point.
[0739] Figure 96 shows the transformation coefficients a0c, a1c, a2c, b0c, b1c, c0c, c1c, c2c, and c3c, which are three-dimensional points a0, a1, a2, b0, b1, c0c, c1c, c2c, and c3c, in correspondence with the hierarchy to which they belong (e.g., LoD).
[0740] For example, the conversion coefficient b0c is calculated by the following equation (29).
[0741] b0c = b0 - b0p (Equation 29)
[0742] Here, the predicted value b0p is the predicted value of point b0. The predicted value b0p of point b0 may be calculated using the normal vectors of points b0, a0, and a1, for example, as described in this embodiment.
[0743] Furthermore, for example, the conversion coefficient c²c is calculated by the following (Equation 30).
[0744] c2c = c2 - c2p (Equation 30)
[0745] Here, the predicted value c2p is the predicted value of point c2. The predicted value c2p of point c2 may be calculated using the normal vectors of points c2, a1, and b1, for example, as described in this embodiment.
[0746] Furthermore, the encoding or decoding device may apply an encoding scheme (lifting transform) that calculates the transformation coefficients of the displacement vectors of three-dimensional points contained in the lower layers of the LoD and feeds these transformation coefficients back to the upper layers for encoding. By applying this lifting transform, the encoding or decoding device can collect the transformation coefficients of the low-frequency components of the displacement vectors in the upper layers and the transformation coefficients of the high-frequency components of the displacement vectors in the lower layers. Then, for example, the encoding efficiency can be improved by reducing the amount of information in the transformation coefficients of the high-frequency components in the lower layers through quantization.
[0747] In the explanation of Figure 96, the conversion coefficient may be read as the predicted residual. That is, the predicted residual b0c of point b0 may be calculated by (Equation 29) above. Here, the predicted value b0p of point b0 may be calculated using the normal vectors of points b0, a0 and a1, for example, as described in this embodiment. Also, the predicted residual c2p of point c2 may be calculated by (Equation 30) above. Here, the predicted value c2p of point c2 may be calculated using the normal vectors of points c2, a1 and b1, for example, as described in this embodiment.
[0748] Figure 97 is a flowchart showing an example of the encoding process in this embodiment.
[0749] The encoding process shown in Figure 97 is a method for encoding three-dimensional point displacement data, performed by the encoding device. The encoding algorithm shown in Figure 97 is executed by the encoding device.
[0750] In step S9701, the encoding device generates three-dimensional points by subdivision using each three-dimensional point of the base mesh (see Figure 37).
[0751] In step S9702, the encoding device generates a displacement vector from the difference between each three-dimensional point of the base mesh and the three-dimensional points generated by subdivision in step S9701, and the corresponding three-dimensional points in the input three-dimensional mesh frame.
[0752] In step S9703, the encoding device calculates the normal vector of each three-dimensional point after sub-division (step SS9701). For example, if three-dimensional points a, b, and c form a triangular mesh, the encoding device may calculate the normal vector from the information of that triangular mesh and assign it to the normal vectors of three-dimensional points a, b, and c. For example, if the vector from three-dimensional point a to three-dimensional point b is vector Vab, and the vector from three-dimensional point a to three-dimensional point c is vector Vac, the surface normal may be calculated by the cross product of vector Vab and vector Vac. Alternatively, the encoding device may calculate the surface normals of multiple triangular meshes to which point a belongs, add these values together to normalize the vector, and use that as the normal vector for point a.
[0753] The method for calculating the normal vector can be anything. For example, if a three-dimensional point d is the vertex of two or more triangular meshes, the normal vector can be calculated from each of those two or more triangular meshes, and the average of those normal vectors can be used as the normal vector of the three-dimensional point d.
[0754] In step S9704, the encoding device assigns the three-dimensional points after sub-partitioning (step SS9701) to one or more LoD levels, and generates conversion coefficients representing components from low-frequency to high-frequency components by applying a lifting transform or the like to the displacement vectors of the three-dimensional points. In this case, as shown in this embodiment, the encoding device may generate predicted values of the displacement vectors using the normal vectors of each three-dimensional point. This can improve encoding efficiency.
[0755] In step S9705, the encoding device can quantize the conversion coefficients generated in step S9704 for each level of the Level of Data (LoD).
[0756] In step S9706, the encoding device can encode the conversion coefficients after quantization in step S9705 using a video codec (also called video encoding) or perform arithmetic encoding.
[0757] Through the series of processes shown in Figure 97, the encoding device can appropriately encode the displacement data of three-dimensional points.
[0758] Figure 98 is a flowchart showing an example of the decoding process in this embodiment.
[0759] The decoding process shown in Figure 98 is a method for decoding three-dimensional point displacement data performed by the decoding device. The decoding algorithm shown in Figure 98 is executed by the decoding device.
[0760] In step S9801, the decoding device can decode the conversion coefficients using a video codec (also called video decoding) or perform arithmetic decoding. The conversion coefficients that are the target of processing in step S9801 may be in an encoded state. If the conversion coefficients that are the target of processing in step S9801 are in an encoded state, the decoding device can decode the conversion coefficients using a video codec (also called video decoding) or perform arithmetic decoding.
[0761] In step S9802, the decoding device can de-quantize the conversion coefficients for each LoD level. The conversion coefficients that are the target of processing in step S9802 may be in a quantized state. If the conversion coefficients that are the target of processing in step S9802 are in a quantized state, the decoding device can de-quantize the conversion coefficients for each LoD level.
[0762] In step S9803, the decoding device generates three-dimensional points by subdivision using each three-dimensional point of the base mesh (see Figure 37).
[0763] In step S9804, the decoding device calculates the normal vector of each three-dimensional point after sub-division (step SS9803). For example, if three-dimensional points a, b, and c form a triangular mesh, the decoding device may calculate the normal vector from the information of that triangular mesh and assign it to the normal vectors of three-dimensional points a, b, and c. For example, if the vector from three-dimensional point a to three-dimensional point b is vector Vab, and the vector from three-dimensional point a to three-dimensional point c is vector Vac, the surface normal may be calculated by the cross product of vector Vab and vector Vac. Alternatively, the decoding device may calculate the surface normals of multiple triangular meshes to which point a belongs, add these values together, and use the normal vector of point a as the normal vector.
[0764] The method for calculating the normal vector can be anything. For example, if a three-dimensional point d is the vertex of two or more triangular meshes, the normal vector can be calculated from each of the two or more triangular meshes, and the average of these normal vectors can be used as the normal vector of the three-dimensional point d.
[0765] In step S9805, the decoding device recovers the value of the displacement vector from the transformation coefficients by applying an inverse lifting transform or the like. In this case, as shown in this embodiment, the decoding device may generate predicted values of the displacement vector using the normal vectors of each three-dimensional point. This allows the decoding device to correctly decode the bitstream in which the encoding device has improved encoding efficiency by generating predicted values of the displacement vector using the normal vectors.
[0766] In step S9806, the decoding device reconstructs the decoded mesh using the recovered displacement vectors and the base mesh.
[0767] Through the series of processes shown in Figure 98, the decoding device can appropriately decode the displacement data of three-dimensional points.
[0768] The above example shows a method in which an encoding or decoding device calculates normal vectors using information from the three-dimensional points constituting the triangular mesh and generates predicted displacement vectors using those normal vectors, but this is not necessarily the only method. For example, if the normal vectors for each three-dimensional point are calculated in advance, the encoding device may use those normal vectors to generate predicted displacement vectors. In that case, the decoding device also needs to calculate the predicted values using the same normal vectors, so the encoding device may encode the normal vectors for each three-dimensional point and add them to the bitstream. This allows the decoding device to decode the normal vectors for each three-dimensional point from the bitstream and use those normal vectors to calculate predicted displacement vectors, thereby generating the same values as the predicted values generated by the encoding device and correctly decoding the bitstream.
[0769] Furthermore, if a normal vector does not exist or cannot be calculated, the encoding or decoding device may switch prediction methods depending on whether the normal vector could be calculated. For example, instead of generating a predicted displacement vector using the normal vector, it may use the average value of the displacement vectors of adjacent points. This could potentially improve encoding efficiency by allowing the encoding or decoding device to calculate the predicted displacement vector using an alternative method even if the normal vector could not be calculated.
[0770] Furthermore, the encoding device may add a method for predicting displacement vectors to the bitstream at the sequence level, frame level, submesh level, or hierarchy level (more specifically, Line of D). This will be explained in more detail later.
[0771] Figure 99 is an explanatory diagram showing an example of the definition of the displacement vector prediction method in this embodiment. Figure 100 is an explanatory diagram showing an example of a header in this embodiment. Figure 101 is an explanatory diagram showing an example of a header in this embodiment.
[0772] For example, the method for predicting the displacement vector may be defined as disp_pred_type as shown in the table in Figure 99. In the table in Figure 99, "disp_pred_type=0" indicates that the mean value (in other words, a simple mean) is used to predict the displacement vector, and "disp_pred_type=1" indicates that a weighted mean value using the normal vector is used to predict the displacement vector. The encoding device may then add disp_pred_type to the header, etc., at the sequence level, frame level, submesh level, or hierarchy level (more specifically, LoD).
[0773] The encoding device may add `disp_pred_type` to the header of a submesh, for example, as shown in the header example in Figure 100. If the encoding device encodes the displacement vectors of three-dimensional points within a submesh using predicted values based on averages, it may add `disp_pred_type=0` to the header related to that submesh. This allows the decoding device to determine that the displacement vectors within that submesh were encoded using predicted values based on averages, and the decoding device can correctly decode the bitstream by generating predicted values based on averages. For example, if the displacement vectors of three-dimensional points within a submesh are encoded using a weighted average with normal vectors as predicted values, as described in this embodiment, `disp_pred_type=1` may be added to the header related to that submesh. This allows the decoding device to determine that the displacement vectors within that submesh were encoded using a weighted average with normal vectors as predicted values, and the decoding device can correctly decode the bitstream by using a weighted average with normal vectors as predicted values.
[0774] Furthermore, the encoding device may, for example, add disp_pred_type to the bitstream for each level of the LoD, as shown in the header example in Figure 101, and switch the displacement vector prediction method for each LoD. For example, the encoding device may compare the code weight when ...
Claims
1. A method for encoding parameters used in the encoding process of three-dimensional points, comprising: obtaining a first parameter set having a first number of layers and a second parameter set having a second number of layers that is referenced in the encoding of the first parameter set; executing a first process when the first number of layers is greater than the second number of layers; and executing a second process different from the first process when the first number of layers is not greater than the second number of layers.
2. The encoding method according to claim 1, wherein the first process includes a process of referring to a second parameter of a higher hierarchy than the first hierarchy included in the second parameter set for a first parameter of a hierarchy included in the first parameter set, and the second process includes a process of referring to a second parameter of the same hierarchy as the first parameter of each hierarchy included in the first parameter set for a first parameter of that hierarchy included in the second parameter set.
3. The encoding method according to claim 2, wherein the first process includes a process of referring to the second parameter of the highest level included in the second parameter set for each first parameter of the level included in the first parameter set.
4. The encoding method according to claim 2 or 3, wherein the first process further includes a process for calculating difference information showing the difference between the first parameter of the first hierarchy and the referenced second parameter, and the second process further includes a process for calculating difference information showing the difference between the first parameter of the first hierarchy and the referenced second parameter.
5. The encoding method according to claim 1, wherein the first parameter set is included in the frame parameter set of the frame to which the three-dimensional point belongs, and the second parameter set is included in the sequence parameter set of the sequence to which the three-dimensional point belongs.
6. The encoding method according to claim 1, wherein the first parameter set is included in the mesh patch data unit of the mesh patch to which the three-dimensional point belongs, and the second parameter set is included in the frame parameter set of the frame to which the three-dimensional point belongs.
7. The encoding method according to any one of claims 1 to 6, wherein the first parameter set includes parameters for quantization processing, and the second parameter set includes parameters for quantization processing.
8. A decoding method for parameters used in decoding a three-dimensional point, comprising: obtaining a first parameter set having a first number of layers and a second parameter set having a second number of layers that is referenced in decoding the first parameter set; executing a first process when the first number of layers is greater than the second number of layers; and executing a second process different from the first process when the first number of layers is not greater than the second number of layers.
9. The decoding method according to claim 8, wherein the first process includes a process of referring to a second parameter of a higher hierarchy than the first hierarchy included in the second parameter set for a first parameter of a hierarchy included in the first parameter set, and the second process includes a process of referring to a second parameter of the same hierarchy as the first parameter of each hierarchy included in the first parameter set for a first parameter of that hierarchy included in the second parameter set.
10. The decoding method according to claim 9, wherein the first process includes a process of referring to the second parameter of the highest level included in the second parameter set for each first parameter of the level included in the first parameter set.
11. The decoding method according to claim 9 or 10, wherein the first process further includes decoding difference information from a bitstream showing the difference between the first parameter of one hierarchy and the referenced second parameter, and the second process further includes decoding difference information from a bitstream showing the difference between the first parameter of one hierarchy and the referenced second parameter, and adding the difference shown in the decoded difference information to the second parameter.
12. The decoding method according to claim 8, wherein the first parameter set is included in the frame parameter set of the frame to which the three-dimensional point belongs, and the second parameter set is included in the sequence parameter set of the sequence to which the three-dimensional point belongs.
13. The decoding method according to claim 8, wherein the first parameter set is included in the mesh patch data unit of the mesh patch to which the three-dimensional point belongs, and the second parameter set is included in the frame parameter set of the frame to which the three-dimensional point belongs.
14. The decoding method according to any one of claims 8 to 13, wherein the first parameter set includes parameters for quantization processing, and the second parameter set includes parameters for quantization processing.
15. An encoding device comprising: a memory; and a circuit that can access the memory, wherein the circuit, in operation, performs a method for encoding parameters used in encoding processing of three-dimensional points; the encoding method obtains a first parameter set having a first number of layers and a second parameter set having a second number of layers that is referenced in encoding the first parameter set; if the first number of layers is greater than the second number of layers, it performs a first process; and if the first number of layers is not greater than the second number of layers, it performs a second process different from the first process.
16. A decoding device comprising: a memory; and a circuit that can access the memory, wherein the circuit, in operation, performs a decoding method for parameters used in decoding a three-dimensional point; the decoding method obtains a first parameter set having a first number of layers and a second parameter set having a second number of layers that is referenced in decoding the first parameter set; if the first number of layers is greater than the second number of layers, it performs a first process; and if the first number of layers is not greater than the second number of layers, it performs a second process different from the first process.
Citation Information
Patent Citations
Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
WO2020196677A1
Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
WO2021066163A1