Encoding method, decoding method, encoding device, and decoding device

The method enhances three-dimensional mesh encoding and decoding by optimizing reference selection and parameter determination, improving efficiency and reducing redundant signaling and data loss.

WO2026094856A1PCT designated stage Publication Date: 2026-05-07PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
Filing Date
2025-10-27
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing methods for encoding and decoding three-dimensional meshes are inefficient and lack consistency in reference selection, leading to unnecessary signaling and potential data loss.

Method used

A method and device that determine whether a data unit is interpredictive, specify different reference destinations for each parameter type, and encode or decode based on these parameters, ensuring consistent interpretation and reducing redundant transmissions.

Benefits of technology

Improves encoding efficiency by selecting optimal references, maintaining data consistency, and reducing bitrate through differential parameter representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025037620_07052026_PF_FP_ABST
    Figure JP2025037620_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an encoding method that is executed by an encoding device (3000), the encoding method comprising: determining whether or not inter prediction is employed as a method for encoding a data unit of interest including a mesh data unit (MDU) (S3081); identifying reference source data units indicated by reference information (S3082) if it is determined that inter prediction is employed (Yes in S3081); determining parameters using different reference sources for different parameter types (S3083); and encoding the data unit of interest on the basis of the determined parameters (S3084).
Need to check novelty before this filing date? Find Prior Art

Description

Symbolization method, decoding method, symbolization device, and decoding device

[0001] The present disclosure relates to a symbolization method and the like.

[0002] In Patent Document 1, a method and a device for symbolizing and decoding three-dimensional mesh data have been proposed.

[0003] Japanese Patent Application Laid-Open No. 2006-187015

[0004] Further improvement is desired for the symbolization or decoding process related to three-dimensional meshes. The purpose of the present disclosure is to improve the symbolization or decoding process related to three-dimensional meshes.

[0005] The symbolization method according to one aspect of the present disclosure is a symbolization method executed by a symbolization device. The method determines whether the symbolization method of a data unit to be symbolized including a mesh data unit (MDU) is inter prediction. When it is determined that it is inter prediction, the reference destination data unit indicated by the reference information is specified, parameters are determined by varying the reference destination for each type of parameter, and the data unit to be symbolized is symbolized based on the determined parameters.

[0006] The decoding method according to one aspect of the present disclosure is a decoding method executed by a decoding device. The method determines whether the decoding method of a data unit to be decoded including a mesh data unit (MDU) is inter prediction. When it is determined that it is inter prediction, the reference destination data unit indicated by the reference information is specified, parameters are determined by varying the reference destination for each type of parameter, and the data unit to be decoded is decoded based on the determined parameters.

[0007] These general or specific aspects may be realized by a system, a device, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be realized by any combination of a system, a device, an integrated circuit, a computer program, and a recording medium.

[0008] The present disclosure can contribute to the improvement of symbolization processing and the like related to three-dimensional meshes.

[0009] This is a conceptual diagram showing a three-dimensional mesh according to Embodiment 1. This is a conceptual diagram showing the basic elements of a three-dimensional mesh according to Embodiment 1. This is a conceptual diagram showing a mapping according to Embodiment 1. This is a block diagram showing an example configuration of an encoding / decoding system according to Embodiment 1. This is a block diagram showing an example configuration of an encoding device according to Embodiment 1. This is a block diagram showing another example configuration of an encoding device according to Embodiment 1. This is a block diagram showing an example configuration of a decoding device according to Embodiment 1. This is a block diagram showing another example configuration of a decoding device according to Embodiment 1. This is a conceptual diagram showing an example configuration of a bitstream according to Embodiment 1. This is a conceptual diagram showing yet another example configuration of a bitstream according to Embodiment 1. This is a block diagram showing a specific example of an encoding / decoding system according to Embodiment 1. This is a conceptual diagram showing an example configuration of point cloud data according to Embodiment 1. This is a conceptual diagram showing an example data file for point cloud data according to Embodiment 1. This is a conceptual diagram showing an example configuration of mesh data according to Embodiment 1. This is a conceptual diagram showing an example data file for mesh data according to Embodiment 1. This is a conceptual diagram showing the types of three-dimensional data according to Embodiment 1. This is a block diagram showing an example configuration of a three-dimensional data encoder according to Embodiment 1. This is a block diagram showing an example configuration of a three-dimensional data decoder according to Embodiment 1. This is a block diagram showing another example configuration of a three-dimensional data encoder according to Embodiment 1. This is a block diagram showing another configuration example of the three-dimensional data decoder according to Embodiment 1. This is a conceptual diagram showing a specific example of the encoding process according to Embodiment 1. This is a conceptual diagram showing a specific example of the decoding process according to Embodiment 1. This is a block diagram showing an implementation example of the encoding device according to Embodiment 1. This is a block diagram showing an implementation example of the decoding device according to Embodiment 1. This is a block diagram showing another configuration example of the encoding / decoding system according to Embodiment 1. This is a block diagram showing another configuration example of the encoding device according to Embodiment 1. This is a block diagram showing yet another configuration example of the decoding device according to Embodiment 1. This is a flow diagram showing the processing of the encoding device according to Embodiment 1.This is an explanatory diagram conceptually showing the encoding of a mesh frame according to Embodiment 1. This is a flowchart showing the processing of a decoding device according to Embodiment 1. This is an explanatory diagram conceptually showing the decoding of a mesh frame according to Embodiment 1. This is a block diagram showing an example of the configuration of a decoding device according to Embodiment 1. This is a block diagram showing an example of the configuration of an encoding device according to Embodiment 1. This is a block diagram showing an example of the configuration of a decoding device according to Embodiment 1. This is an explanatory diagram showing an example of subdivision according to Embodiment 1. This is an explanatory diagram showing an example of vertex displacement after subdivision according to Embodiment 1. This is an explanatory diagram showing an example of vertices of the original mesh according to Embodiment 1. This is an explanatory diagram showing an example of a mesh according to Embodiment 1. This is an explanatory diagram showing an example of subdivision of a mesh according to Embodiment 1. This is a first explanatory diagram showing an example of packing displacement information into an image frame according to Embodiment 1. This is a second explanatory diagram showing an example of packing displacement information into an image frame according to Embodiment 1. This is a third explanatory diagram showing an example of packing displacement information into an image frame according to Embodiment 1. This is a block diagram showing a detailed example of the configuration of a decoding device according to Embodiment 1. This is an explanatory diagram showing the coordinates of vertices in a three-dimensional mesh according to Embodiment 1. This is an explanatory diagram showing prediction information according to Embodiment 1. This is a block diagram showing an example of the configuration of an encoding device according to Embodiment 1. This is a block diagram showing an example configuration of a decoding device according to Embodiment 1. This is a block diagram showing an example configuration of an encoding device according to Embodiment 1. This is a block diagram showing an example configuration of a decoding device according to Embodiment 1. This is an explanatory diagram showing a method for generating LoD in Embodiment 1. This is an explanatory diagram showing a method for generating LoD in Embodiment 1. This is an explanatory diagram showing a method for generating predicted values ​​of displacement vectors in Embodiment 1. This is an explanatory diagram showing an example of calculating predicted values ​​in Embodiment 1. This is an explanatory diagram showing an example of calculating transformation coefficients in Embodiment 1. This is an explanatory diagram showing an example of lifting transformation processing in Embodiment 1. This is an explanatory diagram showing an example of lifting transformation processing in Embodiment 1. This is an explanatory diagram showing an example of encoding by subtracting the lifting offset from the transformation coefficient of the displacement vector of each three-dimensional point according to Embodiment 1. This is a block diagram showing another example configuration of an encoding device according to Embodiment 1.This is a block diagram showing an example of the configuration of a preprocessor according to Embodiment 1. This is a block diagram showing another example of the configuration of a decoding device according to Embodiment 1. This is a block diagram showing a specific example of the configuration of a decoding device according to Embodiment 1. This is a conceptual diagram showing a specific example of a bitstream according to Embodiment 1. This is a block diagram showing yet another example of the configuration of an encoding device according to Embodiment 1. This is a block diagram showing yet another example of the configuration of a decoding device according to Embodiment 1. This is an explanatory diagram showing an example of syntax in Embodiment 1. This is an explanatory diagram showing an example of syntax in Embodiment 1. This is an explanatory diagram showing an example of syntax in Embodiment 1. This is an explanatory diagram showing an example of syntax in Embodiment 1. This is an explanatory diagram showing an example of syntax in Embodiment 1. This is an explanatory diagram showing an example of syntax in Embodiment 1. This is an explanatory diagram showing an example of syntax in Embodiment 1. This is an explanatory diagram showing an example of syntax in Embodiment 1. This is an explanatory diagram showing an example of syntax in Embodiment 1. This is an explanatory diagram showing an example of syntax in Embodiment 1. This figure shows the subdivision-related syntax in the sequence parameter set according to Embodiment 1. This figure shows the subdivision-related syntax in the frame parameter set according to Embodiment 1. This figure shows the subdivision-related syntax in the mesh patch data unit according to Embodiment 1. This figure shows the syntax of Subdivision_method according to Embodiment 1. This flowchart shows the procedure for processing the three-dimensional mesh of the hierarchy to be encoded in the encoding device according to Embodiment 2. This flowchart shows the procedure for processing the three-dimensional mesh of the hierarchy to be decoded in the decoding device according to Embodiment 2. This flowchart shows the encoding procedure for the number of divisions, division method, quantization parameters and transformation parameters in the encoding device according to Embodiment 2.This is a flowchart showing the decoding procedure for the number of divisions, division method, quantization parameters, and transformation parameters in the decoding device according to Embodiment 2. This is a diagram showing the syntax configuration according to Example 1 according to Embodiment 2. This is a diagram showing the syntax configuration according to Example 2 according to Embodiment 2. This is a diagram showing the syntax configuration according to Example 3 according to Embodiment 2. This is a diagram showing the syntax configuration according to Example 4 of this embodiment according to Embodiment 2. This is a diagram showing the syntax configuration according to Example 5 of this embodiment according to Embodiment 2. This is a flowchart showing the procedure for processing the three-dimensional mesh of the hierarchy to be encoded in the encoding device according to Embodiment 2. This is a flowchart showing the procedure for processing the three-dimensional mesh of the hierarchy to be encoded in the decoding device according to Embodiment 2. This is a flowchart showing the first example of the procedure for processing the three-dimensional mesh of the hierarchy to be encoded in the encoding device according to Embodiment 2. This is a flowchart showing the first example of the procedure for processing the three-dimensional mesh of the hierarchy to be decoded in the decoding device according to Embodiment 2. This is a flowchart showing the second example of the procedure for processing the three-dimensional mesh of the hierarchy to be encoded in the encoding device according to Embodiment 2. This is a flowchart showing the syntax configuration according to Example 6 according to Embodiment 2. This figure shows the syntax configuration according to Example 7 in Embodiment 2. This figure shows the syntax configuration according to Example 8 in Embodiment 2. This figure shows the syntax configuration according to Example 9 in Embodiment 2. This figure shows the syntax configuration according to Example 10 in Embodiment 2. This figure shows an example of the configuration of the encoding device in Embodiment 2. This flowchart shows an example of the encoding method by the encoding device in Embodiment 2. This figure shows an example of the configuration of the decoding device in Embodiment 2. This flowchart shows an example of the decoding method by the decoding device in Embodiment 2. This shows the configuration of the mesh sequence, frame and submesh, and the metadata configuration related to decoding and reconstructing the corresponding encoded data in Embodiment 3. This shows the reference configuration of the IMDU in Embodiment 3.This document shows the reference configuration of the MMDU according to Embodiment 3. This document shows the correspondence between the metadata unit of the reference frame and the metadata unit of the frame to be processed, and the positioning of parameters related to subdivision according to Embodiment 3. This document shows an example of processing when subdivision is performed multiple times according to Embodiment 3. This document shows an example of decoding processing when subdivision is used multiple times according to Embodiment 3. This document shows the relationship between parameter copying and signaling in the reference data unit and the data unit to be processed according to Embodiment 3. This document shows an example of the syntax of the intermesh data unit (IMDU / MMDU) according to Embodiment 3. This document shows the encoding procedure of Method A according to Embodiment 3. This document shows the decoding procedure of Method A according to Embodiment 3. This document shows an example of the derivation process of the subdivision method using subdivision_iteration and subdivision_iterationCount according to Embodiment 3. This document shows the syntax configuration related to the parameter derivation process according to Embodiment 3. This document shows the syntax configuration related to the parameter derivation process (Method B) according to Embodiment 3. This document shows a conceptual diagram of Method B according to Embodiment 3. This shows the decoding process procedure for Method B according to Embodiment 3. This is a diagram of the derivation process syntax for the subdivision method according to Embodiment 3. This is a diagram of the concept of switching operations using subdivision_method_override_flag according to Embodiment 3. This is a diagram of an example of the syntax for an intermesh data unit (IMDU / MMDU) according to Embodiment 3. This is a diagram of the derivation process syntax for the subdivision method according to Embodiment 3 (configuration using layer-specific override flags). This is a diagram of a specific example of switching between copying and signaling based on layer-specific override flags (subdivision_method_override_flag[i]) according to Embodiment 3. This is a diagram of an example of the derivation process syntax for the subdivision method according to Embodiment 3. This is a diagram of the syntax operation concept according to Embodiment 3. This is a diagram of the syntax configuration related to the derivation process of the subdivision method according to Embodiment 3. This is a diagram of the operation concept using a fixed reference method according to Embodiment 3.This figure shows an overall example of parameter derivation and overwriting processing according to Embodiment 3. This figure shows an example of the configuration of the encoding device according to Embodiment 3. This flowchart shows an example of the encoding method by the encoding device according to Embodiment 3. This figure shows an example of the configuration of the decoding device according to Embodiment 3. This flowchart shows an example of the decoding method by the decoding device according to Embodiment 3.

[0010] <Summary of this disclosure> For example, a three-dimensional (3D) mesh is used in computer graphics images. For example, a computer graphics image may consist of multiple frames that are different in time, and each frame may be represented by a three-dimensional mesh.

[0011] Furthermore, a three-dimensional mesh consists of vertex information indicating the position of each of several vertices in three-dimensional space, connection information indicating the connections between the multiple vertices, and attribute information indicating the attributes of each vertex or face. Each face is constructed according to the connection relationships of the multiple vertices. Various computer graphics images can be represented using such a three-dimensional mesh.

[0012] Furthermore, efficient encoding and decoding of three-dimensional meshes are expected for the transmission and storage of these meshes. Arithmetic coding and decoding may be used for efficient encoding and decoding of three-dimensional meshes.

[0013] Further improvements are desired in the encoding or decoding of three-dimensional data. This disclosure aims to improve the encoding or decoding of three-dimensional data.

[0014] The following describes examples of inventions that can be obtained from the disclosures in this specification, and explains the effects and other benefits that can be obtained from such inventions.

[0015] An encoding method according to a first aspect of this disclosure is an encoding method performed by an encoding device, which determines whether the encoding method of a data unit to be encoded, including a mesh data unit (MDU), is interpredictive or not, and if it is determined to be interpredictive, identifies a reference data unit indicated by reference information, determines parameters by specifying different references for each type of parameter, and encodes the data unit to be encoded based on the determined parameters.

[0016] This allows for the selection of the optimal reference source for each parameter type, suppressing unnecessary signaling. Furthermore, it ensures consistent interpretation of references and improves the stability of the decoding results. As a result, encoding efficiency can be improved.

[0017] The encoding method according to a second aspect of this disclosure is the encoding method according to a first aspect, wherein, if the data unit to be encoded does not include the parameter, the parameter is determined by referring to the corresponding parameter of the referenced data unit indicated in the reference information.

[0018] This allows the parameter to be determined based on the corresponding parameter in the reference, even if the data unit being encoded does not contain that parameter. This avoids data loss in the bitstream and maintains continuous processing. As a result, consistency can be ensured while reducing redundant transmissions.

[0019] An encoding method relating to a third aspect of this disclosure is an encoding method relating to a first or second aspect, wherein the parameter is represented by the difference between the referenced data unit indicated in the reference information and the corresponding parameter.

[0020] This allows for a reduction in the amount of information transmitted and received by sending and receiving only the differences, thereby lowering the bitrate. Furthermore, it clarifies the relative relationship with the reference source, improving the robustness of parameter interpretation.

[0021] The encoding method according to the fourth aspect of this disclosure is an encoding method according to any one aspect of the first to third aspects, wherein the referenced data unit indicated by the reference information is a data unit that has been processed before the data unit to be encoded.

[0022] Therefore, the reference information allows us to identify data units that were processed earlier in time, thus maintaining consistency in the temporal direction. This helps stabilize predictions and suppress the accumulation of errors.

[0023] The encoding method according to the fifth aspect of this disclosure is an encoding method according to any one aspect of the first to fourth aspects, wherein the data unit to be encoded is an intermesh data unit (IMDU) or a merged mesh data unit (MMDU).

[0024] This allows the process of switching the reference target according to the parameter type to be applied using the same procedure, regardless of whether the data unit being encoded is an intermesh data unit or a merged mesh data unit, thereby simplifying the implementation.

[0025] A decoding method according to a sixth aspect of this disclosure is a decoding method performed by a decoding device, which determines whether the decoding method for a data unit to be decoded, including a mesh data unit (MDU), is interpredictive or not, and if it is determined to be interpredictive, identifies a reference data unit indicated by reference information, determines parameters by specifying different reference units for each type of parameter, and decodes the data unit to be decoded based on the determined parameters.

[0026] This allows for the selection of appropriate references for each parameter type during decoding, suppressing unnecessary signaling. Furthermore, it ensures consistent interpretation of references and improves the stability of the decoded output. As a result, the efficiency of the decoding process can be improved.

[0027] A decoding method according to a seventh aspect of this disclosure is a decoding method according to a sixth aspect, wherein, if the data unit to be decoded does not include the parameter, the parameter is determined by referring to the corresponding parameter of the referenced data unit indicated in the reference information.

[0028] This allows the value to be determined based on the corresponding parameter in the reference, even if the data unit to be decoded does not contain that parameter. This avoids data loss in the bitstream and maintains continuous decoding. As a result, consistency can be ensured while reducing redundant transmissions.

[0029] The decoding method according to the eighth aspect of this disclosure is the decoding method according to the sixth or seventh aspect, wherein the parameter is represented by the difference between the referenced data unit indicated in the reference information and the corresponding parameter.

[0030] This allows for a reduction in the amount of information by using values ​​expressed as the difference with the reference parameter, contributing to a lower bitrate. Furthermore, it clarifies relative relationships and improves the robustness of parameter interpretation.

[0031] A decoding method according to the ninth aspect of this disclosure is a decoding method according to any one aspect of the sixth to eighth aspects, wherein the reference data unit indicated by the reference information is a data unit that has been processed before the data unit to be decoded.

[0032] Therefore, the reference information allows us to refer to previously processed data units, maintaining consistency in the time direction. This helps stabilize predictions and suppress the accumulation of errors.

[0033] A decoding method according to the tenth aspect of this disclosure is a decoding method according to any one aspect of the sixth to ninth aspects, wherein the data unit to be decoded is an intermesh data unit (IMDU) or a merged mesh data unit (MMDU).

[0034] This allows the process of switching the reference target according to the parameter type to be applied using the same procedure, regardless of whether the target of decoding is an intermesh data unit or a merged mesh data unit, thereby simplifying the implementation.

[0035] An encoding device according to an eleventh aspect of the present disclosure comprises a circuit and a memory connected to the circuit, wherein the circuit, in operation, determines whether the encoding method of a data unit to be encoded, including a mesh data unit (MDU), is interpredictive or not, and if it is determined to be interpredictive, identifies a reference data unit indicated by reference information, determines parameters by specifying different references for each type of parameter, and encodes the data unit to be encoded based on the determined parameters.

[0036] This allows for the selection of the optimal reference source for each parameter type, suppressing unnecessary signaling. Furthermore, it ensures consistent interpretation of references and improves the stability of the decoding results. As a result, encoding efficiency can be improved.

[0037] A decoding device according to a twelfth aspect of the present disclosure comprises a circuit and a memory connected to the circuit, wherein the circuit, in operation, determines whether the decoding method for a data unit to be decoded, including a mesh data unit (MDU), is interpredictive or not, and if it is determined to be interpredictive, identifies a reference data unit indicated by reference information, determines parameters by specifying different references for each type of parameter, and decodes the data unit to be decoded based on the determined parameters.

[0038] This allows for the selection of appropriate references for each parameter type during decoding, suppressing unnecessary signaling. Furthermore, it ensures consistent interpretation of references and improves the stability of the decoded output. As a result, the efficiency of the decoding process can be improved.

[0039] Note that these general or specific aspects may be implemented in a system, apparatus, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or may be implemented in any combination of a system, apparatus, integrated circuit, computer program, or recording medium.

[0040] Hereinafter, embodiments will be specifically described with reference to the drawings.

[0041] Note that all of the embodiments described below show general or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present invention. In addition, among the components in the following embodiments, components not described in the independent claims indicating the most general concept are described as optional components.

[0042] (Embodiment 1) In this embodiment, an encoding method, a decoding method, etc. will be described.

[0043] <Expressions and Terms> Here, the following expressions and terms are used.

[0044] (1) Three-dimensional mesh A three-dimensional mesh is a set of a plurality of surfaces and, for example, represents a three-dimensional object. Also, a three-dimensional mesh is mainly composed of vertex information, connection information, and attribute information. A three-dimensional mesh may be expressed as a polygon mesh or a mesh. Also, a three-dimensional mesh may have a temporal change. A three-dimensional mesh may include metadata regarding vertex information, connection information, and attribute information, or may include other additional information.

[0045] (2) Vertex information Vertex information is information indicating a vertex. For example, vertex information indicates the position of a vertex in three-dimensional space. Also, a vertex corresponds to a vertex of a surface constituting a three-dimensional mesh. Vertex information may be expressed as "Geometry", and may also be expressed as position information.

[0046] (3) Connection Information Connection information is information indicating the connection between vertices. For example, the connection information indicates the connection for forming the faces or edges of a three-dimensional mesh. The connection information may be expressed as "Connectivity". Also, the connection information may be expressed as face information.

[0047] (4) Attribute Information Attribute information is information indicating the attributes of vertices or faces. For example, the attribute information indicates attributes such as colors, images, and normal vectors associated with vertices or faces. The attribute information may be expressed as "Texture".

[0048] (5) Faces Faces are elements that make up a three-dimensional mesh. Specifically, a face is a polygon on a plane in three-dimensional space. For example, a face can be defined as a triangle in three-dimensional space.

[0049] (6) Planes A plane is a two-dimensional plane in three-dimensional space. For example, polygons are formed on a plane, and multiple polygons are formed on multiple planes.

[0050] (7) Bitstream A bitstream corresponds to encoded information. A bitstream may also be expressed as a stream, an encoded bitstream, a compressed bitstream, or an encoded signal.

[0051] (8) Encoding and Decoding The expression "encoding" may be replaced with expressions such as storing, including, writing, describing, signaling, sending out, notifying, saving, or compressing, and these expressions may be replaced with each other. For example, encoding information may mean including the information in a bitstream. Also, encoding information into a bitstream may mean encoding the information to generate a bitstream containing the encoded information.

[0052] Furthermore, the expression "decode" may be replaced with expressions such as "read out," "decipher," "read," "load," "derive," "obtain," "receive," "extract," "restore," "reconstruct," "decompress," or "expand," and these expressions may be interchangeable. For example, decoding information may mean obtaining information from a bitstream. Also, decoding information from a bitstream may mean decoding the bitstream and obtaining the information contained in the bitstream.

[0053] (9) In the explanation of ordinal numbers, first and second ordinal numbers may be assigned to components, etc. These ordinal numbers may be rearranged as appropriate. Ordinal numbers may also be newly assigned to components, etc., or removed. Furthermore, these ordinal numbers may be assigned to elements in order to identify them, and may not correspond to a meaningful order.

[0054] <Three-Dimensional Mesh> Figure 1 is a conceptual diagram showing a three-dimensional mesh according to this embodiment. The three-dimensional mesh is composed of multiple faces. For example, each face is a triangle. The vertices of these triangles are defined in three-dimensional space. The three-dimensional mesh represents a three-dimensional object. Each face may have a color or image.

[0055] Figure 2 is a conceptual diagram showing the basic elements of a three-dimensional mesh according to this embodiment. The three-dimensional mesh consists of vertex information, connection information, and attribute information. Vertex information indicates the position of the vertices of a face in three-dimensional space. Connection information indicates the connections between vertices. A face can be identified by the vertex information and connection information. In other words, a colorless three-dimensional object is formed in three-dimensional space by the vertex information and connection information.

[0056] Attribute information may be associated with vertices or faces. Attribute information associated with vertices may be expressed as "Attribute Per Point". Attribute information associated with vertices may indicate the attributes of the vertex itself or the attributes of the faces connected to the vertex.

[0057] For example, a color may be associated with a vertex as attribute information. The color associated with a vertex may be the color of the vertex itself, or the color of the face connected to the vertex. The color of a face may be the average of multiple colors associated with multiple vertices of the face. In addition, a normal vector may be associated with a vertex or face as attribute information. Such a normal vector can represent the front and back of a face.

[0058] Furthermore, a two-dimensional image may be associated with a surface as attribute information. The two-dimensional image associated with a surface is also referred to as a texture image or "Attribute Map". Additionally, information indicating the mapping between the surface and the two-dimensional image may be associated with the surface as attribute information. Such information indicating the mapping may be referred to as mapping information, vertex information of the texture image, texture coordinates, or "Attribute UV Coordinate".

[0059] Furthermore, information such as color, images, and moving images used as attribute information may be expressed as "Parametric Space."

[0060] Such attribute information can be used to reflect textures onto three-dimensional objects. In other words, vertex information, connection information, and attribute information allow a three-dimensional object with color to be formed in three-dimensional space.

[0061] In the above, attribute information is associated with vertices or faces, but it may also be associated with edges.

[0062] Figure 3 is a conceptual diagram illustrating the mapping according to this embodiment. For example, a region of a two-dimensional image in a two-dimensional plane can be mapped to a surface of a three-dimensional mesh in three-dimensional space. Specifically, coordinate information of a region in a two-dimensional image is associated with a surface of a three-dimensional mesh. As a result, the image of the region mapped in the two-dimensional image is reflected on the surface of the three-dimensional mesh.

[0063] By using mapping, a two-dimensional image used as attribute information can be separated from a three-dimensional mesh. For example, in encoding a three-dimensional mesh, the two-dimensional image may be encoded using an image encoding scheme or a video encoding scheme.

[0064] <System Configuration> Figure 4 is a block diagram showing an example configuration of the encoding and decoding system according to this embodiment. In Figure 4, the encoding and decoding system comprises an encoding device 100 and a decoding device 200.

[0065] For example, the encoding device 100 acquires a three-dimensional mesh and encodes the three-dimensional mesh into a bitstream. The encoding device 100 then outputs the bitstream to the network 300. For example, the bitstream includes the encoded three-dimensional mesh and control information for decoding the encoded three-dimensional mesh. By encoding the three-dimensional mesh, the information of the three-dimensional mesh is compressed.

[0066] Network 300 transmits the bitstream from the encoding device 100 to the decoding device 200. Network 300 may be the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. Network 300 is not necessarily limited to bidirectional communication; it may also be a one-way communication network for terrestrial digital broadcasting or satellite broadcasting, etc.

[0067] Furthermore, the network 300 may be replaced by recording media such as DVD (Digital Versatile Disc) or BD (Blu-Ray Disc®).

[0068] The decoding device 200 acquires the bitstream and decodes the three-dimensional mesh from the bitstream. The decoding of the three-dimensional mesh expands the information of the three-dimensional mesh. For example, the decoding device 200 decodes the three-dimensional mesh according to a decoding method that corresponds to the encoding method used by the encoding device 100 to encode the three-dimensional mesh. That is, the encoding device 100 and the decoding device 200 perform encoding and decoding according to their respective encoding and decoding methods.

[0069] The three-dimensional mesh before encoding can also be referred to as the original three-dimensional mesh. Similarly, the three-dimensional mesh after decoding can be referred to as the reconstructed three-dimensional mesh.

[0070] <Encoding Device> Figure 5 is a block diagram showing an example configuration of the encoding device 100 according to this embodiment. For example, the encoding device 100 includes a vertex information encoder 101, a connection information encoder 102, and an attribute information encoder 103.

[0071] The vertex information encoder 101 is an electrical circuit that encodes vertex information. For example, the vertex information encoder 101 encodes vertex information into a bitstream according to a defined format for vertex information.

[0072] The connection information encoder 102 is an electrical circuit that encodes connection information. For example, the connection information encoder 102 encodes connection information into a bitstream according to a specified format for connection information.

[0073] The attribute information encoder 103 is an electrical circuit that encodes attribute information. For example, the attribute information encoder 103 encodes attribute information into a bitstream according to a format defined for the attribute information.

[0074] Variable-length coding or fixed-length coding may be used to encode vertex information, connection information, and attribute information. Variable-length coding may correspond to Huffman coding or context-adaptive binary arithmetic coding (CABAC), etc.

[0075] The vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be integrated into a single unit. Alternatively, each of the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be further subdivided into multiple components.

[0076] Figure 6 is a block diagram showing another configuration example of the encoding device 100 according to this embodiment. For example, the encoding device 100 includes a preprocessor 104 and a postprocessor 105 in addition to the configuration shown in Figure 5.

[0077] The preprocessor 104 is an electrical circuit that performs processing before encoding vertex information, connection information, and attribute information. For example, the preprocessor 104 may perform transformation processing, separation processing, or multiplexing processing on the three-dimensional mesh before encoding. More specifically, for example, the preprocessor 104 may separate vertex information, connection information, and attribute information from the three-dimensional mesh before encoding.

[0078] The post-processor 105 is an electrical circuit that performs processing after encoding the vertex information, connection information, and attribute information. For example, the post-processor 105 may perform conversion processing, separation processing, or multiplexing processing on the encoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 105 may multiplex the encoded vertex information, connection information, and attribute information into a bitstream. Alternatively, for example, the post-processor 105 may further perform variable-length encoding on the encoded vertex information, connection information, and attribute information.

[0079] <Decoding Device> Figure 7 is a block diagram showing an example configuration of the decoding device 200 according to this embodiment. For example, the decoding device 200 includes a vertex information decoder 201, a connection information decoder 202, and an attribute information decoder 203.

[0080] The vertex information decoder 201 is an electrical circuit that decodes vertex information. For example, the vertex information decoder 201 decodes vertex information from a bitstream according to a format defined for vertex information.

[0081] The connection information decoder 202 is an electrical circuit that decodes connection information. For example, the connection information decoder 202 decodes connection information from a bitstream according to a format defined for connection information.

[0082] The attribute information decoder 203 is an electrical circuit that decodes attribute information. For example, the attribute information decoder 203 decodes attribute information from a bitstream according to a format defined for attribute information.

[0083] Variable-length decoding or fixed-length decoding may be used for decoding vertex information, connection information, and attribute information. Variable-length decoding may correspond to Huffman coding or context-adaptive binary arithmetic coding (CABAC), etc.

[0084] The vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be integrated. Alternatively, each of the vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be subdivided into multiple components.

[0085] Figure 8 is a block diagram showing another configuration example of the decoding device 200 according to this embodiment. For example, the decoding device 200 includes a pre-processor 204 and a post-processor 205 in addition to the configuration shown in Figure 7.

[0086] The preprocessor 204 is an electrical circuit that performs processing before decoding the vertex information, connection information, and attribute information. For example, the preprocessor 204 may perform conversion processing, separation processing, or multiplexing processing on the bitstream before decoding the vertex information, connection information, and attribute information.

[0087] More specifically, for example, the preprocessor 204 may separate the bitstream into sub-bitstreams corresponding to vertex information, connection information, and attribute information. Alternatively, for example, the preprocessor 204 may perform variable-length decoding on the bitstream before decoding the vertex information, connection information, and attribute information.

[0088] The post-processor 205 is an electrical circuit that performs processing after decoding the vertex information, connection information, and attribute information. For example, the post-processor 205 may perform conversion processing, separation processing, or multiplexing processing on the decoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 205 may multiplex the decoded vertex information, connection information, and attribute information into a three-dimensional mesh.

[0089] <Bitstream> Vertex information, connection information, and attribute information are encoded and stored in the bitstream. The relationship between this information and the bitstream is shown below.

[0090] Figure 9 is a conceptual diagram showing an example of the bitstream configuration according to this embodiment. In this example, connection information, vertex information, and attribute information are integrated in the bitstream. For example, connection information, vertex information, and attribute information may be included in a single file.

[0091] Furthermore, multiple parts of this information may be stored sequentially, such as the first part of connection information, the first part of vertex information, the first part of attribute information, the second part of connection information, the second part of vertex information, the second part of attribute information, and so on. These multiple parts may correspond to multiple parts that are different in time, multiple parts that are different in space, or multiple faces that are different.

[0092] Furthermore, the storage order of connection information, vertex information, and attribute information is not limited to the example above, and a different storage order may be used.

[0093] Figure 10 is a conceptual diagram showing another example of the bitstream configuration according to this embodiment. In this example, multiple files are included in the bitstream, and connection information, vertex information, and attribute information are stored in different files. Here, a file containing connection information, a file containing vertex information, and a file containing attribute information are shown, but the storage format is not limited to this example. For example, two types of information from connection information, vertex information, and attribute information may be included in one file, and the remaining type of information may be included in another file.

[0094] Alternatively, this information may be divided and stored in more files. For example, multiple parts of connection information may be stored in multiple files, multiple parts of vertex information may be stored in multiple files, and multiple parts of attribute information may be stored in multiple files. These multiple parts may correspond to multiple parts that are different in time, multiple parts that are different in space, or multiple faces that are different.

[0095] Furthermore, the storage order of connection information, vertex information, and attribute information is not limited to the example above, and a different storage order may be used.

[0096] Figure 11 is a conceptual diagram showing another example of the bitstream configuration according to this embodiment. In this example, the bitstream is composed of multiple separable sub-bitstreams, and connection information, vertex information, and attribute information are stored in different sub-bitstreams.

[0097] Here, sub-bitstreams containing connection information, sub-bitstreams containing vertex information, and sub-bitstreams containing attribute information are shown, but the storage format is not limited to these examples.

[0098] For example, two types of information from connection information, vertex information, and attribute information may be included in one sub-bitstream, and the remaining type of information may be included in another sub-bitstream. Specifically, attribute information of a two-dimensional image, etc., may be stored in a sub-bitstream compliant with an image encoding scheme, separate from the sub-bitstreams of connection information and vertex information.

[0099] Furthermore, each sub-bitstream may contain multiple files. Multiple parts of connection information may be stored in multiple files, multiple parts of vertex information may be stored in multiple files, and multiple parts of attribute information may be stored in multiple files.

[0100] Furthermore, the storage order of connection information, vertex information, and attribute information is not limited to the examples in Figures 9, 10, and 11, and a different storage order may be used. For example, they may be stored in the bitstream in the order of vertex information, connection information, and attribute information. Alternatively, they may be stored in the bitstream in any of the following orders: connection information, attribute information, and vertex information; vertex information, attribute information, and connection information; attribute information, connection information, and vertex information; or attribute information, vertex information, and connection information.

[0101] Furthermore, connection information, vertex information, and attribute information may each be divided into multiple data points, and these multiple data points may be stored in a bitstream in a periodic or random order.

[0102] <Specific Example> Figure 12 is a block diagram showing a specific example of the encoding and decoding system according to this embodiment. In Figure 12, the encoding and decoding system comprises a three-dimensional data encoding system 110, a three-dimensional data decoding system 210, and an external connector 310.

[0103] The three-dimensional data encoding system 110 comprises a controller 111, an input / output processor 112, a three-dimensional data encoder 113, a three-dimensional data generator 115, and a system multiplexer 114. The three-dimensional data decoding system 210 comprises a controller 211, an input / output processor 212, a three-dimensional data decoder 213, a system demultiplexer 214, a presenter 215, and a user interface 216.

[0104] In the three-dimensional data encoding system 110, sensor data is input from the sensor terminal to the three-dimensional data generator 115. The three-dimensional data generator 115 generates three-dimensional data, such as point cloud data or mesh data, from the sensor data and inputs it to the three-dimensional data encoder 113.

[0105] For example, the 3D data generator 115 generates vertex information and connection information and attribute information corresponding to the vertex information. The 3D data generator 115 may process the vertex information when generating the connection information and attribute information. For example, the 3D data generator 115 may reduce the amount of data by deleting duplicate vertices or perform transformations on the vertex information (such as position shifting, rotation, or normalization). The 3D data generator 115 may also render the attribute information.

[0106] Furthermore, although the three-dimensional data generator 115 is a component of the three-dimensional data encoding system 110 in Figure 12, it may also be located externally and independently of the three-dimensional data encoding system 110.

[0107] The sensor terminal that provides sensor data for generating three-dimensional data may be, for example, a moving object such as an automobile, a flying object such as an airplane, a mobile terminal, or a camera. Alternatively, distance sensors such as LIDAR, millimeter-wave radar, infrared sensors, or rangefinders, stereo cameras, or combinations of multiple monocular cameras may be used as sensor terminals.

[0108] The sensor data may include the distance (position) of the object, monocular camera images, stereo camera images, color, reflectivity, sensor attitude, orientation, gyroscope, sensing position (GPS information or altitude), speed, acceleration, sensing time, temperature, atmospheric pressure, humidity, or magnetism.

[0109] The three-dimensional data encoder 113 corresponds to the encoding device 100 shown in Figure 5, etc. For example, the three-dimensional data encoder 113 encodes three-dimensional data and generates encoded data. The three-dimensional data encoder 113 also generates control information during the encoding of three-dimensional data. The three-dimensional data encoder 113 then inputs the encoded data along with the control information to the system multiplexer 114.

[0110] The encoding method for three-dimensional data may be a geometry-based encoding method or a video codec-based encoding method. Here, the geometry-based encoding method can also be referred to as a geometry-based encoding method. The video codec-based encoding method can also be referred to as a video-based encoding method.

[0111] The system multiplexer 114 multiplexes the encoded data and control information input from the three-dimensional data encoder 113 and generates multiplexed data using a predetermined multiplexing scheme. The system multiplexer 114 may also multiplex other media such as video, audio, subtitles, application data, or document files, or reference time information, along with the encoded data and control information of the three-dimensional data. Furthermore, the system multiplexer 114 may also multiplex sensor data or attribute information related to the three-dimensional data.

[0112] For example, the multiplexed data may have a file format for storage or a packet format for transmission. ISOBMFF or an ISOBMFF-based format may be used as these formats. Alternatively, MPEG-DASH, MMT, MPEG-2 TS Systems, or RTP may be used.

[0113] The multiplexed data is then output as a transmission signal to the external connector 310 by the input / output processor 112. The multiplexed data may be transmitted as a transmission signal by wire or by wireless. Alternatively, the multiplexed data may be stored in internal memory or storage device. The multiplexed data may be transmitted to a cloud server via the internet or stored in an external storage device.

[0114] For example, the transmission or storage of multiplexed data is carried out in a manner appropriate to the medium for transmission or storage, such as broadcasting or telecommunications. Communication protocols such as HTTP, FTP, TCP, UDP, IP, or combinations thereof may be used. Furthermore, either a pull-type or push-type communication method may be used.

[0115] For wired transmission, Ethernet®, USB, RS-232C, HDMI®, or coaxial cable may be used. For wireless transmission, 3GPP®, IEEE 3G / 4G / 5G, wireless LAN, Wi-Fi, Bluetooth, or millimeter wave may be used. Furthermore, as a broadcasting method, for example, DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 may be used.

[0116] The sensor data may also be input to the three-dimensional data generator 115 or the system multiplexer 114. Alternatively, the three-dimensional data or encoded data may be output directly as a transmission signal to the external connector 310 via the input / output processor 112. The transmission signal output from the three-dimensional data encoding system 110 is input to the three-dimensional data decoding system 210 via the external connector 310.

[0117] Furthermore, each operation of the three-dimensional data encoding system 110 may be controlled by a controller 111 that executes an application program.

[0118] In the three-dimensional data decoding system 210, the transmission signal is input to the input / output processor 212. The input / output processor 212 decodes the transmission signal into multiplexed data in file format or packet format and inputs the multiplexed data to the system demultiplexer 214. The system demultiplexer 214 obtains encoded data and control information from the multiplexed data and inputs them to the three-dimensional data decoder 213. The system demultiplexer 214 may also extract other media or reference time information from the multiplexed data.

[0119] The three-dimensional data decoder 213 corresponds to the decoding device 200 shown in Figure 7, etc. For example, the three-dimensional data decoder 213 decodes three-dimensional data from encoded data based on a predetermined encoding scheme. The three-dimensional data is then presented to the user by the presenter 215.

[0120] In addition, additional information such as sensor data may be input to the display device 215. The display device 215 may present three-dimensional data based on the additional information. Furthermore, user instructions may be input from the user terminal to the user interface 216. The display device 215 may then present three-dimensional data based on the input instructions.

[0121] The input / output processor 212 may also acquire three-dimensional data and encoded data from the external connector 310.

[0122] Furthermore, each operation of the three-dimensional data decoding system 210 may be controlled by a controller 211 that executes an application program.

[0123] Figure 13 is a conceptual diagram showing an example of the configuration of point cloud data according to this embodiment. Point cloud data is data representing a three-dimensional object.

[0124] Specifically, a point cloud consists of multiple points and contains positional information indicating the three-dimensional coordinate position of each point, as well as attribute information indicating the attributes of each point. Positional information is also expressed as geometry.

[0125] The type of attribute information may be, for example, color or reflectance. A single point may be associated with attribute information relating to one type, a single point may be associated with attribute information relating to multiple different types, or a single point may be associated with attribute information having multiple values ​​for the same type.

[0126] Figure 14 is a conceptual diagram showing an example of a data file for point cloud data according to this embodiment. In this example, there is a one-to-one correspondence between location information items and attribute information items, and it shows the location information and attribute information of N points that constitute the point cloud data. In this example, the location information is information that indicates the three-dimensional coordinate position on the three axes x, y, and z, and the attribute information is information that indicates the color in RGB. A PLY file or the like can be used as a typical data file for point cloud data.

[0127] Figure 15 is a conceptual diagram showing an example of the configuration of mesh data according to this embodiment. Mesh data is data used in CG (Computer Graphics), etc., and is three-dimensional mesh data that shows the three-dimensional shape of an object with multiple faces. Each face is also represented as a polygon and has the shape of a polygon such as a triangle or quadrilateral.

[0128] Specifically, a three-dimensional mesh consists of multiple points that make up a point cloud, as well as multiple edges and multiple faces. Each point can also be expressed as a vertex or position. Each edge corresponds to a line segment connected by two vertices. Each face corresponds to a region enclosed by three or more edges.

[0129] Furthermore, a three-dimensional mesh contains positional information indicating the three-dimensional coordinate positions of its vertices. This positional information is also referred to as vertex information or geometry. A three-dimensional mesh also contains connection information indicating the relationships between multiple vertices that constitute an edge or face. This connection information is also referred to as connectivity. Finally, a three-dimensional mesh contains attribute information indicating the attributes of vertices, edges, or faces. This attribute information in a three-dimensional mesh is also referred to as texture.

[0130] For example, attribute information may indicate the color, reflectance, or normal vector for a vertex, edge, or face. The direction of the normal vector can represent the front and back of the face.

[0131] Object files or similar formats may be used as the data file format for mesh data.

[0132] Figure 16 is a conceptual diagram showing an example of a data file for mesh data according to this embodiment. In this example, the data file includes position information G(1) to G(N) for N vertices constituting the three-dimensional mesh, and attribute information A1(1) to A1(N) for the N vertices. In addition, this example includes M pieces of attribute information A2(1) to A2(M). The items of the attribute information do not have to correspond one-to-one with vertices, nor do they have to correspond one-to-one with faces. Furthermore, attribute information may not exist at all.

[0133] Connection information is indicated by a combination of vertex indices. n[1, 3, 4] represents a triangular face composed of three vertices n=1, n=3, and n=4. Also, m[2, 4, 6] indicates that attribute information m=2, m=4, and m=6 correspond to the three vertices, respectively.

[0134] Furthermore, the actual content of the attribute information may be described in a separate file. A pointer to that content may be associated with a vertex or face, etc. For example, attribute information indicating an image for a face may be stored in a two-dimensional attribute map file. The file name of the attribute map and the two-dimensional coordinate values ​​in the attribute map may be described in attribute information A2(1) to A2(M). The method of specifying attribute information for a face is not limited to these methods, and any method may be used.

[0135] Figure 17 is a conceptual diagram showing the types of three-dimensional data according to this embodiment. Point cloud data and mesh data may represent static objects or dynamic objects. Static objects are objects that do not change over time, while dynamic objects are objects that change over time. Static objects may correspond to three-dimensional data for any given point in time.

[0136] For example, point cloud data for any given time point may be referred to as a PCC frame. Similarly, mesh data for any given time point may be referred to as a mesh frame. Furthermore, both PCC frames and mesh frames may simply be referred to as frames.

[0137] Furthermore, the object's area may be limited to a certain range, like in regular video data, or it may not be limited, like in map data. Also, the density of points or surfaces can be determined in various ways. Sparse point cloud data or sparse mesh data may be used, or dense point cloud data or dense mesh data may be used.

[0138] Next, the encoding and decoding of point clouds or three-dimensional meshes will be described. The apparatus, processing, or syntax for encoding and decoding vertex information of a three-dimensional mesh in this disclosure may be applied to the encoding and decoding of point clouds. The apparatus, processing, or syntax for encoding and decoding of point clouds in this disclosure may be applied to the encoding and decoding of vertex information of a three-dimensional mesh.

[0139] Furthermore, the apparatus, processing, or syntax for encoding and decoding point cloud attribute information in this disclosure may also be applied to the encoding and decoding of connection information or attribute information of a three-dimensional mesh.

[0140] Furthermore, at least some of the processing may be shared between the encoding and decoding of point cloud data and the encoding and decoding of mesh data. This can reduce the size of the circuit and software programs.

[0141] Figure 18 is a block diagram showing an example configuration of a three-dimensional data encoder 113 according to this embodiment. In this example, the three-dimensional data encoder 113 comprises a vertex information encoder 121, an attribute information encoder 122, a metadata encoder 123, and a multiplexer 124. The vertex information encoder 121, the attribute information encoder 122, and the multiplexer 124 may correspond to the vertex information encoder 101, the attribute information encoder 103, and the post-processor 105 in Figure 6, etc.

[0142] In this example, the three-dimensional data encoder 113 encodes the three-dimensional data according to a geometry-based encoding scheme. In encoding according to a geometry-based encoding scheme, the three-dimensional structure is taken into consideration. In addition, in encoding according to a geometry-based encoding scheme, attribute information is encoded using the configuration information obtained in encoding the vertex information.

[0143] Specifically, first, vertex information, attribute information, and metadata contained in the three-dimensional data generated from sensor data are input to the vertex information encoder 121, attribute information encoder 122, and metadata encoder 123, respectively. Here, connection information contained in the three-dimensional data may be treated in the same way as attribute information. In the case of point cloud data, position information may be treated as vertex information.

[0144] The vertex information encoder 121 encodes the vertex information into compressed vertex information and outputs the compressed vertex information as encoded data to the multiplexer 124. The vertex information encoder 121 also generates metadata for the compressed vertex information and outputs it to the multiplexer 124. Furthermore, the vertex information encoder 121 generates configuration information and outputs it to the attribute information encoder 122.

[0145] The attribute information encoder 122 uses the configuration information generated by the vertex information encoder 121 to encode the attribute information into compressed attribute information and outputs the compressed attribute information as encoded data to the multiplexer 124. The attribute information encoder 122 also generates metadata for the compressed attribute information and outputs it to the multiplexer 124.

[0146] The metadata encoder 123 encodes compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The metadata encoded by the metadata encoder 123 may be used for encoding vertex information and attribute information.

[0147] The multiplexer 124 multiplexes the compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, the metadata of the compressed attribute information, and the compressed metadata into a bitstream. The multiplexer 124 then inputs the bitstream to the system layer.

[0148] Figure 19 is a block diagram showing an example configuration of a three-dimensional data decoder 213 according to this embodiment. In this example, the three-dimensional data decoder 213 includes a vertex information decoder 221, an attribute information decoder 222, a metadata decoder 223, and a demultiplexer 224. The vertex information decoder 221, attribute information decoder 222, and demultiplexer 224 may correspond to the vertex information decoder 201, attribute information decoder 203, and preprocessor 204 in Figure 8, etc.

[0149] In this example, the three-dimensional data decoder 213 decodes the three-dimensional data according to a geometry-based coding scheme. In decoding according to a geometry-based coding scheme, the three-dimensional structure is taken into consideration. In addition, in decoding according to a geometry-based coding scheme, attribute information is decoded using the configuration information obtained in the decoding of vertex information.

[0150] Specifically, first, the bitstream is input from the system layer to the demultiplexer 224. The demultiplexer 224 separates the compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, the metadata of the compressed attribute information, and the compressed metadata from the bitstream. The compressed vertex information and the metadata of the compressed vertex information are input to the vertex information decoder 221. The compressed attribute information and the metadata of the compressed attribute information are input to the attribute information decoder 222. The metadata is input to the metadata decoder 223.

[0151] The vertex information decoder 221 decodes vertex information from compressed vertex information using the metadata of the compressed vertex information. The vertex information decoder 221 also generates configuration information and outputs it to the attribute information decoder 222. The attribute information decoder 222 decodes attribute information from compressed attribute information using the configuration information generated by the vertex information decoder 221 and the metadata of the compressed attribute information. The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 may be used for decoding vertex information and attribute information.

[0152] Subsequently, vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. For example, this metadata is metadata for vertex information and attribute information, and can be used in application programs.

[0153] Figure 20 is a block diagram showing another configuration example of the three-dimensional data encoder 113 according to this embodiment. In this example, the three-dimensional data encoder 113 includes a vertex image generator 131, an attribute image generator 132, a metadata generator 133, a video encoder 134, a metadata encoder 123, and a multiplexer 124. The vertex image generator 131, the attribute image generator 132, and the video encoder 134 may correspond to the vertex information encoder 101 and the attribute information encoder 103 in Figure 6, etc.

[0154] In this example, the three-dimensional data encoder 113 encodes the three-dimensional data according to a video-based encoding scheme. In encoding according to a video-based encoding scheme, multiple two-dimensional images are generated from the three-dimensional data, and these multiple two-dimensional images are encoded according to a video encoding scheme. Here, the video encoding scheme may be HEVC (High Efficiency Video Coding) or VVC (Versatile Video Coding), etc.

[0155] Specifically, first, vertex information and attribute information contained in the three-dimensional data generated from sensor data are input to the metadata generator 133. The vertex information and attribute information are then input to the vertex image generator 131 and the attribute image generator 132, respectively. Furthermore, metadata contained in the three-dimensional data is input to the metadata encoder 123. Here, connection information contained in the three-dimensional data may be treated similarly to attribute information. In the case of point cloud data, position information may be treated as vertex information.

[0156] The metadata generator 133 generates map information for multiple two-dimensional images from vertex information and attribute information. The metadata generator 133 then inputs the map information to the vertex image generator 131, the attribute image generator 132, and the metadata encoder 123.

[0157] The vertex image generator 131 generates a vertex image based on the vertex information and map information and inputs it to the video encoder 134. The attribute image generator 132 generates an attribute image based on the attribute information and map information and inputs it to the video encoder 134.

[0158] The video encoder 134 encodes the vertex image and attribute image into compressed vertex information and compressed attribute information, respectively, according to the video encoding scheme, and outputs the compressed vertex information and compressed attribute information as encoded data to the multiplexer 124. The video encoder 134 also generates metadata for the compressed vertex information and metadata for the compressed attribute information and outputs them to the multiplexer 124.

[0159] The metadata encoder 123 encodes compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The compressible metadata includes map information. The metadata encoded by the metadata encoder 123 may also be used for encoding vertex information and attribute information.

[0160] The multiplexer 124 multiplexes the compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, the metadata of the compressed attribute information, and the compressed metadata into a bitstream. The multiplexer 124 then inputs the bitstream to the system layer.

[0161] Figure 21 is a block diagram showing another configuration example of the three-dimensional data decoder 213 according to this embodiment. In this example, the three-dimensional data decoder 213 includes a vertex information generator 231, an attribute information generator 232, a video decoder 234, a metadata decoder 223, and a demultiplexer 224. The vertex information generator 231, the attribute information generator 232, and the video decoder 234 may correspond to the vertex information decoder 201 and the attribute information decoder 203 in Figure 8, etc.

[0162] In this example, the three-dimensional data decoder 213 decodes the three-dimensional data according to a video-based coding scheme. In decoding according to a video-based coding scheme, multiple two-dimensional images are decoded according to the video coding scheme, and three-dimensional data is generated from the multiple two-dimensional images. Here, the video coding scheme may be HEVC (High Efficiency Video Coding) or VVC (Versatile Video Coding), etc.

[0163] Specifically, first, the bitstream is input from the system layer to the demultiplexer 224. The demultiplexer 224 separates the compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, the metadata of the compressed attribute information, and the compressed metadata from the bitstream. The compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, and the metadata of the compressed attribute information are input to the video decoder 234. The compressed metadata is input to the metadata decoder 223.

[0164] The video decoder 234 decodes the vertex image according to the video encoding scheme. In doing so, the video decoder 234 decodes the vertex image from the compressed vertex information using the metadata of the compressed vertex information. The video decoder 234 then inputs the vertex image to the vertex information generator 231. The video decoder 234 also decodes the attribute image according to the video encoding scheme. In doing so, the video decoder 234 decodes the attribute image from the compressed attribute information using the metadata of the compressed attribute information. The video decoder 234 then inputs the attribute image to the attribute information generator 232.

[0165] The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 includes map information used for generating vertex information and attribute information. The metadata decoded by the metadata decoder 223 may also be used for decoding vertex images and attribute images.

[0166] The vertex information generator 231 reconstructs vertex information from the vertex image according to the map information contained in the metadata decoded by the metadata decoder 223. The attribute information generator 232 reconstructs attribute information from the attribute image according to the map information contained in the metadata decoded by the metadata decoder 223.

[0167] Subsequently, vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. For example, this metadata is metadata for vertex information and attribute information, and can be used in application programs.

[0168] Figure 22 is a conceptual diagram showing a specific example of the encoding process according to this embodiment. Figure 22 shows a three-dimensional data encoder 113 and a description encoder 148. In this example, the three-dimensional data encoder 113 comprises a two-dimensional data encoder 141 and a mesh data encoder 142. The two-dimensional data encoder 141 comprises a texture encoder 143. The mesh data encoder 142 comprises a vertex information encoder 144 and a connectivity information encoder 145.

[0169] The vertex information encoder 144, the connection information encoder 145, and the texture encoder 143 may correspond to the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 in Figure 6, etc.

[0170] For example, the two-dimensional data encoder 141 operates as a texture encoder 143 and generates a texture file by encoding the texture corresponding to the attribute information as two-dimensional data according to an image encoding scheme or video encoding scheme.

[0171] Furthermore, the mesh data encoder 142 operates as a vertex information encoder 144 and a connection information encoder 145, and generates a mesh file by encoding vertex information and connection information. The mesh data encoder 142 may also encode mapping information for textures. The encoded mapping information may then be included in the mesh file.

[0172] Furthermore, the description encoder 148 generates a description file by encoding the description corresponding to metadata such as text data. The description encoder 148 may encode the description at the system layer. For example, the description encoder 148 may be included in the system multiplexer 114 in Figure 12.

[0173] The above operation generates a bitstream containing texture files, mesh files, and description files. These files may be multiplexed into the bitstream in a file format such as glTF (Graphics Language Transmission Format) or USD (Universal Scene Description).

[0174] The three-dimensional data encoder 113 may also include two mesh data encoders, namely the mesh data encoder 142. For example, one mesh data encoder encodes the vertex and connection information of a static three-dimensional mesh, while the other mesh data encoder encodes the vertex and connection information of a dynamic three-dimensional mesh.

[0175] Correspondingly, two mesh files may be included in the bitstream. For example, one mesh file may correspond to a static 3D mesh, and the other mesh file may correspond to a dynamic 3D mesh.

[0176] Furthermore, a static three-dimensional mesh may be a three-dimensional mesh of an intraframe encoded using intraprediction, and a dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe encoded using interprediction. In addition, as information for the dynamic three-dimensional mesh, the difference information between the vertex information or connection information of the intraframe three-dimensional mesh and the vertex information or connection information of the interframe three-dimensional mesh may be used.

[0177] Figure 23 is a conceptual diagram showing a specific example of the decoding process according to this embodiment. Figure 23 shows a three-dimensional data decoder 213, a description decoder 248, and an presenter 247. In this example, the three-dimensional data decoder 213 comprises a two-dimensional data decoder 241, a mesh data decoder 242, and a mesh reconstructor 246. The two-dimensional data decoder 241 comprises a texture decoder 243. The mesh data decoder 242 comprises a vertex information decoder 244 and a connection information decoder 245.

[0178] The vertex information decoder 244, connection information decoder 245, texture decoder 243, and mesh reconstructor 246 may correspond to the vertex information decoder 201, connection information decoder 202, attribute information decoder 203, and post-processor 205, etc., shown in Figure 8. The presenter 247 may correspond to the presenter 215, etc., shown in Figure 12.

[0179] For example, the two-dimensional data decoder 241 operates as a texture decoder 243 and decodes the texture corresponding to the attribute information from the texture file as two-dimensional data according to the image encoding scheme or video encoding scheme.

[0180] Furthermore, the mesh data decoder 242 operates as a vertex information decoder 244 and a connection information decoder 245, decoding vertex information and connection information from the mesh file. The mesh data decoder 242 may also decode mapping information to textures from the mesh file.

[0181] Furthermore, the description decoder 248 decodes the description corresponding to metadata such as text data from the description file. The description decoder 248 may decode the description at the system layer. For example, the description decoder 248 may be included in the system demultiplexer 214 in Figure 12.

[0182] The mesh reconstructor 246 reconstructs a three-dimensional mesh from vertex information, connection information, and textures according to the description. The presenter 247 renders and outputs the three-dimensional mesh according to the description.

[0183] The above process reconstructs and outputs a 3D mesh from a bitstream containing texture files, mesh files, and description files.

[0184] The three-dimensional data decoder 213 may also include two mesh data decoders, which are mesh data decoders 242. For example, one mesh data decoder decodes the vertex and connection information of a static three-dimensional mesh, while the other mesh data decoder decodes the vertex and connection information of a dynamic three-dimensional mesh.

[0185] Correspondingly, two mesh files may be included in the bitstream. For example, one mesh file may correspond to a static 3D mesh, and the other mesh file may correspond to a dynamic 3D mesh.

[0186] Furthermore, a static three-dimensional mesh may be a three-dimensional mesh of an intraframe encoded using intraprediction, and a dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe encoded using interprediction. In addition, as information for the dynamic three-dimensional mesh, the difference information between the vertex information or connection information of the intraframe three-dimensional mesh and the vertex information or connection information of the interframe three-dimensional mesh may be used.

[0187] A coding scheme for dynamic three-dimensional meshes is sometimes called DMC (Dynamic Mesh Coding). Similarly, a video-based coding scheme for dynamic three-dimensional meshes is sometimes called V-DMC (Video-based Dynamic Mesh Coding).

[0188] The encoding method for point clouds is sometimes called PCC (Point Cloud Compression). The video-based encoding method for point clouds is sometimes called V-PCC (Video-based Point Cloud Compression). The geometry-based encoding method for point clouds is sometimes called G-PCC (Geometry-based Point Cloud Compression).

[0189] <Implementation Example> Figure 24 is a block diagram showing an implementation example of the encoding device 100 according to this embodiment. The encoding device 100 includes a circuit 151 and a memory 152. For example, the multiple components of the encoding device 100 shown in Figure 5, etc., are implemented by the circuit 151 and memory 152 shown in Figure 24.

[0190] Circuit 151 is an information processing circuit and is a circuit that can access memory 152. For example, circuit 151 is a dedicated or general-purpose electrical circuit for encoding a three-dimensional mesh. Circuit 151 may also be a processor such as a CPU. Alternatively, circuit 151 may be a collection of multiple electrical circuits.

[0191] Memory 152 is a dedicated or general-purpose memory in which information for the circuit 151 to encode a three-dimensional mesh is stored. Memory 152 may be an electrical circuit, or it may be connected to circuit 151. Memory 152 may also be included in circuit 151. Memory 152 may also be a collection of multiple electrical circuits. Memory 152 may also be a magnetic disk or an optical disk, or it may be described as storage or a recording medium. Memory 152 may also be a non-volatile memory or a volatile memory.

[0192] For example, memory 152 may store a three-dimensional mesh or a bitstream. Alternatively, memory 152 may store a program for circuit 151 to encode the three-dimensional mesh.

[0193] Furthermore, not all of the multiple components shown in Figure 5, etc., are to be implemented in the encoding device 100, nor are all of the multiple processes shown herein to be performed. Some of the multiple components shown in Figure 5, etc., may be included in other devices, and some of the multiple processes shown herein may be executed by other devices. In addition, the multiple components of this disclosure may be implemented in any combination in the encoding device 100, and the multiple processes of this disclosure may be performed in any combination.

[0194] Figure 25 is a block diagram showing an example of the implementation of the decoding device 200 according to this embodiment. The decoding device 200 includes a circuit 251 and a memory 252. For example, the multiple components of the decoding device 200 shown in Figure 7, etc., are implemented by the circuit 251 and memory 252 shown in Figure 25.

[0195] Circuit 251 is an information processing circuit and is a circuit that can access memory 252. For example, circuit 251 is a dedicated or general-purpose electrical circuit for decoding a three-dimensional mesh. Circuit 251 may also be a processor such as a CPU. Alternatively, circuit 251 may be a collection of multiple electrical circuits.

[0196] Memory 252 is a dedicated or general-purpose memory that stores information for circuit 251 to decode the three-dimensional mesh. Memory 252 may be an electrical circuit and may be connected to circuit 251. Memory 252 may also be included in circuit 251. Memory 252 may also be a collection of multiple electrical circuits. Memory 252 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory 252 may also be a non-volatile memory or a volatile memory.

[0197] For example, memory 252 may store a three-dimensional mesh or a bitstream. Alternatively, memory 252 may store a program for circuit 251 to decode the three-dimensional mesh.

[0198] Furthermore, the decoding device 200 does not need to implement all of the components shown in Figure 7, etc., nor does it need to perform all of the processes shown herein. Some of the components shown in Figure 7, etc., may be included in other devices, and some of the processes shown herein may be performed by other devices. In addition, the decoding device 200 may implement the components of this disclosure in any combination, and the processes of this disclosure may be performed in any combination.

[0199] The encoding and decoding methods, including the steps performed by each component of the encoding device 100 and decoding device 200 of this disclosure, may be performed by any device or system. For example, part or all of the encoding and decoding methods may be performed by a computer equipped with a processor, memory, and input / output circuits, etc. In this case, the encoding and decoding methods may be performed by the computer executing a program that causes the computer to perform the encoding and decoding methods.

[0200] Furthermore, a non-temporary computer-readable recording medium such as a CD-ROM may contain either a program or a bitstream.

[0201] An example of a program may be a bitstream. For instance, a bitstream containing an encoded three-dimensional mesh may include syntax elements to cause the decoding device 200 to decode the three-dimensional mesh. The bitstream then causes the decoding device 200 to decode the three-dimensional mesh according to the syntax elements contained within the bitstream. Therefore, a bitstream can perform a similar role to a program.

[0202] The above bitstream may be an encoded bitstream containing an encoded three-dimensional mesh, or a multiplexed bitstream containing an encoded three-dimensional mesh and other information.

[0203] Furthermore, each component of the encoding device 100 and the decoding device 200 may be made of dedicated hardware, general-purpose hardware that executes the above-mentioned program, or a combination thereof. The general-purpose hardware may also consist of a memory on which the program is stored, and a general-purpose processor that reads the program from the memory and executes it. Here, the memory may be semiconductor memory or a hard disk, and the general-purpose processor may be a CPU.

[0204] Furthermore, dedicated hardware may consist of memory and a dedicated processor, etc. For example, a dedicated processor may refer to memory for recording data and execute an encoding method and a decoding method.

[0205] Furthermore, each component of the encoding device 100 and the decoding device 200 may be an electrical circuit, as described above. These electrical circuits may form a single electrical circuit as a whole, or they may be separate electrical circuits. These electrical circuits may correspond to dedicated hardware, or they may correspond to general-purpose hardware that executes the above-mentioned program, etc. Furthermore, the encoding device 100 and the decoding device 200 may be implemented as integrated circuits.

[0206] Furthermore, the encoding device 100 may be a transmitting device that transmits a three-dimensional mesh. The decoding device 200 may be a receiving device that receives a three-dimensional mesh.

[0207] <Encoding and Decoding of Displacements> Here, the following terms are used as examples.

[0208] (1) Image An image is a data unit composed of a collection of pixels, and includes a picture or a block smaller than a picture. Images include both moving images and still images.

[0209] (2) A picture is an image processing unit composed of a collection of pixels, and is also called a frame or field.

[0210] (3) A block is a processing unit consisting of a specific number of pixels. The term block is also used in the examples shown below. The shape of a block is not particularly limited. A block may be, for example, a rectangular shape of M x N pixels, or a square shape of M x M pixels. A block may also be triangular, circular, or have other shapes. Examples of blocks are as follows.

[0211] - Slices, tiles, or bricks - CTU, superblock, or basic partitioning unit - VPDU, hardware processing partitioning unit - CU, processing block unit, prediction block unit (PU), or orthogonal transformation block unit (TU) - subblocks

[0212] (4) A pixel or sample is the smallest point, or in other words, the smallest unit, of an image. A pixel or sample includes not only pixels at integer positions but also pixels at sub-pixel positions generated from pixels at integer positions.

[0213] (5) Pixel values ​​or sample values ​​Pixel values ​​or sample values ​​are eigenvalues ​​of a pixel. Pixel values ​​or sample values ​​may include luma values, chroma values, or RGB tonal levels, and may also include depth values ​​or binary values ​​of 0 or 1.

[0214] (6) Flags A flag represents one or more bits. A flag is, for example, a parameter or index represented by two or more bits. A flag may also represent a value that is not represented by a binary number, but by a non-binary number.

[0215] (7) A signal is a symbol or encoded information used to transmit information. Signals include discrete digital signals or continuous analog signals.

[0216] (8) Stream or Bitstream A stream or bitstream is a sequence of digital data that represents a flow of digital data. A stream or bitstream may be a single stream or may consist of multiple streams having multiple layers. A stream or bitstream may be transmitted by serial communication using a single transmission path or by packet communication using multiple transmission paths.

[0217] (9) Differences: In the case of scalar quantities, differences can include simple differences (x - y) and difference calculations. Differences can include the absolute value of the difference (|x - y|), the square of the difference (x^2 - y^2), the square root of the difference (√(x - y)), weighted differences (ax - by, where a and b are constants), or offset differences (x - y + a, where a is the offset).

[0218] (10) In the case of scalar quantities, sums can include simple sums (x + y) and addition operations. Sums can include the absolute value of the sum (|x + y|), the sum of squares (x^2 + y^2), the square root of the sum (√(x + y)), a weighted sum (ax + by, where a and b are constants), or an offset sum (x + y + a, where a is the offset).

[0219] (11) "Based on" The expression "based on something" means that other things besides that "something" may also be considered. Also, "based on" can be used when a direct result is obtained, or when a result is obtained after intermediate results.

[0220] (12) "Used" or "Used" The expression "something was used" or "something was used" means that something other than that "something" may also be considered. Also, the expression "used" or "used" can be used when a direct result is obtained, or when a result is obtained after an intermediate result.

[0221] (13) Prohibition "To prohibit" can be rephrased as "not to permit." Also, "not prohibited / prohibited" or "permitted / permitted" does not necessarily mean "obligation."

[0222] (14) "Restriction" or "Limitation" "Restriction" or "Limitation" can be rephrased as "not permitted / not allowed" or "not permitted / permitted." Also, "prohibited / not prohibited" or "not permitted / permitted" does not necessarily mean "obligation." Furthermore, the prohibition, quantitatively or qualitatively, may be partial or entire.

[0223] (15) Chroma The term chroma is an adjective represented by the symbols Cb or Cr, indicating that a sample sequence or a single sample represents one of two color difference signals associated with a primary color. The term chroma is sometimes used instead of the term chrominance.

[0224] (16) Luma The term luma is an adjective represented by the symbol or subscript Y or L, indicating that a sample sequence or single sample represents a monochrome signal relating to a primary color. The term luma is sometimes used as a substitute for the term luminance.

[0225] The coding and decoding system of this embodiment will be described below.

[0226] A typical three-dimensional model (also called a 3D model) digitally represents an object so that the user can explore the model using zoom, pan, and rotate in all three dimensions while rendering it over time. One way to construct such a representation is to build a 3D mesh using triangles. In the above model, the positions of the triangle vertices, the connectivity of the triangle vertices to each other, and their associated attributes (such as normals or UV patches) are stored.

[0227] Storing all this information in an uncompressed format requires a very large amount of memory, and therefore a very large bandwidth for transmission. The triangles that form a mesh often have attributes similar to repeating patterns, especially in temporal and spatial neighborhoods. These repetitions can be used to develop efficient encoding and decoding methods for storage and transmission. One such encoding and decoding method is Video-based Dynamic Mesh Coding (V-DMC).

[0228] Figure 26 is a block diagram showing different configuration examples of the coding and decoding system according to this embodiment. As shown in Figure 26, the coding and decoding system includes an coding device 100 and a decoding device 200.

[0229] The encoding / decoding system accepts a three-dimensional mesh (also called a 3D mesh) as input in the form of three-dimensional coordinates of vertices (vertex information), connectivity (connection information), and associated attributes (attribute information). Note that the 3D mesh can include not only geometry but also texture maps.

[0230] The encoding device 100 captures the input 3D mesh (also called the input 3D mesh or input mesh) in the form of the three-dimensional coordinates of the vertices, connectivity, and associated attributes. The encoding device 100 encodes all the associated information into a stream. The stream may consist of a single bitstream or multiple bitstreams.

[0231] Network 300 transmits the stream generated by the encoding device 100 to the decoding device 200. Network 300 may be the Internet, a WAN (Wide Area Network), a LAN (Local Area Network), or any combination thereof. Furthermore, network 300 is not necessarily limited to a bidirectional communication network, but may also be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Alternatively, a recording medium such as a DVD (Digital Versatile Disc) or a BD (Blue-Ray Disc) on which the stream is recorded may be used instead of network 300.

[0232] The stream is transmitted to the decoder 200 via the network 300. The decoder 200 decodes the bitstream and generates a three-dimensional mesh using the three-dimensional coordinates, connectivity, and associated attributes of the decoded vertices. The decoder 200 outputs the generated three-dimensional mesh (also called the output 3D mesh or output mesh).

[0233] Figure 27 shows another example of the encoding device 100 configuration.

[0234] As shown in Figure 27, the encoding device 100 includes a preprocessor 1103 and a compressor 1106.

[0235] The encoding device 100 reads the input mesh 1101 and attribute map 1102 and passes them to the preprocessor 1103. The preprocessor 1103 processes the input mesh and extracts the base mesh 1104 and displacement data 1105. The attribute map 1102, along with the extracted base mesh 1104 and displacement data 1105, is passed to the compressor 1106.

[0236] Furthermore, the compressor 1106 compresses the base mesh 1104, displacement data 1105, and attribute map 1102 to generate a bitstream 1107. The compressor 1106 can transmit additional information to the decoder 200 by further including metadata 1108 in the bitstream 1107.

[0237] Figure 28 shows another example of the configuration of the decoding device 200.

[0238] As shown in Figure 28, the decoding device 200 includes an expander 2102 and a post-processing unit 2106.

[0239] The decoding device 200 reads the bitstream 2101 and passes it to the decompressor 2102. The decompressor 2102 decompresses the base mesh 2103, displacement data 2104, and attribute map 2108 from the bitstream 2101 and passes them to the post-processor 2106. An example of displacement data 2104 is a displacement vector.

[0240] Furthermore, the post-processor 2106 generates the output mesh 2107 by processing the base mesh 2103 according to the displacement data 2104 and attribute map 2108. The post-processor 2106 may also use information from metadata 2105 to generate the output mesh 2107.

[0241] Figure 29 is a block diagram showing yet another configuration example of the encoding device 100 according to this embodiment.

[0242] In this example, the encoding device 100 includes a volumetric capturer 511, a projector 512, a base mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and optionally one or more other type encoders 516.

[0243] The volumetric capture device 511 captures the content and outputs the captured content to the projector 512.

[0244] The projector 512 projects content onto an input mesh (three-dimensional mesh frame) that includes geometric coordinates (vertex coordinates indicating the positions of vertices), texture coordinates, and connectivity (connection information). This data is output to the base mesh encoder 513, displacement encoder 514, attribute encoder 515, and optionally one or more other type encoders 516. Each encoder compresses the data into a bitstream.

[0245] Figure 30 is a block diagram showing yet another configuration example of the decoding device 200 according to this embodiment.

[0246] In this example, the decoding device 200 comprises a base mesh decoder 613, a displacement decoder 614, an attribute decoder 615, one or more other type decoders 616, and a three-dimensional reconstructor 617.

[0247] The bitstream is sent to the base mesh decoder 613, the displacement decoder 614, the attribute decoder 615, and optionally one or more other type decoders 616. These decoders decode the bitstream to generate data (decoded data) including geometric coordinates, texture coordinates, and connectivity. The decoded data is then sent to the 3D reconstructor 617, where the output mesh (3D mesh frame) is reconstructed.

[0248] The encoding process performed by the encoding device 100 will be described in detail below.

[0249] Figure 31 is a flowchart showing the processing of the encoding device 100. Figure 32 is an explanatory diagram conceptually illustrating the encoding of a mesh frame. The processing of the encoding device 100 will be explained with reference to Figures 31 and 32.

[0250] In step S101, the encoding device 100 reads the input mesh frame, which is a 3D mesh frame, and its attributes. The input mesh frame is the mesh frame input to the encoding device 100. An example of the input mesh frame, a 3D mesh frame, is shown as mesh frame 1301 (see Figure 32).

[0251] In step S102, the encoding device 100 generates a base mesh frame with fewer vertices than the input mesh frame by performing a decimation process on the input mesh frame read in step S101. The base mesh frame generated by decimating the mesh frame 1301 is shown as the base mesh frame 1302 (see Figure 32).

[0252] In step S103, the encoding device 100 calculates displacement information that the decoding device 200 uses to reconstruct the mesh frame. The displacement information corresponds to a displacement vector from the vertices of the base mesh frame generated in step S102 to the vertices of the input mesh frame. One method for calculating the displacement information is to subtract the coordinates of the vertices of the base mesh frame from the coordinates of the vertices of the input mesh frame. The displacement information calculated from the mesh frame 1301 and the base mesh frame 1302 is shown as displacement information 1303 (see Figure 32). The displacement information 1303 is in vector form, or in other words, it is expressed as a displacement vector.

[0253] In step S104, the encoding device 100 encodes the base mesh frame generated in step S102, the displacement information generated in step S103, and the attributes of the input mesh frame into a bitstream (corresponding to a compressed bitstream). An example of a bitstream is shown as bitstream 1304 (see Figure 32).

[0254] Specifically, bitstream 1304 includes a video bitstream containing vertex coordinates and connection information for vertices A, C, E, and F, displacement information, and texture data, as well as a compressed attribute map (see Figure 32). The displacement information includes displacement information for displacing vertices based on vertex coordinates obtained from the subdivided base mesh frame. The compressed attribute map is texture coordinates for applying texture data to the mesh frame reconstructed using the base mesh frame and displacement information.

[0255] The decoding process performed by the decoding device 200 will be described in detail below.

[0256] Figure 33 is a flowchart showing the processing of the decoding device 200. Figure 34 is an explanatory diagram conceptually illustrating the decoding of a mesh frame (3D mesh). The processing of the decoding device 200 will be explained with reference to Figures 33 and 34.

[0257] In step S201, the decoding device 200 decodes the base mesh frame and attributes from the bitstream (corresponding to the compressed bitstream). An example of the decoded base mesh frame (corresponding to the decoded base mesh frame) is shown as the decoded base mesh frame 2301 (see Figure 34).

[0258] In step S202, the decoding device 200 generates subdivided vertices by performing a subdivision process on the base mesh frame decoded in step S201. An example of a base mesh frame (mesh frame) containing subdivided vertices is shown as base mesh frame 2302 (see Figure 34).

[0259] In step S203, the decoding device 200 decodes displacement information from the bitstream (corresponding to a compressed bitstream). An example of the decoded displacement information is shown as displacement information 2303 (see Figure 34). Displacement information 2303 is in vector form, or in other words, it is represented as a displacement vector.

[0260] In step S204, the decoding device 200 reconstructs the shape of the mesh frame by moving the vertices of the base mesh frame, including the sub-divided vertices, to new positions using displacement information, and then restores the mesh frame by applying attribute information. An example of an attribute is a texture. An example of the reconstructed mesh frame is shown as mesh frame 2304 (see Figure 34).

[0261] Figure 35 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.

[0262] Figure 35 shows an example of a typical intra-decryption block diagram.

[0263] The decoding device shown in Figure 35 comprises an inverse multiplexer 1231, a switch 1232, a static mesh decoder 1233, a mesh buffer 1234, a motion decoder 1235, a base mesh reconstructor 1236, an inverse quantizer 1237, a video decoder 1238, an image umpacker 1239, an inverse quantizer 1240, an inverse wavelet converter 1241, a reconstructor 1242, a video decoder 1243, and a color converter 1244.

[0264] The demultiplexer 1231 acquires the compressed bitstream and separates the compressed data for the base mesh, the video containing displacement data (also called the displacement bitstream), and the video containing attribute data (also called the attribute bitstream). The compressed data for the base mesh is passed to the switch 1232. The switch 1232 decides whether to perform intra-decoding or inter-decoding based on the parameters in the bitstream.

[0265] If intra-decoding is selected, the bitstream is passed to a static mesh decoder 1233 that generates a quantized base mesh. The static mesh decoder 1233 is, for example, a decoder that uses an edge breaker algorithm to decode 3D mesh data. The static mesh decoder 1233 generates a quantized base mesh from the bitstream. The quantized base mesh generated by the static mesh decoder 1233 is stored in the mesh buffer 1234 for reference when inter-decoding is selected.

[0266] Switch 1232 passes compressed data about the base mesh to the motion decoder 1235 if inter-decoding is selected. The motion decoder 1235 receives the previously decoded, quantized base mesh and decodes motion data representing the difference in vertex coordinates between the quantized base mesh stored in the mesh buffer 1234 and the current quantized base mesh. The motion data and the quantized base mesh stored in the mesh buffer 1234 are used by the base mesh reconstructor 1236 to reconstruct the current quantized base mesh. The quantized base mesh obtained from inter-decoding or intra-decoding is passed to the inverse quantizer 1237 to obtain the decoded base mesh.

[0267] The video containing displacement data is passed to the video decoder 1238 because the bitstream contains displacement data in an image format having two chroma information and one luma information. The video decoder 1238 decodes the data using the video frame decompression method. In another example, the displacement data is decoded using an arithmetic decoder. This decompressed data is passed to the image umpacker 1239, which extracts wavelet coefficients associated with each vertex from the decompressed image data. The inverse quantizer 1240 inverse quantizes the quantized wavelet coefficients with three components associated with each vertex. The inverse wavelet transformer 1241 inverse transforms the result to obtain the finally decoded displacement data. The decoded displacement data and the decoded base mesh are passed to the reconstructor 1242. The reconstructor 1242 subdivides the edges of the decoded base mesh, displaces the vertices using the decoded displacement data, and obtains the decoded mesh.

[0268] The video containing attribute data is passed to another video decoder 1243 to obtain a decoded attribute bitstream. The decoded attribute bitstream is further processed by a color converter 1244 for color space and color format conversion to obtain a decoded attribute map.

[0269] Figure 36 is a block diagram showing an example configuration of the encoding device according to this embodiment.

[0270] First, the encoding device acquires the base mesh bitstream, the displacement bitstream, and the attribute bitstream obtained from the three-dimensional mesh preprocessing step.

[0271] The encoding device shown in Figure 36 comprises a quantizer 1261, a switch 1262, a static mesh encoder 1263, a mesh buffer 1264, a motion encoder 1265, a base mesh reconstructor 1266, a displacement data updater 1267, a wavelet converter 1268, a quantizer 1269, an image packer 1270, a video encoder 1271, a color converter 1272, and a video encoder 1273.

[0272] The base mesh (specifically, the position information of the multiple vertices that make up the base mesh) is first quantized by the quantizer 1261. The quantized base mesh (base mesh data) is output to the switch 1262, which determines whether to use intra-coding or inter-coding. If intra-coding is selected, the quantized base mesh (base mesh bitstream) is output to the static mesh encoder 1263, which generates the quantized base mesh. An example of the static mesh encoder 1263 is an encoder that uses an edge breaker algorithm for encoding three-dimensional mesh data. This encoded (quantized) base mesh is stored in the mesh buffer 1264 for reference when inter-coding is selected. If inter-coding is selected, the switch 1262 outputs compressed data related to the base mesh to the motion encoder 1265. The motion data and static three-dimensional mesh in the mesh buffer 1264 are used by the base mesh reconstructor 1266 to reconstruct the current quantized base mesh.

[0273] The displacement data is output to the displacement data updater 1267, where it is updated based on the quantized static base mesh or the reconstructed inter-encoded base mesh. Next, the wavelet converter 1268 performs the transformation, followed by quantization in the quantizer 1269. The quantized displacement data is packed into an image by the image packer 1270 and finally encoded by the video encoder 1271. The encoded displacement data is output to the multiplexer 1274.

[0274] Attribute information (e.g., attribute map) is output to the color converter 1272 for conversion of color space and color format. This converted attribute information is encoded by the video encoder 1273 and output to the multiplexer 1274.

[0275] The multiplexer 1274 acquires data relating to the encoded base mesh (compressed data relating to the base mesh), video data including encoded displacement data, and video data including attribute information such as encoded attribute maps, and generates a bitstream (compressed bitstream) containing the acquired data. The generated compressed bitstream is output to, for example, a decoding device.

[0276] Figure 37 is a block diagram showing an example configuration of a decoding device according to this embodiment.

[0277] Figure 37 shows an example of a reconstructor that obtains a decoded 3D mesh 1256 from the decoded base mesh 1251 and the decoded displacement data 1254.

[0278] The decoded base mesh 1251 is passed to the sub-divider 1252.

[0279] The sub-decomposer 1252 subdivides any two connected vertices in the entire 3D mesh by adding a new vertex between them. This process can be repeated several times, including the vertices created in the previous sub-decomposition step, to generate a predetermined number of vertices. Each iteration of sub-decomposition across the entire 3D mesh generates a new level of detail (LoD). The subdivided mesh 1253 and the decoded displacement data 1254 are passed to the displacementr 1255. The displacementr 1255 generates the decoded 3D mesh 1256 by moving each vertex to a new position according to the corresponding displacement data.

[0280] Sub-partitioning will be described below. Sub-partitioning is performed, for example, by a sub-partitioner 1252.

[0281] Figure 38 is an explanatory diagram showing an example of subdivision.

[0282] The base mesh shown in Figure 38(a) includes vertices A, B, and C, and connectivity information indicating their connectivity.

[0283] Figure 38(b) shows the mesh generated by the first subdivision, in other words, the mesh after the first subdivision. In the first subdivision, the subdivision generator generates vertices D, E, and F, and connection information indicating their connectivity. The mesh generated by the subdivision generator is also called LoD1 or the first LoD.

[0284] Vertex D of the mesh after the first subdivision is a vertex generated by the subdivision based on vertices A and B. Similarly, vertex F is a vertex generated by the subdivision based on vertices B and C. Vertex E is a vertex generated by the subdivision based on vertices A and C.

[0285] For example, vertex D could be the midpoint of the line segment AB (or side AB) connecting vertices A and B, which were the source of its generation. Similarly, vertex E could be the midpoint of line segment AC, and vertex F could be the midpoint of line segment BC.

[0286] Figure 38(c) shows the mesh generated by the second subdivision, in other words, the mesh after the second subdivision. In the second subdivision, the subdivision generator generates vertices G, H, I, J, K, L, M, N, and O, and connection information indicating their connectivity. The mesh generated by the subdivision generator is also called LoD2 or the second LoD.

[0287] Vertex G of the mesh after the second subdivision is a vertex generated by the subdivision based on vertices A and D. Similarly, vertex H is a vertex generated by the subdivision based on vertices A and E. Vertex I is a vertex generated by the subdivision based on vertices B and D. Vertex J is a vertex generated by the subdivision based on vertices D and F. Vertex K is a vertex generated by the subdivision based on vertices E and F. Vertex L is a vertex generated by the subdivision based on vertices C and E. Vertex M is a vertex generated by the subdivision based on vertices B and F. Vertex N is a vertex generated by the subdivision based on vertices C and F. Vertex O is a vertex generated by the subdivision based on vertices D and E.

[0288] For example, vertex G could be the midpoint of the line segment AD (or edge AD) connecting vertices A and D, which were the source of its generation. Similarly, vertex H could be the midpoint of line segment AE. Vertex I could be the midpoint of line segment BD. Vertex J could be the midpoint of line segment DF. Vertex K could be the midpoint of line segment EF. Vertex L could be the midpoint of line segment CE. Vertex M could be the midpoint of line segment BF. Vertex N could be the midpoint of line segment CF. Vertex O could be the midpoint of line segment DE.

[0289] In the following sections, the displacement of the vertices will be explained with reference to Figures 39 and 40. The displacement of the vertices is performed by the reconstructor.

[0290] Figure 39 is an explanatory diagram showing an example of vertex displacement after subdivision. Figure 40 is an explanatory diagram showing an example of vertices in the original mesh.

[0291] The base mesh shown in Figure 39(a) includes vertices A, B, C, and Z, and connectivity information indicating their connectivity.

[0292] Figure 39(b) shows the mesh generated by the first subdivision, in other words, the mesh after the first subdivision (i.e., the first LoD). In the first subdivision, the subdivision generator generates vertices S, T, U, X, or Y and connection information indicating their connectivity. Vertices S, T, U, X, or Y are the same as vertices D, E, and F shown in Figure 38(b).

[0293] Figure 39(c) shows the mesh generated by the second subdivision, in other words, the mesh after the second subdivision (i.e., the second LoD). In the second subdivision, the subdivision tool generates vertices D, E, F, G, and H, and connectivity information indicating their connectivity. Vertices D, E, F, G, and H are the same as vertices G, H, I, J, K, L, M, N, or O shown in Figure 38(c).

[0294] Figure 39(d) shows the mesh containing the vertices after subdivision and displacement. The vertices A, B, C, D, E, F, G, H, S, T, U, X, Y, and Z shown in Figure 39(d) are located at positions that have been displaced using displacement information from the positions of those vertices shown in Figure 39(c).

[0295] The original mesh shown in Figure 40 is an example of the mesh input to the encoding device 100, that is, the mesh before encoding.

[0296] The mesh shown in Figure 39 has a shape similar to the original mesh shown in Figure 40. The displacement information is generated by the displacement vector calculator 1207 of the encoding device 100 as information indicating the displacement from the vertices of the base mesh to the vertices of the original mesh. By reconstructing the mesh using the displacement information generated in this way, a mesh with a shape similar to the original mesh is generated.

[0297] The decoding device 200 can output the mesh shown in Figure 39(d).

[0298] Next, we will explain the division of the mesh into submeshes with reference to Figures 41 and 42.

[0299] A mesh can be divided into multiple smaller parts, and each divided part can be encoded. When dividing a mesh, the vertices of the mesh are divided in such a way that the coordinates and connectivity of the vertices included in each part can be encoded independently.

[0300] Figure 41 is an explanatory diagram showing an example of a mesh. Figure 42 is an explanatory diagram showing an example of dividing a mesh into submeshes.

[0301] The mesh shown in Figure 41 is the original mesh, and is sometimes called the full mesh in contrast to the submesh.

[0302] Figure 42 shows how the full mesh shown in Figure 41 is divided into two submeshes. For vertices A, B, and C of the full mesh (see Figure 41), vertex A is duplicated into vertices A1 and A2, vertex B is duplicated into vertices B1 and B2, and vertex C is duplicated into vertices C1 and C2, thereby creating two submeshes (i.e., the first submesh and the second submesh) from the full mesh. The first submesh and the second submesh are meshes that can be decoded independently.

[0303] In the following sections, the packing of displacement information into image frames will be explained with reference to Figures 43, 44, and 45.

[0304] Figures 43, 44, and 45 are explanatory diagrams illustrating examples of packing displacement information into image frames. Note that image frames can also be referred to as video frames.

[0305] Vertex displacement data is encoded as image frame data by mapping it to each component of an image frame in YUV format (i.e., the Y component (Y Plane), U component (U Plane), and V component (V Plane) respectively). This case is explained below as an example. Alternatively, vertex displacement data may be encoded as image frame data by mapping it to each component of an image frame in RGB format (the R component, G component, and B component, respectively).

[0306] The decoding device 200 can use an image coding module to extract displacement data. The displacement data may be in the form of X, Y, or Z components in a global coordinate system (e.g., a Cartesian coordinate system), or normal, tangent, or both tangent components in a local coordinate system. Methods for mapping displacement data to an image frame include the following:

[0307] For example, in the first method, displacement data is arranged in scan order within the image frame. An example of packing displacement data in this case is shown in Figure 43. The displacement data is directly mapped to the image frame according to a predefined scan order.

[0308] Note that since the height and width of the image frame are fixed, the displacement data may not fit perfectly within the frame. In such cases, the remaining portion of the image frame is padded with padding data (also called padded data) (see Figure 43).

[0309] For example, in the second method, displacement data is separated into multiple Lines of Data (LDs) and mapped to the Y, U, and V components of the image frame. An example of packing the displacement data in this case is shown in Figure 44. Here, the displacement data for the image frame of the next LD starts immediately after the displacement data for the previous LD ends. Similar to the first method, if the displacement data does not fit perfectly into the image frame, padding is applied to the end of the image frame (see Figure 44).

[0310] For example, in the third method, the displacement data corresponding to the LoD is mapped to the Y, U, and V components of the image frame in a different manner than in the second method. An example of the packing of the displacement data in this case is shown in Figure 45. In this way, each LoD can be decoded independently. In the third method, intermediate padding is performed on the displacement data of each LoD, and CTU alignment is performed together with padding at the end of the video frame (see Figure 45).

[0311] Figure 46 is a block diagram showing a detailed configuration example of the decoding device 200 according to this embodiment. Specifically, Figure 46 shows an example of the configuration of the geometry coordinate decoder included in the decoding device 200.

[0312] In this example, the decoding device 200 includes a frame header decoder 631, a vertex geometry coordinate predictor 632, a vertex geometry coordinate difference decoder 633, and a reconstructor 634.

[0313] The frame header decoder 631 reads the bitstream, decodes the frame header in the bitstream, and decides whether to intra-decode (intra-predict) or inter-decode (inter-predict) the frame data.

[0314] If inter-decoding is selected, the frame data contained in the bitstream is output to the vertex geometry coordinate predictor 632.

[0315] The vertex geometry coordinate predictor 632 outputs prediction information to the reconstructor 634. An example of prediction information is a motion vector.

[0316] The reconstructor 634 outputs the three-dimensional coordinates of the vertices (vertex geometry coordinates) using prediction information along with the vertex coordinates from previously decoded frames.

[0317] On the other hand, if intra-decoding is selected, the frame data contained in the bitstream is output to the vertex geometry coordinate difference decoder 633.

[0318] The vertex geometry coordinate difference decoder 633 decodes the frame data, which has been encoded as the difference between the coordinates of the vertices contained in the frame, in order to generate vertex coordinates. Only one of the vertex geometry coordinates from the vertex geometry coordinate difference decoder 633 or the reconstructor 634 is used to generate the decoded three-dimensional mesh frame.

[0319] Figure 47 is an explanatory diagram showing the coordinates of vertices in a three-dimensional mesh according to this embodiment. Specifically, Figure 47 shows an example in which the entire frame of a three-dimensional mesh frame is decoded using the actual vertex coordinates (positions) contained in the bitstream.

[0320] The coordinates of vertex A in the three-dimensional mesh frame at time (t) are decoded as (6, 8, 9) using the Cartesian coordinate system (x, y, z), as shown in Figure 47(a). Similarly, the coordinates of vertex B are decoded as (10, 6, 7), and the coordinates of vertex C are decoded as (14, 8, 9). The same applies to vertices D through G.

[0321] Figure 48 is an explanatory diagram showing the prediction information according to this embodiment. Specifically, Figure 48 shows another example in which the entire frame of the three-dimensional mesh frame at time (t) is decoded using the frame at time (t-1) (past frame) and the prediction information contained in the bitstream.

[0322] The coordinates of vertex A in the frame to be decoded (the current frame), which are (6, 8, 9), are decoded by adding the coordinates of vertex A in the past frame, which are (4, 7, 8), and the value for vertex A indicated by the prediction information, which is (2, 1, 1). Similarly, the coordinates of vertex B in the current frame, which are (10, 6, 7), are decoded by adding the coordinates of vertex B in the past frame, which are (8, 6, 7), and the value for vertex B indicated by the prediction information, which is (2, 0, 0).

[0323] Hereafter, an example of the configuration of the encoding device according to this embodiment will be described.

[0324] Figure 49 is a block diagram showing an example of the configuration of the encoding device according to this embodiment.

[0325] The encoding device shown in Figure 49 comprises a decimator 4801, a subdivider 4802, a displacement vector calculator 4803, a wavelet converter 4804, an interpreter 4805, a quantizer 4806, an image packer 4807, a video encoder 4808, an inverse quantizer 4811, a reconstructor 4812, and a reference buffer 4813.

[0326] The decimator 4801 acquires the mesh frame (corresponding to the original 3D mesh frame, also called the original mesh frame or original mesh) input to the encoding device, and generates a base mesh frame (also called the base mesh) by performing a decimation process (in other words, a thinning process) on the acquired mesh frame. Decimation is a process that deletes (in other words, thins out) some of the vertices included in the original mesh. Decimation may include a process that changes the position of at least some of the vertices included in the original mesh, or a process that changes the connectivity of at least some of the vertices included in the original mesh. Decimation is also simply called decimation.

[0327] The base mesh generated by the decimation process has fewer vertices than the original mesh. The vertices of the base mesh may be located in different positions than the vertices of the original mesh. Also, the connectivity of the vertices of the base mesh may differ from that of the vertices of the original mesh. The decimator 4801 provides the generated base mesh frame to the sub-divider 4802.

[0328] The sub-divider 4802 performs sub-division processing on the base mesh frame generated by the decimator 4801. Sub-division processing can be a process of subdividing the base mesh frame into sub-subdivisions. The sub-divider 4802 provides the subdivided base mesh frame to the displacement vector calculator 4803.

[0329] Specifically, the sub-divider 4802 can subdivide a mesh frame by generating a new vertex between two connected vertices within the mesh frame. By repeating the generation of new vertices, the number of vertices in the mesh frame can be set to a predetermined number. Through repeated subdivision across the entire mesh frame (in other words, multiple executions of subdivision), multiple Level of Detail (LoD) hierarchies are generated.

[0330] The displacement vector calculator 4803 obtains the original mesh frame acquired by the encoding device and also obtains the sub-subdivided base mesh frame from the sub-subdivider 4802. The displacement vector calculator 4803 calculates a displacement vector from the vertices of the base mesh frame and the vertices generated by the sub-subdivision of the base mesh frame, pointing to the corresponding vertices of the original mesh frame. The displacement vector calculator 4803 provides the calculated displacement vector to the wavelet converter 4804.

[0331] The wavelet transformer 4804 obtains transformation coefficients (also called wavelet coefficients) by applying a wavelet transform to the displacement vector calculated by the displacement vector calculator 4803. The wavelet transformer 4804 provides the obtained wavelet coefficients to the interpreter 4805. In the wavelet transform, the wavelet transformer 4804 can calculate wavelet coefficients representing various components from low-frequency to high-frequency components by assigning vertices to multiple LoD hierarchies and applying, for example, a lifting transform to the displacement vectors of those vertices.

[0332] The interpreter 4805 calculates the predicted residual of the wavelet coefficients of the displacement vector of the frame to be encoded using interpretation. Specifically, the interpreter 4805 calculates the predicted residual of the wavelet coefficients of the displacement vector of the frame to be encoded by interpreting the wavelet coefficients of the displacement vector of the frame to be encoded using the wavelet coefficients of the displacement vector of the encoded frame (also called the reference frame) stored in the reference buffer 4813.

[0333] The quantizer 4806 quantizes the prediction residuals of the wavelet coefficients calculated by the interpreter 4805. The quantizer 4806 can quantize the prediction residuals of the wavelet coefficients for each LoD level. The quantizer 4806 provides the quantized prediction residuals to the image packer 4807 and the inverse quantizer 4811.

[0334] The image packer 4807 generates an image containing the predicted residuals quantized by the quantizer 4806. The image packer 4807 can generate the above image by mapping the predicted residuals quantized by the quantizer 4806 to pixels of a two-dimensional image format. The image packer 4807 provides the generated image to the video encoder 4808. In the process of mapping the quantized predicted residuals to pixels of a two-dimensional image format, mapping information representing the assignment of the quantized predicted residuals to pixels of a two-dimensional image format may be used.

[0335] The video encoder 4808 encodes the image generated by the image packer 4807 into a bitstream (also called a displacement bitstream) (in other words, it generates a displacement bitstream). The video encoder 4808 outputs the displacement bitstream. The displacement bitstream may be a bitstream that contains displacement information in image format. The image format may be, for example, a format that contains two chroma information and one luma information. The video encoder 4808 can utilize a general-purpose module that has the function of converting an image to a bitstream. By utilizing a highly reliable general-purpose module as the video encoder 4808, the above function can be executed more reliably.

[0336] The inverse quantizer 4811 generates predicted residuals of wavelet coefficients by inverse quantizing the predicted residuals quantized by the quantizer 4806. Specifically, the inverse quantizer 4811 can inverse quantize the predicted residuals quantized by the quantizer 4806 for each LoD level, thereby generating predicted residuals. The inverse quantizer 4811 provides the generated predicted residuals of wavelet coefficients to the reconstructor 4812.

[0337] The reconstructor 4812 reconstructs (also called reconstructing) the wavelet coefficients from the predicted residuals of the wavelet coefficients provided by the inverse quantizer 4811 and the reference frame stored in the reference buffer 4813. The reconstructor 4812 stores the reconstructed wavelet coefficients in the reference buffer 4813.

[0338] The reference buffer 4813 is a memory device that stores, for example, the wavelet coefficients of a reference frame. The wavelet coefficients stored in the reference buffer 4813 can be used for interprediction by the interpreter 4805.

[0339] Figure 50 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.

[0340] The decoding device comprises a video decoder 5001, an image umpacker 5002, an inverse quantizer 5003, a reconstructor 5004, an inverse wavelet converter 5005, a reconstructor 5006, and a reference buffer 5011.

[0341] The video decoder 5001 acquires a displacement bitstream and decodes the acquired displacement bitstream into an image. The image may be an image contained in which quantized wavelet coefficients are mapped to pixels of a two-dimensional image format. The video decoder 5001 provides the image to the image unpacker 5002. The video decoder 5001 can utilize a general-purpose module that has the function of converting a bitstream to an image. By utilizing a highly reliable general-purpose module as the video decoder 5001, the above function can be performed more reliably.

[0342] The image umpacker 5002 extracts quantized wavelet coefficients from the image provided by the video decoder 5001. A mapping that represents the assignment of the quantized wavelet coefficients to pixels in a two-dimensional image format may be used for the process of extracting the quantized wavelet coefficients from the image. The image umpacker 5002 provides the quantized wavelet coefficients extracted from the image to the inverse quantizer 5003.

[0343] The inverse quantizer 5003 generates predicted residuals of wavelet coefficients by inverse quantizing the predicted residuals of the quantized wavelet coefficients provided by the image umpacker 5002. Specifically, the inverse quantizer 5003 can generate predicted residuals of wavelet coefficients by inverse quantizing the predicted residuals of the quantized wavelet coefficients for each LoD hierarchy.

[0344] The reconstructor 5004 reconstructs (also called reconstructing) the transformation coefficients of the frame to be decoded from the predicted residual of the transformation coefficients of the frame to be decoded and the transformation coefficients of the reference frame. The reconstructor 5004 provides the reconstructed transformation coefficients to the inverse wavelet converter 5005 and the reference buffer 5011.

[0345] The inverse wavelet transformer 5005 generates a displacement vector (corresponding to the decoded displacement vector) by applying an inverse wavelet transform to the wavelet coefficients provided by the reconstructor 5004. The inverse wavelet transform is equivalent to the inverse transform of the wavelet transform performed by the wavelet transformer 4804. Specifically, the inverse wavelet transformer 5005 can calculate the vertex displacement vector by applying an inverse lifting transform to the wavelet coefficients in the inverse wavelet transform. The inverse wavelet transformer 5005 provides the generated decoded displacement vector to the reconstructor 5006.

[0346] The reconstructor 5006 reconstructs the mesh (corresponding to the decoded mesh) using the decoded displacement vectors and the decoded base mesh provided by the inverse wavelet transformer 5005. The reconstructor 5006 outputs the reconstructed decoded mesh.

[0347] The reference buffer 5011 is a memory device that stores, for example, the wavelet coefficients of a reference frame. The wavelet coefficients stored in the reference buffer 5011 can be used by the reconstructor 5004 to reconstruct the conversion coefficients of the frame to be decoded.

[0348] Figure 51 is a block diagram showing an example configuration of the encoding device according to this embodiment.

[0349] The encoding device shown in Figure 51 comprises a decimator 5701, a subdivider 5702, a displacement vector calculator 5703, a wavelet converter 5704, an LoD-based inter predictor 5705, a quantizer 5706, a switch 5707, an image packer 5708, a video encoder 5709, an arithmetic encoder 5710, an inverse quantizer 5711, a reconstructor 5712, and a reference buffer 5713.

[0350] The decimator 5701, the sub-divider 5702, and the displacement vector calculator 5703 are the same as the decimator 4801, sub-divider 4802, and displacement vector calculator 4803 shown in Figure 49, respectively.

[0351] The wavelet transformer 5704 obtains transformation coefficients (also called wavelet coefficients) by applying a wavelet transform to the displacement vector calculated by the displacement vector calculator 5703. The wavelet transformer 5704 provides the obtained wavelet coefficients to the LoD-based interpreter 5705. In the wavelet transform, the wavelet transformer 5704 can calculate wavelet coefficients representing various components from low-frequency to high-frequency components by assigning vertices to multiple LoD hierarchy levels and applying, for example, a lifting transform to the displacement vectors of those vertices.

[0352] The LoD-based interpreter 5705 outputs the predicted residual of the wavelet coefficients of the displacement vector of the frame to be encoded using interpretation for each LoD. Specifically, for each LoD, the LoD-based interpreter 5705 outputs the predicted residual of the wavelet coefficients of the displacement vector of the frame to be encoded by interpreting the wavelet coefficients of the displacement vector of the frame to be encoded using the wavelet coefficients of the displacement vector of the encoded frame (also called a reference frame) stored in the reference buffer 5713. The LoD-based interpreter 5705 can determine and switch for each LoD whether or not to encode the transformation coefficients of the displacement vector using interpretation.

[0353] The quantizer 5706 quantizes the predicted residuals of the wavelet coefficients calculated by the LoD-based interpreter 5705. The quantizer 5706 quantizes the predicted residuals of the wavelet coefficients for each LoD level. The quantizer 5706 provides the quantized predicted residuals to the image packer 5708 or the arithmetic encoder 5710 via the switch 5707, and also to the inverse quantizer 5711.

[0354] Switch 5707 is a switch that determines whether the predicted residuals quantized by the quantizer 5706 are provided to the image packer 5708 or to the arithmetic encoder 5710.

[0355] The image packer 5708 and the video encoder 5709 are the same as the image packer 4807 and video encoder 4808 shown in Figure 49, respectively.

[0356] The arithmetic encoder 5710 encodes the prediction residuals quantized by the quantizer 5706 into a bitstream (also called a displacement bitstream) using arithmetic coding (in other words, it generates a displacement bitstream). The arithmetic encoder 5710 outputs the displacement bitstream.

[0357] The inverse quantizer 5711 generates predicted residuals of wavelet coefficients by inverse quantizing the predicted residuals quantized by the quantizer 5706. Specifically, the inverse quantizer 5711 generates predicted residuals by inverse quantizing the predicted residuals quantized for each LoD level by the quantizer 5706, for each LoD level. The inverse quantizer 5711 provides the generated predicted residuals of wavelet coefficients to the reconstructor 5712.

[0358] The reconstructor 5712 reconstructs (also called reconstructing) the wavelet coefficients from the predicted residuals of the wavelet coefficients provided by the inverse quantizer 5711 and the reference frame stored in the reference buffer 5713. The reconstructor 5712 stores the reconstructed wavelet coefficients in the reference buffer 5713.

[0359] The reference buffer 5713 is a memory device that stores, for example, the wavelet coefficients of a reference frame. The wavelet coefficients stored in the reference buffer 5713 can be used for interpretation by the LoD-based interpreter 5705.

[0360] The encoding device shown in Figure 51 can switch between encoding the transformation coefficients of the displacement vector using the video encoder 5709 or encoding them using arithmetic encoding with the arithmetic encoder 5710. This allows the encoding device to efficiently encode the displacement vector using arithmetic encoding even when the video encoder 5709 cannot be used.

[0361] Even when using arithmetic coding, it is possible to encode the displacement vector transformation coefficients by switching whether or not to use interpretation for each LoD. Generally, arithmetic coding becomes more efficient as the change in the input value decreases, so coding efficiency can be improved by suppressing the change in the predicted residual value of the displacement vector transformation coefficients by interpretation for each LoD.

[0362] Furthermore, information indicating whether the transformation coefficients of the displacement vector were encoded using the video encoder 5709 or by arithmetic encoding using the arithmetic encoder 5710 may be added to the header. This allows the decoding device to appropriately switch the decoding method by referring to the header information.

[0363] In Figure 51, an example was given showing the case where the transformation coefficients of the displacement vector are either encoded using the video encoder 5709 or encoded by arithmetic coding using the arithmetic encoder 5710. However, a configuration that always uses arithmetic coding is also possible. With this configuration, it is possible to improve the coding efficiency when arithmetic coding the transformation coefficients of the displacement vector compared to the case where interpretation is used at all LoD levels.

[0364] The decoding device shown in Figure 52 comprises a video decoder 5801, an image umpacker 5802, an arithmetic decoder 5803, a switch 5804, an inverse quantizer 5805, a reconstructor 5806, an inverse wavelet converter 5807, a reconstructor 5808, and a reference buffer 5811.

[0365] The video decoder 5801 and the image ampucker 5802 are the same as the video decoder 5001 and the image ampucker 5002 shown in Figure 50, respectively.

[0366] The arithmetic decoder 5803 acquires the displacement bitstream and arithmetically decodes the predicted residuals contained in the acquired displacement bitstream. The arithmetic decoder 5803 may also decode various header information.

[0367] Switch 5804 is a switch that toggles whether to provide the inverse quantizer 5805 with the predicted residuals provided by the image umpacker 5802, or with the inverse quantizer 5805 with the predicted residuals provided by the arithmetic decoder 5803.

[0368] The inverse quantizer 5805 generates predicted residuals of wavelet coefficients by inverse quantizing the predicted residuals of quantized wavelet coefficients provided from the image umpacker 5802 or arithmetic decoder 5803 via switch 5804. Specifically, the inverse quantizer 5805 generates predicted residuals of wavelet coefficients by inverse quantizing the predicted residuals of quantized wavelet coefficients for each LoD hierarchy.

[0369] The reconstructor 5806 reconstructs (also called reconstructs) the transformation coefficients of the frame to be decoded from the predicted residual of the transformation coefficients of the frame to be decoded and the transformation coefficients of the reference frame. The reconstructor 5806 provides the reconstructed transformation coefficients to the inverse wavelet converter 5807 and the reference buffer 5811. The reconstructor 5806 may determine from the header information whether interpretation was applied for each Line of Data (LoD) and switch the reconstruction method.

[0370] The inverse wavelet transformer 5807 generates a displacement vector (corresponding to the decoded displacement vector) by applying an inverse wavelet transform to the wavelet coefficients provided by the reconstructor 5806. The inverse wavelet transform is equivalent to the inverse transform of the wavelet transform performed by the wavelet transformer 5704. Specifically, the inverse wavelet transformer 5807 can calculate the vertex displacement vector by applying an inverse lifting transform to the wavelet coefficients in the inverse wavelet transform.

[0371] The inverse wavelet converter 5807 provides the generated decoded displacement vector to the reconstructor 5808.

[0372] The reconstructor 5808 reconstructs the mesh (corresponding to the decoded mesh) using the decoded displacement vectors and the decoded base mesh provided by the inverse wavelet transformer 5807. The reconstructor 5808 outputs the reconstructed decoded mesh.

[0373] The reference buffer 5811 is a memory device that stores, for example, the wavelet coefficients of a reference frame. The wavelet coefficients stored in the reference buffer 5811 can be used by the reconstructor 5806 to reconstruct the conversion coefficients of the frame to be decoded.

[0374] The decoding device shown in Figure 52 decodes the header information and may determine whether the encoding device encoded the transformation coefficients of the displacement vector using a video encoder (e.g., video encoder 5709) or arithmetic encoding using an arithmetic encoder (e.g., arithmetic encoder 5710), and switch the decoding method. This allows the decoding device to appropriately decode a bitstream in which the displacement vector has been efficiently encoded using arithmetic encoding even when a video encoder (e.g., video encoder 5709) cannot be used.

[0375] The method for generating the Line of Deposition (LoD) will be explained below.

[0376] Figures 53 and 54 are explanatory diagrams showing the method for generating LoD in this embodiment.

[0377] When encoding displacement vectors of three-dimensional points, encoding devices may classify each three-dimensional point into one or more levels using the positional information of the three-dimensional points before encoding. Here, each level used for classification is called a Level of Detail (LoD). Each LoD is assigned a unique identifier (e.g., a number). For example, the 0th LoD is also called LoD0, the 1st LoD is also called LoD1, the nth LoD is also called LoDn, and the (n-1)th LoD is also called LoD(n-1).

[0378] The method for generating the Line of Data (LoD) will be explained using Figures 53 and 54. If the encoding or decoding device cannot calculate the position or distance information of a three-dimensional point within the frame to be encoded or decoded, the position or distance information of a corresponding three-dimensional point within an already encoded or decoded frame may be used. This may allow for efficient encoding by classifying the three-dimensional points to be encoded or decoded into one or more layers.

[0379] Figure 53 shows the three-dimensional points to be encoded: points a0, a1, a2, b0, b1, b2, c0, c1, and c2. Note that d(x, y) represents the distance between point x and point y.

[0380] By setting the threshold values ​​for each layer of the Line of D to be larger for higher layers (layers closer to Line of D0), higher layers become point groups where the distance between three-dimensional points is greater (also called sparse point groups), and lower layers become point groups where the distance between three-dimensional points is smaller (also called dense point groups). Here, Line of D0 is the top layer (see Figure 54).

[0381] Point y belongs to the same LoD as point x if its distance d(x,y) from point x is greater than the threshold of the LoD to which point x belongs, but less than or equal to the threshold of the LoD above that LoD. Furthermore, if point x belongs to the highest layer, LoD0, then point y belongs to the same LoD as point x if its distance d(x,y) from point x is greater than the threshold of the LoD to which point x belongs.

[0382] First, the encoding device selects point a0 as the initial point and assigns it to LoD0. Next, the encoding device extracts point a1 whose distance from point a0 is greater than the LoD0 threshold Thres_Lod[0] and assigns it to LoD0. Next, the encoding device extracts point a2 whose distance from point a1 is greater than the LoD0 threshold Thres_Lod[0] and assigns it to LoD0. In this way, the encoding device configures LoD0 such that the distance between each point in LoD0 is greater than the threshold Thres_Lod[0].

[0383] Next, the encoding device selects point b0, which has not yet been assigned an LoD, and assigns it to LoD1. Next, the encoding device selects point b1, which is farther from point b0 than the LoD1 threshold Thres_Lod[1] and has not yet been assigned an LoD, and assigns it to LoD1. Next, the encoding device selects point b2, which is farther from point b1 than the LoD1 threshold Thres_Lod[1] and has not yet been assigned an LoD, and assigns it to LoD1. In this way, the encoding device configures LoD1 such that the distance between each point in LoD1 is greater than the threshold Thres_Lod[1].

[0384] Next, the encoding device selects point c0, which has not yet been assigned an LoD, and assigns it to LoD2. Next, the encoding device selects point c1, which is farther from point c0 than the LoD2 threshold Thres_Lod[2] and has not yet been assigned an LoD, and assigns it to LoD2. Next, the encoding device selects point c2, which is farther from point c1 than the LoD2 threshold Thres_Lod[2] and has not yet been assigned an LoD, and assigns it to LoD2. In this way, LoD2 is configured such that the distance between each point in LoD2 is greater than the threshold Thres_Lod[2].

[0385] The threshold values ​​for each LoD may be added to the bitstream header. For example, in Figure 53, the threshold values ​​Thres_Lod[0], Thres_Lod[1], and Thres_Lod[2] may be added to the bitstream header.

[0386] Alternatively, the lowest layer of the LoD may be assigned to all three-dimensional points that have not yet been assigned an LoD. In this case, the threshold of the lowest layer of the LoD is not added to the header, which has the effect of reducing the amount of code in the header. For example, in the case of Figure 53, the encoding device may add the thresholds Thres_Lod[0] and Thres_Lod[1] to the header, but Thres_Lod[2] may not be added to the header, and the decoding device may estimate Thres_Lod[2] to be a value of 0.

[0387] Furthermore, the number of levels in the Level of Deposition (LoD) may be added to the header. This allows the decryption device to determine whether the LoD is at the lowest level.

[0388] Furthermore, if the LoD hierarchy is one level, that is, if the displacement vectors of three-dimensional points are encoded without generating an LoD, the encoding device may omit the LoD generation process described in the above example. Alternatively, the encoding device may apply the LoD generation method described in the above example, assuming the LoD hierarchy is 1. In this case, the encoding device may perform the LoD generation process assuming that all three-dimensional points belong to the same LoD. This allows the encoding device to reduce the processing time for LoD generation.

[0389] The encoding or decoding of displacement vectors described in this embodiment may be applied to methods other than the LoD generation method described above. For example, even when the LoD hierarchy to which the three-dimensional points belong is determined in advance, the encoding efficiency may be improved by applying the displacement vector encoding or decoding method described in this embodiment.

[0390] The method for generating the Level of Data (LoD) is not limited to the method described above. For example, as shown in Figures 35 and 37, the LoD to which a level belongs may be determined according to the number of subdivisions from the base mesh. For example, if the base mesh is subdivided twice, the three-dimensional points included in the base mesh may first be assigned to LoD0, the points generated by one subdivision from the three-dimensional points of the base mesh may be assigned to LoD1, and the points generated by two subdivisions may be assigned to LoD2. This reduces the processing time for LoD generation.

[0391] The method for selecting the initial three-dimensional points when constructing each LoD may depend on the coding order during displacement vector coding. For example, the coding device selects the first three-dimensional point coded during displacement vector coding as the initial point a0 of LoD0, and then selects points a1 and a2 from point a0 to construct LoD0. Then, the coding device may select the three-dimensional point that was coded earliest among the three-dimensional points that do not currently belong to LoD0 as the initial point b0 of LoD1. In other words, the coding device may select the three-dimensional point that was coded earliest among the three-dimensional points that do not belong to the LoDs of the hierarchy below LoD(n-1) as the initial point n0 of LoDn. This allows the same LoD as during coding to be constructed during decoding using the same initial point selection method (specifically, selecting the three-dimensional point that was coded earliest among the three-dimensional points that do not belong to the LoDs of the hierarchy below LoD(n-1) as the initial point n0 of LoDn), and the bitstream to be decoded appropriately.

[0392] Figure 55 is an explanatory diagram showing the method for generating predicted values ​​of the displacement vector in this embodiment.

[0393] The encoding device can generate predicted values ​​for the displacement vectors of three-dimensional points using information from the Line of Deposition (LoD).

[0394] For example, if the encoding device encodes the three-dimensional points in LoD0 sequentially, it may generate LoD1 using the encoded and decoded displacement vectors contained in LoD0 and LoD1. In this way, the encoding device can generate predicted values ​​of the displacement vectors of the three-dimensional points contained in LoDn using the encoded and decoded displacement vectors contained in LoDn' (where n'≦n).

[0395] Furthermore, the predicted displacement vector of a three-dimensional point can be generated by calculating the average of the displacement vectors of a certain number of three-dimensional points that are adjacent to the three-dimensional point to be encoded and have been encoded and decoded. This certain number is, for example, the number of adjacent points to the three-dimensional point to be encoded (e.g., N points). In this case, the value N is added to the bitstream header, etc.

[0396] The value N, which indicates the number of adjacent points (i.e., N points) used to calculate the predicted value, may be added to each three-dimensional point from which a predicted value is generated. This allows the encoding device to select an appropriate N adjacent points for each three-dimensional point from which a predicted value is generated, thereby improving the accuracy of the predicted value and reducing the prediction residual. Alternatively, the encoding device may add the value N to the bitstream header and fix it within the bitstream (in other words, the value N may be used as a common fixed value in encoding the three-dimensional points included in the bitstream). This eliminates the need for the encoding device to encode or decode the value N for each three-dimensional point, thus reducing the processing load. Furthermore, the encoding device may encode the value N separately for each Line of D (LoD). This may allow the encoding device to improve encoding efficiency by selecting an appropriate value N for each LoD.

[0397] Furthermore, the predicted displacement vector of a three-dimensional point may be calculated from the weighted average of N neighboring points that have been encoded and decoded. The encoding device can, for example, perform a weighted average using the distance information between the three-dimensional point to be encoded and each of the N neighboring points. This will be explained using Figure 55.

[0398] When an encoding device encodes using a different value N for each Line of Data (LoD), it may set the value of N to be larger for higher layers of the LoD and smaller for lower layers. In the higher layers of the LoD, the distance between three-dimensional points belonging to that LoD is relatively large, so by setting a large value of N, it may be possible to improve prediction accuracy by selecting and averaging a relatively large number of surrounding three-dimensional points. In the lower layers of the LoD, the distance between three-dimensional points belonging to that LoD is relatively small, so by setting a small value of N, it may be possible to perform efficient prediction while reducing the processing load of averaging.

[0399] The predicted value of point P belonging to LoDN is generated from the reconstructed point P' belonging to LoDN' (where N'≦N). Here, it is assumed that adjacent points are selected from point P' based on connectivity and distance.

[0400] Furthermore, the predicted displacement vector can be calculated from an unweighted average value. This reduces the amount of processing required.

[0401] As shown in Figure 55, point a2 is predicted from points a0 and a1. Similarly, point b2 is predicted from points a0, a1, a2, b0, and b1. The points selected as adjacent points used for prediction may change depending on the number of adjacent points N used for prediction. For example, when N=5, points a0, a1, a2, b0, and b1 may be selected as adjacent points to point b2, and when N=4, points a0, a1, a2, and b1 may be selected based on distance information.

[0402] For example, the three-dimensional points included in the base mesh are included in LoD0, the three-dimensional points generated from the three-dimensional points of the base mesh by one subdivision are included in LoD1, and the three-dimensional points generated from the three-dimensional points of the base mesh by two subdivisions are included in LoD2.

[0403] For example, if the weighted mean of adjacent points is used for prediction, the predicted value a2p for point a2 is calculated by the weighted mean of points a0 and a1 (see Equations 1 and 2). Here, A i This is the value of the displacement vector of point ai.

[0404]

[0405] however,

[0406]

[0407] Furthermore, the predicted value b2p for point b2 is calculated by the weighted average of points a0, a1, a2, b0, and b1 (see Equations 3, 4, and 5). Here, B i This is the value of the displacement vector of point bi.

[0408]

[0409] however,

[0410]

[0411]

[0412] Furthermore, when generating predicted displacement vectors, it is possible to avoid referencing the same hierarchical level. This can reduce the amount of processing required. Also, when generating a 3D point at the midpoint of two 3D points by subdivision, the weight w i This can be fixed at 0.5. This will reduce the processing load.

[0413] The encoding device may encode the displacement vector of a three-dimensional point by calculating the difference between the predicted value generated from adjacent points of the three-dimensional point and the three-dimensional point itself (also called a conversion coefficient, see (Equations 6) and (7) below), and then quantizing the calculated conversion coefficient. Here, the conversion coefficient a2c is the conversion coefficient for point a2, and the conversion coefficient b2c is the conversion coefficient for point b2.

[0414]

[0415]

[0416] For example, an encoding device can perform quantization by dividing the conversion coefficients by the quantization scale. In this case, the smaller the quantization scale, the smaller the error that can occur due to quantization (quantization error), and conversely, the larger the quantization scale, the larger the quantization error.

[0417] The quantized value a2q is obtained by quantizing the conversion coefficient a2c, and the quantized value b2q is obtained by quantizing the conversion coefficient b2r (see Equations 8 and 9 below). QS_LoD0 is the quantization scale of LoD0, and QS_LoD1 is the quantization scale of LoD1.

[0418]

[0419]

[0420] Furthermore, the encoding device may change the quantization scale value for each Level of Derivation (LoD). For example, the quantization scale can be made smaller for higher-level LoDs and larger for lower-level LoDs. Since the displacement vector values ​​of three-dimensional points belonging to higher-level layers may be used as predicted values ​​for the displacement vectors of three-dimensional points belonging to lower-level layers, encoding efficiency can be improved by reducing the quantization scale of higher-level layers to suppress quantization errors that may occur in higher-level layers and increase the accuracy of the predicted values. The encoding device may also add the quantization scale to the header or other elements for each LoD. This allows the encoding device to contribute to the decoding device correctly decoding the quantization scale and appropriately decoding the bitstream.

[0421] The encoding device may convert the converted coefficients after quantization from signed integer values ​​to unsigned integer values. For example, the encoding device may convert the quantized value a2q, which is a signed integer, to the quantized value a2u, which is an unsigned integer, as shown below.

[0422] If the quantized value a²q is less than 0: a²u = -1 - (2 × a²q) Otherwise: a²u = 2 × a²q (Equation 10)

[0423] Alternatively, for example, the encoding device may convert the quantized value b2q, which is a signed integer value, into the quantized value b2u, which is an unsigned integer value, as shown below.

[0424] If the quantized value b²q is less than 0: b²u = -1 - (2 × b²q) Otherwise: b²u = 2 × b²q (Equation 11)

[0425] This has the advantage that the encoding device does not need to consider the occurrence of negative integers when entropy encoding the conversion coefficients.

[0426] Furthermore, the encoding device does not necessarily need to convert signed integer values ​​to unsigned integer values; for example, the sign bit can be separately entropy encoded.

[0427] Furthermore, the encoding method for the conversion coefficients is not limited to this. For example, the encoding device may arithmetically encode the sign bit representing the sign of the conversion coefficient and the binarized data of the absolute value of the conversion coefficient bit by bit using context. This may allow the encoding device to improve the encoding efficiency of the conversion coefficients of the displacement vector.

[0428] If quantization of the transformation coefficients of the displacement vector is not required, this process can be skipped, and the transformation coefficients can be arithmetically coded directly. This can reduce processing time.

[0429] Figure 56 is an explanatory diagram showing an example of how to calculate predicted values ​​in this embodiment.

[0430] Referring to Figure 56, we will explain an example of generating LoD and calculating the predicted displacement vectors for each three-dimensional point.

[0431] In Figure 56, points a0, a1, and a2 are three-dimensional points included in the base mesh and belong to LoD0. Points b0 and b1 are three-dimensional points generated by one subdivision from the three-dimensional points included in the base mesh and belong to LoD1. Points c0, c1, c2, and c3 are three-dimensional points generated by two subdivisions from the three-dimensional points included in the base mesh and belong to LoD2.

[0432] If point b0 is a three-dimensional point generated by subdividing from points a0 and a1, the predicted displacement vector of point b0 can be calculated using points a0 and a1.

[0433] Furthermore, if point c2 is a three-dimensional point generated by subdividing from points a1 and b1, the predicted value of the displacement vector of point c2 can be calculated using points a1 and b1.

[0434] Figure 57 is an explanatory diagram showing an example of the calculation of the conversion coefficient in this embodiment.

[0435] Referring to Figure 57, we will explain an example of calculating the transformation coefficient by subtracting the predicted values ​​from the displacement vectors of each three-dimensional point.

[0436] In FIG. 57, the conversion coefficients a0c, a1c, and a2c are the conversion coefficients included in the base mesh and belong to LoD0. The conversion coefficients b0c and b1c are the three-dimensional points generated by one sub-division from the three-dimensional points included in the base mesh and belong to LoD1. The conversion coefficients c0c, c1c, c2c, and c3c are the three-dimensional points generated by two sub-divisions from the three-dimensional points included in the base mesh and belong to LoD2.

[0437] The conversion coefficient b0c of the point b0 is obtained by subtracting the predicted value b0p of the point b0 from the value of the point b0.

[0438] b0c = b0 - b0p

[0439] The predicted value b0p may be the average value of the value of the point a0 and the value of the point a1.

[0440] b0p = (a0 + a1) / 2

[0441] Also, the conversion coefficient c2c of the point c2 is obtained by subtracting the predicted value c2p of the point c2 from the value of the point c2.

[0442] c2c = c2 - c2p

[0443] The predicted value c2p may be the average value of the value of the point a1 and the value of the point b1.

[0444] c2p = (a1 + b1) / 2

[0445] Note that an encoding method (lifting transform) may be applied in which the conversion coefficients of the displacement vectors of the three-dimensional points included in the lower layer of LoD are calculated, and the conversion coefficients are fed back to the upper layer and encoded. By applying the lifting transform, the conversion coefficients of the low-frequency components of the displacement vector can be collected in the upper layer, and the conversion coefficients of the high-frequency components of the displacement vector can be collected in the lower layer. Thereby, for example, the encoding efficiency can be improved by reducing the amount of information by quantizing the conversion coefficients of the high-frequency components in the lower layer.

[0446] FIG. 58 is an explanatory diagram showing an example of the process of the lifting transform in the present embodiment.

[0447] The lifting conversion process is executed, for example, by a wavelet transformer (e.g., wavelet transformer 4804 (see FIG. 49)).

[0448] FIG. 58 shows a method (lifting conversion) in which, for the displacement vector of three-dimensional points, an encoding device generates LoD by sub-division, performs a prediction process (also simply referred to as prediction) and an update process (also simply referred to as update), and encodes output values for each LoD.

[0449] It can also be said that the encoding device converts coefficients (also referred to as input, input value, or "in") input to the wavelet transformer into other coefficients (also referred to as output, output value, or "out") by lifting conversion. In FIG. 58, the coefficients input to the wavelet transformer are LoD0 in (also denoted as a0), LoD1 in (also denoted as b0), LoD2 in (also denoted as c0), and LoD3 in (also denoted as d0), etc. Further, the coefficients after conversion by the wavelet transformer are LoD0 out, LoD1 out, LoD2 out, and LoD3 out, etc.

[0450] Specifically, prediction at least includes a process of multiplying the coefficients of two points around a three-dimensional point (also referred to as a prediction target point) that is the target of prediction by weights and subtracting them from the coefficient of the prediction target point respectively. Update at least includes a process of multiplying the coefficient of the prediction target point by weights and adding them to the coefficients of the update target points respectively.

[0451] The encoding device performs the above prediction and update processes for each layer. As a result, it is possible to make the difference in coefficients small in prediction and to average the coefficients in update. As a result, the encoding device may be able to improve the encoding efficiency by adjusting the quantization value for each layer such that the lower layer (e.g., LoD3) contains more high frequencies and the upper layer (e.g., LoD0) contains more low frequencies.

[0452] In Figure 58, the prediction and update processes for LoD3 are shown when the target point is point d0, but the illustrations for the prediction and update processes when the target points are points d1, d2, and d3 are omitted. Similarly, the prediction and update processes for LoD2 are shown when the target point is point c0, but the illustrations for the prediction and update processes when the target point is point c1 are omitted.

[0453] While the points of this disclosure show an example of applying the lifting transformation when the number of subdivisions is 3 (in other words, when the number of LoD levels is 4), the points of this disclosure are not necessarily limited to this case, and may also be applied to cases with a different number of subdivisions or a different number of LoD levels.

[0454] Figure 59 is an explanatory diagram showing an example of the lifting transformation process in this embodiment.

[0455] For coefficients referenced in prediction and update, the weights multiplied by the referenced coefficients may be different for prediction and update. Furthermore, the weights multiplied by the referenced coefficients may be different for each hierarchy in both prediction and update. Additionally, the weights multiplied by the referenced coefficients may be different for the two coefficients referenced in both prediction and update.

[0456] More specifically, the predicted output value d0' of the target point d0 in LoD3 at the lowest level of the lifting transformation can be calculated as follows, with the weights of input a0 and c0 being predWeight20 (also written as pw20) and predWeight21 (also written as pw21), respectively.

[0457] d0'= d0 - (predWeight20 * a0 + predWeight21 * c0) (Equation 12)

[0458] Furthermore, the updated output value a0' of target point a0 in LoD0 and the updated output value c0' of target point c0 in LoD2 at the lowest level of the lifting transformation can be calculated as follows, with the weights of input a0 and c0 being updateWeight20 (also written as uw20) and updateWeight21 (also written as uw21), respectively.

[0459] a0' = a0 + updateWeight20 * d0' c0' = c0 + updateWeight21 * d0' (Equation 13)

[0460] The second and subsequent levels from the lowest level of the lifting transformation can be calculated similarly as follows.

[0461] Specifically, the predicted output value c0'' for the target point c0 can be calculated as follows, using predWeight10 (also written as pw10) and predWeight11 (also written as pw11) as the weights of input a0' and b0, respectively.

[0462] c0'' = c0' - (predWeight10 * a0' + predWeight11 * b0) (Equation 14)

[0463] Furthermore, the updated output value a0'' for target point a0 and the updated output value b0'' for target point b0 can be calculated as follows, using updateWeight10 (also written as uw10) and updateWeight11 (also written as uw11) as the weights of input a0'' and b0, respectively.

[0464] a0'' = a0' + updateWeight10 * c0'' b0'' = b0 + updateWeight11 * c0'' (Equation 15)

[0465] Furthermore, the predicted output value b0'' for the target point b0 can be calculated as follows, using predWeight00 (also written as pw00) and predWeight01 (also written as pw01) as the weights of input a0'' and a1, respectively.

[0466] b0''' = b0'' - (predWeight00 * a0'' + predWeight01 * a1) (Equation 16)

[0467] Furthermore, the updated output value a0''' for target point a0 and the updated output value a1''' for target point a1 can be calculated as follows, using updateWeight00 (also written as uw00) and updateWeight01 (also written as uw01) as the weights of input a0'' and a1, respectively.

[0468] a0''' = a0'' + updateWeight00 * b0''' a1''' = a1 + updateWeight01 * b0''' (Equation 17)

[0469] In the above, it can also be said that the encoding device, when the target point is a three-dimensional point d0, calculates a predicted value of the displacement data of the three-dimensional point d0 using the displacement data of a three-dimensional point a0 belonging to a different hierarchy (LoD0) from the target point, and the weight pw20 determined according to the hierarchy to which the three-dimensional point d0 belongs (i.e., LoD3).

[0470] Furthermore, it can be said that when the three-dimensional point d0 is the target point, the encoding device calculates a predicted value of the displacement data of the three-dimensional point d0 using the displacement data of a three-dimensional point c0 belonging to a different hierarchy (LoD2) from the target point, and the weight pw21 determined according to the hierarchy to which the three-dimensional point d0 belongs (i.e., LoD3).

[0471] Furthermore, it can be said that when the three-dimensional point c0 is the target point, the encoding device calculates a predicted value of the displacement data of the three-dimensional point c0 using the displacement data of a three-dimensional point a0 belonging to a different hierarchy (LoD0) from the target point, and a weight pw10 determined according to the hierarchy to which the three-dimensional point c0 belongs (i.e., LoD2).

[0472] Furthermore, it can be said that when the three-dimensional point c0 is the target point, the encoding device calculates a predicted value of the displacement data of the three-dimensional point c0 using the displacement data of a three-dimensional point b0 belonging to a different hierarchy (LoD1) than the target point, and the weight pw11 determined according to the hierarchy to which the three-dimensional point c0 belongs (i.e., LoD2).

[0473] Furthermore, it can be said that when the three-dimensional point b0 is the target point, the encoding device calculates a predicted value of the displacement data of the three-dimensional point b0 using the displacement data of the three-dimensional point a0, which belongs to a different hierarchy (LoD0) than the target point, and the weight pw00 which is determined according to the hierarchy to which the three-dimensional point c0 belongs (i.e., LoD1).

[0474] Furthermore, it can be said that when the three-dimensional point b0 is the target point, the encoding device calculates a predicted value of the displacement data of the three-dimensional point b0 using the displacement data of the three-dimensional point a1, which belongs to a different hierarchy (LoD0) than the target point, and the weight pw11, which is determined according to the hierarchy to which the three-dimensional point b0 belongs (i.e., LoD1).

[0475] Furthermore, the prediction weights predWeight (specifically predWeight20, predWeight21, predWeight10, predWeight11, predWeight00, and predWeight01, and so on) and the update weights updateWeight (specifically updateWeight20, updateWeight21, updateWeight10, updateWeight11, updateWeight00, and updateWeight01, and so on) may be encoded by the encoding device and transmitted to the decoding device. In this way, the decoding device can correctly restore the weights by decoding the prediction weights predWeight and the update weights updateWeight.

[0476] Furthermore, the prediction weight predWeight and the update weight updateWeight may be fixed values, such as 0.5 or 1.0. In that case, the amount of coding required for encoding the weights can be reduced by inferring the prediction weight predWeight and the update weight updateWeight as fixed values ​​predetermined by the decoding device.

[0477] Further, only a part of the prediction weight predWeight and the update weight updateWeight may be transmitted from the encoding device to the decoding device, and the decoding device may infer the remaining part excluding the part as a fixed value.

[0478] Note that the prediction weight predWeight and the update weight updateWeight may be different as two reference coefficients, or may be shared.

[0479] Further, the prediction weight predWeight and the update weight updateWeight may be different for each lifting layer, or may be shared, or only some of the layers may be shared.

[0480] Further, the prediction weight predWeight and the update weight updateWeight may be different, or may be shared, or only some of the layers may be shared.

[0481] Note that when performing lifting conversion, the encoding device may change the weight multiplied by the prediction reference coefficient for each lifting layer. For example, since the encoding device predicts from a spatially close distance in the lowest layer and from the farthest distance in the highest layer, the effect of prediction may be lost as the layer progresses. Also, since prediction or update will be performed from a coefficient obtained by performing prediction or update multiple times as the layer progresses to the upper layer, there is also a possibility that the error accumulates and the effect of prediction is lost. Therefore, the encoding device may reduce the weight so as to reduce the influence of prediction as the layer becomes higher. Thereby, the encoding device can improve the encoding efficiency by suppressing the prediction error in the upper layer.

[0482] For example, the encoding device may reduce the weight so as to reduce the influence of prediction in the highest layer, and specifically may do so as follows.

[0483] predWeight20 = 0.5 predWeight21 = 0.5 predWeight10 = 0.5 predWeight11 = 0.5 predWeight00 = 0.25 predWeight01 = 0.25 (Formula 18)

[0484] This allows the encoding device to mitigate the impact of prediction errors at the highest level, potentially improving encoding efficiency.

[0485] Furthermore, the encoding device may gradually decrease the weights at each layer, or it may use different values ​​for the weights of the two reference coefficients. For example, it may be done as follows.

[0486] predWeight20 = 0.5 predWeight21 = 0.5 predWeight10 = 0.35 predWeight11 = 0.35 predWeight00 = 0.25 predWeight01 = 0.25 (Formula 19)

[0487] This allows the encoding device to mitigate the impact of prediction errors at higher levels, potentially improving encoding efficiency.

[0488] Furthermore, the encoding device may set the top layer to be unpredictable by configuring the weights as follows.

[0489] predWeight00 = 0.0 predWeight01 = 0.0 (Equation 20)

[0490] This allows the encoding device to mitigate the impact of prediction errors at higher levels, potentially improving encoding efficiency.

[0491] Furthermore, if the prediction effect of the top layer is significant, the encoding device may use the same weights as other layers. This may improve encoding efficiency even when the prediction effect is significant at higher layers.

[0492] As described above, encoding devices may be able to reduce the impact of predictions and suppress prediction deterioration by decreasing the weights of the top-level or higher-level layers. Furthermore, the effect of predictions may improve encoding efficiency.

[0493] Furthermore, when performing a lifting transformation, the encoding device may change the weight multiplied by the reference coefficient for prediction for each lifting hierarchy. For example, the encoding device predicts from spatially close distances at the lowest hierarchy and from the furthest distances at the highest hierarchy, so the effectiveness of prediction may decrease as the hierarchy level increases. Also, as the hierarchy level increases, predictions are made from coefficients that have been predicted or updated multiple times, so errors may accumulate and the effectiveness of prediction may decrease.

[0494] Therefore, the encoding device may reduce the weights in the highest layer to minimize the influence of predictions, and specifically, it may do the following:

[0495] predWeight20 = 0.5 predWeight21 = 0.5 predWeight10 = 0.5 predWeight11 = 0.5 predWeight00 = 0.25 predWeight01 = 0.25 (Formula 21)

[0496] Furthermore, the encoding device may change the weights multiplied by the update reference coefficients for each hierarchical level. For example, if the encoding device changes the prediction weights at the top level as described above so that the coefficients become smaller, the predicted coefficients will become larger (i.e., the prediction residuals will become larger). Therefore, the update weights may be reduced in order to return the predicted coefficients to a size closer to the original coefficients (i.e., the coefficients before prediction), and this may be done as follows.

[0497] updateWeight20 = 1.0 updateWeight21 = 1.0 updateWeight10 = 1.0 updateWeight11 = 1.0 updateWeight00 = 0.9 updateWeight01 = 0.9 (Formula 22)

[0498] Furthermore, the encoding device may gradually decrease the weights for both the prediction weight (predWeight) and the update weight (updateWeight) at each hierarchical level, or it may use different values ​​for the weights of the two reference coefficients.

[0499] Alternatively, the top level can be left unpredicted by setting the weights as follows.

[0500] predWeight00 = 0.0 predWeight01 = 0.0 (Equation 23)

[0501] Furthermore, if the prediction effect of the top layer is significant, the encoding device may use the same weights for both prediction (predWeight) and update (updateWeight) as for the other layers.

[0502] Furthermore, the encoding device may be able to reduce the impact of predictions and suppress prediction deterioration by decreasing the weights of the top layer. Additionally, the effect of predictions may improve encoding efficiency.

[0503] Furthermore, when performing a lifting transform, the encoding device may adaptively determine the weights multiplied by the reference coefficients of the prediction.

[0504] Next, I will explain the inverse wavelet transform.

[0505] As described above, the displacement bitstream decoded using the video decoder is subjected to image unfolding and inverse quantization, followed by an inverse wavelet transform. The decoded data after inverse quantization is the value of the initial displacement coefficient corresponding to each vertex (each 3D point) of the 3D mesh frame. These displacement coefficient values ​​are added to the vertex's 3D position in the 3D mesh frame, thereby obtaining the final displacement position of the vertex.

[0506] By further fine-tuning the displacement coefficient using the inverse wavelet transform, the displacement values ​​of each vertex of the three-dimensional mesh frame can be calculated more accurately. For example, the displacement coefficient is fine-tuned using the inverse wavelet transform only when a flag is set to a predetermined value. The decoder 200 performs an inverse wavelet transform using, for example, a weight parameter and a lifting offset, and calculates the displacement parameter of the vertices of the three-dimensional mesh frame as follows.

[0507] (Displacement parameter) = (Weight parameter) × ((Displacement coefficient) + (Lifting offset))

[0508] For example, in the reconstructor, displacement parameters are added to the vertex coordinates of the 3D mesh frame, and the vertex coordinates are moved to the new position. Another example is that weight parameters and lifting offsets are signaled to the bitstream. Yet another example is that the weight parameters and lifting offsets used to perform a lifting transformation to calculate the displacement parameters of the vertices of the 3D mesh frame are derived from the weight parameters and lifting offsets corresponding to different vertices in the currently encoded 3D mesh frame.

[0509] This derivation of weight parameters and lifting offsets is also called the interpretation mode. Furthermore, this derivation of weight parameters and lifting offsets is also called the merge prediction mode. Finally, this derivation of weight parameters and lifting offsets is also called the skip prediction mode.

[0510] Referring to Figure 57, an example of calculating the lifting offset from the transformation coefficients of the displacement vectors of each three-dimensional point will be explained.

[0511] a0c, a1c, a2c, b0c, b1c, c0c, c1c, c2c, and c3c are transformation coefficients calculated by subtracting their respective predicted values ​​from the displacement vectors of each three-dimensional point.

[0512] Note that lifting_offset[n] indicates the lifting offset of the nth (n: non-negative integer) LoD.

[0513] Furthermore, lifting_offset[0] represents the average value of a0c, a1c, and a2c.

[0514] Furthermore, lifting_offset[1] represents the average value of b0c and b1c.

[0515] Furthermore, lifting_offset[2] represents the average value of c0c, c1c, c2c, and c3c.

[0516] The lifting offset may be calculated for each Level of Disorder (LoD). For example, the average value of the transformation coefficients obtained after applying a lifting transformation to the displacement vectors of the three-dimensional points belonging to that LoD may be used as the lifting offset for each LoD. This allows for the calculation of an appropriate lifting offset for each LoD.

[0517] Furthermore, any method other than the average value may be used to calculate the lifting offset. For example, the encoding device 100 may calculate a weighted average, maximum, or minimum value from the transformation coefficients after the lifting transformation is applied, and set the value with the highest encoding efficiency among the calculated values ​​as the lifting offset. This can improve the encoding efficiency.

[0518] Next, we will explain an example in which the encoding device 100 encodes the displacement vectors of each three-dimensional point by subtracting the lifting offset from the transformation coefficients.

[0519] Figure 60 is an explanatory diagram showing an example of the process of encoding by subtracting the lifting offset from the transformation coefficient of the displacement vector of each three-dimensional point according to the embodiment. Note that lo[n] represents lifting_offset[n].

[0520] As described above, the encoding device 100 may calculate a difference value by subtracting the lifting offset value for each LoD from the conversion coefficients after the lifting transformation, and encode the calculated difference value. In this case, when the values ​​of the conversion coefficients after the lifting transformation are close to the values ​​for each LoD, the difference value approaches 0 by subtracting the lifting offset value, which is calculated as the average, from the conversion coefficients after the lifting transformation. Therefore, encoding efficiency can be improved by applying binarization or arithmetic coding.

[0521] Furthermore, the encoding device 100 may add the lifting offset value for each LoD to the bitstream. This allows the decoding device 200 to decode the lifting offset for each LoD from the bitstream and decode the conversion coefficients by adding the lifting offset for each LoD to the difference value of the displacement vector of the binarized or arithmetic-decoded three-dimensional points. As a result, the decoding device 200 can appropriately decode a bitstream in which the encoding efficiency has been improved by subtracting the lifting offset from the conversion coefficients.

[0522] As an example, the lifting offset can be derived by dividing the lifting offset numerator parameter by the lifting offset denominator parameter, as follows:

[0523] (Lifting offset) = (Numerator of lifting offset) / (Denominator of lifting offset)

[0524] For example, the numerator (numerator parameter) of the lifting offset is a positive number. Another example is that the numerator parameter of the lifting offset is a negative number. Also, for example, the denominator (denominator parameter) of the lifting offset is a positive number. Another example is that the denominator parameter of the lifting offset is a negative number.

[0525] Another example is that the numerator and denominator of the lifting offset are signaled to the bitstream. Another example is that the numerator of the lifting offset is derived by adding the numerator of the lifting offset corresponding to a previously encoded vertex and the delta value of the numerator of the lifting offset (delta lifting offset numerator), as follows:

[0526] (Numerator of lifting offset) = (Numerator of lifting offset corresponding to previously coded vertices) + (Delta value of the numerator of lifting offset)

[0527] Another example is that the denominator of the lifting offset is derived by adding the denominator of the lifting offset corresponding to a previously coded (encoded or decoded) vertex to the delta value of the lifting offset denominator (delta lifting offset denominator), as follows:

[0528] (Denominator of the lifting offset) = (Denominator of the lifting offset corresponding to previously coded vertices) + (Delta value of the denominator of the lifting offset)

[0529] Next, we will describe the encoding device 100 and decoding device 200 when the mesh is divided into multiple submeshes.

[0530] Figure 61 shows another example of the configuration of the encoding device 100 according to the embodiment. Specifically, Figure 61 shows the configuration of a submesh encoding device, which is a device that performs encoding processing when the input mesh 1101 is divided into subdivided meshes (multiple submeshes). For example, the submesh encoding device comprises multiple encoding devices 100.

[0531] The input mesh 1101 (full mesh) input to the submesh encoding device is divided into multiple meshes (submeshes). The multiple submeshes are input to, for example, multiple encoding devices 100. Each of the multiple submeshes may be input to any of the multiple encoding devices 100. For example, the submesh encoding device divides the input mesh 1101 into multiple submeshes and inputs the divided multiple submeshes to the multiple encoding devices 100.

[0532] Furthermore, after being divided into multiple submeshes, encoding processing (processing of overlapping submeshes) is performed on the boundaries of the submeshes.

[0533] For example, for each submesh, preprocessing is performed by the preprocessor 1103, generating and encoding the base mesh, displacement data, and metadata.

[0534] Furthermore, the encoding device 100 may be implemented with the configuration of such a submesh encoding device. In other words, the encoding device 100 may be configured to include multiple preprocessors 1103 and compressors 1106, and a predetermined process may be performed on the submesh for each set of multiple preprocessors 1103 and compressors 1106. Also, the number of sets of preprocessors 1103 and compressors 1106 included in the encoding device 100 is arbitrary and not particularly limited.

[0535] Figure 62 shows an example of the configuration of the pre-treatment unit 1103 according to the embodiment.

[0536] The preprocessing unit 1103 includes, for example, a base mesh generator 1401, a sub-divider 1402, and a displacement data generator 1403.

[0537] First, in the preprocessing unit 1103, a base mesh is generated by the base mesh generator 1401.

[0538] Next, the base mesh is subdivided by the sub-divider 1402 in a predetermined manner, and a subdivided mesh (Subdivided Mesh or Subdivided Base Mesh), which is the subdivided base mesh, is generated.

[0539] Displacement data is generated by the displacement data generator 1403 from the sub-divided mesh and the sub-mesh which is the input mesh 1101 after division.

[0540] The displacement data is, for example, the difference vector between the input mesh 1101 and the sub-divided mesh.

[0541] Furthermore, the same method used for sub-partitioning in encoding is employed in decoding.

[0542] Furthermore, for example, the encoding device 100 may transmit the sub-partitioning method in encoding and the parameters used in sub-partitioning to the decoding device 200.

[0543] Figure 63 is a diagram showing another example of the configuration of the decoding device 200 according to the embodiment. Specifically, Figure 63 is a diagram showing the configuration of a submesh decoding device, which is a device that performs decoding processing when a bitstream 2101 contains multiple submeshes. For example, the submesh decoding device comprises a plurality of decoding devices 200. For example, the submesh decoding device comprises a plurality of decoding devices 200 and a coupler 2109.

[0544] The encoded data for each submesh contained in the bitstream 2101 is input to the decompressor 2102 of each decoder 200.

[0545] Furthermore, for example, the post-processing unit 2106 executes the processing of the above-mentioned reconstructor.

[0546] In the post-processing unit 2106, for each decoded submesh, the base mesh is subdivided, a displacement vector is added to the subdivided base mesh, and the submesh is restored.

[0547] In other words, the post-processing unit 2106 performs the above-mentioned reconstruction process for each submesh.

[0548] The combiner 2109 merges the submeshes restored by each of the multiple decoding devices 200 to reconstruct the full mesh (output mesh 2107) before it was divided.

[0549] The decoding device 200 may be implemented with the configuration of such a submesh decoding device. In other words, the decoding device 200 may be configured to include multiple decompressors 2102 and post-processors 2106, and a predetermined process may be performed on the encoded data for each submesh for each set of multiple decompressors 2102 and post-processors 2106. Furthermore, the number of sets of decompressors 2102 and post-processors 2106 provided in the decoding device 200 is arbitrary and not particularly limited.

[0550] Figure 64 is a diagram showing a specific example of the configuration of the decoding device 200 according to the embodiment. Specifically, Figure 64 shows the specific configuration of the post-decoder 2306, which is one of the post-decoder 2306 and decoder 2305 that the decoding device 200 is equipped with.

[0551] The post-decoder 2306 comprises a pre-reconstructor 2307, a reconstructor 2308, a post-reconstructor 2309, and an adaptor 2310.

[0552] The processing in the post-decoder 2306 is optional depending on the application. An example of processing in the post-decoder 2306 (post-decoder processing) is the conversion of decoded data to a nominal format, such as video conversion from YUV space to RGB space. Post-decoder processing can be encapsulated into multiple processes such as pre-reconstruction, reconstruction, post-reconstruction, and adaptation.

[0553] The pre-reconstructor 2307 performs a pre-reconstruction process. The pre-reconstructor 2307 scales the normalized texture coordinates to match the dimensions of the texture image, for example, in the context of video-based dynamic mesh coding.

[0554] The reconstructor 2308 performs a reconstruction process. The reconstruction process is invoked, for example, on the decoded atlas frame, the decoded base mesh frame, the decoded video frame, and the syntax elements associated with the same mesh sequence. The output of the reconstruction process is a series of mesh frames reconstructed before the post-reconstruction process.

[0555] The post-reconstructor 2309 performs post-reconstruction processing. The post-reconstructor 2309 performs a number of smoothing operations on the reconstructed mesh frame, for example, in the context of video-based dynamic mesh coding. Smoothing operations include, for example, folding edges in the mesh or adding new vertices to the mesh.

[0556] The conformer 2310 performs a conformance process. The conformance process is applied by some application, for example, to fit the reconstructed mesh to a predetermined scenario. For example, the vertices of the reconstructed mesh are converted from the 3D model coordinate system to the 3D world coordinate system. The conformer 2310 outputs the reconstructed final mesh frame (final 3D mesh frame 2311).

[0557] <Encoding and Decoding of Attribute Information> Figure 65 is a conceptual diagram showing a specific example of a bitstream. In a bitstream, encoded data is encapsulated in a data unit structure. Specifically, a bitstream of encoded mesh data is also expressed as a DMC (Dynamic Mesh Coding) bitstream and consists of a series of DMC units. Each DMC unit is a data unit and is classified into a base mesh data unit, displacement data unit, metadata unit, or texture map unit, etc.

[0558] For example, encoded base mesh data is stored in a base mesh data unit with a header. Encoded displacement data is stored in a displacement data unit with a header. Encoded metadata is stored in a metadata unit with a header. Encoded texture maps are stored in a texture map unit with a header.

[0559] The header included in a unit is also called the unit header. The unit header stores the unit type, which indicates the type of data stored in the payload.

[0560] Metadata may correspond to a parameter set or SEI (Supplemental Enhancement Information). A parameter set may be a set of parameters common to a frame (data, access units, and samples at the same time, etc.), such as a frame parameter set, or a set of parameters common to a sequence, such as a sequence parameter set. Frame parameters may be expressed as a picture parameter set.

[0561] Metadata may correspond to the bitstream header. That is, the bitstream header may include the headers for each unit, or it may include the metadata unit.

[0562] In the decoding device 200, the bitstream is decoded using the data type and divided into multiple data units.

[0563] In the encoding and decoding of a three-dimensional mesh, surface information, which consists of vertex information and connectivity information, and attribute information for the surface are encoded and decoded, respectively. Examples of attribute information for a surface include color information and reflectance for the surface.

[0564] For example, three-dimensional data has attribute information (e.g., color information for the surface and reflectivity for the surface) for multiple faces of a single mesh. For example, as shown in Figure 3, the image of each face is mapped to a two-dimensional image. Not only the image of each face, but the attribute information of each face can also be mapped to a two-dimensional image. Therefore, for each attribute type, the attribute information is mapped to a two-dimensional image, and multiple two-dimensional images corresponding to multiple attribute types are obtained.

[0565] Furthermore, in order to signal attribute information corresponding to multiple attribute types, multiple two-dimensional images are signaled. However, the overhead of signaling multiple two-dimensional images may increase the amount of coding and processing required. Therefore, in this embodiment, multiple two-dimensional images are packed into units corresponding to a single image frame, such as a single image. This makes it possible to signal multiple two-dimensional images together.

[0566] Here, the two-dimensional image on which attribute information is mapped is an image, and can also be expressed as an attribute map, a two-dimensional attribute information set, or a two-dimensional attribute dataset, etc.

[0567] Furthermore, the mapping information showing the correspondence between multiple faces in a three-dimensional mesh and multiple faces in a two-dimensional image may be the same or different for multiple attribute types. In other words, the attribute information for each attribute type in a three-dimensional mesh may be mapped to a two-dimensional image based on the same correspondence between multiple attribute types, or it may be mapped to a two-dimensional image based on different correspondence between multiple attribute types.

[0568] Furthermore, the mapping information may be signaled from the encoding device 100 to the decoding device 200. In other words, the encoding device 100 may encode the mapping information, and the decoding device 200 may decode the mapping information. The mapping information may be signaled as attribute information or as vertex information.

[0569] Figure 66 is a block diagram showing an example of the configuration of an encoding device 100 that encodes multiple attribute maps. The encoding device 100 includes a preprocessor 1103 and a compressor 1106.

[0570] Multiple attribute maps 1102 are input to the encoding device 100. The preprocessor 1103 of the encoding device 100 stores the multiple attribute maps 1102 in the image 1110. Then, the compressor 1106 encodes the image 1110 into a bitstream 1107 using an image encoding method.

[0571] The encoding device 100 may receive an input mesh 1101. The preprocessor 1103 may obtain a base mesh 1104, displacement data 1105, and metadata 1108 from the input mesh 1101. The compressor 1106 may encode the base mesh 1104, displacement data 1105, and metadata 1108 into a bitstream 1107.

[0572] Alternatively, the preprocessor 1103 may acquire attribute information from the input mesh 1101 and acquire multiple attribute maps 1102 from the attribute information, without acquiring multiple attribute maps 1102 from an external source.

[0573] Furthermore, in the example shown in Figure 66, image 1110 may be considered as a single attribute map containing multiple attribute maps 1102.

[0574] Figure 67 is a block diagram showing an example of the configuration of a decoding device 200 that decodes multiple attribute maps. The decoding device 200 includes an expander 2102 and a post-processor 2106.

[0575] The decoding device 200 decodes the image 2110 from the bitstream 2101. Then, the post-processor 2106 reconstructs multiple attribute maps 2108 from the image 2110.

[0576] The decoding device 200 inputs the restored attribute maps 2108 along with information indicating the attribute type of each attribute map to the renderer 2111. The renderer 2111 is a rendering device that reconstructs a three-dimensional mesh by switching between the multiple attribute maps 2108 according to the application of the three-dimensional mesh.

[0577] The expander 2102 may decode the base mesh 2103, displacement data 2104, and metadata 2105 from the bitstream 2101. The post-processor 2106 may then obtain the output mesh 2107 from the base mesh 2103, displacement data 2104, and metadata 2105, and input the output mesh 2107 to the renderer 2111.

[0578] The renderer 2111 may then reconstruct a three-dimensional mesh based on the output mesh 2107 and a plurality of attribute maps 2108. The renderer 2111 may be included in the decoding device 200.

[0579] Furthermore, in the example shown in Figure 67, image 2110 may be considered as a single attribute map containing multiple attribute maps 2108.

[0580] In this embodiment, an example is shown of packing multiple attribute maps, each having different attribute types such as color and reflectance, into a single image. Packing methods, unpacking methods, encoding methods, decoding methods, and metadata related to this process are disclosed.

[0581] Figure 68 is an explanatory diagram showing an example of the syntax in this embodiment.

[0582] The syntax example shown in Figure 68 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.

[0583] In this specification, information shown in bold indicates information contained in the bitstream generated by the encoding device, while other information indicates information such as the control of calculations or processes when constructing the syntax.

[0584] The syntax shown in Figure 68 includes Sequence_parameter_set, which includes subdivisionNum.

[0585] In the Sequence_parameter_set() header, which is the sequence header, the encoder may call lifting_transform_parameters() to send default values ​​for the main lifting parameters. Similarly, in the Frame_parameter_set() header, which is the frame header, the encoder may call lifting_transform_parameters().

[0586] subdivisionNum indicates the number of subdivisions. Note that subdivisionNum may be added not only to the sequence header, but also to the frame header or any header indicating a smaller unit, updating the subdivisionNum set in the higher-level header. This allows the encoding device to switch the number of subdivisions at the frame level or even smaller. Alternatively, the LoD hierarchy can be calculated as = subdivisionNum + 1.

[0587] Figure 69 is an explanatory diagram showing an example of the syntax in this embodiment.

[0588] The syntax example shown in Figure 69 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.

[0589] The syntax shown in Figure 69 includes Meshpatch_data_unit, which includes lifting_main_param.

[0590] Meshpatch_data_unit can be header information that sends parameters to a unit smaller than a sequence or frame, or it can be a data unit. For example, it can be header information for a unit into which a mesh within a frame has been divided into multiple parts.

[0591] The `lifting_main_param` parameter, when set to 1, indicates that the encoding device will send the main parameters related to lifting, and when set to 0, indicates that the main parameters related to lifting will not be sent. This allows the encoding device to send lifting parameters in small units such as `Meshpatch_data_unit` by setting `lifting_main_param=1`, and to reduce the amount of code in `Meshpatch_data_unit` by setting `lifting_main_param=0` if it does not want to send lifting parameters.

[0592] Figure 70 is an explanatory diagram showing an example of syntax in this embodiment.

[0593] The syntax example shown in Figure 70 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.

[0594] The syntax shown in Figure 70 includes Lifting_transform_parameters. Lifting_transform_parameters includes lifting_adaptive_update_flag, lifting_update_weight_numerator[i], lifting_update_weight_denominator_minus1[i], lifting_adaptive_prediction_flag, lifting_prediction_weight_numerator[i], and lifting_prediction_weight_denominator_minus1[i].

[0595] The `lifting_adaptive_update_flag` flag, when set to 1, indicates that the encoding device will send weights for lifting updates, and when set to 0, indicates that no weights will be sent.

[0596] `lifting_update_weight_numerator` indicates the weight (numerator) for updating the lifting performance.

[0597] `lifting_update_weight_denominatior_minus1` indicates the weight (denominator) - 1 for the lifting update.

[0598] Using the above lifting_update_weight_numerator and lifting_update_weight_denominatior_minus1, the updateWeight for updates can be calculated as follows (Equation 59).

[0599] updateWeight = lifting_update_weight_numerator ÷ (lifting_update_weight_denominatior_minus1 + 1) (Formula 59)

[0600] If lifting_adaptive_update_flag is 0, the encoder may estimate a default value such as updateWeight = 0.5 or 1.0, or it may send a syntax indicating a different updateWeight. If lifting_adaptive_update_flag=0 and lifting_update_weight_numerator and lifting_update_weight_denominatior_minus1 are not included in the bitstream, the encoder may estimate lifting_update_weight_numerator to be 1 and lifting_update_weight_denominatior_minus1 to be 1, thereby calculating updateWeight = 0.5. In this way, the encoder can reduce the amount of code by not encoding lifting_update_weight_numerator and lifting_update_weight_denominatior_minus1 when updateWeight = 0.5 by setting lifting_adaptive_update_flag=0.

[0601] The `lifting_adaptive_prediction_flag` flag, when set to 1, indicates that the encoder will send weights for lifting prediction, and when set to 0, indicates that no weights will be sent.

[0602] The `lifting_prediction_weight_numerator` parameter indicates the weight (numerator) used for predicting lifting.

[0603] `lifting_prediction_weight_denominatior_minus1` indicates a weight (denominator) of -1 for the lifting prediction.

[0604] Using the above lifting_prediction_weight_numerator and lifting_prediction_weight_denominatior_minus1, the predWeight for prediction can be calculated as follows (Equation 60).

[0605] predWeight = lifting_prediction_weight_numerator ÷ (lifting_prediction_weight_denominatior_minus1 + 1) (Formula 60)

[0606] If lifting_adaptive_prediction_flag is 0, a default value such as predWeight = 0.5 or 1.0 may be estimated, or the encoding device may send a syntax indicating a different predWeight. Furthermore, if lifting_adaptive_prediction_flag=0 and lifting_prediction_weight_numerator and lifting_prediction_weight_denominatior_minus1 are not included in the bitstream, lifting_prediction_weight_numerator may be estimated to be 1, and lifting_prediction_weight_denominatior_minus1 may also be estimated to be 1, resulting in the calculation of predWeight = 0.5. This allows the encoding device to reduce the amount of code by setting `lifting_adaptive_prediction_flag=0` when `predWeight = 0.5`, thereby not encoding `lifting_prediction_weight_numerator` and `lifting_prediction_weight_denominatior_minus1`.

[0607] Alternatively, you may include updateWeight or predWeight directly in the syntax instead of *****_weight_numerator or *****_weight_denominatior_minus1. In the above, "*****_weight_numerator" may mean "lifting_update_weight_numerator" or "lifting_prediction_weight_numerator". Also, "*****_weight_denominatior_minus1" may mean "lifting_update_weight_denominator_minus1" or "lifting_prediction_weight_denominator_minus1".

[0608] Alternatively, a syntax such as lifting_update_weight_log2 or lifting_prediction_weight_log2 may be used, in which case the encoding device transmits a value corresponding to the exponent of a power of 2. In this case, updateWeight and predictionWeight may be as shown in (Equation 61) below.

[0609] updateWeight = 1 ÷ (1 << lifting_update_weight_log2) predictionWeight = 1 ÷ (1 << lifting_prediction_weight_log2) (Equation 61)

[0610] Note that in Figure 70, the syntax for update and the syntax for prediction may be in any order.

[0611] Figure 71 is an explanatory diagram showing an example of syntax in this embodiment.

[0612] The syntax example shown in Figure 71 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.

[0613] The syntax shown in Figure 71 includes Sequence_parameter_set. Sequence_parameter_set includes sps_subdivisionNum, sps_transform_method, and sps_lifting_offset_present_flag.

[0614] sps_subdivisionNum indicates the number of subdivisions (subdivisionNum) common to the entire sequence. Note that subdivisionNum may be added not only to the sequence header, but also to the frame header or a header indicating a smaller unit (such as Meshpath_data_unit), updating the subdivisionNum set in the higher-level header. This allows switching the number of subdivisions at the frame level or smaller units. Alternatively, the LoD hierarchy can be calculated as = subdivisionNum + 1. Note that subdivision or the subdivision process is also referred to as the division process.

[0615] `sps_transform_method` is information indicating the transformation method used when encoding or decoding the displacement vector. For example, a value of 0 indicates no transformation, a value of 1 indicates a lifting transformation, and values ​​of 2 or higher may be reserved for extension. Here, if `sps_transform_method` is set to a value of 1, it may indicate that the encoding device has encoded the displacement vector by applying the lifting transformation as described above. This allows the decoding device, when `sps_transform_method=1`, to appropriately decode a bitstream whose encoding efficiency has been improved by predictive coding using lifting transformation with an LoD hierarchy by applying an inverse lifting transformation. Note that if `sps_transform_method` is set to a value of 0, it may indicate that the encoding device has encoded the displacement vector without applying any transformation processing such as a lifting transformation. Specifically, it may indicate that the input displacement vector value was encoded by directly applying quantization or arithmetic coding without applying any transformation processing. This allows the decryption device to properly decrypt a bitstream with reduced processing load by not applying the transformation process when sps_transform_method=0.

[0616] In this embodiment, the transformation method for sps_transform_method is at least selectable as either no transformation or a lifting transformation, but it is not necessarily limited to these, and any transformation process may be applied. For example, the code size may be reduced by applying a frequency transformation such as a discrete cosine transform.

[0617] Note that the lifting transform can be applied to predictive coding when the LoD hierarchy is 2 or higher, but cannot be predicted when the LoD hierarchy is 1. When the LoD hierarchy is 1, it has the same effect as not making a prediction. Therefore, as shown in the syntax of Figure 71, sps_transform_method may be added to the bitstream when sps_subdivisionNum > 0, that is, when the LoD hierarchy is 2 or higher, and sps_transform_method may not be added to the bitstream when sps_subdivisionNum = 0, that is, when the LoD hierarchy is 1. Alternatively, the decoder may estimate the value of sps_transform_method to be 0 (no prediction) if it is not added to the bitstream. In this way, when sps_subdivisionNum = 0, that is, when there is one LoD hierarchy, the encoder does not add sps_transform_method to the bitstream, and the decoder estimates its value to be 0, thereby reducing the amount of code in the header.

[0618] As shown in Figure 71, if the value of sps_transform_method indicates a lifting transform, i.e., if the value is 1, the encoding device may add parameters used for the lifting transform, such as sps_lifting_offset_present_flag or Lifting_transform_parameters(), to the bitstream. If the value of sps_transform_method does not indicate a lifting transform, the encoding device may not add parameters used for the lifting transform to the bitstream. This allows the encoding device to reduce the amount of code in the header when the lifting transform is not used, for example, when no transformation is performed.

[0619] The sps_lifting_offset_present_flag indicates whether the offset value to be added after the lifting transform is included in the bitstream. For example, a value of 1 may indicate that the offset for the lifting transform is added to the bitstream, while a value of 0 may indicate that the offset for the lifting transform is not added to the bitstream. If sps_lifting_offset_present_flag is not added to the bitstream, the decoder may estimate its value to be 0. As a result, if sps_transform_method does not indicate a lifting transform, the value of sps_lifting_offset_present_flag is estimated to be 0, and the code size can be reduced by not adding lifting_offset_parameters to the bitstream.

[0620] Figure 72 is an explanatory diagram showing an example of syntax in this embodiment.

[0621] The syntax example shown in Figure 72 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.

[0622] The syntax shown in Figure 72 includes Lifting_transform_parameters().

[0623] The syntax described above can also be used in Lifting_transform_parameters().

[0624] Figure 73 is an explanatory diagram showing an example of syntax in this embodiment.

[0625] The syntax example shown in Figure 73 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.

[0626] The syntax shown in Figure 73 includes Frame_parameter_set. Frame_parameter_set includes fps_transform_method_overridden_flag, fps_subdivisionNum, and fps_transform_method.

[0627] The `fps_transform_method_overridden_flag` flag indicates whether the `transform_method` should be updated at the frame level. For example, a value of 1 may indicate that the `transform_method` should be updated at the frame level, while a value of 0 may indicate that the `transform_method` should not be updated at the frame level. This allows an encoding device to reduce the amount of code by setting `fps_transform_method_overridden_flag=0` if it does not intend to update the transform method at the frame level, thus eliminating the need to add frame-level transform_method information to the bitstream.

[0628] fps_subdivisionNum indicates the number of subdivisions at the frame level (subdivisionNum). Note that subdivisionNum may be appended not only to the frame header but also to headers indicating smaller units (such as Meshpath_data_unit), updating the subdivisionNum set in the higher-level header. This allows the encoding device to switch the number of subdivisions at a unit smaller than the frame level. Alternatively, the LoD hierarchy can be calculated as = subdivisionNum + 1.

[0629] `fps_transform_method` is information indicating the transformation method used when encoding or decoding the displacement vector. For example, a value of 0 indicates no transformation, a value of 1 indicates a lifting transformation, and values ​​of 2 or higher may be reserved for extension. Here, if `fps_transform_method` is set to a value of 1, it may indicate that the encoding device has encoded the displacement vector by applying the lifting transformation as described above. This allows the decoding device, when `fps_transform_method=1`, to appropriately decode a bitstream whose encoding efficiency has been improved by predictive coding using a lifting transformation with an LoD hierarchy by applying an inverse lifting transformation. Note that if `fps_transform_method` is set to a value of 0, it may indicate that the encoding device has encoded the displacement vector without applying any transformation processing such as a lifting transformation. Specifically, it may indicate that the input displacement vector value was encoded by directly applying quantization or arithmetic coding without applying any transformation processing. This allows the decoding device to properly decode a bitstream with reduced processing load by not applying the transformation process when fps_transform_method=0.

[0630] In this embodiment, the fps_transform_method allows selection of either no transformation or a lifting transformation, but it is not limited to these options, and any transformation process may be applied. For example, the code size may be reduced by applying a frequency transformation such as a discrete cosine transform.

[0631] In lifting transformations, the encoding device can apply predictive coding when the LoD hierarchy is 2 or higher, but cannot predict when the LoD hierarchy is 1. When the LoD hierarchy is 1, the effect is the same as not making a prediction. Therefore, as shown in the syntax of Figure 73, fps_transform_method may be added to the bitstream when fps_transform_method_overridden_flag=1 and fps_subdivisionNum>0, that is, when the LoD hierarchy is 2 or higher. Alternatively, when fps_transform_method_overridden_flag=0 or fps_subdivisionNum=0, that is, when the LoD hierarchy is 1, fps_transform_method may not be added to the bitstream. Furthermore, if fps_transform_method is not added to the bitstream, the decoding device may estimate fps_transform_method according to the flow chart shown in Figure 74.

[0632] Figure 74 is a flowchart showing the estimation process of fps_transform_method in this embodiment. As shown in Figure 74, the decoder determines whether fps_subdivisionNum is 0 or not (step S13501). If it is determined that fps_subdivisionNum is 0 (Yes in step S13501), the decoder estimates that fps_transform_method is 0 (i.e., no prediction) (step S13502). On the other hand, if it is determined that fps_subdivisionNum is not 0 (No in step S13501), the decoder estimates that fps_transform_method is sps_transform_method (step S13511).

[0633] This allows for a reduction in the amount of code in the header by having the decoder estimate the fps_transform_method instead of the encoder adding it to the bitstream when fps_transform_method_overridden_flag=0 or fps_subdivisionNum=0, i.e., when the LoD hierarchy is 1.

[0634] As shown in Figure 73, if the value of fps_transform_method indicates a lifting transformation, i.e., if the value is 1, parameters used for the lifting transformation, such as Lifting_transform_parameters(), may be added to the bitstream. If the value of fps_transform_method does not indicate a lifting transformation, parameters used for the lifting transformation may not be added to the bitstream. This reduces the amount of code in the header when the lifting transformation is not used, for example, when no transformation is performed.

[0635] The syntax described above can also be used in Lifting_transform_parameters().

[0636] Figure 75 is an explanatory diagram showing an example of the syntax in this embodiment.

[0637] The syntax example shown in Figure 75 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.

[0638] The syntax shown in Figure 75 includes a Meshpatch_data_unit. The Meshpatch_data_unit includes mdu_transform_method_overridden_flag, mdu_subdivisionNum, mdu_transform_method, mdu_lifting_offset_present_flag, and lifting_offset_parameters.

[0639] The `mdu_transform_method_overridden_flag` flag indicates whether to update the `transform_method` at the `Meshpatch_data_unit` level. A `Meshpatch_data_unit` is a data unit that corresponds to, for example, a submesh. For example, a value of 1 may indicate that the `transform_method` is updated at the `Meshpatch_data_unit` level, while a value of 0 may indicate that the `transform_method` is not updated at the `Meshpatch_data_unit` level. This means that, for example, if the transform method is not updated at the `Meshpatch_data_unit` level, the encoding device can reduce the amount of code by setting `mdu_transform_method_overridden_flag=0`, eliminating the need to add `Meshpatch_data_unit` level transform_method information to the bitstream.

[0640] mdu_subdivisionNum indicates the number of subdivisions at the Meshpatch_data_unit level (subdivisionNum). Alternatively, the LoD hierarchy can be calculated as subdivisionNum + 1.

[0641] `mdu_transform_method` is information indicating the transformation method used when encoding or decoding the displacement vector. For example, a value of 0 indicates no transformation, a value of 1 indicates a lifting transformation, and values ​​of 2 or higher may be reserved for extension. Here, if `mdu_transform_method` is set to a value of 1, it may indicate that the encoding device has encoded the displacement vector by applying the lifting transformation as described above. This allows the decoding device, when `mdu_transform_method=1`, to appropriately decode a bitstream whose encoding efficiency has been improved by predictive coding using a lifting transformation with an LoD hierarchy by applying an inverse lifting transformation. Note that if `mdu_transform_method` is set to a value of 0, it may indicate that the encoding device has encoded the displacement vector without applying any transformation processing such as a lifting transformation. Specifically, it may indicate that the input displacement vector value was encoded by directly applying quantization or arithmetic coding without applying any transformation processing. This allows the decoding device to properly decode a bitstream with reduced processing load by not applying the transformation process when mdu_transform_method=0.

[0642] In this embodiment, the mdu_transform_method allows selection of at least two methods: no transformation or a lifting transformation. However, it is not limited to these, and any transformation process may be applied. For example, the code size may be reduced by applying a frequency transformation such as a discrete cosine transform.

[0643] Note that the lifting transform can be applied to predictive coding when the LoD hierarchy is 2 or higher, but it cannot be predicted when the LoD hierarchy is 1. When the LoD hierarchy is 1, the effect is the same as not making a prediction. Therefore, as shown in the syntax of Figure 75, mdu_transform_method may be added to the bitstream when mdu_transform_method_overridden_flag=1 and mdu_subdivisionNum>0, that is, when the LoD hierarchy is 2 or higher, and mdu_transform_method may not be added to the bitstream when mdu_transform_method_overridden_flag=0 or mdu_subdivisionNum=0, that is, when the LoD hierarchy is 1. Alternatively, if mdu_transform_method is not added to the bitstream, the decoder may estimate mdu_transform_method according to the flowchart shown in Figure 76.

[0644] Figure 76 is a flowchart showing the estimation process for mdu_transform_method in this embodiment. As shown in Figure 76, the decoder determines whether mdu_subdivisionNum is 0 or not (step S13701). If it is determined that mdu_subdivisionNum is 0 (Yes in step S13701), the decoder estimates that mdu_transform_method is 0 (i.e., no prediction) (step S13702). On the other hand, if it is determined that mdu_subdivisionNum is not 0 (No in step S13701), the decoder estimates that mdu_transform_method is fps_transform_method (step S13711).

[0645] This allows for a reduction in the amount of code in the header when mdu_transform_method_overridden_flag=0 or mdu_subdivisionNum=0, i.e., when the LoD hierarchy is 1, the encoder does not add mdu_transform_method to the bitstream, and the decoder estimates mdu_transform_method.

[0646] As shown in Figure 75, if the value of mdu_transform_method indicates a lifting transformation, i.e., if the value is 1, parameters used for the lifting transformation, such as mdu_lifting_offset_present_flag, lifting_offset_parameters, or Lifting_transform_parameters(), may be added to the bitstream. If the value of mdu_transform_method does not indicate a lifting transformation, the parameters used for the lifting transformation may not be added to the bitstream. This reduces the amount of code in the header when the lifting transformation is not used, for example, when no transformation is performed.

[0647] `mdu_lifting_offset_present_flag` indicates whether the offset value to be added after the lifting transformation is included in the bitstream. For example, a value of 1 may indicate that the offset for the lifting transformation is included in the bitstream, while a value of 0 may indicate that the offset for the lifting transformation is not included in the bitstream. If `mdu_lifting_offset_present_flag` is not included in the bitstream, i.e., `sps_lifting_offset_present_flag=1`, the decoder may estimate `mdu_lifting_offset_present_flag` to be 1. This allows `lifting_offset_parameters` to be added to the bitstream without adding `mdu_lifting_offset_present_flag`, thereby reducing the amount of code.

[0648] `lifting_offset_parameters` contains information about the offset value to be added after the lifting transformation. Note that if `lifting_offset_parameters` is not attached to the bitstream, the decoder may estimate `lifting_offset_parameters` to be 0.

[0649] The syntax described above can also be used in Lifting_transform_parameters().

[0650] Figure 77 is an explanatory diagram showing an example of the syntax in this embodiment.

[0651] The syntax example shown in Figure 77 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.

[0652] The syntax shown in Figure 77 includes Sequence_parameter_set(SPS). Sequence_parameter_set(SPS) includes sps_subdivisionNum and sps_qpparameters_present_flag.

[0653] sps_subdivisionNum indicates the number of subdivisions common to the entire sequence. Note that subdivisionNum may be added not only to the sequence header (Sequence Parameter Set (SPS)), but also to the frame header (Frame Parameter Set (FPS)) or a header indicating a smaller unit (such as Meshpath_data_unit (MDU)), and the subdivisionNum set in the higher-level header may be updated. This allows switching the number of subdivisions at the frame level or smaller units. Alternatively, the LoD hierarchy can be calculated as = subdivisionNum + 1.

[0654] The sps_qpparameters_present_flag indicates whether quantization parameter information common to the entire sequence is attached to the bitstream. For example, a value of 1 may indicate that quantization parameter information common to the entire sequence is attached to the SPS, while a value of 0 may indicate that quantization parameter information common to the entire sequence is not attached to the SPS. Note that if sps_qpparameters_present_flag=0, it is not necessary to attach the quantization parameter information Qp_parameters to the SPS. This can reduce the code size of the SPS. Note that if sps_qpparameters_present_flag=1, quantization parameter information Qp_parameters(sps_subdivisionNum) corresponding to the value of sps_subdivisionNum may be attached to the SPS. This allows the decoding device to be informed of the quantization parameter information that has been appropriately set according to the value of sps_subdivisionNum.

[0655] Figure 78 is an explanatory diagram showing an example of the syntax in this embodiment.

[0656] The syntax example shown in Figure 78 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.

[0657] The syntax shown in Figure 78 includes Frame_parameter_set(FPS). Frame_parameter_set(FPS) includes fps_subdivisionNum and fps_qpparameters_present_flag.

[0658] fps_subdivisionNum indicates the number of subdivisions common to the entire frame. Note that subdivisionNum can be added not only to the frame header (frame parameter set (FPS)) but also to headers indicating smaller units (such as Meshpath_data_unit), and the subdivisionNum set in the higher-level header can be updated. This allows switching the number of subdivisions at a unit smaller than the frame level. Alternatively, the LoD hierarchy can be calculated as subdivisionNum + 1.

[0659] The `fps_qpparameters_present_flag` indicates whether common quantization parameter information is attached to the bitstream across the entire frame. For example, a value of 1 may indicate that common quantization parameter information is attached to the FPS across the entire frame, while a value of 0 may indicate that common quantization parameter information is not attached to the FPS across the entire frame. If no quantization parameter information exists for the entire sequence, i.e., `fps_qpparameters_present_flag=0`, the encoding device may not attach `fps_qpparameters_present_flag` to the FPS, and the decoding device may estimate `fps_qpparameters_present_flag` to be 0. This can reduce the code size of the FPS. Furthermore, if `fps_qpparameters_present_flag=0`, the quantization parameter information `Qp_parameters` does not need to be attached to the FPS. This can also reduce the code size of the FPS. Furthermore, if fps_qpparameters_present_flag=1, quantization parameter information Qp_parameters(fps_subdivisionNum) corresponding to the value of fps_subdivisionNum may be added to the FPS. This allows the encoding device to inform the decoding device of the quantization parameter information that has been appropriately set according to the value of fps_subdivisionNum.

[0660] Figure 79 is an explanatory diagram showing an example of the syntax in this embodiment.

[0661] The syntax example shown in Figure 79 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.

[0662] The syntax shown in Figure 79 includes a Meshpatch_data_unit (MDU). The Meshpatch_data_unit (MDU) includes mdu_subdivisionNum and mdu_qpparameters_present_flag.

[0663] mdu_subdivisionNum indicates the number of common subdivisions (subdivisionNum) in a Mesh Patch Data Unit (MDU). Note that subdivisionNum may be added to a header representing a smaller unit than MDU, and the subdivisionNum set in the higher-level header may be updated. This allows switching the number of subdivisions in units smaller than MDU. Alternatively, the LoD hierarchy can be calculated as = subdivisionNum + 1.

[0664] `mdu_qpparameters_present_flag` indicates whether common quantization parameter information for the MDU is attached to the bitstream. For example, a value of 1 may indicate that common quantization parameter information for the MDU is attached to the MDU, while a value of 0 may indicate that common quantization parameter information for the MDU is not attached to the MDU. If no quantization parameter information exists for the entire sequence, i.e., `sps_qpparameters_present_flag=0`, the encoding device may not attach `mdu_qpparameters_present_flag` to the MDU, and the decoding device may estimate `mdu_qpparameters_present_flag` to be 0. This can reduce the code size of the MDU. Furthermore, if `mdu_qpparameters_present_flag=0`, it is not necessary to attach the quantization parameter information `Qp_parameters` to the MDU. This can also reduce the code size of the MDU. Furthermore, if mdu_qpparameters_present_flag=1, quantization parameter information Qp_parameters(mdu_subdivisionNum) corresponding to the value of mdu_subdivisionNum may be added to the MDU. This allows the decoding device to be informed of the quantization parameter information that has been appropriately set according to the value of mdu_subdivisionNum.

[0665] Figure 80 is an explanatory diagram showing an example of syntax in this embodiment.

[0666] The syntax example shown in Figure 80 illustrates an example of the structure of information contained in a bitstream generated by an encoding device.

[0667] The syntax shown in Figure 80 is such that Qp_parameters(subdivisionNum) includes lod_adaptive_quantization_flag, qp_params, and lod_qp_params[i].

[0668] The `lod_adaptive_quantization_flag` indicates whether the encoding device adds quantization parameter information to the bitstream for each LoD hierarchy and performs adaptive quantization using the quantization parameter information added for each LoD hierarchy. For example, a value of 0 may indicate that the encoding device adds common quantization parameter information `qp_params` to the bitstream for each LoD hierarchy and performs quantization using this common information without adaptively switching the quantization parameter information for each LoD hierarchy. This reduces the amount of code by not adding hierarchy-specific quantization parameter information to the bitstream. Alternatively, a value of 1 may indicate that quantization parameter information `lod_qp_params` for each LoD hierarchy is added to the bitstream for each LoD hierarchy, and quantization is performed by adaptively switching the quantization parameter information for each LoD hierarchy. This improves encoding efficiency by setting appropriate quantization parameter information for each hierarchy. For example, when an encoding device uses a lifting transform, higher levels of the Line of Deposition (LoD) tend to contain more important components. Therefore, encoding efficiency can be improved by setting the quantization parameter information to reduce the quantization values ​​of higher levels.

[0669] Furthermore, if subdivisionNum=0, i.e., the number of LoD layers=1, the encoding device does not need to add quantization parameter information for each layer, and therefore does not need to add lod_adaptive_quantization_flag to the bitstream. This reduces the amount of code. Also, if subdivisionNum=0, the decoding device may assume lod_adaptive_quantization_flag=0. This allows the decoding device to properly decode a bitstream whose code size has been reduced because lod_adaptive_quantization_flag is not added to the bitstream when subdivisionNum=0.

[0670] qp_params is quantization parameter information common to the LoD hierarchy. The decoding device may also set the decoded qp_params value as the quantization parameter information for each LoD hierarchy. That is, the quantization parameter information for LoD hierarchy i, lod_qp_params[i]=qp_params, may be set. This allows quantization processing to be performed using the same quantization parameter information at each hierarchy.

[0671] lod_qp_params[i] is the quantization parameter information for LoD hierarchy i. Note that the quantization parameter information may include any information related to quantization, such as quantization parameters or quantization scales.

[0672] Figure 81 shows the sub-partitioning related syntax in the sequence parameter set according to this embodiment.

[0673] The encoder and decoder indicate sps_subdivisionNum as information indicating the number of subdivisions common to the entire sequence. If sps_subdivisionNum is greater than 0, the encoder and decoder then indicate Subdivision_method as information indicating the subdivision method. If sps_subdivisionNum is 0, no subdivision is performed, and therefore Subdivision_method is not added to the bitstream. The decoder assumes no subdivision if sps_subdivisionNum is 0.

[0674] Figure 82 shows the sub-partitioning related syntax in the frame parameter set according to this embodiment.

[0675] The encoding and decoding devices indicate fps_subdivisionNum as information that shows the number of subdivisions common to the entire frame. If fps_subdivisionNum is greater than 0, the encoding and decoding devices then indicate Subdivision_method as information that shows the subdivision method. fps_subdivisionNum may also be used to update the number of subdivisions set in the higher-level header. This allows for switching the number of subdivisions for units smaller than a frame.

[0676] Figure 83 shows the sub-division related syntax in the mesh patch data unit according to this embodiment.

[0677] The encoding and decoding devices indicate mdu_subdivisionNum as information indicating the common number of subdivisions within the mesh patch. If mdu_subdivisionNum is greater than 0, the encoding and decoding devices then indicate Subdivision_method as information indicating the subdivision method. mdu_subdivisionNum may also be used to update the number of subdivisions set in the higher-level header. This allows for switching the number of subdivisions for units smaller than the mesh patch.

[0678] Figure 84 shows the syntax of Subdivision_method according to this embodiment.

[0679] The encoding and decoding devices switch the signaling pattern for the subdivision method based on the argument subdivisionNum. If subdivisionNum is greater than 1, the decoding device obtains lod_adaptive_subdivision_flag as information indicating whether or not to switch the subdivision method at each level. If lod_adaptive_subdivision_flag is 0, the decoding device obtains subdivision_method, which indicates a subdivision method common to all levels, and sets this value to lod_subdivision_method[i], the subdivision method for each level, and uses it. If lod_adaptive_subdivision_flag is not 0, the decoding device obtains lod_subdivision_method[i] sequentially according to the number of levels and uses it in association with each level. If subdivisionNum is 0, no subdivision is performed, and subdivision_method or lod_subdivision_method[i] can be presumed to be a value indicating no subdivision.

[0680] In the above embodiment, subdivision_method is used to identify multiple methods such as midpoint subdivision, interpolation subdivision, normal vector subdivision, and no subdivision. The encoding device may compare the results encoded by any of the methods and select the most efficient method to add to the bitstream. The decoding device appropriately decodes using the method added to the bitstream. The number of subdivisions may be treated as the number of layers plus 1.

[0681] (Embodiment 2) Figure 85 is a flowchart showing the procedure for processing the three-dimensional mesh of the hierarchy to be encoded in the encoding device.

[0682] The encoding device encodes a second flag into the bitstream indicating whether or not to overwrite the number of divisions for the three-dimensional mesh of the hierarchy to be encoded (S2001).

[0683] The encoding device determines whether the second flag indicates an overwrite. If it indicates an overwrite, it proceeds to S2003; otherwise, it proceeds to S2004 (S2002).

[0684] If the second flag indicates overwriting (Yes in S2002), the encoding device encodes the second parameter indicating the number of divisions into a bitstream (S2003).

[0685] The encoding device may encode a third set of parameters indicating the division method into a bitstream, but only if the number of divisions is not zero. This modification allows for a reduction in the amount of code by omitting the signaling of the division method.

[0686] The encoding device may encode a fourth set of parameters indicating the quantization parameters into a bitstream, but only if the number of divisions is not zero. This modification allows for a reduction in the amount of code by omitting the signaling of the quantization parameters.

[0687] The encoding device may encode the fourth set of parameters into a bitstream only if quantization is present. In this modification, the presence of quantization is indicated by encoding a flag indicating its presence into a bitstream, and the amount of code can be reduced by omitting the signaling of the quantization parameters.

[0688] The encoding device may encode a fifth set of parameters indicating the conversion parameters into a bitstream, but only if the number of divisions is not zero. This modification allows for a reduction in the amount of code by omitting the signaling of the conversion parameters.

[0689] The encoding device may encode the fifth set of parameters into a bitstream only if a conversion process exists (other than NONE). According to this modification, the presence of a conversion process is indicated by encoding a flag indicating its presence into a bitstream, and the amount of encoding can be reduced by omitting the signaling of the conversion parameters.

[0690] If the second flag does not indicate an overwrite (No in S2002), the encoding device derives the number of divisions from the higher layer and encodes a third flag indicating whether or not to encode the division method to be applied to the target layer, a fourth flag indicating whether or not to encode the quantization parameters, and a seventh flag indicating whether or not to encode the transformation parameters into a bitstream (S2004). Then, the encoding device derives the number of divisions to be applied to the target layer.

[0691] The derivation may also be done by copying the value of the number of divisions in the higher layer. The encoding device may encode a third flag indicating the presence or absence of a division method into the bitstream only if the number of divisions is not zero. This modification allows for a reduction in the amount of code by omitting the signaling of the division method.

[0692] The encoding device may infer the third flag to be 0 if the number of divisions is 0.

[0693] The encoding device may encode a fourth flag in the bitstream indicating the presence or absence of quantization parameters, but only if quantization processing is present. According to this modification, the presence of quantization processing is indicated by encoding a flag indicating its presence, and the amount of code can be reduced by omitting the signaling of quantization parameters.

[0694] The encoding device may encode the fourth flag only if the number of divisions is not zero. According to this modification, if the number of divisions is zero, the fourth flag is presumed to be zero, and the amount of code can be reduced by omitting unnecessary signaling.

[0695] The encoding device may encode a seventh flag in the bitstream indicating the presence or absence of conversion parameters, but only if a conversion process exists (other than NONE). According to this modification, the presence of a conversion process is indicated by encoding a flag, and the amount of code can be reduced by omitting the signaling of conversion parameters.

[0696] The encoding device may encode the seventh flag only if the number of divisions is not zero. According to this modification, if the number of divisions is zero, the seventh flag is presumed to be zero, and the amount of code can be reduced by omitting unnecessary signaling.

[0697] The encoding device encodes displacement information for reconstructing the three-dimensional mesh of the hierarchy to be encoded using at least the above number of divisions (S2005).

[0698] The layer to be encoded may mean a mesh sequence, mesh frame, submesh, group of faces, or Line of Deposition (LD). Here, a mesh frame is a lower layer of a mesh sequence, and a submesh is a lower layer of a mesh frame. "Other layers" may mean default values ​​for parameters or groups of parameters, mesh sequences, mesh frames, submesh, group of faces, or LD.

[0699] Figure 86 is a flowchart showing the procedure for processing the three-dimensional mesh of the hierarchy to be decoded in the decoding device.

[0700] The decoding device decodes a second flag from the bitstream that indicates whether or not to overwrite the number of divisions for the three-dimensional mesh of the hierarchy to be decoded (S2011).

[0701] The decoding device determines whether the second flag indicates an overwrite. If it indicates an overwrite, it proceeds to S2013; otherwise, it proceeds to S2014 (S2012).

[0702] If the second flag indicates overwriting (Yes in S2012), the decoding device decodes the second parameter, which indicates the number of divisions, from the bitstream (S2013).

[0703] The decoding device may decode a third set of parameters indicating the division method only if the number of divisions is not zero. According to this modification, the amount of code can be reduced by omitting the signaling of the division method.

[0704] The decoding device may decode the fourth set of parameters representing the quantization parameters only if the number of divisions is not zero. According to this modification, the code size can be reduced by omitting the signaling of the quantization parameters.

[0705] The decoding device may decode the fourth parameter group only if quantization is present. In this modification, the presence of quantization is indicated by decoding a flag from the bitstream, and the code size can be reduced by omitting the signaling of quantization parameters.

[0706] The decoding device may decode the fifth set of parameters indicating the transformation parameters only if the number of divisions is not zero. According to this modification, the code size can be reduced by omitting the signaling of the transformation parameters.

[0707] The decoding device may decode the fifth set of parameters only if a conversion process exists (other than NONE). According to this modification, the presence of a conversion process is indicated by decoding a flag from the bitstream, and the code size can be reduced by omitting the signaling of the conversion parameters.

[0708] If the second flag does not indicate an overwrite (No in S2012), the decoding device derives the number of divisions from the higher layer, and decodes from the bitstream a third flag indicating whether or not to decode the division method to be applied to the target layer, a fourth flag indicating whether or not to decode the quantization parameters, and a seventh flag indicating whether or not to decode the transformation parameters (S2014). Then, the decoding device derives the number of divisions to be applied to the target layer. The derivation may be done by copying the value of the number of divisions in the higher layer.

[0709] The decoding device may decode a third flag indicating the presence or absence of a division method, but only if the number of divisions is not zero. This modification allows for a reduction in the amount of code by omitting the signaling of the division method.

[0710] The decoding device may infer the third flag to be 0 if the number of divisions is 0.

[0711] The decoding device may decode a fourth flag indicating the presence or absence of quantization parameters, but only if quantization processing is present. According to this modification, the presence of quantization processing is indicated by decoding a flag indicating its presence, and the code size can be reduced by omitting the signaling of quantization parameters.

[0712] The decoding device may decode the fourth flag only if the number of divisions is not zero. According to this modification, if the number of divisions is zero, the fourth flag is presumed to be zero, and the amount of coding can be reduced by omitting unnecessary signaling.

[0713] The decoding device may decode a seventh flag indicating the presence or absence of conversion parameters, but only if a conversion process exists (other than NONE). According to this modification, the presence of a conversion process is indicated by decoding a flag indicating its presence, and the code size can be reduced by omitting the signaling of conversion parameters.

[0714] The decoding device may decode the seventh flag only if the number of divisions is not zero. According to this modification, if the number of divisions is zero, the seventh flag is presumed to be zero, and the amount of coding can be reduced by omitting unnecessary signaling.

[0715] The decoding device decodes displacement information for reconstructing the three-dimensional mesh of the hierarchy to be decoded using at least the above number of divisions (S2015).

[0716] The layer to be decoded may mean a mesh sequence, mesh frame, submesh, group of faces, or Line of Deposition (LD). Here, a mesh frame is a lower layer of a mesh sequence, and a submesh is a lower layer of a mesh frame. "Other layers" may mean default values ​​for parameters or groups of parameters, mesh sequences, mesh frames, submesh, group of faces, or LD.

[0717] Figure 87 is a flowchart showing the coding procedure for the number of divisions, division method, quantization parameters, and transformation parameters in an encoding device.

[0718] The encoding device first encodes subdivision_iterationCount_override_flag, which indicates whether or not to overwrite the number of divisions (S2021).

[0719] The encoding device determines whether subdivision_iterationCount_override_flag is TRUE or not (S2022).

[0720] If subdivision_iterationCount_override_flag is TRUE (Yes in S2022), the encoding device encodes the subdivision_iterationCount (S2023).

[0721] The encoding device encodes the subdivision_methods after encoding the subdivision_iterationCount (S2024).

[0722] The encoding device encodes quantization_params after encoding subdivision_methods (S2025).

[0723] The encoding device encodes transform_params after encoding quantization_params (S2026).

[0724] On the other hand, if subdivision_iterationCount_override_flag is FALSE (No in S2022), the encoding device estimates subdivision_iterationCount from higher layers, etc. (S2027).

[0725] The encoding device encodes subdivision_method_override_flag (S2028).

[0726] The encoding device determines whether subdivision_method_override_flag is TRUE or not (S2029).

[0727] The encoding device encodes subdivision_methods (S2030) if subdivision_method_override_flag is TRUE (Yes in S2029).

[0728] If subdivision_method_override_flag is FALSE (No in S2029), the encoding device estimates subdivision_methods (S2031).

[0729] The encoding device encodes quantization_override_flag (S2032).

[0730] The encoding device determines whether quantization_override_flag is TRUE or not (S2033).

[0731] The encoding device encodes quantization_params (S2034) if quantization_override_flag is TRUE (Yes in S2033).

[0732] The encoding device estimates quantization_params (S2035) if quantization_override_flag is FALSE (No in S2033).

[0733] The encoding device encodes transform_params_override_flag (S2036).

[0734] The encoding device determines whether transform_params_override_flag is TRUE or not (S2037).

[0735] The encoding device encodes transform_params (S2038) if transform_params_override_flag is TRUE (Yes in S2037).

[0736] The encoding device estimates transform_params (S2039) if transform_params_override_flag is FALSE (No in S2037).

[0737] Thus, by performing the processing according to the procedure shown in Figure 87, the encoding device can efficiently encode the necessary parameters while omitting unnecessary signaling.

[0738] Figure 88 is a flowchart showing the decoding procedure for the number of divisions, division method, quantization parameters, and transformation parameters in the decoding device.

[0739] The decoding device first decodes the subdivision_iterationCount_override_flag from the bitstream, which indicates whether or not to overwrite the number of divisions (S2041).

[0740] The decoding device determines whether subdivision_iterationCount_override_flag is TRUE or not (S2042).

[0741] If subdivision_iterationCount_override_flag is TRUE (Yes in S2042), the decoder decodes subdivision_iterationCount from the bitstream (S2043).

[0742] The decoder decodes subdivision_methods from the bitstream after encoding subdivision_iterationCount (S2044).

[0743] The decryption device decrypts quantization_params from the bitstream after decrypting subdivision_methods (S2045).

[0744] The decryption device decrypts transform_params from the bitstream after decrypting quantization_params (S2046).

[0745] On the other hand, if subdivision_iterationCount_override_flag is FALSE (No in S2042), the decoder estimates subdivision_iterationCount from higher layers, etc. (S2047).

[0746] The decoder decodes the subdivision_method_override_flag from the bitstream (S2048).

[0747] The decoding device determines whether subdivision_method_override_flag is TRUE or not (S2049).

[0748] If subdivision_method_override_flag is TRUE (Yes in S2049), the decryption device decrypts subdivision_methods from the bitstream (S2050).

[0749] If subdivision_method_override_flag is FALSE (No in S2049), the decoding device estimates subdivision_methods (S2051).

[0750] The decoder decodes quantization_override_flag from the bitstream (S2052).

[0751] The decoding device determines whether quantization_override_flag is TRUE or not (S2053).

[0752] If quantization_override_flag is TRUE (Yes in S2053), the decoder decodes quantization_params from the bitstream (S2054).

[0753] If quantization_override_flag is FALSE (No in S2053), the decoder estimates quantization_params (S2055).

[0754] The decryption device decrypts transform_params_override_flag from the bitstream (S2056).

[0755] The decoding device determines whether transform_params_override_flag is TRUE or not (S2057).

[0756] If transform_params_override_flag is TRUE (Yes in S2057), the decryption device decrypts transform_params from the bitstream (S2058).

[0757] The decoding device estimates transform_params (S2059) if transform_params_override_flag is FALSE (No in S2057).

[0758] Thus, by performing the processing according to the procedure shown in Figure 88, the decoding device can efficiently decode the necessary parameters while omitting unnecessary signaling.

[0759] The correspondence between the names mentioned above will be explained below.

[0760] Here, the second flag, named subdivision_iterationCount_override_flag in the flowchart, indicates whether or not to override the number of divisions for the target hierarchy. The second parameter is subdivision_iterationCount, which indicates the number of divisions applied to the target hierarchy. Furthermore, the third set of parameters is subdivision_methods, which indicates the identifier of the method used for the division process. The fourth set of parameters is quantization_params, which indicates the quantization parameters applied to the target hierarchy. The fifth set of parameters is transform_params, which indicates the lifting transformation parameters applied to the target hierarchy. The third flag is subdivision_method_override_flag, which indicates whether or not a division method exists in the target hierarchy. The fourth flag is quantization_override_flag, which indicates whether or not quantization parameters exist in the target hierarchy. The seventh flag is transform_params_override_flag, which indicates whether or not lifting transformation parameters exist in the target hierarchy.

[0761] Next, each of Examples 1 to 5 in this embodiment will be described.

[0762] (Example 1) Figure 89 is a diagram showing the syntax configuration according to Example 1 of this embodiment.

[0763] The syntax in this example is configured such that when the higher-level control flag parameters_override_flag is 1 (TRUE), the following set of flags are shown in order (each flag is u(1), which is a 1-bit unsigned value). First, subdivision_iterationCount_override_flag is shown, and only when this flag is 0 (FALSE), subdivision_method_override_flag, quantization_override_flag, and transform_params_override_flag are shown in that order. Here, `subdivision_iterationCount_override_flag` indicates whether to override the number of subdivisions (`subdivision_iterationCount`) for the target hierarchy, `subdivision_method_override_flag` indicates whether to override the subdivision method (`subdivision_methods`), `quantization_override_flag` indicates whether to override the quantization parameters (`quantization_params`), and `transform_params_override_flag` indicates whether to override the lifting transformation parameters (`transform_params`). Note that if `parameters_override_flag` is 0, these flags are not shown.

[0764] When subdivision_iterationCount_override_flag is 1, subdivision_method_override_flag, quantization_override_flag, and transform_params_override_flag are each estimated to be 1 (TRUE) (therefore, these three flags are not shown in the description). This estimation reduces the amount of information due to the flags, and the handling (description or estimation) of each parameter (subdivision_methods / quantization_params / transform_params) is carried out according to the procedure shown in Figure 87, etc.

[0765] As a concrete example, (Case 1) If only subdivision_iterationCount is changed, this can be achieved by setting subdivision_iterationCount_override_flag = 1. In this case, subdivision_method_override_flag, quantization_override_flag, and transform_params_override_flag are not shown, thus reducing the amount of information. (Case 2) If only subdivision_method is changed, this can be achieved by setting subdivision_iterationCount_override_flag = 0, subdivision_method_override_flag = 1, quantization_override_flag = 0, and transform_params_override_flag = 0. In this case, quantization parameters and transformation parameters are not required, thus reducing the amount of information.

[0766] (Example 2) Figure 90 is a diagram showing the syntax configuration according to Example 2 of this embodiment.

[0767] The syntax in this example is configured such that, when the higher-level control flag parameters_override_flag is 1 (TRUE), the following set of flags are shown in order (each flag is u(1), which is a 1-bit unsigned value). First, subdivision_iterationCount_override_flag is shown, and only if this flag is 0 (FALSE), subdivision_method_override_flag is shown. Next, quantization_exist_flag is shown only if quantization_exist_flag is 1 (TRUE) and subdivision_iterationCount_override_flag is 0 (FALSE). Subsequently, transform_method_override_flag is shown, and furthermore, transform_params_override_flag is shown only if subdivision_iterationCount_override_flag is 0 (FALSE) and transform_method_override_flag is 0 (FALSE). Note that if parameter_override_flag is 0, the flags will not be displayed.

[0768] When subdivision_iterationCount_override_flag is 1 (TRUE), subdivision_method_override_flag is estimated to be 1 (TRUE). quantization_override_flag is estimated to be 1 (TRUE) if subdivision_iterationCount_override_flag is 1 (TRUE) and quantization exists in the target hierarchy (quantization_exist_flag = 1), and is estimated to be 0 (FALSE) if quantization does not exist in the target hierarchy. transform_params_override_flag is estimated to be 1 (TRUE) if subdivision_iterationCount_override_flag is 1 (TRUE) or if the transformation method is overridden in the target hierarchy (transform_method_override_flag = 1).

[0769] In this Example 2, the advantages of Example 1 (the advantage of being able to estimate other flags and omit signaling when subdivision_iterationCount_override_flag = 1) are directly applied. Furthermore, if there is no quantization process in the target layer, it is not necessary to indicate quantization_override_flag, thus reducing the amount of information. In addition, if the transformation method is overridden in the target layer, it is not necessary to indicate transform_params_override_flag, and the need to send transformation parameters is determined based on the estimation rule, thus reducing the amount of information corresponding to that flag.

[0770] (Example 3) Figure 91 is a diagram showing the syntax configuration according to Example 3 of this embodiment.

[0771] The syntax in this example is configured such that the following elements are shown in order when the higher-level control flag parameter_override_flag is 1 (TRUE). First, subdivision_iterationCount_override_flag (u(1)) is shown, and if this flag is 1 (TRUE), subdivision_iterationCount (u(3)) is shown. subdivision_method_override_flag (u(1)) is shown only if subdivision_iterationCount_override_flag is 0 (FALSE) and subdivision_iterationCount is not 0. Next, quantization_exist_flag (u(1)) is shown only if quantization_exist_flag is 1 (TRUE), subdivision_iterationCount_override_flag is 0 (FALSE), and subdivision_iterationCount is not 0. Furthermore, transform_method_override_flag(u(1)) is shown only if subdivision_iterationCount is not 0, and transform_params_override_flag(u(1)) is shown only if subdivision_iterationCount_override_flag is 0 (FALSE), transform_method_override_flag is 0 (FALSE), and subdivision_iterationCount is not 0. Note that if parameters_override_flag is 0, the set of elements is not shown.

[0772] The subdivision_method_override_flag is estimated to be 1 (TRUE) if subdivision_iterationCount_override_flag is 1 (TRUE) and subdivision_iterationCount is not 0, and is estimated to be 0 (FALSE) if subdivision_iterationCount is 0. The quantization_override_flag is estimated to be 1 (TRUE) if subdivision_iterationCount_override_flag is 1 (TRUE) and subdivision_iterationCount is not 0 and quantization processing exists in the target layer, and is estimated to be 0 (FALSE) if subdivision_iterationCount is 0 or quantization processing does not exist in the target layer. The transform_params_override_flag is estimated to be 1 (TRUE) if subdivision_iterationCount_override_flag is 1 (TRUE) or transform_method_override_flag is 1 (TRUE) and subdivision_iterationCount is not 0, and is estimated to be 0 (FALSE) if subdivision_iterationCount is 0. Furthermore, if subdivision_iterationCount is 0, subdivision_method_override_flag, quantization_override_flag, and transform_params_override_flag are all estimated to be 0 (FALSE).

[0773] In this example 3, the advantages of example 2 are applied, and when subdivision_iterationCount is 0, it becomes unnecessary to signal flags related to the division method, quantization parameters, and transformation parameters, thus further promoting the reduction of information.

[0774] According to this syntax, the encoder encodes a parameters_override_flag (first flag) into the bitstream, indicating whether to override the parameters for encoding the first-level 3D mesh. If the first flag indicates that parameters should be overridden, the encoder encodes a subdivision_iterationCount_override_flag (second flag) into the bitstream, indicating whether to override the number of subdivisions applied to the first-level 3D mesh (indicating the existence of such subdivisions). If the second flag indicates that there are subdivisions, the encoder encodes a subdivision_iterationCount (number of subdivisions) for the first level into the bitstream.

[0775] The encoding device encodes the subdivision_method_override_flag (third flag) into the bitstream if the second flag indicates that there are no subdivisions applied to the first-level three-dimensional mesh.

[0776] The encoding device encodes the subdivision_method_override_flag (third flag) into the bitstream if the second flag indicates that there are no subdivisions applied to the first-level three-dimensional mesh, and subdivision_iterationCount (number of subdivisions) is not zero.

[0777] The encoding device does not include the subdivision_method_override_flag (third flag) in the bitstream if the second flag indicates that there is a number of subdivisions that will be applied to the first-level three-dimensional mesh.

[0778] The encoding device encodes the quantization_override_flag (fourth flag) into the bitstream if the second flag indicates that there are no divisions applied to the first-level three-dimensional mesh.

[0779] The encoding device encodes quantization_override_flag (fourth flag) into a bitstream if the second flag indicates that there are no divisions applied to the first-level three-dimensional mesh, and quantization_exist_flag (fifth flag) indicates that quantization parameters exist for the second-level three-dimensional mesh.

[0780] The encoding device does not include the quantization_override_flag (fourth flag) in the bitstream if the second flag indicates that there are a number of divisions applied to the first-level three-dimensional mesh.

[0781] If the subdivision_iterationCount (number of divisions) for the first layer is not zero, the encoding device encodes transform_method_override_flag (the sixth flag) into the bitstream, indicating that there is a transformation method for the first three-dimensional mesh.

[0782] If the encoding device indicates that there are no divisions applied to the first-level three-dimensional mesh, it encodes the transform_params_override_flag (the seventh flag) into the bitstream, which indicates whether or not to override the transformation parameters of the first three-dimensional mesh (indicating the presence or absence of such transformation parameters).

[0783] The encoding device encodes transform_params_override_flag (flag 7) into a bitstream if the second flag indicates that there are no divisions applied to the first-level three-dimensional mesh, and transform_method_override_flag (flag 6) indicates that there is no transformation method for the first three-dimensional mesh.

[0784] The encoding device does not include the transform_params_override_flag (the seventh flag) in the bitstream if the second flag indicates that there is a number of divisions that will be applied to the first-level three-dimensional mesh.

[0785] The encoding device treats the first layer as a frame or MDU (Meshpatch Data Unit) and performs encoding at that unit.

[0786] The encoding device treats the second layer as a sequence or frame and determines the presence or absence of quantization parameters at that unit.

[0787] If the second flag indicates that there are no divisions applied to the first layer's three-dimensional mesh, the encoding device encodes other flags in the bitstream indicating the presence or absence of other parameters for the first layer's three-dimensional mesh; if the second flag indicates that there are divisions, it omits those other flags from the bitstream.

[0788] Furthermore, following this syntax, the decoder decodes parameters_override_flag (first flag) from the bitstream, indicating whether or not to override the parameters for decoding the first-level three-dimensional mesh; if the first flag indicates an override, it decodes subdivision_iterationCount_override_flag (second flag) from the bitstream, indicating whether or not there is a subdivision count applied to the first-level three-dimensional mesh; and if the second flag indicates that there is a subdivision count, it decodes subdivision_iterationCount (number of subdivisions) from the bitstream.

[0789] The decoder decodes the subdivision_method_override_flag (third flag) from the bitstream if the second flag indicates that there are no divisions applied to the first-level three-dimensional mesh.

[0790] The decoder decodes the subdivision_method_override_flag (third flag) from the bitstream if the second flag indicates no divisions and subdivision_iterationCount (number of divisions) is not 0.

[0791] The decoding device assumes that there is a subdivision method for the first three-dimensional mesh (it assumes that there is a subdivision method) if the second flag indicates that there is a subdivision count and subdivision_iterationCount (number of subdivisions) is not 0.

[0792] The decoding device assumes that there is no subdivision method for the first three-dimensional mesh if subdivision_iterationCount (number of subdivisions) is 0 (it assumes there are no subdivision methods).

[0793] If the second flag indicates no divisions, the decoder decodes the quantization_override_flag (fourth flag) from the bitstream.

[0794] The decoder decodes the quantization_override_flag (fourth flag) from the bitstream if the second flag indicates no division count and the quantization_exist_flag (fifth flag) indicates that quantization parameters exist in the second-level three-dimensional mesh.

[0795] The decoding device assumes that the first three-dimensional mesh has quantization parameters (it assumes that quantization_params (the fourth parameter) exists) if the second flag indicates that there are divisions and the quantization_exist_flag (fifth flag) indicates that quantization parameters exist in the second layer.

[0796] The decoding device assumes that the first three-dimensional mesh has no quantization parameters (i.e., it assumes that quantization_params (the fourth parameter) is missing) if quantization_exist_flag (the fifth flag) indicates that no quantization parameters exist in the second layer.

[0797] If the subdivision_iterationCount (number of divisions) is not zero, the decryption device decrypts the transform_method_override_flag (sixth flag) from the bitstream.

[0798] If the second flag indicates no divisions, the decoder decodes the transform_params_override_flag (the seventh flag) from the bitstream.

[0799] The decoder decodes transform_params_override_flag (flag 7) from the bitstream if the second flag indicates no division count and transform_method_override_flag (flag 6) indicates no transformation method.

[0800] The decoding device assumes that the first three-dimensional mesh has transformation parameters (it assumes that transform_params (fifth parameter) exists) if the second flag indicates that there is a division count, or if transform_method_override_flag (sixth flag) indicates that there is a transformation method and subdivision_iterationCount (division count) is not 0.

[0801] The decoding device assumes that there are no transformation parameters for the first three-dimensional mesh if subdivision_iterationCount (number of divisions) is 0 (it assumes that transform_params (fifth parameter) is missing).

[0802] The decryption device treats the first layer as a frame or MDU (Meshpatch Data Unit) and performs the decryption process on that unit.

[0803] The decoding device treats the second layer as a sequence or frame and determines the presence or absence of quantization parameters at that level.

[0804] If the second flag indicates no division count, the decoder decodes from the bitstream a set of other flags indicating the presence or absence of other parameters for the first three-dimensional mesh for the first layer. If the second flag indicates a division count, the decoder determines the other parameters without decoding them (determining their values ​​based on estimation rules).

[0805] Note that indicating whether or not to override each parameter can be interpreted as indicating the presence or absence of each parameter for overriding.

[0806] (Example 4) Figure 92 is a diagram showing the syntax configuration according to Example 4 of this embodiment.

[0807] The syntax in this example is configured such that when the higher-level control flag parameters_override_flag is 1 (TRUE), the following flags and parameters are shown in order (each flag is defined by a code length such as u(1) and each parameter by a code length such as u(3)). First, subdivision_iterationCount_override_flag is shown, and if this flag is 0 (FALSE), subdivision_method_override_flag is shown. Next, if quantization_exist_flag is 1 (TRUE) and subdivision_iterationCount_override_flag is 0 (FALSE), quantization_override_flag is shown. Subsequently, transform_method_override_flag is shown, and if this flag is also 1 (TRUE), transform_method (a 3-bit parameter indicating the transformation method) is shown. Finally, transform_params_override_flag is shown only if subdivision_iterationCount_override_flag is 0 (FALSE), transform_method_override_flag is 0 (FALSE), and transform_method is not NONE_TRANSFORM. Note that if parameters_override_flag is 0 (FALSE), these flags and parameters are not shown.

[0808] The subdivision_method_override_flag is estimated to be 1 (TRUE) when subdivision_iterationCount_override_flag is 1 (TRUE), and estimated to be 0 (FALSE) when subdivision_iterationCount is 0. The quantization_override_flag is estimated to be 1 (TRUE) when subdivision_iterationCount_override_flag is 1 (TRUE) and quantization exists in the target layer (quantization_exist_flag = 1), and estimated to be 0 (FALSE) when quantization does not exist in the target layer. The transform_params_override_flag is estimated to be 1 (TRUE) if subdivision_iterationCount_override_flag is 1 (TRUE), or if transform_method_override_flag is 1 (TRUE) and transform_method is not NONE_TRANSFORM. On the other hand, if subdivision_iterationCount is 0 or transform_method is NONE_TRANSFORM, transform_params_override_flag is estimated to be 0 (FALSE).

[0809] In this example 4, the advantages of examples 2 and 3 are directly applied. Furthermore, if transform_method is NONE_TRANSFORM in the target hierarchy, it becomes unnecessary to indicate transform_params_override_flag, thus omitting signaling for that flag and reducing the amount of information.

[0810] (Example 5) Figure 93 is a diagram showing the syntax configuration according to Example 5 of this embodiment.

[0811] The syntax in this example is configured such that, when the higher-level control flag parameters_override_flag is 1 (TRUE), the following flags and parameters are shown in order (each flag is u(1), and each parameter is an unsigned value of a specified bit length). First, subdivision_iterationCount_override_flag is shown, and only if this flag is 1 (TRUE), subdivision_iterationCount is shown. Only if subdivision_iterationCount_override_flag is 0 (FALSE) and subdivision_iterationCount is not 0, subdivision_method_override_flag is shown. Furthermore, only if quantization_exist_flag is 1 (TRUE), subdivision_iterationCount_override_flag is 0 (FALSE), and subdivision_iterationCount is not 0, quantization_override_flag is shown. If subdivision_iterationCount is not 0, transform_method_override_flag is shown, and if this flag is 1 (TRUE), transform_method is shown. Finally, transform_params_override_flag is shown only if subdivision_iterationCount_override_flag is 0 (FALSE), transform_method_override_flag is 0 (FALSE), subdivision_iterationCount is not 0, and transform_method is not NONE_TRANSFORM. Note that if parameters_override_flag is 0, these flags and parameters are not shown.

[0812] The subdivision_method_override_flag is estimated to be 1 (TRUE) if subdivision_iterationCount_override_flag is 1 (TRUE) and subdivision_iterationCount is not 0, and is estimated to be 0 (FALSE) if subdivision_iterationCount is 0. The quantization_override_flag is estimated to be 1 (TRUE) if subdivision_iterationCount_override_flag is 1 (TRUE) and quantization processing exists in the target layer and subdivision_iterationCount is not 0, and is estimated to be 0 (FALSE) if quantization processing does not exist in the target layer or if subdivision_iterationCount is 0. The transform_params_override_flag is presumed to be 1 (TRUE) if subdivision_iterationCount_override_flag is 1 (TRUE), or if the transformation method is overridden at the target hierarchy, subdivision_iterationCount is not 0, and transform_method is not NONE_TRANSFORM. If subdivision_iterationCount is 0, subdivision_method_override_flag, quantization_override_flag, and transform_params_override_flag are presumed to be 0 (FALSE).

[0813] In this Example 5, the advantages of Examples 3 and 4 are combined and applied. Specifically, in addition to the advantage of being able to estimate other flags and omit signaling when subdivision_iterationCount_override_flag = 1, it is not necessary to indicate quantization_override_flag when there is no quantization process in the target layer or when subdivision_iterationCount is 0, and furthermore, it is not necessary to indicate transform_params_override_flag when the transformation method is NONE_TRANSFORM. This reduces unnecessary signaling and improves bit efficiency.

[0814] According to this embodiment, it is possible to reduce the size of the bitstream. In particular, when the number of divisions is constant in the target hierarchy and only the division method differs, it is possible to provide flexibility in estimating the quantization parameters and transformation parameters to be applied to that hierarchy. As a result, the encoding of quantization parameters and transformation parameters in the bitstream can be omitted, decoupled from the process of overriding the number of divisions and the division method. Furthermore, when the number of divisions changes, further bit reduction can be achieved by estimating subdivision_method_override_flag as 1 (TRUE).

[0815] Furthermore, the above effects can be combined and applied to achieve additional information reduction by combining various syntax table designs.

[0816] Figure 94 is a flowchart showing the procedure for processing the three-dimensional mesh of the hierarchy to be encoded in the encoding device.

[0817] The encoding device encodes a first parameter indicating the number of divisions for the three-dimensional mesh of the hierarchy to be encoded into a bitstream (S2061). As a modification, the sub-division method applied to the hierarchy may also be encoded. Alternatively, the sub-division method may be encoded only if the number of divisions is not zero. With this modification, the amount of encoding can be reduced by omitting the signaling of the sub-division method.

[0818] The encoding device determines whether the number of divisions differs from the reference number of divisions (S2062). In a modified example, the reference number of divisions may be the number of divisions of other frames, submeshes, or LoDs, or it may be a predefined value.

[0819] If the number of divisions differs from the standard (Yes in S2062), the encoding device encodes a second parameter indicating the quantization parameter for the target hierarchy and a third parameter indicating the transformation parameter into a bitstream (S2063). As a modification, the quantization parameter may be encoded only if the number of divisions is not zero. With this modification, the amount of code can be reduced by omitting the signaling of the quantization parameter. Furthermore, the quantization parameter may be encoded only if a quantization process exists. In this case, the presence or absence of a quantization process is indicated by a flag encoded in the bitstream. Similarly, the transformation parameter may be encoded only if the number of divisions is not zero, or only if a transformation process exists (other than a NONE transformation). In this case, the presence or absence of a transformation process is indicated by a flag encoded in the bitstream.

[0820] The encoding device, if the number of divisions is the same as the standard (No in S2062), encodes a fourth flag in the bitstream indicating whether or not to encode quantization parameters for the three-dimensional mesh of the target hierarchy, and a seventh flag indicating whether or not to encode transformation parameters (S2064). As a modification, the fourth flag may be encoded only if the number of divisions is not zero, and if the number of divisions is zero, the fourth flag may be assumed to be zero. According to this modification, the amount of encoding can be reduced by omitting the signaling of quantization parameters. Furthermore, the fourth flag may be encoded only if quantization processing exists. In this case, the presence or absence of quantization processing is indicated by the flag encoded in the bitstream. Similarly, the seventh flag may be encoded only if the number of divisions is not zero, and if the number of divisions is zero, the seventh flag may be assumed to be zero. Also, the seventh flag may be encoded only if transformation processing exists (other than NONE). In this case, the presence or absence of transformation processing is indicated by the flag encoded in the bitstream.

[0821] The encoding device encodes displacement information for reconstructing the three-dimensional mesh of the hierarchy to be encoded using at least the above number of divisions (S2065).

[0822] Figure 95 is a flowchart showing the procedure for processing the three-dimensional mesh of the hierarchy to be encoded in the decoding device.

[0823] The decoding device decodes a first parameter indicating the number of divisions for the three-dimensional mesh of the target hierarchy from the bitstream (S2071). In this modified case, the decoding device may decode the division method to be applied to the target hierarchy. According to this modified case, it becomes possible to specify the division method to be applied to the hierarchy. In this modified case, the division method may be decoded only if the number of divisions is not zero. According to this modified case, the amount of code can be reduced by omitting the signaling of the division method.

[0824] The decoding device determines whether the number of divisions of the target hierarchy differs from the reference number of divisions (S2072). As an alternative, the reference number of divisions may be the number of divisions of other frames, submeshes, or LoDs. Alternatively, the reference number of divisions may be a predefined value.

[0825] If the number of divisions differs from the standard number of divisions (Yes in S2072), the decoding device decodes a second parameter indicating the quantization parameter for the target hierarchy and a third parameter indicating the transformation parameter from the bitstream (S2073). As a modification, the quantization parameter may be decoded only if the number of divisions is not zero. According to this modification, the code amount can be reduced by omitting the signaling of the quantization parameter. Also, the second parameter may be decoded only if a quantization process exists. According to this modification, the presence of a quantization process is indicated by a flag included in the bitstream, and the code amount can be reduced by omitting unnecessary signaling. Furthermore, the transformation parameter may be decoded only if the number of divisions is not zero. According to this modification, the code amount can be reduced by omitting the signaling of the transformation parameter. Also, the third parameter may be decoded only if a transformation process exists (other than NONE). According to this modification, the presence of a transformation process is indicated by a flag included in the bitstream, and the code amount can be reduced by omitting unnecessary signaling.

[0826] If the number of divisions is the same as the standard number of divisions (No in S2072), the decoding device decodes a fourth flag indicating whether or not to decode quantization parameters and a seventh flag indicating whether or not to decode transformation parameters from the bitstream for the three-dimensional mesh of the target hierarchy (S2074). As a modification, the fourth flag may be decoded only if the number of divisions is not 0. If the number of divisions is 0, the fourth flag is presumed to be 0. According to this modification, the amount of code can be reduced by omitting unnecessary signaling. Also, the fourth flag may be decoded only if quantization processing exists. According to this modification, the presence of quantization processing is indicated by a flag included in the bitstream. Furthermore, the seventh flag may be decoded only if the number of divisions is not 0. If the number of divisions is 0, the seventh flag is presumed to be 0. According to this modification, the amount of code can be reduced by omitting unnecessary signaling. Also, the seventh flag may be decoded only if transformation processing exists (other than NONE). According to this modification, the presence of transformation processing is indicated by a flag included in the bitstream.

[0827] The decoding device decodes the displacement information using at least the number of divisions in order to reconstruct the three-dimensional mesh of the target hierarchy (S2075).

[0828] Figure 96 is a flowchart showing a first example of the procedure for processing the three-dimensional mesh of the hierarchy to be encoded in an encoding device.

[0829] The encoding device encodes subdivision_override_flag into a bitstream (S2081).

[0830] The encoding device determines whether subdivision_override_flag is TRUE or not (S2082).

[0831] If subdivision_override_flag is TRUE (Yes in S2082), the encoding device encodes subdivision_iterationCount into a bitstream (S2083).

[0832] The encoding device determines whether subdivision_iterationCount is different from ref_iterationCount (S2084).

[0833] If subdivision_iterationCount is different from ref_iterationCount (Yes in S2084), the encoding device encodes quantization_params into a bitstream (S2085).

[0834] Subsequently, the encoding device encodes transform_params into a bitstream (S2086).

[0835] If subdivision_override_flag is FALSE (No in S2082), the encoding device estimates subdivision_iterationCount (S2087).

[0836] The encoding device encodes quantization_override_flag into a bitstream (S2088).

[0837] The encoding device determines whether quantization_override_flag is TRUE or not (S2089).

[0838] If quantization_override_flag is TRUE (Yes in S2089), the encoding device encodes quantization_params into a bitstream (S2090).

[0839] If quantization_override_flag is FALSE (No in S2089), the encoding device estimates quantization_params (S2091).

[0840] The encoding device encodes transform_params_override_flag into a bitstream (S2092).

[0841] The encoding device determines whether transform_params_override_flag is TRUE or not (S2093).

[0842] If transform_params_override_flag is TRUE (Yes in S2093), the encoding device encodes transform_params into a bitstream (S2094).

[0843] If transform_params_override_flag is FALSE (No in S2093), the encoding device estimates transform_params (S2095).

[0844] Figure 97 is a flowchart showing a first example of the procedure for processing the three-dimensional mesh of the hierarchy to be decoded in the decoding device.

[0845] The decoder decodes the subdivision_override_flag from the bitstream (S2101).

[0846] The decoding device determines whether subdivision_override_flag is TRUE or not (S2102).

[0847] If subdivision_override_flag is TRUE (Yes in S2102), the decoder decodes subdivision_iterationCount from the bitstream (S2103).

[0848] The decoding device determines whether subdivision_iterationCount is different from ref_iterationCount (S2104).

[0849] If subdivision_iterationCount is different from ref_iterationCount (Yes in S2104), the decoder decodes quantization_params from the bitstream (S2105).

[0850] Subsequently, the decryption device decrypts transform_params from the bitstream (S2106).

[0851] If subdivision_override_flag is FALSE (No in S2102), the decoder estimates subdivision_iterationCount (S2107).

[0852] The decoder decodes quantization_override_flag from the bitstream (S2108).

[0853] The decoding device determines whether quantization_override_flag is TRUE or not (S2109).

[0854] If quantization_override_flag is TRUE (Yes in S2109), the decoder decodes quantization_params from the bitstream (S2110).

[0855] If quantization_override_flag is FALSE (No in S2109), the decoder estimates quantization_params (S2111).

[0856] The decoder decodes transform_params_override_flag from the bitstream (S2112).

[0857] The decoding device determines whether transform_params_override_flag is TRUE or not (S2113).

[0858] If transform_params_override_flag is TRUE (Yes in S2113), the decoder decodes transform_params from the bitstream (S2114).

[0859] If transform_params_override_flag is FALSE (Nos in S2113), the decoder estimates transform_params (S2115).

[0860] Figure 98 is a flowchart showing a second example of the procedure for processing the three-dimensional mesh of the hierarchy to be encoded in an encoding device.

[0861] The encoding device encodes subdivision_override_flag into a bitstream (S2081).

[0862] The encoding device determines whether subdivision_override_flag is TRUE or not (S2082).

[0863] If subdivision_override_flag is TRUE (Yes in S2082), the encoding device encodes subdivision_iterationCount and subdivision_method into a bitstream (S2083a).

[0864] If subdivision_override_flag is FALSE (No in S2082), the encoding device estimates subdivision_iterationCount and subdivision_method (S2087a).

[0865] The encoding device determines whether subdivision_iterationCount is different from ref_iterationCount (S2084).

[0866] If subdivision_iterationCount is different from ref_iterationCount (Yes in S2084), the encoding device encodes quantization_params into a bitstream (S2085).

[0867] Subsequently, the encoding device encodes transform_params into a bitstream (S2086).

[0868] The encoding device encodes quantization_override_flag into a bitstream (S2088).

[0869] The encoding device determines whether quantization_override_flag is TRUE or not (S2089).

[0870] If quantization_override_flag is TRUE (Yes in S2089), the encoding device encodes quantization_params into a bitstream (S2090).

[0871] If quantization_override_flag is FALSE (No in S2089), the encoding device estimates quantization_params (S2091).

[0872] The encoding device encodes transform_params_override_flag into a bitstream (S2092).

[0873] The encoding device determines whether transform_params_override_flag is TRUE or not (S2093).

[0874] If transform_params_override_flag is TRUE (Yes in S2093), the encoding device encodes transform_params into a bitstream (S2094).

[0875] If transform_params_override_flag is FALSE (No in S2093), the encoding device estimates transform_params (S2095).

[0876] Figure 99 is a flowchart showing a second example of the procedure for processing the three-dimensional mesh of the hierarchy to be decoded in the decoding device.

[0877] The decoder decodes the subdivision_override_flag from the bitstream (S2101).

[0878] The decoding device determines whether subdivision_override_flag is TRUE or not (S2102).

[0879] If subdivision_override_flag is TRUE (Yes in S2102), the decoder decodes subdivision_iterationCount and subdivision_method from the bitstream (S2103a).

[0880] If subdivision_override_flag is FALSE (No in S2102), the decoding device estimates subdivision_iterationCount and subdivision_method (S2107a).

[0881] The decoding device determines whether subdivision_iterationCount is different from ref_iterationCount (S2104).

[0882] If subdivision_iterationCount is different from ref_iterationCount (Yes in S2104), the decoder decodes quantization_params from the bitstream (S2105).

[0883] Subsequently, the decryption device decrypts transform_params from the bitstream (S2106).

[0884] The decoder decodes quantization_override_flag from the bitstream (S2108).

[0885] The decoding device determines whether quantization_override_flag is TRUE or not (S2109).

[0886] If quantization_override_flag is TRUE (Yes in S2109), the decoder decodes quantization_params from the bitstream (S2110).

[0887] If quantization_override_flag is FALSE (No in S2109), the decoder estimates quantization_params (S2111).

[0888] The decoder decodes transform_params_override_flag from the bitstream (S2112).

[0889] The decoding device determines whether transform_params_override_flag is TRUE or not (S2113).

[0890] If transform_params_override_flag is TRUE (Yes in S2113), the decoder decodes transform_params from the bitstream (S2114).

[0891] If transform_params_override_flag is FALSE (No in S2113), the decoder estimates transform_params (S2115).

[0892] The correspondence between the names mentioned above will be explained below.

[0893] Here, the second flag is subdivision_override_flag in the flowchart name, and it indicates whether or not a subdivision exists for the target hierarchy. The second set of parameters is quantization_params, which indicates the quantization parameters applied to the target hierarchy. Furthermore, the third set of parameters is transform_params, which indicates the lifting transformation parameters applied to the target hierarchy. The fourth flag is quantization_override_flag, which indicates whether or not quantization parameters exist in the target hierarchy. The seventh flag is transform_params_override_flag, which indicates whether or not lifting transformation parameters exist in the target hierarchy.

[0894] The second flag corresponds to the combination of subdivision_iterationCount_override_flag and subdivision_method_override_flag in the conventional first aspect.

[0895] (Example 6) Figure 100 is a diagram showing the syntax configuration according to Example 6 of this embodiment.

[0896] The syntax in this example is structured so that the following elements are shown in order when the higher-level control flag parameters_override_flag is 1 (TRUE). First, subdivision_override_flag is shown (all flags are unsigned 1-bit values ​​of u(1)), followed by a statement that initializes the internal state subdivision_iterationCount_overridden to FALSE. Next, subdivision_iteration_count is shown only if subdivision_override_flag is 1, and then a statement that sets subdivision_iterationCount_overridden to TRUE appears if subdivision_iteration_count is different from the reference value ref_subdivision_iteration_count. Subsequently, quantization_override_flag is shown only if subdivision_iterationCount_overridden is FALSE, followed by transform_method_override_flag. Finally, transform_params_override_flag is shown only if subdivision_iterationCount_overridden is FALSE.

[0897] Here, subdivision_override_flag indicates whether to override at least subdivision_iteration_count or subdivision_method for the target layer. subdivision_iterationCount_overridden represents an internal state that holds the result of determining whether subdivision_iteration_count is different from the number of reference divisions (ref_subdivision_iteration_count).

[0898] When subdivision_override_flag is 1, it indicates that at least one of subdivision_iteration_count or subdivision_method exists in the target hierarchy. Furthermore, if subdivision_iterationCount_overridden is TRUE (i.e., the number of divisions is different from the reference value), quantization_override_flag and transform_params_override_flag are presumed to be 1 (TRUE), and these flags are not shown syntactically. This presumption allows the presence of quantization parameters and lifting transformation parameters to be communicated without additional flags, reducing the amount of information that would otherwise be required for the flags. On the other hand, if subdivision_iterationCount_overridden is FALSE (i.e., the number of divisions is the same as the reference value), it is not necessarily required to show the quantization parameters and transformation parameters, so the encoder can selectively communicate the presence or absence of these parameters by showing quantization_override_flag and transform_params_override_flag. This selective signaling also results in a reduction in the amount of information.

[0899] As described above, the syntax shown in Figure 100 allows for the omission or selective signaling of flags based on the number of divisions, the matching / mismatch of the reference value, and the estimation rules for each override flag. As a result, it can uniquely communicate to the decoder whether or not the necessary quantization parameters and lifting transformation parameters are present, while reducing the amount of information in the bitstream.

[0900] (Example 7) Figure 101 is a diagram showing the syntax configuration according to Example 7 of this embodiment.

[0901] The syntax in this example is configured such that, when the higher-level control flag parameters_override_flag is 1 (TRUE), the following elements are shown in order (each flag is an unsigned 1-bit value of u(1)). First, subdivision_override_flag is shown, followed by the internal state subdivision_iterationCount_overridden being initialized to FALSE. Then, only if subdivision_override_flag is 1, subdivision_iteration_count is shown, and if subdivision_iteration_count is different from the reference value ref_subdivision_iteration_count, subdivision_iterationCount_overridden is set to TRUE.

[0902] Next, quantization_override_flag is shown only if quantization_exist_flag is 1 and su...

Claims

1. An encoding method performed by an encoding device, wherein the encoding method of a data unit to be encoded, including a mesh data unit (MDU), is determined to be interpredictive, and if it is determined to be interpredictive, a reference data unit indicated by reference information is identified, the reference data unit is determined to differ for each type of parameter, and the parameters are determined, and the data unit to be encoded is encoded based on the determined parameters.

2. The encoding method according to claim 1, wherein, in the determination, if the data unit to be encoded does not include the parameter, the parameter is determined by referring to the corresponding parameter of the referenced data unit indicated in the reference information.

3. The encoding method according to claim 1 or 2, wherein the parameter is represented by the difference between the reference data unit indicated in the reference information and the corresponding parameter.

4. The encoding method according to claim 1 or 2, wherein the reference data unit indicated by the reference information is a data unit that has been processed before the data unit to be encoded.

5. The encoding method according to claim 1 or 2, wherein the data unit to be encoded is an intermesh data unit (IMDU) or a merged mesh data unit (MMDU).

6. A decoding method performed by a decoding device, comprising: determining whether the decoding method for a data unit to be decoded, including a mesh data unit (MDU), is interpredictive or not; if it is determined to be interpredictive, identifying a reference data unit indicated by reference information; determining parameters by specifying different reference units for each type of parameter; and decoding the data unit to be decoded based on the determined parameters.

7. The decoding method according to claim 6, wherein, in the determination, if the data unit to be decoded does not contain the parameter, the parameter is determined by referring to the corresponding parameter of the referenced data unit indicated in the reference information.

8. The decoding method according to claim 6 or 7, wherein the parameter is represented by the difference between the referenced data unit indicated in the reference information and the corresponding parameter.

9. The decoding method according to claim 6 or 7, wherein the referenced data unit indicated by the reference information is a data unit that was processed before the data unit to be decoded.

10. The decoding method according to claim 6 or 7, wherein the data unit to be decoded is an intermesh data unit (IMDU) or a merged mesh data unit (MMDU).

11. An encoding device comprising a circuit and a memory connected to the circuit, wherein the circuit, in operation, determines whether the encoding method of a data unit to be encoded, including a mesh data unit (MDU), is interpredictive or not, identifies a reference data unit indicated by reference information, determines parameters by specifying different references for each type of parameter, and encodes the data unit to be encoded based on the determined parameters.

12. A decoding device comprising a circuit and a memory connected to the circuit, wherein the circuit, in operation, determines whether the decoding method of a data unit to be decoded, including a mesh data unit (MDU), is interpredictive or not, identifies a reference data unit indicated by reference information, determines parameters by specifying different reference units for each type of parameter, and decodes the data unit to be decoded based on the determined parameters.