Encoding method, decoding method, encoding device, and decoding device

By encoding submeshes with dependency information, the method enhances the efficiency of three-dimensional data encoding and decoding, reducing storage and transmission requirements.

WO2025211289A1PCT designated stage Publication Date: 2025-10-09PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/012883
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-05
Filing Date
2025-03-28
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing encoding and decoding processes for three-dimensional data are inefficient and require significant storage and bandwidth due to the repetition of patterns and attributes in 3D meshes, leading to suboptimal compression and transmission.

Method used

The method generates multiple encoded data by encoding submeshes, creating a parameter set that includes identification information about their dependencies, allowing for efficient decoding without strict order dependency and optimizing storage and transmission.

Benefits of technology

This approach reduces the amount of information needed in the bitstream by enabling individual decoding of submeshes based on dependency information, thus improving storage efficiency and transmission speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025012883_09102025_PF_FP_ABST
    Figure JP2025012883_09102025_PF_FP_ABST
Patent Text Reader

Abstract

In an encoding method according to one embodiment of the present disclosure: a plurality of sub-meshes are encoded to generate a plurality of pieces of encoded data (S551); a parameter set indicating the configuration of the plurality of sub-meshes is generated (S552); a bit stream including the plurality of pieces of encoded data and the parameter set is generated; and, in cases where a prescribed condition is satisfied, the parameter set includes first identification information indicating whether each of at least one sub-mesh among the plurality of sub-meshes depends on another sub-mesh among the plurality of sub-meshes.
Need to check novelty before this filing date? Find Prior Art

Description

Encoding method, decoding method, encoding device, and decoding device

[0001] The present disclosure relates to encoding methods and the like.

[0002] In US Pat. No. 6,299,549 a method and apparatus for encoding and decoding three-dimensional mesh data is proposed.

[0003] Japanese Patent Application Laid-Open No. 2006-187015

[0004] Further improvements are desired in the encoding or decoding process for three-dimensional data. The present disclosure aims to improve the encoding or decoding process for three-dimensional data.

[0005] An encoding method according to one aspect of the present disclosure generates multiple encoded data by encoding multiple submeshes, generates a parameter set indicating the configuration of the multiple submeshes, and generates a bit stream including the multiple encoded data and the parameter set, wherein the parameter set includes first identification information indicating whether or not at least one of the multiple submeshes depends on other submeshes among the multiple submeshes when a predetermined condition is met.

[0006] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0007] The present disclosure may contribute to improvements in encoding processes and the like related to three-dimensional data.

[0008] 1 is a conceptual diagram showing a three-dimensional mesh according to an embodiment. FIG. 2 is a conceptual diagram showing basic elements of a three-dimensional mesh according to an embodiment. FIG. 3 is a conceptual diagram showing mapping according to an embodiment. FIG. 4 is a block diagram showing a configuration example of an encoding / decoding system according to an embodiment. FIG. 5 is a block diagram showing a configuration example of an encoding device according to an embodiment. FIG. 6 is a block diagram showing another configuration example of an encoding device according to an embodiment. FIG. 7 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 8 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 9 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 10 is a conceptual diagram showing another configuration example of a bit stream according to an embodiment. FIG. 11 is a conceptual diagram showing yet another configuration example of a bit stream according to an embodiment. FIG. 12 is a block diagram showing a specific example of an encoding / decoding system according to an embodiment. FIG. 13 is a conceptual diagram showing an example configuration of point cloud data according to an embodiment. FIG. 14 is a conceptual diagram showing an example data file of point cloud data according to an embodiment. FIG. 15 is a conceptual diagram showing an example configuration of mesh data according to an embodiment. FIG. 16 is a conceptual diagram showing an example data file of mesh data according to an embodiment. FIG. 17 is a conceptual diagram showing types of three-dimensional data according to an embodiment. FIG. 18 is a block diagram showing an example configuration of a three-dimensional data encoder according to an embodiment. FIG. 19 is a block diagram showing an example configuration of a three-dimensional data decoder according to an embodiment. FIG. 19 is a block diagram showing another configuration example of a three-dimensional data encoder according to an embodiment. FIG. 19 is a block diagram showing another configuration example of a three-dimensional data decoder according to an embodiment. FIG. 1 is a conceptual diagram showing a specific example of encoding processing according to an embodiment. FIG. 2 is a conceptual diagram showing a specific example of decoding processing according to an embodiment. FIG. 3 is a block diagram showing an implementation example of an encoding device according to an embodiment. FIG. 4 is a block diagram showing an implementation example of a decoding device according to an embodiment. FIG. 5 is a block diagram showing another configuration example of an encoding / decoding system according to an embodiment. FIG. 6 is a block diagram showing another configuration example of an encoding device according to an embodiment. FIG. 7 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 8 is a block diagram showing a detailed configuration example of an encoding device according to an embodiment. FIG. 9 is a block diagram showing a modified example of the detailed configuration of an encoding device according to an embodiment. FIG. 10 is a flow diagram showing processing of an encoding device according to an embodiment. FIG. 11 is an explanatory diagram conceptually showing encoding of a mesh frame according to an embodiment.1 is a block diagram showing a detailed configuration example of a decoding device according to an embodiment; FIG. 2 is a block diagram showing a modified example of the detailed configuration of a decoding device according to an embodiment; FIG. 3 is a flow diagram showing processing of a decoding device according to an embodiment; FIG. 4 is an explanatory diagram conceptually showing decoding of a mesh frame according to an embodiment; FIG. 5 is an explanatory diagram showing an example of subdivision according to an embodiment; FIG. 6 is an explanatory diagram showing an example of displacement of vertices after displacement after subdivision according to an embodiment; FIG. 7 is an explanatory diagram showing an example of vertices of an original mesh according to an embodiment; FIG. 8 is an explanatory diagram showing an example of a mesh according to an embodiment; FIG. 9 is an explanatory diagram showing an example of division of a mesh into sub-meshes according to an embodiment; FIG. 10 is a first explanatory diagram showing an example of packing of displacement information into an image frame according to an embodiment; FIG. 11 is an explanatory diagram showing an example of packing of displacement information into an image frame according to an embodiment; FIG. 12 is an explanatory diagram showing an example of packing of displacement information into an image frame according to an embodiment; FIG. 13 is a block diagram showing a detailed configuration example of a decoding device according to an embodiment; FIG. 14 is an explanatory diagram showing coordinates of vertices in a three-dimensional mesh according to an embodiment; FIG. 15 is an explanatory diagram showing prediction information according to an embodiment; FIG. 16 is a flow diagram showing processing of re-initializing a CABAC encoding / decoding engine in accordance with a CABAC initialization flag during encoding or decoding according to an embodiment; FIG. 17 is a block diagram showing the configuration of a first encoding unit included in a three-dimensional data encoding device according to an embodiment; 1 is a block diagram showing the configuration of a division unit according to an embodiment. FIG. 2 is a block diagram showing the configurations of a position information encoding unit and an attribute information encoding unit according to an embodiment. FIG. 3 is a block diagram showing the configuration of a first decoding unit according to an embodiment. FIG. 4 is a block diagram showing the configurations of a position information decoding unit and an attribute information decoding unit according to an embodiment. FIG. 5 is a flow diagram showing an example of processing related to initialization of CABAC in encoding of position information or encoding of attribute information according to an embodiment. FIG. 6 is a diagram showing an example of timing of CABAC initialization in point cloud data converted into a bit stream according to an embodiment. FIG. 7 is a diagram showing the configuration of encoded data according to an embodiment and a method of storing encoded data in an NAL unit. FIG. 8 is a flow diagram showing an example of processing related to initialization of CABAC in decoding of position information or decoding of attribute information according to an embodiment. FIG. 9 is a flow diagram of encoding processing of point cloud data according to an embodiment.1 is a flow diagram showing an example of a process for updating additional information according to an embodiment. FIG. 1 is a flow diagram showing an example of a process for initializing CABAC according to an embodiment. FIG. 2 is a flow diagram showing a process for decoding point cloud data according to an embodiment. FIG. 3 is a flow diagram showing an example of a process for initializing a CABAC decoding unit according to an embodiment. FIG. 4 is a diagram showing examples of tiles and slices according to an embodiment. FIG. 5 is a flow diagram showing an example of a method for initializing CABAC and determining a context initial value according to an embodiment. FIG. 6 is a diagram showing an example of a case where a map in which point cloud data obtained by LiDAR is viewed from above is divided into tiles according to an embodiment. FIG. 7 is a flow diagram showing another example of a method for initializing CABAC and determining a context initial value according to an embodiment. FIG. 8 is a diagram showing an example of a data structure of position information included in each data unit after division according to an embodiment, and an example of the syntax of a header of the position information. FIG. 9 is a flow diagram showing an example of a method for encoding three-dimensional data according to an embodiment. FIG. 10 is a flow diagram showing an example of a method for decoding three-dimensional data according to an embodiment. FIG. 11 is a diagram for explaining initialization of a context when the encoding method is switched according to an embodiment. FIG. 12 is a flow diagram of a process of a three-dimensional data encoding device according to an embodiment. FIG. 13 is a flow diagram of a process of a three-dimensional data decoding device according to an embodiment. FIG. 14 is a diagram showing an example of the syntax of an SPS according to an embodiment. FIG. 1 is a diagram showing an example of the syntax of the header (DividedGeometryHeader) of division position information according to an embodiment. FIG. 2 is a diagram showing an example of the syntax of the header (DividedAttributeHeader) of division attribute information according to an embodiment. FIG. 3 is a diagram showing another example of the syntax of the header (DividedAttributeHeader) of division attribute information according to an embodiment. FIG. 4 is a diagram showing another example of the syntax of the header (DividedGeometryHeader) of division position information according to an embodiment. FIG. 5 is a flow diagram showing an example of a first decision for determining whether to initialize entropy coding in a three-dimensional data encoding device according to an embodiment. FIG. 6 is a flow diagram showing an example of processing for determining whether an entropy coding flag complies with conformance (whether it complies with specifications) in a three-dimensional data decoding device according to an embodiment. FIG. 7 is a flow diagram of processing in a three-dimensional data encoding device according to an embodiment.1 is a flowchart of processing in a three-dimensional data decoding device according to an embodiment. FIG. 2 is a diagram showing an example of the syntax of an SPS according to an embodiment. FIG. 3 is a diagram showing an example of the syntax of an APS according to an embodiment. FIG. 4 is a diagram showing an example of the syntax of a header (DividedGeometryHeader) of division position information according to an embodiment. FIG. 5 is a diagram showing an example of the syntax of a header (DividedAttributeHeader) of division attribute information according to an embodiment. FIG. 6 is a flowchart showing an example of processing for determining whether to continue a context used for entropy encoding in a three-dimensional data encoding device according to an embodiment. FIG. 7 is a diagram for explaining table updating according to an embodiment. FIG. 8 is a flowchart of encoding an occupancy map by a three-dimensional data encoding device according to an embodiment. FIG. 9 is a flowchart of decoding an occupancy map by a three-dimensional data decoding device according to an embodiment. FIG. 10 is a flowchart of processing for switching entropy encoding methods in a three-dimensional data encoding device according to an embodiment. FIG. 11 is a flowchart of a method for continuing byte-based entropy encoding in a three-dimensional data encoding device according to an embodiment. 1 is a flow diagram of a process for switching entropy decoding methods in a three-dimensional data decoding device according to an embodiment. FIG. 2 is a flow diagram of a method for continuing byte-based entropy decoding in a three-dimensional data decoding device according to an embodiment. FIG. 3 is a flow diagram of a process for a three-dimensional data encoding device according to an embodiment. FIG. 4 is a flow diagram of a process for a three-dimensional data decoding device according to an embodiment. FIG. 5 is a diagram showing an example of a three-dimensional point cloud when encoding slices by dividing them into groups according to an embodiment. FIG. 6 is a diagram showing various examples of bitstream configurations according to an embodiment. FIG. 7 is an example in which a slice flag indicates whether to initialize CABAC for each slice, and a tree flag indicates whether to initialize CABAC for each tree within a slice, according to an embodiment. FIG. 8 is a diagram for explaining a method for decoding multiple prediction trees by parallel processing according to an embodiment. FIG. 9 is a diagram showing an example of a three-dimensional data encoding method according to an embodiment. FIG. 10 is a diagram showing an example of a three-dimensional data decoding method according to an embodiment. FIG. 11 is a diagram showing an example of parallel decoding in a three-dimensional data decoding method according to an embodiment.1 is a diagram showing an example of the syntax of a data unit of position information when an initialization flag is stored in the data of position information according to an embodiment. FIG. 2 is a diagram showing an example of header syntax when an initialization flag and offset information are stored in the header of position information according to an embodiment. FIG. 3 is a diagram showing an example of header syntax when an initialization flag and offset information are stored in the header of position information according to an embodiment in random access units. FIG. 4 is a diagram for explaining the relationship between an original mesh and sub-meshes according to an embodiment. FIG. 5 is a block diagram showing another example of the configuration of an encoding device according to an embodiment. FIG. 6 is a block diagram showing another example of the configuration of a decoding device according to an embodiment. FIG. 7 is a flow diagram showing initialization processing of context information according to an embodiment. FIG. 8 is a diagram showing another example of the syntax of an SPS according to an embodiment. FIG. 9 is a diagram showing an example of the syntax of a header of a divided base mesh (DividedBasemeshHeader) according to an embodiment. FIG. 10 is a diagram showing an example of the syntax of a header of a divided displacement vector (DividedDisplacementVectorHeader) according to an embodiment. 1 is a diagram showing another example of the syntax of a divided displacement vector header (DividedDisplacementVectorHeader) according to an embodiment. FIG. 2 is a diagram for explaining low delay encoding according to an embodiment. FIG. 3 is a diagram for explaining continuation of context information according to an embodiment. FIG. 4 is a flow diagram of a submesh encoding process according to an embodiment. FIG. 5 is a flow diagram of a submesh partial decoding process according to an embodiment. FIG. 6 is a diagram showing an encoded submesh and a submesh ID according to an embodiment. FIG. 7 is a diagram showing an example of a syntax in which metadata indicating a submesh structure is signaled according to an embodiment. FIG. 8 is a diagram showing an example of a syntax in which information of a submesh header is signaled according to an embodiment. FIG. 9 is a diagram showing another example of a syntax in which metadata indicating a submesh structure according to an embodiment. FIG. 10 is a flow diagram showing an example of a submesh partial decoding process according to an embodiment. FIG. 11 is a diagram showing a subdivided 3D mesh according to an embodiment. FIG. 12 is a diagram showing an example of a submesh according to an embodiment.1 is a diagram illustrating table information in metadata indicating a structure of a three-dimensional region according to an embodiment; FIG. 2 is a diagram illustrating table information in metadata indicating a structure of a submesh according to an embodiment; FIG. 3 is a flowchart of a submesh encoding process according to an embodiment; FIG. 4 is a flowchart of a submesh partial decoding process according to an embodiment; FIG. 5 is a diagram illustrating table information in metadata indicating a structure of a three-dimensional region according to an embodiment; FIG. 6 is a diagram illustrating table information in metadata indicating a structure of a submesh according to an embodiment; FIG. 7 is a diagram illustrating an example of syntax in which a flag indicating whether or not a sequence according to an embodiment is likely to include a submesh having a dependency relationship is signaled; FIG. 8 is a diagram illustrating an example of syntax in which a flag indicating whether or not a context continuation function is to be used according to an embodiment; FIG. 9 is a diagram illustrating an example of syntax in which metadata indicating a structure of a submesh according to an embodiment; FIG. 10 is a diagram illustrating an example of syntax in which metadata indicating a structure of a submesh according to an embodiment;

[0009] Introduction A typical three-dimensional (3D) model digitally represents an object so that a user can explore the 3D model by zooming, panning, and / or rotating in all three dimensions while rendering it over time. One way to construct such a representation is to use triangles to construct a 3D mesh. A 3D mesh stores the positions of the triangle vertices, their connectivity to each other, and their associated attributes (such as normals or UV patches). Storing all this information in an uncompressed format requires significant storage space. Transmitting all this information also requires significant bandwidth. The triangles that make up a 3D mesh often have repeating patterns and similar attributes, especially in temporal and spatial neighborhoods. These repetitions can be exploited to develop efficient encoding and decoding methods for storage and transmission.

[0010] The present disclosure relates to encoding of multimedia data, and more particularly to systems, components, and methods for encoding and decoding multimedia data. Multimedia data includes any three-dimensional digital representation of an object or surface in computer graphics applications. In particular, the present disclosure includes static and animated three-dimensional models represented by triangular meshes. Multimedia data may also include moving images.

[0011] The encoding process of point position information described in each of the multiple embodiments of the present disclosure can be applied to encoding of 3D point position information in point cloud compression methods such as Video-based PCC (V-PCC) or Geometry-based PCC (G-PCC). Similarly, the processes and syntax structures related to decoding and displaying 3D point position information described in each of the multiple embodiments of the present disclosure can be applied to decoding, displaying, and syntax structures of 3D point position information in point cloud compression methods such as V-PCC and G-PCC.

[0012] Furthermore, the encoding process for the position information and additional information of three-dimensional points described in each of the embodiments of the present disclosure can be used to encode vector groups each composed of vectors having any number of dimensions and dataset groups each composed of datasets having any number of elements. Similarly, the processes and syntax structures related to the decoding and display of position information and additional information of three-dimensional points described in each of the embodiments of the present disclosure can be used to decode, display, and syntactically structure vector groups each composed of vectors having any number of dimensions and dataset groups each composed of datasets having any number of elements. Here, each vector included in the vector group and each dataset included in the dataset group may be in any format composed of values ​​indicating the position of each dimension in a space having any number of dimensions, such as a three-dimensional point, and / or values ​​of one or more corresponding additional information.

[0013] Below, examples of inventions that can be obtained from the disclosure of this specification will be given, and the effects and the like that can be obtained from these inventions will be explained.

[0014] The encoding method of Example 1 generates multiple encoded data by encoding multiple submeshes, generates a parameter set indicating the configuration of the multiple submeshes, and generates a bit stream including the multiple encoded data and the parameter set, and the parameter set includes first identification information indicating whether at least one of the multiple submeshes depends on other submeshes among the multiple submeshes when a predetermined condition is met.

[0015] When encoding multiple submeshes (e.g., position information of each vertex included in the multiple submeshes), the context information used for encoding may be used in succession in the encoding. If the encoded multiple submeshes have a dependency relationship, the encoded multiple submeshes must be decoded in the order in which they were encoded. On the other hand, if the context information is initialized during encoding of the multiple submeshes and is not used in succession, the encoded multiple submeshes can be decoded regardless of the order in which they were encoded. Therefore, by including first identification information indicating whether each of the multiple submeshes depends on other submeshes among the multiple submeshes in the parameter set, a decoding device that acquires a bitstream can easily determine whether each of the encoded multiple submeshes can be individually decoded.

[0016] The encoding method of Example 2 is the encoding method of Example 1, wherein the parameter set may further include information indicating the number of the plurality of sub-meshes.

[0017] This allows the decoding device that has acquired the bitstream to perform processing that uses the number of multiple sub-meshes.

[0018] The encoding method of Example 3 is the encoding method of Example 1 or 2, wherein the parameter set further includes second identification information indicating whether the plurality of submeshes includes a dependent submesh that depends on another submesh among the plurality of submeshes, and the parameter set may include the first identification information if the plurality of submeshes includes the dependent submesh.

[0019] According to this, a decoding device that acquires a bitstream can easily determine, based on the second identification information, whether it is possible to decode some of the encoded submeshes among the multiple encoded submeshes.

[0020] The encoding method of Example 4 is any of the encoding methods of Examples 1 to 3, wherein the parameter set further includes third identification information indicating whether multiple submesh identification numbers that uniquely identify each of the multiple submeshes are consecutive numbers, and if the multiple submesh identification numbers are not consecutive numbers, the parameter set may further include information indicating the multiple submesh identification numbers.

[0021] For example, submesh identification numbers may be assigned consecutively to each of the multiple submeshes in accordance with the encoding order, or they may be assigned randomly regardless of the encoding order. For example, when the submesh identification numbers are assigned based on a specific condition, such as when the multiple submesh identification numbers are consecutive, if the specific condition is shared with the decoding device that acquired the bitstream, the decoding device can perform processing on a specific submesh according to the submesh identification number, even if the bitstream does not contain information indicating the submesh identification number. On the other hand, when the multiple submesh identification numbers are not consecutive, the decoding device requires information indicating the multiple submesh identification numbers when performing processing on a specific submesh. Therefore, when the multiple submesh identification numbers are not consecutive, the parameter set includes information indicating the multiple submesh identification numbers, allowing the decoding device to perform processing on a specific submesh.

[0022] In addition, when multiple sub-mesh identification numbers are consecutive numbers, the parameter set does not need to include information indicating multiple sub-mesh identification numbers.Accordingly, for example, when the parameter set does not include information indicating multiple sub-mesh identification numbers, the decoding device can determine that multiple sub-mesh identification numbers are consecutive numbers.Therefore, the amount of information in the bitstream can be reduced.

[0023] The encoding method of Example 5 is the encoding method of any of Examples 1 to 4, wherein the parameter sets include a first parameter set related to a base mesh corresponding to the plurality of submeshes and a second parameter set related to displacement vectors for displacing vertices included in the plurality of submeshes, and the first identification information may be included in at least one of the first parameter set and the second parameter set.

[0024] With this, a decoding device that has acquired a bitstream can acquire the first identification information from at least one of the first parameter set and the second parameter set.

[0025] The decoding method of Example 6 obtains from a bitstream a plurality of coded data in which each of a plurality of submeshes is coded, and a parameter set indicating the configuration of the plurality of submeshes, and decodes the plurality of coded data based on the parameter set, wherein the parameter set includes first identification information indicating whether or not each of at least one of the plurality of submeshes depends on other submeshes among the plurality of submeshes when a predetermined condition is satisfied.

[0026] According to this, it is possible to easily determine whether each of the multiple coded sub-meshes can be individually decodable by using the first identification information.

[0027] A decoding method of Example 7 is the decoding method of Example 6, wherein the parameter set may further include information indicating the number of the plurality of sub-meshes.

[0028] This allows for processing that uses a plurality of sub-meshes.

[0029] The decoding method of Example 8 is the decoding method of Example 6 or 7, wherein the parameter set further includes second identification information indicating whether the plurality of submeshes includes a dependent submesh that depends on another submesh among the plurality of submeshes, and the parameter set may include the first identification information if the plurality of submeshes includes the dependent submesh.

[0030] This makes it possible to easily determine whether or not some of the coded submeshes among the coded submeshes can be decoded based on the second identification information.

[0031] The decoding method of Example 9 is any of the decoding methods of Examples 6 to 8, wherein the parameter set further includes third identification information indicating whether multiple submesh identification numbers that uniquely identify each of the multiple submeshes are consecutive numbers, and if the multiple submesh identification numbers are not consecutive numbers, the parameter set may further include information indicating the multiple submesh identification numbers.

[0032] This allows processing to be performed on a specific submesh based on the submesh identification number, even if the submesh identification numbers are not consecutive.

[0033] The decoding method of Example 10 is the decoding method of any of Examples 6 to 9, wherein the parameter sets include a first parameter set related to a base mesh corresponding to the plurality of submeshes and a second parameter set related to displacement vectors for displacing vertices included in the plurality of submeshes, and the first identification information may be included in at least one of the first parameter set and the second parameter set.

[0034] This allows the first identification information to be obtained from at least one of the first parameter set and the second parameter set.

[0035] The encoding device of Example 11 includes a processor and a memory, wherein the processor uses the memory to generate multiple encoded data by encoding multiple submeshes, generate a parameter set indicating the configuration of the multiple submeshes, and generate a bit stream including the multiple encoded data and the parameter set, wherein the parameter set includes first identification information indicating whether at least one of the multiple submeshes depends on other submeshes among the multiple submeshes when a predetermined condition is met.

[0036] This provides the same effect as the encoding method according to Example 1.

[0037] The decoding device of Example 12 includes a processor and a memory, and the processor uses the memory to obtain from a bitstream a plurality of encoded data in which each of a plurality of submeshes is encoded and a parameter set indicating the configuration of the plurality of submeshes, and decodes the plurality of encoded data based on the parameter set, and the parameter set includes first identification information indicating whether or not each of at least one of the plurality of submeshes depends on other submeshes among the plurality of submeshes when a predetermined condition is satisfied.

[0038] This provides the same effect as the decoding method according to Example 6.

[0039] Furthermore, these comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0040] <Expressions and Terms> The following expressions and terms are used herein.

[0041] (1) Three-dimensional Mesh A three-dimensional mesh is a collection of multiple faces, and represents, for example, a three-dimensional object. A three-dimensional mesh is mainly composed of vertex information, connectivity information, and attribute information. A three-dimensional mesh may be expressed as a polygon mesh or a mesh. A three-dimensional mesh may also vary over time. A three-dimensional mesh may include metadata related to the vertex information, connectivity information, and attribute information, and may also include other additional information.

[0042] (2) Vertex Information Vertex information is information indicating a vertex. For example, the vertex information indicates the position of a vertex in a three-dimensional space. Furthermore, a vertex corresponds to a vertex of a face that constitutes a three-dimensional mesh. Vertex information may be expressed as "geometry." Furthermore, vertex information may be expressed as position information.

[0043] (3) Connection Information Connection information is information that indicates connections between vertices. For example, connection information indicates connections for forming faces or edges of a three-dimensional mesh. Connection information may be expressed as "Connectivity." Connection information may also be expressed as face information.

[0044] (4) Attribute Information Attribute information is information that indicates attributes of a vertex or a face. For example, attribute information indicates attributes such as a color, an image, and a normal vector associated with a vertex or a face. Attribute information may be expressed as "texture."

[0045] (5) Faces A face is an element that makes up a three-dimensional mesh. Specifically, a face is a polygon on a plane in three-dimensional space. For example, a face can be defined as a triangle in three-dimensional space.

[0046] (6) Plane A plane is a two-dimensional plane in a three-dimensional space. For example, a polygon is formed on a plane, and multiple polygons are formed on multiple planes.

[0047] (7) Bitstream: A bitstream corresponds to coded information. A bitstream may also be referred to as a stream, a coded bitstream, a compressed bitstream, or a coded signal.

[0048] (8) Encoding and Decoding The term encoding may be substituted with terms such as storing, including, writing, describing, signaling, sending, notifying, saving, or compressing, and these terms may be interchangeable. For example, encoding information may mean including the information in a bitstream. Also, encoding information into a bitstream may mean encoding the information to generate a bitstream that includes the encoded information.

[0049] Additionally, the term "decode" may be replaced with terms such as "read," "decode," "read," "load," "derive," "obtain," "receive," "extract," "reconstruct," "reconstruct," "decompress," or "decompress," and these terms may be interchangeable. For example, decoding information may mean obtaining information from a bitstream. Decoding information from a bitstream may mean decoding the bitstream to obtain information contained in the bitstream.

[0050] (9) Ordinal Numbers In the description, ordinal numbers such as first and second may be assigned to components, etc. These ordinal numbers may be changed as appropriate. Furthermore, new ordinal numbers may be assigned to components, etc., or removed. Furthermore, these ordinal numbers may be assigned to elements in order to identify them, and may not correspond to a meaningful order.

[0051] <Three-dimensional mesh> Fig. 1 is a conceptual diagram showing a three-dimensional mesh according to this embodiment. A three-dimensional mesh is composed of multiple faces. For example, each face is a triangle. The vertices of these triangles are defined in three-dimensional space. The three-dimensional mesh then represents a three-dimensional object. Each face may have a color or an image.

[0052] FIG. 2 is a conceptual diagram showing the basic elements of a three-dimensional mesh according to this embodiment. A three-dimensional mesh is composed of vertex information, connection information, and attribute information. The vertex information indicates the positions of the vertices of a face in three-dimensional space. The connection information indicates the connections between the vertices. A face can be identified by the vertex information and connection information. In other words, a colorless three-dimensional object is formed in three-dimensional space by the vertex information and connection information.

[0053] The attribute information may be associated with a vertex or a face. The attribute information associated with a vertex may be expressed as "Attribute Per Point." The attribute information associated with a vertex may indicate an attribute of the vertex itself, or may indicate an attribute of a face connected to the vertex.

[0054] For example, a color may be associated with a vertex as attribute information. The color associated with a vertex may be the color of the vertex itself, or the color of a face connected to the vertex. The color of a face may be the average of multiple colors associated with multiple vertices of the face. Furthermore, a normal vector may be associated with a vertex or a face as attribute information. Such a normal vector can represent the front and back of a face.

[0055] A two-dimensional image may be associated with a surface as attribute information. The two-dimensional image associated with a surface is also expressed as a texture image or an "Attribute Map." Information indicating mapping between the surface and the two-dimensional image may be associated with the surface as attribute information. Information indicating such mapping may be expressed as mapping information, vertex information of a texture image, texture coordinates, or "Attribute UV Coordinate."

[0056] Furthermore, information such as color, image, and moving image used as attribute information may be expressed as "parametric space."

[0057] The attribute information allows texture to be reflected on the three-dimensional object. That is, a three-dimensional object having color is formed in three-dimensional space based on the vertex information, connection information, and attribute information.

[0058] In the above, the attribute information is associated with the vertices or faces, but it may also be associated with the edges.

[0059] 3 is a conceptual diagram illustrating mapping according to this embodiment. For example, a region of a two-dimensional image on a two-dimensional plane can be mapped onto a surface of a three-dimensional mesh in three-dimensional space. Specifically, coordinate information of the region in the two-dimensional image is associated with the surface of the three-dimensional mesh. As a result, an image of the mapped region in the two-dimensional image is reflected on the surface of the three-dimensional mesh.

[0060] By using the mapping, the 2D image used as attribute information can be separated from the 3D mesh. For example, in encoding the 3D mesh, the 2D image may be encoded by an image encoding method or a video encoding method.

[0061] <System Configuration> Fig. 4 is a block diagram showing an example of the configuration of a coding / decoding system according to this embodiment. In Fig. 4, the coding / decoding system includes a coding device 100 and a decoding device 200.

[0062] For example, the encoding device 100 obtains a three-dimensional mesh and encodes the three-dimensional mesh into a bitstream. Then, the encoding device 100 outputs the bitstream to the network 300. For example, the bitstream includes the encoded three-dimensional mesh and control information for decoding the encoded three-dimensional mesh. By encoding the three-dimensional mesh, information about the three-dimensional mesh is compressed.

[0063] The network 300 transmits a bitstream from the encoding device 100 to the decoding device 200. The network 300 may be the Internet, a wide area network (WAN), a local area network (LAN), or a combination of these. The network 300 is not necessarily limited to bidirectional communication, and may be a unidirectional communication network for terrestrial digital broadcasting, satellite broadcasting, or the like.

[0064] Furthermore, the network 300 can be replaced by a recording medium such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc (registered trademark)).

[0065] The decoding device 200 obtains a bitstream and decodes a three-dimensional mesh from the bitstream. By decoding the three-dimensional mesh, information about the three-dimensional mesh is expanded. For example, the decoding device 200 decodes the three-dimensional mesh according to a decoding method corresponding to the encoding method used by the encoding device 100 to encode the three-dimensional mesh. That is, the encoding device 100 and the decoding device 200 perform encoding and decoding according to encoding methods and decoding methods that correspond to each other.

[0066] The 3D mesh before encoding may also be referred to as an original 3D mesh, and the 3D mesh after decoding may also be referred to as a reconstructed 3D mesh.

[0067] 5 is a block diagram showing an example of the configuration of a coding device 100 according to this embodiment. For example, the coding device 100 includes a vertex information encoder 101, a connection information encoder 102, and an attribute information encoder 103.

[0068] The vertex information encoder 101 is an electrical circuit that encodes vertex information. For example, the vertex information encoder 101 encodes the vertex information into a bitstream according to a format defined for the vertex information.

[0069] The connection information encoder 102 is an electrical circuit that encodes the connection information, for example, the connection information encoder 102 encodes the connection information into a bitstream according to a format defined for the connection information.

[0070] The attribute information encoder 103 is an electric circuit that encodes the attribute information. For example, the attribute information encoder 103 encodes the attribute information into a bit stream in accordance with a format defined for the attribute information.

[0071] The vertex information, connectivity information, and attribute information may be coded using variable-length coding or fixed-length coding, such as Huffman coding or context-adaptive binary arithmetic coding (CABAC).

[0072] The vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be integrated together, or each of the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be further subdivided into multiple components.

[0073] 6 is a block diagram showing another example of the configuration of the encoding device 100 according to this embodiment. For example, the encoding device 100 includes a pre-processor 104 and a post-processor 105 in addition to the configuration shown in FIG.

[0074] The preprocessor 104 is an electrical circuit that performs processing before encoding the vertex information, connectivity information, and attribute information. For example, the preprocessor 104 may perform a conversion process, a separation process, a multiplexing process, or the like on the 3D mesh before encoding. More specifically, for example, the preprocessor 104 may separate the vertex information, connectivity information, and attribute information from the 3D mesh before encoding.

[0075] The post-processor 105 is an electrical circuit that performs processing after the vertex information, connection information, and attribute information are encoded. For example, the post-processor 105 may perform conversion processing, separation processing, multiplexing processing, or the like on the encoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 105 may multiplex the encoded vertex information, connection information, and attribute information into a bitstream. Furthermore, for example, the post-processor 105 may further perform variable-length coding on the encoded vertex information, connection information, and attribute information.

[0076] 7 is a block diagram showing an example of the configuration of a decoding device 200 according to this embodiment. For example, the decoding device 200 includes a vertex information decoder 201, a connection information decoder 202, and an attribute information decoder 203.

[0077] The vertex information decoder 201 is an electrical circuit that decodes vertex information. For example, the vertex information decoder 201 decodes vertex information from a bitstream according to a format defined for the vertex information.

[0078] The connection information decoder 202 is an electrical circuit that decodes the connection information, for example, the connection information decoder 202 decodes the connection information from the bitstream according to a format defined for the connection information.

[0079] The attribute information decoder 203 is an electric circuit that decodes the attribute information. For example, the attribute information decoder 203 decodes the attribute information from the bitstream in accordance with a format defined for the attribute information.

[0080] The vertex information, connection information, and attribute information may be decoded using variable length decoding or fixed length decoding, which may correspond to Huffman coding, context-adaptive binary arithmetic coding (CABAC), or the like.

[0081] The vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be integrated together, or each of the vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be further subdivided into multiple components.

[0082] 8 is a block diagram showing another example of the configuration of the decoding device 200 according to this embodiment. For example, the decoding device 200 includes a pre-processor 204 and a post-processor 205 in addition to the configuration shown in FIG.

[0083] The preprocessor 204 is an electrical circuit that performs processing before decoding the vertex information, connection information, and attribute information. For example, the preprocessor 204 may perform conversion processing, separation processing, multiplexing processing, or the like on the bitstream before decoding the vertex information, connection information, and attribute information.

[0084] More specifically, for example, the preprocessor 204 may separate a sub-bitstream corresponding to vertex information, a sub-bitstream corresponding to connectivity information, and a sub-bitstream corresponding to attribute information from the bitstream. Also, for example, the preprocessor 204 may perform variable-length decoding on the bitstream in advance before decoding the vertex information, connectivity information, and attribute information.

[0085] The post-processor 205 is an electrical circuit that performs processing after the vertex information, connection information, and attribute information are decoded. For example, the post-processor 205 may perform conversion processing, separation processing, multiplexing processing, or the like on the decoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 205 may multiplex the decoded vertex information, connection information, and attribute information onto a three-dimensional mesh.

[0086] <Bitstream> Vertex information, connection information, and attribute information are coded and stored in a bitstream. The relationship between this information and the bitstream is shown below.

[0087] 9 is a conceptual diagram showing an example of the configuration of a bitstream according to this embodiment. In this example, connection information, vertex information, and attribute information are integrated in the bitstream. For example, the connection information, vertex information, and attribute information may be included in a single file.

[0088] Furthermore, multiple portions of this information may be stored sequentially, such as a first portion of connection information, a first portion of vertex information, a first portion of attribute information, a second portion of connection information, a second portion of vertex information, a second portion of attribute information, etc. These multiple portions may correspond to multiple portions that are different in time, multiple portions that are different in space, or multiple different faces.

[0089] Furthermore, the storage order of the connection information, vertex information, and attribute information is not limited to the above example, and a storage order different from the above example may be used.

[0090] 10 is a conceptual diagram showing another example of the configuration of a bitstream according to this embodiment. In this example, a plurality of files are included in the bitstream, and connection information, vertex information, and attribute information are stored in different files. Here, a file containing connection information, a file containing vertex information, and a file containing attribute information are shown, but the storage format is not limited to this example. For example, two types of information among the connection information, vertex information, and attribute information may be included in one file, and the remaining type of information may be included in another file.

[0091] Alternatively, the information may be split and stored in more files. For example, multiple pieces of connectivity information may be stored in multiple files, multiple pieces of vertex information may be stored in multiple files, or multiple pieces of attribute information may be stored in multiple files. These multiple pieces may correspond to multiple temporally different pieces, multiple spatially different pieces, or multiple different faces.

[0092] Furthermore, the storage order of the connection information, vertex information, and attribute information is not limited to the above example, and a storage order different from the above example may be used.

[0093] 11 is a conceptual diagram showing another example of the configuration of a bitstream according to this embodiment. In this example, the bitstream is composed of multiple separable sub-bitstreams, and connection information, vertex information, and attribute information are stored in different sub-bitstreams.

[0094] Here, a sub-bitstream containing connection information, a sub-bitstream containing vertex information, and a sub-bitstream containing attribute information are shown, but the storage format is not limited to this example.

[0095] For example, two types of information among the connection information, vertex information, and attribute information may be included in one sub-bitstream, and the remaining type of information may be included in another sub-bitstream. Specifically, attribute information of a two-dimensional image or the like may be stored in a sub-bitstream that complies with an image coding method, separate from the sub-bitstreams of the connection information and vertex information.

[0096] Each sub-bitstream may also include multiple files, and multiple pieces of connectivity information may be stored in multiple files, multiple pieces of vertex information may be stored in multiple files, or multiple pieces of attribute information may be stored in multiple files.

[0097] 9, 10, and 11, and a storage order different from the above examples may be used. For example, the vertex information, connection information, and attribute information may be stored in the bitstream in this order. Alternatively, the connection information, connection information, and attribute information may be stored in the bitstream in any of the following orders: connection information, attribute information, and vertex information; vertex information, attribute information, and connection information; attribute information, connection information, and vertex information; or attribute information, vertex information, and connection information.

[0098] Furthermore, each of the connection information, vertex information, and attribute information may be divided into a plurality of data, and the plurality of data may be stored in a cyclical or random order within the bitstream.

[0099] 12 is a block diagram showing a specific example of an encoding / decoding system according to this embodiment. In FIG. 12, the encoding / decoding system includes a three-dimensional data encoding system 110, a three-dimensional data decoding system 210, and an external connector 310.

[0100] The three-dimensional data encoding system 110 includes a controller 111, an input / output processor 112, a three-dimensional data encoder 113, a three-dimensional data generator 115, and a system multiplexer 114. The three-dimensional data decoding system 210 includes a controller 211, an input / output processor 212, a three-dimensional data decoder 213, a system demultiplexer 214, a presenter 215, and a user interface 216.

[0101] In the three-dimensional data encoding system 110, sensor data is input from a sensor terminal to a three-dimensional data generator 115. The three-dimensional data generator 115 generates three-dimensional data, such as point cloud data or mesh data, from the sensor data and inputs it to a three-dimensional data encoder 113.

[0102] For example, the three-dimensional data generator 115 generates vertex information, and generates connection information and attribute information corresponding to the vertex information. The three-dimensional data generator 115 may process the vertex information when generating the connection information and attribute information. For example, the three-dimensional data generator 115 may reduce the amount of data by deleting duplicate vertices, or may transform the vertex information (such as by shifting its position, rotating it, or normalizing it). The three-dimensional data generator 115 may also render the attribute information.

[0103] Furthermore, although the three-dimensional data generator 115 is a component of the three-dimensional data encoding system 110 in FIG. 12, it may be arranged externally and independently of the three-dimensional data encoding system 110.

[0104] The sensor terminal that provides the sensor data for generating the three-dimensional data may be, for example, a moving body such as an automobile, a flying object such as an airplane, a mobile terminal, a camera, etc. Furthermore, a distance sensor such as a LIDAR, a millimeter wave radar, an infrared sensor, or a range finder, a stereo camera, or a combination of multiple monocular cameras may also be used as the sensor terminal.

[0105] The sensor data may be the distance (position) of the object, monocular camera images, stereo camera images, color, reflectance, sensor attitude, orientation, gyro, sensing position (GPS information or altitude), speed, acceleration, sensing time, temperature, air pressure, humidity, or magnetism.

[0106] The three-dimensional data encoder 113 corresponds to the encoding device 100 shown in FIG. 5 and other figures. For example, the three-dimensional data encoder 113 encodes three-dimensional data to generate encoded data. The three-dimensional data encoder 113 also generates control information when encoding the three-dimensional data. The three-dimensional data encoder 113 then inputs the encoded data together with the control information to the system multiplexer 114.

[0107] The encoding method for the three-dimensional data may be an encoding method using geometry or an encoding method using a video codec. Here, the encoding method using geometry may also be referred to as a geometry-based encoding method. The encoding method using a video codec may also be referred to as a video-based encoding method.

[0108] The system multiplexer 114 multiplexes the encoded data and control information input from the 3D data encoder 113 to generate multiplexed data using a specified multiplexing method. The system multiplexer 114 may multiplex other media such as video, audio, subtitles, application data, or document files, or reference time information, along with the encoded data and control information of the 3D data. Furthermore, the system multiplexer 114 may multiplex attribute information related to the sensor data or the 3D data.

[0109] For example, the multiplexed data may have a file format for storage or a packet format for transmission. As these formats, ISOBMFF or a format based on ISOBMFF may be used. Also, MPEG-DASH, MMT, MPEG-2 TS Systems, RTP, or the like may be used.

[0110] The multiplexed data is then output as a transmission signal to the external connector 310 by the input / output processor 112. The multiplexed data may be transmitted as a transmission signal by wire or wirelessly. Alternatively, the multiplexed data is stored in an internal memory or a storage device. The multiplexed data may be transmitted to a cloud server via the Internet or may be stored in an external storage device.

[0111] For example, the transmission or storage of the multiplexed data is performed by a method according to the medium for transmission or storage, such as broadcasting or communication. The communication protocol may be http, ftp, TCP, UDP, IP, or a combination thereof. Furthermore, a pull-type communication method or a push-type communication method may be used.

[0112] For wired transmission, Ethernet (registered trademark), USB, RS-232C, HDMI (registered trademark), coaxial cable, etc. may be used. For wireless transmission, 3GPP (registered trademark), 3G / 4G / 5G defined by IEEE, wireless LAN, Wi-Fi, Bluetooth, or millimeter wave may be used. For broadcasting, for example, DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 may be used.

[0113] The sensor data may be input to the three-dimensional data generator 115 or the system multiplexer 114. The three-dimensional data or encoded data may be output as a transmission signal directly to the external connector 310 via the input / output processor 112. The transmission signal output from the three-dimensional data encoding system 110 is input to the three-dimensional data decoding system 210 via the external connector 310.

[0114] Furthermore, each operation of the three-dimensional data encoding system 110 may be controlled by a controller 111 that executes an application program.

[0115] In the three-dimensional data decoding system 210, a transmission signal is input to an input / output processor 212. The input / output processor 212 decodes multiplexed data having a file format or a packet format from the transmission signal and inputs the multiplexed data to a system demultiplexer 214. The system demultiplexer 214 obtains coded data and control information from the multiplexed data and inputs them to a three-dimensional data decoder 213. The system demultiplexer 214 may extract other media or reference time information from the multiplexed data.

[0116] The three-dimensional data decoder 213 corresponds to the decoding device 200 shown in Fig. 7 etc. For example, the three-dimensional data decoder 213 decodes three-dimensional data from the encoded data based on a predefined encoding method. The three-dimensional data is then presented to the user by the presenter 215.

[0117] Additionally, additional information such as sensor data may be input to the presenter 215. The presenter 215 may present three-dimensional data based on the additional information. Additionally, a user instruction may be input from a user terminal to the user interface 216. Then, the presenter 215 may present three-dimensional data based on the input instruction.

[0118] The input / output processor 212 may acquire the three-dimensional data and the encoded data from the external connector 310 .

[0119] Furthermore, each operation of the three-dimensional data decoding system 210 may be controlled by a controller 211 that executes an application program.

[0120] 13 is a conceptual diagram showing an example of the configuration of point cloud data according to this embodiment. The point cloud data is data of a group of points representing a three-dimensional object.

[0121] Specifically, a point cloud is made up of a plurality of points, and has position information indicating the three-dimensional coordinate position of each point and attribute information indicating the attribute of each point. The position information is also expressed as geometry.

[0122] The type of attribute information may be, for example, color, reflectance, etc. One point may be associated with attribute information of one type, one point may be associated with attribute information of multiple different types, or one point may be associated with attribute information having multiple values ​​for the same type.

[0123] 14 is a conceptual diagram showing an example of a data file of point cloud data according to this embodiment. This example shows a case where there is a one-to-one correspondence between position information items and attribute information items, and shows position information and attribute information for N points that make up the point cloud data. In this example, the position information is information indicating a three-dimensional coordinate position using three axes, x, y, and z, and the attribute information is information indicating a color using RGB. A PLY file or the like can be used as a representative data file for point cloud data.

[0124] 15 is a conceptual diagram showing an example of the configuration of mesh data according to this embodiment. Mesh data is data used in CG (Computer Graphics) and the like, and is three-dimensional mesh data that shows the three-dimensional shape of an object using multiple surfaces. Each surface is also expressed as a polygon, and has a polygonal shape such as a triangle or a rectangle.

[0125] Specifically, a 3D mesh is composed of a plurality of points constituting a point cloud, as well as a plurality of edges and a plurality of faces. Each point is also expressed as a vertex or a position. Each edge corresponds to a line segment connected by two vertices. Each face corresponds to an area surrounded by three or more edges.

[0126] Furthermore, a three-dimensional mesh has position information indicating the three-dimensional coordinate positions of vertices. The position information is also expressed as vertex information or geometry. A three-dimensional mesh also has connection information indicating the relationship between multiple vertices that make up an edge or a face. The connection information is also expressed as connectivity. A three-dimensional mesh also has attribute information indicating the attributes of the vertices, edges, or faces. The attribute information in a three-dimensional mesh is also expressed as texture.

[0127] For example, the attribute information may indicate the color, reflectance, or normal vector for a vertex, edge, or face. The direction of the normal vector may represent the front and back of the face.

[0128] The mesh data may be stored in a data file format such as an object file.

[0129] 16 is a conceptual diagram showing an example of a data file of mesh data according to this embodiment. In this example, the data file includes position information G(1) to G(N) of N vertices that make up the three-dimensional mesh, and attribute information A1(1) to A1(N) of the N vertices. Also, in this example, M pieces of attribute information A2(1) to A2(M) are included. The attribute information items do not need to correspond one-to-one to vertices or faces. Furthermore, attribute information need not exist.

[0130] The connection information is represented by a combination of vertex indices. n[1, 3, 4] indicates a triangular face formed by three vertices, n=1, n=3, and n=4. Also, m[2, 4, 6] indicates that the attribute information of m=2, m=4, and m=6 corresponds to the three vertices, respectively.

[0131] Furthermore, the actual contents of the attribute information may be written in a separate file. A pointer to that content may be associated with a vertex, a face, or the like. For example, attribute information indicating an image for a face may be stored in a two-dimensional attribute map file. The file name of the attribute map and two-dimensional coordinate values ​​in the attribute map may be written in attribute information A2(1) to A2(M). The method of specifying attribute information for a face is not limited to these methods, and any method may be used.

[0132] 17 is a conceptual diagram showing types of three-dimensional data according to this embodiment. Point cloud data and mesh data may represent static objects or dynamic objects. A static object is an object that does not change over time, and a dynamic object is an object that changes over time. A static object may correspond to three-dimensional data for any point in time.

[0133] For example, point cloud data for a given point in time may be referred to as a PCC frame, mesh data for a given point in time may be referred to as a mesh frame, and PCC frames and mesh frames may be simply referred to as frames.

[0134] The area of ​​the object may be limited to a certain range, as in normal video data, or may not be limited, as in map data. The density of points or surfaces may be determined in various ways. Sparse point cloud data or sparse mesh data may be used, or dense point cloud data or dense mesh data may be used.

[0135] Next, encoding and decoding of a point cloud or a three-dimensional mesh will be described. The device, process, or syntax for encoding and decoding vertex information of a three-dimensional mesh in the present disclosure may be applied to encoding and decoding of a point cloud. The device, process, or syntax for encoding and decoding of a point cloud in the present disclosure may be applied to encoding and decoding vertex information of a three-dimensional mesh.

[0136] Furthermore, a device, process, or syntax for encoding and decoding attribute information of a point cloud in the present disclosure may be applied to encoding and decoding connectivity information or attribute information of a three-dimensional mesh.Furthermore, a device, process, or syntax for encoding and decoding connectivity information or attribute information of a three-dimensional mesh in the present disclosure may be applied to encoding and decoding attribute information of a point cloud.

[0137] Furthermore, at least some of the processing may be shared between the encoding and decoding of point cloud data and the encoding and decoding of mesh data, thereby reducing the scale of the circuit and software program.

[0138] 18 is a block diagram showing an example configuration of a three-dimensional data encoder 113 according to this embodiment. In this example, the three-dimensional data encoder 113 includes a vertex information encoder 121, an attribute information encoder 122, a metadata encoder 123, and a multiplexer 124. The vertex information encoder 121, the attribute information encoder 122, and the multiplexer 124 may correspond to the vertex information encoder 101, the attribute information encoder 103, the post-processor 105, etc. in FIG.

[0139] In this example, the three-dimensional data encoder 113 encodes the three-dimensional data according to a geometry-based encoding method, which takes into account the three-dimensional structure. In addition, in the geometry-based encoding method, attribute information is encoded using configuration information obtained in encoding the vertex information.

[0140] Specifically, first, vertex information, attribute information, and metadata included in three-dimensional data generated from sensor data are input to a vertex information encoder 121, an attribute information encoder 122, and a metadata encoder 123, respectively. Here, connectivity information included in the three-dimensional data may be treated in the same way as attribute information. In addition, in the case of point cloud data, position information may be treated as vertex information.

[0141] The vertex information encoder 121 encodes the vertex information into compressed vertex information and outputs the compressed vertex information as encoded data to the multiplexer 124. The vertex information encoder 121 also generates metadata for the compressed vertex information and outputs it to the multiplexer 124. The vertex information encoder 121 also generates configuration information and outputs it to the attribute information encoder 122.

[0142] The attribute information encoder 122 uses the configuration information generated by the vertex information encoder 121 to encode the attribute information into compressed attribute information and outputs the compressed attribute information as encoded data to the multiplexer 124. The attribute information encoder 122 also generates metadata of the compressed attribute information and outputs it to the multiplexer 124.

[0143] The metadata encoder 123 encodes compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The metadata encoded by the metadata encoder 123 may be used to encode vertex information and attribute information.

[0144] The multiplexer 124 multiplexes the compressed vertex information, the compressed vertex information metadata, the compressed attribute information, the compressed attribute information metadata, and the compressed metadata into a bitstream, and then inputs the bitstream to the system layer.

[0145] 19 is a block diagram showing an example configuration of a three-dimensional data decoder 213 according to this embodiment. In this example, the three-dimensional data decoder 213 includes a vertex information decoder 221, an attribute information decoder 222, a metadata decoder 223, and a demultiplexer 224. The vertex information decoder 221, the attribute information decoder 222, and the demultiplexer 224 may correspond to the vertex information decoder 201, the attribute information decoder 203, the preprocessor 204, and the like in FIG.

[0146] In this example, the three-dimensional data decoder 213 decodes three-dimensional data according to a geometry-based encoding method. The three-dimensional structure is taken into consideration in the decoding according to the geometry-based encoding method. Furthermore, in the decoding according to the geometry-based encoding method, attribute information is decoded using configuration information obtained in decoding vertex information.

[0147] Specifically, first, a bitstream is input from the system layer to a demultiplexer 224. The demultiplexer 224 separates compressed vertex information, compressed vertex information metadata, compressed attribute information, compressed attribute information metadata, and compressed metadata from the bitstream. The compressed vertex information and compressed vertex information metadata are input to a vertex information decoder 221. The compressed attribute information and compressed attribute information metadata are input to an attribute information decoder 222. The metadata is input to a metadata decoder 223.

[0148] The vertex information decoder 221 decodes vertex information from the compressed vertex information using metadata of the compressed vertex information. The vertex information decoder 221 also generates configuration information and outputs it to the attribute information decoder 222. The attribute information decoder 222 decodes attribute information from the compressed attribute information using the configuration information generated by the vertex information decoder 221 and the metadata of the compressed attribute information. The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 may be used to decode the vertex information and the attribute information.

[0149] Thereafter, the vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. Note that, for example, this metadata is metadata of the vertex information and attribute information, and can be used in an application program.

[0150] 20 is a block diagram showing another example configuration of the three-dimensional data encoder 113 according to the present embodiment. In this example, the three-dimensional data encoder 113 includes a vertex image generator 131, an attribute image generator 132, a metadata generator 133, a video encoder 134, a metadata encoder 123, and a multiplexer 124. The vertex image generator 131, the attribute image generator 132, and the video encoder 134 may correspond to the vertex information encoder 101 and the attribute information encoder 103 in FIG. 6 , etc.

[0151] In this example, the 3D data encoder 113 encodes the 3D data according to a video-based encoding method. In encoding according to the video-based encoding method, multiple 2D images are generated from the 3D data, and the multiple 2D images are encoded according to a video encoding method. Here, the video encoding method may be High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or the like.

[0152] Specifically, first, vertex information and attribute information included in three-dimensional data generated from sensor data are input to a metadata generator 133. The vertex information and attribute information are then input to a vertex image generator 131 and an attribute image generator 132, respectively. The metadata included in the three-dimensional data is then input to a metadata encoder 123. Here, connectivity information included in the three-dimensional data may be treated in the same way as attribute information. In the case of point cloud data, position information may be treated as vertex information.

[0153] The metadata generator 133 generates map information of a plurality of two-dimensional images from the vertex information and attribute information, and inputs the map information to the vertex image generator 131, the attribute image generator 132, and the metadata encoder 123.

[0154] The vertex image generator 131 generates a vertex image based on the vertex information and map information, and inputs the generated image to the video encoder 134. The attribute image generator 132 generates an attribute image based on the attribute information and map information, and inputs the generated image to the video encoder 134.

[0155] The video encoder 134 encodes the vertex images and attribute images into compressed vertex information and compressed attribute information, respectively, in accordance with a video encoding method, and outputs the compressed vertex information and compressed attribute information as encoded data to the multiplexer 124. The video encoder 134 also generates metadata for the compressed vertex information and metadata for the compressed attribute information, and outputs them to the multiplexer 124.

[0156] The metadata encoder 123 encodes the compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The compressible metadata includes map information. The metadata encoded by the metadata encoder 123 may also be used to encode vertex information and attribute information.

[0157] The multiplexer 124 multiplexes the compressed vertex information, the compressed vertex information metadata, the compressed attribute information, the compressed attribute information metadata, and the compressed metadata into a bitstream, and then inputs the bitstream to the system layer.

[0158] 21 is a block diagram showing another example configuration of the 3D data decoder 213 according to this embodiment. In this example, the 3D data decoder 213 includes a vertex information generator 231, an attribute information generator 232, a video decoder 234, a metadata decoder 223, and a demultiplexer 224. The vertex information generator 231, the attribute information generator 232, and the video decoder 234 may correspond to the vertex information decoder 201 and the attribute information decoder 203 in FIG. 8, etc.

[0159] In this example, the 3D data decoder 213 decodes the 3D data according to a video-based coding method. In the decoding according to the video-based coding method, a plurality of 2D images are decoded according to a video coding method, and 3D data is generated from the plurality of 2D images. Here, the video coding method may be High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or the like.

[0160] Specifically, first, a bitstream is input from the system layer to the demultiplexer 224. The demultiplexer 224 separates compressed vertex information, compressed vertex information metadata, compressed attribute information, compressed attribute information metadata, and compressed metadata from the bitstream. The compressed vertex information, compressed vertex information metadata, compressed attribute information, and compressed attribute information metadata are input to the video decoder 234. The compressed metadata is input to the metadata decoder 223.

[0161] The video decoder 234 decodes the vertex images in accordance with the video encoding method. At this time, the video decoder 234 decodes the vertex images from the compressed vertex information using the metadata of the compressed vertex information. Then, the video decoder 234 inputs the vertex images to the vertex information generator 231. The video decoder 234 also decodes the attribute images in accordance with the video encoding method. At this time, the video decoder 234 decodes the attribute images from the compressed attribute information using the metadata of the compressed attribute information. Then, the video decoder 234 inputs the attribute images to the attribute information generator 232.

[0162] The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 includes map information used to generate vertex information and attribute information. The metadata decoded by the metadata decoder 223 may also be used to decode vertex images and attribute images.

[0163] The vertex information generator 231 reproduces vertex information from the vertex image in accordance with the map information included in the metadata decoded by the metadata decoder 223. The attribute information generator 232 reproduces attribute information from the attribute image in accordance with the map information included in the metadata decoded by the metadata decoder 223.

[0164] Thereafter, the vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. Note that, for example, this metadata is metadata of the vertex information and attribute information, and can be used in an application program.

[0165] Fig. 22 is a conceptual diagram showing a specific example of encoding processing according to this embodiment. Fig. 22 shows a three-dimensional data encoder 113 and a description encoder 148. In this example, the three-dimensional data encoder 113 includes a two-dimensional data encoder 141 and a mesh data encoder 142. The two-dimensional data encoder 141 includes a texture encoder 143. The mesh data encoder 142 includes a vertex information encoder 144 and a connection information encoder 145.

[0166] The vertex information encoder 144, the connection information encoder 145, and the texture encoder 143 may correspond to the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 in FIG.

[0167] For example, the two-dimensional data encoder 141 operates as a texture encoder 143 and generates a texture file by encoding the texture corresponding to the attribute information as two-dimensional data according to an image encoding method or a video encoding method.

[0168] The mesh data encoder 142 also operates as a vertex information encoder 144 and a connectivity information encoder 145, and generates a mesh file by encoding the vertex information and connectivity information. The mesh data encoder 142 may further encode mapping information for textures. The encoded mapping information may then be included in the mesh file.

[0169] The description encoder 148 also generates a description file by encoding a description corresponding to metadata such as text data. The description encoder 148 may encode the description at the system layer. For example, the description encoder 148 may be included in the system multiplexer 114 of FIG. 12 .

[0170] The above operations generate a bitstream containing texture files, mesh files, and description files, which may be multiplexed into the bitstream in file formats such as glTF (Graphics Language Transmission Format) or USD (Universal Scene Description).

[0171] The three-dimensional data encoder 113 may include two mesh data encoders as the mesh data encoder 142. For example, one mesh data encoder encodes vertex information and connectivity information of a static three-dimensional mesh, and the other mesh data encoder encodes vertex information and connectivity information of a dynamic three-dimensional mesh.

[0172] Correspondingly, two mesh files may then be included in the bitstream: for example, one mesh file corresponding to a static 3D mesh and another mesh file corresponding to a dynamic 3D mesh.

[0173] Furthermore, the static three-dimensional mesh may be a three-dimensional mesh of an intraframe coded using intraprediction, and the dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe coded using interprediction. Furthermore, information on the dynamic three-dimensional mesh may be differential information between vertex information or connectivity information of the three-dimensional mesh of an intraframe and vertex information or connectivity information of the three-dimensional mesh of an interframe.

[0174] Fig. 23 is a conceptual diagram showing a specific example of the decoding process according to this embodiment. Fig. 23 shows a three-dimensional data decoder 213, a description decoder 248, and a renderer 247. In this example, the three-dimensional data decoder 213 includes a two-dimensional data decoder 241, a mesh data decoder 242, and a mesh reconstructor 246. The two-dimensional data decoder 241 includes a texture decoder 243. The mesh data decoder 242 includes a vertex information decoder 244 and a connectivity information decoder 245.

[0175] The vertex information decoder 244, the connection information decoder 245, the texture decoder 243, and the mesh reconstructor 246 may correspond to the vertex information decoder 201, the connection information decoder 202, the attribute information decoder 203, and the post-processor 205 in Fig. 8. The presenter 247 may correspond to the presenter 215 in Fig. 12.

[0176] For example, the two-dimensional data decoder 241 operates as a texture decoder 243, and decodes the texture corresponding to the attribute information from the texture file as two-dimensional data in accordance with an image coding method or a video coding method.

[0177] The mesh data decoder 242 also operates as a vertex information decoder 244 and a connectivity information decoder 245 to decode vertex information and connectivity information from the mesh file. The mesh data decoder 242 may further decode mapping information for textures from the mesh file.

[0178] The description decoder 248 also decodes descriptions corresponding to metadata such as text data from the description file. The description decoder 248 may decode the descriptions at the system layer. For example, the description decoder 248 may be included in the system demultiplexer 214 of FIG. 12 .

[0179] The mesh reconstructor 246 reconstructs a 3D mesh from the vertex information, connectivity information, and textures according to the description. The renderer 247 renders and outputs the 3D mesh according to the description.

[0180] Through the above operations, a 3D mesh is reconstructed and output from a bitstream containing a texture file, a mesh file, and a description file.

[0181] The three-dimensional data decoder 213 may include two mesh data decoders as the mesh data decoder 242. For example, one mesh data decoder decodes vertex information and connectivity information of a static three-dimensional mesh, and the other mesh data decoder decodes vertex information and connectivity information of a dynamic three-dimensional mesh.

[0182] Correspondingly, two mesh files may then be included in the bitstream: for example, one mesh file corresponding to a static 3D mesh and another mesh file corresponding to a dynamic 3D mesh.

[0183] Furthermore, the static three-dimensional mesh may be a three-dimensional mesh of an intraframe coded using intraprediction, and the dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe coded using interprediction. Furthermore, information on the dynamic three-dimensional mesh may be differential information between vertex information or connectivity information of the three-dimensional mesh of an intraframe and vertex information or connectivity information of the three-dimensional mesh of an interframe.

[0184] A dynamic 3D mesh coding method is sometimes called DMC (Dynamic Mesh Coding), and a video-based dynamic 3D mesh coding method is sometimes called V-DMC (Video-based Dynamic Mesh Coding).

[0185] The point cloud encoding method is sometimes called PCC (Point Cloud Compression). The point cloud video-based encoding method is sometimes called V-PCC (Video-based Point Cloud Compression). The point cloud geometry-based encoding method is sometimes called G-PCC (Geometry-based Point Cloud Compression).

[0186] <Implementation Example> Fig. 24 is a block diagram showing an implementation example of the encoding device 100 according to this embodiment. The encoding device 100 includes a circuit 151 and a memory 152. For example, multiple components of the encoding device 100 shown in Fig. 5 etc. are implemented by the circuit 151 and memory 152 shown in Fig. 24.

[0187] The circuit 151 is a circuit that performs information processing and is a circuit that can access the memory 152. For example, the circuit 151 is a dedicated or general-purpose electric circuit that encodes a three-dimensional mesh. The circuit 151 may be a processor such as a CPU. Alternatively, the circuit 151 may be a collection of multiple electric circuits.

[0188] The memory 152 is a dedicated or general-purpose memory that stores information used by the circuit 151 to encode the three-dimensional mesh. The memory 152 may be an electric circuit and may be connected to the circuit 151. The memory 152 may also be included in the circuit 151. The memory 152 may also be a collection of multiple electric circuits. The memory 152 may also be a magnetic disk, an optical disk, or the like, and may also be expressed as a storage, a recording medium, or the like. The memory 152 may also be a non-volatile memory or a volatile memory.

[0189] For example, the memory 152 may store a three-dimensional mesh or a bitstream, or may store a program for the circuit 151 to encode the three-dimensional mesh.

[0190] Note that the encoding device 100 does not necessarily have to implement all of the components shown in Figure 5 and the like, and does not necessarily have to perform all of the processes shown here. Some of the components shown in Figure 5 and the like may be included in another device, and some of the processes shown here may be executed by another device. Furthermore, the encoding device 100 may implement any combination of the components of the present disclosure, and may perform any combination of the processes of the present disclosure.

[0191] Fig. 25 is a block diagram showing an example implementation of a decoding device 200 according to this embodiment. The decoding device 200 includes a circuit 251 and a memory 252. For example, multiple components of the decoding device 200 shown in Fig. 7 and other figures are implemented by the circuit 251 and memory 252 shown in Fig. 25.

[0192] The circuit 251 is a circuit that performs information processing and is a circuit that can access the memory 252. For example, the circuit 251 is a dedicated or general-purpose electric circuit that decodes a three-dimensional mesh. The circuit 251 may be a processor such as a CPU. Alternatively, the circuit 251 may be a collection of multiple electric circuits.

[0193] The memory 252 is a dedicated or general-purpose memory that stores information for the circuit 251 to decode the 3D mesh. The memory 252 may be an electric circuit and may be connected to the circuit 251. The memory 252 may also be included in the circuit 251. The memory 252 may also be a collection of multiple electric circuits. The memory 252 may also be a magnetic disk, an optical disk, or the like, and may also be expressed as a storage, a recording medium, or the like. The memory 252 may also be a non-volatile memory or a volatile memory.

[0194] For example, the memory 252 may store a three-dimensional mesh or a bitstream, or may store a program for the circuit 251 to decode the three-dimensional mesh.

[0195] Note that the decoding device 200 does not necessarily have to implement all of the components shown in Figure 7 and the like, and does not necessarily have to perform all of the processes shown here. Some of the components shown in Figure 7 and the like may be included in another device, and some of the processes shown here may be executed by another device. Furthermore, the decoding device 200 may implement any combination of the components of the present disclosure, and may perform any combination of the processes of the present disclosure.

[0196] The encoding method and the decoding method including the steps performed by each component of the encoding device 100 and the decoding device 200 of the present disclosure may be executed by any device or system. For example, part or all of the encoding method and the decoding method may be executed by a computer including a processor, a memory, an input / output circuit, etc. In this case, the encoding method and the decoding method may be executed by the computer executing a program for causing the computer to execute the encoding method and the decoding method.

[0197] Alternatively, the program or the bitstream may be recorded on a non-transitory computer-readable recording medium such as a CD-ROM.

[0198] An example of a program may be a bitstream. For example, a bitstream including an encoded three-dimensional mesh includes syntax elements for causing the decoding device 200 to decode the three-dimensional mesh. The bitstream then causes the decoding device 200 to decode the three-dimensional mesh according to the syntax elements included in the bitstream. Thus, the bitstream may play a role similar to that of a program.

[0199] The bitstream may be an encoded bitstream containing the encoded 3D mesh, or may be a multiplexed bitstream containing the encoded 3D mesh and other information.

[0200] Furthermore, each component of the encoding device 100 and the decoding device 200 may be configured with dedicated hardware, general-purpose hardware that executes the above-mentioned programs, or a combination of these. The general-purpose hardware may be configured with a memory in which the programs are recorded and a general-purpose processor that reads and executes the programs from the memory. Here, the memory may be a semiconductor memory or a hard disk, and the general-purpose processor may be a CPU.

[0201] Furthermore, the dedicated hardware may be configured with a memory, a dedicated processor, etc. For example, the dedicated processor may execute the encoding method and the decoding method by referring to a memory for recording data.

[0202] Furthermore, as described above, each component of the encoding device 100 and the decoding device 200 may be an electric circuit. These electric circuits may form a single electric circuit as a whole, or each may be a separate electric circuit. Furthermore, these electric circuits may correspond to dedicated hardware, or may correspond to general-purpose hardware that executes the above-mentioned programs, etc. Furthermore, the encoding device 100 and the decoding device 200 may be implemented as an integrated circuit.

[0203] Furthermore, the encoding device 100 may be a transmitting device that transmits the three-dimensional mesh, and the decoding device 200 may be a receiving device that receives the three-dimensional mesh.

[0204] Displacement Encoding and Decoding The following terminology is used here by way of example:

[0205] (1) Image An image is a data unit made up of a set of pixels, and includes a picture or a block smaller than a picture. Images include both moving images and still images.

[0206] (2) Picture A picture is a unit of image processing that is made up of a set of pixels, and is also called a frame or field.

[0207] (3) Block A block is a processing unit consisting of a specific number of pixels. The term "block" shown in the following example is also used. The shape of a block is not particularly limited. A block may be, for example, a rectangular shape of M x N pixels or a square shape of M x M pixels. A block may also be a triangular shape, a circular shape, or another shape. Examples of blocks are as follows:

[0208] Slice, tile, or brick CTU, superblock, or basic division unit VPDU, processing division unit for hardware CU, processing block unit, prediction block unit (PU), or orthogonal transform block unit (TU) Sub-block

[0209] (4) Pixel or Sample A pixel or sample is the smallest point of an image, in other words, the smallest unit. Pixels or samples include not only pixels at integer positions, but also pixels at sub-pixel positions generated based on pixels at integer positions.

[0210] (5) Pixel Value or Sample Value: A pixel value or sample value is a unique value of a pixel. The pixel value or sample value may include a luma value, a chroma value, or an RGB gradation level, and may also include a depth value or a binary value of 0 or 1.

[0211] (6) Flags A flag indicates one or more bits. A flag is, for example, a parameter or index represented by two or more bits. A flag may indicate not only a value represented by a binary number, but also a value represented by a number other than a binary number.

[0212] (7) Signal: A signal is something that is symbolized or coded to transmit information. A signal includes a discrete digital signal or a continuous analog signal.

[0213] (8) Stream or Bit Stream A stream or bit stream is a digital data sequence that indicates the flow of digital data. A stream or bit stream may be a single stream, or may be configured to include multiple streams with multiple layers. A stream or bit stream may be transmitted by serial communication using a single transmission path, or may be transmitted by packet communication using multiple transmission paths.

[0214] (9) Difference: For scalar quantities, the difference can include simple difference (x-y) and difference calculations. The difference can include absolute difference (|x-y|), squared difference (x^2-y^2), square root difference (√(x-y)), weighted difference (ax-by, where a and b are constants), or offset difference (x-y+a, where a is an offset).

[0215] (10) Sum: For scalar quantities, sums can include simple sum (x + y) and addition calculations. The sum can also include absolute sum (|x + y|), sum of squares (x^2 + y^2), square root of the sum (√(x + y)), weighted sum (ax + by, where a and b are constants), or offset sum (x + y + a, where a is an offset).

[0216] (11) "Based on" The expression "based on something" means that something other than that "something" may be taken into consideration. Also, "based on" can be used both when a direct result is obtained and when a result is obtained through an intermediate result.

[0217] (12) "Used" or "Using" The phrases "something was used" or "used something" mean that something other than the "something" may be taken into consideration. The phrases "used" or "used" may be used both in cases where a direct result is obtained and in cases where a result is obtained via an intermediate result.

[0218] (13) Prohibition "Prohibit" can be rephrased as "not permitted." Also, "not prohibited / prohibited" or "permitted / permitted" does not necessarily mean "obligation."

[0219] (14) "Restriction" or "Limitation" "Restriction" or "Limitation" can be rephrased as "not permitted / not allowed" or "not permitted / permitted." Furthermore, "prohibited / not prohibited" or "not permitted / permitted" does not necessarily mean "obligation." Furthermore, what is prohibited quantitatively or qualitatively may be either partial or total.

[0220] (15) Chroma The term chroma is an adjective, represented by the symbols Cb or Cr, that indicates that a sample array or a single sample represents one of the two color difference signals associated with a primary color. The term chroma is sometimes used instead of the term chrominance.

[0221] (16) Luma The term luma is an adjective, denoted by the symbols or subscripts Y or L, that indicates that a sample array or a single sample represents a monochrome signal for a primary color. The term luma is sometimes used instead of the term luminance.

[0222] The encoding / decoding system of this embodiment will be described below.

[0223] A typical three-dimensional model (also called a 3D model) digitally represents an object so that a user can explore the model using zoom, pan, and rotation in all three dimensions while it is rendered over time. One way to construct such a representation is to build a 3D mesh using triangles. The model stores the positions of the triangle vertices, their connectivity to each other, and their associated attributes (such as normals or UV patches).

[0224] Storing all this information in uncompressed form requires a very large storage space and therefore a very large bandwidth for transmission. The triangles that form the mesh often have repeating patterns and similar properties, especially in temporal and spatial neighborhoods. These repetitions can be exploited to develop efficient encoding and decoding methods for storage and transmission. One such encoding and decoding method is Video-based Dynamic Mesh Coding (V-DMC).

[0225] 26 is a block diagram showing another example of the configuration of the encoding / decoding system according to this embodiment. As shown in FIG. 26, the encoding / decoding system includes an encoding device 100 and a decoding device 200.

[0226] The encoding / decoding system accepts input three-dimensional meshes (also called 3D meshes) in the form of three-dimensional coordinates of vertices (vertex information), connectivity (connection information) and associated attributes (attribute information), which may include texture maps as well as geometry.

[0227] The encoding device 100 takes an input 3D mesh (also referred to as an input 3D mesh or input mesh) in the form of 3D coordinates of vertices, connectivity, and associated attributes. The encoding device 100 encodes all associated information into a stream. The stream may consist of a single bitstream or multiple bitstreams.

[0228] The network 300 transmits the stream generated by the encoding device 100 to the decoding device 200. The network 300 may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof. Furthermore, the network 300 is not necessarily limited to a two-way communication network, but may also be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Instead of the network 300, a recording medium such as a digital versatile disc (DVD) or a blue-ray disc (BD) on which a stream is recorded may be used.

[0229] The stream is transmitted to a decoding device 200 via a network 300. The decoding device 200 decodes the bitstream and generates a 3D mesh using the 3D coordinates, connectivity, and associated attributes of the decoded vertices. The decoding device 200 outputs the generated 3D mesh (also referred to as an output 3D mesh or output mesh).

[0230] FIG. 27 is a diagram showing another example of the configuration of the encoding device 100.

[0231] As shown in FIG. 27, the encoding device 100 includes a preprocessor 1103 and a compressor 1106 .

[0232] The encoding device 100 reads an input mesh 1101 and an attribute map 1102 and passes them to a preprocessor 1103. The preprocessor 1103 processes the input mesh to extract a base mesh 1104 and displacement data 1105. The attribute map 1102, along with the extracted base mesh 1104 and displacement data 1105, are passed to a compressor 1106.

[0233] The compressor 1106 also compresses the base mesh 1104, the displacement data 1105, and the attribute map 1102 to generate a bitstream 1107. The compressor 1106 can transmit additional information to the decoding device 200 by further including metadata 1108 in the bitstream 1107.

[0234] FIG. 28 is a diagram showing another example of the configuration of the decoding device 200.

[0235] As shown in FIG. 28, the decoding device 200 includes a decompressor 2102 and a post-processor 2106 .

[0236] The decoding device 200 reads a bitstream 2101 and passes it to a decompressor 2102. The decompressor 2102 decompresses a base mesh 2103, displacement data 2104, and an attribute map 2108 from the bitstream 2101 and passes them to a post-processor 2106. An example of the displacement data 2104 is a displacement vector.

[0237] The post-processor 2106 also processes the base mesh 2103 according to the displacement data 2104 and the attribute map 2108 to generate an output mesh 2107. The post-processor 2106 may further use information from the metadata 2105 to generate the output mesh 2107.

[0238] The detailed configuration of the encoding device 100 will be described below.

[0239] FIG. 29 is a block diagram showing a detailed configuration example of the encoding device 100.

[0240] As shown in FIG. 29, the encoding device 100 includes a decimator 1201, a quantizer 1202, a base mesh encoder 1203, a base mesh decoder 1204, an inverse quantizer 1205, a subdivision unit 1206, a displacement vector calculator 1207, a wavelet transformer 1208, a quantizer 1209, an image packer 1210, a video encoder 1211, a color converter 1212, a video encoder 1213, and a multiplexer 1214.

[0241] The decimator 1201 acquires the mesh input to the encoding device 100 (corresponding to the input mesh 1101) as an original mesh, and generates a base mesh by performing a decimation process (in other words, a thinning process) on the acquired original mesh. The decimation process is a process of deleting (in other words, thinning) some of the vertices included in the original mesh. The decimation process may include a process of changing the positions of at least some of the vertices included in the original mesh, or a process of changing the connectivity of at least some of the vertices included in the original mesh. The decimation process is also simply referred to as decimation.

[0242] The base mesh generated by the decimation process has fewer vertices than the original mesh. The vertices of the base mesh may be located at different positions than the vertices of the original mesh. Also, the vertex connectivity of the base mesh may be different from the vertex connectivity of the original mesh. The decimator 1201 provides the generated base mesh to the quantizer 1202.

[0243] The quantizer 1202 quantizes the base mesh generated by the decimator 1201. The quantizer 1202 provides the quantized base mesh to the base mesh encoder 1203.

[0244] The base mesh encoder 1203 encodes the base mesh quantized by the quantizer 1202 into a bitstream (also called a base mesh bitstream) (in other words, generates a base mesh bitstream). The base mesh encoder 1203 provides the base mesh bitstream to the base mesh decoder 1204 and the multiplexer 1214.

[0245] The base mesh decoder 1204 obtains a quantized base mesh by decoding the base mesh bitstream provided by the base mesh encoder 1203. The base mesh decoder 1204 provides the quantized base mesh to the inverse quantizer 1205.

[0246] The inverse quantizer 1205 generates a base mesh (also referred to as a decoded base mesh) by inverse quantizing the quantized base mesh provided by the base mesh decoder 1204. The inverse quantizer 1205 provides the decoded base mesh to the subdivision unit 1206. There may be differences between the decoded base mesh generated by the inverse quantizer 1205 and the base mesh generated by the decimator 1201 due to the quantization and inverse quantization processes.

[0247] The subdivider 1206 performs a subdivision process on the decoded base mesh generated by the inverse quantizer 1205. The subdivision process may be a process of subdividing the faces of the decoded base mesh to make it smaller. The subdivider 1206 provides the subdivided decoded base mesh to the displacement vector calculator 1207.

[0248] Specifically, the subdivider 1206 subdivides the mesh by generating a new vertex between two connected vertices in the mesh. By repeating this process, the number of vertices in the mesh can be increased to a predetermined number. By repeating the subdivision process throughout the mesh (i.e., by performing the subdivision multiple times), multiple levels of detail (LoD) are generated.

[0249] The displacement vector calculator 1207 receives the original mesh obtained by the encoding device 100 and also receives the subdivided decoded base mesh from the subdivider 1206. The displacement vector calculator 1207 calculates a vector from a vertex of the subdivided decoded base mesh to a vertex, face, or edge of the original mesh as a displacement vector. The displacement vector calculator 1207 provides the displacement vector to the wavelet transformer 1208.

[0250] The wavelet transformer 1208 obtains wavelet coefficients by performing wavelet transform processing on the displacement vectors calculated by the displacement vector calculator 1207. The wavelet transformer 1208 provides the wavelet coefficients to the quantizer 1209. In the wavelet transform, the wavelet transformer 1208 assigns vertices to multiple LoD layers and applies, for example, a lifting transform to the displacement vectors of the vertices, thereby calculating wavelet coefficients that represent various components from low-frequency components to high-frequency components.

[0251] The quantizer 1209 quantizes the wavelet coefficients acquired by the wavelet transformer 1208. The quantizer 1209 can quantize the wavelet coefficients for each LoD layer. The quantizer 1209 provides the quantized wavelet coefficients to the image packer 1210.

[0252] The image packer 1210 generates an image containing wavelet coefficients quantized by the quantizer 1209. The image packer 1210 can generate the image by mapping the wavelet coefficients quantized by the quantizer 1209 to pixels in a two-dimensional image format. The image packer 1210 provides the generated image to the video encoder 1211. The process of mapping the quantized wavelet coefficients to pixels in the two-dimensional image format can use mapping information that represents the assignment of the quantized wavelet coefficients to pixels in the two-dimensional image format.

[0253] The video encoder 1211 encodes the image generated by the image packer 1210 into a bitstream (also called a displacement bitstream) (in other words, generates a displacement bitstream). The video encoder 1211 provides the displacement bitstream to the multiplexer 1214. The displacement bitstream may be a bitstream containing displacement information in an image format. The image format may be, for example, a format containing two chroma information pieces and one luma information piece.

[0254] The color converter 1212 obtains the attribute map obtained by the encoding device 100 as an original attribute map and performs color conversion processing on the original attribute map. The color conversion processing may include conversion processing of the color representation format or color space. The color converter 1212 provides the attribute map after color conversion processing to the video encoder 1213. Note that, although the description here takes as an example a case where the original attribute map is input to the color converter 1212, if the number or positions of vertices differ between the decoded mesh and the original mesh, the feature map may be converted to match the structure of the decoded mesh.

[0255] The video encoder 1213 encodes the attribute map converted by the color converter 1212 into a bitstream (also called an attribute bitstream) (in other words, generates an attribute bitstream). The video encoder 1213 provides the attribute bitstream to the multiplexer 1214.

[0256] The multiplexer 1214 obtains the base mesh bitstream from the base mesh encoder 1203, the displacement bitstream from the video encoder 1211, and the attribute bitstream from the video encoder 1213, and multiplexes these bitstreams to generate and output a compressed bitstream. The output of the compressed bitstream by the multiplexer 1214 may correspond to the output of the bitstream by the encoding device 100.

[0257] The process of encoding wavelet coefficients into a displacement bitstream, which is performed by the image packer 1210 and the video encoder 1211, may be performed by arithmetic coding. Alternatively, the image packer 1210 and the video encoder 1211 may be configured to select whether the process is performed by the image packer 1210 and the video encoder 1211 (also referred to as video coding) or by arithmetic coding. An example of such a configuration is described below.

[0258] Fig. 30 is a block diagram showing a modification of the detailed configuration of the encoding device 100. Fig. 30 shows a modification of the functional blocks enclosed by the dashed line in Fig. 29.

[0259] The displacement vector calculator 1207, wavelet transformer 1208, quantizer 1209, image packer 1210, and video encoder 1211 shown in FIG. 30 are the same as those shown in FIG.

[0260] As shown in FIG. 30, the encoding device 100 further includes a switch 1221, a switch 1222, and an arithmetic encoder 1223.

[0261] The switch 1221 and the switch 1222 are switchers that respectively switch whether the process of encoding wavelet coefficients into a displacement bitstream is performed by the image packer 1210 and the video coder 1211, or by the arithmetic coder 1223.

[0262] The switches 1221 and 1222 may dynamically switch the components that perform the above processing between the image packer 1210 and the video encoder 1211, and the arithmetic encoder 1223. Furthermore, the switches 1221 and 1222 may always (in other words, fixedly) use the image packer 1210 and the video encoder 1211 as the components that perform the above processing, or may always (in other words, fixedly) use the arithmetic encoder 1223.

[0263] The arithmetic coder 1223 performs the process of encoding the wavelet coefficients into a displacement bitstream by arithmetic coding.

[0264] The encoding device 100 may add information to the header information indicating whether the process of encoding the wavelet coefficients into the displacement bitstream was performed by the image packer 1210 and the video encoder 1211 (in other words, by a video encoding process) or by the arithmetic encoder 1223 (in other words, by an arithmetic encoding process). In this way, the decoding device 200 that receives the bitstream encoded as described above can appropriately decode the bitstream by referring to the header information and switching the decoding method for decoding the bitstream.

[0265] The encoding process performed by the encoding device 100 will be described in detail below.

[0266] Fig. 31 is a flow diagram showing the processing of the encoding device 100. Fig. 32 is an explanatory diagram conceptually showing the encoding of mesh frames. The processing of the encoding device 100 will be described with reference to Figs. 31 and 32.

[0267] In step S101, the encoding device 100 reads a 3D mesh frame, which is an input mesh frame, and its attributes. The input mesh frame is a mesh frame input to the encoding device 100. An example of the 3D mesh frame that is an input mesh frame is shown as mesh frame 1301 (see FIG. 32 ).

[0268] In step S102, the encoding device 100 performs a decimation process on the input mesh frame read in step S101 to generate a base mesh frame having fewer vertices than the input mesh frame. The base mesh frame generated by decimating the mesh frame 1301 is shown as a base mesh frame 1302 (see FIG. 32).

[0269] In step S103, the encoding device 100 calculates displacement information that the decoding device 200 uses to reconstruct a mesh frame. The displacement information corresponds to a displacement vector directed from a vertex of the base mesh frame generated in step S102 to a vertex of the input mesh frame. One method for calculating the displacement information is to subtract the coordinates of the vertex of the base mesh frame from the coordinates of the vertex of the input mesh frame. The displacement information calculated from the mesh frame 1301 and the base mesh frame 1302 is shown as displacement information 1303 (see FIG. 32). The displacement information 1303 is in vector format, in other words, expressed as a displacement vector.

[0270] In step S104, the encoding device 100 encodes the base mesh frame generated in step S102, the displacement information generated in step S103, and the attributes of the input mesh frame into a bitstream (corresponding to a compressed bitstream). An example of the bitstream is shown as bitstream 1304 (see FIG. 32).

[0271] Specifically, the bitstream 1304 includes vertex coordinates and connectivity information for vertices A, C, E, and F, displacement information, a video bitstream including texture data, and a compressed attribute map (see FIG. 32). The displacement information includes displacement information for displacing vertices based on vertex coordinates obtained from the subdivided base mesh frame. The compressed attribute map is texture coordinates for applying texture data to a mesh frame reconstructed using the base mesh frame and the displacement information.

[0272] The detailed configuration of the decoding device 200 will be described below.

[0273] FIG. 33 is a block diagram showing a detailed configuration example of the decoding device 200.

[0274] As shown in Figure 33, the decoding device 200 comprises a demultiplexer 2201, a base mesh decoder 2202, an inverse quantizer 2203, a subdivision unit 2204, a video decoder 2205, an image unpacker 2206, an inverse quantizer 2207, an inverse wavelet transformer 2208, a reconstructor 2209, a video decoder 2210, and a color transformer 2211.

[0275] The demultiplexer 2201 receives the compressed bitstream input to the decoding device 200 and separates it into a base mesh bitstream, a displacement bitstream, and an attribute bitstream. The demultiplexer 2201 provides the base mesh bitstream to the base mesh decoder 2202, the displacement bitstream to the video decoder 2205, and the attribute bitstream to the video decoder 2210. The compressed bitstream input to the decoding device 200 may be, for example, a compressed bitstream output by the encoding device 100, and this case will be described as an example.

[0276] The base mesh decoder 2202 obtains a quantized base mesh by decoding the base mesh bitstream provided by the demultiplexer 2201. The base mesh decoder 2202 provides the quantized base mesh to the inverse quantizer 2203.

[0277] The inverse quantizer 2203 generates a base mesh (also called a decoded base mesh) by inverse quantizing the quantized base mesh provided by the base mesh decoder 2202. The inverse quantizer 2203 provides the decoded base mesh to the subdivision unit 2204.

[0278] The subdivider 2204 performs a subdivision process on the decoded base mesh generated by the inverse quantizer 2203. The subdivision process is similar to the subdivision process performed by the subdivider 1206. The subdivider 2204 provides the subdivided decoded base mesh to the reconstructor 2209.

[0279] The video decoder 2205 decodes the displacement bitstream provided by the demultiplexer 2201 into an image, which may be an image stored by mapping quantized wavelet coefficients to pixels in a two-dimensional image format, and provides the image to the image unpacker 2206.

[0280] The image unpacker 2206 extracts quantized wavelet coefficients from the image provided by the video decoder 2205. The process of extracting the quantized wavelet coefficients from the image may use a mapping that represents the assignment of the quantized wavelet coefficients to pixels in a two-dimensional image format. The image unpacker 2206 provides the quantized wavelet coefficients extracted from the image to the inverse quantizer 2207.

[0281] The inverse quantizer 2207 generates wavelet coefficients by inverse quantizing the quantized wavelet coefficients provided by the image unpacker 2206 .

[0282] The inverse wavelet transformer 2208 generates a displacement vector (corresponding to a decoded displacement vector) by performing an inverse wavelet transform process on the wavelet coefficients provided by the inverse quantizer 2207. The inverse wavelet transform process corresponds to the inverse transform of the wavelet transform process performed by the wavelet transformer 1208. The inverse wavelet transformer 2208 provides the generated decoded displacement vector to the reconstructor 2209.

[0283] The reconstructor 2209 reconstructs a mesh (corresponding to a decoded mesh frame) using the subdivided decoded base mesh provided by the subdivider 2204 and the decoded displacement vectors provided by the inverse wavelet transformer 2208. The reconstructor 2209 also outputs the reconstructed decoded mesh as the output mesh 2107.

[0284] The video decoder 2210 decodes the attribute bitstream provided by the demultiplexer 2201 into an attribute map (corresponding to a decoded attribute map). The video decoder 2210 provides the decoded attribute map to a color converter 2211.

[0285] The color converter 2211 performs color conversion processing on the decoded attribute map provided by the video decoder 2210. The color conversion processing corresponds to the inverse conversion of the color conversion processing performed by the color converter 1212, and may include conversion processing of color representation formats or color spaces. The color converter 2211 outputs the decoded attribute map after the color conversion processing.

[0286] The process of decoding the displacement bitstream into wavelet coefficients, performed by the video decoder 2205 and the image unpacker 2206, may be performed by arithmetic coding. Furthermore, the configuration may be such that it is selectable whether the process is performed by the video decoder 2205 and the image unpacker 2206 (also referred to as video decoding) or by arithmetic coding. An example of such a configuration is described below.

[0287] Fig. 34 is a block diagram showing a modification of the detailed configuration of the decoding device 200. Fig. 34 shows a modification of the functional blocks enclosed by the dashed line in Fig. 33.

[0288] The video decoder 2205, image unpacker 2206, inverse quantizer 2207, inverse wavelet transformer 2208, and reconstructor 2209 shown in FIG. 34 are the same as those shown in FIG.

[0289] As shown in FIG. 34, the decoding device 200 further includes a switch 2221, a switch 2222, and an arithmetic decoder 2223.

[0290] The switch 2221 and the switch 2222 are switchers that respectively switch whether the process of decoding the displacement bitstream into wavelet coefficients is performed by the video decoder 2205 and image unpacker 2206, or by the arithmetic decoder 2223.

[0291] The switch 2221 and the switch 2222 may dynamically switch the components that perform the above processing between the video decoder 2205 and the image unpacker 2206 and the arithmetic decoder 2223. Furthermore, the switch 2221 and the switch 2222 may always (in other words, fixedly) use the video decoder 2205 and the image unpacker 2206 as the components that perform the above processing, or may always (in other words, fixedly) use the arithmetic decoder 2223.

[0292] The arithmetic decoder 2223 performs arithmetic decoding to decode the displaced bitstream into wavelet coefficients.

[0293] Note that the header information may include information indicating whether the process of decoding the displaced bitstream into wavelet coefficients was performed by the video decoder 2205 and the image unpacker 2206 (in other words, by a video decoding process) or by the arithmetic decoder 2223 (in other words, by an arithmetic decoding process). In this case, the decoding device 200 can appropriately decode the bitstream by switching the decoding method for the bitstream by referring to the header information.

[0294] The decoding process performed by the decoding device 200 will be described in detail below.

[0295] Fig. 35 is a flow diagram showing the processing of the decoding device 200. Fig. 36 is an explanatory diagram conceptually showing the decoding of a mesh frame (3D mesh). The processing of the decoding device 200 will be described with reference to Figs. 35 and 36.

[0296] In step S201, the decoding device 200 decodes a base mesh frame and attributes from a bitstream (corresponding to a compressed bitstream). An example of the decoded base mesh frame (corresponding to a decoded base mesh frame) is shown as a decoded base mesh frame 2301 (see FIG. 36).

[0297] In step S202, the decoding device 200 generates subdivided vertices by performing a subdivision process on the base mesh frame decoded in step S201. An example of a base mesh frame (mesh frame) including subdivided vertices is shown as a base mesh frame 2302 (see FIG. 36).

[0298] In step S203, the decoding device 200 decodes the disparity information from the bitstream (corresponding to the compressed bitstream). An example of the decoded disparity information is shown as disparity information 2303 (see FIG. 36). The disparity information 2303 is in vector format, in other words, expressed as a disparity vector.

[0299] In step S204, the decoding device 200 reconstructs the shape of the mesh frame by moving the vertices of the base mesh frame, including the subdivided vertices, to new positions using the displacement information, and then restores the mesh frame by applying attribute information. An example of the attribute is texture. An example of the reconstructed mesh frame is shown as mesh frame 2304 (see FIG. 36 ).

[0300] The subdivision is described below and is performed by a subdivider (specifically subdivider 1206 or subdivider 2204).

[0301] FIG. 37 is an explanatory diagram showing an example of subdivision.

[0302] The base mesh shown in FIG. 37(a) includes vertices A, B, and C and connectivity information indicating their connectivity.

[0303] 37(b) shows a mesh generated by the first subdivision, in other words, the mesh after the first subdivision. In the first subdivision, the subdivider generates vertices D, E, and F and connectivity information indicating their connectivity. The mesh generated by the subdivider is also referred to as LoD1 or first LoD.

[0304] Vertex D of the mesh after the first subdivision is a vertex generated by subdivision based on vertices A and B. Similarly, vertex F is a vertex generated by subdivision based on vertices B and C. Vertex E is a vertex generated by subdivision based on vertices A and C.

[0305] As an example, vertex D may be the midpoint of line segment AB (in other words, side AB) connecting vertices A and B that were the basis for its generation. Similarly, vertex E may be the midpoint of line segment AC. Vertex F may be the midpoint of line segment BC.

[0306] 37(c) shows the mesh generated by the second subdivision, i.e., the mesh after the second subdivision. In the second subdivision, the subdivider generates vertices G, H, I, J, K, L, M, N, and O and connectivity information indicating their connectivity. The mesh generated by the subdivider is also called LoD2 or second LoD.

[0307] Vertex G of the mesh after the second subdivision is a vertex generated by subdivision based on vertices A and D. Similarly, vertex H is a vertex generated by subdivision based on vertices A and E. Vertex I is a vertex generated by subdivision based on vertices B and D. Vertex J is a vertex generated by subdivision based on vertices D and F. Vertex K is a vertex generated by subdivision based on vertices E and F. Vertex L is a vertex generated by subdivision based on vertices C and E. Vertex M is a vertex generated by subdivision based on vertices B and F. Vertex N is a vertex generated by subdivision based on vertices C and F. Vertex O is a vertex generated by subdivision based on vertices D and E.

[0308] As an example, vertex G may be the midpoint of line segment AD (in other words, side AD) connecting vertices A and D, which were the source of its generation. Similarly, vertex H may be the midpoint of line segment AE. vertex I may be the midpoint of line segment BD. vertex J may be the midpoint of line segment DF. vertex K may be the midpoint of line segment EF. vertex L may be the midpoint of line segment CE. vertex M may be the midpoint of line segment BF. vertex N may be the midpoint of line segment CF. vertex O may be the midpoint of line segment DE.

[0309] The displacement of vertices will be described below with reference to Figures 38 and 39. The displacement of vertices is performed by the reconstructor 2209.

[0310] Fig. 38 is an explanatory diagram showing an example of displacement of vertices after subdivision, and Fig. 39 is an explanatory diagram showing an example of vertices of an original mesh.

[0311] The base mesh shown in FIG. 38(a) includes vertices A, B, C, and Z and connectivity information indicating their connectivity.

[0312] 38(b) shows a mesh generated by the first subdivision, in other words, a mesh after the first subdivision (i.e., the first LoD). In the first subdivision, the subdivider generates vertices S, T, U, X, or Y and connectivity information indicating their connectivity. The vertices S, T, U, X, or Y are similar to the vertices D, E, and F shown in FIG. 37(b).

[0313] 38(c) shows a mesh generated by the second subdivision, in other words, a mesh after the second subdivision (i.e., the second LoD). In the second subdivision, the subdivider generates vertices D, E, F, G, and H and connectivity information indicating their connectivity. Vertices D, E, F, G, and H are similar to vertices G, H, I, J, K, L, M, N, and O shown in FIG. 37(c).

[0314] Figure 38(d) shows a mesh including the vertices after they have been displaced after subdivision, with vertices A, B, C, D, E, F, G, H, S, T, U, X, Y, and Z shown in Figure 38(d) being located at positions displaced using displacement information from the positions of the vertices shown in Figure 38(c).

[0315] The original mesh shown in FIG. 39 is an example of the mesh input to the encoding device 100, that is, the mesh before encoding.

[0316] The mesh shown in Fig. 38 has a shape similar to that of the original mesh shown in Fig. 39. The displacement information is generated by the displacement vector calculator 1207 of the encoding device 100 as information indicating the displacement from the vertices of the base mesh to the vertices of the original mesh, and therefore, by reconstructing the mesh using the displacement information thus generated, a mesh having a shape similar to that of the original mesh is generated.

[0317] The decoding device 200 can output the mesh shown in FIG.

[0318] Next, the division of a mesh into sub-meshes will be described with reference to FIGS.

[0319] A mesh can be divided into smaller parts and coded separately, with the vertices of the mesh being divided in such a way that the coordinates and connectivity of the vertices in each part can be coded independently.

[0320] Fig. 40 is an explanatory diagram showing an example of a mesh, and Fig. 41 is an explanatory diagram showing an example of dividing a mesh into sub-meshes.

[0321] The mesh shown in FIG. 40 is the original mesh, which is sometimes called a full mesh in contrast to a sub-mesh.

[0322] Figure 41 shows how the full mesh shown in Figure 40 is divided into two sub-meshes. For vertices A, B, and C of the full mesh (see Figure 40), vertex A is duplicated to vertices A1 and A2, vertex B is duplicated to vertices B1 and B2, and vertex C is duplicated to vertices C1 and C2, thereby creating two sub-meshes (i.e., a first sub-mesh and a second sub-mesh) from the full mesh. The first sub-mesh and the second sub-mesh are each independently decodable meshes.

[0323] Packing of displacement information into image frames will be described below with reference to FIGS.

[0324] 42, 43 and 44 are explanatory diagrams showing examples of packing of displacement information into image frames. Note that image frames can also be called video frames.

[0325] The vertex displacement data is encoded as image frame data by being mapped to each component of a YUV format image frame (i.e., each of the Y component (Y Plane), U component (U Plane), and V component (V Plane)). This case will be described below as an example. As another example, the vertex displacement data may be encoded as image frame data by being mapped to each component of an RGB format image frame (each of the R component, G component, and B component).

[0326] The decoding device 200 can use an image encoding module to extract the displacement data. The displacement data can be in the form of X, Y, or Z components in a global coordinate system (e.g., a Cartesian coordinate system), or normal, tangential, or both tangential components in a local coordinate system. Methods for mapping the displacement data to an image frame include the following:

[0327] For example, in the first method, the displacement data is arranged in the image frame in scan order, and an example of packing the displacement data in this case is shown in Figure 42. The displacement data is directly mapped onto the image frame according to a predefined scan order.

[0328] Note that since an image frame has a fixed height and width, it may happen that the displacement data does not fit perfectly in the frame, in which case the remaining part of the image frame is padded with padding data (see Figure 42).

[0329] For example, in the second method, the displacement data is separated into multiple LoDs and mapped to the Y, U, and V components of the image frame. An example of packing of the displacement data in this case is shown in Figure 43. Here, the displacement data of the image frame of the next LoD starts immediately after the displacement data of the previous LoD ends. As in the first method, if the displacement data does not fit exactly into the image frame, padding is performed at the end of the image frame (see Figure 43).

[0330] For example, in the third method, displacement data corresponding to the LoD is mapped to the Y component, U component, and V component of the image frame in a manner different from that in the second method. An example of packing of the displacement data in this case is shown in Figure 44. In this way, each LoD can be decoded independently. In the third method, middle padding is performed on the displacement data of each LoD, and CTU alignment is performed together with padding at the end of the video frame (see Figure 44).

[0331] Fig. 45 is a block diagram showing a detailed configuration example of the decoding device 200 according to this embodiment. Specifically, Fig. 45 shows an example of the configuration of a geometry coordinate decoder included in the decoding device 200.

[0332] In this example, the decoding device 200 comprises a frame header decoder 631 , a vertex geometry coordinate predictor 632 , a vertex geometry coordinate difference decoder 633 , and a reconstructor 634 .

[0333] The frame header decoder 631 reads the bitstream and decodes the frame headers in the bitstream to determine whether the frame data is to be intra-decoded (intra-predicted) or inter-decoded (inter-predicted).

[0334] If inter-decoding is selected, the frame data contained in the bitstream is output to a vertex geometry coordinate predictor 632 .

[0335] The vertex geometry coordinate predictor 632 outputs prediction information to the reconstructor 634. An example of prediction information is a motion vector.

[0336] The reconstructor 634 uses prediction information along with vertex coordinates from previously decoded frames to output the three-dimensional coordinates of the vertices (vertex geometry coordinates).

[0337] On the other hand, if intra-decoding is selected, the frame data included in the bitstream is output to the vertex geometry coordinate difference decoder 633 .

[0338] The vertex geometry coordinate differential decoder 633 decodes frame data encoded as differences between the coordinates of the vertices contained in the frame to generate vertex coordinates. Only one of the vertex geometry coordinates from the vertex geometry coordinate differential decoder 633 and the reconstructor 634 is used to generate the decoded 3D mesh frame.

[0339] Fig. 46 is an explanatory diagram showing the coordinates of vertices in a 3D mesh according to this embodiment. Specifically, Fig. 46 shows an example in which the entire 3D mesh frame is decoded using the coordinates (positions) of the actual vertices included in the bitstream.

[0340] The coordinates of vertex A included in the 3D mesh frame at time (t) are decoded as (6, 8, 9) using the Cartesian coordinate system (x, y, z) as shown in (a) of Figure 46. Similarly, the coordinates of vertex B are decoded as (10, 6, 7), and the coordinates of vertex C are decoded as (14, 8, 9). The same is true for vertices D to G.

[0341] Fig. 47 is an explanatory diagram showing prediction information according to this embodiment. Specifically, Fig. 47 shows another example in which the entire 3D mesh frame at time (t) is decoded using a frame at time (t-1) (a past frame) and prediction information included in the bitstream.

[0342] The coordinates (6, 8, 9) of vertex A in the frame to be decoded (current frame) are decoded by adding the coordinates (4, 7, 8) of vertex A in the past frame to the value (2, 1, 1) for vertex A indicated by the prediction information. Similarly, the coordinates (10, 6, 7) of vertex B in the current frame are decoded by adding the coordinates (8, 6, 7) of vertex B in the past frame to the value (2, 0, 0) for vertex B indicated by the prediction information.

[0343] One way to encode a 3D mesh frame is to divide the original 3D mesh (original mesh) into several smaller meshes (sub-meshes) so that each sub-mesh can be coded independently. The vertices of the 3D mesh frame are divided so that the coordinates and connectivity information of the vertices within each partition can be coded independently. Each divided mesh is called a sub-mesh.

[0344] Next, encoding and decoding using CABAC will be described. Note that the three-dimensional data encoding device described below is a specific example of the encoding device 100, and the three-dimensional data decoding device described below is a specific example of the decoding device 200.

[0345] To divide point cloud data into tiles and slices and efficiently encode or decode the divided data, appropriate control is required on the encoding and decoding sides. By encoding and decoding the divided data independently without any dependency between the divided data, the divided data can be processed in parallel on each thread / core using a multi-threaded or multi-core processor, improving performance.

[0346] There are various methods for dividing point cloud data into tiles and slices, including methods for dividing based on characteristics such as the attributes of objects in the point cloud data, such as road surfaces, or color information, such as green, in the point cloud data.

[0347] CABAC is an abbreviation for Context-Based Adaptive Binary Arithmetic Coding, and is a coding method that improves the accuracy of probability by sequentially updating the context (a model that estimates the occurrence probability of input binary symbols) based on encoded information, thereby achieving arithmetic coding (entropy coding) with a high compression rate.

[0348] In order to process divided data such as tiles or slices in parallel, it is necessary to be able to encode or decode each divided data independently. However, in order to make CABAC independent between divided data, it is necessary to initialize CABAC at the beginning of the divided data in encoding and decoding, but there is no mechanism for doing so.

[0349] CABAC The CABAC initialization flag is used to initialize CABAC in CABAC encoding and decoding.

[0350] FIG. 48 is a flow diagram showing the process of initializing CABACCABAC in accordance with the CABAC initialization flag during encoding or decoding.

[0351] The three-dimensional data encoding device or three-dimensional data decoding device determines whether the CABAC initialization flag is 1 during encoding or decoding (S5201).

[0352] If the CABAC initialization flag is 1 (Yes in S5201), the three-dimensional data encoding device or three-dimensional data decoding device initializes the CABAC encoding unit / decoding unit to the default state (S5202) and continues encoding or decoding.

[0353] If the CABAC initialization flag is not 1 (No in S5201), the three-dimensional data encoding device or three-dimensional data decoding device continues encoding or decoding without initialization.

[0354] That is, when initializing CABAC, CABAC_init_flag = 1 is set, and the CABAC encoding unit or CABAC decoding unit is initialized or reinitialized. When initializing, the initial value (default state) of the context used in CABAC processing is set.

[0355] The encoding process will now be described. Fig. 49 is a block diagram showing the configuration of a first encoding unit 5200 included in the three-dimensional data encoding device according to this embodiment. Fig. 50 is a block diagram showing the configuration of a dividing unit 5201 according to this embodiment. Fig. 51 is a block diagram showing the configurations of a position information encoding unit 5202 and an attribute information encoding unit 5203 according to this embodiment.

[0356] The first encoding unit 5200 generates encoded data (encoded stream) by encoding the point cloud data using a first encoding method (GPCC (Geometry based PCC)). The first encoding unit 5200 includes a dividing unit 5201, a plurality of position information encoding units 5202, a plurality of attribute information encoding units 5203, an additional information encoding unit 5204, and a multiplexing unit 5205.

[0357] The dividing unit 5201 divides the point cloud data to generate a plurality of divided data. Specifically, the dividing unit 5201 divides the space of the point cloud data into a plurality of subspaces to generate a plurality of divided data. Here, a subspace is one of a tile and a slice, or a combination of a tile and a slice. More specifically, the point cloud data includes position information, attribute information, and additional information. The dividing unit 5201 divides the position information into a plurality of divided position information pieces, and divides the attribute information into a plurality of divided attribute information pieces. The dividing unit 5201 also generates additional information related to the division.

[0358] 50 , the dividing unit 5201 includes a tile dividing unit 5211 and a slice dividing unit 5212. For example, the tile dividing unit 5211 divides a point cloud into tiles. The tile dividing unit 5211 may determine, as tile additional information, a quantization value to be used for each divided tile.

[0359] The slice division unit 5212 further divides the tiles obtained by the tile division unit 5211 into slices. The slice division unit 5212 may determine a quantization value to be used for each divided slice as slice additional information.

[0360] The position information encoding units 5202 encode the plurality of pieces of divided position information to generate a plurality of pieces of encoded position information. For example, the position information encoding units 5202 process the plurality of pieces of divided position information in parallel.

[0361] 51 , the position information encoding unit 5202 includes a CABAC initialization unit 5221 and an entropy encoding unit 5222. The CABAC initialization unit 5221 initializes or re-initializes CABAC according to a CABAC initialization flag. The entropy encoding unit 5222 encodes the split position information using CABAC.

[0362] The attribute information encoding units 5203 encode the divided attribute information to generate the coded attribute information, for example, the attribute information encoding units 5203 process the divided attribute information in parallel.

[0363] 51 , the attribute information encoding unit 5203 includes a CABAC initialization unit 5231 and an entropy encoding unit 5232. The CABAC initialization unit 5221 initializes or re-initializes CABAC according to a CABAC initialization flag. The entropy encoding unit 5232 encodes the divided attribute information using CABAC.

[0364] The additional information encoding unit 5204 generates encoded additional information by encoding the additional information included in the point cloud data and the additional information related to the data division generated by the division unit 5201 at the time of division.

[0365] The multiplexing unit 5205 multiplexes a plurality of pieces of encoding position information, a plurality of pieces of encoding attribute information, and encoding additional information to generate encoded data (encoded stream), and transmits the generated encoded data. The encoded additional information is also used during decoding.

[0366] 49 shows an example in which there are two position information encoding units 5202 and two attribute information encoding units 5203, but the number of position information encoding units 5202 and two attribute information encoding units 5203 may each be one, or three or more. Furthermore, the multiple pieces of divided data may be processed in parallel within the same chip, like multiple cores within a CPU, or may be processed in parallel by cores on multiple chips, or may be processed in parallel by multiple cores on multiple chips.

[0367] Next, the decoding process will be described. Fig. 52 is a block diagram showing the configuration of the first decoding unit 5240. Fig. 53 is a block diagram showing the configurations of the position information decoding unit 5242 and the attribute information decoding unit 5243.

[0368] The first decoding unit 5240 restores the point cloud data by decoding the coded data (coded stream) generated by coding the point cloud data using the first coding method (GPCC). The first decoding unit 5240 includes a demultiplexing unit 5241, a plurality of position information decoding units 5242, a plurality of attribute information decoding units 5243, an additional information decoding unit 5244, and a combining unit 5245.

[0369] The demultiplexing unit 5241 demultiplexes the coded data (coded stream) to generate a plurality of pieces of coding position information, a plurality of pieces of coding attribute information, and coded additional information.

[0370] The position information decoding units 5242 generate a plurality of pieces of quantized position information by decoding the plurality of pieces of encoded position information. For example, the position information decoding units 5242 process the plurality of pieces of encoded position information in parallel.

[0371] 53 , the position information decoding unit 5242 includes a CABAC initialization unit 5251 and an entropy decoding unit 5252. The CABAC initialization unit 5251 initializes or re-initializes CABAC according to a CABAC initialization flag. The entropy decoding unit 5252 decodes the position information using CABAC.

[0372] The attribute information decoding units 5243 generate a plurality of pieces of divided attribute information by decoding the plurality of pieces of encoded attribute information. For example, the attribute information decoding units 5243 process the plurality of pieces of encoded attribute information in parallel.

[0373] 53 , the attribute information decoding unit 5243 includes a CABAC initialization unit 5261 and an entropy decoding unit 5262. The CABAC initialization unit 5261 initializes or re-initializes CABAC according to a CABAC initialization flag. The entropy decoding unit 5262 decodes the attribute information using CABAC.

[0374] The plurality of additional information decoders 5244 generate additional information by decoding the encoded additional information.

[0375] The combining unit 5245 generates position information by combining multiple pieces of split position information using the additional information. The combining unit 5245 generates attribute information by combining multiple pieces of split attribute information using the additional information. For example, the combining unit 5245 first generates point cloud data corresponding to a tile by combining decoded point cloud data for a slice using the slice additional information. Next, the combining unit 5245 restores the original point cloud data by combining the point cloud data corresponding to the tile using the tile additional information.

[0376] 52 shows an example in which there are two position information decoding units 5242 and two attribute information decoding units 5243, but the number of position information decoding units 5242 and two attribute information decoding units 5243 may each be one, or three or more. Furthermore, multiple pieces of divided data may be processed in parallel within the same chip, like multiple cores within a CPU, or may be processed in parallel by cores on multiple chips, or may be processed in parallel by multiple cores on multiple chips.

[0377] FIG. 54 is a flow diagram showing an example of a process related to initialization of CABAC in encoding of position information or encoding of attribute information.

[0378] First, the three-dimensional data encoding device determines, for each slice, based on a predetermined condition, whether or not to perform CABAC initialization when encoding the position information of that slice (S5201).

[0379] When the three-dimensional data encoding device determines to perform CABAC initialization (Yes in S5202), it determines an initial context value to be used for encoding the position information (S5203). The initial context value is set to an initial value that takes into account encoding characteristics. The initial value may be a predetermined value, or may be adaptively determined according to the characteristics of the data in the slice.

[0380] Next, the three-dimensional data encoding device sets the CABAC initialization flag of the position information to 1 and sets a context initial value (S5204). When performing CABAC initialization, initialization processing is performed using the context initial value in encoding of the position information.

[0381] On the other hand, if the three-dimensional data encoding device determines not to perform CABAC initialization (No in S5202), it sets the CABAC initialization flag of the position information to 0 (S5205).

[0382] Next, the three-dimensional data encoding device determines, for each slice, based on a predetermined condition, whether or not to perform CABAC initialization when encoding the attribute information of that slice (S5206).

[0383] When the three-dimensional data encoding device determines to perform CABAC initialization (Yes in S5207), it determines an initial context value to be used for encoding the attribute information (S5208). The initial context value is set to an initial value that takes into account encoding characteristics. The initial value may be a predetermined value, or may be adaptively determined according to the characteristics of the data in the slice.

[0384] Next, the three-dimensional data encoding device sets the CABAC initialization flag of the attribute information to 1 and sets a context initial value (S5209). When performing CABAC initialization, initialization processing is performed using the context initial value in encoding of the attribute information.

[0385] On the other hand, if the three-dimensional data encoding device determines not to perform CABAC initialization (No in S5207), it sets the CABAC initialization flag in the attribute information to 0 (S5210).

[0386] In the flowchart of FIG. 54, the processing order of the processing related to the position information and the processing related to the attribute information may be reversed, or may be performed in parallel.

[0387] Note that, although the flow diagram in Fig. 54 illustrates processing in units of slices as an example, processing in units of tiles or other data units can also be performed in the same manner as in units of slices. In other words, the word "slice" in the flow diagram in Fig. 54 can be read as "tile" or other data unit.

[0388] Furthermore, the predetermined condition may be the same for the location information and the attribute information, or may be different for each.

[0389] FIG. 55 is a diagram showing an example of the timing of CABAC initialization in point cloud data converted into a bit stream.

[0390] Point cloud data includes position information and zero or more pieces of attribute information. That is, point cloud data may have no attribute information or may have multiple pieces of attribute information.

[0391] For example, the attribute information for one three-dimensional point may include color information, color information and reflection information, or one or more pieces of color information each associated with one or more pieces of viewpoint information.

[0392] The method described in this embodiment can be applied to either configuration.

[0393] Next, the conditions for determining whether to initialize CABAC will be described.

[0394] If the following conditions are met, CABAC in encoding position information or attribute information may be initialized.

[0395] For example, the CABAC may be initialized with the leading data of the position information or attribute information (if there is more than one, each piece of attribute information). For example, the CABAC may be initialized with the leading data of the data constituting an independently decodable PCC frame. That is, as shown in (a) of Figure 55, if the PCC frame can be decodable on a frame-by-frame basis, the CABAC may be initialized with the leading data of the PCC frame.

[0396] Also, for example, as shown in (b) of Figure 55, if a frame cannot be decoded independently, such as when inter-prediction is used between PCC frames, CABAC may be initialized with the first data of a random access unit (e.g., GOF).

[0397] Also, for example, as shown in (c) of Figure 55, CABAC may be initialized at the beginning of slice data divided into one or more pieces, the beginning of tile data divided into one or more pieces, or the beginning of other divided data.

[0398] Although (c) in Fig. 55 shows an example of a tile, the same applies to a slice. Initialization may or may not be required at the beginning of a tile or slice.

[0399] FIG. 56 shows the structure of coded data and a method of storing coded data in an NAL unit.

[0400] The initialization information may be stored in a header of the encoded data, or in metadata, or may be stored in both the header and metadata. The initialization information may be, for example, caba_init_flag, a CABAC initial value, or an index of a table that can identify the initial value.

[0401] In this embodiment, the description that the information is stored in the metadata may be interpreted as being stored in the header of the encoded data, and vice versa.

[0402] When the initialization information is stored in the header of the encoded data, it may be stored in, for example, the first NAL unit in the encoded data. The position information stores initialization information for encoding the position information, and the attribute information stores initialization information for encoding the attribute information.

[0403] The cabac_init_flag for encoding the attribute information and the cabac_init_flag for encoding the position information may be the same value or different values. If the same value is used, the cabac_init_flag for the position information and the attribute information may be the same. If different values ​​are used, the cabac_init_flag for the position information and the attribute information will each indicate a different value.

[0404] The initialization information may be stored in metadata common to the location information and the attribute information, or may be stored in at least one of the individual metadata for the location information and the individual metadata for the attribute information, or may be stored in both the common metadata and the individual metadata. Also, a flag may be used to indicate whether the initialization information is described in the individual metadata for the location information, the individual metadata for the attribute information, or the common metadata.

[0405] FIG. 57 is a flow diagram showing an example of a process related to initialization of CABAC in decoding of position information or attribute information.

[0406] The three-dimensional data decoding device analyzes the encoded data, and obtains the CABAC initialization flag of the position information, the CABAC initialization flag of the attribute information, and the context initial value (S5211).

[0407] Next, the three-dimensional data decoding device determines whether the CABAC initialization flag of the position information is 1 (S5212).

[0408] If the CABAC initialization flag for the position information is 1 (Yes in S5212), the three-dimensional data decoding device initializes the CABAC decoding of the position information encoding using the context initial value for the position information encoding (S5213).

[0409] On the other hand, if the CABAC initialization flag for the position information is 0 (No in S5212), the three-dimensional data decoding device does not initialize CABAC decoding in the position information encoding (S5214).

[0410] Next, the three-dimensional data decoding device determines whether the CABAC initialization flag in the attribute information is 1 (S5215).

[0411] If the CABAC initialization flag of the attribute information is 1 (Yes in S5215), the three-dimensional data decoding device initializes the CABAC decoding of the attribute information encoding using the context initial value of the attribute information encoding (S5216).

[0412] On the other hand, if the CABAC initialization flag in the attribute information is 0 (No in S5215), the three-dimensional data decoding device does not initialize CABAC decoding in the attribute information encoding (S5217).

[0413] In the flowchart of FIG. 57, the processing order of the processing relating to the position information and the processing relating to the attribute information may be reversed, or may be performed in parallel.

[0414] The flow chart in FIG. 57 can be applied to both the case of slice division and the case of tile division.

[0415] Next, the flow of the encoding process and decoding process of point cloud data according to this embodiment will be described. Fig. 58 is a flow diagram of the encoding process of point cloud data according to this embodiment.

[0416] First, the three-dimensional data encoding device determines the division method to be used (S5221). This division method includes whether or not to perform tile division and whether or not to perform slice division. The division method may also include the number of divisions when performing tile division or slice division, and the type of division. The type of division may be a method based on the object shape as described above, a method based on map information or location information, or a method based on the amount of data or processing. The division method may be predetermined.

[0417] If tile division is to be performed (Yes in S5222), the three-dimensional data encoding device divides the position information and attribute information into tiles to generate multiple pieces of tile position information and multiple pieces of tile attribute information (S5223). The three-dimensional data encoding device also generates tile additional information related to the tile division.

[0418] If slice division is performed (Yes in S5224), the three-dimensional data encoding device divides the plurality of pieces of tile position information and the plurality of pieces of tile attribute information (or the position information and the attribute information) to generate a plurality of pieces of division position information and a plurality of pieces of division attribute information (S5225). The three-dimensional data encoding device also generates position slice additional information and attribute slice additional information related to the slice division.

[0419] Next, the three-dimensional data encoding device generates a plurality of pieces of encoding position information and a plurality of pieces of encoding attribute information by encoding each of the plurality of pieces of division position information and the plurality of pieces of division attribute information (S5226). The three-dimensional data encoding device also generates dependency relationship information.

[0420] Next, the three-dimensional data encoding device generates encoded data (encoded stream) by NAL unitizing (multiplexing) the plurality of pieces of encoding position information, the plurality of pieces of encoding attribute information, and the additional information (S5227).The three-dimensional data encoding device also transmits the generated encoded data.

[0421] FIG. 59 is a flowchart showing an example of a process of determining the value of the CABAC initialization flag and updating additional information in the division into tiles (S5222) or the division into slices (S5225).

[0422] In steps S5222 and S5225, the position information and attribute information of the tiles and / or slices may be divided independently and individually by the respective methods, or may be divided together in common, thereby generating additional information divided for each tile and / or slice.

[0423] At this time, the three-dimensional data encoding device determines whether to set the CABAC initialization flag to 1 or 0 (S5231).

[0424] Then, the three-dimensional data encoding device updates the additional information so that the determined CABAC initialization flag is included (S5232).

[0425] FIG. 60 is a flowchart showing an example of the CABAC initialization process in the encoding process (S5226).

[0426] The three-dimensional data encoding device determines whether the CABAC initialization flag is 1 (S5241).

[0427] If the CABAC initialization flag is 1 (Yes in S5241), the three-dimensional data encoding device re-initializes the CABAC encoding unit to the default state (S5242).

[0428] Then, the three-dimensional data encoding device continues the encoding process until a condition for stopping the encoding process is satisfied, for example, until there is no more data to be encoded (S5243).

[0429] 61 is a flow diagram of a decoding process for point cloud data according to this embodiment. First, the 3D data decoding device determines the division method by analyzing additional information related to the division method (tile additional information, position slice additional information, and attribute slice additional information) included in the coded data (coded stream) (S5251). This division method includes whether or not to perform tile division and whether or not to perform slice division. The division method may also include the number of divisions into tiles or slices, the type of division, etc.

[0430] Next, the three-dimensional data decoding device generates split position information and split attribute information by decoding the multiple pieces of coded position information and multiple pieces of coded attribute information contained in the coded data using the dependency information contained in the coded data (S5252).

[0431] If the additional information indicates that slice division has been performed (Yes in S5253), the three-dimensional data decoding device generates multiple tile position information and multiple tile attribute information by combining multiple division position information and multiple division attribute information based on the position slice additional information and attribute slice additional information (S5254).

[0432] If the additional information indicates that tile division has been performed (Yes in S5255), the three-dimensional data decoding device generates position information and attribute information by combining multiple tile position information and multiple tile attribute information (multiple division position information and multiple division attribute information) based on the tile additional information (S5256).

[0433] FIG. 62 is a flow diagram showing an example of processing for initializing the CABAC decoding unit in combining information divided for each slice (S5254) or combining information divided for each tile (S5256).

[0434] The position information and attribute information of the slices or tiles may be combined using different methods or may be combined using the same method.

[0435] The three-dimensional data decoding device decodes the CABAC initialization flag from the additional information of the coded stream (S5261).

[0436] Next, the three-dimensional data decoding device determines whether the CABAC initialization flag is 1 (S5262).

[0437] If the CABAC initialization flag is 1 (Yes in S5262), the three-dimensional data decoding device re-initializes the CABAC decoding unit to the default state (S5263).

[0438] On the other hand, if the CABAC initialization flag is not 1 (No in S5262), the three-dimensional data decoding device proceeds to step S5264 without re-initializing the CABAC decoding unit.

[0439] Then, the three-dimensional data decoding device continues the decoding process until a condition for stopping the decoding process is satisfied, for example, until there is no more data to be decoded (S5264).

[0440] Next, other determination conditions for CABAC initialization will be described.

[0441] Whether to initialize the encoding of the position information or the encoding of the attribute information may be determined in consideration of the encoding efficiency of data units such as tiles or slices. In this case, CABAC may be initialized in the first data of a tile or slice that satisfies a predetermined condition.

[0442] Next, the determination conditions for CABAC initialization in encoding position information will be described.

[0443] For example, the three-dimensional data encoding device may determine the density of point cloud data for each slice, i.e., the number of points per unit area belonging to the slice, compare the data density of the slice with that of other slices, and if the change in data density does not exceed a predetermined condition, determine that not initializing CABAC will result in better encoding efficiency, and decide not to initialize CABAC.On the other hand, if the change in data density does not satisfy the predetermined condition, the three-dimensional data encoding device may determine that initialization will result in better encoding efficiency, and decide to initialize CABAC.

[0444] Here, the other slice may be, for example, the slice immediately preceding the current slice in the decoding process order or a spatially adjacent slice. Furthermore, the three-dimensional data encoding device may determine whether to perform CABAC initialization depending on whether the data density of the current slice is a predetermined data density, without comparing it with the data densities of other slices.

[0445] When it is determined that CABAC initialization is to be performed, the three-dimensional data encoding device determines a context initial value to be used for encoding the position information. The context initial value is set to an initial value that has good encoding characteristics according to the data density. The three-dimensional data encoding device may store a table of initial values ​​for each data density in advance and select an optimal initial value from the table.

[0446] The three-dimensional data encoding device may determine whether to perform CABAC initialization based on the number of points, the distribution of points, the bias of points, or the like, without being limited to the example of slice density. Alternatively, the three-dimensional data encoding device may determine whether to perform CABAC initialization based on feature amounts or the number of feature points obtained from point information, or a recognized object. In this case, the determination criteria may be stored in advance in memory as a table associated with feature amounts or the number of feature points obtained from point information, or an object recognized based on point information.

[0447] The three-dimensional data encoding device may, for example, determine an object in the location information of the map information and determine whether to perform CABAC initialization based on the object based on the location information, or may determine whether to perform CABAC initialization based on information or features obtained by projecting three-dimensional data into two dimensions.

[0448] Next, the determination conditions for CABAC initialization in encoding attribute information will be described.

[0449] The three-dimensional data encoding device may, for example, compare the color characteristics of the previous slice with the color characteristics of the current slice, and if the change in the color characteristics satisfies a predetermined condition, determine that not initializing the CABAC will result in better coding efficiency, and may decide not to initialize the CABAC. On the other hand, if the change in the color characteristics does not satisfy the predetermined condition, the three-dimensional data encoding device may determine that initializing the CABAC will result in better coding efficiency, and may decide to initialize the CABAC. The color characteristics include, for example, luminance, chromaticity, saturation, their histograms, color continuity, etc.

[0450] Here, the other slice may be, for example, the slice immediately preceding the current slice in the decoding process order or a spatially adjacent slice. Furthermore, the three-dimensional data encoding device may determine whether to perform CABAC initialization depending on whether the data density of the current slice is a predetermined data density, without comparing it with the data densities of other slices.

[0451] When it is determined that CABAC initialization is to be performed, the three-dimensional data encoding device determines a context initial value to be used for encoding the attribute information. The context initial value is set to an initial value that has good encoding characteristics according to the data density. The three-dimensional data encoding device may store a table of initial values ​​for each data density in advance and select an optimal initial value from the table.

[0452] When the attribute information is reflectance, the three-dimensional data encoding device may determine whether to perform CABAC initialization according to information based on the reflectance.

[0453] When a three-dimensional point has multiple pieces of attribute information, the three-dimensional data encoding device may determine initialization information based on each piece of attribute information independently for each piece of attribute information, or may determine initialization information for multiple pieces of attribute information based on one piece of attribute information, or may determine initialization information for the multiple pieces of attribute information using multiple pieces of attribute information.

[0454] An example has been described in which the initialization information for location information is determined based on location information, and the initialization information for attribute information is determined based on attribute information, but the initialization information for location information and attribute information may also be determined based on location information, or the initialization information for location information and attribute information may also be determined based on attribute information, or the initialization information for location information and attribute information may also be determined based on both information.

[0455] The three-dimensional data encoding device may determine initialization information based on the results of a prior simulation of encoding efficiency, for example, by setting cabac_init_flag to on or off, or by using one or more initial values ​​from an initial value table.

[0456] When the three-dimensional data encoding device determines the method of dividing data into slices, tiles, etc. based on position information or attribute information, it may determine the initialization information based on the same information as the information used to determine the division method.

[0457] FIG. 63 is a diagram showing examples of tiles and slices.

[0458] For example, slices in a tile that have a portion of the PCC data are identified as shown in the legend. The CABAC initialization flag can be used to determine whether context reinitialization is required for subsequent slices. For example, in Figure 63, if a tile contains slice data divided by object (moving object, sidewalk, building, tree, and other object), the CABAC initialization flags for the moving object, sidewalk, and tree slices are set to 1, and the CABAC initialization flags for the building and other slices are set to 0. This means that, for example, if both sidewalks and buildings are dense permanent structures and may have similar coding efficiency, coding efficiency may be improved by not reinitializing CABAC between sidewalk and building slices. On the other hand, if the density and coding efficiency of buildings and trees may be significantly different, coding efficiency may be improved by initializing CABAC between building and tree slices.

[0459] FIG. 64 is a flow diagram showing an example of a method for initializing CABAC and determining a context initial value.

[0460] First, the three-dimensional data encoding device divides the point cloud data into slices based on the object determined from the position information (S5271).

[0461] Next, the three-dimensional data encoding device determines, for each slice, whether to perform CABAC initialization for encoding position information and attribute information, based on the data density of the object in that slice (S5272). That is, the three-dimensional data encoding device determines CABAC initialization information (CABAC initialization flag) for encoding position information and attribute information, based on the position information. The three-dimensional data encoding device determines initialization with good encoding efficiency, for example, based on the point cloud data density. Note that the CABAC initialization information may be indicated in a common cabac_init_flag for both position information and attribute information.

[0462] Next, if it is determined that CABAC initialization is to be performed (Yes in S5273), the three-dimensional data encoding device determines the initial context value for encoding the position information (S5274).

[0463] Next, the three-dimensional data encoding device determines the initial context value for encoding the attribute information (S5275).

[0464] Next, the three-dimensional data encoding device sets the CABAC initialization flag for the position information to 1, sets a context initial value for the position information, and also sets the CABAC initialization flag for the attribute information to 1, and sets a context initial value for the attribute information (S5276). Note that when performing CABAC initialization, the three-dimensional data encoding device performs initialization processing using the context initial value in each of encoding the position information and encoding the attribute information.

[0465] On the other hand, if the three-dimensional data encoding device determines not to perform CABAC initialization (No in S5273), it sets the CABAC initialization flag for the position information to 0 and sets the CABAC initialization flag for the attribute information to 0 (S5277).

[0466] Fig. 65 is a diagram showing an example of a case where a map obtained by LiDAR and viewed from above is divided into tiles. Fig. 66 is a flow diagram showing another example of a method for CABAC initialization and determining a context initial value.

[0467] The three-dimensional data encoding device divides point cloud data into one or more tiles in large-scale map data using a two-dimensional division method in a top view based on position information (S5281). The three-dimensional data encoding device may divide the data into square regions, for example, as shown in Fig. 65. The three-dimensional data encoding device may also divide the point cloud data into tiles of various shapes and sizes. The division into tiles may be performed using one or more predetermined methods, or may be performed adaptively.

[0468] Next, the three-dimensional data encoding device determines, for each tile, an object within the tile, and determines whether to initialize CABAC when encoding the position information or attribute information of the tile (S5282). In the case of slice division, the three-dimensional data encoding device recognizes objects (trees, people, moving objects, buildings), and determines slice division and initial values ​​according to the objects.

[0469] When it is determined that CABAC initialization is to be performed (Yes in S5283), the three-dimensional data encoding device determines the initial context value for encoding the position information (S5284).

[0470] Next, the three-dimensional data encoding device determines the initial context value for encoding the attribute information (S5285).

[0471] In steps S5284 and S5285, the initial values ​​of tiles having specific coding characteristics may be stored as initial values, and may be used as initial values ​​of tiles having the same coding characteristics.

[0472] Next, the three-dimensional data encoding device sets the CABAC initialization flag for the position information to 1, sets a context initial value for the position information, and also sets the CABAC initialization flag for the attribute information to 1, and sets a context initial value for the attribute information (S5286). Note that when performing CABAC initialization, the three-dimensional data encoding device performs initialization processing using the context initial value in each of encoding the position information and encoding the attribute information.

[0473] On the other hand, if the three-dimensional data encoding device determines not to perform CABAC initialization (No in S5283), it sets the CABAC initialization flag for the position information to 0 and sets the CABAC initialization flag for the attribute information to 0 (S5287).

[0474] In the CABAC (Context-Based Adaptive Binary Arithmetic Coding) of the above embodiment, the three-dimensional data encoding device may encode multiple three-dimensional points included in a data unit using one of multiple encoding methods that are different from each other. That is, for each data unit, the three-dimensional data encoding device determines the encoding method for encoding the multiple three-dimensional points included in the data unit as the encoding method suitable for that data unit from among the multiple encoding methods. For example, in encoding the position information of the three-dimensional points, the multiple encoding methods include octet encoding, which is an encoding method using an octet tree, and predictive tree encoding, which is an encoding method using a predictive tree.

[0475] Signaling of a CABAC initialization flag (hereinafter also referred to as initialization information or identification information) in such CABAC encoding will be described.

[0476] [Signaling] Initialization information is stored in the header of encoded data. The initialization information is, for example, caba_init_flag, a CABAC initial value, an index of a table that can identify the initial value, etc. The initialization information is used to initialize CABAC in CABAC encoding and decoding. In other words, the initialization information (identification information) is information indicating whether or not to continue using the context used for encoding.

[0477] The three-dimensional data encoding device may store the initialization information in the metadata, or may describe it in both the header and the metadata. Note that, in this embodiment, storing in the metadata may be interpreted as storing in the header of the encoded data, and conversely, storing in the header of the encoded data may be interpreted as storing in the metadata.

[0478] The three-dimensional data encoding device may apply the initialization information to either encoding of position information or encoding of attribute information. When storing initialization information in the header of encoded data, the three-dimensional data encoding device may store initialization information for encoding of position information in the position information, and may store initialization information for attribute information in the attribute information.

[0479] [Location Information Header] CABAC stands for Context-Based Adaptive Binary Arithmetic Coding, and is a coding method that improves the accuracy of probability by sequentially updating the context (a model that estimates the occurrence probability of an input binary symbol) based on coded information, thereby achieving arithmetic coding (entropy coding) with a high compression rate. In order to process multiple data units (multiple divided data) obtained by dividing point cloud data, such as tiles or slices, in parallel, it is necessary to be able to encode or decode each data unit independently. However, in order to make CABAC independent between data units, it is necessary to initialize CABAC at the beginning of the data unit during encoding and decoding. The CABAC initialization flag is used to initialize CABAC during CABAC encoding and decoding.

[0480] FIG. 67 is a diagram showing an example of the data structure of the position information included in each data unit after division and the syntax of the header of the position information.

[0481] The three-dimensional data encoding device may apply the initialization information to either one of encoding methods (encoding methods) such as octet-tree encoding or predictive tree encoding when encoding the position information, or to both. Octet-tree encoding and predictive tree encoding are encoding methods that use different tree structures.

[0482] In the case of octree coding using an octree structure, the three-dimensional data coding device stores a context used in the octree coding (i.e., a context for octree coding). In addition, in the case of predictive tree coding using a predictive tree structure, the three-dimensional data coding device stores a context used in the predictive tree coding (i.e., a context for predictive tree coding).

[0483] The three-dimensional data encoding device stores initialization information in the header of each data unit of divided position information, thereby enabling switching, for each divided data unit, whether or not to initialize the context used for encoding. In other words, the three-dimensional data encoding device stores identification information in the header of each data unit of divided position information, thereby enabling switching, for each divided data unit, whether or not to continue using the context used for encoding.

[0484] SPS_ID indicates the identifier of the SPS (parameter set) referenced by the data unit. GPS_ID indicates the identifier of the GPS (location parameter set) referenced by the data unit. tile_id indicates the identifier of the tile to which the data unit belongs (divided data identifier 1). slice_id indicates the identifier of the slice to which the data unit belongs (divided data identifier 2).

[0485] The tree_mode indicates a tree structure used in encoding the position information of the data unit. When there are two types of tree structures, the tree_mode may be a flag. For example, the tree_mode may indicate an octree when the flag is 0, and a predictive tree (predtree) when the flag is 1. Note that when the tree_mode is indicated in the GPS, it does not need to be indicated in the slice header.

[0486] The three-dimensional data encoding device may switch and signal the structure of the metadata used for each encoding based on tree_mode.

[0487] For example, when the tree structure is an octree (tree_mode=='octree'), the three-dimensional data encoding device signals a parameter (octree_information) to be used for octree encoding. Furthermore, a flag (cabac_init_flag) indicating whether to initialize a context in octree encoding, in other words, identification information indicating whether to continue using the context, may be indicated.

[0488] Also, for example, when the tree structure is a predictive tree (tree_mode=='predtree'), the three-dimensional data encoding device signals a parameter (predtree_information) to be used for predictive tree encoding. Furthermore, a flag (cabac_init_flag) indicating whether to initialize a context in predictive tree encoding, in other words, identification information indicating whether to continue using the context, may be indicated.

[0489] It should be noted that the amount of information to be signaled may be reduced and compression efficiency may be improved by using the following method: The three-dimensional data encoding device may set cabac_init_flag as a flag common to multiple encoding methods and signal it before the conditional branch of tree_mode.

[0490] Furthermore, the three-dimensional data encoding device may apply context initialization to some tree structures and not to other tree structures. In this case, the three-dimensional data encoding device may generate a syntax header that does not include initialization information for tree structures to which context initialization is not applied, and that includes initialization information for tree structures to which context initialization is applied. For example, when performing initialization for all divided data units, the three-dimensional data encoding device may indicate initialization information commonly to a higher-level parameter set, such as SPS or GPS, and may not indicate initialization information for each data unit.

[0491] 68 is a flow diagram showing an example of a three-dimensional data encoding method, in which encoding of position information of a plurality of three-dimensional points included in a data unit is described.

[0492] The three-dimensional data encoding device determines the encoding method for the data unit to be processed, and determines whether to continue CABAC in encoding the position information of the first three-dimensional point of the data unit to be processed (S11401). In other words, the three-dimensional data encoding device determines the encoding method for the data unit to be either octree encoding or predictive tree encoding, and determines whether to continue encoding using the context used for encoding.

[0493] Next, if the three-dimensional data encoding device determines to continue using the context (Yes in S11402), it sets cabac_init_flag to false (S11403). That is, the three-dimensional data encoding device sets the identification information to indicate that the context used for encoding will continue to be used. The three-dimensional data encoding device sets the identification information so that the identification information indicates the result of the determination in step S11402.

[0494] Next, when the three-dimensional data encoding device determines that the encoding method is octet encoding (octet in S11404), it continues to use the context used in the octet encoding to encode the data unit using the octet (S11405). The context used in the octet encoding is the context used in the octet encoding of the data unit immediately before the data unit to be processed. This context is temporarily stored, for example, in the memory of the three-dimensional data encoding device, and the three-dimensional data encoding device reads out the context stored in the memory and uses it to encode the data unit to be processed.

[0495] On the other hand, when the three-dimensional data encoding device determines that the encoding method is predictive tree encoding (predictive tree in S11404), it continues to use the context used in predictive tree encoding to perform predictive tree encoding (S11406). The context used in predictive tree encoding is the context used in predictive tree encoding of the data unit immediately preceding the data unit to be processed. This context is, for example, temporarily stored in the memory of the three-dimensional data encoding device, and the three-dimensional data encoding device reads the context stored in the memory and uses it to encode the data unit to be processed.

[0496] As shown in steps S11405 and S11406, the three-dimensional data encoding device continues to use the context used in the encoding method determined in step S11401 from among the multiple encoding methods to perform encoding.

[0497] When a context is continuously used, the three-dimensional data encoding device changes the value of the continuing context depending on the encoding method of the position information (octree or predictive tree). For example, a context for occupancy tree encoding is a context for entropy encoding occupancy codes, quantization values, overlapping points in leaf nodes, etc., and a context for predictive tree encoding is a context for entropy encoding the number of nodes, prediction mode, etc.

[0498] Furthermore, if the three-dimensional data encoding device determines not to continue using the context (No in S11402), that is, if it determines to initialize the context, it sets cabac_init_flag to true (S11407). That is, the three-dimensional data encoding device sets the identification information to indicate that the context used for encoding will not be continued. The three-dimensional data encoding device sets the identification information so that the identification information indicates the result of the determination in step S11402.

[0499] Next, the three-dimensional data encoding device encodes the position information of the first three-dimensional point of the data unit using the initialized context for the encoding method determined in step S11401 (S11408).

[0500] As described above, the three-dimensional data encoding device performs encoding using a context for octet-tree encoding when performing octet-tree encoding, and performs encoding using a context for predictive tree encoding when performing predictive tree encoding. In other words, the three-dimensional data encoding method changes the context to be used continuously for encoding depending on the encoding method for position information.

[0501] 69 is a flow diagram showing an example of a three-dimensional data decoding method, in which decoding of position information of a plurality of three-dimensional points included in a data unit is described.

[0502] The three-dimensional data decoding device analyzes the header of the coded data unit (coded data) to be processed, and analyzes the cabac_init_flag (S11411).

[0503] The three-dimensional data decoding device determines whether or not cabac_init_flag indicates that the context is to be continuously used (S11412).

[0504] If cabac_init_flag indicates that the context will continue to be used (Yes in S11412), that is, if cabac_init_flag is set to false, the three-dimensional data decoding device determines the encoding method of the encoded data to be processed (S11413).

[0505] If the encoding method of the encoded data to be processed is octet encoding (octet in S11413), the three-dimensional data decoding device continues to use the context used in the octet encoding, performs entropy decoding as the initial value of the context to be used in the octet encoding, and reconstructs and decodes the octet (S11414).

[0506] If the encoding method of the encoded data to be processed is predictive tree encoding (predictive tree in S11413), the three-dimensional data decoding device continues to use the context used in the predictive tree encoding, performs entropy decoding as the initial value of the context to be used in the predictive tree encoding, and reconstructs and decodes the predictive tree (S11415).

[0507] In this way, when the cabac_init_flag (identification information) indicates that the context used for encoding will continue to be used, the three-dimensional data decoding device decodes the encoded data by continuing to use the context used in the encoding method of the encoded data.

[0508] In addition, if the cabac_init_flag indicates that the context will not be used continuously (No in S11412), that is, if the cabac_init_flag is set to true, the three-dimensional data decoding device initializes the context for the specified encoding method, performs entropy decoding, and decodes using a decoding method corresponding to the specified encoding method (S11416).

[0509] In the above embodiment, a method for changing the continuing context depending on the encoding method of the position information (octree or prediction tree) has been described, but this method can also be applied to the encoding method of the attribute information. The encoding method of the attribute information includes, for example, an LoD-based encoding method and a Transform-based encoding method. In this case, the three-dimensional data encoding device may change the continuing context depending on the encoding method of the attribute information. That is, when performing LoD-based encoding, the three-dimensional data encoding device performs encoding using a context for LoD-based encoding, and when performing Transform-based encoding, the three-dimensional data encoding device performs encoding using a context for Transform-based encoding.

[0510] When signaling cabac_init_flag in encoding attribute information, it may be signaled independently for each of the LoD-based encoding method and the Transform-based encoding method, or may be signaled in common. That is, the three-dimensional data encoding device may store cabac_init_flag for each encoding method in the header, or may store cabac_init_flag common to multiple encoding methods in the header. When using one encoding method among multiple encoding methods, the three-dimensional data encoding device can reduce the amount of information to be signaled by commonizing (i.e., unifying) cabac_init_flag.

[0511] Furthermore, the cabac_init_flag for encoding the attribute information and the cabac_init_flag for encoding the position information may be set to the same value or different values.

[0512] When the cabac_init_flag for encoding the attribute information and the cabac_init_flag for encoding the position information are set to the same value, the cabac_init_flag for encoding the attribute information and the cabac_init_flag for encoding the position information may be common and stored in sequence common metadata such as SPS. In this case, when the three-dimensional data encoding device determines to continue using the context used in encoding in encoding, (i) it continues to use the context used in the encoding method of the multiple three-dimensional points among the multiple encoding methods to encode the position information of the multiple three-dimensional points, and (ii) it continues to use the context used in the encoding method of the attribute information to encode the attribute information of the multiple three-dimensional points. Conversely, when the three-dimensional data encoding device determines not to continue using the context used for encoding in encoding, it (i) encodes the position information of the multiple three-dimensional points using an initialized context for the encoding method of the multiple three-dimensional points among the multiple encoding methods, and (ii) encodes the attribute information of the multiple three-dimensional points using an initialized context for the encoding method of the attribute information.

[0513] In this case, when the identification information indicates that the context used for encoding will continue to be used, the three-dimensional data decoding device (i) calculates position information of the multiple three-dimensional points by continuing to use the context used for encoding the position information of the multiple three-dimensional points among the multiple encoding methods, and (ii) calculates attribute information of the multiple three-dimensional points by continuing to use the context used in the encoding method of the attribute information. Conversely, when the identification information indicates that the context used for encoding will not continue to be used, the three-dimensional data decoding device (i) calculates position information of the multiple encoded three-dimensional points by decoding using an initialized context for the encoding method used for encoding the position information of the multiple three-dimensional points among the multiple encoding methods, and (ii) decodes the attribute information of the multiple three-dimensional points by decoding using the initialized context for the encoding method of the attribute information.

[0514] When the cabac_init_flag for encoding attribute information and the cabac_init_flag for encoding position information are set to different values, the three-dimensional data encoding device stores the cabac_init_flag for encoding attribute information and the cabac_init_flag for encoding position information in the APS, GPS, data unit header, etc., respectively.

[0515] The three-dimensional data encoding device may store the cabac_init_flag for encoding the attribute information and the cabac_init_flag for encoding the position information in metadata common to the position information and the attribute information, or in either or both of the individual metadata, or in both the common metadata and the individual metadata. Furthermore, the three-dimensional data encoding device may use a flag indicating whether or not the cabac_init_flag for encoding the attribute information and the cabac_init_flag for encoding the position information are described.

[0516] In addition, when the three-dimensional data encoding device switches the encoding method between data units when encoding position information, it may decide to initialize the context of the data unit that is first encoded after the encoding method is switched rather than continuing it.

[0517] 70 is a diagram for explaining initialization of a context when the encoding method is switched. Fig. 70 shows an example in which the data unit of slice #1 is encoded by octree encoding, and the data units of slice #2 and slice #3 are encoded by predictive tree encoding.

[0518] The three-dimensional data encoding device sets the initialization flag (cabac_init_flag) used for encoding the position information of the first data unit (slice #1) of octree encoding to ON (true).The three-dimensional data encoding device then sets the initialization flag (cabac_init_flag) used for encoding the position information of the first data unit (slice #2) of predictive tree encoding to ON (true).The initialization flag for slice #3 may be set to ON or OFF.

[0519] In this way, when the encoding method of the first data unit and the encoding method of the second data unit to be encoded after the first data unit are different, the 3D data encoding device determines not to continue using the context used for encoding, and encodes the multiple 3D points of the second data unit using the initialized context for the encoding method of the second data unit from among the multiple encoding methods. In this case, the identification information (second identification information) corresponding to the second data unit is set to indicate that the context used for encoding will not be continued.

[0520] As described above, a three-dimensional data encoding device according to one aspect of the present embodiment performs the processing shown in FIG. 71 . The three-dimensional data encoding device acquires a first data unit including a plurality of first three-dimensional points (S11421). Next, the three-dimensional data encoding device encodes the plurality of first three-dimensional points included in the acquired first data unit using one of a plurality of different encoding methods (S11422). Next, the three-dimensional data encoding device generates a bitstream including first encoded data in which the plurality of first three-dimensional points are encoded and first identification information (S11423). In the encoding (S11422), it determines whether to continue encoding using the context used for encoding, and encodes the plurality of first three-dimensional points using a context corresponding to the result of the determination, from among the contexts used in the encoding method of the plurality of encoding methods. The first identification information includes the result of the determination.

[0521] This allows for improved coding efficiency by determining whether or not to continue the context used for coding, and also allows for appropriate decoding by the three-dimensional data decoding device by generating a bitstream including the first identification information.

[0522] For example, if it is determined in the encoding (S11422) that the context used for the encoding will continue to be used, the encoding (S11422) will encode the plurality of first three-dimensional points by continuing to use the context used in the encoding method of the plurality of first three-dimensional points among the plurality of encoding methods, and the first identification information indicates that the context used for the encoding will continue to be used for encoding.

[0523] For example, if it is determined in the encoding (S11422) that the context used for the encoding will not be continued, the encoding (S11422) encodes the first three-dimensional points using an initialized context for the encoding method of the first three-dimensional points from among the plurality of encoding methods, and the first identification information indicates that the context used for the encoding will not be continued.

[0524] For example, each of the plurality of first 3D points includes position information of the first 3D point and attribute information of the first 3D point. The plurality of encoding methods are encoding methods for position information. In the encoding (S11422), the attribute information of the plurality of first 3D points is encoded using another encoding method. In the encoding (S11422), if it is determined that a context used in the encoding is to be continued, (i) the position information of the plurality of first 3D points is encoded using the context used in the encoding method of the plurality of first 3D points among the plurality of encoding methods, and (ii) the attribute information of the plurality of first 3D points is encoded using the context used in the other encoding method.

[0525] For example, in the encoding (S11422), if it is determined not to continue using the context used for encoding in the encoding (S11422), (i) the position information of the first three-dimensional points is encoded using an initialized context for the encoding method of the first three-dimensional points among the multiple encoding methods, and (ii) the attribute information of the first three-dimensional points is encoded using an initialized context for the other encoding method.

[0526] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0527] Furthermore, a three-dimensional data decoding device according to one aspect of the present embodiment performs the processing shown in Fig. 72. The three-dimensional data decoding device acquires a bitstream including first coded data in which a plurality of first three-dimensional points are coded, and first identification information indicating whether or not the context used for coding will continue to be used (S11431). Next, the three-dimensional data decoding device decodes the first coded data using a decoding method corresponding to the coding method used for coding the first coded data, from among a plurality of different coding methods (S11432). In the decoding (S11432), the first coded data is decoded using a context corresponding to the first identification information.

[0528] This makes it possible to calculate a plurality of appropriate first 3D points by decoding the first coded data in accordance with the first identification information included in the bitstream.

[0529] For example, in the decoding (S11432), if the first identification information indicates that the context used for encoding will continue to be used, the first encoded data is decoded by continuing to use the context used in the encoding method corresponding to the decoding method.

[0530] For example, in the decoding (S11432), if the first identification information indicates that the context used for encoding will not continue to be used, the first encoded data is decoded using an initialized context for the encoding method used to encode the first encoded data among the multiple encoding methods.

[0531] For example, the first encoded data includes encoded position information of the plurality of first 3D points and attribute information of the plurality of first 3D points. The plurality of encoding methods are encoding methods for the encoded position information of the plurality of first 3D points. The encoded attribute information of the plurality of first 3D points is encoded using another encoding method. In the decoding (S11432), if the first identification information indicates that a context used for encoding is to be continuously used, (i) the position information of the plurality of first 3D points is calculated by decoding using a context used for encoding the encoded position information of the plurality of first 3D points among the plurality of encoding methods, and (ii) the attribute information of the plurality of first 3D points is calculated by decoding using a context used in the other encoding method.

[0532] For example, in the decoding (S11432), if the first identification information indicates that the context used for encoding will not be continued, (i) the position information of the encoded first three-dimensional points is calculated by decoding using an initialized context for an encoding method among the plurality of encoding methods that is used to encode the position information of the encoded first three-dimensional points, and (ii) the attribute information of the encoded first three-dimensional points is decoded by decoding using an initialized context for the other encoding method.

[0533] For example, the bitstream further includes second encoded data in which a plurality of second 3D points are encoded, and second identification information indicating whether or not to continue using a context used for encoding. The plurality of second 3D points are encoded next to the plurality of first 3D points. The second identification information indicates that the context used for encoding will not be continued.

[0534] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above processing using the memory.

[0535] When indicating whether to continue a context in entropy coding, it is preferable that, when continuing a context, it be ensured that the corresponding context is continued in the 3D data decoding device. For example, if continuing a context is indicated when the context is not saved or is not available, the 3D data decoding device will not be able to decode the 3D point cloud. Therefore, when continuing a context, the following constraints may be imposed to ensure that the 3D point cloud can be decoded in the 3D data decoding device.

[0536] [Coding Constraints] Examples of cases where the context used in coding changes include when switching coding methods such as predictive tree or octree, as well as when switching parameters in coding.

[0537] When parameters in encoding are switched, for example, in octet-tree encoding, the number of tree structures may be changed, such as when switching from an octet to a quad-tree or a binary tree. Also, when parameters in encoding are switched, the context for an occupancy code may be switched, such as when switching between bit-by-bit encoding in which a context is assigned to each bit of an 8-bit occupancy code and byte-by-byte encoding in which a context is assigned to the entire occupancy code. Also, when parameters in encoding are switched, in octet-tree encoding, the context for an occupancy code may be switched between referencing the same node and referencing adjacent nodes.

[0538] The flag indicating whether or not to switch these encoding parameters is indicated in the SPS (sequence parameter set), GPS (location information parameter set), data unit header, etc.

[0539] In this way, when the context to be used is switched, the three-dimensional data encoding device may reset the context first in the entropy encoding used for the data unit. In other words, the three-dimensional data encoding device may perform encoding using a reset (initialized) context rather than using a saved context.

[0540] As with the encoding of position information, when the context to be used is switched in encoding attribute information, the three-dimensional data encoding device may reset the context in the entropy encoding of the data unit.

[0541] Furthermore, in the encoding process, if entropy encoding is not reset, the encoding parameters may be restricted not to be switched.

[0542] For example, in encoding, the parameters related to encoding may be specified to be the same in a parameter set (e.g., GPS1) referenced by a "previous data unit DU_prev" that preserves the context and a parameter set (GPS2) referenced by a "current data unit DU_cur" that starts entropy encoding using the context of DU_prev.

[0543] Alternatively, the encoding may be specified so that the geometry parameter sets of GPS1 and GPS2 are the same and describe the same content, i.e., the identifier GPS_id in the parameter set is the same.

[0544] For example, in encoding, a parameter set (e.g., APS1) referred to by a "previous data unit DU_prev" that saves a context and a parameter set (APS2) referred to by a "current data unit DU_cur" that starts entropy encoding using the context of DU_prev may be specified so that the parameters related to encoding are the same.

[0545] Alternatively, in the encoding, the geometry parameter sets of APS1 and APS2 may be defined to be the same and to describe the same content, i.e., the identifier APS_id in the parameter set may be defined to be the same.

[0546] When continuing a context, the ID of the slice to which DU_cur belongs and the ID of the slice to which DU_prev belongs may be stored in the data unit header of DU_cur.

[0547] [Decoding Constraints] Furthermore, the following constraints may be imposed on the three-dimensional data decoding device.

[0548] When the context saved in DU_prev is used in the entropy coding of DU_cur, i.e., when cabac_init_flag = 0, the three-dimensional data decoding device determines whether the slice ID of the data unit decoded before decoding DU_cur matches DU_prev, and may decode it if they match. In other words, the three-dimensional data decoding device confirms that the target slice for which the context is saved on the three-dimensional data encoding device side matches the target slice for which the context is saved on the three-dimensional data decoding device side. Then, if the result of the confirmation shows that these slices do not match, the three-dimensional data decoding device may determine that there is a conformance violation (or specification non-compliance). Note that the bitstream conformance specification is a specification that defines the necessary conditions for the bitstream created by the encoder so that the bitstream can be properly decoded by the decoder, that is, it can also be said to be a specification for the decoder to check the constraints of the encoding method in the encoder. Therefore, by defining the encoding method that will allow the decoder to decode it correctly as the conformance specification for the bitstream, and determining whether the bitstream conforms to the specification in the decoder, it is possible to determine whether the bitstream can be decoded correctly.

[0549] Furthermore, when the context used changes due to switching of encoding parameters between the parameter set referenced by DU_cur and the parameter set referenced by DU_prev, the three-dimensional data decoding device may determine that there is a conformance violation (or non-compliance with the specifications). When determining that there is a conformance violation (or non-compliance with the specifications), the three-dimensional data decoding device may stop decoding or may implement specific avoidance processing.

[0550] Alternatively, the three-dimensional data decoding device may determine that there is a conformance violation (or specification non-compliance) if the GPS_id or APS_id of the parameter set referenced by DU_cur and the parameter set referenced by DU_prev are different.

[0551] 73 to 77 are diagrams showing an example of syntax. FIG. 73 is a diagram showing an example of the syntax of an SPS. FIG. 74 is a diagram showing an example of the syntax of the header (DividedGeometryHeader) of division position information. FIG. 75 is a diagram showing an example of the syntax of the header (DividedAttributeHeader) of division attribute information. FIG. 76 is a diagram showing another example of the syntax of the header (DividedAttributeHeader) of division attribute information. FIG. 77 is a diagram showing another example of the syntax of the header (DividedGeometryHeader) of division position information.

[0552] Previously, the flag indicating whether to initialize the context used for encoding was expressed as cabac_init_flag, but here it is expressed as entropy_continue_flag. The definition of the flag is reversed for entropy_continue_flag. That is, the initialization flag (entropy_continue_flag) is a flag indicating whether to continue entropy encoding without initialization, i.e., a flag indicating that the context used in encoding the previous data unit is saved and applied to the next data unit.

[0553] 73, the SPS indicates a flag (entropy_continue_enable) indicating whether or not a function for continuing a context between data units is provided (used). The flag (entropy_continue_enable) is an example of third identification information.

[0554] As shown in FIG. 74, the header of the divided position information (DividedGeometryHeader) includes a GPS identifier (gps_id), a tile identifier (tile_id), and a frame identifier (frame_id) referenced by the data unit including the divided position information.

[0555] Furthermore, if entropy_continue_enable indicated in the SPS is valid (that is, if entropy_continue_enable indicates that the SPS has the function of continuing a context), the header of the division position information (DividedGeometryHeader) indicates a flag (geom_du_entropy_continue_flag) indicating whether or not to continue using the context used in encoding the data unit preceding the data unit containing the division position information in encoding the data unit containing the division position information. That is, in this case, the header of the division position information (DividedGeometryHeader) includes the flag (geom_du_entropy_continue_flag). Furthermore, if the geom_du_entropy_continue_flag is valid, the header (DividedGeometryHeader) of the division position information indicates the slice ID (slice2_id) to which the previous data unit belongs. That is, the header (DividedGeometryHeader) of the division position information includes the slice ID (slice2_id). The flag (geom_du_entropy_continue_flag) is an example of first identification information.

[0556] As shown in Figure 75, the header of the divided attribute information (DividedAttributeHeader) includes an identifier (aps_id) of the APS (attribute information parameter set) referenced by the data unit containing the divided attribute information, a number (attr_index) indicating the attribute information in the order of the attribute information described in the SPS, and a slice ID (geom_slice_id) of the position information to which the attribute information corresponds.

[0557] Furthermore, if entropy_continue_enable indicated in the SPS is valid (that is, if entropy_continue_enable indicates that the SPS has the function of continuing a context), the header of the divided attribute information (DividedAttributeHeader) indicates a flag (attr_du_entropy_continue_flag) indicating whether or not to continue using the context used in encoding the data unit preceding the data unit containing the divided attribute information in encoding the data unit. That is, in this case, the header of the divided attribute information (DividedAttributeHeader) includes the flag (attr_du_entropy_continue_flag). Also, if the attr_du_entropy_continue_flag is valid, the header (DividedAttributeHeader) of the division attribute information indicates the slice ID (slice2_id) to which the previous data unit belongs. That is, the header (DividedAttributeHeader) of the division attribute information includes the slice ID (slice2_id). The flag (geom_du_entropy_continue_flag) is an example of first identification information.

[0558] As described above, if entropy_continue_enable indicated in the SPS is valid (that is, if entropy_continue_enable indicates that the SPS has the function of continuing a context), the header of the division position information (DividedGeometryHeader) includes a flag (geom_du_entropy_continue_flag), and the header of the division attribute information (DividedAttributeHeader) includes a flag (attr_du_entropy_continue_flag). In other words, the three-dimensional data encoding device determines whether to perform encoding using a context (second decision), and, if it has decided to perform encoding using a context, may determine whether to continue using the context used in encoding the data unit preceding the data unit to be processed in encoding the data unit to be processed (first decision). In this case, the 3D data encoding device generates a bitstream including entropy_continue_enable (third identification information) indicating whether or not to perform encoding using a context. Furthermore, geom_du_entropy_continue_flag (first identification information) or attr_du_entropy_continue_flag (first identification information) is indicated in the header of a data unit when entropy_continue_enable (third identification information) indicates that encoding using a context is to be performed.

[0559] Furthermore, since the header of the divided position information and the header of the divided attribute information include the geom_du_entropy_continue_flag and the attr_du_entropy_continue_flag, it is possible to control whether to continue a context individually for the position information and the attribute information, thereby enabling flexible control.

[0560] Also, as shown in FIG. 76, a flag (attr_du_entropy_continue_flag) indicating whether to continue the context used to encode attribute information may be defined as valid when entropy encoding of position information is continuing, and if valid, whether entropy encoding of the data unit of the attribute information is valid may be indicated in the header of the divided attribute information (DividedAttributeHeader).

[0561] In other words, when the three-dimensional data encoding device determines to continue using the context used in encoding the previous data unit, it may (i) continue to use the context used in encoding the position information of the previous data unit to encode the position information of the data unit to be processed, and (ii) continue to use the context used in encoding the attribute information of the previous data unit to encode the attribute information of the data unit to be processed. In this case, when the first identification information indicates that the context used in encoding the previous data unit will be continued to be used, the three-dimensional data decoding device (i) calculates position information of multiple three-dimensional points included in the data unit by continuing to use the context used in encoding the position information of the previous data unit for decoding, and (ii) calculates attribute information of the multiple three-dimensional points included in the data unit by continuing to use the context used in encoding the attribute information of the previous data unit for decoding.

[0562] Alternatively, if entropy_continue_enable is true, attr_du_entropy_continue_flag may be set in the header of the divided attribute information (DividedAttributeHeader), and if geom_du_entropy_continue_flag is false, attr_du_entropy_continue_flag may be set to false regardless of the value of attr_du_entropy_continue_flag.

[0563] Alternatively, it may be specified that a case where geom_du_entropy_continue_flag is false and attr_du_entropy_continue_flag is true is a conformance violation (or specification non-compliance).

[0564] Note that control of encoding of position information and control of encoding of attribute information may be unified. That is, geom_du_entropy_continue_flag and attr_du_entropy_continue_flag may be merged. In this case, attr_du_entropy_continue_flag does not need to be indicated in the header of the divided attribute information. In entropy encoding of the divided attribute information, whether to continue the context used in encoding the previous data unit is determined according to du_entropy_continue_flag indicated in the data unit header of the position information associated with geom_slice_id.

[0565] Note that if DU_cur and DU_prev refer to the same parameter set, DU_cur does not need to indicate a parameter set ID (GPS_id or APS_id). The three-dimensional data decoding device refers to DU_prev indicated in the header of DU_cur, and refers to the parameter set having the parameter set ID indicated in the DU_prev header.

[0566] Note that, if DU_cur and DU_prev are determined to belong to the same tile, tile_id does not need to be indicated. In other words, the header of the division position information does not need to include tile_id. This reduces the process of determining whether DU_cur and DU_prev are the same. It is also possible to prevent confusion such as indicating that entropy coding is continuing despite the context being switched.

[0567] It may also be specified that when du_entropy_continue_flag is false, at least one of gps_id and tile_id is indicated, and when du_entropy_continue_flag is true, gps_id and tile_id are not indicated.

[0568] Note that geom_entropy_continue_enable_flag is a flag indicating whether to continue the entropy coding context. It may be specified that the following condition 1 or condition 2 must be satisfied for geom_entropy_continue_enable_flag to be set to true. Condition 1 is to allow coding dependencies between slices (data units). Condition 2 is to not allow the order of slices (data units) to be changed.

[0569] The order of slices (data units) may be indicated, for example, by the ID (slice ID) of the data unit. The ID of the data unit may be an ID (number) for identifying the data unit on a frame-by-frame basis, or may be an ID (number) for identifying the data unit on a random access unit basis. The first data unit of a random access unit is a data unit assigned a predetermined ID (predetermined number). In this way, each data unit is assigned its order in the random access unit. In other words, if the three-dimensional data encoding device determines to continue using the context used in encoding the previous data unit to satisfy condition 2, it may not rearrange the order of multiple data units. In this case, the three-dimensional data encoding device may further generate a bitstream including second identification information indicating whether rearranging the order of multiple data units in a random access unit is permitted. Furthermore, the three-dimensional data decoding device may determine that the acquired bitstream complies with the conformance if the du_entropy_continue_flag (first identification information) indicates that the context is not continued and the second identification information indicates that reordering is not permitted. Furthermore, the three-dimensional data decoding device may determine that the acquired bitstream does not comply with the conformance if the du_entropy_continue_flag (first identification information) indicates that the context is not continued or if the second identification information indicates that reordering is permitted.

[0570] In this case, whether geom_entropy_continue_enable_flag is valid may be indicated when condition 1 or condition 2 is satisfied. Alternatively, when geom_entropy_continue_enable_flag is true and condition 1 or condition 2 is not satisfied, a conformance violation (or specification non-compliance) may be indicated. Furthermore, geom_entropy_continue_enable_flag and the flag indicating condition 1 or condition 2 may be merged and replaced with either flag.

[0571] Furthermore, if the data unit to be processed is the first data unit of a random access unit, it may be determined that du_entropy_continue_flag is set to false for the random access unit. That is, in this case, the three-dimensional data encoding device may determine not to continue using the context used in encoding the previous data unit. Furthermore, if the data unit to be processed is not the first data unit of a random access unit, the three-dimensional data encoding device may determine to continue using the context used in encoding the previous data unit.

[0572] Furthermore, the three-dimensional data decoding device may determine that the bitstream conforms if the du_entropy_continue_flag (first identification information) indicates that the context used in encoding the previous data unit is not continued and the data unit is the first data unit of a random access unit. In other words, the three-dimensional data decoding device determines that, when the data unit is the first data unit of a random access unit, the conformance condition of the bitstream is satisfied, that the first identification information indicates that the context used in encoding the previous data unit is not continued (i.e., the conformance condition of the bitstream is satisfied). Furthermore, the three-dimensional data decoding device may determine that the bitstream does not conform to the conformance if the du_entropy_continue_flag (first identification information) indicates that the context used in encoding the previous data unit is continued or if the data unit is not the first data unit of a random access unit.

[0573] For example, the random access unit may be a single frame. In this case, the three-dimensional data encoding device may determine that the context used to encode the previous data unit is not continued to be used to encode the first slice (data unit) of a frame. Furthermore, the three-dimensional data decoding device may determine that the bitstream complies with the conformance when the du_entropy_continue_flag (first identification information) indicates that the context used to encode the previous data unit is not continued to be used and the data unit is the first data unit of a frame.

[0574] Furthermore, for example, the random access unit may be a group of frames (GOF). In this case, the three-dimensional data encoding device may determine that the context used to encode the previous data unit is not continued when encoding the first slice (data unit) of the GOF. Furthermore, the three-dimensional data decoding device may determine that the bitstream complies with the conformance when the du_entropy_continue_flag (first identification information) indicates that the context used to encode the previous data unit is not continued and the data unit is the first data unit of the GOF.

[0575] Alternatively, the random access unit may be a single tile. In this case, the three-dimensional data encoding device may determine that the context used in encoding the previous data unit is not continued to be used in encoding the first slice (data unit) of a tile. Furthermore, the three-dimensional data decoding device may determine that the bitstream complies with the conformance when du_entropy_continue_flag (first identification information) indicates that the context used in encoding the previous data unit is not continued to be used and the data unit is the first data unit of a tile.

[0576] In addition, a random access point flag indicating whether the data unit is the first data unit (random access point) may be set in the parameter set or header referenced by the data unit. In this case, the du_entropy_continue_flag may be enabled in the header (i.e., may be included in the header) when the random access point flag indicates that the data unit is the first data unit. Furthermore, the du_entropy_continue_flag and the random access point flag may be merged.

[0577] FIG. 78 is a flow diagram showing an example of a first decision for determining whether to initialize entropy coding in a three-dimensional data coding device.

[0578] The three-dimensional data encoding device determines whether to continue entropy encoding (S11901).

[0579] Next, the three-dimensional data encoding device executes processing for each slice (data unit) (S11902).

[0580] Next, the three-dimensional data encoding device determines whether the slice to be processed is a random access point (S11903). A random access point is the first slice of a frame if random access is possible on a frame-by-frame basis, the first slice of a GOF if random access is possible on a multi-frame basis (GOF), or the first slice of a tile if random access is possible on a tile-by-tile basis.

[0581] When the three-dimensional data encoding device determines that the slice to be processed is not a random access point (No in S11903), it determines whether the context to be used for encoding the slice to be processed is the same as the context to be used for encoding the previous slice (S11904). Note that the three-dimensional data encoding device may determine whether the context to be used for encoding changes (i.e., the context is different) based on, for example, a flag indicating whether the tree structure is an octtree or a predictive tree, a flag indicating whether a 2- or 4-ary tree is used, or a flag indicating whether a bit-wise context is used.

[0582] If the three-dimensional data encoding device determines that the context used to encode the slice to be processed is the same as the context used to encode the previous slice (Yes in S11904), it determines whether to initialize the context (S11905).

[0583] When it is determined that the context should not be initialized (No in S11905), the three-dimensional data encoding device determines to continue using the context without initializing it (S11906).

[0584] On the other hand, the three-dimensional data encoding device decides to initialize the context (i.e., not continue) (S11907) if it determines that the slice to be processed is a random access point (Yes in S11903), if it determines that the context to be used to encode the slice to be processed is not the same as the context used to encode the previous slice (No in S11904), or if it determines that the context should be initialized (Yes in S11905).

[0585] FIG. 79 is a flow diagram showing an example of processing for determining whether an entropy coding flag complies with conformance (whether it complies with specifications) in a three-dimensional data decoding device.

[0586] The three-dimensional data decoding device analyzes the header of each slice (data unit) (S11911).

[0587] Next, the three-dimensional data decoding device determines whether or not du_entropy_continue_flag (first identification information) is true (S11912).

[0588] If the three-dimensional data decoding device determines that du_entropy_continue_flag (first identification information) is true (Yes in S11912), it determines whether the GPS identifier (gps_id) referenced by the slice to be processed is the same as the GPS identifier (gps_id) referenced by the previous slice (S11913).

[0589] If the three-dimensional data decoding device determines that the GPS identifier (gps_id) referenced by the slice to be processed is not the same as the GPS identifier (gps_id) referenced by the previous slice (No in S11913), it determines whether the slice to be processed is the beginning of a random access unit (S11914). If the frame unit is random access, the three-dimensional data decoding device may detect the frame boundary when it detects the presence of a data unit indicating a frame boundary or when it detects a change in frame index.

[0590] If the three-dimensional data decoding device determines that the slice to be processed is not the first slice in the random access unit (No in S11914), it determines that the slice violates conformance (does not comply with specifications) (S11915).

[0591] If the three-dimensional data decoding device determines that du_entropy_continue_flag (first identification information) is not true (i.e., false) (No in S11912), if it determines that the GPS identifier (gps_id) referenced by the slice to be processed is the same as the GPS identifier (gps_id) referenced by the previous slice (Yes in S11913), or if it determines that the slice to be processed is the beginning of a random access unit (Yes in S11914), it determines that the slice is in conformance (specification compliant) (S11916).

[0592] As described above, the three-dimensional data encoding device according to one aspect of the present embodiment performs the processing shown in FIG. 80 . The three-dimensional data encoding device acquires a data unit including a plurality of three-dimensional points (S11921). Next, for each of the data units, the three-dimensional data encoding device encodes the plurality of three-dimensional points included in the data unit (S11922). The three-dimensional data encoding device generates a bitstream including encoded data resulting from encoding the data units (S11925). During the encoding (S11922), the three-dimensional data encoding device makes a first determination to determine whether to continue to use the context used to encode the data unit preceding the data unit (S11923). Then, following S11923, the three-dimensional data encoding device encodes the data unit using a context according to the result of the first determination (S11924). In the first decision (S11923), the three-dimensional data encoding device decides not to continue using the context used to encode the previous data unit if the data unit is the first data unit of a random access unit.

[0593] According to this, if the data unit to be processed is the first data unit of a random access unit, it is possible to determine not to continue using the context used to encode the previous data unit, thereby allowing the three-dimensional data decoding device to properly decode the bitstream.

[0594] For example, the random access unit is one frame unit.

[0595] For example, the random access unit is a unit of a plurality of frames.

[0596] For example, the random access unit is one tile unit.

[0597] For example, the data unit is assigned with the order of the data unit in the random access unit. In the first decision (S11923), if the three-dimensional data encoding device determines to continue using the context used in encoding the previous data unit, it does not rearrange the order of the multiple data units in the random access unit.

[0598] For example, in the encoding, the three-dimensional data encoding device further makes a second decision to determine whether or not to perform encoding using the context. If it is determined in the second decision that encoding using the context is to be performed, the three-dimensional data encoding device makes the first decision (S11923).

[0599] For example, each of the plurality of 3D points includes position information of the respective 3D points and attribute information of the respective 3D points. In the encoding (S11922), if the 3D data encoding device determines in the first determination (S11923) to continue to use the context used in encoding the previous data unit, it (i) encodes the position information of the data unit by continuing to use the context used in encoding the position information of the previous data unit, and (ii) encodes the attribute information of the data unit by continuing to use the context used in encoding the attribute information of the previous data unit.

[0600] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0601] Furthermore, a three-dimensional data decoding device according to one aspect of the present embodiment performs the processing shown in Fig. 81. The three-dimensional data decoding device acquires a bitstream including coded data in which a data unit including a plurality of three-dimensional points is coded, and first identification information indicating whether or not the context used in coding the data unit preceding the data unit is to be continued in coding the data unit (S11931). The three-dimensional data decoding device decodes the coded data using a context according to the first identification information (S11932). In the decoding (S11932), if the data unit is the first data unit of a random access unit, the three-dimensional data decoding device determines that the conformance condition of the bitstream is satisfied (i.e., the conformance condition of the bitstream is satisfied) when the first identification information indicates that the context used in coding the previous data unit is not to be continued.

[0602] Therefore, if the three-dimensional data decoding device determines that the bitstream conforms to the conformance, it can appropriately decode the bitstream by, for example, continuing the decoding. Also, if the three-dimensional data decoding device determines that the bitstream does not conform to the conformance, it can prevent the bitstream from being inappropriately decoded by, for example, stopping the decoding or executing a specific avoidance process.

[0603] For example, the random access unit is one frame unit.

[0604] For example, the random access unit is a unit of a plurality of frames.

[0605] For example, the random access unit is one tile unit.

[0606] For example, the data units are assigned with an order of the data units in the random access unit. The bitstream further includes second identification information indicating whether or not rearranging the order of the multiple data units in the random access unit is permitted. In the decoding (S11932), the 3D data decoding device determines that the bitstream complies with the conformance if the first identification information indicates that the context is not continuously used and the second identification information indicates that rearranging the order is not permitted.

[0607] For example, the bitstream further includes third identification information indicating whether or not the encoding using the context is to be performed, and the first identification information is indicated in a header of the data unit when the third identification information indicates that the encoding using the context is to be performed.

[0608] For example, the encoded data includes position information of the encoded three-dimensional points and attribute information of the encoded three-dimensional points. In the decoding (S11932), if the first identification information indicates that the context used in encoding the previous data unit is to be continued to be used, the three-dimensional data decoding device (i) calculates the position information of the multiple three-dimensional points included in the data unit by continuing to use the context used in encoding the position information of the previous data unit while decoding, and (ii) calculates the attribute information of the multiple three-dimensional points included in the data unit by continuing to use the context used in encoding the attribute information of the previous data unit while decoding.

[0609] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above processing using the memory.

[0610] In this embodiment, another example of entropy coding will be described. Fig. 82 is a diagram showing an example of the syntax of SPS. Fig. 83 is a diagram showing an example of the syntax of APS. Fig. 84 is a diagram showing an example of the syntax of the header (DividedGeometryHeader) of division position information. Fig. 85 is a diagram showing an example of the syntax of the header (DividedAttributeHeader) of division attribute information.

[0611] In entropy coding, a switching flag (identification information) for switching whether or not to enable the continuation function of the context used in entropy coding may be provided for each data unit of attribute information.

[0612] The switching flag may be provided in a parameter set (SPS or APS) referenced by the attribute information. Specifically, the switching flag (entropy_continue_attr_enable_flag) may be indicated in the SPS as shown in FIG. 82, or may be indicated in the APS as shown in FIG. 83. The entropy_continue_attr_enable_flag is a flag indicating whether or not to enable the continuation function of the context used in encoding the attribute information. The entropy_continue_attr_enable_flag may be indicated when the entropy_continue_enable_flag indicated in the SPS is true. The entropy_continue_enable_flag is a flag indicating whether or not to enable the continuation function of the context. The Sps. entropy_continue_enable_flag means the entropy_continue_enable_flag signaled in the SPS.

[0613] Also, as shown in Fig. 84, the header of the division position information (DividedGeometryHeader) may indicate a flag (geom_du_entropy_continue_flag) indicating whether or not to continue the context for each data unit. Also, as shown in Fig. 85, the header of the division attribute information (DividedAttributeHeader) does not need to indicate information regarding the continuation of the context.

[0614] Using the syntax shown in Figures 82 to 85, the continuation flag for each data unit (position information and attribute information) may be derived as shown in the flow diagram of Figure 86. Figure 86 is a flow diagram showing an example of processing for determining whether or not to continue a context used for entropy encoding in a three-dimensional data encoding device.

[0615] First, the three-dimensional data encoding device determines whether or not the entropy_continue_enable_flag of the SPS indicates true (S12101).

[0616] If the entropy_continue_enable_flag of the SPS indicates true (Yes in S12101), the three-dimensional data encoding device determines whether the geom_du_entropy_continue_flag of the header (DividedGeometryHeader) of the division position information is true (S12102).

[0617] If the geom_du_entropy_continue_flag in the header (DividedGeometryHeader) of the divided position information indicates true (Yes in S12102), the three-dimensional data encoding device determines, at the start of encoding the data unit of position information, to use the context saved (stored or saved) in the storage device (memory) for the data unit of previous position information (S12103). In other words, in this case, the three-dimensional data encoding device determines to continue using the context used for encoding the data unit of previous position information.

[0618] Next, the three-dimensional data encoding device determines whether the entropy_continue_attr_enable_flag of the SPS or APS is true (S12104).

[0619] If the entropy_continue_attr_enable_flag of the SPS or APS indicates true (Yes in S12104), the three-dimensional data encoding device determines to use the context stored in the storage device (memory) for the previous attribute information data unit when starting to encode the attribute information data unit (S12105). That is, in this case, the three-dimensional data encoding device determines to continue to use the context used for encoding the previous attribute information data unit.

[0620] If the entropy_continue_enable_flag of the SPS indicates false (No in S12101), or if the geom_du_entropy_continue_flag of the header (DividedGeometryHeader) of the divided position information indicates false (No in S12102), the three-dimensional data encoding device decides to initialize the context at the start of encoding the data unit of position information (S12106). In other words, in this case, the three-dimensional data encoding device decides to initialize the context without continuing to use the context used in encoding the previous data unit of position information.

[0621] Then, if the entropy_continue_attr_enable_flag of the SPS or APS indicates false (No in S12104), or after step S12106, the three-dimensional data encoding device determines to initialize the context at the start of encoding the data unit of attribute information (S12107). That is, in this case, the three-dimensional data encoding device determines to initialize the context without continuing to use the context used in encoding the previous data unit of attribute information.

[0622] When the entropy_continue_attr_enable_flag is set in the SPS, the three-dimensional data encoding device can commonly control all attribute components such as colors, reflectances, etc. When the entropy_continue_attr_enable_flag is set in the APS, the three-dimensional data encoding device can individually control all attribute components such as colors, reflectances, etc.

[0623] The determination of whether to continue the context of the data unit of attribute information may depend on the determination result of whether to continue the context of each data unit of location information. In this case, the combination (Gometry DU, Attribute DU) of whether to continue the context of the data unit of location information (ON / OFF) and whether to continue the context of the data unit of attribute information (ON / OFF) can be any combination of (ON, ON), (ON, OFF), and (OFF, OFF).

[0624] In entropy coding, the previous data unit or the context used in coding the data unit is stored in a storage device (memory), and the stored context is applied to the coding of the next data unit, thereby improving coding performance. When entropy coding is performed bit-wise and a bit-wise context is used, there is an entropy context to continue in the bit-wise entropy coding. However, when entropy coding is performed byte-wise, there is no byte-wise context, and therefore there is no entropy context to continue in the byte-wise entropy coding. Therefore, the methods described so far do not provide a way to continue byte-wise entropy coding.

[0625] In this embodiment, a method for continuing entropy coding when an occupancy code is coded in units of bytes in coding position information using an N-ary tree (N is an integer equal to or greater than 2, for example, an occupancy tree) will be described. Also, a method for switching the method for continuing entropy coding depending on the coding method (bit-by-bit coding or byte-by-bit coding) will be described.

[0626] In byte-by-byte encoding, the three-dimensional data encoding device uses a lookup table to convert the occupancy code into table index information and encodes the converted index information. Here, as described in the above embodiment, the occupancy code is 8-bit information that indicates, when a three-dimensional point cloud is represented as an occupancy tree, at what position after division at a certain node is the next node or leaf located. Hereinafter, the occupancy code may also be referred to as an occupancy map. Hereinafter, the occupancy code will be referred to as an occupancy map.

[0627] FIG. 87 is a diagram for explaining table updating.

[0628] The three-dimensional data encoding device converts the occupancy map into a table showing the relationship between a histogram indicating the number of times the occupancy map occurs and a dictionary index indicating the order of the number of times the occupancy map occurs, using table 12101 shown in Fig. 87. In Fig. 87, the occupancy map is indicated by m, the dictionary index is indicated by d, and the histogram is indicated by h. Also, Fig. 87 shows an example of a table with three types of occupancy maps, 15, 25, and 35, but the table is not limited to this and may have four or more types of occupancy maps.

[0629] The three-dimensional data encoding device updates the table each time it encodes an occupancy map. When encoding an occupancy map, the three-dimensional data encoding device adds 1 to the histogram value of the corresponding occupancy map. For example, as shown in (a) of Fig. 87, if the input occupancy map indicates 25, the three-dimensional data encoding device adds 1 to the histogram value corresponding to the occupancy map of 25.

[0630] Next, the three-dimensional data encoding device updates the dictionary index based on the updated histogram. That is, the three-dimensional data encoding device assigns (sets) the order of occurrence counts as a dictionary index based on the occurrence counts of the updated occupancy map. Note that the dictionary index may indicate either the occurrence count of the occupancy map or the occurrence frequency of the occupancy map.

[0631] For example, as shown in (a) of Figure 87, when the three-dimensional data encoding device updates the histogram values, it assigns indexes in descending order of the updated histogram values, so that the index corresponding to an occupancy map of 15 is 2, the index corresponding to an occupancy map of 25 is 1, and the index corresponding to an occupancy map of 35 is 3. Note that the histogram for the occupancy map of 15 and the histogram for the occupancy map of 35 have the same value, but in this case the index for the smaller occupancy map may be assigned a smaller value. Note that this is not a limitation, and the index for the smaller occupancy map may also be assigned a larger value.

[0632] In (b) of Figure 87, 25 occupancy maps are input, in each of (c) to (f) of Figure 87, 35 occupancy maps are input, and in (g) of Figure 87, 15 occupancy maps are input. When each occupancy map is input, the three-dimensional data encoding device adds 1 to the histogram value corresponding to the input occupancy map, and assigns dictionary indices in descending order of the calculated histogram values. Therefore, a smaller dictionary index value indicates a higher occurrence frequency (input frequency) of the occupancy map.

[0633] Next, the three-dimensional data encoding device encodes index information indicating the set dictionary index. In byte-by-byte encoding, the three-dimensional data encoding device can reduce the amount of code by converting the occupancy map into index information.

[0634] In this way, the three-dimensional data encoding device derives index information for the occupancy map using the table, encodes the derived index information, and then updates the table stored in the storage device (memory) to a table indicating the relationship between the updated histogram and the set index.

[0635] The three-dimensional data decoding device decodes the index information contained in the encoded bitstream from the encoded bitstream, and derives an occupancy map using a table stored in a storage device (memory) in the same manner as the three-dimensional data encoding device.Then, the table is updated in the same manner as the three-dimensional data encoding device.

[0636] The histogram and index information in the table may be updated until the encoding of the occupancy map for a slice (data unit) is completed, and then initialized at the beginning of the next slice. Alternatively, the table used at the end of a slice may be stored, and the stored table may continue to be used in the encoding of the next slice. Byte-wise entropy coding can be expected to improve encoding by carrying over the table learned in the previous slice to the encoding of the next slice. Thus, bit-wise coding differs from byte-wise coding in that the context is carried over in bit-wise coding, while the table is carried over in byte-wise coding.

[0637] Although the above description takes byte-by-byte encoding as an example, the present invention is not limited to this. The present invention can be similarly applied to cases where a table is used for a histogram (such as an occupancy map) that counts the occurrence of values ​​without using a context, or to additional information used for encoding, such as other learning parameters. This method can also be applied to encoding attribute information in addition to encoding position information. Even when additional information is stored in a storage device (memory) and the additional information stored in the next data unit is applied, an improvement in the encoding rate can be expected.

[0638] FIG. 88 is a flow diagram of the encoding of an occupancy map by a three-dimensional data encoding device.

[0639] The three-dimensional data encoding device converts three-dimensional points into an N-ary tree representation for each data unit and starts encoding an occupancy map for each node (S12111). The three-dimensional data encoding device generates an occupancy map by converting the position information of each of multiple three-dimensional points included in the data unit to be processed into an N-ary tree representation (e.g., an occupancy map representation). In other words, the three-dimensional data encoding device converts multiple pieces of position information of multiple three-dimensional points included in the data unit to be encoded into multiple occupancy maps using an octree.

[0640] Next, the three-dimensional data encoding device generates index information corresponding to the occupancy map using the table, and encodes the generated index information (S12112). The three-dimensional data encoding device converts each of the multiple occupancy maps into an index using a table that indicates the correspondence between the occupancy map and the index, and encodes the index to generate encoded data.

[0641] Next, the three-dimensional data encoding device updates the histogram and index using the occupancy map, thereby updating the table stored in the storage device (memory) (S12113). The three-dimensional data encoding device updates the table according to the converted index and stores it in the memory.

[0642] FIG. 89 is a flow diagram of decoding an occupancy map by a three-dimensional data decoding device.

[0643] The three-dimensional data decoding device starts a decoding process for each data unit in the bit stream (S12121). The bit stream includes, for example, encoded data in which a data unit including a plurality of three-dimensional points is encoded, and first identification information indicating whether or not to initialize and use, in encoding the data unit, a table used in encoding a data unit preceding the data unit.

[0644] Next, the three-dimensional data decoding device decodes the coded index information included in the bitstream, and derives an occupancy map for the index indicated by the index information using the decoded index information and a table stored in a storage device (memory) (S12122). In other words, the three-dimensional data decoding device calculates the position information of the three-dimensional point by deriving an occupancy map in the table that corresponds to the index obtained by decoding the coded data.

[0645] Next, the three-dimensional data decoding device updates the histogram and index using the occupancy map, thereby updating the table stored in the storage device (memory) (S12123).

[0646] FIG. 90 is a flowchart showing the process of switching the entropy encoding method in the three-dimensional data encoding device.

[0647] The three-dimensional data encoding device determines whether the encoding method is bit-wise or byte-wise (S12131).

[0648] When the three-dimensional data encoding device determines that the encoding method is bit-wise ("bit-wise" in S12131), it performs encoding using a bit-wise entropy encoding continuation method (S12132).

[0649] The three-dimensional data encoding device sets a flag (bit-wise_flag) indicating whether the encoding method is bit-wise or byte-wise to true (S12133). The three-dimensional data encoding device generates a bitstream including the flag and transmits it to the three-dimensional data decoding device.

[0650] When the three-dimensional data encoding device determines that the encoding method is byte-based ("byte-based" in S12131), it performs encoding using a byte-based entropy encoding continuation method (S12134).

[0651] The three-dimensional data encoding device sets a flag (bit-wise_flag) indicating whether the encoding method is bit-wise or byte-wise to false (S12133). The three-dimensional data encoding device generates a bitstream including the flag and transmits it to the three-dimensional data decoding device.

[0652] FIG. 91 is a flow diagram of a method for continuing byte-by-byte entropy encoding in a three-dimensional data encoding device.

[0653] The three-dimensional data encoding device determines whether to initialize entropy encoding (S12141). The three-dimensional data encoding device determines whether to initialize a table used for byte-by-byte entropy encoding, that is, whether to continue using the table.

[0654] If the three-dimensional data encoding device determines not to initialize entropy encoding (No in S12142), that is, if it determines to continue using the table, it continues to use the table stored in the storage device (memory) when encoding the previous data unit to perform encoding (S12142).

[0655] Next, the three-dimensional data encoding device sets cabac_init_flag to false (S12143). That is, the three-dimensional data encoding device sets the flag (identification information) indicating whether or not the table has been continued to a value indicating that the table has been continued.

[0656] If the three-dimensional data encoding device determines to initialize entropy encoding (Yes in S12142), that is, if it determines that the table will not continue to be used, it initializes the table stored in the storage device (memory) when encoding the previous data unit and performs encoding (S12144).

[0657] Next, the three-dimensional data encoding device sets cabac_init_flag to true (S12145). That is, the three-dimensional data encoding device sets the flag (identification information) indicating whether or not the table was continued to a value indicating that the table was not continued.

[0658] The three-dimensional data encoding device updates the table in accordance with the entropy encoding, and stores the updated table in the storage device (memory) (S12146).

[0659] FIG. 92 is a flowchart showing the process of switching the entropy decoding method in the three-dimensional data decoding device.

[0660] The three-dimensional data decoding device analyzes the bit-wise_flag corresponding to the data unit to be decoded that is included in the bitstream (S12151).

[0661] Next, the three-dimensional data decoding device determines, based on the analysis result, whether the encoding method of the data unit to be decoded is bit-wise or byte-wise (S12152). In other words, the three-dimensional data decoding device determines whether the bit-wise_flag corresponding to the data unit to be decoded indicates true.

[0662] If the encoding method of the data unit to be decoded is bitwise, that is, if bit-wise_flag indicates true, the three-dimensional data decoding device performs decoding using a bitwise entropy encoding continuation method (S12153).

[0663] If the encoding method of the data unit to be decoded is byte-wise, that is, if bit-wise_flag indicates false, the three-dimensional data decoding device performs decoding using a continuation method of byte-wise entropy encoding (S12154).

[0664] FIG. 93 is a flow diagram of a method for continuing entropy decoding in units of bytes in a three-dimensional data decoding device.

[0665] The three-dimensional data decoding device analyzes the cabac_init_flag included in the bitstream (S12161).

[0666] Next, the three-dimensional data decoding device determines whether or not cabac_init_flag indicates true (S12162).

[0667] If cabac_init_flag indicates false (No in S12162), the three-dimensional data decoding device continues to use the table stored in the storage device (memory) in the decoding of the previous data unit to perform decoding (S12163).

[0668] If cabac_init_flag indicates true (Yes in S12162), the three-dimensional data decoding device initializes the table saved in the storage device (memory) in the decoding of the previous data unit, and executes decoding (S12164).

[0669] The three-dimensional data decoding device updates the table in accordance with the entropy decoding, and stores the updated table in the storage device (memory) (S12165).

[0670] As described above, a three-dimensional data encoding device according to one aspect of the present embodiment performs the processing shown in Fig. 94. The three-dimensional data encoding device acquires a plurality of data units, each of which includes a plurality of three-dimensional points (S12171). Next, the three-dimensional data encoding device encodes a plurality of three-dimensional points included in each of the plurality of data units (S12172). Next, the three-dimensional data encoding device generates a bitstream including encoded data in which the plurality of three-dimensional points have been encoded (S12173). In the encoding (S12172), the 3D data encoding device converts multiple pieces of position information of multiple 3D points included in the data unit to be encoded into multiple occupancy maps using an N-ary tree, converts each of the multiple occupancy maps into an index using a table indicating the correspondence between the occupancy map and the index, generates the encoded data by encoding the index, updates the table according to the converted index and stores it in memory, and when encoding the first 3D point of the data unit following the data unit to be encoded, determines whether to initialize the table stored in the memory, and if it is determined not to initialize, starts encoding the next data unit using the table stored in the memory. The bitstream further includes first identification information indicating the result of the determination.

[0671] According to this, when encoding an occupancy map with converted position information, the index obtained using the table is encoded and a bitstream containing first identification information indicating whether or not to initialize the table used for encoding is generated, thereby allowing the three-dimensional data decoding device to appropriately decode the bitstream.

[0672] For example, the index indicates either the number of occurrences of the occupancy map or the frequency of occurrence of the occupancy map. Therefore, for example, the coding efficiency can be improved by setting the index to a smaller value as the number of occurrences and the frequency of occurrence are larger.

[0673] For example, if the first identification information indicates that the table is to be initialized, it indicates that the context of the previous data unit is to be initialized and the multiple attribute information of the multiple 3D points is to be encoded, and if the first identification information indicates that the table is not to be initialized, it indicates that the context of the previous data unit is to continue to be used to encode the multiple attribute information.

[0674] For example, the bitstream further includes second identification information indicating whether or not a function for continuing entropy between the plurality of data units is used, and the first identification information is indicated when the second identification information indicates that a function for continuing entropy between the plurality of data units is used.

[0675] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0676] Furthermore, a three-dimensional data decoding device according to one aspect of the present embodiment performs the processing shown in FIG. 95 . The three-dimensional data decoding device acquires a bitstream including coded data in which a data unit including a plurality of three-dimensional points is coded, and first identification information indicating whether or not the table used to code the data unit before the data unit is initialized and used to code the data unit (S12181). The three-dimensional data encoding device decodes the coded data using a table corresponding to the first identification information (S12182). The table indicates a correspondence between an occupancy map, in which position information of three-dimensional points is expressed using an N-ary tree, and an index. The coded data includes the coded index. In the decoding step (S12182), the three-dimensional data decoding device calculates position information of the three-dimensional points by deriving an occupancy map in the table corresponding to the index obtained by decoding the coded data.

[0677] This allows an occupancy map to be derived using an index obtained by decoding the encoded data and a table corresponding to the first identification information, thereby allowing the three-dimensional data decoding device to appropriately decode the bit stream.

[0678] For example, the index indicates either the number of occurrences of the occupancy map or the frequency of occurrence of the occupancy map.

[0679] For example, if the first identification information indicates that the table is to be initialized, it indicates that the context of the previous data unit is to be initialized and the multiple attribute information of the multiple 3D points is to be encoded, and if the first identification information indicates that the table is not to be initialized, it indicates that the context of the previous data unit is to continue to be used to encode the multiple attribute information.

[0680] For example, the bitstream further includes second identification information indicating whether or not a function for continuing entropy between the plurality of data units is used, and the first identification information is indicated when the second identification information indicates that a function for continuing entropy between the plurality of data units is used.

[0681] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above processing using the memory.

[0682] When point cloud data is divided into multiple data units (slices), if CABAC is initialized, the data units are not dependent on each other and can be encoded or decoded independently. On the other hand, the current data structure of the data units (slices) does not support a function for processing data in parallel.

[0683] Therefore, in predictive tree coding, a function is provided that allows multiple data units within a data unit (slice) to be processed in parallel by adding a function to initialize the context in units of predictive trees within the data of each data unit (slice) and information to access the units of predictive trees.

[0684] Fig. 96 is a diagram showing an example of a 3D point cloud when encoding slices divided into groups. Fig. 97 is a diagram showing various examples of bitstream configurations.

[0685] 96, the three-dimensional point cloud may be divided into a plurality of data units 11401 to 11404 and 11411 to 11413. Furthermore, among the plurality of data units 11401 to 11404 and 11411 to 11413, a plurality of data units 11401 to 11404 may be classified into group 1, and a plurality of data units 11411 to 11413 may be classified into group 2.

[0686] Next, the relationship between slices and prediction trees will be described with reference to FIG.

[0687] The three-dimensional data encoding device may encode the data units of one slice with one prediction tree as in bit stream 1, or may encode the data units of one slice with multiple prediction trees as in bit stream 2. Furthermore, for example, if it is possible to cluster or group point clouds based on the characteristics of the point clouds, the three-dimensional data encoding device may encode each group by dividing them into slices as in bit stream 4, or may encode without dividing them into slices as in bit stream 3. When encoding without dividing them into bit stream slices, the three-dimensional data encoding device may rearrange the point clouds so that they are in the order of each group, and encode each group with a prediction tree.

[0688] Fig. 98 shows an example in which a slice flag (slice_cabac_init_flag) indicates whether or not to initialize CABAC for each slice, and a tree flag (tree_cabac_init_flag) indicates whether or not to initialize CABAC for each tree within the slice. Note that bitstreams 1 to 4 in Fig. 98 are the same as bitstreams 1 to 4 shown in Fig. 97, respectively.

[0689] When initializing CABAC at the beginning of each processing unit, the three-dimensional data encoding device sets slice_cabac_init_flag or tree_cabac_init_flag to 1 and transmits the slice_cabac_init_flag or tree_cabac_init_flag set to 1 as metadata. Note that slice_cabac_init_flag or tree_cabac_init_flag set to 1 indicates that CABAC is initialized at the beginning of each processing unit. slice_cabac_init_flag is an initialization flag for controlling the initialization of CABAC in slice units. tree_cabac_init_flag is an initialization flag for controlling the initialization of CABAC in tree structure units.

[0690] The three-dimensional data decoding device analyzes the metadata and initializes CABAC when slice_cabac_init_flag or tree_cabac_init_flag is 1. In Fig. 98, when tree_cabac_init_flag indicates 1 or when slice_cabac_init_flag indicates 1, it indicates that CABAC has been initialized, and when tree_cabac_init_flag indicates 0 or when slice_cabac_init_flag indicates 0, it indicates that CABAC is not initialized and the context is continuous (that is, the context is continuously used).

[0691] Setting tree_cabac_init_flag enables initialization of CABAC in units of prediction trees. Setting tree_cabac_init_flag enables resetting at the beginning of any prediction tree. For example, tree_cabac_init_flag may be set to initialize CABAC at the beginning of each group. Alternatively, tree_cabac_init_flag may be set to initialize CABAC at a boundary where encoding parameters of a prediction tree change. Note that when an initialization flag is indicated for each slice, an initialization flag for the tree structure unit at the beginning of the slice does not need to be indicated.

[0692] Since prediction tables in a particular group have similar characteristics, continuing CABAC is likely to improve coding efficiency. Therefore, in bitstream 3, an initialization flag may be set so that CABAC is initialized at the beginning of each group so that CABAC is continued within the same group.

[0693] FIG. 99 is a diagram for explaining a method for decoding a plurality of prediction trees by parallel processing.

[0694] FIG. 99 shows a bitstream in which CABAC is initialized at the beginning of prediction trees 1, 2, 5, and 7 within one slice. With such a bitstream, the three-dimensional data decoding device can independently process the decoding process for prediction tree 1, the decoding process for prediction trees 2 to 4, the decoding process for prediction trees 5 and 6, and the decoding process for prediction trees 7 and 8. In order for the three-dimensional data decoding device to perform parallel processing, the three-dimensional data decoding device must directly access the memory storage locations of the independently decodable data units. Therefore, the three-dimensional data encoding device includes offset information (information indicating the storage location) of the beginning of the encoded data in the encoded data. The offset information may be, for example, byte information from the beginning of the slice. In FIG. 99, the offset information indicated by offset 2 is the number of bytes from the beginning of the slice to the encoded data of prediction tree 2. The offset information may be indicated for each prediction tree, or for each unit of one or more prediction trees that can be processed independently. Furthermore, as in offset D_56, the offset information may be indicated by the number of bytes difference from the beginning of prediction tree 5, which immediately precedes prediction tree 6.

[0695] FIG. 100 is a diagram showing an example of a three-dimensional data encoding method.

[0696] The three-dimensional data encoding device performs predictive tree encoding for each slice (S11441).

[0697] Next, the three-dimensional data encoding device generates a prediction tree and performs entropy encoding for each prediction tree (S11442).

[0698] Next, the three-dimensional data encoding device determines whether or not to continue the context at the top of the tree structure (prediction tree) (S11443).

[0699] When the three-dimensional data encoding device determines not to continue the context at the top of the tree structure (prediction tree) (No in S11443), it initializes the context and sets tree_cabac_init_flag to 1 (S11444).

[0700] Next, the three-dimensional data encoding device stores the offset information (information indicating the storage location) at the beginning of the tree structure (S11445).

[0701] On the other hand, if the three-dimensional data encoding device determines to continue the context at the top of the tree structure (prediction tree) (Yes in S11443), it continues the context and sets tree_cabac_init_flag to 0 (S11446).

[0702] Next, the three-dimensional data encoding device signals at least tree_cabac_init_flag out of tree_cabac_init_flag and offset information using a predetermined method (S11447).

[0703] FIG. 101 is a diagram showing an example of a three-dimensional data decoding method.

[0704] The three-dimensional data decoding device analyzes tree_cabac_init_flag (S11451).

[0705] Next, the three-dimensional data decoding device determines whether or not tree_cabac_init_flag indicates that the context is to be continued at the top of the tree structure (prediction tree) (S11452).

[0706] If tree_cabac_init_flag indicates that the context will not be continued at the top of the tree structure (prediction tree) (No in S11452), the three-dimensional data decoding device initializes the context and performs entropy decoding (S11453).

[0707] If tree_cabac_init_flag indicates that the context is to be continued at the top of the tree structure (prediction tree) (Yes in S11452), the three-dimensional data decoding device continues to use the context and performs entropy decoding (S11454).

[0708] FIG. 102 is a diagram showing an example of parallel decoding in the three-dimensional data decoding method.

[0709] The three-dimensional data decoding device determines whether or not to perform parallel decoding (S11461).

[0710] When the three-dimensional data decoding device determines to perform parallel decoding (Yes in S11461), it accesses parallel coding units based on the offset information and decodes multiple coding units in parallel (S11462).

[0711] In this way, the three-dimensional data encoding device can perform independent processing by initializing CABAC and eliminating dependency in the tree structure. Furthermore, since offset information at the beginning of the tree structure is indicated, the three-dimensional data decoding device can randomly access multiple pieces of encoded data encoded using multiple prediction trees, allowing for independent decoding processing and parallel decoding processing. Furthermore, since the three-dimensional data encoding device and the three-dimensional data decoding device initialize CABAC based on tree_cabac_init_flag, the initialization timing during encoding and decoding can be made the same.

[0712] FIG. 103 is a diagram showing an example of the syntax of a data unit of location information when an initialization flag is stored in the data of the location information.

[0713] In the data unit of position information, the coded data of the predictive tree coding may indicate node information, such as a prediction mode (pred_mode), within a loop for each 3D point. When pred_mode = 0 (direct mode), this indicates that the node is a root node. The root node is the top node (3D point) of the prediction tree, and when the processing target is the root node, an initialization flag (tree_cabac_init_flag) indicating whether or not CABAC has been initialized at the root node may be indicated. Note that, instead of the initialization flag, a random access flag may indicate whether or not CABAC has been initialized. For example, when the random access flag is ON, CABAC may always be considered to be initialized.

[0714] FIG. 104 is a diagram showing an example of header syntax when an initialization flag and offset information are stored in the header of position information.

[0715] The initialization flag and offset information may be collectively indicated in the data unit header of the position information. The data unit header of the position information may indicate the number of prediction trees (num_predtree_minus2) included in the data unit of position information, or may indicate tree_cabac_init_flag for each prediction tree. Furthermore, when tree_cabac_init_flag is set to 1, the data unit header indicates offset information. The offset information may be an offset (difference information) from the beginning of the data unit, or may be an offset (difference information) from the beginning of the previous prediction tree. Note that num_predtree_minus2 may be set so that information about the first prediction tree is not included in the header, and num_predtree_minus1 may be set so that information about the first prediction tree is included in the header.

[0716] FIG. 105 is a diagram showing an example of header syntax when an initialization flag and offset information are stored in the header of position information in units of random access.

[0717] num_rap indicates the number of units that can be decoded in parallel (random access). Offset information may be indicated for each unit that can be decoded in parallel. Note that tree_cabac_init_flag may not be indicated, and CABAC may be initialized at the beginning of the prediction tree indicated by the offset information.

[0718] Alternatively, an identifier (tree_id) for each prediction tree may be indicated in the position information data, and the identifier (tree_id) of the randomly accessible prediction tree may be indicated in the header. In this way, the order of the prediction trees can be determined (identified) by clearly indicating the numbers of the prediction trees.

[0719] The offset information must be indicated in the header, and the tree_cabac_init_flag may be indicated in either the data or the header. The offset information may be indicated in the header, and the tree_cabac_init_flag may be indicated in the data.

[0720] When initializing CABAC, the initial value of CABAC may be set to a predetermined value or may be signaled in the same manner as cabac_init_flag or offset.

[0721] Parallel processing of attribute information can also be achieved by using a method similar to that for position information. An initialization flag or offset information may be indicated using a similar signaling method. The initialization flag may be included in the header or data of the attribute information.

[0722] The parallel-decodable unit may be common to the position information and the attribute information. In this case, the information on the parallel-decodable unit of the attribute information and the initialization flag are common to the position information and may be indicated in the header of the position information, and the offset information in the attribute information may be indicated in the header of the attribute information.

[0723] Although the above example shows the application of the initialization flag (entropy_continue_flag) to slices and tiles of 3D points, the application of the initialization flag is not necessarily limited to the above. For example, the initialization flag may be applied to 3D meshes and sub-meshes.

[0724] Fig. 106 is a diagram for explaining the relationship between an original mesh and sub-meshes according to an embodiment. Specifically, Fig. 106 shows an original mesh and two sub-meshes (a first sub-mesh and a second sub-mesh) generated by dividing the original mesh into sub-meshes.

[0725] For example, when the encoding device 100 or the decoding device 200 sequentially encodes or decodes pre-divided submeshes A and B as shown in (b) of Figure 106, if entropy_continue_flag = 1, the encoding device 100 or the decoding device 200 may set the context information (also simply referred to as entropy coding context information) used for entropy encoding or entropy decoding after encoding or decoding of submesh A to the initial value of the context information at the start of encoding or decoding of the next submesh B, and then encode or decode submesh B.

[0726] As a result, for example, when the appearance trends of the encoding target data of submeshes A and B are similar, the encoding device 100 can improve the encoding efficiency when entropy encoding the encoding target data of submesh B by using the context information after entropy encoding of submesh A as the initial value of the context information (also referred to as the context initial value) used when encoding submesh B. Furthermore, for example, when decoding (e.g., entropy decoding) submesh B, the decoding device 200 can appropriately decode the decoding target data of submesh B by using the context information after entropy decoding of submesh A as the initial value of the context information used for decoding submesh B.

[0727] For example, encoding a submesh means encoding information about the submesh, such as the position information of the vertices of the submesh and attribute information. Also, decoding a submesh means decoding encoded information about the submesh, such as the encoded position information of the vertices of the submesh and encoded attribute information. The same applies to encoding and decoding of items other than submeshes.

[0728] Furthermore, if entropy_continue_flag=0, the encoding device 100 or the decoding device 200 may initialize the context information at the start of encoding or decoding of submesh B.

[0729] This allows the encoding device 100 or the decoding device 200 to independently encode or decode the submesh A and the submesh B. Therefore, for example, parallel processing can be applied, thereby improving the processing speed.

[0730] Here, the initialization flag (entropy_continue_flag) is a flag indicating whether to continue entropy encoding or entropy decoding without initializing context information, i.e., whether to save the context of the previous data unit and apply it to the next data unit.

[0731] For example, entropy_continue_flag=1 indicates that entropy encoding or entropy decoding is continued without initializing context information between sub-meshes, and entropy_continue_flag=0 indicates that entropy encoding or entropy decoding is performed with context information initialized between sub-meshes.

[0732] Note that the expression cabac_init_flag used above may be used instead of entropy_continue_flag. In this case, the definition of the flag is opposite to that of entropy_continue_flag. That is, for example, cabac_init_flag=0 indicates that entropy encoding or entropy decoding is continued without initializing context information between sub-meshes. Also, for example, cabac_init_flag=1 indicates that context information is initialized between sub-meshes and that entropy encoding or entropy decoding is performed.

[0733] If entropy_continue_flag=1, for example, the encoding device 100 may use the context information at the end of encoding submesh A as the initial value of the context when encoding submesh B.

[0734] This can improve the coding efficiency of the entropy coding of sub-mesh B.

[0735] Alternatively, the context information may be exchanged via a memory. For example, the encoding device 100 may store the context information after encoding of the submesh A in a memory, and when encoding of the submesh B starts, read the context information from the memory and set it as the initial value of the context information.

[0736] This allows the context information to be properly transferred between submeshes. That is, the encoding device 100 can continue to use the context information properly. Similarly to the encoding device 100, the decoding device 200 may store the context information after decoding of the submesh A in a memory, and read the context information from the memory when starting to decode the submesh B, and use the read context information as the initial value of the context information.

[0737] Furthermore, for example, if entropy_continue_flag = 0, the encoding device 100 may initialize the context information after encoding submesh A and at the start of encoding submesh B. Similarly, for example, if entropy_continue_flag = 0, the decoding device 200 may initialize the context information after decoding submesh A and at the start of decoding submesh B.

[0738] This allows the encoding device 100 or the decoding device 200 to independently encode or decode the submesh A and the submesh B. Therefore, the processing speed can be improved by applying parallel processing, for example.

[0739] The sub-mesh (sub-mesh information) includes multiple components such as a base mesh (base mesh information), displacement vector information, and attribute information including texture information of the 3D mesh. The base mesh includes information such as geometry coordinate information, texture coordinate information, and / or connectivity data. In this case, the entropy_continue_flag may be set for each component.

[0740] Also, for example, an entropy_continue_flag (for the base mesh) may be prepared. For example, if entropy_continue_flag = 1, the encoding device 100 or the decoding device 200 may set the entropy coding context information after encoding the base mesh of submesh A (after base mesh coding) or after decoding the encoded base mesh (after base mesh decoding) to the initial value of the entropy coding context information at the start of base mesh coding or base mesh decoding of the next submesh B of submesh A, and may encode or decode the base mesh of submesh B.

[0741] As a result, for example, when the appearance trends of the encoding target data of the base mesh of submesh A and the base mesh of submesh B are similar, the encoding device 100 can improve the encoding efficiency when entropy encoding the encoding target data of the base mesh of submesh B by using the context information after the base mesh encoding of submesh A as the initial value for the base mesh encoding of submesh B. Furthermore, during decoding, for example, the decoding device 200 can appropriately decode the decoding target data of the base mesh of submesh B by using the context information after the base mesh decoding of submesh A as the initial value for the base mesh decoding of submesh B.

[0742] Note that a mechanism similar to the above-described base mesh encoding and base mesh decoding may be applied to encoding or decoding of displacement vectors or attribute information.

[0743] By providing an entropy_continue_flag for each component constituting a submesh, it is possible to select whether to improve coding efficiency by passing context information between submeshes for each component, or to initialize context information between submeshes to enable parallel processing, etc. This makes it possible to balance coding efficiency and processing volume.

[0744] For example, the encoding device 100 may define an sps_entropy_continue_flag that is valid for all components that make up a submesh in a higher-level syntax such as SPS. Furthermore, for example, when sps_entropy_continue_flag = 1, the encoding device 100 may estimate the value of entropy_continue_flag for each component as 1 without adding it to the header of each component. Furthermore, for example, when sps_entropy_continue_flag = 0, the encoding device 100 may add the entropy_continue_flag of each component to the header of each component to control the passing of context information between submeshes.

[0745] This makes it possible to reduce the amount of information contained in the header (header information).

[0746] FIG. 107 is a block diagram showing another example configuration of the encoding device 100 according to the embodiment.

[0747] In this example, the encoding device 100 comprises a sub-mesh divider 517, a projector 512, a base mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and optionally one or more other type encoders 516.

[0748] The submesh divider 517 acquires a three-dimensional mesh, divides the acquired three-dimensional mesh into submeshes, and outputs the multiple submeshes generated by the submesh division to the content projector 512 .

[0749] The projector 512 projects the content onto an input mesh (a 3D mesh frame) that includes geometry coordinates (vertex coordinates indicating the positions of the vertices), texture coordinates, and connectivity (connectivity information). The data is output to a base mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and optionally one or more other type encoders 516. Each encoder compresses the data into a bitstream.

[0750] For example, the submesh divider 517 determines whether or not the entropy_continue_flag is set (is it zero) in the input three-dimensional mesh.

[0751] Furthermore, for example, each encoder determines whether to initialize and use a context (for entropy coding) for entropy coding or to continue using it without initializing it, according to entropy_continue_flag.

[0752] In addition, when a video codec is used to encode or decode attribute information such as displacement vectors or texture information of three-dimensional meshes, control such as entropy_continue_flag may be realized by a function possessed by the video codec.

[0753] For example, when encoding a displacement vector of a submesh or texture information of a submesh using a video codec, the encoding device 100 may map each component of the submesh to a slice of an image. For example, when entropy_continue_flag=1, the encoding device 100 may continue to use context information between slices by using a dependent slice mechanism, which is a function of the video codec.

[0754] This improves the coding efficiency of an image onto which each component of a submesh is mapped. More specifically, for example, when information about submesh A is mapped onto an image as slice A and information about submesh B is mapped onto the image as slice B, the coding device 100 can code slice A and slice B using the dependent slice mechanism, thereby improving the coding efficiency.

[0755] In this way, the method of mapping texture information and / or displacement vector information input to the video codec onto an image may be switched, or the video codec settings may be switched, depending on the value of entropy_continue_flag.

[0756] This makes it possible to improve the coding efficiency of the video codec according to the value of entropy_continue_flag, thereby improving the coding efficiency of the entire coding process.

[0757] In addition, when a low-delay mode (low-delay transmission mode) profile is specified in an international standard such as MPEG, and when a low-delay mode profile is used, it is acceptable to limit entropy_continue_flag to 1 and to limit the dependent slice mechanism to on in the video codec.

[0758] As a result, when the standard is set to low latency mode, context information is continued between sub-meshes when encoding each component, thereby improving encoding efficiency.

[0759] In addition, when a low latency mode profile is used and the video codec's dependent slice mechanism is turned off when entropy_continue_flag = 1, the decoding device 200 may output information indicating a standard conformance violation, etc.

[0760] This allows the user to determine whether the bitstream conforms to the standard.

[0761] FIG. 108 is a block diagram showing another example of the configuration of the decoding device 200 according to this embodiment.

[0762] In this example, the decoding device 200 comprises a base mesh decoder 613 , a displacement decoder 614 , an attribute decoder 615 , one or more other type decoders 616 , and a 3D reconstructor 617 .

[0763] The bitstream is sent to a base mesh decoder 613, a displacement decoder 614, an attribute decoder 615, and optionally one or more other type decoders 616. These decoders decode the bitstream to generate data (decoded data) including geometry coordinates, texture coordinates, connectivity, etc. The decoded data is then sent to a 3D reconstructor 617, which reconstructs an output mesh (a 3D mesh frame). For example, the 3D reconstructor 617 merges multiple sub-meshes to reconstruct a 3D mesh.

[0764] Each decoder initializes the context information for entropy coding according to, for example, entropy_continue_flag.

[0765] Next, a method for initializing context information for entropy coding and determining an initial context value will be described.

[0766] [Method of determining entropy_continue_flag] Fig. 109 is a flow diagram showing initialization processing of context information according to an embodiment. Specifically, Fig. 109 is a flow diagram showing processing related to initialization of context information in base mesh coding or displacement vector coding (coding of displacement vectors used in base meshes).

[0767] First, the encoding device 100 determines, for each submesh, based on a predetermined condition, whether or not to initialize context information (context information of the base mesh) when encoding the base mesh of that submesh (S401).

[0768] Next, when the encoding device 100 determines that the context information is to be initialized when performing base mesh encoding (Yes in S402), it determines the initial context value to be used for base mesh encoding (S403). The initial context value is set to an initial value that takes into account the encoding characteristics, for example. The initial context value may be a predetermined value, or may be adaptively determined according to the characteristics of the data included in the sub-mesh.

[0769] Next, the encoding device 100 sets the entropy_continue_flag of the base mesh to 0, and sets the context initial value (S404). For example, the encoding device 100 signals the information indicating the entropy_continue_flag=0 of the base mesh and the information indicating the context initial value to the bitstream.

[0770] On the other hand, for example, when the encoding device 100 determines that the context information is not initialized when encoding the base mesh (No in S402), it sets the entropy_continue_flag of the base mesh to 1 (S405). For example, the encoding device 100 signals information indicating that the entropy_continue_flag of the base mesh is 1 to the bitstream.

[0771] After step S404 or step S405, the encoding device 100 determines, for each submesh, based on predetermined conditions, whether or not to initialize context information (context information of the displacement vector) when encoding the displacement vector of that submesh (S406).

[0772] When the encoding device 100 determines that the context information is to be initialized when encoding the displacement vector (Yes in S407), it determines an initial context value to be used for encoding the displacement vector (S408). The initial context value is set to, for example, an initial value that takes into account encoding characteristics. The initial context value may be a predetermined value, or may be adaptively determined according to the characteristics of the data included in the submesh.

[0773] Next, the encoding device 100 sets the entropy_continue_flag (for the displacement vector) of the displacement vector to 0, and sets the context initial value (S409). For example, the encoding device 100 signals information indicating entropy_continue_flag=0 of the displacement vector and information indicating the context initial value to the bitstream.

[0774] On the other hand, for example, when the encoding device 100 determines not to initialize the context information when encoding the disparity vector (No in S407), the encoding device 100 sets the entropy_continue_flag of the disparity vector to 1 (S410). For example, the encoding device 100 signals information indicating that the entropy_continue_flag of the disparity vector is 1 to the bitstream.

[0775] For example, when initializing context information, the encoding device 100 performs initialization processing using a context initial value in base mesh encoding in step S404. Also, when initializing context information, the encoding device 100 performs initialization processing using a context initial value in displacement vector encoding in step S409.

[0776] The processing relating to the base mesh and the processing relating to the displacement vector may be performed in the reverse order to that described above, or may be performed in parallel.

[0777] Furthermore, although the above description has been given taking the example of processing in units of sub-meshes, the same processing is also applicable to processing in units of tiles, slices, or other data units.

[0778] Furthermore, the predetermined conditions may be the same for the processing related to the base mesh and the processing related to the displacement vector, or may be different.

[0779] In addition, in the above example, the initialization process of the context information is performed using the base mesh and the displacement vector that constitute the sub-mesh. However, the initialization process is not limited to this combination and may be applied to other combinations of components. For example, the above process may be applied to the base mesh and the attribute information. Alternatively, the above process may be applied to the base mesh, the displacement vector, and the attribute information.

[0780] This allows switching whether or not to initialize the context information for each component, thereby ensuring a balance between coding efficiency and...

Claims

1. An encoding method comprising: generating a plurality of encoded data by encoding a plurality of submeshes; generating a parameter set indicating the configuration of the plurality of submeshes; and generating a bit stream including the plurality of encoded data and the parameter set, wherein the parameter set includes first identification information indicating whether at least one of the plurality of submeshes depends on other submeshes among the plurality of submeshes when a predetermined condition is satisfied.

2. The encoding method according to claim 1, wherein the parameter set further includes information indicating the number of the plurality of sub-meshes.

3. The encoding method of claim 1, wherein the parameter set further includes second identification information indicating whether the plurality of submeshes includes a dependent submesh that depends on another submesh among the plurality of submeshes, and the parameter set includes the first identification information if the plurality of submeshes includes the dependent submesh.

4. The encoding method of claim 1, wherein the parameter set further includes third identification information indicating whether multiple submesh identification numbers that uniquely identify each of the multiple submeshes are consecutive numbers, and when the multiple submesh identification numbers are not consecutive numbers, the parameter set further includes information indicating the multiple submesh identification numbers.

5. An encoding method according to any one of claims 1 to 4, wherein the parameter sets include a first parameter set relating to a base mesh corresponding to the plurality of sub-meshes and a second parameter set relating to displacement vectors for displacing vertices included in the plurality of sub-meshes, and the first identification information is included in at least one of the first parameter set and the second parameter set.

6. A decoding method comprising: obtaining, from a bitstream, a plurality of coded data in which each of a plurality of submeshes is coded, and a parameter set indicating the configuration of the plurality of submeshes; decoding the plurality of coded data based on the parameter set; and, when a predetermined condition is satisfied, the parameter set including first identification information indicating whether or not each of at least one of the plurality of submeshes depends on another submesh among the plurality of submeshes.

7. The decoding method according to claim 6, wherein the parameter set further includes information indicating the number of the plurality of sub-meshes.

8. The decoding method of claim 6, wherein the parameter set further includes second identification information indicating whether the plurality of submeshes includes a dependent submesh that depends on another submesh among the plurality of submeshes, and the parameter set includes the first identification information if the plurality of submeshes includes the dependent submesh.

9. The decoding method of claim 6, wherein the parameter set further includes third identification information indicating whether multiple submesh identification numbers that uniquely identify each of the multiple submeshes are consecutive numbers, and when the multiple submesh identification numbers are not consecutive numbers, the parameter set further includes information indicating the multiple submesh identification numbers.

10. A decoding method according to any one of claims 6 to 9, wherein the parameter sets include a first parameter set relating to a base mesh corresponding to the plurality of submeshes and a second parameter set relating to displacement vectors for displacing vertices included in the plurality of submeshes, and the first identification information is included in at least one of the first parameter set and the second parameter set.

11. An encoding device comprising: a processor; and a memory, wherein the processor uses the memory to generate a plurality of encoded data by encoding a plurality of submeshes; generate a parameter set indicating the configuration of the plurality of submeshes; and generate a bit stream including the plurality of encoded data and the parameter set, wherein the parameter set includes, when a predetermined condition is satisfied, first identification information indicating whether at least one of the plurality of submeshes depends on another submesh among the plurality of submeshes.

12. A decoding device comprising: a processor; and a memory, wherein the processor uses the memory to obtain, from a bitstream, a plurality of coded data in which each of a plurality of submeshes is coded, and a parameter set indicating the configuration of the plurality of submeshes; and decodes the plurality of coded data based on the parameter set, wherein the parameter set includes, when a predetermined condition is satisfied, first identification information indicating whether or not each of at least one of the plurality of submeshes depends on another submesh among the plurality of submeshes.

Citation Information

Patent Citations

  • A Method for Motion Estimation Using Deformable Meshes

    JP2008514073A

  • Methods for instance-based mesh coding

    WO2024025638A1