Decoding method and decoding device

The decoding method ensures accurate three-dimensional mesh reconstruction by processing submesh data with consistent attribute information, addressing inefficiencies and errors in existing encoding and decoding processes.

WO2026018820A1PCT designated stage Publication Date: 2026-01-22PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/025208
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-17
Filing Date
2025-07-14
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing encoding and decoding processes for three-dimensional data are inefficient and prone to errors in reconstructing accurate three-dimensional meshes due to inconsistencies in attribute information configurations.

Method used

A decoding method that processes multiple encoded data from submeshes to generate submesh data with consistent attribute information configurations, ensuring correct reconstruction of three-dimensional meshes by verifying and outputting data with matching attribute information.

Benefits of technology

Prevents incorrect combinations of attribute information during mesh reconstruction, enabling accurate and reliable decoding of three-dimensional data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025025208_22012026_PF_FP_ABST
    Figure JP2025025208_22012026_PF_FP_ABST
Patent Text Reader

Abstract

In a decoding method according to one aspect of the present invention, a plurality of items of encoded data concerning a plurality of sub-meshes are acquired (S421); the plurality of items of encoded data are decoded to generate a plurality of items of sub-mesh data concerning the plurality of sub-meshes (S422); and the plurality of items of sub-mesh data that have been generated are output (S423). The plurality of items of sub-mesh data that have been output have the same composition of attribute information.
Need to check novelty before this filing date? Find Prior Art

Description

Decoding method and decoding device

[0001] The present disclosure relates to a decoding method and the like.

[0002] In US Pat. No. 6,299,549 a method and apparatus for encoding and decoding three-dimensional mesh data is proposed.

[0003] Japanese Patent Application Laid-Open No. 2006-187015

[0004] Further improvements are desired in the encoding or decoding process for three-dimensional data. The present disclosure aims to improve the encoding or decoding process for three-dimensional data.

[0005] A decoding method according to one aspect of the present disclosure obtains multiple encoded data relating to multiple submeshes, decodes the multiple encoded data to generate multiple submesh data relating to the multiple submeshes, and outputs the generated multiple submesh data, wherein the output multiple submesh data have the same attribute information configuration.

[0006] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0007] The present disclosure may contribute to improvements in decoding processes and the like related to three-dimensional data.

[0008] 1 is a conceptual diagram showing a three-dimensional mesh according to an embodiment. FIG. 2 is a conceptual diagram showing basic elements of a three-dimensional mesh according to an embodiment. FIG. 3 is a conceptual diagram showing mapping according to an embodiment. FIG. 4 is a block diagram showing a configuration example of an encoding / decoding system according to an embodiment. FIG. 5 is a block diagram showing a configuration example of an encoding device according to an embodiment. FIG. 6 is a block diagram showing another configuration example of an encoding device according to an embodiment. FIG. 7 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 8 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 9 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 10 is a conceptual diagram showing another configuration example of a bit stream according to an embodiment. FIG. 11 is a conceptual diagram showing yet another configuration example of a bit stream according to an embodiment. FIG. 12 is a block diagram showing a specific example of an encoding / decoding system according to an embodiment. FIG. 13 is a conceptual diagram showing an example configuration of point cloud data according to an embodiment. FIG. 14 is a conceptual diagram showing an example data file of point cloud data according to an embodiment. FIG. 15 is a conceptual diagram showing an example configuration of mesh data according to an embodiment. FIG. 16 is a conceptual diagram showing an example data file of mesh data according to an embodiment. FIG. 17 is a conceptual diagram showing types of three-dimensional data according to an embodiment. FIG. 18 is a block diagram showing an example configuration of a three-dimensional data encoder according to an embodiment. FIG. 19 is a block diagram showing an example configuration of a three-dimensional data decoder according to an embodiment. FIG. 19 is a block diagram showing another configuration example of a three-dimensional data encoder according to an embodiment. FIG. 19 is a block diagram showing another configuration example of a three-dimensional data decoder according to an embodiment. FIG. 1 is a conceptual diagram showing a specific example of encoding processing according to an embodiment. FIG. 2 is a conceptual diagram showing a specific example of decoding processing according to an embodiment. FIG. 3 is a block diagram showing an implementation example of an encoding device according to an embodiment. FIG. 4 is a block diagram showing an implementation example of a decoding device according to an embodiment. FIG. 5 is a block diagram showing another configuration example of an encoding / decoding system according to an embodiment. FIG. 6 is a block diagram showing another configuration example of an encoding device according to an embodiment. FIG. 7 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 8 is a block diagram showing yet another configuration example of an encoding device according to an embodiment. FIG. 9 is a block diagram showing yet another configuration example of a decoding device according to an embodiment. FIG. 10 is a flow diagram showing processing of an encoding device according to an embodiment. FIG. 11 is an explanatory diagram conceptually showing encoding of a mesh frame according to an embodiment.1 is a flow diagram showing processing of a decoding device according to an embodiment. FIG. 2 is an explanatory diagram conceptually showing decoding of a mesh frame according to an embodiment. FIG. 3 is a block diagram showing an example configuration of a decoding device according to an embodiment. FIG. 4 is a block diagram showing an example configuration of an encoding device according to an embodiment. FIG. 5 is a block diagram showing an example configuration of a decoding device according to an embodiment. FIG. 6 is an explanatory diagram showing an example of subdivision according to an embodiment. FIG. 7 is an explanatory diagram showing an example of displacement of vertices after displacement after subdivision according to an embodiment. FIG. 8 is an explanatory diagram showing example vertices of an original mesh according to an embodiment. FIG. 9 is an explanatory diagram showing an example mesh according to an embodiment. FIG. 10 is an explanatory diagram showing an example of division of a mesh into sub-meshes according to an embodiment. FIG. 11 is a first explanatory diagram showing an example of packing of displacement information into an image frame according to an embodiment. FIG. 12 is a second explanatory diagram showing an example of packing of displacement information into an image frame according to an embodiment. FIG. 13 is a third explanatory diagram showing an example of packing of displacement information into an image frame according to an embodiment. FIG. 14 is a block diagram showing another example configuration of an encoding device according to an embodiment. FIG. 15 is a block diagram showing an example configuration of a preprocessor according to an embodiment. FIG. 16 is a block diagram showing another example configuration of a decoding device according to an embodiment. FIG. 17 is a block diagram showing a specific example configuration of a decoding device according to an embodiment. FIG. 18 is a diagram showing the data structure of a bitstream according to an embodiment. FIG. 19 is a diagram showing attribute information of a sub-mesh according to an embodiment. FIG. 19 is a diagram for explaining reconstruction processing of attribute information according to an embodiment. FIG. 1 is a flow diagram showing a submesh encoding method according to an embodiment. FIG. 2 is a diagram showing a syntax for signaling metadata according to an embodiment. FIG. 3 is a diagram showing a syntax for signaling metadata according to an embodiment. FIG. 4 is a diagram showing attribute information of a submesh according to an embodiment. FIG. 5 is a diagram showing attribute information of a submesh according to an embodiment. FIG. 6 is a diagram for explaining the processing of a reconstructor according to an embodiment. FIG. 7 is a flow diagram showing a submesh decoding method according to an embodiment. FIG. 8 is a flow diagram showing a reconstruction process according to an embodiment. FIG. 9 is a diagram showing a comparison process between submesh level attribute information and frame level attribute information according to an embodiment. FIG. 10 is a diagram showing an example of the reconstruction process according to an embodiment.10 is a diagram showing another example of a reconstruction process according to an embodiment; FIG. 11 is a diagram showing another example of a reconstruction process according to an embodiment; FIG. 12 is a diagram showing details of the reconstruction process according to an embodiment; FIG. 13 is a diagram showing syntax for setting default values ​​according to an embodiment; FIG. 14 is a flow diagram showing a process for setting default values ​​according to an embodiment; FIG. 15 is a diagram for explaining parameters for each sub-mesh according to an embodiment; FIG. 16 is a diagram showing syntax for signaling a flag indicating whether a parameter is common to each sub-mesh according to an embodiment; FIG. 17 is a block diagram showing an example configuration of a base mesh frame decoder according to an embodiment; FIG. 18 is a diagram showing syntax for signaling a base mesh parameter set according to an embodiment; FIG. 19 is a diagram showing syntax for signaling a sub-mesh structure set according to an embodiment; FIG. 19 is a flow diagram showing an example process of a verifier according to an embodiment; FIG. 19 is a flow diagram showing another example process of a verifier according to an embodiment; FIG. 19 is a diagram showing attribute information of a sub-mesh according to an embodiment; FIG. 19 is a flow diagram showing an example of a basic decoding process according to an embodiment;

[0009] Introduction Three-dimensional (3D) meshes are used in computer graphics images, which may be composed of multiple temporally distinct frames, each of which may be represented by a 3D mesh.

[0010] A 3D mesh is composed of vertex information indicating the positions of each of the vertices in 3D space, connectivity information indicating the connections between the vertices, and attribute information indicating the attributes of each vertex or face. Each face is constructed according to the connectivity between the vertices. Various computer graphics images can be expressed using such 3D meshes.

[0011] Furthermore, for transmission and storage of the 3D mesh, efficient encoding and decoding of the 3D mesh is expected. For efficient encoding and decoding of the 3D mesh, arithmetic coding and decoding may be used.

[0012] Further improvements are desired in the encoding or decoding process for three-dimensional data. The present disclosure aims to improve the encoding or decoding process for three-dimensional data.

[0013] Below, examples of inventions that can be obtained from the disclosure of this specification will be given, and the effects and the like that can be obtained from these inventions will be explained.

[0014] The decoding method of Example 1 obtains multiple coded data relating to multiple submeshes, decodes the multiple coded data to generate multiple submesh data relating to the multiple submeshes, and outputs the generated multiple submesh data, and the output multiple submesh data have the same attribute information configuration.

[0015] According to this, since the plurality of sub-mesh data have the same attribute information configuration, when the plurality of sub-mesh data are used to combine the plurality of sub-meshes to reconstruct a three-dimensional mesh, it is possible to prevent the combination of attribute information from being incorrectly combined, thereby preventing the reconstruction of an incorrect three-dimensional mesh.

[0016] The decoding method of Example 2 may be the decoding method of Example 1, further comprising determining a decoding method for each of the plurality of encoded data, and generating the plurality of sub-mesh data by decoding each of the plurality of encoded data using the determined decoding method.

[0017] This allows each of the plurality of coded data to be decoded by, for example, a decoding method corresponding to the coding method of the coded data.

[0018] The decoding method of Example 3 may be the decoding method of Example 1 or Example 2, further determining whether the generated plurality of sub-mesh data have the same attribute information configuration.

[0019] This makes it possible to determine whether a three-dimensional mesh can be correctly reconstructed based on the result of determining whether a plurality of sub-mesh data have the same attribute information configuration.

[0020] The decoding method of Example 4 is the decoding method of Example 3, and if it is determined that the generated plurality of sub-mesh data have the same attribute information configuration, the generated plurality of sub-mesh data may be output to a reconstructor that reconstructs a three-dimensional mesh using the plurality of sub-mesh data.

[0021] This allows the reconstructor to correctly reconstruct a three-dimensional mesh using multiple sub-mesh data.

[0022] The decoding method of Example 5 is a decoding method of Examples 1 to 4, and the configuration of the attribute information of the output multiple sub-mesh data may be the same as the configuration of the attribute information indicated in the metadata of the base mesh related to the multiple sub-meshes.

[0023] This makes it possible to output a plurality of sub-mesh data having the same structure as the structure of the attribute information indicated in the metadata of the base mesh.

[0024] The decoding device of Example 6 comprises a circuit and a memory connected to the circuit, and in operation, the circuit acquires multiple encoded data relating to multiple submeshes, decodes the multiple encoded data to generate multiple submesh data relating to the multiple submeshes, and outputs the generated multiple submesh data, and the output multiple submesh data have the same attribute information configuration.

[0025] This provides the same effect as the decoding method of Example 1.

[0026] Furthermore, these comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0027] <Expressions and Terms> The following expressions and terms are used herein.

[0028] (1) Three-dimensional Mesh A three-dimensional mesh is a collection of multiple faces, and represents, for example, a three-dimensional object. A three-dimensional mesh is mainly composed of vertex information, connectivity information, and attribute information. A three-dimensional mesh may be expressed as a polygon mesh or a mesh. A three-dimensional mesh may also vary over time. A three-dimensional mesh may include metadata related to the vertex information, connectivity information, and attribute information, and may also include other additional information.

[0029] (2) Vertex Information Vertex information is information indicating a vertex. For example, the vertex information indicates the position of a vertex in a three-dimensional space. Furthermore, a vertex corresponds to a vertex of a face that constitutes a three-dimensional mesh. Vertex information may be expressed as "geometry." Furthermore, vertex information may be expressed as position information.

[0030] (3) Connection Information Connection information is information that indicates connections between vertices. For example, connection information indicates connections for forming faces or edges of a three-dimensional mesh. Connection information may be expressed as "Connectivity." Connection information may also be expressed as face information.

[0031] (4) Attribute Information Attribute information is information that indicates attributes of a vertex or a face. For example, attribute information indicates attributes such as a color, an image, and a normal vector associated with a vertex or a face. Attribute information may be expressed as "texture."

[0032] (5) Faces A face is an element that constitutes a three-dimensional mesh. Specifically, a face is a polygon on a plane in three-dimensional space. For example, a face can be defined as a triangle in three-dimensional space.

[0033] (6) Plane A plane is a two-dimensional plane in a three-dimensional space. For example, a polygon is formed on a plane, and multiple polygons are formed on multiple planes.

[0034] (7) Bitstream: A bitstream corresponds to coded information. A bitstream may also be referred to as a stream, a coded bitstream, a compressed bitstream, or a coded signal.

[0035] (8) Encoding and Decoding The term encoding may be substituted with terms such as storing, including, writing, describing, signaling, sending, notifying, saving, or compressing, and these terms may be interchangeable. For example, encoding information may mean including the information in a bitstream. Also, encoding information into a bitstream may mean encoding the information to generate a bitstream that includes the encoded information.

[0036] Additionally, the term "decode" may be replaced with terms such as "read," "decode," "read," "load," "derive," "obtain," "receive," "extract," "reconstruct," "reconstruct," "decompress," or "decompress," and these terms may be interchangeable. For example, decoding information may mean obtaining information from a bitstream. Decoding information from a bitstream may mean decoding the bitstream to obtain information contained in the bitstream.

[0037] (9) Ordinal Numbers In the description, ordinal numbers such as first and second may be assigned to components, etc. These ordinal numbers may be changed as appropriate. Furthermore, new ordinal numbers may be assigned to components, etc., or removed. Furthermore, these ordinal numbers may be assigned to elements in order to identify them, and may not correspond to a meaningful order.

[0038] <Three-dimensional mesh> Fig. 1 is a conceptual diagram showing a three-dimensional mesh according to this embodiment. A three-dimensional mesh is composed of multiple faces. For example, each face is a triangle. The vertices of these triangles are defined in three-dimensional space. The three-dimensional mesh then represents a three-dimensional object. Each face may have a color or an image.

[0039] FIG. 2 is a conceptual diagram showing the basic elements of a three-dimensional mesh according to this embodiment. A three-dimensional mesh is composed of vertex information, connection information, and attribute information. The vertex information indicates the positions of the vertices of a face in three-dimensional space. The connection information indicates the connections between the vertices. A face can be identified by the vertex information and connection information. In other words, a colorless three-dimensional object is formed in three-dimensional space by the vertex information and connection information.

[0040] The attribute information may be associated with a vertex or a face. The attribute information associated with a vertex may be expressed as "Attribute Per Point." The attribute information associated with a vertex may indicate an attribute of the vertex itself, or may indicate an attribute of a face connected to the vertex.

[0041] For example, a color may be associated with a vertex as attribute information. The color associated with a vertex may be the color of the vertex itself, or the color of a face connected to the vertex. The color of a face may be the average of multiple colors associated with multiple vertices of the face. Furthermore, a normal vector may be associated with a vertex or a face as attribute information. Such a normal vector can represent the front and back of a face.

[0042] A two-dimensional image may be associated with a surface as attribute information. The two-dimensional image associated with a surface is also expressed as a texture image or an "Attribute Map." Information indicating mapping between the surface and the two-dimensional image may be associated with the surface as attribute information. Information indicating such mapping may be expressed as mapping information, vertex information of a texture image, texture coordinates, or "Attribute UV Coordinate."

[0043] Furthermore, information such as color, image, and moving image used as attribute information may be expressed as "parametric space."

[0044] The attribute information allows texture to be reflected on the three-dimensional object. That is, a three-dimensional object having color is formed in three-dimensional space based on the vertex information, connection information, and attribute information.

[0045] In the above, the attribute information is associated with the vertices or faces, but it may also be associated with the edges.

[0046] 3 is a conceptual diagram illustrating mapping according to this embodiment. For example, a region of a two-dimensional image on a two-dimensional plane can be mapped onto a surface of a three-dimensional mesh in three-dimensional space. Specifically, coordinate information of the region in the two-dimensional image is associated with the surface of the three-dimensional mesh. As a result, an image of the mapped region in the two-dimensional image is reflected on the surface of the three-dimensional mesh.

[0047] By using the mapping, the 2D image used as attribute information can be separated from the 3D mesh. For example, in encoding the 3D mesh, the 2D image may be encoded by an image encoding method or a video encoding method.

[0048] <System Configuration> Fig. 4 is a block diagram showing an example of the configuration of a coding / decoding system according to this embodiment. In Fig. 4, the coding / decoding system includes a coding device 100 and a decoding device 200.

[0049] For example, the encoding device 100 obtains a three-dimensional mesh and encodes the three-dimensional mesh into a bitstream. Then, the encoding device 100 outputs the bitstream to the network 300. For example, the bitstream includes the encoded three-dimensional mesh and control information for decoding the encoded three-dimensional mesh. By encoding the three-dimensional mesh, information about the three-dimensional mesh is compressed.

[0050] The network 300 transmits a bitstream from the encoding device 100 to the decoding device 200. The network 300 may be the Internet, a wide area network (WAN), a local area network (LAN), or a combination of these. The network 300 is not necessarily limited to bidirectional communication, and may be a unidirectional communication network for terrestrial digital broadcasting, satellite broadcasting, or the like.

[0051] Furthermore, the network 300 can be replaced by a recording medium such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc (registered trademark)).

[0052] The decoding device 200 obtains a bitstream and decodes a three-dimensional mesh from the bitstream. By decoding the three-dimensional mesh, information about the three-dimensional mesh is expanded. For example, the decoding device 200 decodes the three-dimensional mesh according to a decoding method corresponding to the encoding method used by the encoding device 100 to encode the three-dimensional mesh. That is, the encoding device 100 and the decoding device 200 perform encoding and decoding according to encoding methods and decoding methods that correspond to each other.

[0053] The 3D mesh before encoding may also be referred to as an original 3D mesh, and the 3D mesh after decoding may also be referred to as a reconstructed 3D mesh.

[0054] 5 is a block diagram showing an example of the configuration of a coding device 100 according to this embodiment. For example, the coding device 100 includes a vertex information encoder 101, a connection information encoder 102, and an attribute information encoder 103.

[0055] The vertex information encoder 101 is an electrical circuit that encodes vertex information. For example, the vertex information encoder 101 encodes the vertex information into a bitstream according to a format defined for the vertex information.

[0056] The connection information encoder 102 is an electrical circuit that encodes the connection information, for example, the connection information encoder 102 encodes the connection information into a bitstream according to a format defined for the connection information.

[0057] The attribute information encoder 103 is an electric circuit that encodes the attribute information. For example, the attribute information encoder 103 encodes the attribute information into a bit stream in accordance with a format defined for the attribute information.

[0058] The vertex information, connectivity information, and attribute information may be coded using variable-length coding or fixed-length coding, such as Huffman coding or context-adaptive binary arithmetic coding (CABAC).

[0059] The vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be integrated together, or each of the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be further subdivided into multiple components.

[0060] 6 is a block diagram showing another example of the configuration of the encoding device 100 according to this embodiment. For example, the encoding device 100 includes a pre-processor 104 and a post-processor 105 in addition to the configuration shown in FIG.

[0061] The preprocessor 104 is an electrical circuit that performs processing before encoding the vertex information, connectivity information, and attribute information. For example, the preprocessor 104 may perform a conversion process, a separation process, a multiplexing process, or the like on the 3D mesh before encoding. More specifically, for example, the preprocessor 104 may separate the vertex information, connectivity information, and attribute information from the 3D mesh before encoding.

[0062] The post-processor 105 is an electrical circuit that performs processing after the vertex information, connection information, and attribute information are encoded. For example, the post-processor 105 may perform conversion processing, separation processing, multiplexing processing, or the like on the encoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 105 may multiplex the encoded vertex information, connection information, and attribute information into a bitstream. Furthermore, for example, the post-processor 105 may further perform variable-length coding on the encoded vertex information, connection information, and attribute information.

[0063] 7 is a block diagram showing an example of the configuration of a decoding device 200 according to this embodiment. For example, the decoding device 200 includes a vertex information decoder 201, a connection information decoder 202, and an attribute information decoder 203.

[0064] The vertex information decoder 201 is an electrical circuit that decodes vertex information. For example, the vertex information decoder 201 decodes vertex information from a bitstream according to a format defined for the vertex information.

[0065] The connection information decoder 202 is an electrical circuit that decodes the connection information, for example, the connection information decoder 202 decodes the connection information from the bitstream according to a format defined for the connection information.

[0066] The attribute information decoder 203 is an electric circuit that decodes the attribute information. For example, the attribute information decoder 203 decodes the attribute information from the bitstream in accordance with a format defined for the attribute information.

[0067] The vertex information, connection information, and attribute information may be decoded using variable length decoding or fixed length decoding, which may correspond to Huffman coding, context-adaptive binary arithmetic coding (CABAC), or the like.

[0068] The vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be integrated together, or each of the vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be further subdivided into multiple components.

[0069] 8 is a block diagram showing another example of the configuration of the decoding device 200 according to this embodiment. For example, the decoding device 200 includes a pre-processor 204 and a post-processor 205 in addition to the configuration shown in FIG.

[0070] The preprocessor 204 is an electrical circuit that performs processing before decoding the vertex information, connection information, and attribute information. For example, the preprocessor 204 may perform conversion processing, separation processing, multiplexing processing, or the like on the bitstream before decoding the vertex information, connection information, and attribute information.

[0071] More specifically, for example, the preprocessor 204 may separate a sub-bitstream corresponding to vertex information, a sub-bitstream corresponding to connectivity information, and a sub-bitstream corresponding to attribute information from the bitstream. Also, for example, the preprocessor 204 may perform variable-length decoding on the bitstream in advance before decoding the vertex information, connectivity information, and attribute information.

[0072] The post-processor 205 is an electrical circuit that performs processing after the vertex information, connection information, and attribute information are decoded. For example, the post-processor 205 may perform conversion processing, separation processing, multiplexing processing, or the like on the decoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 205 may multiplex the decoded vertex information, connection information, and attribute information onto a three-dimensional mesh.

[0073] <Bitstream> Vertex information, connection information, and attribute information are coded and stored in a bitstream. The relationship between this information and the bitstream is shown below.

[0074] 9 is a conceptual diagram showing an example of the configuration of a bitstream according to this embodiment. In this example, connection information, vertex information, and attribute information are integrated in the bitstream. For example, the connection information, vertex information, and attribute information may be included in a single file.

[0075] Furthermore, multiple portions of this information may be stored sequentially, such as a first portion of connection information, a first portion of vertex information, a first portion of attribute information, a second portion of connection information, a second portion of vertex information, a second portion of attribute information, etc. These multiple portions may correspond to multiple portions that are different in time, multiple portions that are different in space, or multiple different faces.

[0076] Furthermore, the storage order of the connection information, vertex information, and attribute information is not limited to the above example, and a storage order different from the above example may be used.

[0077] 10 is a conceptual diagram showing another example of the configuration of a bitstream according to this embodiment. In this example, a plurality of files are included in the bitstream, and connection information, vertex information, and attribute information are stored in different files. Here, a file containing connection information, a file containing vertex information, and a file containing attribute information are shown, but the storage format is not limited to this example. For example, two types of information among the connection information, vertex information, and attribute information may be included in one file, and the remaining type of information may be included in another file.

[0078] Alternatively, the information may be split and stored in more files. For example, multiple pieces of connectivity information may be stored in multiple files, multiple pieces of vertex information may be stored in multiple files, or multiple pieces of attribute information may be stored in multiple files. These multiple pieces may correspond to multiple temporally different pieces, multiple spatially different pieces, or multiple different faces.

[0079] Furthermore, the storage order of the connection information, vertex information, and attribute information is not limited to the above example, and a storage order different from the above example may be used.

[0080] 11 is a conceptual diagram showing another example of the configuration of a bitstream according to this embodiment. In this example, the bitstream is composed of multiple separable sub-bitstreams, and connection information, vertex information, and attribute information are stored in different sub-bitstreams.

[0081] Here, a sub-bitstream containing connection information, a sub-bitstream containing vertex information, and a sub-bitstream containing attribute information are shown, but the storage format is not limited to this example.

[0082] For example, two types of information among the connection information, vertex information, and attribute information may be included in one sub-bitstream, and the remaining type of information may be included in another sub-bitstream. Specifically, attribute information of a two-dimensional image or the like may be stored in a sub-bitstream that complies with an image coding method, separate from the sub-bitstreams of the connection information and vertex information.

[0083] Each sub-bitstream may also include multiple files, and multiple pieces of connectivity information may be stored in multiple files, multiple pieces of vertex information may be stored in multiple files, or multiple pieces of attribute information may be stored in multiple files.

[0084] 9, 10, and 11, and a storage order different from the above examples may be used. For example, the vertex information, connection information, and attribute information may be stored in the bitstream in this order. Alternatively, the connection information, connection information, and attribute information may be stored in the bitstream in any of the following orders: connection information, attribute information, and vertex information; vertex information, attribute information, and connection information; attribute information, connection information, and vertex information; or attribute information, vertex information, and connection information.

[0085] Furthermore, each of the connection information, vertex information, and attribute information may be divided into a plurality of data, and the plurality of data may be stored in a cyclical or random order within the bitstream.

[0086] 12 is a block diagram showing a specific example of an encoding / decoding system according to this embodiment. In FIG. 12, the encoding / decoding system includes a three-dimensional data encoding system 110, a three-dimensional data decoding system 210, and an external connector 310.

[0087] The three-dimensional data encoding system 110 includes a controller 111, an input / output processor 112, a three-dimensional data encoder 113, a three-dimensional data generator 115, and a system multiplexer 114. The three-dimensional data decoding system 210 includes a controller 211, an input / output processor 212, a three-dimensional data decoder 213, a system demultiplexer 214, a presenter 215, and a user interface 216.

[0088] In the three-dimensional data encoding system 110, sensor data is input from a sensor terminal to a three-dimensional data generator 115. The three-dimensional data generator 115 generates three-dimensional data, such as point cloud data or mesh data, from the sensor data and inputs it to a three-dimensional data encoder 113.

[0089] For example, the three-dimensional data generator 115 generates vertex information, and generates connection information and attribute information corresponding to the vertex information. The three-dimensional data generator 115 may process the vertex information when generating the connection information and attribute information. For example, the three-dimensional data generator 115 may reduce the amount of data by deleting duplicate vertices, or may transform the vertex information (such as by shifting its position, rotating it, or normalizing it). The three-dimensional data generator 115 may also render the attribute information.

[0090] Furthermore, although the three-dimensional data generator 115 is a component of the three-dimensional data encoding system 110 in FIG. 12, it may be arranged externally and independently of the three-dimensional data encoding system 110.

[0091] The sensor terminal that provides the sensor data for generating the three-dimensional data may be, for example, a moving body such as an automobile, a flying object such as an airplane, a mobile terminal, a camera, etc. Furthermore, a distance sensor such as a LIDAR, a millimeter wave radar, an infrared sensor, or a range finder, a stereo camera, or a combination of multiple monocular cameras may also be used as the sensor terminal.

[0092] The sensor data may be the distance (position) of the object, monocular camera images, stereo camera images, color, reflectance, sensor attitude, orientation, gyro, sensing position (GPS information or altitude), speed, acceleration, sensing time, temperature, air pressure, humidity, or magnetism.

[0093] The three-dimensional data encoder 113 corresponds to the encoding device 100 shown in FIG. 5 and other figures. For example, the three-dimensional data encoder 113 encodes three-dimensional data to generate encoded data. The three-dimensional data encoder 113 also generates control information when encoding the three-dimensional data. The three-dimensional data encoder 113 then inputs the encoded data together with the control information to the system multiplexer 114.

[0094] The encoding method for the three-dimensional data may be an encoding method using geometry or an encoding method using a video codec. Here, the encoding method using geometry may also be referred to as a geometry-based encoding method. The encoding method using a video codec may also be referred to as a video-based encoding method.

[0095] The system multiplexer 114 multiplexes the encoded data and control information input from the 3D data encoder 113 to generate multiplexed data using a specified multiplexing method. The system multiplexer 114 may multiplex other media such as video, audio, subtitles, application data, or document files, or reference time information, along with the encoded data and control information of the 3D data. Furthermore, the system multiplexer 114 may multiplex attribute information related to the sensor data or the 3D data.

[0096] For example, the multiplexed data may have a file format for storage or a packet format for transmission. As these formats, ISOBMFF or a format based on ISOBMFF may be used. Also, MPEG-DASH, MMT, MPEG-2 TS Systems, RTP, or the like may be used.

[0097] The multiplexed data is then output as a transmission signal to the external connector 310 by the input / output processor 112. The multiplexed data may be transmitted as a transmission signal by wire or wirelessly. Alternatively, the multiplexed data is stored in an internal memory or a storage device. The multiplexed data may be transmitted to a cloud server via the Internet or may be stored in an external storage device.

[0098] For example, the transmission or storage of the multiplexed data is performed by a method according to the medium for transmission or storage, such as broadcasting or communication. The communication protocol may be http, ftp, TCP, UDP, IP, or a combination thereof. Furthermore, a pull-type communication method or a push-type communication method may be used.

[0099] For wired transmission, Ethernet (registered trademark), USB, RS-232C, HDMI (registered trademark), coaxial cable, etc. may be used. For wireless transmission, 3GPP (registered trademark), 3G / 4G / 5G defined by IEEE, wireless LAN, Wi-Fi, Bluetooth, or millimeter wave may be used. For broadcasting, for example, DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 may be used.

[0100] The sensor data may be input to the three-dimensional data generator 115 or the system multiplexer 114. The three-dimensional data or encoded data may be output as a transmission signal directly to the external connector 310 via the input / output processor 112. The transmission signal output from the three-dimensional data encoding system 110 is input to the three-dimensional data decoding system 210 via the external connector 310.

[0101] Furthermore, each operation of the three-dimensional data encoding system 110 may be controlled by a controller 111 that executes an application program.

[0102] In the three-dimensional data decoding system 210, a transmission signal is input to an input / output processor 212. The input / output processor 212 decodes multiplexed data having a file format or a packet format from the transmission signal and inputs the multiplexed data to a system demultiplexer 214. The system demultiplexer 214 obtains coded data and control information from the multiplexed data and inputs them to a three-dimensional data decoder 213. The system demultiplexer 214 may extract other media or reference time information from the multiplexed data.

[0103] The three-dimensional data decoder 213 corresponds to the decoding device 200 shown in Fig. 7 etc. For example, the three-dimensional data decoder 213 decodes three-dimensional data from the encoded data based on a predefined encoding method. The three-dimensional data is then presented to the user by the presenter 215.

[0104] Additionally, additional information such as sensor data may be input to the presenter 215. The presenter 215 may present three-dimensional data based on the additional information. Additionally, a user instruction may be input from a user terminal to the user interface 216. Then, the presenter 215 may present three-dimensional data based on the input instruction.

[0105] The input / output processor 212 may acquire the three-dimensional data and the encoded data from the external connector 310 .

[0106] Furthermore, each operation of the three-dimensional data decoding system 210 may be controlled by a controller 211 that executes an application program.

[0107] 13 is a conceptual diagram showing an example of the configuration of point cloud data according to this embodiment. The point cloud data is data of a group of points representing a three-dimensional object.

[0108] Specifically, a point cloud is made up of a plurality of points, and has position information indicating the three-dimensional coordinate position of each point and attribute information indicating the attribute of each point. The position information is also expressed as geometry.

[0109] The type of attribute information may be, for example, color, reflectance, etc. One point may be associated with attribute information of one type, one point may be associated with attribute information of multiple different types, or one point may be associated with attribute information having multiple values ​​for the same type.

[0110] 14 is a conceptual diagram showing an example of a data file of point cloud data according to this embodiment. This example shows a case where there is a one-to-one correspondence between position information items and attribute information items, and shows position information and attribute information for N points that make up the point cloud data. In this example, the position information is information indicating a three-dimensional coordinate position using three axes, x, y, and z, and the attribute information is information indicating a color using RGB. A PLY file or the like can be used as a representative data file for point cloud data.

[0111] 15 is a conceptual diagram showing an example of the configuration of mesh data according to this embodiment. Mesh data is data used in CG (Computer Graphics) and the like, and is three-dimensional mesh data that shows the three-dimensional shape of an object using multiple surfaces. Each surface is also expressed as a polygon, and has a polygonal shape such as a triangle or a rectangle.

[0112] Specifically, a 3D mesh is composed of a plurality of points constituting a point cloud, as well as a plurality of edges and a plurality of faces. Each point is also expressed as a vertex or a position. Each edge corresponds to a line segment connected by two vertices. Each face corresponds to an area surrounded by three or more edges.

[0113] Furthermore, a three-dimensional mesh has position information indicating the three-dimensional coordinate positions of vertices. The position information is also expressed as vertex information or geometry. A three-dimensional mesh also has connection information indicating the relationship between multiple vertices that make up an edge or a face. The connection information is also expressed as connectivity. A three-dimensional mesh also has attribute information indicating the attributes of the vertices, edges, or faces. The attribute information in a three-dimensional mesh is also expressed as texture.

[0114] For example, the attribute information may indicate the color, reflectance, or normal vector for a vertex, edge, or face. The direction of the normal vector may represent the front and back of the face.

[0115] The mesh data may be stored in a data file format such as an object file.

[0116] 16 is a conceptual diagram showing an example of a data file of mesh data according to this embodiment. In this example, the data file includes position information G(1) to G(N) of N vertices that make up the three-dimensional mesh, and attribute information A1(1) to A1(N) of the N vertices. Also, in this example, M pieces of attribute information A2(1) to A2(M) are included. The attribute information items do not need to correspond one-to-one to vertices or faces. Furthermore, attribute information need not exist.

[0117] The connection information is represented by a combination of vertex indices. n[1, 3, 4] indicates a triangular face formed by three vertices, n=1, n=3, and n=4. Also, m[2, 4, 6] indicates that the attribute information of m=2, m=4, and m=6 corresponds to the three vertices, respectively.

[0118] Furthermore, the actual contents of the attribute information may be written in a separate file. A pointer to that content may be associated with a vertex, a face, or the like. For example, attribute information indicating an image for a face may be stored in a two-dimensional attribute map file. The file name of the attribute map and two-dimensional coordinate values ​​in the attribute map may be written in attribute information A2(1) to A2(M). The method of specifying attribute information for a face is not limited to these methods, and any method may be used.

[0119] 17 is a conceptual diagram showing types of three-dimensional data according to this embodiment. Point cloud data and mesh data may represent static objects or dynamic objects. A static object is an object that does not change over time, and a dynamic object is an object that changes over time. A static object may correspond to three-dimensional data for any point in time.

[0120] For example, point cloud data for a given point in time may be referred to as a PCC frame, mesh data for a given point in time may be referred to as a mesh frame, and PCC frames and mesh frames may be simply referred to as frames.

[0121] The area of ​​the object may be limited to a certain range, as in normal video data, or may not be limited, as in map data. The density of points or surfaces may be determined in various ways. Sparse point cloud data or sparse mesh data may be used, or dense point cloud data or dense mesh data may be used.

[0122] Next, encoding and decoding of a point cloud or a three-dimensional mesh will be described. The device, process, or syntax for encoding and decoding vertex information of a three-dimensional mesh in the present disclosure may be applied to encoding and decoding of a point cloud. The device, process, or syntax for encoding and decoding of a point cloud in the present disclosure may be applied to encoding and decoding vertex information of a three-dimensional mesh.

[0123] Furthermore, a device, process, or syntax for encoding and decoding attribute information of a point cloud in the present disclosure may be applied to encoding and decoding connectivity information or attribute information of a three-dimensional mesh.Furthermore, a device, process, or syntax for encoding and decoding connectivity information or attribute information of a three-dimensional mesh in the present disclosure may be applied to encoding and decoding attribute information of a point cloud.

[0124] Furthermore, at least some of the processing may be shared between the encoding and decoding of point cloud data and the encoding and decoding of mesh data, thereby reducing the scale of the circuit and software program.

[0125] 18 is a block diagram showing an example configuration of a three-dimensional data encoder 113 according to this embodiment. In this example, the three-dimensional data encoder 113 includes a vertex information encoder 121, an attribute information encoder 122, a metadata encoder 123, and a multiplexer 124. The vertex information encoder 121, the attribute information encoder 122, and the multiplexer 124 may correspond to the vertex information encoder 101, the attribute information encoder 103, the post-processor 105, etc. in FIG.

[0126] In this example, the three-dimensional data encoder 113 encodes the three-dimensional data according to a geometry-based encoding method, which takes into account the three-dimensional structure. In addition, in the geometry-based encoding method, attribute information is encoded using configuration information obtained in encoding the vertex information.

[0127] Specifically, first, vertex information, attribute information, and metadata included in three-dimensional data generated from sensor data are input to a vertex information encoder 121, an attribute information encoder 122, and a metadata encoder 123, respectively. Here, connectivity information included in the three-dimensional data may be treated in the same way as attribute information. In addition, in the case of point cloud data, position information may be treated as vertex information.

[0128] The vertex information encoder 121 encodes the vertex information into compressed vertex information and outputs the compressed vertex information as encoded data to the multiplexer 124. The vertex information encoder 121 also generates metadata for the compressed vertex information and outputs it to the multiplexer 124. The vertex information encoder 121 also generates configuration information and outputs it to the attribute information encoder 122.

[0129] The attribute information encoder 122 uses the configuration information generated by the vertex information encoder 121 to encode the attribute information into compressed attribute information and outputs the compressed attribute information as encoded data to the multiplexer 124. The attribute information encoder 122 also generates metadata of the compressed attribute information and outputs it to the multiplexer 124.

[0130] The metadata encoder 123 encodes compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The metadata encoded by the metadata encoder 123 may be used to encode vertex information and attribute information.

[0131] The multiplexer 124 multiplexes the compressed vertex information, the compressed vertex information metadata, the compressed attribute information, the compressed attribute information metadata, and the compressed metadata into a bitstream, and then inputs the bitstream to the system layer.

[0132] 19 is a block diagram showing an example configuration of a three-dimensional data decoder 213 according to this embodiment. In this example, the three-dimensional data decoder 213 includes a vertex information decoder 221, an attribute information decoder 222, a metadata decoder 223, and a demultiplexer 224. The vertex information decoder 221, the attribute information decoder 222, and the demultiplexer 224 may correspond to the vertex information decoder 201, the attribute information decoder 203, the preprocessor 204, and the like in FIG.

[0133] In this example, the three-dimensional data decoder 213 decodes three-dimensional data according to a geometry-based encoding method. The three-dimensional structure is taken into consideration in the decoding according to the geometry-based encoding method. Furthermore, in the decoding according to the geometry-based encoding method, attribute information is decoded using configuration information obtained in decoding vertex information.

[0134] Specifically, first, a bitstream is input from the system layer to a demultiplexer 224. The demultiplexer 224 separates compressed vertex information, compressed vertex information metadata, compressed attribute information, compressed attribute information metadata, and compressed metadata from the bitstream. The compressed vertex information and compressed vertex information metadata are input to a vertex information decoder 221. The compressed attribute information and compressed attribute information metadata are input to an attribute information decoder 222. The metadata is input to a metadata decoder 223.

[0135] The vertex information decoder 221 decodes vertex information from the compressed vertex information using metadata of the compressed vertex information. The vertex information decoder 221 also generates configuration information and outputs it to the attribute information decoder 222. The attribute information decoder 222 decodes attribute information from the compressed attribute information using the configuration information generated by the vertex information decoder 221 and the metadata of the compressed attribute information. The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 may be used to decode the vertex information and the attribute information.

[0136] Thereafter, the vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. Note that, for example, this metadata is metadata of the vertex information and attribute information, and can be used in an application program.

[0137] 20 is a block diagram showing another example configuration of the three-dimensional data encoder 113 according to the present embodiment. In this example, the three-dimensional data encoder 113 includes a vertex image generator 131, an attribute image generator 132, a metadata generator 133, a video encoder 134, a metadata encoder 123, and a multiplexer 124. The vertex image generator 131, the attribute image generator 132, and the video encoder 134 may correspond to the vertex information encoder 101 and the attribute information encoder 103 in FIG. 6 , etc.

[0138] In this example, the 3D data encoder 113 encodes the 3D data according to a video-based encoding method. In encoding according to the video-based encoding method, multiple 2D images are generated from the 3D data, and the multiple 2D images are encoded according to a video encoding method. Here, the video encoding method may be High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or the like.

[0139] Specifically, first, vertex information and attribute information included in three-dimensional data generated from sensor data are input to a metadata generator 133. The vertex information and attribute information are then input to a vertex image generator 131 and an attribute image generator 132, respectively. The metadata included in the three-dimensional data is then input to a metadata encoder 123. Here, connectivity information included in the three-dimensional data may be treated in the same way as attribute information. In the case of point cloud data, position information may be treated as vertex information.

[0140] The metadata generator 133 generates map information of a plurality of two-dimensional images from the vertex information and attribute information, and inputs the map information to the vertex image generator 131, the attribute image generator 132, and the metadata encoder 123.

[0141] The vertex image generator 131 generates a vertex image based on the vertex information and map information, and inputs the generated image to the video encoder 134. The attribute image generator 132 generates an attribute image based on the attribute information and map information, and inputs the generated image to the video encoder 134.

[0142] The video encoder 134 encodes the vertex images and attribute images into compressed vertex information and compressed attribute information, respectively, in accordance with a video encoding method, and outputs the compressed vertex information and compressed attribute information as encoded data to the multiplexer 124. The video encoder 134 also generates metadata for the compressed vertex information and metadata for the compressed attribute information, and outputs them to the multiplexer 124.

[0143] The metadata encoder 123 encodes the compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The compressible metadata includes map information. The metadata encoded by the metadata encoder 123 may also be used to encode vertex information and attribute information.

[0144] The multiplexer 124 multiplexes the compressed vertex information, the compressed vertex information metadata, the compressed attribute information, the compressed attribute information metadata, and the compressed metadata into a bitstream, and then inputs the bitstream to the system layer.

[0145] 21 is a block diagram showing another example configuration of the 3D data decoder 213 according to this embodiment. In this example, the 3D data decoder 213 includes a vertex information generator 231, an attribute information generator 232, a video decoder 234, a metadata decoder 223, and a demultiplexer 224. The vertex information generator 231, the attribute information generator 232, and the video decoder 234 may correspond to the vertex information decoder 201 and the attribute information decoder 203 in FIG. 8, etc.

[0146] In this example, the 3D data decoder 213 decodes the 3D data according to a video-based coding method. In the decoding according to the video-based coding method, a plurality of 2D images are decoded according to a video coding method, and 3D data is generated from the plurality of 2D images. Here, the video coding method may be High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or the like.

[0147] Specifically, first, a bitstream is input from the system layer to the demultiplexer 224. The demultiplexer 224 separates compressed vertex information, compressed vertex information metadata, compressed attribute information, compressed attribute information metadata, and compressed metadata from the bitstream. The compressed vertex information, compressed vertex information metadata, compressed attribute information, and compressed attribute information metadata are input to the video decoder 234. The compressed metadata is input to the metadata decoder 223.

[0148] The video decoder 234 decodes the vertex images in accordance with the video encoding method. At this time, the video decoder 234 decodes the vertex images from the compressed vertex information using the metadata of the compressed vertex information. Then, the video decoder 234 inputs the vertex images to the vertex information generator 231. The video decoder 234 also decodes the attribute images in accordance with the video encoding method. At this time, the video decoder 234 decodes the attribute images from the compressed attribute information using the metadata of the compressed attribute information. Then, the video decoder 234 inputs the attribute images to the attribute information generator 232.

[0149] The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 includes map information used to generate vertex information and attribute information. The metadata decoded by the metadata decoder 223 may also be used to decode vertex images and attribute images.

[0150] The vertex information generator 231 reproduces vertex information from the vertex image in accordance with the map information included in the metadata decoded by the metadata decoder 223. The attribute information generator 232 reproduces attribute information from the attribute image in accordance with the map information included in the metadata decoded by the metadata decoder 223.

[0151] Thereafter, the vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. Note that, for example, this metadata is metadata of the vertex information and attribute information, and can be used in an application program.

[0152] Fig. 22 is a conceptual diagram showing a specific example of encoding processing according to this embodiment. Fig. 22 shows a three-dimensional data encoder 113 and a description encoder 148. In this example, the three-dimensional data encoder 113 includes a two-dimensional data encoder 141 and a mesh data encoder 142. The two-dimensional data encoder 141 includes a texture encoder 143. The mesh data encoder 142 includes a vertex information encoder 144 and a connection information encoder 145.

[0153] The vertex information encoder 144, the connection information encoder 145, and the texture encoder 143 may correspond to the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 in FIG.

[0154] For example, the two-dimensional data encoder 141 operates as a texture encoder 143 and generates a texture file by encoding the texture corresponding to the attribute information as two-dimensional data according to an image encoding method or a video encoding method.

[0155] The mesh data encoder 142 also operates as a vertex information encoder 144 and a connectivity information encoder 145, and generates a mesh file by encoding the vertex information and connectivity information. The mesh data encoder 142 may further encode mapping information for textures. The encoded mapping information may then be included in the mesh file.

[0156] The description encoder 148 also generates a description file by encoding a description corresponding to metadata such as text data. The description encoder 148 may encode the description at the system layer. For example, the description encoder 148 may be included in the system multiplexer 114 of FIG. 12 .

[0157] The above operations generate a bitstream containing texture files, mesh files, and description files, which may be multiplexed into the bitstream in file formats such as glTF (Graphics Language Transmission Format) or USD (Universal Scene Description).

[0158] The three-dimensional data encoder 113 may include two mesh data encoders as the mesh data encoder 142. For example, one mesh data encoder encodes vertex information and connectivity information of a static three-dimensional mesh, and the other mesh data encoder encodes vertex information and connectivity information of a dynamic three-dimensional mesh.

[0159] Correspondingly, two mesh files may then be included in the bitstream: for example, one mesh file corresponding to a static 3D mesh and another mesh file corresponding to a dynamic 3D mesh.

[0160] Furthermore, the static three-dimensional mesh may be a three-dimensional mesh of an intraframe coded using intraprediction, and the dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe coded using interprediction. Furthermore, information on the dynamic three-dimensional mesh may be differential information between vertex information or connectivity information of the three-dimensional mesh of an intraframe and vertex information or connectivity information of the three-dimensional mesh of an interframe.

[0161] Fig. 23 is a conceptual diagram showing a specific example of the decoding process according to this embodiment. Fig. 23 shows a three-dimensional data decoder 213, a description decoder 248, and a renderer 247. In this example, the three-dimensional data decoder 213 includes a two-dimensional data decoder 241, a mesh data decoder 242, and a mesh reconstructor 246. The two-dimensional data decoder 241 includes a texture decoder 243. The mesh data decoder 242 includes a vertex information decoder 244 and a connectivity information decoder 245.

[0162] The vertex information decoder 244, the connection information decoder 245, the texture decoder 243, and the mesh reconstructor 246 may correspond to the vertex information decoder 201, the connection information decoder 202, the attribute information decoder 203, and the post-processor 205 in Fig. 8. The presenter 247 may correspond to the presenter 215 in Fig. 12.

[0163] For example, the two-dimensional data decoder 241 operates as a texture decoder 243, and decodes the texture corresponding to the attribute information from the texture file as two-dimensional data in accordance with an image coding method or a video coding method.

[0164] The mesh data decoder 242 also operates as a vertex information decoder 244 and a connectivity information decoder 245 to decode vertex information and connectivity information from the mesh file. The mesh data decoder 242 may further decode mapping information for textures from the mesh file.

[0165] The description decoder 248 also decodes descriptions corresponding to metadata such as text data from the description file. The description decoder 248 may decode the descriptions at the system layer. For example, the description decoder 248 may be included in the system demultiplexer 214 of FIG. 12 .

[0166] The mesh reconstructor 246 reconstructs a 3D mesh from the vertex information, connectivity information, and textures according to the description. The renderer 247 renders and outputs the 3D mesh according to the description.

[0167] Through the above operations, a 3D mesh is reconstructed and output from a bitstream containing a texture file, a mesh file, and a description file.

[0168] The three-dimensional data decoder 213 may include two mesh data decoders as the mesh data decoder 242. For example, one mesh data decoder decodes vertex information and connectivity information of a static three-dimensional mesh, and the other mesh data decoder decodes vertex information and connectivity information of a dynamic three-dimensional mesh.

[0169] Correspondingly, two mesh files may then be included in the bitstream: for example, one mesh file corresponding to a static 3D mesh and another mesh file corresponding to a dynamic 3D mesh.

[0170] Furthermore, the static three-dimensional mesh may be a three-dimensional mesh of an intraframe coded using intraprediction, and the dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe coded using interprediction. Furthermore, information on the dynamic three-dimensional mesh may be differential information between vertex information or connectivity information of the three-dimensional mesh of an intraframe and vertex information or connectivity information of the three-dimensional mesh of an interframe.

[0171] A dynamic 3D mesh coding method is sometimes called DMC (Dynamic Mesh Coding), and a video-based dynamic 3D mesh coding method is sometimes called V-DMC (Video-based Dynamic Mesh Coding).

[0172] The point cloud encoding method is sometimes called PCC (Point Cloud Compression). The point cloud video-based encoding method is sometimes called V-PCC (Video-based Point Cloud Compression). The point cloud geometry-based encoding method is sometimes called G-PCC (Geometry-based Point Cloud Compression).

[0173] <Implementation Example> Fig. 24 is a block diagram showing an implementation example of the encoding device 100 according to this embodiment. The encoding device 100 includes a circuit 151 and a memory 152. For example, multiple components of the encoding device 100 shown in Fig. 5 etc. are implemented by the circuit 151 and memory 152 shown in Fig. 24.

[0174] The circuit 151 is a circuit that performs information processing and is a circuit that can access the memory 152. For example, the circuit 151 is a dedicated or general-purpose electric circuit that encodes a three-dimensional mesh. The circuit 151 may be a processor such as a CPU. Alternatively, the circuit 151 may be a collection of multiple electric circuits.

[0175] The memory 152 is a dedicated or general-purpose memory that stores information used by the circuit 151 to encode the three-dimensional mesh. The memory 152 may be an electric circuit and may be connected to the circuit 151. The memory 152 may also be included in the circuit 151. The memory 152 may also be a collection of multiple electric circuits. The memory 152 may also be a magnetic disk, an optical disk, or the like, and may also be expressed as a storage, a recording medium, or the like. The memory 152 may also be a non-volatile memory or a volatile memory.

[0176] For example, the memory 152 may store a three-dimensional mesh or a bitstream, or may store a program for the circuit 151 to encode the three-dimensional mesh.

[0177] Note that the encoding device 100 does not necessarily have to implement all of the components shown in Figure 5 and the like, and does not necessarily have to perform all of the processes shown here. Some of the components shown in Figure 5 and the like may be included in another device, and some of the processes shown here may be executed by another device. Furthermore, the encoding device 100 may implement any combination of the components of the present disclosure, and may perform any combination of the processes of the present disclosure.

[0178] Fig. 25 is a block diagram showing an example implementation of a decoding device 200 according to this embodiment. The decoding device 200 includes a circuit 251 and a memory 252. For example, multiple components of the decoding device 200 shown in Fig. 7 and other figures are implemented by the circuit 251 and memory 252 shown in Fig. 25.

[0179] The circuit 251 is a circuit that performs information processing and is a circuit that can access the memory 252. For example, the circuit 251 is a dedicated or general-purpose electric circuit that decodes a three-dimensional mesh. The circuit 251 may be a processor such as a CPU. Alternatively, the circuit 251 may be a collection of multiple electric circuits.

[0180] The memory 252 is a dedicated or general-purpose memory that stores information for the circuit 251 to decode the 3D mesh. The memory 252 may be an electric circuit and may be connected to the circuit 251. The memory 252 may also be included in the circuit 251. The memory 252 may also be a collection of multiple electric circuits. The memory 252 may also be a magnetic disk, an optical disk, or the like, and may also be expressed as a storage, a recording medium, or the like. The memory 252 may also be a non-volatile memory or a volatile memory.

[0181] For example, the memory 252 may store a three-dimensional mesh or a bitstream, or may store a program for the circuit 251 to decode the three-dimensional mesh.

[0182] Note that the decoding device 200 does not necessarily have to implement all of the components shown in Figure 7 and the like, and does not necessarily have to perform all of the processes shown here. Some of the components shown in Figure 7 and the like may be included in another device, and some of the processes shown here may be executed by another device. Furthermore, the decoding device 200 may implement any combination of the components of the present disclosure, and may perform any combination of the processes of the present disclosure.

[0183] The encoding method and the decoding method including the steps performed by each component of the encoding device 100 and the decoding device 200 of the present disclosure may be executed by any device or system. For example, part or all of the encoding method and the decoding method may be executed by a computer including a processor, a memory, an input / output circuit, etc. In this case, the encoding method and the decoding method may be executed by the computer executing a program for causing the computer to execute the encoding method and the decoding method.

[0184] Alternatively, the program or the bitstream may be recorded on a non-transitory computer-readable recording medium such as a CD-ROM.

[0185] An example of a program may be a bitstream. For example, a bitstream including an encoded three-dimensional mesh includes syntax elements for causing the decoding device 200 to decode the three-dimensional mesh. The bitstream then causes the decoding device 200 to decode the three-dimensional mesh according to the syntax elements included in the bitstream. Thus, the bitstream may play a role similar to that of a program.

[0186] The bitstream may be an encoded bitstream containing the encoded 3D mesh, or may be a multiplexed bitstream containing the encoded 3D mesh and other information.

[0187] Furthermore, each component of the encoding device 100 and the decoding device 200 may be configured with dedicated hardware, general-purpose hardware that executes the above-mentioned programs, or a combination of these. The general-purpose hardware may be configured with a memory in which the programs are recorded and a general-purpose processor that reads and executes the programs from the memory. Here, the memory may be a semiconductor memory or a hard disk, and the general-purpose processor may be a CPU.

[0188] Furthermore, the dedicated hardware may be configured with a memory, a dedicated processor, etc. For example, the dedicated processor may execute the encoding method and the decoding method by referring to a memory for recording data.

[0189] Furthermore, as described above, each component of the encoding device 100 and the decoding device 200 may be an electric circuit. These electric circuits may form a single electric circuit as a whole, or each may be a separate electric circuit. Furthermore, these electric circuits may correspond to dedicated hardware, or may correspond to general-purpose hardware that executes the above-mentioned programs, etc. Furthermore, the encoding device 100 and the decoding device 200 may be implemented as an integrated circuit.

[0190] Furthermore, the encoding device 100 may be a transmitting device that transmits the three-dimensional mesh, and the decoding device 200 may be a receiving device that receives the three-dimensional mesh.

[0191] Displacement Encoding and Decoding The following terminology is used here by way of example:

[0192] (1) Image An image is a data unit made up of a set of pixels, and includes a picture or a block smaller than a picture. Images include both moving images and still images.

[0193] (2) Picture A picture is a unit of image processing that is made up of a set of pixels, and is also called a frame or field.

[0194] (3) Block A block is a processing unit consisting of a specific number of pixels. The term shown in the following example is also used for a block. The shape of a block is not particularly limited. A block may be, for example, a rectangular shape of M×N pixels or a square shape of M×M pixels. A block may also be a triangular shape, a circular shape, or another shape. Examples of blocks are as follows:

[0195] Slice, tile, or brick CTU, superblock, or basic division unit VPDU, processing division unit for hardware CU, processing block unit, prediction block unit (PU), or orthogonal transform block unit (TU) Sub-block

[0196] (4) Pixel or Sample A pixel or sample is the smallest point of an image, in other words, the smallest unit. Pixels or samples include not only pixels at integer positions, but also pixels at sub-pixel positions generated based on pixels at integer positions.

[0197] (5) Pixel Value or Sample Value: A pixel value or sample value is a unique value of a pixel. The pixel value or sample value may include a luma value, a chroma value, or an RGB gradation level, and may also include a depth value or a binary value of 0 or 1.

[0198] (6) Flags A flag indicates one or more bits. A flag is, for example, a parameter or index represented by two or more bits. A flag may indicate not only a value represented by a binary number, but also a value represented by a number other than a binary number.

[0199] (7) Signal: A signal is something that is symbolized or coded to transmit information. A signal includes a discrete digital signal or a continuous analog signal.

[0200] (8) Stream or Bit Stream A stream or bit stream is a digital data sequence that indicates the flow of digital data. A stream or bit stream may be a single stream, or may be configured to include multiple streams with multiple layers. A stream or bit stream may be transmitted by serial communication using a single transmission path, or may be transmitted by packet communication using multiple transmission paths.

[0201] (9) Difference: For scalar quantities, the difference can include simple difference (x-y) and difference calculations. The difference can include absolute difference (|x-y|), squared difference (x^2-y^2), square root difference (√(x-y)), weighted difference (ax-by, where a and b are constants), or offset difference (x-y+a, where a is an offset).

[0202] (10) Sum: For scalar quantities, sums can include simple sum (x + y) and addition calculations. The sum can also include absolute sum (|x + y|), sum of squares (x^2 + y^2), square root of the sum (√(x + y)), weighted sum (ax + by, where a and b are constants), or offset sum (x + y + a, where a is an offset).

[0203] (11) "Based on" The expression "based on something" means that something other than that "something" may be taken into consideration. Also, "based on" can be used both when a direct result is obtained and when a result is obtained through an intermediate result.

[0204] (12) "Used" or "Using" The phrases "something was used" or "used something" mean that something other than the "something" may be taken into consideration. The phrases "used" or "used" may be used both in cases where a direct result is obtained and in cases where a result is obtained via an intermediate result.

[0205] (13) Prohibition "Prohibit" can be rephrased as "not permitted." Also, "not prohibited / prohibited" or "permitted / permitted" does not necessarily mean "obligation."

[0206] (14) "Restriction" or "Limitation" "Restriction" or "Limitation" can be rephrased as "not permitted / not allowed" or "not permitted / permitted." Furthermore, "prohibited / not prohibited" or "not permitted / permitted" does not necessarily mean "obligation." Furthermore, what is prohibited quantitatively or qualitatively may be either partial or total.

[0207] (15) Chroma The term chroma is an adjective, represented by the symbols Cb or Cr, that indicates that a sample array or a single sample represents one of the two color difference signals associated with a primary color. The term chroma is sometimes used instead of the term chrominance.

[0208] (16) Luma The term luma is an adjective, denoted by the symbols or subscripts Y or L, that indicates that a sample array or a single sample represents a monochrome signal for a primary color. The term luma is sometimes used instead of the term luminance.

[0209] The encoding / decoding system of this embodiment will be described below.

[0210] A typical three-dimensional model (also called a 3D model) digitally represents an object so that a user can explore the model using zoom, pan, and rotation in all three dimensions while it is rendered over time. One way to construct such a representation is to build a 3D mesh using triangles. The model stores the positions of the triangle vertices, their connectivity to each other, and their associated attributes (such as normals or UV patches).

[0211] Storing all this information in uncompressed form requires a very large storage space and therefore a very large bandwidth for transmission. The triangles that form the mesh often have repeating patterns and similar properties, especially in temporal and spatial neighborhoods. These repetitions can be exploited to develop efficient encoding and decoding methods for storage and transmission. One such encoding and decoding method is Video-based Dynamic Mesh Coding (V-DMC).

[0212] 26 is a block diagram showing another example of the configuration of the encoding / decoding system according to this embodiment. As shown in FIG. 26, the encoding / decoding system includes an encoding device 100 and a decoding device 200.

[0213] The encoding / decoding system accepts input three-dimensional meshes (also called 3D meshes) in the form of three-dimensional coordinates of vertices (vertex information), connectivity (connection information) and associated attributes (attribute information), which may include texture maps as well as geometry.

[0214] The encoding device 100 takes an input 3D mesh (also referred to as an input 3D mesh or input mesh) in the form of three-dimensional coordinates of vertices, connectivity, and associated attributes. The encoding device 100 encodes all associated information into a stream. The stream may consist of a single bitstream or multiple bitstreams.

[0215] The network 300 transmits the stream generated by the encoding device 100 to the decoding device 200. The network 300 may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof. Furthermore, the network 300 is not necessarily limited to a two-way communication network, but may also be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Instead of the network 300, a recording medium such as a digital versatile disc (DVD) or a blue-ray disc (BD) on which a stream is recorded may be used.

[0216] The stream is transmitted to a decoding device 200 via a network 300. The decoding device 200 decodes the bitstream and generates a 3D mesh using the 3D coordinates, connectivity, and associated attributes of the decoded vertices. The decoding device 200 outputs the generated 3D mesh (also referred to as an output 3D mesh or output mesh).

[0217] FIG. 27 is a diagram showing another example of the configuration of the encoding device 100.

[0218] As shown in FIG. 27, the encoding device 100 includes a preprocessor 1103 and a compressor 1106 .

[0219] The encoding device 100 reads an input mesh 1101 and an attribute map 1102 and passes them to a preprocessor 1103. The preprocessor 1103 processes the input mesh to extract a base mesh 1104 and displacement data 1105. The attribute map 1102, along with the extracted base mesh 1104 and displacement data 1105, are passed to a compressor 1106.

[0220] The compressor 1106 also compresses the base mesh 1104, the displacement data 1105, and the attribute map 1102 to generate a bitstream 1107. The compressor 1106 can transmit additional information to the decoding device 200 by further including metadata 1108 in the bitstream 1107.

[0221] FIG. 28 is a diagram showing another example of the configuration of the decoding device 200.

[0222] As shown in FIG. 28, the decoding device 200 includes a decompressor 2102 and a post-processor 2106 .

[0223] The decoding device 200 reads a bitstream 2101 and passes it to a decompressor 2102. The decompressor 2102 decompresses a base mesh 2103, displacement data 2104, and an attribute map 2108 from the bitstream 2101 and passes them to a post-processor 2106. An example of the displacement data 2104 is a displacement vector.

[0224] The post-processor 2106 also processes the base mesh 2103 according to the displacement data 2104 and the attribute map 2108 to generate an output mesh 2107. The post-processor 2106 may further use information from the metadata 2105 to generate the output mesh 2107.

[0225] FIG. 29 is a block diagram showing yet another example configuration of the encoding device 100 according to this embodiment.

[0226] In this example, the encoding device 100 comprises a volumetric capturer 511, a projector 512, a base mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and optionally one or more other type encoders 516.

[0227] The volumetric capturer 511 captures content and outputs the captured content to the projector 512 .

[0228] The projector 512 projects the content onto an input mesh (a 3D mesh frame) that includes geometry coordinates (vertex coordinates indicating the positions of the vertices), texture coordinates, and connectivity (connectivity information). The data is output to a base mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and optionally one or more other type encoders 516. Each encoder compresses the data into a bitstream.

[0229] FIG. 30 is a block diagram showing yet another example configuration of the decoding device 200 according to this embodiment.

[0230] In this example, the decoding device 200 comprises a base mesh decoder 613 , a displacement decoder 614 , an attribute decoder 615 , one or more other type decoders 616 , and a 3D reconstructor 617 .

[0231] The bitstream is sent to a base mesh decoder 613, a displacement decoder 614, an attribute decoder 615, and optionally one or more other type decoders 616. These decoders decode the bitstream to generate data (decoded data) including geometry coordinates, texture coordinates, connectivity, etc. The decoded data is then sent to a 3D reconstructor 617, which reconstructs an output mesh (a 3D mesh frame).

[0232] The encoding process performed by the encoding device 100 will be described in detail below.

[0233] Fig. 31 is a flow diagram showing the processing of the encoding device 100. Fig. 32 is an explanatory diagram conceptually showing the encoding of mesh frames. The processing of the encoding device 100 will be described with reference to Figs. 31 and 32.

[0234] In step S101, the encoding device 100 reads a 3D mesh frame, which is an input mesh frame, and its attributes. The input mesh frame is a mesh frame input to the encoding device 100. An example of the 3D mesh frame that is an input mesh frame is shown as mesh frame 1301 (see FIG. 32 ).

[0235] In step S102, the encoding device 100 performs a decimation process on the input mesh frame read in step S101 to generate a base mesh frame having fewer vertices than the input mesh frame. The base mesh frame generated by decimating the mesh frame 1301 is shown as a base mesh frame 1302 (see FIG. 32).

[0236] In step S103, the encoding device 100 calculates displacement information that the decoding device 200 uses to reconstruct a mesh frame. The displacement information corresponds to a displacement vector directed from a vertex of the base mesh frame generated in step S102 to a vertex of the input mesh frame. One method for calculating the displacement information is to subtract the coordinates of the vertex of the base mesh frame from the coordinates of the vertex of the input mesh frame. The displacement information calculated from the mesh frame 1301 and the base mesh frame 1302 is shown as displacement information 1303 (see FIG. 32). The displacement information 1303 is in vector format, in other words, expressed as a displacement vector.

[0237] In step S104, the encoding device 100 encodes the base mesh frame generated in step S102, the displacement information generated in step S103, and the attributes of the input mesh frame into a bitstream (corresponding to a compressed bitstream). An example of the bitstream is shown as bitstream 1304 (see FIG. 32).

[0238] Specifically, the bitstream 1304 includes vertex coordinates and connectivity information for vertices A, C, E, and F, displacement information, a video bitstream including texture data, and a compressed attribute map (see FIG. 32). The displacement information includes displacement information for displacing vertices based on vertex coordinates obtained from the subdivided base mesh frame. The compressed attribute map is texture coordinates for applying texture data to a mesh frame reconstructed using the base mesh frame and the displacement information.

[0239] The decoding process performed by the decoding device 200 will be described in detail below.

[0240] Fig. 33 is a flow diagram showing the processing of the decoding device 200. Fig. 34 is an explanatory diagram conceptually showing the decoding of a mesh frame (3D mesh). The processing of the decoding device 200 will be described with reference to Figs. 33 and 34.

[0241] In step S201, the decoding device 200 decodes a base mesh frame and attributes from a bitstream (corresponding to a compressed bitstream). An example of the decoded base mesh frame (corresponding to a decoded base mesh frame) is shown as a decoded base mesh frame 2301 (see FIG. 34).

[0242] In step S202, the decoding device 200 generates subdivided vertices by performing a subdivision process on the base mesh frame decoded in step S201. An example of a base mesh frame (mesh frame) including subdivided vertices is shown as a base mesh frame 2302 (see FIG. 34).

[0243] In step S203, the decoding device 200 decodes the disparity information from the bitstream (corresponding to the compressed bitstream). An example of the decoded disparity information is shown as disparity information 2303 (see FIG. 34). The disparity information 2303 is in vector format, in other words, expressed as a disparity vector.

[0244] In step S204, the decoding device 200 reconstructs the shape of the mesh frame by moving the vertices of the base mesh frame, including the subdivided vertices, to new positions using the displacement information, and then restores the mesh frame by applying attribute information. An example of the attribute is texture. An example of the reconstructed mesh frame is shown as mesh frame 2304 (see FIG. 34 ).

[0245] FIG. 35 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.

[0246] FIG. 35 shows an example of a block diagram of a general intra-decoding system.

[0247] The decoding device shown in FIG. 35 comprises a demultiplexer 1231, a switch 1232, a static mesh decoder 1233, a mesh buffer 1234, a motion decoder 1235, a base mesh reconstructor 1236, an inverse quantizer 1237, a video decoder 1238, an image unpacker 1239, an inverse quantizer 1240, an inverse wavelet transformer 1241, a reconstructor 1242, a video decoder 1243, and a color converter 1244.

[0248] The demultiplexer 1231 receives the compressed bitstream and separates it into compressed data for the base mesh, video containing displacement data (also called displacement bitstream), and video containing attribute data (also called attribute bitstream). The compressed data for the base mesh is passed to a switch 1232. The switch 1232 determines whether to perform intra-decoding or inter-decoding based on parameters in the bitstream.

[0249] If an intra-decoding process is selected, the bitstream is passed to a static mesh decoder 1233, which generates a quantized base mesh. The static mesh decoder 1233 is, for example, a decoder that uses an edge breaker algorithm to decode 3D mesh data. The static mesh decoder 1233 generates a quantized base mesh from the bitstream. The quantized base mesh generated by the static mesh decoder 1233 is stored in a mesh buffer 1234 for reference when an inter-decoding process is selected.

[0250] If inter-decoding is selected, switch 1232 passes compressed data for the base mesh to motion decoder 1235. Motion decoder 1235 receives a previously decoded quantized base mesh and decodes motion data representing the differences in vertex coordinates between the quantized base mesh stored in mesh buffer 1234 and the current quantized base mesh. The motion data and the quantized base mesh stored in mesh buffer 1234 are used by base mesh reconstructor 1236 to reconstruct the current quantized base mesh. The quantized base mesh resulting from either inter-decoding or intra-decoding is passed to inverse quantizer 1237 to obtain a decoded base mesh.

[0251] The video containing the displacement data is passed to a video decoder 1238, since the bitstream contains the displacement data in an image format with two chroma information and one luma information. The video decoder 1238 decodes the data using a video frame decompression method. Alternatively, the displacement data can be decoded using an arithmetic decoder. This decompressed data is passed to an image unpacker 1239, which extracts wavelet coefficients associated with each vertex from the image-format decompressed data. An inverse quantizer 1240 dequantizes the quantized wavelet coefficients into the three components associated with each vertex. An inverse wavelet transformer 1241 inversely transforms the result to finally obtain decoded displacement data. The decoded displacement data and the decoded base mesh are passed to a reconstructor 1242, which performs edge refinement on the decoded base mesh and displaces the vertices using the decoded displacement data to obtain a decoded mesh.

[0252] The video containing the attribute data is passed to another video decoder 1243 to obtain a decoded attribute bitstream, which is further processed in a color converter 1244 for color space and color format conversion to obtain a decoded attribute map.

[0253] FIG. 36 is a block diagram showing an example of the configuration of an encoding device according to this embodiment.

[0254] First, the encoding device obtains the base mesh bitstream, the displacement bitstream, and the attribute bitstream resulting from the 3D mesh preprocessing step.

[0255] The encoding device shown in FIG. 36 comprises a quantizer 1261, a switch 1262, a static mesh encoder 1263, a mesh buffer 1264, a motion encoder 1265, a base mesh reconstructor 1266, a displacement data updater 1267, a wavelet transformer 1268, a quantizer 1269, an image packer 1270, a video encoder 1271, a color converter 1272, and a video encoder 1273.

[0256] The base mesh (specifically, the position information of the multiple vertices that make up the base mesh) is first quantized by a quantizer 1261. The quantized base mesh (base mesh data) is output to a switch 1262, which determines whether intra-coding or inter-coding is to be performed. If intra-coding is selected, the quantized base mesh (base mesh bitstream) is output to a static mesh encoder 1263, which generates a quantized base mesh. An example of the static mesh encoder 1263 is an encoder that uses an edge breaker algorithm to encode 3D mesh data. This encoded (quantized) base mesh is stored in a mesh buffer 1264 for reference when inter-coding is selected. If inter-coding is selected, the switch 1262 outputs compressed data related to the base mesh to a motion encoder 1265. The motion data and static 3D mesh in the mesh buffer 1264 are used by a base mesh reconstructor 1266 to reconstruct the currently quantized base mesh.

[0257] The displacement data is output to a displacement data updater 1267, where it is updated based on the quantized static base mesh or the reconstructed inter-coded base mesh. Next, a wavelet transformer 1268 performs a transformation process, followed by quantization in a quantizer 1269. The quantized displacement data is packed into an image in an image packer 1270, and finally encoded in a video coder 1271. The encoded displacement data is output to a multiplexer 1274.

[0258] The attribute information (e.g., attribute map) is output to a color converter 1272 for color space and color format conversion, and the converted attribute information is coded by a video coder 1273 and output to a multiplexer 1274.

[0259] The multiplexer 1274 acquires data related to the encoded base mesh (compressed data related to the base mesh), video data including encoded displacement data, and video data including attribute information such as an encoded attribute map, and generates a bitstream (compressed bitstream) including these acquired data. The generated compressed bitstream is output to, for example, a decoding device.

[0260] FIG. 37 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.

[0261] FIG. 37 illustrates an example of a reconstructor that obtains a decoded 3D mesh 1256 from a decoded base mesh 1251 and decoded displacement data 1254.

[0262] The decoded base mesh 1251 is passed to a subdivider 1252 .

[0263] The subdivision unit 1252 subdivides any two connected vertices in the entire 3D mesh by adding a new vertex between them. This process can be repeated several times to include vertices created in previous subdivision steps to generate a predefined number of vertices. Each subdivision iteration across the 3D mesh generates a new level of detail (LoD). The subdivided mesh 1253 and the decoded displacement data 1254 are passed to a displacer 1255. The displacer 1255 generates a decoded 3D mesh 1256 by moving each vertex to a new position according to the corresponding displacement data.

[0264] The subdivision is described below and is performed, for example, by subdivision unit 1252.

[0265] FIG. 38 is an explanatory diagram showing an example of subdivision.

[0266] The base mesh shown in FIG. 38(a) includes vertices A, B, and C and connectivity information indicating their connectivity.

[0267] 38(b) shows a mesh generated by the first subdivision, in other words, the mesh after the first subdivision. In the first subdivision, the subdivider generates vertices D, E, and F and connectivity information indicating their connectivity. The mesh generated by the subdivider is also referred to as LoD1 or first LoD.

[0268] Vertex D of the mesh after the first subdivision is a vertex generated by subdivision based on vertices A and B. Similarly, vertex F is a vertex generated by subdivision based on vertices B and C. Vertex E is a vertex generated by subdivision based on vertices A and C.

[0269] As an example, vertex D may be the midpoint of line segment AB (in other words, side AB) connecting vertices A and B that were the basis for its generation. Similarly, vertex E may be the midpoint of line segment AC. Vertex F may be the midpoint of line segment BC.

[0270] 38(c) shows the mesh generated by the second subdivision, i.e., the mesh after the second subdivision. In the second subdivision, the subdivider generates vertices G, H, I, J, K, L, M, N, and O and connectivity information indicating their connectivity. The mesh generated by the subdivider is also called LoD2 or second LoD.

[0271] Vertex G of the mesh after the second subdivision is a vertex generated by subdivision based on vertices A and D. Similarly, vertex H is a vertex generated by subdivision based on vertices A and E. Vertex I is a vertex generated by subdivision based on vertices B and D. Vertex J is a vertex generated by subdivision based on vertices D and F. Vertex K is a vertex generated by subdivision based on vertices E and F. Vertex L is a vertex generated by subdivision based on vertices C and E. Vertex M is a vertex generated by subdivision based on vertices B and F. Vertex N is a vertex generated by subdivision based on vertices C and F. Vertex O is a vertex generated by subdivision based on vertices D and E.

[0272] As an example, vertex G may be the midpoint of line segment AD (in other words, side AD) connecting vertices A and D, which were the source of its generation. Similarly, vertex H may be the midpoint of line segment AE. vertex I may be the midpoint of line segment BD. vertex J may be the midpoint of line segment DF. vertex K may be the midpoint of line segment EF. vertex L may be the midpoint of line segment CE. vertex M may be the midpoint of line segment BF. vertex N may be the midpoint of line segment CF. vertex O may be the midpoint of line segment DE.

[0273] The displacement of vertices will be described below with reference to Figures 39 and 40. The displacement of vertices is performed by the reconstructor.

[0274] Fig. 39 is an explanatory diagram showing an example of displacement of vertices after subdivision, and Fig. 40 is an explanatory diagram showing an example of vertices of an original mesh.

[0275] The base mesh shown in FIG. 39(a) includes vertices A, B, C, and Z, and connectivity information indicating their connectivity.

[0276] 39(b) shows a mesh generated by the first subdivision, in other words, a mesh after the first subdivision (i.e., the first LoD). In the first subdivision, the subdivider generates vertices S, T, U, X, or Y and connectivity information indicating their connectivity. The vertices S, T, U, X, or Y are similar to the vertices D, E, and F shown in FIG. 38(b).

[0277] 39(c) shows a mesh generated by the second subdivision, in other words, a mesh after the second subdivision (i.e., the second LoD). In the second subdivision, the subdivider generates vertices D, E, F, G, and H and connectivity information indicating their connectivity. Vertices D, E, F, G, and H are similar to vertices G, H, I, J, K, L, M, N, and O shown in FIG. 38(c).

[0278] Figure 39(d) shows a mesh including the vertices after they have been displaced after subdivision, with vertices A, B, C, D, E, F, G, H, S, T, U, X, Y, and Z shown in Figure 39(d) being located at positions displaced using displacement information from the positions of the vertices shown in Figure 39(c).

[0279] The original mesh shown in FIG. 40 is an example of the mesh input to the encoding device 100, that is, the mesh before encoding.

[0280] The mesh shown in Fig. 39 has a shape similar to that of the original mesh shown in Fig. 40. The displacement information is generated by the displacement vector calculator 1207 of the encoding device 100 as information indicating the displacement from the vertices of the base mesh to the vertices of the original mesh, and therefore, by reconstructing the mesh using the displacement information thus generated, a mesh having a shape similar to that of the original mesh is generated.

[0281] The decoding device 200 can output the mesh shown in FIG.

[0282] Next, the division of a mesh into sub-meshes will be described with reference to FIGS.

[0283] A mesh can be divided into smaller parts and coded separately, with the vertices of the mesh being divided in such a way that the coordinates and connectivity of the vertices in each part can be coded independently.

[0284] Fig. 41 is an explanatory diagram showing an example of a mesh, and Fig. 42 is an explanatory diagram showing an example of dividing a mesh into sub-meshes.

[0285] The mesh shown in FIG. 41 is the original mesh, which is sometimes called a full mesh in contrast to a sub-mesh.

[0286] Figure 42 shows how the full mesh shown in Figure 41 is divided into two sub-meshes. For vertices A, B, and C of the full mesh (see Figure 41), vertex A is duplicated to vertices A1 and A2, vertex B is duplicated to vertices B1 and B2, and vertex C is duplicated to vertices C1 and C2, thereby creating two sub-meshes (i.e., a first sub-mesh and a second sub-mesh) from the full mesh. The first sub-mesh and the second sub-mesh are each independently decodable meshes.

[0287] Packing of displacement information into image frames will be described below with reference to FIGS. 43, 44 and 45.

[0288] 43, 44, and 45 are explanatory diagrams showing examples of packing displacement information into image frames. Note that image frames can also be called video frames.

[0289] The vertex displacement data is encoded as image frame data by being mapped to each component of a YUV format image frame (i.e., each of the Y component (Y Plane), U component (U Plane), and V component (V Plane)). This case will be described below as an example. As another example, the vertex displacement data may be encoded as image frame data by being mapped to each component of an RGB format image frame (each of the R component, G component, and B component).

[0290] The decoding device 200 can use an image encoding module to extract the displacement data. The displacement data can be in the form of X, Y, or Z components in a global coordinate system (e.g., a Cartesian coordinate system), or normal, tangential, or both tangential components in a local coordinate system. Methods for mapping the displacement data to an image frame include the following:

[0291] For example, in the first method, the displacement data is arranged in the image frame in scan order. An example of packing of the displacement data in this case is shown in Figure 43. The displacement data is directly mapped to the image frame according to a predefined scan order.

[0292] Note that since an image frame has a fixed height and width, it may happen that the displacement data does not fit perfectly in the frame, in which case the remaining part of the image frame is padded with padding data (see Figure 43).

[0293] For example, in the second method, the displacement data is separated into multiple LoDs and mapped to the Y, U, and V components of the image frame. An example of packing of the displacement data in this case is shown in Figure 44. Here, the displacement data of the image frame of the next LoD starts immediately after the displacement data of the previous LoD ends. As in the first method, if the displacement data does not fit exactly into the image frame, padding is performed at the end of the image frame (see Figure 44).

[0294] For example, in the third method, displacement data corresponding to the LoD is mapped to the Y component, U component, and V component of the image frame in a manner different from that in the second method. An example of packing of the displacement data in this case is shown in Figure 45. In this way, each LoD can be decoded independently. In the third method, middle padding is performed on the displacement data of each LoD, and CTU alignment is performed together with padding at the end of the video frame (see Figure 45).

[0295] Next, the encoding device 100 and the decoding device 200 when a mesh is divided into a plurality of sub-meshes will be described.

[0296] Fig. 46 is a diagram showing another example of the configuration of the encoding device 100 according to an embodiment. Specifically, Fig. 46 is a diagram showing the configuration of a submesh encoding device that performs encoding processing when an input mesh 1101 is divided into divided meshes (plurality of submeshes). For example, the submesh encoding device includes a plurality of encoding devices 100.

[0297] An input mesh 1101 (full mesh) input to the submesh encoding device is divided into multiple meshes (submeshes). The multiple submeshes are input to, for example, multiple encoding devices 100. Each of the multiple submeshes may be input to one of the multiple encoding devices 100. For example, the submesh encoding device divides the input mesh 1101 into multiple submeshes and inputs the divided multiple submeshes to the multiple encoding devices 100.

[0298] After the image is divided into a plurality of sub-meshes, coding processing is performed on the boundaries of the sub-meshes (processing of overlapping sub-meshes).

[0299] For example, for each sub-mesh, preprocessing is performed by the preprocessor 1103, and a base mesh, displacement data, and metadata are generated and encoded.

[0300] The encoding device 100 may be realized by such a submesh encoding device configuration. That is, the encoding device 100 may be configured to include a plurality of preprocessors 1103 and compressors 1106, and may perform a predetermined process on the submesh for each of a plurality of pairs of preprocessors 1103 and compressors 1106. Furthermore, the number of pairs of preprocessors 1103 and compressors 1106 included in the encoding device 100 may be any number and is not particularly limited.

[0301] FIG. 47 is a diagram illustrating an example of the configuration of the preprocessor 1103 according to the embodiment.

[0302] The preprocessor 1103 includes, for example, a base mesh generator 1401 , a subdivision unit 1402 , and a displacement data generator 1403 .

[0303] First, in the preprocessor 1103, a base mesh is generated by the base mesh generator 1401.

[0304] The base mesh is then subdivided in a predetermined manner by a subdivider 1402 to generate a subdivided mesh (or a subdivided base mesh) that is the subdivided base mesh.

[0305] Displacement data is generated by a displacement data generator 1403 from the sub-divided mesh and the sub-mesh that is the input mesh 1101 after division.

[0306] The displacement data is, for example, a difference vector between the input mesh 1101 and the sub-division mesh.

[0307] The same subdivision method as that used in encoding is used in decoding as well.

[0308] Furthermore, for example, the encoding device 100 may transmit to the decoding device 200 a subdivision method for encoding and parameters used in the subdivision.

[0309] Fig. 48 is a diagram showing another example of the configuration of the decoding device 200 according to the embodiment. Specifically, Fig. 48 is a diagram showing the configuration of a submesh decoding device, which is a device that performs decoding processing when a bitstream 2101 includes multiple submeshes. For example, the submesh decoding device includes multiple decoding devices 200. For example, the submesh decoding device includes multiple decoding devices 200 and a combiner 2109.

[0310] The coded data for each submesh included in the bit stream 2101 is input to the decompressor 2102 of each decoding device 200 .

[0311] Also, for example, the post-processor 2106 performs the processing of the reconstructor described above.

[0312] In the post-processor 2106, for each decoded sub-mesh, the base mesh is subdivided and the displacement vector is added to the subdivided base mesh to reconstruct the sub-mesh.

[0313] That is, the post-processor 2106 performs the above-described reconstructor processing for each sub-mesh.

[0314] The combiner 2109 combines (merges) the sub-meshes restored by the respective decoding devices 200 to reconstruct the full mesh (output mesh 2107) before division.

[0315] Note that the decoding device 200 may be realized by such a submesh decoding device configuration. That is, the decoding device 200 may be configured to include a plurality of decompressors 2102 and post-processors 2106, and each of the plurality of pairs of decompressors 2102 and post-processors 2106 may perform a predetermined process on the encoded data for each submesh. Furthermore, the number of pairs of decompressors 2102 and post-processors 2106 included in the decoding device 200 may be any number and is not particularly limited.

[0316] Fig. 49 is a diagram showing a specific example of the configuration of the decoding device 200 according to the embodiment. Specifically, Fig. 49 shows a specific configuration of the post-decoder 2306 out of the decoder 2305 and post-decoder 2306 included in the decoding device 200.

[0317] The post-decoder 2306 comprises a pre-reconstructor 2307 , a reconstructor 2308 , a post-reconstructor 2309 and an adaptor 2310 .

[0318] The processing in the post-decoder 2306 is optional depending on the application. An example of the processing in the post-decoder 2306 (post-decoding processing) is conversion of the decoded data to a nominal format, such as video conversion from YUV space to RGB space. The post-decoding processing can be encapsulated into multiple processes, such as pre-reconstruction, reconstruction, post-reconstruction, and adaptation.

[0319] The pre-reconstructor 2307 performs a pre-reconstruction process, for example, upscaling the normalized texture coordinates to match the dimensions of the texture image in the context of video-based dynamic mesh coding.

[0320] The reconstructor 2308 performs a reconstruction process, which is invoked, for example, on the decoded atlas frame, the decoded base mesh frame, the decoded video frame, and syntax elements associated with the same mesh sequence. The output of the reconstruction process is a series of reconstructed mesh frames prior to a post-reconstruction process.

[0321] The post-reconstructor 2309 performs post-reconstruction operations, e.g., in the context of video-based dynamic mesh coding, which perform a number of smoothing operations on the reconstructed mesh frames, such as collapsing edges in the mesh or adding new vertices in the mesh.

[0322] The adaptor 2310 performs a fitting process, which may be applied by some application to fit the reconstructed mesh to a given scenario. For example, the vertices of the reconstructed mesh are transformed from the 3D model coordinate system to the 3D world coordinate system. The adaptor 2310 outputs a final reconstructed mesh frame (final 3D mesh frame 2311).

[0323] Next, the data structure of the bitstream will be described.

[0324] FIG. 50 is a diagram showing the data structure of a bitstream according to an embodiment.

[0325] In a bitstream, the coded data is encapsulated in a data unit structure.

[0326] A bitstream containing coded data of a three-dimensional mesh (coded mesh data) is composed of a series of units (DMC (Dynamic Mesh Coding) units). In the example shown in Figure 50, the bitstream (DMC bitstream) includes four data units (DMC units). Examples of the data units include a base mesh data unit, a displacement vector data unit, a metadata data unit, and a texture map data unit (texture map video NAL unit). Each data unit includes, for example, a header ("H" shown in Figure 50) and a payload. The header of the data unit stores a unit type (unit type information) indicating the type of data stored in the payload. The decoding device 200 analyzes the bitstream using the unit type (i.e., data type) to divide the data included in the bitstream into multiple data units.

[0327] The number of data units included in a bitstream is not particularly limited.

[0328] The base mesh data unit stores encoded data about the base mesh, specifically, the encoded base mesh (specifically, the encoded data about the base mesh) is stored in the base mesh data unit with a header.

[0329] The displacement vector data unit stores encoded data related to the displacement vector. Specifically, the encoded data related to the displacement vector is stored in a displacement data unit with a header.

[0330] The metadata data unit stores encoded metadata. Specifically, the encoded metadata, parameter set, or SEI (Supplemental Enhancement Information) is stored in the metadata data unit with a header.

[0331] A parameter set is, for example, a parameter set that is common to each frame (e.g., the same timing data, access units, and samples, etc.), such as a frame parameter set (FPS), or a parameter set that is common to each sequence, such as a sequence parameter set (SPS).

[0332] The texture map data unit stores, for example, an encoded texture map.

[0333] When a complete 3D mesh is divided into multiple submeshes and encoded, the decoding device 200 decodes the multiple submeshes (specifically, data related to the multiple submeshes) and, in a reconstruction process, merges the decoded submeshes to reconstruct a 3D mesh (in other words, a frame). When multiple submeshes are combined into one frame, all attribute information, including vertices (specifically, the 3D coordinates of 3D points), texture coordinate information, connection information, and normal information (normal vector information), is combined.

[0334] Here, in the conventional coding of three-dimensional meshes, it is assumed that the number of attribute information items constituting a submesh (e.g., the number of types of attribute information items) is fixed for each submesh. Furthermore, conventionally, there has been no function for varying the number of attribute information items for each submesh, and it has not been possible to code or decode three-dimensional meshes in which the number of attribute information items varies for each submesh.

[0335] Furthermore, in the past, the reconstruction process has not been defined for cases where the number of pieces of attribute information varies for each submesh, and therefore the decoding device 200 cannot perform the reconstruction process if the number of pieces of attribute information varies for each submesh. Furthermore, in the past, the process of combining attribute information such as texture coordinate information, normal information, color information, material ID information, transparency information, and reflectance information has not been fully defined.

[0336] Therefore, below we will explain a method of making the number of attribute information variable for each submesh and signaling the number of attribute information that differs for each submesh, and a method for detecting in the decoding device 200 that the number of attribute information differs for each submesh.

[0337] In addition, the decoding device 200 checks the number and type (category) of attribute information defined in the frame, and if attribute information of the type defined in the frame does not exist in the submesh, performs special processing such as padding the absent attribute information with a default value. In other words, the decoding device 200 compares the frame information defined as information commonly related to multiple submeshes with the submesh information, and performs specific processing if these pieces of information differ. Furthermore, for example, if the type of attribute information of the submesh does not exist in the types of attribute information defined in the frame, the decoding device 200 performs special processing such as padding the attribute information of the submesh with a default value.

[0338] This allows encoding, decoding, and reconstruction of 3D meshes even when the frames or sub-meshes do not have specific types of attribute information.

[0339] Next, an example of dividing a complete three-dimensional mesh into a plurality of sub-meshes and encoding the divided sub-meshes will be described. Specifically, an example of encoding attribute information of each of a plurality of sub-meshes divided from the three-dimensional mesh will be described.

[0340] Fig. 51 is a diagram showing attribute information of sub-meshes according to an embodiment. Specifically, Fig. 51 shows attribute information of sub-meshes 1 and 2 divided from the same three-dimensional mesh.

[0341] A 3D mesh is composed of, for example, position information (geometry) and attribute information. The position information includes, for example, information about polygonal faces. The face information includes, for example, position information indicating the positions of multiple 3D points and connection information connecting each 3D point (i.e., information indicating the connection relationships between each 3D point).

[0342] The attribute information includes attribute information corresponding to three-dimensional points and attribute information corresponding to faces, which are encoded in the base mesh encoding process, for example.

[0343] The attribute information corresponding to the three-dimensional points and the attribute information corresponding to the faces include, for example, texture coordinate information, connection information, normal vector information, face ID information, transparency information, and other ID information. The attribute information may include any attribute information defined by the user. This attribute information is encoded by, for example, the encoding device 100.

[0344] Texture coordinate information is, for example, information indicating coordinates in a texture map corresponding to a three-dimensional point and / or a face. Connection information is, for example, information indicating the three-dimensional points at both ends of an edge that constitutes a face. Normal information is, for example, information indicating the normal vector of a three-dimensional point and / or a face. Face ID information is, for example, information indicating an identifier uniquely assigned to each face included in a submesh. Transparency information is, for example, information indicating the transmittance of light corresponding to a three-dimensional point and / or a face.

[0345] In the example shown in FIG. 51, one frame (three-dimensional mesh) is divided into two sub-meshes (sub-mesh 1 and sub-mesh 2).

[0346] Submesh 1 has three pieces of attribute information (also called Attr). Submesh 2 has two pieces of attribute information. For example, Attr1 is color information, Attr2 is normal information, and Attr3 is face ID information. The color information is, for example, information indicating the color of a 3D point and / or face.

[0347] This example shows a case where submesh 1 has all attribute information, while submesh 2 does not have normal information. In other words, submesh 1 has all of the multiple attribute information defined in the frame, and submesh 2 has only some of the multiple attribute information defined in the frame. For example, if expressive rendering is required for the space of submesh 1 relative to the space of submesh 2, in order to improve the expressiveness of submesh 1, the normal information of submesh 1 may be transmitted to the decoding device 200, but the normal information of submesh 2 may not be transmitted. Furthermore, if the sensing method or data generation method differs between submeshes 1 and 2, a specific submesh may not have specific attribute information.

[0348] In this embodiment, a method for showing the structure of a three-dimensional mesh when the attribute information (more specifically, the type of attribute information) held by each sub-mesh is different will be described.

[0349] The number and types of attribute information shown in this embodiment are merely an example and are not particularly limited.

[0350] In the reconstruction process in the decoding device 200, even if the type of attribute information differs for each submesh, it is necessary to reconstruct a frame (i.e., a three-dimensional mesh) by combining attribute information of the same type for multiple submeshes. However, if the configuration of the attribute information for each submesh is not correctly notified to the decoding device 200, for example, different attribute information may be combined, resulting in an incorrect three-dimensional mesh being reconstructed.

[0351] FIG. 52 is a diagram for explaining the reconstruction process according to the embodiment.

[0352] In the reconstruction process, the decoding device 200 reconstructs a frame by combining attribute information of multiple submeshes. In the example shown in FIG. 52, the decoding device 200 combines submeshes 1 and 2 shown in (a) of FIG. 52. More specifically, the decoding device 200 combines attribute information included in submeshes 1 and 2. Note that the term "combining" here means, for example, generating data in which values ​​(more specifically, numerical values) indicated by multiple pieces of attribute information are arranged consecutively. For example, the decoding device 200 combines Attr1 included in submeshes 1 and 2 by linearly arranging the numerical value of Attr1 included in submeshes 1 and 2 (in other words, storing the numerical value in a one-dimensional array data structure). Furthermore, Attr1, Attr2, and Attr3 are different types of attribute information. Also, in this example, submesh 2 does not have Attr2.

[0353] As shown in (b) of Figure 52, in an ideal example of the reconstruction process, attribute information of the same type corresponding to each submesh is combined. That is, ideally, Attr1 of submesh 1 is combined with Attr1 of submesh 2, and Attr3 of submesh 1 is combined with Attr3 of submesh 2. As a result, Frm Attr1 is generated, which includes Attr1 of submesh 1 and Attr1 of submesh 2, Frm Attr2 is generated, which includes Attr2 of submesh 1, and Frm Attr3 is generated, which includes Attr3 of submesh 1 and Attr3 of submesh 2. Frm Attr1, Frm Attr2, and Frm Attr3 are attribute information corresponding to frames. In this way, when the reconstruction process is performed appropriately, attribute information of the same type is combined.

[0354] On the other hand, as shown in (c) of Figure 52, in an example of a failure in the reconstruction process, attribute information of different types is combined. In this example, Attr1 of submesh 1 is combined with Attr1 of submesh 2, and Attr2 of submesh 1 is combined with Attr3 of submesh 2. As a result, Frm Attr1 is generated, which includes Attr1 of submesh 1 and Attr1 of submesh 2, Frm Attr2 is generated, which includes Attr2 of submesh 1 and Attr3 of submesh 2, and Frm Attr3 is generated, which includes Attr3 of submesh 1.

[0355] In this way, in the failure example, Attr2 of submesh 1 and Attr3 of submesh 2 are erroneously combined (packed), causing an error in Frm Attr2 and Frm Attr3. Therefore, an incorrect frame (3D mesh) is rendered from Frm Attr1, Frm Attr2, and Frm Attr3 generated in this failure example.

[0356] 53 is a flow diagram showing a method for encoding sub-meshes according to an embodiment of the present invention, specifically, a method for encoding sub-meshes having different numbers of attribute information (more specifically, different types of attribute information).

[0357] First, the encoding device 100 encodes a first parameter indicating the number of submeshes in a frame (S301). Specifically, the encoding device 100 signals, in metadata, information indicating the number of submeshes included in the frame (i.e., the number of submeshes divided from a 3D mesh).

[0358] Next, the encoding device 100 encodes a second parameter indicating the maximum number of attribute information items defined for the frame, a first parameter set indicating the types of attribute information items defined for the frame, and a second parameter set indicating the number of dimensions of each attribute information item (S302). Specifically, the encoding device 100 signals, in the metadata, the maximum number of attribute information items held by each of the multiple submeshes included in the frame. For example, the encoding device 100 signals, in the bitstream, information indicating the maximum number, which is the number of attribute information items held by the submesh having the most attribute information items among the multiple submeshes. Furthermore, if the number or types of attribute information items differ for each submesh, the encoding device 100 may set the total number of all types of attribute information items for all submeshes included in the frame as the maximum number, and signal information indicating this maximum number in the metadata. The information indicating the maximum number is a parameter common to all submeshes (i.e., a parameter defined for the frame).

[0359] Next, the encoding device 100 encodes each submesh (S303). Specifically, the encoding device 100 encodes the number of submeshes indicated by the first parameter. More specifically, the encoding device 100 encodes each piece of attribute information included in each submesh. Each submesh includes, for example, multiple pieces of attribute information including coordinate values ​​and connection information indicating the connection relationship between the coordinates of these pieces of attribute information. The encoding device 100 encodes, for example, the multiple pieces of attribute information and the connection information.

[0360] The number of pieces of attribute information may differ for each submesh.

[0361] Next, the encoding device 100 encodes a third parameter set indicating the number of pieces of attribute information for each submesh and a fourth parameter set indicating the type of the attribute information for each submesh (S304). That is, the encoding device 100 signals the number and type of attribute information for each submesh (more specifically, the number of pieces of attribute information for each type) in the metadata.

[0362] Parameters common to each submesh are stored in a parameter set common to each sequence, such as an ASPS (Atlas Frame Parameter Set), and / or a parameter set common to each frame, such as an AFPS (Atlas Frame Parameter Set).

[0363] Furthermore, the parameters for each submesh are stored in the header of each submesh or in metadata (such as a patch data unit header or a base mesh header).

[0364] Figures 54 to 56 are diagrams showing syntax for signaling metadata according to an embodiment. Specifically, Figure 54 shows syntax for storing information indicating the number of submeshes in a sequence parameter set, Figure 55 shows syntax for storing information indicating a submesh ID uniquely determined for each submesh in a submesh header, and Figure 56 shows syntax for storing information indicating the number of submeshes in a submesh parameter set.

[0365] submeshCount indicates the number of submeshes contained in the frame (more specifically, the current frame).

[0366] bmesh_data_attribute_count indicates the maximum number of attribute information items in a frame (more specifically, the current frame). Note that each submesh cannot have more types of attribute information items than the value indicated by bmesh_data_attribute_count.

[0367] AttributeType indicates the type of attribute information, such as "color" or "normal."

[0368] AttributeDimension indicates the number of dimensions of the attribute information. For example, color information may include information indicating each of the three elements r, g, and b. For example, if the color information includes information on these three elements, the number of dimensions of the color information is 3. That is, in this case, attributeDimension for the color information = 3.

[0369] Fig. 55 shows an example of metadata added to the header (SubmeshHeader) of data for each submesh. Fig. 56 shows an example of metadata (specifically, Submesh Parameter Set) including data for each submesh in all submeshes.

[0370] SubmeshID indicates the identifier of a submesh included in a three-dimensional mesh.

[0371] attrCount_Submesh indicates the number of attribute information items included in the submesh.

[0372] AttributeType_Submesh indicates the type of attribute information included in the submesh.

[0373] AttributeDimension_Submesh indicates the number of dimensions of the attribute information included in the submesh.

[0374] Note that parameters common to each submesh may be signaled to the SPS or FPS. Although the above describes an example of a method in which parameters specific to each submesh are signaled to parameters for each submesh, a syntax or semantics may be adopted in which parameters common to each submesh are signaled to the SPS or FPS and the parameters for each submesh are used to overwrite the parameters common to each submesh.

[0375] Next, specific values ​​to be signaled will be described.

[0376] 57 is a diagram showing attribute information of sub-meshes according to an embodiment. First, a signaling example 1 will be described in which information relating to sub-meshes having attribute information as shown in FIG.

[0377] In signaling example 1, submesh 1 has color information ("Color" shown in FIG. 57), normal information ("Normal" shown in FIG. 57), and face ID information ("FaceID" shown in FIG. 57) as attribute information. On the other hand, submesh 2 has color information ("Color" shown in FIG. 57) and face ID information ("FaceID" shown in FIG. 57) as attribute information. In other words, submesh 2 does not have normal information.

[0378] Under these conditions, the number of attribute information items (the number of types of attribute information items) defined for a frame is 3. The number of attribute information items for submesh 1 is 3. The number of attribute information items for submesh 2 is 2.

[0379] Among these pieces of information, information relating to frames is signaled to the SPS, for example, as follows.

[0380] submeshCount=2 bmesh_data_attribute_count=3 attributeType[0]='Color' attributeType[1]='Normal' attributeType[2]='FaceID'

[0381] Of these pieces of information, the information relating to submesh 1 is signaled, for example, in the header in which the information on submesh 1 is stored, as follows:

[0382] attrCount_Submesh=3 attributeType_Submesh[0]='Color' attributeType_Submesh[1]='Normal' attributeType_Submesh[2]='FaceID'

[0383] Of these pieces of information, the information relating to submesh 2 is signaled, for example, in the header in which the information on submesh 2 is stored, as follows:

[0384] attrCount_Submesh=2 attributeType_Submesh[0]='Color' attributeType_Submesh[1]='FaceID'

[0385] 58 is a diagram showing attribute information of sub-meshes according to an embodiment. Next, a signaling example 2 will be described in which information relating to sub-meshes having attribute information as shown in FIG.

[0386] In signaling example 2, submesh 1 has color information and normal information as attribute information. On the other hand, submesh 2 has color information and face ID information as attribute information. In other words, submesh 1 does not have face ID information, and submesh 2 does not have normal information.

[0387] Under these conditions, the number of pieces of attribute information defined for a frame is 3. The number of pieces of attribute information for submesh 1 is 2. The number of pieces of attribute information for submesh 2 is 2.

[0388] Among these pieces of information, information relating to frames is signaled to the SPS, for example, as follows.

[0389] submeshCount=2 bmesh_data_attribute_count=3 attributeType[0]='Color' attributeType[1]='Normal' attributeType[2]='FaceID'

[0390] Of these pieces of information, the information relating to submesh 1 is signaled, for example, in the header in which the information on submesh 1 is stored, as follows:

[0391] attrCount_Submesh=2; attributeType_Submesh[0]='Color' attributeType_Submesh[1]='Normal'

[0392] Of these pieces of information, the information relating to submesh 2 is signaled, for example, in the header in which the information on submesh 2 is stored, as follows:

[0393] attrCount_Submesh=2; attributeType_Submesh[0]='Color' attributeType_Submesh[1]='FaceID'

[0394] FIG. 59 is a diagram for explaining the processing of the reconstructor 2402 according to the embodiment.

[0395] For example, the decoding device 200 includes a sub-mesh decoder 2401 and a reconstructor 2402 .

[0396] The submesh decoder 2401, for example, decodes encoded data included in the bitstream to obtain attribute information of submesh 1 and attribute information of submesh 2. In the example shown in FIG. 59, the submesh decoder 2401 obtains Attr1, Attr2, and Attr3 as attribute information of submesh 1, and Attr1 and Attr3 as attribute information of submesh 2. The submesh decoder 2401 outputs the attribute information of submesh 1 (specifically, Attr1, Attr2, and Attr3) and the attribute information of submesh 2 (specifically, Attr1 and Attr3) to the reconstructor 2402. More specifically, the submesh decoder 2401 outputs parameters such as information indicating the maximum number of attribute information of submeshes in a frame, information indicating the type of attribute information included in the frame, information indicating the number of attribute information for each submesh, and information indicating the type of attribute information for each submesh to the reconstructor 2402. This information may be included in the bitstream, or may be calculated from the attribute information of submesh 1 and the attribute information of submesh 2.

[0397] In this way, the submesh decoder 2401 outputs a plurality of decoded submeshes.

[0398] The reconstructor 2402 combines attribute information using various information related to the input submeshes. In this example, the reconstructor 2402 generates Frm Attr1 by combining Attr1 of submesh 1 with Attr1 of submesh 2, sets Attr2 of submesh 1 as Frm Attr2, and generates Frm Attr3 by combining Attr3 of submesh 1 with Attr3 of submesh 2. This generates frame attribute information (i.e., frame-level attribute information). The reconstructor 2402 outputs Frm Attr1, Frm Attr2, and Frm Attr3. In this way, the frame (mesh frame) reconstructed by the reconstructor 2402, i.e., the frame attribute information, is output from the reconstructor 2402.

[0399] The reconstructor 2402 receives, for example, submeshCount, submeshes, and vps_ext_bmesh_data_attribute_count as input.

[0400] The submeshCount is information indicating the number of submeshes in the current frame.

[0401] The submeshes is information about the submeshes decoded in the current frame. Specifically, the submesh includes information indicating the positions of the vertices (3D points) of the submesh, information indicating the faces (connectivity) of the submesh, attrValues, attrFaces, and the like.

[0402] vps_ext_bmesh_data_attribute_count is information indicating the number of attribute information items in the current atlas (frame).

[0403] Moreover, the reconstructor 2402 outputs a full mesh.

[0404] The fullMesh is the fully reconstructed current frame. Specifically, the fullMesh contains information indicating the positions of the vertices of the frame (the reconstructed 3D mesh), information indicating the faces (connectivity) of the frame, attrValues, attrFaces, etc.

[0405] For example, vps_ext_bmesh_data_attribute_count indicates the maximum number of attribute information in the current frame (i.e., the number of attribute information defined in the frame). Each submesh cannot have more attribute information than this value.

[0406] FIG. 60 is a flow diagram showing a submesh decoding method according to an embodiment.

[0407] First, the decoding device 200 decodes a first parameter indicating the number of submeshes in a frame (S311). Specifically, the decoding device 200 obtains the first parameter by decoding the coded first parameter included in the bitstream.

[0408] Next, the decoding device 200 decodes a second parameter indicating the maximum number of attribute information items defined for the frame, a first parameter set indicating the type of attribute information items defined for the frame, and a second parameter set indicating the number of dimensions of each attribute information item (S312). Specifically, the decoding device 200 obtains the second parameter, the first parameter set, and the second parameter set by decoding the coded second parameter, the coded first parameter set, and the coded second parameter set included in the bitstream.

[0409] Next, the decoding device 200 decodes each submesh (S313). Specifically, the decoding device 200 decodes each piece of coded attribute information for each submesh included in the bitstream. Each submesh includes, for example, multiple pieces of attribute information including coordinate values ​​and connection information indicating the connection relationship between the coordinates of these pieces of attribute information. The decoding device 200, for example, decodes the coded pieces of attribute information and the coded connection information.

[0410] Next, the decoding device 200 decodes a third parameter set indicating the number of pieces of attribute information for each submesh and a fourth parameter set indicating the type of attribute information for each submesh (S314). Specifically, the decoding device 200 obtains the third parameter set and the fourth parameter set by decoding the coded third parameter set and the coded fourth parameter set included in the bitstream.

[0411] 61 is a flowchart showing reconstruction processing according to an embodiment. The processing procedure shown in FIG. 61 is executed by the decoding device 200, for example, following the processing procedure shown in FIG.

[0412] First, the decoding device 200 compares the second parameter with the third and fourth parameter sets (S321). That is, the decoding device 200 compares the second parameter (the number of types of attribute information) at the frame level with the third and fourth parameter sets (the number of types of attribute information) at the submesh level.

[0413] Next, the decoding device 200 determines whether the number and types of attribute information of the sub-mesh are different from the number and types of attribute information of the frame (S322). That is, the decoding device 200 compares the second parameter with the third parameter set and the fourth parameter set to determine for each sub-mesh whether the number of attribute information of each type defined for the frame is different from the number of attribute information of each type defined for each sub-mesh.

[0414] When the decoding device 200 determines that the number and types of attribute information of the submesh are different from the number and types of attribute information of the frame (Yes in S323), it identifies attribute information that exists in the attribute information defined in the frame but does not exist in the attribute information defined in the submesh (S324). That is, the decoding device 200 identifies attribute information that exists at the frame level but does not exist at the submesh level. For example, the decoding device 200 identifies attribute information that exists in the type of attribute information defined in the frame but does not exist in the type of attribute information defined in the submesh. When multiple numerical values ​​are defined for the same type, it may be determined whether the number of attribute information differs between the frame and the submesh for each type.

[0415] The decoding device 200 generates the value of the identified attribute information, that is, the attribute information that does not exist in the submesh, by padding or the like, and executes the reconstruction process using the generated value (S325).

[0416] For example, when submesh 1 has Attr1, Attr2, and Attr3, and submesh 2 has Attr1 and Attr3 but does not have Attr2, the decoding device 200 generates a value of attribute information corresponding to Attr2 of submesh 2. This makes the number of attribute information pieces in submesh 1 and submesh 2 appear to be the same, and attribute information pieces of the same type are appropriately combined.

[0417] The value of the attribute information generated by padding or the like may be any value and is not particularly limited.

[0418] On the other hand, if the decoding device 200 determines that the number and type of attribute information of the submesh are not different from the number and type of attribute information of the frame (No in S323), it performs reconstruction processing without performing processing such as padding (S326).

[0419] 62 is a diagram showing a process of comparing attribute information at a submesh level with attribute information at a frame level according to an embodiment. Specifically, FIG. 62 shows a program for executing a method for determining (detecting) whether the number and types of attribute information of a submesh are different from the number and types of attribute information defined for a frame.

[0420] The decoding device 200 (e.g., submesh decoder 2401) detects whether the attribute information of the submesh contains the attribute information defined in the frame by comparing, for each submesh, the number and type of attribute information defined in the frame signaled in the bitstream with the number and type of attribute information of the submesh, and outputs the detection result (found).

[0421] It should be noted that the number of pieces of attribute information for each submesh may be set to be equal to or less than the number of pieces of attribute information set for a frame as a condition for conformance of the bitstream.

[0422] Alternatively, it may be determined as a condition for bitstream conformance that the type of attribute information for each submesh is included in the type of attribute information indicated in a parameter set defined for a frame.

[0423] The decoding device 200 may determine that there is a conformance error if the conformance conditions are not met (if the number of attribute information of the submesh is greater than the number of attribute information specified in the frame, or if the type of attribute information of the submesh is not included in the type of attribute information indicated in the frame parameter set).

[0424] The conformance conditions may be arbitrarily defined, and are not particularly limited. For example, since it is not reasonable for a sub-mesh to have more attribute information than the number of attribute information defined in the frame, if the number of attribute information in the sub-mesh is greater than the number of attribute information defined in the frame, it is determined to be a conformance error.

[0425] Fig. 63 is a diagram showing an example of a reconstruction process according to an embodiment. Specifically, Fig. 63 shows an example of a program for executing a process in which the decoding device 200 reconstructs multiple decoded data into a one-dimensional array (a one-dimensional array data structure) and stores the decoded data of the one-dimensional array in the reconstructed data. The number of multiple decoded data is, for example, submeshCount. That is, submeshCount is information indicating the number of submeshes. Furthermore, the one-dimensional array of decoded data is, for example, refinedSubmeshFrames. That is, refinedSubmeshFrames indicates attribute information of each submesh. More specifically, refinedSubmeshFrames indicates the numerical value of the attribute information of each submesh expressed in a one-dimensional array. Furthermore, the reconstructed data is, for example, recMeshFrame. That is, the attribute information of the submeshes is stored in a one-dimensional array (refinedSubmeshFrames), and the frame finally reconstructed by combining the values ​​indicated by the attribute information of each submesh is recMeshFrame.

[0426] In the example shown in Figure 63, each sub-mesh has the same number of attribute information. attrCount is the number of attribute information defined in the frame, and is the number (maximum number) of attribute information common to each sub-mesh.

[0427] The decoding device 200 executes the program shown in FIG. 63, for example, to combine the numerical values ​​of the attribute information of each sub-mesh for each type of attribute information, thereby generating attribute information of the frame.

[0428] Fig. 64 is a diagram showing another example of the reconstruction process according to the embodiment. Specifically, Fig. 64 shows an example of a program in the case where the number of faces of the sub-meshes corresponding to the attribute information differs for each piece of attribute information.

[0429] By executing the program shown in Figure 64, even if the number of faces of the sub-meshes corresponding to the attribute information differs for each attribute information, the numerical values ​​of the attribute information of each sub-mesh are combined for each type of attribute information to generate attribute information for the frame.

[0430] As described above, for example, the submesh decoder 2401 decodes the encoded data to input the number of submeshes (submeshCount) and one-dimensional array data (refinedSubmeshFrames) storing information corresponding to each submesh to the reconstructor 2402. The reconstructor 2402 sets the number of types of attribute information (AttrCount) of the initial reconstructed mesh frame (recMeshFrame) based on, for example, the number of types of attribute information (AttrCount) of the first element of refinedSubmeshFrames. Furthermore, for data relating to each submesh (i = 2 to submeshCount) stored in refinedSubmeshFrames, the reconstructor 2402 concatenates (merges) values ​​corresponding to each attribute information item (AttrCount) for the number of types of attribute information of the submesh with attribute information items (AttrCount) for the number of types of attribute information of the corresponding recMeshFrame. Furthermore, for example, the reconstructor 2402 generates attribute information for the frame by cumulatively updating the number of concatenated attribute information items (count value). Note that the attrIndex shown in FIG. 63 or 64 may be data determined for each attribute information item or for each type of attribute information item.

[0431] The recombination process may be performed by the encoding device 100. The encoding device 100 may generate a plurality of sub-meshes by, for example, dividing a three-dimensional mesh, and may generate a recMeshFrame by performing the process shown in Fig. 63 or 64 on the generated sub-meshes. The encoding device 100 may determine encoding parameters for data related to the plurality of sub-meshes (e.g., attribute information of each sub-mesh) based on, for example, a comparison result between the original three-dimensional mesh and the generated recMeshFrame.

[0432] Fig. 65 is a diagram showing another example of the reconstruction processing according to the embodiment. Specifically, Fig. 65 shows an example of a program for the reconstruction processing when specific attribute information is not included in the attribute information of any of the sub-meshes among the sub-meshes.

[0433] The fullMesh.attrCount is information indicating the number of attribute information items (attribute information items for the entire frame) included in the SPS or FPS and defined for the frame. In the loop for the number of attribute information items for the entire frame in the program of this example, if attribute information exists in the submesh, the decoding device 200 performs a reconstruction process using the values ​​of the decoded attribute information items. If attribute information does not exist in the submesh, the decoding device 200 pads the values ​​of the attribute information items for the three-dimensional points or faces of the submesh and performs a reconstruction process using the padded values.

[0434] FIG. 66 is a diagram showing details of the reconstruction process according to the embodiment.

[0435] For example, if there is no attribute information of a submesh, the decoding device 200 pads the submesh using a default value.

[0436] The default value may be a predetermined value, or may be included in a parameter set and transmitted from the encoding device 100 to the decoding device 200 .

[0437] Alternatively, the decoding device 200 may pad using values ​​of the geometry or attribute information of the three-dimensional points or faces of other sub-meshes.

[0438] The default value may be, for example, 0, or may be pow(2, attr_bitDepth)-1 calculated based on the bit information (attr_Depth) of the attribute information.

[0439] For example, when attribute information of a plurality of three-dimensional points or surfaces has a constant value, the encoding device 100 may not encode the attribute information and may notify (transmit) a default value to the decoding device 200. The decoding device 200 may perform the reconstruction process using the default value.

[0440] FIG. 67 is a diagram showing syntax for setting default values ​​according to an embodiment.

[0441] defaultValue_flag is a flag (flag information) indicating whether or not a default value exists in the bitstream.

[0442] defaultValue is information indicating the default value of the attribute information. An individual value may be set for each dimension of the attribute information, or a common value may be set for each dimension of the attribute information.

[0443] FIG. 68 is a flowchart showing a process for setting default values ​​according to an embodiment.

[0444] First, the decoding device 200 compares, for each submesh, the number and type of attribute information defined in the frame with the number and type of attribute information of the submesh (S331).

[0445] Next, the decoding device 200 determines whether the number and types of attribute information of the submesh are sufficient (S331). Specifically, the decoding device 200 compares the number and types of attribute information defined in the frame with the number and types of attribute information of the submesh to determine whether the number and types of attribute information of the submesh are the same as the number and types of attribute information defined in the frame. If the decoding device 200 determines that the number and types of attribute information defined in the frame are the same as the number and types of attribute information of the submesh, it determines that the number and types of attribute information of the submesh are sufficient (Yes in step S332). On the other hand, if the decoding device 200 determines that the number and types of attribute information defined in the frame are different from the number and types of attribute information of the submesh, it determines that the number and types of attribute information of the submesh are insufficient (No in step S332).

[0446] When the decoding device 200 determines that the number and types of attribute information of the sub-mesh are insufficient (No in S332), the decoding device 200 pads the values ​​of the attribute information that are insufficient in the sub-mesh with default values ​​(S333). Specifically, the decoding device 200 assigns to the sub-mesh attribute information that is of a type that the sub-mesh does not have, among the multiple types of attribute information defined in the frame, and indicates a default value.

[0447] For example, if a default value is signaled in the bitstream, the decoding device 200 pads using the signaled default value, and if a default value is not signaled in the bitstream, the decoding device 200 pads using a predetermined value as the default value.

[0448] The predetermined value may be arbitrarily determined and is not particularly limited. Information indicating the predetermined value may be pre-stored in the decoding device 200, or information indicating a method for calculating the predetermined value may be pre-stored in the decoding device 200.

[0449] After step S333, or when it is determined that the number and types of attribute information of the sub-mesh are sufficient (Yes in S332), the decoding device 200 performs a reconstruction process (S334). Specifically, the decoding device 200 combines values ​​indicated by multiple attribute information for each type of attribute information (in other words, stores them in a one-dimensional array data structure).

[0450] In addition, when encoding multiple pieces of attribute information of the same type, the encoding device 100 may signal an identifier (attrIDforSameAttributeType) to the bitstream indicating that the attribute information is of the same type but different, in order to identify that the attribute information is of the same type but different.

[0451] Note that the method of signaling information about submeshes, such as attribute information, is not limited to the above-mentioned Signaling Example 1 and Signaling Example 2. Signaling Example 3, which is a signaling method different from Signaling Example 1, is described below in which information about submeshes having attribute information as shown in Figure 57 is signaled.

[0452] In signaling example 3, as shown in Figure 57, submesh 1 has color information, normal information, and face ID information as attribute information. On the other hand, submesh 2 has color information and face ID information as attribute information. In other words, submesh 2 does not have normal information.

[0453] Under these conditions, the number of pieces of attribute information defined for the frame is 3. Also, the number of pieces of attribute information for submesh 1 is 3.

[0454] Here, the encoding device 100 may set the number of attribute information in submesh 2 to 3, and may newly define 'None' as the type of attributeType_Submesh for the second attribute information of submesh 2 (attribute information corresponding to the normal information of submesh 1), thereby explicitly indicating in the bitstream that the attribute information is not included in the bitstream.

[0455] Methods for explicitly indicating in the bitstream that attribute information is not included include indicating such information in attributeType, defining a data unit that does not include attribute information, and defining a data unit type and indicating this in the bitstream.

[0456] In this example, the information relating to the frame among these pieces of information is signaled to the SPS as follows, for example.

[0457] submeshCount=2 bmesh_data_attribute_count=3

[0458] Of these pieces of information, the information relating to submesh 1 is signaled, for example, in the header in which the information on submesh 1 is stored, as follows:

[0459] attrCount_Submesh=3 attributeType_Submesh[0]='Color' attributeType_Submesh[1]='Normal' attributeType_Submesh[2]='FaceID'

[0460] Of these pieces of information, the information relating to submesh 2 is signaled, for example, in the header in which the information on submesh 2 is stored, as follows:

[0461] attrCount_Submesh=3 attributeType_Submesh[0]='Color' attributeType_Submesh[1]='None' attributeType_Submesh[2]='FaceID'

[0462] When signaling example 3 is used, it is explicitly stated that a certain type of attribute information is not included, so the decoding device 200 can identify the type of attribute information that is not included in the bitstream without comparing the type of attribute information indicated in the SPS with the type of attribute information indicated for each submesh.

[0463] FIG. 69 is a diagram for explaining parameters for each submesh according to the embodiment.

[0464] The 3D mesh coding standard is a standard for dividing a frame into sub-meshes and encoding each sub-mesh. Furthermore, the 3D mesh coding standard is a standard in which various different standards, such as MEB (MPEG Edge Breaker) or Draco, may be used for encoding each sub-mesh. In other words, the sub-mesh is encoded using an independent codec layer.

[0465] In this case, various parameters can be set for each submesh in the codec layer. For example, if Param1 is information indicating the number of pieces of attribute information, the number of pieces of attribute information can be set for each submesh. However, if all parameters can be set for each submesh in the coding layer of the three-dimensional mesh, the standard becomes redundant, and the design of the decoding device 200 becomes complicated.

[0466] Therefore, metadata that can identify parameters that may be variable for each submesh and parameters that are not variable but are restricted to be common to each submesh (Codec Dependent Param Restriction) may be indicated in a frame-level parameter set (e.g., FPS).

[0467] FIG. 70 is a diagram showing syntax for signaling a flag indicating whether a parameter in an embodiment is common to each submesh.

[0468] For example, when Param1_frame_common_flag is set to true, it is assumed that Param1 of submesh 1 and submesh 2 is common, and the decoding device 200 determines that an error has occurred if different values ​​are set for Param1 of submesh 1 and Param1 of submesh 2.

[0469] Also, for example, if Param3_frame_common_flag is set to true and a parameter (frame_Param3) common to each submesh is indicated in the frame-level parameter set, the decoding device 200 determines that there is an error if Param3 indicated in submesh 1 and submesh 2 is different from the frame-level parameter.

[0470] Furthermore, for example, when it is determined that the number of pieces of attribute information for each submesh is to be the same, NumAttr_frame_common_flag=True is set, and the number of pieces of attribute information indicated for each submesh is restricted to be the same.

[0471] Among the parameters described in the codec layer, parameters that are common to each sub-mesh and parameters that may have different values ​​in each sub-mesh may be defined in advance by a standard.

[0472] For example, a plurality of profiles may be prepared, and each profile may be configured to switch between restricting commonality among sub-meshes. For example, a profile may be divided into a profile in which Param3_frame_common_flag=true and a profile in which Param3_frame_common_flag=false.

[0473] By using the method described above, even if the number of attribute information for each submesh differs during the reconstruction process, it is possible to combine (integrate) each submesh (more specifically, the attribute information of the submeshes) and reconstruct the frame.

[0474] This makes it possible to encode, decode, and reconstruct data of three-dimensional meshes in which the number of attribute information pieces varies for each sub-mesh, and by reconstructing the complete three-dimensional mesh, it becomes possible to render the three-dimensional mesh.

[0475] <Summary of the present disclosure> [Encoding device] The encoding device 100 is an encoding device that inputs two or more submeshes with different numbers of attribute information and encodes them, and indicates the number or type of attribute information that constitutes a frame in metadata common to the two or more submeshes, and indicates the number or type of attribute information for each submesh in individual metadata sets for the two or more submeshes.

[0476] For example, the number of pieces of attribute information for each submesh is smaller than the number of pieces of attribute information that make up a frame.

[0477] Also, for example, if the number of pieces of attribute information for each submesh is smaller than the number of pieces of attribute information that constitute a frame, the encoding device 100 signals the default values ​​of the attribute information in the bitstream.

[0478] [Decoding device] The decoding device 200 is a decoding device that inputs and decodes two or more submeshes each having a different number of pieces of attribute information. For each submesh, the decoding device decodes the respective number of pieces of attribute information. If the attribute information that constitutes a frame is not included in a particular submesh, the decoding device generates the attribute information of the submesh by padding a predetermined value into the value of the submesh, and reconstructs the frame.

[0479] For example, the predetermined value is 0 or a value calculated from the number of bits of the attribute information.

[0480] Also, for example, the predetermined value is data included in a parameter set of the sequence, and the decoding device 200 decodes the parameter set and pads it with the predetermined value indicated by the data included in the decoded parameter set.

[0481] Also, for example, the decoding device 200 obtains information indicating the number or type of attribute information that constitutes the frame from metadata common to two or more submeshes, obtains information indicating the number or type of attribute information for each submesh from metadata sets individual to each submesh, and uses the obtained information to identify attribute information that constitutes the frame that is not included in the submeshes.

[0482] Also, for example, the attribute information is either attribute information for each three-dimensional point or attribute information for each surface, and the specified value signaled in the bitstream indicates either attribute information for each three-dimensional point or attribute information for each surface depending on the type of attribute information.

[0483] [Encoding method] In the encoding method, the bitstream includes (1) for each frame or sequence including multiple submeshes, first unit numerical information, which is information on the number of attribute information pieces of the submesh that includes the most attribute information, and (2) second unit numerical information, which is information on the number of attribute information pieces included in each submesh.

[0484] For example, the encoding method determines whether there is a submesh among a plurality of submeshes in which the number information of the second unit is less than the number information of the first unit, and includes information on the determination result in the bitstream.

[0485] Also, for example, in the encoding method, for submeshes among multiple submeshes that do not satisfy the numerical information of the first unit, attribute information is complemented to a number that satisfies the numerical information of the first unit, and the complemented attribute information is included in the bitstream.

[0486] Furthermore, for example, the attribute information used for complementation is a predetermined initial value.

[0487] Furthermore, for example, the attribute information used for complementation is calculated using attribute information or position information of other sub-meshes or three-dimensional points.

[0488] Also, for example, in the encoding method, the number information of the first units is included in the bitstream as metadata (control information).

[0489] Furthermore, for example, the number information of the first unit is stored in the ASPS, which is a parameter set common to sequences, or the AFPS, which is a parameter set common to frames.

[0490] Also, for example, the number information of the second units is stored in the header or metadata of the sub-mesh unit (patch data unit header or base mesh header, etc.).

[0491] [Decoding method] The decoding method is a decoding method that decodes a bit stream (encoded data) to decode a mesh including a first submesh and a second submesh, wherein the mesh includes multiple attribute information, and the bit stream includes control information including the number of attribute information included in the mesh.

[0492] [Encoding method] In the encoding method, the bit stream includes (1) first information, which is information indicating the number of attribute information pieces included in the first submesh or the type of attribute information, and (2) second information, which is information indicating the number of attribute information pieces included in the second submesh or the type of attribute information.

[0493] For example, in the encoding method, the first information and the second information are compared, it is determined whether the first information and the second information are different, and information on the determination result is included in the bitstream.

[0494] Also, for example, in the encoding method, if the first information and the second information are different, the attribute information of the first submesh or the second submesh is complemented so that the first information and the second information are the same.

[0495] Also, for example, each of the first information and the second information is stored in a header or metadata (patch data unit header or base mesh header, etc.) for each submesh of the first submesh and the second submesh.

[0496] Furthermore, for example, the attribute information used for complementation is a predetermined initial value.

[0497] Furthermore, for example, the attribute information used for complementation is calculated using attribute information or position information of other sub-meshes or three-dimensional points.

[0498] [Decoding method] The decoding method is a decoding method that decodes a bit stream (encoded data) to decode a mesh including a first submesh and a second submesh, wherein the mesh includes attribute information and the bit stream includes control information including information indicating the number or type of attribute information included in the mesh.

[0499] <Codec Parameter Constraints> As described above, the three-dimensional mesh coding standard (mesh coding standard) is a standard for dividing a frame (three-dimensional mesh) into a plurality of sub-meshes and encoding each sub-mesh.

[0500] Here, in the mesh coding standard, in coding for each submesh, various different coding standards, such as MEB (MPEG Edge Breaker) or Draco, may be used for intra-coded submeshes (static submeshes). In other words, intra-coding of a submesh does not depend on the coding method of the base mesh, meaning that any coding method can be selected for intra-coding. Since such intra-coded submeshes (specifically, the geometry information and attribute information of the submesh) are processed independently of the coding method of the base mesh, the coding method for intra-coding is a coding method that has no selection or limitation regarding the type of codec, that is, a codec-independent coding method. It is also sometimes called a codec-agnostic coding method.

[0501] Here, when the hierarchical structure of a mesh encoding method (or encoding standard) is expressed as a protocol stack, the logical unit of the structural layer is called a layer. For example, there are layers that handle structural information of a base mesh, layers that handle geometry and attribute information of sub-meshes, etc. Furthermore, a codec-independent layer (Codec Independent Layer) or a codec-agnostic layer (Codec Agonistic Layer) is a hierarchical encoding method that does not depend on a specific codec, and data encoded using an arbitrary encoding method, such as an encoding standard defined by an external organization, is stored in a data structure defined by each encoding standard. Furthermore, so that data encoded using an arbitrary encoding method can be stored, a codec type indicating which encoding method was used is indicated in a higher-level encoding method layer. For example, if the arbitrary encoding standard used in the codec-independent layer has specifications that allow various parameters to be set for each sub-mesh, various parameters can be set for each sub-mesh.

[0502] On the other hand, the inter-coding performed on the sub-meshes is the processing of the base mesh layer defined by the base mesh standard. In other words, the processing of the base mesh layer uses a standard dependent on the base mesh standard, and the base mesh layer is called a layer that is not codec independent. In the inter-coding, parameters defined by the base mesh standard are used for each sub-mesh (geometry information and attribute information of the sub-mesh).

[0503] Note that a submesh that is intra-coded and a submesh that has been intra-coded are also simply referred to as an intra submesh, and a submesh that is inter-coded and a submesh that has been inter-coded are also referred to as an inter submesh.

[0504] 71 is a block diagram showing an example of the configuration of a base mesh frame decoder 2501 according to an embodiment. The base mesh frame decoder 2501 includes a metadata acquirer 2502, an intra sub-mesh decoder 2503, a verifier 2504, and an inter sub-mesh decoder 2505.

[0505] In the inter-submesh decoder 2505, the inter-submesh is decoded using the decoded intra-submesh and the decoded difference information.

[0506] Here, because the intra-submesh decoder 2503 uses the coding standard and parameters of codec-independent layer to decode coded data, it is not guaranteed that the output of the intra-submesh decoder 2503 conforms to the output defined in the base mesh layer.In other words, it is not guaranteed that the data structure (data configuration) of the data output from the intra-submesh decoder 2503 is the same as the data structure defined in the base mesh layer.In other words, it is not guaranteed that the standard of the data output by the codec-independent intra-submesh decoder 2503 is the same as the standard of the base mesh.

[0507] For example, the inter-submesh decoder 2505 has the problem that if the data structure of the data output by the intra-submesh decoder 2503 does not conform to the data structure of the parameters specified in the base mesh layer, it cannot process correctly and cannot decode (inter-decode).

[0508] Therefore, the checker 2504 included in the decoding device 200 checks whether the data processed in the codec-independent layer by the intra-submesh decoder 2503 or the like complies with the base mesh specifications.

[0509] Furthermore, by checking the data structure of each submesh before integrating the plurality of submeshes, it becomes possible to prevent inconsistencies when integrating the plurality of submeshes.

[0510] The decoding device 200 may include one or more sets of the intra submesh decoder 2503, the verifier 2504, and the inter submesh decoder 2505. When the decoding device 200 includes more than one set, the coded submesh may be input to each set. For example, coded submeshes included in the bitstream that are different from each other are input to each set.

[0511] FIG. 72 illustrates syntax for signaling a base mesh parameter set (i.e., parameters of the base mesh) according to an embodiment.

[0512] The metadata of the base mesh may be common to the sequence or the frame. Here, the parameters common to the base mesh are shown. In other words, the parameters common to the frame are shown. For example, if the base mesh is composed of multiple sub-meshes, the metadata of the base mesh are parameters common to the multiple sub-meshes.

[0513] The base mesh metadata may be defined as parameters common to the sub-mesh data output from the base mesh frame decoder 2501, parameters common to the sub-mesh data provided (output) to the reconstructor 2506, and / or parameters common to the sub-mesh data input to the reconstructor 2506. That is, it is defined that the data output from the base mesh frame decoder 2501 and / or the data handled by the base mesh frame decoder 2501 must match these parameters.

[0514] Here, the data structure of the data output from the intra-submesh decoder 2503 may not conform to the base mesh standard. For example, even if the base mesh standard specifies common parameters for each submesh, if an arbitrary encoding standard used in a codec-independent layer allows various parameters to be set for each submesh, different parameters may be used for each submesh. In this case, the parameters of the data output from the intra-submesh decoder 2503 do not match the parameters specified in the base mesh layer. Therefore, before the reconstructor 2506 performs processing, or while the reconstructor 2506 is performing processing, the decoding device 200 detects whether the data output from the intra-submesh decoder 2503 matches the metadata of the base mesh. In other words, the decoding device 200 checks (determines) whether the data output from the codec-independent intra-submesh decoder matches the metadata of the base mesh.

[0515] Here, since the intra-submesh has a standard independent from that of the base mesh layer, the base mesh cannot be involved in the processing in the intra-submesh decoder 2503. Therefore, the verifier 2504 performs processing by referring to the data output from the intra-submesh decoder 2503. The base mesh frame decoder 2501 is characterized by verifying differences in data structure despite processing restrictions.

[0516] The configuration information regarding the data structure output from the intra sub-mesh decoder will be described.

[0517] In a first method for determining whether the data output from the intra-submesh decoder 2503 conforms to the standard, information about the data structure is output from the intra-submesh decoder 2503. A verifier 2504 compares the information about the data structure output from the intra-submesh decoder 2503 with the metadata of the base mesh.

[0518] FIG. 73 is a diagram illustrating syntax for signaling a sub-mesh structure set according to an embodiment.

[0519] The information about the data structure output from the intra sub-mesh decoder 2503 has a syntax such as that shown in FIG. 73, for example, and is similar to the metadata of the base mesh.

[0520] The verifier 2504 compares, for each parameter, the parameters contained in the data output from the intra-submesh decoder 2503 with the parameters contained in the metadata of the base mesh, and determines whether these parameters match.

[0521] For example, the standard specifies that if these parameters do not match, the decoding device 200 will execute a process that is determined to be a conformance violation, indicating that the bitstream is not in conformance with the standard. If the base mesh frame decoder 2501 determines that a conformance violation has occurred, it will, for example, stop the process as an error. Note that the base mesh frame decoder 2501 does not need to stop the process if there is no problem even if it does not stop the process.

[0522] In this way, if these parameters do not match, the base mesh frame decoder 2501 performs processing assuming that the data output from the intra sub-mesh decoder 2503 is non-conformant data.

[0523] If a three-dimensional mesh is divided into multiple submeshes, processing is performed for each submesh. If a three-dimensional mesh is divided into multiple submeshes, processing is performed for each submesh. Specifically, it is determined whether parameters, such as the type and number of dimensions of attribute information included in each submesh, match corresponding parameters included in the metadata of the base mesh. In other words, it is also determined whether the parameters in all submeshes match each other. If, as a result of these determinations, a parameter does not match in any submesh or if a parameter does not match between submeshes, the parameter is determined to be a conformance violation. These processes may be specified as part of the compatibility requirements for the decoder / decoding device 200 in the standard. If any of the multiple submeshes is a conformance violation, processing may be performed on the non-conformant data for the submesh determined to be a conformance violation, or all submeshes may be determined to be a conformance violation.

[0524] FIG. 74 is a flowchart showing an example of processing by the verifying unit 2504 according to the embodiment.

[0525] First, the checker 2504 analyzes the metadata of the base mesh (S401). For example, the checker 2504 acquires the metadata included in the base mesh sub-bitstream from the metadata acquirer 2502 and checks the data structure of the acquired metadata.

[0526] Next, the verifying unit 2504 analyzes the metadata output from the intra sub-mesh decoder 2503 (S402). For example, the verifying unit 2504 obtains the metadata indicating the data structure output from the intra sub-mesh decoder 2503.

[0527] Next, the verifying unit 2504 determines whether the parameters of the base mesh metadata and the metadata output from the intra-submesh decoder 2503 match (S403). Specifically, the verifying unit 2504 determines whether the data structure of the base mesh metadata and the data structure indicated by the metadata output from the intra-submesh decoder 2503 match.

[0528] If the verifying unit 2504 determines that the parameters of the base mesh metadata and the metadata output from the intra-submesh decoder 2503 match (Yes in S403), the process ends. In other words, in this case, the decoding device 200 determines that the data output from the intra-submesh decoder 2503 complies with the standard, and continues the process.

[0529] On the other hand, if the checker 2504 determines that the parameters of the base mesh metadata and the metadata output from the intra-submesh decoder 2503 do not match (No in S403), it determines that the data output from the intra-submesh decoder 2503 violates conformance (S404). In this case, the decoding device 200 executes a predetermined process when a conformance violation occurs, or stops the process.

[0530] It should be noted that the intra sub-mesh decoder 2503 does not necessarily output information relating to the data structure.

[0531] In a second method for determining whether the data output from the intra-submesh decoder 2503 complies with the standard, if the intra-submesh decoder 2503 does not output information regarding the data structure, the verifying unit 2504 verifies the data output from the intra-submesh decoder 2503 to determine whether the data output from the intra-submesh decoder 2503 matches the parameters included in the metadata of the base mesh.

[0532] In the process of checking whether the data structure of the attribute information matches, the checker 2504 first determines whether the data output from the intra-submesh decoder 2503 contains the same number of pieces of attribute information as the number indicated in bmesh_attrCount. If the data output from the intra-submesh decoder 2503 does not contain the same number of pieces of attribute information as the number indicated in bmesh_attrCount, the checker 2504 determines that there is a conformance violation.

[0533] On the other hand, if the number of attribute information included in the data output from the intra-submesh decoder 2503 is correct, i.e., if the number indicated in bmesh_attrCount matches the number of attribute information included in the data, the checker 2504 determines whether the attributeType, attributeDimension, etc. are correct for each attribute information included in the data. That is, the checker 2504 determines whether the type and number of dimensions of each attribute information included in the data match the parameters included in the metadata of the base mesh. If the type and / or number of dimensions of the attribute information included in the data are incorrect, the checker 2504 determines that the data violates conformance. If the three-dimensional mesh is divided into multiple submeshes, processing is performed for each submesh.

[0534] If a three-dimensional mesh is divided into multiple submeshes, processing is performed for each submesh. Specifically, it is determined whether parameters such as the type and number of dimensions of attribute information included in each submesh match the corresponding parameters included in the metadata of the base mesh. In other words, it is also determined whether the parameters in all submeshes match each other. If, as a result of these determinations, the parameters do not match in any submesh or if the parameters do not match between submeshes, the parameters are determined to be in conformance violation. These processes may be specified as part of the compatibility requirements for the decoder / decoding device 200 in the standard.

[0535] FIG. 75 is a flowchart showing another example of the process of the verifying unit 2504 according to the embodiment.

[0536] First, the checker 2504 analyzes the metadata of the base mesh (S411). For example, the checker 2504 acquires the metadata included in the base mesh sub-bitstream from the metadata acquirer 2502 and checks the data structure of the acquired metadata.

[0537] Next, the checker 2504 analyzes the metadata output from the intra-submesh decoder 2503 based on the metadata of the base mesh (S412). For example, the checker 2504 determines whether the data output from the intra-submesh decoder 2503 contains the same number of attribute information as the number indicated by bmesh_attrCount included in the metadata of the base mesh.

[0538] Next, the verifying unit 2504 determines whether the parameters of the base mesh metadata and the metadata output from the intra-submesh decoder 2503 match (S413). For example, if the data output from the intra-submesh decoder 2503 contains the same number of attribute information as the number indicated by bmesh_attrCount in the base mesh metadata, the verifying unit 2504 determines that the parameters match. On the other hand, for example, if the data output from the intra-submesh decoder 2503 does not contain the same number of attribute information as the number indicated by bmesh_attrCount in the base mesh metadata, the verifying unit 2504 determines that the parameters do not match.

[0539] If the verifying unit 2504 determines that the parameters of the base mesh metadata and the metadata output from the intra-submesh decoder 2503 match (Yes in S413), the process ends. In other words, in this case, the decoding device 200 determines that the data output from the intra-submesh decoder 2503 complies with the standard, and continues the process.

[0540] On the other hand, if the checker 2504 determines that the parameters of the base mesh metadata and the metadata output from the intra-submesh decoder 2503 do not match (No in S413), it determines that the data output from the intra-submesh decoder 2503 violates conformance (S414). In this case, the decoding device 200 executes a predetermined process when a conformance violation occurs, or stops the process.

[0541] Next, a specific example of a method for checking the number of pieces of attribute information will be described. Figures 76 and 77 are diagrams showing attribute information of submeshes according to an embodiment. More specifically, Figures 76 and 77 are diagrams showing specific examples of attribute information included in data output from the intra-submesh decoder 2503. In the example shown in Figure 76, submesh 1 has color information (Color) and normal information (Normal), and submesh 2 has only color information (Color). In the example shown in Figure 77, submesh 1 has color information (Color) and normal information (Normal), and submesh 2 has color information (Color) and UV coordinate information (UVCoord).

[0542] For example, suppose that bmesh_attriCount = 2 is set in the base mesh parameter set. As a result, the number of attribute information items processed by the base mesh frame decoder 2501 (specifically, the number of types of attribute information items) or the number of attribute information items output from the base mesh frame decoder 2501 can be determined to be 2. More specifically, the number of attribute information items for each sub-mesh output from the intra sub-mesh decoder 2503 can be determined to be 2.

[0543] Also, the number of submeshes is assumed to be 2. In this case, if both submesh 1 and submesh 2 belong to the same base mesh, the number of attribute information pieces of data for each submesh output from the intra-submesh decoder 2503 is expected to be 2.

[0544] For example, as shown in the example of Figure 76, when the number of attribute information of submesh 2 is 1, the checker 2504 determines that the data structure of the data output from the intra submesh decoder 2503 does not match the data structure of the metadata of the base mesh, and determines that there is a conformance violation. Note that the checker 2504 may also determine that there is a conformance violation when the number of attribute information of submesh 1 and the number of attribute information of submesh 2 do not match.

[0545] Furthermore, for example, when color information and normal information are set as attributeType (type of attribute information) in a base mesh parameter set, as in the example shown in Figure 77, when the attributeType of submesh 2 is different from the attributeType indicated in the base mesh parameter set, the verifier 2504 determines that the data structure of the data output from the intra-submesh decoder 2503 does not match the data structure of the metadata of the base mesh, and determines that there is a conformance violation.

[0546] This method is a method for specifying in a standard to prevent inconsistencies in the data structure when combining each sub-mesh when codec-independent encoding and decoding processes are included and when it is not guaranteed that the data structure (data configuration) of the output data resulting from the codec-independent encoding process is the same as the data structure specified in the base mesh, and a method for checking whether the bitstream complies with the standard.

[0547] It is clearly specified that the data structure of the data specified in the base mesh is the same as the data structure of the data processed within the base mesh processing, and / or the data structure of the data output from the base mesh frame decoder 2501, and if the data structures do not match, it is specified as a violation of the bitstream specifications (conformance violation).

[0548] In the conformance violation determination process, whether or not there is a conformance violation is determined by comparing the metadata of the base mesh with the metadata or output data output from the codec-independent layer.

[0549] These provisions and processes enable integrated processing (interframe decoding and sub-mesh combining processing (reconstruction processing)) using data output from a codec-independent layer whose data structure is not guaranteed and a base mesh whose data structure is guaranteed.

[0550] Furthermore, by performing the conformance violation determination process for each sub-mesh, it is possible to deal with cases where the data structures differ between sub-meshes, and to prevent inconsistencies in the sub-mesh joining process.

[0551] Furthermore, by providing such a provision in the base mesh layer after the output of the codec-independent layer and comparing it with the metadata of the base layer, it becomes possible to encode and decode the base mesh, including codec-independent encoding and decoding.

[0552] <Representative Example> Figure 78 is a flow diagram showing an example of basic decoding processing according to this embodiment. For example, the decoding device 200 shown in Figure 25 includes a circuit 251 and a memory 252 connected to the circuit 251. In the decoding device 200, the circuit 251 performs the following processing during operation.

[0553] First, the decoding device 200 acquires a plurality of coded data items related to a plurality of submeshes (S421). The coded data items are, for example, coded attribute information items for the respective submeshes. The decoding device 200 acquires, for example, a bitstream output from the coding device 100, and acquires the coded data items included in the acquired bitstream.

[0554] Next, the decoding device 200 generates a plurality of sub-mesh data relating to a plurality of sub-meshes by decoding the plurality of coded data (S422). The sub-mesh data is, for example, attribute information of the sub-meshes.

[0555] Next, the decoding device 200 outputs the generated submesh data (S423). The decoding device 200 outputs the submesh data to, for example, a reconstructor that uses the submesh data to generate (reconstruct) a three-dimensional mesh by combining the submeshes. The decoding device 200 includes, for example, a base mesh frame decoder 2501. The decoding device 200 outputs the submesh data to, for example, a reconstructor 2506.

[0556] Here, the output sub-mesh data have the same attribute information configuration. Sub-mesh data having the same attribute information configuration is, for example, sub-mesh data in which the types of attribute information included in the sub-mesh data and the attribute information for each type are the same. In other words, sub-mesh data having the same attribute information configuration is, for example, data in which the number and types of attribute information are the same.

[0557] According to this, since the plurality of sub-mesh data have the same attribute information configuration, when the plurality of sub-mesh data are used to combine the plurality of sub-meshes to reconstruct a three-dimensional mesh, it is possible to prevent the combination of attribute information from being incorrectly combined, thereby preventing the reconstruction of an incorrect three-dimensional mesh.

[0558] Furthermore, for example, the decoding device 200 may further determine a decoding method for each of the plurality of coded data, and in generating the plurality of sub-mesh data (S422), generate the plurality of sub-mesh data by decoding each of the plurality of coded data using the determined decoding method. In other words, the decoding device 200 independently determines a decoding method for each of the plurality of coded data, and decodes the data using the determined decoding method.

[0559] This allows each of the plurality of coded data to be decoded by, for example, a decoding method corresponding to the coding method of the coded data.

[0560] The decoding methods may be the same or different from each other. The decoding methods may be determined by any method, and are not particularly limited. The decoding methods may be determined in advance, or information indicating the decoding method may be included in a bitstream containing multiple pieces of coded data.

[0561] Furthermore, for example, the decoding device 200 may further determine whether the generated plurality of sub-mesh data have the same attribute information configuration.

[0562] According to this, it is possible to determine whether a 3D mesh can be correctly reconstructed based on the determination result of whether multiple submesh data have the same attribute information configuration. Note that the decoding device 200 may determine whether to output multiple submesh data based on the determination result of whether multiple submesh data have the same attribute information configuration. For example, if the decoding device 200 determines that multiple submesh data have the same attribute information configuration, it may output the multiple submesh data, and if it determines that the multiple submesh data do not have the same attribute information configuration, it may not output the multiple submesh data. The decoding device 200 may output only the submesh data determined to have the same attribute information configuration among the multiple submesh data. Furthermore, for example, if the decoding device 200 determines that the multiple submesh data do not have the same attribute information configuration, it may output information indicating a conformance violation.

[0563] Furthermore, for example, if the decoding device 200 determines that the generated plurality of submesh data have the same attribute information configuration, it may output the generated plurality of submesh data to a reconstructor (e.g., reconstructor 2506) that reconstructs a three-dimensional mesh using the plurality of submesh data. On the other hand, for example, if the decoding device 200 determines that the plurality of submesh data do not have the same attribute information configuration, it may not output the generated plurality of submesh data to the reconstructor. In this way, for example, if the attribute information configuration of each submesh is common, the reconstruction process is performed.

[0564] This allows the reconstructor to correctly reconstruct a three-dimensional mesh using the input sub-mesh data.

[0565] Furthermore, for example, the configuration of the attribute information of the output sub-mesh data may be the same as the configuration of the attribute information indicated in the metadata of the base mesh related to the sub-meshes. The metadata of the base mesh is, for example, metadata common to the sub-meshes. The metadata is, for example, included in a bitstream including the multiple encoded data. The decoding device 200 acquires the metadata from the bitstream, for example.

[0566] This makes it possible to output a plurality of sub-mesh data having the same structure as the structure of the attribute information indicated in the metadata of the base mesh.

[0567] <Other Examples> Although aspects of the encoding device 100 and the decoding device 200 have been described above according to the embodiments, the aspects of the encoding device 100 and the decoding device 200 are not limited to the embodiments. Modifications conceivable by those skilled in the art may be applied to the embodiments, and multiple components in the embodiments may be combined in any manner.

[0568] For example, a process performed by a specific component in the embodiment may be performed by another component instead of the specific component. Also, the order of multiple processes may be changed, or multiple processes may be performed in parallel.

[0569] Furthermore, as described above, at least some of the configurations of the present disclosure may be implemented as an integrated circuit. At least some of the processes of the present disclosure may be used as an encoding method or a decoding method. A program for causing a computer to execute the encoding method or the decoding method may be used. A non-transitory computer-readable recording medium on which the program is recorded may be used. A bitstream for causing the decoding device 200 to perform a decoding process may be used.

[0570] Furthermore, at least some of the configurations and processes of the present disclosure may be used as a transmitting device, a receiving device, a transmitting method, or a receiving method. A program for causing a computer to execute the transmitting method or the receiving method may be used. Furthermore, a non-transitory computer-readable recording medium on which the program is recorded may be used.

[0571] The present disclosure is useful, for example, in encoding devices, decoding devices, transmitting devices, receiving devices, etc. related to three-dimensional meshes, and is applicable to computer graphics systems, three-dimensional data display systems, etc.

[0572] 100 Encoding device 101, 121, 144 Vertex information encoder 102, 145 Connection information encoder 103, 122 Attribute information encoder 104, 204, 1103 Preprocessor 105, 205, 2106 Postprocessor 110 Three-dimensional data encoding system 111, 211 Controller 112, 212 Input / output processor 113 Three-dimensional data encoder 114 System multiplexer 115 Three-dimensional data generator 123 Metadata encoder 124, 1274 Multiplexer 131 Vertex image generator 132 Attribute image generator 133 Metadata generator 134 Video encoder 141 Two-dimensional data encoder 142 Mesh data encoder 143 Texture encoder 148 Description encoder 151, 251 Circuit 152, 252 Memory 200 Decoding device 201, 221, 244 Vertex information decoder 202, 245 Connection information decoder 203, 222 Attribute information decoder 210 3D data decoding system 213 3D data decoder 214 System demultiplexer 215, 247 Presentation device 216 User interface 223 Metadata decoder 224 Demultiplexer 231 Vertex information generator 232 Attribute information generator 234 Video decoder 241 2D data decoder 242 Mesh data decoder 243 Texture decoder 246 Mesh reconstructor 248 Description decoder 300 Network 310 External connector 511 Volumetric capture device 512 Projector 513 Base mesh encoder 514 Displacement encoder 515 Attribute Encoder 516 Other Type Encoder 613 Base Mesh Decoder 614 Displacement Decoder 615 Attribute Decoder 616 Other Type Decoder 617 3D Reconstructor 1101 Input Mesh 1102, 2108 Attribute Map 1104, 2103 Base Mesh 1105, 2104 Displacement Data 1106 Compressor 1107, 1304, 2101 Bitstream1108, 2105 Metadata 1231 Demultiplexer 1232, 1262 Switch 1233 Static Mesh Decoder 1234, 1264 Mesh Buffer 1235 Motion Decoder 1236, 1266 Base Mesh Reconstructor 1237, 1240 Inverse Quantizer 1238, 1243 Video Decoder 1239 Image Unpacker 1241 Inverse Wavelet Transformer 1242 Reconstructor 1244, 1272 Color Transformer 1251 Decoded Base Mesh 1252, 1402 Subdivider 1253 Subdivided Mesh 1254 Decoded Displacement Data 1255 Displacer 1256 Decoded 3D Mesh 1261, 1269 Quantizer 1263 Static Mesh Encoder 1265 Motion Encoder 1267 Displacement data updater 1268 Wavelet transformer 1270 Image packer 1271, 1273 Video encoder 1301, 2304 Mesh frame 1302, 2301, 2302 Base mesh frame 1303, 2303 Displacement information 1401 Base mesh generator 1403 Displacement data generator 2102 Decompressor 2107 Output mesh 2109 Combiner 2305 Decoder 2306 Post-decoder 2307 Pre-reconstructor 2308, 2402, 2506 Reconstructor 2309 Post-reconstructor 2310 Adapter 2311 Final 3D mesh frame 2401 Sub-mesh decoder 2501 Base mesh frame decoder 2502 Metadata obtainer 2503 Intra-sub-mesh decoder 2504 Verifier 2505 Inter-submesh decoder

Claims

1. A decoding method comprising: acquiring a plurality of coded data relating to a plurality of submeshes; decoding the plurality of coded data to generate a plurality of submesh data relating to the plurality of submeshes; outputting the generated plurality of submesh data; and outputting the plurality of output submesh data having the same attribute information configuration.

2. The decoding method according to claim 1, further comprising determining a decoding method for each of the plurality of coded data, and generating the plurality of sub-mesh data by decoding each of the plurality of coded data using the determined decoding method.

3. The decoding method according to claim 1, further comprising determining whether the generated plurality of sub-mesh data have the same attribute information configuration.

4. A decoding method as described in claim 3, wherein, if it is determined that the generated plurality of sub-mesh data have the same attribute information configuration, the generated plurality of sub-mesh data is output to a reconstructor that reconstructs a three-dimensional mesh using the plurality of sub-mesh data.

5. A decoding method according to any one of claims 1 to 4, wherein the configuration of the attribute information of the output sub-mesh data is identical to the configuration of the attribute information indicated in the metadata of the base mesh relating to the sub-meshes.

6. A decoding device comprising: a circuit; and a memory connected to the circuit, wherein the circuit, in operation, obtains a plurality of coded data relating to a plurality of submeshes; generates a plurality of submesh data relating to the plurality of submeshes by decoding the plurality of coded data; and outputs the generated plurality of submesh data, wherein the output plurality of submesh data have the same attribute information configuration.

Citation Information

Patent Citations

  • Encoding device, decoding device, encoding method, and decoding method

    WO2024075608A1

  • Information processing device and method

    WO2024084952A1