Encoding method, decoding method, encoding device, and decoding device

By packing multiple two-dimensional attribute information sets into a single image frame unit and using image encoding, the method addresses the challenge of increased overhead in encoding three-dimensional data, achieving efficient encoding and decoding.

WO2025216144A1PCT designated stage Publication Date: 2025-10-16PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/013498
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-11
Filing Date
2025-04-02
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Existing encoding methods for three-dimensional data, such as three-dimensional meshes, face challenges in efficiently handling various attribute information with different types, leading to increased coding and processing overhead.

Method used

The method involves packing multiple two-dimensional attribute information sets of three-dimensional data into a single image frame unit and encoding it using an image encoding method, allowing for collective encoding of attribute information sets as a single image frame.

Benefits of technology

This approach reduces the overhead in encoding and decoding processes by collectively handling multiple two-dimensional attribute information sets, thereby minimizing the amount of code and processing required.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025013498_16102025_PF_FP_ABST
    Figure JP2025013498_16102025_PF_FP_ABST
Patent Text Reader

Abstract

This encoding method involves that: a plurality of two-dimensional attribute information sets, which are a plurality of two-dimensional attribute information sets of three-dimensional data and of which attribute types are different from each other, are packed in one image frame unit which is a unit corresponding to one image frame (S501); and the one image frame unit is encoded by using an image encoding method which is a method used for encoding an image (S502).
Need to check novelty before this filing date? Find Prior Art

Description

Encoding method, decoding method, encoding device, and decoding device

[0001] The present disclosure relates to encoding methods and the like.

[0002] In US Pat. No. 6,299,549 a method and apparatus for encoding and decoding three-dimensional mesh data is proposed.

[0003] Japanese Patent Application Laid-Open No. 2006-187015

[0004] Further improvements are desired in the encoding or decoding process for three-dimensional data. The present disclosure aims to improve the encoding process and the like for three-dimensional data.

[0005] An encoding method according to one aspect of the present disclosure packs a plurality of two-dimensional attribute information sets of three-dimensional data, each having different attribute types, into a single image frame unit, which is a unit corresponding to a single image frame, and encodes the single image frame unit using an image encoding method that is a method used for encoding images.

[0006] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0007] The present disclosure may contribute to improvements in encoding processes and the like related to three-dimensional data.

[0008] 1 is a conceptual diagram showing a three-dimensional mesh according to an embodiment. FIG. 2 is a conceptual diagram showing basic elements of a three-dimensional mesh according to an embodiment. FIG. 3 is a conceptual diagram showing mapping according to an embodiment. FIG. 4 is a block diagram showing a configuration example of an encoding / decoding system according to an embodiment. FIG. 5 is a block diagram showing a configuration example of an encoding device according to an embodiment. FIG. 6 is a block diagram showing another configuration example of an encoding device according to an embodiment. FIG. 7 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 8 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 9 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 10 is a conceptual diagram showing another configuration example of a bit stream according to an embodiment. FIG. 11 is a conceptual diagram showing yet another configuration example of a bit stream according to an embodiment. FIG. 12 is a block diagram showing a specific example of an encoding / decoding system according to an embodiment. FIG. 13 is a conceptual diagram showing an example configuration of point cloud data according to an embodiment. FIG. 14 is a conceptual diagram showing an example data file of point cloud data according to an embodiment. FIG. 15 is a conceptual diagram showing an example configuration of mesh data according to an embodiment. FIG. 16 is a conceptual diagram showing an example data file of mesh data according to an embodiment. FIG. 17 is a conceptual diagram showing types of three-dimensional data according to an embodiment. FIG. 18 is a block diagram showing an example configuration of a three-dimensional data encoder according to an embodiment. FIG. 19 is a block diagram showing an example configuration of a three-dimensional data decoder according to an embodiment. FIG. 19 is a block diagram showing another configuration example of a three-dimensional data encoder according to an embodiment. FIG. 19 is a block diagram showing another configuration example of a three-dimensional data decoder according to an embodiment. FIG. 1 is a conceptual diagram showing a specific example of encoding processing according to an embodiment. FIG. 2 is a conceptual diagram showing a specific example of decoding processing according to an embodiment. FIG. 3 is a block diagram showing an implementation example of an encoding device according to an embodiment. FIG. 4 is a block diagram showing an implementation example of a decoding device according to an embodiment. FIG. 5 is a block diagram showing another configuration example of an encoding / decoding system according to an embodiment. FIG. 6 is a block diagram showing another configuration example of an encoding device according to an embodiment. FIG. 7 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 8 is a block diagram showing yet another configuration example of an encoding device according to an embodiment. FIG. 9 is a block diagram showing yet another configuration example of a decoding device according to an embodiment. FIG. 10 is a flow diagram showing processing of an encoding device according to an embodiment. FIG. 11 is an explanatory diagram conceptually showing encoding of a mesh frame according to an embodiment.1 is a flow diagram showing processing of a decoding device according to an embodiment. FIG. 2 is an explanatory diagram conceptually showing decoding of a mesh frame according to an embodiment. FIG. 3 is a block diagram showing an example configuration of a decoding device according to an embodiment. FIG. 4 is a block diagram showing an example configuration of an encoding device according to an embodiment. FIG. 5 is a block diagram showing an example configuration of a decoding device according to an embodiment. FIG. 6 is an explanatory diagram showing an example of subdivision according to an embodiment. FIG. 7 is an explanatory diagram showing an example of displacement of vertices after displacement after subdivision according to an embodiment. FIG. 8 is an explanatory diagram showing example vertices of an original mesh according to an embodiment. FIG. 9 is an explanatory diagram showing an example mesh according to an embodiment. FIG. 10 is an explanatory diagram showing an example of division of a mesh into sub-meshes according to an embodiment. FIG. 11 is a first explanatory diagram showing an example of packing of displacement information into an image frame according to an embodiment. FIG. 12 is a second explanatory diagram showing an example of packing of displacement information into an image frame according to an embodiment. FIG. 13 is a block diagram showing another example configuration of an encoding device according to an embodiment. FIG. 14 is a block diagram showing an example configuration of a preprocessor according to an embodiment. FIG. 15 is a block diagram showing another example configuration of a decoding device according to an embodiment. FIG. 16 is a block diagram showing a specific example configuration of a decoding device according to an embodiment. FIG. 17 is a conceptual diagram showing a specific example of a bitstream according to an embodiment. FIG. 18 is a block diagram showing yet another example configuration of an encoding device according to an embodiment. FIG. 19 is a block diagram showing yet another example configuration of a decoding device according to an embodiment. FIG. 1 is a flow diagram showing another example of processing by an encoding device according to an embodiment. FIG. 2 is a flow diagram showing another example of processing by a decoding device according to an embodiment. FIG. 3 is a conceptual diagram showing an example of the number of attribute maps and storage locations of attribute types in a bitstream according to an embodiment. FIG. 4 is a conceptual diagram showing another example of the number of attribute maps and storage locations of attribute types in a bitstream according to an embodiment. FIG. 5 is an explanatory diagram showing an example of attribute types and their explanations according to an embodiment. FIG. 6 is a conceptual diagram showing a first packing example of a plurality of attribute maps according to an embodiment. FIG. 7 is a syntax diagram showing an example syntax structure in the first packing example according to an embodiment. FIG. 8 is a syntax diagram showing another example syntax structure according to an embodiment. FIG. 9 is a syntax diagram showing yet another example syntax structure according to an embodiment.FIG. 1 is a conceptual diagram showing a second packing example of multiple attribute maps according to an embodiment. FIG. 2 is a syntax diagram showing an example of a syntax structure in the second packing example according to the embodiment. FIG. 3 is an explanatory diagram showing an example of a video subcomponent ID in the second packing example according to the embodiment. FIG. 4 is a conceptual diagram showing a third packing example of multiple attribute maps according to the embodiment. FIG. 5 is a syntax diagram showing an example of a syntax structure in the third packing example according to the embodiment. FIG. 6 is a conceptual diagram showing an example of unpacking of multiple attribute maps according to the embodiment. FIG. 7 is a conceptual diagram showing an example of an attribute map request and response according to the embodiment. FIG. 8 is a conceptual diagram showing an example of rendering using multiple attributes according to the embodiment. FIG. 9 is a flow diagram showing an example of a basic encoding process according to the embodiment. FIG. 10 is a flow diagram showing an example of a basic decoding process according to the embodiment.

[0009] Introduction Three-dimensional (3D) meshes are used in computer graphics images, which may be composed of multiple temporally distinct frames, each of which may be represented by a 3D mesh.

[0010] A 3D mesh is composed of vertex information indicating the positions of each of the vertices in 3D space, connectivity information indicating the connections between the vertices, and attribute information indicating the attributes of each vertex or face. Each face is constructed according to the connectivity between the vertices. Various computer graphics images can be expressed using such 3D meshes.

[0011] However, when encoding three-dimensional data such as a three-dimensional mesh, various attribute information having various attribute types may be signaled, which may increase the amount of coding and processing.

[0012] Therefore, the encoding method of Example 1 packs multiple two-dimensional attribute information sets of three-dimensional data, each having different attribute types, into one image frame unit, which is a unit corresponding to one image frame, and encodes the one image frame unit using an image encoding method that is a method used for encoding images.

[0013] This may enable multiple two-dimensional attribute information sets with different attribute types to be collectively encoded as a single image frame unit using an image encoding method. Therefore, it may be possible to reduce the overhead in encoding multiple two-dimensional attribute information sets. Therefore, it may be possible to reduce the amount of code and the amount of processing.

[0014] The encoding method of Example 2 may be the encoding method of Example 1, in which the one image frame unit is the one image frame, and the multiple two-dimensional attribute information sets are packed into the one image frame.

[0015] This may enable encoding a plurality of two-dimensional attribute information sets as a single image frame using an image encoding method, thereby reducing overhead and reducing the amount of coding and processing.

[0016] The encoding method of Example 3 may be the encoding method of Example 1, in which the one image frame unit is a plurality of components of the one image frame, and the plurality of two-dimensional attribute information sets are packed into the plurality of components of the one image frame.

[0017] This may enable multiple sets of two-dimensional attribute information to be coded as multiple components of one image frame using an image coding method, thereby reducing overhead and reducing the amount of coding and processing.

[0018] The encoding method of Example 4 may be the encoding method of Example 1, in which the one image frame unit is a plurality of layers corresponding to the one image frame, and the plurality of two-dimensional attribute information sets are packed into the plurality of layers corresponding to the one image frame.

[0019] This may enable encoding multiple sets of two-dimensional attribute information as multiple layers of one image frame using an image encoding method, thereby reducing overhead and reducing the amount of coding and processing.

[0020] Furthermore, the encoding method of Example 5 may be any of the encoding methods of Examples 1 to 4, in which the plurality of two-dimensional attribute information sets include, as two-dimensional attribute information sets, images onto which attributes of each of the plurality of faces of the three-dimensional data are mapped.

[0021] This may enable encoding the image onto which the attributes of each surface are mapped as one of multiple sets of two-dimensional attribute information, thus enabling efficient encoding of the attributes of each surface.

[0022] Furthermore, the encoding method of Example 6 may be any of the encoding methods of Examples 1 to 5, in which the plurality of two-dimensional attribute information sets include an image onto which the three-dimensional data is projected as a two-dimensional attribute information set.

[0023] This may enable encoding of the projection image of the three-dimensional data as one of a plurality of two-dimensional attribute information sets, and therefore may enable efficient encoding of the projection image of the three-dimensional data.

[0024] Furthermore, the encoding method of Example 7 may be any of the encoding methods of Examples 1 to 6, further encoding, for each of the plurality of two-dimensional attribute information sets, placement information indicating the placement of the two-dimensional attribute information set in the one image frame unit.

[0025] This may enable adaptive arrangement and efficient identification of two-dimensional attribute information sets in units of one image frame.

[0026] Furthermore, the encoding method of Example 8 may be the encoding method of Example 7, in which the placement information includes information indicating the position and size of the two-dimensional attribute information set in the image included in the one image frame unit.

[0027] This may allow adaptive placement and efficient identification of sets of two-dimensional attribute information in an image.

[0028] Furthermore, the encoding method of Example 9 may be the encoding method of Example 7 or 8, in which the placement information includes information for identifying an image that includes the two-dimensional attribute information set among the multiple images included in the one image frame unit.

[0029] This may allow adaptive placement and efficient identification of sets of two-dimensional attribute information across multiple images.

[0030] Furthermore, the encoding method of Example 10 may be any of the encoding methods of Examples 1 to 9, further encoding related information indicating the number of the plurality of two-dimensional attribute information sets and at least one of the attribute types of each of the plurality of two-dimensional attribute information sets.

[0031] This may make it possible to appropriately share the number of multiple two-dimensional attribute information sets or the information of each attribute type of multiple two-dimensional attribute information sets.

[0032] In addition, the decoding method of Example 11 is a decoding method that uses an image decoding method, which is a method used to decode images, to decode one image frame unit, which is a unit corresponding to one image frame, and unpacks multiple two-dimensional attribute information sets of three-dimensional data, each having different attribute types, from the one image frame unit.

[0033] This may enable multiple two-dimensional attribute information sets with different attribute types to be decoded together as a single image frame unit using an image decoding method. Therefore, it may be possible to reduce the overhead in decoding multiple two-dimensional attribute information sets. Therefore, it may be possible to reduce the amount of coding and the amount of processing.

[0034] Furthermore, the decoding method of Example 12 may be the decoding method of Example 11, in which the one image frame unit is the one image frame, and the multiple two-dimensional attribute information sets are unpacked from the one image frame.

[0035] This may enable decoding of multiple two-dimensional attribute information sets as a single image frame using an image decoding method, thereby reducing overhead and reducing the amount of coding and processing.

[0036] Furthermore, the decoding method of Example 13 may be the decoding method of Example 11, in which the one image frame unit is a plurality of components of the one image frame, and the plurality of two-dimensional attribute information sets are unpacked from the plurality of components of the one image frame.

[0037] This may enable decoding of multiple two-dimensional attribute information sets as multiple components of one image frame using an image decoding method, thereby reducing overhead and reducing the amount of coding and processing.

[0038] Furthermore, the decoding method of Example 14 may be the decoding method of Example 11, in which the one image frame unit is a plurality of layers corresponding to the one image frame, and the plurality of two-dimensional attribute information sets are unpacked from the plurality of layers corresponding to the one image frame.

[0039] This may enable decoding of multiple two-dimensional attribute information sets as multiple layers of one image frame using an image decoding method, thereby reducing overhead and reducing the amount of coding and processing.

[0040] Furthermore, the decoding method of Example 15 may be any of the decoding methods of Examples 11 to 14, in which the plurality of two-dimensional attribute information sets include, as two-dimensional attribute information sets, images onto which attributes of each of the plurality of surfaces of the three-dimensional data are mapped.

[0041] This may enable the image onto which the attributes of each surface are mapped to be decoded as one of a plurality of sets of two-dimensional attribute information, thereby enabling the attributes of each surface to be decoded efficiently.

[0042] Furthermore, the decoding method of Example 16 may be any of the decoding methods of Examples 11 to 15, in which the plurality of two-dimensional attribute information sets include an image onto which the three-dimensional data is projected as a two-dimensional attribute information set.

[0043] This may enable the projection image of the three-dimensional data to be decoded as one of a plurality of two-dimensional attribute information sets, thereby enabling the projection image of the three-dimensional data to be decoded efficiently.

[0044] Furthermore, the decoding method of Example 17 may be any of the decoding methods of Examples 11 to 16, further comprising decoding, for each of the plurality of two-dimensional attribute information sets, placement information indicating the placement of the two-dimensional attribute information set in the one image frame unit.

[0045] This may enable adaptive arrangement and efficient identification of two-dimensional attribute information sets in units of one image frame.

[0046] Furthermore, the decoding method of Example 18 may be the decoding method of Example 17, in which the placement information includes information indicating the position and size of the two-dimensional attribute information set in the image included in the one image frame unit.

[0047] This may allow adaptive placement and efficient identification of sets of two-dimensional attribute information in an image.

[0048] Furthermore, the decoding method of Example 19 may be the decoding method of Example 17 or 18, in which the placement information includes information for identifying an image that includes the two-dimensional attribute information set from among the multiple images included in the one image frame unit.

[0049] This may allow adaptive placement and efficient identification of sets of two-dimensional attribute information across multiple images.

[0050] Furthermore, the decoding method of Example 20 may be any of the decoding methods of Examples 11 to 19, further decoding related information indicating the number of the plurality of two-dimensional attribute information sets and at least one of the attribute types of each of the plurality of two-dimensional attribute information sets.

[0051] This may make it possible to appropriately share the number of multiple two-dimensional attribute information sets or the information of each attribute type of multiple two-dimensional attribute information sets.

[0052] In addition, the encoding device of Example 21 includes a circuit and a memory accessible by the circuit, and the circuit packs a plurality of two-dimensional attribute information sets of three-dimensional data, the plurality of two-dimensional attribute information sets having different attribute types, into one image frame unit, which is a unit corresponding to one image frame, and encodes the one image frame unit using an image encoding method, which is a method used for encoding images.

[0053] This may enable multiple two-dimensional attribute information sets with different attribute types to be collectively encoded as a single image frame unit using an image encoding method. Therefore, it may be possible to reduce the overhead in encoding multiple two-dimensional attribute information sets. Therefore, it may be possible to reduce the amount of code and the amount of processing.

[0054] In addition, the decoding device of Example 22 includes a circuit and a memory accessible by the circuit, and the circuit uses an image decoding method that is a method used for decoding images to decode one image frame unit, which is a unit corresponding to one image frame, and unpacks from the one image frame unit a plurality of two-dimensional attribute information sets of three-dimensional data having different attribute types.

[0055] This may enable multiple two-dimensional attribute information sets with different attribute types to be decoded together as a single image frame unit using an image decoding method. Therefore, it may be possible to reduce the overhead in decoding multiple two-dimensional attribute information sets. Therefore, it may be possible to reduce the amount of coding and the amount of processing.

[0056] Furthermore, these comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0057] <Expressions and Terms> The following expressions and terms are used herein.

[0058] (1) Three-dimensional Mesh A three-dimensional mesh is a collection of multiple faces, and represents, for example, a three-dimensional object. A three-dimensional mesh is mainly composed of vertex information, connectivity information, and attribute information. A three-dimensional mesh may be expressed as a polygon mesh or a mesh. A three-dimensional mesh may also vary over time. A three-dimensional mesh may include metadata related to the vertex information, connectivity information, and attribute information, and may also include other additional information.

[0059] (2) Vertex Information Vertex information is information indicating a vertex. For example, the vertex information indicates the position of a vertex in a three-dimensional space. Furthermore, a vertex corresponds to a vertex of a face that constitutes a three-dimensional mesh. Vertex information may be expressed as "geometry." Furthermore, vertex information may be expressed as position information.

[0060] (3) Connection Information Connection information is information that indicates connections between vertices. For example, connection information indicates connections for forming faces or edges of a three-dimensional mesh. Connection information may be expressed as "Connectivity." Connection information may also be expressed as face information.

[0061] (4) Attribute Information Attribute information is information that indicates attributes of a vertex or a face. For example, attribute information indicates attributes such as a color, an image, and a normal vector associated with a vertex or a face. Attribute information may be expressed as "texture."

[0062] (5) Faces A face is an element that makes up a three-dimensional mesh. Specifically, a face is a polygon on a plane in three-dimensional space. For example, a face can be defined as a triangle in three-dimensional space.

[0063] (6) Plane A plane is a two-dimensional plane in a three-dimensional space. For example, a polygon is formed on a plane, and multiple polygons are formed on multiple planes.

[0064] (7) Bitstream: A bitstream corresponds to coded information. A bitstream may also be referred to as a stream, a coded bitstream, a compressed bitstream, or a coded signal.

[0065] (8) Encoding and Decoding The term encoding may be substituted with terms such as storing, including, writing, describing, signaling, sending, notifying, saving, or compressing, and these terms may be interchangeable. For example, encoding information may mean including the information in a bitstream. Also, encoding information into a bitstream may mean encoding the information to generate a bitstream that includes the encoded information.

[0066] Additionally, the term "decode" may be replaced with terms such as "read," "decode," "read," "load," "derive," "obtain," "receive," "extract," "reconstruct," "reconstruct," "decompress," or "decompress," and these terms may be interchangeable. For example, decoding information may mean obtaining information from a bitstream. Decoding information from a bitstream may mean decoding the bitstream to obtain information contained in the bitstream.

[0067] (9) Ordinal Numbers In the description, ordinal numbers such as first and second may be assigned to components, etc. These ordinal numbers may be changed as appropriate. Furthermore, new ordinal numbers may be assigned to components, etc., or removed. Furthermore, these ordinal numbers may be assigned to elements in order to identify them, and may not correspond to a meaningful order.

[0068] <Three-dimensional mesh> Fig. 1 is a conceptual diagram showing a three-dimensional mesh according to this embodiment. A three-dimensional mesh is composed of multiple faces. For example, each face is a triangle. The vertices of these triangles are defined in three-dimensional space. The three-dimensional mesh then represents a three-dimensional object. Each face may have a color or an image.

[0069] FIG. 2 is a conceptual diagram showing the basic elements of a three-dimensional mesh according to this embodiment. A three-dimensional mesh is composed of vertex information, connection information, and attribute information. The vertex information indicates the positions of the vertices of a face in three-dimensional space. The connection information indicates the connections between the vertices. A face can be identified by the vertex information and connection information. In other words, a colorless three-dimensional object is formed in three-dimensional space by the vertex information and connection information.

[0070] The attribute information may be associated with a vertex or a face. The attribute information associated with a vertex may be expressed as "Attribute Per Point." The attribute information associated with a vertex may indicate an attribute of the vertex itself, or may indicate an attribute of a face connected to the vertex.

[0071] For example, a color may be associated with a vertex as attribute information. The color associated with a vertex may be the color of the vertex itself, or the color of a face connected to the vertex. The color of a face may be the average of multiple colors associated with multiple vertices of the face. Furthermore, a normal vector may be associated with a vertex or a face as attribute information. Such a normal vector can represent the front and back of a face.

[0072] A two-dimensional image may be associated with a surface as attribute information. The two-dimensional image associated with a surface is also expressed as a texture image or an "Attribute Map." Information indicating mapping between the surface and the two-dimensional image may be associated with the surface as attribute information. Information indicating such mapping may be expressed as mapping information, vertex information of a texture image, texture coordinates, or "Attribute UV Coordinate."

[0073] Furthermore, information such as color, image, and moving image used as attribute information may be expressed as "parametric space."

[0074] The attribute information allows texture to be reflected on the three-dimensional object. That is, a three-dimensional object having color is formed in three-dimensional space based on the vertex information, connection information, and attribute information.

[0075] In the above, the attribute information is associated with the vertices or faces, but it may also be associated with the edges.

[0076] 3 is a conceptual diagram illustrating mapping according to this embodiment. For example, a region of a two-dimensional image on a two-dimensional plane can be mapped onto a surface of a three-dimensional mesh in three-dimensional space. Specifically, coordinate information of the region in the two-dimensional image is associated with the surface of the three-dimensional mesh. As a result, an image of the mapped region in the two-dimensional image is reflected on the surface of the three-dimensional mesh.

[0077] By using the mapping, the 2D image used as attribute information can be separated from the 3D mesh. For example, in encoding the 3D mesh, the 2D image may be encoded by an image encoding method or a video encoding method.

[0078] <System Configuration> Fig. 4 is a block diagram showing an example of the configuration of a coding / decoding system according to this embodiment. In Fig. 4, the coding / decoding system includes a coding device 100 and a decoding device 200.

[0079] For example, the encoding device 100 obtains a three-dimensional mesh and encodes the three-dimensional mesh into a bitstream. Then, the encoding device 100 outputs the bitstream to the network 300. For example, the bitstream includes the encoded three-dimensional mesh and control information for decoding the encoded three-dimensional mesh. By encoding the three-dimensional mesh, information about the three-dimensional mesh is compressed.

[0080] The network 300 transmits a bitstream from the encoding device 100 to the decoding device 200. The network 300 may be the Internet, a wide area network (WAN), a local area network (LAN), or a combination of these. The network 300 is not necessarily limited to bidirectional communication, and may be a unidirectional communication network for terrestrial digital broadcasting, satellite broadcasting, or the like.

[0081] Furthermore, the network 300 can be replaced by a recording medium such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc (registered trademark)).

[0082] The decoding device 200 obtains a bitstream and decodes a three-dimensional mesh from the bitstream. By decoding the three-dimensional mesh, information about the three-dimensional mesh is expanded. For example, the decoding device 200 decodes the three-dimensional mesh according to a decoding method corresponding to the encoding method used by the encoding device 100 to encode the three-dimensional mesh. That is, the encoding device 100 and the decoding device 200 perform encoding and decoding according to encoding methods and decoding methods that correspond to each other.

[0083] The 3D mesh before encoding may also be referred to as an original 3D mesh, and the 3D mesh after decoding may also be referred to as a reconstructed 3D mesh.

[0084] 5 is a block diagram showing an example of the configuration of a coding device 100 according to this embodiment. For example, the coding device 100 includes a vertex information encoder 101, a connection information encoder 102, and an attribute information encoder 103.

[0085] The vertex information encoder 101 is an electrical circuit that encodes vertex information. For example, the vertex information encoder 101 encodes the vertex information into a bitstream according to a format defined for the vertex information.

[0086] The connection information encoder 102 is an electrical circuit that encodes the connection information, for example, the connection information encoder 102 encodes the connection information into a bitstream according to a format defined for the connection information.

[0087] The attribute information encoder 103 is an electric circuit that encodes the attribute information. For example, the attribute information encoder 103 encodes the attribute information into a bit stream in accordance with a format defined for the attribute information.

[0088] The vertex information, connectivity information, and attribute information may be coded using variable-length coding or fixed-length coding, such as Huffman coding or context-adaptive binary arithmetic coding (CABAC).

[0089] The vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be integrated together, or each of the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be further subdivided into multiple components.

[0090] 6 is a block diagram showing another example of the configuration of the encoding device 100 according to this embodiment. For example, the encoding device 100 includes a pre-processor 104 and a post-processor 105 in addition to the configuration shown in FIG.

[0091] The preprocessor 104 is an electrical circuit that performs processing before encoding the vertex information, connectivity information, and attribute information. For example, the preprocessor 104 may perform a conversion process, a separation process, a multiplexing process, or the like on the 3D mesh before encoding. More specifically, for example, the preprocessor 104 may separate the vertex information, connectivity information, and attribute information from the 3D mesh before encoding.

[0092] The post-processor 105 is an electrical circuit that performs processing after the vertex information, connection information, and attribute information are encoded. For example, the post-processor 105 may perform conversion processing, separation processing, multiplexing processing, or the like on the encoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 105 may multiplex the encoded vertex information, connection information, and attribute information into a bitstream. Furthermore, for example, the post-processor 105 may further perform variable-length coding on the encoded vertex information, connection information, and attribute information.

[0093] 7 is a block diagram showing an example of the configuration of a decoding device 200 according to this embodiment. For example, the decoding device 200 includes a vertex information decoder 201, a connection information decoder 202, and an attribute information decoder 203.

[0094] The vertex information decoder 201 is an electrical circuit that decodes vertex information. For example, the vertex information decoder 201 decodes vertex information from a bitstream according to a format defined for the vertex information.

[0095] The connection information decoder 202 is an electrical circuit that decodes the connection information, for example, the connection information decoder 202 decodes the connection information from the bitstream according to a format defined for the connection information.

[0096] The attribute information decoder 203 is an electric circuit that decodes the attribute information. For example, the attribute information decoder 203 decodes the attribute information from the bitstream in accordance with a format defined for the attribute information.

[0097] The vertex information, connection information, and attribute information may be decoded using variable length decoding or fixed length decoding, which may correspond to Huffman coding, context-adaptive binary arithmetic coding (CABAC), or the like.

[0098] The vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be integrated together, or each of the vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be further subdivided into multiple components.

[0099] 8 is a block diagram showing another example of the configuration of the decoding device 200 according to this embodiment. For example, the decoding device 200 includes a pre-processor 204 and a post-processor 205 in addition to the configuration shown in FIG.

[0100] The preprocessor 204 is an electrical circuit that performs processing before decoding the vertex information, connection information, and attribute information. For example, the preprocessor 204 may perform conversion processing, separation processing, multiplexing processing, or the like on the bitstream before decoding the vertex information, connection information, and attribute information.

[0101] More specifically, for example, the preprocessor 204 may separate a sub-bitstream corresponding to vertex information, a sub-bitstream corresponding to connectivity information, and a sub-bitstream corresponding to attribute information from the bitstream. Also, for example, the preprocessor 204 may perform variable-length decoding on the bitstream in advance before decoding the vertex information, connectivity information, and attribute information.

[0102] The post-processor 205 is an electrical circuit that performs processing after the vertex information, connection information, and attribute information are decoded. For example, the post-processor 205 may perform conversion processing, separation processing, multiplexing processing, or the like on the decoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 205 may multiplex the decoded vertex information, connection information, and attribute information onto a three-dimensional mesh.

[0103] <Bitstream> Vertex information, connection information, and attribute information are coded and stored in a bitstream. The relationship between this information and the bitstream is shown below.

[0104] 9 is a conceptual diagram showing an example of the configuration of a bitstream according to this embodiment. In this example, connection information, vertex information, and attribute information are integrated in the bitstream. For example, the connection information, vertex information, and attribute information may be included in a single file.

[0105] Furthermore, multiple portions of this information may be stored sequentially, such as a first portion of connection information, a first portion of vertex information, a first portion of attribute information, a second portion of connection information, a second portion of vertex information, a second portion of attribute information, etc. These multiple portions may correspond to multiple portions that are different in time, multiple portions that are different in space, or multiple different faces.

[0106] Furthermore, the storage order of the connection information, vertex information, and attribute information is not limited to the above example, and a storage order different from the above example may be used.

[0107] 10 is a conceptual diagram showing another example of the configuration of a bitstream according to this embodiment. In this example, a plurality of files are included in the bitstream, and connection information, vertex information, and attribute information are stored in different files. Here, a file containing connection information, a file containing vertex information, and a file containing attribute information are shown, but the storage format is not limited to this example. For example, two types of information among the connection information, vertex information, and attribute information may be included in one file, and the remaining type of information may be included in another file.

[0108] Alternatively, the information may be split and stored in more files. For example, multiple pieces of connectivity information may be stored in multiple files, multiple pieces of vertex information may be stored in multiple files, or multiple pieces of attribute information may be stored in multiple files. These multiple pieces may correspond to multiple temporally different pieces, multiple spatially different pieces, or multiple different faces.

[0109] Furthermore, the storage order of the connection information, vertex information, and attribute information is not limited to the above example, and a storage order different from the above example may be used.

[0110] 11 is a conceptual diagram showing another example of the configuration of a bitstream according to this embodiment. In this example, the bitstream is composed of multiple separable sub-bitstreams, and connection information, vertex information, and attribute information are stored in different sub-bitstreams.

[0111] Here, a sub-bitstream containing connection information, a sub-bitstream containing vertex information, and a sub-bitstream containing attribute information are shown, but the storage format is not limited to this example.

[0112] For example, two types of information among the connection information, vertex information, and attribute information may be included in one sub-bitstream, and the remaining type of information may be included in another sub-bitstream. Specifically, attribute information of a two-dimensional image or the like may be stored in a sub-bitstream that complies with an image coding method, separate from the sub-bitstreams of the connection information and vertex information.

[0113] Also, each sub-bitstream may include multiple files, and multiple pieces of connectivity information may be stored in multiple files, multiple pieces of vertex information may be stored in multiple files, or multiple pieces of attribute information may be stored in multiple files.

[0114] 9, 10, and 11, and a storage order different from the above examples may be used. For example, the vertex information, connection information, and attribute information may be stored in the bitstream in this order. Alternatively, the connection information, connection information, and attribute information may be stored in the bitstream in any of the following orders: connection information, attribute information, and vertex information; vertex information, attribute information, and connection information; attribute information, connection information, and vertex information; or attribute information, vertex information, and connection information.

[0115] Furthermore, each of the connection information, vertex information, and attribute information may be divided into a plurality of data, and the plurality of data may be stored in a cyclical or random order within the bitstream.

[0116] 12 is a block diagram showing a specific example of an encoding / decoding system according to this embodiment. In FIG. 12, the encoding / decoding system includes a three-dimensional data encoding system 110, a three-dimensional data decoding system 210, and an external connector 310.

[0117] The three-dimensional data encoding system 110 includes a controller 111, an input / output processor 112, a three-dimensional data encoder 113, a three-dimensional data generator 115, and a system multiplexer 114. The three-dimensional data decoding system 210 includes a controller 211, an input / output processor 212, a three-dimensional data decoder 213, a system demultiplexer 214, a presenter 215, and a user interface 216.

[0118] In the three-dimensional data encoding system 110, sensor data is input from a sensor terminal to a three-dimensional data generator 115. The three-dimensional data generator 115 generates three-dimensional data, such as point cloud data or mesh data, from the sensor data and inputs it to a three-dimensional data encoder 113.

[0119] For example, the three-dimensional data generator 115 generates vertex information, and generates connection information and attribute information corresponding to the vertex information. The three-dimensional data generator 115 may process the vertex information when generating the connection information and attribute information. For example, the three-dimensional data generator 115 may reduce the amount of data by deleting duplicate vertices, or may transform the vertex information (such as by shifting its position, rotating it, or normalizing it). The three-dimensional data generator 115 may also render the attribute information.

[0120] Furthermore, although the three-dimensional data generator 115 is a component of the three-dimensional data encoding system 110 in FIG. 12, it may be arranged externally and independently of the three-dimensional data encoding system 110.

[0121] The sensor terminal that provides the sensor data for generating the three-dimensional data may be, for example, a moving body such as an automobile, a flying object such as an airplane, a mobile terminal, a camera, etc. Furthermore, a distance sensor such as a LIDAR, a millimeter wave radar, an infrared sensor, or a range finder, a stereo camera, or a combination of multiple monocular cameras may also be used as the sensor terminal.

[0122] The sensor data may be the distance (position) of the object, monocular camera images, stereo camera images, color, reflectance, sensor attitude, orientation, gyro, sensing position (GPS information or altitude), speed, acceleration, sensing time, temperature, air pressure, humidity, or magnetism.

[0123] The three-dimensional data encoder 113 corresponds to the encoding device 100 shown in FIG. 5 and other figures. For example, the three-dimensional data encoder 113 encodes three-dimensional data to generate encoded data. The three-dimensional data encoder 113 also generates control information when encoding the three-dimensional data. The three-dimensional data encoder 113 then inputs the encoded data together with the control information to the system multiplexer 114.

[0124] The encoding method for the three-dimensional data may be an encoding method using geometry or an encoding method using a video codec. Here, the encoding method using geometry may also be referred to as a geometry-based encoding method. The encoding method using a video codec may also be referred to as a video-based encoding method.

[0125] The system multiplexer 114 multiplexes the encoded data and control information input from the 3D data encoder 113 to generate multiplexed data using a specified multiplexing method. The system multiplexer 114 may multiplex other media such as video, audio, subtitles, application data, or document files, or reference time information, along with the encoded data and control information of the 3D data. Furthermore, the system multiplexer 114 may multiplex attribute information related to the sensor data or the 3D data.

[0126] For example, the multiplexed data may have a file format for storage or a packet format for transmission. As these formats, ISOBMFF or a format based on ISOBMFF may be used. Also, MPEG-DASH, MMT, MPEG-2 TS Systems, RTP, or the like may be used.

[0127] The multiplexed data is then output as a transmission signal to the external connector 310 by the input / output processor 112. The multiplexed data may be transmitted as a transmission signal by wire or wirelessly. Alternatively, the multiplexed data is stored in an internal memory or a storage device. The multiplexed data may be transmitted to a cloud server via the Internet or may be stored in an external storage device.

[0128] For example, the transmission or storage of the multiplexed data is performed by a method according to the medium for transmission or storage, such as broadcasting or communication. The communication protocol may be http, ftp, TCP, UDP, IP, or a combination thereof. Furthermore, a pull-type communication method or a push-type communication method may be used.

[0129] For wired transmission, Ethernet (registered trademark), USB, RS-232C, HDMI (registered trademark), coaxial cable, etc. may be used. For wireless transmission, 3GPP (registered trademark), 3G / 4G / 5G defined by IEEE, wireless LAN, Wi-Fi, Bluetooth, or millimeter wave may be used. For broadcasting, for example, DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 may be used.

[0130] The sensor data may be input to the three-dimensional data generator 115 or the system multiplexer 114. The three-dimensional data or encoded data may be output as a transmission signal directly to the external connector 310 via the input / output processor 112. The transmission signal output from the three-dimensional data encoding system 110 is input to the three-dimensional data decoding system 210 via the external connector 310.

[0131] Furthermore, each operation of the three-dimensional data encoding system 110 may be controlled by a controller 111 that executes an application program.

[0132] In the three-dimensional data decoding system 210, a transmission signal is input to an input / output processor 212. The input / output processor 212 decodes multiplexed data having a file format or a packet format from the transmission signal and inputs the multiplexed data to a system demultiplexer 214. The system demultiplexer 214 obtains coded data and control information from the multiplexed data and inputs them to a three-dimensional data decoder 213. The system demultiplexer 214 may extract other media or reference time information from the multiplexed data.

[0133] The three-dimensional data decoder 213 corresponds to the decoding device 200 shown in Fig. 7 etc. For example, the three-dimensional data decoder 213 decodes three-dimensional data from the encoded data based on a predefined encoding method. The three-dimensional data is then presented to the user by the presenter 215.

[0134] Additionally, additional information such as sensor data may be input to the presenter 215. The presenter 215 may present three-dimensional data based on the additional information. Additionally, a user instruction may be input from a user terminal to the user interface 216. Then, the presenter 215 may present three-dimensional data based on the input instruction.

[0135] The input / output processor 212 may acquire the three-dimensional data and the encoded data from the external connector 310 .

[0136] Furthermore, each operation of the three-dimensional data decoding system 210 may be controlled by a controller 211 that executes an application program.

[0137] 13 is a conceptual diagram showing an example of the configuration of point cloud data according to this embodiment. The point cloud data is data of a group of points representing a three-dimensional object.

[0138] Specifically, a point cloud is made up of a plurality of points, and has position information indicating the three-dimensional coordinate position of each point and attribute information indicating the attribute of each point. The position information is also expressed as geometry.

[0139] The type of attribute information may be, for example, color, reflectance, etc. One point may be associated with attribute information of one type, one point may be associated with attribute information of multiple different types, or one point may be associated with attribute information having multiple values ​​for the same type.

[0140] 14 is a conceptual diagram showing an example of a data file of point cloud data according to this embodiment. This example shows a case where there is a one-to-one correspondence between position information items and attribute information items, and shows position information and attribute information for N points that make up the point cloud data. In this example, the position information is information indicating a three-dimensional coordinate position using three axes, x, y, and z, and the attribute information is information indicating a color using RGB. A PLY file or the like can be used as a representative data file for point cloud data.

[0141] 15 is a conceptual diagram showing an example of the configuration of mesh data according to this embodiment. Mesh data is data used in CG (Computer Graphics) and the like, and is three-dimensional mesh data that shows the three-dimensional shape of an object using multiple surfaces. Each surface is also expressed as a polygon, and has a polygonal shape such as a triangle or a rectangle.

[0142] Specifically, a 3D mesh is composed of a plurality of points constituting a point cloud, as well as a plurality of edges and a plurality of faces. Each point is also expressed as a vertex or a position. Each edge corresponds to a line segment connected by two vertices. Each face corresponds to an area surrounded by three or more edges.

[0143] Furthermore, a three-dimensional mesh has position information indicating the three-dimensional coordinate positions of vertices. The position information is also expressed as vertex information or geometry. A three-dimensional mesh also has connection information indicating the relationship between multiple vertices that make up an edge or a face. The connection information is also expressed as connectivity. A three-dimensional mesh also has attribute information indicating the attributes of the vertices, edges, or faces. The attribute information in a three-dimensional mesh is also expressed as texture.

[0144] For example, the attribute information may indicate the color, reflectance, or normal vector for a vertex, edge, or face. The direction of the normal vector may represent the front and back of the face.

[0145] The mesh data may be stored in a data file format such as an object file.

[0146] 16 is a conceptual diagram showing an example of a data file of mesh data according to this embodiment. In this example, the data file includes position information G(1) to G(N) of N vertices that make up the three-dimensional mesh, and attribute information A1(1) to A1(N) of the N vertices. Also, in this example, M pieces of attribute information A2(1) to A2(M) are included. The attribute information items do not need to correspond one-to-one to vertices or faces. Furthermore, attribute information need not exist.

[0147] The connection information is represented by a combination of vertex indices. n[1, 3, 4] indicates a triangular face formed by three vertices, n=1, n=3, and n=4. Also, m[2, 4, 6] indicates that the attribute information of m=2, m=4, and m=6 corresponds to the three vertices, respectively.

[0148] Furthermore, the actual contents of the attribute information may be written in a separate file. A pointer to that content may be associated with a vertex, a face, or the like. For example, attribute information indicating an image for a face may be stored in a two-dimensional attribute map file. The file name of the attribute map and two-dimensional coordinate values ​​in the attribute map may be written in attribute information A2(1) to A2(M). The method of specifying attribute information for a face is not limited to these methods, and any method may be used.

[0149] 17 is a conceptual diagram showing types of three-dimensional data according to this embodiment. Point cloud data and mesh data may represent static objects or dynamic objects. A static object is an object that does not change over time, and a dynamic object is an object that changes over time. A static object may correspond to three-dimensional data for any point in time.

[0150] For example, point cloud data for a given point in time may be referred to as a PCC frame, mesh data for a given point in time may be referred to as a mesh frame, and PCC frames and mesh frames may be simply referred to as frames.

[0151] The area of ​​the object may be limited to a certain range, as in normal video data, or may not be limited, as in map data. The density of points or surfaces may be determined in various ways. Sparse point cloud data or sparse mesh data may be used, or dense point cloud data or dense mesh data may be used.

[0152] Next, encoding and decoding of a point cloud or a three-dimensional mesh will be described. The device, process, or syntax for encoding and decoding vertex information of a three-dimensional mesh in the present disclosure may be applied to encoding and decoding of a point cloud. The device, process, or syntax for encoding and decoding of a point cloud in the present disclosure may be applied to encoding and decoding vertex information of a three-dimensional mesh.

[0153] Furthermore, a device, process, or syntax for encoding and decoding attribute information of a point cloud in the present disclosure may be applied to encoding and decoding connectivity information or attribute information of a three-dimensional mesh.Furthermore, a device, process, or syntax for encoding and decoding connectivity information or attribute information of a three-dimensional mesh in the present disclosure may be applied to encoding and decoding attribute information of a point cloud.

[0154] Furthermore, at least some of the processing may be shared between the encoding and decoding of point cloud data and the encoding and decoding of mesh data, thereby reducing the scale of the circuit and software program.

[0155] 18 is a block diagram showing an example configuration of a three-dimensional data encoder 113 according to this embodiment. In this example, the three-dimensional data encoder 113 includes a vertex information encoder 121, an attribute information encoder 122, a metadata encoder 123, and a multiplexer 124. The vertex information encoder 121, the attribute information encoder 122, and the multiplexer 124 may correspond to the vertex information encoder 101, the attribute information encoder 103, the post-processor 105, etc. in FIG.

[0156] In this example, the three-dimensional data encoder 113 encodes the three-dimensional data according to a geometry-based encoding method, which takes into account the three-dimensional structure. In addition, in the geometry-based encoding method, attribute information is encoded using configuration information obtained in encoding the vertex information.

[0157] Specifically, first, vertex information, attribute information, and metadata included in three-dimensional data generated from sensor data are input to a vertex information encoder 121, an attribute information encoder 122, and a metadata encoder 123, respectively. Here, connectivity information included in the three-dimensional data may be treated in the same way as attribute information. In addition, in the case of point cloud data, position information may be treated as vertex information.

[0158] The vertex information encoder 121 encodes the vertex information into compressed vertex information and outputs the compressed vertex information as encoded data to the multiplexer 124. The vertex information encoder 121 also generates metadata for the compressed vertex information and outputs it to the multiplexer 124. The vertex information encoder 121 also generates configuration information and outputs it to the attribute information encoder 122.

[0159] The attribute information encoder 122 uses the configuration information generated by the vertex information encoder 121 to encode the attribute information into compressed attribute information and outputs the compressed attribute information as encoded data to the multiplexer 124. The attribute information encoder 122 also generates metadata of the compressed attribute information and outputs it to the multiplexer 124.

[0160] The metadata encoder 123 encodes compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The metadata encoded by the metadata encoder 123 may be used to encode vertex information and attribute information.

[0161] The multiplexer 124 multiplexes the compressed vertex information, the compressed vertex information metadata, the compressed attribute information, the compressed attribute information metadata, and the compressed metadata into a bitstream, and then inputs the bitstream to the system layer.

[0162] 19 is a block diagram showing an example configuration of a three-dimensional data decoder 213 according to this embodiment. In this example, the three-dimensional data decoder 213 includes a vertex information decoder 221, an attribute information decoder 222, a metadata decoder 223, and a demultiplexer 224. The vertex information decoder 221, the attribute information decoder 222, and the demultiplexer 224 may correspond to the vertex information decoder 201, the attribute information decoder 203, the preprocessor 204, and the like in FIG.

[0163] In this example, the three-dimensional data decoder 213 decodes three-dimensional data according to a geometry-based encoding method. The three-dimensional structure is taken into consideration in the decoding according to the geometry-based encoding method. Furthermore, in the decoding according to the geometry-based encoding method, attribute information is decoded using configuration information obtained in decoding vertex information.

[0164] Specifically, first, a bitstream is input from the system layer to a demultiplexer 224. The demultiplexer 224 separates compressed vertex information, compressed vertex information metadata, compressed attribute information, compressed attribute information metadata, and compressed metadata from the bitstream. The compressed vertex information and compressed vertex information metadata are input to a vertex information decoder 221. The compressed attribute information and compressed attribute information metadata are input to an attribute information decoder 222. The metadata is input to a metadata decoder 223.

[0165] The vertex information decoder 221 decodes vertex information from the compressed vertex information using metadata of the compressed vertex information. The vertex information decoder 221 also generates configuration information and outputs it to the attribute information decoder 222. The attribute information decoder 222 decodes attribute information from the compressed attribute information using the configuration information generated by the vertex information decoder 221 and the metadata of the compressed attribute information. The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 may be used to decode the vertex information and the attribute information.

[0166] Thereafter, the vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. Note that, for example, this metadata is metadata of the vertex information and attribute information, and can be used in an application program.

[0167] 20 is a block diagram showing another example configuration of the three-dimensional data encoder 113 according to the present embodiment. In this example, the three-dimensional data encoder 113 includes a vertex image generator 131, an attribute image generator 132, a metadata generator 133, a video encoder 134, a metadata encoder 123, and a multiplexer 124. The vertex image generator 131, the attribute image generator 132, and the video encoder 134 may correspond to the vertex information encoder 101 and the attribute information encoder 103 in FIG. 6 , etc.

[0168] In this example, the 3D data encoder 113 encodes the 3D data according to a video-based encoding method. In encoding according to the video-based encoding method, multiple 2D images are generated from the 3D data, and the multiple 2D images are encoded according to a video encoding method. Here, the video encoding method may be High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or the like.

[0169] Specifically, first, vertex information and attribute information included in three-dimensional data generated from sensor data are input to a metadata generator 133. The vertex information and attribute information are then input to a vertex image generator 131 and an attribute image generator 132, respectively. The metadata included in the three-dimensional data is then input to a metadata encoder 123. Here, connectivity information included in the three-dimensional data may be treated in the same way as attribute information. In the case of point cloud data, position information may be treated as vertex information.

[0170] The metadata generator 133 generates map information of a plurality of two-dimensional images from the vertex information and attribute information, and inputs the map information to the vertex image generator 131, the attribute image generator 132, and the metadata encoder 123.

[0171] The vertex image generator 131 generates a vertex image based on the vertex information and map information, and inputs the generated image to the video encoder 134. The attribute image generator 132 generates an attribute image based on the attribute information and map information, and inputs the generated image to the video encoder 134.

[0172] The video encoder 134 encodes the vertex images and attribute images into compressed vertex information and compressed attribute information, respectively, in accordance with a video encoding method, and outputs the compressed vertex information and compressed attribute information as encoded data to the multiplexer 124. The video encoder 134 also generates metadata for the compressed vertex information and metadata for the compressed attribute information, and outputs them to the multiplexer 124.

[0173] The metadata encoder 123 encodes the compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The compressible metadata includes map information. The metadata encoded by the metadata encoder 123 may also be used to encode vertex information and attribute information.

[0174] The multiplexer 124 multiplexes the compressed vertex information, the compressed vertex information metadata, the compressed attribute information, the compressed attribute information metadata, and the compressed metadata into a bitstream, and then inputs the bitstream to the system layer.

[0175] 21 is a block diagram showing another example configuration of the 3D data decoder 213 according to this embodiment. In this example, the 3D data decoder 213 includes a vertex information generator 231, an attribute information generator 232, a video decoder 234, a metadata decoder 223, and a demultiplexer 224. The vertex information generator 231, the attribute information generator 232, and the video decoder 234 may correspond to the vertex information decoder 201 and the attribute information decoder 203 in FIG. 8, etc.

[0176] In this example, the 3D data decoder 213 decodes the 3D data according to a video-based coding method. In the decoding according to the video-based coding method, a plurality of 2D images are decoded according to a video coding method, and 3D data is generated from the plurality of 2D images. Here, the video coding method may be High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or the like.

[0177] Specifically, first, a bitstream is input from the system layer to the demultiplexer 224. The demultiplexer 224 separates compressed vertex information, compressed vertex information metadata, compressed attribute information, compressed attribute information metadata, and compressed metadata from the bitstream. The compressed vertex information, compressed vertex information metadata, compressed attribute information, and compressed attribute information metadata are input to the video decoder 234. The compressed metadata is input to the metadata decoder 223.

[0178] The video decoder 234 decodes the vertex images in accordance with the video encoding method. At this time, the video decoder 234 decodes the vertex images from the compressed vertex information using the metadata of the compressed vertex information. Then, the video decoder 234 inputs the vertex images to the vertex information generator 231. The video decoder 234 also decodes the attribute images in accordance with the video encoding method. At this time, the video decoder 234 decodes the attribute images from the compressed attribute information using the metadata of the compressed attribute information. Then, the video decoder 234 inputs the attribute images to the attribute information generator 232.

[0179] The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 includes map information used to generate vertex information and attribute information. The metadata decoded by the metadata decoder 223 may also be used to decode vertex images and attribute images.

[0180] The vertex information generator 231 reproduces vertex information from the vertex image in accordance with the map information included in the metadata decoded by the metadata decoder 223. The attribute information generator 232 reproduces attribute information from the attribute image in accordance with the map information included in the metadata decoded by the metadata decoder 223.

[0181] Thereafter, the vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. Note that, for example, this metadata is metadata of the vertex information and attribute information, and can be used in an application program.

[0182] Fig. 22 is a conceptual diagram showing a specific example of encoding processing according to this embodiment. Fig. 22 shows a three-dimensional data encoder 113 and a description encoder 148. In this example, the three-dimensional data encoder 113 includes a two-dimensional data encoder 141 and a mesh data encoder 142. The two-dimensional data encoder 141 includes a texture encoder 143. The mesh data encoder 142 includes a vertex information encoder 144 and a connection information encoder 145.

[0183] The vertex information encoder 144, the connection information encoder 145, and the texture encoder 143 may correspond to the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 in FIG.

[0184] For example, the two-dimensional data encoder 141 operates as a texture encoder 143 and generates a texture file by encoding the texture corresponding to the attribute information as two-dimensional data according to an image encoding method or a video encoding method.

[0185] The mesh data encoder 142 also operates as a vertex information encoder 144 and a connectivity information encoder 145, and generates a mesh file by encoding the vertex information and connectivity information. The mesh data encoder 142 may further encode mapping information for textures. The encoded mapping information may then be included in the mesh file.

[0186] The description encoder 148 also generates a description file by encoding a description corresponding to metadata such as text data. The description encoder 148 may encode the description at the system layer. For example, the description encoder 148 may be included in the system multiplexer 114 of FIG. 12 .

[0187] The above operations generate a bitstream containing texture files, mesh files, and description files, which may be multiplexed into the bitstream in file formats such as glTF (Graphics Language Transmission Format) or USD (Universal Scene Description).

[0188] The three-dimensional data encoder 113 may include two mesh data encoders as the mesh data encoder 142. For example, one mesh data encoder encodes vertex information and connectivity information of a static three-dimensional mesh, and the other mesh data encoder encodes vertex information and connectivity information of a dynamic three-dimensional mesh.

[0189] Correspondingly, two mesh files may then be included in the bitstream, for example one mesh file corresponding to a static 3D mesh and another mesh file corresponding to a dynamic 3D mesh.

[0190] Furthermore, the static three-dimensional mesh may be a three-dimensional mesh of an intraframe coded using intraprediction, and the dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe coded using interprediction. Furthermore, information on the dynamic three-dimensional mesh may be differential information between vertex information or connectivity information of the three-dimensional mesh of an intraframe and vertex information or connectivity information of the three-dimensional mesh of an interframe.

[0191] Fig. 23 is a conceptual diagram showing a specific example of the decoding process according to this embodiment. Fig. 23 shows a three-dimensional data decoder 213, a description decoder 248, and a renderer 247. In this example, the three-dimensional data decoder 213 includes a two-dimensional data decoder 241, a mesh data decoder 242, and a mesh reconstructor 246. The two-dimensional data decoder 241 includes a texture decoder 243. The mesh data decoder 242 includes a vertex information decoder 244 and a connectivity information decoder 245.

[0192] The vertex information decoder 244, the connection information decoder 245, the texture decoder 243, and the mesh reconstructor 246 may correspond to the vertex information decoder 201, the connection information decoder 202, the attribute information decoder 203, and the post-processor 205 in Fig. 8. The presenter 247 may correspond to the presenter 215 in Fig. 12.

[0193] For example, the two-dimensional data decoder 241 operates as a texture decoder 243, and decodes the texture corresponding to the attribute information from the texture file as two-dimensional data in accordance with an image coding method or a video coding method.

[0194] The mesh data decoder 242 also operates as a vertex information decoder 244 and a connectivity information decoder 245 to decode vertex information and connectivity information from the mesh file. The mesh data decoder 242 may further decode mapping information for textures from the mesh file.

[0195] The description decoder 248 also decodes descriptions corresponding to metadata such as text data from the description file. The description decoder 248 may decode the descriptions at the system layer. For example, the description decoder 248 may be included in the system demultiplexer 214 of FIG. 12 .

[0196] The mesh reconstructor 246 reconstructs a 3D mesh from the vertex information, connectivity information, and textures according to the description. The renderer 247 renders and outputs the 3D mesh according to the description.

[0197] Through the above operations, a 3D mesh is reconstructed and output from a bitstream containing a texture file, a mesh file, and a description file.

[0198] The three-dimensional data decoder 213 may include two mesh data decoders as the mesh data decoder 242. For example, one mesh data decoder decodes vertex information and connectivity information of a static three-dimensional mesh, and the other mesh data decoder decodes vertex information and connectivity information of a dynamic three-dimensional mesh.

[0199] Correspondingly, two mesh files may then be included in the bitstream, for example one mesh file corresponding to a static 3D mesh and another mesh file corresponding to a dynamic 3D mesh.

[0200] Furthermore, the static three-dimensional mesh may be a three-dimensional mesh of an intraframe coded using intraprediction, and the dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe coded using interprediction. Furthermore, information on the dynamic three-dimensional mesh may be differential information between vertex information or connectivity information of the three-dimensional mesh of an intraframe and vertex information or connectivity information of the three-dimensional mesh of an interframe.

[0201] A dynamic 3D mesh coding method is sometimes called DMC (Dynamic Mesh Coding), and a video-based dynamic 3D mesh coding method is sometimes called V-DMC (Video-based Dynamic Mesh Coding).

[0202] The point cloud encoding method is sometimes called PCC (Point Cloud Compression). The point cloud video-based encoding method is sometimes called V-PCC (Video-based Point Cloud Compression). The point cloud geometry-based encoding method is sometimes called G-PCC (Geometry-based Point Cloud Compression).

[0203] <Implementation Example> Fig. 24 is a block diagram showing an implementation example of the encoding device 100 according to this embodiment. The encoding device 100 includes a circuit 151 and a memory 152. For example, multiple components of the encoding device 100 shown in Fig. 5 etc. are implemented by the circuit 151 and memory 152 shown in Fig. 24.

[0204] The circuit 151 is a circuit that performs information processing and is a circuit that can access the memory 152. For example, the circuit 151 is a dedicated or general-purpose electric circuit that encodes a three-dimensional mesh. The circuit 151 may be a processor such as a CPU. Alternatively, the circuit 151 may be a collection of multiple electric circuits.

[0205] The memory 152 is a dedicated or general-purpose memory that stores information used by the circuit 151 to encode the three-dimensional mesh. The memory 152 may be an electric circuit and may be connected to the circuit 151. The memory 152 may also be included in the circuit 151. The memory 152 may also be a collection of multiple electric circuits. The memory 152 may also be a magnetic disk, an optical disk, or the like, and may also be expressed as a storage, a recording medium, or the like. The memory 152 may also be a non-volatile memory or a volatile memory.

[0206] For example, the memory 152 may store a three-dimensional mesh or a bitstream, or may store a program for the circuit 151 to encode the three-dimensional mesh.

[0207] Note that the encoding device 100 does not necessarily have to implement all of the components shown in Figure 5 and the like, and does not necessarily have to perform all of the processes shown here. Some of the components shown in Figure 5 and the like may be included in another device, and some of the processes shown here may be executed by another device. Furthermore, the encoding device 100 may implement any combination of the components of the present disclosure, and may perform any combination of the processes of the present disclosure.

[0208] Fig. 25 is a block diagram showing an example implementation of a decoding device 200 according to this embodiment. The decoding device 200 includes a circuit 251 and a memory 252. For example, multiple components of the decoding device 200 shown in Fig. 7 and other figures are implemented by the circuit 251 and memory 252 shown in Fig. 25.

[0209] The circuit 251 is a circuit that performs information processing and is a circuit that can access the memory 252. For example, the circuit 251 is a dedicated or general-purpose electric circuit that decodes a three-dimensional mesh. The circuit 251 may be a processor such as a CPU. Alternatively, the circuit 251 may be a collection of multiple electric circuits.

[0210] The memory 252 is a dedicated or general-purpose memory that stores information for the circuit 251 to decode the 3D mesh. The memory 252 may be an electric circuit and may be connected to the circuit 251. The memory 252 may also be included in the circuit 251. The memory 252 may also be a collection of multiple electric circuits. The memory 252 may also be a magnetic disk, an optical disk, or the like, and may also be expressed as a storage, a recording medium, or the like. The memory 252 may also be a non-volatile memory or a volatile memory.

[0211] For example, the memory 252 may store a three-dimensional mesh or a bitstream, or may store a program for the circuit 251 to decode the three-dimensional mesh.

[0212] Note that the decoding device 200 does not necessarily have to implement all of the components shown in Figure 7 and the like, and does not necessarily have to perform all of the processes shown here. Some of the components shown in Figure 7 and the like may be included in another device, and some of the processes shown here may be executed by another device. Furthermore, the decoding device 200 may implement any combination of the components of the present disclosure, and may perform any combination of the processes of the present disclosure.

[0213] The encoding method and the decoding method including the steps performed by each component of the encoding device 100 and the decoding device 200 of the present disclosure may be executed by any device or system. For example, part or all of the encoding method and the decoding method may be executed by a computer including a processor, a memory, an input / output circuit, etc. In this case, the encoding method and the decoding method may be executed by the computer executing a program for causing the computer to execute the encoding method and the decoding method.

[0214] Alternatively, the program or the bitstream may be recorded on a non-transitory computer-readable recording medium such as a CD-ROM.

[0215] An example of a program may be a bitstream. For example, a bitstream including an encoded three-dimensional mesh includes syntax elements for causing the decoding device 200 to decode the three-dimensional mesh. The bitstream then causes the decoding device 200 to decode the three-dimensional mesh according to the syntax elements included in the bitstream. Thus, the bitstream may play a role similar to that of a program.

[0216] The bitstream may be an encoded bitstream containing the encoded 3D mesh, or may be a multiplexed bitstream containing the encoded 3D mesh and other information.

[0217] Furthermore, each component of the encoding device 100 and the decoding device 200 may be configured with dedicated hardware, general-purpose hardware that executes the above-mentioned programs, or a combination of these. The general-purpose hardware may be configured with a memory in which the programs are recorded and a general-purpose processor that reads and executes the programs from the memory. Here, the memory may be a semiconductor memory or a hard disk, and the general-purpose processor may be a CPU.

[0218] Furthermore, the dedicated hardware may be configured with a memory, a dedicated processor, etc. For example, the dedicated processor may execute the encoding method and the decoding method by referring to a memory for recording data.

[0219] Furthermore, as described above, each component of the encoding device 100 and the decoding device 200 may be an electric circuit. These electric circuits may form a single electric circuit as a whole, or each may be a separate electric circuit. Furthermore, these electric circuits may correspond to dedicated hardware, or may correspond to general-purpose hardware that executes the above-mentioned programs, etc. Furthermore, the encoding device 100 and the decoding device 200 may be implemented as an integrated circuit.

[0220] Furthermore, the encoding device 100 may be a transmitting device that transmits the three-dimensional mesh, and the decoding device 200 may be a receiving device that receives the three-dimensional mesh.

[0221] Displacement Encoding and Decoding The following terminology is used here by way of example:

[0222] (1) Image An image is a data unit made up of a set of pixels, and includes a picture or a block smaller than a picture. Images include both moving images and still images.

[0223] (2) Picture A picture is a unit of image processing that is made up of a set of pixels, and is also called a frame or field.

[0224] (3) Block A block is a processing unit consisting of a specific number of pixels. The term shown in the following example is also used for a block. The shape of a block is not particularly limited. A block may be, for example, a rectangular shape of M×N pixels or a square shape of M×M pixels. A block may also be a triangular shape, a circular shape, or another shape. Examples of blocks are as follows:

[0225] Slice, tile, or brick CTU, superblock, or basic division unit VPDU, processing division unit for hardware CU, processing block unit, prediction block unit (PU), or orthogonal transform block unit (TU) Sub-block

[0226] (4) Pixel or Sample A pixel or sample is the smallest point of an image, in other words, the smallest unit. Pixels or samples include not only pixels at integer positions, but also pixels at sub-pixel positions generated based on pixels at integer positions.

[0227] (5) Pixel Value or Sample Value: A pixel value or sample value is a unique value of a pixel. The pixel value or sample value may include a luma value, a chroma value, or an RGB gradation level, and may also include a depth value or a binary value of 0 or 1.

[0228] (6) Flags A flag indicates one or more bits. A flag is, for example, a parameter or index represented by two or more bits. A flag may indicate not only a value represented by a binary number, but also a value represented by a number other than a binary number.

[0229] (7) Signal: A signal is something that is symbolized or coded to transmit information. A signal includes a discrete digital signal or a continuous analog signal.

[0230] (8) Stream or Bit Stream A stream or bit stream is a digital data sequence that indicates the flow of digital data. A stream or bit stream may be a single stream, or may be configured to include multiple streams with multiple layers. A stream or bit stream may be transmitted by serial communication using a single transmission path, or may be transmitted by packet communication using multiple transmission paths.

[0231] (9) Difference: For scalar quantities, the difference can include simple difference (x-y) and difference calculations. The difference can include absolute difference (|x-y|), squared difference (x^2-y^2), square root difference (√(x-y)), weighted difference (ax-by, where a and b are constants), or offset difference (x-y+a, where a is an offset).

[0232] (10) Sum: For scalar quantities, sums can include simple sum (x + y) and addition calculations. The sum can also include absolute sum (|x + y|), sum of squares (x^2 + y^2), square root of the sum (√(x + y)), weighted sum (ax + by, where a and b are constants), or offset sum (x + y + a, where a is an offset).

[0233] (11) "Based on" The expression "based on something" means that something other than that "something" may be taken into consideration. Also, "based on" can be used both when a direct result is obtained and when a result is obtained through an intermediate result.

[0234] (12) "Used" or "Using" The phrases "something was used" or "used something" mean that something other than the "something" may be taken into consideration. The phrases "used" or "used" may be used both in cases where a direct result is obtained and in cases where a result is obtained via an intermediate result.

[0235] (13) Prohibition "Prohibit" can be rephrased as "not permitted." Also, "not prohibited / prohibited" or "permitted / permitted" does not necessarily mean "obligation."

[0236] (14) "Restriction" or "Limitation" "Restriction" or "Limitation" can be rephrased as "not permitted / not allowed" or "not permitted / permitted." Furthermore, "prohibited / not prohibited" or "not permitted / permitted" does not necessarily mean "obligation." Furthermore, what is prohibited quantitatively or qualitatively may be either partial or total.

[0237] (15) Chroma The term chroma is an adjective, represented by the symbols Cb or Cr, that indicates that a sample array or a single sample represents one of the two color difference signals associated with a primary color. The term chroma is sometimes used instead of the term chrominance.

[0238] (16) Luma The term luma is an adjective, denoted by the symbols or subscripts Y or L, that indicates that a sample array or a single sample represents a monochrome signal for a primary color. The term luma is sometimes used instead of the term luminance.

[0239] The encoding / decoding system of this embodiment will be described below.

[0240] A typical three-dimensional model (also called a 3D model) digitally represents an object so that a user can explore the model using zoom, pan, and rotation in all three dimensions while it is rendered over time. One way to construct such a representation is to build a 3D mesh using triangles. The model stores the positions of the triangle vertices, their connectivity to each other, and their associated attributes (such as normals or UV patches).

[0241] Storing all this information in uncompressed form requires a very large storage space and therefore a very large bandwidth for transmission. The triangles that form the mesh often have repeating patterns and similar properties, especially in temporal and spatial neighborhoods. These repetitions can be exploited to develop efficient encoding and decoding methods for storage and transmission. One such encoding and decoding method is Video-based Dynamic Mesh Coding (V-DMC).

[0242] 26 is a block diagram showing another example of the configuration of the encoding / decoding system according to this embodiment. As shown in FIG. 26, the encoding / decoding system includes an encoding device 100 and a decoding device 200.

[0243] The encoding / decoding system accepts input three-dimensional meshes (also called 3D meshes) in the form of three-dimensional coordinates of vertices (vertex information), connectivity (connection information) and associated attributes (attribute information), which may include texture maps as well as geometry.

[0244] The encoding device 100 takes an input 3D mesh (also referred to as an input 3D mesh or input mesh) in the form of 3D coordinates of vertices, connectivity, and associated attributes. The encoding device 100 encodes all associated information into a stream. The stream may consist of a single bitstream or multiple bitstreams.

[0245] The network 300 transmits the stream generated by the encoding device 100 to the decoding device 200. The network 300 may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof. Furthermore, the network 300 is not necessarily limited to a two-way communication network, but may also be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Instead of the network 300, a recording medium such as a digital versatile disc (DVD) or a blue-ray disc (BD) on which a stream is recorded may be used.

[0246] The stream is transmitted to a decoding device 200 via a network 300. The decoding device 200 decodes the bitstream and generates a 3D mesh using the 3D coordinates, connectivity, and associated attributes of the decoded vertices. The decoding device 200 outputs the generated 3D mesh (also referred to as an output 3D mesh or output mesh).

[0247] FIG. 27 is a diagram showing another example of the configuration of the encoding device 100.

[0248] As shown in FIG. 27, the encoding device 100 includes a preprocessor 1103 and a compressor 1106 .

[0249] The encoding device 100 reads an input mesh 1101 and an attribute map 1102 and passes them to a preprocessor 1103. The preprocessor 1103 processes the input mesh to extract a base mesh 1104 and displacement data 1105. The attribute map 1102, along with the extracted base mesh 1104 and displacement data 1105, are passed to a compressor 1106.

[0250] The compressor 1106 also compresses the base mesh 1104, the displacement data 1105, and the attribute map 1102 to generate a bitstream 1107. The compressor 1106 can transmit additional information to the decoding device 200 by further including metadata 1108 in the bitstream 1107.

[0251] FIG. 28 is a diagram showing another example of the configuration of the decoding device 200.

[0252] As shown in FIG. 28, the decoding device 200 includes a decompressor 2102 and a post-processor 2106 .

[0253] The decoding device 200 reads a bitstream 2101 and passes it to a decompressor 2102. The decompressor 2102 decompresses a base mesh 2103, displacement data 2104, and an attribute map 2108 from the bitstream 2101 and passes them to a post-processor 2106. An example of the displacement data 2104 is a displacement vector.

[0254] The post-processor 2106 also processes the base mesh 2103 according to the displacement data 2104 and the attribute map 2108 to generate an output mesh 2107. The post-processor 2106 may further use information from the metadata 2105 to generate the output mesh 2107.

[0255] FIG. 29 is a block diagram showing yet another example configuration of the encoding device 100 according to this embodiment.

[0256] In this example, the encoding device 100 comprises a volumetric capturer 511, a projector 512, a base mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and optionally one or more other type encoders 516.

[0257] The volumetric capturer 511 captures content and outputs the captured content to the projector 512 .

[0258] The projector 512 projects the content onto an input mesh (a 3D mesh frame) that includes geometry coordinates (vertex coordinates indicating the positions of the vertices), texture coordinates, and connectivity (connectivity information). The data is output to a base mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and optionally one or more other type encoders 516. Each encoder compresses the data into a bitstream.

[0259] FIG. 30 is a block diagram showing yet another example configuration of the decoding device 200 according to this embodiment.

[0260] In this example, the decoding device 200 comprises a base mesh decoder 613 , a displacement decoder 614 , an attribute decoder 615 , one or more other type decoders 616 , and a 3D reconstructor 617 .

[0261] The bitstream is sent to a base mesh decoder 613, a displacement decoder 614, an attribute decoder 615, and optionally one or more other type decoders 616. These decoders decode the bitstream to generate data (decoded data) including geometry coordinates, texture coordinates, connectivity, etc. The decoded data is then sent to a 3D reconstructor 617, which reconstructs an output mesh (a 3D mesh frame).

[0262] The encoding process performed by the encoding device 100 will be described in detail below.

[0263] Fig. 31 is a flow diagram showing the processing of the encoding device 100. Fig. 32 is an explanatory diagram conceptually showing the encoding of mesh frames. The processing of the encoding device 100 will be described with reference to Figs. 31 and 32.

[0264] In step S101, the encoding device 100 reads a 3D mesh frame, which is an input mesh frame, and its attributes. The input mesh frame is a mesh frame input to the encoding device 100. An example of the 3D mesh frame that is an input mesh frame is shown as mesh frame 1301 (see FIG. 32 ).

[0265] In step S102, the encoding device 100 performs a decimation process on the input mesh frame read in step S101 to generate a base mesh frame having fewer vertices than the input mesh frame. The base mesh frame generated by decimating the mesh frame 1301 is shown as a base mesh frame 1302 (see FIG. 32).

[0266] In step S103, the encoding device 100 calculates displacement information that the decoding device 200 uses to reconstruct a mesh frame. The displacement information corresponds to a displacement vector directed from a vertex of the base mesh frame generated in step S102 to a vertex of the input mesh frame. One method for calculating the displacement information is to subtract the coordinates of the vertex of the base mesh frame from the coordinates of the vertex of the input mesh frame. The displacement information calculated from the mesh frame 1301 and the base mesh frame 1302 is shown as displacement information 1303 (see FIG. 32). The displacement information 1303 is in vector format, in other words, expressed as a displacement vector.

[0267] In step S104, the encoding device 100 encodes the base mesh frame generated in step S102, the displacement information generated in step S103, and the attributes of the input mesh frame into a bitstream (corresponding to a compressed bitstream). An example of the bitstream is shown as bitstream 1304 (see FIG. 32).

[0268] Specifically, the bitstream 1304 includes vertex coordinates and connectivity information for vertices A, C, E, and F, displacement information, a video bitstream including texture data, and a compressed attribute map (see FIG. 32). The displacement information includes displacement information for displacing vertices based on vertex coordinates obtained from the subdivided base mesh frame. The compressed attribute map is texture coordinates for applying texture data to a mesh frame reconstructed using the base mesh frame and the displacement information.

[0269] The decoding process performed by the decoding device 200 will be described in detail below.

[0270] Fig. 33 is a flow diagram showing the processing of the decoding device 200. Fig. 34 is an explanatory diagram conceptually showing the decoding of a mesh frame (3D mesh). The processing of the decoding device 200 will be described with reference to Figs. 33 and 34.

[0271] In step S201, the decoding device 200 decodes a base mesh frame and attributes from a bitstream (corresponding to a compressed bitstream). An example of the decoded base mesh frame (corresponding to a decoded base mesh frame) is shown as a decoded base mesh frame 2301 (see FIG. 34).

[0272] In step S202, the decoding device 200 generates subdivided vertices by performing a subdivision process on the base mesh frame decoded in step S201. An example of a base mesh frame (mesh frame) including subdivided vertices is shown as a base mesh frame 2302 (see FIG. 34).

[0273] In step S203, the decoding device 200 decodes the disparity information from the bitstream (corresponding to the compressed bitstream). An example of the decoded disparity information is shown as disparity information 2303 (see FIG. 34). The disparity information 2303 is in vector format, in other words, expressed as a disparity vector.

[0274] In step S204, the decoding device 200 reconstructs the shape of the mesh frame by moving the vertices of the base mesh frame, including the subdivided vertices, to new positions using the displacement information, and then restores the mesh frame by applying attribute information. An example of the attribute is texture. An example of the reconstructed mesh frame is shown as mesh frame 2304 (see FIG. 34 ).

[0275] FIG. 35 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.

[0276] FIG. 35 shows an example of a block diagram of a general intra-decoding system.

[0277] The decoding device shown in FIG. 35 comprises a demultiplexer 1231, a switch 1232, a static mesh decoder 1233, a mesh buffer 1234, a motion decoder 1235, a base mesh reconstructor 1236, an inverse quantizer 1237, a video decoder 1238, an image unpacker 1239, an inverse quantizer 1240, an inverse wavelet transformer 1241, a reconstructor 1242, a video decoder 1243, and a color converter 1244.

[0278] The demultiplexer 1231 receives the compressed bitstream and separates it into compressed data for the base mesh, video containing displacement data (also called displacement bitstream), and video containing attribute data (also called attribute bitstream). The compressed data for the base mesh is passed to a switch 1232. The switch 1232 determines whether to perform intra-decoding or inter-decoding based on parameters in the bitstream.

[0279] If an intra-decoding process is selected, the bitstream is passed to a static mesh decoder 1233, which generates a quantized base mesh. The static mesh decoder 1233 is, for example, a decoder that uses an edge breaker algorithm to decode 3D mesh data. The static mesh decoder 1233 generates a quantized base mesh from the bitstream. The quantized base mesh generated by the static mesh decoder 1233 is stored in a mesh buffer 1234 for reference when an inter-decoding process is selected.

[0280] If inter-decoding is selected, switch 1232 passes compressed data for the base mesh to motion decoder 1235. Motion decoder 1235 receives a previously decoded quantized base mesh and decodes motion data representing the differences in vertex coordinates between the quantized base mesh stored in mesh buffer 1234 and the current quantized base mesh. The motion data and the quantized base mesh stored in mesh buffer 1234 are used by base mesh reconstructor 1236 to reconstruct the current quantized base mesh. The quantized base mesh resulting from either inter-decoding or intra-decoding is passed to inverse quantizer 1237 to obtain a decoded base mesh.

[0281] The video containing the displacement data is passed to a video decoder 1238, since the bitstream contains the displacement data in an image format with two chroma information and one luma information. The video decoder 1238 decodes the data using a video frame decompression method. Alternatively, the displacement data can be decoded using an arithmetic decoder. This decompressed data is passed to an image unpacker 1239, which extracts wavelet coefficients associated with each vertex from the image-format decompressed data. An inverse quantizer 1240 dequantizes the quantized wavelet coefficients into the three components associated with each vertex. An inverse wavelet transformer 1241 inversely transforms the result to finally obtain decoded displacement data. The decoded displacement data and the decoded base mesh are passed to a reconstructor 1242, which performs edge refinement on the decoded base mesh and displaces the vertices using the decoded displacement data to obtain a decoded mesh.

[0282] The video containing the attribute data is passed to another video decoder 1243 to obtain a decoded attribute bitstream, which is further processed in a color converter 1244 for color space and color format conversion to obtain a decoded attribute map.

[0283] FIG. 36 is a block diagram showing an example of the configuration of an encoding device according to this embodiment.

[0284] First, the encoding device obtains the base mesh bitstream, the displacement bitstream, and the attribute bitstream resulting from the 3D mesh preprocessing step.

[0285] The encoding device shown in FIG. 36 comprises a quantizer 1261, a switch 1262, a static mesh encoder 1263, a mesh buffer 1264, a motion encoder 1265, a base mesh reconstructor 1266, a displacement data updater 1267, a wavelet transformer 1268, a quantizer 1269, an image packer 1270, a video encoder 1271, a color converter 1272, and a video encoder 1273.

[0286] The base mesh (specifically, the position information of the multiple vertices that make up the base mesh) is first quantized by a quantizer 1261. The quantized base mesh (base mesh data) is output to a switch 1262, which determines whether intra-coding or inter-coding is to be performed. If intra-coding is selected, the quantized base mesh (base mesh bitstream) is output to a static mesh encoder 1263, which generates a quantized base mesh. An example of the static mesh encoder 1263 is an encoder that uses an edge breaker algorithm to encode 3D mesh data. This encoded (quantized) base mesh is stored in a mesh buffer 1264 for reference when inter-coding is selected. If inter-coding is selected, the switch 1262 outputs compressed data related to the base mesh to a motion encoder 1265. The motion data and static 3D mesh in the mesh buffer 1264 are used by a base mesh reconstructor 1266 to reconstruct the currently quantized base mesh.

[0287] The displacement data is output to a displacement data updater 1267, where it is updated based on the quantized static base mesh or the reconstructed inter-coded base mesh. Next, a wavelet transformer 1268 performs a transformation process, followed by quantization in a quantizer 1269. The quantized displacement data is packed into an image in an image packer 1270, and finally encoded in a video coder 1271. The encoded displacement data is output to a multiplexer 1274.

[0288] The attribute information (e.g., attribute map) is output to a color converter 1272 for color space and color format conversion, and the converted attribute information is coded by a video coder 1273 and output to a multiplexer 1274.

[0289] The multiplexer 1274 acquires data related to the encoded base mesh (compressed data related to the base mesh), video data including encoded displacement data, and video data including attribute information such as an encoded attribute map, and generates a bitstream (compressed bitstream) including these acquired data. The generated compressed bitstream is output to, for example, a decoding device.

[0290] FIG. 37 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.

[0291] FIG. 37 illustrates an example of a reconstructor that obtains a decoded 3D mesh 1256 from a decoded base mesh 1251 and decoded displacement data 1254.

[0292] The decoded base mesh 1251 is passed to a subdivider 1252 .

[0293] The subdivision unit 1252 subdivides any two connected vertices in the entire 3D mesh by adding a new vertex between them. This process can be repeated several times to include vertices created in previous subdivision steps to generate a predefined number of vertices. Each subdivision iteration across the 3D mesh generates a new level of detail (LoD). The subdivided mesh 1253 and the decoded displacement data 1254 are passed to a displacer 1255. The displacer 1255 generates a decoded 3D mesh 1256 by moving each vertex to a new position according to the corresponding displacement data.

[0294] The subdivision is described below and is performed, for example, by subdivision unit 1252.

[0295] FIG. 38 is an explanatory diagram showing an example of subdivision.

[0296] The base mesh shown in FIG. 38(a) includes vertices A, B, and C and connectivity information indicating their connectivity.

[0297] 38(b) shows a mesh generated by the first subdivision, in other words, the mesh after the first subdivision. In the first subdivision, the subdivider generates vertices D, E, and F and connectivity information indicating their connectivity. The mesh generated by the subdivider is also referred to as LoD1 or first LoD.

[0298] Vertex D of the mesh after the first subdivision is a vertex generated by subdivision based on vertices A and B. Similarly, vertex F is a vertex generated by subdivision based on vertices B and C. Vertex E is a vertex generated by subdivision based on vertices A and C.

[0299] As an example, vertex D may be the midpoint of line segment AB (in other words, side AB) connecting vertices A and B that were the basis for its generation. Similarly, vertex E may be the midpoint of line segment AC. Vertex F may be the midpoint of line segment BC.

[0300] 38(c) shows the mesh generated by the second subdivision, i.e., the mesh after the second subdivision. In the second subdivision, the subdivider generates vertices G, H, I, J, K, L, M, N, and O and connectivity information indicating their connectivity. The mesh generated by the subdivider is also called LoD2 or second LoD.

[0301] Vertex G of the mesh after the second subdivision is a vertex generated by subdivision based on vertices A and D. Similarly, vertex H is a vertex generated by subdivision based on vertices A and E. Vertex I is a vertex generated by subdivision based on vertices B and D. Vertex J is a vertex generated by subdivision based on vertices D and F. Vertex K is a vertex generated by subdivision based on vertices E and F. Vertex L is a vertex generated by subdivision based on vertices C and E. Vertex M is a vertex generated by subdivision based on vertices B and F. Vertex N is a vertex generated by subdivision based on vertices C and F. Vertex O is a vertex generated by subdivision based on vertices D and E.

[0302] As an example, vertex G may be the midpoint of line segment AD (in other words, side AD) connecting vertices A and D, which were the source of its generation. Similarly, vertex H may be the midpoint of line segment AE. vertex I may be the midpoint of line segment BD. vertex J may be the midpoint of line segment DF. vertex K may be the midpoint of line segment EF. vertex L may be the midpoint of line segment CE. vertex M may be the midpoint of line segment BF. vertex N may be the midpoint of line segment CF. vertex O may be the midpoint of line segment DE.

[0303] The displacement of vertices will be described below with reference to Figures 39 and 40. The displacement of vertices is performed by the reconstructor.

[0304] Fig. 39 is an explanatory diagram showing an example of displacement of vertices after subdivision, and Fig. 40 is an explanatory diagram showing an example of vertices of an original mesh.

[0305] The base mesh shown in FIG. 39(a) includes vertices A, B, C, and Z, and connectivity information indicating their connectivity.

[0306] 39(b) shows a mesh generated by the first subdivision, in other words, a mesh after the first subdivision (i.e., the first LoD). In the first subdivision, the subdivider generates vertices S, T, U, X, or Y and connectivity information indicating their connectivity. The vertices S, T, U, X, or Y are similar to the vertices D, E, and F shown in FIG. 38(b).

[0307] 39(c) shows a mesh generated by the second subdivision, in other words, a mesh after the second subdivision (i.e., the second LoD). In the second subdivision, the subdivider generates vertices D, E, F, G, and H and connectivity information indicating their connectivity. Vertices D, E, F, G, and H are similar to vertices G, H, I, J, K, L, M, N, and O shown in FIG. 38(c).

[0308] Figure 39(d) shows a mesh including the vertices after they have been displaced after subdivision, with vertices A, B, C, D, E, F, G, H, S, T, U, X, Y, and Z shown in Figure 39(d) being located at positions displaced using displacement information from the positions of the vertices shown in Figure 39(c).

[0309] The original mesh shown in FIG. 40 is an example of the mesh input to the encoding device 100, that is, the mesh before encoding.

[0310] The mesh shown in Fig. 39 has a shape similar to that of the original mesh shown in Fig. 40. The displacement information is generated by the encoding device 100 as information indicating the displacement from the vertices of the base mesh to the vertices of the original mesh, and therefore, by reconstructing the mesh using the displacement information thus generated, a mesh having a shape similar to that of the original mesh is generated.

[0311] The decoding device 200 can output the mesh shown in FIG.

[0312] Next, the division of a mesh into sub-meshes will be described with reference to FIGS.

[0313] A mesh can be divided into smaller parts and coded separately, with the vertices of the mesh being divided in such a way that the coordinates and connectivity of the vertices in each part can be coded independently.

[0314] Fig. 41 is an explanatory diagram showing an example of a mesh, and Fig. 42 is an explanatory diagram showing an example of dividing a mesh into sub-meshes.

[0315] The mesh shown in FIG. 41 is the original mesh, which is sometimes called a full mesh in contrast to a sub-mesh.

[0316] Figure 42 shows how the full mesh shown in Figure 41 is divided into two sub-meshes. For vertices A, B, and C of the full mesh (see Figure 41), vertex A is duplicated to vertices A1 and A2, vertex B is duplicated to vertices B1 and B2, and vertex C is duplicated to vertices C1 and C2, thereby creating two sub-meshes (i.e., a first sub-mesh and a second sub-mesh) from the full mesh. The first sub-mesh and the second sub-mesh are each independently decodable meshes.

[0317] Packing of displacement information into image frames will be described below with reference to FIGS. 43, 44 and 45.

[0318] 43, 44, and 45 are explanatory diagrams showing examples of packing displacement information into image frames. Note that image frames can also be called video frames.

[0319] The vertex displacement data is encoded as image frame data by being mapped to each component of a YUV format image frame (i.e., each of the Y component (Y Plane), U component (U Plane), and V component (V Plane)). This case will be described below as an example. As another example, the vertex displacement data may be encoded as image frame data by being mapped to each component of an RGB format image frame (each of the R component, G component, and B component).

[0320] The decoding device 200 can use an image encoding module to extract the displacement data. The displacement data can be in the form of X, Y, or Z components in a global coordinate system (e.g., a Cartesian coordinate system), or normal, tangential, or both tangential components in a local coordinate system. Methods for mapping the displacement data to an image frame include the following:

[0321] For example, in the first method, the displacement data is arranged in the image frame in scan order. An example of packing the displacement data in this case is shown in Figure 43. The displacement data is directly mapped to the image frame according to a predefined scan order.

[0322] Note that since an image frame has a fixed height and width, it may happen that the displacement data does not fit perfectly in the frame, in which case the remaining part of the image frame is padded with padding data (see Figure 43).

[0323] For example, in the second method, the displacement data is separated into multiple LoDs and mapped to the Y, U, and V components of the image frame. An example of packing of the displacement data in this case is shown in Figure 44. Here, the displacement data of the image frame of the next LoD starts immediately after the displacement data of the previous LoD ends. As in the first method, if the displacement data does not fit exactly into the image frame, padding is performed at the end of the image frame (see Figure 44).

[0324] For example, in the third method, displacement data corresponding to the LoD is mapped to the Y component, U component, and V component of the image frame in a manner different from that in the second method. An example of packing of the displacement data in this case is shown in Figure 45. In this way, each LoD can be decoded independently. In the third method, middle padding is performed on the displacement data of each LoD, and CTU alignment is performed together with padding at the end of the video frame (see Figure 45).

[0325] Next, the encoding device 100 and the decoding device 200 when a mesh is divided into a plurality of sub-meshes will be described.

[0326] Fig. 46 is a diagram showing another example of the configuration of the encoding device 100 according to an embodiment. Specifically, Fig. 46 is a diagram showing the configuration of a submesh encoding device that performs encoding processing when an input mesh 1101 is divided into divided meshes (plurality of submeshes). For example, the submesh encoding device includes a plurality of encoding devices 100.

[0327] An input mesh 1101 (full mesh) input to the submesh encoding device is divided into multiple meshes (submeshes). The multiple submeshes are input to, for example, multiple encoding devices 100. Each of the multiple submeshes may be input to one of the multiple encoding devices 100. For example, the submesh encoding device divides the input mesh 1101 into multiple submeshes and inputs the divided multiple submeshes to the multiple encoding devices 100.

[0328] After the image is divided into a plurality of sub-meshes, coding processing is performed on the boundaries of the sub-meshes (processing of overlapping sub-meshes).

[0329] For example, for each sub-mesh, preprocessing is performed by the preprocessor 1103, and a base mesh, displacement data, and metadata are generated and encoded.

[0330] The encoding device 100 may be realized by such a submesh encoding device configuration. That is, the encoding device 100 may be configured to include a plurality of preprocessors 1103 and compressors 1106, and may perform a predetermined process on the submesh for each of a plurality of pairs of preprocessors 1103 and compressors 1106. Furthermore, the number of pairs of preprocessors 1103 and compressors 1106 included in the encoding device 100 may be any number and is not particularly limited.

[0331] FIG. 47 is a diagram illustrating an example of the configuration of the preprocessor 1103 according to the embodiment.

[0332] The preprocessor 1103 includes, for example, a base mesh generator 1401 , a subdivision unit 1402 , and a displacement data generator 1403 .

[0333] First, in the preprocessor 1103, a base mesh is generated by the base mesh generator 1401.

[0334] The base mesh is then subdivided in a predetermined manner by a subdivider 1402 to generate a subdivided mesh (or a subdivided base mesh), which is the subdivided base mesh.

[0335] Displacement data is generated by a displacement data generator 1403 from the sub-divided mesh and the sub-mesh that is the input mesh 1101 after division.

[0336] The displacement data is, for example, a difference vector between the input mesh 1101 and the sub-division mesh.

[0337] The same subdivision method as that used in encoding is used in decoding as well.

[0338] Furthermore, for example, the encoding device 100 may transmit to the decoding device 200 a subdivision method for encoding and parameters used in the subdivision.

[0339] Fig. 48 is a diagram showing another example of the configuration of the decoding device 200 according to the embodiment. Specifically, Fig. 48 is a diagram showing the configuration of a submesh decoding device, which is a device that performs decoding processing when a bitstream 2101 includes multiple submeshes. For example, the submesh decoding device includes multiple decoding devices 200. For example, the submesh decoding device includes multiple decoding devices 200 and a combiner 2109.

[0340] The coded data for each submesh included in the bit stream 2101 is input to the decompressor 2102 of each decoding device 200 .

[0341] Also, for example, the post-processor 2106 performs the processing of the reconstructor described above.

[0342] In the post-processor 2106, for each decoded sub-mesh, the base mesh is subdivided and the displacement vector is added to the subdivided base mesh to reconstruct the sub-mesh.

[0343] That is, the post-processor 2106 performs the above-described reconstructor processing for each sub-mesh.

[0344] The combiner 2109 combines (merges) the sub-meshes restored by the respective decoding devices 200 to reconstruct the full mesh (output mesh 2107) before division.

[0345] Note that the decoding device 200 may be realized by such a submesh decoding device configuration. That is, the decoding device 200 may be configured to include a plurality of decompressors 2102 and post-processors 2106, and each of the plurality of pairs of decompressors 2102 and post-processors 2106 may perform a predetermined process on the encoded data for each submesh. Furthermore, the number of pairs of decompressors 2102 and post-processors 2106 included in the decoding device 200 may be any number and is not particularly limited.

[0346] Fig. 49 is a diagram showing a specific example of the configuration of the decoding device 200 according to the embodiment. Specifically, Fig. 49 shows a specific configuration of the post-decoder 2306 out of the decoder 2305 and post-decoder 2306 included in the decoding device 200.

[0347] The post-decoder 2306 comprises a pre-reconstructor 2307 , a reconstructor 2308 , a post-reconstructor 2309 and an adaptor 2310 .

[0348] The processing in the post-decoder 2306 is optional depending on the application. An example of the processing in the post-decoder 2306 (post-decoding processing) is conversion of the decoded data to a nominal format, such as video conversion from YUV space to RGB space. The post-decoding processing can be encapsulated into multiple processes, such as pre-reconstruction, reconstruction, post-reconstruction, and adaptation.

[0349] The pre-reconstructor 2307 performs a pre-reconstruction process, for example, upscaling the normalized texture coordinates to match the dimensions of the texture image in the context of video-based dynamic mesh coding.

[0350] The reconstructor 2308 performs a reconstruction process, which is invoked, for example, on the decoded atlas frame, the decoded base mesh frame, the decoded video frame, and syntax elements associated with the same mesh sequence. The output of the reconstruction process is a series of reconstructed mesh frames prior to a post-reconstruction process.

[0351] The post-reconstructor 2309 performs post-reconstruction operations, e.g., in the context of video-based dynamic mesh coding, which perform a number of smoothing operations on the reconstructed mesh frames, such as collapsing edges in the mesh or adding new vertices in the mesh.

[0352] The adaptor 2310 performs a fitting process, which may be applied by some application to fit the reconstructed mesh to a given scenario. For example, the vertices of the reconstructed mesh are transformed from the 3D model coordinate system to the 3D world coordinate system. The adaptor 2310 outputs a final reconstructed mesh frame (final 3D mesh frame 2311).

[0353] <Encoding and Decoding of Attribute Information> Figure 50 is a conceptual diagram showing a specific example of a bitstream. In a bitstream, encoded data is encapsulated in a data unit structure. Specifically, a bitstream of encoded mesh data is also expressed as a DMC (Dynamic Mesh Coding) bitstream and is composed of a series of DMC units. Each DMC unit is a data unit and is classified into a base mesh data unit, a displacement data unit, a metadata unit, a texture map unit, or the like.

[0354] For example, encoded base meshes are stored in base mesh data units with headers, encoded displacement data are stored in displacement data units with headers, encoded metadata are stored in metadata units with headers, and encoded texture maps are stored in texture map units with headers.

[0355] The header included in the unit is also called a unit header. The unit header stores a unit type that indicates the type of data stored in the payload.

[0356] The metadata may correspond to a parameter set or SEI (Supplemental Enhancement Information). The parameter set may be, for example, a parameter set common to frames (data with the same timing, access units, samples, etc.), such as a frame parameter set, or a parameter set common to a sequence, such as a sequence parameter set. Frame parameters may be expressed as picture parameter sets.

[0357] The metadata may correspond to the header of the bitstream, i.e. the header of the bitstream may contain the headers of each unit and may also contain the metadata units.

[0358] In the decoding device 200, the bitstream is decoded using the data type and divided into a plurality of data units.

[0359] In encoding and decoding of a 3D mesh, face information consisting of vertex information and connectivity information, and attribute information for the faces are encoded and decoded, respectively. The attribute information for the faces includes, for example, color information for the faces and reflectance information for the faces.

[0360] For example, the three-dimensional data has attribute information (e.g., color information for each face and reflectance for each face) for multiple faces of one mesh. For example, as shown in FIG. 3, an image of each face is mapped onto a two-dimensional image. Not only the image of each face, but also the attribute information of each face can be mapped onto the two-dimensional image. Therefore, attribute information is mapped onto the two-dimensional image for each attribute type, and multiple two-dimensional images corresponding to the multiple attribute types are obtained.

[0361] Then, in order to signal attribute information corresponding to multiple attribute types, multiple 2D images are signaled. However, the overhead of signaling multiple 2D images may increase the amount of code and processing. Therefore, in this embodiment, multiple 2D images are packed into a unit corresponding to one image frame, such as one image. This makes it possible to signal multiple 2D images collectively.

[0362] Here, the two-dimensional image onto which attribute information is mapped is an image, and may also be expressed as an attribute map, a two-dimensional attribute information set, or a two-dimensional attribute data set.

[0363] The mapping information indicating the correspondence between the multiple faces in the 3D mesh and the multiple faces in the 2D image may be the same or different for multiple attribute types. That is, the attribute information of each attribute type of the 3D mesh may be mapped to the 2D image based on the same correspondence among the multiple attribute types, or may be mapped to the 2D image based on different correspondence among the multiple attribute types.

[0364] Then, the mapping information may be signaled from encoding apparatus 100 to decoding apparatus 200. That is, encoding apparatus 100 may encode the mapping information, and decoding apparatus 200 may decode the mapping information. The mapping information may be signaled as attribute information or as vertex information.

[0365] 51 is a block diagram showing an example of the configuration of an encoding device 100 that encodes a plurality of attribute maps. The encoding device 100 includes a preprocessor 1103 and a compressor 1106.

[0366] A plurality of attribute maps 1102 are input to the encoding device 100. A preprocessor 1103 of the encoding device 100 stores the plurality of attribute maps 1102 in an image 1110. A compressor 1106 then encodes the image 1110 into a bitstream 1107 using an image coding method.

[0367] Note that an input mesh 1101 may be input to the encoding device 100. Then, the preprocessor 1103 may obtain a base mesh 1104, displacement data 1105, and metadata 1108 from the input mesh 1101. Then, the compressor 1106 may encode the base mesh 1104, displacement data 1105, and metadata 1108 into a bitstream 1107.

[0368] Alternatively, the preprocessor 1103 may acquire attribute information from the input mesh 1101 and acquire the attribute maps 1102 from the attribute information, instead of acquiring the attribute maps 1102 from an external source.

[0369] Also, in the example of FIG. 51, the image 1110 may be considered as one attribute map in which multiple attribute maps 1102 are stored.

[0370] 52 is a block diagram showing an example of the configuration of a decoding device 200 that decodes a plurality of attribute maps. The decoding device 200 includes a decompressor 2102 and a post-processor 2106.

[0371] The decoding device 200 decodes an image 2110 from a bitstream 2101. Then, a post-processor 2106 reconstructs a plurality of attribute maps 2108 from the image 2110.

[0372] The decoding device 200 inputs the restored attribute maps 2108 together with information indicating the attribute type of each attribute map to a renderer 2111. The renderer 2111 is a rendering device that reconstructs a 3D mesh by switching between the attribute maps 2108 depending on the application use of the 3D mesh.

[0373] The decompressor 2102 may decode the base mesh 2103, the displacement data 2104, and the metadata 2105 from the bitstream 2101. The post-processor 2106 may then obtain an output mesh 2107 from the base mesh 2103, the displacement data 2104, and the metadata 2105, and input the output mesh 2107 to the renderer 2111.

[0374] The renderer 2111 may then reconstruct a 3D mesh based on the output mesh 2107 and the plurality of attribute maps 2108. The renderer 2111 may be included in the decoding device 200.

[0375] Also, in the example of FIG. 52, the image 2110 may be considered as one attribute map in which multiple attribute maps 2108 are stored.

[0376] In this embodiment, an example is shown in which a plurality of attribute maps, each having different attribute types such as color and reflectance, are packed into one image, etc. Then, a packing method, an unpacking method, an encoding method, a decoding method, metadata, etc., related to this processing are disclosed.

[0377] 53 is a flow diagram showing an example of processing performed by the encoding device 100 to encode multiple attribute maps. In this example, the preprocessor 1103 of the encoding device 100 packs multiple attribute maps into an image (S301). Then, the preprocessor 1103 generates placement information indicating the placement of each attribute map in the image (S302).

[0378] The compressor 1106 of the encoding device 100 encodes the image and encodes metadata including the number of attribute maps, the attribute type of each attribute map, and placement information (S303), and then transmits a bitstream including the encoded image and encoded metadata (S303).

[0379] FIG. 54 is a flowchart showing an example of processing by the decoding device 200 to decode a plurality of attribute maps.

[0380] In this example, the decompressor 2102 of the decoding device 200 decodes an image from a bitstream (S401). The decompressor 2102 may decode the image according to a video coding standard such as VVC or HEVC. The decompressor 2102 decodes the image according to an image coding standard such as JPEG or PNG. These video coding standards and image coding standards correspond to methods used to encode or decode the image (i.e., image coding methods or image decoding methods).

[0381] The image format may be a YUV420 color difference format, a YUV444 color difference format, or a YUV400 color difference format.

[0382] The decompressor 2102 also decodes from the bitstream the number of attribute maps and the attribute type of each of the attribute maps (S402), where the number of attribute maps is the number of attribute maps of the 3D mesh, and each attribute map is associated with one attribute type.

[0383] The decompressor 2102 also decodes placement information indicating the placement of each attribute map in the image from the bitstream (S403). The placement information may indicate the position and size of each attribute map in the image. The placement information may also be expressed as position information.

[0384] Next, the decompressor 2102 unpacks the attribute maps from the image using the placement information (S404), and the renderer 2111 uses the unpacked attribute maps to render a 3D mesh (S405).

[0385] 55 is a conceptual diagram showing an example of the storage locations of the number of attribute maps and the attribute type of each attribute map in a bitstream. In this example, the number of attribute maps and the attribute type of each attribute map are stored in the main body of the bitstream, i.e., in the data area of ​​the bitstream. The data area may be the payload area.

[0386] 56 is a conceptual diagram showing another example of the storage location of the number of attribute maps and the attribute type of each attribute map in a bitstream. In this example, the number of attribute maps and the attribute type of each attribute map are stored in the header area.

[0387] The header area may include a sequence parameter set, a picture parameter set, a frame parameter set, SEI (Supplementary Enhancement Information), a header of each data unit, etc. The number of attribute maps and the attribute type of each attribute map may be stored in the header of the data unit, or may be stored in the sequence parameter set, frame parameter set, SEI, etc.

[0388] 57 is an explanatory diagram showing an example of attribute types and their descriptions. An example of the attribute type is texture, in which case the attribute map indicates the surface color of the 3D mesh, specifically, the color of each position on the surface.

[0389] Another example of an attribute type is material ID, in which case the attribute map indicates a different identifier for each material on the surface of the 3D mesh, specifically a material ID for each position on the surface. Another example of an attribute type is transparency, in which case the attribute map indicates the transparency of the surface of the 3D mesh, specifically the transparency for each position on the surface.

[0390] Another example of an attribute type is reflectance, in which case the attribute map indicates the reflectance of the surface of the 3D mesh, specifically the reflectance at each position on the surface. The reflectance may be ambient reflectance, diffuse reflectance, or specular reflectance. Two or more types of reflectance may be used as two or more attribute types. The reflectance of the surface of the 3D mesh may represent light and shadow on the surface of the 3D mesh.

[0391] Another example of an attribute type is normals, where the attribute map indicates the surface normals of the 3D mesh, specifically the normals at each location on the surface. The surface normals of the 3D mesh may represent height differences between different parts of the surface of the 3D mesh. Another example of an attribute type is texture coordinates, where the attribute map indicates the surface texture coordinates of the 3D mesh, specifically the texture coordinates at each location on the surface.

[0392] Another example of an attribute type is a face group ID, in which case the attribute map indicates the face group IDs of the surfaces of the 3D mesh, specifically the face group IDs of each position on the surface.

[0393] Another example of an attribute type is transmission filter color, where the attribute map indicates a transmission filter color for the surface of the 3D mesh. Another example of an attribute type is lighting model, where the attribute map is a lighting model to use for the surface of the 3D mesh. Another example of an attribute type is focus, where the attribute map indicates the focus of a specular highlight to use for the surface of the 3D mesh. Another example of an attribute type is sharpness, where the attribute map indicates the sharpness of reflections for the surface of the 3D mesh.

[0394] Another example of an attribute type is optical density, where the attribute map indicates the optical density of a surface of the 3D mesh. Another example of an attribute type is surface roughness, where the attribute map indicates a scalar value for deforming the surface of the 3D mesh to create the surface roughness. Another example of an attribute type is surface identifier, where the attribute map indicates identifiers for different surfaces of the 3D mesh. Another example of an attribute type is direction, where the attribute map indicates the direction in which a surface of the 3D mesh is facing.

[0395] Another example of an attribute type is light and shadow, where the attribute map indicates light and shadow on the surface of the 3D mesh. Another example of an attribute type is elevation difference, where the attribute map indicates elevation difference in different parts of the surface of the 3D mesh.

[0396] The attribute types of the attribute maps packed into the image may include one or more of the attribute types listed above.

[0397] FIG. 58 is a conceptual diagram illustrating a first example of packing multiple attribute maps. In this example, multiple attribute maps with different attribute types are stored (i.e., embedded) in a single image. Specifically, an attribute map of attribute type A, an attribute map of attribute type B, and an attribute map of attribute type C are packed in a single image. All attribute maps of a 3D mesh may be packed in the same image. In this example, each attribute map is stored in a rectangular area within a single image.

[0398] Each rectangular area may include an area where no attribute information is mapped, and the area may include padding data. Also, the image may include an area where no attribute map is stored, and the area may include padding data.

[0399] 59 is a syntax diagram showing an example of a syntax structure in packing example 1. In this example, num_attributes indicates the number of attribute maps. The number of attribute maps may be two or more.

[0400] The video_component_id indicates an identifier of a video component, specifically, indicates which video component each attribute map belongs to. The video component may be the video itself. That is, the video_component_id may indicate an identifier of the video, and may indicate which video each attribute map belongs to. The video component may also be an image included in the video.

[0401] The attribute_type_id indicates an identifier of the attribute type. That is, the attribute_type_id indicates the attribute type of each attribute map. The start_x, start_y, attribute_width, and attribute_height are examples of the layout information of the attribute map.

[0402] Specifically, start_x indicates the horizontal position of the attribute map in the image (more specifically, the position of the left edge of the attribute map), start_y indicates the vertical position of the attribute map in the image (more specifically, the position of the top edge of the attribute map), attribute_width indicates the width of the attribute map in the image, and attribute_height indicates the height of the attribute map in the image.

[0403] For example, multiple attribute maps may be stored in different parts of an image. Multiple attribute maps may be stored in the luminance Y component, the chrominance Cb or U component, and the chrominance Cr or V component of an image. Alternatively, multiple attribute maps may be stored in multiple color components of red R, green G, and blue B. Alternatively, multiple attribute maps may be stored in only one component, such as the luminance Y component. Also, each attribute map may be stored in consecutive samples in the image.

[0404] 60 is a syntax diagram showing another example of a syntax structure. Multiple attribute maps may be stored in multiple video components. In this example, the number of attribute maps included in each video component, the attribute type of each attribute map, and location information of each attribute map are signaled as information about the video components of the coded image.

[0405] Specifically, num_attributes_video_component indicates the number of video components in which multiple attribute maps are stored, and num_attributes_in_video_component indicates the number of attribute maps stored in each video component.

[0406] The attribute_type_id field indicates the attribute type of each attribute map by an identifier, the start_x field indicates the horizontal position of each attribute map, the start_y field indicates the vertical position of each attribute map, the attribute_width field indicates the width of each attribute map, and the attribute_height field indicates the height of each attribute map.

[0407] Figure 61 is a syntax diagram illustrating yet another example of a syntax structure, in which a syntax structure for signaling an attribute component is shown, where the attribute component corresponds to attribute information, and is not limited to an attribute map stored in a video component.

[0408] In this example, num_attribute_component indicates the number of attribute components to be processed, component_type indicates whether each attribute component is stored in a video component, and component_attribute_type_id indicates the attribute type of the attribute component by an identifier.

[0409] When an attribute component is stored in a video component, a packing_attribute_info_flag is signaled to indicate whether packing information of the attribute component is signaled. If packing_attribute_info_flag=1, the packing information of the attribute component is signaled by packing_attribute_info.

[0410] Packing_attribute_info corresponds to packing information and includes num_packing_attribute and packing_attribute_type_id, where num_packing_attribute indicates the number of attribute maps packed into the video component, and packing_attribute_type_id indicates the attribute type of each attribute map.

[0411] That is, an attribute component may include multiple attribute maps, and the attribute types of the attribute maps may be more specific than the attribute types of the attribute component.

[0412] Although the configuration information is omitted in the example of Fig. 61 , the configuration information may be signaled. Specifically, start_x, start_y, attribute_width, and attribute_height may be signaled, similarly to the examples of Fig. 59 etc.

[0413] 62 is a conceptual diagram showing a second example of packing of multiple attribute maps. In this example, multiple attribute maps with different attribute types are stored in multiple sub-components of a single image.

[0414] Specifically, in this example, multiple attribute maps of a 3D mesh are packed into a luma Y component, a chroma U component, and a chroma V component. More specifically, in this example, an attribute map of attribute type A is stored in the luma Y component, an attribute map of attribute type B is stored in the chroma U component, and an attribute map of attribute type C is stored in the chroma V component. The luma Y component, the chroma U component, and the chroma V component are different components (also expressed as video sub-components) of the same image.

[0415] Also, in this example, there are three attribute maps packed into three components, the Y component, the U component, and the V component of the image. The number of attribute maps packed into multiple components of the image may be two, or may be four or more. The number of components of the image into which multiple attribute maps are packed may be two, or may be four or more. Furthermore, two or more of the multiple attribute maps may be stored in one of the multiple components of the image. Alternatively, multiple attribute maps may be stored in a one-to-one correspondence with each of the multiple components of the image.

[0416] 63 is a syntax diagram showing an example of a syntax structure in the second packing example. In this example, num_attributes indicates the number of attribute maps. The number of attribute maps may be two or more. attribute_type_id indicates the attribute type of each attribute map with an identifier.

[0417] The video_component_id indicates an identifier of a video component, specifically, indicates which video component each attribute map belongs to. The video component may be the video itself. That is, the video_component_id may indicate an identifier of a video, and may indicate which video each attribute map belongs to. The video component may also be an image included in the video.

[0418] "video_sub_component_id" indicates the identifier of a video subcomponent (video subcomponent ID), specifically, to which video subcomponent each attribute map belongs. A video subcomponent may correspond to any one of the luminance Y component, the chrominance U component, and the chrominance V component. That is, "video_component_id" may indicate any one of the luminance Y component, the chrominance U component, and the chrominance V component with an identifier, and may indicate to which component each attribute map belongs.

[0419] In the example of Figure 63, start_x, start_y, attribute_width, and attribute_height may be signaled, similar to the examples of Figure 59 and the like.

[0420] Fig. 64 is an explanatory diagram showing an example of video subcomponent IDs in the second packing example. In this example, the video subcomponent ID (video_sub_component_id) can take three values: 0, 1, and 2. When the video subcomponent ID is 0, it indicates the luminance Y component. When the video subcomponent ID is 1, it indicates the chrominance Cb (or U) component. When the video subcomponent ID is 2, it indicates the chrominance Cr (or V) component.

[0421] The video subcomponents are not limited to this example and may correspond to color components R, G, and B. Furthermore, the video subcomponents may be substantially video components or image components.

[0422] 65 is a conceptual diagram showing a third example of packing of multiple attribute maps. In this example, multiple attribute maps with different attribute types are stored in multiple different hierarchical levels (layers) corresponding to a single image.

[0423] That is, in this example, multiple attribute maps of different attribute types are packed into multiple images in different layers in an access unit (AU). Also, in this example, each layer includes only one attribute map. Specifically, in this example, an attribute map of attribute type A is stored in the first layer, an attribute map of attribute type B is stored in the second layer, and an attribute map of attribute type C is stored in the third layer. The first layer, second layer, and third layer are different layers corresponding to the same image.

[0424] As mentioned above, in this example, each layer includes only one attribute map, but each layer may also include multiple attribute maps of different attribute types.

[0425] The multiple layers may be, for example, multiple layers in scalable coding. The multiple layers may include a base layer and one or more enhancement layers. The multiple layers may correspond to multiple different resolutions. Alternatively, the multiple layers may correspond to multiple different views including a base view and one or more enhanced views. The multiple images in the multiple layers may have different sizes or the same size.

[0426] FIG. 66 is a syntax diagram showing an example of a syntax structure in the third packing example. In this example, num_attributes indicates the number of attribute maps. The number of attribute maps may be two or more. attribute_type_id indicates the attribute type of each attribute map with an identifier. layer indicates which layer each attribute map belongs to. attribute_width indicates the width of the attribute map in the image, and attribute_height indicates the height of the attribute map in the image.

[0427] Two or more of the first packing example, the second packing example, and the third packing example may be combined. For example, a plurality of attribute maps may be packed into each component of each layer corresponding to an image, and placement information indicating the layer, component, position, and size of each attribute map may be encoded.

[0428] In the example of Fig. 66, start_x and start_y may be signaled, similar to the examples of Fig. 59 etc. Also, in the example of Fig. 66, video_component_id may be signaled, or video_sub_component_id may be signaled, similar to the examples of Fig. 63 etc.

[0429] 67 is a conceptual diagram illustrating an example of unpacking multiple attribute maps. For example, the decoding device 200 unpacks multiple attribute maps from an image by dividing the attribute maps from the image, the attribute maps having different attribute types. Here, each attribute map is unpacked using placement information.

[0430] The decoding device 200 may decode only some of the attribute maps corresponding to some of the attribute types, rather than decoding all of the attribute maps corresponding to all of the attribute types. In other words, the decoding device 200 may extract and output some of the attribute maps from the image.

[0431] 68 is a conceptual diagram showing an example of an attribute map request and response. For example, the application device 2112 requests an attribute map by specifying an attribute type. The decoding device 200 responds with an attribute map for the specified attribute type.

[0432] The application device 2112 may specify a large (coarse) unit attribute type, and the decoding device 200 may respond with a plurality of small (fine) unit attribute types included in the large (coarse) unit attribute type and a plurality of attribute maps for the plurality of small (fine) unit attribute types.

[0433] The application apparatus 2112 may be the renderer 2111. The decoding apparatus 200 may include the application apparatus 2112, or the application apparatus 2112 may include the decoding apparatus 200.

[0434] 69 is a conceptual diagram illustrating an example of rendering using multiple attribute maps. The renderer 2111 renders a 3D mesh using multiple attribute maps of multiple attribute types. For example, the renderer 2111 renders the 3D mesh using a combination of a first attribute in a first attribute map of a first attribute type and a second attribute in a second attribute map of a second attribute type.

[0435] The renderer 2111 may render a 3D mesh using all attribute types, or may render a 3D mesh using one or more of all attribute types. That is, the renderer 2111 may render a 3D mesh using all attribute maps for all attribute types, or may render a 3D mesh using one or more attribute maps for one or more attribute types. The attribute type determines how the attributes in the attribute maps are reflected in the rendering of the 3D mesh.

[0436] In one example, the multiple attribute types are used together to render a 3D mesh by clipping sample values ​​obtained by adding multiple sample values ​​reflecting multiple attributes of multiple different attribute types. In another example, the multiple attribute types are used together to render a 3D mesh by applying a mathematical function to multiple sample values ​​reflecting multiple attributes of multiple different attribute types based on an attribute identifier.

[0437] In the above examples, multiple attribute maps of a 3D mesh are aggregated into an image frame unit, which corresponds to one image frame. This may reduce overhead related to coded data headers, metadata, etc., in signaling multiple attribute maps of a 3D mesh compared to encoding and decoding multiple attribute maps individually. Also, because multiple attribute maps of a 3D mesh are decoded simultaneously by decoding one image frame unit, reduced complexity may be achieved.

[0438] Furthermore, since it becomes possible to encode and decode a plurality of attribute maps collectively, it becomes possible to suppress the implementation scale and complexity of the hardware and software.

[0439] Of the multiple aspects of the present disclosure, one aspect may be adopted, or two or more aspects may be combined and adopted. Also, of the multiple processing elements, multiple components, and multiple syntax elements of the present disclosure, one element may be adopted, or two or more elements may be combined and adopted. In particular, all elements of the present disclosure are not necessarily required, and only some elements of the present disclosure may be adopted.

[0440] Furthermore, processing corresponding to processing performed by the encoding device 100 may be performed by the decoding device 200 , and processing corresponding to processing performed by the decoding device 200 may be performed by the encoding device 100 .

[0441] In the above examples, multiple attribute maps for the faces of a 3D mesh are packed into one image frame unit for encoding a 3D mesh. However, the present invention is not limited to the attribute maps of a 3D mesh, and multiple attribute maps of other 3D models may be packed into one image frame unit.

[0442] For example, a plurality of attribute maps of a 3D model such as point cloud data, Neural Radiance Fields (NeRF), or 3D Gaussian Splatting (3DGS) may be packed into a single image frame. Specifically, in transmitting a projected image generated by projecting the 3D model onto an image, the projected image may be packed into a single image frame as an attribute map.

[0443] Specifically, by projecting the attributes of the 3D model onto an image for each of a plurality of attribute types, such as color, transparency, and reflectance, a plurality of projected images corresponding to the plurality of attribute types may be generated as a plurality of attribute maps corresponding to the plurality of attribute types.

[0444] The attribute type may correspond to a projection condition such as a viewpoint, a projection direction, or a projection surface. A plurality of projection images based on a plurality of different projection conditions may be treated as a plurality of attribute maps having different attribute types. Alternatively, a single attribute type, "projection image," may be assigned to a plurality of attribute maps corresponding to a plurality of projection images regardless of the projection conditions.

[0445] For example, attribute information of a 3D model defined using a method such as DMC (Dynamic Mesh Coding) or V3C (ISO / IES23090-5) or an attribute map of a projected image may be packed into one image frame unit.

[0446] The attribute information of the present disclosure is information associated with points or surfaces that constitute a three-dimensional object (three-dimensional model) to be processed, and is different from geometric information relating to the coordinates, size, shape, etc. of the points or surfaces in space. The attribute information may include not only image information expressed using color or brightness, but also other information expressed by numerical values, variables, or vectors. The attribute information may also include information indicating the correspondence between the attribute information and the geometric information.

[0447] In the present disclosure, a three-dimensional object (three-dimensional model) to be processed is associated with multiple attribute information sets. Multiple attribute information sets with different attribute types may be used, or multiple attribute information sets with the same attribute type may be used. For example, multiple images corresponding to the same three-dimensional object (three-dimensional model) may be used as multiple attribute information sets.

[0448] The attribute information may also be information to which a different correspondence is applied to the three-dimensional object (three-dimensional model) to be processed than that applied to the geometry information. For example, multiple images created by projecting the three-dimensional object (three-dimensional model) onto multiple planes may be used as multiple attribute information sets. Also, for example, multiple images created by projecting the three-dimensional object (three-dimensional model) onto a plane at multiple layers, such as multiple resolutions, may be used as multiple attribute information sets.

[0449] A plurality of attribute information sets may be included in one image frame unit by any of a plurality of packing examples in the present disclosure.

[0450] Specifically, as in the first packing example, a plurality of attribute information sets may be included in one image frame. Alternatively, as in the second packing example, a plurality of attribute information sets may be included in a plurality of different components constituting one image frame. Alternatively, as in the third packing example, a plurality of attribute information sets may be included in a plurality of different layers corresponding to one image frame.

[0451] Alternatively, a plurality of attribute information sets may be packed in one image frame unit by a combination of any two or more packing examples among the first packing example, the second packing example, and the third packing example.

[0452] Furthermore, for example, multiple different layers corresponding to one image frame may correspond to multiple corresponding image frames in multiple different layers. Specifically, multiple different layers corresponding to one image frame may correspond to multiple image frames that have the same metadata or the same time as one image frame. In other words, multiple different layers corresponding to one image frame may correspond to one access unit.

[0453] That is, one image frame unit corresponding to one image frame may be one access unit, or one image frame unit may be one picture unit.

[0454] In this disclosure, an image frame may correspond to a picture of an image or a field of an image. The representation of an image frame may be replaced by an image.

[0455] Furthermore, an image frame may be configured with a plurality of sub-image frames corresponding to a plurality of components, or may be configured with a plurality of sub-image frames corresponding to a plurality of layers, and a plurality of attribute information sets may be packed into the plurality of sub-image frames constituting the image frame.

[0456] Furthermore, an attribute value in an attribute information set included in a plurality of attribute information sets may be expressed as a difference value with respect to an attribute value in another attribute information set.

[0457] Each of the plurality of attribute information sets may be stored in a processing unit that allows parallel processing in an image frame. Specifically, each of the plurality of attribute information sets may be stored in a tile, a brick, a slice, or the like.

[0458] The same packing method may be applied to a plurality of three-dimensional objects (three-dimensional models) such as a plurality of three-dimensional mesh frames, or a plurality of different packing methods may be applied to the plurality of three-dimensional objects (three-dimensional models).

[0459] The encoding device 100 may transmit an encoded stream including multiple attribute information sets generated by the encoding method of the present disclosure. The decoding device 200 may receive the encoded stream. The encoding device 100 and the decoding device 200 may encode, transmit, receive, and decode only an attribute information set from the multiple attribute information sets that is specified, for example, by metadata included in the encoded stream or by the encoding device 100 or the decoding device 200.

[0460] The encoding device 100 and the decoding device 200 may share the attribute type and arrangement information for each of the multiple attribute maps in advance, and signaling of the attribute type and arrangement information may be omitted.

[0461] <Representative Example> Fig. 70 is a flow diagram showing an example of basic encoding processing according to this embodiment. For example, the circuit 151 of the encoding device 100 shown in Fig. 24 performs the encoding processing shown in Fig. 70 using the memory 152 in operation.

[0462] Specifically, the circuit 151 packs a plurality of two-dimensional attribute information sets of three-dimensional data having different attribute types into an image frame unit, which corresponds to one image frame (S501), and then encodes the image frame unit using an image encoding method that is a method used for encoding images (S502).

[0463] This may enable multiple two-dimensional attribute information sets with different attribute types to be collectively encoded as a single image frame unit using an image encoding method. Therefore, it may be possible to reduce the overhead in encoding multiple two-dimensional attribute information sets. Therefore, it may be possible to reduce the amount of code and the amount of processing.

[0464] For example, one image frame unit may be one image frame. Then, multiple two-dimensional attribute information sets may be packed into one image frame. This may make it possible to encode multiple two-dimensional attribute information sets as one image frame using an image encoding method. Therefore, it may be possible to reduce overhead, and it may be possible to reduce the amount of coding and the amount of processing.

[0465] In addition, multiple two-dimensional attribute information sets may be packed into one of multiple components of one image frame, or may be packed into one of multiple layers corresponding to one image frame.

[0466] Furthermore, for example, one image frame unit may be multiple components of one image frame. Then, multiple two-dimensional attribute information sets may be packed into multiple components of one image frame. This may make it possible to encode multiple two-dimensional attribute information sets as multiple components of one image frame using an image encoding method. Therefore, it may be possible to reduce overhead, and it may be possible to reduce the amount of coding and processing.

[0467] Furthermore, for example, one image frame unit may be multiple layers corresponding to one image frame. Then, multiple two-dimensional attribute information sets may be packed into multiple layers corresponding to one image frame. This may make it possible to encode multiple two-dimensional attribute information sets as multiple layers of one image frame using an image encoding method. Therefore, it may be possible to reduce overhead, and it may be possible to reduce the amount of code and the amount of processing.

[0468] Furthermore, for example, the multiple two-dimensional attribute information sets may include an image onto which attributes of each of multiple surfaces of the three-dimensional data are mapped as the two-dimensional attribute information set. This may make it possible to encode the image onto which attributes of each surface are mapped as one of the multiple two-dimensional attribute information sets. Therefore, it may be possible to efficiently encode the attributes of each surface.

[0469] Furthermore, for example, the plurality of two-dimensional attribute information sets may include an image onto which three-dimensional data is projected as a two-dimensional attribute information set. This may make it possible to encode the projected image of the three-dimensional data as one of the plurality of two-dimensional attribute information sets. Therefore, it may be possible to efficiently encode the projected image of the three-dimensional data.

[0470] Furthermore, for example, the circuit 151 may encode, for each of a plurality of two-dimensional attribute information sets, arrangement information indicating the arrangement of the two-dimensional attribute information set in one image frame unit. This may enable adaptive arrangement and efficient identification of the two-dimensional attribute information sets in one image frame unit. Here, the expression of arrangement may be replaced with arrangement location.

[0471] Furthermore, for example, the arrangement information may include information indicating the position and size of the two-dimensional attribute information set in the image included in one image frame unit, which may enable adaptive arrangement and efficient identification of the two-dimensional attribute information set in the image.

[0472] Note that an image included in one image frame unit may be an image corresponding to one component of one image frame, or may be an image corresponding to one layer corresponding to one image frame. Also, an image included in one image frame unit may be an image corresponding to a combination of one component of one image frame and one layer corresponding to one image frame.

[0473] Furthermore, for example, the arrangement information may include information for identifying an image including a two-dimensional attribute information set among a plurality of images included in one image frame unit, which may enable adaptive arrangement and efficient identification of the two-dimensional attribute information sets in the plurality of images.

[0474] The multiple images included in one image frame unit may be multiple images corresponding to multiple components of one image frame, or multiple images corresponding to multiple layers corresponding to one image frame. Also, the multiple images included in one image frame unit may be multiple images corresponding to a combination of multiple components of one image frame and multiple layers corresponding to one image frame.

[0475] Furthermore, for example, the circuit 151 may encode related information indicating the number of the multiple two-dimensional attribute information sets and at least one of the attribute types of each of the multiple two-dimensional attribute information sets. This may make it possible to appropriately share the number of the multiple two-dimensional attribute information sets or the information on the attribute types of each of the multiple two-dimensional attribute information sets.

[0476] Furthermore, for example, the circuit 151 may map the attributes of each of multiple surfaces of the three-dimensional data onto an image, thereby generating the attribute-mapped image as a two-dimensional attribute information set.

[0477] Furthermore, for example, the circuit 151 may project the three-dimensional data onto an image, thereby generating the image onto which the three-dimensional data is projected as a two-dimensional attribute information set.

[0478] In the above, packing a plurality of two-dimensional attribute information sets into one image frame unit may mean storing a plurality of two-dimensional attribute information sets in one image frame unit.

[0479] The three-dimensional data may be a three-dimensional model or a three-dimensional mesh.

[0480] Furthermore, although the circuit 151 of the encoding device 100 performs each process in the above, the attribute information encoder 103 of the encoding device 100 may perform each of the above processes. Alternatively, the preprocessor 1103 and compressor 1106 of the encoding device 100 may perform each of the above processes. Alternatively, other components may perform each of the above processes.

[0481] Fig. 71 is a flow diagram showing an example of basic decoding processing according to this embodiment. For example, the circuit 251 of the decoding device 200 shown in Fig. 25 performs the decoding processing shown in Fig. 71 using the memory 252 in operation.

[0482] Specifically, the circuit 251 decodes one image frame unit, which is a unit corresponding to one image frame, using an image decoding method that is a method used for decoding images (S601).Then, the circuit 251 unpacks multiple two-dimensional attribute information sets of three-dimensional data, each having different attribute types, from one image frame unit (S602).

[0483] This may enable multiple two-dimensional attribute information sets with different attribute types to be decoded together as a single image frame unit using an image decoding method. Therefore, it may be possible to reduce the overhead in decoding multiple two-dimensional attribute information sets. Therefore, it may be possible to reduce the amount of coding and the amount of processing.

[0484] For example, one image frame unit may be one image frame. Then, multiple two-dimensional attribute information sets may be unpacked from one image frame. This may make it possible to decode multiple two-dimensional attribute information sets as one image frame using an image decoding method. Therefore, it may be possible to reduce overhead, and it may be possible to reduce the amount of coding and processing.

[0485] Note that the multiple 2D attribute information sets may be unpacked from one of the multiple components of one image frame, or may be unpacked from one of the multiple layers corresponding to one image frame.

[0486] Furthermore, for example, one image frame unit may be multiple components of one image frame. Then, multiple two-dimensional attribute information sets may be unpacked from the multiple components of one image frame. This may make it possible to decode the multiple two-dimensional attribute information sets as multiple components of one image frame using an image decoding method. Therefore, it may be possible to reduce overhead, and it may be possible to reduce the amount of coding and processing.

[0487] Furthermore, for example, one image frame unit may be multiple layers corresponding to one image frame. Then, multiple two-dimensional attribute information sets may be unpacked from multiple layers corresponding to one image frame. This may make it possible to decode multiple two-dimensional attribute information sets as multiple layers of one image frame using an image decoding method. Therefore, it may be possible to reduce overhead, and it may be possible to reduce the amount of coding and processing.

[0488] Furthermore, for example, the multiple two-dimensional attribute information sets may include an image onto which attributes of each of multiple surfaces of the three-dimensional data are mapped as the two-dimensional attribute information set. This may make it possible to decode the image onto which the attributes of each surface are mapped as one of the multiple two-dimensional attribute information sets. Therefore, it may be possible to efficiently decode the attributes of each surface.

[0489] Furthermore, for example, the plurality of two-dimensional attribute information sets may include an image onto which three-dimensional data is projected as a two-dimensional attribute information set. This may make it possible to decode the projected image of the three-dimensional data as one of the plurality of two-dimensional attribute information sets. Therefore, it may be possible to efficiently decode the projected image of the three-dimensional data.

[0490] Furthermore, for example, the circuit 251 may decode, for each of a plurality of two-dimensional attribute information sets, arrangement information indicating the arrangement of the two-dimensional attribute information set in one image frame unit. This may enable adaptive arrangement and efficient identification of the two-dimensional attribute information set in one image frame unit. Here, the expression of arrangement may be replaced with arrangement location.

[0491] Furthermore, for example, the placement information may include information indicating the position and size of the two-dimensional attribute information set in an image included in one image frame unit, which may enable adaptive placement and efficient identification of the two-dimensional attribute information set in the image.

[0492] Note that an image included in one image frame unit may be an image corresponding to one component of one image frame, or may be an image corresponding to one layer corresponding to one image frame. Also, an image included in one image frame unit may be an image corresponding to a combination of one component of one image frame and one layer corresponding to one image frame.

[0493] Furthermore, for example, the arrangement information may include information for identifying an image including a two-dimensional attribute information set among a plurality of images included in one image frame unit, which may enable adaptive arrangement and efficient identification of the two-dimensional attribute information sets in the plurality of images.

[0494] The multiple images included in one image frame unit may be multiple images corresponding to multiple components of one image frame, or multiple images corresponding to multiple layers corresponding to one image frame. Also, the multiple images included in one image frame unit may be multiple images corresponding to a combination of multiple components of one image frame and multiple layers corresponding to one image frame.

[0495] Furthermore, for example, the circuit 251 may decode related information indicating the number of the multiple two-dimensional attribute information sets and at least one of the attribute types of each of the multiple two-dimensional attribute information sets, which may make it possible to appropriately share the number of the multiple two-dimensional attribute information sets or the information on the attribute type of each of the multiple two-dimensional attribute information sets.

[0496] Furthermore, for example, the circuit 251 may render three-dimensional data using a plurality of two-dimensional attribute information sets. In this case, the circuit 251 may select a two-dimensional attribute information set from the plurality of two-dimensional attribute information sets and render the three-dimensional data using the selected two-dimensional attribute information set.

[0497] In the above, unpacking a plurality of two-dimensional attribute information sets from one image frame unit may be extracting a plurality of two-dimensional attribute information sets from one image frame unit.

[0498] The three-dimensional data may be a three-dimensional model or a three-dimensional mesh.

[0499] In the above, the circuit 251 of the decoding device 200 performs each process, but the attribute information decoder 203 of the decoding device 200 may perform each of the above processes. Alternatively, the decompressor 2102 and post-processor 2106 of the decoding device 200 may perform each of the above processes. Alternatively, other components may perform each of the above processes.

[0500] <Other Examples> Although aspects of the encoding device 100 and the decoding device 200 have been described above according to the embodiments, the aspects of the encoding device 100 and the decoding device 200 are not limited to the embodiments. Modifications conceivable by those skilled in the art may be applied to the embodiments, and multiple components in the embodiments may be combined in any manner.

[0501] For example, a process performed by a specific component in the embodiment may be performed by another component instead of the specific component. Also, the order of multiple processes may be changed, or multiple processes may be performed in parallel.

[0502] Furthermore, as described above, at least some of the configurations of the present disclosure may be implemented as an integrated circuit. At least some of the processes of the present disclosure may be used as an encoding method or a decoding method. A program for causing a computer to execute the encoding method or the decoding method may be used. A non-transitory computer-readable recording medium on which the program is recorded may be used. A bitstream for causing the decoding device 200 to perform a decoding process may be used.

[0503] Furthermore, at least some of the configurations and processes of the present disclosure may be used as a transmitting device, a receiving device, a transmitting method, or a receiving method. A program for causing a computer to execute the transmitting method or the receiving method may be used. Furthermore, a non-transitory computer-readable recording medium on which the program is recorded may be used.

[0504] Also, for example, a phrase "at least one of" a first element, a second element, and a third element corresponds to the first element, the second element, the third element, or any combination thereof.

[0505] The present disclosure is useful, for example, in encoding devices, decoding devices, transmitting devices, receiving devices, etc. related to three-dimensional meshes, and is applicable to computer graphics systems, three-dimensional data display systems, etc.

[0506] 100 Encoding device 101, 121, 144 Vertex information encoder 102, 145 Connection information encoder 103, 122 Attribute information encoder 104, 204, 1103 Preprocessor 105, 205, 2106 Postprocessor 110 Three-dimensional data encoding system 111, 211 Controller 112, 212 Input / output processor 113 Three-dimensional data encoder 114 System multiplexer 115 Three-dimensional data generator 123 Metadata encoder 124, 1274 Multiplexer 131 Vertex image generator 132 Attribute image generator 133 Metadata generator 134 Video encoder 141 Two-dimensional data encoder 142 Mesh data encoder 143 Texture encoder 148 Description encoder 151, 251 Circuit 152, 252 Memory 200 Decoding device 201, 221, 244 Vertex information decoder 202, 245 Connection information decoder 203, 222 Attribute information decoder 210 3D data decoding system 213 3D data decoder 214 System demultiplexer 215, 247 Presentation device 216 User interface 223 Metadata decoder 224, 1231 Demultiplexer 231 Vertex information generator 232 Attribute information generator 234 Video decoder 241 2D data decoder 242 Mesh data decoder 243 Texture decoder 246 Mesh reconstructor 248 Description decoder 300 Network 310 External connector 511 Volumetric capture device 512 Projector 513 Base mesh encoder 514 Displacement encoder 515 Attribute encoder 516 Other type encoder 613 Base mesh decoder 614 Displacement decoder 615 Attribute decoder 616 Other type decoder 617 3D reconstructor 1101 Input mesh 1102, 2108 Attribute map 1104, 2103 Base mesh 1105, 2104 Displacement data 1106 Compressor 1107, 1304, 2101 Bitstream1108, 2105 Metadata 1110, 2110 Image 1232, 1262 Switch 1233 Static Mesh Decoder 1234, 1264 Mesh Buffer 1235 Motion Decoder 1236, 1266 Base Mesh Reconstructor 1237, 1240 Inverse Quantizer 1238, 1243 Video Decoder 1239 Image Unpacker 1241 Inverse Wavelet Transformer 1242, 2308 Reconstructor 1244, 1272 Color Converter 1251 Decoded Base Mesh 1252, 1402 Subdivider 1253 Subdivided Mesh 1254 Decoded Displacement Data 1255 Displacer 1256 Decoded 3D Mesh 1261, 1269 Quantizer 1263 Static Mesh Encoder 1265 Motion Encoder 1267 Displacement Data Updater 1268 Wavelet Transformer 1270 Image Packer 1271, 1273 Video Encoder 1301, 2304 Mesh Frame 1302, 2301, 2302 Base Mesh Frame 1303, 2303 Displacement Information 1401 Base Mesh Generator 1403 Displacement Data Generator 2102 Decompressor 2107 Output Mesh 2109 Combiner 2111 Renderer 2112 Application Device 2305 Decoder 2306 Post-Decoder 2307 Pre-Reconstructor 2309 Post-Reconstructor 2310 Adapter 2311 Final 3D Mesh Frame

Claims

1. An encoding method comprising packing a plurality of two-dimensional attribute information sets of three-dimensional data, each having different attribute types, into an image frame unit, which corresponds to one image frame, and encoding the one image frame unit using an image encoding method used for encoding images.

2. The encoding method according to claim 1, wherein the one image frame unit is the one image frame, and the plurality of two-dimensional attribute information sets are packed into the one image frame.

3. The encoding method according to claim 1, wherein the one image frame unit is a plurality of components of the one image frame, and the plurality of two-dimensional attribute information sets are packed into the plurality of components of the one image frame.

4. The encoding method according to claim 1, wherein the one image frame unit is a plurality of layers corresponding to the one image frame, and the plurality of two-dimensional attribute information sets are packed into the plurality of layers corresponding to the one image frame.

5. The encoding method according to any one of claims 1 to 4, wherein the plurality of two-dimensional attribute information sets include an image onto which attributes of each of the plurality of faces of the three-dimensional data are mapped as a two-dimensional attribute information set.

6. The encoding method according to any one of claims 1 to 4, wherein the plurality of two-dimensional attribute information sets includes an image onto which the three-dimensional data is projected as a two-dimensional attribute information set.

7. The encoding method according to any one of claims 1 to 4, further comprising encoding, for each of the plurality of two-dimensional attribute information sets, arrangement information indicating the arrangement of the two-dimensional attribute information set in the one image frame unit.

8. The encoding method according to claim 7, wherein the arrangement information includes information indicating the position and size of the two-dimensional attribute information set in the image included in the one image frame unit.

9. The encoding method according to claim 7, wherein the arrangement information includes information for identifying an image that includes the two-dimensional attribute information set from among a plurality of images included in one image frame unit.

10. The encoding method according to any one of claims 1 to 4, further comprising encoding related information indicating the number of the plurality of two-dimensional attribute information sets and at least one of the attribute types of each of the plurality of two-dimensional attribute information sets.

11. A decoding method comprising: using an image decoding method used for decoding images to decode one image frame unit, which is a unit corresponding to one image frame; and unpacking from said one image frame unit a plurality of two-dimensional attribute information sets of three-dimensional data, the plurality of two-dimensional attribute information sets having different attribute types.

12. The decoding method according to claim 11, wherein the one image frame unit is the one image frame, and the plurality of two-dimensional attribute information sets are unpacked from the one image frame.

13. The decoding method according to claim 11, wherein the one image frame unit is a plurality of components of the one image frame, and the plurality of two-dimensional attribute information sets are unpacked from the plurality of components of the one image frame.

14. The decoding method according to claim 11, wherein the one image frame unit is a plurality of layers corresponding to the one image frame, and the plurality of two-dimensional attribute information sets are unpacked from the plurality of layers corresponding to the one image frame.

15. A decoding method according to any one of claims 11 to 14, wherein the plurality of two-dimensional attribute information sets include an image onto which attributes of each of the plurality of faces of the three-dimensional data are mapped as a two-dimensional attribute information set.

16. A decoding method according to any one of claims 11 to 14, wherein the plurality of two-dimensional attribute information sets include an image onto which the three-dimensional data is projected as a two-dimensional attribute information set.

17. A decoding method according to any one of claims 11 to 14, further comprising decoding, for each of the plurality of two-dimensional attribute information sets, arrangement information indicating the arrangement of the two-dimensional attribute information set in the one image frame unit.

18. The decoding method according to claim 17, wherein the arrangement information includes information indicating the position and size of the two-dimensional attribute information set in the image included in the one image frame unit.

19. The decoding method according to claim 17, wherein the arrangement information includes information for identifying an image that includes the two-dimensional attribute information set from among the multiple images included in the one image frame unit.

20. A decoding method according to any one of claims 11 to 14, further comprising decoding related information indicating the number of the plurality of two-dimensional attribute information sets and at least one of the attribute types of each of the plurality of two-dimensional attribute information sets.

21. An encoding device comprising: a circuit; and a memory accessible by the circuit, wherein the circuit packs a plurality of two-dimensional attribute information sets of three-dimensional data, the plurality of two-dimensional attribute information sets having different attribute types, into one image frame unit, which is a unit corresponding to one image frame, and encodes the one image frame unit using an image encoding method that is a method used for encoding images.

22. A decoding device comprising: a circuit; and a memory accessible by the circuit, wherein the circuit decodes one image frame unit, which is a unit corresponding to one image frame, using an image decoding method that is a method used for decoding images; and unpacks from the one image frame unit a plurality of two-dimensional attribute information sets of three-dimensional data, the plurality of two-dimensional attribute information sets having mutually different attribute types.

Citation Information

Patent Citations

  • Image / video-based mesh compression

    US20230290008A1

  • Information processing device and method

    WO2024071283A1