Encoding method, decoding method, encoding device, and decoding device
By encoding or decoding a flag based on the number of attribute videos and sharing region information efficiently, the method addresses inefficiencies in three-dimensional mesh data encoding and decoding, reducing resource demands and improving processing efficiency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
- Filing Date
- 2025-10-14
- Publication Date
- 2026-05-07
AI Technical Summary
Existing methods for encoding and decoding three-dimensional mesh data are inefficient due to increased coding and processing requirements when signaling information in multiple attribute videos, leading to excessive resource utilization.
A method that encodes or decodes a flag indicating whether region information is the same among attribute videos based on the number of attribute videos, omitting flags when unnecessary, and encoding shared region information efficiently when applicable, thereby reducing coding and processing demands.
This approach reduces the amount of coding and processing required by optimizing the encoding and decoding of three-dimensional mesh data, enhancing efficiency and resource utilization.
Smart Images

Figure JP2025036144_07052026_PF_FP_ABST
Abstract
Description
Symbolization method, decoding method, symbolization device, and decoding device
[0001] The present disclosure relates to a symbolization method and the like.
[0002] In Patent Document 1, methods and apparatuses for encoding and decoding three-dimensional mesh data have been proposed.
[0003] Japanese Patent Application Laid-Open No. 2006-187015
[0004] Further improvement in the processing of encoding or decoding for three-dimensional data is desired. The present disclosure aims to improve encoding processing and the like for three-dimensional data.
[0005] The symbolization method according to one aspect of the present disclosure includes symbolizing a flag indicating whether region information for defining one or more regions in an image region of an attribute video is the same among attribute videos according to the number of attribute videos to be symbolized in three-dimensional data. When the number of attribute videos is 2 or more, the flag is symbolized. When the number of attribute videos is 1 or less, the flag is not symbolized.
[0006] These general or specific aspects may be implemented by a system, apparatus, method, integrated circuit, computer program, or a non-temporary recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, apparatus, method, integrated circuit, computer program, and recording medium.
[0007] The present disclosure can contribute to the improvement of encoding processing and the like for three-dimensional data.
[0008] This is a conceptual diagram showing a three-dimensional mesh according to the embodiment. This is a conceptual diagram showing the basic elements of a three-dimensional mesh according to the embodiment. This is a conceptual diagram showing mapping according to the embodiment. This is a block diagram showing an example configuration of an encoding / decoding system according to the embodiment. This is a block diagram showing an example configuration of an encoding device according to the embodiment. This is a block diagram showing another example configuration of an encoding device according to the embodiment. This is a block diagram showing an example configuration of a decoding device according to the embodiment. This is a block diagram showing another example configuration of a decoding device according to the embodiment. This is a conceptual diagram showing an example configuration of a bitstream according to the embodiment. This is a conceptual diagram showing yet another example configuration of a bitstream according to the embodiment. This is a block diagram showing a concrete example of an encoding / decoding system according to the embodiment. This is a conceptual diagram showing an example configuration of point cloud data according to the embodiment. This is a conceptual diagram showing an example data file for point cloud data according to the embodiment. This is a conceptual diagram showing an example configuration of mesh data according to the embodiment. This is a conceptual diagram showing an example data file for mesh data according to the embodiment. This is a conceptual diagram showing the types of three-dimensional data according to the embodiment. This is a block diagram showing an example configuration of a three-dimensional data encoder according to the embodiment. This is a block diagram showing an example configuration of a three-dimensional data decoder according to the embodiment. This is a block diagram showing another example configuration of a three-dimensional data encoder according to the embodiment. This is a block diagram showing another example configuration of a three-dimensional data decoder according to the embodiment. This is a conceptual diagram showing a specific example of the encoding process according to the embodiment. This is a conceptual diagram showing a specific example of the decoding process according to the embodiment. This is a block diagram showing an implementation example of the encoding device according to the embodiment. This is a block diagram showing an implementation example of the decoding device according to the embodiment. This is a block diagram showing another configuration example of the encoding / decoding system according to the embodiment. This is a block diagram showing another configuration example of the encoding device according to the embodiment. This is a block diagram showing yet another configuration example of the decoding device according to the embodiment. This is a flow diagram showing the processing of the encoding device according to the embodiment. This is an explanatory diagram conceptually showing the encoding of a mesh frame according to the embodiment.This is a flowchart showing the processing of the decoding device according to the embodiment. This is an explanatory diagram conceptually showing the decoding of a mesh frame according to the embodiment. This is a block diagram showing an example of the configuration of the decoding device according to the embodiment. This is a block diagram showing an example of the configuration of the encoding device according to the embodiment. This is a block diagram showing an example of the configuration of the decoding device according to the embodiment. This is an explanatory diagram showing an example of subdivision according to the embodiment. This is an explanatory diagram showing an example of vertex displacement after displacement following subdivision according to the embodiment. This is an explanatory diagram showing an example of the vertices of the original mesh according to the embodiment. This is an explanatory diagram showing an example of a mesh according to the embodiment. This is an explanatory diagram showing an example of subdivision of a mesh according to the embodiment. This is a first explanatory diagram showing an example of packing displacement information into an image frame according to the embodiment. This is a second explanatory diagram showing an example of packing displacement information into an image frame according to the embodiment. This is a third explanatory diagram showing an example of packing displacement information into an image frame according to the embodiment. This is another example of the configuration of the encoding device according to the embodiment. This is a block diagram showing an example of the configuration of the preprocessor according to the embodiment. This is another example of the configuration of the decoding device according to the embodiment. This is a block diagram showing a specific example of the configuration of the decoding device according to the embodiment. This is a conceptual diagram showing a specific example of a bitstream according to the embodiment. This is yet another example of the configuration of the encoding device according to the embodiment. This is yet another example of the configuration of the decoding device according to the embodiment. This is a flowchart showing another example of the processing of the encoding device according to the embodiment. This is a flowchart showing another example of the processing of the decoding device according to the embodiment. This is a conceptual diagram showing an example of the number of attribute maps and the storage location of attribute types in a bitstream according to the embodiment. This is a conceptual diagram showing another example of the number of attribute maps and the storage location of attribute types in a bitstream according to the embodiment. This is an explanatory diagram showing an example of attribute types and their descriptions according to the embodiment. This is a conceptual diagram showing a first packing example of multiple attribute maps according to the embodiment. This is a syntax diagram showing an example of the syntax structure in the first packing example according to the embodiment. This is a syntax diagram showing another example of the syntax structure according to the embodiment. This is a syntax diagram showing yet another example of the syntax structure according to the embodiment.This is a conceptual diagram showing a second example of packing multiple attribute maps according to the embodiment. This is a syntax diagram showing an example of the syntax structure in the second example of packing according to the embodiment. This is an explanatory diagram showing an example of a video subcomponent ID in the second example of packing according to the embodiment. This is a conceptual diagram showing a third example of packing multiple attribute maps according to the embodiment. This is a syntax diagram showing an example of the syntax structure in the third example of packing according to the embodiment. This is a conceptual diagram showing an example of unpacking multiple attribute maps according to the embodiment. This is a conceptual diagram showing an example of request and response for attribute maps according to the embodiment. This is a conceptual diagram showing an example of rendering using multiple attributes according to the embodiment. This is a conceptual diagram showing an example of tiles defined in the image area of an attribute video according to the embodiment. This is a conceptual diagram showing an example of the relationship between multiple types of attribute information and multiple attribute videos according to the embodiment. This is a conceptual diagram showing an example of the relationship between multiple attribute videos and tile information according to the embodiment. This is a flowchart showing yet another example of the processing of an encoding device according to the embodiment. This is a flowchart showing yet another example of the processing of a decoding device according to the embodiment. This is a syntax diagram showing an example of the syntax structure for signaling tile information. This is a syntax diagram showing another example of the syntax structure for signaling tile information. This is a syntax diagram showing yet another example of a syntax structure for signaling tile information. This is a syntax diagram showing yet another example of a syntax structure for signaling tile information. This is a syntax diagram showing yet another example of a syntax structure for signaling tile information. This is a flowchart showing an example of a basic encoding process according to the embodiment. This is a flowchart showing an example of a basic decoding process according to the embodiment.
[0009] <Introduction> For example, three-dimensional (3D) meshes are used in computer graphics images. For example, computer graphics images consist of multiple frames that are different in time, and each frame may be represented by a three-dimensional mesh.
[0010] Furthermore, a three-dimensional mesh consists of vertex information indicating the position of each of several vertices in three-dimensional space, connection information indicating the connections between the multiple vertices, and attribute information indicating the attributes of each vertex or face. Each face is constructed according to the connection relationships of the multiple vertices. Various computer graphics images can be represented using such a three-dimensional mesh.
[0011] For example, in the encoding of three-dimensional data such as a three-dimensional mesh, multiple types of attribute information may be embedded in one or more attribute videos for signaling. In this case, one or more regions may be defined in the image area of each attribute video, and attribute information may be embedded in at least one of these regions.
[0012] However, signaling information in these regions may increase the amount of coding and processing power required.
[0013] Therefore, the encoding method of Example 1 includes encoding a flag indicating whether or not the region information for defining one or more regions in the image region of the attribute video is the same among the attribute videos, depending on the number of attribute videos to be encoded in the three-dimensional data, wherein the flag is encoded when the number of attribute videos is two or more, and the flag is not encoded when the number of attribute videos is one or less.
[0014] This may allow for the omission of encoding a flag indicating whether or not region information is identical between attribute videos, depending on the number of attribute videos. Consequently, it may be possible to reduce the amount of coding and processing required.
[0015] Furthermore, the encoding method of Example 2 may be the encoding method of Example 1, wherein if the number of attribute videos is two or more and the region information is the same among the attribute videos, the region information shared by multiple attribute videos is encoded, and if the number of attribute videos is two or more and the region information is not the same among the attribute videos, the region information is encoded for each attribute video.
[0016] This may allow for efficient encoding of region information when there are two or more attribute videos, depending on whether the region information is identical between the attribute videos. Consequently, it may be possible to reduce the amount of encoding and processing required.
[0017] Furthermore, the encoding method in Example 3 may be the encoding method in Example 1 or 2, wherein the region information is encoded when the number of attribute videos is 1, and the region information is not encoded when the number of attribute videos is 0.
[0018] This may allow for the omission of encoding region information depending on the number of attribute videos. Consequently, it may be possible to reduce the amount of encoding and processing required.
[0019] Furthermore, the encoding method in Example 4 may be any of the encoding methods in Examples 1 to 3, wherein if the number of attribute videos is 1 or more, one or more attribute videos are encoded.
[0020] This may make it possible to encode attribute videos and region information in accordance with the number of attribute videos. Therefore, it may be possible to encode attribute videos and region information efficiently.
[0021] Furthermore, the encoding method in Example 5 may be any of the encoding methods in Examples 1 to 4, wherein if the number of attribute videos is two or more and the region information is the same among the attribute videos, an index for referencing the region information is encoded.
[0022] This can make it possible to efficiently identify and reference identical region information between attribute videos when there are two or more attribute videos. Therefore, it may be possible to efficiently define one or more regions within the image region of an attribute video.
[0023] Furthermore, the encoding method of Example 6 may be the encoding method of Example 5, wherein the index is not encoded if the number of attribute videos is 1 or less, or if the area information is not the same among the attribute videos.
[0024] This may allow for the omission of index encoding depending on the number of attribute videos and whether the region information is identical across attribute videos. Consequently, it may be possible to reduce the amount of coding and processing required.
[0025] Furthermore, the encoding method in Example 7 may be any of the encoding methods in Examples 1 to 6, wherein the flag is encoded for each layer of the three-dimensional data according to the number of attribute videos.
[0026] This may allow for the omission of encoding a flag indicating whether or not the region information is identical between attribute videos, depending on the number of attribute videos in each layer. Consequently, it may be possible to reduce the amount of coding and processing required.
[0027] Furthermore, the encoding method of Example 8 may be any of the encoding methods of Examples 1 to 7, wherein the region information is tile information for defining one or more tiles in the image region.
[0028] This may allow for the omission of encoding a flag indicating whether tile information is identical between attribute videos, depending on the number of attribute videos. Consequently, it may be possible to reduce the amount of coding and processing required.
[0029] Furthermore, the decoding method of Example 9 includes decoding a flag indicating whether or not the region information for defining one or more regions in the image region of an attribute video is the same among the attribute videos, depending on the number of attribute videos to be decoded in the three-dimensional data, wherein the flag is decoded when the number of attribute videos is two or more, and the flag is not decoded when the number of attribute videos is one or less.
[0030] This may allow for the omission of decoding a flag indicating whether or not region information is identical between attribute videos, depending on the number of attribute videos. Consequently, it may be possible to reduce the amount of coding and processing required.
[0031] Furthermore, the decoding method in Example 10 may be the decoding method in Example 9, wherein if the number of attribute videos is two or more and the region information is the same among the attribute videos, the region information shared by multiple attribute videos is decoded, and if the number of attribute videos is two or more and the region information is not the same among the attribute videos, the region information is decoded for each attribute video.
[0032] This may allow for efficient decoding of region information when there are two or more attribute videos, depending on whether the region information is identical between the attribute videos. Consequently, it may be possible to reduce the amount of coding and processing required.
[0033] Furthermore, the decoding method in Example 11 may be the decoding method in Example 9 or 10, wherein the region information is decoded when the number of attribute videos is 1, and the region information is not decoded when the number of attribute videos is 0.
[0034] This may allow for the omission of decoding region information depending on the number of attribute videos. Consequently, it may be possible to reduce the amount of coding and processing required.
[0035] Furthermore, the decoding method in Example 12 may be any of the decoding methods in Examples 9 to 11, wherein if the number of attribute videos is 1 or more, one or more attribute videos are decoded.
[0036] This may make it possible to decode attribute videos and region information simultaneously, depending on the number of attribute videos. Therefore, it may be possible to efficiently decode attribute videos and region information.
[0037] Furthermore, the decoding method in Example 13 may be any of the decoding methods in Examples 9 to 12, wherein if the number of attribute videos is two or more and the region information is the same among the attribute videos, an index for referencing the region information is decoded.
[0038] As a result, when the number of attribute videos is 2 or more, it may be possible to efficiently identify and refer to the same region information among the attribute videos. Therefore, it may be possible to efficiently define one or more regions in the image region of the attribute video.
[0039] Further, the decoding method of Example 14 may be the decoding method of Example 13, in which the index is not decoded when the number of attribute videos is 1 or less or the region information is not the same among the attribute videos.
[0040] As a result, it may be possible to omit the decoding of the index according to the number of attribute videos and whether or not the region information is the same among the attribute videos. Therefore, it may be possible to reduce the amount of code and the processing amount.
[0041] Further, the decoding method of Example 15 may be the decoding method of any of Examples 9 to 14, in which the flag is decoded according to the number of attribute videos for each layer of the three-dimensional data.
[0042] As a result, it may be possible to omit the decoding of the flag indicating whether or not the region information is the same among the attribute videos according to the number of attribute videos for each layer. Therefore, it may be possible to reduce the amount of code and the processing amount.
[0043] Further, the decoding method of Example 16 may be the decoding method of any of Examples 9 to 15, in which the region information is tile information for defining one or more tiles in the image region.
[0044] As a result, it may be possible to omit the decoding of the flag indicating whether or not the tile information is the same among the attribute videos according to the number of attribute videos. Therefore, it may be possible to reduce the amount of code and the processing amount.
[0045] Further, the encoding device of Example 17 includes a circuit and a memory accessible by the circuit. In operation, the circuit encodes a flag indicating whether region information for defining one or more regions in the image region of an attribute video is the same among attribute videos according to the number of attribute videos to be encoded in three-dimensional data. When the number of attribute videos is 2 or more, the flag is encoded. When the number of attribute videos is 1 or less, the flag is not encoded.
[0046] Thus, depending on the number of attribute videos, it may be possible to omit encoding of a flag indicating whether region information is the same among attribute videos. Therefore, it may be possible to reduce the amount of coding and the processing amount.
[0047] Further, the decoding device of Example 18 includes a circuit and a memory accessible by the circuit. In operation, the circuit decodes a flag indicating whether region information for defining one or more regions in the image region of an attribute video is the same among attribute videos according to the number of attribute videos to be decoded in three-dimensional data. When the number of attribute videos is 2 or more, the flag is decoded. When the number of attribute videos is 1 or less, the flag is not decoded.
[0048] Thus, depending on the number of attribute videos, it may be possible to omit decoding of a flag indicating whether region information is the same among attribute videos. Therefore, it may be possible to reduce the amount of coding and the processing amount.
[0049] Furthermore, these general or specific aspects may be implemented by a system, a device, a method, an integrated circuit, a computer program, or a non-temporary recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, a device, a method, an integrated circuit, a computer program, and a recording medium.
[0050] <Expressions and Terms> Here, the following expressions and terms are used.
[0051] (1) Three-dimensional mesh A three-dimensional mesh is a collection of multiple faces, for example, representing a three-dimensional object. A three-dimensional mesh mainly consists of vertex information, connection information, and attribute information. A three-dimensional mesh may be expressed as a polygon mesh or a mesh. A three-dimensional mesh may also have temporal changes. A three-dimensional mesh may contain metadata related to vertex information, connection information, and attribute information, as well as other additional information.
[0052] (2) Vertex Information Vertex information is information that indicates a vertex. For example, vertex information indicates the position of a vertex in three-dimensional space. Also, a vertex corresponds to the vertices of the faces that make up a three-dimensional mesh. Vertex information is sometimes expressed as "Geometry". Vertex information is also sometimes expressed as position information.
[0053] (3) Connection Information Connection information is information that indicates the connections between vertices. For example, connection information indicates the connections that make up the faces or edges of a three-dimensional mesh. Connection information is sometimes expressed as "Connectivity". Connection information is also sometimes expressed as face information.
[0054] (4) Attribute Information Attribute information is information that indicates the attributes of a vertex or face. For example, attribute information indicates attributes such as color, image, and normal vector associated with a vertex or face. Attribute information is sometimes expressed as "Texture".
[0055] (5) A face is an element that makes up a three-dimensional mesh. Specifically, a face is a polygon on a plane in three-dimensional space. For example, a face can be defined as a triangle in three-dimensional space.
[0056] (6) A plane is a two-dimensional plane in three-dimensional space. For example, a polygon is formed on a plane, and multiple polygons are formed on multiple planes.
[0057] (7) Bitstream A bitstream corresponds to encoded information. A bitstream can also be expressed as a stream, encoded bitstream, compressed bitstream, or encoded signal.
[0058] (8) The expressions "encode" and "decode" may be replaced with expressions such as "store," "include," "write," "describe," "signalize," "transmit," "notify," "save," or "compress," and these expressions may be interchangeable. For example, encoding information may mean including information in a bitstream. Also, encoding information into a bitstream may mean encoding information and generating a bitstream that contains the encoded information.
[0059] Furthermore, the expression "decode" may be replaced with expressions such as "read out," "decipher," "read," "load," "derive," "obtain," "receive," "extract," "restore," "reconstruct," "decompress," or "expand," and these expressions may be interchangeable. For example, decoding information may mean obtaining information from a bitstream. Also, decoding information from a bitstream may mean decoding the bitstream and obtaining the information contained in the bitstream.
[0060] (9) In the explanation of ordinal numbers, first and second ordinal numbers may be assigned to components, etc. These ordinal numbers may be rearranged as appropriate. Ordinal numbers may also be newly assigned to components, etc., or removed. Furthermore, these ordinal numbers may be assigned to elements in order to identify them, and may not correspond to a meaningful order.
[0061] <Three-Dimensional Mesh> Figure 1 is a conceptual diagram showing a three-dimensional mesh according to this embodiment. The three-dimensional mesh is composed of multiple faces. For example, each face is a triangle. The vertices of these triangles are defined in three-dimensional space. The three-dimensional mesh represents a three-dimensional object. Each face may have a color or image.
[0062] Figure 2 is a conceptual diagram showing the basic elements of a three-dimensional mesh according to this embodiment. The three-dimensional mesh consists of vertex information, connection information, and attribute information. Vertex information indicates the position of the vertices of a face in three-dimensional space. Connection information indicates the connections between vertices. A face can be identified by the vertex information and connection information. In other words, a colorless three-dimensional object is formed in three-dimensional space by the vertex information and connection information.
[0063] Attribute information may be associated with vertices or faces. Attribute information associated with vertices may be expressed as "Attribute Per Point". Attribute information associated with vertices may indicate the attributes of the vertex itself or the attributes of the faces connected to the vertex.
[0064] For example, a color may be associated with a vertex as attribute information. The color associated with a vertex may be the color of the vertex itself, or the color of the face connected to the vertex. The color of a face may be the average of multiple colors associated with multiple vertices of the face. In addition, a normal vector may be associated with a vertex or face as attribute information. Such a normal vector can represent the front and back of a face.
[0065] Furthermore, a two-dimensional image may be associated with a surface as attribute information. The two-dimensional image associated with a surface is also referred to as a texture image or "Attribute Map". Additionally, information indicating the mapping between the surface and the two-dimensional image may be associated with the surface as attribute information. Such information indicating the mapping may be referred to as mapping information, vertex information of the texture image, texture coordinates, or "Attribute UV Coordinate".
[0066] Furthermore, information such as color, images, and moving images used as attribute information may be expressed as "Parametric Space."
[0067] Such attribute information can be used to reflect textures onto three-dimensional objects. In other words, vertex information, connection information, and attribute information allow a three-dimensional object with color to be formed in three-dimensional space.
[0068] In the above, attribute information is associated with vertices or faces, but it may also be associated with edges.
[0069] Figure 3 is a conceptual diagram illustrating the mapping according to this embodiment. For example, a region of a two-dimensional image in a two-dimensional plane can be mapped to a surface of a three-dimensional mesh in three-dimensional space. Specifically, coordinate information of a region in a two-dimensional image is associated with a surface of a three-dimensional mesh. As a result, the image of the region mapped in the two-dimensional image is reflected on the surface of the three-dimensional mesh.
[0070] By using mapping, a two-dimensional image used as attribute information can be separated from a three-dimensional mesh. For example, in encoding a three-dimensional mesh, the two-dimensional image may be encoded using an image encoding scheme or a video encoding scheme.
[0071] <System Configuration> Figure 4 is a block diagram showing an example configuration of the encoding and decoding system according to this embodiment. In Figure 4, the encoding and decoding system comprises an encoding device 100 and a decoding device 200.
[0072] For example, the encoding device 100 acquires a three-dimensional mesh and encodes the three-dimensional mesh into a bitstream. The encoding device 100 then outputs the bitstream to the network 300. For example, the bitstream includes the encoded three-dimensional mesh and control information for decoding the encoded three-dimensional mesh. By encoding the three-dimensional mesh, the information of the three-dimensional mesh is compressed.
[0073] Network 300 transmits the bitstream from the encoding device 100 to the decoding device 200. Network 300 may be the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. Network 300 is not necessarily limited to bidirectional communication; it may also be a one-way communication network for terrestrial digital broadcasting or satellite broadcasting, etc.
[0074] Furthermore, the network 300 may be replaced by recording media such as DVD (Digital Versatile Disc) or BD (Blu-Ray Disc®).
[0075] The decoding device 200 acquires the bitstream and decodes the three-dimensional mesh from the bitstream. The decoding of the three-dimensional mesh expands the information of the three-dimensional mesh. For example, the decoding device 200 decodes the three-dimensional mesh according to a decoding method that corresponds to the encoding method used by the encoding device 100 to encode the three-dimensional mesh. That is, the encoding device 100 and the decoding device 200 perform encoding and decoding according to their respective encoding and decoding methods.
[0076] The three-dimensional mesh before encoding can also be referred to as the original three-dimensional mesh. Similarly, the three-dimensional mesh after decoding can be referred to as the reconstructed three-dimensional mesh.
[0077] <Encoding Device> Figure 5 is a block diagram showing an example configuration of the encoding device 100 according to this embodiment. For example, the encoding device 100 includes a vertex information encoder 101, a connection information encoder 102, and an attribute information encoder 103.
[0078] The vertex information encoder 101 is an electrical circuit that encodes vertex information. For example, the vertex information encoder 101 encodes vertex information into a bitstream according to a defined format for vertex information.
[0079] The connection information encoder 102 is an electrical circuit that encodes connection information. For example, the connection information encoder 102 encodes connection information into a bitstream according to a specified format for connection information.
[0080] The attribute information encoder 103 is an electrical circuit that encodes attribute information. For example, the attribute information encoder 103 encodes attribute information into a bitstream according to a format defined for the attribute information.
[0081] Variable-length coding or fixed-length coding may be used to encode vertex information, connection information, and attribute information. Variable-length coding may correspond to Huffman coding or context-adaptive binary arithmetic coding (CABAC), etc.
[0082] The vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be integrated into a single unit. Alternatively, each of the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be further subdivided into multiple components.
[0083] Figure 6 is a block diagram showing another configuration example of the encoding device 100 according to this embodiment. For example, the encoding device 100 includes a preprocessor 104 and a postprocessor 105 in addition to the configuration shown in Figure 5.
[0084] The preprocessor 104 is an electrical circuit that performs processing before encoding vertex information, connection information, and attribute information. For example, the preprocessor 104 may perform transformation processing, separation processing, or multiplexing processing on the three-dimensional mesh before encoding. More specifically, for example, the preprocessor 104 may separate vertex information, connection information, and attribute information from the three-dimensional mesh before encoding.
[0085] The post-processor 105 is an electrical circuit that performs processing after encoding the vertex information, connection information, and attribute information. For example, the post-processor 105 may perform conversion processing, separation processing, or multiplexing processing on the encoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 105 may multiplex the encoded vertex information, connection information, and attribute information into a bitstream. Alternatively, for example, the post-processor 105 may further perform variable-length encoding on the encoded vertex information, connection information, and attribute information.
[0086] <Decoding Device> Figure 7 is a block diagram showing an example configuration of the decoding device 200 according to this embodiment. For example, the decoding device 200 includes a vertex information decoder 201, a connection information decoder 202, and an attribute information decoder 203.
[0087] The vertex information decoder 201 is an electrical circuit that decodes vertex information. For example, the vertex information decoder 201 decodes vertex information from a bitstream according to a format defined for vertex information.
[0088] The connection information decoder 202 is an electrical circuit that decodes connection information. For example, the connection information decoder 202 decodes connection information from a bitstream according to a format defined for connection information.
[0089] The attribute information decoder 203 is an electrical circuit that decodes attribute information. For example, the attribute information decoder 203 decodes attribute information from a bitstream according to a format defined for attribute information.
[0090] Variable-length decoding or fixed-length decoding may be used for decoding vertex information, connection information, and attribute information. Variable-length decoding may correspond to Huffman coding or context-adaptive binary arithmetic coding (CABAC), etc.
[0091] The vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be integrated. Alternatively, each of the vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be subdivided into multiple components.
[0092] Figure 8 is a block diagram showing another configuration example of the decoding device 200 according to this embodiment. For example, the decoding device 200 includes a pre-processor 204 and a post-processor 205 in addition to the configuration shown in Figure 7.
[0093] The preprocessor 204 is an electrical circuit that performs processing before decoding the vertex information, connection information, and attribute information. For example, the preprocessor 204 may perform conversion processing, separation processing, or multiplexing processing on the bitstream before decoding the vertex information, connection information, and attribute information.
[0094] More specifically, for example, the preprocessor 204 may separate the bitstream into sub-bitstreams corresponding to vertex information, connection information, and attribute information. Alternatively, for example, the preprocessor 204 may perform variable-length decoding on the bitstream before decoding the vertex information, connection information, and attribute information.
[0095] The post-processor 205 is an electrical circuit that performs processing after decoding the vertex information, connection information, and attribute information. For example, the post-processor 205 may perform conversion processing, separation processing, or multiplexing processing on the decoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 205 may multiplex the decoded vertex information, connection information, and attribute information into a three-dimensional mesh.
[0096] <Bitstream> Vertex information, connection information, and attribute information are encoded and stored in the bitstream. The relationship between this information and the bitstream is shown below.
[0097] Figure 9 is a conceptual diagram showing an example of the bitstream configuration according to this embodiment. In this example, connection information, vertex information, and attribute information are integrated in the bitstream. For example, connection information, vertex information, and attribute information may be included in a single file.
[0098] Furthermore, multiple parts of this information may be stored sequentially, such as the first part of connection information, the first part of vertex information, the first part of attribute information, the second part of connection information, the second part of vertex information, the second part of attribute information, and so on. These multiple parts may correspond to multiple parts that are different in time, multiple parts that are different in space, or multiple faces that are different.
[0099] Furthermore, the storage order of connection information, vertex information, and attribute information is not limited to the example above, and a different storage order may be used.
[0100] Figure 10 is a conceptual diagram showing another example of the bitstream configuration according to this embodiment. In this example, multiple files are included in the bitstream, and connection information, vertex information, and attribute information are stored in different files. Here, a file containing connection information, a file containing vertex information, and a file containing attribute information are shown, but the storage format is not limited to this example. For example, two types of information from connection information, vertex information, and attribute information may be included in one file, and the remaining type of information may be included in another file.
[0101] Alternatively, this information may be divided and stored in more files. For example, multiple parts of connection information may be stored in multiple files, multiple parts of vertex information may be stored in multiple files, and multiple parts of attribute information may be stored in multiple files. These multiple parts may correspond to multiple parts that are different in time, multiple parts that are different in space, or multiple faces that are different.
[0102] Furthermore, the storage order of connection information, vertex information, and attribute information is not limited to the example above, and a different storage order may be used.
[0103] Figure 11 is a conceptual diagram showing another example of the bitstream configuration according to this embodiment. In this example, the bitstream is composed of multiple separable sub-bitstreams, and connection information, vertex information, and attribute information are stored in different sub-bitstreams.
[0104] Here, sub-bitstreams containing connection information, sub-bitstreams containing vertex information, and sub-bitstreams containing attribute information are shown, but the storage format is not limited to these examples.
[0105] For example, two types of information from connection information, vertex information, and attribute information may be included in one sub-bitstream, and the remaining type of information may be included in another sub-bitstream. Specifically, attribute information of a two-dimensional image, etc., may be stored in a sub-bitstream compliant with an image encoding scheme, separate from the sub-bitstreams of connection information and vertex information.
[0106] Furthermore, each sub-bitstream may contain multiple files. Multiple parts of connection information may be stored in multiple files, multiple parts of vertex information may be stored in multiple files, and multiple parts of attribute information may be stored in multiple files.
[0107] Furthermore, the storage order of connection information, vertex information, and attribute information is not limited to the examples in Figures 9, 10, and 11, and a different storage order may be used. For example, they may be stored in the bitstream in the order of vertex information, connection information, and attribute information. Alternatively, they may be stored in the bitstream in any of the following orders: connection information, attribute information, and vertex information; vertex information, attribute information, and connection information; attribute information, connection information, and vertex information; or attribute information, vertex information, and connection information.
[0108] Furthermore, connection information, vertex information, and attribute information may each be divided into multiple data points, and these multiple data points may be stored in a bitstream in a periodic or random order.
[0109] <Specific Example> Figure 12 is a block diagram showing a specific example of the encoding and decoding system according to this embodiment. In Figure 12, the encoding and decoding system comprises a three-dimensional data encoding system 110, a three-dimensional data decoding system 210, and an external connector 310.
[0110] The three-dimensional data encoding system 110 comprises a controller 111, an input / output processor 112, a three-dimensional data encoder 113, a three-dimensional data generator 115, and a system multiplexer 114. The three-dimensional data decoding system 210 comprises a controller 211, an input / output processor 212, a three-dimensional data decoder 213, a system demultiplexer 214, a presenter 215, and a user interface 216.
[0111] In the three-dimensional data encoding system 110, sensor data is input from the sensor terminal to the three-dimensional data generator 115. The three-dimensional data generator 115 generates three-dimensional data, such as point cloud data or mesh data, from the sensor data and inputs it to the three-dimensional data encoder 113.
[0112] For example, the 3D data generator 115 generates vertex information and connection information and attribute information corresponding to the vertex information. The 3D data generator 115 may process the vertex information when generating the connection information and attribute information. For example, the 3D data generator 115 may reduce the amount of data by deleting duplicate vertices or perform transformations on the vertex information (such as position shifting, rotation, or normalization). The 3D data generator 115 may also render the attribute information.
[0113] Furthermore, although the three-dimensional data generator 115 is a component of the three-dimensional data encoding system 110 in Figure 12, it may also be located externally and independently of the three-dimensional data encoding system 110.
[0114] The sensor terminal that provides sensor data for generating three-dimensional data may be, for example, a moving object such as an automobile, a flying object such as an airplane, a mobile terminal, or a camera. Alternatively, distance sensors such as LIDAR, millimeter-wave radar, infrared sensors, or rangefinders, stereo cameras, or combinations of multiple monocular cameras may be used as sensor terminals.
[0115] The sensor data may include the distance (position) of the object, monocular camera images, stereo camera images, color, reflectivity, sensor attitude, orientation, gyroscope, sensing position (GPS information or altitude), speed, acceleration, sensing time, temperature, atmospheric pressure, humidity, or magnetism.
[0116] The three-dimensional data encoder 113 corresponds to the encoding device 100 shown in Figure 5, etc. For example, the three-dimensional data encoder 113 encodes three-dimensional data and generates encoded data. The three-dimensional data encoder 113 also generates control information during the encoding of three-dimensional data. The three-dimensional data encoder 113 then inputs the encoded data along with the control information to the system multiplexer 114.
[0117] The encoding method for three-dimensional data may be a geometry-based encoding method or a video codec-based encoding method. Here, the geometry-based encoding method can also be referred to as a geometry-based encoding method. The video codec-based encoding method can also be referred to as a video-based encoding method.
[0118] The system multiplexer 114 multiplexes the encoded data and control information input from the three-dimensional data encoder 113 and generates multiplexed data using a predetermined multiplexing scheme. The system multiplexer 114 may also multiplex other media such as video, audio, subtitles, application data, or document files, or reference time information, along with the encoded data and control information of the three-dimensional data. Furthermore, the system multiplexer 114 may also multiplex sensor data or attribute information related to the three-dimensional data.
[0119] For example, the multiplexed data may have a file format for storage or a packet format for transmission. ISOBMFF or an ISOBMFF-based format may be used as these formats. Alternatively, MPEG-DASH, MMT, MPEG-2 TS Systems, or RTP may be used.
[0120] The multiplexed data is then output as a transmission signal to the external connector 310 by the input / output processor 112. The multiplexed data may be transmitted as a transmission signal by wire or by wireless. Alternatively, the multiplexed data may be stored in internal memory or storage device. The multiplexed data may be transmitted to a cloud server via the internet or stored in an external storage device.
[0121] For example, the transmission or storage of multiplexed data is carried out in a manner appropriate to the medium for transmission or storage, such as broadcasting or telecommunications. Communication protocols such as HTTP, FTP, TCP, UDP, IP, or combinations thereof may be used. Furthermore, either a pull-type or push-type communication method may be used.
[0122] For wired transmission, Ethernet®, USB, RS-232C, HDMI®, or coaxial cable may be used. For wireless transmission, 3GPP®, IEEE 3G / 4G / 5G, wireless LAN, Wi-Fi, Bluetooth, or millimeter wave may be used. Furthermore, as a broadcasting method, for example, DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 may be used.
[0123] The sensor data may also be input to the three-dimensional data generator 115 or the system multiplexer 114. Alternatively, the three-dimensional data or encoded data may be output directly as a transmission signal to the external connector 310 via the input / output processor 112. The transmission signal output from the three-dimensional data encoding system 110 is input to the three-dimensional data decoding system 210 via the external connector 310.
[0124] Furthermore, each operation of the three-dimensional data encoding system 110 may be controlled by a controller 111 that executes an application program.
[0125] In the three-dimensional data decoding system 210, the transmission signal is input to the input / output processor 212. The input / output processor 212 decodes the transmission signal into multiplexed data in file format or packet format and inputs the multiplexed data to the system demultiplexer 214. The system demultiplexer 214 obtains encoded data and control information from the multiplexed data and inputs them to the three-dimensional data decoder 213. The system demultiplexer 214 may also extract other media or reference time information from the multiplexed data.
[0126] The three-dimensional data decoder 213 corresponds to the decoding device 200 shown in Figure 7, etc. For example, the three-dimensional data decoder 213 decodes three-dimensional data from encoded data based on a predetermined encoding scheme. The three-dimensional data is then presented to the user by the presenter 215.
[0127] In addition, additional information such as sensor data may be input to the display device 215. The display device 215 may present three-dimensional data based on the additional information. Furthermore, user instructions may be input from the user terminal to the user interface 216. The display device 215 may then present three-dimensional data based on the input instructions.
[0128] The input / output processor 212 may also acquire three-dimensional data and encoded data from the external connector 310.
[0129] Furthermore, each operation of the three-dimensional data decoding system 210 may be controlled by a controller 211 that executes an application program.
[0130] Figure 13 is a conceptual diagram showing an example of the configuration of point cloud data according to this embodiment. Point cloud data is data representing a three-dimensional object.
[0131] Specifically, a point cloud consists of multiple points and contains positional information indicating the three-dimensional coordinate position of each point, as well as attribute information indicating the attributes of each point. Positional information is also expressed as geometry.
[0132] The type of attribute information may be, for example, color or reflectance. A single point may be associated with attribute information relating to one type, a single point may be associated with attribute information relating to multiple different types, or a single point may be associated with attribute information having multiple values for the same type.
[0133] Figure 14 is a conceptual diagram showing an example of a data file for point cloud data according to this embodiment. In this example, there is a one-to-one correspondence between location information items and attribute information items, and it shows the location information and attribute information of N points that constitute the point cloud data. In this example, the location information is information that indicates the three-dimensional coordinate position on the three axes x, y, and z, and the attribute information is information that indicates the color in RGB. A PLY file or the like can be used as a typical data file for point cloud data.
[0134] Figure 15 is a conceptual diagram showing an example of the configuration of mesh data according to this embodiment. Mesh data is data used in CG (Computer Graphics), etc., and is three-dimensional mesh data that shows the three-dimensional shape of an object with multiple faces. Each face is also represented as a polygon and has the shape of a polygon such as a triangle or quadrilateral.
[0135] Specifically, a three-dimensional mesh consists of multiple points that make up a point cloud, as well as multiple edges and multiple faces. Each point can also be expressed as a vertex or position. Each edge corresponds to a line segment connected by two vertices. Each face corresponds to a region enclosed by three or more edges.
[0136] Furthermore, a three-dimensional mesh contains positional information indicating the three-dimensional coordinate positions of its vertices. This positional information is also referred to as vertex information or geometry. A three-dimensional mesh also contains connection information indicating the relationships between multiple vertices that constitute an edge or face. This connection information is also referred to as connectivity. Finally, a three-dimensional mesh contains attribute information indicating the attributes of vertices, edges, or faces. This attribute information in a three-dimensional mesh is also referred to as texture.
[0137] For example, attribute information may indicate the color, reflectance, or normal vector for a vertex, edge, or face. The direction of the normal vector can represent the front and back of the face.
[0138] Object files or similar formats may be used as the data file format for mesh data.
[0139] Figure 16 is a conceptual diagram showing an example of a data file for mesh data according to this embodiment. In this example, the data file includes position information G(1) to G(N) for N vertices constituting the three-dimensional mesh, and attribute information A1(1) to A1(N) for the N vertices. In addition, this example includes M pieces of attribute information A2(1) to A2(M). The items of the attribute information do not have to correspond one-to-one with vertices, nor do they have to correspond one-to-one with faces. Furthermore, attribute information may not exist at all.
[0140] Connection information is indicated by a combination of vertex indices. n[1, 3, 4] represents a triangular face composed of three vertices n=1, n=3, and n=4. Also, m[2, 4, 6] indicates that attribute information m=2, m=4, and m=6 correspond to the three vertices, respectively.
[0141] Furthermore, the actual content of the attribute information may be described in a separate file. A pointer to that content may be associated with a vertex or face, etc. For example, attribute information indicating an image for a face may be stored in a two-dimensional attribute map file. The file name of the attribute map and the two-dimensional coordinate values in the attribute map may be described in attribute information A2(1) to A2(M). The method of specifying attribute information for a face is not limited to these methods, and any method may be used.
[0142] Figure 17 is a conceptual diagram showing the types of three-dimensional data according to this embodiment. Point cloud data and mesh data may represent static objects or dynamic objects. Static objects are objects that do not change over time, while dynamic objects are objects that change over time. Static objects may correspond to three-dimensional data for any given point in time.
[0143] For example, point cloud data for any given time point may be referred to as a PCC frame. Similarly, mesh data for any given time point may be referred to as a mesh frame. Furthermore, both PCC frames and mesh frames may simply be referred to as frames.
[0144] Furthermore, the object's area may be limited to a certain range, like in regular video data, or it may not be limited, like in map data. Also, the density of points or surfaces can be determined in various ways. Sparse point cloud data or sparse mesh data may be used, or dense point cloud data or dense mesh data may be used.
[0145] Next, the encoding and decoding of point clouds or three-dimensional meshes will be described. The apparatus, processing, or syntax for encoding and decoding vertex information of a three-dimensional mesh in this disclosure may be applied to the encoding and decoding of point clouds. The apparatus, processing, or syntax for encoding and decoding of point clouds in this disclosure may be applied to the encoding and decoding of vertex information of a three-dimensional mesh.
[0146] Furthermore, the apparatus, processing, or syntax for encoding and decoding point cloud attribute information in this disclosure may also be applied to the encoding and decoding of connection information or attribute information of a three-dimensional mesh.
[0147] Furthermore, at least some of the processing may be shared between the encoding and decoding of point cloud data and the encoding and decoding of mesh data. This can reduce the size of the circuit and software programs.
[0148] Figure 18 is a block diagram showing an example configuration of a three-dimensional data encoder 113 according to this embodiment. In this example, the three-dimensional data encoder 113 comprises a vertex information encoder 121, an attribute information encoder 122, a metadata encoder 123, and a multiplexer 124. The vertex information encoder 121, the attribute information encoder 122, and the multiplexer 124 may correspond to the vertex information encoder 101, the attribute information encoder 103, and the post-processor 105 in Figure 6, etc.
[0149] In this example, the three-dimensional data encoder 113 encodes the three-dimensional data according to a geometry-based encoding scheme. In encoding according to a geometry-based encoding scheme, the three-dimensional structure is taken into consideration. In addition, in encoding according to a geometry-based encoding scheme, attribute information is encoded using the configuration information obtained in encoding the vertex information.
[0150] Specifically, first, vertex information, attribute information, and metadata contained in the three-dimensional data generated from sensor data are input to the vertex information encoder 121, attribute information encoder 122, and metadata encoder 123, respectively. Here, connection information contained in the three-dimensional data may be treated in the same way as attribute information. In the case of point cloud data, position information may be treated as vertex information.
[0151] The vertex information encoder 121 encodes the vertex information into compressed vertex information and outputs the compressed vertex information as encoded data to the multiplexer 124. The vertex information encoder 121 also generates metadata for the compressed vertex information and outputs it to the multiplexer 124. Furthermore, the vertex information encoder 121 generates configuration information and outputs it to the attribute information encoder 122.
[0152] The attribute information encoder 122 uses the configuration information generated by the vertex information encoder 121 to encode the attribute information into compressed attribute information and outputs the compressed attribute information as encoded data to the multiplexer 124. The attribute information encoder 122 also generates metadata for the compressed attribute information and outputs it to the multiplexer 124.
[0153] The metadata encoder 123 encodes compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The metadata encoded by the metadata encoder 123 may be used for encoding vertex information and attribute information.
[0154] The multiplexer 124 multiplexes the compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, the metadata of the compressed attribute information, and the compressed metadata into a bitstream. The multiplexer 124 then inputs the bitstream to the system layer.
[0155] Figure 19 is a block diagram showing an example configuration of a three-dimensional data decoder 213 according to this embodiment. In this example, the three-dimensional data decoder 213 includes a vertex information decoder 221, an attribute information decoder 222, a metadata decoder 223, and a demultiplexer 224. The vertex information decoder 221, attribute information decoder 222, and demultiplexer 224 may correspond to the vertex information decoder 201, attribute information decoder 203, and preprocessor 204 in Figure 8, etc.
[0156] In this example, the three-dimensional data decoder 213 decodes the three-dimensional data according to a geometry-based coding scheme. In decoding according to a geometry-based coding scheme, the three-dimensional structure is taken into consideration. In addition, in decoding according to a geometry-based coding scheme, attribute information is decoded using the configuration information obtained in the decoding of vertex information.
[0157] Specifically, first, the bitstream is input from the system layer to the demultiplexer 224. The demultiplexer 224 separates the compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, the metadata of the compressed attribute information, and the compressed metadata from the bitstream. The compressed vertex information and the metadata of the compressed vertex information are input to the vertex information decoder 221. The compressed attribute information and the metadata of the compressed attribute information are input to the attribute information decoder 222. The metadata is input to the metadata decoder 223.
[0158] The vertex information decoder 221 decodes vertex information from compressed vertex information using the metadata of the compressed vertex information. The vertex information decoder 221 also generates configuration information and outputs it to the attribute information decoder 222. The attribute information decoder 222 decodes attribute information from compressed attribute information using the configuration information generated by the vertex information decoder 221 and the metadata of the compressed attribute information. The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 may be used for decoding vertex information and attribute information.
[0159] Subsequently, vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. For example, this metadata is metadata for vertex information and attribute information, and can be used in application programs.
[0160] Figure 20 is a block diagram showing another configuration example of the three-dimensional data encoder 113 according to this embodiment. In this example, the three-dimensional data encoder 113 includes a vertex image generator 131, an attribute image generator 132, a metadata generator 133, a video encoder 134, a metadata encoder 123, and a multiplexer 124. The vertex image generator 131, the attribute image generator 132, and the video encoder 134 may correspond to the vertex information encoder 101 and the attribute information encoder 103 in Figure 6, etc.
[0161] In this example, the three-dimensional data encoder 113 encodes the three-dimensional data according to a video-based encoding scheme. In encoding according to a video-based encoding scheme, multiple two-dimensional images are generated from the three-dimensional data, and these multiple two-dimensional images are encoded according to a video encoding scheme. Here, the video encoding scheme may be HEVC (High Efficiency Video Coding) or VVC (Versatile Video Coding), etc.
[0162] Specifically, first, vertex information and attribute information contained in the three-dimensional data generated from sensor data are input to the metadata generator 133. The vertex information and attribute information are then input to the vertex image generator 131 and the attribute image generator 132, respectively. Furthermore, metadata contained in the three-dimensional data is input to the metadata encoder 123. Here, connection information contained in the three-dimensional data may be treated similarly to attribute information. In the case of point cloud data, position information may be treated as vertex information.
[0163] The metadata generator 133 generates map information for multiple two-dimensional images from vertex information and attribute information. The metadata generator 133 then inputs the map information to the vertex image generator 131, the attribute image generator 132, and the metadata encoder 123.
[0164] The vertex image generator 131 generates a vertex image based on the vertex information and map information and inputs it to the video encoder 134. The attribute image generator 132 generates an attribute image based on the attribute information and map information and inputs it to the video encoder 134.
[0165] The video encoder 134 encodes the vertex image and attribute image into compressed vertex information and compressed attribute information, respectively, according to the video encoding scheme, and outputs the compressed vertex information and compressed attribute information as encoded data to the multiplexer 124. The video encoder 134 also generates metadata for the compressed vertex information and metadata for the compressed attribute information and outputs them to the multiplexer 124.
[0166] The metadata encoder 123 encodes compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The compressible metadata includes map information. The metadata encoded by the metadata encoder 123 may also be used for encoding vertex information and attribute information.
[0167] The multiplexer 124 multiplexes the compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, the metadata of the compressed attribute information, and the compressed metadata into a bitstream. The multiplexer 124 then inputs the bitstream to the system layer.
[0168] Figure 21 is a block diagram showing another configuration example of the three-dimensional data decoder 213 according to this embodiment. In this example, the three-dimensional data decoder 213 includes a vertex information generator 231, an attribute information generator 232, a video decoder 234, a metadata decoder 223, and a demultiplexer 224. The vertex information generator 231, the attribute information generator 232, and the video decoder 234 may correspond to the vertex information decoder 201 and the attribute information decoder 203 in Figure 8, etc.
[0169] In this example, the three-dimensional data decoder 213 decodes the three-dimensional data according to a video-based coding scheme. In decoding according to a video-based coding scheme, multiple two-dimensional images are decoded according to the video coding scheme, and three-dimensional data is generated from the multiple two-dimensional images. Here, the video coding scheme may be HEVC (High Efficiency Video Coding) or VVC (Versatile Video Coding), etc.
[0170] Specifically, first, the bitstream is input from the system layer to the demultiplexer 224. The demultiplexer 224 separates the compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, the metadata of the compressed attribute information, and the compressed metadata from the bitstream. The compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, and the metadata of the compressed attribute information are input to the video decoder 234. The compressed metadata is input to the metadata decoder 223.
[0171] The video decoder 234 decodes the vertex image according to the video encoding scheme. In doing so, the video decoder 234 decodes the vertex image from the compressed vertex information using the metadata of the compressed vertex information. The video decoder 234 then inputs the vertex image to the vertex information generator 231. The video decoder 234 also decodes the attribute image according to the video encoding scheme. In doing so, the video decoder 234 decodes the attribute image from the compressed attribute information using the metadata of the compressed attribute information. The video decoder 234 then inputs the attribute image to the attribute information generator 232.
[0172] The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 includes map information used for generating vertex information and attribute information. The metadata decoded by the metadata decoder 223 may also be used for decoding vertex images and attribute images.
[0173] The vertex information generator 231 reconstructs vertex information from the vertex image according to the map information contained in the metadata decoded by the metadata decoder 223. The attribute information generator 232 reconstructs attribute information from the attribute image according to the map information contained in the metadata decoded by the metadata decoder 223.
[0174] Subsequently, vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. For example, this metadata is metadata for vertex information and attribute information, and can be used in application programs.
[0175] Figure 22 is a conceptual diagram showing a specific example of the encoding process according to this embodiment. Figure 22 shows a three-dimensional data encoder 113 and a description encoder 148. In this example, the three-dimensional data encoder 113 comprises a two-dimensional data encoder 141 and a mesh data encoder 142. The two-dimensional data encoder 141 comprises a texture encoder 143. The mesh data encoder 142 comprises a vertex information encoder 144 and a connectivity information encoder 145.
[0176] The vertex information encoder 144, the connection information encoder 145, and the texture encoder 143 may correspond to the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 in Figure 6, etc.
[0177] For example, the two-dimensional data encoder 141 operates as a texture encoder 143 and generates a texture file by encoding the texture corresponding to the attribute information as two-dimensional data according to an image encoding scheme or video encoding scheme.
[0178] Furthermore, the mesh data encoder 142 operates as a vertex information encoder 144 and a connection information encoder 145, and generates a mesh file by encoding vertex information and connection information. The mesh data encoder 142 may also encode mapping information for textures. The encoded mapping information may then be included in the mesh file.
[0179] Furthermore, the description encoder 148 generates a description file by encoding the description corresponding to metadata such as text data. The description encoder 148 may encode the description at the system layer. For example, the description encoder 148 may be included in the system multiplexer 114 in Figure 12.
[0180] The above operation generates a bitstream containing texture files, mesh files, and description files. These files may be multiplexed into the bitstream in a file format such as glTF (Graphics Language Transmission Format) or USD (Universal Scene Description).
[0181] The three-dimensional data encoder 113 may also include two mesh data encoders, namely the mesh data encoder 142. For example, one mesh data encoder encodes the vertex and connection information of a static three-dimensional mesh, while the other mesh data encoder encodes the vertex and connection information of a dynamic three-dimensional mesh.
[0182] Correspondingly, two mesh files may be included in the bitstream. For example, one mesh file may correspond to a static 3D mesh, and the other mesh file may correspond to a dynamic 3D mesh.
[0183] Furthermore, a static three-dimensional mesh may be a three-dimensional mesh of an intraframe encoded using intraprediction, and a dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe encoded using interprediction. In addition, as information for the dynamic three-dimensional mesh, the difference information between the vertex information or connection information of the intraframe three-dimensional mesh and the vertex information or connection information of the interframe three-dimensional mesh may be used.
[0184] Figure 23 is a conceptual diagram showing a specific example of the decoding process according to this embodiment. Figure 23 shows a three-dimensional data decoder 213, a description decoder 248, and an presenter 247. In this example, the three-dimensional data decoder 213 comprises a two-dimensional data decoder 241, a mesh data decoder 242, and a mesh reconstructor 246. The two-dimensional data decoder 241 comprises a texture decoder 243. The mesh data decoder 242 comprises a vertex information decoder 244 and a connection information decoder 245.
[0185] The vertex information decoder 244, connection information decoder 245, texture decoder 243, and mesh reconstructor 246 may correspond to the vertex information decoder 201, connection information decoder 202, attribute information decoder 203, and post-processor 205, etc., shown in Figure 8. The presenter 247 may correspond to the presenter 215, etc., shown in Figure 12.
[0186] For example, the two-dimensional data decoder 241 operates as a texture decoder 243 and decodes the texture corresponding to the attribute information from the texture file as two-dimensional data according to the image encoding scheme or video encoding scheme.
[0187] Furthermore, the mesh data decoder 242 operates as a vertex information decoder 244 and a connection information decoder 245, decoding vertex information and connection information from the mesh file. The mesh data decoder 242 may also decode mapping information to textures from the mesh file.
[0188] Furthermore, the description decoder 248 decodes the description corresponding to metadata such as text data from the description file. The description decoder 248 may decode the description at the system layer. For example, the description decoder 248 may be included in the system demultiplexer 214 in Figure 12.
[0189] The mesh reconstructor 246 reconstructs a three-dimensional mesh from vertex information, connection information, and textures according to the description. The presenter 247 renders and outputs the three-dimensional mesh according to the description.
[0190] The above process reconstructs and outputs a 3D mesh from a bitstream containing texture files, mesh files, and description files.
[0191] The three-dimensional data decoder 213 may also include two mesh data decoders, which are mesh data decoders 242. For example, one mesh data decoder decodes the vertex and connection information of a static three-dimensional mesh, while the other mesh data decoder decodes the vertex and connection information of a dynamic three-dimensional mesh.
[0192] Correspondingly, two mesh files may be included in the bitstream. For example, one mesh file may correspond to a static 3D mesh, and the other mesh file may correspond to a dynamic 3D mesh.
[0193] Furthermore, a static three-dimensional mesh may be a three-dimensional mesh of an intraframe encoded using intraprediction, and a dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe encoded using interprediction. In addition, as information for the dynamic three-dimensional mesh, the difference information between the vertex information or connection information of the intraframe three-dimensional mesh and the vertex information or connection information of the interframe three-dimensional mesh may be used.
[0194] A coding scheme for dynamic three-dimensional meshes is sometimes called DMC (Dynamic Mesh Coding). Similarly, a video-based coding scheme for dynamic three-dimensional meshes is sometimes called V-DMC (Video-based Dynamic Mesh Coding).
[0195] The encoding method for point clouds is sometimes called PCC (Point Cloud Compression). The video-based encoding method for point clouds is sometimes called V-PCC (Video-based Point Cloud Compression). The geometry-based encoding method for point clouds is sometimes called G-PCC (Geometry-based Point Cloud Compression).
[0196] <Implementation Example> Figure 24 is a block diagram showing an implementation example of the encoding device 100 according to this embodiment. The encoding device 100 includes a circuit 151 and a memory 152. For example, the multiple components of the encoding device 100 shown in Figure 5, etc., are implemented by the circuit 151 and memory 152 shown in Figure 24.
[0197] Circuit 151 is an information processing circuit and is a circuit that can access memory 152. For example, circuit 151 is a dedicated or general-purpose electrical circuit for encoding a three-dimensional mesh. Circuit 151 may also be a processor such as a CPU. Alternatively, circuit 151 may be a collection of multiple electrical circuits.
[0198] Memory 152 is a dedicated or general-purpose memory in which information for the circuit 151 to encode a three-dimensional mesh is stored. Memory 152 may be an electrical circuit, or it may be connected to circuit 151. Memory 152 may also be included in circuit 151. Memory 152 may also be a collection of multiple electrical circuits. Memory 152 may also be a magnetic disk or an optical disk, or it may be described as storage or a recording medium. Memory 152 may also be a non-volatile memory or a volatile memory.
[0199] For example, memory 152 may store a three-dimensional mesh or a bitstream. Alternatively, memory 152 may store a program for circuit 151 to encode the three-dimensional mesh.
[0200] Furthermore, not all of the multiple components shown in Figure 5, etc., are to be implemented in the encoding device 100, nor are all of the multiple processes shown herein to be performed. Some of the multiple components shown in Figure 5, etc., may be included in other devices, and some of the multiple processes shown herein may be executed by other devices. In addition, the multiple components of this disclosure may be implemented in any combination in the encoding device 100, and the multiple processes of this disclosure may be performed in any combination.
[0201] Figure 25 is a block diagram showing an example of the implementation of the decoding device 200 according to this embodiment. The decoding device 200 includes a circuit 251 and a memory 252. For example, the multiple components of the decoding device 200 shown in Figure 7, etc., are implemented by the circuit 251 and memory 252 shown in Figure 25.
[0202] Circuit 251 is an information processing circuit and is a circuit that can access memory 252. For example, circuit 251 is a dedicated or general-purpose electrical circuit for decoding a three-dimensional mesh. Circuit 251 may also be a processor such as a CPU. Alternatively, circuit 251 may be a collection of multiple electrical circuits.
[0203] Memory 252 is a dedicated or general-purpose memory that stores information for circuit 251 to decode the three-dimensional mesh. Memory 252 may be an electrical circuit and may be connected to circuit 251. Memory 252 may also be included in circuit 251. Memory 252 may also be a collection of multiple electrical circuits. Memory 252 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory 252 may also be a non-volatile memory or a volatile memory.
[0204] For example, memory 252 may store a three-dimensional mesh or a bitstream. Alternatively, memory 252 may store a program for circuit 251 to decode the three-dimensional mesh.
[0205] Furthermore, the decoding device 200 does not need to implement all of the components shown in Figure 7, etc., nor does it need to perform all of the processes shown herein. Some of the components shown in Figure 7, etc., may be included in other devices, and some of the processes shown herein may be performed by other devices. In addition, the decoding device 200 may implement the components of this disclosure in any combination, and the processes of this disclosure may be performed in any combination.
[0206] The encoding and decoding methods, including the steps performed by each component of the encoding device 100 and decoding device 200 of this disclosure, may be performed by any device or system. For example, part or all of the encoding and decoding methods may be performed by a computer equipped with a processor, memory, and input / output circuits, etc. In this case, the encoding and decoding methods may be performed by the computer executing a program that causes the computer to perform the encoding and decoding methods.
[0207] Furthermore, a non-temporary computer-readable recording medium such as a CD-ROM may contain either a program or a bitstream.
[0208] An example of a program may be a bitstream. For instance, a bitstream containing an encoded three-dimensional mesh may include syntax elements to cause the decoding device 200 to decode the three-dimensional mesh. The bitstream then causes the decoding device 200 to decode the three-dimensional mesh according to the syntax elements contained within the bitstream. Therefore, a bitstream can perform a similar role to a program.
[0209] The above bitstream may be an encoded bitstream containing an encoded three-dimensional mesh, or a multiplexed bitstream containing an encoded three-dimensional mesh and other information.
[0210] Furthermore, each component of the encoding device 100 and the decoding device 200 may be made of dedicated hardware, general-purpose hardware that executes the above-mentioned program, or a combination thereof. The general-purpose hardware may also consist of a memory on which the program is stored, and a general-purpose processor that reads the program from the memory and executes it. Here, the memory may be semiconductor memory or a hard disk, and the general-purpose processor may be a CPU.
[0211] Furthermore, dedicated hardware may consist of memory and a dedicated processor, etc. For example, a dedicated processor may refer to memory for recording data and execute an encoding method and a decoding method.
[0212] Furthermore, each component of the encoding device 100 and the decoding device 200 may be an electrical circuit, as described above. These electrical circuits may form a single electrical circuit as a whole, or they may be separate electrical circuits. These electrical circuits may correspond to dedicated hardware, or they may correspond to general-purpose hardware that executes the above-mentioned program, etc. Furthermore, the encoding device 100 and the decoding device 200 may be implemented as integrated circuits.
[0213] Furthermore, the encoding device 100 may be a transmitting device that transmits a three-dimensional mesh. The decoding device 200 may be a receiving device that receives a three-dimensional mesh.
[0214] <Encoding and Decoding of Displacements> Here, the following terms are used as examples.
[0215] (1) Image An image is a data unit composed of a collection of pixels, and includes a picture or a block smaller than a picture. Images include both moving images and still images.
[0216] (2) A picture is an image processing unit composed of a collection of pixels, and is also called a frame or field.
[0217] (3) A block is a processing unit consisting of a specific number of pixels. The term block is also used in the examples shown below. The shape of a block is not particularly limited. A block may be, for example, a rectangular shape of M x N pixels, or a square shape of M x M pixels. A block may also be triangular, circular, or have other shapes. Examples of blocks are as follows.
[0218] - Slices, tiles, or bricks - CTU, superblock, or basic partitioning unit - VPDU, hardware processing partitioning unit - CU, processing block unit, prediction block unit (PU), or orthogonal transformation block unit (TU) - subblocks
[0219] (4) A pixel or sample is the smallest point, or in other words, the smallest unit, of an image. A pixel or sample includes not only pixels at integer positions but also pixels at sub-pixel positions generated from pixels at integer positions.
[0220] (5) Pixel values or sample values Pixel values or sample values are eigenvalues of a pixel. Pixel values or sample values may include luma values, chroma values, or RGB tonal levels, and may also include depth values or binary values of 0 or 1.
[0221] (6) Flags A flag represents one or more bits. A flag is, for example, a parameter or index represented by two or more bits. A flag may also represent a value that is not represented by a binary number, but by a non-binary number.
[0222] (7) A signal is a symbol or encoded information used to transmit information. Signals include discrete digital signals or continuous analog signals.
[0223] (8) Stream or Bitstream A stream or bitstream is a sequence of digital data that represents a flow of digital data. A stream or bitstream may be a single stream or may consist of multiple streams having multiple layers. A stream or bitstream may be transmitted by serial communication using a single transmission path or by packet communication using multiple transmission paths.
[0224] (9) Differences: In the case of scalar quantities, differences can include simple differences (x - y) and difference calculations. Differences can include the absolute value of the difference (|x - y|), the square of the difference (x^2 - y^2), the square root of the difference (√(x - y)), weighted differences (ax - by, where a and b are constants), or offset differences (x - y + a, where a is the offset).
[0225] (10) In the case of scalar quantities, sums can include simple sums (x + y) and addition operations. Sums can include the absolute value of the sum (|x + y|), the sum of squares (x^2 + y^2), the square root of the sum (√(x + y)), a weighted sum (ax + by, where a and b are constants), or an offset sum (x + y + a, where a is the offset).
[0226] (11) "Based on" The expression "based on something" means that other things besides that "something" may also be considered. Also, "based on" can be used when a direct result is obtained, or when a result is obtained after intermediate results.
[0227] (12) "Used" or "Used" The expression "something was used" or "something was used" means that something other than that "something" may also be considered. Also, the expression "used" or "used" can be used when a direct result is obtained, or when a result is obtained after an intermediate result.
[0228] (13) Prohibition "To prohibit" can be rephrased as "not to permit." Also, "not prohibited / prohibited" or "permitted / permitted" does not necessarily mean "obligation."
[0229] (14) "Restriction" or "Limitation" "Restriction" or "Limitation" can be rephrased as "not permitted / not allowed" or "not permitted / permitted." Also, "prohibited / not prohibited" or "not permitted / permitted" does not necessarily mean "obligation." Furthermore, the prohibition, quantitatively or qualitatively, may be partial or entire.
[0230] (15) Chroma The term chroma is an adjective represented by the symbols Cb or Cr, indicating that a sample sequence or a single sample represents one of two color difference signals associated with a primary color. The term chroma is sometimes used instead of the term chrominance.
[0231] (16) Luma The term luma is an adjective represented by the symbol or subscript Y or L, indicating that a sample sequence or single sample represents a monochrome signal relating to a primary color. The term luma is sometimes used as a substitute for the term luminance.
[0232] The coding and decoding system of this embodiment will be described below.
[0233] A typical three-dimensional model (also called a 3D model) digitally represents an object so that the user can explore the model using zoom, pan, and rotate in all three dimensions while rendering it over time. One way to construct such a representation is to build a 3D mesh using triangles. In the above model, the positions of the triangle vertices, the connectivity of the triangle vertices to each other, and their associated attributes (such as normals or UV patches) are stored.
[0234] Storing all this information in an uncompressed format requires a very large amount of memory, and therefore a very large bandwidth for transmission. The triangles that form a mesh often have attributes similar to repeating patterns, especially in temporal and spatial neighborhoods. These repetitions can be used to develop efficient encoding and decoding methods for storage and transmission. One such encoding and decoding method is Video-based Dynamic Mesh Coding (V-DMC).
[0235] Figure 26 is a block diagram showing different configuration examples of the coding and decoding system according to this embodiment. As shown in Figure 26, the coding and decoding system includes an coding device 100 and a decoding device 200.
[0236] The encoding / decoding system accepts a three-dimensional mesh (also called a 3D mesh) as input in the form of three-dimensional coordinates of vertices (vertex information), connectivity (connection information), and associated attributes (attribute information). Note that the 3D mesh can include not only geometry but also texture maps.
[0237] The encoding device 100 captures the input 3D mesh (also called the input 3D mesh or input mesh) in the form of vertex 3D coordinates, connectivity, and associated attributes. The encoding device 100 encodes all associated information into a stream. The stream may consist of a single bitstream or multiple bitstreams.
[0238] Network 300 transmits the stream generated by the encoding device 100 to the decoding device 200. Network 300 may be the Internet, a WAN (Wide Area Network), a LAN (Local Area Network), or any combination thereof. Furthermore, network 300 is not necessarily limited to a bidirectional communication network, but may also be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Alternatively, a recording medium such as a DVD (Digital Versatile Disc) or a BD (Blue-Ray Disc) on which the stream is recorded may be used instead of network 300.
[0239] The stream is transmitted to the decoder 200 via the network 300. The decoder 200 decodes the bitstream and generates a three-dimensional mesh using the three-dimensional coordinates, connectivity, and associated attributes of the decoded vertices. The decoder 200 outputs the generated three-dimensional mesh (also called the output 3D mesh or output mesh).
[0240] Figure 27 shows another example of the encoding device 100 configuration.
[0241] As shown in Figure 27, the encoding device 100 includes a preprocessor 1103 and a compressor 1106.
[0242] The encoding device 100 reads the input mesh 1101 and attribute map 1102 and passes them to the preprocessor 1103. The preprocessor 1103 processes the input mesh and extracts the base mesh 1104 and displacement data 1105. The attribute map 1102, along with the extracted base mesh 1104 and displacement data 1105, is passed to the compressor 1106.
[0243] Furthermore, the compressor 1106 compresses the base mesh 1104, displacement data 1105, and attribute map 1102 to generate a bitstream 1107. The compressor 1106 can transmit additional information to the decoder 200 by further including metadata 1108 in the bitstream 1107.
[0244] Figure 28 shows another example of the configuration of the decoding device 200.
[0245] As shown in Figure 28, the decoding device 200 includes an expander 2102 and a post-processing unit 2106.
[0246] The decoding device 200 reads the bitstream 2101 and passes it to the decompressor 2102. The decompressor 2102 decompresses the base mesh 2103, displacement data 2104, and attribute map 2108 from the bitstream 2101 and passes them to the post-processor 2106. An example of displacement data 2104 is a displacement vector.
[0247] Furthermore, the post-processor 2106 generates the output mesh 2107 by processing the base mesh 2103 according to the displacement data 2104 and attribute map 2108. The post-processor 2106 may also use information from metadata 2105 to generate the output mesh 2107.
[0248] Figure 29 is a block diagram showing yet another configuration example of the encoding device 100 according to this embodiment.
[0249] In this example, the encoding device 100 includes a volumetric capturer 511, a projector 512, a base mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and optionally one or more other type encoders 516.
[0250] The volumetric capture device 511 captures the content and outputs the captured content to the projector 512.
[0251] The projector 512 projects content onto an input mesh (three-dimensional mesh frame) that includes geometric coordinates (vertex coordinates indicating the positions of vertices), texture coordinates, and connectivity (connection information). This data is output to the base mesh encoder 513, displacement encoder 514, attribute encoder 515, and optionally one or more other type encoders 516. Each encoder compresses the data into a bitstream.
[0252] Figure 30 is a block diagram showing yet another configuration example of the decoding device 200 according to this embodiment.
[0253] In this example, the decoding device 200 comprises a base mesh decoder 613, a displacement decoder 614, an attribute decoder 615, one or more other type decoders 616, and a three-dimensional reconstructor 617.
[0254] The bitstream is sent to the base mesh decoder 613, the displacement decoder 614, the attribute decoder 615, and optionally one or more other type decoders 616. These decoders decode the bitstream to generate data (decoded data) including geometric coordinates, texture coordinates, and connectivity. The decoded data is then sent to the 3D reconstructor 617, where the output mesh (3D mesh frame) is reconstructed.
[0255] The encoding process performed by the encoding device 100 will be described in detail below.
[0256] Figure 31 is a flowchart showing the processing of the encoding device 100. Figure 32 is an explanatory diagram conceptually illustrating the encoding of a mesh frame. The processing of the encoding device 100 will be explained with reference to Figures 31 and 32.
[0257] In step S101, the encoding device 100 reads the input mesh frame, which is a 3D mesh frame, and its attributes. The input mesh frame is the mesh frame input to the encoding device 100. An example of the input mesh frame, a 3D mesh frame, is shown as mesh frame 1301 (see Figure 32).
[0258] In step S102, the encoding device 100 generates a base mesh frame with fewer vertices than the input mesh frame by performing a decimation process on the input mesh frame read in step S101. The base mesh frame generated by decimating the mesh frame 1301 is shown as the base mesh frame 1302 (see Figure 32).
[0259] In step S103, the encoding device 100 calculates displacement information that the decoding device 200 uses to reconstruct the mesh frame. The displacement information corresponds to a displacement vector from the vertices of the base mesh frame generated in step S102 to the vertices of the input mesh frame. One method for calculating the displacement information is to subtract the coordinates of the vertices of the base mesh frame from the coordinates of the vertices of the input mesh frame. The displacement information calculated from the mesh frame 1301 and the base mesh frame 1302 is shown as displacement information 1303 (see Figure 32). The displacement information 1303 is in vector form, or in other words, it is expressed as a displacement vector.
[0260] In step S104, the encoding device 100 encodes the base mesh frame generated in step S102, the displacement information generated in step S103, and the attributes of the input mesh frame into a bitstream (corresponding to a compressed bitstream). An example of a bitstream is shown as bitstream 1304 (see Figure 32).
[0261] Specifically, bitstream 1304 includes a video bitstream containing vertex coordinates and connection information for vertices A, C, E, and F, displacement information, and texture data, as well as a compressed attribute map (see Figure 32). The displacement information includes displacement information for displacing vertices based on vertex coordinates obtained from the subdivided base mesh frame. The compressed attribute map is texture coordinates for applying texture data to the mesh frame reconstructed using the base mesh frame and displacement information.
[0262] The decoding process performed by the decoding device 200 will be described in detail below.
[0263] Figure 33 is a flowchart showing the processing of the decoding device 200. Figure 34 is an explanatory diagram conceptually illustrating the decoding of a mesh frame (3D mesh). The processing of the decoding device 200 will be explained with reference to Figures 33 and 34.
[0264] In step S201, the decoding device 200 decodes the base mesh frame and attributes from the bitstream (corresponding to the compressed bitstream). An example of the decoded base mesh frame (corresponding to the decoded base mesh frame) is shown as the decoded base mesh frame 2301 (see Figure 34).
[0265] In step S202, the decoding device 200 generates subdivided vertices by performing a subdivision process on the base mesh frame decoded in step S201. An example of a base mesh frame (mesh frame) containing subdivided vertices is shown as base mesh frame 2302 (see Figure 34).
[0266] In step S203, the decoding device 200 decodes displacement information from the bitstream (corresponding to a compressed bitstream). An example of the decoded displacement information is shown as displacement information 2303 (see Figure 34). Displacement information 2303 is in vector form, or in other words, it is represented as a displacement vector.
[0267] In step S204, the decoding device 200 reconstructs the shape of the mesh frame by moving the vertices of the base mesh frame, including the sub-divided vertices, to new positions using displacement information, and then restores the mesh frame by applying attribute information. An example of an attribute is a texture. An example of the reconstructed mesh frame is shown as mesh frame 2304 (see Figure 34).
[0268] Figure 35 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.
[0269] Figure 35 shows an example of a typical intra-decryption block diagram.
[0270] The decoding device shown in Figure 35 comprises an inverse multiplexer 1231, a switch 1232, a static mesh decoder 1233, a mesh buffer 1234, a motion decoder 1235, a base mesh reconstructor 1236, an inverse quantizer 1237, a video decoder 1238, an image umpacker 1239, an inverse quantizer 1240, an inverse wavelet converter 1241, a reconstructor 1242, a video decoder 1243, and a color converter 1244.
[0271] The demultiplexer 1231 acquires the compressed bitstream and separates the compressed data for the base mesh, the video containing displacement data (also called the displacement bitstream), and the video containing attribute data (also called the attribute bitstream). The compressed data for the base mesh is passed to the switch 1232. The switch 1232 decides whether to perform intra-decoding or inter-decoding based on the parameters in the bitstream.
[0272] If intra-decoding is selected, the bitstream is passed to a static mesh decoder 1233 that generates a quantized base mesh. The static mesh decoder 1233 is, for example, a decoder that uses an edge breaker algorithm to decode 3D mesh data. The static mesh decoder 1233 generates a quantized base mesh from the bitstream. The quantized base mesh generated by the static mesh decoder 1233 is stored in the mesh buffer 1234 for reference when inter-decoding is selected.
[0273] Switch 1232 passes compressed data about the base mesh to the motion decoder 1235 if inter-decoding is selected. The motion decoder 1235 receives the previously decoded, quantized base mesh and decodes motion data representing the difference in vertex coordinates between the quantized base mesh stored in the mesh buffer 1234 and the current quantized base mesh. The motion data and the quantized base mesh stored in the mesh buffer 1234 are used by the base mesh reconstructor 1236 to reconstruct the current quantized base mesh. The quantized base mesh obtained from inter-decoding or intra-decoding is passed to the inverse quantizer 1237 to obtain the decoded base mesh.
[0274] The video containing displacement data is passed to the video decoder 1238 because the bitstream contains displacement data in an image format having two chroma information and one luma information. The video decoder 1238 decodes the data using the video frame decompression method. In another example, the displacement data is decoded using an arithmetic decoder. This decompressed data is passed to the image umpacker 1239, which extracts wavelet coefficients associated with each vertex from the decompressed image data. The inverse quantizer 1240 inverse quantizes the quantized wavelet coefficients with three components associated with each vertex. The inverse wavelet transformer 1241 inverse transforms the result to obtain the finally decoded displacement data. The decoded displacement data and the decoded base mesh are passed to the reconstructor 1242. The reconstructor 1242 subdivides the edges of the decoded base mesh, displaces the vertices using the decoded displacement data, and obtains the decoded mesh.
[0275] The video containing attribute data is passed to another video decoder 1243 to obtain a decoded attribute bitstream. The decoded attribute bitstream is further processed by a color converter 1244 for color space and color format conversion to obtain a decoded attribute map.
[0276] Figure 36 is a block diagram showing an example configuration of the encoding device according to this embodiment.
[0277] First, the encoding device acquires the base mesh bitstream, the displacement bitstream, and the attribute bitstream obtained from the three-dimensional mesh preprocessing step.
[0278] The encoding device shown in Figure 36 comprises a quantizer 1261, a switch 1262, a static mesh encoder 1263, a mesh buffer 1264, a motion encoder 1265, a base mesh reconstructor 1266, a displacement data updater 1267, a wavelet converter 1268, a quantizer 1269, an image packer 1270, a video encoder 1271, a color converter 1272, and a video encoder 1273.
[0279] The base mesh (specifically, the position information of the multiple vertices that make up the base mesh) is first quantized by the quantizer 1261. The quantized base mesh (base mesh data) is output to the switch 1262, which determines whether to use intra-coding or inter-coding. If intra-coding is selected, the quantized base mesh (base mesh bitstream) is output to the static mesh encoder 1263, which generates the quantized base mesh. An example of the static mesh encoder 1263 is an encoder that uses an edge breaker algorithm for encoding three-dimensional mesh data. This encoded (quantized) base mesh is stored in the mesh buffer 1264 for reference when inter-coding is selected. If inter-coding is selected, the switch 1262 outputs compressed data related to the base mesh to the motion encoder 1265. The motion data and static three-dimensional mesh in the mesh buffer 1264 are used by the base mesh reconstructor 1266 to reconstruct the current quantized base mesh.
[0280] The displacement data is output to the displacement data updater 1267, where it is updated based on the quantized static base mesh or the reconstructed inter-encoded base mesh. Next, the wavelet converter 1268 performs the transformation, followed by quantization in the quantizer 1269. The quantized displacement data is packed into an image by the image packer 1270 and finally encoded by the video encoder 1271. The encoded displacement data is output to the multiplexer 1274.
[0281] Attribute information (e.g., attribute map) is output to the color converter 1272 for conversion of color space and color format. This converted attribute information is encoded by the video encoder 1273 and output to the multiplexer 1274.
[0282] The multiplexer 1274 acquires data relating to the encoded base mesh (compressed data relating to the base mesh), video data including encoded displacement data, and video data including attribute information such as encoded attribute maps, and generates a bitstream (compressed bitstream) containing the acquired data. The generated compressed bitstream is output to, for example, a decoding device.
[0283] Figure 37 is a block diagram showing an example configuration of a decoding device according to this embodiment.
[0284] Figure 37 shows an example of a reconstructor that obtains a decoded 3D mesh 1256 from the decoded base mesh 1251 and the decoded displacement data 1254.
[0285] The decoded base mesh 1251 is passed to the sub-divider 1252.
[0286] The sub-decomposer 1252 subdivides any two connected vertices in the entire 3D mesh by adding a new vertex between them. This process can be repeated several times, including the vertices created in the previous sub-decomposition step, to generate a predetermined number of vertices. Each iteration of sub-decomposition across the entire 3D mesh generates a new level of detail (LoD). The subdivided mesh 1253 and the decoded displacement data 1254 are passed to the displacementr 1255. The displacementr 1255 generates the decoded 3D mesh 1256 by moving each vertex to a new position according to the corresponding displacement data.
[0287] Sub-partitioning will be described below. Sub-partitioning is performed, for example, by a sub-partitioner 1252.
[0288] Figure 38 is an explanatory diagram showing an example of subdivision.
[0289] The base mesh shown in Figure 38(a) includes vertices A, B, and C, and connectivity information indicating their connectivity.
[0290] Figure 38(b) shows the mesh generated by the first subdivision, in other words, the mesh after the first subdivision. In the first subdivision, the subdivision generator generates vertices D, E, and F, and connection information indicating their connectivity. The mesh generated by the subdivision generator is also called LoD1 or the first LoD.
[0291] Vertex D of the mesh after the first subdivision is a vertex generated by the subdivision based on vertices A and B. Similarly, vertex F is a vertex generated by the subdivision based on vertices B and C. Vertex E is a vertex generated by the subdivision based on vertices A and C.
[0292] For example, vertex D could be the midpoint of the line segment AB (or side AB) connecting vertices A and B, which were the source of its generation. Similarly, vertex E could be the midpoint of line segment AC, and vertex F could be the midpoint of line segment BC.
[0293] Figure 38(c) shows the mesh generated by the second subdivision, in other words, the mesh after the second subdivision. In the second subdivision, the subdivision generator generates vertices G, H, I, J, K, L, M, N, and O, and connection information indicating their connectivity. The mesh generated by the subdivision generator is also called LoD2 or the second LoD.
[0294] Vertex G of the mesh after the second subdivision is a vertex generated by the subdivision based on vertices A and D. Similarly, vertex H is a vertex generated by the subdivision based on vertices A and E. Vertex I is a vertex generated by the subdivision based on vertices B and D. Vertex J is a vertex generated by the subdivision based on vertices D and F. Vertex K is a vertex generated by the subdivision based on vertices E and F. Vertex L is a vertex generated by the subdivision based on vertices C and E. Vertex M is a vertex generated by the subdivision based on vertices B and F. Vertex N is a vertex generated by the subdivision based on vertices C and F. Vertex O is a vertex generated by the subdivision based on vertices D and E.
[0295] For example, vertex G could be the midpoint of the line segment AD (or edge AD) connecting vertices A and D, which were the source of its generation. Similarly, vertex H could be the midpoint of line segment AE. Vertex I could be the midpoint of line segment BD. Vertex J could be the midpoint of line segment DF. Vertex K could be the midpoint of line segment EF. Vertex L could be the midpoint of line segment CE. Vertex M could be the midpoint of line segment BF. Vertex N could be the midpoint of line segment CF. Vertex O could be the midpoint of line segment DE.
[0296] In the following sections, the displacement of the vertices will be explained with reference to Figures 39 and 40. The displacement of the vertices is performed by the reconstructor.
[0297] Figure 39 is an explanatory diagram showing an example of vertex displacement after subdivision. Figure 40 is an explanatory diagram showing an example of vertices in the original mesh.
[0298] The base mesh shown in Figure 39(a) includes vertices A, B, C, and Z, and connectivity information indicating their connectivity.
[0299] Figure 39(b) shows the mesh generated by the first subdivision, in other words, the mesh after the first subdivision (i.e., the first LoD). In the first subdivision, the subdivision generator generates vertices S, T, U, X, or Y and connection information indicating their connectivity. Vertices S, T, U, X, or Y are the same as vertices D, E, and F shown in Figure 38(b).
[0300] Figure 39(c) shows the mesh generated by the second subdivision, in other words, the mesh after the second subdivision (i.e., the second LoD). In the second subdivision, the subdivision tool generates vertices D, E, F, G, and H, and connectivity information indicating their connectivity. Vertices D, E, F, G, and H are the same as vertices G, H, I, J, K, L, M, N, or O shown in Figure 38(c).
[0301] Figure 39(d) shows the mesh containing the vertices after subdivision and displacement. The vertices A, B, C, D, E, F, G, H, S, T, U, X, Y, and Z shown in Figure 39(d) are located at positions that have been displaced using displacement information from the positions of those vertices shown in Figure 39(c).
[0302] The original mesh shown in Figure 40 is an example of the mesh input to the encoding device 100, that is, the mesh before encoding.
[0303] The mesh shown in Figure 39 has a shape similar to the original mesh shown in Figure 40. Since the displacement information is generated by the encoding device 100 as information indicating the displacement from the vertices of the base mesh to the vertices of the original mesh, a mesh with a shape similar to the original mesh is generated by reconstructing the mesh using the displacement information generated in this way.
[0304] The decoding device 200 can output the mesh shown in Figure 39(d).
[0305] Next, we will explain the division of the mesh into submeshes with reference to Figures 41 and 42.
[0306] A mesh can be divided into multiple smaller parts, and each divided part can be encoded. When dividing a mesh, the vertices of the mesh are divided in such a way that the coordinates and connectivity of the vertices included in each part can be encoded independently.
[0307] Figure 41 is an explanatory diagram showing an example of a mesh. Figure 42 is an explanatory diagram showing an example of dividing a mesh into submeshes.
[0308] The mesh shown in Figure 41 is the original mesh, and is sometimes called the full mesh in contrast to the submesh.
[0309] Figure 42 shows how the full mesh shown in Figure 41 is divided into two submeshes. For vertices A, B, and C of the full mesh (see Figure 41), vertex A is duplicated into vertices A1 and A2, vertex B is duplicated into vertices B1 and B2, and vertex C is duplicated into vertices C1 and C2, thereby creating two submeshes (i.e., the first submesh and the second submesh) from the full mesh. The first submesh and the second submesh are meshes that can be decoded independently.
[0310] In the following sections, the packing of displacement information into image frames will be explained with reference to Figures 43, 44, and 45.
[0311] Figures 43, 44, and 45 are explanatory diagrams illustrating examples of packing displacement information into image frames. Note that image frames can also be referred to as video frames.
[0312] Vertex displacement data is encoded as image frame data by mapping it to each component of an image frame in YUV format (i.e., the Y component (Y Plane), U component (U Plane), and V component (V Plane) respectively). This case is explained below as an example. Alternatively, vertex displacement data may be encoded as image frame data by mapping it to each component of an image frame in RGB format (the R component, G component, and B component, respectively).
[0313] The decoding device 200 can use an image coding module to extract displacement data. The displacement data may be in the form of X, Y, or Z components in a global coordinate system (e.g., a Cartesian coordinate system), or normal, tangent, or both tangent components in a local coordinate system. Methods for mapping displacement data to an image frame include the following:
[0314] For example, in the first method, displacement data is arranged in scan order within the image frame. An example of packing displacement data in this case is shown in Figure 43. The displacement data is directly mapped to the image frame according to a predefined scan order.
[0315] Note that since the height and width of the image frame are fixed, the displacement data may not fit perfectly within the frame. In such cases, the remaining portion of the image frame is padded with padding data (also called padded data) (see Figure 43).
[0316] For example, in the second method, displacement data is separated into multiple Lines of Data (LDs) and mapped to the Y, U, and V components of the image frame. An example of packing the displacement data in this case is shown in Figure 44. Here, the displacement data for the image frame of the next LD starts immediately after the displacement data for the previous LD ends. Similar to the first method, if the displacement data does not fit perfectly into the image frame, padding is applied to the end of the image frame (see Figure 44).
[0317] For example, in the third method, the displacement data corresponding to the LoD is mapped to the Y, U, and V components of the image frame in a different manner than in the second method. An example of the packing of the displacement data in this case is shown in Figure 45. In this way, each LoD can be decoded independently. In the third method, intermediate padding is performed on the displacement data of each LoD, and CTU alignment is performed together with padding at the end of the video frame (see Figure 45).
[0318] Next, we will describe the encoding device 100 and decoding device 200 when the mesh is divided into multiple submeshes.
[0319] Figure 46 shows another example of the configuration of the encoding device 100 according to the embodiment. Specifically, Figure 46 shows the configuration of a submesh encoding device, which is a device that performs encoding processing when the input mesh 1101 is divided into subdivided meshes (multiple submeshes). For example, the submesh encoding device comprises multiple encoding devices 100.
[0320] The input mesh 1101 (full mesh) input to the submesh encoding device is divided into multiple meshes (submeshes). The multiple submeshes are input to, for example, multiple encoding devices 100. Each of the multiple submeshes may be input to any of the multiple encoding devices 100. For example, the submesh encoding device divides the input mesh 1101 into multiple submeshes and inputs the divided multiple submeshes to the multiple encoding devices 100.
[0321] Furthermore, after being divided into multiple submeshes, encoding processing (processing of overlapping submeshes) is performed on the boundaries of the submeshes.
[0322] For example, for each submesh, preprocessing is performed by the preprocessor 1103, generating and encoding the base mesh, displacement data, and metadata.
[0323] Furthermore, the encoding device 100 may be implemented with the configuration of such a submesh encoding device. In other words, the encoding device 100 may be configured to include multiple preprocessors 1103 and compressors 1106, and a predetermined process may be performed on the submesh for each set of multiple preprocessors 1103 and compressors 1106. Also, the number of sets of preprocessors 1103 and compressors 1106 included in the encoding device 100 is arbitrary and not particularly limited.
[0324] Figure 47 shows an example of the configuration of the pre-treatment unit 1103 according to the embodiment.
[0325] The preprocessing unit 1103 includes, for example, a base mesh generator 1401, a sub-divider 1402, and a displacement data generator 1403.
[0326] First, in the preprocessing unit 1103, a base mesh is generated by the base mesh generator 1401.
[0327] Next, the base mesh is subdivided by the sub-divider 1402 in a predetermined manner, and a subdivided mesh (Subdivided Mesh or Subdivided Base Mesh), which is the subdivided base mesh, is generated.
[0328] Displacement data is generated by the displacement data generator 1403 from the sub-divided mesh and the sub-mesh which is the input mesh 1101 after division.
[0329] The displacement data is, for example, the difference vector between the input mesh 1101 and the sub-divided mesh.
[0330] Furthermore, the same method used for sub-partitioning in encoding is employed in decoding.
[0331] Furthermore, for example, the encoding device 100 may transmit the sub-partitioning method in encoding and the parameters used in sub-partitioning to the decoding device 200.
[0332] Figure 48 shows another example of the configuration of the decoding device 200 according to the embodiment. Specifically, Figure 48 shows the configuration of a submesh decoding device, which is a device that performs decoding processing when a bitstream 2101 contains multiple submeshes. For example, the submesh decoding device comprises a plurality of decoding devices 200. For example, the submesh decoding device comprises a plurality of decoding devices 200 and a coupler 2109.
[0333] The encoded data for each submesh contained in the bitstream 2101 is input to the decompressor 2102 of each decoder 200.
[0334] Furthermore, for example, the post-processing unit 2106 executes the processing of the above-mentioned reconstructor.
[0335] In the post-processing unit 2106, for each decoded submesh, the base mesh is subdivided, a displacement vector is added to the subdivided base mesh, and the submesh is restored.
[0336] In other words, the post-processing unit 2106 performs the above-mentioned reconstruction process for each submesh.
[0337] The combiner 2109 merges the submeshes restored by each of the multiple decoding devices 200 to reconstruct the full mesh (output mesh 2107) before it was divided.
[0338] The decoding device 200 may be implemented with the configuration of such a submesh decoding device. In other words, the decoding device 200 may be configured to include multiple decompressors 2102 and post-processors 2106, and a predetermined process may be performed on the encoded data for each submesh for each set of multiple decompressors 2102 and post-processors 2106. Furthermore, the number of sets of decompressors 2102 and post-processors 2106 provided in the decoding device 200 is arbitrary and not particularly limited.
[0339] Figure 49 is a diagram showing a specific example of the configuration of the decoding device 200 according to the embodiment. Specifically, Figure 49 shows the specific configuration of the post-decoder 2306, which is one of the post-decoder 2306 and decoder 2305 that the decoding device 200 is equipped with.
[0340] The post-decoder 2306 comprises a pre-reconstructor 2307, a reconstructor 2308, a post-reconstructor 2309, and an adaptor 2310.
[0341] The processing in the post-decoder 2306 is optional depending on the application. An example of processing in the post-decoder 2306 (post-decoder processing) is the conversion of decoded data to a nominal format, such as video conversion from YUV space to RGB space. Post-decoder processing can be encapsulated into multiple processes such as pre-reconstruction, reconstruction, post-reconstruction, and adaptation.
[0342] The pre-reconstructor 2307 performs a pre-reconstruction process. The pre-reconstructor 2307 scales the normalized texture coordinates to match the dimensions of the texture image, for example, in the context of video-based dynamic mesh coding.
[0343] The reconstructor 2308 performs a reconstruction process. The reconstruction process is invoked, for example, on the decoded atlas frame, the decoded base mesh frame, the decoded video frame, and the syntax elements associated with the same mesh sequence. The output of the reconstruction process is a series of mesh frames reconstructed before the post-reconstruction process.
[0344] The post-reconstructor 2309 performs post-reconstruction processing. The post-reconstructor 2309 performs a number of smoothing operations on the reconstructed mesh frame, for example, in the context of video-based dynamic mesh coding. Smoothing operations include, for example, folding edges in the mesh or adding new vertices to the mesh.
[0345] The conformer 2310 performs a conformance process. The conformance process is applied by some application, for example, to fit the reconstructed mesh to a predetermined scenario. For example, the vertices of the reconstructed mesh are converted from the 3D model coordinate system to the 3D world coordinate system. The conformer 2310 outputs the reconstructed final mesh frame (final 3D mesh frame 2311).
[0346] <Encoding and Decoding of Attribute Information> Figure 50 is a conceptual diagram showing a specific example of a bitstream. In a bitstream, encoded data is encapsulated in a data unit structure. Specifically, a bitstream of encoded mesh data is also expressed as a DMC (Dynamic Mesh Coding) bitstream and consists of a series of DMC units. Each DMC unit is a data unit and is classified into a base mesh data unit, displacement data unit, metadata unit, or texture map unit, etc.
[0347] For example, encoded base mesh data is stored in a base mesh data unit with a header. Encoded displacement data is stored in a displacement data unit with a header. Encoded metadata is stored in a metadata unit with a header. Encoded texture maps are stored in a texture map unit with a header.
[0348] The header included in a unit is also called the unit header. The unit header stores the unit type, which indicates the type of data stored in the payload.
[0349] Metadata may correspond to a parameter set or SEI (Supplemental Enhancement Information). A parameter set may be a set of parameters common to a frame (data, access units, and samples at the same time, etc.), such as a frame parameter set, or a set of parameters common to a sequence, such as a sequence parameter set. Frame parameters may be expressed as a picture parameter set.
[0350] Metadata may correspond to the bitstream header. That is, the bitstream header may include the headers for each unit, or it may include the metadata unit.
[0351] In the decoding device 200, the bitstream is decoded using the data type and divided into multiple data units.
[0352] In the encoding and decoding of a three-dimensional mesh, surface information, which consists of vertex information and connectivity information, and attribute information for the surface are encoded and decoded, respectively. Examples of attribute information for a surface include color information and reflectance for the surface.
[0353] For example, three-dimensional data has attribute information (e.g., color information for the surface and reflectivity for the surface) for multiple faces of a single mesh. For example, as shown in Figure 3, the image of each face is mapped to a two-dimensional image. Not only the image of each face, but the attribute information of each face can also be mapped to a two-dimensional image. Therefore, for each attribute type, the attribute information is mapped to a two-dimensional image, and multiple two-dimensional images corresponding to multiple attribute types are obtained.
[0354] Furthermore, in order to signal attribute information corresponding to multiple attribute types, multiple two-dimensional images are signaled. However, the overhead of signaling multiple two-dimensional images may increase the amount of coding and processing required. Therefore, in this embodiment, multiple two-dimensional images are packed into units corresponding to a single image frame, such as a single image. This makes it possible to signal multiple two-dimensional images together.
[0355] Here, the two-dimensional image on which attribute information is mapped is an image, and can also be expressed as an attribute map, a two-dimensional attribute information set, or a two-dimensional attribute dataset, etc.
[0356] Furthermore, the mapping information showing the correspondence between multiple faces in a three-dimensional mesh and multiple faces in a two-dimensional image may be the same or different for multiple attribute types. In other words, the attribute information for each attribute type in a three-dimensional mesh may be mapped to a two-dimensional image based on the same correspondence between multiple attribute types, or it may be mapped to a two-dimensional image based on different correspondence between multiple attribute types.
[0357] Furthermore, the mapping information may be signaled from the encoding device 100 to the decoding device 200. In other words, the encoding device 100 may encode the mapping information, and the decoding device 200 may decode the mapping information. The mapping information may be signaled as attribute information or as vertex information.
[0358] Figure 51 is a block diagram showing an example of the configuration of an encoding device 100 that encodes multiple attribute maps. The encoding device 100 includes a preprocessor 1103 and a compressor 1106.
[0359] Multiple attribute maps 1102 are input to the encoding device 100. The preprocessor 1103 of the encoding device 100 stores the multiple attribute maps 1102 in the image 1110. Then, the compressor 1106 encodes the image 1110 into a bitstream 1107 using an image encoding method.
[0360] The encoding device 100 may receive an input mesh 1101. The preprocessor 1103 may obtain a base mesh 1104, displacement data 1105, and metadata 1108 from the input mesh 1101. The compressor 1106 may encode the base mesh 1104, displacement data 1105, and metadata 1108 into a bitstream 1107.
[0361] Alternatively, the preprocessor 1103 may acquire attribute information from the input mesh 1101 and acquire multiple attribute maps 1102 from the attribute information, without acquiring multiple attribute maps 1102 from an external source.
[0362] Furthermore, in the example shown in Figure 51, image 1110 may be considered as a single attribute map containing multiple attribute maps 1102.
[0363] Figure 52 is a block diagram showing an example of the configuration of a decoding device 200 that decodes multiple attribute maps. The decoding device 200 includes an expander 2102 and a post-processor 2106.
[0364] The decoding device 200 decodes the image 2110 from the bitstream 2101. Then, the post-processor 2106 reconstructs multiple attribute maps 2108 from the image 2110.
[0365] The decoding device 200 inputs the restored attribute maps 2108 along with information indicating the attribute type of each attribute map to the renderer 2111. The renderer 2111 is a rendering device that reconstructs a three-dimensional mesh by switching between the multiple attribute maps 2108 according to the application of the three-dimensional mesh.
[0366] The expander 2102 may decode the base mesh 2103, displacement data 2104, and metadata 2105 from the bitstream 2101. The post-processor 2106 may then obtain the output mesh 2107 from the base mesh 2103, displacement data 2104, and metadata 2105, and input the output mesh 2107 to the renderer 2111.
[0367] The renderer 2111 may then reconstruct a three-dimensional mesh based on the output mesh 2107 and a plurality of attribute maps 2108. The renderer 2111 may be included in the decoding device 200.
[0368] Furthermore, in the example shown in Figure 52, image 2110 may be considered as a single attribute map containing multiple attribute maps 2108.
[0369] In this embodiment, an example is shown of packing multiple attribute maps, each having different attribute types such as color and reflectance, into a single image. Packing methods, unpacking methods, encoding methods, decoding methods, and metadata related to this process are disclosed.
[0370] Figure 53 is a flowchart showing an example of the processing of an encoding device 100 that encodes multiple attribute maps. In this example, the preprocessor 1103 of the encoding device 100 packs the multiple attribute maps into an image (S301). The preprocessor 1103 then generates arrangement information indicating the arrangement of each attribute map in the image (S302).
[0371] Furthermore, the compressor 1106 of the encoding device 100 encodes the image and encodes metadata including the number of attribute maps, the attribute type of each attribute map, and placement information (S303). Then, the compressor 1106 transmits a bitstream containing the encoded image and encoded metadata (S303).
[0372] Figure 54 is a flowchart showing an example of the processing of a decoding device 200 that decodes multiple attribute maps.
[0373] In this example, the decompressor 2102 of the decoding device 200 decodes an image from the bitstream (S401). The decompressor 2102 may decode the image according to a video encoding standard such as VVC or HEVC. The decompressor 2102 decodes the image according to an image encoding standard such as JPEG or PNG. These video encoding standards and image encoding standards correspond to methods used for encoding or decoding images (i.e., image encoding methods or image decoding methods).
[0374] The image format can be YUV420 chromatic difference format, YUV444 chromatic difference format, or YUV400 chromatic difference format.
[0375] Furthermore, the decompressor 2102 decodes the number of attribute maps and the attribute types of each of the multiple attribute maps from the bitstream (S402). Here, the number of attribute maps is the number of attribute maps of the 3D mesh. Also, each attribute map is associated with one attribute type.
[0376] Furthermore, the decompressor 2102 decodes arrangement information from the bitstream, indicating the placement of each attribute map in the image (S403). The arrangement information may indicate the position and size of each attribute map in the image. The arrangement information can also be expressed as position information.
[0377] Next, the decompressor 2102 unpacks multiple attribute maps from the image using the placement information (S404). Then, the renderer 2111 renders a 3D mesh using the unpacked attribute maps (S405).
[0378] Figure 55 is a conceptual diagram showing an example of the storage location of the number of attribute maps and the attribute types of each attribute map in a bitstream. In this example, the number of attribute maps and the attribute types of each attribute map are stored in the body of the bitstream, i.e., the data area of the bitstream. The data area may also be the payload area.
[0379] Figure 56 is a conceptual diagram showing another example of the storage location of the number of attribute maps and the attribute types of each attribute map in a bitstream. In this example, the number of attribute maps and the attribute types of each attribute map are stored in the header area.
[0380] The header area may include the sequence parameter set, picture parameter set, frame parameter set, SEI (Supplementary Enhancement Information), and the headers of each data unit. The number of attribute maps and the attribute types of each attribute map may be stored, for example, in the header of the data unit, or in the sequence parameter set, frame parameter set, or SEI.
[0381] Figure 57 is an explanatory diagram showing an example of attribute types and their descriptions. An example of an attribute type is texture, in which case the attribute map shows the color of the surface of the 3D mesh, specifically the color of each position on the surface.
[0382] Another example of an attribute type is material ID, in which case the attribute map shows a different identifier for each material on the surface of the 3D mesh, specifically the material ID for each location on the surface. Another example of an attribute type is transparency, in which case the attribute map shows the transparency of the surface of the 3D mesh, specifically the transparency for each location on the surface.
[0383] Another example of an attribute type is reflectance, in which case the attribute map shows the reflectance of the surface of the 3D mesh, specifically the reflectance at each location on the surface. The reflectance may be ambient reflectance, diffuse reflectance, or specular reflectance. Two or more types of reflectance may be used as two or more attribute types. The reflectance of the surface of the 3D mesh may represent the light and shadow on the surface of the 3D mesh.
[0384] Another example of an attribute type is normals, in which case the attribute map shows the normals of the surface of the 3D mesh, specifically the normals of each location on the surface. The normals of the surface of the 3D mesh may represent the height differences of different parts of the surface of the 3D mesh. Another example of an attribute type is texture coordinates, in which case the attribute map shows the texture coordinates of the surface of the 3D mesh, specifically the texture coordinates of each location on the surface.
[0385] Another example of an attribute type is a face group ID, in which case the attribute map shows the face group IDs of the surface of the 3D mesh, specifically the face group IDs of each location on the surface.
[0386] Another example of an attribute type is the transmission filter color, in which case the attribute map shows the transmission filter color of the 3D mesh surface. Another example of an attribute type is the illumination model, in which case the attribute map shows the illumination model used on the surface of the 3D mesh. Another example of an attribute type is the focus, in which case the attribute map shows the focus of the specular reflection highlights used on the surface of the 3D mesh. Another example of an attribute type is the sharpness, in which case the attribute map shows the sharpness of the reflections on the surface of the 3D mesh.
[0387] Another example of an attribute type is optical density, in which case the attribute map shows the optical density of the surface of the 3D mesh. Another example of an attribute type is surface roughness, in which case the attribute map shows a scalar value for deforming the surface of the 3D mesh to create surface roughness. Another example of an attribute type is surface identifier, in which case the attribute map shows identifiers for different surfaces of the 3D mesh. Another example of an attribute type is orientation, in which case the attribute map shows the direction in which the surface of the 3D mesh is facing.
[0388] Another example of attribute types is light and shadow, in which case the attribute map shows the light and shadow on the surface of the 3D mesh. Another example of attribute types is elevation, in which case the attribute map shows the elevation differences of different parts of the surface of the 3D mesh.
[0389] The multiple attribute types of the multiple attribute maps packed into the image may include one or more of the multiple attribute types shown above.
[0390] Figure 58 is a conceptual diagram showing a first example of packing multiple attribute maps. In this example, multiple attribute maps of different attribute types are stored (i.e., embedded) in a single image. Specifically, an attribute map of attribute type A, an attribute map of attribute type B, and an attribute map of attribute type C are packed into a single image. All attribute maps of a 3D mesh may also be packed into the same image. In this example, each attribute map is stored in a rectangular area within a single image.
[0391] Each rectangular region may contain areas where attribute information is not mapped, and these areas may contain padding data. Furthermore, the image may contain areas where no attribute maps are stored, and these areas may contain padding data.
[0392] Figure 59 is a syntax diagram showing an example of the syntax structure in the first packing example. In this example, num_attributes indicates the number of attribute maps. The number of attribute maps may be two or more.
[0393] The `video_component_id` indicates the identifier of a video component, specifically indicating which video component each attribute map belongs to. A video component can be the video itself. In other words, `video_component_id` may indicate the video identifier, or it may indicate which video each attribute map belongs to. Furthermore, a video component can be an image contained within a video.
[0394] `attribute_type_id` indicates the attribute type identifier. In other words, `attribute_type_id` indicates the attribute type of each attribute map. `start_x`, `start_y`, `attribute_width`, and `attribute_height` are examples of attribute map placement information.
[0395] Specifically, `start_x` indicates the horizontal position of the attribute map in the image (more specifically, the position of the left edge of the attribute map), and `start_y` indicates the vertical position of the attribute map in the image (more specifically, the position of the top edge of the attribute map). Additionally, `attribute_width` indicates the width of the attribute map in the image, and `attribute_height` indicates the height of the attribute map in the image.
[0396] For example, multiple attribute maps may be stored in multiple distinct parts of an image. Multiple attribute maps may be stored in the luminance Y component, chrominance Cb or U component, and chrominance Cr or V component of the image. Alternatively, multiple attribute maps may be stored in multiple color components such as red R, green G, and blue B. Or, multiple attribute maps may be stored in only one component, such as the luminance Y component. Furthermore, each attribute map may be stored in a sequence of samples in the image.
[0397] Figure 60 is a syntax diagram showing another example of the syntax structure. Multiple attribute maps may be stored in multiple video components. In this example, the number of attribute maps contained in each video component, the attribute type of each attribute map, and the placement information of each attribute map are signaled as information about the video components of the encoded image.
[0398] Specifically, `num_attributes_video_component` indicates the number of video components that store multiple attribute maps, and `num_attributes_in_video_component` indicates the number of attribute maps stored in each video component.
[0399] Furthermore, attribute_type_id indicates the attribute type of each attribute map by its identifier. Additionally, start_x indicates the horizontal position of each attribute map, start_y indicates the vertical position of each attribute map, attribute_width indicates the width of each attribute map, and attribute_height indicates the height of each attribute map.
[0400] Figure 61 is a syntax diagram showing yet another example of the syntax structure. This example shows a syntax structure for signaling attribute components, where attribute components are not limited to attribute maps that correspond to attribute information and are stored in the video component.
[0401] In this example, num_attribute_component indicates the number of attribute components to be processed. component_type indicates whether each attribute component is stored in a video component. component_attribute_type_id indicates the attribute type of the attribute component by identifier.
[0402] When an attribute component is stored in a video component, the packing_attribute_info_flag is signaled, indicating whether or not the packing information of the attribute component is signaled. If packing_attribute_info_flag = 1, the packing information of the attribute component is signaled by packing_attribute_info.
[0403] packing_attribute_info corresponds to packing information and includes num_packing_attribute and packing_attribute_type_id. num_packing_attribute indicates the number of attribute maps packed into the video component, and packing_attribute_type_id indicates the attribute type of each attribute map.
[0404] In other words, an attribute component may contain multiple attribute maps, and the attribute types of the attribute maps may be more granular than the attribute types of the attribute component.
[0405] In the example in Figure 61, the placement information is omitted, but the placement information may be signaled. Specifically, as in the example in Figure 59, start_x, start_y, attribute_width, and attribute_height may be signaled.
[0406] Figure 62 is a conceptual diagram showing a second example of packing multiple attribute maps. In this example, multiple attribute maps with different attribute types are stored in multiple subcomponents of a single image.
[0407] Specifically, in this example, multiple attribute maps of a 3D mesh are packed into the luminance Y component, chrominance U component, and chrominance V component. More specifically, in this example, the attribute map of attribute type A is stored in the luminance Y component, the attribute map of attribute type B is stored in the chrominance U component, and the attribute map of attribute type C is stored in the chrominance V component. The luminance Y component, chrominance U component, and chrominance V component are different components (also referred to as video subcomponents) of the same image.
[0408] In this example, there are three attribute maps packed into the three components of the image: the Y, U, and V components. The number of attribute maps packed into multiple components of an image may be two or more. The number of components of an image into which multiple attribute maps are packed may be two or more. Also, two or more of the multiple attribute maps may be stored in one of the multiple components of an image. Alternatively, multiple attribute maps may be stored in a one-to-one correspondence between each of the multiple components of an image.
[0409] Figure 63 is a syntax diagram showing an example of the syntax structure in the second packing example. In this example, num_attributes indicates the number of attribute maps. The number of attribute maps may be two or more. attribute_type_id indicates the attribute type of each attribute map by identifier.
[0410] The `video_component_id` indicates the identifier of a video component, specifically indicating which video component each attribute map belongs to. The video component may be the video itself. In other words, `video_component_id` may indicate the identifier of a video, or it may indicate which video each attribute map belongs to. Furthermore, the video component may be an image contained within a video.
[0411] video_sub_component_id indicates the identifier of a video subcomponent (video subcomponent ID), specifically indicating which video subcomponent each attribute map belongs to. A video subcomponent may correspond to one of the luminance Y component, chrominance U component, or chrominance V component. In other words, video_component_id may be an identifier representing one of the luminance Y component, chrominance U component, or chrominance V component, and may indicate which component each attribute map belongs to.
[0412] In the example in Figure 63, start_x, start_y, attribute_width, and attribute_height may be signaled, similar to the examples in Figure 59, etc.
[0413] Figure 64 is an explanatory diagram showing an example of a video subcomponent ID in the second packing example. In this example, the video subcomponent ID (video_sub_component_id) can take three values: 0, 1, and 2. When the video subcomponent ID is 0, it represents the luminance Y component. When the video subcomponent ID is 1, it represents the chrominance Cb (or U) component. When the video subcomponent ID is 2, it represents the chrominance Cr (or V) component.
[0414] The video subcomponents are not limited to this example and may correspond to the R, G, and B color components. Furthermore, the video subcomponents may substantially be components of video or components of an image.
[0415] Figure 65 is a conceptual diagram showing a third example of packing multiple attribute maps. In this example, multiple attribute maps with different attribute types are stored in multiple different layers corresponding to a single image.
[0416] In other words, in this example, multiple attribute maps of different attribute types are packed into multiple images on multiple different layers in an access unit (AU). Also, in this example, each layer contains only one attribute map. Specifically, in this example, an attribute map of attribute type A is stored in the first layer, an attribute map of attribute type B is stored in the second layer, and an attribute map of attribute type C is stored in the third layer. The first, second, and third layers are different layers that correspond to the same image.
[0417] As described above, in this example, each layer contains only one attribute map. However, each layer may contain multiple attribute maps with different attribute types.
[0418] Note that the multiple layers are, for example, multiple layers in scalable coding. The multiple layers may include a base layer and one or more extension layers. The multiple layers may also correspond to multiple different resolutions. Alternatively, the multiple layers may correspond to multiple different views, including a base view and one or more extension views. The multiple images in the multiple layers may have different sizes or may have the same size.
[0419] Figure 66 is a syntax diagram showing an example of the syntax structure in the third packing example. In this example, num_attributes indicates the number of attribute maps. The number of attribute maps may be two or more. attribute_type_id indicates the attribute type of each attribute map by identifier. layer indicates which layer each attribute map belongs to. attribute_width indicates the width of the attribute map in the image, and attribute_height indicates the height of the attribute map in the image.
[0420] Two or more packing examples from the first packing example, the second packing example, and the third packing example may be combined. For example, multiple attribute maps may be packed into each component of each layer corresponding to an image, and placement information indicating the layer, component, position, and size of each attribute map may be encoded.
[0421] In the example in Figure 66, start_x and start_y may be signaled, similar to the examples in Figure 59, etc. Also, in the example in Figure 66, video_component_id may be signaled, similar to the examples in Figure 63, etc., and video_sub_component_id may also be signaled.
[0422] Figure 67 is a conceptual diagram illustrating an example of unpacking multiple attribute maps. For example, the decoding device 200 unpacks multiple attribute maps from an image by separating the multiple attribute maps from an image that contains multiple attribute maps of different attribute types. Here, the unpacking of each attribute map is performed using placement information.
[0423] Furthermore, the decoding device 200 may decode only some attribute maps corresponding to some attribute types, rather than decoding all attribute maps corresponding to all attribute types. In other words, the decoding device 200 may extract and output some attribute maps from the image.
[0424] Figure 68 is a conceptual diagram showing an example of an attribute map request and response. For example, application device 2112 requests an attribute map by specifying the attribute type. Decoder 200 responds with an attribute map for the specified attribute type.
[0425] The application device 2112 may specify an attribute type in a large (coarse) unit. The decoding device 200 may then respond with multiple attribute types in a small (fine) unit that are included in the attribute type in the large (coarse) unit, and multiple attribute maps for the multiple attribute types in the small (fine) unit.
[0426] The application device 2112 may also be a renderer 2111. The decoding device 200 may include the application device 2112, or the application device 2112 may include the decoding device 200.
[0427] Figure 69 is a conceptual diagram illustrating an example of rendering using multiple attribute maps. The renderer 2111 renders a 3D mesh using multiple attribute maps of multiple attribute types. For example, the renderer 2111 renders a 3D mesh using a combination of the first attribute in the first attribute map of the first attribute type and the second attribute in the second attribute map of the second attribute type.
[0428] The renderer 2111 may render the 3D mesh using all attribute types, or it may render the 3D mesh using one or more of all attribute types. In other words, the renderer 2111 may render the 3D mesh using all attribute maps for all attribute types, or it may render the 3D mesh using one or more attribute maps for one or more attribute types. The attribute type determines how the attributes in the attribute map are reflected in the rendering of the 3D mesh.
[0429] In one example, a 3D mesh is rendered using multiple attribute types together by clipping sample values obtained by adding multiple sample values, each reflecting multiple attributes of multiple different attribute types. In another example, a 3D mesh is rendered using multiple attribute types together by applying a mathematical function based on attribute identifiers to multiple sample values, each reflecting multiple attributes of multiple different attribute types.
[0430] In the examples above, multiple attribute maps of a 3D mesh are combined into a single image frame unit, which corresponds to one image frame. This can potentially reduce the overhead related to the header and metadata of the encoded data in signaling multiple attribute maps of a 3D mesh compared to encoding and decoding multiple attribute maps individually. Furthermore, since multiple attribute maps of a 3D mesh are decoded simultaneously by decoding a single image frame unit, a reduction in complexity can also be achieved.
[0431] Furthermore, since it becomes possible to encode and decode multiple attribute maps together, the size and complexity of the hardware and software implementation can be reduced.
[0432] One of the multiple aspects of this disclosure may be adopted, or two or more aspects may be adopted in combination. Furthermore, one of the multiple processing elements, multiple components, and multiple syntax elements of this disclosure may be adopted, or two or more elements may be adopted in combination. In particular, not all elements of this disclosure are necessarily required, and only some elements of this disclosure may be adopted.
[0433] Furthermore, processing corresponding to the processing performed by the encoding device 100 may be performed by the decoding device 200, and processing corresponding to the processing performed by the decoding device 200 may be performed by the encoding device 100.
[0434] In the examples above, with regard to the encoding of the 3D mesh, multiple attribute maps for the faces of the 3D mesh are packed into a single image frame. However, this is not limited to attribute maps of a 3D mesh; multiple attribute maps of other 3D models may also be packed into a single image frame.
[0435] For example, multiple attribute maps of a 3D model, such as point cloud data, NeRF (Neural Radiance Fields), or 3DGS (3D Gaussian Platting), may be packed into a single image frame. Specifically, in the transmission of a projected image generated by projecting the 3D model onto an image, the projected image may be packed as attribute maps into a single image frame.
[0436] Specifically, for each of the multiple attribute types such as color, transparency, and reflectance, the attributes of the 3D model may be projected onto the image, thereby generating multiple projected images corresponding to the multiple attribute types as multiple attribute maps corresponding to the multiple attribute types.
[0437] Furthermore, attribute types may correspond to projection conditions such as viewpoint, projection direction, or projection plane. Multiple projection images based on multiple different projection conditions may be treated as multiple attribute maps with different attribute types. Alternatively, multiple attribute maps corresponding to multiple projection images, regardless of projection conditions, may be assigned a single attribute type: projection image.
[0438] For example, attribute information of a 3D model or attribute maps such as projected images, defined using methods such as DMC (Dynamic Mesh Coding) or V3C (ISO / IES23090-5), may be packed into a single image frame.
[0439] The attribute information in this disclosure is information associated with points or faces that constitute a three-dimensional object (three-dimensional model) to be processed, and is different from geometric information such as the coordinates, size, or shape of the point or face in space. The attribute information may include not only image information expressed using color or brightness, but also other information represented by numbers, variables, or vectors. Furthermore, the attribute information may include information indicating the correspondence between the attribute information and the geometric information.
[0440] The three-dimensional object (three-dimensional model) to be processed in this disclosure is associated with multiple attribute information sets. Multiple attribute information sets with different attribute types may be used, or multiple attribute information sets with the same attribute type may be used. For example, multiple images corresponding to the same three-dimensional object (three-dimensional model) may be used as multiple attribute information sets.
[0441] Furthermore, attribute information may be information to which a different mapping from the geometric information is applied to the three-dimensional object (three-dimensional model) being processed. For example, multiple images created by projecting a three-dimensional object (three-dimensional model) onto multiple planes may be used as multiple sets of attribute information. Alternatively, multiple images created by projecting a three-dimensional object (three-dimensional model) onto a plane with multiple layers, such as multiple resolutions, may be used as multiple sets of attribute information.
[0442] Multiple attribute information sets may be included in a single image frame unit by any of the multiple packing examples provided in this disclosure.
[0443] Specifically, as in the first packing example, multiple attribute information sets may be included in a single image frame. Alternatively, as in the second packing example, multiple attribute information sets may be included in multiple distinct components that make up a single image frame. Alternatively, as in the third packing example, multiple attribute information sets may be included in multiple distinct layers corresponding to a single image frame.
[0444] Alternatively, multiple attribute information sets may be packed into a single image frame unit by any combination of two or more packing examples from the first packing example, the second packing example, and the third packing example described above.
[0445] Further, for example, a plurality of different layers corresponding to one image frame may correspond to a plurality of image frames corresponding to each other in a plurality of different layers. Specifically, a plurality of different layers corresponding to one image frame may correspond to a plurality of image frames having the same metadata or the same time as one image frame. In other words, a plurality of different layers corresponding to one image frame may correspond to one access unit.
[0446] That is, one image frame unit corresponding to one image frame may be one access unit. Alternatively, one image frame unit may be a picture unit.
[0447] In the present disclosure, the image frame may correspond to a picture of an image or a field of an image. The expression of the image frame may be replaced with an image.
[0448] Further, the image frame may be composed of a plurality of sub-image frames corresponding to a plurality of components, or may be composed of a plurality of sub-image frames corresponding to a plurality of layers. And a plurality of attribute information sets may be packed into a plurality of sub-image frames constituting the image frame.
[0449] Further, the attribute values in the attribute information sets included in the plurality of attribute information sets may be expressed as difference values with respect to the attribute values in other attribute information sets.
[0450] Each of the plurality of attribute information sets may be stored in a processing unit capable of parallel processing in the image frame. Specifically, each of the plurality of attribute information sets may be stored in a tile, a brick, a slice, or the like.
[0451] The same packing method may be applied to a plurality of three-dimensional objects (three-dimensional models) such as a plurality of three-dimensional mesh frames, or a plurality of different packing methods may be applied.
[0452] The encoding device 100 may transmit an encoded stream including a plurality of attribute information sets generated by the encoding method in the present disclosure. The decoding device 200 may receive the encoded stream. The encoding device 100 and the decoding device 200 may encode, transmit, receive, and decode only the attribute information set specified by, for example, metadata included in the encoded stream or by the encoding device 100 or the decoding device 200, among the plurality of attribute information sets.
[0453] The encoding device 100 and the decoding device 200 may share in advance the attribute type and arrangement information for each of the plurality of attribute maps. Then, the signaling of the attribute type and arrangement information may be omitted.
[0454] <Attribute Video Encoding and Decoding> In V-DMC, a video codec can be used for encoding various types of attribute information that make up a mesh. For example, the attribute information is embedded in the image area of the video. Then, by encoding the video with the attribute information embedded in the image area, the attribute information is encoded. Also, by decoding the video with the attribute information embedded in the image area, the attribute information is decoded.
[0455] Embedding the attribute information in the image area of the video can also be expressed as mapping or packing the attribute information in the image area of the video. Also, the video with the attribute information embedded in the image area can also be expressed as an attribute video. The attribute video corresponds to the above-mentioned attribute map and the like. That is, the expressions of the attribute video and the attribute map and the like may be replaced with each other.
[0456] Note that the attribute video includes a plurality of frames corresponding to a plurality of time points. And for each of the plurality of time points, the attribute information corresponding to that time point is mapped to the frame corresponding to that time point.
[0457] Two or more types of attribute information may be mapped to a single attribute video, or one type of attribute information may be mapped to two or more attribute videos. For example, when multiple types of attribute information are mapped to a single attribute video, the image area of the single attribute video is divided into multiple areas, and one type of attribute information is mapped to each area. Each area may be an area that is represented as a tile in the video codec standard.
[0458] For example, a tile is a region obtained by dividing the image area of an attribute video into one or more rows and one or more columns. However, tiles may be defined by other division methods. Also, while tiles are primarily shown here as examples of each region, each region is not limited to tiles. Each region may be a region expressed as a slice or subpicture in the video codec standard, or it may be any other region.
[0459] Furthermore, for example, one type of attribute information may be mapped to multiple areas. On the other hand, multiple types of attribute information are not mapped to a single area, but rather each is mapped to multiple different areas.
[0460] Figure 70 is a conceptual diagram showing an example of tiles defined in the image region of an attribute video. In this example, three tiles are defined in the image region of the attribute video. Three types of attribute information may each be mapped to one of the three tiles. Alternatively, one type of attribute information may be divided and mapped across three tiles. In the example of Figure 70, attribute maps corresponding to the attribute types shown in Figure 58 can be implemented by the tiles.
[0461] Here, three tiles are defined in the image region, but one tile, two tiles, or four or more tiles may be defined in the image region.
[0462] Note that start_x and start_y correspond to the tile position, attribute_width corresponds to the tile width, and attribute_height corresponds to the tile height.
[0463] Figure 71 is a conceptual diagram illustrating an example of the relationship between multiple types of attribute information and multiple attribute videos. In this example, four types of attribute information—first, second, third, and fourth—are mapped to the first and second attribute videos. The two attribute videos, the first and second attribute videos, may be encoded into two bitstreams and transmitted via these two bitstreams. These two bitstreams may also be two sub-bitstreams constituting a single bitstream.
[0464] As described above, Figure 71 shows an example of the relationship between multiple types of attribute information and multiple attribute videos. The relationship between multiple types of attribute information and multiple attribute videos is not limited to the example in Figure 71. In particular, the number of types of attribute information and the number of attribute videos are not limited to the example in Figure 71.
[0465] When one or more tiles are defined in the image area of an attribute video, tile information for defining each tile is encoded as metadata and included in the bitstream. For example, the tile information indicates the number and arrangement of the tiles. The tile arrangement may include the position and size of the tiles, etc., where the tile position may be the coordinates of the top-left corner of the tile in the image area, and the tile size may be the width and height of the tile. The tile information may also be expressed as tiling information or atlas frame attribute tiling information.
[0466] Figure 72 is a conceptual diagram illustrating an example of the relationship between multiple attribute videos and tile information. When multiple attribute videos exist, the tile information may or may not be the same across the attribute videos.
[0467] Here, if the tile information is identical between attribute videos, a single, identical tile information is generated and transmitted between the attribute videos. Then, the same tile information is referenced between the attribute videos. For example, the tile information corresponding to the ref_idxth attribute video among multiple attribute videos is referenced by multiple attribute videos. This can reduce the amount of coding and processing required.
[0468] On the other hand, if the tile information is not the same across attribute videos, tile information is generated and transmitted for each attribute video. This makes it possible to change the tile information for each attribute video.
[0469] For example, whether tile information is identical between attribute videos corresponds to whether tile information such as the number and arrangement of tiles is consistent between attribute videos, and also corresponds to whether the same single tile information is referenced between attribute videos. These expressions may be interchangeable.
[0470] Figure 73 is a flowchart illustrating yet another example of the processing performed by the encoding device 100. For example, the encoding device 100 encodes three-dimensional data into a bitstream. In doing so, the encoding device 100 encodes attribute video defined in the three-dimensional data into a bitstream. Figure 73 shows the processing related to the encoding of attribute video.
[0471] In this example, first, the encoding device 100 encodes a parameter indicating the number of attribute videos into a bitstream (S501). The number of attribute videos is the number of attribute videos to be encoded. The parameter indicating the number of attribute videos may also be expressed as the first parameter or count parameter. Examples of attribute information encoded via attribute videos may include information such as texture, color, surface normal, vertex normal, reflectance, or transmittance.
[0472] The 3D mesh may include non-video attribute information, which is attribute information not encoded via attribute video. The non-video attribute information may be encoded in the base mesh subbitstream.
[0473] Furthermore, if the 3D mesh does not have attribute video, the encoding of the first parameter may be omitted. In this case, the value of the first parameter may be assumed to be equal to 0. This can reduce the amount of coding required.
[0474] The encoding device 100 then determines whether the number of attribute videos is greater than 1 (S502). If the number of attribute videos is greater than 1 (Yes in S502), the encoding device 100 encodes a flag into the bitstream indicating whether the tile information is identical between the attribute videos (S503). This flag may also be expressed as a second flag or identical information flag.
[0475] Then, the encoding device 100 encodes the tile information into a bitstream depending on whether the tile information is the same between attribute videos (S504).
[0476] Tile information may be represented by a parameter set consisting of one or more parameters. The encoding device 100 may encode the tile information by encoding the parameter set representing the tile information. This parameter set may also be referred to as a third parameter set or a tile parameter set. The tile information indicates the number and arrangement of tiles and may be used to combine attribute information with 3D mesh geometry.
[0477] For example, in the encoding of the second flag (S503), a second flag may be encoded with a value (1: True) indicating that the tile information is identical between attribute videos. In this case, the tile information is identical between attribute videos, and the same tile information is encoded and used for all videos. This may reduce the amount of encoding required. In this case, the index of the attribute video to which the tile information is associated is encoded, and the tile information associated with that attribute video is also used for at least one other attribute video.
[0478] On the other hand, for example, in the encoding of the second flag (S503), a second flag may be encoded with a value (0: False) indicating that the tile information is not the same between attribute videos. In this case, the tile information is not the same between attribute videos, and individual tile information is encoded and used for each attribute video. In other words, in this case, tile information is encoded for each attribute video. As a result, unique tile information suitable for the attribute information is determined for each attribute video, which can lead to the efficient embedding of attribute information in each attribute video and a reduction in the amount of encoding required.
[0479] If the number of attribute videos is not greater than 1 (No in S502), the encoding device 100 encodes the tile information according to the number of attribute videos (S505).
[0480] If the number of attribute videos is not greater than 1, the number of attribute videos is 0 or 1. Therefore, in this case, the tile information used for one attribute video is not used for other attribute videos. Thus, the second flag, which indicates whether the tile information is the same between attribute videos, is not encoded. In this case, the second flag may be considered to have a value (0: False) indicating that the tile information is not the same between attribute videos.
[0481] Furthermore, for example, if the number of attribute videos is 0, the encoding device 100 does not encode tile information. If the number of attribute videos is 1, the encoding device 100 encodes tile information associated with one attribute video.
[0482] Then, the encoding device 100 encodes the attribute videos according to the number of attribute videos (S506).
[0483] In the above determination process (S502), the encoding device 100 may determine whether the number of attribute videos is greater than 0, instead of whether the number of attribute videos is greater than 1. If the number of attribute videos is greater than 0, the second flag may be encoded.
[0484] Also, the encoding process may be performed layer by layer. That is, when encoding the three-dimensional data into a bitstream, the encoding device 100 may encode each of the plurality of layers constituting the three-dimensional data into a sub-bitstream constituting the bitstream.
[0485] For example, for each layer, the encoding device 100 encodes a parameter indicating the number of attribute videos, which is the number of videos to be encoded in the layer, into the sub-bitstream. Also, when the number of attribute videos in a layer is greater than 1, the encoding device 100 encodes a flag indicating whether the region information is the same among the attribute videos in the layer into the sub-bitstream.
[0486] Also, for each layer, the encoding device 100 encodes tile information into the sub-bitstream according to the number of attribute videos in the layer and encodes the attribute videos in the layer into the sub-bitstream.
[0487] Here, the layer may be defined based on a mesh sequence, a mesh frame, a sub-mesh, a group of faces, or LoD.
[0488] The attribute video may be encoded into a video sub-bitstream constituting the bitstream. The video sub-bitstream may be defined as a stream of attribute information encoded by a video codec among the components of the encoded data of V-DMC.
[0489] Also, the fact that the number of attribute videos is greater than 1 corresponds to the number of attribute videos being 2 or more and corresponds to the existence of a plurality of attribute videos. In this case, a second flag indicating whether the tile information is the same among the attribute videos is signaled.
[0490] When the tile information is the same among the attribute videos, the same and single tile information is signaled among the attribute videos, and index information for referring to the tile information is signaled. When the tile information is not the same among the attribute videos, the tile information of each of the two or more attribute videos is signaled.
[0491] If the number of attribute videos is 1 or less (i.e., 0 or 1), tile information is signaled according to the number of attribute videos.
[0492] Figure 74 is a flowchart illustrating yet another example of the processing of the decoding device 200. For example, the decoding device 200 decodes three-dimensional data from a bitstream. In doing so, the decoding device 200 decodes attribute video defined in the three-dimensional data from the bitstream. Figure 74 shows the processing related to the decoding of attribute video.
[0493] In this example, first, the decoding device 200 decodes a parameter indicating the number of attribute videos from the bitstream (S601). The number of attribute videos is the number of attribute videos to be decoded. The parameter indicating the number of attribute videos may also be expressed as the first parameter or count parameter. Examples of attribute information decoded via attribute videos may include information such as texture, color, surface normal, vertex normal, reflectance, or transmittance.
[0494] The 3D mesh may contain non-video attribute information, which is attribute information that is not decoded via attribute video. The non-video attribute information may be decoded from the base mesh subbitstream.
[0495] Furthermore, if the 3D mesh does not have attribute video, decoding of the first parameter may be omitted. In this case, the value of the first parameter may be estimated to be equal to 0. This can reduce the amount of coding required.
[0496] The decoding device 200 then determines whether the number of attribute videos is greater than 1 (S602). If the number of attribute videos is greater than 1 (Yes in S602), the decoding device 200 decodes a flag from the bitstream indicating whether the tile information is identical between the attribute videos (S603). This flag may also be expressed as a second flag or identical information flag.
[0497] Then, the decoding device 200 decodes the tile information from the bitstream depending on whether the tile information is the same between attribute videos (S604).
[0498] The tile information may be represented by a parameter set consisting of one or more parameters. The decoding device 200 may decode the tile information by decoding the parameter set representing the tile information. This parameter set may also be referred to as a third parameter set or a tile parameter set. The tile information indicates the number and arrangement of tiles and may be used to combine attribute information with 3D mesh geometry.
[0499] For example, in the decoding of the second flag (S603), the second flag may be decoded if it is set to a value (1: True) indicating that the tile information is the same across attribute videos. In this case, the tile information is the same across attribute videos, and the same tile information is decoded and used for all videos. This may reduce the amount of coding required. In this case, the index of the attribute video to which the tile information is associated is decoded, and the tile information associated with that attribute video is also used for at least one other attribute video.
[0500] On the other hand, for example, in the decoding of the second flag (S603), the second flag may be decoded with a value (0: False) that indicates that the tile information is not the same between attribute videos. In this case, the tile information is not the same between attribute videos, and individual tile information is decoded and used for each attribute video. In other words, in this case, tile information is decoded for each attribute video. As a result, unique tile information suitable for the attribute information is determined for each attribute video, which may allow attribute information to be efficiently embedded in each attribute video and reduce the amount of coding.
[0501] If the number of attribute videos is not greater than 1 (No in S602), the decoding device 200 decodes the tile information according to the number of attribute videos (S605).
[0502] If the number of attribute videos is not greater than 1, the number of attribute videos is 0 or 1. Therefore, in this case, the tile information used for one attribute video is not used for other attribute videos. Thus, the second flag, which indicates whether the tile information is the same between attribute videos, is not decoded. In this case, the second flag may be considered to have a value (0: False) indicating that the tile information is not the same between attribute videos.
[0503] Furthermore, for example, if the number of attribute videos is 0, the decoding device 200 does not decode the tile information. If the number of attribute videos is 1, the decoding device 200 decodes the tile information associated with one attribute video.
[0504] Then, the decoding device 200 decodes the attribute videos according to the number of attribute videos (S606).
[0505] In the above determination process (S602), the decoding device 200 may determine whether the number of attribute videos is greater than 0, instead of whether the number of attribute videos is greater than 1. If the number of attribute videos is greater than 0, the second flag may be decoded.
[0506] Furthermore, the decoding process may be performed for each layer. In other words, when the decoding device 200 decodes three-dimensional data from a bitstream, it may decode each of the multiple layers constituting the three-dimensional data from the sub-bitstreams constituting the bitstream.
[0507] For example, the decoding device 200 decodes a parameter from the sub-bitstream for each layer, which is the number of attribute videos to be decoded in that layer. Also, for each layer, if the number of attribute videos in that layer is greater than 1, the decoding device 200 decodes a flag from the sub-bitstream indicating whether or not the region information is the same between the attribute videos in that layer.
[0508] Furthermore, the decoding device 200 decodes tile information from the sub-bitstream for each layer according to the number of attribute videos in that layer, and decodes attribute videos in that layer from the sub-bitstream.
[0509] Here, layers may be defined based on mesh sequences, mesh frames, submeshes, groups of faces, or Line of D.
[0510] Figure 75 is a syntax diagram showing an example of a syntax structure for signaling tile information. For example, this syntax structure is included in a Frame Parameter Set (FPS). However, this syntax structure may also be included in a Sequence Parameter Set (SPS) or in other data unit headers.
[0511] Here, `layer_attribute_video_count` is a parameter that indicates the number of attribute videos. In other words, `layer_attribute_video_count` indicates the number of attribute videos signaled via the video subbitstream.
[0512] Additionally, `layer_attribute_video_atlas_consistent_tiling_flag` is a flag that indicates whether or not the tile information is identical between attribute videos.
[0513] Specifically, if `layer_attribute_video_atlas_consistent_tiling_flag` is equal to 1, it indicates that the tile information is identical across attribute videos. In this case, `layer_reference_attribute_index` exists, and the same tile information is used for all attribute videos.
[0514] On the other hand, if layer_attribute_video_atlas_consistent_tiling_flag is equal to 0, it indicates that the tile information is not identical between attribute videos. In this case, layer_reference_attribute_index does not exist, and different tile information is used for each attribute video.
[0515] Furthermore, layer_attribute_video_atlas_consistent_tiling_flag is presumed to have a value of 0 (False) if layer_attribute_video_count is less than 2. In other words, in this case, it is presumed that the tile information is not identical between attribute videos.
[0516] Additionally, if present, `layer_reference_attribute_index` indicates the index of the attribute video associated with the referenced tile information. For example, tile information of an attribute video having the same index as `layer_reference_attribute_index` may be used for an attribute video having a different index than `layer_reference_attribute_index`.
[0517] Furthermore, the value of layer_reference_attribute_index is within the range of 0 and layer_attribute_video_count-1. Also, if layer_reference_attribute_index does not exist, it is assumed to have a value of 0. In other words, in this case, the tile information of the attribute video with index 0 is referenced.
[0518] Furthermore, decode_atlas_layer_attribute_tiling_information(i) is a parameter set that indicates the tile information associated with the attribute video whose index is i.
[0519] As shown in Figure 75, if the number of attribute videos is greater than 1, a flag is signaled indicating whether or not the tile information is identical between the attribute videos.
[0520] Then, if the tile information is identical between attribute videos, an index for referencing the tile information is signaled. Then, the tile information associated with the signaled index is signaled.
[0521] If the tile information is not identical between attribute videos, the tile information associated with the index of that attribute video is signaled for each attribute video.
[0522] Figure 76 is a syntax diagram showing another example of a syntax structure for signaling tile information. Figure 76 uses the same syntax elements as Figure 75. In this example, each syntax element is processed when the number of attribute videos is greater than one. That is, not only is the signaling of flags indicating whether tile information is identical between attribute videos performed, but other processing is also performed when the number of attribute videos is greater than one. Therefore, the amount of coding and processing is reduced.
[0523] Figure 77 is a syntax diagram showing another example of a syntax structure for signaling tile information. Figure 77 uses the same syntax elements as Figure 75. In this example, the `layer_attribute_video_atlas_consistent_tiling_flag` is signaled regardless of the number of attribute videos. For example, if the flag exists in another header such as an SPS, or if it is referenced earlier in the processing of a higher layer, the flag is always signaled.
[0524] In this example, if the number of attribute videos is greater than 1 and layer_attribute_video_atlas_consistent_tiling_flag is 1 (True), then layer_reference_attribute_index is signaled. This makes it possible to adaptively omit the signaling of layer_reference_attribute_index, thereby reducing the number of signaling bits.
[0525] Alternatively, layer_reference_attribute_index may always be signaled. This could make it easier to change the placement of the parameter and to refer to it. In this case as well, if the number of attribute videos is two or more, layer_attribute_video_atlas_consistent_tiling_flag is signaled, which can reduce the number of signaling bits.
[0526] Figure 78 is a syntax diagram showing another example of a syntax structure for signaling tile information. Figure 78 uses the same syntax elements as Figure 75. In this example, if the number of attribute videos is one or more, flags such as `layer_attribute_video_atlas_consistent_tiling_flag` are signaled.
[0527] For example, in a frame sequence, the number of attribute videos may vary from frame to frame. In this case, even if the number of attribute videos for a given frame is 1, the signaling of `layer_attribute_video_atlas_consistent_tiling_flag` may enable efficient processing in a common manner.
[0528] Furthermore, even in this case, it becomes possible to reduce the number of signaling bits with respect to layer_attribute_video_atlas_consistent_tiling_flag and layer_reference_attribute_index.
[0529] Figure 79 is a syntax diagram showing another example of a syntax structure for signaling tile information. Figure 79 uses the same syntax elements as Figure 75. While the syntax structure of this example differs from that of Figure 75, the actual processing is the same. Note that the initial value of att_idx, which indicates the index, may be 0 or other values.
[0530] In some of the above examples, it may be possible to reduce the amount of coding. For example, if the number of attribute videos is 0 or 1, a flag indicating whether the tile information is the same between attribute videos is not signaled. The flag is then presumed to have a value (0: False) indicating that the tile information is not the same between attribute videos.
[0531] Furthermore, if the number of attribute videos is 0, the signaling for the reference index of a non-existent attribute video is omitted because no attribute videos exist. This can help prevent errors caused by references to non-existent attribute videos.
[0532] The syntax tables provided above may be combined, or partial modifications may be applied. This may further reduce the amount of code.
[0533] <Representative Examples> The following are representative examples of the encoding and decoding processes described above. Specific processes described above may be added to these representative examples.
[0534] Figure 80 is a flowchart showing an example of basic encoding processing according to this embodiment. For example, the circuit 151 of the encoding device 100 shown in Figure 24 uses the memory 152 to perform the encoding processing shown in Figure 80 during operation.
[0535] Specifically, circuit 151 encodes a flag indicating whether the region information for defining one or more regions in the image region of an attribute video is the same among the attribute videos, depending on the number of attribute videos to be encoded in the three-dimensional data (S701). Here, if the number of attribute videos is two or more, the flag is encoded, and if the number of attribute videos is one or less, the flag is not encoded.
[0536] This may allow for the omission of encoding a flag indicating whether or not region information is identical between attribute videos, depending on the number of attribute videos. Consequently, it may be possible to reduce the amount of coding and processing required.
[0537] For example, if there are two or more attribute videos and the region information is the same across all attribute videos, the region information shared by multiple attribute videos may be encoded. Conversely, if there are two or more attribute videos and the region information is not the same across all attribute videos, the region information may be encoded for each attribute video. This makes it possible to efficiently encode the region information depending on whether the region information is the same across all attribute videos when there are two or more attribute videos. Consequently, it may be possible to reduce the amount of encoding and processing required.
[0538] Furthermore, for example, if the number of attribute videos is 1, the region information may be encoded. Conversely, if the number of attribute videos is 0, the region information does not need to be encoded. This makes it possible to omit the encoding of region information depending on the number of attribute videos. Consequently, it may be possible to reduce the amount of encoding and processing required.
[0539] Furthermore, for example, if the number of attribute videos is one or more, one or more attribute videos may be encoded. This may make it possible to encode the attribute videos and region information according to the number of attribute videos. Therefore, it may be possible to encode attribute videos and region information efficiently.
[0540] Furthermore, for example, if the number of attribute videos is two or more and the region information is the same across the attribute videos, an index for referencing the region information may be encoded. This may make it possible to efficiently identify and reference the same region information across attribute videos when the number of attribute videos is two or more. Therefore, it may be possible to efficiently define one or more regions within the image region of an attribute video.
[0541] Furthermore, for example, if the number of attribute videos is one or less, or if the region information is not the same across attribute videos, the index does not need to be encoded. This makes it possible to omit the encoding of the index depending on the number of attribute videos and whether the region information is the same across attribute videos. Consequently, it may be possible to reduce the amount of coding and processing required.
[0542] Furthermore, for example, if the number of attribute videos is one or less, or if the region information is not the same across attribute videos, the index may be assigned sequentially from 0 for each attribute video. This may make it possible to efficiently identify and reference the region information assigned to each attribute video. Therefore, it may be possible to efficiently define one or more regions within the image region of an attribute video.
[0543] Furthermore, for example, a flag may be encoded for each layer of three-dimensional data according to the number of attribute videos. This may make it possible to omit the encoding of a flag indicating whether or not the region information is identical between attribute videos for each layer, according to the number of attribute videos. Therefore, it may be possible to reduce the amount of coding and processing required.
[0544] Furthermore, for example, the region information may be tile information for defining one or more tiles in an image region. This may make it possible to omit the encoding of a flag indicating whether the tile information is identical between attribute videos, depending on the number of attribute videos. Consequently, it may be possible to reduce the amount of coding and processing required.
[0545] Furthermore, for example, circuit 151 may map the attribute information of the three-dimensional data to one or more regions of one or more attribute videos. This may make it possible to generate one or more attribute videos in which the attribute information is mapped to one or more regions.
[0546] Furthermore, the fact that the region information is identical across attribute videos corresponds to the region information being commonly defined and not individually defined for each attribute video. Conversely, the fact that the region information is not identical across attribute videos corresponds to the region information not being commonly defined and being individually defined for each attribute video. Therefore, whether or not the region information is identical across attribute videos can be replaced with whether or not the region information is commonly defined, or with whether or not the region information is individually defined for each attribute video.
[0547] Furthermore, for example, if a flag indicating whether or not the region information is the same across attribute videos is not encoded, the region information may be considered to be different across attribute videos. In other words, in this case, the region information may be considered not to be commonly defined, but to be defined individually for each attribute video. And in this case, the index for referencing the region information does not need to be encoded.
[0548] Furthermore, the three-dimensional data mentioned above may be either a three-dimensional model or a three-dimensional mesh.
[0549] Furthermore, although the circuit 151 of the encoding device 100 performs each process as described above, the attribute information encoder 103 of the encoding device 100 may also perform each of the above processes. Alternatively, the preprocessor 1103 and compressor 1106 of the encoding device 100 may perform each of the above processes. Alternatively, other components may perform each of the above processes.
[0550] Figure 81 is a flowchart showing an example of a basic decoding process according to this embodiment. For example, the circuit 251 of the decoding device 200 shown in Figure 25 uses the memory 252 to perform the decoding process shown in Figure 81 during operation.
[0551] Specifically, circuit 251 decodes a flag indicating whether the region information for defining one or more regions in the image region of an attribute video is the same among the attribute videos, depending on the number of attribute videos to be decoded in the three-dimensional data (S801). Here, if the number of attribute videos is two or more, the flag is decoded, and if the number of attribute videos is one or less, the flag is not decoded.
[0552] This may allow for the omission of decoding a flag indicating whether or not region information is identical between attribute videos, depending on the number of attribute videos. Consequently, it may be possible to reduce the amount of coding and processing required.
[0553] For example, if there are two or more attribute videos and the region information is the same across all attribute videos, the region information shared by multiple attribute videos may be decoded. Conversely, if there are two or more attribute videos and the region information is not the same across all attribute videos, the region information may be decoded for each individual attribute video. This allows for efficient decoding of region information depending on whether the region information is the same across all attribute videos when there are two or more attribute videos. Consequently, it may be possible to reduce the amount of coding and processing required.
[0554] Furthermore, for example, if the number of attribute videos is 1, the region information may be decoded. Conversely, if the number of attribute videos is 0, the region information does not need to be decoded. This makes it possible to omit the decoding of region information depending on the number of attribute videos. Consequently, it may be possible to reduce the amount of coding and processing required.
[0555] Furthermore, for example, if the number of attribute videos is one or more, one or more attribute videos may be decoded. This may make it possible to decode the attribute videos and the region information in accordance with the number of attribute videos. Therefore, it may be possible to efficiently decode both attribute videos and region information.
[0556] Furthermore, for example, if the number of attribute videos is two or more and the region information is the same across the attribute videos, the index for referencing the region information may be decoded. This may make it possible to efficiently identify and reference the same region information across attribute videos when the number of attribute videos is two or more. Therefore, it may be possible to efficiently define one or more regions within the image region of the attribute video.
[0557] Furthermore, for example, if the number of attribute videos is one or less, or if the region information is not the same across attribute videos, the index does not need to be decoded. This makes it possible to omit index decoded depending on the number of attribute videos and whether the region information is the same across attribute videos. Consequently, it may be possible to reduce the amount of coding and processing required.
[0558] Furthermore, for example, if the number of attribute videos is one or less, or if the region information is not the same across attribute videos, the index may be assigned sequentially from 0 for each attribute video. This may make it possible to efficiently identify and reference the region information assigned to each attribute video. Therefore, it may be possible to efficiently define one or more regions within the image region of an attribute video.
[0559] Furthermore, for example, a flag may be decoded for each layer of three-dimensional data according to the number of attribute videos. This may make it possible to omit the decoding of a flag indicating whether or not the region information is identical between attribute videos for each layer, according to the number of attribute videos. Therefore, it may be possible to reduce the amount of coding and processing required.
[0560] Furthermore, for example, the region information may be tile information for defining one or more tiles in an image region. This may make it possible to omit the decoding of a flag indicating whether the tile information is identical between attribute videos, depending on the number of attribute videos. Therefore, it may be possible to reduce the amount of coding and processing required.
[0561] Furthermore, for example, circuit 251 may extract attribute information of three-dimensional data from one or more regions of one or more attribute videos. This may make it possible to derive attribute information from one or more attribute videos in which attribute information is mapped to one or more regions.
[0562] Furthermore, the fact that the region information is identical across attribute videos corresponds to the region information being commonly defined and not individually defined for each attribute video. Conversely, the fact that the region information is not identical across attribute videos corresponds to the region information not being commonly defined and being individually defined for each attribute video. Therefore, whether or not the region information is identical across attribute videos can be replaced with whether or not the region information is commonly defined, or with whether or not the region information is individually defined for each attribute video.
[0563] Furthermore, for example, if a flag indicating whether or not the region information is the same across attribute videos is not decoded, the region information may be considered to be different across attribute videos. In other words, in this case, the region information may be considered not to be commonly defined, but rather to be defined individually for each attribute video. And in this case, the index for referencing the region information may not be decoded.
[0564] Furthermore, the three-dimensional data mentioned above may be either a three-dimensional model or a three-dimensional mesh.
[0565] Furthermore, although the circuit 251 of the decoding device 200 performs each process as described above, the attribute information decoder 203 of the decoding device 200 may also perform each of the above processes. Alternatively, the expander 2102 and post-processor 2106 of the decoding device 200 may perform each of the above processes. Alternatively, other components may perform each of the above processes.
[0566] <Other Examples> Although the embodiments of the encoding device 100 and the decoding device 200 have been described above according to the embodiments, the embodiments of the encoding device 100 and the decoding device 200 are not limited to the embodiments. Modifications that a person skilled in the art can conceive of may be made to the embodiments, and the multiple components in the embodiments may be combined arbitrarily.
[0567] For example, in the embodiment, a process performed by a specific component may be performed by another component instead of that specific component. Also, the order of multiple processes may be changed, or multiple processes may be executed in parallel.
[0568] Furthermore, as described above, at least some of the configurations of this disclosure can be implemented as an integrated circuit. At least some of the processes of this disclosure may be used as an encoding method or a decoding method. A program for causing a computer to execute the encoding method or the decoding method may be used. A non-temporary computer-readable recording medium on which the program is recorded may also be used. A bitstream for causing the decoding device 200 to perform the decoding process may also be used.
[0569] Furthermore, at least some of the configurations and processes of this disclosure may be used as transmitting devices, receiving devices, transmitting methods, and receiving methods. A program for causing a computer to execute such transmitting method or receiving method may be used. A non-temporary computer-readable recording medium on which such program is recorded may also be used.
[0570] Furthermore, for example, the expression "at least one (or more) of the first element, second element, and third element" corresponds to the first element, second element, third element, or any combination thereof.
[0571] This disclosure is useful, for example, for encoding devices, decoding devices, transmitting devices, and receiving devices related to three-dimensional meshes, and is applicable to computer graphics systems and three-dimensional data display systems.
[0572] 100 Encoding Device 101, 121, 144 Vertex Information Encoder 102, 145 Connection Information Encoder 103, 122 Attribute Information Encoder 104, 204, 1103 Preprocessor 105, 205, 2106 Postprocessor 110 Three-Dimensional Data Encoding System 111, 211 Controller 112, 212 Input / Output Processor 113 Three-Dimensional Data Encoder 114 System Multiplexer 115 Three-Dimensional Data Generator 123 Metadata Encoder 124, 1274 Multiplexer 131 Vertex Image Generator 132 Attribute Image Generator 133 Metadata Generator 134 Video Encoder 141 Two-Dimensional Data Encoder 142 Mesh Data Encoder 143 Texture Encoder 148 Description Encoder 151, 251 Circuit 152, 252 Memory 200 Decoder 201, 221, 244 Vertex Information Decoder 202, 245 Connection Information Decoder 203, 222 Attribute Information Decoder 210 Three-Dimensional Data Decoder System 213 Three-Dimensional Data Decoder 214 System Demultiplexer 215, 247 Presenter 216 User Interface 223 Metadata Decoder 224, 1231 Demultiplexer 231 Vertex Information Generator 232 Attribute Information Generator 234 Video Decoder 241 Two-Dimensional Data Decoder 242 Mesh Data Decoder 243 Texture Decoder 246 Mesh Reconstructor 248 Description Decoder 300 Network 310 External Connector 511 Volumetric Capturer 512 Projector 513 Base Mesh Encoder 514 Displacement Encoder 515 Attribute Encoder 516 Other Type Encoder 613 Base Mesh Decoder 614 Displacement Decoder 615 Attribute Decoder 616 Other Type Decoder 617 Three-Dimensional Reconstructor 1101 Input Mesh 1102, 2108 Attribute Map 1104, 2103 Base Mesh 1105, 2104 Displacement Data 1106 Compressor 1107, 1304, 2101 Bitstream1108, 2105 Metadata 1110, 2110 Images 1232, 1262 Switches 1233 Static Mesh Decoder 1234, 1264 Mesh Buffer 1235 Motion Decoder 1236, 1266 Base Mesh Reconstructor 1237, 1240 Inverse Quantizer 1238, 1243 Video Decoder 1239 Image Unpacker 1241 Inverse Wavelet Converter 1242, 2308 Reconstructor 1244, 1272 Color Converter 1251 Decoded Base Mesh 1252, 1402 Subpartitioner 1253 Subpartitioned Mesh 1254 Decoded Displacement Data 1255 Displacer 1256 Decoded 3D Mesh 1261, 1269 Quantizer 1263 Static Mesh Encoder 1265 Motion Encoder 1267 Displacement Data Updater 1268 Wavelet Converter 1270 Image Packer 1271, 1273 Video Encoder 1301, 2304 Mesh Frame 1302, 2301, 2302 Base Mesh Frame 1303, 2303 Displacement Information 1401 Base Mesh Generator 1403 Displacement Data Generator 2102 Extensor 2107 Output Mesh 2109 Coupler 2111 Renderer 2112 Application Device 2305 Decoder 2306 Post-Decoder 2307 Pre-Reconstructor 2309 Post-Reconstructor 2310 Adapter 2311 Final 3D Mesh Frame
Claims
1. An encoding method comprising encoding a flag indicating whether or not region information for defining one or more regions in the image region of an attribute video is the same among attribute videos, depending on the number of attribute videos to be encoded in three-dimensional data, wherein the flag is encoded when the number of attribute videos is two or more, and the flag is not encoded when the number of attribute videos is one or less.
2. The encoding method according to claim 1, wherein if the number of attribute videos is two or more and the region information is the same among the attribute videos, the region information shared by the multiple attribute videos is encoded, and if the number of attribute videos is two or more and the region information is not the same among the attribute videos, the region information is encoded for each attribute video.
3. The encoding method according to claim 1 or 2, wherein the region information is encoded when the number of attribute videos is 1, and the region information is not encoded when the number of attribute videos is 0.
4. The encoding method according to claim 1 or 2, wherein if the number of attribute videos is one or more, one or more attribute videos are encoded.
5. The encoding method according to claim 1 or 2, wherein if the number of attribute videos is two or more and the region information is the same among the attribute videos, an index for referencing the region information is encoded.
6. The encoding method according to claim 5, wherein the index is not encoded if the number of attribute videos is one or less, or if the area information is not the same among the attribute videos.
7. The encoding method according to claim 1 or 2, wherein the flag is encoded for each layer of the three-dimensional data according to the number of attribute videos.
8. The encoding method according to claim 1 or 2, wherein the region information is tile information for defining one or more tiles in the image region.
9. A decoding method comprising decoding a flag indicating whether or not region information for defining one or more regions in the image region of an attribute video is the same among the attribute videos, depending on the number of attribute videos to be decoded in three-dimensional data, wherein the flag is decoded when the number of attribute videos is two or more, and the flag is not decoded when the number of attribute videos is one or less.
10. The decoding method according to claim 9, wherein if the number of attribute videos is two or more and the region information is the same among the attribute videos, the region information shared by the multiple attribute videos is decoded, and if the number of attribute videos is two or more and the region information is not the same among the attribute videos, the region information is decoded for each attribute video.
11. The decoding method according to claim 9 or 10, wherein the region information is decoded when the number of attribute videos is 1, and the region information is not decoded when the number of attribute videos is 0.
12. The decoding method according to claim 9 or 10, wherein if the number of attribute videos is one or more, one or more attribute videos are decoded.
13. The decoding method according to claim 9 or 10, wherein if the number of attribute videos is two or more and the region information is the same among the attribute videos, an index for referencing the region information is decoded.
14. The decoding method according to claim 13, wherein the index is not decoded if the number of attribute videos is one or less, or if the area information is not the same among the attribute videos.
15. The decoding method according to claim 9 or 10, wherein the flag is decoded for each layer of the three-dimensional data according to the number of attribute videos.
16. The decoding method according to claim 9 or 10, wherein the region information is tile information for defining one or more tiles in the image region.
17. An encoding device comprising a circuit and a memory accessible by the circuit, wherein the circuit, in operation, encodes a flag indicating whether or not the region information for defining one or more regions in the image region of an attribute video is the same among the attribute videos, depending on the number of attribute videos to be encoded in the three-dimensional data, the flag is encoded when the number of attribute videos is two or more, and the flag is not encoded when the number of attribute videos is one or less.
18. A decoding device comprising a circuit and a memory accessible by the circuit, wherein the circuit, in operation, decodes a flag indicating whether or not the region information for defining one or more regions in the image region of an attribute video is the same among the attribute videos, according to the number of attribute videos to be decoded in the three-dimensional data, the flag is decoded when the number of attribute videos is two or more, and the flag is not decoded when the number of attribute videos is one or less.
Citation Information
Patent Citations
Three-dimensional data coding method, three-dimensional data decoding method, three-dimensional data coding device, and three-dimensional data decoding device
JP2024023802A
Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
JP2024055947A