Encoding method, decoding method, encoding device, and decoding device

By encoding submeshes with threshold information for each combination, the method optimizes the encoding and decoding of three-dimensional data, reducing processing requirements and improving efficiency in combining three-dimensional points.

WO2025216315A1PCT designated stage Publication Date: 2025-10-16PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/014504
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-12
Filing Date
2025-04-11
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Existing encoding and decoding processes for three-dimensional data are inefficient, leading to excessive processing requirements and unnecessary operations in combining three-dimensional points of submeshes.

Method used

The proposed method generates multiple encoded data by encoding submeshes and sets threshold information for each combination of submeshes based on the positions of three-dimensional points, allowing for appropriate combination of points and reducing processing by omitting unnecessary operations through specific threshold values.

Benefits of technology

This approach reduces the amount of processing required for decoding by ensuring appropriate combination of three-dimensional points and minimizing unnecessary search processes, thereby optimizing the encoding and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025014504_16102025_PF_FP_ABST
    Figure JP2025014504_16102025_PF_FP_ABST
Patent Text Reader

Abstract

In an encoding method according to one aspect of the present disclosure, a plurality of items of encoded data are generated (S721) by encoding a plurality of sub-meshes, threshold value information indicating threshold values that are determined for each combination of two sub-meshes selected from the plurality of sub-meshes is generated (S722) on the basis of the locations of a plurality of three-dimensional points included in the plurality of sub-meshes, and a bitstream including the plurality of items of encoded data and the threshold value information is generated (S723).
Need to check novelty before this filing date? Find Prior Art

Description

Encoding method, decoding method, encoding device, and decoding device

[0001] The present disclosure relates to encoding methods and the like.

[0002] In US Pat. No. 6,299,549 a method and apparatus for encoding and decoding three-dimensional mesh data is proposed.

[0003] Japanese Patent Application Laid-Open No. 2006-187015

[0004] Further improvements are desired in the encoding or decoding process for three-dimensional data. The present disclosure aims to improve the encoding or decoding process for three-dimensional data.

[0005] An encoding method according to one aspect of the present disclosure generates multiple encoded data by encoding multiple submeshes, generates threshold information indicating a threshold value determined for each combination of two submeshes selected from the multiple submeshes based on the positions of multiple three-dimensional points included in the multiple submeshes, and generates a bitstream including the multiple encoded data and the threshold information.

[0006] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0007] The present disclosure may contribute to improvements in encoding processes and the like related to three-dimensional data.

[0008] 1 is a conceptual diagram showing a three-dimensional mesh according to an embodiment. FIG. 2 is a conceptual diagram showing basic elements of a three-dimensional mesh according to an embodiment. FIG. 3 is a conceptual diagram showing mapping according to an embodiment. FIG. 4 is a block diagram showing a configuration example of an encoding / decoding system according to an embodiment. FIG. 5 is a block diagram showing a configuration example of an encoding device according to an embodiment. FIG. 6 is a block diagram showing another configuration example of an encoding device according to an embodiment. FIG. 7 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 8 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 9 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 10 is a conceptual diagram showing another configuration example of a bit stream according to an embodiment. FIG. 11 is a conceptual diagram showing yet another configuration example of a bit stream according to an embodiment. FIG. 12 is a block diagram showing a specific example of an encoding / decoding system according to an embodiment. FIG. 13 is a conceptual diagram showing an example configuration of point cloud data according to an embodiment. FIG. 14 is a conceptual diagram showing an example data file of point cloud data according to an embodiment. FIG. 15 is a conceptual diagram showing an example configuration of mesh data according to an embodiment. FIG. 16 is a conceptual diagram showing an example data file of mesh data according to an embodiment. FIG. 17 is a conceptual diagram showing types of three-dimensional data according to an embodiment. FIG. 18 is a block diagram showing an example configuration of a three-dimensional data encoder according to an embodiment. FIG. 19 is a block diagram showing an example configuration of a three-dimensional data decoder according to an embodiment. FIG. 19 is a block diagram showing another configuration example of a three-dimensional data encoder according to an embodiment. FIG. 19 is a block diagram showing another configuration example of a three-dimensional data decoder according to an embodiment. FIG. 1 is a conceptual diagram showing a specific example of encoding processing according to an embodiment. FIG. 2 is a conceptual diagram showing a specific example of decoding processing according to an embodiment. FIG. 3 is a block diagram showing an implementation example of an encoding device according to an embodiment. FIG. 4 is a block diagram showing an implementation example of a decoding device according to an embodiment. FIG. 5 is a block diagram showing another configuration example of an encoding / decoding system according to an embodiment. FIG. 6 is a block diagram showing another configuration example of an encoding device according to an embodiment. FIG. 7 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 8 is a block diagram showing yet another configuration example of an encoding device according to an embodiment. FIG. 9 is a block diagram showing yet another configuration example of a decoding device according to an embodiment. FIG. 10 is a flow diagram showing processing of an encoding device according to an embodiment. FIG. 11 is an explanatory diagram conceptually showing encoding of a mesh frame according to an embodiment.1 is a flow diagram showing processing of a decoding device according to an embodiment. It is an explanatory diagram conceptually showing decoding of a mesh frame according to an embodiment. It is a block diagram showing an example of a configuration of a decoding device according to an embodiment. It is a block diagram showing an example of a configuration of an encoding device according to an embodiment. It is a block diagram showing an example of a configuration of a decoding device according to an embodiment. It is a block diagram showing an example of a configuration of a decoding device according to an embodiment. It is an explanatory diagram showing an example of subdivision according to an embodiment. It is an explanatory diagram showing an example of displacement of vertices after displacement after subdivision according to an embodiment. It is an explanatory diagram showing an example of vertices of an original mesh according to an embodiment. It is an explanatory diagram showing an example of a mesh according to an embodiment. It is an explanatory diagram showing an example of division of a mesh into sub-meshes according to an embodiment. It is a first explanatory diagram showing an example of packing of displacement information into an image frame according to an embodiment. It is a second explanatory diagram showing an example of packing of displacement information into an image frame according to an embodiment. It is a third explanatory diagram showing an example of packing of displacement information into an image frame according to an embodiment. It is a block diagram showing another example of the configuration of an encoding device according to an embodiment. It is a block diagram showing an example of the configuration of a preprocessor according to an embodiment. It is a block diagram showing another example of the configuration of a decoding device according to an embodiment. It is a block diagram showing a specific example of the configuration of a decoding device according to an embodiment. It is a diagram showing an example of two subdivided sub-meshes having boundary edges according to an embodiment. It is a flow diagram showing processing of a decoding device according to an embodiment. It is a diagram for explaining examples of boundary edges and non-boundary edges according to an embodiment. FIG. 1 is a diagram illustrating another example of boundary edges and non-boundary edges according to an embodiment. FIG. 2 is a diagram illustrating an example of syntax for signaling different subdivision types and subdivision counts in a header according to an embodiment. FIG. 3 is a diagram illustrating an example of syntax for signaling the subdivision type and subdivision count using a sequence parameter set according to an embodiment. FIG. 4 is a diagram illustrating an example of syntax for determining the subdivision type and subdivision count using a sequence parameter set according to an embodiment, and checking whether an edge is on a boundary. FIG. 5 is a flow diagram illustrating an example of a process for dividing multiple edges that constitute a submesh according to an embodiment. FIG. 6 is a diagram illustrating a first example of a process for dividing multiple edges that constitute a submesh according to an embodiment.10 is a diagram for explaining a second example of a process for dividing a plurality of edges that constitute a submesh according to an embodiment. FIG. 11 is a diagram for explaining a third example of a process for dividing a plurality of edges that constitute a submesh according to an embodiment. FIG. 12 is a diagram for explaining a fourth example of a process for dividing a plurality of edges that constitute a submesh according to an embodiment. FIG. 13 is a diagram for explaining a fifth example of a process for dividing a plurality of edges that constitute a submesh according to an embodiment. FIG. 14 is a diagram for explaining a sixth example of a process for dividing a plurality of edges that constitute a submesh according to an embodiment. FIG. 15 is a flow diagram showing a boundary vertex determination process performed by an encoding device according to an embodiment. FIG. 16 is a flow diagram showing a submesh overlap determination process performed by a decoding device according to an embodiment. FIG. 17 is a flow diagram showing a vertex modification process performed by a decoding device according to an embodiment. FIG. 18 is a diagram showing a specific example of a vertex modification process according to an embodiment. FIG. 19 is a diagram showing a specific example of a location where flag information according to an embodiment is stored. FIG. 19 is a diagram for explaining an example of a syntax in which flag information according to an embodiment is signaled. FIG. 19 is a diagram for explaining another example of a syntax in which flag information according to an embodiment is signaled. FIG. 19 is a diagram showing a specific example of a location where method information according to an embodiment is stored. FIG. 19 is a diagram for explaining an example of a syntax in which method information according to an embodiment is signaled. FIG. 1 is a diagram for explaining another example of a syntax in which method information according to an embodiment is signaled. FIG. 2 is a diagram for explaining a specific example of a vertex modification method according to an embodiment. FIG. 3 is a diagram for explaining an example of a relationship between the number of subdivisions of a plurality of submeshes and flag information according to an embodiment. FIG. 4 is a diagram for explaining another example of the relationship between the number of subdivisions of a plurality of submeshes and flag information according to an embodiment. FIG. 5 is a diagram for explaining an example of a syntax in which information indicating the type of overlap of a plurality of submeshes according to an embodiment is signaled. FIG. 6 is a diagram for explaining LoD according to an embodiment. FIG. 7 is a flow diagram for explaining vertex merging processing in a decoding device according to an embodiment. FIG. 8 is an explanatory diagram showing an example of vertex merging processing according to an embodiment. FIG. 9 is an explanatory diagram showing another example of vertex merging processing according to an embodiment.1 is an explanatory diagram showing another example of vertex joining processing according to an embodiment. FIG. 2 is a flow diagram explaining vertex joining processing for each LoD in a decoding device according to an embodiment. FIG. 3 is a flow diagram explaining vertex joining processing for LoD0 in a decoding device according to an embodiment. FIG. 4 is a diagram explaining another example of syntax in which method information according to an embodiment is signaled. FIG. 5 is an explanatory diagram showing another example of vertex joining processing according to an embodiment. FIG. 6 is an explanatory diagram showing another example of vertex joining processing according to an embodiment. FIG. 7 is a block diagram showing example configurations of an encoding device and a decoding device according to an embodiment. FIG. 8 is a flow diagram showing matching point search processing according to an embodiment. FIG. 9 is a diagram explaining the positional relationship of a plurality of submeshes according to an embodiment. FIG. 10 is a diagram explaining matching point search processing for a plurality of submeshes according to an embodiment. FIG. 11 is a diagram explaining matching point search processing for a plurality of submeshes according to an embodiment. FIG. 12 is a flow diagram showing matching point search processing according to an embodiment. FIG. 13 is a diagram explaining matching point search processing for a plurality of submeshes according to an embodiment. FIG. 14 is a diagram explaining matching point search processing for a plurality of submeshes according to an embodiment. FIG. 15 is a diagram explaining submesh joining metadata according to an embodiment. Fig. 1 is a diagram for explaining an example of a syntax in which sub-mesh combining metadata according to an embodiment is signaled. Fig. 2 is a diagram for explaining an example of a syntax in which sub-mesh combining metadata according to an embodiment is signaled. Fig. 3 is a diagram for explaining an example of a syntax in which information indicating a combination of sub-meshes to be combined according to an embodiment is signaled. Fig. 4 is a flow diagram showing an example of a basic encoding process according to an embodiment. Fig. 5 is a flow diagram showing an example of a basic decoding process according to an embodiment.

[0009] Introduction Three-dimensional (3D) meshes are used in computer graphics images, for example. For example, a computer graphics image may be composed of multiple temporally distinct frames, and each frame may be represented by a 3D mesh.

[0010] A 3D mesh is composed of vertex information indicating the positions of each of the vertices in 3D space, connectivity information indicating the connections between the vertices, and attribute information indicating the attributes of each vertex or face. Each face is constructed according to the connectivity between the vertices. Various computer graphics images can be expressed using such 3D meshes.

[0011] Furthermore, for transmission and storage of the 3D mesh, efficient encoding and decoding of the 3D mesh is expected. For efficient encoding and decoding of the 3D mesh, arithmetic coding and decoding may be used.

[0012] Further improvements are desired in the encoding or decoding process for three-dimensional data. The present disclosure aims to improve the encoding or decoding process for three-dimensional data.

[0013] Below, examples of inventions that can be obtained from the disclosure of this specification will be given, and the effects and the like that can be obtained from these inventions will be explained.

[0014] The encoding method of Example 1 generates multiple encoded data by encoding multiple submeshes, generates threshold information indicating a threshold value determined for each combination of two submeshes selected from the multiple submeshes based on the positions of multiple three-dimensional points included in the multiple submeshes, and generates a bitstream including the multiple encoded data and the threshold information.

[0015] The threshold is used, for example, by a decoding device that has acquired a bitstream to combine 3D points of multiple submeshes. Here, if the threshold is set based solely on the distance from the 3D point, for example, when there are multiple submeshes that can be combined with a submesh having the 3D point, there is a possibility that the 3D point will be combined with a 3D point of a submesh other than the submesh having the 3D point that should be combined. Therefore, by setting a threshold for each combination of submeshes, 3D points can be appropriately combined. Furthermore, for two submeshes that are not combined, a specific value, such as 0, can be set as the threshold information, allowing the decoding device to omit the process of determining whether or not to combine each 3D point for those two submeshes. This reduces the amount of processing required by the decoding device.

[0016] The encoding method of Example 2 is the encoding method of Example 1, and the combination may be determined based on a division method for dividing a three-dimensional mesh into the plurality of sub-meshes.

[0017] Depending on the division method (specifically, how the division is performed), multiple submeshes may only be combined with specific submeshes. In such cases, multiple submeshes can be appropriately combined without setting a threshold for every combination of multiple submeshes. This reduces the amount of processing required to determine the combinations. It also reduces the amount of information in the bitstream.

[0018] The encoding method of Example 3 may be the encoding method of Example 2, in which the combination is determined based on whether the division method is a first division method in which division is performed between two submeshes to which adjacent submesh identification numbers are assigned among the submesh identification numbers assigned consecutively to each of the plurality of submeshes, or a second division method other than the first division method.

[0019] When a three-dimensional mesh is divided into multiple submeshes, multiple submeshes may be generated by repeatedly dividing the three-dimensional mesh at predetermined intervals in a direction perpendicular to a single line, such as by repeatedly dividing the three-dimensional mesh at predetermined intervals in a predetermined direction. In such cases, the multiple submeshes are assigned submesh identification numbers consecutively in the order in which they were divided from the three-dimensional mesh. With this division method, a given submesh can only be combined with submeshes that are adjacent to the given submesh. In other words, with this division method, a given submesh can only be combined with submeshes that have an identification number adjacent to the identification number assigned to the given submesh.

[0020] The encoding method of Example 4 may be any of the encoding methods of Examples 1 to 3, and may include calculating a distance between each pair of two three-dimensional points selected from a plurality of three-dimensional points included in a first submesh of the two submeshes of the combination and a plurality of three-dimensional points included in a second submesh of the two submeshes, and calculating the longest distance between the calculated distances for each pair as the threshold for the combination.

[0021] This allows the first submesh and the second submesh to be appropriately combined using one threshold value.

[0022] The decoding method of Example 5 acquires a bitstream containing multiple coded data generated by coding multiple submeshes and threshold information indicating a threshold value determined for each combination of two submeshes selected from the multiple submeshes, decodes the multiple coded data, and performs a joining process to join multiple three-dimensional points included in the multiple submeshes based on the threshold information.

[0023] According to this, by setting a threshold for each combination of submeshes, 3D points can be appropriately combined. Furthermore, for two submeshes that are not combined, a specific value such as 0 is set as the threshold information, so that the process of determining whether or not to combine each 3D point for those two submeshes can be omitted. This reduces the amount of processing.

[0024] The decoding method of Example 6 is the decoding method of Example 5, and in the combining process, for each combination, if a predetermined condition is satisfied, a search process is performed to search for a pair of two three-dimensional points to be combined from among the multiple three-dimensional points included in the multiple submeshes, and the two three-dimensional points of the searched pair may be combined based on the threshold information.

[0025] This allows the search process to connect two suitable 3D points.

[0026] The decoding method of Example 7 is the decoding method of Example 6, wherein the combining process determines, for each combination, whether or not there is a pair of two three-dimensional points to be combined among the multiple three-dimensional points included in the two submeshes of the combination, and when it is determined that there is a pair of two three-dimensional points to be combined among the multiple three-dimensional points included in the two submeshes of the combination, the search process is performed, and when it is determined that there is not a pair of two three-dimensional points to be combined among the multiple three-dimensional points included in the two submeshes of the combination, the search process does not have to be performed.

[0027] According to this, when the search process is unnecessary, the search process is omitted, thereby reducing the amount of processing.

[0028] The decoding method of Example 8 may be the decoding method of Example 7, and may determine, based on the threshold information, whether or not there is a pair of two 3D points to be combined among the plurality of 3D points included in the two submeshes of the combination.

[0029] For two submeshes that are not to be joined, a specific value such as 0 is set as the threshold information, and thus it is possible to appropriately determine whether or not to join the two submeshes based on whether or not the threshold indicated by the threshold information is a specific value.

[0030] The encoding device of Example 9 includes a processor and a memory, and the processor uses the memory to generate multiple encoded data by encoding multiple submeshes, generates threshold information indicating a threshold value determined for each combination of two submeshes selected from the multiple submeshes based on the positions of multiple three-dimensional points included in the multiple submeshes, and generates a bitstream including the multiple encoded data and the threshold information.

[0031] This provides the same effects as the encoding method according to the first technique.

[0032] The decoding device of Example 10 includes a processor and a memory, and the processor uses the memory to acquire a bitstream including multiple encoded data generated by encoding multiple submeshes and threshold information indicating a threshold value determined for each combination of two submeshes selected from the multiple submeshes, decodes the multiple encoded data, and performs a joining process to join multiple three-dimensional points included in the multiple submeshes based on the threshold information.

[0033] This provides the same effect as the encoding method according to Technique 5.

[0034] Furthermore, these comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0035] <Expressions and Terms> The following expressions and terms are used herein.

[0036] (1) Three-dimensional Mesh A three-dimensional mesh is a collection of multiple faces, and represents, for example, a three-dimensional object. A three-dimensional mesh is mainly composed of vertex information, connectivity information, and attribute information. A three-dimensional mesh may be expressed as a polygon mesh or a mesh. A three-dimensional mesh may also vary over time. A three-dimensional mesh may include metadata related to the vertex information, connectivity information, and attribute information, and may also include other additional information.

[0037] (2) Vertex Information Vertex information is information indicating a vertex. For example, the vertex information indicates the position of a vertex in a three-dimensional space. Furthermore, a vertex corresponds to a vertex of a face that constitutes a three-dimensional mesh. Vertex information may be expressed as "geometry." Furthermore, vertex information may be expressed as position information.

[0038] (3) Connection Information Connection information is information that indicates connections between vertices. For example, connection information indicates connections for forming faces or edges of a three-dimensional mesh. Connection information may be expressed as "Connectivity." Connection information may also be expressed as face information.

[0039] (4) Attribute Information Attribute information is information that indicates attributes of a vertex or a face. For example, attribute information indicates attributes such as a color, an image, and a normal vector associated with a vertex or a face. Attribute information may be expressed as "texture."

[0040] (5) Faces A face is an element that constitutes a three-dimensional mesh. Specifically, a face is a polygon on a plane in three-dimensional space. For example, a face can be defined as a triangle in three-dimensional space.

[0041] (6) Plane A plane is a two-dimensional plane in a three-dimensional space. For example, a polygon is formed on a plane, and multiple polygons are formed on multiple planes.

[0042] (7) Bitstream: A bitstream corresponds to coded information. A bitstream may also be referred to as a stream, a coded bitstream, a compressed bitstream, or a coded signal.

[0043] (8) Encoding and Decoding The term encoding may be substituted with terms such as storing, including, writing, describing, signaling, sending, notifying, saving, or compressing, and these terms may be interchangeable. For example, encoding information may mean including the information in a bitstream. Also, encoding information into a bitstream may mean encoding the information to generate a bitstream that includes the encoded information.

[0044] Additionally, the term "decode" may be replaced with terms such as "read," "decode," "read," "load," "derive," "obtain," "receive," "extract," "reconstruct," "reconstruct," "decompress," or "decompress," and these terms may be interchangeable. For example, decoding information may mean obtaining information from a bitstream. Decoding information from a bitstream may mean decoding the bitstream to obtain information contained in the bitstream.

[0045] (9) Ordinal Numbers In the description, ordinal numbers such as first and second may be assigned to components, etc. These ordinal numbers may be changed as appropriate. Furthermore, new ordinal numbers may be assigned to components, etc., or removed. Furthermore, these ordinal numbers may be assigned to elements in order to identify them, and may not correspond to a meaningful order.

[0046] <Three-dimensional mesh> Fig. 1 is a conceptual diagram showing a three-dimensional mesh according to this embodiment. A three-dimensional mesh is composed of multiple faces. For example, each face is a triangle. The vertices of these triangles are defined in three-dimensional space. The three-dimensional mesh then represents a three-dimensional object. Each face may have a color or an image.

[0047] FIG. 2 is a conceptual diagram showing the basic elements of a three-dimensional mesh according to this embodiment. A three-dimensional mesh is composed of vertex information, connection information, and attribute information. The vertex information indicates the positions of the vertices of a face in three-dimensional space. The connection information indicates the connections between the vertices. A face can be identified by the vertex information and connection information. In other words, a colorless three-dimensional object is formed in three-dimensional space by the vertex information and connection information.

[0048] The attribute information may be associated with a vertex or a face. The attribute information associated with a vertex may be expressed as "Attribute Per Point." The attribute information associated with a vertex may indicate an attribute of the vertex itself, or may indicate an attribute of a face connected to the vertex.

[0049] For example, a color may be associated with a vertex as attribute information. The color associated with a vertex may be the color of the vertex itself, or the color of a face connected to the vertex. The color of a face may be the average of multiple colors associated with multiple vertices of the face. Furthermore, a normal vector may be associated with a vertex or a face as attribute information. Such a normal vector can represent the front and back of a face.

[0050] A two-dimensional image may be associated with a surface as attribute information. The two-dimensional image associated with a surface is also expressed as a texture image or an "Attribute Map." Information indicating mapping between the surface and the two-dimensional image may be associated with the surface as attribute information. Information indicating such mapping may be expressed as mapping information, vertex information of a texture image, texture coordinates, or "Attribute UV Coordinate."

[0051] Furthermore, information such as color, image, and moving image used as attribute information may be expressed as "parametric space."

[0052] The attribute information allows texture to be reflected on the three-dimensional object. That is, a three-dimensional object having color is formed in three-dimensional space based on the vertex information, connection information, and attribute information.

[0053] In the above, the attribute information is associated with the vertices or faces, but it may also be associated with the edges.

[0054] 3 is a conceptual diagram illustrating mapping according to this embodiment. For example, a region of a two-dimensional image on a two-dimensional plane can be mapped onto a surface of a three-dimensional mesh in three-dimensional space. Specifically, coordinate information of the region in the two-dimensional image is associated with the surface of the three-dimensional mesh. As a result, an image of the mapped region in the two-dimensional image is reflected on the surface of the three-dimensional mesh.

[0055] By using the mapping, the 2D image used as attribute information can be separated from the 3D mesh. For example, in encoding the 3D mesh, the 2D image may be encoded by an image encoding method or a video encoding method.

[0056] <System Configuration> Fig. 4 is a block diagram showing an example of the configuration of a coding / decoding system according to this embodiment. In Fig. 4, the coding / decoding system includes a coding device 100 and a decoding device 200.

[0057] For example, the encoding device 100 obtains a three-dimensional mesh and encodes the three-dimensional mesh into a bitstream. Then, the encoding device 100 outputs the bitstream to the network 300. For example, the bitstream includes the encoded three-dimensional mesh and control information for decoding the encoded three-dimensional mesh. By encoding the three-dimensional mesh, information about the three-dimensional mesh is compressed.

[0058] The network 300 transmits a bitstream from the encoding device 100 to the decoding device 200. The network 300 may be the Internet, a wide area network (WAN), a local area network (LAN), or a combination of these. The network 300 is not necessarily limited to bidirectional communication, and may be a unidirectional communication network for terrestrial digital broadcasting, satellite broadcasting, or the like.

[0059] Furthermore, the network 300 can be replaced by a recording medium such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc (registered trademark)).

[0060] The decoding device 200 obtains a bitstream and decodes a three-dimensional mesh from the bitstream. By decoding the three-dimensional mesh, information about the three-dimensional mesh is expanded. For example, the decoding device 200 decodes the three-dimensional mesh according to a decoding method corresponding to the encoding method used by the encoding device 100 to encode the three-dimensional mesh. That is, the encoding device 100 and the decoding device 200 perform encoding and decoding according to encoding methods and decoding methods that correspond to each other.

[0061] The 3D mesh before encoding may also be referred to as an original 3D mesh, and the 3D mesh after decoding may also be referred to as a reconstructed 3D mesh.

[0062] 5 is a block diagram showing an example of the configuration of a coding device 100 according to this embodiment. For example, the coding device 100 includes a vertex information encoder 101, a connection information encoder 102, and an attribute information encoder 103.

[0063] The vertex information encoder 101 is an electrical circuit that encodes vertex information. For example, the vertex information encoder 101 encodes the vertex information into a bitstream according to a format defined for the vertex information.

[0064] The connection information encoder 102 is an electrical circuit that encodes the connection information, for example, the connection information encoder 102 encodes the connection information into a bitstream according to a format defined for the connection information.

[0065] The attribute information encoder 103 is an electric circuit that encodes the attribute information. For example, the attribute information encoder 103 encodes the attribute information into a bit stream in accordance with a format defined for the attribute information.

[0066] The vertex information, connectivity information, and attribute information may be coded using variable-length coding or fixed-length coding, such as Huffman coding or context-adaptive binary arithmetic coding (CABAC).

[0067] The vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be integrated together, or each of the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be further subdivided into multiple components.

[0068] 6 is a block diagram showing another example of the configuration of the encoding device 100 according to this embodiment. For example, the encoding device 100 includes a pre-processor 104 and a post-processor 105 in addition to the configuration shown in FIG.

[0069] The preprocessor 104 is an electrical circuit that performs processing before encoding the vertex information, connectivity information, and attribute information. For example, the preprocessor 104 may perform a conversion process, a separation process, a multiplexing process, or the like on the 3D mesh before encoding. More specifically, for example, the preprocessor 104 may separate the vertex information, connectivity information, and attribute information from the 3D mesh before encoding.

[0070] The post-processor 105 is an electrical circuit that performs processing after the vertex information, connection information, and attribute information are encoded. For example, the post-processor 105 may perform conversion processing, separation processing, multiplexing processing, or the like on the encoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 105 may multiplex the encoded vertex information, connection information, and attribute information into a bitstream. Furthermore, for example, the post-processor 105 may further perform variable-length coding on the encoded vertex information, connection information, and attribute information.

[0071] 7 is a block diagram showing an example of the configuration of a decoding device 200 according to this embodiment. For example, the decoding device 200 includes a vertex information decoder 201, a connection information decoder 202, and an attribute information decoder 203.

[0072] The vertex information decoder 201 is an electrical circuit that decodes vertex information. For example, the vertex information decoder 201 decodes vertex information from a bitstream according to a format defined for the vertex information.

[0073] The connection information decoder 202 is an electrical circuit that decodes the connection information, for example, the connection information decoder 202 decodes the connection information from the bitstream according to a format defined for the connection information.

[0074] The attribute information decoder 203 is an electric circuit that decodes the attribute information. For example, the attribute information decoder 203 decodes the attribute information from the bitstream in accordance with a format defined for the attribute information.

[0075] The vertex information, connection information, and attribute information may be decoded using variable length decoding or fixed length decoding, which may correspond to Huffman coding, context-adaptive binary arithmetic coding (CABAC), or the like.

[0076] The vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be integrated together, or each of the vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be further subdivided into multiple components.

[0077] 8 is a block diagram showing another example of the configuration of the decoding device 200 according to this embodiment. For example, the decoding device 200 includes a pre-processor 204 and a post-processor 205 in addition to the configuration shown in FIG.

[0078] The preprocessor 204 is an electrical circuit that performs processing before decoding the vertex information, connection information, and attribute information. For example, the preprocessor 204 may perform conversion processing, separation processing, multiplexing processing, or the like on the bitstream before decoding the vertex information, connection information, and attribute information.

[0079] More specifically, for example, the preprocessor 204 may separate a sub-bitstream corresponding to vertex information, a sub-bitstream corresponding to connectivity information, and a sub-bitstream corresponding to attribute information from the bitstream. Also, for example, the preprocessor 204 may perform variable-length decoding on the bitstream in advance before decoding the vertex information, connectivity information, and attribute information.

[0080] The post-processor 205 is an electrical circuit that performs processing after the vertex information, connection information, and attribute information are decoded. For example, the post-processor 205 may perform conversion processing, separation processing, multiplexing processing, or the like on the decoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 205 may multiplex the decoded vertex information, connection information, and attribute information onto a three-dimensional mesh.

[0081] <Bitstream> Vertex information, connection information, and attribute information are coded and stored in a bitstream. The relationship between this information and the bitstream is shown below.

[0082] 9 is a conceptual diagram showing an example of the configuration of a bitstream according to this embodiment. In this example, connection information, vertex information, and attribute information are integrated in the bitstream. For example, the connection information, vertex information, and attribute information may be included in a single file.

[0083] Furthermore, multiple portions of this information may be stored sequentially, such as a first portion of connection information, a first portion of vertex information, a first portion of attribute information, a second portion of connection information, a second portion of vertex information, a second portion of attribute information, etc. These multiple portions may correspond to multiple portions that are different in time, multiple portions that are different in space, or multiple different faces.

[0084] Furthermore, the storage order of the connection information, vertex information, and attribute information is not limited to the above example, and a storage order different from the above example may be used.

[0085] 10 is a conceptual diagram showing another example of the configuration of a bitstream according to this embodiment. In this example, a plurality of files are included in the bitstream, and connection information, vertex information, and attribute information are stored in different files. Here, a file containing connection information, a file containing vertex information, and a file containing attribute information are shown, but the storage format is not limited to this example. For example, two types of information among the connection information, vertex information, and attribute information may be included in one file, and the remaining type of information may be included in another file.

[0086] Alternatively, the information may be split and stored in more files. For example, multiple pieces of connectivity information may be stored in multiple files, multiple pieces of vertex information may be stored in multiple files, or multiple pieces of attribute information may be stored in multiple files. These multiple pieces may correspond to multiple temporally different pieces, multiple spatially different pieces, or multiple different faces.

[0087] Furthermore, the storage order of the connection information, vertex information, and attribute information is not limited to the above example, and a storage order different from the above example may be used.

[0088] 11 is a conceptual diagram showing another example of the configuration of a bitstream according to this embodiment. In this example, the bitstream is composed of multiple separable sub-bitstreams, and connection information, vertex information, and attribute information are stored in different sub-bitstreams.

[0089] Here, a sub-bitstream containing connection information, a sub-bitstream containing vertex information, and a sub-bitstream containing attribute information are shown, but the storage format is not limited to this example.

[0090] For example, two types of information among the connection information, vertex information, and attribute information may be included in one sub-bitstream, and the remaining type of information may be included in another sub-bitstream. Specifically, attribute information of a two-dimensional image or the like may be stored in a sub-bitstream that complies with an image coding method, separate from the sub-bitstreams of the connection information and vertex information.

[0091] Also, each sub-bitstream may include multiple files, and multiple pieces of connectivity information may be stored in multiple files, multiple pieces of vertex information may be stored in multiple files, or multiple pieces of attribute information may be stored in multiple files.

[0092] 9, 10, and 11, and a storage order different from the above examples may be used. For example, the vertex information, connection information, and attribute information may be stored in the bitstream in this order. Alternatively, the connection information, connection information, and attribute information may be stored in the bitstream in any of the following orders: connection information, attribute information, and vertex information; vertex information, attribute information, and connection information; attribute information, connection information, and vertex information; or attribute information, vertex information, and connection information.

[0093] Furthermore, each of the connection information, vertex information, and attribute information may be divided into a plurality of data, and the plurality of data may be stored in a cyclical or random order within the bitstream.

[0094] 12 is a block diagram showing a specific example of an encoding / decoding system according to this embodiment. In FIG. 12, the encoding / decoding system includes a three-dimensional data encoding system 110, a three-dimensional data decoding system 210, and an external connector 310.

[0095] The three-dimensional data encoding system 110 includes a controller 111, an input / output processor 112, a three-dimensional data encoder 113, a three-dimensional data generator 115, and a system multiplexer 114. The three-dimensional data decoding system 210 includes a controller 211, an input / output processor 212, a three-dimensional data decoder 213, a system demultiplexer 214, a presenter 215, and a user interface 216.

[0096] In the three-dimensional data encoding system 110, sensor data is input from a sensor terminal to a three-dimensional data generator 115. The three-dimensional data generator 115 generates three-dimensional data, such as point cloud data or mesh data, from the sensor data and inputs it to a three-dimensional data encoder 113.

[0097] For example, the three-dimensional data generator 115 generates vertex information, and generates connection information and attribute information corresponding to the vertex information. The three-dimensional data generator 115 may process the vertex information when generating the connection information and attribute information. For example, the three-dimensional data generator 115 may reduce the amount of data by deleting duplicate vertices, or may transform the vertex information (such as by shifting its position, rotating it, or normalizing it). The three-dimensional data generator 115 may also render the attribute information.

[0098] Furthermore, although the three-dimensional data generator 115 is a component of the three-dimensional data encoding system 110 in FIG. 12, it may be arranged externally and independently of the three-dimensional data encoding system 110.

[0099] The sensor terminal that provides the sensor data for generating the three-dimensional data may be, for example, a moving body such as an automobile, a flying object such as an airplane, a mobile terminal, a camera, etc. Furthermore, a distance sensor such as a LIDAR, a millimeter wave radar, an infrared sensor, or a range finder, a stereo camera, or a combination of multiple monocular cameras may also be used as the sensor terminal.

[0100] The sensor data may be the distance (position) of the object, monocular camera images, stereo camera images, color, reflectance, sensor attitude, orientation, gyro, sensing position (GPS information or altitude), speed, acceleration, sensing time, temperature, air pressure, humidity, or magnetism.

[0101] The three-dimensional data encoder 113 corresponds to the encoding device 100 shown in FIG. 5 and other figures. For example, the three-dimensional data encoder 113 encodes three-dimensional data to generate encoded data. The three-dimensional data encoder 113 also generates control information when encoding the three-dimensional data. The three-dimensional data encoder 113 then inputs the encoded data together with the control information to the system multiplexer 114.

[0102] The encoding method for the three-dimensional data may be an encoding method using geometry or an encoding method using a video codec. Here, the encoding method using geometry may also be referred to as a geometry-based encoding method. The encoding method using a video codec may also be referred to as a video-based encoding method.

[0103] The system multiplexer 114 multiplexes the encoded data and control information input from the 3D data encoder 113 to generate multiplexed data using a specified multiplexing method. The system multiplexer 114 may multiplex other media such as video, audio, subtitles, application data, or document files, or reference time information, along with the encoded data and control information of the 3D data. Furthermore, the system multiplexer 114 may multiplex attribute information related to the sensor data or the 3D data.

[0104] For example, the multiplexed data may have a file format for storage or a packet format for transmission. As these formats, ISOBMFF or a format based on ISOBMFF may be used. Also, MPEG-DASH, MMT, MPEG-2 TS Systems, RTP, or the like may be used.

[0105] The multiplexed data is then output as a transmission signal to the external connector 310 by the input / output processor 112. The multiplexed data may be transmitted as a transmission signal by wire or wirelessly. Alternatively, the multiplexed data is stored in an internal memory or a storage device. The multiplexed data may be transmitted to a cloud server via the Internet or may be stored in an external storage device.

[0106] For example, the transmission or storage of the multiplexed data is performed by a method according to the medium for transmission or storage, such as broadcasting or communication. The communication protocol may be http, ftp, TCP, UDP, IP, or a combination thereof. Furthermore, a pull-type communication method or a push-type communication method may be used.

[0107] For wired transmission, Ethernet (registered trademark), USB, RS-232C, HDMI (registered trademark), coaxial cable, etc. may be used. For wireless transmission, 3GPP (registered trademark), 3G / 4G / 5G defined by IEEE, wireless LAN, Wi-Fi, Bluetooth, or millimeter wave may be used. For broadcasting, for example, DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 may be used.

[0108] The sensor data may be input to the three-dimensional data generator 115 or the system multiplexer 114. The three-dimensional data or encoded data may be output as a transmission signal directly to the external connector 310 via the input / output processor 112. The transmission signal output from the three-dimensional data encoding system 110 is input to the three-dimensional data decoding system 210 via the external connector 310.

[0109] Furthermore, each operation of the three-dimensional data encoding system 110 may be controlled by a controller 111 that executes an application program.

[0110] In the three-dimensional data decoding system 210, a transmission signal is input to an input / output processor 212. The input / output processor 212 decodes multiplexed data having a file format or a packet format from the transmission signal and inputs the multiplexed data to a system demultiplexer 214. The system demultiplexer 214 obtains coded data and control information from the multiplexed data and inputs them to a three-dimensional data decoder 213. The system demultiplexer 214 may extract other media or reference time information from the multiplexed data.

[0111] The three-dimensional data decoder 213 corresponds to the decoding device 200 shown in Fig. 7 etc. For example, the three-dimensional data decoder 213 decodes three-dimensional data from the encoded data based on a predefined encoding method. The three-dimensional data is then presented to the user by the presenter 215.

[0112] Additionally, additional information such as sensor data may be input to the presenter 215. The presenter 215 may present three-dimensional data based on the additional information. Additionally, a user instruction may be input from a user terminal to the user interface 216. Then, the presenter 215 may present three-dimensional data based on the input instruction.

[0113] The input / output processor 212 may acquire the three-dimensional data and the encoded data from the external connector 310 .

[0114] Furthermore, each operation of the three-dimensional data decoding system 210 may be controlled by a controller 211 that executes an application program.

[0115] 13 is a conceptual diagram showing an example of the configuration of point cloud data according to this embodiment. The point cloud data is data of a group of points representing a three-dimensional object.

[0116] Specifically, a point cloud is made up of a plurality of points, and has position information indicating the three-dimensional coordinate position of each point and attribute information indicating the attribute of each point. The position information is also expressed as geometry.

[0117] The type of attribute information may be, for example, color, reflectance, etc. One point may be associated with attribute information of one type, one point may be associated with attribute information of multiple different types, or one point may be associated with attribute information having multiple values ​​for the same type.

[0118] 14 is a conceptual diagram showing an example of a data file of point cloud data according to this embodiment. This example shows a case where there is a one-to-one correspondence between position information items and attribute information items, and shows position information and attribute information for N points that make up the point cloud data. In this example, the position information is information indicating a three-dimensional coordinate position using three axes, x, y, and z, and the attribute information is information indicating a color using RGB. A PLY file or the like can be used as a representative data file for point cloud data.

[0119] 15 is a conceptual diagram showing an example of the configuration of mesh data according to this embodiment. Mesh data is data used in CG (Computer Graphics) and the like, and is three-dimensional mesh data that shows the three-dimensional shape of an object using multiple surfaces. Each surface is also expressed as a polygon, and has a polygonal shape such as a triangle or a rectangle.

[0120] Specifically, a 3D mesh is composed of a plurality of points constituting a point cloud, as well as a plurality of edges and a plurality of faces. Each point is also expressed as a vertex or a position. Each edge corresponds to a line segment connected by two vertices. Each face corresponds to an area surrounded by three or more edges.

[0121] Furthermore, a three-dimensional mesh has position information indicating the three-dimensional coordinate positions of vertices. The position information is also expressed as vertex information or geometry. A three-dimensional mesh also has connection information indicating the relationship between multiple vertices that make up an edge or a face. The connection information is also expressed as connectivity. A three-dimensional mesh also has attribute information indicating the attributes of the vertices, edges, or faces. The attribute information in a three-dimensional mesh is also expressed as texture.

[0122] For example, the attribute information may indicate the color, reflectance, or normal vector for a vertex, edge, or face. The direction of the normal vector may represent the front and back of the face.

[0123] The mesh data may be stored in a data file format such as an object file.

[0124] 16 is a conceptual diagram showing an example of a data file of mesh data according to this embodiment. In this example, the data file includes position information G(1) to G(N) of N vertices that make up the three-dimensional mesh, and attribute information A1(1) to A1(N) of the N vertices. Also, in this example, M pieces of attribute information A2(1) to A2(M) are included. The attribute information items do not need to correspond one-to-one to vertices or faces. Furthermore, attribute information need not exist.

[0125] The connection information is represented by a combination of vertex indices. n[1, 3, 4] indicates a triangular face formed by three vertices, n=1, n=3, and n=4. Also, m[2, 4, 6] indicates that the attribute information of m=2, m=4, and m=6 corresponds to the three vertices, respectively.

[0126] Furthermore, the actual contents of the attribute information may be written in a separate file. A pointer to that content may be associated with a vertex, a face, or the like. For example, attribute information indicating an image for a face may be stored in a two-dimensional attribute map file. The file name of the attribute map and two-dimensional coordinate values ​​in the attribute map may be written in attribute information A2(1) to A2(M). The method of specifying attribute information for a face is not limited to these methods, and any method may be used.

[0127] 17 is a conceptual diagram showing types of three-dimensional data according to this embodiment. Point cloud data and mesh data may represent static objects or dynamic objects. A static object is an object that does not change over time, and a dynamic object is an object that changes over time. A static object may correspond to three-dimensional data for any point in time.

[0128] For example, point cloud data for a given point in time may be referred to as a PCC frame, mesh data for a given point in time may be referred to as a mesh frame, and PCC frames and mesh frames may be simply referred to as frames.

[0129] The area of ​​the object may be limited to a certain range, as in normal video data, or may not be limited, as in map data. The density of points or surfaces may be determined in various ways. Sparse point cloud data or sparse mesh data may be used, or dense point cloud data or dense mesh data may be used.

[0130] Next, encoding and decoding of a point cloud or a three-dimensional mesh will be described. The device, process, or syntax for encoding and decoding vertex information of a three-dimensional mesh in the present disclosure may be applied to encoding and decoding of a point cloud. The device, process, or syntax for encoding and decoding of a point cloud in the present disclosure may be applied to encoding and decoding vertex information of a three-dimensional mesh.

[0131] Furthermore, a device, process, or syntax for encoding and decoding attribute information of a point cloud in the present disclosure may be applied to encoding and decoding connectivity information or attribute information of a three-dimensional mesh.Furthermore, a device, process, or syntax for encoding and decoding connectivity information or attribute information of a three-dimensional mesh in the present disclosure may be applied to encoding and decoding attribute information of a point cloud.

[0132] Furthermore, at least some of the processing may be shared between the encoding and decoding of point cloud data and the encoding and decoding of mesh data, thereby reducing the scale of the circuit and software program.

[0133] 18 is a block diagram showing an example configuration of a three-dimensional data encoder 113 according to this embodiment. In this example, the three-dimensional data encoder 113 includes a vertex information encoder 121, an attribute information encoder 122, a metadata encoder 123, and a multiplexer 124. The vertex information encoder 121, the attribute information encoder 122, and the multiplexer 124 may correspond to the vertex information encoder 101, the attribute information encoder 103, the post-processor 105, etc. in FIG.

[0134] In this example, the three-dimensional data encoder 113 encodes the three-dimensional data according to a geometry-based encoding method, which takes into account the three-dimensional structure. In addition, in the geometry-based encoding method, attribute information is encoded using configuration information obtained in encoding the vertex information.

[0135] Specifically, first, vertex information, attribute information, and metadata included in three-dimensional data generated from sensor data are input to a vertex information encoder 121, an attribute information encoder 122, and a metadata encoder 123, respectively. Here, connectivity information included in the three-dimensional data may be treated in the same way as attribute information. In addition, in the case of point cloud data, position information may be treated as vertex information.

[0136] The vertex information encoder 121 encodes the vertex information into compressed vertex information and outputs the compressed vertex information as encoded data to the multiplexer 124. The vertex information encoder 121 also generates metadata for the compressed vertex information and outputs it to the multiplexer 124. The vertex information encoder 121 also generates configuration information and outputs it to the attribute information encoder 122.

[0137] The attribute information encoder 122 uses the configuration information generated by the vertex information encoder 121 to encode the attribute information into compressed attribute information and outputs the compressed attribute information as encoded data to the multiplexer 124. The attribute information encoder 122 also generates metadata of the compressed attribute information and outputs it to the multiplexer 124.

[0138] The metadata encoder 123 encodes compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The metadata encoded by the metadata encoder 123 may be used to encode vertex information and attribute information.

[0139] The multiplexer 124 multiplexes the compressed vertex information, the compressed vertex information metadata, the compressed attribute information, the compressed attribute information metadata, and the compressed metadata into a bitstream, and then inputs the bitstream to the system layer.

[0140] 19 is a block diagram showing an example configuration of a three-dimensional data decoder 213 according to this embodiment. In this example, the three-dimensional data decoder 213 includes a vertex information decoder 221, an attribute information decoder 222, a metadata decoder 223, and a demultiplexer 224. The vertex information decoder 221, the attribute information decoder 222, and the demultiplexer 224 may correspond to the vertex information decoder 201, the attribute information decoder 203, the preprocessor 204, and the like in FIG.

[0141] In this example, the three-dimensional data decoder 213 decodes three-dimensional data according to a geometry-based encoding method. The three-dimensional structure is taken into consideration in the decoding according to the geometry-based encoding method. Furthermore, in the decoding according to the geometry-based encoding method, attribute information is decoded using configuration information obtained in decoding vertex information.

[0142] Specifically, first, a bitstream is input from the system layer to a demultiplexer 224. The demultiplexer 224 separates compressed vertex information, compressed vertex information metadata, compressed attribute information, compressed attribute information metadata, and compressed metadata from the bitstream. The compressed vertex information and compressed vertex information metadata are input to a vertex information decoder 221. The compressed attribute information and compressed attribute information metadata are input to an attribute information decoder 222. The metadata is input to a metadata decoder 223.

[0143] The vertex information decoder 221 decodes vertex information from the compressed vertex information using metadata of the compressed vertex information. The vertex information decoder 221 also generates configuration information and outputs it to the attribute information decoder 222. The attribute information decoder 222 decodes attribute information from the compressed attribute information using the configuration information generated by the vertex information decoder 221 and the metadata of the compressed attribute information. The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 may be used to decode the vertex information and the attribute information.

[0144] Thereafter, the vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. Note that, for example, this metadata is metadata of the vertex information and attribute information, and can be used in an application program.

[0145] 20 is a block diagram showing another example configuration of the three-dimensional data encoder 113 according to the present embodiment. In this example, the three-dimensional data encoder 113 includes a vertex image generator 131, an attribute image generator 132, a metadata generator 133, a video encoder 134, a metadata encoder 123, and a multiplexer 124. The vertex image generator 131, the attribute image generator 132, and the video encoder 134 may correspond to the vertex information encoder 101 and the attribute information encoder 103 in FIG. 6 , etc.

[0146] In this example, the 3D data encoder 113 encodes the 3D data according to a video-based encoding method. In encoding according to the video-based encoding method, multiple 2D images are generated from the 3D data, and the multiple 2D images are encoded according to a video encoding method. Here, the video encoding method may be High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or the like.

[0147] Specifically, first, vertex information and attribute information included in three-dimensional data generated from sensor data are input to a metadata generator 133. The vertex information and attribute information are then input to a vertex image generator 131 and an attribute image generator 132, respectively. The metadata included in the three-dimensional data is then input to a metadata encoder 123. Here, connectivity information included in the three-dimensional data may be treated in the same way as attribute information. In the case of point cloud data, position information may be treated as vertex information.

[0148] The metadata generator 133 generates map information of a plurality of two-dimensional images from the vertex information and attribute information, and inputs the map information to the vertex image generator 131, the attribute image generator 132, and the metadata encoder 123.

[0149] The vertex image generator 131 generates a vertex image based on the vertex information and map information, and inputs the generated image to the video encoder 134. The attribute image generator 132 generates an attribute image based on the attribute information and map information, and inputs the generated image to the video encoder 134.

[0150] The video encoder 134 encodes the vertex images and attribute images into compressed vertex information and compressed attribute information, respectively, in accordance with a video encoding method, and outputs the compressed vertex information and compressed attribute information as encoded data to the multiplexer 124. The video encoder 134 also generates metadata for the compressed vertex information and metadata for the compressed attribute information, and outputs them to the multiplexer 124.

[0151] The metadata encoder 123 encodes the compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The compressible metadata includes map information. The metadata encoded by the metadata encoder 123 may also be used to encode vertex information and attribute information.

[0152] The multiplexer 124 multiplexes the compressed vertex information, the compressed vertex information metadata, the compressed attribute information, the compressed attribute information metadata, and the compressed metadata into a bitstream, and then inputs the bitstream to the system layer.

[0153] 21 is a block diagram showing another example configuration of the 3D data decoder 213 according to this embodiment. In this example, the 3D data decoder 213 includes a vertex information generator 231, an attribute information generator 232, a video decoder 234, a metadata decoder 223, and a demultiplexer 224. The vertex information generator 231, the attribute information generator 232, and the video decoder 234 may correspond to the vertex information decoder 201 and the attribute information decoder 203 in FIG. 8, etc.

[0154] In this example, the 3D data decoder 213 decodes the 3D data according to a video-based coding method. In the decoding according to the video-based coding method, a plurality of 2D images are decoded according to a video coding method, and 3D data is generated from the plurality of 2D images. Here, the video coding method may be High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or the like.

[0155] Specifically, first, a bitstream is input from the system layer to the demultiplexer 224. The demultiplexer 224 separates compressed vertex information, compressed vertex information metadata, compressed attribute information, compressed attribute information metadata, and compressed metadata from the bitstream. The compressed vertex information, compressed vertex information metadata, compressed attribute information, and compressed attribute information metadata are input to the video decoder 234. The compressed metadata is input to the metadata decoder 223.

[0156] The video decoder 234 decodes the vertex images in accordance with the video encoding method. At this time, the video decoder 234 decodes the vertex images from the compressed vertex information using the metadata of the compressed vertex information. Then, the video decoder 234 inputs the vertex images to the vertex information generator 231. The video decoder 234 also decodes the attribute images in accordance with the video encoding method. At this time, the video decoder 234 decodes the attribute images from the compressed attribute information using the metadata of the compressed attribute information. Then, the video decoder 234 inputs the attribute images to the attribute information generator 232.

[0157] The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 includes map information used to generate vertex information and attribute information. The metadata decoded by the metadata decoder 223 may also be used to decode vertex images and attribute images.

[0158] The vertex information generator 231 reproduces vertex information from the vertex image in accordance with the map information included in the metadata decoded by the metadata decoder 223. The attribute information generator 232 reproduces attribute information from the attribute image in accordance with the map information included in the metadata decoded by the metadata decoder 223.

[0159] Thereafter, the vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. Note that, for example, this metadata is metadata of the vertex information and attribute information, and can be used in an application program.

[0160] Fig. 22 is a conceptual diagram showing a specific example of encoding processing according to this embodiment. Fig. 22 shows a three-dimensional data encoder 113 and a description encoder 148. In this example, the three-dimensional data encoder 113 includes a two-dimensional data encoder 141 and a mesh data encoder 142. The two-dimensional data encoder 141 includes a texture encoder 143. The mesh data encoder 142 includes a vertex information encoder 144 and a connection information encoder 145.

[0161] The vertex information encoder 144, the connection information encoder 145, and the texture encoder 143 may correspond to the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 in FIG.

[0162] For example, the two-dimensional data encoder 141 operates as a texture encoder 143 and generates a texture file by encoding the texture corresponding to the attribute information as two-dimensional data according to an image encoding method or a video encoding method.

[0163] The mesh data encoder 142 also operates as a vertex information encoder 144 and a connectivity information encoder 145, and generates a mesh file by encoding the vertex information and connectivity information. The mesh data encoder 142 may further encode mapping information for textures. The encoded mapping information may then be included in the mesh file.

[0164] The description encoder 148 also generates a description file by encoding a description corresponding to metadata such as text data. The description encoder 148 may encode the description at the system layer. For example, the description encoder 148 may be included in the system multiplexer 114 of FIG. 12 .

[0165] The above operations generate a bitstream containing texture files, mesh files, and description files, which may be multiplexed into the bitstream in file formats such as glTF (Graphics Language Transmission Format) or USD (Universal Scene Description).

[0166] The three-dimensional data encoder 113 may include two mesh data encoders as the mesh data encoder 142. For example, one mesh data encoder encodes vertex information and connectivity information of a static three-dimensional mesh, and the other mesh data encoder encodes vertex information and connectivity information of a dynamic three-dimensional mesh.

[0167] Correspondingly, two mesh files may then be included in the bitstream, for example one mesh file corresponding to a static 3D mesh and another mesh file corresponding to a dynamic 3D mesh.

[0168] Furthermore, the static three-dimensional mesh may be a three-dimensional mesh of an intraframe coded using intraprediction, and the dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe coded using interprediction. Furthermore, information on the dynamic three-dimensional mesh may be differential information between vertex information or connectivity information of the three-dimensional mesh of an intraframe and vertex information or connectivity information of the three-dimensional mesh of an interframe.

[0169] Fig. 23 is a conceptual diagram showing a specific example of the decoding process according to this embodiment. Fig. 23 shows a three-dimensional data decoder 213, a description decoder 248, and a renderer 247. In this example, the three-dimensional data decoder 213 includes a two-dimensional data decoder 241, a mesh data decoder 242, and a mesh reconstructor 246. The two-dimensional data decoder 241 includes a texture decoder 243. The mesh data decoder 242 includes a vertex information decoder 244 and a connectivity information decoder 245.

[0170] The vertex information decoder 244, the connection information decoder 245, the texture decoder 243, and the mesh reconstructor 246 may correspond to the vertex information decoder 201, the connection information decoder 202, the attribute information decoder 203, and the post-processor 205 in Fig. 8. The presenter 247 may correspond to the presenter 215 in Fig. 12.

[0171] For example, the two-dimensional data decoder 241 operates as a texture decoder 243, and decodes the texture corresponding to the attribute information from the texture file as two-dimensional data in accordance with an image coding method or a video coding method.

[0172] The mesh data decoder 242 also operates as a vertex information decoder 244 and a connectivity information decoder 245 to decode vertex information and connectivity information from the mesh file. The mesh data decoder 242 may further decode mapping information for textures from the mesh file.

[0173] The description decoder 248 also decodes descriptions corresponding to metadata such as text data from the description file. The description decoder 248 may decode the descriptions at the system layer. For example, the description decoder 248 may be included in the system demultiplexer 214 of FIG. 12 .

[0174] The mesh reconstructor 246 reconstructs a 3D mesh from the vertex information, connectivity information, and textures according to the description. The renderer 247 renders and outputs the 3D mesh according to the description.

[0175] Through the above operations, a 3D mesh is reconstructed and output from a bitstream containing a texture file, a mesh file, and a description file.

[0176] The three-dimensional data decoder 213 may include two mesh data decoders as the mesh data decoder 242. For example, one mesh data decoder decodes vertex information and connectivity information of a static three-dimensional mesh, and the other mesh data decoder decodes vertex information and connectivity information of a dynamic three-dimensional mesh.

[0177] Correspondingly, two mesh files may then be included in the bitstream, for example one mesh file corresponding to a static 3D mesh and another mesh file corresponding to a dynamic 3D mesh.

[0178] Furthermore, the static three-dimensional mesh may be a three-dimensional mesh of an intraframe coded using intraprediction, and the dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe coded using interprediction. Furthermore, information on the dynamic three-dimensional mesh may be differential information between vertex information or connectivity information of the three-dimensional mesh of an intraframe and vertex information or connectivity information of the three-dimensional mesh of an interframe.

[0179] A dynamic 3D mesh coding method is sometimes called DMC (Dynamic Mesh Coding), and a video-based dynamic 3D mesh coding method is sometimes called V-DMC (Video-based Dynamic Mesh Coding).

[0180] The point cloud encoding method is sometimes called PCC (Point Cloud Compression). The point cloud video-based encoding method is sometimes called V-PCC (Video-based Point Cloud Compression). The point cloud geometry-based encoding method is sometimes called G-PCC (Geometry-based Point Cloud Compression).

[0181] <Implementation Example> Fig. 24 is a block diagram showing an implementation example of the encoding device 100 according to this embodiment. The encoding device 100 includes a circuit 151 and a memory 152. For example, multiple components of the encoding device 100 shown in Fig. 5 etc. are implemented by the circuit 151 and memory 152 shown in Fig. 24.

[0182] The circuit 151 is a circuit that performs information processing and is a circuit that can access the memory 152. For example, the circuit 151 is a dedicated or general-purpose electric circuit that encodes a three-dimensional mesh. The circuit 151 may be a processor such as a CPU. Alternatively, the circuit 151 may be a collection of multiple electric circuits.

[0183] The memory 152 is a dedicated or general-purpose memory that stores information used by the circuit 151 to encode the three-dimensional mesh. The memory 152 may be an electric circuit and may be connected to the circuit 151. The memory 152 may also be included in the circuit 151. The memory 152 may also be a collection of multiple electric circuits. The memory 152 may also be a magnetic disk, an optical disk, or the like, and may also be expressed as a storage, a recording medium, or the like. The memory 152 may also be a non-volatile memory or a volatile memory.

[0184] For example, the memory 152 may store a three-dimensional mesh or a bitstream, or may store a program for the circuit 151 to encode the three-dimensional mesh.

[0185] Note that the encoding device 100 does not necessarily have to implement all of the components shown in Figure 5 and the like, and does not necessarily have to perform all of the processes shown here. Some of the components shown in Figure 5 and the like may be included in another device, and some of the processes shown here may be executed by another device. Furthermore, the encoding device 100 may implement any combination of the components of the present disclosure, and may perform any combination of the processes of the present disclosure.

[0186] Fig. 25 is a block diagram showing an example implementation of a decoding device 200 according to this embodiment. The decoding device 200 includes a circuit 251 and a memory 252. For example, multiple components of the decoding device 200 shown in Fig. 7 and other figures are implemented by the circuit 251 and memory 252 shown in Fig. 25.

[0187] The circuit 251 is a circuit that performs information processing and is a circuit that can access the memory 252. For example, the circuit 251 is a dedicated or general-purpose electric circuit that decodes a three-dimensional mesh. The circuit 251 may be a processor such as a CPU. Alternatively, the circuit 251 may be a collection of multiple electric circuits.

[0188] The memory 252 is a dedicated or general-purpose memory that stores information for the circuit 251 to decode the 3D mesh. The memory 252 may be an electric circuit and may be connected to the circuit 251. The memory 252 may also be included in the circuit 251. The memory 252 may also be a collection of multiple electric circuits. The memory 252 may also be a magnetic disk, an optical disk, or the like, and may also be expressed as a storage, a recording medium, or the like. The memory 252 may also be a non-volatile memory or a volatile memory.

[0189] For example, the memory 252 may store a three-dimensional mesh or a bitstream, or may store a program for the circuit 251 to decode the three-dimensional mesh.

[0190] Note that the decoding device 200 does not necessarily have to implement all of the components shown in Figure 7 and the like, and does not necessarily have to perform all of the processes shown here. Some of the components shown in Figure 7 and the like may be included in another device, and some of the processes shown here may be executed by another device. Furthermore, the decoding device 200 may implement any combination of the components of the present disclosure, and may perform any combination of the processes of the present disclosure.

[0191] The encoding method and the decoding method including the steps performed by each component of the encoding device 100 and the decoding device 200 of the present disclosure may be executed by any device or system. For example, part or all of the encoding method and the decoding method may be executed by a computer including a processor, a memory, an input / output circuit, etc. In this case, the encoding method and the decoding method may be executed by the computer executing a program for causing the computer to execute the encoding method and the decoding method.

[0192] Alternatively, the program or the bitstream may be recorded on a non-transitory computer-readable recording medium such as a CD-ROM.

[0193] An example of a program may be a bitstream. For example, a bitstream including an encoded three-dimensional mesh includes syntax elements for causing the decoding device 200 to decode the three-dimensional mesh. The bitstream then causes the decoding device 200 to decode the three-dimensional mesh according to the syntax elements included in the bitstream. Thus, the bitstream may play a role similar to that of a program.

[0194] The bitstream may be an encoded bitstream containing the encoded 3D mesh, or may be a multiplexed bitstream containing the encoded 3D mesh and other information.

[0195] Furthermore, each component of the encoding device 100 and the decoding device 200 may be configured with dedicated hardware, general-purpose hardware that executes the above-mentioned programs, or a combination of these. The general-purpose hardware may be configured with a memory in which the programs are recorded and a general-purpose processor that reads and executes the programs from the memory. Here, the memory may be a semiconductor memory or a hard disk, and the general-purpose processor may be a CPU.

[0196] Furthermore, the dedicated hardware may be configured with a memory, a dedicated processor, etc. For example, the dedicated processor may execute the encoding method and the decoding method by referring to a memory for recording data.

[0197] Furthermore, as described above, each component of the encoding device 100 and the decoding device 200 may be an electric circuit. These electric circuits may form a single electric circuit as a whole, or each may be a separate electric circuit. Furthermore, these electric circuits may correspond to dedicated hardware, or may correspond to general-purpose hardware that executes the above-mentioned programs, etc. Furthermore, the encoding device 100 and the decoding device 200 may be implemented as an integrated circuit.

[0198] Furthermore, the encoding device 100 may be a transmitting device that transmits the three-dimensional mesh, and the decoding device 200 may be a receiving device that receives the three-dimensional mesh.

[0199] Displacement Encoding and Decoding The following terminology is used here by way of example:

[0200] (1) Image An image is a data unit made up of a set of pixels, and includes a picture or a block smaller than a picture. Images include both moving images and still images.

[0201] (2) Picture A picture is a unit of image processing that is made up of a set of pixels, and is also called a frame or field.

[0202] (3) Block A block is a processing unit consisting of a specific number of pixels. The term shown in the following example is also used for a block. The shape of a block is not particularly limited. A block may be, for example, a rectangular shape of M×N pixels or a square shape of M×M pixels. A block may also be a triangular shape, a circular shape, or another shape. Examples of blocks are as follows:

[0203] Slice, tile, or brick CTU, superblock, or basic division unit VPDU, processing division unit for hardware CU, processing block unit, prediction block unit (PU), or orthogonal transform block unit (TU) Sub-block

[0204] (4) Pixel or Sample A pixel or sample is the smallest point of an image, in other words, the smallest unit. Pixels or samples include not only pixels at integer positions, but also pixels at sub-pixel positions generated based on pixels at integer positions.

[0205] (5) Pixel Value or Sample Value: A pixel value or sample value is a unique value of a pixel. The pixel value or sample value may include a luma value, a chroma value, or an RGB gradation level, and may also include a depth value or a binary value of 0 or 1.

[0206] (6) Flags A flag indicates one or more bits. A flag is, for example, a parameter or index represented by two or more bits. A flag may indicate not only a value represented by a binary number, but also a value represented by a number other than a binary number.

[0207] (7) Signal: A signal is something that is symbolized or coded to transmit information. A signal includes a discrete digital signal or a continuous analog signal.

[0208] (8) Stream or Bit Stream A stream or bit stream is a digital data sequence that indicates the flow of digital data. A stream or bit stream may be a single stream, or may be configured to include multiple streams with multiple layers. A stream or bit stream may be transmitted by serial communication using a single transmission path, or may be transmitted by packet communication using multiple transmission paths.

[0209] (9) Difference: For scalar quantities, the difference can include simple difference (x-y) and difference calculations. The difference can include absolute difference (|x-y|), squared difference (x^2-y^2), square root difference (√(x-y)), weighted difference (ax-by, where a and b are constants), or offset difference (x-y+a, where a is an offset).

[0210] (10) Sum: For scalar quantities, sums can include simple sum (x + y) and addition calculations. The sum can also include absolute sum (|x + y|), sum of squares (x^2 + y^2), square root of the sum (√(x + y)), weighted sum (ax + by, where a and b are constants), or offset sum (x + y + a, where a is an offset).

[0211] (11) "Based on" The expression "based on something" means that something other than that "something" may be taken into consideration. Also, "based on" can be used both when a direct result is obtained and when a result is obtained through an intermediate result.

[0212] (12) "Used" or "Using" The phrases "something was used" or "used something" mean that something other than the "something" may be taken into consideration. The phrases "used" or "used" may be used both in cases where a direct result is obtained and in cases where a result is obtained via an intermediate result.

[0213] (13) Prohibition "Prohibit" can be rephrased as "not permitted." Also, "not prohibited / prohibited" or "permitted / permitted" does not necessarily mean "obligation."

[0214] (14) "Restriction" or "Limitation" "Restriction" or "Limitation" can be rephrased as "not permitted / not allowed" or "not permitted / permitted." Furthermore, "prohibited / not prohibited" or "not permitted / permitted" does not necessarily mean "obligation." Furthermore, what is prohibited quantitatively or qualitatively may be either partial or total.

[0215] (15) Chroma The term chroma is an adjective, represented by the symbols Cb or Cr, that indicates that a sample array or a single sample represents one of the two color difference signals associated with a primary color. The term chroma is sometimes used instead of the term chrominance.

[0216] (16) Luma The term luma is an adjective, denoted by the symbols or subscripts Y or L, that indicates that a sample array or a single sample represents a monochrome signal for a primary color. The term luma is sometimes used instead of the term luminance.

[0217] The encoding / decoding system of this embodiment will be described below.

[0218] A typical three-dimensional model (also called a 3D model) digitally represents an object so that a user can explore the model using zoom, pan, and rotation in all three dimensions while it is rendered over time. One way to construct such a representation is to build a 3D mesh using triangles. The model stores the positions of the triangle vertices, their connectivity to each other, and their associated attributes (such as normals or UV patches).

[0219] Storing all this information in uncompressed form requires a very large storage space and therefore a very large bandwidth for transmission. The triangles that form the mesh often have repeating patterns and similar properties, especially in temporal and spatial neighborhoods. These repetitions can be exploited to develop efficient encoding and decoding methods for storage and transmission. One such encoding and decoding method is Video-based Dynamic Mesh Coding (V-DMC).

[0220] 26 is a block diagram showing another example of the configuration of the encoding / decoding system according to this embodiment. As shown in FIG. 26, the encoding / decoding system includes an encoding device 100 and a decoding device 200.

[0221] The encoding / decoding system accepts input three-dimensional meshes (also called 3D meshes) in the form of three-dimensional coordinates of vertices (vertex information), connectivity (connection information) and associated attributes (attribute information), which may include texture maps as well as geometry.

[0222] The encoding device 100 takes an input 3D mesh (also referred to as an input 3D mesh or input mesh) in the form of 3D coordinates of vertices, connectivity, and associated attributes. The encoding device 100 encodes all associated information into a stream. The stream may consist of a single bitstream or multiple bitstreams.

[0223] The network 300 transmits the stream generated by the encoding device 100 to the decoding device 200. The network 300 may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof. Furthermore, the network 300 is not necessarily limited to a two-way communication network, but may also be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Instead of the network 300, a recording medium such as a digital versatile disc (DVD) or a blue-ray disc (BD) on which a stream is recorded may be used.

[0224] The stream is transmitted to a decoding device 200 via a network 300. The decoding device 200 decodes the bitstream and generates a 3D mesh using the 3D coordinates, connectivity, and associated attributes of the decoded vertices. The decoding device 200 outputs the generated 3D mesh (also referred to as an output 3D mesh or output mesh).

[0225] FIG. 27 is a diagram showing another example of the configuration of the encoding device 100.

[0226] As shown in FIG. 27, the encoding device 100 includes a preprocessor 1103 and a compressor 1106 .

[0227] The encoding device 100 reads an input mesh 1101 and an attribute map 1102 and passes them to a preprocessor 1103. The preprocessor 1103 processes the input mesh to extract a base mesh 1104 and displacement data 1105. The attribute map 1102, along with the extracted base mesh 1104 and displacement data 1105, are passed to a compressor 1106.

[0228] The compressor 1106 also compresses the base mesh 1104, the displacement data 1105, and the attribute map 1102 to generate a bitstream 1107. The compressor 1106 can transmit additional information to the decoding device 200 by further including metadata 1108 in the bitstream 1107.

[0229] FIG. 28 is a diagram showing another example of the configuration of the decoding device 200.

[0230] As shown in FIG. 28, the decoding device 200 includes a decompressor 2102 and a post-processor 2106 .

[0231] The decoding device 200 reads a bitstream 2101 and passes it to a decompressor 2102. The decompressor 2102 decompresses a base mesh 2103, displacement data 2104, and an attribute map 2108 from the bitstream 2101 and passes them to a post-processor 2106. An example of the displacement data 2104 is a displacement vector.

[0232] The post-processor 2106 also processes the base mesh 2103 according to the displacement data 2104 and the attribute map 2108 to generate an output mesh 2107. The post-processor 2106 may further use information from the metadata 2105 to generate the output mesh 2107.

[0233] FIG. 29 is a block diagram showing yet another example configuration of the encoding device 100 according to this embodiment.

[0234] In this example, the encoding device 100 comprises a volumetric capturer 511, a projector 512, a base mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and optionally one or more other type encoders 516.

[0235] The volumetric capturer 511 captures content and outputs the captured content to the projector 512 .

[0236] The projector 512 projects the content onto an input mesh (a 3D mesh frame) that includes geometry coordinates (vertex coordinates indicating the positions of the vertices), texture coordinates, and connectivity (connectivity information). The data is output to a base mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and optionally one or more other type encoders 516. Each encoder compresses the data into a bitstream.

[0237] FIG. 30 is a block diagram showing yet another example configuration of the decoding device 200 according to this embodiment.

[0238] In this example, the decoding device 200 comprises a base mesh decoder 613 , a displacement decoder 614 , an attribute decoder 615 , one or more other type decoders 616 , and a 3D reconstructor 617 .

[0239] The bitstream is sent to a base mesh decoder 613, a displacement decoder 614, an attribute decoder 615, and optionally one or more other type decoders 616. These decoders decode the bitstream to generate data (decoded data) including geometry coordinates, texture coordinates, connectivity, etc. The decoded data is then sent to a 3D reconstructor 617, which reconstructs an output mesh (a 3D mesh frame).

[0240] The encoding process performed by the encoding device 100 will be described in detail below.

[0241] Fig. 31 is a flow diagram showing the processing of the encoding device 100. Fig. 32 is an explanatory diagram conceptually showing the encoding of mesh frames. The processing of the encoding device 100 will be described with reference to Figs. 31 and 32.

[0242] In step S101, the encoding device 100 reads a 3D mesh frame, which is an input mesh frame, and its attributes. The input mesh frame is a mesh frame input to the encoding device 100. An example of the 3D mesh frame that is an input mesh frame is shown as mesh frame 1301 (see FIG. 32 ).

[0243] In step S102, the encoding device 100 performs a decimation process on the input mesh frame read in step S101 to generate a base mesh frame having fewer vertices than the input mesh frame. The base mesh frame generated by decimating the mesh frame 1301 is shown as a base mesh frame 1302 (see FIG. 32).

[0244] In step S103, the encoding device 100 calculates displacement information that the decoding device 200 uses to reconstruct a mesh frame. The displacement information corresponds to a displacement vector directed from a vertex of the base mesh frame generated in step S102 to a vertex of the input mesh frame. One method for calculating the displacement information is to subtract the coordinates of the vertex of the base mesh frame from the coordinates of the vertex of the input mesh frame. The displacement information calculated from the mesh frame 1301 and the base mesh frame 1302 is shown as displacement information 1303 (see FIG. 32). The displacement information 1303 is in vector format, in other words, expressed as a displacement vector.

[0245] In step S104, the encoding device 100 encodes the base mesh frame generated in step S102, the displacement information generated in step S103, and the attributes of the input mesh frame into a bitstream (corresponding to a compressed bitstream). An example of the bitstream is shown as bitstream 1304 (see FIG. 32).

[0246] Specifically, the bitstream 1304 includes vertex coordinates and connectivity information for vertices A, C, E, and F, displacement information, a video bitstream including texture data, and a compressed attribute map (see FIG. 32). The displacement information includes displacement information for displacing vertices based on vertex coordinates obtained from the subdivided base mesh frame. The compressed attribute map is texture coordinates for applying texture data to a mesh frame reconstructed using the base mesh frame and the displacement information.

[0247] The decoding process performed by the decoding device 200 will be described in detail below.

[0248] Fig. 33 is a flow diagram showing the processing of the decoding device 200. Fig. 34 is an explanatory diagram conceptually showing the decoding of a mesh frame (3D mesh). The processing of the decoding device 200 will be described with reference to Figs. 33 and 34.

[0249] In step S201, the decoding device 200 decodes a base mesh frame and attributes from a bitstream (corresponding to a compressed bitstream). An example of the decoded base mesh frame (corresponding to a decoded base mesh frame) is shown as a decoded base mesh frame 2301 (see FIG. 34).

[0250] In step S202, the decoding device 200 generates subdivided vertices by performing a subdivision process on the base mesh frame decoded in step S201. An example of a base mesh frame (mesh frame) including subdivided vertices is shown as a base mesh frame 2302 (see FIG. 34).

[0251] In step S203, the decoding device 200 decodes the disparity information from the bitstream (corresponding to the compressed bitstream). An example of the decoded disparity information is shown as disparity information 2303 (see FIG. 34). The disparity information 2303 is in vector format, in other words, expressed as a disparity vector.

[0252] In step S204, the decoding device 200 reconstructs the shape of the mesh frame by moving the vertices of the base mesh frame, including the subdivided vertices, to new positions using the displacement information, and then restores the mesh frame by applying attribute information. An example of the attribute is texture. An example of the reconstructed mesh frame is shown as mesh frame 2304 (see FIG. 34 ).

[0253] FIG. 35 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.

[0254] FIG. 35 shows an example of a block diagram of a general intra-decoding system.

[0255] The decoding device shown in FIG. 35 comprises a demultiplexer 1231, a switch 1232, a static mesh decoder 1233, a mesh buffer 1234, a motion decoder 1235, a base mesh reconstructor 1236, an inverse quantizer 1237, a video decoder 1238, an image unpacker 1239, an inverse quantizer 1240, an inverse wavelet transformer 1241, a reconstructor 1242, a video decoder 1243, and a color converter 1244.

[0256] The demultiplexer 1231 receives the compressed bitstream and separates it into compressed data for the base mesh, video containing displacement data (also called displacement bitstream), and video containing attribute data (also called attribute bitstream). The compressed data for the base mesh is passed to a switch 1232. The switch 1232 determines whether to perform intra-decoding or inter-decoding based on parameters in the bitstream.

[0257] If an intra-decoding process is selected, the bitstream is passed to a static mesh decoder 1233, which generates a quantized base mesh. The static mesh decoder 1233 is, for example, a decoder that uses an edge breaker algorithm to decode 3D mesh data. The static mesh decoder 1233 generates a quantized base mesh from the bitstream. The quantized base mesh generated by the static mesh decoder 1233 is stored in a mesh buffer 1234 for reference when an inter-decoding process is selected.

[0258] If inter-decoding is selected, switch 1232 passes compressed data for the base mesh to motion decoder 1235. Motion decoder 1235 receives a previously decoded quantized base mesh and decodes motion data representing the differences in vertex coordinates between the quantized base mesh stored in mesh buffer 1234 and the current quantized base mesh. The motion data and the quantized base mesh stored in mesh buffer 1234 are used by base mesh reconstructor 1236 to reconstruct the current quantized base mesh. The quantized base mesh resulting from either inter-decoding or intra-decoding is passed to inverse quantizer 1237 to obtain a decoded base mesh.

[0259] The video containing the displacement data is passed to a video decoder 1238, since the bitstream contains the displacement data in an image format with two chroma information and one luma information. The video decoder 1238 decodes the data using a video frame decompression method. Alternatively, the displacement data can be decoded using an arithmetic decoder. This decompressed data is passed to an image unpacker 1239, which extracts wavelet coefficients associated with each vertex from the image-format decompressed data. An inverse quantizer 1240 dequantizes the quantized wavelet coefficients into the three components associated with each vertex. An inverse wavelet transformer 1241 inversely transforms the result to finally obtain decoded displacement data. The decoded displacement data and the decoded base mesh are passed to a reconstructor 1242, which performs edge refinement on the decoded base mesh and displaces the vertices using the decoded displacement data to obtain a decoded mesh.

[0260] The video containing the attribute data is passed to another video decoder 1243 to obtain a decoded attribute bitstream, which is further processed in a color converter 1244 for color space and color format conversion to obtain a decoded attribute map.

[0261] FIG. 36 is a block diagram showing an example of the configuration of an encoding device according to this embodiment.

[0262] First, the encoding device obtains the base mesh bitstream, the displacement bitstream, and the attribute bitstream resulting from the 3D mesh preprocessing step.

[0263] The encoding device shown in FIG. 36 comprises a quantizer 1261, a switch 1262, a static mesh encoder 1263, a mesh buffer 1264, a motion encoder 1265, a base mesh reconstructor 1266, a displacement data updater 1267, a wavelet transformer 1268, a quantizer 1269, an image packer 1270, a video encoder 1271, a color converter 1272, and a video encoder 1273.

[0264] The base mesh (specifically, the position information of the multiple vertices that make up the base mesh) is first quantized by a quantizer 1261. The quantized base mesh (base mesh data) is output to a switch 1262, which determines whether intra-coding or inter-coding is to be performed. If intra-coding is selected, the quantized base mesh (base mesh bitstream) is output to a static mesh encoder 1263, which generates a quantized base mesh. An example of the static mesh encoder 1263 is an encoder that uses an edge breaker algorithm to encode 3D mesh data. This encoded (quantized) base mesh is stored in a mesh buffer 1264 for reference when inter-coding is selected. If inter-coding is selected, the switch 1262 outputs compressed data related to the base mesh to a motion encoder 1265. The motion data and static 3D mesh in the mesh buffer 1264 are used by a base mesh reconstructor 1266 to reconstruct the currently quantized base mesh.

[0265] The displacement data is output to a displacement data updater 1267, where it is updated based on the quantized static base mesh or the reconstructed inter-coded base mesh. Next, a wavelet transformer 1268 performs a transformation process, followed by quantization in a quantizer 1269. The quantized displacement data is packed into an image in an image packer 1270, and finally encoded in a video coder 1271. The encoded displacement data is output to a multiplexer 1274.

[0266] The attribute information (e.g., attribute map) is output to a color converter 1272 for color space and color format conversion, and the converted attribute information is coded by a video encoder 1273 and output to a multiplexer 1274.

[0267] The multiplexer 1274 acquires data related to the encoded base mesh (compressed data related to the base mesh), video data including encoded displacement data, and video data including attribute information such as an encoded attribute map, and generates a bitstream (compressed bitstream) including these acquired data. The generated compressed bitstream is output to, for example, a decoding device.

[0268] FIG. 37 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.

[0269] FIG. 37 illustrates an example of a reconstructor that obtains a decoded 3D mesh 1256 from a decoded base mesh 1251 and decoded displacement data 1254.

[0270] The decoded base mesh 1251 is passed to a subdivider 1252 .

[0271] The subdivision unit 1252 subdivides any two connected vertices in the entire 3D mesh by adding a new vertex between them. This process can be repeated several times to include vertices created in previous subdivision steps to generate a predefined number of vertices. Each subdivision iteration across the 3D mesh generates a new level of detail (LoD). The subdivided mesh 1253 and the decoded displacement data 1254 are passed to a displacer 1255. The displacer 1255 generates a decoded 3D mesh 1256 by moving each vertex to a new position according to the corresponding displacement data.

[0272] The subdivision is described below and is performed, for example, by subdivision unit 1252.

[0273] FIG. 38 is an explanatory diagram showing an example of subdivision.

[0274] The base mesh shown in FIG. 38(a) includes vertices A, B, and C and connectivity information indicating their connectivity.

[0275] 38(b) shows a mesh generated by the first subdivision, in other words, the mesh after the first subdivision. In the first subdivision, the subdivider generates vertices D, E, and F and connectivity information indicating their connectivity. The mesh generated by the subdivider is also referred to as LoD1 or first LoD.

[0276] Vertex D of the mesh after the first subdivision is a vertex generated by subdivision based on vertices A and B. Similarly, vertex F is a vertex generated by subdivision based on vertices B and C. Vertex E is a vertex generated by subdivision based on vertices A and C.

[0277] As an example, vertex D may be the midpoint of line segment AB (in other words, side AB) connecting vertices A and B that were the basis for its generation. Similarly, vertex E may be the midpoint of line segment AC. Vertex F may be the midpoint of line segment BC.

[0278] 38(c) shows the mesh generated by the second subdivision, i.e., the mesh after the second subdivision. In the second subdivision, the subdivider generates vertices G, H, I, J, K, L, M, N, and O and connectivity information indicating their connectivity. The mesh generated by the subdivider is also called LoD2 or second LoD.

[0279] Vertex G of the mesh after the second subdivision is a vertex generated by subdivision based on vertices A and D. Similarly, vertex H is a vertex generated by subdivision based on vertices A and E. Vertex I is a vertex generated by subdivision based on vertices B and D. Vertex J is a vertex generated by subdivision based on vertices D and F. Vertex K is a vertex generated by subdivision based on vertices E and F. Vertex L is a vertex generated by subdivision based on vertices C and E. Vertex M is a vertex generated by subdivision based on vertices B and F. Vertex N is a vertex generated by subdivision based on vertices C and F. Vertex O is a vertex generated by subdivision based on vertices D and E.

[0280] As an example, vertex G may be the midpoint of line segment AD (in other words, side AD) connecting vertices A and D, which were the source of its generation. Similarly, vertex H may be the midpoint of line segment AE. vertex I may be the midpoint of line segment BD. vertex J may be the midpoint of line segment DF. vertex K may be the midpoint of line segment EF. vertex L may be the midpoint of line segment CE. vertex M may be the midpoint of line segment BF. vertex N may be the midpoint of line segment CF. vertex O may be the midpoint of line segment DE.

[0281] The displacement of vertices will be described below with reference to Figures 39 and 40. The displacement of vertices is performed by the reconstructor.

[0282] Fig. 39 is an explanatory diagram showing an example of displacement of vertices after subdivision, and Fig. 40 is an explanatory diagram showing an example of vertices of an original mesh.

[0283] The base mesh shown in FIG. 39(a) includes vertices A, B, C, and Z, and connectivity information indicating their connectivity.

[0284] 39(b) shows a mesh generated by the first subdivision, in other words, a mesh after the first subdivision (i.e., the first LoD). In the first subdivision, the subdivider generates vertices S, T, U, X, or Y and connectivity information indicating their connectivity. The vertices S, T, U, X, or Y are similar to the vertices D, E, and F shown in FIG. 38(b).

[0285] 39(c) shows a mesh generated by the second subdivision, in other words, a mesh after the second subdivision (i.e., the second LoD). In the second subdivision, the subdivider generates vertices D, E, F, G, and H and connectivity information indicating their connectivity. Vertices D, E, F, G, and H are similar to vertices G, H, I, J, K, L, M, N, and O shown in FIG. 38(c).

[0286] Figure 39(d) shows a mesh including the vertices after they have been displaced after subdivision, with vertices A, B, C, D, E, F, G, H, S, T, U, X, Y, and Z shown in Figure 39(d) being located at positions displaced using displacement information from the positions of the vertices shown in Figure 39(c).

[0287] The original mesh shown in FIG. 40 is an example of the mesh input to the encoding device 100, that is, the mesh before encoding.

[0288] The mesh shown in Fig. 39 has a shape similar to that of the original mesh shown in Fig. 40. The displacement information is generated by the displacement vector calculator 1207 of the encoding device 100 as information indicating the displacement from the vertices of the base mesh to the vertices of the original mesh, and therefore, by reconstructing the mesh using the displacement information thus generated, a mesh having a shape similar to that of the original mesh is generated.

[0289] The decoding device 200 can output the mesh shown in FIG.

[0290] Next, the division of a mesh into sub-meshes will be described with reference to FIGS.

[0291] A mesh can be divided into smaller parts and coded separately, with the vertices of the mesh being divided in such a way that the coordinates and connectivity of the vertices in each part can be coded independently.

[0292] Fig. 41 is an explanatory diagram showing an example of a mesh, and Fig. 42 is an explanatory diagram showing an example of dividing a mesh into sub-meshes.

[0293] The mesh shown in FIG. 41 is the original mesh, which is sometimes called a full mesh in contrast to a sub-mesh.

[0294] Figure 42 shows how the full mesh shown in Figure 41 is divided into two sub-meshes. For vertices A, B, and C of the full mesh (see Figure 41), vertex A is duplicated to vertices A1 and A2, vertex B is duplicated to vertices B1 and B2, and vertex C is duplicated to vertices C1 and C2, thereby creating two sub-meshes (i.e., a first sub-mesh and a second sub-mesh) from the full mesh. The first sub-mesh and the second sub-mesh are each independently decodable meshes.

[0295] Packing of displacement information into image frames will be described below with reference to FIGS. 43, 44 and 45.

[0296] 43, 44, and 45 are explanatory diagrams showing examples of packing displacement information into image frames. Note that image frames can also be called video frames.

[0297] The vertex displacement data is encoded as image frame data by being mapped to each component of a YUV format image frame (i.e., each of the Y component (Y Plane), U component (U Plane), and V component (V Plane)). This case will be described below as an example. As another example, the vertex displacement data may be encoded as image frame data by being mapped to each component of an RGB format image frame (each of the R component, G component, and B component).

[0298] The decoding device 200 can use an image encoding module to extract the displacement data. The displacement data can be in the form of X, Y, or Z components in a global coordinate system (e.g., a Cartesian coordinate system), or normal, tangential, or both tangential components in a local coordinate system. Methods for mapping the displacement data to an image frame include the following:

[0299] For example, in the first method, the displacement data is arranged in the image frame in scan order. An example of packing the displacement data in this case is shown in Figure 43. The displacement data is directly mapped to the image frame according to a predefined scan order.

[0300] Note that since an image frame has a fixed height and width, it may happen that the displacement data does not fit perfectly in the frame, in which case the remaining part of the image frame is padded with padding data (see Figure 43).

[0301] For example, in the second method, the displacement data is separated into multiple LoDs and mapped to the Y, U, and V components of the image frame. An example of packing of the displacement data in this case is shown in Figure 44. Here, the displacement data of the image frame of the next LoD starts immediately after the displacement data of the previous LoD ends. As in the first method, if the displacement data does not fit exactly into the image frame, padding is performed at the end of the image frame (see Figure 44).

[0302] For example, in the third method, displacement data corresponding to the LoD is mapped to the Y component, U component, and V component of the image frame in a manner different from that in the second method. An example of packing of the displacement data in this case is shown in Figure 45. In this way, each LoD can be decoded independently. In the third method, middle padding is performed on the displacement data of each LoD, and CTU alignment is performed together with padding at the end of the video frame (see Figure 45).

[0303] Next, the encoding device 100 and the decoding device 200 when a mesh is divided into a plurality of sub-meshes will be described.

[0304] Fig. 46 is a diagram showing another example of the configuration of the encoding device 100 according to an embodiment. Specifically, Fig. 46 is a diagram showing the configuration of a submesh encoding device that performs encoding processing when an input mesh 1101 is divided into divided meshes (plurality of submeshes). For example, the submesh encoding device includes a plurality of encoding devices 100.

[0305] An input mesh 1101 (full mesh) input to the submesh encoding device is divided into multiple meshes (submeshes). The multiple submeshes are input to, for example, multiple encoding devices 100. Each of the multiple submeshes may be input to one of the multiple encoding devices 100. For example, the submesh encoding device divides the input mesh 1101 into multiple submeshes and inputs the divided multiple submeshes to the multiple encoding devices 100.

[0306] After the image is divided into a plurality of sub-meshes, coding processing is performed on the boundaries of the sub-meshes (processing of overlapping sub-meshes).

[0307] For example, for each sub-mesh, preprocessing is performed by the preprocessor 1103, and a base mesh, displacement data, and metadata are generated and encoded.

[0308] The encoding device 100 may be realized by such a submesh encoding device configuration. That is, the encoding device 100 may be configured to include a plurality of preprocessors 1103 and compressors 1106, and may perform a predetermined process on the submesh for each of a plurality of pairs of preprocessors 1103 and compressors 1106. Furthermore, the number of pairs of preprocessors 1103 and compressors 1106 included in the encoding device 100 may be any number and is not particularly limited.

[0309] FIG. 47 is a diagram illustrating an example of the configuration of the preprocessor 1103 according to the embodiment.

[0310] The preprocessor 1103 includes, for example, a base mesh generator 1401 , a subdivision unit 1402 , and a displacement data generator 1403 .

[0311] First, in the preprocessor 1103, a base mesh is generated by the base mesh generator 1401.

[0312] The base mesh is then subdivided in a predetermined manner by a subdivider 1402 to generate a subdivided mesh (or a subdivided base mesh) that is the subdivided base mesh.

[0313] Displacement data is generated by a displacement data generator 1403 from the sub-divided mesh and the sub-mesh that is the input mesh 1101 after division.

[0314] The displacement data is, for example, a difference vector between the input mesh 1101 and the sub-division mesh.

[0315] The same subdivision method as that used in encoding is used in decoding as well.

[0316] Furthermore, for example, the encoding device 100 may transmit to the decoding device 200 a subdivision method for encoding and parameters used in the subdivision.

[0317] Fig. 48 is a diagram showing another example of the configuration of the decoding device 200 according to the embodiment. Specifically, Fig. 48 is a diagram showing the configuration of a submesh decoding device, which is a device that performs decoding processing when a bitstream 2101 includes multiple submeshes. For example, the submesh decoding device includes multiple decoding devices 200. For example, the submesh decoding device includes multiple decoding devices 200 and a combiner 2109.

[0318] The coded data for each submesh included in the bit stream 2101 is input to the decompressor 2102 of each decoding device 200 .

[0319] Also, for example, the post-processor 2106 performs the processing of the reconstructor described above.

[0320] In the post-processor 2106, for each decoded sub-mesh, the base mesh is subdivided and the displacement vector is added to the subdivided base mesh to reconstruct the sub-mesh.

[0321] That is, the post-processor 2106 performs the above-described reconstructor processing for each sub-mesh.

[0322] The combiner 2109 combines (merges) the sub-meshes restored by the respective decoding devices 200 to reconstruct the full mesh (output mesh 2107) before division.

[0323] Note that the decoding device 200 may be realized by such a submesh decoding device configuration. That is, the decoding device 200 may be configured to include a plurality of decompressors 2102 and post-processors 2106, and each of the plurality of pairs of decompressors 2102 and post-processors 2106 may perform a predetermined process on the encoded data for each submesh. Furthermore, the number of pairs of decompressors 2102 and post-processors 2106 included in the decoding device 200 may be any number and is not particularly limited.

[0324] Fig. 49 is a diagram showing a specific example of the configuration of the decoding device 200 according to the embodiment. Specifically, Fig. 49 shows a specific configuration of the post-decoder 2306 out of the decoder 2305 and post-decoder 2306 included in the decoding device 200.

[0325] The post-decoder 2306 comprises a pre-reconstructor 2307 , a reconstructor 2308 , a post-reconstructor 2309 and an adaptor 2310 .

[0326] The processing in the post-decoder 2306 is optional depending on the application. An example of the processing in the post-decoder 2306 (post-decoding processing) is conversion of the decoded data to a nominal format, such as video conversion from YUV space to RGB space. The post-decoding processing can be encapsulated into multiple processes, such as pre-reconstruction, reconstruction, post-reconstruction, and adaptation.

[0327] The pre-reconstructor 2307 performs a pre-reconstruction process, for example, upscaling the normalized texture coordinates to match the dimensions of the texture image in the context of video-based dynamic mesh coding.

[0328] The reconstructor 2308 performs a reconstruction process, which is invoked, for example, on the decoded atlas frame, the decoded base mesh frame, the decoded video frame, and syntax elements associated with the same mesh sequence. The output of the reconstruction process is a series of reconstructed mesh frames prior to a post-reconstruction process.

[0329] The post-reconstructor 2309 performs post-reconstruction operations, e.g., in the context of video-based dynamic mesh coding, which perform a number of smoothing operations on the reconstructed mesh frames, such as collapsing edges in the mesh or adding new vertices in the mesh.

[0330] The adaptor 2310 performs a fitting process, which may be applied by some application to fit the reconstructed mesh to a given scenario. For example, the vertices of the reconstructed mesh are transformed from the 3D model coordinate system to the 3D world coordinate system. The adaptor 2310 outputs a final reconstructed mesh frame (final 3D mesh frame 2311).

[0331] The details of subdivision are explained below.

[0332] <Subdivision> Figure 50 is a diagram showing an example of two subdivided submeshes having boundary edges according to an embodiment. Specifically, (a) of Figure 50 shows an example of a submesh, and (b) of Figure 50 shows another example of a submesh. (c) of Figure 50 shows a 3D mesh in which the submesh shown in (a) of Figure 50 and the submesh shown in (b) of Figure 50 are merged (i.e., combined).

[0333] The decoding device 200 subdivides each edge of the submesh based on position information of each vertex of the submesh and connection information indicating the connection relationship between the vertices, i.e., information on the multiple edges of the submesh. In the subdivision, for example, new vertices (three-dimensional points) are generated on the edges. The generated vertices are connected by new edges, for example. As a result, for example, each of the multiple faces included in the submesh is subdivided into multiple faces. For example, a triangular face enclosed by three vertices included in the submesh is subdivided into four faces.

[0334] The decoding device 200 performs subdivision for each submesh and merges the resulting submeshes to reconstruct a three-dimensional mesh corresponding to the original mesh.

[0335] Here, the encoding device 100 generates multiple submeshes from an original mesh and encodes, for each submesh, the position information, displacement information, etc., of the vertices constituting each submesh. Therefore, the position information, displacement information, etc., of the vertices constituting each submesh may be encoded using different encoding parameters for each submesh. As a result, when the decoding device 200 performs subdivision based on this encoded information, an edge (also called a boundary edge) shared by two or more submeshes may be subdivision at a different position for each submesh, or the number of times each submesh is subdivision may differ. This may result in a problem in which the submeshes cannot be merged properly.

[0336] For example, in the example shown in Figure 50, side CD is a boundary edge. When the submesh shown in Figure 50(a) is subdivided, for example, vertex F is generated on side CD. Also, when the submesh shown in Figure 50(b) is subdivided, for example, vertex M is generated on side CD. If the positions of vertex F and vertex M are different, or if the number of vertices generated on side CD is different, that is, if the number of times side CD is subdivided is different, when the submesh shown in Figure 50(a) and the submesh shown in Figure 50(b) are merged, gaps will be created around side CD, and the submesh will not be able to merge properly.

[0337] To solve this problem caused by using different subdivision schemes in adjacent submeshes (i.e., submeshes having the same boundary edge), the present application imposes a constraint that all boundary edges have the same type of subdivision scheme and the same number of iterations. That is, for example, when subdividing a boundary edge, the subdivision is performed in each submesh using the same subdivision method and the same number of subdivisions. As a result, if the displacement information (i.e., the value of the displacement vector) of vertex F and vertex M is the same, the vertices generated by the subdivision are displaced, and after the submeshes are merged, there can be no gaps around the boundary edge, as shown in (c) of FIG.

[0338] Note that the subdivision of the boundary edge may be performed using a predetermined subdivision method and / or a predetermined number of subdivisions, or the encoding device 100 may determine the predetermined subdivision method and / or the predetermined number of subdivisions and signal the determined information in the bitstream. Also, for example, the alignment of vertex F with vertex M may be performed when subdividing a submesh or when merging multiple submeshes.

[0339] [First Aspect] Next, the first aspect of subdivision will be described.

[0340] Fig. 51 is a flowchart showing the process of the decoding device 200 according to the embodiment. Fig. 51 is a flowchart showing an example in which a predetermined subdivision method is used.

[0341] First, the decoding device 200 decodes, from the bit stream, a first vertex and a second vertex that are connected via an edge (S301). Specifically, the decoding device 200 acquires, from the bit stream, position information of the first vertex and the second vertex and connection information indicating whether the first vertex and the second vertex are connected.

[0342] Next, the decoding device 200 determines whether the edge connecting the first vertex and the second vertex is a boundary edge (S302).

[0343] When the decoding device 200 determines that the edge connecting the first vertex and the second vertex is a boundary edge (Yes in S302), it derives a third vertex from only the first vertex and the second vertex (S303). That is, the position of the third vertex is derived based only on the positions of the first vertex and the second vertex, and the third vertex is complemented (added).

[0344] On the other hand, if the decoding device 200 determines that the edge connecting the first vertex and the second vertex is not a boundary edge (No in S302), it derives a fourth vertex from the first vertex, the second vertex, and at least one fifth vertex (S304). That is, the position of the fourth vertex is derived based on the positions of the first vertex, the second vertex, and at least one fifth vertex, and the fourth vertex is complemented.

[0345] The number of fifth vertices may be one or more, that is, the fourth vertex may be derived from three or more vertices including the first vertex and the second vertex.

[0346] The processes in steps S303 and S304 are merely examples, and other methods may be used.

[0347] 52 and 53 are diagrams for explaining an example of a boundary edge and a non-boundary edge, respectively, according to an embodiment.

[0348] A boundary edge is, for example, an edge that is shared by multiple sub-meshes.

[0349] On the other hand, a non-boundary edge is an edge that is not shared by multiple sub-meshes, i.e., an edge that is included in only one of the multiple sub-meshes included in the base mesh.

[0350] In the example shown in FIG. 52 , only one face Y formed by side AB and vertex F is connected to side AB. Therefore, side AB is a boundary edge. Also, for example, only one face Z formed by side EF and vertex D is connected to side EF. Therefore, side EF is a boundary edge.

[0351] Also, for example, side BD is connected to two faces: face W formed by side BD and vertex C, and face X formed by side BD and vertex F. Therefore, side BD is not a boundary edge but a non-boundary edge. Also, for example, side DF is connected to two faces: face X formed by side DF and vertex B, and face Z formed by side DF and vertex E. Therefore, side DF is not a boundary edge but a non-boundary edge.

[0352] For example, in step S303, if the edge connecting the first vertex and the second vertex is a boundary edge, the decoding device 200 derives the third vertex from only the first vertex and the second vertex. For example, in FIG. 52, vertex A is assumed to be the first vertex and vertex C is assumed to be the second vertex. Furthermore, vertex A and vertex C are assumed to be on the boundary edge. In this case, vertex B, which is the third vertex, is derived from only the first vertex and the second vertex. As an example of deriving the coordinates (position) of B, there is a method of assigning the midpoint between vertex A, which is the first vertex, and vertex C, which is the second vertex, to the coordinates. In other words, the coordinates of the midpoint between vertex A and vertex C may be calculated as the coordinates of intersection B.

[0353] Also, for example, in step S304, if the edge connecting the first vertex and the second vertex is not a boundary edge, decoding device 200 derives a fourth vertex from the first vertex, the second vertex, and at least one fifth vertex. For example, in FIG. 53 , vertex B is the first vertex, and vertex D is the second vertex. Furthermore, vertices B and D are not on a boundary edge. In other words, edge BD is a non-boundary edge. In this case, vertex H, which is the fourth vertex, is derived from vertex B, which is the first vertex, vertex D, which is the second vertex, and another vertex, vertex F, which is a fifth vertex. As an example of deriving the coordinates of vertex H, there is a method of assigning the coordinates of vertex H when vertex F, which is the fifth vertex, is orthogonally projected onto edge BD.

[0354] For example, one or more parameters are decoded from a bitstream. Alternatively, the one or more parameters may be decoded from a header of the bitstream. Alternatively, the one or more parameters may include boundary edge information indicating which of the multiple edges constituting the submesh are boundary edges. The decoding device 200 may identify the boundary edges using the decoded boundary edge information.

[0355] FIG. 54 is a diagram illustrating an example of syntax for signaling different sub-division types (also referred to as sub-division methods) and the number of sub-divisions in a header according to an embodiment.

[0356] The encoding device 100 generates a bitstream that includes, for example, information (subdivision_type) indicating the method of subdividing each edge that constitutes the submesh (also simply referred to as the submesh subdivision method) and information (subdivision_num_iteration) indicating the number of times that each edge that constitutes the submesh is subdivided (also simply referred to as the submesh subdivision number) in the submesh header.

[0357] FIG. 55 is a diagram illustrating an example of syntax for signaling a subdivision type and the number of subdivisions using a sequence parameter set (SPS) according to an embodiment.

[0358] The encoding device 100 generates a bitstream whose sequence parameter set includes, for example, information indicating whether an edge is a boundary edge (boundary_subdivision_flag), information indicating a method for subdividing the boundary edge (also simply referred to as the boundary edge subdivision method) (boundary_subdivision_type), and information indicating the number of times the boundary edge is subdivided (also simply referred to as the number of times the boundary edge is subdivisioned) (boundary_subdivision_num_iteration).

[0359] FIG. 56 is a diagram showing an example of syntax for determining the subdivision type and the number of subdivisions using a sequence parameter set according to an embodiment, and for checking whether an edge is on a boundary.

[0360] For example, if an edge is a boundary edge, the decoding device 200 subdivides the edge based on information indicating a boundary edge subdivision method and information indicating the number of times the boundary edge is subdivisioned, which are included in the SPS of the bitstream acquired from the encoding device 100. On the other hand, if the edge is not a boundary edge, the decoding device 200 subdivides the edge based on information indicating a subdivision method of the submesh and information indicating the number of times the submesh is subdivisioned, which are included in the header of the submesh of the bitstream acquired from the encoding device 100.

[0361] In addition, information indicating the subdivision method of the submesh and / or information indicating the number of times the submesh is subdivision may be stored in a parameter set common to each frame if it is common to each frame, or may be stored in a parameter set common to the sequence if it is common to the sequence.

[0362] In addition, if this information is common to each sub-mesh, this information does not need to be signaled in the header of the sub-mesh. Also, if the sub-division method of the sub-mesh and / or the number of sub-divisions of the sub-mesh are not signaled in any of the headers of adjacent sub-meshes, it may be determined that the sub-division method of the sub-mesh and / or the number of sub-divisions of the sub-mesh are common to each sub-mesh, and it may be determined not to use this method.

[0363] In the above, an example is shown in which information indicating the subdivision method of the boundary edge and information indicating the number of subdivisions of the boundary edge are stored in a parameter set common to the sequence (specifically, SPS), but these pieces of information may also be stored in a parameter set common to the frame.

[0364] Furthermore, the subdivision method and / or the number of subdivisions indicated by the information stored in the parameter set common to the sequence may be used for the subdivision method of the boundary edge and / or the number of subdivisions of the boundary edge. That is, when the subdivision method and / or the number of subdivisions are signaled both in the parameter set common to the sequence and in the header of the submesh, the decoding device 200 may use the value of the submesh header for the subdivision of the submesh and the value of the parameter set common to the sequence for the subdivision of the boundary edge.

[0365] For example, if at least one of the multiple edges constituting a submesh is a boundary edge and the number of subdivisions of the boundary edge is different from the number of subdivisions of the submesh, it is necessary to define how to subdivide the multiple edges constituting the submesh including the boundary edge and how to connect the vertices to generate a new mesh. Note that the number of subdivisions of the boundary edge of a submesh may be different from the number of subdivisions of the submesh, or the number of subdivisions of the boundary edge of a submesh to be merged may be different from the number of subdivisions of the submesh.

[0366] FIG. 57 is a flowchart showing an example of a process for dividing a plurality of edges that form a submesh according to an embodiment.

[0367] First, the decoding device 200 determines whether or not at least one of the multiple edges constituting a mesh (specifically, a submesh) is a boundary edge (S401).

[0368] If the decoding device 200 determines that at least one of the multiple edges that make up the submesh is a boundary edge (Yes in S401), it compares the number of subdivisions of the boundary edge with the number of subdivisions of the submesh (S402).

[0369] Next, the decoding device 200 performs subdivision based on the comparison result in step S402 (S403).

[0370] For example, the boundary edge is sub-divided A times, and the sub-mesh is sub-divided B times.

[0371] For example, when A=B, the decoding device 200 performs subdivision using a method A described later. Also, for example, when A<B, the decoding device 200 performs subdivision using a method B (method B1 or method B2) described later. Also, for example, when A>B, the decoding device 200 performs subdivision using a method C (method C1 or method C2) described later.

[0372] In addition, if the decoding device 200 determines that at least one of the multiple edges that make up the submesh is not a boundary edge (No in S401), it subdivides the multiple edges using the submesh subdivision method and the number of times the submesh is subdivisioned (S404).

[0373] [Method A] When subdivision is performed using Method A, that is, when the number of subdivisions of the boundary edge (A times) is the same as the number of subdivisions of the submesh (B times) (A = B), the decoding device 200 generates a new mesh by subdividing each edge constituting the submesh and connecting the vertices generated by the subdivision. The decoding device 200 repeats this process A (= B) times.

[0374] When the decoding device 200 performs subdivision using method B, that is, when the number of subdivisions of the boundary edge (A times) is less than the number of subdivisions of the submesh (B times) (A < B), the decoding device 200 performs subdivision using the following method B1 or method B2.

[0375] [Method B1] Figure 58 is a diagram illustrating a first example of a process for dividing multiple edges constituting a submesh according to an embodiment. In the example shown in Figure 58, polygon ABCD is the submesh, side CD (side CF' and side F'D) and side EC are boundary edges, and the other edges are non-boundary edges. Also, in the example shown in Figure 58, the submesh is sub-divided twice, and the boundary edge is sub-divided once.

[0376] In the example shown in step S303, the decoding device 200 decodes a parameter from the bitstream and derives the number of subdivisions of the boundary edge based on the parameter, which may be a value indicating the exact number of subdivisions of the boundary edge or a value indicating the difference from the number of subdivisions of the sub-mesh.

[0377] When the number of subdivisions of a boundary edge is less than the number of subdivisions of a submesh, the decoding device 200 performs subdivision on the boundary edge (edge ​​CD) the number of times as shown in FIG. 58 . In this example, for example, the decoding device 200 generates a vertex F′ by subdividing edge CD once. The decoding device 200 then subdivides each of the edges that are not boundary edges (e.g., edges AB, BC, AD, and BD). Furthermore, the decoding device 200 subdivides the edges that are not boundary edges and newly generated edges by connecting the vertices newly generated by the subdivision. As a result, the edges that are not boundary edges are subdivided twice. The decoding device 200 also creates an edge CW by connecting vertex C and the newly generated vertex W, and creates an edge DT by connecting vertex D and the newly generated vertex T. The decoding device 200 then creates an edge TL.

[0378] In this way, the decoding device 200 determines, for example, whether the number of subdivisions of a submesh is different from the number of subdivisions of a boundary edge and whether (the number of subdivisions of a submesh)>(the number of subdivisions of a boundary edge) is true.

[0379] If the determination is Yes, the decoding device 200 first subdivides each edge constituting the submesh using a conventional method (connecting the vertices after division to generate a new mesh) up to the number of times the boundary edge has been subdivisioned. Furthermore, if the number of times the boundary edge has been subdivisioned exceeds the number of times the submesh has been subdivisioned, the decoding device 200 repeats the following process until the number of times the submesh has been subdivisioned.

[0380] The decoding device 200 determines whether at least one edge constituting a mesh to be sub-divided (e.g., a polygonal face included in the sub-mesh) is a boundary edge of the sub-mesh. If there is no boundary edge, the decoding device 200 performs normal sub-division. If there is a boundary edge, the decoding device 200 sub-divides the non-boundary edge.

[0381] Furthermore, when the number of boundary edges in a mesh to be subdivided is one and the number of non-boundary edges is two, the decoding device 200 connects vertices generated by subdividing at least the non-boundary edges, and generates a new mesh using one of two edges formed by connecting the generated vertex to each of the two vertices constituting the boundary edges. Furthermore, for example, the decoding device 200 determines a priority based on whether the edge to which the vertex is connected has been subdivided, and determines which edge to use according to the determined priority.

[0382] Alternatively, a new mesh may be generated using both of the two edges.

[0383] Furthermore, for example, the plurality of edges constituting a submesh may include edges that are not sub-divided, and for example, information indicating the edges that are not sub-divided may be included in the bitstream.

[0384] Method B2: In Method B1, when performing subdivision, if the number of subdivisions of a boundary edge is exceeded, if there is no boundary edge, the normal subdivision method is performed, otherwise the non-boundary edge is subdivisioned.

[0385] In contrast, method B2 performs the normal subdivision method if there are no boundary edges, and does not perform subdivision if there are boundary edges.

[0386] 59 is a diagram illustrating a second example of a process for dividing a plurality of edges constituting a submesh according to an embodiment. In the example shown in FIG. 59, polygon ABCD is the submesh, edges CD (edges CF' and F'D) and EC are boundary edges, and the other edges are non-boundary edges. In the example shown in FIG. 59, the submesh is sub-divided twice, and the boundary edge is sub-divided once.

[0387] In another example of step S303, the decoding device 200 decodes a parameter from the bitstream and derives the number of subdivisions of the boundary edge based on the parameter. The parameter is, for example, a value indicating the exact number of subdivisions of the boundary edge, or a value indicating the difference from the number of subdivisions of the submesh. If the value indicating the number of subdivisions of the boundary edge is smaller than the value indicating the number of subdivisions of the submesh, the decoding device 200 subdivides each edge constituting the submesh a number of times corresponding to the value indicating the number of subdivisions of the boundary edge, as shown in FIG. 59 . Thereafter, the decoding device 200 does not explicitly shift vertex W and vertex T. Instead, the decoding device 200 shifts vertex F, for example, based on the shift information, and then projects vertex W onto edge EF and vertex T onto edge GF, thereby subdividing edge EF and edge GF. In this way, even after the face (specifically, each vertex constituting the face) is shifted, vertices E, W, and F′ and vertices G, T, and F′ are aligned on a straight line. Next, the decoding device 200 creates the sides CW and DT, and further creates the side TL.

[0388] [Method C1] Figure 60 is a diagram illustrating a third example of a process for dividing multiple edges constituting a submesh according to an embodiment. In the example shown in Figure 60, polygon CDU is the submesh, and edge CD is a boundary edge. Edges CU and DU are non-boundary edges. In the example shown in Figure 60, the submesh is sub-divided once, and the boundary edge is sub-divided twice.

[0389] If the number of subdivisions of the boundary edge is greater than the number of subdivisions of the submesh, the decoding device 200 first subdivides all edges including the boundary edge the number of times equal to the number of subdivisions of the submesh, as shown in Figure 60. Next, the decoding device 200 subdivides the boundary edge until the number of times the boundary edge has been subdivided equals the number of subdivisions of the boundary edge.

[0390] In addition, if the number of subdivisions of a submesh is less than the number of subdivisions of a boundary edge, the number of subdivisions of the boundary edge may be regarded as the number of subdivisions of the submesh. For example, if (the number of subdivisions of a submesh) < (the number of subdivisions of a boundary edge), the decoding device 200 may regard (the number of subdivisions of a submesh) = (the number of subdivisions of a boundary edge), so that the number of subdivisions of the boundary edge may be defined so as not to exceed the number of subdivisions of the submesh.

[0391] [Method C2] Figure 61 is a diagram illustrating a fourth example of a process for dividing multiple edges constituting a submesh according to an embodiment. In the example shown in Figure 61, polygon CDU is the submesh, and edge CD is a boundary edge. In the example shown in Figure 61, edges CU and DU are non-boundary edges. In the example shown in Figure 61, the submesh is sub-divided once, and the boundary edge is sub-divided three times.

[0392] If the number of subdivisions of the boundary edge is greater than the number of subdivisions of the submesh, the decoding device 200 first subdivides the boundary edge a number of times equal to the number of subdivisions of the boundary edge minus the number of subdivisions of the submesh to generate sides UQ, UM, and UR. Next, the decoding device 200 subdivides all edges including the boundary edge a number of times equal to the number of subdivisions of the submesh, as shown in FIG.

[0393] 62 is a diagram illustrating a fifth example of a process for dividing a plurality of edges constituting a submesh according to an embodiment of the present invention, specifically, a method for subdividing an edge BD, which is a non-boundary edge.

[0394] In the example of step S304, the decoding device 200 determines that the side BD is a non-boundary edge. Then, the decoding device 200 derives a vertex G using the vertices B, D, and A. For example, the decoding device 200 determines the intersection of the bisector of the angle BAD and the side BD as the vertex G. For example, the decoding device 200 subdivides the side BD by generating the vertex G in this way.

[0395] When an edge is subdivided, the subdivision method and the number of subdivisions may be set to be the same.

[0396] For example, the decoding device 200 may determine whether a submesh to be subdivided includes a boundary edge, and if the submesh does not include a boundary edge, may perform subdivision using the submesh subdivision method and the submesh subdivision count. On the other hand, for example, if the submesh includes a boundary edge, the decoding device 200 may perform subdivision using either or both of (i) the submesh subdivision method and the submesh subdivision count, and (ii) the boundary edge subdivision method and the boundary edge subdivision count.

[0397] A predetermined subdivision method and number of subdivisions may be used for the subdivision. Alternatively, the encoding device 100 may determine the subdivision method and number of subdivisions, and transmit the determined subdivision method and number of subdivisions in a bitstream.

[0398] Furthermore, for example, a method may be adopted in which either the information regarding the subdivision of boundary edges (e.g., the subdivision method of boundary edges and the number of times the boundary edges are subdivisioned) or the information regarding the subdivision of submeshes (e.g., the subdivision method of submeshes and the number of times the submeshes are subdivisioned) is determined in a predetermined manner, and the other is notified by signaling.

[0399] Also, the subdivision of submeshes may be performed using a subdivision method that is suitable for subdivision of submeshes, and the subdivision of boundary edges may be performed using a subdivision method that is suitable for subdivision of boundary edges. Different subdivision methods may be used for boundary edges and non-boundary edges if suitable subdivision methods exist, or the same subdivision method may be used for boundary edges and non-boundary edges if suitable subdivision methods do not exist.

[0400] The subdivision method may be switched depending on the result of comparing the number of subdivisions of the boundary edges with the number of subdivisions of the sub-meshes.

[0401] The subdivision of a mesh to be subdivided may use the boundary edge subdivision number for boundary edges and the submesh subdivision number for non-boundary edges, or may use either the boundary edge subdivision number or the submesh subdivision number.

[0402] The order in which the above determination processes are performed may be changed arbitrarily.

[0403] [Effects of the first aspect] In the configuration according to the first aspect, it is possible to merge sub-meshes coded using different sub-division methods or sub-division counts, thereby improving the subjective quality of the full mesh (the three-dimensional mesh after merging multiple sub-meshes).

[0404] [Combination with Other Aspects] The decoding device 200 of this aspect can be implemented by combining it with at least a part of other aspects of the present disclosure. Furthermore, this aspect may be implemented by combining a part of the processing shown in any of the flowcharts according to this aspect, a part of the configuration of any of the devices, and / or a part of the syntax, etc., with other aspects.

[0405] The above-described processing of the decoding device 200 can also be similarly performed in the encoding device 100. Furthermore, not all of the components shown in this aspect are always necessary, and only some of the components of the first aspect may be included.

[0406] [Second Aspect] Fig. 63 is a diagram for explaining a sixth example of the process for dividing a plurality of edges constituting a submesh according to an embodiment. Specifically, Fig. 63 is a diagram for explaining another example of the process of step S303.

[0407] In step S303, for example, the encoding device 100 encodes into the bitstream a parameter used by the decoding device 200 to derive the number of subdivisions of the boundary edge. The parameter may be a value indicating the exact number of subdivisions of the boundary edge, or a value indicating the difference from the number of subdivisions of the submesh. If the value indicating the number of subdivisions of the boundary edge (derived value) is greater than the value indicating the number of subdivisions of the submesh, the encoding device 100 performs subdivision of the boundary edge corresponding to the derived value. For example, if the derived value is 3, three vertices (vertex Q, vertex M, and vertex R) are generated along the boundary edge, and their connectivity is added to the base mesh of the submesh, as shown in FIG. 63 . The decoding device 200 performs subdivision corresponding to the number of subdivisions of the submesh. As a result, the decoding device 200 performs uniform subdivision along all edges, including the boundary edge, in the modified base mesh corresponding to the submesh that has been subdivided by the number of subdivisions of the submesh.

[0408] [Effects of the Second Aspect] In the configuration according to the second aspect, sub-meshes coded using different sub-division methods or different sub-division counts can be merged, thereby improving the subjective quality of the full mesh.

[0409] [Combination with Other Aspects] This aspect of the encoding device 100 can be implemented by combining it with at least a part of other aspects of the present disclosure. Furthermore, this aspect may be implemented by combining a part of the processing shown in any of the flowcharts according to this aspect, a part of the configuration of any of the devices, and / or a part of the syntax, etc., with other aspects.

[0410] The above-described processing of the encoding device 100 can also be similarly executed in the decoding device 200. Furthermore, not all of the components described in this aspect are always required, and only some of the components of the first aspect may be included.

[0411] In the third aspect, the encoding device 100 does not perform processing to make identical vertices on edges shared by multiple submeshes (also called boundary edges between submeshes). On the other hand, in the third aspect, the encoding device 100 transmits to the decoding device 200 metadata that enables the decoding device 200 to modify vertices on boundary edges (vertices located on boundary edges).

[0412] For example, the encoding device 100 signals hint information (e.g., a flag, etc.) in the metadata indicating that the number of subdivisions or the subdivision methods are different between submeshes (i.e., between adjacent submeshes), or that as a result, the vertices after subdivision will not be identical (or may not be identical), and transmits this to the decoding device 200.

[0413] The decoding device 200 determines whether to modify the vertices at the boundary edge between the sub-meshes based on the metadata. If the decoding device 200 determines to modify the vertices, the decoding device 200 modifies the number of vertices at the boundary edge by increasing or decreasing the number of vertices at the boundary edge so that the number of vertices at the boundary edge is the same for each sub-mesh that includes the boundary edge.

[0414] Alternatively, the encoding device 100 may signal, in the metadata, information indicating whether or not the decoding device 200 should modify a vertex on a boundary edge and / or information indicating a method for modifying the vertex, rather than hint information. In this case, for example, the decoding device 200 determines whether or not to modify the vertex based on the information.

[0415] The encoding device 100 may signal information indicating whether multiple submeshes overlap and / or information indicating how multiple submeshes overlap. Furthermore, the encoding device 100 may perform processing related to the overlap (e.g., processing such as modifying vertices) when multiple submeshes may overlap.

[0416] FIG. 64 is a flowchart showing the boundary vertex determination process executed by the encoding device 100 according to the embodiment.

[0417] First, the encoding device 100 performs subdivision for each submesh based on the number of subdivisions (S501).

[0418] Next, the encoding device 100 determines whether the boundaries between the sub-meshes overlap (S502). For example, the encoding device 100 determines whether the sub-meshes have boundary edges (i.e., edges shared with other sub-meshes).

[0419] When the encoding device 100 determines that the boundaries between the submeshes overlap (Yes in S503), it indicates in the metadata that the boundaries between the submeshes overlap (S504). That is, the encoding device 100 includes information indicating that the boundaries between the submeshes overlap in the metadata. In other words, the encoding device 100 generates metadata including information indicating that the boundaries between the submeshes overlap.

[0420] Next, the encoding device 100 determines whether the boundary vertices (specifically, the vertices at the boundary edge, in other words, the vertices located on the boundary edge) after subdivision are identical at the boundary edge (overlapping edge between sub-meshes) (S505).

[0421] If the encoding device 100 determines that the boundary vertices after the subdivision are not the same (No in S505), the encoding device 100 indicates in the metadata that the boundary vertices after the subdivision are not the same for each submesh (S506). That is, the encoding device 100 includes information indicating that the boundary vertices are not the same in the metadata. In other words, the encoding device 100 generates metadata that also includes information indicating that the boundary vertices are not the same.

[0422] On the other hand, if the encoding device 100 determines that the boundary vertices after the subdivision are the same (Yes in S505), the encoding device 100 indicates in the metadata that the boundary vertices after the subdivision are the same for each submesh (S507). That is, the encoding device 100 includes information indicating that the boundary vertices are the same in the metadata. In other words, the encoding device 100 generates metadata that also includes information indicating that the boundary vertices are the same.

[0423] Furthermore, if the encoding device 100 determines that the boundaries between the submeshes do not overlap (No in S503), it indicates in the metadata that the boundaries between the submeshes do not overlap (S508). That is, the encoding device 100 includes information indicating that the boundaries between the submeshes do not overlap in the metadata. In other words, the encoding device 100 generates metadata including information indicating that the boundaries between the submeshes do not overlap.

[0424] The encoding device 100 may further determine how each submesh overlaps, and indicate the determination result in the metadata.

[0425] For example, the encoding device 100 checks the number of subdivisions of two or more submeshes, determines whether they are identical, and sets a flag indicating the determination result. For example, the encoding device 100 generates a flag (flag information) indicating whether they are identical or not based on the determination result.

[0426] For example, the encoding device 100 checks the number of subdivisions of each of two or more submeshes. For example, if the number of subdivisions of two or more submeshes is different, the encoding device 100 sets a flag (modification flag) to ON. On the other hand, for example, if the number of subdivisions of two or more submeshes is the same, the encoding device 100 sets the flag to OFF. That is, for example, the flag indicates that the number of subdivisions between the two or more submeshes is different. In other words, it can also be said to be a flag indicating that the vertices at the boundary edges of the submeshes after subdivision are not identical or may not be identical.

[0427] Furthermore, for example, the encoding device 100 compares the coordinates of all boundary vertices after subdivision of two or more submeshes, determines whether the boundary vertices are identical, and sets a flag based on the determination result.

[0428] Furthermore, for example, the encoding device 100 compares the coordinates of the boundary vertices of two or more submeshes and determines the value of the flag based on the comparison result. For example, the encoding device 100 sets the flag to ON if there is at least one boundary vertex that is not common to two or more submeshes. On the other hand, for example, the encoding device 100 sets the flag to OFF if all the boundary vertices of two or more submeshes are the same. In other words, the flag indicates whether the boundary vertices of two or more submeshes after subdivision n will be the same or not.

[0429] This flag may be set for the entire mesh frame. If the flag is set for the entire mesh frame, the boundary vertex correction process applies to the boundary edges of all sub-meshes associated with the current frame. If the flag is not set for the entire mesh frame, the flag may be set per sub-mesh. If the flag is set per sub-mesh, the correction process applies to the sub-mesh associated with the current sub-mesh ID.

[0430] [Method of Signaling a Modification Method] In addition to the flag, the encoding device 100 may determine how the decoding device 200 modifies a vertex (boundary vertex) and signal the type of the vertex modification method. The decoding device 200 may modify the vertex based on the type of modification method.

[0431] FIG. 65 is a flowchart showing the process of determining whether sub-meshes overlap, which is executed by the decoding device 200 according to the embodiment.

[0432] First, the decoding device 200 analyzes the metadata to determine whether or not the boundaries between the sub-meshes overlap (S511).

[0433] If the decoding device 200 determines that the boundaries between submeshes overlap (Yes in S512), it determines whether the decoding device 200 or an application for performing decoding-related processing will perform processing related to the overlap (processing of overlapping submeshes) (S513).

[0434] When the decoding device 200 or the application determines to process the overlapping submeshes (Yes in S513), the decoding device 200 processes the overlapping submeshes (S514). As the process of the overlapping submeshes, the decoding device 200 performs, for example, subdivision, detects boundaries between the submeshes (specifically, boundary edges), and performs predetermined processing on the boundaries between the submeshes after the subdivision, such as modifying boundary vertices.

[0435] On the other hand, for example, if the decoding device 200 determines that the boundaries between submeshes do not overlap (No in S512), or if the decoding device 200 or the application determines that it will not process the overlapping submeshes (No in S513), it does not process the overlapping submeshes, i.e., does not process the overlapping submeshes, that is, does not process the submeshes when they overlap, and terminates the processing.

[0436] 66 is a flowchart showing the vertex modification process executed by the decoding device 200 according to the embodiment. Specifically, FIG. 66 shows a specific example of the process of step S514.

[0437] For example, before starting the process shown in Fig. 66, the decoding device 200 decodes and analyzes data and metadata from the bitstream to subdivide multiple sub-meshes. Then, the decoding device 200 executes the process shown in Fig. 66.

[0438] First, the decoding device 200 decodes two or more sub-meshes, flag information (flag), and method information (method) from a bitstream (encoded bitstream) (S521).

[0439] The method information is information that indicates the method of modifying the boundary vertex (the manner of modification or the type of modification method).

[0440] Next, the decoding device 200 determines whether the flag information indicates that a correction is to be applied (S522), that is, the decoding device 200 determines whether the flag is on or off.

[0441] If the flag information indicates that a correction is to be applied (Yes in S522), the decoding device 200 determines whether the method information indicates that a new vertex is to be added (S523). Specifically, the decoding device 200 determines whether the method information indicates that a new vertex is to be added to a boundary edge in at least one of the two or more submeshes. In this way, the decoding device 200 executes the processing indicated by the method information on the two or more submeshes based on the method information.

[0442] For example, when the decoding device 200 determines that the method information indicates that a new vertex is to be added (Yes in S523), the decoding device 200 introduces one or more vertices to the boundary edge based on the method information (S524). Specifically, the decoding device 200 adds one or more vertices to the boundary edge of at least one of the two or more submeshes based on the method information. In this way, the decoding device 200 applies the boundary vertex modification to the two or more submeshes based on the method information.

[0443] On the other hand, for example, when the decoding device 200 determines that the method information does not indicate that a new vertex is to be added (No in S523), the decoding device 200 removes at least one of the one or more vertices in the boundary edge based on the method information (S525). Specifically, the decoding device 200 removes at least one of the one or more vertices added (generated) by the subdivision in the boundary edge of at least one of the two or more submeshes based on the method information.

[0444] If the flag information does not indicate that modification is to be applied (No in S522), the decoding device 200 ends the process without modifying the vertices on the boundary edges.

[0445] Furthermore, the decoding device 200 may or may not execute processing based on the method information.

[0446] Furthermore, the decoding device 200 may determine whether to add or delete a vertex at a boundary edge by a method other than the above, such as by analyzing the number of times a submesh is sub-divided.

[0447] As described above, in the third aspect, flag information may be decoded from the bitstream, and the flag information indicates whether to modify the number of vertices at boundary edges when two or more sub-meshes are combined.

[0448] For example, the decoding device 200 applies the subdivision process to all edges of a given submesh according to a specified number of subdivisions.

[0449] Furthermore, for example, the decoding device 200 adds or deletes a vertex to or from a submesh when the decoded flag information indicates that the number of vertices on the boundary edge is to be equalized.

[0450] For example, the flag information is signaled as an SEI message. After completing the subdivision process and combining two or more adjacent sub-meshes, the decoding device 200 adjusts (modifies) the number of vertices. When the number of vertices is changed, the topology of each sub-mesh is updated accordingly.

[0451] It should be noted that the triangles in the sub-mesh may be reconstructed after the vertices in the boundary edges are modified (added or removed).

[0452] 67 is a diagram illustrating a specific example of a vertex modification process according to an embodiment of the present invention, specifically, a diagram illustrating a process of adding a vertex to a boundary edge based on method information.

[0453] Note that whether to add or delete a vertex may be determined by any method, such as by the decoding device 200 analyzing the number of times a submesh is sub-divided, without using the method information.

[0454] For example, the decoding device 200 adds (generates) new vertices to the boundary edges by copying the coordinates of the vertices from the adjacent sub-meshes.

[0455] In this example, we will use submesh 1 having vertices A, B, C, D, and E, and submesh 2 having vertices B1, C1, X, Y, and Z. Submesh 1 and submesh 2 are adjacent submeshes in the base mesh. Vertices B and B1 are the same vertex in the base mesh, and vertices C and C1 are the same vertex in the base mesh. The edge connecting vertices B and C in submesh 1 and the edge connecting vertices B1 and C1 in submesh 2 are boundary edges.

[0456] Note that an internal edge is an edge that is not shared (not overlapped) with other sub-meshes.

[0457] In this example, submesh 1 is subdivided zero times, and submesh 2 is subdivided once. For example, the decoding device 200 generates submesh 1 shown in FIG. 67(b) by subdividing submesh 1 shown in FIG. 67(a) zero times. In other words, if the number of subdivisions is zero, the decoding device 200 does not subdivide submesh 1 shown in FIG. 67(a), as shown in FIG. 67(b). For example, the decoding device 200 generates submesh 2 shown in FIG. 67(f) by subdividing submesh 2 shown in FIG. 67(e) once. As shown in FIG. 67(f), new vertices Q1, R, S, T, U, V, and W are added (generated) to submesh 2 by subdivision, and new edges (additional edges) connecting these vertices are added (generated). In this way, the decoding device 200 creates triangles within submesh 2 by subdividing submesh 2.

[0458] For example, when the flag information indicates that the vertex is not to be modified (i.e., the flag is off or flag=no), the decoding device 200 does not perform processing to modify the vertex on the submesh 1 shown in (b) of Figure 67, as shown in (c) of Figure 67. On the other hand, when the flag information indicates that the vertex is to be modified (i.e., the flag is on (flag=ON) or flag=yes) and the method information indicates that a vertex is to be added (i.e., method=ADD_VERTEX), the decoding device 200 adds (generates) a new vertex Q to the submesh 1, as shown in (d) of Figure 67.

[0459] A specific example of the method information (method) will be described later (see, for example, FIG. 75).

[0460] Furthermore, the decoding device 200 uses the new vertex Q to create a triangle included in the submesh 1. Specifically, the decoding device 200 uses the new vertex Q to divide the triangle ABC into two triangles, the triangle BAQ and the triangle QAC.

[0461] For example, the decoding device 200 compares submesh 1 and submesh 2, and adds vertex Q corresponding to vertex Q1 to the boundary edges in submesh 1.

[0462] For example, the encoding device 100 determines whether a vertex on a boundary edge, which is an edge shared by a first submesh (e.g., submesh 1) and a second submesh (e.g., submesh 2), is the same in a first mesh formed by subdividing the first submesh a first number of times and a second mesh formed by subdividing the second submesh a second number of times. Furthermore, based on the determination result, the encoding device 100 encodes, into the bitstream, combining information related to the process of combining the first mesh and the second mesh, such as flag information and / or method information. The flag information is, for example, information indicating whether or not the vertices (specifically, the number of vertices) on the boundary edge of at least one of the first mesh and the second mesh need to be modified. The method information is information indicating how to modify the vertices on the boundary edge of at least one of the first mesh and the second mesh (modification method). In this example, the encoding device 100 encodes, into the bitstream, method information including information (ADD_VERTEX) indicating the addition of a vertex to the boundary edge of at least one of the first mesh and the second mesh. The decoding device 200 decodes such connection information from the bitstream and determines, based on the decoded connection information, whether the vertices at the boundary edges of a first mesh formed by subdividing a first submesh a first number of times are the same as those of a second mesh formed by subdividing a second submesh a second number of times. Here, for example, if the decoding device 200 determines that the vertices are not the same, it modifies the vertices at at least one of the boundary edges based on the method information included in the connection information, modifies a triangle included in at least one of the meshes based on the modified vertices, and combines the first mesh and the second mesh with the modified triangle included in at least one of the meshes.

[0463] 68 is a diagram illustrating a specific example of a vertex modification process according to an embodiment of the present invention, specifically, a diagram for explaining a process of removing a vertex from a boundary edge based on method information.

[0464] In this example, similar to the example shown in FIG. 67, sub-meshes 1 and 2 are used for explanation.

[0465] Also, in this example, submesh 1 is subdivided only once, and submesh 2 is subdivided zero times. For example, the decoding device 200 generates submesh 1 shown in FIG. 68(b) by subdividing submesh 1 shown in FIG. 68(a) only once. As shown in FIG. 68(b), new vertices Q, R, S, T, U, V, and W are added (generated) to submesh 1 by subdivision, and new edges (additional edges) connecting these vertices are added (generated). In this way, the decoding device 200 creates triangles within submesh 1 by subdividing submesh 1. Also, for example, the decoding device 200 generates submesh 2 shown in FIG. 68(f) by subdividing submesh 2 shown in FIG. 68(e) zero times. In other words, if the number of subdivisions is zero, the decoding device 200 does not subdivide submesh 1 shown in FIG. 68(b), as shown in FIG. 68(f).

[0466] For example, when the flag information indicates that the vertex is not to be modified (i.e., the flag is off or flag=no), the decoding device 200 does not perform processing to modify the vertex on the submesh 1 shown in (b) of FIG. 68, as shown in (c) of FIG. 68. On the other hand, when the flag information indicates that the vertex is to be modified (i.e., flag=ON or flag=yes) and the method information indicates that the vertex is to be removed (i.e., method=REMOVE_VERTEX), the decoding device 200 removes the vertex on the boundary edge. In this example, the decoding device 200 removes the vertex Q to generate the submesh 1 shown in (d) of FIG. 68.

[0467] In this way, for example, the decoding device 200 removes vertices on boundary edges by merging two or more vertices into one. In this example, the decoding device 200 removes vertex Q, which does not have a corresponding vertex in submesh 2, from among the multiple vertices (vertices B, C, and Q) on the boundary edge of submesh 1.

[0468] Furthermore, the decoding device 200 recreates a triangle included in submesh 1 based on the removed vertex Q. That is, since vertex Q has been removed, the decoding device 200 modifies submesh 1 so that the edge connected to vertex Q is deleted and a new triangle is created within submesh 1. Specifically, as shown in (d) of Figure 68, the decoding device 200 creates triangles UTB and BTC by adding an edge connecting vertex B and vertex T. In this way, for example, triangles UQT, BUQ, and QTC are replaced with triangles UTB and BTC.

[0469] For example, the decoding device 200 compares submesh 1 with submesh 2 to remove vertex Q, which does not have a corresponding vertex in submesh 2, from the boundary edges in submesh 1.

[0470] In addition, in this example, the encoding device 100 encodes method information into the bitstream, including information (REMOVE_VERTEX) indicating the removal of at least one of one or more vertices on the boundary edge of at least one of the first mesh and the second mesh.

[0471] FIG. 69 is a diagram showing a specific example of a location where flag information according to an embodiment is stored.

[0472] As shown in Figure 69, for example, flag information is signaled within an SEI message in the bitstream.

[0473] Note that the flag information may be signaled in the header of the bitstream.

[0474] FIG. 70 is a diagram illustrating an example of syntax in which flag information related to an embodiment is signaled.

[0475] As shown in Figure 70, the value indicated by the flag information (flag) is, for example, a binary value. The value indicated by the flag information is signaled at the frame level. For example, when the value indicated by the flag information is equal to 1, the flag information indicates that the modification process is applied to all sub-meshes related to the current frame. On the other hand, when the value indicated by the flag information is equal to 0, the flag information indicates that the modification process is not applied. If the flag information does not exist, the value at the location where the value of the flag information is signaled is set to 0.

[0476] FIG. 71 is a diagram illustrating another example of syntax in which flag information related to an embodiment is signaled.

[0477] Flag information may be signaled for each submesh, as shown in Figure 71. When flag (flag[i]) is equal to 1, it indicates that the modification process is applied to the submesh associated with submesh_id. On the other hand, when flag[i] is equal to 0, it indicates that the modification process is not applied to the submesh associated with submesh_id. If flag[i] is not present, flag[i] is set to 0.

[0478] Although the flag information indicates whether or not to apply the correction process, the flag may indicate whether or not the vertices of the boundary edges match. The decoding device 200 may determine whether or not to apply the correction process based on the flag.

[0479] Furthermore, the encoding device 100 may not transmit the submesh_id. In this case, it may be specified that the loop order is determined based on the submesh_id.

[0480] FIG. 72 is a diagram showing a specific example of a location where method information according to the embodiment is stored.

[0481] As shown in FIG. 72, for example, method information (method) is signaled in an SEI message in the bitstream.

[0482] Note that the method information may be signaled in the header of the bitstream.

[0483] FIG. 73 is a diagram illustrating an example of syntax in which method information related to an embodiment is signaled.

[0484] As shown in FIG. 73, the method information indicates the identifier of the modification process to be applied to the sub-mesh associated with the current frame.

[0485] FIG. 74 is a diagram illustrating another example of syntax in which method information related to an embodiment is signaled.

[0486] As shown in FIG. 74, the method information may indicate an identifier for a modification process to be applied to the submesh associated with the submesh_id.

[0487] Fig. 75 is a diagram showing a specific example of a method for modifying a vertex according to an embodiment. Specifically, Fig. 75 is a list showing an example of the relationship between an identifier indicated by method information and the type of modification method (the name of the modification method).

[0488] For example, when the method information indicates 0, this indicates that the decoding device 200 does not perform the correction process. Also, when the method information indicates 1, this indicates that the decoding device 200 performs the process of adding a vertex in the correction process. For example, when the method information indicates 2, this indicates that the decoding device 200 performs the process of removing a vertex in the correction process.

[0489] Note that the modification process may be derived by the decoding device 200 without the method information being signaled in the bitstream.

[0490] For example, the decoding device 200 may compare the subdivision counts of the two submeshes, for example by calculating the difference in the subdivision counts between the two submeshes, and determine whether to add or remove a vertex. For example, if submesh 1 has fewer subdivision counts than submesh 2, a modification process that adds a vertex is applied to submesh 1. On the other hand, if submesh 1 has more subdivision counts than submesh 2, a modification process that removes a vertex is applied to submesh 1. Also, for example, if submesh 1 has the same subdivision count as submesh 2, no vertices are added or removed.

[0491] Also, for example, if the method information indicates a process of adding a vertex in the modification process, and if submesh 1 has been sub-divided fewer times than submesh 2, a modification process of adding a vertex is applied to submesh 1. On the other hand, for example, if the method information indicates a process of deleting a vertex in the modification process, and if submesh 1 has been sub-divided more times than submesh 2, a modification process of removing a vertex is applied to submesh 1.

[0492] The example shown in FIG. 69 and the example shown in FIG. 72 may be realized in combination.

[0493] Furthermore, the example shown in FIG. 70, the example shown in FIG. 71, the example shown in FIG. 73, and the example shown in FIG. 74 may be realized in combination.

[0494] Next, an example of a plurality of sub-meshes and adjacent sub-meshes will be described.

[0495] FIG. 76 is a diagram illustrating an example of the relationship between the number of subdivisions of a plurality of submeshes and flag information according to an embodiment.

[0496] In the example shown in Figure 76, submesh 3 is adjacent to submesh 1 and submesh 2, and their boundaries overlap. Furthermore, submesh 1 has been subdivision once (iterations), submesh 2 has been subdivision twice, and submesh 3 has been subdivision twice. Therefore, the vertices at the boundary edge shared by submeshes 1 and 3 are likely to be different between the vertices generated by subdividing submesh 1 and the vertices generated by subdividing submesh 3. Furthermore, the vertices at the boundary edge shared by submeshes 2 and 3 are likely to be the same between the vertices generated by subdividing submesh 2 and the vertices generated by subdividing submesh 3.

[0497] For example, the encoding device 100 sets flag information corresponding to submesh 1 to flag=ON because submesh 1 shares a boundary edge only with submesh 3 and the subdivision counts are not the same.

[0498] Also, for example, since submesh 2 shares a boundary edge only with submesh 3 and the subdivision count is the same, the encoding device 100 sets flag information corresponding to submesh 2 to flag=OFF.

[0499] Submesh 3 also shares boundary edges with submeshes 1 and 2. The number of subdivisions at the boundary edges of submesh 3 is the same as that of submesh 2, but is not the same as that of submesh 1.

[0500] For example, the encoding device 100 sets flag = ON if the number of subdivisions of any one of multiple submeshes with overlapping boundaries (in other words, multiple submeshes that share a boundary edge) is different, or if the vertices after subdivision of the submeshes are not identical.

[0501] FIG. 77 is a diagram for explaining another example of the relationship between the number of subdivisions of a plurality of submeshes and flag information according to an embodiment.

[0502] In the example shown in Figure 77, submesh 3 is adjacent to submesh 1 and submesh 2, and their boundaries overlap. Furthermore, submesh 1 has been subdivision once (iterations), submesh 2 has been subdivision three times, and submesh 3 has been subdivision twice. Therefore, the vertices at the boundary edge shared by submeshes 1 and 3 are likely to be different between the vertices generated by subdivision of submesh 1 and the vertices generated by subdivision of submesh 3. Furthermore, the vertices at the boundary edge shared by submeshes 2 and 3 are likely to be different between the vertices generated by subdivision of submesh 2 and the vertices generated by subdivision of submesh 3.

[0503] For example, the encoding device 100 sets the flag corresponding to submesh 1 to ON because submesh 1 shares a boundary only with submesh 3 and the subdivision counts are not the same.

[0504] Furthermore, for example, the encoding device 100 sets the flag corresponding to submesh 2 to ON because submesh 2 shares a boundary edge only with submesh 3 and the subdivision counts are not the same.

[0505] Also, for example, the encoding device 100 sets the flag corresponding to submesh 3 to ON because submesh 3 shares boundary edges with submesh 1 and submesh 2 and has the same number of subdivisions as either submesh 1 or submesh 2.

[0506] When the encoding device 100 indicates method information (method) in metadata, the encoding device 100 may indicate the method information for each submesh by indicating information on a plurality of submeshes adjacent to one submesh.

[0507] For example, in the loop for submesh 3 in the syntax, the encoding device 100 may indicate that flag=ON and method=add corresponding to submesh 1 and flag=ON and method=remove corresponding to submesh 2.

[0508] As described above, for example, the encoding device 100 determines whether the vertices of a boundary edge, which is an edge shared by a first submesh (e.g., submesh 1) and a second submesh (e.g., submesh 2), are the same in a first mesh formed by subdividing the first submesh a first number of times and a second mesh formed by subdividing the second submesh a second number of times. Here, for example, if the first and second numbers are the same, the encoding device 100 determines that the vertices of the boundary edge are the same in the first mesh and the second mesh; and if the first and second numbers are not the same, the encoding device 100 determines that the vertices of the boundary edge are not the same in the first mesh and the second mesh. Furthermore, for example, if the encoding device 100 determines that the vertices of the boundary edge are not the same in the first mesh and the second mesh, it encodes the connection information including the method information into a bitstream. For example, if the encoding device 100 determines that the vertices at the boundary edge are the same in the first mesh and the second mesh, the encoding device 100 may generate a bitstream that does not include method information.

[0509] In addition, the encoding device 100 may indicate in the bitstream (specifically, metadata included in the bitstream) a flag indicating whether or not the boundaries between submeshes overlap, and if the boundaries between submeshes overlap, the type of overlap.

[0510] Figure 78 is a diagram for explaining an example of syntax in which information indicating the type of overlap between multiple sub-meshes in an embodiment is signaled.

[0511] The submesh_duplicate_flag is a flag indicating whether a submesh constituting a mesh (e.g., a base mesh) overlaps with another submesh. For example, if submesh_duplicate_flag = 1, it indicates that there is overlap. On the other hand, if submesh_duplicate_flag = 0, it indicates that there is no overlap.

[0512] submesh_duplicate_type is information indicating the type of submesh duplication (how the submeshes are duplicated). For example, when submesh_duplicate_type = 0, submesh_duplicate_type indicates that the boundary edges of the outer periphery of the submesh and the vertices of the outer periphery of the submesh are duplicated, and the edges, vertices, and faces located inside the periphery are not duplicated. On the other hand, when submesh_duplicate_type = 1, submesh_duplicate_type indicates that not only the outer periphery of the submesh but also the edges, vertices, and faces located inside the periphery are duplicated.

[0513] In addition, the encoding device 100 may indicate in the bitstream (more specifically, the metadata included in the bitstream) that there is a possibility of submeshes overlapping, instead of a flag indicating that submeshes overlap.

[0514] Furthermore, such a flag (parameter set) may be stored in a sequence-level parameter set, a frame-level parameter set, or both. If stored in both parameter sets, for example, the encoding device 100 overwrites the sequence-level parameter set with the frame-level parameter set.

[0515] Such parameter sets may also be indicated in the SEI.

[0516] Furthermore, it may be specified that the encoding device 100 signals a flag indicating whether boundary vertices after subdivision at boundary edges between submeshes are the same or whether boundary vertices after subdivision of a submesh are adjusted when submesh_duplicate_flag is ON, and does not signal when submesh_duplicate_flag is OFF. It may also be specified that the encoding device 100 must not signal when submesh_duplicate_flag is OFF. In this case, if the flag is signaled, it may be specified that the flag set by the signaling is invalid.

[0517] Furthermore, if submesh_duplicate_flag is ON, the decoding device 200 may execute a process of determining whether or not to modify the boundary vertices after subdivision, and may determine not to modify them if the flag is OFF.

[0518] According to this, the encoding device 100 includes the above flag in the higher-level metadata, so that the decoding device 200 can determine whether or not there is a possibility that a submesh overlaps with another submesh. Therefore, the decoding device 200 can determine, from the flag, whether or not a predetermined process (correction process) needs to be performed on the overlapping boundary edge.

[0519] If the decoding device 200 does not need to perform such processing, the processing of detecting boundary information and adjusting boundary vertices after boundary subdivision is not required, and the processing of receiving, decoding, and analyzing related information is not required, thereby reducing the amount of processing.

[0520] In addition, when the encoding device 100 transmits only the above flag, the decoding device 200 determines that there is a possibility of overlapping boundaries, and further analyzes the number of subdivisions for each submesh to determine whether the boundary vertices after subdivision will be identical, and may determine whether the boundary vertices need to be modified.

[0521] [Effect of the Third Aspect] With the configuration according to the third aspect, the subjective quality of a full mesh is improved by merging sub-meshes that have been coded using different sub-division counts or sub-division methods (sub-division schemes).

[0522] [Combination with Other Aspects] The encoding device 100 according to this aspect can be implemented by combining at least a part of other aspects of the present disclosure. Furthermore, this aspect may be implemented by combining a part of the processing shown in any of the flowcharts according to this aspect, a part of the configuration of any of the devices, and / or a part of the syntax, etc., with other aspects.

[0523] The above processing of the encoding device 100 may also be executed in the decoding device 200 in the same manner.

[0524] Furthermore, not all of the components described in this embodiment are always necessary, and only some of the components of the third embodiment may be included.

[0525] <Hierarchical Structure Based on Subdivision> Next, the Level of Detail (LoD) hierarchy in subdivision will be described.

[0526] FIG. 79 is a diagram for explaining LoD according to an embodiment.

[0527] The triangle ABC shown in Figure 79 is a mesh before subdivision. Specifically, the triangle ABC is a base mesh. The triangle ABC (i.e., the base mesh) is defined as having an LoD layer of 0. In other words, each vertex of the triangle ABC is defined as a vertex belonging to LoD0.

[0528] Next, each vertex generated after the base mesh is subdivided once is referred to as LoD1. In the example shown in Figure 79, the vertices indicated by squares (vertices D, E, and F) belong to LoD1.

[0529] Furthermore, the vertex generated after the second subdivision is designated as LoD2, which in the example shown in Fig. 79 is the vertex indicated by a triangle (such as vertex G).

[0530] For example, vertex F of LoD1 is generated using vertices A and C of LoD0. Here, vertices A and C are sometimes called parent vertices of vertex F. A parent vertex is a vertex that belongs to the next higher LoD hierarchy. In other words, for example, vertices A and C that belong to LoD0 are parent vertices of vertex F that belongs to LoD1.

[0531] Also, for example, vertex A belonging to LoD0 and vertex F belonging to LoD1 are parent vertices of vertex G. In other words, vertex G is a child vertex of vertex A belonging to LoD0 and vertex F belonging to LoD1. In this way, it is sufficient that at least one of the parent vertices is one level higher in the LoD hierarchy than its child vertex.

[0532] Here, the LoD layer may also be referred to as an LoD index. For example, the LoD index of LoD0 is 0.

[0533] When the number of subdivisions between submeshes is different, i.e., when the number of LoD hierarchies is different between submeshes, the above method shows an example in which, when combining (merging) vertices at the overlapping boundary of multiple submeshes, vertices are generated and combined by adding and / or removing vertices based on information on whether to add or remove (delete) a vertex.

[0534] Another method is to calculate the distance between adjacent vertices (adjacent points) when joining vertices at the overlapping boundary of multiple sub-meshes, and join vertices whose calculated distance is within a predetermined threshold.

[0535] Although it has been described above that points whose distance to adjacent points is within a predetermined threshold are joined, it is also possible to associate a vertex of one sub-mesh on the overlapping boundary with the vertex of the other sub-mesh that is closest in distance, and then join the two associated vertices. The predetermined threshold may be set arbitrarily and is not particularly limited.

[0536] However, these methods may generate incorrect vertices, and the incorrect vertices may be connected, resulting in densely packed points or holes in the mesh, which may result in an improperly shaped mesh.

[0537] Therefore, in order to solve the above problem, in the following, the encoding device 100 and the decoding device 200 perform processing that takes into account the LoD hierarchy of the vertices when connecting vertices (also called boundary vertices) at the boundaries between multiple sub-meshes.

[0538] In the following, an example of a method for combining a plurality of divided sub-meshes decoded in a sub-mesh combining unit included in a decoding device in this embodiment will be described.

[0539] For example, a subdivided boundary vertex is joined only if its "parent vertex" is already joined.

[0540] Also, for example, the encoding device 100 signals a flag in the bitstream indicating that LoD boundary vertices that do not have corresponding LoDs should be removed from adjacent sub-meshes, or that new corresponding boundary vertices should be artificially generated from adjacent sub-meshes.

[0541] Also, for example, the encoding device 100 and / or the decoding device 200 create new triangles in the 3D geometry and UV maps of adjacent sub-meshes.

[0542] Furthermore, for example, the decoding device 200 combines vertices with the same LoD hierarchy on overlapping edges between adjacent submeshes that are to be combined (i.e., vertices that are to be combined with a certain vertex), and does not combine points with different LoD hierarchy.

[0543] Furthermore, for example, if two vertices to be joined are in the same LoD hierarchy and the distance between the vertices is within a predetermined threshold, the two vertices may be joined.

[0544] Furthermore, for example, when joining two vertices that are far apart, the encoding device 100 and the decoding device 200 calculate the position coordinates of the joining point (the vertex generated by joining the two vertices) using a predetermined method.

[0545] Furthermore, for example, the calculated position coordinates (position coordinates of the connecting point) may be the midpoint or center of gravity of the two points, or the position coordinates may be made closer to one of the vertices by a weighting coefficient.

[0546] Also, for example, when a vertex in a higher hierarchical layer (i.e., a layer with a smaller LoD index than a certain vertex) is combined, the encoding device 100 and the decoding device 200 determine whether or not to combine a vertex in a lower hierarchical layer (i.e., a layer with a larger LoD index) than the higher hierarchical layer.

[0547] Furthermore, for example, when the number of LoD layers differs between submeshes, the decoding device 200 generates a vertex (e.g., performs further subdivision) at the boundary of one of the submeshes (the boundary edge of the submesh) to match the number of vertices (the number of LoD layers) between the submeshes, and then combines the vertices.

[0548] Furthermore, for example, the decoding device 200 removes a vertex at one of the submesh boundaries to match the number of vertices and the number of layers of the vertices between the submeshes, and then combines the vertices.

[0549] Whether to add or remove a vertex may be determined based on information signaled in the bitstream, or may be determined by the decoding device 200 in a predetermined manner.

[0550] According to the above conditions, it may be possible to solve the problem of meshes not being generated with an appropriate shape, such as when vertices are generated with high density or when holes are created in the mesh.

[0551] Figure 80 is a flow diagram illustrating vertex merging processing in the decoding device 200 according to an embodiment. Note that the merging processing is also referred to as merging or zipping. The flowchart shown in Figure 80 illustrates a specific example of processing in which, for example, the decoding device 200 acquires boundary vertices (specifically, coded information related to the boundary vertices, such as position information) from a bitstream, decodes the boundary vertices of submeshes (specifically, the coded information related to the boundary vertices), and further subdivides two submeshes and merges the two resulting submeshes.

[0552] First, the decoding device 200 obtains a flag (flag information) indicating the addition or removal of a LoD vertex that does not correspond to an adjacent submesh from the bitstream (S601). Specifically, the decoding device 200 obtains the coded flag information from the bitstream and decodes it.

[0553] Next, the decoding device 200 determines whether or not there is a corresponding LoD index for each boundary vertex of one submesh in the adjacent submesh (the other submesh) (S602). That is, the decoding device 200 determines whether or not the number of LoD layers of the two submeshes is the same.

[0554] When the decoding device 200 determines that there is a corresponding LoD index (Yes in S602), it combines the boundary vertices (more specifically, boundary vertices with the same LoD) of one submesh and the other adjacent submesh (S603). In other words, since the number of LoD layers (more specifically, the number of subdivisions) is the same, the decoding device 200 combines the vertices without adding or removing vertices.

[0555] On the other hand, if it is determined that there is no corresponding LoD index (No in S602), the decoding device 200 determines whether the decoded flag indicates that a new vertex is to be added (S604).

[0556] When the decoding device 200 determines that the decoded flag indicates that a new vertex is to be added (Yes in S604), for each boundary vertex, if the parent vertex of the boundary vertex is connected to one or more vertices of the adjacent sub-meshes, the decoding device 200 generates a new boundary vertex in the adjacent sub-meshes (S605). Note that, if the parent vertex of the boundary vertex is not connected to one or more vertices of the adjacent sub-meshes, the decoding device 200 does not need to generate a new boundary vertex in the adjacent sub-meshes.

[0557] Next, the decoding device 200 combines the boundary vertices of one submesh generated based on the flag with the vertices generated by the subdivision of the other submesh (S606).

[0558] On the other hand, if the decoding device 200 determines that the decoded flag does not indicate the addition of a new vertex (No in S604), it regards the flag as indicating the removal of a vertex, and removes some of the boundary vertices of the current submesh using a predetermined method (S607). Note that the predetermined method may be determined arbitrarily and is not particularly limited.

[0559] Next, the decoding device 200 combines boundary vertices without using the removed vertex (S608).

[0560] At the beginning of the process, the decoding device 200 determines whether the LoD layer numbers between the submeshes match, and if they match, the vertices may be joined without decoding and analyzing the flags.

[0561] In addition, the decoding device 200 may first determine whether the number of LoD layers is 0 (base mesh), skip this process if it is the base mesh, and perform this process if the number of LoD layers is 1 or more.

[0562] Furthermore, for example, when the numbers of layers match, the decoding device 200 combines the vertices without adding or removing vertices because a corresponding layer exists. On the other hand, when the numbers of layers do not match, for example, the decoding device 200 adds or removes vertices from one of the two submeshes to be combined, matches the number of vertices of the boundary edges, and then combines the vertices.

[0563] Next, a method for connecting vertices between sub-meshes with LoD constraints will be described.

[0564] FIG. 81 is an explanatory diagram showing an example of vertex merging processing according to an embodiment. Specifically, FIG. 81 shows an example of merging (merging) vertices of the same LoD in two different submeshes, thereby merging (zipping) the boundary between two submeshes. Note that vertices are located at both ends of each line segment between submeshes 1 and 2 shown in FIG. 81. Vertices belonging to LoD0 are located at both ends of the line segment shown by solid lines, vertices belonging to LoD1 are located at both ends of the line segment shown by dashed lines, and vertices belonging to LoD2 are located at both ends of the line segment shown by dotted lines. The same applies to the following FIGS. 82 to 84.

[0565] Figure 81 shows an example of combining two submeshes with different numbers of subdivisions (specifically, submesh 1, which has been subdivision twice, and submesh 2, which has been subdivision once). In the example shown in Figure 81, if there are vertices with the same LoD, the vertices are combined, and if there are no vertices with the same LoD, the vertices are not combined.

[0566] Vertices A, B, C, D, E, and F are vertices of the base mesh of LoD0, vertices X, Y, Z, L, M, and N are vertices of LoD1, and the remaining vertices are vertices of LoD2.

[0567] Also, each vertex is a vertex after being corrected by a displacement vector after subdivision.

[0568] When the boundaries of two sub-meshes are joined, vertices with the same LoD in different sub-meshes are joined.

[0569] In this example, the vertices of LoD0 and LoD1 of submesh 1 are connected to the vertices of LoD0 and LoD1 of submesh 2, respectively. The coordinates of the connected vertex (connection point) between the two vertices are calculated as a weighted sum, such as their average. For example, the coordinates of the connection point are calculated as M(A, D) = w1 x A + w2 x D.

[0570] Note that M(A, D) is the coordinate of the connection point between vertex A and vertex B, which is generated between vertex A and vertex D. A is the coordinate of vertex A. D is the coordinate of vertex D. w1 is a weighting coefficient by which the coordinate of vertex A is multiplied. w2 is a weighting coefficient by which the coordinate of vertex D is multiplied.

[0571] For example, a vertex of LoD2 of submesh 1 is not connected to a vertex of submesh 2. The position coordinates of the connection point may be determined based on the position coordinates of the two vertices before connection. For example, the position coordinates may be the midpoint of the coordinates of the two vertices, or a weight may be applied to the coordinates of one of the two vertices.

[0572] Furthermore, the coordinates of the connection point may be determined based on the position information of a vertex in a higher layer rather than the two vertices before the connection, or may be determined based on the coordinates of the two vertices before the connection and the vertex in a layer above those two vertices. The same applies to the coordinates (three-dimensional coordinates) of the position (three-dimensional position) of the connection point and the coordinates (two-dimensional coordinates) of the attribute map (UV map) of the connection point. For example, the position coordinates of the connection point may be determined based on the position information of the two vertices before the connection and the vertex in a layer above those two vertices rather than the two vertices before the connection. The coordinates of the attribute map (e.g., attribute information) of the connection point may be determined based on the coordinates (e.g., attribute information) of the attribute map of the connection point rather than the two vertices before the connection and the vertex in a layer above those two vertices.

[0573] This method can be applied to any combination of the number of subdivisions of two submeshes as long as the number of subdivisions of the two submeshes is equal to or greater than 0. For example, the difference in the number of subdivisions of the two submeshes may be equal to or greater than 2, and it can also be applied to combining submeshes with the same difference.

[0574] 82 is an explanatory diagram showing another example of vertex merging processing according to an embodiment. Specifically, FIG. 82 shows an example in which, of two submeshes, vertices with LoDs (non-matching LoDs) in one submesh with a larger number of subdivisions that do not exist in the other submesh are removed, thereby equalizing the number of vertices in the two submeshes at the boundary edge. That is, in this example, the decoding device 200 removes vertices with non-matching LoDs in the two submeshes.

[0575] When the decoding device 200 combines two submeshes with different numbers of subdivisions (specifically, submesh 1 that has been subdivision twice and submesh 2 that has been subdivision once), for example, if there are vertices with the same LoD, the vertices are combined, and if there are no vertices with the same LoD, the vertices are not combined.

[0576] In addition, if there are no vertices with the same LoD, the decoding device 200 removes the vertex with the LoD from the submesh with the most number of subdivisions out of the two submeshes, and reconstructs the mesh by combining the two submeshes.

[0577] After vertices with the same LoD in two different sub-meshes are combined, the method removes vertices with inconsistent LoDs from the sub-meshes. In this example, vertices R and Q on the boundary edge of the inconsistent LoD in sub-mesh 1 are removed. Then, the neighborhood of the removed vertices is remeshed. For example, after vertex R is removed, triangles PAR, PRS, and SRX are replaced with triangles PAS and SAX. Similarly, after vertex Q is removed, triangles TXQ, TQK, and KQB are replaced with triangles TXK and KXB.

[0578] If an attribute map is present, the decoding device 200 also performs similar remeshing on the triangles in the attribute map. The attribute map may be, for example, a UV map. Therefore, the number of vertices on boundary edges is equalized, and artifacts are removed after merging. UV coordinates (two-dimensional UV coordinates) are coordinates that indicate the location of the texture corresponding to the mesh in the UV map (two-dimensional map). One UV coordinate exists for one point. When a vertex is removed, the UV coordinate (information indicating the UV coordinate) is also removed at the same time. Furthermore, information indicating a new mesh generated by removing the vertex includes information indicating a set of three 3D points and their connection relationship, and information indicating the texture located at the three UV coordinates for the three 3D points becomes information indicating the new texture.

[0579] Note that if no vertices are removed when vertices are combined, the original UV coordinates of each vertex may be used as they are. For example, when vertices A and D are combined, even if the position coordinates of vertices A and D are shifted, the UV coordinates of vertices A and D may be used as they are. Furthermore, the UV coordinates may be corrected using a predetermined method.

[0580] Furthermore, even when two vertices are combined, the decoding device 200 may retain the position information and UV coordinate information of each of the two vertices without combining the information of the two vertices. In other words, the decoding device 200 may treat the combined point as vertex A and vertex D, which have the same position information but different UV coordinate information.

[0581] 83 is an explanatory diagram showing another example of vertex merging processing according to an embodiment. Specifically, Fig. 83 shows an example in which the number of vertices at boundary edges is equalized by generating a vertex in one of two submeshes that has been divided into submeshes less times. In other words, in Fig. 83, a vertex is generated for a non-match LoD.

[0582] When the decoding device 200 combines two submeshes with different numbers of subdivisions (specifically, submesh 1 that has been subdivision twice and submesh 2 that has been subdivision once), for example, if there are vertices with the same LoD, the vertices are combined, and if there are no vertices with the same LoD, the vertices are not combined.

[0583] Furthermore, if there are no vertices with the same LoD, the decoding device 200 adds the vertex with the LoD to a submesh with a smaller number of subdivisions, and reconstructs the mesh by combining the two submeshes.

[0584] After vertices with the same LoD from different sub-meshes are combined, the method generates new vertices for the inconsistent LoDs in the sub-meshes with fewer subdivisions, for example, new vertices G and H are generated on the boundary edges of sub-mesh 2 to correspond to vertices R and Q of sub-mesh 1.

[0585] As an example, this method first finds the parent vertex of vertex R. In this example, the parent vertices of vertex R are vertex A and vertex X. Also, the vertices of submesh 2 connected to vertex A and vertex X are vertex D and vertex L, respectively.

[0586] Next, in this method, vertex G is generated between vertex D and vertex L of submesh 2. The coordinates of the vertex located between vertex R and vertex G and generated by combining vertex R and vertex G are calculated as a weighted sum, such as the average, of the coordinates of vertex R and vertex G.

[0587] Then, the neighborhood of the generated vertices is remeshed. For example, after vertex G is generated, triangle DML is replaced with triangle DMG, GML. Similarly, after vertex H is generated, triangle LNF is replaced with triangle LNH, HNF. Similar remeshing is also performed on triangles in the attribute map, if one exists. The attribute map is, for example, a UV map. Therefore, the number of vertices on the boundary edges is equalized, and artifacts are removed after merging.

[0588] Furthermore, for example, when generating a vertex, the decoding device 200 generates the position coordinates of the vertex and simultaneously generates the UV coordinates (UV coordinate information) of the texture for the vertex. For example, when generating a vertex G at the midpoint between vertices D and L, the decoding device 200 sets the midpoint between the UV coordinates of vertex D and the UV coordinates of vertex L as the UV coordinates of the new vertex G. A point other than the midpoint may be used to calculate these UV coordinates. There is a possibility that quality may be improved by calculating the UV coordinates using a formula similar to the method for generating the position coordinates of vertices D and L. Furthermore, even if the position of vertex G is corrected when combining vertex G and vertex R, the UV coordinates calculated previously may be used as is without correction. In other words, for example, when generating vertices within the same submesh, the decoding device 200 may generate UV coordinates, but when combining vertices between different submeshes, the UV coordinates may not need to be corrected.

[0589] 84 is an explanatory diagram showing another example of vertex merging processing according to an embodiment. Specifically, FIG. 84 shows an example of zipping the boundary between two partial meshes by merging vertices based on the parent-child relationship of the vertices. That is, FIG. 84 shows an example of vertex merging using the parent-child relationship of the vertices.

[0590] In this example, the border zipper method is used. In the border zipper method, first, LoD0 vertices of different submeshes are connected. As an example, the decoding device 200 finds connecting vertices using the nearest neighbor principle. Next, vertices of higher LoD indices are connected only if their parent vertices are connected to each other. For example, vertex L is connected to vertex X because vertex D and vertex F, which are parents of vertex L, are connected to vertex A and vertex B, which are parents of vertex X. For example, vertex W is the closest vertex to vertex L among the LoD1 vertices of submesh 1, but is not connected to vertex L because the parent vertices of vertex W and vertex L are not connected to each other. Instead, vertex W is connected to vertex K.

[0591] In this way, for example, the decoding device 200 checks whether the parent vertex of the vertex to be combined is combined, and if so, determines the child vertex of the combined parent vertex as the vertex to be combined. This process allows combining without deriving distance information or performing distance-based search processing, thereby reducing the processing load.

[0592] In addition to the above process, the decoding device 200 may search for a distance using a vertex with the same LoD as the vertex to be combined, and determine whether to combine the vertex based on the searched distance using a predetermined method. This increases the processing load, but may enable mesh combination with high quality and accuracy.

[0593] FIG. 85 is a flow diagram illustrating the process of joining vertices for each LoD in the decoding device 200 according to the embodiment.

[0594] First, the decoding device 200 acquires metadata indicating the number of subdivisions of each of the two submeshes from the bitstream, extracts the number of subdivisions of each of the two submeshes from the acquired metadata, and compares them (S611). Here, the number of subdivisions of one of the two submeshes is set to N, and the number of subdivisions of the other submesh is set to M. Both N and M are integers equal to or greater than 0.

[0595] Next, the decoding device 200 subdivides each of the two submeshes based on the extracted number of subdivisions, and connects the LoD0 vertices of the two submeshes (i.e., the vertices of the base meshes of the two submeshes) (S612).

[0596] Next, the decoding device 200 combines the vertices from LoD1 to LoDX of the two submeshes (S613). Here, X is min(N, M). That is, the decoding device 200 performs the combining process up to the vertices of the LoD hierarchy that are common to the two submeshes. In other words, if there is an LoD hierarchy to be combined, the decoding device 200 performs the combining process as is without adding or removing vertices.

[0597] Next, the decoding device 200 combines the vertices from LoDX+1 to LoDY of the two submeshes (S614). Here, Y is max(N, M). That is, the decoding device 200 performs combining processing on vertices of LoD hierarchies that do not exist in one of the two submeshes. Specifically, if there is no LoD hierarchies to be combined, the decoding device 200 adds or removes vertices as described above to perform combining processing. The decoding device 200 determines whether to add or remove vertices by, for example, analyzing metadata (flags), that is, based on the metadata.

[0598] 86 is a flowchart illustrating the LoD0 vertex joining process in the decoding device 200 according to the embodiment. Specifically, FIG. 86 is a flowchart illustrating the details of the process in step S612.

[0599] First, the decoding device 200 searches for one or more LoD0 vertices from among the boundary vertices of one of the two submeshes (the corresponding submesh) to be combined with the other submesh, and determines, among the one or more LoD0 vertices found, the vertex that is closest to the LoD0 vertex of the one submesh as the vertex to be combined with the LoD0 vertex of the one submesh (S621). In other words, the decoding device 200 combines the LoD0 vertices of the two submeshes that are closest in distance.

[0600] Next, the decoding device 200 determines the position coordinates of the joining point based on the position coordinates of the two vertices to be joined (S622).

[0601] Next, the decoding device 200 determines the UV coordinates of the joining point based on the UV coordinates of the two vertices to be joined (S623).

[0602] Next, the decoding device 200 reconstructs a mesh (a mesh formed by joining two sub-meshes) based on the position coordinates and UV coordinates of the joining points (S624).

[0603] The above process is performed similarly for vertices other than LoD0.

[0604] In the above, X is the minimum value of N and M, and Y is the maximum value of N and M, but other values ​​may be used for X and Y.

[0605] In addition, in the processing of LoD1 to LoDX, the joining process (determining the position coordinates of the joining points, shifting the position coordinates of the joining points, determining the UV coordinates of the joining points, and reconstructing the mesh) may be performed for each LoD, or the joining process may be performed for multiple LoDs together.

[0606] In addition, in the processing of LoDX+1 to LoDY, combining processing (adding vertices, removing vertices, determining the position coordinates of connecting points, shifting the position coordinates of connecting points, determining the UV coordinates of connecting points, and reconstructing the mesh) may be performed for each LoD, or combining processing may be performed for multiple LoDs at once.

[0607] Additionally, the subdivision described above as being performed by the decoding device 200 may instead be performed by the encoding device 100. Additionally, the encoding device 100 may or may not perform the subdivision of the base mesh. Additionally, the encoding device 100 may or may not combine the sub-divided sub-meshes after subdividing the sub-meshes.

[0608] The method information may be coded into the bitstream, for example, by signaling it within an SEI message in the bitstream, as shown in FIG.

[0609] 87 is a diagram illustrating another example of syntax in which method information according to an embodiment is signaled. Specifically, FIG. 87 illustrates an example of syntax in which parameters are set in a mesh frame.

[0610] The method information is encoded for each frame using, for example, ue(v) exponential Golomb coding. For example, if the value indicated by the method information (method value) is equal to 0, the method information indicates that no correction of mismatched LoDs is performed. Furthermore, for example, if the value indicated by the method information (method value) is equal to 1, the method information indicates that new vertices are generated on the submesh by performing fewer subdivisions (i.e., adding vertices). Furthermore, for example, if the value indicated by the method information (method value) is equal to 2, the method information indicates that vertices of mismatched LoDs are removed. If the method information is not present, for example, the method value is considered to be 0.

[0611] For example, as shown in Fig. 76, when the method information indicates 0, this indicates that the decoding device 200 does not perform the modification process. Also, when the method information indicates 1, this indicates that the decoding device 200 performs the process of adding a vertex in the modification process. When the method information indicates 2, this indicates that the decoding device 200 performs the process of removing a vertex in the modification process.

[0612] Note that the modification process may be derived by the decoding device 200 without the method information being signaled in the bitstream.

[0613] The method information may also indicate an identifier of a modification operation to be applied to the submesh associated with the current frame, for example, as shown in Figure 75. For example, the method information may indicate an identifier of a modification operation to be applied to the submesh associated with submesh_id.

[0614] Next, variations in the joining method will be described.

[0615] In the following, the midpoint of the two vertices to be joined will be described as an example, but weighting may be applied, or other calculation formulas may be used.

[0616] FIG. 88 is an explanatory diagram showing another example of the vertex joining process according to the embodiment.

[0617] In this example, the decoding device 200 generates vertex I of the LoDN by combining vertex A of the LoDN of submesh 1 with vertex X of the LoDN of submesh 2. The decoding device 200 also generates vertex K of the LoDN by combining vertex C of the LoDN of submesh 1 with vertex Z of the LoDN of submesh 2. The decoding device 200 also generates vertex J by combining vertex B of the LoDN+1 of submesh 1 with vertex Y of submesh 2. Vertices Y and J correspond to the vertices of LoDN+1.

[0618] For example, the decoding device 200 determines the position coordinates of the joining point of LoDN and LoDN+1 using two vertices in the same LoD hierarchy. For example, the decoding device 200 determines the position coordinates of the joining point of LoDN by calculating I = (A + X) / 2 and K = (C + Z) / 2, where I, A, and X are the position coordinates of vertices I, A, and X, respectively.

[0619] Furthermore, the decoding device 200 determines the position coordinates of the vertex J by, for example, calculating J=(B+Y) / 2, where J, B, and Y are the position coordinates of the vertices J, B, and Y, respectively.

[0620] Here, if Y does not exist in submesh 2, the decoding device 200 calculates the position coordinate of vertex Y, for example, based on the position coordinates of the vertices (e.g., the two parent vertices of vertex Y) of submesh 2. For example, the decoding device 200 calculates the position coordinate of vertex Y by calculating Y=(X+Z) / 2.

[0621] The position coordinates of vertex J may be determined based on the position coordinates of a vertex of the LoD (for example, LoDN) in the higher hierarchy than vertex J.

[0622] FIG. 89 is an explanatory diagram showing another example of the vertex joining process according to the embodiment.

[0623] 89 , the decoding device 200 calculates the position coordinates of vertex J by shifting the position coordinates of vertex B without generating vertex Y. Here, the vertex that serves as the starting point of the vector indicating the amount of shift is, for example, the vertex of the submesh that has been sub-divided the most (in this example, submesh 1) out of submeshes 1 and 2.

[0624] First, the decoding device 200 calculates the vector AI = I - A (for example, I is the midpoint between A and X). The decoding device 200 also calculates the vector CK = K - C (for example, K is the midpoint between C and Z). Next, the decoding device 200 uses the vector AI and the vector CK to calculate the shift amount (vector BJ) of the vertex B in order to calculate the position coordinates of the vertex J. Specifically, the decoding device 200 calculates the vector BJ = (vector AI + vector CK) / 2 and J = B + vector BJ.

[0625] According to this method, the decoding device 200 can determine the position coordinates of the joining points in LoDN+1 without performing calculations using the position coordinates of the vertices (vertex X and vertex Z) of LoDN.

[0626] Furthermore, the calculation amount can be reduced by using information on vectors that have already been calculated.

[0627] Furthermore, when comparing the shapes of submesh 1 and submesh 2, which has been sub-divided fewer times than submesh 1, there is a high possibility that the shape of submesh 1 is more accurate than that of submesh 2. Therefore, decoding device 200 sets the submesh with the greater number of sub-divisions as the starting point (here, submesh 1) of the two submeshes to be combined, and shifts the vertices by the vectors determined from the higher LoD layer, thereby making it possible to more accurately represent the shape after the shift.

[0628] Although this embodiment has been described using submeshes as an example of a plurality of divided meshes and connecting the submeshes, it can also be applied to mesh data other than submeshes. For example, this embodiment can also be applied to units obtained by further dividing a submesh (e.g., mesh patches). Furthermore, the method of this embodiment can be applied when the number of subdivisions differs between mesh patches. In this case, information on whether to add or remove a vertex for a boundary edge whose LoD layer number does not match in mesh patch units can be signaled in the bitstream.

[0629] Also, in the above description, the encoding device 100 determines a method for adding or removing a vertex to or from a submesh, stores information indicating the method in metadata and signals it in the bitstream, and the decoding device 200 determines whether to add or remove a vertex based on the metadata. The encoding device 100 may not signal such information in the bitstream, and the decoding device 200 may be predetermined to execute either the addition or removal method. Alternatively, the encoding device 100 and the decoding device 200 may be predetermined to switch between the methods using a predetermined method. Alternatively, the decoding device 200 may determine the predetermined method.

[0630] For example, the decoding device 200 may measure the quality of the mesh after the combination and determine whether to add or remove a vertex based on the measured quality, which may be measured based on geometric quality such as the positions of the vertices, the shape of the mesh, and whether the mesh has holes, as well as the evaluation result of the entire mesh including the texture.

[0631] Furthermore, the decoding device 200 may combine multiple sub-meshes based on the metadata, or may output the metadata to an application without performing the combining process, and the application may combine multiple sub-meshes.

[0632] <Summary of the Disclosure> Next, an example of an overview of the technology obtained from the disclosure of this specification will be given.

[0633] For example, in the decoding method disclosed herein, for multiple submeshes that make up a three-dimensional mesh, position information of vertices included in polygons that make up the submeshes and connection information regarding the connection relationships of the vertices are obtained from the encoded bitstream (i.e., decoding of information), the polygons are generated using the position information and the connection information (i.e., decoding of faces), a determination is made as to whether an edge that makes up the polygon is a boundary of the submesh (i.e., edge condition determination), a division process for the edge is determined based on the determination result (i.e., determination of subdivision process based on the determination result), and the edge is divided (subdivision) using the division process.

[0634] Specific examples and variations of the division process are described below.

[0635] In the division process, a method for dividing the edge (division method) may be specified.

[0636] In the division process, the number of times to divide the edge (number of divisions) may be specified.

[0637] In the division process, a method and number of times to divide the edge (division method and number of divisions) may be specified.

[0638] The division process may be a process of generating a new vertex based on position information of a plurality of vertices that form the edge (another definition of the division process).

[0639] The decoding process may decode parameters specifying the segmentation process from the coded bitstream (segmentation process signaling).

[0640] At least one of the division processes may be defined in advance (predetermined division process).

[0641] The splitting process includes not splitting the edge (no split option).

[0642] Specific examples and modifications of the determination process will be described below.

[0643] In the determination process, it may be determined whether or not one of the sides constituting the polygon to be processed is the boundary of the sub-mesh (determination process for each side).

[0644] The determination process may determine whether or not any of the edges constituting the polygon is a boundary edge of the sub-mesh (determining whether or not the polygon includes a boundary edge).

[0645] The submesh boundary may be an edge that includes at both ends a plurality of vertices that constitute a plurality of submeshes (definition of boundary).

[0646] A specific example of the relationship between the determination result and the division process will be described below.

[0647] In the division process, when the side to be processed is the submesh boundary, the side may be divided using a first division process.

[0648] In the division process, when the side to be processed is not the submesh boundary, the side may be divided using a second division process.

[0649] In the division process, when the polygon includes an edge that is the submesh boundary, the edge that is the submesh boundary may be divided using a first division process, and the edge that is not the submesh boundary may be divided using a second division process.

[0650] In the division process, when the polygon does not include an edge that is the boundary of the sub-mesh, all edges included in the polygon may be divided using a second division process.

[0651] In the division process, when the polygon includes an edge that is the submesh boundary, it may be determined whether the first division process and the second division process have a predetermined relationship, and the division process may be determined based on the determination result. For example, the division process may be switched based on a result of comparing the number of divisions specified by the first division process and the second division process.

[0652] Specific examples of the first division process (boundary division process) and the second division process (non-boundary division process) will be described below.

[0653] The first division process and the second division process may be different processes (different division processes may be selected).

[0654] The first division process may be selected from a first division process group, and the second division process may be selected from a second division process group. The first division process group and the second division process group may include different division processes (selected from a plurality of division processes, with different options).

[0655] The first division process may select the same process common to a plurality of sub-meshes, or the division process may select processes from the same group for a plurality of sub-meshes.

[0656] The first division process may be determined on a sequence-by-sequence or frame-by-frame basis. Furthermore, parameters used in the first division process may be coded into the coded bitstream.

[0657] The second division process may be determined for each sub-mesh. Furthermore, parameters used in the second division process may be coded into the coded bitstream.

[0658] Note that "different processing" may mean processing in which at least one of the number of divisions or the division method is different.

[0659] Also, for example, an encoding device of the present disclosure includes a circuit and a memory connected to the circuit, wherein the circuit, during operation, encodes a first submesh to generate first encoded data, encodes a second submesh to generate second encoded data, and in a decoding device, generates control information used to select a set of vertices including a first vertex included in the first submesh and a second vertex included in the second submesh, and generates a bitstream including the first encoded data, the second encoded data, and the control information, wherein the first vertex and the second vertex have the same level of detail (LoD) index value.

[0660] Also, for example, a decoding device of the present disclosure includes a circuit and a memory connected to the circuit, wherein the circuit, during operation, generates a first submesh and a second submesh, selects a set of vertices including a first vertex included in the first submesh and a second vertex included in the second submesh, modifies the positions of the first vertex and the second vertex to modified positions generated from the set of vertices, and the first vertex and the second vertex have the same level of detail (LoD) index value.

[0661] Also, for example, the encoding method of the present disclosure encodes a first submesh to generate first encoded data, encodes a second submesh to generate second encoded data, generates control information in a decoding device used to select a set of vertices including a first vertex included in the first submesh and a second vertex included in the second submesh, and generates a bitstream including the first encoded data, the second encoded data, and the control information, wherein the first vertex and the second vertex have the same level of detail (LoD) index value.

[0662] Also, for example, the decoding method of the present disclosure generates a first submesh and a second submesh, selects a set of vertices including a first vertex included in the first submesh and a second vertex included in the second submesh, modifies the positions of the first vertex and the second vertex to modified positions generated from the set of vertices, and the first vertex and the second vertex have the same level of detail (LoD) index value.

[0663] In this disclosure, to solve the problem of using different subdivision methods in adjacent sub-meshes, we enforce the constraint to have the same type of subdivision method and the same number of subdivisions along all boundary edges, as shown in Figure 50. Therefore, after the subdivided vertices are displaced and the sub-meshes are joined, there are no holes around the boundary edges, as shown in Figure 50(c), where the values ​​of the displacement vectors of vertex F and vertex M are the same.

[0664] A predetermined subdivision method and / or a predetermined number of subdivisions may be used for subdivision of the boundary edge, or the encoding device may determine the subdivision method and / or the number of subdivisions and signal information in the bitstream indicating the determined subdivision method and / or the number of subdivisions.

[0665] In the field of multimedia data coding technology, it is desirable to propose new methods for improving coding efficiency, improving image quality, and reducing circuit scale.

[0666] Each of the embodiments, some of the components, and each of the methods in the present disclosure enables at least one of, for example, improved coding efficiency, improved image quality, reduced encoding / decoding processing volume, reduced circuit size, and improved encoding / decoding processing speed. Alternatively, each of the embodiments, some of the components, and each of the methods in the present disclosure enables appropriate selection of any of elements or operations, such as filters, block sizes, motion vectors, reference pictures, and reference blocks, in encoding and decoding. The present disclosure includes disclosure of configurations and methods that can provide advantages other than those described above. Examples of such configurations and methods include configurations and methods that improve coding efficiency while suppressing an increase in processing volume.

[0667] Additional value and advantages of aspects of the present disclosure will become apparent from the specification and drawings, which may be obtained individually through various embodiments and features of the specification and drawings, not all of which need be provided to obtain one or more of such value and / or advantages.

[0668] These general or specific aspects may be implemented using a system, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or any combination of a system, a method, an integrated circuit, a computer program, or a recording medium.

[0669] <Submesh Combination Threshold> When zipping multiple submeshes, specifically when zipping the submeshes (more specifically, when zipping the vertices of the submeshes (more specifically, when zipping the boundary vertices) of the submeshes), the decoding device 200 searches for the vertices of each submesh and determines whether they are vertices to be combined (also called matching points). In this case, for each vertex of one submesh, the decoding device 200 searches for multiple vertices of the other submesh and determines whether the distance from the vertex of the one submesh is within a predetermined threshold range.

[0670] Note that a boundary vertex is a vertex on a boundary edge, in other words, a vertex located on a boundary edge. For example, boundary vertices are vertices located on both ends of a boundary edge. A boundary edge is an edge (side) that overlaps between sub-meshes. Specifically, a boundary edge is an edge that is common to two sub-meshes when a three-dimensional mesh is divided into two sub-meshes.

[0671] The threshold may be set per sequence, per frame, per submesh, and / or per vertex.

[0672] Here, the threshold value is set to, for example, the distance from a vertex of a certain submesh.

[0673] However, when there are two or more submeshes to be combined with a certain submesh, that is, when a certain submesh is combined with multiple submeshes, it may not be possible to set an optimal threshold. For example, the optimal threshold may differ for each vertex. In such a case, if a threshold is set for each submesh, the decoding device 200 may not be able to properly search for the vertex to be combined and may make an incorrect combination, resulting in a failure to obtain a high-quality 3D mesh shape.

[0674] Therefore, for example, the encoding device 100 transmits information (combination information) indicating a combination of submeshes to be combined (submesh-pair) to the decoding device 200. In addition, the encoding device 100 transmits information (threshold information) indicating a threshold for each combination of submeshes to be combined to the decoding device 200.

[0675] By transmitting information indicating the threshold value for each combination of submeshes to be combined to the decoding device 200, even when a submesh is combined with multiple submeshes, the decoding device 200 can properly search for the vertices to be combined, thereby preventing incorrect combinations and obtaining a high-quality three-dimensional mesh shape.

[0676] In particular, if the coding parameters, such as base mesh quantization, displacement vector quantization, and lifting transform, are different for each sub-mesh, the distance between the sub-meshes will change, which will result in a change in the optimal threshold for each vertex. Even in such cases, a high-quality connected 3D mesh can be reproduced.

[0677] Furthermore, by setting a threshold for each combination, it is not necessary to set a threshold for each vertex, and therefore the amount of information can be reduced.

[0678] Furthermore, when a large number of submeshes are combined into one submesh, information indicating the optimal threshold for each combination is transmitted to the decoding device 200, thereby enabling a high-quality three-dimensional mesh shape to be obtained.

[0679] In addition, by transmitting the submesh combination information to the decoding device 200, it is possible to omit transmitting threshold information corresponding to submeshes that are not connected to any submeshes, thereby reducing the amount of information in the threshold information.

[0680] Furthermore, by transmitting the submesh combination information to the decoding device 200, the decoding device 200 can search only the vertices of the submeshes to be combined. This reduces the amount of processing required for the vertex search process by the decoding device 200. It also makes it possible to avoid erroneous combination with submeshes other than the combination target.

[0681] The transmission of threshold information and the process of searching for matching points using a threshold in combining sub-meshes will be described below.

[0682] FIG. 90 is a block diagram showing an example of the configuration of the encoding device 100 and the decoding device 200 according to the embodiment.

[0683] The encoding device 100 includes, for example, a submesh divider 2401 and a submesh encoder 2402 .

[0684] The submesh divider 2401 divides an input 3D mesh into a plurality of submeshes. The submesh divider 2401 also generates hint information (submesh hint information metadata) that indicates useful hints for combining submeshes with high quality in the decoding device 200. The hint information is metadata that includes, for example, threshold information (distance information).

[0685] The submesh divider 2401, for example, calculates a threshold value for searching for vertices between submeshes when combining the submeshes, and stores threshold information indicating the calculated threshold value in metadata such as SEI.

[0686] The sub-mesh encoder 2402 encodes the SEI, and includes (stores) the encoded SEI in a bitstream together with encoded data that encodes information about the sub-mesh, and transmits the encoded SEI to the decoding device 200.

[0687] The decoding device 200 includes, for example, a sub-mesh decoder 2403 and a sub-mesh combiner 2404 .

[0688] The submesh decoder 2403 decodes the coded data contained in the bitstream (more specifically, information relating to the coded submeshes).

[0689] The submesh combiner 2404 obtains threshold information from metadata such as SEI contained in the bitstream, searches for vertices to be combined with vertices of a certain submesh based on the threshold information, and combines the found vertices with vertices of a certain submesh.

[0690] 91 is a flow diagram showing a search process for matching points according to an embodiment. Before the search process is performed, the decoding device 200, for example, obtains multiple submeshes (specifically, position information for each boundary vertex of the multiple submeshes) by decoding encoded data included in the bitstream. In this example, a threshold is set for each boundary vertex, and the threshold information for each boundary vertex is included in the bitstream. In this example, the threshold for boundary vertices that are not connected to other boundary vertices is set to 0.

[0691] First, the decoding device 200 determines whether maxDistance, which is the maximum value of the threshold indicated by each piece of threshold information included in the bitstream, is 0 (S701). In other words, the decoding device 200 determines whether there is a matching point at each boundary vertex of the multiple sub-meshes based on maxDistance. For example, if maxDistance = 0, the decoding device 200 determines that there is no matching point at any of the boundary vertices of the multiple sub-meshes. In this case, the decoding device 200 does not perform the matching process described below. On the other hand, for example, if maxDistance = 0, the decoding device 200 determines that there is a matching point at at least one of the boundary vertices of the multiple sub-meshes.

[0692] If the decoding device 200 determines that the maxDistance included in the bitstream is 0 (Yes in S701), it determines that there is no matching point (S709) and ends the process.

[0693] On the other hand, if the decoding device 200 determines that the maxDistance included in the bitstream is not 0 (No in S701), it determines that a matching point exists and starts matching processing (S702). Specifically, it executes the processing from step S703 onwards for each boundary vertex of multiple sub-meshes.

[0694] The decoding device 200 calculates a threshold value for each boundary vertex in the frame (for example, a plurality of sub-meshes) based on the SEI included in the bitstream (S703).

[0695] Next, the decoding device 200 performs processing for each boundary vertex in the frame (S704).

[0696] Next, the decoding device 200 calculates the distances to all boundary vertices of all submeshes other than the submesh containing the boundary vertex (target boundary vertex) for which a matching point is searched (S705).

[0697] Next, the decoding device 200 determines whether the distance between the target boundary vertex and each boundary vertex is less than a threshold value (S706).

[0698] If the decoding device 200 determines that the distance between the target boundary vertex and each boundary vertex is not less than the threshold value (No in S706), that is, if it determines that the distances between the target boundary vertex and each boundary vertex are all greater than or equal to the threshold value, it determines that no matching point exists at the target boundary vertex (S709).

[0699] On the other hand, if the decoding device 200 determines that the distance between the target boundary vertex and each boundary vertex is less than the threshold (Yes in S706), that is, if it determines that the distances between the target boundary vertex and each boundary vertex include distances less than the threshold, it determines the boundary vertex with the smallest distance among the distances less than the threshold as the matching point (S707).

[0700] Furthermore, the decoding device 200 determines not to search for a matching point for the boundary vertex determined as the matching point (S708).

[0701] After searching for a matching point for the target boundary vertex, the decoding device 200 performs a joining process between the target boundary vertex and the matching point. For example, when joining two boundary vertices that are far apart, the decoding device 200 calculates the position coordinates of the joining point (a new vertex generated by joining the two boundary vertices) using a predetermined method. The calculated position coordinates (position coordinates of the joining point) may be, for example, the midpoint or center of gravity of the two points, or the position coordinates may be brought closer to one of the boundary vertices using a weighting factor. For example, a three-dimensional point is generated at the position coordinates calculated in this way, and the two boundary vertices are deleted, thereby joining the two boundary vertices.

[0702] FIG. 92 is a diagram for explaining the positional relationship of a plurality of sub-meshes according to an embodiment.

[0703] In the example shown in Figure 92, multiple sub-meshes are generated by cutting and dividing a 3D mesh in a certain direction (horizontally in this example) (linear segmentation). With this division method, the maximum number of sub-meshes adjacent to one sub-mesh is two. For example, sub-mesh 1 is not adjacent to sub-mesh 3.

[0704] In the encoding process, the positions (coordinates) of the vertices of the submeshes may be moved. For example, suppose that the division of the three-dimensional mesh results in the maximum distance between the boundary vertices of submeshes 1 and 2 being 5 units, and the maximum distance between the boundary vertices of submeshes 2 and 3 being 20 units. Note that the unit of distance may be any unit and is not particularly limited.

[0705] For example, when a threshold is set for each submesh, for the three submeshes (submeshes 1, 2, and 3) shown in Figure 92, the threshold for submesh 1 is set to 5 (Threshold(submesh_1) = 5), the threshold for submesh 2 is set to 20 (Threshold(submesh_2) = 20), and the threshold for submesh 3 is set to 20 (Threshold(submesh_3) = 20).

[0706] 92 is used and it is guaranteed that the submeshes can be processed in, for example, the encoding order. Therefore, even if threshold information indicating a threshold for each submesh is transmitted, the threshold for submesh 1 will be the threshold used in the joining process between submeshes 1 and 2, and the threshold for submesh 2 will be the threshold used in the joining process between submeshes 2 and 3. Therefore, even if a threshold is set for each submesh, the set threshold may be treated as the threshold for each combination of submeshes.

[0707] However, if submeshes 1 and 3 are located close to each other, there is a possibility that submeshes 1 and 3 will be joined even if submeshes 1 are not a target for joining with submeshes 3.

[0708] Here, for example, the threshold value for submeshes 1 and 2 is set to 5, the threshold value for submeshes 2 and 3 is set to 20, and the threshold value for submeshes 1 and 3 is set to 0. In other words, a threshold value is set for each combination of submeshes. For example, if the threshold value for a combination is 0, the decoding device 200 does not combine the submeshes of that combination. In this example, the decoding device 200 does not combine submeshes 1 and 3.

[0709] In this way, by including in the bit stream information indicating that submesh 1 is not a target for joining with submesh 3, it is possible to prevent erroneous joining.

[0710] FIG. 93 is a diagram for explaining the process of searching for matching points of a plurality of sub-meshes according to an embodiment.

[0711] In this example, submesh 1 contains boundary vertices A, B, C, L, M, and N; submesh 2 contains boundary vertices D, E, F, G, H, and K; and submesh 3 contains boundary vertices X, Y, and Z.

[0712] The algorithm for merging boundary vertices (the merge algorithm) uses a set threshold (threshold distance) to search for matching points for boundary vertices. Figure 93 shows the search radius associated with the search for matching points for each boundary vertex of submesh 2. The search radius corresponds, for example, to a threshold value set for each boundary vertex. Other boundary vertices within a circle whose distance (radius) from the target boundary vertex is the threshold value become the search range for matching points for the target boundary vertex.

[0713] 93 indicates the search radius of boundary vertex E. The dashed line in FIG. 93 indicates the search radius of boundary vertex D. The dashed line in FIG. 93 indicates the search radius of boundary vertex H. The dotted line in FIG. 93 indicates the search radius of boundary vertex F.

[0714] Boundary vertices D, H, F, and E are boundary vertices of submesh 2 that are connected to boundary vertices of other submeshes. Boundary vertex D is connected to boundary vertex A, boundary vertex H is connected to boundary vertex L, boundary vertex F is connected to boundary vertex B, and boundary vertex E is connected to boundary vertex X. On the other hand, boundary vertices G and K of submesh 2 are not connected to any boundary vertices of other submeshes.

[0715] For example, if a threshold is set for each submesh, the maximum value of the distance between each target boundary vertex of the submesh and each matching point is set as the threshold for the submesh. Therefore, for submesh 2, the threshold associated with boundary vertex E, which has the largest search radius, is set.

[0716] When a threshold is set for each submesh, in this example, the distance between submeshes 1 and 2 is close, but the distance between submeshes 2 and 3 is far, so in order to join boundary vertex E and boundary vertex X, the distance between submeshes 2 and 3 (i.e., the search radius of boundary vertex E) needs to be set as the threshold.

[0717] 94 is a diagram for explaining the process of searching for matching points of a plurality of sub-meshes according to an embodiment. Note that the two-dot chain line in FIG. 94 indicates the search radius of the boundary vertex G.

[0718] The decoding device 200, using the threshold set by the encoding device 100, searches for matching points for all boundary vertices of submesh 2.

[0719] Because the set threshold (distance) is greater than or equal to the threshold (distance) that is originally sufficient for the boundary vertices D, H, F, and E, the boundary vertices D, H, F, and E are combined with the appropriate matching points.

[0720] For example, when a threshold (search radius) set for boundary vertex G is used, boundary vertex A of submesh 1 is found within the search range. Therefore, boundary vertex G is combined with boundary vertex A. Then, boundary vertices A, D, and G may be combined to form a degenerate triangle of face DGH.

[0721] Furthermore, for example, even if a threshold value (search radius) set for boundary vertex K is used, boundary vertices of submeshes 1 and 3 may not be found within the range in which matching points are searched.

[0722] Therefore, for example, the decoding device 200 performs the following process in searching for matching points.

[0723] 95 is a flow diagram showing a search process for matching points according to an embodiment. Before the search process is performed, the decoding device 200, for example, decodes coded data included in the bitstream to obtain multiple submeshes (specifically, position information for each boundary vertex of the multiple submeshes). In this example, a threshold is set for each combination of submeshes, and the threshold information for each combination is included in the bitstream. In this example, the threshold for combinations of submeshes that are not combined is set to 0.

[0724] First, the decoding device 200 determines whether maxDistance, which is the maximum value of the thresholds indicated by the threshold information included in the bitstream, is 0 (S711). In other words, the decoding device 200 determines whether there is a matching point at each boundary vertex of the multiple sub-meshes based on maxDistance. For example, if maxDistance = 0, the decoding device 200 determines that there is no matching point at any of the boundary vertices of the multiple sub-meshes. In this case, the decoding device 200 does not perform the matching process described below. On the other hand, for example, if maxDistance = 0, the decoding device 200 determines that there is a matching point at at least one of the boundary vertices of the multiple sub-meshes.

[0725] If the decoding device 200 determines that the maxDistance included in the bitstream is 0 (Yes in S711), it determines that there is no matching point (S719) and ends the process.

[0726] On the other hand, if the decoding device 200 determines that the maxDistance included in the bitstream is not 0 (No in S711), it determines that a matching point exists and starts matching processing (S712). Specifically, it executes the processing from step S713 onwards for each boundary vertex of multiple sub-meshes.

[0727] The decoding device 200 calculates a threshold for each combination of submeshes (connected pairs of submeshes) within a frame (e.g., multiple submeshes) based on the SEI contained in the bitstream, and calculates a threshold for each boundary vertex based on the threshold for each combination (S713).

[0728] Next, the decoding device 200 performs processing for each boundary vertex in the frame (S714).

[0729] Next, the decoding device 200 calculates the distances to all boundary vertices of all submeshes for which combinations are defined (S715).

[0730] Next, the decoding device 200 determines whether the distance between the target boundary vertex and each boundary vertex is less than a threshold value (S716).

[0731] If the decoding device 200 determines that the distance between the target boundary vertex and each boundary vertex is not less than the threshold value (No in S716), that is, if it determines that the distances between the target boundary vertex and each boundary vertex are all greater than or equal to the threshold value, it determines that no matching point exists at the target boundary vertex (S719).

[0732] On the other hand, if the decoding device 200 determines that the distance between the target boundary vertex and each boundary vertex is less than the threshold (Yes in S716), that is, if it determines that the distances between the target boundary vertex and each boundary vertex include distances less than the threshold, it determines the boundary vertex with the smallest distance among the distances less than the threshold as the matching point (S717).

[0733] Furthermore, the decoding device 200 determines not to search for a matching point for the boundary vertex determined as the matching point (S718).

[0734] Fig. 96 is a diagram for explaining the search process for matching points of multiple sub-meshes according to an embodiment. Note that the two-dot chain line in Fig. 96 indicates the search radius of boundary vertex E. The dashed line in Fig. 96 indicates the search radius of boundary vertex D. The one-dot chain line in Fig. 96 indicates the search radius of boundary vertex H. The dotted line in Fig. 96 indicates the search radius of boundary vertex F.

[0735] When a threshold is set for each combination (pair) of submeshes as described above, if the search radius of boundary vertex H is larger than the search radius of boundary vertices D and F, as shown in Figure 96, the threshold for the combination of submesh 1 and submesh 2 is set to the search radius of boundary vertex H.

[0736] 97 is a diagram for explaining the process of searching for matching points of multiple sub-meshes according to an embodiment. Note that the dashed-dotted line in Fig. 97 indicates the search radius of the boundary vertex G used for sub-mesh 1. Also, the dashed-two-dotted line in Fig. 97 indicates the search radius of the boundary vertex G used for sub-mesh 3.

[0737] The decoding device 200 searches for boundary vertices of submesh 1 that are included in the search range for matching points of boundary vertex G, for example, using a threshold value associated with the combination of submesh 1 and submesh 2.

[0738] In addition, the decoding device 200 searches for boundary vertices of submesh 3 that are included in the search range of matching points for boundary vertex G, for example, using a threshold value associated with the combination of submesh 2 and submesh 3.

[0739] Boundary vertex A is included in submesh 1, and boundary vertex G is included in submesh 2. Furthermore, the distance between boundary vertex G and boundary vertex A is closer than the search radius of boundary vertex G used for submesh 3, but is farther than the search radius of boundary vertex G used for submesh 1. Therefore, boundary vertex G is not connected to boundary vertex A.

[0740] Fig. 98 is a diagram for explaining an example of syntax in which sub-mesh combining metadata according to an embodiment is signaled. Specifically, Fig. 98 is a diagram showing syntax for transmitting threshold information indicating a threshold (maximum distance threshold for searching for matching points in combining processing) to the decoding device 200. Fig. 98 is a first example of the syntax.

[0741] The sub-mesh combining metadata exists for each instance and is a syntax indicating the kth (kth integer) metadata. That is, each sub-mesh combining metadata has k elements. In this example, the index k is omitted in each syntax.

[0742] In this example, the metadata for submesh combination includes distance_per_submesh_flag (a flag (flag information) indicating whether or not to indicate a threshold value for each submesh), and when distance_per_submesh_flag = 1 (i.e., when a threshold value for each submesh is indicated), it includes information indicating the threshold value for each submesh.

[0743] On the other hand, if distance_per_submesh_flag=0 (that is, if a threshold value for each submesh is not indicated), the submesh combining metadata includes distance_per_submesh_pair_flag (a flag indicating whether or not a threshold value for each combination of submeshes is indicated).

[0744] If distance_per_submesh_pair_flag=1 (that is, if a threshold value for each combination of submeshes is indicated), the submesh combining metadata includes threshold value information indicating the threshold value for each combination of submeshes.

[0745] Here, if the threshold value indicated by the threshold value information is a predetermined value (for example, 0), the sub-meshes may be considered not to be adjacent to each other.

[0746] 99 is a diagram for explaining an example of syntax in which sub-mesh combining metadata according to an embodiment is signaled, and is a second example of the syntax.

[0747] In this example, the type of sub-mesh division method (division type / segmentation_type) is also indicated.

[0748] For example, if the division method of a 3D mesh is linear segmentation (linear_segmentation), it indicates that the 3D mesh is divided by cutting in a certain direction (for example, the horizontal direction).

[0749] The division method may be indicated by a flag instead of the division type.

[0750] Furthermore, if a division method other than linear segmentation exists, the method of transmitting threshold information may be switched depending on the division method. For example, when the division method of 3D meshes is linear segmentation, the encoding device 100 does not include threshold information indicating the threshold for each combination of submeshes in the bitstream, but includes threshold information for the number of submeshes or the number of submeshes minus 1 in the bitstream.

[0751] Note that distance_per_submesh_flag and distance_per_submesh_pair_flag may be included independently in the bitstream.

[0752] Furthermore, distance_per_submesh_pair may be included in the bitstream as a one-dimensional array instead of a two-dimensional array.

[0753] In a syntax loop indicating a combination of submeshes, the second loop indicates information about submeshes from the pth submesh (p is an integer greater than or equal to 1) in the first loop onwards, thereby reducing the amount of information.

[0754] For example, distance_per_submesh_flag[k] = 1 indicates that for multiple submeshes (zippering instances) corresponding to the metadata of index k, a threshold for each combination of submeshes is set to the boundary vertex. For example, distance_per_submesh_flag[k] = 0 indicates that for multiple submeshes (zippering instances) corresponding to the metadata of index k, a threshold for each combination of submeshes is not set to the boundary vertex. The default value of distance_per_submesh_pair_flag[k] is, for example, 0.

[0755] Also, for example, segmentation_type[k]=1 indicates that submeshes having an index (submeshIdx, for example, a submesh identification number) in the range from 1 to MaxNumSubmeshes[frameIdx]-2 can have mutual boundary vertices (i.e., have boundary vertices that can be joined) only with submeshes having an index of submeshIdx-1 and an index of submeshIdx+1. For example, when segmentation_type[k]=0, a submesh having an index of 0 and a submesh having an index of MaxNumSubmeshes[frameIdx]-1 can have mutual boundary vertices only with a submesh having an index of 1 and a submesh having an index of MaxNumSubmeshes[frameIdx]-2, respectively. Also, for example, segmentation_type[k] = 0 indicates that each submesh can have a mutual boundary vertex with a submesh different from the above. In other words, when segmentation_type[k] = 0, it indicates that a submesh having an index of submeshIdx may have a mutual boundary vertex with submeshes other than those having an index of submeshIdx-1 and an index of submeshIdx+1. The default value of segmentation_type[k] is, for example, 1.

[0756] Also, for example, segmentation_type[k] = 1 indicates that a submesh having an index (submeshIdx) ranging from 0 to MaxNumSubmeshes[frameIdx] - 1 can have mutual boundary vertices with the submeshes with indices submeshIdx - 1 and submeshIdx + 1 when submeshIdx - 1 is 0 or greater and submeshIdx + 1 is MaxNumSubmeshes[frameIdx] - 2 or less. When segmentation_type[k] = 0, it indicates that each submesh may have mutual boundary vertices with a submesh different from the above. The default value of segmentation_type[k] is, for example, 1.

[0757] distance_per_submesh_pair_flag[k][p][n] indicates the value of a variable (zipperingMaxMatchDistancePerPatchPair[k][p][n]) used to process a combination of a submesh with index p and a submesh with index n+p+1 that is different from the submesh with index p and that is divided from the same 3D mesh as the submesh (zippering instance) for multiple submeshes (zippering instances) corresponding to the metadata of index [k] when a zipping process is used. For example, the length of the syntax element distance_per_submesh_pair_flag[k][p][n] is Ceil(Log2(max_match_distance[k])) bits.

[0758] A specific example of syntax setting using an example of a threshold determination method in the encoding device 100 will be described.

[0759] For example, suppose the values ​​of each variable are given as follows:

[0760] 1. distance_per_submesh_pair_flag[k]=True 2. distance_per_submesh_flag[k]=False 3. number_of_submesh[k]=4 4. segmentation_type[k]=1 (Linear Segmentation)

[0761] These variables cause the following code to be executed:

[0762] for (p=0; p<numSubmeshes-1; p++) { distance_per_submesh_pair_flag[k][p][0]}

[0763] If numSubmeshes is replaced with a signed value, the above code can be written as follows:

[0764] for (p=0; p<3; p++) { distance_per_submesh_pair_flag[k][p][0]}

[0765] Then, from the bitstream, (i) distance_per_submesh_pair_flag[k][0][0] indicating the threshold (distance) between submesh 0 and submesh 1, (ii) distance_per_submesh_pair_flag[k][1][0] indicating the threshold (distance) between submesh 1 and submesh 2, and (iii) distance_per_submesh_pair_flag[k][2][0] indicating the threshold (distance) between submesh 2 and submesh 3 are obtained.

[0766] Here, the threshold value for the combination of submesh 1 and submesh 2 is the same as the threshold value for the combination of submesh 2 and submesh 1, so two pieces of information about the same value do not need to be included in the bit stream.

[0767] Next, another example of syntax setting using an example of a threshold determination method in the encoding device 100 will be described.

[0768] For example, suppose the values ​​of each variable are given as follows:

[0769] 1. distance_per_submesh_pair_flag[k]=True 2. distance_per_submesh_flag[k]=False 3. number_of_submesh[k]=4 4. segmentation_type[k]=0

[0770] These variables cause the following code to be executed:

[0771] for (p=0; p<numSubmeshes-1; p++) { for (n=p+1; n<numSubmeshes; n++) { distance_per_submesh_pair_flag[k][p][n-p-1]}}

[0772] If numSubmeshes is replaced with a signed value, the above code can be written as follows:

[0773] for (p=0; p<3; p++) { for (n=p+1; n<4; n++) { distance_per_submesh_pair_flag[k][p][n-p-1]}}

[0774] Then, from the bitstream, (i) distance_per_submesh_pair_flag[k][0][0] indicating the threshold (distance) between submesh 0 and submesh 1, (ii) distance_per_submesh_pair_flag[k][0][1] indicating the threshold (distance) between submesh 0 and submesh 2, and (iii) distance_per_submesh_pair_flag[k][0] indicating the threshold (distance) between submesh 0 and submesh 3. [2], (iv) distance_per_submesh_pair_flag[k][1][0] indicating the threshold (distance) between submesh 1 and submesh 2, (v) distance_per_submesh_pair_flag[k][1][1] indicating the threshold (distance) between submesh 1 and submesh 3, and (vi) distance_per_submesh_pair_flag[k][2][0] indicating the threshold (distance) between submesh 2 and submesh 3 are obtained.

[0775] Here, the threshold value for the combination of submeshes 1 and 3 is the same as the threshold value for the combination of submeshes 3 and 0, so two pieces of information about the same value do not need to be included in the bit stream.

[0776] Next, another example of the syntax of the sub-mesh connection metadata will be shown.

[0777] Figure 100 is a diagram for explaining an example of a syntax in which submesh combination metadata according to an embodiment is signaled. Figure 101 is a diagram for explaining an example of a syntax in which information indicating a combination of submeshes to be combined according to an embodiment is signaled. Specifically, Figure 101 shows a syntax (connected_submesh_info) indicating combination information indicating whether multiple submeshes are combined.

[0778] The information indicated by connected_submesh_info may be included in the submesh combining metadata, or may be included in an SEI different from the submesh combining metadata and transmitted to the decoding device 200 via the bitstream.

[0779] Based on combination information indicating whether the sub-meshes are combined or not, threshold information indicating the threshold value of the combination is transmitted if the combination is combined, and if the combination is not combined, the threshold information does not need to be transmitted.

[0780] As described above, threshold information indicating the threshold for each submesh may be included in the bit stream, or threshold information indicating the threshold for each combination of submeshes may be included in the bit stream.

[0781] The information indicating the threshold value may be signaled in the bitstream for each submesh, or for each combination of vertices (specifically, boundary vertices) to be combined. Information on whether a vertex is a boundary vertex may also be signaled in the bitstream.

[0782] For example, the encoding device 100 determines a method for setting the threshold, includes information indicating the determined method and threshold information in metadata, and transmits a bitstream including the metadata to the decoding device 200.

[0783] The decoding device 200 performs a search for matching points in the combination of sub-meshes based on, for example, information indicating the method and threshold information included in the bitstream.

[0784] When threshold information indicating a threshold for each combination of sub-meshes is transmitted, the decoding device 200 may determine that the sub-meshes are not combined and skip the search process for matching points using the threshold if, for example, the threshold indicated by the threshold information is 0. In other words, in this case, the search process for matching points does not need to be performed.

[0785] Furthermore, for example, if the threshold value indicated by the threshold value information is other than 0, the decoding device 200 determines that the sub-meshes are to be combined, and executes a search process using the threshold value.

[0786] Furthermore, the encoding device 100 may transmit, for example, for each combination of submeshes, a bitstream including information indicating whether the submeshes are combined to the decoding device 200. Furthermore, when the submeshes are combined, the encoding device 100 may transmit, to the decoding device 200, a bitstream including threshold information.

[0787] Furthermore, in the process of searching for matching points using a threshold value, the decoding device 200 executes the search process in two submeshes to be combined using a threshold value corresponding to the combination.

[0788] Furthermore, the decoding device 200 determines whether or not submeshes are to be combined, for example, based on threshold information or information indicating whether or not the submeshes are to be combined, and performs processing on combinations of submeshes that are determined to be combined, and does not perform processing on combinations of submeshes that are determined not to be combined. Note that the processing referred to here may be a search process for matching points, other processing in combining submeshes, or a combining process of attribute information of the submeshes and attribute information of the faces of the submeshes.

[0789] Furthermore, the encoding device 100 may switch between transmitting a bitstream including threshold information for each submesh and transmitting a bitstream including threshold information for each combination of submeshes, depending on the submesh division method.

[0790] Also, a third threshold transmission method may be used depending on the sub-mesh division method.

[0791] <Representative Example> Fig. 102 is a flowchart showing an example of basic encoding processing according to this embodiment. For example, the encoding device 100 shown in Fig. 24 includes a circuit 151 and a memory 152 connected to the circuit 151. In the encoding device 100, the circuit 151 performs the following processing during operation.

[0792] First, the encoding device 100 generates a plurality of encoded data by encoding a plurality of submeshes (S721). The encoding device 100 encodes the plurality of submeshes by encoding, for example, vertex information and connection information of each of the plurality of submeshes (specifically, base mesh information of each of the plurality of submeshes), as well as information about the plurality of submeshes, such as displacement vector information.

[0793] Next, the encoding device 100 generates threshold information indicating a threshold value determined for each combination of two submeshes selected from the plurality of submeshes, based on the positions of a plurality of 3D points included in the plurality of submeshes (S722). The 3D points are, for example, the boundary vertices described above. The threshold information is, for example, the distance_per_submesh_pair described above.

[0794] Next, the encoding device 100 generates a bitstream including a plurality of pieces of encoded data and threshold information (S723).

[0795] The threshold is used, for example, by the decoding device 200 that has acquired the bitstream to combine 3D points of multiple submeshes. Here, if the threshold is set based solely on the distance from the 3D point, for example, when there are multiple submeshes that can be combined with a submesh having the 3D point, there is a possibility that the 3D point will be combined with a 3D point of a submesh other than the submesh having the 3D point that should be combined. Therefore, by setting a threshold for each combination of submeshes, 3D points can be appropriately combined. Furthermore, for two submeshes that are not combined, a specific value, such as 0, can be set as the threshold information, allowing the decoding device 200 to omit the process of determining whether or not to combine each 3D point for those two submeshes. This reduces the amount of processing by the decoding device 200.

[0796] Furthermore, for example, the encoding device 100 determines the combination based on a division method for dividing a three-dimensional mesh into a plurality of sub-meshes.

[0797] Depending on the division method (specifically, how the division is performed), multiple submeshes may only be combined with specific submeshes. In such cases, multiple submeshes can be appropriately combined without setting a threshold for every combination of multiple submeshes. This reduces the amount of processing required to determine the combinations. It also reduces the amount of information in the bitstream.

[0798] Also, for example, the encoding device 100 determines the combination based on whether the division method is a first division method in which division is performed between two submeshes to which adjacent submesh identification numbers are assigned among the submesh identification numbers assigned consecutively to each of the multiple submeshes, or a second division method other than the first division method.

[0799] The submesh identification number is, for example, the above-mentioned submesh ID. The first division method is, for example, the above-mentioned linear segmentation (linear_segmentation).

[0800] When a three-dimensional mesh is divided into multiple submeshes, multiple submeshes may be generated by repeatedly dividing the three-dimensional mesh at predetermined intervals in a direction perpendicular to a single line, such as by repeatedly dividing the three-dimensional mesh at predetermined intervals in a predetermined direction. In such cases, the multiple submeshes are assigned submesh identification numbers consecutively in the order in which they were divided from the three-dimensional mesh. With this division method, a given submesh can only be combined with submeshes that are adjacent to the given submesh. In other words, with this division method, a given submesh can only be combined with submeshes that have an identification number adjacent to the identification number assigned to the given submesh.

[0801] Furthermore, for example, the encoding device 100 calculates the distance between each pair of two three-dimensional points, one selected from each of the plurality of three-dimensional points included in a first submesh of the two submeshes of the combination and one selected from each of the plurality of three-dimensional points included in a second submesh of the two submeshes, and calculates the longest distance between the calculated pairs as the threshold value for the combination.

[0802] This allows the first submesh and the second submesh to be appropriately combined using one threshold value.

[0803] Fig. 103 is a flowchart showing an example of basic decoding processing according to this embodiment. For example, a decoding device 200 shown in Fig. 25 includes a circuit 251 and a memory 252 connected to the circuit 251. In the decoding device 200, the circuit 251 performs the following processing during operation.

[0804] First, the decoding device 200 obtains a bitstream including multiple encoded data generated by encoding multiple submeshes and threshold information indicating a threshold value determined for each combination of two submeshes selected from the multiple submeshes (S731).

[0805] Next, the decoding device 200 decodes the plurality of coded data (S732).

[0806] Next, the decoding device 200 performs a combining process for combining a plurality of 3D points included in a plurality of sub-meshes based on the threshold information (S733).

[0807] According to this, by setting a threshold for each combination of submeshes, 3D points can be appropriately combined. Furthermore, for two submeshes that are not combined, a specific value such as 0 is set as the threshold information, so that the process of determining whether or not to combine each 3D point for those two submeshes can be omitted. This reduces the amount of processing.

[0808] Furthermore, for example, in the combining process (S733), the decoding device 200 performs a search process for searching for a pair of two 3D points to be combined among the multiple 3D points included in the multiple submeshes if a predetermined condition is satisfied for each combination, and combines the two 3D points of the searched pair based on threshold information. The predetermined condition may be determined arbitrarily and is not particularly limited.

[0809] This allows the search process to connect two suitable 3D points.

[0810] Furthermore, for example, in the combining process, the decoding device 200 determines, for each combination, whether or not there is a pair of two combined 3D points among the multiple 3D points included in the two submeshes of the combination, and if it is determined that there is a pair of two combined 3D points among the multiple 3D points included in the two submeshes of the combination, it performs a search process, and if it is determined that there is not a pair of two combined 3D points among the multiple 3D points included in the two submeshes of the combination, it does not perform the search process. In other words, for example, if it is determined that there is a pair of two combined 3D points among the multiple 3D points included in the two submeshes of the combination, the decoding device 200 determines that a predetermined condition is satisfied and performs the search process.

[0811] According to this, when the search process is unnecessary, the search process is omitted, thereby reducing the amount of processing.

[0812] Furthermore, for example, the decoding device 200 determines, based on threshold information, whether or not there is a pair of two 3D points to be combined among the multiple 3D points included in the two submeshes of the combination.

[0813] For two submeshes that are not to be joined, a specific value such as 0 is set as the threshold information, and thus it is possible to appropriately determine whether or not to join the two submeshes based on whether or not the threshold indicated by the threshold information is a specific value.

[0814] <Other Examples> Although aspects of the encoding device 100 and the decoding device 200 have been described above according to the embodiments, the aspects of the encoding device 100 and the decoding device 200 are not limited to the embodiments. Modifications conceivable by those skilled in the art may be applied to the embodiments, and multiple components in the embodiments may be combined in any manner.

[0815] For example, a process performed by a specific component in the embodiment may be performed by another component instead of the specific component. Also, the order of multiple processes may be changed, or multiple processes may be performed in parallel.

[0816] Furthermore, as described above, at least some of the configurations of the present disclosure may be implemented as an integrated circuit. At least some of the processes of the present disclosure may be used as an encoding method or a decoding method. A program for causing a computer to execute the encoding method or the decoding method may be used. A non-transitory computer-readable recording medium on which the program is recorded may be used. A bitstream for causing the decoding device 200 to perform a decoding process may be used.

[0817] Furthermore, at least some of the configurations and processes of the present disclosure may be used as a transmitting device, a receiving device, a transmitting method, or a receiving method. A program for causing a computer to execute the transmitting method or the receiving method may be used. Furthermore, a non-transitory computer-readable recording medium on which the program is recorded may be used.

[0818] The present disclosure is useful, for example, in encoding devices, decoding devices, transmitting devices, receiving devices, etc. related to three-dimensional meshes, and is applicable to computer graphics systems, three-dimensional data display systems, etc.

[0819] 100 Encoding device 101, 121, 144 Vertex information encoder 102, 145 Connection information encoder 103, 122 Attribute information encoder 104, 204, 1103 Preprocessor 105, 205, 2106 Postprocessor 110 Three-dimensional data encoding system 111, 211 Controller 112, 212 Input / output processor 113 Three-dimensional data encoder 114 System multiplexer 115 Three-dimensional data generator 123 Metadata encoder 124, 1274 Multiplexer 131 Vertex image generator 132 Attribute image generator 133 Metadata generator 134 Video encoder 141 Two-dimensional data encoder 142 Mesh data encoder 143 Texture encoder 148 Description encoder 151, 251 Circuit 152, 252 Memory 200 Decoding device 201, 221, 244 Vertex information decoder 202, 245 Connection information decoder 203, 222 Attribute information decoder 210 3D data decoding system 213 3D data decoder 214 System demultiplexer 215, 247 Presentation device 216 User interface 223 Metadata decoder 224 Demultiplexer 231 Vertex information generator 232 Attribute information generator 234 Video decoder 241 2D data decoder 242 Mesh data decoder 243 Texture decoder 246 Mesh reconstructor 248 Description decoder 300 Network 310 External connector 511 Volumetric capture device 512 Projector 513 Base mesh encoder 514 Displacement encoder 515 Attribute Encoder 516 Other Type Encoder 613 Base Mesh Decoder 614 Displacement Decoder 615 Attribute Decoder 616 Other Type Decoder 617 3D Reconstructor 1101 Input Mesh 1102, 2108 Attribute Map 1104, 2103 Base Mesh 1105, 2104 Displacement Data 1106 Compressor 1107, 1304, 2101 Bitstream1108, 2105 Metadata 1231 Demultiplexer 1232, 1262 Switch 1233 Static Mesh Decoder 1234, 1264 Mesh Buffer 1235 Motion Decoder 1236, 1266 Base Mesh Reconstructor 1237, 1240 Inverse Quantizer 1238, 1243 Video Decoder 1239 Image Unpacker 1241 Inverse Wavelet Transformer 1242 Reconstructor 1244, 1272 Color Transformer 1251 Decoded Base Mesh 1252, 1402 Subdivider 1253 Subdivided Mesh 1254 Decoded Displacement Data 1255 Displacer 1256 Decoded 3D Mesh 1261, 1269 Quantizer 1263 Static Mesh Encoder 1265 Motion Encoder 1267 Displacement data updater 1268 Wavelet transformer 1270 Image packer 1271, 1273 Video encoder 1301, 2304 Mesh frame 1302, 2301, 2302 Base mesh frame 1303, 2303 Displacement information 1401 Base mesh generator 1403 Displacement data generator 2102 Decompressor 2107 Output mesh 2109 Combiner 2305 Decoder 2306 Post-decoder 2307 Pre-reconstructor 2308 Reconstructor 2309 Post-reconstructor 2310 Fitter 2311 Final 3D mesh frame 2401 Sub-mesh divider 2402 Sub-mesh encoder 2403 Sub-mesh decoder 2404 Sub-mesh combiner

Claims

1. An encoding method comprising: generating a plurality of encoded data by encoding a plurality of submeshes; generating threshold information indicating a threshold determined for each combination of two submeshes selected from the plurality of submeshes based on the positions of a plurality of three-dimensional points included in the plurality of submeshes; and generating a bitstream including the plurality of encoded data and the threshold information.

2. The encoding method according to claim 1, wherein the combination is determined based on a division method for dividing a three-dimensional mesh into the plurality of sub-meshes.

3. The encoding method according to claim 2, wherein the combination is determined based on whether the division method is a first division method in which division is performed between two submeshes to which adjacent submesh identification numbers are assigned among the submesh identification numbers assigned consecutively to each of the plurality of submeshes, or a second division method other than the first division method.

4. The encoding method according to any one of claims 1 to 3, wherein for each pair of two 3D points selected from each of a plurality of 3D points included in a first submesh of the two submeshes of the combination and a plurality of 3D points included in a second submesh of the two submeshes, the distance between the two 3D points is calculated, and the longest distance among the calculated distances for each pair is calculated as the threshold for the combination.

5. A decoding method comprising: obtaining a bitstream containing a plurality of coded data generated by coding a plurality of submeshes and threshold information indicating a threshold value determined for each combination of two submeshes selected from the plurality of submeshes; decoding the plurality of coded data; and performing a combining process to combine a plurality of three-dimensional points included in the plurality of submeshes based on the threshold information.

6. The decoding method according to claim 5, wherein the combining process performs a search process for each of the combinations to search for a pair of two 3D points to be combined from among the multiple 3D points included in the multiple sub-meshes if a predetermined condition is satisfied, and combines the two 3D points of the searched pair based on the threshold information.

7. The decoding method according to claim 6, wherein the combining process determines, for each combination, whether or not there is a pair of two combined three-dimensional points among the multiple three-dimensional points included in the two submeshes of the combination; performs the search process if it is determined that there is a pair of two combined three-dimensional points among the multiple three-dimensional points included in the two submeshes of the combination; and does not perform the search process if it is determined that there is no pair of two combined three-dimensional points among the multiple three-dimensional points included in the two submeshes of the combination.

8. The decoding method according to claim 7, further comprising determining whether or not there is a pair of two 3D points to be connected among the plurality of 3D points included in the two sub-meshes of the combination based on the threshold information.

9. An encoding device comprising: a processor; and a memory, wherein the processor uses the memory to generate a plurality of encoded data by encoding a plurality of submeshes; generate threshold information indicating a threshold determined for each combination of two submeshes selected from the plurality of submeshes based on the positions of a plurality of three-dimensional points included in the plurality of submeshes; and generate a bitstream including the plurality of encoded data and the threshold information.

10. A decoding device comprising: a processor; and a memory, wherein the processor uses the memory to obtain a bit stream including a plurality of coded data generated by coding a plurality of submeshes and threshold information indicating a threshold value determined for each combination of two submeshes selected from the plurality of submeshes; decodes the plurality of coded data; and performs a combining process to combine a plurality of three-dimensional points included in the plurality of submeshes based on the threshold information.

Citation Information

Patent Citations

  • Information processing device and method

    WO2022269944A1

  • Mesh zippering

    WO2023180844A1