Encoding method, decoding method, encoding device, and decoding device
The method symbolizes displacement vectors within a rectangular region of a picture to generate a bitstream with position information, enhancing the encoding and decoding efficiency of three-dimensional mesh data, addressing inefficiencies in existing processes.
Patent Information
- Application Number
- PCT/JP2024/045580
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-27
- Filing Date
- 2024-12-24
- Publication Date
- 2025-07-03
AI Technical Summary
Existing encoding and decoding processes for three-dimensional mesh data, particularly related to displacement vectors, are inefficient and require improvements for effective transmission and storage.
A method and device that symbolize displacement vectors as samples within a rectangular region of a picture, generating a bitstream with position information associated with a hierarchy, allowing for accurate decoding by a decoding device.
Enables efficient encoding and decoding of three-dimensional mesh data by correctly identifying and decoding displacement vectors within a picture, reducing processing load and improving data transmission efficiency.
Smart Images

Figure JP2024045580_03072025_PF_FP_ABST
Abstract
Description
Encoding method, decoding method, encoding device, and decoding device
[0001] The present disclosure relates to encoding methods and the like.
[0002] In US Pat. No. 6,299,549 a method and apparatus for encoding and decoding three-dimensional mesh data is proposed.
[0003] Japanese Patent Application Laid-Open No. 2006-187015
[0004] Further improvements are desired in the encoding or decoding process for displacement vectors.The present disclosure aims to improve the encoding or decoding process for displacement vectors.
[0005] An encoding method according to one aspect of the present disclosure is an encoding method executed by an encoding device, which encodes displacement vectors as samples of a rectangular area within a picture, and generates a bitstream including the encoded samples and position information indicating the position of the rectangular area within the picture, the position information being associated with a layer corresponding to the rectangular area.
[0006] A decoding method according to one aspect of the present disclosure is a decoding method executed by a decoding device, which obtains a bitstream, obtains position information from the bitstream indicating the position of a rectangular area within a picture, and decodes a displacement vector indicated by a sample of the rectangular area located at a position indicated by the position information within the picture, wherein the position information is associated with a layer corresponding to the rectangular area.
[0007] These comprehensive or specific aspects may be realized as a system, an apparatus, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, an apparatus, an integrated circuit, a computer program, and a recording medium.
[0008] The present disclosure may contribute to improving encoding processes and the like related to displacement vectors.
[0009] 1 is a conceptual diagram showing a three-dimensional mesh according to an embodiment. FIG. 2 is a conceptual diagram showing basic elements of a three-dimensional mesh according to an embodiment. FIG. 3 is a conceptual diagram showing mapping according to an embodiment. FIG. 4 is a block diagram showing a configuration example of an encoding / decoding system according to an embodiment. FIG. 5 is a block diagram showing a configuration example of an encoding device according to an embodiment. FIG. 6 is a block diagram showing another configuration example of an encoding device according to an embodiment. FIG. 7 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 8 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 9 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 10 is a conceptual diagram showing another configuration example of a bit stream according to an embodiment. FIG. 11 is a conceptual diagram showing yet another configuration example of a bit stream according to an embodiment. FIG. 12 is a block diagram showing a specific example of an encoding / decoding system according to an embodiment. FIG. 13 is a conceptual diagram showing an example configuration of point cloud data according to an embodiment. FIG. 14 is a conceptual diagram showing an example data file of point cloud data according to an embodiment. FIG. 15 is a conceptual diagram showing an example configuration of mesh data according to an embodiment. FIG. 16 is a conceptual diagram showing an example data file of mesh data according to an embodiment. FIG. 17 is a conceptual diagram showing types of three-dimensional data according to an embodiment. FIG. 18 is a block diagram showing an example configuration of a three-dimensional data encoder according to an embodiment. FIG. 19 is a block diagram showing an example configuration of a three-dimensional data decoder according to an embodiment. FIG. 19 is a block diagram showing another configuration example of a three-dimensional data encoder according to an embodiment. FIG. 19 is a block diagram showing another configuration example of a three-dimensional data decoder according to an embodiment. 1 is a conceptual diagram showing a specific example of encoding processing according to an embodiment. FIG. 2 is a conceptual diagram showing a specific example of decoding processing according to an embodiment. FIG. 3 is a block diagram showing an implementation example of an encoding device according to an embodiment. FIG. 4 is a block diagram showing an implementation example of a decoding device according to an embodiment. FIG. 5 is a block diagram showing another configuration example of an encoding / decoding system according to an embodiment. FIG. 6 is a block diagram showing another configuration example of an encoding device according to an embodiment. FIG. 7 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 8 is a block diagram showing another configuration example of an encoding device according to an embodiment. FIG. 9 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 10 is a flow diagram showing processing of an encoding device according to an embodiment. FIG. 11 is an explanatory diagram conceptually showing encoding of a mesh frame according to an embodiment. FIG. 12 is a flow diagram showing processing of a decoding device according to an embodiment.1 is an explanatory diagram conceptually illustrating decoding of a mesh frame according to an embodiment. FIG. 2 is a block diagram showing an example of a configuration of a decoding device according to an embodiment. FIG. 3 is a block diagram showing an example of a configuration of a decoding device according to an embodiment. FIG. 4 is an explanatory diagram showing an example of subdivision according to an embodiment. FIG. 5 is an explanatory diagram showing an example of displacement of vertices after displacement after subdivision according to an embodiment. FIG. 6 is an explanatory diagram showing an example of vertices of an original mesh according to an embodiment. FIG. 7 is an explanatory diagram showing an example of a mesh according to an embodiment. FIG. 8 is an explanatory diagram showing an example of division of a mesh into submeshes according to an embodiment. FIG. 9 is a first explanatory diagram showing an example of packing of displacement information into an image frame according to an embodiment. FIG. 10 is a second explanatory diagram showing an example of packing of displacement information into an image frame according to an embodiment. FIG. 11 is a third explanatory diagram showing an example of packing of displacement information into an image frame according to an embodiment. FIG. 12 is a diagram showing an example of a configuration of an encoding device in the case of division into multiple submeshes according to an embodiment. FIG. 13 is a diagram showing an example of a configuration of a preprocessor according to an embodiment. FIG. 14 is a diagram showing an example of a configuration of a decoding device in the case of division into multiple submeshes according to an embodiment. FIG. 15 is a diagram for explaining post-decoding processing according to an embodiment. FIG. 16 is a flowchart showing an example of an encoding method by an encoding device according to an embodiment. FIG. 17 is a diagram showing an example of sample positions according to an embodiment. FIG. 18 is a diagram showing an example of sample positions in absolute values according to an embodiment. FIG. 19 is a diagram showing an example of sample positions in relative values according to an embodiment. FIG. 1 shows an example where sample positions are signaled in the body of a bitstream according to an embodiment. FIG. 2 shows an example where sample positions are signaled in a header of a bitstream according to an embodiment. FIG. 3 shows an example of LoD indices associated with rectangular regions according to an embodiment. FIG. 4 is a flowchart showing an example of a decoding method by a decoding device according to an embodiment. FIG. 5 shows a syntax for decoding sample positions according to an embodiment. FIG. 6 shows a syntax for decoding sample positions including height and width of a rectangular region according to an embodiment. FIG. 7 is a diagram for explaining an example of a decoding process for a rectangular region according to an embodiment. FIG. 8 is a diagram for explaining another example of a decoding process for a rectangular region according to an embodiment. FIG. 9 shows an example where different decoded rectangular regions are arranged row-wise according to an embodiment.1 is a diagram showing an example in which different decoded rectangular areas are arranged in a mixed arrangement of rows and columns according to an embodiment. 2 is a diagram showing an example in which rectangular areas form an inverted L-shape in the horizontal direction according to an embodiment. 3 is a diagram showing an example of parameters notified by syntax according to an embodiment. 4 is a diagram showing an example in which different types of rectangular areas are included according to an embodiment. 5 is a diagram showing an example in which different types of rectangular areas are included in multiple rectangular slices according to an embodiment. 6 is a diagram showing an example of samples converted into displacement vectors according to an embodiment. 7 is a diagram showing the syntax of header example 1 according to an embodiment. 8 is a diagram showing the syntax of header example 2 according to an embodiment. 9 is an example of syntax in which position information is indicated according to an embodiment. 10 is a flowchart showing an example of lifting transform performed by an encoding device according to an embodiment. 11 is a flowchart showing an example of lifting transform performed by a decoding device according to an embodiment. 12 is a diagram showing an example of the configuration of an encoding device according to an embodiment. 13 is a flowchart showing an example of an encoding method performed by an encoding device according to an embodiment. 14 is a diagram showing an example of the configuration of a decoding device according to an embodiment. 15 is a flowchart showing an example of a decoding method performed by a decoding device according to an embodiment.
[0010] <Summary of the Disclosure> Three-dimensional (3D) meshes are used in computer graphics images, for example, which may be composed of multiple temporally distinct frames, each of which may be represented by a 3D mesh.
[0011] A 3D mesh is composed of vertex information indicating the positions of each of the vertices in 3D space, connectivity information indicating the connections between the vertices, and attribute information indicating the attributes of each vertex or face. Each face is constructed according to the connectivity between the vertices. Various computer graphics images can be expressed using such 3D meshes.
[0012] Furthermore, for transmission and storage of the 3D mesh, efficient encoding and decoding of the 3D mesh is expected. For efficient encoding and decoding of the 3D mesh, arithmetic coding and decoding may be used.
[0013] Further improvements are desired in the encoding or decoding process for three-dimensional data. The present disclosure aims to improve the encoding or decoding process for three-dimensional data.
[0014] Below, examples of inventions that can be obtained from the disclosure of this specification will be given, and the effects and the like that can be obtained from these inventions will be explained.
[0015] An encoding method according to a first aspect of the present disclosure is an encoding method executed by an encoding device, which encodes displacement vectors as samples of a rectangular area within a picture, and generates a bitstream including the encoded samples and position information indicating the position of the rectangular area within the picture, the position information being associated with a layer corresponding to the rectangular area.
[0016] This generates a bitstream including samples with coded displacement vectors and position information indicating the positions of rectangular areas within a picture, thereby informing a decoding device of the positions of the samples within the picture, thereby generating a bitstream that enables the decoding device to correctly decode the displacement vector of a desired sample within a picture based on the position information.
[0017] An encoding method according to a second aspect of the present disclosure is the encoding method according to the first aspect, wherein the position information indicates the top left position of the rectangular area.
[0018] This allows the decoding device to know the rectangular area within the picture in which the samples are placed.
[0019] An encoding method according to a third aspect of the present disclosure is the encoding method according to the first or second aspect, wherein the bitstream further includes the size of the rectangular region.
[0020] This allows the decoding device to know the rectangular area within the picture in which the samples are placed.
[0021] An encoding method according to a fourth aspect of the present disclosure is the encoding method according to the third aspect, wherein the size includes a width and a height of the rectangular area.
[0022] This allows the decoding device to know the rectangular area within the picture in which the samples are placed.
[0023] An encoding method according to a fifth aspect of the present disclosure is an encoding method according to any one of the first to fourth aspects, wherein the bitstream includes identification information for identifying the layer.
[0024] This allows the decoding device to know the layer to which the sample corresponds.
[0025] An encoding method according to a sixth aspect of the present disclosure is an encoding method according to any one of the first to fifth aspects, wherein the picture has a plurality of rectangular areas including the rectangular area, the plurality of rectangular areas each corresponding to a plurality of hierarchies, and the plurality of hierarchies includes the hierarchies.
[0026] An encoding method according to a seventh aspect of the present disclosure is an encoding method according to the sixth aspect, wherein the bitstream includes a plurality of pieces of position information including the position information, and the plurality of pieces of position information each indicate the positions of the plurality of rectangular areas.
[0027] This allows the decoding device to be informed of the positions of the multiple layers and multiple rectangular areas to which the multiple samples correspond.
[0028] An encoding method according to an eighth aspect of the present disclosure is an encoding method according to the sixth or seventh aspect, wherein the bitstream includes a plurality of sizes including the size, and the plurality of sizes respectively correspond to the plurality of rectangular regions.
[0029] This allows the decoding device to know the rectangular areas in which the samples are arranged within the picture.
[0030] An encoding method according to a ninth aspect of the present disclosure is an encoding method according to any one of the first to eighth aspects, wherein the rectangular area includes one or more CTUs (Coding Tree Units) within the picture.
[0031] An encoding method according to a tenth aspect of the present disclosure is an encoding method according to any one of the first to eighth aspects, wherein the rectangular region includes one or more tiles within the picture.
[0032] An encoding method according to an eleventh aspect of the present disclosure is an encoding method according to any one of the first to eighth aspects, wherein the rectangular region includes one or more slices within the picture.
[0033] An encoding method according to a twelfth aspect of the present disclosure is an encoding method according to any one of the first to eighth aspects, wherein the rectangular region includes one or more sub-pictures within the picture.
[0034] A decoding method according to a thirteenth aspect of the present disclosure is a decoding method executed by a decoding device, which obtains a bitstream, obtains position information from the bitstream indicating the position of a rectangular area within a picture, and decodes a displacement vector indicated by a sample of the rectangular area located at a position indicated by the position information within the picture, wherein the position information is associated with a layer corresponding to the rectangular area.
[0035] According to this, position information indicating the position of a rectangular area within a picture is obtained from the bitstream, so that a desired sample within a picture can be correctly decoded based on the position information.
[0036] A decoding method according to a fourteenth aspect of the present disclosure is the decoding method according to the thirteenth aspect, wherein the position information indicates the top left position of the rectangular area.
[0037] This allows the decoding device to correctly decode the displacement vector of the desired sample based on the specific position of the rectangular region.
[0038] A decoding method according to a fifteenth aspect of the present disclosure is the decoding method according to the thirteenth or fourteenth aspect, in which the bitstream further includes a size of the rectangular region.
[0039] This allows the decoding device to identify the rectangular area in which the sample is placed within the picture, and correctly decode the displacement vector of the desired sample.
[0040] A decoding method according to a sixteenth aspect of the present disclosure is the decoding method according to the fifteenth aspect, wherein the size includes a width and a height of the rectangular area.
[0041] This allows the decoding device to identify the rectangular area in which the sample is placed within the picture, and correctly decode the displacement vector of the desired sample.
[0042] A decoding method according to a seventeenth aspect of the present disclosure is a decoding method according to any one of the thirteenth to sixteenth aspects, in which the bitstream includes identification information for identifying the layer.
[0043] Therefore, the decoding device can identify the rectangular area corresponding to the layer to which the sample corresponds, and can correctly decode the displacement vector of the desired sample.
[0044] A decoding method according to an 18th aspect of the present disclosure is a decoding method according to any one of the 13th to 17th aspects, wherein the picture has a plurality of rectangular areas including the rectangular area, the plurality of rectangular areas each corresponding to a plurality of layers, and the plurality of layers includes the layer.
[0045] A decoding method according to a 19th aspect of the present disclosure is a decoding method according to the 18th aspect, wherein the bitstream includes a plurality of pieces of position information including the position information, and the plurality of pieces of position information each indicate the positions of the plurality of rectangular areas.
[0046] This allows the decoding device to identify the positions of multiple layers and multiple rectangular areas to which multiple samples correspond.
[0047] A decoding method according to the 20th aspect of the present disclosure is a decoding method according to the 18th or 19th aspect, wherein the bitstream includes a plurality of sizes including the size, and the plurality of sizes respectively correspond to the plurality of rectangular regions.
[0048] This allows the decoding device to identify multiple rectangular areas within a picture in which multiple samples are placed.
[0049] A decoding method according to a 21st aspect of the present disclosure is a decoding method according to any one of the 13th to 20th aspects, wherein the rectangular area includes one or more CTUs (Coding Tree Units) within the picture.
[0050] A decoding method according to a 22nd aspect of the present disclosure is a decoding method according to any one of the 13th to 20th aspects, wherein the rectangular region includes one or more tiles within the picture.
[0051] A decoding method according to a 23rd aspect of the present disclosure is a decoding method according to any one of the 13th to 20th aspects, wherein the rectangular area includes one or more slices within the picture.
[0052] A decoding method according to a 24th aspect of the present disclosure is the decoding method according to any one of the 13th to 20th aspects, wherein the rectangular area includes one or more sub-pictures within the picture.
[0053] A decoding method according to a 25th aspect of the present disclosure is a decoding method according to a 17th aspect, further comprising determining whether the hierarchical level indicated by the identification information is equal to or lower than a predetermined hierarchical level indicated by predetermined identification information, and performing the decoding based on the determination result that the hierarchical level indicated by the identification information is equal to or lower than the predetermined hierarchical level.
[0054] This allows the decoding device to decode rectangular areas at a predetermined hierarchy level or lower, thereby reducing the processing load.
[0055] A coding device according to a 26th aspect of the present disclosure comprises a circuit and a memory connected to the circuit, wherein the circuit, in operation, encodes a displacement vector as a sample of a rectangular area within a picture, and generates a bitstream including the encoded samples and position information indicating the position of the rectangular area within the picture, the position information being associated with a layer corresponding to the rectangular area.
[0056] This generates a bitstream including samples with coded displacement vectors and position information indicating the positions of rectangular areas within a picture, thereby informing a decoding device of the positions of the samples within the picture, thereby generating a bitstream that enables the decoding device to correctly decode the displacement vector of a desired sample within a picture based on the position information.
[0057] A decoding device according to a 27th aspect of the present disclosure comprises a circuit and a memory connected to the circuit, wherein the circuit, in operation, acquires a bitstream, acquires position information from the bitstream indicating the position of a rectangular area within a picture, and decodes a displacement vector indicated by a sample of the rectangular area located at a position indicated by the position information within the picture, wherein the position information is associated with a layer corresponding to the rectangular area.
[0058] According to this, position information indicating the position of a rectangular area within a picture is obtained from the bitstream, so that a desired sample within a picture can be correctly decoded based on the position information.
[0059] These comprehensive or specific aspects may be realized as a system, an apparatus, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, an integrated circuit, a computer program, or a recording medium.
[0060] Hereinafter, the embodiments will be specifically described with reference to the drawings.
[0061] The embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present invention. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concepts are described as optional components.
[0062] (Embodiment) In this embodiment, an encoding method, a decoding method, etc. will be described.
[0063] <Expressions and Terms> The following expressions and terms are used herein.
[0064] (1) Three-dimensional Mesh A three-dimensional mesh is a collection of multiple faces, and represents, for example, a three-dimensional object. A three-dimensional mesh is mainly composed of vertex information, connectivity information, and attribute information. A three-dimensional mesh may be expressed as a polygon mesh or a mesh. A three-dimensional mesh may also vary over time. A three-dimensional mesh may include metadata related to the vertex information, connectivity information, and attribute information, and may also include other additional information.
[0065] (2) Vertex Information Vertex information is information indicating a vertex. For example, the vertex information indicates the position of a vertex in a three-dimensional space. Furthermore, a vertex corresponds to a vertex of a face that constitutes a three-dimensional mesh. Vertex information may be expressed as "geometry." Furthermore, vertex information may be expressed as position information.
[0066] (3) Connection Information Connection information is information that indicates connections between vertices. For example, connection information indicates connections for forming faces or edges of a three-dimensional mesh. Connection information may be expressed as "Connectivity." Connection information may also be expressed as face information.
[0067] (4) Attribute Information Attribute information is information that indicates attributes of a vertex or a face. For example, attribute information indicates attributes such as a color, an image, and a normal vector associated with a vertex or a face. Attribute information may be expressed as "texture."
[0068] (5) Faces A face is an element that constitutes a three-dimensional mesh. Specifically, a face is a polygon on a plane in three-dimensional space. For example, a face can be defined as a triangle in three-dimensional space.
[0069] (6) Plane A plane is a two-dimensional plane in a three-dimensional space. For example, a polygon is formed on a plane, and multiple polygons are formed on multiple planes.
[0070] (7) Bitstream: A bitstream corresponds to coded information. A bitstream may also be referred to as a stream, a coded bitstream, a compressed bitstream, or a coded signal.
[0071] (8) Encoding and Decoding The term encoding may be substituted with terms such as storing, including, writing, describing, signaling, sending, notifying, saving, or compressing, and these terms may be interchangeable. For example, encoding information may mean including the information in a bitstream. Also, encoding information into a bitstream may mean encoding the information to generate a bitstream that includes the encoded information.
[0072] Additionally, the term "decode" may be replaced with terms such as "read," "decode," "read," "load," "derive," "obtain," "receive," "extract," "reconstruct," "reconstruct," "decompress," or "decompress," and these terms may be interchangeable. For example, decoding information may mean obtaining information from a bitstream. Decoding information from a bitstream may mean decoding the bitstream to obtain information contained in the bitstream.
[0073] (9) Ordinal Numbers In the description, ordinal numbers such as first and second may be assigned to components, etc. These ordinal numbers may be changed as appropriate. Furthermore, new ordinal numbers may be assigned to components, etc., or removed. Furthermore, these ordinal numbers may be assigned to elements in order to identify them, and may not correspond to a meaningful order.
[0074] <Three-dimensional mesh> Fig. 1 is a conceptual diagram showing a three-dimensional mesh according to this embodiment. A three-dimensional mesh is composed of multiple faces. For example, each face is a triangle. The vertices of these triangles are defined in three-dimensional space. The three-dimensional mesh then represents a three-dimensional object. Each face may have a color or an image.
[0075] FIG. 2 is a conceptual diagram showing the basic elements of a three-dimensional mesh according to this embodiment. A three-dimensional mesh is composed of vertex information, connection information, and attribute information. The vertex information indicates the positions of the vertices of a face in three-dimensional space. The connection information indicates the connections between the vertices. A face can be identified by the vertex information and connection information. In other words, a colorless three-dimensional object is formed in three-dimensional space by the vertex information and connection information.
[0076] The attribute information may be associated with a vertex or a face. The attribute information associated with a vertex may be expressed as "Attribute Per Point." The attribute information associated with a vertex may indicate an attribute of the vertex itself, or may indicate an attribute of a face connected to the vertex.
[0077] For example, a color may be associated with a vertex as attribute information. The color associated with a vertex may be the color of the vertex itself, or the color of a face connected to the vertex. The color of a face may be the average of multiple colors associated with multiple vertices of the face. Furthermore, a normal vector may be associated with a vertex or a face as attribute information. Such a normal vector can represent the front and back of a face.
[0078] A two-dimensional image may be associated with a surface as attribute information. The two-dimensional image associated with a surface is also expressed as a texture image or an "Attribute Map." Information indicating mapping between the surface and the two-dimensional image may be associated with the surface as attribute information. Information indicating such mapping may be expressed as mapping information, vertex information of a texture image, texture coordinates, or "Attribute UV Coordinate."
[0079] Furthermore, information such as color, image, and moving image used as attribute information may be expressed as "parametric space."
[0080] The attribute information allows texture to be reflected on the three-dimensional object. That is, a three-dimensional object having color is formed in three-dimensional space based on the vertex information, connection information, and attribute information.
[0081] In the above, the attribute information is associated with the vertices or faces, but it may also be associated with the edges.
[0082] 3 is a conceptual diagram illustrating mapping according to this embodiment. For example, a region of a two-dimensional image on a two-dimensional plane can be mapped onto a surface of a three-dimensional mesh in three-dimensional space. Specifically, coordinate information of the region in the two-dimensional image is associated with the surface of the three-dimensional mesh. As a result, an image of the mapped region in the two-dimensional image is reflected on the surface of the three-dimensional mesh.
[0083] By using the mapping, the 2D image used as attribute information can be separated from the 3D mesh. For example, in encoding the 3D mesh, the 2D image may be encoded by an image encoding method or a video encoding method.
[0084] <System Configuration> Fig. 4 is a block diagram showing an example of the configuration of a coding / decoding system according to this embodiment. In Fig. 4, the coding / decoding system includes a coding device 100 and a decoding device 200.
[0085] For example, the encoding device 100 obtains a three-dimensional mesh and encodes the three-dimensional mesh into a bitstream. Then, the encoding device 100 outputs the bitstream to the network 300. For example, the bitstream includes the encoded three-dimensional mesh and control information for decoding the encoded three-dimensional mesh. By encoding the three-dimensional mesh, information about the three-dimensional mesh is compressed.
[0086] The network 300 transmits a bitstream from the encoding device 100 to the decoding device 200. The network 300 may be the Internet, a wide area network (WAN), a local area network (LAN), or a combination of these. The network 300 is not necessarily limited to bidirectional communication, and may be a unidirectional communication network for terrestrial digital broadcasting, satellite broadcasting, or the like.
[0087] Furthermore, the network 300 can be replaced by a recording medium such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc (registered trademark)).
[0088] The decoding device 200 obtains a bitstream and decodes a three-dimensional mesh from the bitstream. By decoding the three-dimensional mesh, information about the three-dimensional mesh is expanded. For example, the decoding device 200 decodes the three-dimensional mesh according to a decoding method corresponding to the encoding method used by the encoding device 100 to encode the three-dimensional mesh. That is, the encoding device 100 and the decoding device 200 perform encoding and decoding according to encoding methods and decoding methods that correspond to each other.
[0089] The 3D mesh before encoding may also be referred to as an original 3D mesh, and the 3D mesh after decoding may also be referred to as a reconstructed 3D mesh.
[0090] 5 is a block diagram showing an example of the configuration of a coding device 100 according to this embodiment. For example, the coding device 100 includes a vertex information encoder 101, a connection information encoder 102, and an attribute information encoder 103.
[0091] The vertex information encoder 101 is an electrical circuit that encodes vertex information. For example, the vertex information encoder 101 encodes the vertex information into a bitstream according to a format defined for the vertex information.
[0092] The connection information encoder 102 is an electrical circuit that encodes the connection information, for example, the connection information encoder 102 encodes the connection information into a bitstream according to a format defined for the connection information.
[0093] The attribute information encoder 103 is an electric circuit that encodes the attribute information. For example, the attribute information encoder 103 encodes the attribute information into a bit stream in accordance with a format defined for the attribute information.
[0094] The vertex information, connectivity information, and attribute information may be coded using variable-length coding or fixed-length coding, such as Huffman coding or context-adaptive binary arithmetic coding (CABAC).
[0095] The vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be integrated together, or each of the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be further subdivided into multiple components.
[0096] 6 is a block diagram showing another example of the configuration of the encoding device 100 according to this embodiment. For example, the encoding device 100 includes a pre-processor 104 and a post-processor 105 in addition to the configuration shown in FIG.
[0097] The preprocessor 104 is an electrical circuit that performs processing before encoding the vertex information, connectivity information, and attribute information. For example, the preprocessor 104 may perform a conversion process, a separation process, a multiplexing process, or the like on the 3D mesh before encoding. More specifically, for example, the preprocessor 104 may separate the vertex information, connectivity information, and attribute information from the 3D mesh before encoding.
[0098] The post-processor 105 is an electrical circuit that performs processing after the vertex information, connection information, and attribute information are encoded. For example, the post-processor 105 may perform conversion processing, separation processing, multiplexing processing, or the like on the encoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 105 may multiplex the encoded vertex information, connection information, and attribute information into a bitstream. Furthermore, for example, the post-processor 105 may further perform variable-length coding on the encoded vertex information, connection information, and attribute information.
[0099] 7 is a block diagram showing an example of the configuration of a decoding device 200 according to this embodiment. For example, the decoding device 200 includes a vertex information decoder 201, a connection information decoder 202, and an attribute information decoder 203.
[0100] The vertex information decoder 201 is an electrical circuit that decodes vertex information. For example, the vertex information decoder 201 decodes vertex information from a bitstream according to a format defined for the vertex information.
[0101] The connection information decoder 202 is an electrical circuit that decodes the connection information, for example, the connection information decoder 202 decodes the connection information from the bitstream according to a format defined for the connection information.
[0102] The attribute information decoder 203 is an electric circuit that decodes the attribute information. For example, the attribute information decoder 203 decodes the attribute information from the bitstream in accordance with a format defined for the attribute information.
[0103] The vertex information, connection information, and attribute information may be decoded using variable length decoding or fixed length decoding, which may correspond to Huffman coding, context-adaptive binary arithmetic coding (CABAC), or the like.
[0104] The vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be integrated together, or each of the vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be further subdivided into multiple components.
[0105] 8 is a block diagram showing another example of the configuration of the decoding device 200 according to this embodiment. For example, the decoding device 200 includes a pre-processor 204 and a post-processor 205 in addition to the configuration shown in FIG.
[0106] The preprocessor 204 is an electrical circuit that performs processing before decoding the vertex information, connection information, and attribute information. For example, the preprocessor 204 may perform conversion processing, separation processing, multiplexing processing, or the like on the bitstream before decoding the vertex information, connection information, and attribute information.
[0107] More specifically, for example, the preprocessor 204 may separate a sub-bitstream corresponding to vertex information, a sub-bitstream corresponding to connectivity information, and a sub-bitstream corresponding to attribute information from the bitstream. Also, for example, the preprocessor 204 may perform variable-length decoding on the bitstream in advance before decoding the vertex information, connectivity information, and attribute information.
[0108] The post-processor 205 is an electrical circuit that performs processing after the vertex information, connection information, and attribute information are decoded. For example, the post-processor 205 may perform conversion processing, separation processing, multiplexing processing, or the like on the decoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 205 may multiplex the decoded vertex information, connection information, and attribute information onto a three-dimensional mesh.
[0109] <Bitstream> Vertex information, connection information, and attribute information are coded and stored in a bitstream. The relationship between this information and the bitstream is shown below.
[0110] 9 is a conceptual diagram showing an example of the configuration of a bitstream according to this embodiment. In this example, connection information, vertex information, and attribute information are integrated in the bitstream. For example, the connection information, vertex information, and attribute information may be included in a single file.
[0111] Furthermore, multiple portions of this information may be stored sequentially, such as a first portion of connection information, a first portion of vertex information, a first portion of attribute information, a second portion of connection information, a second portion of vertex information, a second portion of attribute information, etc. These multiple portions may correspond to multiple portions that are different in time, multiple portions that are different in space, or multiple different faces.
[0112] Furthermore, the storage order of the connection information, vertex information, and attribute information is not limited to the above example, and a storage order different from the above example may be used.
[0113] 10 is a conceptual diagram showing another example of the configuration of a bitstream according to this embodiment. In this example, a plurality of files are included in the bitstream, and connection information, vertex information, and attribute information are stored in different files. Here, a file containing connection information, a file containing vertex information, and a file containing attribute information are shown, but the storage format is not limited to this example. For example, two types of information among the connection information, vertex information, and attribute information may be included in one file, and the remaining type of information may be included in another file.
[0114] Alternatively, the information may be split and stored in more files. For example, multiple pieces of connectivity information may be stored in multiple files, multiple pieces of vertex information may be stored in multiple files, or multiple pieces of attribute information may be stored in multiple files. These multiple pieces may correspond to multiple temporally different pieces, multiple spatially different pieces, or multiple different faces.
[0115] Furthermore, the storage order of the connection information, vertex information, and attribute information is not limited to the above example, and a storage order different from the above example may be used.
[0116] 11 is a conceptual diagram showing another example of the configuration of a bitstream according to this embodiment. In this example, the bitstream is composed of multiple separable sub-bitstreams, and connection information, vertex information, and attribute information are stored in different sub-bitstreams.
[0117] Here, a sub-bitstream containing connection information, a sub-bitstream containing vertex information, and a sub-bitstream containing attribute information are shown, but the storage format is not limited to this example.
[0118] For example, two types of information among the connection information, vertex information, and attribute information may be included in one sub-bitstream, and the remaining type of information may be included in another sub-bitstream. Specifically, attribute information of a two-dimensional image or the like may be stored in a sub-bitstream that complies with an image coding method, separate from the sub-bitstreams of the connection information and vertex information.
[0119] Also, each sub-bitstream may include multiple files, and multiple pieces of connectivity information may be stored in multiple files, multiple pieces of vertex information may be stored in multiple files, or multiple pieces of attribute information may be stored in multiple files.
[0120] 9, 10, and 11, and a storage order different from the above examples may be used. For example, the vertex information, connection information, and attribute information may be stored in the bitstream in this order. Alternatively, the connection information, connection information, and attribute information may be stored in the bitstream in any of the following orders: connection information, attribute information, and vertex information; vertex information, attribute information, and connection information; attribute information, connection information, and vertex information; or attribute information, vertex information, and connection information.
[0121] Furthermore, each of the connection information, vertex information, and attribute information may be divided into a plurality of data, and the plurality of data may be stored in a cyclical or random order within the bitstream.
[0122] 12 is a block diagram showing a specific example of an encoding / decoding system according to this embodiment. In FIG. 12, the encoding / decoding system includes a three-dimensional data encoding system 110, a three-dimensional data decoding system 210, and an external connector 310.
[0123] The three-dimensional data encoding system 110 includes a controller 111, an input / output processor 112, a three-dimensional data encoder 113, a three-dimensional data generator 115, and a system multiplexer 114. The three-dimensional data decoding system 210 includes a controller 211, an input / output processor 212, a three-dimensional data decoder 213, a system demultiplexer 214, a presenter 215, and a user interface 216.
[0124] In the three-dimensional data encoding system 110, sensor data is input from a sensor terminal to a three-dimensional data generator 115. The three-dimensional data generator 115 generates three-dimensional data, such as point cloud data or mesh data, from the sensor data and inputs it to a three-dimensional data encoder 113.
[0125] For example, the three-dimensional data generator 115 generates vertex information, and generates connection information and attribute information corresponding to the vertex information. The three-dimensional data generator 115 may process the vertex information when generating the connection information and attribute information. For example, the three-dimensional data generator 115 may reduce the amount of data by deleting duplicate vertices, or may transform the vertex information (such as by shifting its position, rotating it, or normalizing it). The three-dimensional data generator 115 may also render the attribute information.
[0126] Furthermore, although the three-dimensional data generator 115 is a component of the three-dimensional data encoding system 110 in FIG. 12, it may be arranged externally and independently of the three-dimensional data encoding system 110.
[0127] The sensor terminal that provides the sensor data for generating the three-dimensional data may be, for example, a moving body such as an automobile, a flying object such as an airplane, a mobile terminal, a camera, etc. Furthermore, a distance sensor such as a LIDAR, a millimeter wave radar, an infrared sensor, or a range finder, a stereo camera, or a combination of multiple monocular cameras may also be used as the sensor terminal.
[0128] The sensor data may be the distance (position) of the object, monocular camera images, stereo camera images, color, reflectance, sensor attitude, orientation, gyro, sensing position (GPS information or altitude), speed, acceleration, sensing time, temperature, air pressure, humidity, or magnetism.
[0129] The three-dimensional data encoder 113 corresponds to the encoding device 100 shown in FIG. 5 and other figures. For example, the three-dimensional data encoder 113 encodes three-dimensional data to generate encoded data. The three-dimensional data encoder 113 also generates control information when encoding the three-dimensional data. The three-dimensional data encoder 113 then inputs the encoded data together with the control information to the system multiplexer 114.
[0130] The encoding method for the three-dimensional data may be an encoding method using geometry or an encoding method using a video codec. Here, the encoding method using geometry may also be referred to as a geometry-based encoding method. The encoding method using a video codec may also be referred to as a video-based encoding method.
[0131] The system multiplexer 114 multiplexes the encoded data and control information input from the 3D data encoder 113 to generate multiplexed data using a specified multiplexing method. The system multiplexer 114 may multiplex other media such as video, audio, subtitles, application data, or document files, or reference time information, along with the encoded data and control information of the 3D data. Furthermore, the system multiplexer 114 may multiplex attribute information related to the sensor data or the 3D data.
[0132] For example, the multiplexed data may have a file format for storage or a packet format for transmission. As these formats, ISOBMFF or a format based on ISOBMFF may be used. Also, MPEG-DASH, MMT, MPEG-2 TS Systems, RTP, or the like may be used.
[0133] The multiplexed data is then output as a transmission signal to the external connector 310 by the input / output processor 112. The multiplexed data may be transmitted as a transmission signal by wire or wirelessly. Alternatively, the multiplexed data is stored in an internal memory or a storage device. The multiplexed data may be transmitted to a cloud server via the Internet or may be stored in an external storage device.
[0134] For example, the transmission or storage of the multiplexed data is performed by a method according to the medium for transmission or storage, such as broadcasting or communication. The communication protocol may be http, ftp, TCP, UDP, IP, or a combination thereof. Furthermore, a pull-type communication method or a push-type communication method may be used.
[0135] For wired transmission, Ethernet (registered trademark), USB, RS-232C, HDMI (registered trademark), coaxial cable, etc. may be used. For wireless transmission, 3GPP (registered trademark), 3G / 4G / 5G defined by IEEE, wireless LAN, Wi-Fi, Bluetooth, or millimeter wave may be used. For broadcasting, for example, DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 may be used.
[0136] The sensor data may be input to the three-dimensional data generator 115 or the system multiplexer 114. The three-dimensional data or encoded data may be output as a transmission signal directly to the external connector 310 via the input / output processor 112. The transmission signal output from the three-dimensional data encoding system 110 is input to the three-dimensional data decoding system 210 via the external connector 310.
[0137] Furthermore, each operation of the three-dimensional data encoding system 110 may be controlled by a controller 111 that executes an application program.
[0138] In the three-dimensional data decoding system 210, a transmission signal is input to an input / output processor 212. The input / output processor 212 decodes multiplexed data having a file format or a packet format from the transmission signal and inputs the multiplexed data to a system demultiplexer 214. The system demultiplexer 214 obtains coded data and control information from the multiplexed data and inputs them to a three-dimensional data decoder 213. The system demultiplexer 214 may extract other media or reference time information from the multiplexed data.
[0139] The three-dimensional data decoder 213 corresponds to the decoding device 200 shown in Fig. 7 etc. For example, the three-dimensional data decoder 213 decodes three-dimensional data from the encoded data based on a predefined encoding method. The three-dimensional data is then presented to the user by the presenter 215.
[0140] Additionally, additional information such as sensor data may be input to the presenter 215. The presenter 215 may present three-dimensional data based on the additional information. Additionally, a user instruction may be input from a user terminal to the user interface 216. Then, the presenter 215 may present three-dimensional data based on the input instruction.
[0141] The input / output processor 212 may acquire the three-dimensional data and the encoded data from the external connector 310 .
[0142] Furthermore, each operation of the three-dimensional data decoding system 210 may be controlled by a controller 211 that executes an application program.
[0143] 13 is a conceptual diagram showing an example of the configuration of point cloud data according to this embodiment. The point cloud data is data of a group of points representing a three-dimensional object.
[0144] Specifically, a point cloud is made up of a plurality of points, and has position information indicating the three-dimensional coordinate position of each point and attribute information indicating the attribute of each point. The position information is also expressed as geometry.
[0145] The type of attribute information may be, for example, color, reflectance, etc. One point may be associated with attribute information of one type, one point may be associated with attribute information of multiple different types, or one point may be associated with attribute information having multiple values for the same type.
[0146] 14 is a conceptual diagram showing an example of a data file of point cloud data according to this embodiment. This example shows a case where there is a one-to-one correspondence between position information items and attribute information items, and shows position information and attribute information for N points that make up the point cloud data. In this example, the position information is information indicating a three-dimensional coordinate position using three axes, x, y, and z, and the attribute information is information indicating a color using RGB. A PLY file or the like can be used as a representative data file for point cloud data.
[0147] 15 is a conceptual diagram showing an example of the configuration of mesh data according to this embodiment. Mesh data is data used in CG (Computer Graphics) and the like, and is three-dimensional mesh data that shows the three-dimensional shape of an object using multiple surfaces. Each surface is also expressed as a polygon, and has a polygonal shape such as a triangle or a rectangle.
[0148] Specifically, a 3D mesh is composed of a plurality of points constituting a point cloud, as well as a plurality of edges and a plurality of faces. Each point is also expressed as a vertex or a position. Each edge corresponds to a line segment connected by two vertices. Each face corresponds to an area surrounded by three or more edges.
[0149] Furthermore, a three-dimensional mesh has position information indicating the three-dimensional coordinate positions of vertices. The position information is also expressed as vertex information or geometry. A three-dimensional mesh also has connection information indicating the relationship between multiple vertices that make up an edge or a face. The connection information is also expressed as connectivity. A three-dimensional mesh also has attribute information indicating the attributes of the vertices, edges, or faces. The attribute information in a three-dimensional mesh is also expressed as texture.
[0150] For example, the attribute information may indicate the color, reflectance, or normal vector for a vertex, edge, or face. The direction of the normal vector may represent the front and back of the face.
[0151] The mesh data may be stored in a data file format such as an object file.
[0152] 16 is a conceptual diagram showing an example of a data file of mesh data according to this embodiment. In this example, the data file includes position information G(1) to G(N) of N vertices that make up the three-dimensional mesh, and attribute information A1(1) to A1(N) of the N vertices. Also, in this example, M pieces of attribute information A2(1) to A2(M) are included. The attribute information items do not need to correspond one-to-one to vertices or faces. Furthermore, attribute information need not exist.
[0153] The connection information is represented by a combination of vertex indices. n[1, 3, 4] indicates a triangular face formed by three vertices, n=1, n=3, and n=4. Also, m[2, 4, 6] indicates that the attribute information of m=2, m=4, and m=6 corresponds to the three vertices, respectively.
[0154] Furthermore, the actual contents of the attribute information may be written in a separate file. A pointer to that content may be associated with a vertex, a face, or the like. For example, attribute information indicating an image for a face may be stored in a two-dimensional attribute map file. The file name of the attribute map and two-dimensional coordinate values in the attribute map may be written in attribute information A2(1) to A2(M). The method of specifying attribute information for a face is not limited to these methods, and any method may be used.
[0155] 17 is a conceptual diagram showing types of three-dimensional data according to this embodiment. Point cloud data and mesh data may represent static objects or dynamic objects. A static object is an object that does not change over time, and a dynamic object is an object that changes over time. A static object may correspond to three-dimensional data for any point in time.
[0156] For example, point cloud data for a given point in time may be referred to as a PCC frame, mesh data for a given point in time may be referred to as a mesh frame, and PCC frames and mesh frames may be simply referred to as frames.
[0157] The area of the object may be limited to a certain range, as in normal video data, or may not be limited, as in map data. The density of points or surfaces may be determined in various ways. Sparse point cloud data or sparse mesh data may be used, or dense point cloud data or dense mesh data may be used.
[0158] Next, encoding and decoding of a point cloud or a three-dimensional mesh will be described. The device, process, or syntax for encoding and decoding vertex information of a three-dimensional mesh in the present disclosure may be applied to encoding and decoding of a point cloud. The device, process, or syntax for encoding and decoding of a point cloud in the present disclosure may be applied to encoding and decoding vertex information of a three-dimensional mesh.
[0159] Furthermore, a device, process, or syntax for encoding and decoding attribute information of a point cloud in the present disclosure may be applied to encoding and decoding connectivity information or attribute information of a three-dimensional mesh.Furthermore, a device, process, or syntax for encoding and decoding connectivity information or attribute information of a three-dimensional mesh in the present disclosure may be applied to encoding and decoding attribute information of a point cloud.
[0160] Furthermore, at least some of the processing may be shared between the encoding and decoding of point cloud data and the encoding and decoding of mesh data, thereby reducing the scale of the circuit and software program.
[0161] 18 is a block diagram showing an example configuration of a three-dimensional data encoder 113 according to this embodiment. In this example, the three-dimensional data encoder 113 includes a vertex information encoder 121, an attribute information encoder 122, a metadata encoder 123, and a multiplexer 124. The vertex information encoder 121, the attribute information encoder 122, and the multiplexer 124 may correspond to the vertex information encoder 101, the attribute information encoder 103, the post-processor 105, etc. in FIG.
[0162] In this example, the three-dimensional data encoder 113 encodes the three-dimensional data according to a geometry-based encoding method, which takes into account the three-dimensional structure. In addition, in the geometry-based encoding method, attribute information is encoded using configuration information obtained in encoding the vertex information.
[0163] Specifically, first, vertex information, attribute information, and metadata included in three-dimensional data generated from sensor data are input to a vertex information encoder 121, an attribute information encoder 122, and a metadata encoder 123, respectively. Here, connectivity information included in the three-dimensional data may be treated in the same way as attribute information. In addition, in the case of point cloud data, position information may be treated as vertex information.
[0164] The vertex information encoder 121 encodes the vertex information into compressed vertex information and outputs the compressed vertex information as encoded data to the multiplexer 124. The vertex information encoder 121 also generates metadata for the compressed vertex information and outputs it to the multiplexer 124. The vertex information encoder 121 also generates configuration information and outputs it to the attribute information encoder 122.
[0165] The attribute information encoder 122 uses the configuration information generated by the vertex information encoder 121 to encode the attribute information into compressed attribute information and outputs the compressed attribute information as encoded data to the multiplexer 124. The attribute information encoder 122 also generates metadata of the compressed attribute information and outputs it to the multiplexer 124.
[0166] The metadata encoder 123 encodes compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The metadata encoded by the metadata encoder 123 may be used to encode vertex information and attribute information.
[0167] The multiplexer 124 multiplexes the compressed vertex information, the compressed vertex information metadata, the compressed attribute information, the compressed attribute information metadata, and the compressed metadata into a bitstream, and then inputs the bitstream to the system layer.
[0168] 19 is a block diagram showing an example configuration of a three-dimensional data decoder 213 according to this embodiment. In this example, the three-dimensional data decoder 213 includes a vertex information decoder 221, an attribute information decoder 222, a metadata decoder 223, and a demultiplexer 224. The vertex information decoder 221, the attribute information decoder 222, and the demultiplexer 224 may correspond to the vertex information decoder 201, the attribute information decoder 203, the preprocessor 204, and the like in FIG.
[0169] In this example, the three-dimensional data decoder 213 decodes three-dimensional data according to a geometry-based encoding method. The three-dimensional structure is taken into consideration in the decoding according to the geometry-based encoding method. Furthermore, in the decoding according to the geometry-based encoding method, attribute information is decoded using configuration information obtained in decoding vertex information.
[0170] Specifically, first, a bitstream is input from the system layer to a demultiplexer 224. The demultiplexer 224 separates compressed vertex information, compressed vertex information metadata, compressed attribute information, compressed attribute information metadata, and compressed metadata from the bitstream. The compressed vertex information and compressed vertex information metadata are input to a vertex information decoder 221. The compressed attribute information and compressed attribute information metadata are input to an attribute information decoder 222. The metadata is input to a metadata decoder 223.
[0171] The vertex information decoder 221 decodes vertex information from the compressed vertex information using metadata of the compressed vertex information. The vertex information decoder 221 also generates configuration information and outputs it to the attribute information decoder 222. The attribute information decoder 222 decodes attribute information from the compressed attribute information using the configuration information generated by the vertex information decoder 221 and the metadata of the compressed attribute information. The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 may be used to decode the vertex information and the attribute information.
[0172] Thereafter, the vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. Note that, for example, this metadata is metadata of the vertex information and attribute information, and can be used in an application program.
[0173] 20 is a block diagram showing another example configuration of the three-dimensional data encoder 113 according to the present embodiment. In this example, the three-dimensional data encoder 113 includes a vertex image generator 131, an attribute image generator 132, a metadata generator 133, a video encoder 134, a metadata encoder 123, and a multiplexer 124. The vertex image generator 131, the attribute image generator 132, and the video encoder 134 may correspond to the vertex information encoder 101 and the attribute information encoder 103 in FIG. 6 , etc.
[0174] In this example, the 3D data encoder 113 encodes the 3D data according to a video-based encoding method. In encoding according to the video-based encoding method, multiple 2D images are generated from the 3D data, and the multiple 2D images are encoded according to a video encoding method. Here, the video encoding method may be High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or the like.
[0175] Specifically, first, vertex information and attribute information included in three-dimensional data generated from sensor data are input to a metadata generator 133. The vertex information and attribute information are then input to a vertex image generator 131 and an attribute image generator 132, respectively. The metadata included in the three-dimensional data is then input to a metadata encoder 123. Here, connectivity information included in the three-dimensional data may be treated in the same way as attribute information. In the case of point cloud data, position information may be treated as vertex information.
[0176] The metadata generator 133 generates map information of a plurality of two-dimensional images from the vertex information and attribute information, and inputs the map information to the vertex image generator 131, the attribute image generator 132, and the metadata encoder 123.
[0177] The vertex image generator 131 generates a vertex image based on the vertex information and map information, and inputs the generated image to the video encoder 134. The attribute image generator 132 generates an attribute image based on the attribute information and map information, and inputs the generated image to the video encoder 134.
[0178] The video encoder 134 encodes the vertex images and attribute images into compressed vertex information and compressed attribute information, respectively, in accordance with a video encoding method, and outputs the compressed vertex information and compressed attribute information as encoded data to the multiplexer 124. The video encoder 134 also generates metadata for the compressed vertex information and metadata for the compressed attribute information, and outputs them to the multiplexer 124.
[0179] The metadata encoder 123 encodes the compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The compressible metadata includes map information. The metadata encoded by the metadata encoder 123 may also be used to encode vertex information and attribute information.
[0180] The multiplexer 124 multiplexes the compressed vertex information, the compressed vertex information metadata, the compressed attribute information, the compressed attribute information metadata, and the compressed metadata into a bitstream, and then inputs the bitstream to the system layer.
[0181] 21 is a block diagram showing another example configuration of the 3D data decoder 213 according to this embodiment. In this example, the 3D data decoder 213 includes a vertex information generator 231, an attribute information generator 232, a video decoder 234, a metadata decoder 223, and a demultiplexer 224. The vertex information generator 231, the attribute information generator 232, and the video decoder 234 may correspond to the vertex information decoder 201 and the attribute information decoder 203 in FIG. 8, etc.
[0182] In this example, the 3D data decoder 213 decodes the 3D data according to a video-based coding method. In the decoding according to the video-based coding method, a plurality of 2D images are decoded according to a video coding method, and 3D data is generated from the plurality of 2D images. Here, the video coding method may be High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or the like.
[0183] Specifically, first, a bitstream is input from the system layer to the demultiplexer 224. The demultiplexer 224 separates compressed vertex information, compressed vertex information metadata, compressed attribute information, compressed attribute information metadata, and compressed metadata from the bitstream. The compressed vertex information, compressed vertex information metadata, compressed attribute information, and compressed attribute information metadata are input to the video decoder 234. The compressed metadata is input to the metadata decoder 223.
[0184] The video decoder 234 decodes the vertex images in accordance with the video encoding method. At this time, the video decoder 234 decodes the vertex images from the compressed vertex information using the metadata of the compressed vertex information. Then, the video decoder 234 inputs the vertex images to the vertex information generator 231. The video decoder 234 also decodes the attribute images in accordance with the video encoding method. At this time, the video decoder 234 decodes the attribute images from the compressed attribute information using the metadata of the compressed attribute information. Then, the video decoder 234 inputs the attribute images to the attribute information generator 232.
[0185] The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 includes map information used to generate vertex information and attribute information. The metadata decoded by the metadata decoder 223 may also be used to decode vertex images and attribute images.
[0186] The vertex information generator 231 reproduces vertex information from the vertex image in accordance with the map information included in the metadata decoded by the metadata decoder 223. The attribute information generator 232 reproduces attribute information from the attribute image in accordance with the map information included in the metadata decoded by the metadata decoder 223.
[0187] Thereafter, the vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. Note that, for example, this metadata is metadata of the vertex information and attribute information, and can be used in an application program.
[0188] Fig. 22 is a conceptual diagram showing a specific example of encoding processing according to this embodiment. Fig. 22 shows a three-dimensional data encoder 113 and a description encoder 148. In this example, the three-dimensional data encoder 113 includes a two-dimensional data encoder 141 and a mesh data encoder 142. The two-dimensional data encoder 141 includes a texture encoder 143. The mesh data encoder 142 includes a vertex information encoder 144 and a connection information encoder 145.
[0189] The vertex information encoder 144, the connection information encoder 145, and the texture encoder 143 may correspond to the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 in FIG.
[0190] For example, the two-dimensional data encoder 141 operates as a texture encoder 143 and generates a texture file by encoding the texture corresponding to the attribute information as two-dimensional data according to an image encoding method or a video encoding method.
[0191] The mesh data encoder 142 also operates as a vertex information encoder 144 and a connectivity information encoder 145, and generates a mesh file by encoding the vertex information and connectivity information. The mesh data encoder 142 may further encode mapping information for textures. The encoded mapping information may then be included in the mesh file.
[0192] The description encoder 148 also generates a description file by encoding a description corresponding to metadata such as text data. The description encoder 148 may encode the description at the system layer. For example, the description encoder 148 may be included in the system multiplexer 114 of FIG. 12 .
[0193] The above operations generate a bitstream containing texture files, mesh files, and description files, which may be multiplexed into the bitstream in file formats such as glTF (Graphics Language Transmission Format) or USD (Universal Scene Description).
[0194] The three-dimensional data encoder 113 may include two mesh data encoders as the mesh data encoder 142. For example, one mesh data encoder encodes vertex information and connectivity information of a static three-dimensional mesh, and the other mesh data encoder encodes vertex information and connectivity information of a dynamic three-dimensional mesh.
[0195] Correspondingly, two mesh files may then be included in the bitstream: for example, one mesh file corresponding to a static 3D mesh and another mesh file corresponding to a dynamic 3D mesh.
[0196] Furthermore, the static three-dimensional mesh may be a three-dimensional mesh of an intraframe coded using intraprediction, and the dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe coded using interprediction. Furthermore, information on the dynamic three-dimensional mesh may be differential information between vertex information or connectivity information of the three-dimensional mesh of an intraframe and vertex information or connectivity information of the three-dimensional mesh of an interframe.
[0197] Fig. 23 is a conceptual diagram showing a specific example of the decoding process according to this embodiment. Fig. 23 shows a three-dimensional data decoder 213, a description decoder 248, and a renderer 247. In this example, the three-dimensional data decoder 213 includes a two-dimensional data decoder 241, a mesh data decoder 242, and a mesh reconstructor 246. The two-dimensional data decoder 241 includes a texture decoder 243. The mesh data decoder 242 includes a vertex information decoder 244 and a connectivity information decoder 245.
[0198] The vertex information decoder 244, the connection information decoder 245, the texture decoder 243, and the mesh reconstructor 246 may correspond to the vertex information decoder 201, the connection information decoder 202, the attribute information decoder 203, and the post-processor 205 in Fig. 8. The presenter 247 may correspond to the presenter 215 in Fig. 12.
[0199] For example, the two-dimensional data decoder 241 operates as a texture decoder 243, and decodes the texture corresponding to the attribute information from the texture file as two-dimensional data in accordance with an image coding method or a video coding method.
[0200] The mesh data decoder 242 also operates as a vertex information decoder 244 and a connectivity information decoder 245 to decode vertex information and connectivity information from the mesh file. The mesh data decoder 242 may further decode mapping information for textures from the mesh file.
[0201] The description decoder 248 also decodes descriptions corresponding to metadata such as text data from the description file. The description decoder 248 may decode the descriptions at the system layer. For example, the description decoder 248 may be included in the system demultiplexer 214 of FIG. 12 .
[0202] The mesh reconstructor 246 reconstructs a 3D mesh from the vertex information, connectivity information, and textures according to the description. The renderer 247 renders and outputs the 3D mesh according to the description.
[0203] Through the above operations, a 3D mesh is reconstructed and output from a bitstream containing a texture file, a mesh file, and a description file.
[0204] The three-dimensional data decoder 213 may include two mesh data decoders as the mesh data decoder 242. For example, one mesh data decoder decodes vertex information and connectivity information of a static three-dimensional mesh, and the other mesh data decoder decodes vertex information and connectivity information of a dynamic three-dimensional mesh.
[0205] Correspondingly, two mesh files may then be included in the bitstream: for example, one mesh file corresponding to a static 3D mesh and another mesh file corresponding to a dynamic 3D mesh.
[0206] Furthermore, the static three-dimensional mesh may be a three-dimensional mesh of an intraframe coded using intraprediction, and the dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe coded using interprediction. Furthermore, information on the dynamic three-dimensional mesh may be differential information between vertex information or connectivity information of the three-dimensional mesh of an intraframe and vertex information or connectivity information of the three-dimensional mesh of an interframe.
[0207] A dynamic 3D mesh coding method is sometimes called DMC (Dynamic Mesh Coding), and a video-based dynamic 3D mesh coding method is sometimes called V-DMC (Video-based Dynamic Mesh Coding).
[0208] The point cloud encoding method is sometimes called PCC (Point Cloud Compression). The point cloud video-based encoding method is sometimes called V-PCC (Video-based Point Cloud Compression). The point cloud geometry-based encoding method is sometimes called G-PCC (Geometry-based Point Cloud Compression).
[0209] <Implementation Example> Fig. 24 is a block diagram showing an implementation example of the encoding device 100 according to this embodiment. The encoding device 100 includes a circuit 151 and a memory 152. For example, multiple components of the encoding device 100 shown in Fig. 5 etc. are implemented by the circuit 151 and memory 152 shown in Fig. 24.
[0210] The circuit 151 is a circuit that performs information processing and is a circuit that can access the memory 152. For example, the circuit 151 is a dedicated or general-purpose electric circuit that encodes a three-dimensional mesh. The circuit 151 may be a processor such as a CPU. Alternatively, the circuit 151 may be a collection of multiple electric circuits.
[0211] The memory 152 is a dedicated or general-purpose memory that stores information used by the circuit 151 to encode the three-dimensional mesh. The memory 152 may be an electric circuit and may be connected to the circuit 151. The memory 152 may also be included in the circuit 151. The memory 152 may also be a collection of multiple electric circuits. The memory 152 may also be a magnetic disk, an optical disk, or the like, and may also be expressed as a storage, a recording medium, or the like. The memory 152 may also be a non-volatile memory or a volatile memory.
[0212] For example, the memory 152 may store a three-dimensional mesh or a bitstream, or may store a program for the circuit 151 to encode the three-dimensional mesh.
[0213] Note that the encoding device 100 does not necessarily have to implement all of the components shown in Figure 5 and the like, and does not necessarily have to perform all of the processes shown here. Some of the components shown in Figure 5 and the like may be included in another device, and some of the processes shown here may be executed by another device. Furthermore, the encoding device 100 may implement any combination of the components of the present disclosure, and may perform any combination of the processes of the present disclosure.
[0214] Fig. 25 is a block diagram showing an example implementation of a decoding device 200 according to this embodiment. The decoding device 200 includes a circuit 251 and a memory 252. For example, multiple components of the decoding device 200 shown in Fig. 7 and other figures are implemented by the circuit 251 and memory 252 shown in Fig. 25.
[0215] The circuit 251 is a circuit that performs information processing and is a circuit that can access the memory 252. For example, the circuit 251 is a dedicated or general-purpose electric circuit that decodes a three-dimensional mesh. The circuit 251 may be a processor such as a CPU. Alternatively, the circuit 251 may be a collection of multiple electric circuits.
[0216] The memory 252 is a dedicated or general-purpose memory that stores information for the circuit 251 to decode the 3D mesh. The memory 252 may be an electric circuit and may be connected to the circuit 251. The memory 252 may also be included in the circuit 251. The memory 252 may also be a collection of multiple electric circuits. The memory 252 may also be a magnetic disk, an optical disk, or the like, and may also be expressed as a storage, a recording medium, or the like. The memory 252 may also be a non-volatile memory or a volatile memory.
[0217] For example, the memory 252 may store a three-dimensional mesh or a bitstream, or may store a program for the circuit 251 to decode the three-dimensional mesh.
[0218] Note that the decoding device 200 does not necessarily have to implement all of the components shown in Figure 7 and the like, and does not necessarily have to perform all of the processes shown here. Some of the components shown in Figure 7 and the like may be included in another device, and some of the processes shown here may be executed by another device. Furthermore, the decoding device 200 may implement any combination of the components of the present disclosure, and may perform any combination of the processes of the present disclosure.
[0219] The encoding method and the decoding method including the steps performed by each component of the encoding device 100 and the decoding device 200 of the present disclosure may be executed by any device or system. For example, part or all of the encoding method and the decoding method may be executed by a computer including a processor, a memory, an input / output circuit, etc. In this case, the encoding method and the decoding method may be executed by the computer executing a program for causing the computer to execute the encoding method and the decoding method.
[0220] Alternatively, the program or the bitstream may be recorded on a non-transitory computer-readable recording medium such as a CD-ROM.
[0221] An example of a program may be a bitstream. For example, a bitstream including an encoded three-dimensional mesh includes syntax elements for causing the decoding device 200 to decode the three-dimensional mesh. The bitstream then causes the decoding device 200 to decode the three-dimensional mesh according to the syntax elements included in the bitstream. Thus, the bitstream may play a role similar to that of a program.
[0222] The bitstream may be an encoded bitstream containing the encoded 3D mesh, or may be a multiplexed bitstream containing the encoded 3D mesh and other information.
[0223] Furthermore, each component of the encoding device 100 and the decoding device 200 may be configured with dedicated hardware, general-purpose hardware that executes the above-mentioned programs, or a combination of these. The general-purpose hardware may be configured with a memory in which the programs are recorded and a general-purpose processor that reads and executes the programs from the memory. Here, the memory may be a semiconductor memory or a hard disk, and the general-purpose processor may be a CPU.
[0224] Furthermore, the dedicated hardware may be configured with a memory, a dedicated processor, etc. For example, the dedicated processor may execute the encoding method and the decoding method by referring to a memory for recording data.
[0225] Furthermore, as described above, each component of the encoding device 100 and the decoding device 200 may be an electric circuit. These electric circuits may form a single electric circuit as a whole, or each may be a separate electric circuit. Furthermore, these electric circuits may correspond to dedicated hardware, or may correspond to general-purpose hardware that executes the above-mentioned programs, etc. Furthermore, the encoding device 100 and the decoding device 200 may be implemented as an integrated circuit.
[0226] Furthermore, the encoding device 100 may be a transmitting device that transmits the three-dimensional mesh, and the decoding device 200 may be a receiving device that receives the three-dimensional mesh.
[0227] Displacement Encoding and Decoding The following terminology is used here by way of example:
[0228] (1) Image An image is a data unit made up of a set of pixels, and includes a picture or a block smaller than a picture. Images include both moving images and still images.
[0229] (2) Picture A picture is a unit of image processing that is made up of a set of pixels, and is also called a frame or field.
[0230] (3) Block A block is a processing unit consisting of a specific number of pixels. The term shown in the following example is also used for a block. The shape of a block is not particularly limited. A block may be, for example, a rectangular shape of M×N pixels or a square shape of M×M pixels. A block may also be a triangular shape, a circular shape, or another shape. Examples of blocks are as follows:
[0231] Slice, tile, or brick CTU, superblock, or basic division unit VPDU, processing division unit for hardware CU, processing block unit, prediction block unit (PU), or orthogonal transform block unit (TU) Sub-block
[0232] (4) Pixel or Sample A pixel or sample is the smallest point of an image, in other words, the smallest unit. Pixels or samples include not only pixels at integer positions, but also pixels at sub-pixel positions generated based on pixels at integer positions.
[0233] (5) Pixel Value or Sample Value: A pixel value or sample value is a unique value of a pixel. The pixel value or sample value may include a luma value, a chroma value, or an RGB gradation level, and may also include a depth value or a binary value of 0 or 1.
[0234] (6) Flags A flag indicates one or more bits. A flag is, for example, a parameter or index represented by two or more bits. A flag may indicate not only a value represented by a binary number, but also a value represented by a number other than a binary number.
[0235] (7) Signal: A signal is something that is symbolized or coded to transmit information. A signal includes a discrete digital signal or a continuous analog signal.
[0236] (8) Stream or Bit Stream A stream or bit stream is a digital data sequence that indicates the flow of digital data. A stream or bit stream may be a single stream, or may be configured to include multiple streams with multiple layers. A stream or bit stream may be transmitted by serial communication using a single transmission path, or may be transmitted by packet communication using multiple transmission paths.
[0237] (9) Difference: For scalar quantities, difference can include simple difference (x - y) and difference calculations, such as absolute difference (|x - y|), squared difference (x^2 - y^2), square root difference (√(x - y)), weighted difference (ax - b, where a and b are constants), or offset difference (x - y + a, where a is an offset).
[0238] (10) Sum. For scalar quantities, sums can include simple sum (x + y) and addition operations. Sum can also include absolute sum (|x + y|), sum of squares (x^2 + y^2), square root of sum (√(x + y)), weighted sum (ax + by, where a and b are constants), or offset sum (x + y + a, where a is an offset).
[0239] (11) "Based on" The expression "based on something" means that something other than that "something" may be taken into consideration. Also, "based on" can be used both when a direct result is obtained and when a result is obtained through an intermediate result.
[0240] (12) "Used" or "Using" The phrases "something was used" or "used something" mean that something other than the "something" may be taken into consideration. The phrases "used" or "used" may be used both in cases where a direct result is obtained and in cases where a result is obtained via an intermediate result.
[0241] (13) Prohibition "Prohibit" can be rephrased as "not permitted." Also, "not prohibited / prohibited" or "permitted / permitted" does not necessarily mean "obligation."
[0242] (14) "Restriction" or "Limitation" "Restriction" or "Limitation" can be rephrased as "not permitted / not allowed" or "not permitted / permitted." Furthermore, "prohibited / not prohibited" or "not permitted / permitted" does not necessarily mean "obligation." Furthermore, what is prohibited quantitatively or qualitatively may be either partial or total.
[0243] (15) Chroma The term chroma is an adjective, represented by the symbols Cb or Cr, that indicates that a sample array or a single sample represents one of the two color-difference signals associated with a primary color. The term chroma is sometimes used instead of the term chrominance.
[0244] (16) Luma The term luma is an adjective, represented by the symbols or subscripts Y or L, that indicates that a sample array or a single sample represents a monochrome signal for a primary color. The term luma is sometimes used instead of the term luminance.
[0245] The encoding / decoding system of this embodiment will be described below.
[0246] A typical three-dimensional model (also called a 3D model) digitally represents an object so that a user can explore the model using zoom, pan, and rotation in all three dimensions while it is rendered over time. One way to construct such a representation is to build a 3D mesh using triangles. The model stores the positions of the triangle vertices, their connectivity to each other, and their associated attributes (such as normals or UV patches).
[0247] Storing all this information in uncompressed form requires a very large storage space and therefore a very large bandwidth for transmission. The triangles that form the mesh often have repeating patterns and similar properties, especially in temporal and spatial neighborhoods. These repetitions can be exploited to develop efficient encoding and decoding methods for storage and transmission. One such encoding and decoding method is Video-based Dynamic Mesh Coding (V-DMC).
[0248] 26 is a block diagram showing another example of the configuration of the encoding / decoding system according to this embodiment. As shown in FIG. 26, the encoding / decoding system includes an encoding device 100 and a decoding device 200.
[0249] The encoding / decoding system accepts input three-dimensional meshes (also called 3D meshes) in the form of three-dimensional coordinates of vertices (vertex information), connectivity (connection information) and associated attributes (attribute information), which may include texture maps as well as geometry.
[0250] The encoding device 100 takes an input 3D mesh (also referred to as an input 3D mesh or input mesh) in the form of three-dimensional coordinates of vertices, connectivity, and associated attributes. The encoding device 100 encodes all associated information into a stream. The stream may consist of a single bitstream or multiple bitstreams.
[0251] The network 300 transmits the stream generated by the encoding device to the decoding device 200. The network 300 may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof. Furthermore, the network 300 is not necessarily limited to a two-way communication network, but may also be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Instead of the network 300, a recording medium such as a digital versatile disc (DVD) or a blue-ray disc (BD) on which a stream is recorded may be used.
[0252] The stream is transmitted to a decoding device 200 via a network 300. The decoding device 200 decodes the bitstream and generates a 3D mesh using the 3D coordinates, connectivity, and associated attributes of the decoded vertices. The decoding device 200 outputs the generated 3D mesh (also referred to as an output 3D mesh or output mesh).
[0253] FIG. 27 is a diagram showing another example of the configuration of the encoding device 100.
[0254] As shown in FIG. 27, the encoding device 100 includes a preprocessor 1103 and a compressor 1106 .
[0255] The encoding device 100 reads an input mesh 1101 and an attribute map 1102 and passes them to a preprocessor 1103. The preprocessor 1103 processes the input mesh to extract a base mesh 1104 and displacement data 1105. The attribute map 1102, along with the extracted base mesh 1104 and displacement data 1105, are passed to a compressor 1106.
[0256] The compressor 1106 also compresses the base mesh 1104, the displacement data 1105, and the attribute map 1102 to generate a bitstream 1107. The compressor 1106 can transmit additional information to the decoding device 200 by further including metadata 1108 in the bitstream 1107.
[0257] FIG. 28 is a diagram showing another example of the configuration of the decoding device 200.
[0258] As shown in FIG. 28, the decoding device 200 includes a decompressor 2102 and a post-processor 2106 .
[0259] The decoding device 200 reads a bitstream 2101 and passes it to a decompressor 2102. The decompressor 2102 decompresses a base mesh 2103, displacement data 2104, and an attribute map 2108 from the bitstream 2101 and passes them to a post-processor 2106. An example of the displacement data 2104 is a displacement vector.
[0260] The post-processor 2106 also processes the base mesh 2103 according to the displacement data 2104 and the attribute map 2108 to generate an output mesh 2107. The post-processor 2106 may further use information from the metadata 2105 to generate the output mesh 2107.
[0261] FIG. 29 is a block diagram showing yet another example configuration of the encoding device 100 according to this embodiment.
[0262] In this example, the encoding device 100 comprises a volumetric capturer 511, a projector 512, a base mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and optionally one or more other type encoders 516.
[0263] The volumetric capturer 511 captures content and outputs the captured content to the projector 512 .
[0264] The projector 512 projects the content onto a 3D mesh frame containing vertex geometry coordinates, texture coordinates, and connectivity data. The data is output to a base mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and optionally one or more other type encoders 516. Each encoder compresses the data into a bitstream.
[0265] FIG. 30 is a block diagram showing yet another example configuration of the decoding device 200 according to this embodiment.
[0266] In this example, the decoding device 200 comprises a base mesh decoder 613 , a displacement decoder 614 , an attribute decoder 615 , one or more other type decoders 616 , and a 3D reconstructor 617 .
[0267] The bitstream is sent to a base mesh decoder 613, a displacement decoder 614, an attribute decoder 615, and optionally one or more other type decoders 616. These decoders decode the bitstream to generate decoded data including vertex geometry coordinates, texture coordinates, and connectivity data. The decoded data is then sent to a 3D reconstructor 617, which reconstructs a 3D mesh frame.
[0268] The encoding process performed by the encoding device 100 will be described in detail below.
[0269] Fig. 31 is a flow diagram showing the processing of the encoding device 100. Fig. 32 is an explanatory diagram conceptually showing the encoding of mesh frames. The processing of the encoding device 100 will be described with reference to Figs. 31 and 32.
[0270] In step S101, the encoding device 100 reads a 3D mesh frame, which is an input mesh frame, and its attributes. The input mesh frame is a mesh frame input to the encoding device 100. An example of the 3D mesh frame that is an input mesh frame is shown as mesh frame 1301 (see FIG. 32 ).
[0271] In step S102, the encoding device 100 performs a decimation process on the input mesh frame read in step S101 to generate a base mesh frame having fewer vertices than the input mesh frame. The base mesh frame generated by decimating the mesh frame 1301 is shown as a base mesh frame 1302 (see FIG. 32).
[0272] In step S103, the encoding device 100 calculates displacement information that the decoding device 200 uses to reconstruct a mesh frame. The displacement information corresponds to a displacement vector directed from a vertex of the base mesh frame generated in step S102 to a vertex of the input mesh frame. One method for calculating the displacement information is to subtract the coordinates of the vertex of the base mesh frame from the coordinates of the vertex of the input mesh frame. The displacement information calculated from the mesh frame 1301 and the base mesh frame 1302 is shown as displacement information 1303 (see FIG. 32). The displacement information 1303 is in vector format, in other words, expressed as a displacement vector.
[0273] In step S104, the encoding device 100 encodes the base mesh frame generated in step S102, the displacement information generated in step S103, and the attributes of the input mesh frame into a bitstream (corresponding to a compressed bitstream). An example of the bitstream is shown as bitstream 1304 (see FIG. 32).
[0274] Specifically, the bitstream 1304 includes vertex coordinates and connectivity information for vertices A, C, E, and F, displacement information, a video bitstream including texture data, and a compressed attribute map (see FIG. 32). The displacement information includes displacement information for displacing vertices based on vertex coordinates obtained from the subdivided base mesh frame. The compressed attribute map is texture coordinates for applying texture data to a mesh frame reconstructed using the base mesh frame and the displacement information.
[0275] The decoding process performed by the decoding device 200 will be described in detail below.
[0276] Fig. 33 is a flow diagram showing the processing of the decoding device 200. Fig. 34 is an explanatory diagram conceptually showing the decoding of a 3D mesh. The processing of the decoding device 200 will be described with reference to Figs. 33 and 34.
[0277] In step S201, the decoding device 200 decodes a base mesh frame and attributes from a bitstream (corresponding to a compressed bitstream). An example of the decoded base mesh frame (corresponding to a decoded base mesh frame) is shown as a decoded base mesh frame 2301 (see FIG. 34).
[0278] In step S202, the decoding device 200 generates subdivided vertices by performing a subdivision process on the base mesh frame decoded in step S201. An example of a base mesh frame including subdivided vertices is shown as base mesh frame 2302 (see FIG. 34).
[0279] In step S203, the decoding device 200 decodes the disparity information from the bitstream (corresponding to the compressed bitstream). An example of the decoded disparity information is shown as disparity information 2303 (see FIG. 34). The disparity information 2303 is in vector format, in other words, expressed as a disparity vector.
[0280] In step S204, the decoding device 200 reconstructs the shape of the mesh frame by moving the vertices of the base mesh frame, including the subdivided vertices, to new positions using the displacement information, and then restores the mesh frame by applying attribute information. An example of the attribute is texture. An example of the reconstructed mesh frame is shown as mesh frame 2304 (see FIG. 34 ).
[0281] FIG. 35 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.
[0282] FIG. 35 shows an example of a block diagram of a general intra-decoding system.
[0283] The decoding device shown in FIG. 35 comprises a demultiplexer 1231, a switch 1232, a static mesh decoder 1233, a mesh buffer 1234, a motion decoder 1235, a base mesh reconstructor 1236, an inverse quantizer 1237, a video decoder 1238, an image unpacker 1239, an inverse quantizer 1240, an inverse wavelet transformer 1241, a reconstructor 1242, a video decoder 1243, and a color converter 1244.
[0284] The demultiplexer 1231 receives the compressed bitstream and separates it into compressed data for the base mesh, video containing displacement data (also called displacement bitstream), and video containing attribute data (also called attribute bitstream). The compressed data for the base mesh is passed to a switch 1232. The switch 1232 determines whether to perform intra-decoding or inter-decoding based on parameters in the bitstream.
[0285] If an intra-decoding process is selected, the bitstream is passed to a static mesh decoder 1233, which generates a quantized base mesh. The static mesh decoder 1233 is, for example, a decoder that uses an edge breaker algorithm to decode 3D mesh data. The static mesh decoder 1233 generates a quantized base mesh from the bitstream. The quantized base mesh generated by the static mesh decoder 1233 is stored in a mesh buffer 1234 for reference when an inter-decoding process is selected.
[0286] If inter-decoding is selected, switch 1232 passes compressed data for the base mesh to motion decoder 1235. Motion decoder 1235 receives a previously decoded quantized base mesh and decodes motion data representing the differences in vertex coordinates between the quantized base mesh stored in mesh buffer 1234 and the current quantized base mesh. The motion data and the quantized base mesh stored in mesh buffer 1234 are used by base mesh reconstructor 1236 to reconstruct the current quantized base mesh. The quantized base mesh resulting from either inter-decoding or intra-decoding is passed to inverse quantizer 1237 to obtain a decoded base mesh.
[0287] The video containing the displacement data is passed to a video decoder 1238, since the bitstream contains the displacement data in an image format with two chroma information and one luma information. The video decoder 1238 decodes the data using a video frame decompression method. Alternatively, the displacement data can be decoded using an arithmetic decoder. This decompressed data is passed to an image unpacker 1239, which extracts wavelet coefficients associated with each vertex from the image-format decompressed data. An inverse quantizer 1240 dequantizes the quantized wavelet coefficients into the three components associated with each vertex. An inverse wavelet transformer 1241 inversely transforms the result to finally obtain decoded displacement data. The decoded displacement data and the decoded base mesh are passed to a reconstructor 1242, which performs edge refinement on the decoded base mesh and displaces the vertices using the decoded displacement data to obtain a decoded mesh.
[0288] The video containing the attribute data is passed to another video decoder 1243 to obtain a decoded attribute bitstream, which is further processed in a color converter 1244 for color space and color format conversion to obtain a decoded attribute map.
[0289] FIG. 36 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.
[0290] FIG. 36 illustrates an example of a reconstructor that obtains a decoded 3D mesh 1256 from a decoded base mesh 1251 and decoded displacement data 1254.
[0291] The decoded base mesh 1251 is passed to a subdivider 1252 .
[0292] The subdivision unit 1252 subdivides any two connected vertices in the entire 3D mesh by adding a new vertex between them. This process can be repeated several times to include vertices created in previous subdivision steps to generate a predefined number of vertices. Each subdivision iteration across the 3D mesh generates a new level of detail (LoD). The subdivided mesh 1253 and the decoded displacement data 1254 are passed to a displacer 1255. The displacer 1255 generates a decoded 3D mesh 1256 by moving each vertex to a new position according to the corresponding displacement data.
[0293] The subdivision is described below and is performed by a subdivider (specifically subdivider 1206 or subdivider 2204).
[0294] FIG. 37 is an explanatory diagram showing an example of subdivision.
[0295] The base mesh shown in FIG. 37(a) includes vertices A, B, and C and connectivity information indicating their connectivity.
[0296] 37(b) shows a mesh generated by the first subdivision, in other words, the mesh after the first subdivision. In the first subdivision, the subdivider generates vertices D, E, and F and connectivity information indicating their connectivity. The mesh generated by the subdivider is also referred to as LoD1 or first LoD.
[0297] Vertex D of the mesh after the first subdivision is a vertex generated by subdivision based on vertices A and B. Similarly, vertex E is a vertex generated by subdivision based on vertices B and C. Vertex F is a vertex generated by subdivision based on vertices A and C.
[0298] As an example, vertex D may be the midpoint of line segment AB (in other words, side AB) connecting vertices A and B that were the basis for its generation. Similarly, vertex E may be the midpoint of line segment AC. Vertex F may be the midpoint of line segment BC.
[0299] 37(c) shows the mesh generated by the second subdivision, i.e., the mesh after the second subdivision. In the second subdivision, the subdivider generates vertices G, H, I, J, K, L, M, N, and O and connectivity information indicating their connectivity. The mesh generated by the subdivider is also called LoD2 or second LoD.
[0300] Vertex G of the mesh after the second subdivision is a vertex generated by subdivision based on vertices A and D. Similarly, vertex H is a vertex generated by subdivision based on vertices A and E. Vertex I is a vertex generated by subdivision based on vertices B and D. Vertex J is a vertex generated by subdivision based on vertices D and F. Vertex K is a vertex generated by subdivision based on vertices E and F. Vertex L is a vertex generated by subdivision based on vertices C and E. Vertex M is a vertex generated by subdivision based on vertices B and F. Vertex N is a vertex generated by subdivision based on vertices C and F. Vertex O is a vertex generated by subdivision based on vertices D and E.
[0301] As an example, vertex G may be the midpoint of line segment AD (in other words, side AD) connecting vertices A and D, which were the source of its generation. Similarly, vertex H may be the midpoint of line segment AE. vertex I may be the midpoint of line segment BD. vertex J may be the midpoint of line segment DF. vertex K may be the midpoint of line segment EF. vertex L may be the midpoint of line segment CE. vertex M may be the midpoint of line segment BF. vertex N may be the midpoint of line segment CF. vertex O may be the midpoint of line segment DE.
[0302] The displacement of vertices will be described below with reference to Figures 38 and 39. The displacement of vertices is performed by the reconstructor 2209.
[0303] Fig. 38 is an explanatory diagram showing an example of displacement of vertices after subdivision, and Fig. 39 is an explanatory diagram showing an example of vertices of an original mesh.
[0304] The base mesh shown in FIG. 38(a) includes vertices A, B, C, and Z and connectivity information indicating their connectivity.
[0305] 38(b) shows a mesh generated by the first subdivision, in other words, a mesh after the first subdivision (i.e., the first LoD). In the first subdivision, the subdivider generates vertices S, T, U, X, or Y and connectivity information indicating their connectivity. The vertices S, T, U, X, or Y are similar to the vertices D, E, and F shown in FIG. 37(b).
[0306] 38(c) shows a mesh generated by the second subdivision, in other words, a mesh after the second subdivision (i.e., the second LoD). In the second subdivision, the subdivider generates vertices D, E, F, G, and H and connectivity information indicating their connectivity. Vertices D, E, F, G, and H are the same as vertices G, H, I, J, K, L, M, N, and O shown in FIG. 37(c).
[0307] Figure 38(d) shows a mesh including the vertices after they have been displaced after subdivision, with vertices A, B, C, D, E, F, G, H, S, T, U, X, Y, and Z shown in Figure 38(d) being located at positions displaced using displacement information from the positions of the vertices shown in Figure 38(c).
[0308] The original mesh shown in FIG. 39 is an example of the mesh input to the encoding device 100, that is, the mesh before encoding.
[0309] The mesh shown in Fig. 38 has a shape similar to that of the original mesh shown in Fig. 39. The displacement information is generated by the displacement vector calculator 1207 of the encoding device 100 as information indicating the displacement from the vertices of the base mesh to the vertices of the original mesh, and therefore, by reconstructing the mesh using the displacement information thus generated, a mesh having a shape similar to that of the original mesh is generated.
[0310] The decoding device 200 can output the mesh shown in FIG.
[0311] Next, the division of a mesh into sub-meshes will be described with reference to FIGS.
[0312] A mesh can be divided into smaller parts and coded separately, with the vertices of the mesh being divided in such a way that the coordinates and connectivity of the vertices in each part can be coded independently.
[0313] Fig. 40 is an explanatory diagram showing an example of a mesh, and Fig. 41 is an explanatory diagram showing an example of dividing a mesh into sub-meshes.
[0314] The mesh shown in FIG. 40 is the original mesh, which is sometimes called a full mesh in contrast to a sub-mesh.
[0315] Figure 41 shows how the full mesh shown in Figure 40 is divided into two sub-meshes. For vertices A, B, and C of the full mesh (see Figure 40), vertex A is duplicated to vertices A1 and A2, vertex B is duplicated to vertices B1 and B2, and vertex C is duplicated to vertices C1 and C2, thereby creating two sub-meshes (i.e., a first sub-mesh and a second sub-mesh) from the full mesh. The first sub-mesh and the second sub-mesh are each independently decodable meshes.
[0316] Packing of displacement information into image frames will be described below with reference to FIGS.
[0317] 42, 43 and 44 are explanatory diagrams showing examples of packing of displacement information into image frames. Note that image frames can also be called video frames.
[0318] The vertex displacement data is encoded as image frame data by being mapped to each component of a YUV format image frame (i.e., each of the Y component (Y Plane), U component (U Plane), and V component (V Plane)). This case will be described below as an example. As another example, the vertex displacement data may be encoded as image frame data by being mapped to each component of an RGB format image frame (each of the R component, G component, and B component).
[0319] The decoding device 200 can use an image encoding module to extract the displacement data. The displacement data can be in the form of X, Y, or Z components in a global coordinate system (e.g., a Cartesian coordinate system), or normal, tangential, or both tangential components in a local coordinate system. Methods for mapping the displacement data to an image frame include the following:
[0320] For example, in the first method, the displacement data is arranged in the image frame in scan order, and an example of packing the displacement data in this case is shown in Figure 42. The displacement data is directly mapped onto the image frame according to a predefined scan order.
[0321] Note that since an image frame has a fixed height and width, it may happen that the displacement data does not fit perfectly in the frame, in which case the remaining part of the image frame is padded with padding data (see Figure 42).
[0322] For example, in the second method, the displacement data is separated into multiple LoDs and mapped to the Y, U, and V components of the image frame. An example of packing of the displacement data in this case is shown in Figure 43. Here, the displacement data of the image frame of the next LoD starts immediately after the displacement data of the previous LoD ends. As in the first method, if the displacement data does not fit exactly into the image frame, padding is performed at the end of the image frame (see Figure 43).
[0323] For example, in the third method, displacement data corresponding to the LoD is mapped to the Y component, U component, and V component of the image frame in a manner different from that in the second method. An example of packing of the displacement data in this case is shown in Figure 44. In this way, each LoD can be decoded independently. In the third method, middle padding is performed on the displacement data of each LoD, and CTU alignment is performed together with padding at the end of the video frame (see Figure 44).
[0324] <Other Examples> Although the aspects of the encoding device and the decoding device have been described above according to the embodiments, the aspects of the encoding device and the decoding device are not limited to the embodiments. Modifications that a person skilled in the art can conceive of may be applied to the embodiments, and multiple components in the embodiments may be combined in any manner.
[0325] For example, a process performed by a specific component in the embodiment may be performed by another component instead of the specific component. Also, the order of multiple processes may be changed, or multiple processes may be performed in parallel.
[0326] Furthermore, as described above, at least some of the configurations of the present disclosure may be implemented as an integrated circuit. At least some of the processes of the present disclosure may be used as an encoding method or a decoding method. A program for causing a computer to execute the encoding method or the decoding method may be used. A non-transitory computer-readable recording medium on which the program is recorded may be used. A bitstream for causing a decoding device to perform a decoding process may be used.
[0327] Furthermore, at least some of the configurations and processes of the present disclosure may be used as a transmitting device, a receiving device, a transmitting method, or a receiving method. A program for causing a computer to execute the transmitting method or the receiving method may be used. Furthermore, a non-transitory computer-readable recording medium on which the program is recorded may be used.
[0328] In the above embodiments, each component may be configured with dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory. Here, the software that realizes the encoding device and the like in the above embodiments is the following program.
[0329] In other words, the software is a program that causes a computer to execute an encoding method that obtains a first frame to be encoded, encodes first data having one or more layers contained in the first frame by referencing second data having one or more layers contained in a second frame, and when encoding the first data, determines one of a plurality of processes for each of the one or more layers contained in the first data using a value indicating the layer and a value related to the one or more layers contained in the second data, and executes the determined one process on the first data of that layer.
[0330] The software for realizing the decoding device and the like according to the above-described embodiment is the following program.
[0331] In other words, the software is a program that causes a computer to execute a decoding method in which it obtains a first frame to be decoded, decodes first data having one or more layers contained in the first frame by referring to second data having one or more layers contained in a second frame, and when decoding the first data, determines one of a plurality of processes for each of the one or more layers contained in the first data using a value indicating the layer and a value related to the one or more layers contained in the second data, and executes the determined one process on the first data of that layer.
[0332] Next, the configuration of an encoding device when a mesh is divided into a plurality of sub-meshes will be described.
[0333] Fig. 45 is a diagram showing an example of the configuration of an encoding device when dividing into a plurality of sub-meshes. Fig. 46 is a diagram showing an example of the configuration of a pre-processor.
[0334] An input mesh 1401 is divided into a plurality of sub-meshes 1402. The input mesh 1401 is, for example, a full mesh. Each of the sub-meshes 1402 is input to an encoding device 1400 shown in FIG. 45 . That is, the encoding process by the encoding device 1400 is performed after the input mesh 1401 is divided into the plurality of sub-meshes 1402.
[0335] As shown in FIG. 45, the encoding device 1400 includes a preprocessor 1404 and a compressor 1408 .
[0336] The encoding device 1400 reads one of the submeshes 1402 that corresponds to the encoding device 1400 and an attribute map 1403, and passes them to a preprocessor 1404. The preprocessor 1404 processes the submesh 1402 to extract a base mesh 1405 and displacement data 1406.
[0337] As shown in FIG. 46, the preprocessor 1404 includes a base mesh generator 1411 , a subdivision unit 1412 , and a displacement generator 1413 .
[0338] The base mesh generator 1411 generates a base mesh based on the mesh or sub-mesh 1402 .
[0339] The subdivision unit 1412 subdivides the base mesh using a predetermined method to generate a subdivided mesh (Subdivided Mesh, Subdivided Base Mesh).
[0340] The displacement generator 1413 generates displacement data 1406 based on the subdivided mesh and the submesh 1402. The displacement data is, for example, a difference vector (displacement vector) between the input mesh (submesh 1402) and the subdivided mesh. The displacement vector corresponds to the vertices of the input mesh.
[0341] Note that the decoding device also uses the same subdivision method as that used in the encoding device 1400. The subdivision method or parameters used in the encoding device 1400 may be transmitted to the decoding device 1420. That is, the subdivision method or parameters may be included in the bitstream.
[0342] The attribute map 1403 is passed to a compressor 1408 along with a base mesh 1405 and displacement data 1406 generated by a preprocessor 1404 .
[0343] The compressor 1408 compresses the base mesh 1405, the displacement data 1406, and the attribute map 1403 to generate a bitstream 1409. The compressor 1408 can also include metadata 1407 in the bitstream 1409 to transmit additional information to the decoder 1420.
[0344] Next, the configuration of the decoding device when divided into a plurality of sub-meshes will be described.
[0345] FIG. 47 is a diagram showing an example of the configuration of a decoding device when divided into a plurality of sub-meshes.
[0346] As shown in FIG. 47, the decoding device 1420 includes a decompressor 1422 and a post-processor 1427 .
[0347] The decoder 1420 reads the bitstream 1421 and passes it to the decompressor 1422. The decompressor 1422 decompresses the base mesh 1423, the displacement data 1424, and the attribute map 1426 from the bitstream 1421 and passes them to the post-processor 1427. An example of the displacement data 1424 is a displacement vector.
[0348] A post-processor 1427 processes the base mesh 1423 according to the displacement data 1424 and the attribute map 1426 to generate a merged sub-mesh 1428. The post-processor 1427 may further use information from the metadata 1425 to generate the merged sub-mesh 1428. An output mesh 1429 is then generated based on the merged sub-mesh 1428.
[0349] The post-processor 1427 reconstructs the deformed mesh (subdivision, displacement unit). In the subdivision, the post-processor 1427 subdivides the base mesh for each decoded sub-mesh, adds a displacement vector to the subdivided base mesh, and restores the sub-mesh. That is, the post-processor 1427 processes the reconstructed deformed mesh for each sub-mesh. The restored sub-meshes are merged to reconstruct the mesh before division.
[0350] FIG. 48 is a diagram illustrating post-decoding processing.
[0351] The decoder 1432 acquires the bitstream 1431 and performs a decoding process. After the decoding process, a post-decoding unit 1433 may perform a process.
[0352] The processing in the post-decoding unit 1433 is optional depending on the application. The post-decoding unit 1433 converts the decoded data into a nominal format, such as video conversion from YUV space to RGB space. The post-decoding unit 1433 includes a pre-reconstruction unit 1434, a reconstruction unit 1435, a post-reconstruction unit 1436, and an adaptation unit 1437. In other words, post-decoding encapsulates multiple processes performed by the pre-reconstruction unit 1434, the reconstruction unit 1435, the post-reconstruction unit 1436, and the adaptation unit 1437.
[0353] The pre-reconstructor 1434 scales the normalized texture coordinates to match the dimensions of the texture image, for example, in the context of video-based dynamic mesh coding.
[0354] The reconstructor 1435 performs a reconstruction process, for example, on the decoded atlas frames, the decoded base mesh frames, the decoded video frames, and syntax elements associated with the same mesh sequence, the output of which is a sequence of reconstructed mesh frames prior to a post-reconstruction step.
[0355] In the context of video-based dynamic mesh coding, the post-reconstructor 1436 may, for example, perform a number of smoothing operations on the reconstructed mesh frames, including collapsing edges or adding new vertices.
[0356] The adaptation unit 1437 is applied by some application to adapt the reconstructed mesh to a given scenario, for example, the vertices of the reconstructed mesh are transformed from the 3D model coordinate system to the 3D world coordinate system.
[0357] FIG. 49 is a flowchart showing an example of an encoding method performed by the encoding device.
[0358] The encoding device 1400 encodes the position of a sample in a bitstream (S1401). The encoded sample position indicates the position of the upper left corner of a rectangular region in a picture. For example, the sample position may be used by the decoding device 1420 to identify the rectangular region of the picture for decoding. The sample position may also be used by a server to identify the rectangular region of the picture for transmission.
[0359] FIG. 50 is a diagram showing an example of sample positions.
[0360] In picture 1440, the position of a sample in a rectangular region is indicated by {x, y}, which is located to the upper right of a previously decoded rectangular region. For example, the sample position {x, y} may be the position of the first sample in a CU or CTU. Also, for example, the sample position {x, y} may be the position of a sample in a frame that is different from the first sample in a CU or CTU. The sample position may be indicated by an absolute position (absolute coordinates) in the picture, or by a relative position (relative coordinates).
[0361] FIG. 51 is a diagram showing an example in which the sample positions are absolute values.
[0362] As shown in Figure 51, the position of a sample may be absolute: the position of the sample in the top left corner of a rectangular region of picture 1441 is {192, 0}, and is independent of the samples of other rectangular regions.
[0363] FIG. 52 is a diagram showing an example in which the sample positions are relative values.
[0364] As shown in FIG. 52 , the position of a sample may be indicated by coordinates relative to the position of the sample in the previous rectangular area (relative coordinates). That is, the position of a sample may be indicated by coordinates of an offset relative to the sample in the previous rectangular area. Here, in picture 1441, the position of the sample in the upper left corner of rectangular area 1441a is {128,0}, calculated based on the position {64,0} of the sample in the previous rectangular area 1441b. That is, the position {128,0} of the sample in rectangular area 1441a is the difference obtained by subtracting the absolute position of the sample in rectangular area 1441b from the absolute position of the sample in rectangular area 1441b from the origin {0,0} of picture 1441.
[0365] Figure 53 shows an example where the location of the samples is signaled within the body of the bitstream.
[0366] As shown in Figure 53, the sample positions may be included in the main body (data) of the bitstream.
[0367] Figure 54 shows an example where the sample locations are signaled in the header of the bitstream.
[0368] The sample locations may be included in the bitstream header, as shown in Figure 54.
[0369] Also, for example, the sample positions may be signaled as Supplementary Enhancement Information (SEI). Also, for example, the picture may be encoded in YUV420 chroma format. Also, for example, the picture may be encoded in YUV444 chroma format. Also, for example, the picture may be encoded in YUV400 chroma format.
[0370] Returning to the explanation of Figure 49.
[0371] The encoding device 1400 encodes parameters for calculating a Level of Details (LoD) index into a bitstream (S1402). The encoded LoD index indicates an LoD identifier associated with a rectangular region.
[0372] The parameter for calculating the LoD index may be, for example, subdividion_iteration_count. For example, when subdividion_iteration_count=1, LoD0 and LoD1 are generated, and the LoD index value can be 0 or 1. For example, when subdividion_iteration_count=2, LoD0, LoD1, and LoD2 are generated, and the LoD index value can be 0, 1, or 2. By adding subdividion_iteration_count to the bitstream on the encoding device 1400 side in this way, the decoding device 1420 can calculate the LoD index values that can be taken by decoding subdividion_iteration_count of the bitstream, and can appropriately decode the bitstream.
[0373] FIG. 55 is a diagram showing an example of LoD indexes associated with rectangular regions.
[0374] As shown in Figure 55, the LoD index associated with a rectangular region of picture 1442 is 1. For example, the coded value of the LoD index may be calculated using one of the iterations in encoding device 1400. Also, for example, the association between the LoD index and the rectangular region may be determined using the syntax used to code sample positions x[i][j] and y[i][j]. For the rectangular region of picture 1442, i is equal to 1, indicating that the LoD index is 1.
[0375] Returning to the explanation of Figure 49.
[0376] The encoding device 1400 converts the displacement vector into samples of a rectangular area in the picture based on a predetermined scan order (S1403). For example, the encoding device 1400 may add padding data after the samples converted from the displacement vector to form the rectangular area.
[0377] The encoding device 1400 encodes the rectangular region into a picture (S1404).
[0378] FIG. 56 is a flowchart showing an example of a decoding method performed by the decoding device.
[0379] The decoding device 1420 decodes the sample positions from the bitstream (S1411). The decoded sample positions indicate the positions of the upper left corners of rectangular regions within a picture. For example, the decoding device 1420 may identify rectangular regions of a picture based on the sample positions and decode the identified rectangular regions. Alternatively, for example, a server may identify rectangular regions of a picture based on the sample positions and transmit the identified rectangular regions.
[0380] FIG. 57 illustrates the syntax for decoding sample positions.
[0381] The syntax shown in Figure 57 is used by the decoder 1420 to decode the sample positions using syntax elements x[i][j] and y[i][j] associated with LoD i and rectangular region j, where i=0 indicates the LoD of the base mesh.
[0382] num_rectangular_region[i] indicates the number of rectangular regions corresponding to the i-th LoD index, where the LoD is generated by the encoding device 1400 by executing the subdivision method indicated by the subdivision_method parameter the number of times specified by subdivision_iteration_count. x[i][j] and y[i][j] indicate the position of the sample at the top left corner of the j-th rectangular region associated with the i-th LoD index.
[0383] Figure 58 shows the syntax for decoding the position of a sample, including the height and width of a rectangular area.
[0384] The syntax can also be placed in other header locations. For example, these parameters are signaled at the sequence level. For example, these parameters are signaled at the frame level. For example, these parameters are signaled for each patch of a 3D mesh. For example, these parameters are signaled using Supplemental Enhancement Information (SEI). In this case, if the decoder 1420 chooses not to use and ignore the SEI, it decodes the entire image together rather than decoding the rectangular regions individually.
[0385] For example, the decoded num_rectangular_region[i], x[i][j], y[i][j], height[i][j] and width[i][j] are used by the decoder 1420 to determine the identifier of the vertex associated with LoD index i.
[0386] Also, for example, the decoded num_rectangular_region[i], x[i][j], y[i][j], height[i][j] and width[i][j] are used by the decoding device 1420 to determine the location of the block containing the sample associated with the particular vertex identifier.
[0387] The decoding device 1420 decodes the parameters for calculating the LoD index from the bitstream (S1412). The decoded LoD index indicates the LoD identifier associated with the rectangular region.
[0388] For example, the server may select a rectangular region associated with a particular LoD index from among all rectangular regions encoded by the encoding device 1400 and send the selected rectangular region to the decoding device 1420. In this case, the server may modify the encoded values of the height and width of the image to represent the rectangular region as an image.
[0389] For example, the server may select a rectangle to send based on available network bandwidth, or based on a received request.
[0390] When the server selects an LoD to transmit and transmits data of a rectangular area related to the selected LoD to the decoding device 1420, the value of subdivision_iteration_count added to the header may be changed depending on the number of layers of the LoD to be transmitted. For example, if the original value of subdivision_iteration_count is 3, the number of LoD layers is 4. When data related to two layers, LoD0 and LoD1, of the four layers of LoD are transmitted to the decoding device 1420, the value of subdivision_iteration_count added to the header may be changed from 3 to 1. In this way, the decoding device 1420 determines from the header that data for two layers is included in the bitstream and decodes data for two LoD layers.
[0391] The decoding device 1420 derives the number of vertices associated with the decoded LoD index (S1413). This derived number of vertices indicates the number of vertices of the 3D mesh to be displaced using the displacement vector. The derived number of vertices includes the number of vertices in the base mesh and the number of vertices generated by the subdivision process.
[0392] The decoding device 1420 determines whether the decoded LoD index is less than or equal to a predetermined LoD index (S1414), which may be user-defined or derived from other sources.
[0393] If the decoding device 1420 determines that the decoded LoD index is equal to or less than the predetermined LoD index (Yes in S1414), it determines the position of a rectangular region within the picture using the position of the decoded sample and decodes the corresponding rectangular region (S1415). Here, the decoded rectangular region is combined with multiple other decoded rectangular regions to form a larger rectangular region of the picture, and in this case, the LoD indexes of all the rectangular regions are equal to or less than the predetermined LoD index.
[0394] FIG. 59 is a diagram illustrating an example of a decoding process for a rectangular area.
[0395] Here, subdivision_iteration_count is 2. Therefore, in the position information 1446, the LoD indices for the rectangular area are 0, 1, and 2 in the picture 1447. The position information 1446 indicates the position (for example, the position of the upper left corner) of the rectangular area in the layer corresponding to the LoD index.
[0396] In the example of Figure 59, num_rectangular_region[i] is always 1. Therefore, there is only one rectangular region corresponding to each LoD index. Also, the sample position is an absolute value.
[0397] When the LoD index is 0 and the rectangular region is 0, i is 0 and j is 0. That is, x[0][0] and y[0][0] are 0 and 0, respectively, and indicate the position of the decoded sample. Similarly, when the LoD index is 1 and the rectangular region is 0, i is 1 and j is 0. That is, x[1][0] and y[1][0] are 64 and 0, respectively. When the LoD index is 2 and the rectangular region is 0, i is 2 and j is 0. That is, x[2][0] and y[2][0] are 0 and 128, respectively. Rectangular region 0 has already been decoded. Therefore, in picture 1447, the rectangular region with LoD index 1 is decoded. This decoded rectangular region is combined with the previously decoded rectangular regions to form a larger rectangular region.
[0398] Note that the specified LoD index 1445 is set to 1, so in this example, rectangular areas with LoD indexes of 1 or less (i.e., rectangular areas with LoD indexes of 0 and 1) are decoded, and rectangular areas with LoD indexes of 2 are not decoded.
[0399] The decoder 1420 derives the height and width of the rectangular region and uses this in combination with the position of the decoded sample to determine its position within the picture.
[0400] FIG. 60 is a diagram for explaining another example of the decoding process for a rectangular area.
[0401] In the example of Fig. 60, the position information 1446a indicates the position (e.g., the position of the upper left corner) of a rectangular area of a layer corresponding to the LoD index. The position information 1446a further includes the height (division height) and width (division width) of each rectangular area. In this example, the overall height and width of the picture 1447a are 192 pixels.
[0402] FIG. 61 is a diagram showing an example in which different decoded rectangular areas are arranged in rows.
[0403] For example, rectangular regions corresponding to different decoded LoD indices are arranged row by row as shown in FIG. 61 . Alternatively, for example, rectangular regions corresponding to different decoded LoD indices are arranged column by column, and each rectangular region may include any padding to form a rectangular shape. In this example, bold lines represent the boundaries of rectangular regions, normal lines represent the boundaries of coding tree units (CTUs), and dotted lines represent the boundaries of coding units (CUs). Each rectangular region may include one or more tiles or slices corresponding to the LoD. Each tile has multiple CTUs in a row, and each CTU includes multiple CUs. A slice may include one or more tiles, and one tile may be shared by multiple slices.
[0404] In the padding process, for example, in a rectangular area corresponding to an arbitrary LoD index shown in Fig. 61, if the number of samples generated from the displacement vector is smaller than the number of samples that can be stored in the rectangular area minus the number of samples that can be stored in one CTU, the rectangular area will include a CTU in which only padding samples are stored. Specifically, for example, if the above conditions are satisfied for each of LoD index 1 and LoD index 3, the rectangular area of LoD index 1 and the rectangular area of LoD index 3 will include CTUs in which only padding samples are stored.
[0405] Note that a picture does not have to be composed only of rectangular areas corresponding to any LoD. For example, a picture may include a rectangular area composed only of CTUs that do not include samples generated from a displacement vector, i.e., a rectangular area that is not used to derive a displacement vector. For such rectangular areas that do not include samples generated from a displacement vector, information such as the position and size of the rectangular area may not be stored, or information such as the position and size of the rectangular area may be stored, along with information indicating that no LoD is associated with the rectangular area.
[0406] FIG. 62 is a diagram showing an example in which different decoded rectangular areas are arranged in a mixed arrangement of rows and columns.
[0407] For example, rectangular regions corresponding to different decoded LoD indices are arranged in a mixed arrangement of rows and columns as shown in Figure 62. In this example, thick lines represent boundaries of rectangular regions, normal lines represent boundaries of CTUs (Coding Tree Units), and dotted lines represent boundaries of CUs (Coding Units). Each rectangular region can include one or more tiles or slices corresponding to the LoD. Each tile has multiple CTUs in a row, and each CTU includes multiple CUs. A slice can include one or more tiles, and one tile may be shared by multiple slices.
[0408] As in the above example, the encoding device 1400 can adaptively map data for each LoD of the displacement vector to a picture. As a result, when encoding the mapped picture using a video codec, for example, the encoding device 1400 or the decoding device 1420 can refer to LoD1_CTU0, LoD1_CTU1, or LoD1_CTU2 in the same layer when intra-predicting LoD1_CTU3, and generate a predicted value. When the values of the displacement vector data in the same layer are close, the accuracy of the predicted value is improved, thereby improving encoding efficiency. Furthermore, for example, the encoding device 1400 or the decoding device 1420 can refer to LoD2_CTU0, LoD2_CTU1, LoD2_CTU2, or LoD2_CTU3 in the same layer when intra-predicting LoD2_CTU4, and generate a predicted value. When the data values of the displacement vectors in the same layer are close to each other, the accuracy of the predicted values improves, and as a result, the coding efficiency can be improved.
[0409] FIG. 63 is a diagram showing an example in which there are multiple rectangular areas corresponding to one LoD index, and these rectangular areas form an inverted L shape in the horizontal direction.
[0410] For example, there may be multiple rectangular regions corresponding to one LoD index, and these rectangular regions form an inverted L-shape in the horizontal direction as shown in FIG. 63. In this example, bold lines represent the boundaries of the rectangular regions, normal lines represent the boundaries of CTUs (Coding Tree Units), and dotted lines represent the boundaries of CUs (Coding Units). Each rectangular region may include one or more tiles or slices corresponding to the LoD. Each tile has multiple CTUs in a row, and each CTU includes multiple CUs. A slice may include one or more tiles, and one tile may be shared by multiple slices.
[0411] For example, the rectangular regions corresponding to the LoD indices may be determined by the image unpacker 1239 during the decoding process shown in Figure 35. Also, for example, the decoding device 1420 uses the rectangular regions corresponding to the LoD indices to reconstruct a 3D mesh up to a particular LoD.
[0412] 64 is a diagram showing an example of parameters notified by the syntax. These parameters 1446b are included in the syntax when there are multiple rectangular areas corresponding to one LoD index and these rectangular areas form an inverted L shape in the horizontal direction.
[0413] In this example, the first line indicates the first rectangular area corresponding to LoD index 0, with the position of the top left sample of this rectangular area being "0,0", the height being 64, and the width being 128. Similarly, the second line indicates the first rectangular area corresponding to LoD index 1, with the position of the top left sample of this rectangular area being "128,0", the height being 64, and the width being 64. Furthermore, the third line indicates the second rectangular area corresponding to LoD index 1, with the position of the top left sample of this rectangular area being "0,64", the height being 64, and the width being 192. The same process continues.
[0414] In Figure 64, three parameters, Location, Height, and Width, are notified in units of the number of samples to notify the position and size of a rectangular area, but the method of notifying the position and size of a rectangular area is not limited to this. For example, if the boundary of the rectangular area is not set at a position other than a CTU boundary, the Location, Height, and Width may be expressed as integer multiples of the CTU size or the luminance CTB size. For example, in the example of Figure 64, if the CTU size is 64, the LoD index 0 will have a Location of "0,0", a Height of "1", and a Width of "2", and the LoD index 1 will have a Location of "2,0", a Height of "1", and a Width of "1".
[0415] Although the description here has been given using an example in which the CTU size is used, implementation is similarly possible using rectangular block units other than CTUs. For example, the Location, Height, and Width may be expressed as integer multiples of a tile, an integer multiple of a slice, an integer multiple of a subpicture, or an integer multiple of a rectangular block size different from any of the tile, slice, and subpicture. Note that when the size of a CTU, tile, slice, or subpicture is used to notify the position and size of a rectangular area, it is preferable, but not necessary, that the size of the CTU, tile, slice, or subpicture used is fixed within the picture. For example, when the CTU size is used to notify the position and size of a rectangular area and the CTU size is fixed within the picture, the picture size, i.e., the Height and Width of the picture, may be limited to integer multiples of the CTU size.
[0416] In addition, the position of the rectangular area may not only be specified by the vertical and horizontal positions of the upper left corner of the rectangular area, but may also be notified by the index of the CTU at the upper left corner of the rectangular area based on the CTU index assigned to the CTUs in the picture in raster scan order.
[0417] In addition, when the size of a CTU, tile, slice, or subpicture is used to notify the position and size of a rectangular area, information indicating the size of the CTU, tile, slice, or subpicture to be used may be stored in the initial bitstream and notified.
[0418] Fig. 65 is a diagram showing an example including different types of rectangular regions, where some rectangular regions are arranged in rows, other rectangular regions are arranged in combinations of rows and columns, and some rectangular regions correspond to the same LoD index.
[0419] In this example, there are different types of rectangular regions, some arranged in rows and others arranged in a combination of rows and columns. Also, some rectangular regions correspond to the same LoD index. In this example, bold lines represent the boundaries of rectangular regions, normal lines represent the boundaries of CTUs (coding tree units), and dotted lines represent the boundaries of CUs (coding units). Also, each rectangular region can have one or more tiles and slices corresponding to the LoD. Each tile contains multiple CTUs in a row, and each CTU contains multiple CUs. A slice can contain one or more tiles, and tiles may be shared between slices.
[0420] Figure 66 is a diagram showing an example in which rectangular slices contain different types of rectangular regions, where some rectangular regions are arranged in rows, other rectangular regions are arranged in combinations of rows and columns, and some rectangular regions correspond to the same LoD index.
[0421] For example, the video encoding device 1400 uses a codec such as VVC to encode samples of different rectangular regions into multiple rectangular slices. Each rectangular slice may include multiple rectangular regions, each corresponding to a different LoD index. Figure 66 shows such an example. In this example, the encoding device 1400 divides a picture into a certain number of tile rows and columns. The rectangular slices include tiles with different numbers of rows and columns. Therefore, one rectangular slice may include different rectangular regions. In this example, the rectangular slice located in the upper left corner includes three rectangular regions, one of which corresponds to LoD index 0 and the other two correspond to LoD index 1.
[0422] Returning to the explanation of Figure 56.
[0423] If the decoding device 1420 determines that the decoded LoD index is equal to or less than the predetermined LoD index, it converts samples from the rectangular region of the picture into displacement vectors based on a predetermined scan order (S1416). At this time, the total number of samples converted into displacement vectors in the predetermined scan order is equal to the derived number of vertices, and the remaining samples that are not converted in the rectangular region are padding samples and are not used to displace vertices of the 3D mesh.
[0424] FIG. 67 is a diagram showing an example of samples converted into displacement vectors.
[0425] In this example, the number of samples contained in the rectangular region is 8192, and the number of derived vertices is 7168. Therefore, only 7168 samples from the rectangular region are converted into displacement vectors.
[0426] In the prior art, samples used to determine LoD displacement vectors were not limited to a rectangular area. This restriction made it impossible to independently obtain displacement vectors for specific LoDs. This caused problems in scenarios where a server only transmits up to a specific LoD due to bandwidth constraints, or where a client (the decoding device 1420) requests a specific LoD. The objective of the present disclosure is to enable the decoding device 1420 to accurately decode picture samples without errors.
[0427] (Variation) In a variation, a rectangular region may be padded to align with a CU, CTU, tile, slice, or subpicture.
[0428] Alternatively, in a variant, samples within a rectangular region corresponding to different displacement vectors of different LoDs may be packed in reverse order (starting from the bottom right of the rectangular region).
[0429] Note that information indicating whether the displacement vector data is packed in reverse order (hereinafter referred to as "reverse_packing information") may be added to the header of the bitstream, etc. This information allows the decoding device to match the unpacking method of the displacement vector with the packing method of the encoding device by decoding the reverse_packing information in the bitstream header, thereby enabling the bitstream to be decoded appropriately.
[0430] Alternatively, in a variant, the rectangular regions may be tiles, and all tiles in a picture may share a common header indicating the LoD containing samples corresponding to the displacement vector.
[0431] Also, in a variant, every tile may have a different header indicating the LoD containing the samples corresponding to the displacement vector.
[0432] Alternatively, in a variant, all rectangular regions within a picture may have the same aspect ratio.
[0433] In a variant, at least one rectangular region may have a different aspect ratio than the other rectangular regions.
[0434] Note that if the bitstream contains information related to LoD mapping (such as position, width, height, tile boundary, subdivision_iteration_count, etc.), this information allows the server to understand the LoD mapping information without decoding the bitstream.
[0435] This gives the server the ability to generate a modified video bitstream without decoding the original bitstream. Adjustments such as image resizing based on client requirements and bandwidth are seamlessly incorporated into the new bitstream and transmitted. This innovative approach allows the receiver (decoder) to process the modified bitstream as video frames. With the rectangles and all necessary decoding data transmitted from the server, the decoder can accurately decode the video without errors.
[0436] Next, the mapping rules for generating displacement vectors and placing samples will be described.
[0437] In the present disclosure, the mapping order of samples generated from a displacement vector, or the demapping order or scan order when generating a displacement vector from a sample, can be defined by combining three-stage mapping rules. In the following, the "mapping order" will be described as an example, but it may also be read as the "demapping order" or the "scan order."
[0438] The first stage specifies the order in which rectangular regions are arranged, which may be raster scan order, reverse raster scan order, or some other order. The second stage specifies the order in which samples are arranged within rectangular regions based on CTUs, CUs, or other block sizes, which may again be raster scan order, reverse raster scan order, or some other order. The third stage specifies the order in which samples are arranged within rectangular regions, CTUs, or CUs, which may again be raster scan order, reverse raster scan order, or some other order.
[0439] Other scan orders include, for example, zigzag scan order or reverse zigzag scan order. Also, unlike raster scan, which prioritizes horizontal scanning from the top left to the bottom right, it is possible to adopt a scan order that prioritizes vertical scanning from the top left to the bottom right, or a scan order that prioritizes vertical scanning from the bottom right to the top left.
[0440] Furthermore, the mapping orders for the first to third stages can be selected independently. For example, the order of arrangement of rectangular regions in the first stage and the order of arrangement of samples in CUs within the rectangular region in the second stage can be reverse raster scan order, and the order of arrangement of samples in CUs in the third stage can be raster scan order. Alternatively, the order of arrangement of samples in CTUs within a rectangular region, the order of arrangement of samples in CUs within a CTU, and the order of arrangement of samples within a CU may be specified separately. When samples are arranged for the entire rectangular region without considering CTUs, CUs, or other block units, the rules for the third stage are applied to the entire rectangular region, ignoring the block units. In this case, for example, in the raster scan order, samples in the top row of the rectangular region are arranged in order from left to right, and after reaching the right end, the process proceeds to the next row.
[0441] Furthermore, when combining rectangular areas to form an L-shape or an entire picture, if it is difficult to fix the arrangement order of the rectangular areas in the first stage, it is possible to specify the second and third stages, or only the third stage, without specifying the first stage. Even in this case, it is possible to define a rule, such as arranging the rectangular area with LoD index 0 in the lower right corner of the picture.
[0442] Fig. 68 is a diagram showing the syntax of Header Example 1. Fig. 69 is a diagram showing the syntax of Header Example 2. Fig. 70 is an example of syntax indicating location information. Fig. 70 is another example of the syntax of Fig. 58.
[0443] LoDPartialDecodeEnable is information indicating that the LoD can be partially decoded. For example, a value of 1 may indicate that the LoD can be partially decoded, and a value of 0 may indicate that the LoD cannot be partially decoded. Specifically, when information regarding four layers of LoD is included in a bitstream, if LoDPartialDecodeEnable=1, for example, the decoding device may be able to decode only LoD0, or may be able to decode only LoD0 and LoD1. In this way, by adding information indicating whether the LoD can be partially decoded to the bitstream, the decoding device can know whether the LoD can be partially decoded, and can perform partial decoding depending on the use case.
[0444] As an example of a use case, for example, when displaying three-dimensional mesh data as thumbnails, the thumbnails can be displayed with low processing load by partially decoding and displaying LoD0 and LoD1 out of the four LoD layers, and when displaying three-dimensional mesh data selected on the thumbnail display on one screen, the three-dimensional mesh data can be displayed at high resolution by decoding the four layers from LoD0 to LoD3.
[0445] LoDLiftingSkipUpdateFlag is information that indicates whether to skip the process of updating values in a higher layer using prediction residuals in a lower layer when assigning displacement vector data to an LoD layer and applying a lifting transform or the like to the displacement vector r to generate transform coefficients representing components from low to high frequency. For example, a value of 1 indicates that the update process is skipped, and a value of 0 indicates that the update process is applied. An example of the processing flow for lifting transforms using this information will be explained later on the next slide.
[0446] When adding each piece of information to a bitstream as in Header Example 1, if LoDPartialDecodeEnable=1, a standard specification or the like may restrict it to LoDLiftingSkipUpdateFlag=1. This ensures that partial decoding of LoD is possible when LoDPartialDecodeEnable=1 is set. If LoDLiftingSkipUpdateFlag=0 is set when LoDPartialDecodeEnable=1, the decoding device may output a standard compliance error.
[0447] Note that each piece of information may be added to the bitstream as in Header Example 2. In this case, when LoDPartialDecodeEnable=0, LoDLiftingSkipUpdateFlag is added, and when LoDPartialDecodeEnable=1, LoDLiftingSkipUpdateFlag does not need to be added to the bitstream m. Note that when LoDLiftingSkipUpdateFlag is not added to the bitstream, the decoding device may set its value to 1 and proceed with the decoding process. As a result, when LoDPartialDecodeEnable=1, LoDLiftingSkipUpdateFlag=1 is always set, and partial decoding of LoD can be guaranteed.
[0448] Note that the decoding device may add LoDLiftingSkipUpdateFlag to the bitstream instead of adding LoDPartialDecodeEnable to the bitstream, and determine that LoD can be partially decoded when LoDLiftingSkipUpdateFlag=1. This allows the header code amount to be reduced.
[0449] subdivision_method, subdivision_iteration_count, num_rectangular_region, x, y, height, and width are the information described above, and among them, the information num_rectangular_region, x, y, height, and width may be added to the bitstream when LoDPartialDecodeEnable = 1, and not added to the bitstream when LoDPartialDecodeEnable = 0. This makes it possible to reduce the amount of header code.
[0450] Note that, in the present disclosure, whether or not partial decoding of the LoD is possible is indicated using LoDPartialDecodeEnable, but this is not necessarily limited thereto. For example, information regarding a layer among the LoD layers that is partially decodable (hereinafter referred to as LoDPartialDecodeEnableLayer) may be added. For example, if the total number of LoD layers is four and LoDPartialDecodeEnableLayer=2, it may be possible to indicate that partial decoding is not possible up to LoD0 and LoD1, but that partial decoding is possible from LoD2 onwards. This allows the decoding device to know that partial decoding is possible from LoDPartialDecodeEnableLayer onwards, and to switch the decoding layer from LoDPartialDecodeEnableLayer onwards depending on the use case.
[0451] In the present disclosure, LoDLiftingSkipUpdateFlag is used to indicate whether or not to skip the update process of the lifting transform. However, this is not necessarily limited to this. For example, information regarding a layer among the LoD layers for which the update process is to be skipped (hereinafter, LoDLiftingSkipUpdateLayer) may be added. For example, if the total number of LoD layers is four and LoDLiftingSkipUpdateLayer=2, the update process may not be skipped up to LoD0 and LoD1, but may be skipped from LoD2 onwards. This allows the decoding device to know that the update process of the lifting transform is skipped from LoDLiftingSkipUpdateLayer onwards, and therefore the layers from LoDLiftingSkipUpdateLayer onwards can be partially decoded, and the decoding device can switch the decoding layers from LoDLiftingSkipUpdateLayer onwards depending on the use case.
[0452] It should be noted that the above-described configuration for notifying the information indicating that the LoD can be partially decoded or the information indicating whether to skip the process of updating the value of the upper layer using the prediction residual of the lower layer can be implemented without being combined with the configuration for notifying the position or size of the rectangular area and the LoD index corresponding to the rectangular area described in the present disclosure. For example, the above-described configuration for notifying the information indicating that the LoD can be partially decoded or the information indicating whether to skip the process of updating the value of the upper layer using the prediction residual of the lower layer may be implemented in combination with a method in which samples generated from a displacement vector corresponding to a predetermined LoD index are mapped to a non-rectangular area, or may be implemented in combination with a method in which a displacement vector corresponding to a predetermined LoD index is encoded without generating image format data from the displacement vector corresponding to the predetermined LoD index.
[0453] FIG. 71 is a flowchart showing an example of lifting transform performed by the encoding device.
[0454] The encoding device 1400 generates a predicted value of the value of the displacement vector belonging to the LoD of layer i (hereinafter referred to as LoD[i]) from the value of the displacement vector belonging to the LoD of layer i-1 (hereinafter referred to as LoD[i-1]), which is the higher layer, by using an average value or the like, and calculates the prediction residual of the displacement vector of layer i (hereinafter referred to as resi[i]) (S1421). The encoding device 1400 may add the encoding result of this resi[i] to the bitstream.
[0455] Next, the encoding device 1400 determines whether LoDLiftingSkipUpdateFlag=1 (S1422).
[0456] When LoDLiftingSkipUpdateFlag=0, the encoding device 1400 updates the value of LoD[i-1] using the prediction residual resi[i] (S1423) and proceeds to the next layer. For example, the encoding device 1400 may update the value of LoD[i-1] by adding a value obtained by multiplying resi[i] by a weighting coefficient to LoD[i-1]. In this way, when the prediction residual becomes large due to prediction, the value of LoD[i-1], which is the basis of the prediction, is corrected according to the value of the prediction residual, thereby improving the coding efficiency when encoding LoD[i-1].
[0457] Furthermore, when LoDLiftingSkipUpdateFlag=1, the encoding device 1400 skips the process of updating the value of LoD[i-1] using the prediction residual resi[i], and moves on to the process of the next layer. Here, by skipping the process of updating the value of LoD[i-1] using the prediction residual resi[i], the LoD value of the upper layer does not depend on the resi value of the lower layer, and it becomes possible to decode the LoD value of the upper layer without using the resi value of the lower layer. In other words, it becomes possible to partially decode up to a certain upper layer.
[0458] For example, when the number of LoD layers is 4, the encoding device 1400 may start processing from the lowest layer, LoD3. Specifically, the above flow may be executed in the order of i=3, 2, and 1. Note that when i=0, that is, when LoD[0] (corresponding to the base mesh), there are no higher layers, so when LoD[0] is used, a predicted value may not be generated from the higher layers, and resi[0] may be set to the value of LoD[0].
[0459] FIG. 72 is a flowchart showing an example of lifting transformation performed by the decoding device.
[0460] The decoding device 1420 determines whether LoDLiftingSkipUpdateFlag=1 (S1431).
[0461] When LoDLiftingSkipUpdateFlag=0, the decoding device 1420 updates the value of LoD[i-1] using the prediction residual resi[i] (S1432). For example, the decoding device 1420 may update the value of LoD[i-1] by adding a value obtained by multiplying resi[i] by a weighting coefficient to LoD[i-1]. In this way, when the prediction residual becomes large due to prediction, the value of LoD[i-1], which is the basis of the prediction, is corrected according to the value of the prediction residual, thereby enabling application of the same processing as the encoding device 1400, which improves the encoding efficiency when encoding LoD[i-1].
[0462] Furthermore, when LoDLiftingSkipUpdateFlag=1, the decoding device 1420 skips the process of updating the value of LoD[i-1] using the prediction residual resi[i]. Here, by skipping the process of updating the value of LoD[i-1] using the prediction residual resi[i] in the same way as the encoding device 1400, the LoD value of the upper layer does not depend on the resi value of the lower layer, and it becomes possible to decode the LoD value of the upper layer without using the resi value of the lower layer. In other words, it becomes possible to partially decode up to a certain upper layer.
[0463] Next, the decoding device 1420 generates a predicted value of LoD[i] from LoD[i-1] using the same process as the encoding device, such as averaging, and adds it to the predicted residual resi[i] decoded from the bitstream to restore the value of the displacement vector of LoD[i] (S1433), and then moves on to processing at the next layer.
[0464] As described above, by setting LoDLiftingSkipUpdateFlag in the lifting transform to a value of 0, the value of LoD[i-1] that is the source of prediction can be corrected in accordance with the value of the prediction residual, thereby improving coding efficiency, and by setting LoDLiftingSkipUpdateFlag to a value of 1, the LoD value of the upper layer becomes independent of the resi value of the lower layer, making it possible to partially decode up to a certain upper layer.
[0465] For example, when the number of LoD layers is 4, the decoding device 1420 may start processing from the upper layer, LoD1. Specifically, the above flow may be executed in the order of i=1, 2, and 3. Note that when i=0, that is, when LoD[0] (corresponding to the base mesh), there are no higher layers, so in the case of LoD[0], a predicted value may not be generated from the higher layers, and the value of LoD[0] may be set to resi[0].
[0466] (Configuration Example) Fig. 73 is a diagram showing an example of the configuration of an encoding device in an embodiment. Fig. 74 is a flowchart showing an example of an encoding method by the encoding device in an embodiment.
[0467] The encoding device 1450 includes a circuit 1451 and a memory 1452 connected to the circuit 1451. The encoding device 1450 is a device that realizes the encoding device 1400.
[0468] The circuit 1451 performs the following operations.
[0469] The circuit 1451 encodes the displacement vectors as samples of a rectangular area in a picture (S1451). The displacement vectors correspond to the vertices of a mesh. The circuit 1451 generates a bitstream including the encoded samples and position information indicating the position of the rectangular area in the picture (S1452). The position information is associated with the layer corresponding to the rectangular area (to which the vertices of the mesh belong).
[0470] This generates a bitstream including samples with coded displacement vectors and position information indicating the positions of rectangular areas within a picture, thereby informing a decoding device of the positions of the samples within the picture, thereby generating a bitstream that enables the decoding device to correctly decode the displacement vector of a desired sample within a picture based on the position information.
[0471] For example, the position information indicates the position of the top left corner of the rectangular area, so that the specific position of the rectangular area can be notified to the decoding device.
[0472] For example, the bitstream may further include the size of the rectangular region, thereby informing the decoder of the rectangular region within the picture in which the samples are placed.
[0473] For example, the size includes the width and height of the rectangular area, thereby informing the decoder of the rectangular area in which the samples are placed within the picture.
[0474] For example, the bitstream may include identification information for identifying the layer, so that the decoder can know which layer a sample corresponds to.
[0475] The identification information may be the layer identifier itself, or may be the layer identifier or information that can identify the layer (the number of subdivisions).
[0476] For example, a picture may have a plurality of rectangular regions, each of which corresponds to a plurality of layers. The plurality of layers may include a layer. Also, for example, a bitstream may include a plurality of pieces of position information, each of which includes position information. Each of the pieces of position information indicates the positions of the plurality of rectangular regions. Therefore, the decoding device may be informed of the layers to which the plurality of samples correspond and the positions of the plurality of rectangular regions.
[0477] For example, the bitstream may include multiple sizes including the above-mentioned sizes, each corresponding to a different rectangular region, thereby informing the decoding device of the different rectangular regions in which the samples are arranged within the picture.
[0478] For example, a rectangular region may include one or more coding tree units (CTUs) in a picture, one or more tiles in a picture, one or more slices in a picture, or one or more sub-pictures in a picture.
[0479] Fig. 75 is a diagram showing an example of the configuration of a decoding device according to an embodiment. Fig. 76 is a flowchart showing an example of a decoding method performed by the decoding device according to an embodiment.
[0480] The decoding device 1460 includes a circuit 1461 and a memory 1462 connected to the circuit 1461. The decoding device 1460 is a device that realizes the decoding device 1420.
[0481] The circuit 1461 performs the following operations.
[0482] The circuit 1461 obtains a bitstream (S1461). The circuit 1461 obtains position information indicating the position of a rectangular area within a picture from the bitstream (S1462). The circuit 1461 decodes a displacement vector indicated by a sample of the rectangular area located at the position indicated by the position information within the picture (S1463). The displacement vector corresponds to a vertex of a mesh. The position information is associated with the layer corresponding to the rectangular area (to which the vertex of the mesh belongs).
[0483] According to this, position information indicating the position of a rectangular area within a picture is obtained from the bitstream, and therefore, the displacement vector of a desired sample within a picture can be correctly decoded based on the position information.
[0484] For example, the position information indicates the position of the top left corner of the rectangular region, so that the decoding device 1460 can correctly decode the displacement vector of the desired sample based on the specific position of the rectangular region.
[0485] For example, the bitstream may further include the size of the rectangular region, so that the decoding device 1460 can identify the rectangular region in the picture where the sample is located and correctly decode the displacement vector of the desired sample.
[0486] For example, the size includes the width and height of the rectangular area, so that the decoding device 1460 can identify the rectangular area in which the sample is located within the picture and correctly decode the displacement vector of the desired sample.
[0487] For example, the bitstream includes identification information for identifying the layer, so that the decoding device 1460 can identify the rectangular area corresponding to the layer to which the sample corresponds and correctly decode the displacement vector of the desired sample.
[0488] The identification information may be the layer identifier itself, or may be the layer identifier or information that can identify the layer (the number of subdivisions).
[0489] For example, a picture may have a plurality of rectangular regions, each of which corresponds to a plurality of layers. The plurality of layers may include layers. Also, for example, a bitstream may include a plurality of pieces of position information, each of which includes position information. Each of the pieces of position information indicates the positions of the plurality of rectangular regions. Therefore, the decoding device 1460 may identify the positions of the plurality of layers and the plurality of rectangular regions to which the plurality of samples correspond.
[0490] For example, the bitstream may include multiple sizes, including the above sizes, each corresponding to a different rectangular region, allowing the decoding device 1460 to identify multiple rectangular regions in which multiple samples are placed within a picture.
[0491] For example, a rectangular region may include one or more coding tree units (CTUs) in a picture, one or more tiles in a picture, one or more slices in a picture, or one or more sub-pictures in a picture.
[0492] For example, the circuit 1461 further determines whether the layer indicated by the identification information is equal to or lower than a predetermined layer indicated by the predetermined identification information, and the decoding by the circuit 1461 is performed based on the determination result that the layer indicated by the identification information is equal to or lower than the predetermined layer.
[0493] Therefore, the decoding device 1460 can decode rectangular areas at a predetermined layer or lower, thereby reducing the processing load.
[0494] Furthermore, the encoding device and decoding device of the present disclosure may be able to provide flexibility in the decoding process of displacement data, and may also be able to reduce the complexity of the decoding process.
[0495] The encoding device and decoding device of the present disclosure may also be implemented by combining at least a part of other aspects. Furthermore, the configuration or processing of the encoding device and decoding device of the present disclosure may be realized by combining a part of the processing shown in any of the flowcharts according to this aspect, a part of the configuration of any of the devices, a part of the syntax, or the like with other aspects.
[0496] Furthermore, the processing performed by the decoding device of the present disclosure may also be performed in the encoding device in the same manner.
[0497] Furthermore, not all of the components described in this embodiment are always necessary, and only some of the components of the encoding device and decoding device of the present disclosure may be included.
[0498] The present disclosure is useful, for example, in encoding devices, decoding devices, transmitting devices, receiving devices, etc. related to three-dimensional meshes, and is applicable to computer graphics systems, three-dimensional data display systems, etc.
[0499] 100 Encoding device 101, 121, 144 Vertex information encoder 102, 145 Connection information encoder 103, 122 Attribute information encoder 104, 204, 1103 Preprocessor 105, 205, 2106 Postprocessor 110 Three-dimensional data encoding system 111, 211 Controller 112, 212 Input / output processor 113 Three-dimensional data encoder 114 System multiplexer 115 Three-dimensional data generator 123 Metadata encoder 124 Multiplexer 131 Vertex image generator 132 Attribute image generator 133 Metadata generator 134 Video encoder 141 Two-dimensional data encoder 142 Mesh data encoder 143 Texture encoder 148 Description encoder 151, 251 Circuit 152, 252 Memory 200 Decoding device 201, 221, 244 Vertex information decoder 202, 245 Connection information decoder 203, 222 Attribute information decoder 210 3D data decoding system 213 3D data decoder 214 System demultiplexer 215, 247 Presentation device 216 User interface 223 Metadata decoder 224 Demultiplexer 231 Vertex information generator 232 Attribute information generator 234 Video decoder 241 2D data decoder 242 Mesh data decoder 243 Texture decoder 246 Mesh reconstructor 248 Description decoder 300 Network 310 External connection device 613 Base mesh decoder 614 Displacement decoder 615 Attribute decoder 616 Other type decoder 617 3D reconstructor 631 Frame header decoder 632 Vertex geometry coordinate predictor 633 Vertex geometry coordinate differential decoder 634 Reconstructor 1231 Demultiplexer 1232 Switch 1233 Static mesh decoder 1234 Mesh buffer 1235 Motion decoder 1236 Base mesh reconstructor 1237 Inverse quantizer 1238, 1243 Video decoder1239 Image unpacker 1240 Inverse quantizer 1241 Inverse wavelet transformer 1242 Reconstructor 1244 Color transformer 1251 Decoded base mesh 1252 Subdivider 1253 Subdivided mesh 1254 Decoded displacement data 1255 Displacer 1256 Decoded 3D mesh 1400 Encoder 1401 Input mesh 1402 Submeshes 1403, 1426 Attribute map 1404 Preprocessor 1405, 1423 Base mesh 1406, 1424 Displacement data 1407, 1425 Metadata 1408 Compressor 1409, 1421, 1431 Bitstream 1411 Base mesh generator 1412 Subdivider 1413 Displacement generator 1420 Decoder 1422 Decompressor 1427 Post-processor 1428 Merged sub-meshes 1429 Output mesh 1432 Decoder 1433 Post-decoding unit 1434 Pre-reconstruction unit 1435 Reconstruction unit 1436 Post-reconstruction unit 1437 Adaptation unit 1438 Final 3D mesh frame 1440-1442, 1447, 1447a Picture 1445 Predetermined LoD index 1450 Encoder 1451, 1461 Circuit 1452, 1462 Memory 1460 Decoder 4801, 5701, 11001 Decimator 4802, 5702, 11002 Subdivider 4803, 5703, 11003 Displacement vector calculator 4804, 5704, 11004 Wavelet transformer 4805 Inter predictor 4806, 5706, 11006 Quantizer 4807, 5708, 11008 Image packer 4808, 5709, 11009 Video encoder 4811, 5003, 5711, 5805, 11011 Inverse quantizer 4812, 5004, 5006, 5712, 5806, 5808, 11012 Reconstructor 4813, 5713, 11013, 5811 Reference buffer5001, 5801, 11021 Video decoder 5002, 5802, 11022 Image unpacker 5005, 5807 Inverse wavelet transformer 5705, 11005 LoD-based inter predictor 5707, 5804, 11007, 11024 Switch 5710, 11010 Arithmetic encoder 5803, 11023 Arithmetic decoder
Claims
1. A coding method executed by an encoding device, comprising: encoding a displacement vector as samples of a rectangular region within a picture; generating a bitstream including the encoded samples and position information indicating the position of the rectangular region within the picture, wherein the position information is associated with a hierarchy corresponding to the rectangular region.
2. The coding method according to claim 1, wherein the position information indicates the upper left position of the rectangular region.
3. The coding method according to claim 1, wherein the bitstream further includes the size of the rectangular region.
4. The coding method according to claim 3, wherein the size includes the width and height of the rectangular region.
5. The coding method according to any one of claims 1 to 4, wherein the bitstream includes identification information for identifying the hierarchy.
6. The coding method according to claim 3, wherein the picture has a plurality of rectangular regions including the rectangular region, each of the plurality of rectangular regions corresponds to a plurality of hierarchies, and the plurality of hierarchies includes the hierarchy.
7. The coding method according to claim 6, wherein the bitstream includes a plurality of position information including the position information, and each of the plurality of position information indicates the position of the plurality of rectangular regions.
8. The coding method according to claim 6 or 7, wherein the bitstream includes a plurality of sizes including the size, and each of the plurality of sizes corresponds to the plurality of rectangular regions.
9. The coding method according to any one of claims 1 to 4, wherein the rectangular region includes one or more CTUs (Coding Tree Units) within the picture.
10. The coding method according to any one of claims 1 to 4, wherein the rectangular region includes one or more tiles within the picture.
11. The coding method according to any one of claims 1 to 4, wherein the rectangular region includes one or more slices within the picture.
12. The coding method according to any one of claims 1 to 4, wherein the rectangular region includes one or more sub-pictures within the picture.
13. A decoding method executed by a decoding device, the method comprising: obtaining a bitstream; obtaining position information indicating the position of a rectangular region within a picture from the bitstream; decoding a displacement vector indicated by samples of the rectangular region at the position indicated by the position information within the picture, wherein the position information is associated with a hierarchy corresponding to the rectangular region.
14. The decoding method according to claim 13, wherein the position information indicates the upper left position of the rectangular region.
15. The decoding method according to claim 13, wherein the bitstream further includes the size of the rectangular region.
16. The decoding method according to claim 15, wherein the size includes the width and height of the rectangular region.
17. The decoding method according to any one of claims 13 to 16, wherein the bitstream includes identification information for identifying the hierarchy.
18. The decoding method according to claim 15, wherein the picture has a plurality of rectangular regions including the rectangular region, each of the plurality of rectangular regions corresponds to a plurality of hierarchies, and the plurality of hierarchies includes the hierarchy.
19. The decoding method according to claim 18, wherein the bitstream includes a plurality of position information including the position information, and each of the plurality of position information indicates the position of the plurality of rectangular regions.
20. The decoding method according to claim 18 or 19, wherein the bitstream includes a plurality of sizes including the size, and each of the plurality of sizes corresponds to the plurality of rectangular regions.
21. The decoding method according to any one of claims 13 to 16, wherein the rectangular region includes one or more CTUs (Coding Tree Units) within the picture.
22. The decoding method according to any one of claims 13 to 16, wherein the rectangular region includes one or more tiles within the picture.
23. The decoding method according to any one of claims 13 to 16, wherein the rectangular region includes one or more slices within the picture.
24. The decoding method according to any one of claims 13 to 16, wherein the rectangular region includes one or more sub-pictures within the picture.
25. Further, it is determined whether or not the hierarchy indicated by the identification information is equal to or lower than a predetermined hierarchy indicated by predetermined identification information, and the decoding is executed based on a determination result that the hierarchy indicated by the identification information is equal to or lower than the predetermined hierarchy. The decoding method according to claim 17.
26. An encoding apparatus comprising: a circuit; and a memory connected to the circuit, wherein in operation, the circuit encodes a displacement vector as samples of a rectangular region within a picture, generates a bit stream including the encoded samples and position information indicating a position of the rectangular region within the picture, and the position information is associated with a hierarchy corresponding to the rectangular region.
27. A decoding apparatus comprising: a circuit; and a memory connected to the circuit, wherein in operation, the circuit acquires a bit stream, acquires position information indicating a position of a rectangular region within a picture from the bit stream, decodes a displacement vector indicated by samples of the rectangular region at a position indicated by the position information within the picture, and the position information is associated with a hierarchy corresponding to the rectangular region.
Citation Information
Patent Citations
Progressive three-dimensional mesh information coding / decoding method, and apparatus therefor
JP2006187015A