Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
The proposed three-dimensional data encoding method improves encoding efficiency by using flags to determine quantization and hierarchical data division, addressing inefficiencies in existing processes.
Patent Information
- Application Number
- PCT/JP2025/024341
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-09
- Filing Date
- 2025-07-07
- Publication Date
- 2026-01-15
AI Technical Summary
Existing encoding and decoding processes for three-dimensional mesh data, particularly in displacement vectors, are inefficient and require improvements for better compression and transmission efficiency.
A three-dimensional data encoding method that generates a bitstream with flags to determine whether quantization is performed, allowing for hierarchical division of data into sequence, frame, and patch levels, and includes flags to indicate quantization parameters, reducing information processing and improving encoding efficiency.
The method reduces processing load by selectively applying quantization, enhancing encoding and decoding efficiency through optimized flag usage and hierarchical data division.
Smart Images

Figure JP2025024341_15012026_PF_FP_ABST
Abstract
Description
Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
[0001] The present disclosure relates to a three-dimensional data encoding method and the like.
[0002] In US Pat. No. 6,299,549 a method and apparatus for encoding and decoding three-dimensional mesh data is proposed.
[0003] Japanese Patent Application Laid-Open No. 2006-187015
[0004] Further improvements are desired in the encoding or decoding process for displacement vectors.The present disclosure aims to improve the encoding or decoding process for displacement vectors.
[0005] A three-dimensional data encoding method according to one aspect of the present invention generates a bitstream including data to be encoded and one or more flags for determining whether or not to perform a quantization process on the data to be encoded.
[0006] These comprehensive or specific aspects may be realized as a system, an apparatus, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, an apparatus, an integrated circuit, a computer program, and a recording medium.
[0007] The present disclosure may contribute to improving encoding processes and the like related to displacement vectors.
[0008] 1 is a conceptual diagram showing a three-dimensional mesh according to an embodiment. FIG. 2 is a conceptual diagram showing basic elements of a three-dimensional mesh according to an embodiment. FIG. 3 is a conceptual diagram showing mapping according to an embodiment. FIG. 4 is a block diagram showing a configuration example of an encoding / decoding system according to an embodiment. FIG. 5 is a block diagram showing a configuration example of an encoding device according to an embodiment. FIG. 6 is a block diagram showing another configuration example of an encoding device according to an embodiment. FIG. 7 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 8 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 9 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 10 is a conceptual diagram showing another configuration example of a bit stream according to an embodiment. FIG. 11 is a conceptual diagram showing yet another configuration example of a bit stream according to an embodiment. FIG. 12 is a block diagram showing a specific example of an encoding / decoding system according to an embodiment. FIG. 13 is a conceptual diagram showing an example configuration of point cloud data according to an embodiment. FIG. 14 is a conceptual diagram showing an example data file of point cloud data according to an embodiment. FIG. 15 is a conceptual diagram showing an example configuration of mesh data according to an embodiment. FIG. 16 is a conceptual diagram showing an example data file of mesh data according to an embodiment. FIG. 17 is a conceptual diagram showing types of three-dimensional data according to an embodiment. FIG. 18 is a block diagram showing an example configuration of a three-dimensional data encoder according to an embodiment. FIG. 19 is a block diagram showing an example configuration of a three-dimensional data decoder according to an embodiment. FIG. 19 is a block diagram showing another configuration example of a three-dimensional data encoder according to an embodiment. FIG. 19 is a block diagram showing another configuration example of a three-dimensional data decoder according to an embodiment. 1 is a conceptual diagram showing a specific example of encoding processing according to an embodiment. FIG. 2 is a conceptual diagram showing a specific example of decoding processing according to an embodiment. FIG. 3 is a block diagram showing an implementation example of an encoding device according to an embodiment. FIG. 4 is a block diagram showing an implementation example of a decoding device according to an embodiment. FIG. 5 is a block diagram showing another configuration example of an encoding / decoding system according to an embodiment. FIG. 6 is a block diagram showing another configuration example of an encoding device according to an embodiment. FIG. 7 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 8 is a block diagram showing another configuration example of an encoding device according to an embodiment. FIG. 9 is a block diagram showing another configuration example of a decoding device according to an embodiment. FIG. 10 is a flow diagram showing processing of an encoding device according to an embodiment. FIG. 11 is an explanatory diagram conceptually showing encoding of a mesh frame according to an embodiment. FIG. 12 is a flow diagram showing processing of a decoding device according to an embodiment.1 is an explanatory diagram conceptually illustrating decoding of a mesh frame according to an embodiment. FIG. 2 is a block diagram illustrating an example of a configuration of a decoding device according to an embodiment. FIG. 3 is a block diagram illustrating an example of a configuration of a decoding device according to an embodiment. FIG. 4 is an explanatory diagram illustrating an example of subdivision according to an embodiment. FIG. 5 is an explanatory diagram illustrating an example of displacement of vertices after displacement after subdivision according to an embodiment. FIG. 6 is an explanatory diagram illustrating an example of vertices of an original mesh according to an embodiment. FIG. 7 is an explanatory diagram illustrating an example of a mesh according to an embodiment. FIG. 8 is an explanatory diagram illustrating an example of division of a mesh into sub-meshes according to an embodiment. FIG. 9 is a first explanatory diagram illustrating an example of packing of displacement information into an image frame according to an embodiment. FIG. 10 is an explanatory diagram illustrating an example of packing of displacement information into an image frame according to an embodiment. FIG. 11 is an explanatory diagram illustrating an example of packing of displacement information into an image frame according to an embodiment. FIG. 12 is an explanatory diagram illustrating an example of packing of displacement information into an image frame according to an embodiment. FIG. 13 is a block diagram illustrating a detailed example of a configuration of a decoding device according to an embodiment. FIG. 14 is an explanatory diagram illustrating coordinates of vertices in a three-dimensional mesh according to an embodiment. FIG. 15 is an explanatory diagram illustrating prediction information according to an embodiment. FIG. 16 is a block diagram illustrating an example of a configuration of an encoding device according to an embodiment. FIG. 17 is a flow diagram illustrating a specific example of encoding processing according to an embodiment. FIG. 18 is a block diagram illustrating an example of a configuration of a decoding device according to an embodiment. FIG. 19 is a flow diagram illustrating a specific example of decoding processing according to an embodiment. 1 is an explanatory diagram showing an example of syntax according to an embodiment. FIG. 1 is an explanatory diagram showing an example of syntax according to an embodiment. FIG. 2 is an explanatory diagram showing an example of syntax according to an embodiment. FIG. 3 is a block diagram showing an example of a configuration of an encoding device according to an embodiment. FIG. 4 is a block diagram showing an example of a configuration of a decoding device according to an embodiment. FIG. 5 is an explanatory diagram showing a positional relationship of three-dimensional points according to an embodiment. FIG. 6 is an explanatory diagram showing a method of generating LoD according to an embodiment. FIG. 7 is an explanatory diagram showing a method of generating LoD according to an embodiment. FIG. 8 is an explanatory diagram showing a method of generating a predicted value of a displacement vector according to an embodiment. FIG. 9 is an explanatory diagram showing an example of calculation of a predicted value according to an embodiment. FIG. 10 is an explanatory diagram showing an example of calculation of transform coefficients according to an embodiment. FIG. 11 is an explanatory diagram showing an example of inter-prediction of transform coefficients according to an embodiment. FIG. 12 is an explanatory diagram showing an example of syntax according to an embodiment.1 is an explanatory diagram showing an example of syntax in an embodiment; FIG. 1 is an explanatory diagram showing an example of calculation of a prediction residual in an embodiment; FIG. 1 is an explanatory diagram showing an example of calculation of a prediction residual in an embodiment; FIG. 1 is an explanatory diagram showing an example of predicted value information of a displacement vector in an embodiment; FIG. 2 is an explanatory diagram showing a method for generating a predicted value of a displacement vector in an embodiment; FIG. 2 is an explanatory diagram showing an example of predicted value information of a displacement vector in an embodiment; FIG. 3 is an explanatory diagram showing an example of predicted value information of a displacement vector in an embodiment; FIG. 4 is an explanatory diagram showing a method for generating a predicted value of a displacement vector in an embodiment; FIG. 5 is an explanatory diagram showing an example of predicted value information of a displacement vector in an embodiment; FIG. 6 is an explanatory diagram showing an example of predicted value information of a motion vector in an embodiment; FIG. 7 is an explanatory diagram showing an example of predicted value information of a motion vector in an embodiment; FIG. 8 is an explanatory diagram showing an example of predicted value information of a motion vector in an embodiment; FIG. 1 is an explanatory diagram showing an example of a reference destination of a DVG in an embodiment. FIG. 2 is an explanatory diagram showing an example of a syntax in an embodiment. FIG. 3 is an explanatory diagram showing an example of a reference destination of a DVG in an embodiment. FIG. 4 is an explanatory diagram showing an example of a reference destination of a DVG in an embodiment. FIG. 5 is an explanatory diagram showing an example of a syntax in an embodiment. FIG. 6 is an explanatory diagram showing an example of a syntax in an embodiment. FIG. 7 is a flow diagram showing an example of an encoding process in an embodiment. FIG. 8 is a flow diagram showing an example of a decoding process in an embodiment. FIG. 9 is an explanatory diagram showing an example of a frame structure in an embodiment. FIG. 10 is an explanatory diagram showing an example of a process for determining whether or not to apply inter prediction in an embodiment. FIG. 11 is a flow diagram showing a specific example of an encoding process in an embodiment. FIG. 12 is a flow diagram showing a specific example of an encoding process in an embodiment.1 is a flow diagram showing a specific example of encoding processing according to an embodiment. FIG. 2 is a flow diagram showing a specific example of encoding processing according to an embodiment. FIG. 3 is a flow diagram showing a specific example of decoding processing according to an embodiment. FIG. 4 is a flow diagram showing a specific example of decoding processing according to an embodiment. FIG. 5 is a flow diagram showing a specific example of decoding processing according to an embodiment. FIG. 6 is a flow diagram showing a specific example of decoding processing according to an embodiment. FIG. 7 is a flow diagram showing a specific example of decoding processing according to an embodiment. FIG. 8 is a flow diagram showing a specific example of decoding processing according to an embodiment. FIG. 9 is an explanatory diagram showing an example of syntax in an embodiment. FIG. 10 is an explanatory diagram showing an example of syntax in an embodiment. FIG. 11 is a block diagram showing an example configuration of an encoding device according to an embodiment. FIG. 12 is an explanatory diagram showing an example of an I full mesh according to an embodiment. FIG. 13 is an explanatory diagram showing an example of a P full mesh according to an embodiment. FIG. 14 is an explanatory diagram showing an example of header information common to sequences according to an embodiment. FIG. 15 is an explanatory diagram showing an example of an I full mesh according to an embodiment. FIG. 16 is an explanatory diagram showing an example of a P full mesh according to an embodiment. FIG. 17 is an explanatory diagram showing an example of a P full mesh according to an embodiment. FIG. 18 is an explanatory diagram showing an example of header information common to sequences according to an embodiment. FIG. 19 is an explanatory diagram showing an example of the syntax of a reference frame list according to an embodiment. FIG. 19 is an explanatory diagram showing an example of a reference frame list according to an embodiment. FIG. 1 is an explanatory diagram showing an example of header information corresponding to a full mesh frame according to an embodiment. FIG. 2 is an explanatory diagram showing an example of header information corresponding to a sub-mesh or coding unit according to an embodiment. FIG. 3 is an explanatory diagram showing an example of header information corresponding to a sub-mesh or coding unit according to an embodiment. FIG. 4 is an explanatory diagram showing an example of header information common to sequences according to an embodiment. FIG. 5 is an explanatory diagram showing an example of header information corresponding to a full mesh frame according to an embodiment. FIG. 6 is an explanatory diagram showing an example of header information corresponding to a sub-mesh or coding unit according to an embodiment. FIG. 7 is an explanatory diagram showing an example of header information common to sequences according to an embodiment. FIG. 8 is an explanatory diagram showing an example of header information corresponding to a full mesh frame according to an embodiment.1 is an explanatory diagram showing an example of header information corresponding to a sub-mesh or a coding unit according to an embodiment. FIG. 2 is a flow diagram showing an example of an encoding process according to an embodiment. FIG. 3 is a flow diagram showing an example of a decoding process according to an embodiment. FIG. 4 is a block diagram showing an example of a configuration of an encoding device according to an embodiment. FIG. 5 is a flow diagram showing an example of an encoding process according to an embodiment. FIG. 6 is a block diagram showing an example of a configuration of a decoding device according to an embodiment. FIG. 7 is a block diagram showing an example of a configuration of an encoding device according to an embodiment. FIG. 8 is a block diagram showing an example of a configuration of a decoding device according to an embodiment. FIG. 9 is an explanatory diagram showing an example of a syntax according to an embodiment. FIG. 10 is an explanatory diagram showing an example of a syntax according to an embodiment. FIG. 11 is an explanatory diagram showing an example of a syntax according to an embodiment. FIG. 12 is an explanatory diagram showing an example of a syntax according to an embodiment. FIG. 13 is an explanatory diagram showing an example of a syntax according to an embodiment. FIG. 14 is an explanatory diagram showing an example of a syntax according to an embodiment. FIG. 15 is an explanatory diagram showing an example of a syntax according to an embodiment. 1 is a flow diagram showing an example of encoding processing according to an embodiment. 2 is a flow diagram showing an example of decoding processing according to an embodiment. 3 is a block diagram showing an example of a configuration of an encoding device according to an embodiment. 4 is a block diagram showing an example of a configuration of a decoding device according to an embodiment. 5 is a flow diagram showing an example of encoding processing according to an embodiment. 6 is a flow diagram showing an example of decoding processing according to an embodiment. 7 is an explanatory diagram showing an example of syntax in an embodiment. 8 is an explanatory diagram showing an example of syntax in an embodiment. 9 is an explanatory diagram showing an example of syntax in an embodiment. 10 is an explanatory diagram explaining a method for determining a quantization parameter according to an embodiment. 11 is a flow diagram showing a specific example of encoding processing according to an embodiment. 12 is a flow diagram showing a specific example of decoding processing according to an embodiment. 13 is an explanatory diagram showing an example of syntax in an embodiment.FIG. 1 is an explanatory diagram showing an example of syntax in an embodiment. FIG. 2 is an explanatory diagram showing an example of syntax in an embodiment. FIG. 3 is a flow diagram showing a specific example of encoding processing in an embodiment. FIG. 4 is an explanatory diagram showing an example of syntax in an embodiment. FIG. 5 is an explanatory diagram showing an example of syntax in an embodiment. FIG. 6 is an explanatory diagram showing an example of syntax in an embodiment. FIG. 7 is a flow diagram showing a specific example of encoding processing in an embodiment. FIG. 8 is a flow diagram showing a specific example of encoding processing in an embodiment. FIG. 9 is a flow diagram showing an example of encoding processing in an embodiment. FIG. 10 is a flow diagram showing an example of decoding processing in an embodiment.
[0009] <Summary of the Disclosure> Three-dimensional (3D) meshes are used in computer graphics images, for example, which may be composed of multiple temporally distinct frames, each of which may be represented by a 3D mesh.
[0010] A 3D mesh is composed of vertex information indicating the positions of each of the vertices in 3D space, connectivity information indicating the connections between the vertices, and attribute information indicating the attributes of each vertex or face. Each face is constructed according to the connectivity between the vertices. Various computer graphics images can be expressed using such 3D meshes.
[0011] Furthermore, for transmission and storage of the 3D mesh, efficient encoding and decoding of the 3D mesh is expected. For efficient encoding and decoding of the 3D mesh, arithmetic coding and decoding may be used.
[0012] Further improvements are desired in the encoding or decoding process for three-dimensional data. The present disclosure aims to improve the encoding or decoding process for three-dimensional data.
[0013] Below, examples of inventions that can be obtained from the disclosure of this specification will be given, and the effects and the like that can be obtained from these inventions will be explained.
[0014] (1) A three-dimensional data encoding method for generating a bitstream including data to be encoded and one or more flags for determining whether or not to perform a quantization process on the data to be encoded.
[0015] According to the above aspect, the three-dimensional data encoding method can generate a bitstream including one or more flags, which can indicate whether or not quantization is to be performed, and whether or not quantization is to be performed is determined according to the one or more flags. As a result, when video encoding is used after the quantization process to perform lossless encoding, it can be determined according to the one or more flags that quantization will not be performed, thereby reducing the amount of processing compared to when quantization is performed. Furthermore, when arithmetic encoding is used after the quantization process to improve encoding efficiency through quantization, it can be determined according to the one or more flags that quantization will be performed, thereby easily improving encoding efficiency. Furthermore, since the bitstream includes one or more flags, the three-dimensional data decoding device can determine whether or not the three-dimensional data encoding device has performed quantization based on the one or more flags, contributing to the three-dimensional data decoding device appropriately performing decoding (e.g., inverse quantization). In this way, the above three-dimensional data encoding method can contribute to improving encoding processes.
[0016] (2) The three-dimensional data encoding method according to (1), wherein each of the one or more flags is assigned to each of a plurality of levels into which the encoding target data is hierarchically divided.
[0017] According to the above aspect, whether or not quantization processing can be performed can be indicated for each of the divided multiple levels, and the amount of information included in the bitstream can be reduced compared to when whether or not quantization processing can be performed is indicated for the entire data to be encoded, which may lead to an improvement in the encoding processing. In this way, the above three-dimensional data encoding method can contribute to the improvement of the encoding processing.
[0018] (3) A three-dimensional data encoding method as described in (2), wherein the plurality of levels include a sequence level including the entire data to be encoded, a frame level into which the sequence level is divided in time, and a patch level into which the frame level is divided in space.
[0019] According to the above aspect, it is possible to realize a three-dimensional data encoding method in which the plurality of levels includes a sequence level, a frame level, and a patch level.
[0020] (4) The three-dimensional data encoding method according to (2), wherein the one or more flags include a first flag indicating whether or not a quantization parameter exists.
[0021] According to the above aspect, the presence or absence of a quantization parameter can be indicated by the first flag, and the amount of information included in the generated bitstream can be reduced compared to when the presence or absence of a quantization parameter is indicated by information other than the first flag, which may improve the encoding process. In this way, the above three-dimensional data encoding method can contribute to improving the encoding process.
[0022] (5) The three-dimensional data encoding method according to (4), wherein, when the quantization process is skipped, the first flag assigned to each of the plurality of levels is set to 0.
[0023] According to the above aspect, it is possible to realize a three-dimensional data encoding method in which the first flag assigned to each of the plurality of levels is set to 0 when the quantization process is skipped.
[0024] (6) The three-dimensional data encoding method according to (2), wherein the one or more flags include a second flag indicating whether or not the quantization process is to be skipped.
[0025] According to the above aspect, the second flag can indicate whether or not the quantization process can be skipped, and the amount of information included in the generated bitstream can be reduced compared to when the ability to skip the quantization process is indicated by information other than the second flag, which may improve the encoding process. In this way, the above three-dimensional data encoding method can contribute to improving the encoding process.
[0026] (7) The three-dimensional data encoding method according to (2), wherein whether or not to perform the quantization process is determined based on at least one of the flags assigned to each of the plurality of levels.
[0027] According to the above aspect, whether or not the quantization process is to be performed can be indicated by at least one flag, and compared to a case where a large number of flags are required to indicate whether or not the quantization process is to be performed, the amount of information processing performed by the encoding process can be reduced, which may lead to an improvement in the encoding process. In this way, the above three-dimensional data encoding method can contribute to the improvement of the encoding process.
[0028] (8) The three-dimensional data encoding method according to (7), wherein whether or not to perform the quantization process is determined based on a combination of the flags assigned to each of the plurality of levels.
[0029] According to the above aspect, it is possible to realize a three-dimensional data encoding method in which it is determined whether or not to perform quantization processing based on a combination of flags assigned to each of a plurality of levels.
[0030] (9) The three-dimensional data encoding method according to (7), wherein it is determined to skip the quantization process when all of the flags assigned to each of the plurality of levels have the same value.
[0031] According to the above aspect, it is possible to realize a three-dimensional data encoding method in which it is determined to skip the quantization process when all of the flags assigned to the respective multiple levels have the same value.
[0032] (10) The three-dimensional data encoding method according to (1), wherein, when the quantization process is performed, information signaling a quantization parameter is included in the bitstream.
[0033] According to the above aspect, the three-dimensional data encoding method can transmit the information to the three-dimensional data decoding device simply by outputting a bitstream, and there is no need to output other information, so the amount of information processing performed by the encoding process can be reduced. This may potentially improve the encoding process. Furthermore, the three-dimensional data decoding device can determine a quantization parameter based on the information, contributing to appropriate decoding (e.g., inverse quantization). In this way, the above three-dimensional data encoding method can contribute to improving the encoding process.
[0034] (11) The three-dimensional data encoding method according to (1), wherein the encoding target data includes information corresponding to each of a plurality of vertices in three-dimensional space.
[0035] According to the above aspect, it is possible to realize a three-dimensional data encoding method in which the encoding target data includes information corresponding to each of a plurality of vertices in a three-dimensional space.
[0036] (12) A three-dimensional data decoding method that acquires a bitstream, acquires from the acquired bitstream data to be decoded and one or more flags for determining whether or not to perform inverse quantization processing on the data to be decoded, performs a decoding process to decode the data to be decoded, and determines whether or not to perform inverse quantization processing in the decoding process based on the one or more flags.
[0037] According to the above aspect, the three-dimensional data decoding method can appropriately determine whether or not to perform inverse quantization processing on the data to be decoded, that is, can appropriately perform inverse quantization processing. As a result, when the three-dimensional data decoding method does not perform inverse quantization processing based on one or more flags, it can reduce the processing load due to the execution of the inverse quantization processing, which can contribute to improving the decoding processing.
[0038] (13) The three-dimensional data decoding method according to (12), wherein each of the one or more flags is assigned to each of a plurality of levels into which the decoding target data is hierarchically divided.
[0039] According to the above aspect, whether or not to perform the inverse quantization process can be indicated for each of the divided levels, and the inverse quantization process can be performed more appropriately than when whether or not to perform the inverse quantization process is indicated for the entire data to be decoded, which may lead to an improvement in the decoding process. In this way, the above three-dimensional data decoding method can contribute to the improvement of the decoding process.
[0040] (14) A three-dimensional data decoding method as described in (13), wherein the plurality of levels include a sequence level including the entire data to be decoded, a frame level into which the sequence level is time-divided, and a patch level into which the frame level is spatially divided.
[0041] According to the above aspect, it is possible to realize a three-dimensional data decoding method in which the plurality of levels includes a sequence level, a frame level, and a patch level.
[0042] (15) The three-dimensional data decoding method according to (13), wherein the one or more flags include a first flag indicating whether or not an inverse quantization parameter exists.
[0043] According to the above aspect, the presence or absence of an inverse quantization parameter can be indicated by the first flag, and the amount of information included in the acquired bitstream can be reduced compared to when the presence or absence of an inverse quantization parameter is indicated by information other than the first flag, which may result in an improvement in the decoding process. In this way, the above three-dimensional data decoding method can contribute to the improvement of the decoding process.
[0044] (16) The three-dimensional data decoding method according to (15), wherein, when the inverse quantization process is skipped, the first flag assigned to each of the plurality of levels is set to 0.
[0045] According to the above aspect, it is possible to realize a three-dimensional data decoding method in which the first flag assigned to each of the plurality of levels is set to 0 when the inverse quantization process is skipped.
[0046] (17) The three-dimensional data decoding method according to (13), wherein the one or more flags include a second flag indicating whether or not the inverse quantization process is to be skipped.
[0047] According to the above aspect, the second flag can indicate whether or not the inverse quantization process can be skipped, and the amount of information included in the acquired bitstream can be reduced compared to when the inverse quantization process can be skipped by information other than the second flag, which may improve the decoding process. In this way, the above three-dimensional data decoding method can contribute to improving the decoding process.
[0048] (18) The three-dimensional data decoding method according to (13), wherein it is determined whether or not to perform the inverse quantization process based on at least one of the flags assigned to each of the plurality of levels.
[0049] According to the above aspect, whether or not the inverse quantization process is to be performed can be indicated by at least one flag, and compared to a case where a large number of flags are required to indicate whether or not the inverse quantization process is to be performed, the amount of information processing performed by the decoding process can be reduced, which may lead to an improvement in the decoding process. In this way, the above three-dimensional data decoding method can contribute to the improvement of the decoding process.
[0050] (19) The three-dimensional data decoding method according to (18), wherein whether or not to perform the inverse quantization process is determined based on a combination of the flags assigned to each of the plurality of levels.
[0051] According to the above aspect, it is possible to realize a three-dimensional data decoding method in which it is determined whether or not to perform inverse quantization processing based on a combination of flags assigned to each of a plurality of levels.
[0052] (20) The three-dimensional data decoding method according to (18), wherein it is determined to skip the inverse quantization process if all of the flags assigned to the respective levels have the same value.
[0053] According to the above aspect, it is possible to realize a three-dimensional data decoding method that determines to skip the inverse quantization process when all of the flags assigned to the respective levels have the same value.
[0054] (21) The three-dimensional data decoding method according to (12), wherein the data to be decoded includes information corresponding to each of a plurality of vertices in a three-dimensional space.
[0055] According to the above aspect, it is possible to realize a three-dimensional data decoding method in which the data to be decoded includes information corresponding to each of a plurality of vertices in a three-dimensional space.
[0056] (22) A three-dimensional data encoding device comprising a memory and a circuit accessible to the memory, the circuit generating, in operation, a bitstream including data to be encoded and one or more flags for determining whether or not to perform a quantization process on the data to be encoded.
[0057] According to the above aspect, the three-dimensional data encoding device achieves the same effects as the above three-dimensional data encoding method.
[0058] (23) A three-dimensional data decoding device comprising a memory and a circuit that can access the memory, wherein the circuit, in operation, acquires a bitstream, acquires from the acquired bitstream data to be decoded and one or more flags for determining whether or not to perform inverse quantization processing on the data to be decoded, executes a decoding process to decode the data to be decoded, and determines, in the decoding process, whether or not to perform inverse quantization processing based on the one or more flags.
[0059] According to the above aspect, the three-dimensional data decoding device achieves the same effects as the above three-dimensional data decoding method.
[0060] These comprehensive or specific aspects may be realized as a system, an apparatus, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, an integrated circuit, a computer program, or a recording medium.
[0061] Hereinafter, the embodiments will be specifically described with reference to the drawings.
[0062] The embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present invention. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concepts are described as optional components.
[0063] (Embodiment) In this embodiment, an encoding method, a decoding method, etc. will be described.
[0064] <Expressions and Terms> The following expressions and terms are used herein.
[0065] (1) Three-dimensional Mesh A three-dimensional mesh is a collection of multiple faces, and represents, for example, a three-dimensional object. A three-dimensional mesh is mainly composed of vertex information, connectivity information, and attribute information. A three-dimensional mesh may be expressed as a polygon mesh or a mesh. A three-dimensional mesh may also vary over time. A three-dimensional mesh may include metadata related to the vertex information, connectivity information, and attribute information, and may also include other additional information.
[0066] (2) Vertex Information Vertex information is information indicating a vertex. For example, the vertex information indicates the position of a vertex in a three-dimensional space. Furthermore, a vertex corresponds to a vertex of a face that constitutes a three-dimensional mesh. Vertex information may be expressed as "geometry." Furthermore, vertex information may be expressed as position information.
[0067] (3) Connection Information Connection information is information that indicates connections between vertices. For example, connection information indicates connections for forming faces or edges of a three-dimensional mesh. Connection information may be expressed as "Connectivity." Connection information may also be expressed as face information.
[0068] (4) Attribute Information Attribute information is information that indicates attributes of a vertex or a face. For example, attribute information indicates attributes such as a color, an image, and a normal vector associated with a vertex or a face. Attribute information may be expressed as "texture."
[0069] (5) Faces A face is an element that constitutes a three-dimensional mesh. Specifically, a face is a polygon on a plane in three-dimensional space. For example, a face can be defined as a triangle in three-dimensional space.
[0070] (6) Plane A plane is a two-dimensional plane in a three-dimensional space. For example, a polygon is formed on a plane, and multiple polygons are formed on multiple planes.
[0071] (7) Bitstream: A bitstream corresponds to coded information. A bitstream may also be referred to as a stream, a coded bitstream, a compressed bitstream, or a coded signal.
[0072] (8) Encoding and Decoding The term encoding may be substituted with terms such as storing, including, writing, describing, signaling, sending, notifying, saving, or compressing, and these terms may be interchangeable. For example, encoding information may mean including the information in a bitstream. Also, encoding information into a bitstream may mean encoding the information to generate a bitstream that includes the encoded information.
[0073] Additionally, the term "decode" may be replaced with terms such as "read," "decode," "read," "load," "derive," "obtain," "receive," "extract," "reconstruct," "reconstruct," "decompress," or "decompress," and these terms may be interchangeable. For example, decoding information may mean obtaining information from a bitstream. Decoding information from a bitstream may mean decoding the bitstream to obtain information contained in the bitstream.
[0074] (9) Ordinal Numbers In the description, ordinal numbers such as first and second may be assigned to components, etc. These ordinal numbers may be changed as appropriate. Furthermore, new ordinal numbers may be assigned to components, etc., or removed. Furthermore, these ordinal numbers may be assigned to elements in order to identify them, and may not correspond to a meaningful order.
[0075] <Three-dimensional mesh> Fig. 1 is a conceptual diagram showing a three-dimensional mesh according to this embodiment. A three-dimensional mesh is composed of multiple faces. For example, each face is a triangle. The vertices of these triangles are defined in three-dimensional space. The three-dimensional mesh then represents a three-dimensional object. Each face may have a color or an image.
[0076] FIG. 2 is a conceptual diagram showing the basic elements of a three-dimensional mesh according to this embodiment. A three-dimensional mesh is composed of vertex information, connection information, and attribute information. The vertex information indicates the positions of the vertices of a face in three-dimensional space. The connection information indicates the connections between the vertices. A face can be identified by the vertex information and connection information. In other words, a colorless three-dimensional object is formed in three-dimensional space by the vertex information and connection information.
[0077] The attribute information may be associated with a vertex or a face. The attribute information associated with a vertex may be expressed as "Attribute Per Point." The attribute information associated with a vertex may indicate an attribute of the vertex itself, or may indicate an attribute of a face connected to the vertex.
[0078] For example, a color may be associated with a vertex as attribute information. The color associated with a vertex may be the color of the vertex itself, or the color of a face connected to the vertex. The color of a face may be the average of multiple colors associated with multiple vertices of the face. Furthermore, a normal vector may be associated with a vertex or a face as attribute information. Such a normal vector can represent the front and back of a face.
[0079] A two-dimensional image may be associated with a surface as attribute information. The two-dimensional image associated with a surface is also expressed as a texture image or an "Attribute Map." Information indicating mapping between the surface and the two-dimensional image may be associated with the surface as attribute information. Information indicating such mapping may be expressed as mapping information, vertex information of a texture image, texture coordinates, or "Attribute UV Coordinate."
[0080] Furthermore, information such as color, image, and moving image used as attribute information may be expressed as "parametric space."
[0081] The attribute information allows texture to be reflected on the three-dimensional object. That is, a three-dimensional object having color is formed in three-dimensional space based on the vertex information, connection information, and attribute information.
[0082] In the above, the attribute information is associated with the vertices or faces, but it may also be associated with the edges.
[0083] 3 is a conceptual diagram illustrating mapping according to this embodiment. For example, a region of a two-dimensional image on a two-dimensional plane can be mapped onto a surface of a three-dimensional mesh in three-dimensional space. Specifically, coordinate information of the region in the two-dimensional image is associated with the surface of the three-dimensional mesh. As a result, an image of the mapped region in the two-dimensional image is reflected on the surface of the three-dimensional mesh.
[0084] By using the mapping, the 2D image used as attribute information can be separated from the 3D mesh. For example, in encoding the 3D mesh, the 2D image may be encoded by an image encoding method or a video encoding method.
[0085] <System Configuration> Fig. 4 is a block diagram showing an example of the configuration of a coding / decoding system according to this embodiment. In Fig. 4, the coding / decoding system includes a coding device 100 and a decoding device 200.
[0086] For example, the encoding device 100 obtains a three-dimensional mesh and encodes the three-dimensional mesh into a bitstream. Then, the encoding device 100 outputs the bitstream to the network 300. For example, the bitstream includes the encoded three-dimensional mesh and control information for decoding the encoded three-dimensional mesh. By encoding the three-dimensional mesh, information about the three-dimensional mesh is compressed.
[0087] The network 300 transmits a bitstream from the encoding device 100 to the decoding device 200. The network 300 may be the Internet, a wide area network (WAN), a local area network (LAN), or a combination of these. The network 300 is not necessarily limited to bidirectional communication, and may be a unidirectional communication network for terrestrial digital broadcasting, satellite broadcasting, or the like.
[0088] Furthermore, the network 300 can be replaced by a recording medium such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc (registered trademark)).
[0089] The decoding device 200 obtains a bitstream and decodes a three-dimensional mesh from the bitstream. By decoding the three-dimensional mesh, information about the three-dimensional mesh is expanded. For example, the decoding device 200 decodes the three-dimensional mesh according to a decoding method corresponding to the encoding method used by the encoding device 100 to encode the three-dimensional mesh. That is, the encoding device 100 and the decoding device 200 perform encoding and decoding according to encoding methods and decoding methods that correspond to each other.
[0090] The 3D mesh before encoding may also be referred to as an original 3D mesh, and the 3D mesh after decoding may also be referred to as a reconstructed 3D mesh.
[0091] 5 is a block diagram showing an example of the configuration of a coding device 100 according to this embodiment. For example, the coding device 100 includes a vertex information encoder 101, a connection information encoder 102, and an attribute information encoder 103.
[0092] The vertex information encoder 101 is an electrical circuit that encodes vertex information. For example, the vertex information encoder 101 encodes the vertex information into a bitstream according to a format defined for the vertex information.
[0093] The connection information encoder 102 is an electrical circuit that encodes the connection information, for example, the connection information encoder 102 encodes the connection information into a bitstream according to a format defined for the connection information.
[0094] The attribute information encoder 103 is an electric circuit that encodes the attribute information. For example, the attribute information encoder 103 encodes the attribute information into a bit stream in accordance with a format defined for the attribute information.
[0095] The vertex information, connectivity information, and attribute information may be coded using variable-length coding or fixed-length coding, such as Huffman coding or context-adaptive binary arithmetic coding (CABAC).
[0096] The vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be integrated together, or each of the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be further subdivided into multiple components.
[0097] 6 is a block diagram showing another example of the configuration of the encoding device 100 according to this embodiment. For example, the encoding device 100 includes a pre-processor 104 and a post-processor 105 in addition to the configuration shown in FIG.
[0098] The preprocessor 104 is an electrical circuit that performs processing before encoding the vertex information, connectivity information, and attribute information. For example, the preprocessor 104 may perform a conversion process, a separation process, a multiplexing process, or the like on the 3D mesh before encoding. More specifically, for example, the preprocessor 104 may separate the vertex information, connectivity information, and attribute information from the 3D mesh before encoding.
[0099] The post-processor 105 is an electrical circuit that performs processing after the vertex information, connection information, and attribute information are encoded. For example, the post-processor 105 may perform conversion processing, separation processing, multiplexing processing, or the like on the encoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 105 may multiplex the encoded vertex information, connection information, and attribute information into a bitstream. Furthermore, for example, the post-processor 105 may further perform variable-length coding on the encoded vertex information, connection information, and attribute information.
[0100] 7 is a block diagram showing an example of the configuration of a decoding device 200 according to this embodiment. For example, the decoding device 200 includes a vertex information decoder 201, a connection information decoder 202, and an attribute information decoder 203.
[0101] The vertex information decoder 201 is an electrical circuit that decodes vertex information. For example, the vertex information decoder 201 decodes vertex information from a bitstream according to a format defined for the vertex information.
[0102] The connection information decoder 202 is an electrical circuit that decodes the connection information, for example, the connection information decoder 202 decodes the connection information from the bitstream according to a format defined for the connection information.
[0103] The attribute information decoder 203 is an electric circuit that decodes the attribute information. For example, the attribute information decoder 203 decodes the attribute information from the bitstream in accordance with a format defined for the attribute information.
[0104] The vertex information, connection information, and attribute information may be decoded using variable length decoding or fixed length decoding, which may correspond to Huffman coding, context-adaptive binary arithmetic coding (CABAC), or the like.
[0105] The vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be integrated together, or each of the vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be further subdivided into multiple components.
[0106] 8 is a block diagram showing another example of the configuration of the decoding device 200 according to this embodiment. For example, the decoding device 200 includes a pre-processor 204 and a post-processor 205 in addition to the configuration shown in FIG.
[0107] The preprocessor 204 is an electrical circuit that performs processing before decoding the vertex information, connection information, and attribute information. For example, the preprocessor 204 may perform conversion processing, separation processing, multiplexing processing, or the like on the bitstream before decoding the vertex information, connection information, and attribute information.
[0108] More specifically, for example, the preprocessor 204 may separate a sub-bitstream corresponding to vertex information, a sub-bitstream corresponding to connectivity information, and a sub-bitstream corresponding to attribute information from the bitstream. Also, for example, the preprocessor 204 may perform variable-length decoding on the bitstream in advance before decoding the vertex information, connectivity information, and attribute information.
[0109] The post-processor 205 is an electrical circuit that performs processing after the vertex information, connection information, and attribute information are decoded. For example, the post-processor 205 may perform conversion processing, separation processing, multiplexing processing, or the like on the decoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 205 may multiplex the decoded vertex information, connection information, and attribute information onto a three-dimensional mesh.
[0110] <Bitstream> Vertex information, connection information, and attribute information are coded and stored in a bitstream. The relationship between this information and the bitstream is shown below.
[0111] 9 is a conceptual diagram showing an example of the configuration of a bitstream according to this embodiment. In this example, connection information, vertex information, and attribute information are integrated in the bitstream. For example, the connection information, vertex information, and attribute information may be included in a single file.
[0112] Furthermore, multiple portions of this information may be stored sequentially, such as a first portion of connection information, a first portion of vertex information, a first portion of attribute information, a second portion of connection information, a second portion of vertex information, a second portion of attribute information, etc. These multiple portions may correspond to multiple portions that are different in time, multiple portions that are different in space, or multiple different faces.
[0113] Furthermore, the storage order of the connection information, vertex information, and attribute information is not limited to the above example, and a storage order different from the above example may be used.
[0114] 10 is a conceptual diagram showing another example of the configuration of a bitstream according to this embodiment. In this example, a plurality of files are included in the bitstream, and connection information, vertex information, and attribute information are stored in different files. Here, a file containing connection information, a file containing vertex information, and a file containing attribute information are shown, but the storage format is not limited to this example. For example, two types of information among the connection information, vertex information, and attribute information may be included in one file, and the remaining type of information may be included in another file.
[0115] Alternatively, the information may be split and stored in more files. For example, multiple pieces of connectivity information may be stored in multiple files, multiple pieces of vertex information may be stored in multiple files, or multiple pieces of attribute information may be stored in multiple files. These multiple pieces may correspond to multiple temporally different pieces, multiple spatially different pieces, or multiple different faces.
[0116] Furthermore, the storage order of the connection information, vertex information, and attribute information is not limited to the above example, and a storage order different from the above example may be used.
[0117] 11 is a conceptual diagram showing another example of the configuration of a bitstream according to this embodiment. In this example, the bitstream is composed of multiple separable sub-bitstreams, and connection information, vertex information, and attribute information are stored in different sub-bitstreams.
[0118] Here, a sub-bitstream containing connection information, a sub-bitstream containing vertex information, and a sub-bitstream containing attribute information are shown, but the storage format is not limited to this example.
[0119] For example, two types of information among the connection information, vertex information, and attribute information may be included in one sub-bitstream, and the remaining type of information may be included in another sub-bitstream. Specifically, attribute information of a two-dimensional image or the like may be stored in a sub-bitstream that complies with an image coding method, separate from the sub-bitstreams of the connection information and vertex information.
[0120] Also, each sub-bitstream may include multiple files, and multiple pieces of connectivity information may be stored in multiple files, multiple pieces of vertex information may be stored in multiple files, or multiple pieces of attribute information may be stored in multiple files.
[0121] 9, 10, and 11, and a storage order different from the above examples may be used. For example, the vertex information, connection information, and attribute information may be stored in the bitstream in this order. Alternatively, the connection information, connection information, and attribute information may be stored in the bitstream in any of the following orders: connection information, attribute information, and vertex information; vertex information, attribute information, and connection information; attribute information, connection information, and vertex information; or attribute information, vertex information, and connection information.
[0122] Furthermore, each of the connection information, vertex information, and attribute information may be divided into a plurality of data, and the plurality of data may be stored in a cyclical or random order within the bitstream.
[0123] 12 is a block diagram showing a specific example of an encoding / decoding system according to this embodiment. In FIG. 12, the encoding / decoding system includes a three-dimensional data encoding system 110, a three-dimensional data decoding system 210, and an external connector 310.
[0124] The three-dimensional data encoding system 110 includes a controller 111, an input / output processor 112, a three-dimensional data encoder 113, a three-dimensional data generator 115, and a system multiplexer 114. The three-dimensional data decoding system 210 includes a controller 211, an input / output processor 212, a three-dimensional data decoder 213, a system demultiplexer 214, a presenter 215, and a user interface 216.
[0125] In the three-dimensional data encoding system 110, sensor data is input from a sensor terminal to a three-dimensional data generator 115. The three-dimensional data generator 115 generates three-dimensional data, such as point cloud data or mesh data, from the sensor data and inputs it to a three-dimensional data encoder 113.
[0126] For example, the three-dimensional data generator 115 generates vertex information, and generates connection information and attribute information corresponding to the vertex information. The three-dimensional data generator 115 may process the vertex information when generating the connection information and attribute information. For example, the three-dimensional data generator 115 may reduce the amount of data by deleting duplicate vertices, or may transform the vertex information (such as by shifting its position, rotating it, or normalizing it). The three-dimensional data generator 115 may also render the attribute information.
[0127] Furthermore, although the three-dimensional data generator 115 is a component of the three-dimensional data encoding system 110 in FIG. 12, it may be arranged externally and independently of the three-dimensional data encoding system 110.
[0128] The sensor terminal that provides the sensor data for generating the three-dimensional data may be, for example, a moving body such as an automobile, a flying object such as an airplane, a mobile terminal, a camera, etc. Furthermore, a distance sensor such as a LIDAR, a millimeter wave radar, an infrared sensor, or a range finder, a stereo camera, or a combination of multiple monocular cameras may also be used as the sensor terminal.
[0129] The sensor data may be the distance (position) of the object, monocular camera images, stereo camera images, color, reflectance, sensor attitude, orientation, gyro, sensing position (GPS information or altitude), speed, acceleration, sensing time, temperature, air pressure, humidity, or magnetism.
[0130] The three-dimensional data encoder 113 corresponds to the encoding device 100 shown in FIG. 5 and other figures. For example, the three-dimensional data encoder 113 encodes three-dimensional data to generate encoded data. The three-dimensional data encoder 113 also generates control information when encoding the three-dimensional data. The three-dimensional data encoder 113 then inputs the encoded data together with the control information to the system multiplexer 114.
[0131] The encoding method for the three-dimensional data may be an encoding method using geometry or an encoding method using a video codec. Here, the encoding method using geometry may also be referred to as a geometry-based encoding method. The encoding method using a video codec may also be referred to as a video-based encoding method.
[0132] The system multiplexer 114 multiplexes the encoded data and control information input from the 3D data encoder 113 to generate multiplexed data using a specified multiplexing method. The system multiplexer 114 may multiplex other media such as video, audio, subtitles, application data, or document files, or reference time information, along with the encoded data and control information of the 3D data. Furthermore, the system multiplexer 114 may multiplex attribute information related to the sensor data or the 3D data.
[0133] For example, the multiplexed data may have a file format for storage or a packet format for transmission. As these formats, ISOBMFF or a format based on ISOBMFF may be used. Also, MPEG-DASH, MMT, MPEG-2 TS Systems, RTP, or the like may be used.
[0134] The multiplexed data is then output as a transmission signal to the external connector 310 by the input / output processor 112. The multiplexed data may be transmitted as a transmission signal by wire or wirelessly. Alternatively, the multiplexed data is stored in an internal memory or a storage device. The multiplexed data may be transmitted to a cloud server via the Internet or may be stored in an external storage device.
[0135] For example, the transmission or storage of the multiplexed data is performed by a method according to the medium for transmission or storage, such as broadcasting or communication. The communication protocol may be http, ftp, TCP, UDP, IP, or a combination thereof. Furthermore, a pull-type communication method or a push-type communication method may be used.
[0136] For wired transmission, Ethernet (registered trademark), USB, RS-232C, HDMI (registered trademark), coaxial cable, etc. may be used. For wireless transmission, 3GPP (registered trademark), 3G / 4G / 5G defined by IEEE, wireless LAN, Wi-Fi, Bluetooth, or millimeter wave may be used. For broadcasting, for example, DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 may be used.
[0137] The sensor data may be input to the three-dimensional data generator 115 or the system multiplexer 114. The three-dimensional data or encoded data may be output as a transmission signal directly to the external connector 310 via the input / output processor 112. The transmission signal output from the three-dimensional data encoding system 110 is input to the three-dimensional data decoding system 210 via the external connector 310.
[0138] Furthermore, each operation of the three-dimensional data encoding system 110 may be controlled by a controller 111 that executes an application program.
[0139] In the three-dimensional data decoding system 210, a transmission signal is input to an input / output processor 212. The input / output processor 212 decodes multiplexed data having a file format or a packet format from the transmission signal and inputs the multiplexed data to a system demultiplexer 214. The system demultiplexer 214 obtains coded data and control information from the multiplexed data and inputs them to a three-dimensional data decoder 213. The system demultiplexer 214 may extract other media or reference time information from the multiplexed data.
[0140] The three-dimensional data decoder 213 corresponds to the decoding device 200 shown in Fig. 7 etc. For example, the three-dimensional data decoder 213 decodes three-dimensional data from the encoded data based on a predefined encoding method. The three-dimensional data is then presented to the user by the presenter 215.
[0141] Additionally, additional information such as sensor data may be input to the presenter 215. The presenter 215 may present three-dimensional data based on the additional information. Additionally, a user instruction may be input from a user terminal to the user interface 216. Then, the presenter 215 may present three-dimensional data based on the input instruction.
[0142] The input / output processor 212 may acquire the three-dimensional data and the encoded data from the external connector 310 .
[0143] Furthermore, each operation of the three-dimensional data decoding system 210 may be controlled by a controller 211 that executes an application program.
[0144] 13 is a conceptual diagram showing an example of the configuration of point cloud data according to this embodiment. The point cloud data is data of a group of points representing a three-dimensional object.
[0145] Specifically, a point cloud is made up of a plurality of points, and has position information indicating the three-dimensional coordinate position of each point and attribute information indicating the attribute of each point. The position information is also expressed as geometry.
[0146] The type of attribute information may be, for example, color, reflectance, etc. One point may be associated with attribute information of one type, one point may be associated with attribute information of multiple different types, or one point may be associated with attribute information having multiple values for the same type.
[0147] 14 is a conceptual diagram showing an example of a data file of point cloud data according to this embodiment. This example shows a case where there is a one-to-one correspondence between position information items and attribute information items, and shows position information and attribute information for N points that make up the point cloud data. In this example, the position information is information indicating a three-dimensional coordinate position using three axes, x, y, and z, and the attribute information is information indicating a color using RGB. A PLY file or the like can be used as a representative data file for point cloud data.
[0148] 15 is a conceptual diagram showing an example of the configuration of mesh data according to this embodiment. Mesh data is data used in CG (Computer Graphics) and the like, and is three-dimensional mesh data that shows the three-dimensional shape of an object using multiple surfaces. Each surface is also expressed as a polygon, and has a polygonal shape such as a triangle or a rectangle.
[0149] Specifically, a 3D mesh is composed of a plurality of points constituting a point cloud, as well as a plurality of edges and a plurality of faces. Each point is also expressed as a vertex or a position. Each edge corresponds to a line segment connected by two vertices. Each face corresponds to an area surrounded by three or more edges.
[0150] Furthermore, a three-dimensional mesh has position information indicating the three-dimensional coordinate positions of vertices. The position information is also expressed as vertex information or geometry. A three-dimensional mesh also has connection information indicating the relationship between multiple vertices that make up an edge or a face. The connection information is also expressed as connectivity. A three-dimensional mesh also has attribute information indicating the attributes of the vertices, edges, or faces. The attribute information in a three-dimensional mesh is also expressed as texture.
[0151] For example, the attribute information may indicate the color, reflectance, or normal vector for a vertex, edge, or face. The direction of the normal vector may represent the front and back of the face.
[0152] The mesh data may be stored in a data file format such as an object file.
[0153] 16 is a conceptual diagram showing an example of a data file of mesh data according to this embodiment. In this example, the data file includes position information G(1) to G(N) of N vertices that make up the three-dimensional mesh, and attribute information A1(1) to A1(N) of the N vertices. Also, in this example, M pieces of attribute information A2(1) to A2(M) are included. The attribute information items do not need to correspond one-to-one to vertices or faces. Furthermore, attribute information need not exist.
[0154] The connection information is represented by a combination of vertex indices. n[1, 3, 4] indicates a triangular face formed by three vertices, n=1, n=3, and n=4. Also, m[2, 4, 6] indicates that the attribute information of m=2, m=4, and m=6 corresponds to the three vertices, respectively.
[0155] Furthermore, the actual contents of the attribute information may be written in a separate file. A pointer to that content may be associated with a vertex, a face, or the like. For example, attribute information indicating an image for a face may be stored in a two-dimensional attribute map file. The file name of the attribute map and two-dimensional coordinate values in the attribute map may be written in attribute information A2(1) to A2(M). The method of specifying attribute information for a face is not limited to these methods, and any method may be used.
[0156] 17 is a conceptual diagram showing types of three-dimensional data according to this embodiment. Point cloud data and mesh data may represent static objects or dynamic objects. A static object is an object that does not change over time, and a dynamic object is an object that changes over time. A static object may correspond to three-dimensional data for any point in time.
[0157] For example, point cloud data for a given point in time may be referred to as a PCC frame, mesh data for a given point in time may be referred to as a mesh frame, and PCC frames and mesh frames may be simply referred to as frames.
[0158] The area of the object may be limited to a certain range, as in normal video data, or may not be limited, as in map data. The density of points or surfaces may be determined in various ways. Sparse point cloud data or sparse mesh data may be used, or dense point cloud data or dense mesh data may be used.
[0159] Next, encoding and decoding of a point cloud or a three-dimensional mesh will be described. The device, process, or syntax for encoding and decoding vertex information of a three-dimensional mesh in the present disclosure may be applied to encoding and decoding of a point cloud. The device, process, or syntax for encoding and decoding of a point cloud in the present disclosure may be applied to encoding and decoding vertex information of a three-dimensional mesh.
[0160] Furthermore, a device, process, or syntax for encoding and decoding attribute information of a point cloud in the present disclosure may be applied to encoding and decoding connectivity information or attribute information of a three-dimensional mesh.Furthermore, a device, process, or syntax for encoding and decoding connectivity information or attribute information of a three-dimensional mesh in the present disclosure may be applied to encoding and decoding attribute information of a point cloud.
[0161] Furthermore, at least some of the processing may be shared between the encoding and decoding of point cloud data and the encoding and decoding of mesh data, thereby reducing the scale of the circuit and software program.
[0162] 18 is a block diagram showing an example configuration of a three-dimensional data encoder 113 according to this embodiment. In this example, the three-dimensional data encoder 113 includes a vertex information encoder 121, an attribute information encoder 122, a metadata encoder 123, and a multiplexer 124. The vertex information encoder 121, the attribute information encoder 122, and the multiplexer 124 may correspond to the vertex information encoder 101, the attribute information encoder 103, the post-processor 105, etc. in FIG.
[0163] In this example, the three-dimensional data encoder 113 encodes the three-dimensional data according to a geometry-based encoding method, which takes into account the three-dimensional structure. In addition, in the geometry-based encoding method, attribute information is encoded using configuration information obtained in encoding the vertex information.
[0164] Specifically, first, vertex information, attribute information, and metadata included in three-dimensional data generated from sensor data are input to a vertex information encoder 121, an attribute information encoder 122, and a metadata encoder 123, respectively. Here, connectivity information included in the three-dimensional data may be treated in the same way as attribute information. In addition, in the case of point cloud data, position information may be treated as vertex information.
[0165] The vertex information encoder 121 encodes the vertex information into compressed vertex information and outputs the compressed vertex information as encoded data to the multiplexer 124. The vertex information encoder 121 also generates metadata for the compressed vertex information and outputs it to the multiplexer 124. The vertex information encoder 121 also generates configuration information and outputs it to the attribute information encoder 122.
[0166] The attribute information encoder 122 uses the configuration information generated by the vertex information encoder 121 to encode the attribute information into compressed attribute information and outputs the compressed attribute information as encoded data to the multiplexer 124. The attribute information encoder 122 also generates metadata of the compressed attribute information and outputs it to the multiplexer 124.
[0167] The metadata encoder 123 encodes compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The metadata encoded by the metadata encoder 123 may be used to encode vertex information and attribute information.
[0168] The multiplexer 124 multiplexes the compressed vertex information, the compressed vertex information metadata, the compressed attribute information, the compressed attribute information metadata, and the compressed metadata into a bitstream, and then inputs the bitstream to the system layer.
[0169] 19 is a block diagram showing an example configuration of a three-dimensional data decoder 213 according to this embodiment. In this example, the three-dimensional data decoder 213 includes a vertex information decoder 221, an attribute information decoder 222, a metadata decoder 223, and a demultiplexer 224. The vertex information decoder 221, the attribute information decoder 222, and the demultiplexer 224 may correspond to the vertex information decoder 201, the attribute information decoder 203, the preprocessor 204, and the like in FIG.
[0170] In this example, the three-dimensional data decoder 213 decodes three-dimensional data according to a geometry-based encoding method. The three-dimensional structure is taken into consideration in the decoding according to the geometry-based encoding method. Furthermore, in the decoding according to the geometry-based encoding method, attribute information is decoded using configuration information obtained in decoding vertex information.
[0171] Specifically, first, a bitstream is input from the system layer to a demultiplexer 224. The demultiplexer 224 separates compressed vertex information, compressed vertex information metadata, compressed attribute information, compressed attribute information metadata, and compressed metadata from the bitstream. The compressed vertex information and compressed vertex information metadata are input to a vertex information decoder 221. The compressed attribute information and compressed attribute information metadata are input to an attribute information decoder 222. The metadata is input to a metadata decoder 223.
[0172] The vertex information decoder 221 decodes vertex information from the compressed vertex information using metadata of the compressed vertex information. The vertex information decoder 221 also generates configuration information and outputs it to the attribute information decoder 222. The attribute information decoder 222 decodes attribute information from the compressed attribute information using the configuration information generated by the vertex information decoder 221 and the metadata of the compressed attribute information. The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 may be used to decode the vertex information and the attribute information.
[0173] Thereafter, the vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. Note that, for example, this metadata is metadata of the vertex information and attribute information, and can be used in an application program.
[0174] 20 is a block diagram showing another example configuration of the three-dimensional data encoder 113 according to the present embodiment. In this example, the three-dimensional data encoder 113 includes a vertex image generator 131, an attribute image generator 132, a metadata generator 133, a video encoder 134, a metadata encoder 123, and a multiplexer 124. The vertex image generator 131, the attribute image generator 132, and the video encoder 134 may correspond to the vertex information encoder 101 and the attribute information encoder 103 in FIG. 6 , etc.
[0175] In this example, the 3D data encoder 113 encodes the 3D data according to a video-based encoding method. In encoding according to the video-based encoding method, multiple 2D images are generated from the 3D data, and the multiple 2D images are encoded according to a video encoding method. Here, the video encoding method may be High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or the like.
[0176] Specifically, first, vertex information and attribute information included in three-dimensional data generated from sensor data are input to a metadata generator 133. The vertex information and attribute information are then input to a vertex image generator 131 and an attribute image generator 132, respectively. The metadata included in the three-dimensional data is then input to a metadata encoder 123. Here, connectivity information included in the three-dimensional data may be treated in the same way as attribute information. In the case of point cloud data, position information may be treated as vertex information.
[0177] The metadata generator 133 generates map information of a plurality of two-dimensional images from the vertex information and attribute information, and inputs the map information to the vertex image generator 131, the attribute image generator 132, and the metadata encoder 123.
[0178] The vertex image generator 131 generates a vertex image based on the vertex information and map information, and inputs the generated image to the video encoder 134. The attribute image generator 132 generates an attribute image based on the attribute information and map information, and inputs the generated image to the video encoder 134.
[0179] The video encoder 134 encodes the vertex images and attribute images into compressed vertex information and compressed attribute information, respectively, in accordance with a video encoding method, and outputs the compressed vertex information and compressed attribute information as encoded data to the multiplexer 124. The video encoder 134 also generates metadata for the compressed vertex information and metadata for the compressed attribute information, and outputs them to the multiplexer 124.
[0180] The metadata encoder 123 encodes the compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The compressible metadata includes map information. The metadata encoded by the metadata encoder 123 may also be used to encode vertex information and attribute information.
[0181] The multiplexer 124 multiplexes the compressed vertex information, the compressed vertex information metadata, the compressed attribute information, the compressed attribute information metadata, and the compressed metadata into a bitstream, and then inputs the bitstream to the system layer.
[0182] 21 is a block diagram showing another example configuration of the 3D data decoder 213 according to this embodiment. In this example, the 3D data decoder 213 includes a vertex information generator 231, an attribute information generator 232, a video decoder 234, a metadata decoder 223, and a demultiplexer 224. The vertex information generator 231, the attribute information generator 232, and the video decoder 234 may correspond to the vertex information decoder 201 and the attribute information decoder 203 in FIG. 8, etc.
[0183] In this example, the 3D data decoder 213 decodes the 3D data according to a video-based coding method. In the decoding according to the video-based coding method, a plurality of 2D images are decoded according to a video coding method, and 3D data is generated from the plurality of 2D images. Here, the video coding method may be High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or the like.
[0184] Specifically, first, a bitstream is input from the system layer to the demultiplexer 224. The demultiplexer 224 separates compressed vertex information, compressed vertex information metadata, compressed attribute information, compressed attribute information metadata, and compressed metadata from the bitstream. The compressed vertex information, compressed vertex information metadata, compressed attribute information, and compressed attribute information metadata are input to the video decoder 234. The compressed metadata is input to the metadata decoder 223.
[0185] The video decoder 234 decodes the vertex images in accordance with the video encoding method. At this time, the video decoder 234 decodes the vertex images from the compressed vertex information using the metadata of the compressed vertex information. Then, the video decoder 234 inputs the vertex images to the vertex information generator 231. The video decoder 234 also decodes the attribute images in accordance with the video encoding method. At this time, the video decoder 234 decodes the attribute images from the compressed attribute information using the metadata of the compressed attribute information. Then, the video decoder 234 inputs the attribute images to the attribute information generator 232.
[0186] The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 includes map information used to generate vertex information and attribute information. The metadata decoded by the metadata decoder 223 may also be used to decode vertex images and attribute images.
[0187] The vertex information generator 231 reproduces vertex information from the vertex image in accordance with the map information included in the metadata decoded by the metadata decoder 223. The attribute information generator 232 reproduces attribute information from the attribute image in accordance with the map information included in the metadata decoded by the metadata decoder 223.
[0188] Thereafter, the vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. Note that, for example, this metadata is metadata of the vertex information and attribute information, and can be used in an application program.
[0189] Fig. 22 is a conceptual diagram showing a specific example of encoding processing according to this embodiment. Fig. 22 shows a three-dimensional data encoder 113 and a description encoder 148. In this example, the three-dimensional data encoder 113 includes a two-dimensional data encoder 141 and a mesh data encoder 142. The two-dimensional data encoder 141 includes a texture encoder 143. The mesh data encoder 142 includes a vertex information encoder 144 and a connection information encoder 145.
[0190] The vertex information encoder 144, the connection information encoder 145, and the texture encoder 143 may correspond to the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 in FIG.
[0191] For example, the two-dimensional data encoder 141 operates as a texture encoder 143 and generates a texture file by encoding the texture corresponding to the attribute information as two-dimensional data according to an image encoding method or a video encoding method.
[0192] The mesh data encoder 142 also operates as a vertex information encoder 144 and a connectivity information encoder 145, and generates a mesh file by encoding the vertex information and connectivity information. The mesh data encoder 142 may further encode mapping information for textures. The encoded mapping information may then be included in the mesh file.
[0193] The description encoder 148 also generates a description file by encoding a description corresponding to metadata such as text data. The description encoder 148 may encode the description at the system layer. For example, the description encoder 148 may be included in the system multiplexer 114 of FIG. 12 .
[0194] The above operations generate a bitstream containing texture files, mesh files, and description files, which may be multiplexed into the bitstream in file formats such as glTF (Graphics Language Transmission Format) or USD (Universal Scene Description).
[0195] The three-dimensional data encoder 113 may include two mesh data encoders as the mesh data encoder 142. For example, one mesh data encoder encodes vertex information and connectivity information of a static three-dimensional mesh, and the other mesh data encoder encodes vertex information and connectivity information of a dynamic three-dimensional mesh.
[0196] Correspondingly, two mesh files may then be included in the bitstream, for example one mesh file corresponding to a static 3D mesh and another mesh file corresponding to a dynamic 3D mesh.
[0197] Furthermore, the static three-dimensional mesh may be a three-dimensional mesh of an intraframe coded using intraprediction, and the dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe coded using interprediction. Furthermore, information on the dynamic three-dimensional mesh may be differential information between vertex information or connectivity information of the three-dimensional mesh of an intraframe and vertex information or connectivity information of the three-dimensional mesh of an interframe.
[0198] Fig. 23 is a conceptual diagram showing a specific example of the decoding process according to this embodiment. Fig. 23 shows a three-dimensional data decoder 213, a description decoder 248, and a renderer 247. In this example, the three-dimensional data decoder 213 includes a two-dimensional data decoder 241, a mesh data decoder 242, and a mesh reconstructor 246. The two-dimensional data decoder 241 includes a texture decoder 243. The mesh data decoder 242 includes a vertex information decoder 244 and a connectivity information decoder 245.
[0199] The vertex information decoder 244, the connection information decoder 245, the texture decoder 243, and the mesh reconstructor 246 may correspond to the vertex information decoder 201, the connection information decoder 202, the attribute information decoder 203, and the post-processor 205 in Fig. 8. The presenter 247 may correspond to the presenter 215 in Fig. 12.
[0200] For example, the two-dimensional data decoder 241 operates as a texture decoder 243, and decodes the texture corresponding to the attribute information from the texture file as two-dimensional data in accordance with an image coding method or a video coding method.
[0201] The mesh data decoder 242 also operates as a vertex information decoder 244 and a connectivity information decoder 245 to decode vertex information and connectivity information from the mesh file. The mesh data decoder 242 may further decode mapping information for textures from the mesh file.
[0202] The description decoder 248 also decodes descriptions corresponding to metadata such as text data from the description file. The description decoder 248 may decode the descriptions at the system layer. For example, the description decoder 248 may be included in the system demultiplexer 214 of FIG. 12 .
[0203] The mesh reconstructor 246 reconstructs a 3D mesh from the vertex information, connectivity information, and textures according to the description. The renderer 247 renders and outputs the 3D mesh according to the description.
[0204] Through the above operations, a 3D mesh is reconstructed and output from a bitstream containing a texture file, a mesh file, and a description file.
[0205] The three-dimensional data decoder 213 may include two mesh data decoders as the mesh data decoder 242. For example, one mesh data decoder decodes vertex information and connectivity information of a static three-dimensional mesh, and the other mesh data decoder decodes vertex information and connectivity information of a dynamic three-dimensional mesh.
[0206] Correspondingly, two mesh files may then be included in the bitstream, for example one mesh file corresponding to a static 3D mesh and another mesh file corresponding to a dynamic 3D mesh.
[0207] Furthermore, the static three-dimensional mesh may be a three-dimensional mesh of an intraframe coded using intraprediction, and the dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe coded using interprediction. Furthermore, information on the dynamic three-dimensional mesh may be differential information between vertex information or connectivity information of the three-dimensional mesh of an intraframe and vertex information or connectivity information of the three-dimensional mesh of an interframe.
[0208] A dynamic 3D mesh coding method is sometimes called DMC (Dynamic Mesh Coding), and a video-based dynamic 3D mesh coding method is sometimes called V-DMC (Video-based Dynamic Mesh Coding).
[0209] The point cloud encoding method is sometimes called PCC (Point Cloud Compression). The point cloud video-based encoding method is sometimes called V-PCC (Video-based Point Cloud Compression). The point cloud geometry-based encoding method is sometimes called G-PCC (Geometry-based Point Cloud Compression).
[0210] <Implementation Example> Fig. 24 is a block diagram showing an implementation example of the encoding device 100 according to this embodiment. The encoding device 100 includes a circuit 151 and a memory 152. For example, multiple components of the encoding device 100 shown in Fig. 5 etc. are implemented by the circuit 151 and memory 152 shown in Fig. 24.
[0211] The circuit 151 is a circuit that performs information processing and is a circuit that can access the memory 152. For example, the circuit 151 is a dedicated or general-purpose electric circuit that encodes a three-dimensional mesh. The circuit 151 may be a processor such as a CPU. Alternatively, the circuit 151 may be a collection of multiple electric circuits.
[0212] The memory 152 is a dedicated or general-purpose memory that stores information used by the circuit 151 to encode the three-dimensional mesh. The memory 152 may be an electric circuit and may be connected to the circuit 151. The memory 152 may also be included in the circuit 151. The memory 152 may also be a collection of multiple electric circuits. The memory 152 may also be a magnetic disk, an optical disk, or the like, and may also be expressed as a storage, a recording medium, or the like. The memory 152 may also be a non-volatile memory or a volatile memory.
[0213] For example, the memory 152 may store a three-dimensional mesh or a bitstream, or may store a program for the circuit 151 to encode the three-dimensional mesh.
[0214] Note that the encoding device 100 does not necessarily have to implement all of the components shown in Figure 5 and the like, and does not necessarily have to perform all of the processes shown here. Some of the components shown in Figure 5 and the like may be included in another device, and some of the processes shown here may be executed by another device. Furthermore, the encoding device 100 may implement any combination of the components of the present disclosure, and may perform any combination of the processes of the present disclosure.
[0215] Fig. 25 is a block diagram showing an example implementation of a decoding device 200 according to this embodiment. The decoding device 200 includes a circuit 251 and a memory 252. For example, multiple components of the decoding device 200 shown in Fig. 7 and other figures are implemented by the circuit 251 and memory 252 shown in Fig. 25.
[0216] The circuit 251 is a circuit that performs information processing and is a circuit that can access the memory 252. For example, the circuit 251 is a dedicated or general-purpose electric circuit that decodes a three-dimensional mesh. The circuit 251 may be a processor such as a CPU. Alternatively, the circuit 251 may be a collection of multiple electric circuits.
[0217] The memory 252 is a dedicated or general-purpose memory that stores information for the circuit 251 to decode the 3D mesh. The memory 252 may be an electric circuit and may be connected to the circuit 251. The memory 252 may also be included in the circuit 251. The memory 252 may also be a collection of multiple electric circuits. The memory 252 may also be a magnetic disk, an optical disk, or the like, and may also be expressed as a storage, a recording medium, or the like. The memory 252 may also be a non-volatile memory or a volatile memory.
[0218] For example, the memory 252 may store a three-dimensional mesh or a bitstream, or may store a program for the circuit 251 to decode the three-dimensional mesh.
[0219] Note that the decoding device 200 does not necessarily have to implement all of the components shown in Figure 7 and the like, and does not necessarily have to perform all of the processes shown here. Some of the components shown in Figure 7 and the like may be included in another device, and some of the processes shown here may be executed by another device. Furthermore, the decoding device 200 may implement any combination of the components of the present disclosure, and may perform any combination of the processes of the present disclosure.
[0220] The encoding method and the decoding method including the steps performed by each component of the encoding device 100 and the decoding device 200 of the present disclosure may be executed by any device or system. For example, part or all of the encoding method and the decoding method may be executed by a computer including a processor, a memory, an input / output circuit, etc. In this case, the encoding method and the decoding method may be executed by the computer executing a program for causing the computer to execute the encoding method and the decoding method.
[0221] Alternatively, the program or the bitstream may be recorded on a non-transitory computer-readable recording medium such as a CD-ROM.
[0222] An example of a program may be a bitstream. For example, a bitstream including an encoded three-dimensional mesh includes syntax elements for causing the decoding device 200 to decode the three-dimensional mesh. The bitstream then causes the decoding device 200 to decode the three-dimensional mesh according to the syntax elements included in the bitstream. Thus, the bitstream may play a role similar to that of a program.
[0223] The bitstream may be an encoded bitstream containing the encoded 3D mesh, or may be a multiplexed bitstream containing the encoded 3D mesh and other information.
[0224] Furthermore, each component of the encoding device 100 and the decoding device 200 may be configured with dedicated hardware, general-purpose hardware that executes the above-mentioned programs, or a combination of these. The general-purpose hardware may be configured with a memory in which the programs are recorded and a general-purpose processor that reads and executes the programs from the memory. Here, the memory may be a semiconductor memory or a hard disk, and the general-purpose processor may be a CPU.
[0225] Furthermore, the dedicated hardware may be configured with a memory, a dedicated processor, etc. For example, the dedicated processor may execute the encoding method and the decoding method by referring to a memory for recording data.
[0226] Furthermore, as described above, each component of the encoding device 100 and the decoding device 200 may be an electric circuit. These electric circuits may form a single electric circuit as a whole, or each may be a separate electric circuit. Furthermore, these electric circuits may correspond to dedicated hardware, or may correspond to general-purpose hardware that executes the above-mentioned programs, etc. Furthermore, the encoding device 100 and the decoding device 200 may be implemented as an integrated circuit.
[0227] Furthermore, the encoding device 100 may be a transmitting device that transmits the three-dimensional mesh, and the decoding device 200 may be a receiving device that receives the three-dimensional mesh.
[0228] Displacement Encoding and Decoding The following terminology is used here by way of example:
[0229] (1) Image An image is a data unit made up of a set of pixels, and includes a picture or a block smaller than a picture. Images include both moving images and still images.
[0230] (2) Picture A picture is a unit of image processing that is made up of a set of pixels, and is also called a frame or field.
[0231] (3) Block A block is a processing unit consisting of a specific number of pixels. The term "block" shown in the following example is also used. The shape of a block is not particularly limited. A block may be, for example, a rectangular shape of M x N pixels or a square shape of M x M pixels. A block may also be a triangular shape, a circular shape, or another shape. Examples of blocks are as follows:
[0232] Slice, tile, or brick CTU, superblock, or basic division unit VPDU, processing division unit for hardware CU, processing block unit, prediction block unit (PU), or orthogonal transform block unit (TU) Sub-block
[0233] (4) Pixel or Sample A pixel or sample is the smallest point of an image, in other words, the smallest unit. Pixels or samples include not only pixels at integer positions, but also pixels at sub-pixel positions generated based on pixels at integer positions.
[0234] (5) Pixel Value or Sample Value: A pixel value or sample value is a unique value of a pixel. The pixel value or sample value may include a luma value, a chroma value, or an RGB gradation level, and may also include a depth value or a binary value of 0 or 1.
[0235] (6) Flags A flag indicates one or more bits. A flag is, for example, a parameter or index represented by two or more bits. A flag may indicate not only a value represented by a binary number, but also a value represented by a number other than a binary number.
[0236] (7) Signal: A signal is something that is symbolized or coded to transmit information. A signal includes a discrete digital signal or a continuous analog signal.
[0237] (8) Stream or Bit Stream A stream or bit stream is a digital data sequence that indicates the flow of digital data. A stream or bit stream may be a single stream, or may be configured to include multiple streams with multiple layers. A stream or bit stream may be transmitted by serial communication using a single transmission path, or may be transmitted by packet communication using multiple transmission paths.
[0238] (9) Difference: For scalar quantities, difference can include simple difference (x - y) and difference calculations, such as absolute difference (|x - y|), squared difference (x^2 - y^2), square root difference (√(x - y)), weighted difference (ax - b, where a and b are constants), or offset difference (x - y + a, where a is an offset).
[0239] (10) Sum. For scalar quantities, sums can include simple sum (x + y) and addition operations. Sum can also include absolute sum (|x + y|), sum of squares (x^2 + y^2), square root of sum (√(x + y)), weighted sum (ax + by, where a and b are constants), or offset sum (x + y + a, where a is an offset).
[0240] (11) "Based on" The expression "based on something" means that something other than that "something" may be taken into consideration. Also, "based on" can be used both when a direct result is obtained and when a result is obtained through an intermediate result.
[0241] (12) "Used" or "Using" The phrases "something was used" or "used something" mean that something other than the "something" may be taken into consideration. The phrases "used" or "used" may be used both in cases where a direct result is obtained and in cases where a result is obtained via an intermediate result.
[0242] (13) Prohibition "Prohibit" can be rephrased as "not permitted." Also, "not prohibited / prohibited" or "permitted / permitted" does not necessarily mean "obligation."
[0243] (14) "Restriction" or "Limitation" "Restriction" or "Limitation" can be rephrased as "not permitted / not allowed" or "not permitted / permitted." Furthermore, "prohibited / not prohibited" or "not permitted / permitted" does not necessarily mean "obligation." Furthermore, what is prohibited quantitatively or qualitatively may be either partial or total.
[0244] (15) Chroma The term chroma is an adjective, represented by the symbols Cb or Cr, that indicates that a sample array or a single sample represents one of the two color-difference signals associated with a primary color. The term chroma is sometimes used instead of the term chrominance.
[0245] (16) Luma The term luma is an adjective, represented by the symbols or subscripts Y or L, that indicates that a sample array or a single sample represents a monochrome signal for a primary color. The term luma is sometimes used instead of the term luminance.
[0246] The encoding / decoding system of this embodiment will be described below.
[0247] A typical three-dimensional model (also called a 3D model) digitally represents an object so that a user can explore the model using zoom, pan, and rotation in all three dimensions while it is rendered over time. One way to construct such a representation is to build a 3D mesh using triangles. The model stores the positions of the triangle vertices, their connectivity to each other, and their associated attributes (such as normals or UV patches).
[0248] Storing all this information in uncompressed form requires a very large storage space and therefore a very large bandwidth for transmission. The triangles that form the mesh often have repeating patterns and similar properties, especially in temporal and spatial neighborhoods. These repetitions can be exploited to develop efficient encoding and decoding methods for storage and transmission. One such encoding and decoding method is Video-based Dynamic Mesh Coding (V-DMC).
[0249] 26 is a block diagram showing another example of the configuration of the encoding / decoding system according to this embodiment. As shown in FIG. 26, the encoding / decoding system includes an encoding device 100 and a decoding device 200.
[0250] The encoding / decoding system accepts input three-dimensional meshes (also called 3D meshes) in the form of three-dimensional coordinates of vertices (vertex information), connectivity (connection information) and associated attributes (attribute information), which may include texture maps as well as geometry.
[0251] The encoding device 100 takes an input 3D mesh (also referred to as an input 3D mesh or input mesh) in the form of three-dimensional coordinates of vertices, connectivity, and associated attributes. The encoding device 100 encodes all associated information into a stream. The stream may consist of a single bitstream or multiple bitstreams.
[0252] The network 300 transmits the stream generated by the encoding device to the decoding device 200. The network 300 may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof. Furthermore, the network 300 is not necessarily limited to a two-way communication network, but may also be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Instead of the network 300, a recording medium such as a digital versatile disc (DVD) or a blue-ray disc (BD) on which a stream is recorded may be used.
[0253] The stream is transmitted to a decoding device 200 via a network 300. The decoding device 200 decodes the bitstream and generates a 3D mesh using the 3D coordinates, connectivity, and associated attributes of the decoded vertices. The decoding device 200 outputs the generated 3D mesh (also referred to as an output 3D mesh or output mesh).
[0254] FIG. 27 is a diagram showing another example of the configuration of the encoding device 100.
[0255] As shown in FIG. 27, the encoding device 100 includes a preprocessor 1103 and a compressor 1106 .
[0256] The encoding device 100 reads an input mesh 1101 and an attribute map 1102 and passes them to a preprocessor 1103. The preprocessor 1103 processes the input mesh to extract a base mesh 1104 and displacement data 1105. The attribute map 1102, along with the extracted base mesh 1104 and displacement data 1105, are passed to a compressor 1106.
[0257] The compressor 1106 also compresses the base mesh 1104, the displacement data 1105, and the attribute map 1102 to generate a bitstream 1107. The compressor 1106 can transmit additional information to the decoding device 200 by further including metadata 1108 in the bitstream 1107.
[0258] FIG. 28 is a diagram showing another example of the configuration of the decoding device 200.
[0259] As shown in FIG. 28, the decoding device 200 includes a decompressor 2102 and a post-processor 2106 .
[0260] The decoding device 200 reads a bitstream 2101 and passes it to a decompressor 2102. The decompressor 2102 decompresses a base mesh 2103, displacement data 2104, and an attribute map 2108 from the bitstream 2101 and passes them to a post-processor 2106. An example of the displacement data 2104 is a displacement vector.
[0261] The post-processor 2106 also processes the base mesh 2103 according to the displacement data 2104 and the attribute map 2108 to generate an output mesh 2107. The post-processor 2106 may further use information from the metadata 2105 to generate the output mesh 2107.
[0262] FIG. 29 is a block diagram showing yet another example configuration of the encoding device 100 according to this embodiment.
[0263] In this example, the encoding device 100 comprises a volumetric capturer 511, a projector 512, a base mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and optionally one or more other type encoders 516.
[0264] The volumetric capturer 511 captures content and outputs the captured content to the projector 512 .
[0265] The projector 512 projects the content onto a 3D mesh frame containing vertex geometry coordinates, texture coordinates, and connectivity data. The data is output to a base mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and optionally one or more other type encoders 516. Each encoder compresses the data into a bitstream.
[0266] FIG. 30 is a block diagram showing yet another example configuration of the decoding device 200 according to this embodiment.
[0267] In this example, the decoding device 200 comprises a base mesh decoder 613 , a displacement decoder 614 , an attribute decoder 615 , one or more other type decoders 616 , and a 3D reconstructor 617 .
[0268] The bitstream is sent to a base mesh decoder 613, a displacement decoder 614, an attribute decoder 615, and optionally one or more other type decoders 616. These decoders decode the bitstream to generate decoded data including vertex geometry coordinates, texture coordinates, and connectivity data. The decoded data is then sent to a 3D reconstructor 617, which reconstructs a 3D mesh frame.
[0269] The encoding process performed by the encoding device 100 will be described in detail below.
[0270] Fig. 31 is a flow diagram showing the processing of the encoding device 100. Fig. 32 is an explanatory diagram conceptually showing the encoding of mesh frames. The processing of the encoding device 100 will be described with reference to Figs. 31 and 32.
[0271] In step S101, the encoding device 100 reads a 3D mesh frame, which is an input mesh frame, and its attributes. The input mesh frame is a mesh frame input to the encoding device 100. An example of the 3D mesh frame that is an input mesh frame is shown as mesh frame 1301 (see FIG. 32 ).
[0272] In step S102, the encoding device 100 performs a decimation process on the input mesh frame read in step S101 to generate a base mesh frame having fewer vertices than the input mesh frame. The base mesh frame generated by decimating the mesh frame 1301 is shown as a base mesh frame 1302 (see FIG. 32).
[0273] In step S103, the encoding device 100 calculates displacement information that the decoding device 200 uses to reconstruct a mesh frame. The displacement information corresponds to a displacement vector directed from a vertex of the base mesh frame generated in step S102 to a vertex of the input mesh frame. One method for calculating the displacement information is to subtract the coordinates of the vertex of the base mesh frame from the coordinates of the vertex of the input mesh frame. The displacement information calculated from the mesh frame 1301 and the base mesh frame 1302 is shown as displacement information 1303 (see FIG. 32). The displacement information 1303 is in vector format, in other words, expressed as a displacement vector.
[0274] In step S104, the encoding device 100 encodes the base mesh frame generated in step S102, the displacement information generated in step S103, and the attributes of the input mesh frame into a bitstream (corresponding to a compressed bitstream). An example of the bitstream is shown as bitstream 1304 (see FIG. 32).
[0275] Specifically, the bitstream 1304 includes vertex coordinates and connectivity information for vertices A, C, E, and F, displacement information, a video bitstream including texture data, and a compressed attribute map (see FIG. 32). The displacement information includes displacement information for displacing vertices based on vertex coordinates obtained from the subdivided base mesh frame. The compressed attribute map is texture coordinates for applying texture data to a mesh frame reconstructed using the base mesh frame and the displacement information.
[0276] The decoding process performed by the decoding device 200 will be described in detail below.
[0277] Fig. 33 is a flow diagram showing the processing of the decoding device 200. Fig. 34 is an explanatory diagram conceptually showing the decoding of a 3D mesh. The processing of the decoding device 200 will be described with reference to Figs. 33 and 34.
[0278] In step S201, the decoding device 200 decodes a base mesh frame and attributes from a bitstream (corresponding to a compressed bitstream). An example of the decoded base mesh frame (corresponding to a decoded base mesh frame) is shown as a decoded base mesh frame 2301 (see FIG. 34).
[0279] In step S202, the decoding device 200 generates subdivided vertices by performing a subdivision process on the base mesh frame decoded in step S201. An example of a base mesh frame including subdivided vertices is shown as base mesh frame 2302 (see FIG. 34).
[0280] In step S203, the decoding device 200 decodes the disparity information from the bitstream (corresponding to the compressed bitstream). An example of the decoded disparity information is shown as disparity information 2303 (see FIG. 34). The disparity information 2303 is in vector format, in other words, expressed as a disparity vector.
[0281] In step S204, the decoding device 200 reconstructs the shape of the mesh frame by moving the vertices of the base mesh frame, including the subdivided vertices, to new positions using the displacement information, and then restores the mesh frame by applying attribute information. An example of the attribute is texture. An example of the reconstructed mesh frame is shown as mesh frame 2304 (see FIG. 34 ).
[0282] FIG. 35 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.
[0283] FIG. 35 shows an example of a block diagram of a general intra-decoding system.
[0284] The decoding device shown in FIG. 35 comprises a demultiplexer 1231, a switch 1232, a static mesh decoder 1233, a mesh buffer 1234, a motion decoder 1235, a base mesh reconstructor 1236, an inverse quantizer 1237, a video decoder 1238, an image unpacker 1239, an inverse quantizer 1240, an inverse wavelet transformer 1241, a reconstructor 1242, a video decoder 1243, and a color converter 1244.
[0285] The demultiplexer 1231 receives the compressed bitstream and separates it into compressed data for the base mesh, video containing displacement data (also called displacement bitstream), and video containing attribute data (also called attribute bitstream). The compressed data for the base mesh is passed to a switch 1232. The switch 1232 determines whether to perform intra-decoding or inter-decoding based on parameters in the bitstream.
[0286] If an intra-decoding process is selected, the bitstream is passed to a static mesh decoder 1233, which generates a quantized base mesh. The static mesh decoder 1233 is, for example, a decoder that uses an edge breaker algorithm to decode 3D mesh data. The static mesh decoder 1233 generates a quantized base mesh from the bitstream. The quantized base mesh generated by the static mesh decoder 1233 is stored in a mesh buffer 1234 for reference when an inter-decoding process is selected.
[0287] If inter-decoding is selected, switch 1232 passes compressed data for the base mesh to motion decoder 1235. Motion decoder 1235 receives a previously decoded quantized base mesh and decodes motion data representing the differences in vertex coordinates between the quantized base mesh stored in mesh buffer 1234 and the current quantized base mesh. The motion data and the quantized base mesh stored in mesh buffer 1234 are used by base mesh reconstructor 1236 to reconstruct the current quantized base mesh. The quantized base mesh resulting from either inter-decoding or intra-decoding is passed to inverse quantizer 1237 to obtain a decoded base mesh.
[0288] The video containing the displacement data is passed to a video decoder 1238, since the bitstream contains the displacement data in an image format with two chroma information and one luma information. The video decoder 1238 decodes the data using a video frame decompression method. Alternatively, the displacement data can be decoded using an arithmetic decoder. This decompressed data is passed to an image unpacker 1239, which extracts wavelet coefficients associated with each vertex from the image-format decompressed data. An inverse quantizer 1240 dequantizes the quantized wavelet coefficients into the three components associated with each vertex. An inverse wavelet transformer 1241 inversely transforms the result to finally obtain decoded displacement data. The decoded displacement data and the decoded base mesh are passed to a reconstructor 1242, which performs edge refinement on the decoded base mesh and displaces the vertices using the decoded displacement data to obtain a decoded mesh.
[0289] The video including the attribute data is passed to another video decoder 1243 to obtain a decoded attribute bitstream, which is further processed in a color converter 1244 for color space and color format conversion to obtain a decoded attribute map.
[0290] FIG. 36 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.
[0291] FIG. 36 illustrates an example of a reconstructor that obtains a decoded 3D mesh 1256 from a decoded base mesh 1251 and decoded displacement data 1254.
[0292] The decoded base mesh 1251 is passed to a subdivider 1252 .
[0293] The subdivision unit 1252 subdivides any two connected vertices in the entire 3D mesh by adding a new vertex between them. This process can be repeated several times to include vertices created in previous subdivision steps to generate a predefined number of vertices. Each subdivision iteration across the 3D mesh generates a new level of detail (LoD). The subdivided mesh 1253 and the decoded displacement data 1254 are passed to a displacer 1255. The displacer 1255 generates a decoded 3D mesh 1256 by moving each vertex to a new position according to the corresponding displacement data.
[0294] The subdivision is described below and is performed by a subdivider (specifically subdivider 1206 or subdivider 2204).
[0295] FIG. 37 is an explanatory diagram showing an example of subdivision.
[0296] The base mesh shown in FIG. 37(a) includes vertices A, B, and C and connectivity information indicating their connectivity.
[0297] 37(b) shows a mesh generated by the first subdivision, in other words, the mesh after the first subdivision. In the first subdivision, the subdivider generates vertices D, E, and F and connectivity information indicating their connectivity. The mesh generated by the subdivider is also referred to as LoD1 or first LoD.
[0298] Vertex D of the mesh after the first subdivision is a vertex generated by subdivision based on vertices A and B. Similarly, vertex E is a vertex generated by subdivision based on vertices B and C. Vertex F is a vertex generated by subdivision based on vertices A and C.
[0299] As an example, vertex D may be the midpoint of line segment AB (in other words, side AB) connecting vertices A and B that were the basis for its generation. Similarly, vertex E may be the midpoint of line segment AC. Vertex F may be the midpoint of line segment BC.
[0300] 37(c) shows the mesh generated by the second subdivision, i.e., the mesh after the second subdivision. In the second subdivision, the subdivider generates vertices G, H, I, J, K, L, M, N, and O and connectivity information indicating their connectivity. The mesh generated by the subdivider is also called LoD2 or second LoD.
[0301] Vertex G of the mesh after the second subdivision is a vertex generated by subdivision based on vertices A and D. Similarly, vertex H is a vertex generated by subdivision based on vertices A and E. Vertex I is a vertex generated by subdivision based on vertices B and D. Vertex J is a vertex generated by subdivision based on vertices D and F. Vertex K is a vertex generated by subdivision based on vertices E and F. Vertex L is a vertex generated by subdivision based on vertices C and E. Vertex M is a vertex generated by subdivision based on vertices B and F. Vertex N is a vertex generated by subdivision based on vertices C and F. Vertex O is a vertex generated by subdivision based on vertices D and E.
[0302] As an example, vertex G may be the midpoint of line segment AD (in other words, side AD) connecting vertices A and D, which were the source of its generation. Similarly, vertex H may be the midpoint of line segment AE. vertex I may be the midpoint of line segment BD. vertex J may be the midpoint of line segment DF. vertex K may be the midpoint of line segment EF. vertex L may be the midpoint of line segment CE. vertex M may be the midpoint of line segment BF. vertex N may be the midpoint of line segment CF. vertex O may be the midpoint of line segment DE.
[0303] The displacement of vertices will be described below with reference to Figures 38 and 39. The displacement of vertices is performed by the reconstructor 2209.
[0304] Fig. 38 is an explanatory diagram showing an example of displacement of vertices after subdivision, and Fig. 39 is an explanatory diagram showing an example of vertices of an original mesh.
[0305] The base mesh shown in FIG. 38(a) includes vertices A, B, C, and Z and connectivity information indicating their connectivity.
[0306] 38(b) shows a mesh generated by the first subdivision, in other words, a mesh after the first subdivision (i.e., the first LoD). In the first subdivision, the subdivider generates vertices S, T, U, X, or Y and connectivity information indicating their connectivity. The vertices S, T, U, X, or Y are similar to the vertices D, E, and F shown in FIG. 37(b).
[0307] 38(c) shows a mesh generated by the second subdivision, in other words, a mesh after the second subdivision (i.e., the second LoD). In the second subdivision, the subdivider generates vertices D, E, F, G, and H and connectivity information indicating their connectivity. Vertices D, E, F, G, and H are the same as vertices G, H, I, J, K, L, M, N, and O shown in FIG. 37(c).
[0308] Figure 38(d) shows a mesh including the vertices after they have been displaced after subdivision, with vertices A, B, C, D, E, F, G, H, S, T, U, X, Y, and Z shown in Figure 38(d) being located at positions displaced using displacement information from the positions of the vertices shown in Figure 38(c).
[0309] The original mesh shown in FIG. 39 is an example of the mesh input to the encoding device 100, that is, the mesh before encoding.
[0310] The mesh shown in Fig. 38 has a shape similar to that of the original mesh shown in Fig. 39. The displacement information is generated by the displacement vector calculator 1207 of the encoding device 100 as information indicating the displacement from the vertices of the base mesh to the vertices of the original mesh, and therefore, by reconstructing the mesh using the displacement information thus generated, a mesh having a shape similar to that of the original mesh is generated.
[0311] The decoding device 200 can output the mesh shown in FIG.
[0312] Next, the division of a mesh into sub-meshes will be described with reference to FIGS.
[0313] A mesh can be divided into smaller parts and coded separately, with the mesh vertices being divided in such a way that the coordinates and connectivity of the vertices in each part can be coded independently.
[0314] Fig. 40 is an explanatory diagram showing an example of a mesh, and Fig. 41 is an explanatory diagram showing an example of dividing a mesh into sub-meshes.
[0315] The mesh shown in FIG. 40 is the original mesh, which is sometimes called a full mesh in contrast to a sub-mesh.
[0316] Figure 41 shows how the full mesh shown in Figure 40 is divided into two sub-meshes. For vertices A, B, and C of the full mesh (see Figure 40), vertex A is duplicated to vertices A1 and A2, vertex B is duplicated to vertices B1 and B2, and vertex C is duplicated to vertices C1 and C2, thereby creating two sub-meshes (i.e., a first sub-mesh and a second sub-mesh) from the full mesh. The first sub-mesh and the second sub-mesh are each independently decodable meshes.
[0317] Packing of displacement information into image frames will be described below with reference to FIGS.
[0318] 42, 43 and 44 are explanatory diagrams showing examples of packing of displacement information into image frames. Note that image frames can also be called video frames.
[0319] The vertex displacement data is encoded as image frame data by being mapped to each component of a YUV format image frame (i.e., each of the Y component (Y Plane), U component (U Plane), and V component (V Plane)). This case will be described below as an example. As another example, the vertex displacement data may be encoded as image frame data by being mapped to each component of an RGB format image frame (each of the R component, G component, and B component).
[0320] The decoding device 200 can use an image encoding module to extract the displacement data. The displacement data can be in the form of X, Y, or Z components in a global coordinate system (e.g., a Cartesian coordinate system), or normal, tangential, or both tangential components in a local coordinate system. Methods for mapping the displacement data to an image frame include the following:
[0321] For example, in the first method, the displacement data is arranged in the image frame in scan order, and an example of packing the displacement data in this case is shown in Figure 42. The displacement data is directly mapped onto the image frame according to a predefined scan order.
[0322] Note that since an image frame has a fixed height and width, it may happen that the displacement data does not fit perfectly in the frame, in which case the remaining part of the image frame is padded with padding data (see Figure 42).
[0323] For example, in the second method, the displacement data is separated into multiple LoDs and mapped to the Y, U, and V components of the image frame. An example of packing of the displacement data in this case is shown in Figure 43. Here, the displacement data of the image frame of the next LoD starts immediately after the displacement data of the previous LoD ends. As in the first method, if the displacement data does not fit exactly into the image frame, padding is performed at the end of the image frame (see Figure 43).
[0324] For example, in the third method, displacement data corresponding to the LoD is mapped to the Y component, U component, and V component of the image frame in a manner different from that in the second method. An example of packing of the displacement data in this case is shown in Figure 44. In this way, each LoD can be decoded independently. In the third method, middle padding is performed on the displacement data of each LoD, and CTU alignment is performed together with padding at the end of the video frame (see Figure 44).
[0325] Fig. 45 is a block diagram showing a detailed configuration example of the decoding device 200 according to this embodiment. Specifically, Fig. 45 shows an example of the configuration of a geometry coordinate decoder included in the decoding device 200.
[0326] In this example, the decoding device 200 comprises a frame header decoder 631 , a vertex geometry coordinate predictor 632 , a vertex geometry coordinate difference decoder 633 , and a reconstructor 634 .
[0327] The frame header decoder 631 reads the bitstream and decodes the frame headers in the bitstream to determine whether the frame data should be intra-decoded (intra-predicted) or inter-decoded (inter-predicted).
[0328] If inter-decoding is selected, the frame data contained in the bitstream is output to a vertex geometry coordinate predictor 632 .
[0329] The vertex geometry coordinate predictor 632 outputs prediction information to the reconstructor 634. An example of prediction information is a motion vector.
[0330] The reconstructor 634 uses prediction information along with vertex coordinates from previously decoded frames to output the three-dimensional coordinates of the vertices (vertex geometry coordinates).
[0331] On the other hand, if intra-decoding is selected, the frame data included in the bitstream is output to the vertex geometry coordinate difference decoder 633 .
[0332] The vertex geometry coordinate differential decoder 633 decodes frame data encoded as differences between the coordinates of the vertices contained in the frame to generate vertex coordinates. Only one of the vertex geometry coordinates from the vertex geometry coordinate differential decoder 633 and the reconstructor 634 is used to generate the decoded 3D mesh frame.
[0333] Fig. 46 is an explanatory diagram showing the coordinates of vertices in a 3D mesh according to this embodiment. Specifically, Fig. 46 shows an example in which the entire 3D mesh frame is decoded using the coordinates (positions) of the actual vertices included in the bitstream.
[0334] The coordinates of vertex A included in the 3D mesh frame at time (t) are decoded as (6, 8, 9) using the Cartesian coordinate system (x, y, z) as shown in (a) of Figure 46. Similarly, the coordinates of vertex B are decoded as (10, 6, 7), and the coordinates of vertex C are decoded as (14, 8, 9). The same is true for vertices D to G.
[0335] Fig. 47 is an explanatory diagram showing prediction information according to this embodiment. Specifically, Fig. 47 shows another example in which the entire 3D mesh frame at time (t) is decoded using a frame at time (t-1) (a past frame) and prediction information included in the bitstream.
[0336] The coordinates (6, 8, 9) of vertex A in the frame to be decoded (current frame) are decoded by adding the coordinates (4, 7, 8) of vertex A in the past frame to the value (2, 1, 1) for vertex A indicated by the prediction information. Similarly, the coordinates (10, 6, 7) of vertex B in the current frame are decoded by adding the coordinates (8, 6, 7) of vertex B in the past frame to the value (2, 0, 0) for vertex B indicated by the prediction information.
[0337] An example of the configuration of the encoding device according to this embodiment will be described below.
[0338] FIG. 48 is a block diagram showing an example of the configuration of an encoding device according to this embodiment.
[0339] The encoding device shown in Figure 48 comprises a decimator 4801, a sub-divider 4802, a displacement vector calculator 4803, a wavelet transformer 4804, an inter-predictor 4805, a quantizer 4806, an image packer 4807, a video encoder 4808, an inverse quantizer 4811, a reconstructor 4812, and a reference buffer 4813.
[0340] The decimator 4801 acquires a mesh frame (corresponding to an original 3D mesh frame, also referred to as an original mesh frame or original mesh) input to the encoding device and generates a base mesh frame (also referred to as a base mesh) by performing a decimation process (in other words, a thinning process) on the acquired mesh frame. The decimation process is a process of deleting (in other words, thinning) some of the vertices included in the original mesh. The decimation process may include a process of changing the positions of at least some of the vertices included in the original mesh, or a process of changing the connectivity of at least some of the vertices included in the original mesh. The decimation process is also simply referred to as decimation.
[0341] The base mesh generated by the decimation process is a mesh with fewer vertices than the original mesh. The vertices of the base mesh may be located at different positions than the vertices of the original mesh. Also, the vertex connectivity of the base mesh may be different from the vertex connectivity of the original mesh. The decimator 4801 provides the generated base mesh frame to the subdivider 4802.
[0342] The subdivider 4802 performs a subdivision process on the base mesh frame generated by the decimator 4801. The subdivision process may be a process of subdividing the base mesh frame into smaller pieces. The subdivider 4802 provides the subdivided base mesh frame to the displacement vector calculator 4803.
[0343] Specifically, the subdivider 4802 can subdivide the mesh frame by generating a new vertex between two connected vertices in the mesh frame. By repeating this process of generating new vertices, the number of vertices in the mesh frame can be set to a predetermined number. By repeating the subdivision throughout the mesh frame (in other words, by performing the subdivision multiple times), multiple Level of Detail (LoD) hierarchies are generated.
[0344] The displacement vector calculator 4803 receives the original mesh frame obtained by the encoding device and also receives the subdivided base mesh frame from the subdivider 4802. The displacement vector calculator 4803 calculates, as a displacement vector, a vector directed from a vertex of the base mesh frame and a vertex generated by subdividing the base mesh frame to a corresponding vertex of the original mesh frame. The displacement vector calculator 4803 provides the calculated displacement vector to the wavelet transformer 4804.
[0345] The wavelet transformer 4804 acquires transform coefficients (also referred to as wavelet coefficients) by performing wavelet transform processing on the displacement vectors calculated by the displacement vector calculator 4803. The wavelet transformer 4804 provides the acquired wavelet coefficients to the inter predictor 4805. In the wavelet transform, the wavelet transformer 4804 assigns vertices to multiple LoD layers and applies, for example, a lifting transform to the displacement vectors of the vertices, thereby being able to calculate wavelet coefficients representing various components from low frequency components to high frequency components.
[0346] The inter predictor 4805 calculates prediction residuals of wavelet coefficients of displacement vectors of a frame to be coded using inter prediction. Specifically, the inter predictor 4805 calculates prediction residuals of wavelet coefficients of displacement vectors of a frame to be coded by inter predicting the wavelet coefficients of the displacement vectors of the frame to be coded using wavelet coefficients of displacement vectors of an encoded frame (also referred to as a reference frame) stored in the reference buffer 4813.
[0347] The quantizer 4806 quantizes the prediction residual of the wavelet coefficients calculated by the inter predictor 4805. The quantizer 4806 can quantize the prediction residual of the wavelet coefficients for each LoD layer. The quantizer 4806 provides the quantized prediction residual to the image packer 4807 and the inverse quantizer 4811.
[0348] The image packer 4807 generates an image containing the prediction residual quantized by the quantizer 4806. The image packer 4807 can generate the image by mapping the prediction residual quantized by the quantizer 4806 to pixels in a two-dimensional image format. The image packer 4807 provides the generated image to the video encoder 4808. The process of mapping the quantized prediction residual to pixels in the two-dimensional image format can use mapping information that represents the assignment of the quantized prediction residual to pixels in the two-dimensional image format.
[0349] The video encoder 4808 encodes the image generated by the image packer 4807 into a bitstream (also referred to as a displacement bitstream) (in other words, generates a displacement bitstream). The video encoder 4808 outputs the displacement bitstream. The displacement bitstream may be a bitstream containing displacement information in an image format. The image format may be, for example, a format containing two pieces of chroma information and one piece of luma information. Note that the video encoder 4808 can use a general-purpose module that has the function of converting images into a bitstream. By using a highly reliable general-purpose module as the video encoder 4808, the above function can be executed more reliably.
[0350] The inverse quantizer 4811 generates a prediction residual of wavelet coefficients by inverse quantizing the prediction residual quantized by the quantizer 4806. Specifically, the inverse quantizer 4811 can inverse quantize the prediction residual quantized by the quantizer 4806 for each LoD layer, thereby generating a prediction residual. The inverse quantizer 4811 provides the generated prediction residual of wavelet coefficients to the reconstructor 4812.
[0351] The reconstructor 4812 restores (also referred to as reconstructing) wavelet coefficients from the prediction residuals of the wavelet coefficients provided by the inverse quantizer 4811 and the reference frame stored in the reference buffer 4813. The reconstructor 4812 stores the restored wavelet coefficients in the reference buffer 4813.
[0352] The reference buffer 4813 is a storage device that stores, for example, wavelet coefficients of a reference frame. The wavelet coefficients stored in the reference buffer 4813 can be used for inter prediction by the inter predictor 4805.
[0353] FIG. 49 is a flowchart showing a specific example of the encoding process according to this embodiment.
[0354] In step S4901, the inter predictor 4805 calculates the sum of the transformation coefficients of the displacement vector within the frame (also called sum_nointer).
[0355] In step S4902, the inter predictor 4805 calculates the sum of prediction residuals within a frame (also referred to as sum_inter) when inter prediction is applied to the transform coefficients of the displacement vector.
[0356] In addition, the prediction residual of inter-prediction can be calculated by subtracting the transformation coefficient of the displacement vector of a three-dimensional point (also called a reference point) in the reference frame corresponding to the three-dimensional point to be encoded from the transformation coefficient of the displacement vector of the three-dimensional point to be encoded in the frame to be encoded.
[0357] In step S4903, the inter predictor 4805 determines whether sum_inter calculated in step S4902 is smaller than sum_nointer calculated in step S4901. If it is determined that sum_inter is smaller than sum_nointer (Yes in step S4903), the process proceeds to step S4904; otherwise (No in step S4903), the process proceeds to step S4911.
[0358] In step S4904, the inter predictor 4805 determines that the transform coefficients of the displacement vectors within the frame are to be coded using inter prediction, and outputs the prediction residual to the quantizer 4806.
[0359] In step S4905, information indicating that the transform coefficients of the displacement vectors in the frame have been coded using inter prediction is added to a header (e.g., a stream header). For example, by setting the information disp_frame_inter_mode included in the header, which indicates that the transform coefficients of the displacement vectors have been coded using inter prediction, to 1, it is possible to indicate that the transform coefficients of the displacement vectors in the frame have been coded using inter prediction.
[0360] In step S4911, it is determined that the transform coefficients of the displacement vectors within the frame are coded without using inter prediction, and the transform coefficients are output to the quantizer 4806.
[0361] In step S4912, information indicating that the transform coefficients of the displacement vectors in the frame have been coded without using inter prediction is added to the header. For example, by setting the information disp_frame_inter_mode included in the header, which indicates that the transform coefficients of the displacement vectors have been coded using inter prediction, to 0, it is possible to indicate that the transform coefficients of the displacement vectors in the frame have been coded without using inter prediction.
[0362] FIG. 50 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.
[0363] The decoding device comprises a video decoder 5001 , an image unpacker 5002 , an inverse quantizer 5003 , a reconstructor 5004 , an inverse wavelet transformer 5005 , a reconstructor 5006 , and a reference buffer 5011 .
[0364] The video decoder 5001 acquires a displaced bitstream and decodes the acquired displaced bitstream into an image. The image may be an image stored by mapping quantized wavelet coefficients to pixels in a two-dimensional image format. The video decoder 5001 provides the image to the image unpacker 5002. Note that the video decoder 5001 may use a general-purpose module having a function for converting a bitstream into an image. By using a highly reliable general-purpose module as the video decoder 5001, the above function can be more reliably executed.
[0365] The image unpacker 5002 extracts quantized wavelet coefficients from the image provided by the video decoder 5001. The process of extracting the quantized wavelet coefficients from the image may use a mapping that represents the assignment of the quantized wavelet coefficients to pixels in a two-dimensional image format. The image unpacker 5002 provides the quantized wavelet coefficients extracted from the image to the inverse quantizer 5003.
[0366] The inverse quantizer 5003 generates prediction residuals of wavelet coefficients by inverse quantizing the prediction residuals of the quantized wavelet coefficients provided from the image unpacker 5002. Specifically, the inverse quantizer 5003 can generate prediction residuals of wavelet coefficients by inverse quantizing the prediction residuals of the quantized wavelet coefficients for each LoD layer.
[0367] The reconstructor 5004 restores (reconstructs) the transform coefficients of the frame to be decoded from the prediction residuals of the transform coefficients of the frame to be decoded and the transform coefficients of the reference frame, and provides the restored transform coefficients to the inverse wavelet transformer 5005 and the reference buffer 5011.
[0368] The inverse wavelet transformer 5005 generates a displacement vector (corresponding to a decoded displacement vector) by performing an inverse wavelet transform process on the wavelet coefficients provided by the reconstructor 5004. The inverse wavelet transform process corresponds to the inverse transform of the wavelet transform process performed by the wavelet transformer 4804. Specifically, the inverse wavelet transformer 5005 can calculate the displacement vectors of the vertices by applying an inverse lifting transform to the wavelet coefficients in the inverse wavelet transform. The inverse wavelet transformer 5005 provides the generated decoded displacement vector to the reconstructor 5006.
[0369] The reconstructor 5006 reconstructs a mesh (corresponding to a decoded mesh) using the decoded displacement vectors and the decoded base mesh provided by the inverse wavelet transformer 5005. The reconstructor 5006 outputs the reconstructed decoded mesh.
[0370] The reference buffer 5011 is a storage device that stores, for example, wavelet coefficients of a reference frame. The wavelet coefficients stored in the reference buffer 5011 can be used by the reconstructor 5004 to reconstruct the transform coefficients of the frame to be decoded.
[0371] FIG. 51 is a flowchart showing a specific example of the decoding process according to this embodiment.
[0372] In step S5101, the reconstructor 5004 determines whether information indicating that the transform coefficients of the displacement vectors in the frame have been coded using inter prediction is added to the header. The information is, for example, information indicating that disp_frame_inter_mode is 1. If it is determined that the information is added to the header (Yes in step S5101), the process proceeds to step S5102; if not (No in step S5101), the process proceeds to step S5111.
[0373] In step S5102, the reconstructor 5004 determines that the transform coefficients of the displacement vectors in the frame have been coded using inter-prediction, and adds the transform coefficients of the reference frame to the prediction residual provided by the inverse quantizer 5003 to decode and output the transform coefficients.
[0374] In step S5111, the reconstructor 5004 determines that the transform coefficients of the displacement vectors within the frame were coded without using inter prediction, and outputs the transform coefficients provided by the inverse quantizer 5003.
[0375] FIG. 52 is a flowchart showing a specific example of the encoding process according to this embodiment.
[0376] In the encoding process shown in Figure 52, an example is shown in which the unit of inter prediction of the displacement vector is determined for each LoD, and information disp_lod_inter_mode indicating whether inter prediction has been applied for each LoD is added to the header. This makes it possible to switch inter prediction on or off for each LoD, thereby improving encoding efficiency. For example, there are cases in which inter prediction is more likely to be accurate for small movements and less likely to be accurate for large movements. In such cases, turning on inter prediction in a lower layer of the LoD layer where high-frequency components are concentrated and turning off inter prediction in a higher layer where low-frequency components are concentrated can contribute to improving encoding efficiency.
[0377] In step S5201, the inter predictor 4805 performs start processing of loop A, which repeatedly executes the processing of steps S5202 to S5206 and steps S5211 to S5212, which will be described later. In loop A, attention is paid to each of one or more LoDs, and processing is executed for the focused LoD, and control is exercised so that processing is ultimately performed for all LoDs. The focused LoD is also referred to as the focused LoD. Loop A may also be referred to as an LoD loop.
[0378] In step S5202, the inter predictor 4805 calculates the sum (also referred to as sum_lod_nointer) of the transform coefficients of the displacement vector within the LoD of interest.
[0379] In step S5203, the inter predictor 4805 calculates the sum (also referred to as sum_lod_inter) of prediction residuals within the LoD of interest when inter prediction is applied to the transform coefficients of the displacement vector.
[0380] In step S5204, the inter predictor 4805 determines whether sum_lod_inter calculated in step S5203 is smaller than sum_lod_nointer calculated in step S5202. If it is determined that sum_lod_inter is smaller than sum_lod_nointer (Yes in step S5204), the process proceeds to step S5205; otherwise (No in step S5204), the process proceeds to step S5211.
[0381] In step S5205, the inter predictor 4805 determines to encode the transform coefficients of the displacement vector within the LoD of interest using inter prediction, and outputs the prediction residual to the quantizer 4806.
[0382] In step S5206, the inter predictor 4805 adds information indicating that the transform coefficients of the displacement vectors in the LoD of interest have been coded using inter prediction to a header (e.g., a stream header). For example, by setting the information disp_lod_inter_mode(i) included in the header, which indicates that the transform coefficients of the displacement vectors have been coded using inter prediction, to 1, it is possible to indicate that the transform coefficients of the displacement vectors in the LoD of interest have been coded using inter prediction. Note that i is the ordinal number of the LoD of interest, indicating the ordinal number of the LoD of interest. The same applies hereinafter.
[0383] In step S5211, the inter predictor 4805 determines to encode the transform coefficients of the displacement vectors within the LoD of interest without using inter prediction, and outputs the transform coefficients to the quantizer 4806.
[0384] In step S5212, the inter predictor 4805 adds information indicating that the transform coefficients of the displacement vectors in the LoD of interest have been coded without using inter prediction to the header. For example, by setting the information disp_lod_inter_mode(i), which is included in the header and indicates that the transform coefficients of the displacement vectors have been coded using inter prediction, to 0, it is possible to indicate that the transform coefficients of the displacement vectors in the LoD of interest have been coded without using inter prediction.
[0385] In step S5207, the inter predictor 4805 performs processing to end loop A. Specifically, the inter predictor 4805 determines whether or not the processing of steps S5202 to S5206 and steps S5211 to S5212 (however, only one of the processing of steps S5205 and S5206 and the processing of steps S5211 and S5212 is performed depending on the determination result of step S5204) has been performed for all LoDs, and if not, controls to perform processing focusing on the LoDs that have not yet been performed.
[0386] FIG. 53 is a flowchart showing a specific example of the decoding process according to this embodiment.
[0387] 53 shows an example of decoding a bitstream that is coded by determining the unit of inter prediction of a displacement vector for each LoD and adding information disp_lod_inter_mode indicating whether inter prediction is applied for each LoD to a header. This allows inter prediction to be switched on or off for each LoD, which can contribute to appropriately decoding a bitstream with improved coding efficiency.
[0388] In step S5301, the reconstructor 5004 starts loop A, which repeatedly executes steps S5302 to S5203 and step S5311, which will be described later. In loop A, the reconstructor 5004 focuses on each of one or more LoDs, executes processing for the focused LoD, and controls the processing so that all LoDs are ultimately processed. The focused LoD is also referred to as the focused LoD. Loop A may also be called an LoD loop.
[0389] In step S5302, the reconstructor 5004 determines whether information indicating that the transform coefficients of the displacement vectors in the LoD of interest have been coded using inter prediction is added to the header. The information is, for example, information indicating that disp_lod_inter_mode(i) is 1. If it is determined that the information is added to the header (Yes in step S5302), the process proceeds to step S5303; if not (No in step S5302), the process proceeds to step S5311.
[0390] In step S5303, the reconstructor 5004 determines that the transform coefficients of the displacement vector within the target LoD have been coded using inter-prediction, and adds the transform coefficients of the reference frame to the prediction residual provided by the inverse quantizer 5003 to decode and output the transform coefficients.
[0391] In step S5311, the reconstructor 5004 determines that the transform coefficients of the displacement vector within the LoD of interest have been coded without using inter prediction, and outputs the transform coefficients provided by the inverse quantizer 5003.
[0392] In step S5304, the reconstructor 5004 performs the end process of loop A. Specifically, the reconstructor 5004 determines whether the processes of steps S5302 to S5203 and step S5311 (however, only one of the processes of step S5203 and step S5311 is performed depending on the determination result of step S5302) have been performed for all LoDs, and if not, controls to perform the processes focusing on the LoDs that have not yet been performed.
[0393] FIG. 54 is an explanatory diagram showing an example of syntax according to this embodiment.
[0394] The example syntax shown in Figure 54 shows an example of the structure of information contained in a bitstream generated by an encoding device.
[0395] The syntax shown in Figure 54 includes a displacement vector_header, which includes a disp_frame_inter_mode.
[0396] "disp_frame_inter_mode" is information indicating whether the displacement vector in the frame is coded using inter prediction. For example, a value of 1 may indicate that the displacement vector in the frame is coded using inter prediction, and a value of 0 may indicate that the displacement vector in the frame is coded without using inter prediction. This allows the decoding device to determine whether the displacement vector in the frame is coded using inter prediction and to appropriately decode the bitstream.
[0397] Note that the unit for adding disp_frame_inter_mode is not limited to frame units. For example, disp_frame_inter_mode may be added on a submesh basis. This allows for improved coding efficiency by switching between using and not using inter prediction on a submesh basis. For example, coding efficiency can be improved by using inter prediction for submeshes with small motion such as background objects and not using inter prediction for submeshes with large motion such as foreground objects.
[0398] Furthermore, disp_frame_inter_mode may be added on a sequence-by-sequence basis. This allows the header code amount to be reduced. For example, when encoding a sequence with small overall motion, inter prediction is used for the entire sequence, and when encoding a sequence with large overall motion, inter prediction is not used for the entire sequence, thereby reducing the header code amount and improving coding efficiency. Note that when inter prediction is used for the entire sequence, disp_frame_inter_mode and disp_lod_inter_mode do not need to be added to the header. This allows the header code amount to be reduced.
[0399] FIG. 55 is an explanatory diagram showing an example of syntax according to this embodiment.
[0400] The syntax shown in Figure 55 includes a displacement vector_header, which includes disp_lod_inter_mode[i].
[0401] disp_lod_inter_mode[i] is information indicating whether or not the displacement vector belonging to the i-th LoD is coded using inter prediction when generating LoDs and coding the displacement vector. For example, a value of 1 may indicate that the displacement vector of the 3D point belonging to the i-th LoD is coded using inter prediction, and a value of 0 may indicate that the displacement vector of the 3D point belonging to the i-th LoD is coded without using inter prediction. This allows the decoding device to determine whether or not the displacement vector of the 3D point belonging to the i-th LoD is coded using inter prediction, and to appropriately decode the bitstream.
[0402] FIG. 56 is an explanatory diagram showing an example of syntax according to this embodiment.
[0403] The syntax shown in Figure 56 includes a displacement vector_header, which includes disp_frame_inter_mode, disp_lod_inter_mode_present, and disp_lod_inter_mode[i].
[0404] The disp_frame_inter_mode is the same as the disp_frame_inter_mode shown in FIG.
[0405] disp_lod_inter_mode_present indicates whether disp_lod_inter_mode is included in the header. For example, a value of 1 may indicate that disp_lod_inter_mode is included in the header, and a value of 0 may indicate that disp_lod_inter_mode is not included in the header. Note that when disp_frame_inter_mode=1, inter prediction is performed on a frame-by-frame basis, so the value of disp_lod_inter_mode_present is estimated to be 0, and disp_lod_inter_mode does not need to be added to the header. This allows the amount of header coding to be reduced.
[0406] disp_lod_inter_mode[i] is the same as disp_lod_inter_mode[i] shown in FIG.
[0407] Note that when disp_frame_inter_mode=0, the value of disp_lod_inter_mode_present may be omitted, and disp_lod_inter_mode[i] may be included in the header regardless of the value of disp_lod_inter_mode_present.
[0408] Note that disp_frame_inter_mode, disp_lod_inter_mode, or disp_lod_inter_mode_present may be entropy coded and added to the header. For example, each value may be binarized and arithmetically coded. Alternatively, to reduce the amount of processing, they may be coded at a fixed length.
[0409] FIG. 57 is a block diagram showing an example of the configuration of an encoding device according to this embodiment.
[0410] The encoding device shown in Figure 57 comprises a decimator 5701, a sub-divider 5702, a displacement vector calculator 5703, a wavelet transformer 5704, a LoD-based inter predictor 5705, a quantizer 5706, a switch 5707, an image packer 5708, a video coder 5709, an arithmetic coder 5710, an inverse quantizer 5711, a reconstructor 5712, and a reference buffer 5713.
[0411] The decimator 5701, the subdivider 5702, and the displacement vector calculator 5703 are the same as the decimator 4801, the subdivider 4802, and the displacement vector calculator 4803 shown in FIG. 48, respectively.
[0412] The wavelet transformer 5704 acquires transform coefficients (also referred to as wavelet coefficients) by performing wavelet transform processing on the displacement vectors calculated by the displacement vector calculator 5703. The wavelet transformer 5704 provides the acquired wavelet coefficients to the LoD-based inter predictor 5705. In the wavelet transform, the wavelet transformer 5704 assigns vertices to multiple LoD layers and applies, for example, a lifting transform to the displacement vectors of the vertices, thereby being able to calculate wavelet coefficients representing various components from low frequency components to high frequency components.
[0413] The LoD-based inter predictor 5705 outputs, for each LoD, a prediction residual of the wavelet coefficient of the displacement vector of the frame to be coded using inter prediction. Specifically, the LoD-based inter predictor 5705 outputs a prediction residual of the wavelet coefficient of the displacement vector of the frame to be coded by inter predicting, for each LoD, the wavelet coefficient of the displacement vector of the frame to be coded, using the wavelet coefficient of the displacement vector of an encoded frame (also referred to as a reference frame) stored in the reference buffer 5713. The LoD-based inter predictor 5705 can determine, for each LoD, whether to code the transform coefficient of the displacement vector using inter prediction, and switch between these two methods.
[0414] The quantizer 5706 quantizes the prediction residual of the wavelet coefficients calculated by the LoD-based inter predictor 5705. The quantizer 5706 quantizes the prediction residual of the wavelet coefficients for each LoD layer. The quantizer 5706 provides the quantized prediction residual to the image packer 5708 or the arithmetic encoder 5710 via the switch 5707, and also to the inverse quantizer 5711.
[0415] A switch 5707 is a switch that switches whether the prediction residual quantized by the quantizer 5706 is provided to an image packer 5708 or an arithmetic encoder 5710 .
[0416] The image packer 5708 and video encoder 5709 are the same as the image packer 4807 and video encoder 4808 shown in FIG. 48, respectively.
[0417] The arithmetic encoder 5710 uses arithmetic coding to encode the prediction residual quantized by the quantizer 5706 into a bitstream (also referred to as a displaced bitstream) (in other words, generates a displaced bitstream). The arithmetic encoder 5710 outputs the displaced bitstream.
[0418] The inverse quantizer 5711 generates a prediction residual of wavelet coefficients by inverse quantizing the prediction residual quantized by the quantizer 5706. Specifically, the inverse quantizer 5711 generates a prediction residual by inverse quantizing, for each LoD layer, the prediction residual quantized for each LoD layer by the quantizer 5706. The inverse quantizer 5711 provides the generated prediction residual of wavelet coefficients to the reconstructor 5712.
[0419] The reconstructor 5712 restores (also referred to as reconstructing) wavelet coefficients from the prediction residuals of the wavelet coefficients provided by the inverse quantizer 5711 and the reference frame stored in the reference buffer 5713. The reconstructor 5712 stores the restored wavelet coefficients in the reference buffer 5713.
[0420] The reference buffer 5713 is a storage device that stores, for example, wavelet coefficients of a reference frame. The wavelet coefficients stored in the reference buffer 5713 can be used for inter prediction by the LoD-based inter predictor 5705.
[0421] The encoding device shown in Figure 57 can switch between encoding the transform coefficients of the displacement vectors using the video encoder 5709 and encoding them by arithmetic coding using the arithmetic encoder 5710. This allows the encoding device to efficiently encode the displacement vectors using arithmetic coding even when the video encoder 5709 cannot be used.
[0422] In addition, even when using arithmetic coding, the transform coefficients of the displacement vector may be coded by switching whether to use inter prediction for each LoD. Generally, the smaller the change in the input value, the more efficiently the arithmetic coding can compress, so that the coding efficiency can be improved by suppressing the change in the value of the prediction residual of the transform coefficients of the displacement vector by inter prediction for each LoD.
[0423] Note that information indicating whether the transformation coefficients of the displacement vectors have been coded using the video coder 5709 or arithmetically coded using the arithmetic coder 5710 may be added to the header. This allows the decoding device to appropriately switch the decoding method by referring to the header information.
[0424] 57 has been described with reference to an example in which it is switched between encoding the transform coefficients of the displacement vectors using the video encoder 5709 and encoding them by arithmetic coding using the arithmetic encoder 5710, but it may also be configured to always use arithmetic coding. This configuration may potentially improve coding efficiency compared to the case in which inter prediction is used in all LoD layers when arithmetically coding the transform coefficients of the displacement vectors.
[0425] The decoding device shown in FIG. 58 comprises a video decoder 5801, an image unpacker 5802, an arithmetic decoder 5803, a switch 5804, an inverse quantizer 5805, a reconstructor 5806, an inverse wavelet transformer 5807, a reconstructor 5808, and a reference buffer 5811.
[0426] The video decoder 5801 and image unpacker 5802 are similar to the video decoder 5001 and image unpacker 5002 shown in FIG. 50, respectively.
[0427] The arithmetic decoder 5803 obtains the displaced bitstream and arithmetically decodes the prediction residual included in the obtained displaced bitstream. Note that the arithmetic decoder 5803 may also decode various types of header information.
[0428] The switch 5804 is a switch that switches between providing the prediction residual provided by the image unpacker 5802 to the inverse quantizer 5805 and providing the prediction residual provided by the arithmetic decoder 5803 to the inverse quantizer 5805.
[0429] The inverse quantizer 5805 generates prediction residuals of wavelet coefficients by inverse quantizing prediction residuals of quantized wavelet coefficients provided from the image unpacker 5802 or the arithmetic decoder 5803 via the switch 5804. Specifically, the inverse quantizer 5805 generates prediction residuals of wavelet coefficients by inverse quantizing prediction residuals of quantized wavelet coefficients for each LoD layer.
[0430] The reconstructor 5806 restores (also referred to as reconstructing) the transform coefficients of the frame to be decoded from the prediction residuals of the transform coefficients of the frame to be decoded and the transform coefficients of the reference frame. The reconstructor 5806 provides the restored transform coefficients to the inverse wavelet transformer 5807 and the reference buffer 5811. The reconstructor 5806 may determine from header information whether inter prediction has been applied for each LoD and switch the restoration method.
[0431] The inverse wavelet transformer 5807 generates displacement vectors (corresponding to decoded displacement vectors) by performing an inverse wavelet transform process on the wavelet coefficients provided by the reconstructor 5806. The inverse wavelet transform process corresponds to the inverse transform of the wavelet transform process performed by the wavelet transformer 5704. Specifically, the inverse wavelet transformer 5807 can calculate the displacement vectors of the vertices by applying an inverse lifting transform to the wavelet coefficients in the inverse wavelet transform.
[0432] The inverse wavelet transformer 5807 provides the generated decoded displacement vectors to the reconstructor 5808 .
[0433] The reconstructor 5808 reconstructs a mesh (corresponding to the decoded mesh) using the decoded displacement vectors and the decoded base mesh provided by the inverse wavelet transformer 5807. The reconstructor 5808 outputs the reconstructed decoded mesh.
[0434] The reference buffer 5811 is a storage device that stores, for example, wavelet coefficients of a reference frame. The wavelet coefficients stored in the reference buffer 5811 can be used by the reconstructor 5806 to reconstruct the transform coefficients of the frame to be decoded.
[0435] The decoding device shown in Figure 58 may decode header information, determine whether the encoding device has coded the transform coefficients of the displacement vectors using a video encoder (e.g., video encoder 5709) or arithmetically coded using an arithmetic encoder (e.g., arithmetic encoder 5710), and switch the decoding method accordingly. This allows the decoding device to properly decode a bitstream in which the displacement vectors have been efficiently coded using arithmetic coding even when the video encoder (e.g., video encoder 5709) is unavailable.
[0436] FIG. 59 is an explanatory diagram showing the positional relationship of three-dimensional points according to this embodiment.
[0437] A possible method for encoding a displacement vector of a three-dimensional point is to calculate a predicted value of the displacement vector of a certain three-dimensional point and encode the difference (prediction residual) between the value of the original displacement vector and the predicted value. For example, if the value of the displacement vector of a certain three-dimensional point p is Ap and the predicted value is Pp, the encoding device 100 encodes the absolute difference value Diffp = |Ap - Pp| indicating the absolute value of the difference and information indicating the positive or negative sign of (Ap - Pp). In this case, if the predicted value Pp can be generated with high accuracy, the value of the absolute difference value Diffp will be small. Therefore, for example, the encoding device 100 can reduce the amount of code by performing entropy encoding using a coding table (or context) in which the smaller the value, the smaller the number of generated bits.
[0438] As a method for the encoding device 100 to generate a predicted value of a displacement vector, it is possible to use the displacement vector of another 3D point surrounding the 3D point to be encoded. Here, a 3D point surrounding a 3D point refers to another 3D point that is within a predetermined distance (within a predetermined range) from the 3D point. For example, when there are 3D points p=(x1, y1, z1) and 3D points q=(x2, y2, z2) that are the 3D points to be encoded, the Euclidean distance d(p, q) between the 3D points p and q is √((x1-y1) 2 +(x2-y2) 2 +(x3-y3) 2 ) is smaller than a certain threshold THd, the encoding device 100 determines that the position of the 3D point q is close to the position of the 3D point p, and determines to use the value of the displacement vector of the 3D point q to generate a predicted value of the displacement vector of the 3D point p.
[0439] The distance calculation method may be another method, such as Mahalanobis distance.
[0440] Furthermore, for example, the encoding device 100 may determine that a 3D point that is farther away from the 3D point to be encoded than a predetermined distance (outside a predetermined range) is not to be used for prediction. For example, if a 3D point r exists and the distance d(p, r) between the 3D point p and the 3D point r is equal to or greater than a threshold value THd, the encoding device 100 may determine that the 3D point r is not to be used for prediction. Furthermore, the predetermined distance may be determined arbitrarily and is not particularly limited.
[0441] The encoding device 100 may add the value of the threshold THd to the header of the bitstream.
[0442] For example, when encoding the displacement vector of a three-dimensional point to be encoded using a predicted value, the encoding device 100 uses displacement vectors of surrounding three-dimensional points used to generate the predicted value, and uses displacement vectors that have already been encoded or displacement vectors that have already been decoded.
[0443] In addition, when decoding the displacement vector of the three-dimensional point to be decoded using a predicted value, the decoding device 200 uses displacement vectors that have already been decoded when using the displacement vectors of surrounding three-dimensional points to generate the predicted value.
[0444] This allows the same predicted value to be generated during encoding and decoding, thereby enabling the decoding device 200 to correctly decode the bit stream of 3D points generated by the encoding device 100.
[0445] Although it has been described that the points surrounding a three-dimensional point refer to other three-dimensional points within a predetermined range from the three-dimensional point, this is not necessarily limited to this. For example, in the case of three-dimensional point D (i.e., vertex D) shown in FIG. 47, three-dimensional points A, B, C, E, F, and G exist as surrounding three-dimensional points, but the surrounding three-dimensional points (in other words, adjacent points) may be selected according to one or more of the following conditions A and B. In other words, adjacent points are points selected according to certain conditions and are points that are referenced to predict the information of the three-dimensional point to be encoded. An adjacent point may also be referred to as a reference three-dimensional point, a reference point, or a reference vertex.
[0446] Condition A: A 3D point that has connectivity with the target 3D point. Condition B: A 3D point that has been encoded or decoded before the target 3D point.
[0447] For example, when a three-dimensional point that satisfies the above conditions A and B is selected as an adjacent point, and when three-dimensional point D and its adjacent points are encoded or decoded in the order of three-dimensional point A, three-dimensional point C, three-dimensional point E, three-dimensional point F, three-dimensional point D, three-dimensional point B, and three-dimensional point G, three-dimensional point A, three-dimensional point C, three-dimensional point E, and three-dimensional point F may be selected as adjacent points of three-dimensional point D. Since three-dimensional point A, three-dimensional point C, three-dimensional point E, and three-dimensional point F have connectivity with three-dimensional point D, it is highly likely that their displacement vectors are also close in value. Furthermore, since three-dimensional point A, three-dimensional point C, three-dimensional point E, and three-dimensional point F have been encoded or decoded before three-dimensional point D, the displacement vectors of three-dimensional point A, three-dimensional point C, three-dimensional point E, and three-dimensional point F can be used to calculate the predicted value of the displacement vector of three-dimensional point D. This can improve the accuracy of the predicted value of the displacement vector of three-dimensional point D, thereby improving encoding efficiency.
[0448] In addition to the above conditions A and B, the number of neighboring points of a three-dimensional point may be limited to a predetermined value (NumNeiCnt) or less. For example, by setting NumNeiCnt = 3, the number of neighboring points of a three-dimensional point may be limited to three or less. This reduces the memory capacity required to store information on neighboring points of a three-dimensional point, and also reduces the amount of processing required when calculating a predicted displacement vector. The predetermined value may be set arbitrarily and is not particularly limited.
[0449] Furthermore, for example, the encoding device 100 may add the above-mentioned predetermined value, in other words, NumNeiCnt indicating the maximum number of adjacent points, to the header of the data unit and encode it to add it to the bitstream.
[0450] As a result, the decoding device 200 can properly decode a bitstream in which the maximum number of adjacent points is limited to NumNeiCnt or less by decoding the header of the bitstream.
[0451] In addition, when there are more three-dimensional points satisfying the above conditions A and B than NumNeiCnt as adjacent points, the adjacent points may be selected in order of proximity to the three-dimensional point to be encoded or decoded. For example, when NumNeiCnt = 3, there are four three-dimensional points satisfying the above conditions A and B as adjacent points of three-dimensional point D, namely, three-dimensional point A, three-dimensional point C, three-dimensional point E and three-dimensional point F, and when the three-dimensional points A, three-dimensional point C, three-dimensional point E and three-dimensional point F are closest in distance to the three-dimensional point D in this order, three-dimensional point A, three-dimensional point C and three-dimensional point E may be selected as adjacent points of three-dimensional point D. Because three-dimensional point A, three-dimensional point C and three-dimensional point E have connectivity with three-dimensional point D and are close in distance to three-dimensional point D, the values of their respective displacement vectors are likely to be close to the value of the displacement vector of three-dimensional point D. Furthermore, the 3D points A, C, and E have been encoded or decoded before the 3D point D. Therefore, the displacement vectors of the 3D points A, C, and E can be used to calculate a predicted value of the displacement vector of the 3D point D.
[0452] This improves the accuracy of the predicted value of the displacement vector of the 3D point D. Furthermore, by limiting the number of adjacent points, it is possible to reduce the memory capacity required to store information about adjacent points of the 3D point and reduce the amount of processing required to calculate the predicted displacement vector.
[0453] Note that the method of selecting neighboring points of a 3D point is not limited to the above. For example, if the 3D point to be encoded is 3D point Z generated by subdivision from 3D points X and Y that constitute a base mesh, 3D points X and Y of the base mesh may be set as neighboring points of 3D point Z. Basically, when 3D point Z is generated by subdivision from 3D points X and Y, 3D point Z exists on the line connecting 3D point X and 3D point Y, and therefore 3D point X and 3D point Y become neighboring points of 3D point Z. Furthermore, the displacement vector of 3D point Z is likely to be close in value to the displacement vectors of 3D points X and Y. Therefore, encoding efficiency can be improved by using 3D points X and Y as neighboring points to calculate a predicted value of the displacement vector of 3D point Z.
[0454] The method for generating LoD will be described below.
[0455] 60 and 61 are explanatory diagrams showing a method for generating LoD in this embodiment.
[0456] When encoding the displacement vector of a three-dimensional point, the encoding device may classify each three-dimensional point into one or more layers using the position information of the three-dimensional point and then encode it. Here, each layer used for classification is called LoD (Level of Detail). An identifier (e.g., a number) that uniquely identifies the LoD is assigned to the LoD. For example, the 0th LoD is also called LoD0, the 1st LoD is also called LoD1, the nth LoD is also called LoDn, and the (n-1)th LoD is also called LoD(n-1).
[0457] The LoD generation method will be described with reference to Figures 60 and 61. In addition, if the encoding device or decoding device cannot calculate the position information or distance information of a 3D point in a frame to be encoded or decoded, the position information or distance information of a 3D point corresponding to the 3D point in a frame that has already been encoded or decoded may be used. This may allow the 3D points to be encoded or decoded to be classified into one or more layers and coded efficiently.
[0458] 60 shows three-dimensional points to be encoded, namely, points a0, a1, a2, b0, b1, b2, c0, c1, and c2, where d(x, y) indicates the distance between point x and point y.
[0459] By setting the threshold value of each LoD layer larger for higher layers (layers closer to LoD0), the higher the layer, the greater the distance between three-dimensional points (also called sparse point cloud), and the lower the layer, the closer the distance between three-dimensional points (also called dense point cloud). Here, LoD0 is the top layer (see FIG. 61).
[0460] Point y belongs to the same LoD as point x if the distance d(x, y) from point x is greater than the threshold of the LoD to which point x belongs and is equal to or less than the threshold of the LoD higher than that LoD. Note that if point x belongs to LoD0, which is the highest layer, point y belongs to the same LoD as point x if the distance d(x, y) from point x is greater than the threshold of the LoD to which point x belongs.
[0461] First, the encoding device selects point a0 as the initial point and assigns it to LoD0. Next, the encoding device extracts point a1 whose distance from point a0 is greater than the threshold Thres_Lod[0] of LoD0 and assigns it to LoD0. Next, the encoding device extracts point a2 whose distance from point a1 is greater than the threshold Thres_Lod[0] of LoD0 and assigns it to LoD0. In this way, the encoding device configures LoD0 so that the distance between each point in LoD0 is greater than the threshold Thres_Lod[0].
[0462] Next, the encoding device selects point b0, which has not yet been assigned an LoD, and assigns it to LoD1. Next, the encoding device selects point b1, whose distance from point b0 is greater than the threshold Thres_Lod[1] for LoD1 and has not yet been assigned an LoD, and assigns it to LoD1. Next, the encoding device selects point b2, whose distance from point b1 is greater than the threshold Thres_Lod[1] for LoD1 and has not yet been assigned an LoD, and assigns it to LoD1. In this way, the encoding device configures LoD1 so that the distance between each point in LoD1 is greater than the threshold Thres_Lod[1].
[0463] Next, the encoding device selects point c0, which has not yet been assigned an LoD, and assigns it to LoD2. Next, the encoding device selects point c1, whose distance from point c0 is greater than the LoD2 threshold Thres_Lod[2] and has not yet been assigned an LoD, and assigns it to LoD2. Next, the encoding device selects point c2, whose distance from point c1 is greater than the LoD2 threshold Thres_Lod[2] and has not yet been assigned an LoD, and assigns it to LoD2. In this way, LoD2 is constructed so that the distance between each point in LoD2 is greater than the threshold Thres_Lod[2].
[0464] The thresholds for each LoD may be added to the header of the bitstream. For example, in the case of Fig. 60, the thresholds Thres_Lod[0], Thres_Lod[1], and Thres_Lod[2] may be added to the header of the bitstream.
[0465] Alternatively, all three-dimensional points that have not yet been assigned an LoD may be assigned to the lowest layer of the LoD. In this case, the threshold value of the lowest layer of the LoD is not added to the header, which has the effect of reducing the amount of coding for the header. For example, in the case of Figure 60, the encoding device may add thresholds Thres_Lod[0] and Thres_Lod[1] to the header, and Thres_Lod[2] may not be added to the header, so that the decoding device may estimate Thres_Lod[2] to be a value of 0.
[0466] Furthermore, the number of layers of the LoD may be added to the header, which allows the decoding device to determine whether the LoD is the lowest layer.
[0467] In addition, when the LoD layer is one layer, that is, when encoding the displacement vector of a 3D point without generating the LoD, the encoding device may omit the LoD generation process described in the above example. Alternatively, the encoding device may apply the LoD generation method described in the above example with the LoD layer set to 1. In this case, the encoding device may perform the LoD generation process assuming that all 3D points belong to the same LoD. This allows the encoding device to reduce the processing time for generating the LoD.
[0468] The encoding or decoding of the displacement vector described in this embodiment may be applied to other than the LoD generation method described above. For example, even when the LoD layer to which the three-dimensional point belongs is determined in advance, the encoding efficiency may be improved by applying the displacement vector encoding method or decoding method described in this embodiment.
[0469] The LoD generation method is not limited to the above method, and the LoD to which a point belongs may be determined according to the number of times a base mesh is sub-divided, as shown in Figures 35 and 36. For example, if the base mesh is sub-divided twice, the 3D points included in the base mesh may be assigned to LoD0, the points generated by one sub-division from the 3D points of the base mesh may be assigned to LoD1, and the points generated by two sub-divisions may be assigned to LoD2. This reduces the processing time for LoD generation.
[0470] The method of selecting the initial 3D point when constructing each LoD may depend on the encoding order when encoding the displacement vector. For example, the encoding device selects the 3D point that was first coded when coding the displacement vector as the initial point a0 of LoD0, and then selects points a1 and a2 using point a0 as the base point to construct LoD0. Then, the encoding device may select, as the initial point b0 of LoD1, the 3D point whose displacement vector is coded earliest among the 3D points that do not currently belong to LoD0. In other words, the encoding device may select, as the initial point n0 of LoDn, the 3D point whose displacement vector is coded earliest among the 3D points that do not belong to the LoDs in the LoD(n-1) or lower hierarchy. This allows the same LoD as when encoding to be constructed using a similar initial point selection method (specifically, a method of selecting, as the initial point n0 of LoDn, the 3D point whose displacement vector is decoded earliest among the 3D points that do not belong to the LoDs in the LoD(n-1) or lower hierarchy), and the bitstream can be decoded appropriately.
[0471] FIG. 62 is an explanatory diagram showing a method for generating predicted values of displacement vectors in this embodiment.
[0472] The encoding device can generate predicted values of displacement vectors of 3D points using information of the LoD.
[0473] For example, when encoding three-dimensional points included in LoD0 in order, the encoding device may generate LoD1 using the encoded and decoded displacement vectors included in LoD0 and LoD1. In this way, the encoding device can generate predicted values of displacement vectors of three-dimensional points included in LoDn using the encoded and decoded displacement vectors included in LoDn' (where n'≦n).
[0474] Furthermore, the predicted value of the displacement vector of a 3D point can be generated by calculating the average of the displacement vectors of a certain number of 3D points that are coded and decoded neighbors of the 3D point to be coded. The certain number is, for example, the number of neighbors of the 3D point to be coded (e.g., N). In this case, the value N is added to the header of the bitstream, etc.
[0475] Note that a value N indicating the number of neighboring points (i.e., N) used to calculate a predicted value may be added to each 3D point for which a predicted value is generated. This allows the encoding device to select an appropriate number of N neighboring points for each 3D point for which a predicted value is generated, thereby improving the accuracy of the predicted value and reducing the prediction residual. Alternatively, the encoding device may add the value N to a bitstream header and fix it within the bitstream (in other words, the value N may be commonly used as a fixed value in encoding the 3D points included in the bitstream). This eliminates the need for the encoding device to encode or decode the value N for each 3D point, thereby reducing the amount of processing. Alternatively, the encoding device may encode the value N separately for each LoD. This may potentially improve encoding efficiency by selecting an appropriate value N for each LoD.
[0476] Furthermore, the predicted value of the displacement vector of a 3D point may be calculated from a weighted average value of N adjacent points that have already been coded and decoded. For example, the coding device may perform a weighted average using distance information between the 3D point to be coded and each of the N adjacent points. This will be described with reference to FIG. 62 .
[0477] When encoding using a different value N for each LoD, the encoding device may, for example, set the value of N to be larger for higher LoD layers and set the value of N to be smaller for lower LoD layers. In higher LoD layers, the distance between 3D points belonging to the LoD is relatively large, so by setting the value of N to be large, it is possible to improve prediction accuracy by selecting and averaging a relatively large number of surrounding 3D points. In lower LoD layers, the distance between 3D points belonging to the LoD is relatively small, so setting the value of N to be small may reduce the amount of averaging processing and enable efficient prediction.
[0478] A prediction of a point P belonging to LoDN is generated from a reconstructed point P′ belonging to LoDN′ (where N′≦N), where neighbors are selected to point P′ based on connectivity and distance.
[0479] The predicted value of the displacement vector may be calculated from an unweighted average value, which reduces the amount of processing.
[0480] As shown in FIG. 62 , point a2 is predicted from point a0 and point a1. Furthermore, point b2 is predicted from point a0, point a1, point a2, point b0, and point b1. Note that the points selected as adjacent points used for prediction may vary depending on the number N of adjacent points used for prediction. For example, when N=5, points a0, a1, point a2, point b0, and point b1 may be selected as adjacent points of point b2, and when N=4, points a0, a1, point a2, and point b1 may be selected based on distance information.
[0481] As an example, three-dimensional points contained in the base mesh are included in LoD0, three-dimensional points generated from the three-dimensional points of the base mesh by one subdivision are included in LoD1, and three-dimensional points generated from the three-dimensional points of the base mesh by two subdivisions are included in LoD2.
[0482] For example, when a weighted average value of adjacent points is used for prediction, the predicted value a2p of point a2 is calculated by the weighted average of points a0 and a1 (see (Equation 1) and (Equation 2)). iis the value of the displacement vector of point ai.
[0483]
[0484] however,
[0485]
[0486] The predicted value b2p of point b2 is calculated by the weighted average of points a0, a1, a2, b0, and b1 (see (Equation 3), (Equation 4), and (Equation 5)). i is the value of the displacement vector of point bi.
[0487]
[0488] however,
[0489]
[0490]
[0491] It is also possible to avoid referencing the same layer when generating a predicted value of a displacement vector. This reduces the amount of processing. In addition, when a 3D point is generated by subdivision at the midpoint between two 3D points, the weight w i may be fixed at 0.5, which can reduce the amount of processing.
[0492] When encoding the value of a displacement vector of a three-dimensional point, the encoding device may calculate a difference value between the three-dimensional point and a predicted value generated from adjacent points of the three-dimensional point (also called a transform coefficient, see Equations 6 and 7 below), and then quantize the calculated transform coefficient for encoding. Here, transform coefficient a2c is the transform coefficient of point a2, and transform coefficient b2c is the transform coefficient of point b2.
[0493]
[0494]
[0495] For example, the encoding device may quantize the transform coefficients by dividing them by a quantization scale. In this case, the smaller the quantization scale, the smaller the error that may occur due to quantization (quantization error), and conversely, the larger the quantization scale, the larger the quantization error.
[0496] The value obtained by quantizing the transform coefficient a2c is defined as the quantized value a2q, and the value obtained by quantizing the transform coefficient b2r is defined as the quantized value b2q (see (Equation 8) and (Equation 9) below). QS_LoD0 is the quantization scale for LoD0, and QS_LoD1 is the quantization scale for LoD1.
[0497]
[0498]
[0499] The encoding device may change the value of the quantization scale for each LoD. For example, the quantization scale may be made smaller for higher LoD layers and larger for lower LoD layers. Since the value of the displacement vector of a three-dimensional point belonging to a higher layer may be used as a predicted value of the displacement vector of a three-dimensional point belonging to a lower layer, by reducing the quantization scale of the higher layer, it is possible to suppress quantization errors that may occur in the higher layer and increase the accuracy of the predicted value, thereby improving encoding efficiency. The encoding device may also add the quantization scale for each LoD to a header, etc. This allows the encoding device to contribute to the decoding device correctly decoding the quantization scale and appropriately decoding the bitstream.
[0500] The encoding device may convert the quantized transform coefficients from signed integer values to unsigned integer values. For example, the encoding device may convert the signed integer quantized value a2q to the unsigned integer quantized value a2u as follows:
[0501] If the quantized value a2q is smaller than 0, a2u = -1 - (2 x a2q) Otherwise, a2u = 2 x a2q (Equation 10)
[0502] Furthermore, for example, the encoding device may convert the quantized value b2q, which is a signed integer value, into the quantized value b2u, which is an unsigned integer value, as follows:
[0503] If the quantized value b2q is smaller than 0, then b2u = -1 - (2 × b2q). Otherwise, b2u = 2 × b2q (Equation 11).
[0504] This has the advantage that the coding device does not need to take into account the occurrence of negative integers when entropy coding the transform coefficients.
[0505] It should be noted that the encoding device does not necessarily have to convert from a signed integer value to an unsigned integer value, and may instead entropy encode the sign bit separately, for example.
[0506] The encoding method of the transform coefficients is not limited to this, and for example, the encoding device may perform arithmetic encoding for each bit of a sign bit indicating the positive or negative sign of the transform coefficient and binarized data of the absolute value of the transform coefficient, etc., using a context. This may enable the encoding device to improve the encoding efficiency of the transform coefficients of the displacement vector.
[0507] If quantization of the transformation coefficients of the displacement vectors is not necessary, this process may be skipped and the transformation coefficients may be arithmetically coded as is, thereby reducing the processing time.
[0508] FIG. 63 is an explanatory diagram showing an example of calculation of a predicted value in this embodiment.
[0509] An example of generating LoD and calculating predicted values of displacement vectors of each 3D point will be described with reference to FIG.
[0510] In Figure 63, points a0, a1, and a2 are 3D points included in the base mesh and belong to LoD0. Points b0 and b1 are 3D points generated by one subdivision from 3D points included in the base mesh and belong to LoD1. Points c0, c1, c2, and c3 are 3D points generated by two subdivisions from 3D points included in the base mesh and belong to LoD2.
[0511] If point b0 is a 3D point generated by subdivision from points a0 and a1, then a predicted value of the displacement vector of point b0 may be calculated using points a0 and a1.
[0512] Also, if point c2 is a 3D point generated by subdivision from points a1 and b1, the predicted value of the displacement vector of point c2 can be calculated using points a1 and b1.
[0513] FIG. 64 is an explanatory diagram showing an example of calculation of conversion coefficients in this embodiment.
[0514] An example of calculating a conversion coefficient by subtracting each predicted value from the displacement vector of each three-dimensional point will be described with reference to FIG.
[0515] In Figure 64, conversion coefficients a0c, a1c, and a2c are conversion coefficients included in the base mesh and belong to LoD0. Conversion coefficients b0c and b1c are 3D points generated by one subdivision from 3D points included in the base mesh and belong to LoD1. Conversion coefficients c0c, c1c, c2c, and c3c are 3D points generated by two subdivisions from 3D points included in the base mesh and belong to LoD2.
[0516] The transform coefficient b0c of the point b0 is obtained by subtracting the predicted value b0p of the point b0 from the value of the point b0.
[0517] b0c=b0-b0p
[0518] The predicted value b0p may be the average value of the value at point a0 and the value at point a1.
[0519] b0p=(a0+a1) / 2
[0520] Furthermore, the conversion coefficient c2c of the point c2 is obtained by subtracting the predicted value c2p of the point c2 from the value of the point c2.
[0521] c2c=c2-c2p
[0522] The predicted value c2p may be the average value of the value at point a1 and the value at point b1.
[0523] c2p=(a1+b1) / 2
[0524] In addition, a coding method (lifting transform) may be applied in which transform coefficients of displacement vectors of three-dimensional points included in a lower layer of the LoD are calculated and the transform coefficients are fed back to a higher layer for coding. By applying the lifting transform, transform coefficients of low-frequency components of the displacement vectors can be collected in the higher layer, and transform coefficients of high-frequency components of the displacement vectors can be collected in the lower layer. This makes it possible to improve coding efficiency by reducing the amount of information of the transform coefficients of high-frequency components in the lower layer through quantization, for example.
[0525] FIG. 65 is an explanatory diagram showing an example of inter prediction of transform coefficients in this embodiment.
[0526] The transform coefficients of the displacement vectors of the three-dimensional points may be inter-predicted using the transform coefficients of the displacement vectors of a frame temporally different from the current frame to be coded. This will be described with reference to FIG. 65 .
[0527] The transform coefficients of the displacement vectors of the three-dimensional points may be inter-predicted using, for example, the transform coefficients of the displacement vectors of the frame that was coded or decoded immediately before.
[0528] For example, when Frame(t), which is a frame at time t, is the frame to be coded, inter-prediction may be performed using the transform coefficients of the displacement vector of the frame coded or decoded immediately before, i.e., Frame(t-1), which is a frame at time t-1.
[0529] Specifically, it is possible to encode a0c, which is the transform coefficient of the displacement vector of 3D point a0 in Frame(t), using inter prediction with a0c', which is the transform coefficient of the displacement vector of a0' corresponding to 3D point a0 in Frame(t-1). More specifically, it is possible to encode the value obtained by subtracting a0c' from a0c. If the 3D point a0 in Frame(t) and the 3D point a0' in Frame(t-1) correspond to each other between frames, it is highly likely that the displacement vector values are close. Therefore, by subtracting a0c' from a0c, it is possible to further reduce the transform coefficient, thereby improving the encoding efficiency by entropy coding.
[0530] Note that inter prediction is not limited to referencing the most recently coded frame, and any frame may be referenced. In this case, information about the referenced frame may be added to the header. This allows the decoding device to properly decode the bitstream by referring to the same frame as the frame referenced by the coding device. Inter prediction may also be performed by referencing multiple frames. For example, coding efficiency can be improved by using bi-prediction using two reference frames. Note that when bi-prediction is used, the average value of the transform coefficients of each displacement vector of the two reference frames may be used as the inter prediction value. This allows prediction values to be generated with high accuracy, improving coding efficiency.
[0531] In addition, information disp_inter_mode indicating whether inter prediction is applied to the transform coefficients of the displacement vector may be added to the header. Thereby, for example, when there is a large change in motion between frames and correspondence between three-dimensional points between frames cannot be established, inter prediction of the transform coefficients of the displacement vector is turned off (e.g., disp_inter_mode=0), and when there is a small change in motion between frames and correspondence between three-dimensional points between frames can be established, inter prediction is turned on (e.g., disp_inter_mode=1), thereby improving coding efficiency by adaptively controlling on / off of inter prediction.
[0532] If adaptive switching of inter prediction is performed on a frame-by-frame basis, disp_frame_inter_mode may be added to a header storing frame information, and if adaptive switching is performed on a sequence-by-sequence basis, disp_seq_inter_mode may be added to a header storing sequence information. This makes it possible to control whether inter prediction of transform coefficients of displacement vectors is on or off on a frame-by-frame or sequence-by-sequence basis, thereby improving coding efficiency.
[0533] Alternatively, information disp_lod_inter_mode indicating whether inter prediction is applied to the transform coefficients of the displacement vector may be prepared for each LoD, and whether inter prediction is applied may be switched for each LoD. For example, the encoding device may compare the amount of code generated when inter prediction is applied to the transform coefficients of the displacement vector and when it is not applied, select the smaller amount of code, and add the information to the header. The decoding device decodes the transform coefficients of the displacement vector according to the information added to the header. This allows for improved coding efficiency by switching whether inter prediction is applied for each LoD.
[0534] The encoding device can decode the quantized transform coefficients by inverse quantization and reconstruction, and use them for prediction of subsequent three-dimensional points to be encoded. Specifically, the encoding device calculates inverse quantization values by multiplying the quantized transform coefficients by a quantization scale, and then adds the inverse quantization values to the predicted values to obtain decoded values. For example, the encoding device can calculate the inverse quantization value a2iq from the quantization value a2q as follows, and can calculate the inverse quantization value b2iq from the quantization value b2q as follows:
[0535] a2iq=a2q×QS_LoD0 b2iq=b2q×QS_LoD1 (Formula 12)
[0536] Furthermore, the encoding device can calculate the reconstructed value a2rec from the inverse quantized value a2iq as follows, and can calculate the reconstructed value b2rec from the inverse quantized value b2iq as follows.
[0537] a2rec=a2iq+a2p b2rec=b2iq+b2p (Formula 13)
[0538] In this embodiment, a method is shown in which an encoding device constructs one or more LoDs and generates a predicted value of a displacement vector of a three-dimensional point, but this is not necessarily limited to this. For example, it may be applied when constructing a one-layer LoD and generating a predicted value of a displacement vector of a three-dimensional point, or when generating a predicted value of a displacement vector of a three-dimensional point without generating an LoD.
[0539] In this case, since all the 3D points belong to the same LoD (for example, LoD0), when encoding or decoding in order starting from the 3D points included in LoD0, the encoding device may generate predicted values of the 3D points belonging to LoD0 using the encoded and decoded displacement vectors included in LoD0. In this way, the encoding device may be able to reduce processing time by encoding without generating multiple layers of LoD.
[0540] If quantization of the transformation coefficients of the displacement vector is not necessary, the encoding device may skip the quantization and inverse quantization processes and add the arithmetically decoded transformation coefficients directly to the predicted value to obtain the decoded value, thereby reducing processing time.
[0541] FIG. 66 is an explanatory diagram showing an example of syntax in this embodiment.
[0542] The example syntax shown in Figure 66 shows an example of the structure of information contained in a bitstream generated by an encoding device.
[0543] The syntax shown in Figure 66 includes a displacement vector_header, which includes NumLoD, NumOfPoint[i], Thres_Lod[i], NumNeiCnt[i], THd[i], and QS[i].
[0544] NumLoD indicates the number of layers of the LoD.
[0545] NumOfPoint[i] indicates the number of 3D points belonging to layer i. Note that if the encoding device adds the total number of 3D points, AllNumOfPoint, to a separate header, NumOfPoint[NumLoD-1] (i.e., the number of 3D points belonging to the lowest layer) may not be added to the header. In this case, NumOfPoint[NumLoD-1] can be calculated using the following (Equation 14).
[0546]
[0547] Thres_Lod[i] indicates the LoD threshold of layer i. The encoding device configures LoDi so that the distance between each point in LoDi is greater than the threshold Thres_Lod[i]. Note that the value of Thres_Lod[NumLoD-1] (i.e., the LoD threshold of the lowest layer) may not be added to the header. In this case, Thres_Lod[NumLoD-1] may be estimated to be 0. This allows the amount of coding for the header to be reduced.
[0548] NumNeiCnt[i] indicates the upper limit of the number of neighboring points used to generate a predicted value of a 3D point belonging to layer i. If the number of neighboring points M is less than NumNeiCnt[i] (i.e., M<NumNeiCnt[i]), the encoding device may calculate a predicted value using M neighboring points. Furthermore, if it is not necessary to change the value of NumNeiCnt[i] for each LoD, the encoding device may add one NumNeiCnt to the header.
[0549] THd[i] indicates the upper limit of the distance of a 3D point used to predict a 3D point to be coded or decoded in layer i. The coding device may not use 3D points whose distance from the 3D point to be coded or decoded is greater than THd[i] for prediction. Note that if it is not necessary to change the value of THd[i] for each LoD, one THd may be added to the header.
[0550] QS[i] denotes the quantization scale of layer i.
[0551] The encoding device may entropy-encode NumLoD, Thres_Lod[i], NumNeiCnt[i], THd[i], or QS[i] and add it to the header. For example, the encoding device may binarize each value and perform arithmetic encoding. Alternatively, the encoding device may perform fixed-length encoding to reduce the amount of processing.
[0552] It should be noted that the encoding device does not necessarily need to add NumLoD, Thres_Lod[i], NumNeiCnt[i], THd[i], or QS[i] to the header, and may be defined, for example, by a profile or level of a standard, etc. This allows the number of bits in the header to be reduced.
[0553] FIG. 67 is an explanatory diagram showing an example of syntax in this embodiment.
[0554] The example syntax shown in Figure 67 shows an example of the structure of information contained in a bitstream generated by an encoding device.
[0555] The syntax shown in Figure 67 includes displacement vector_data, which may include dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k] for each of the 0th to NumLoDth layers (also referred to as the jth layer) of LoD.
[0556] dispd_is_zero[k] is information indicating whether the absolute value of the transformation coefficient of the k-th component of the displacement vector of the i-th 3D point (i.e., vertex[i]) included in the j-th layer of LoD is 0. A value of 1 may indicate that the absolute value of the transformation coefficient of the k-th component is 0, and a value of 0 may indicate that the absolute value of the transformation coefficient of the k-th component is 1 or greater.
[0557] dispd_is_one[k] is information indicating whether the absolute value of the transform coefficient of the k-th component of the displacement vector of the ith 3D point (i.e., vertex[i]) included in the j-th layer of the LoD is 1. A value of 1 may indicate that the absolute value of the prediction residual of the k-th component is 1, and a value of 0 may indicate that the absolute value of the prediction residual of the k-th component is 2 or more.
[0558] If dispd_is_one[k] is not included in the bitstream, the decoding device may estimate its value to be 0. This makes it possible to prevent dispd_is_one[k] from being set to an indefinite value during decoding, and to perform decoding processing appropriately.
[0559] dispd_minus2[k] is information indicating the value obtained by subtracting the value 2 from the absolute value of the transformation coefficient of the kth component of the displacement vector of the i-th 3D point (vertex[i]) included in the j-th layer of LoD.
[0560] If dispd_minus2[k] is not included in the bitstream, the decoding device may estimate its value to be 0. This prevents dispd_minus2[k] from being set to an indefinite value during decoding, allowing appropriate decoding processing.
[0561] dispd_sign[k] indicates the sign bit of the k-th component of the displacement vector of the i-th 3D point (vertex[i]) included in the j-th layer of the LoD. A value of 1 may indicate that the k-th component of the transform coefficient is negative, and a value of 0 may indicate that the k-th component of the transform coefficient is positive.
[0562] For the k-th component, if the displacement vector is represented in a Cartesian coordinate system, the first component may represent the x-component, the second component may represent the y-component, and the third component may represent the z-component. Furthermore, if the displacement vector is represented in a local coordinate system, the first component may represent the normal component, the second component may represent the tangential component, and the third component may represent the binormal component. This allows a common syntax structure to be used regardless of whether the displacement vector is represented in a Cartesian coordinate system or a local coordinate system.
[0563] In addition, the transformation coefficient dispd[k] of the kth component of the displacement vector of the i-th three-dimensional point (i.e., vertex[i]) may be calculated by the calculation process shown in Figure 68 using the above information.
[0564] By introducing the syntax configuration shown in Fig. 67, when encoding a transform coefficient that is likely to have dispd[k] = 0, the encoding device can reduce the frequency of encoding dispd_is_one[k], dispd_minus2[k], and dispd_sign[k] and adding them to the bitstream, which may improve encoding efficiency. Also, when encoding a prediction residual that is likely to have dispd[k] = 1 or 0, the encoding device can reduce the frequency of encoding dispd_minus2[k] and adding it to the bitstream, which may improve encoding efficiency.
[0565] In this embodiment, an example has been shown in which a case where dispd[k] is likely to be 0 or 1 for a transform coefficient has been assumed, but this is not necessarily limited to this, and similar processing may be applied to any dispd[k]. For example, when encoding a transform coefficient that is likely to have dispd[k]=2, dispd_is_two[k] and dispd_minus3[k] may be newly introduced. This makes it possible to reduce the frequency of encoding dispd_minus3[k] and adding it to the bitstream when encoding a transform coefficient that is likely to have dispd[k]=2, and as a result, it is possible to improve coding efficiency. In this case, dispd[k] may be calculated by the calculation processing shown in FIG.
[0566] The encoding device may binarize at least one of dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k], and apply arithmetic coding using a context. For example, since dispd_is_zero[k], dispd_is_one[k], and dispd_sign[k] are each 1 bit, the encoding device may assign one context to each of them, and perform encoding while updating the occurrence probability based on the occurrence frequency of 0 and 1. This may improve encoding efficiency. Alternatively, the encoding device may binarize dispd_minus2[k] using Exponential Golomb, assign a context to each bit, and perform encoding while updating the occurrence probability based on the occurrence frequency of 0 and 1. This may improve encoding efficiency.
[0567] The encoding device may assign different contexts to dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k] for each component of dispd. This may improve encoding efficiency when the dispd value differs for each component. The encoding device may assign the same context to dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k] for each component of dispd. This may improve encoding efficiency when the values of the components of dispd are close to each other.
[0568] The decoding device may convert the decoded quantized transform coefficients from unsigned integers to signed integers in a manner opposite to that used by the encoding device, thereby enabling proper decoding of the generated bitstream without considering the occurrence of negative integers when entropy coding the transform coefficients.
[0569] Note that it is not necessary to convert from unsigned integer values to signed integer values. For example, when decoding a bitstream generated by separately entropy-encoding the sign bits, the decoding device may decode the sign bits. Note that the decoding method of the transform coefficients by the decoding device is not limited to this. For example, the decoding device may arithmetically decode the sign bits indicating the positive or negative sign of the transform coefficients and the binarized data of the absolute values of the transform coefficients, etc., for each bit using a context. This allows the decoding device to appropriately decode a bitstream with improved coding efficiency for the transform coefficients of the displacement vector.
[0570] The decoding device decodes the quantized transform coefficients converted into signed integer values by inverse quantization and reconstruction, and uses them for prediction of the three-dimensional point to be decoded and beyond. Specifically, the decoding device multiplies the quantized transform coefficients by the decoded quantization scale to calculate inverse quantization values, and adds the inverse quantization values and the predicted values to obtain decoded values.
[0571] For example, the decoded unsigned quantized value a2u is converted to a signed value a2q as follows: where ">>" indicates a bit shift operation.
[0572] If the LSB (least significant bit) of a2u is 1: a2q = -((a2u + 1) >> 1) Otherwise: a2q = (a2u >> 1) (Equation 15)
[0573] Also, for example, the decoded unsigned quantized value b2u is converted into a signed value b2q as follows.
[0574] If the LSB of b2u is 1, then b2q = -((b2u + 1) >> 1). Otherwise, b2q = (b2u >> 1). (Equation 16)
[0575] The decoder calculates reconstructed values after inverse quantization, which can be used for prediction of subsequent three-dimensional points to be decoded.
[0576] For example, the decoding device can calculate the inverse quantized value a2iq from the quantized value a2q as follows, and can calculate the inverse quantized value b2iq from the quantized value b2q as follows.
[0577] a2iq=a2q×QS_LoD0 b2iq=b2q×QS_LoD1 (Formula 17)
[0578] Furthermore, the decoding device can calculate the reconstructed value a2rec from the inverse quantized value a2iq as follows, and can calculate the reconstructed value b2rec from the inverse quantized value b2iq as follows.
[0579] a2rec=a2iq+a2p b2rec=b2iq+b2p (Formula 18)
[0580] In the above, an example was shown in which the encoding device calculates the average of the displacement vectors of a certain number of three-dimensional points among the coded and decoded adjacent points of the three-dimensional point to be encoded to generate a predicted value of the displacement vector of a three-dimensional point, but this is not necessarily limited to this, and the predicted value can be generated using other methods.
[0581] For example, the encoding device may use the displacement vector of the closest three-dimensional point among the three-dimensional points that have been coded and decoded adjacent to the three-dimensional point to be coded as the predicted value. The encoding device may also assign a prediction mode value (PredMode) to each three-dimensional point, allowing the prediction value to be selected. For example, the encoding device may set a total number M of prediction modes, assign the average value to prediction mode 0, assign the displacement vector of three-dimensional point A to prediction mode 1, ..., assign the displacement vector of three-dimensional point Z to prediction mode M-1, and add the prediction mode used for prediction to the bitstream for each three-dimensional point. The three-dimensional points A to Z, to which displacement vectors are assigned from prediction mode 1 to prediction mode M-1, may be used in order of proximity to the three-dimensional point to be coded among the three-dimensional points that have been coded and decoded adjacent to the three-dimensional point to be coded.
[0582] Fig. 70 is an explanatory diagram showing an example of predicted value information of a displacement vector in this embodiment. Fig. 71 is an explanatory diagram showing a method of generating a predicted value of a displacement vector in this embodiment.
[0583] 70 shows an example of prediction value information used to predict point b2 when the number N of adjacent three-dimensional points used for prediction is 4 and the number M of prediction modes is 5. The prediction value information includes information indicating, for each of one or more prediction modes, the predicted value to be used in that prediction mode. The example of prediction value information shown in FIG. 70 is a table indicating, for each of one or more prediction modes, the predicted value to be used in that prediction mode.
[0584] In the example shown in Fig. 70, the predicted values used to predict point b2 are, for example, the adjacent three-dimensional points a0, a1, a2, and b1 (see Fig. 71). Correspondingly, the "average value of points a0, a1, a2, and b1" is assigned as the predicted value for prediction mode 0.
[0585] In FIG. 70 , a "point b1" is assigned as a predicted value to prediction mode 1. A "point b2" is assigned as a predicted value to prediction mode 2. A "point a1" is assigned as a predicted value to prediction mode 3. A "point a0" is assigned as a predicted value to prediction mode 4.
[0586] Note that a numerical value that uniquely indicates a prediction mode is also referred to as a prediction mode value. Here, the following description will be given assuming that the prediction mode value of prediction mode m is m. For example, the prediction mode values are used in ascending order of integer values.
[0587] The assignment of prediction mode values may be determined in order of distance from the 3D point to be coded. For example, the coding device can assign a relatively smaller prediction mode value to a 3D point that is closer to the 3D point to be coded. In the above example, the 3D point that is closest to the 3D point b2 to be coded (i.e., closest to the 3D point b2) may be point b1, the 3D point that is next closest to the 3D point b2 may be point a2, the 3D point that is next closest to the 3D point b2 may be point a1, and the 3D point that is next closest to the 3D point b2 may be point a0.
[0588] This allows a smaller prediction mode value to be assigned to a point that is relatively likely to be selected as a prediction value due to a small distance and therefore a relatively small difference between the displacement vector and the prediction value, thereby reducing the number of bits required to encode the prediction mode value. Also, a smaller prediction mode value may be preferentially assigned to a 3D point that belongs to the same LoD as the 3D point to be encoded.
[0589] FIG. 72 shows an example of predicted value information used to predict point a2 when the number N of adjacent three-dimensional points used for prediction is 2 and the number M of prediction modes is 5.
[0590] In the example of predicted value information shown in Fig. 72 , the predicted values used to predict point a2 are, for example, adjacent three-dimensional points a0 and a1. Correspondingly, in Fig. 72 , "the average value of points a0 and a1" is assigned as the predicted value for prediction mode 0.
[0591] Furthermore, "point a1" is assigned as a predicted value to prediction mode 1. "point a0" is assigned as a predicted value to prediction mode 2.
[0592] Note that, if the number of adjacent points is less than four, information indicating that the prediction mode to which a prediction value is not assigned may be set to be unavailable (denoted as "not available" in the figure).
[0593] An example of predicted value information when the displacement vector is represented in an orthogonal coordinate system (XYZ coordinate system) is shown in FIG.
[0594] In the example shown in FIG. 73, the values used to predict point b2 are, for example, the adjacent three-dimensional points a0, a1, a2, and b1 (see FIG. 71). Correspondingly, in FIG. 73, the coordinates (Xave, Yave, Zave) of the "average value of points a0, a1, a2, and b1" are assigned as the predicted value of prediction mode 0. Here, Xave can be calculated as the average or weighted average of Xb1, Xa2, Xa1, and Xa0. Yave can be calculated as the average or weighted average of Yb1, Yb2, Ya1, and Ya0. Zave can be calculated as the average or weighted average of Zb1, Zb2, Za1, and Za0.
[0595] Furthermore, the coordinates (Xb1, Yb1, Zb1) of "point b1" are assigned as a predicted value for prediction mode 1. The coordinates (Xa2, Ya2, Za2) of "point b2" are assigned as a predicted value for prediction mode 2. The coordinates (Xa1, Ya1, Za1) of "point a1" are assigned as a predicted value for prediction mode 3. The coordinates (Xa0, Ya0, Za0) of "point a0" are assigned as a predicted value for prediction mode 4.
[0596] For example, the encoding device may select prediction mode 2 (i.e., prediction mode value 2) and encode the X, Y, and Z components of the displacement vector of the three-dimensional point to be encoded using predicted values Xa2, Ya2, and Za2, respectively. In this case, the encoding device adds the prediction mode value 2 to the bitstream.
[0597] In the above example, the displacement vector is in a Cartesian coordinate system, but this is not necessarily limited to this, and the invention may be applied to a displacement vector expressed in a local coordinate system, for example.
[0598] The number of prediction modes M may be added to the bitstream. Alternatively, the number of prediction modes M may not be added to the bitstream, but may be defined by a standard profile or level. Alternatively, the number of prediction modes M may be a value calculated from the number of three-dimensional points N used for prediction (for example, M=N+1).
[0599] When an encoding device generates a predicted value of a displacement vector of a three-dimensional point by adding a prediction mode value (PredMode) to each three-dimensional point, an example of a method for assigning a predicted value to each prediction mode has been shown in which the displacement vector of an adjacent point is assigned as a predicted value using distance information from the three-dimensional point to be encoded, but this is not necessarily limited to this, and the method for assigning a predicted value to a prediction mode may be changed in some way.
[0600] For example, the encoding device may calculate a median from the predicted values assigned to each prediction mode and assign the calculated median to prediction mode 0. In this way, the encoding device may assign the median as a predicted value to a prediction mode having a small prediction mode value. This allows the encoding device to generate predicted value candidates that prioritize the median of the displacement vectors of adjacent points, thereby improving encoding efficiency.
[0601] The change in allocation of predicted values using the median will be described with reference to FIGS.
[0602] Fig. 74 is an explanatory diagram showing an example of predicted value information of a displacement vector in an embodiment. Fig. 75 is an explanatory diagram showing a method of generating a predicted value of a displacement vector in this embodiment. Fig. 76 is an explanatory diagram showing an example of predicted value information of a displacement vector in this embodiment.
[0603] In the example shown in Fig. 75, the number of 3D points used for prediction N = 4, and the number of prediction modes M = 4. Point a2 is predicted from points a0 and a1. Point b2 is predicted from points a0, a1, a2, b0, and b1.
[0604] Here, the order of distance to the 3D point to be coded is assumed to be point b1, point a2, point a1, and point a0, and the displacement vector of the 3D point with the closest distance is assigned to the prediction mode with the smallest prediction mode value. The magnitudes of the prediction values are assumed to be b1 > a1 > a0 > a2.
[0605] The encoding device calculates the median of the predicted values of the prediction modes. For example, the encoding device can sort the n predicted values assigned to the prediction modes in ascending or descending order, and determine the (n / 2)th value as the median. Note that the calculation method of the median may be switched depending on whether the value of n is an odd number or an even number.
[0606] For example, when n is an odd number, the encoding device may set the (n / 2)th predicted value (where the decimal point is discarded) of the 0th to n-1th predicted values after sorting as the median. Also, when n is an even number, the encoding device may set the (n / 2-1)th predicted value and the n / 2th predicted value of the 0th to n-1th predicted values after sorting as median candidates A and B, and use either A or B as the median in some way. For example, the one of A and B that is closest to the 3D point to be encoded may be set as the median.
[0607] In the example shown in Figure 75, n = 4, so the median can be calculated using the method for calculating the median when n is an even number. For example, if b1, a1, a0, and a2 are sorted in ascending order, the result will be a2, a0, a1, b1. In this case, the (n / 2-1)th predicted value is a0, and the n / 2th is a1, which are set as median candidates A and B. Then, since a1 is closer to the 3D point to be encoded than a0, a1 is selected as the median.
[0608] In this case, as shown in Fig. 76 , the encoding device assigns predicted value a1 selected as the median value to prediction mode 0, and assigns predicted value b1, which was originally assigned to prediction mode 0, to prediction mode 2, which was originally assigned to predicted value a1. In other words, the encoding device swaps the predicted values of prediction mode 0 and prediction mode 2. This allows the encoding device to generate predicted value candidates that prioritize the median value of the displacement vectors of adjacent points, thereby improving encoding efficiency.
[0609] Although the example in which the median is used as a method for assigning predicted values to prediction modes has been described, this is not necessarily limited to this. For example, the encoding device may calculate an average value from the predicted values assigned to each prediction mode and assign a predicted value close to the average value to prediction mode 0. This makes it possible to generate predicted value candidates that prioritize displacement vectors close to the average of displacement vectors of adjacent points, thereby improving encoding efficiency.
[0610] In addition, the encoding device may first calculate the median from the displacement vectors of adjacent points and assign it to prediction mode 0, and then assign the displacement vectors of surrounding three-dimensional points other than the median to prediction modes from prediction mode 1 onwards using the distance information of those three-dimensional points.
[0611] The encoding device may also add information indicating whether to prioritize the median (also referred to as median priority information) to a header or the like. If the median priority information indicates that the median is prioritized, the encoding device may assign the median to prediction mode 0 using the above method, and if not, assign a predicted value to the prediction mode regardless of the median. This allows the encoding device to adaptively switch between cases where the median is prioritized and cases where it is not, potentially improving encoding efficiency. Furthermore, the decoding device can appropriately decode the bitstream based on the median priority information added to the header or the like.
[0612] As an example of changing the allocation of predicted values with priority given to the median, an example has been given in which the median is assigned to prediction mode 0 and the predicted value originally assigned to prediction mode 0 is replaced with the prediction mode to which the median was assigned, but this is not necessarily limited to this.
[0613] FIG. 77 is an explanatory diagram showing an example of predicted value information of a displacement vector in this embodiment.
[0614] For example, as shown in Fig. 77 , the encoding device may shift the predicted values assigned to each prediction mode until a value is reassigned to the prediction mode to which the median value was originally assigned, such as assigning the median to prediction mode 0, assigning the predicted value originally assigned to prediction mode 0 to prediction mode 1, and assigning the predicted value originally assigned to prediction mode 1 to prediction mode 2. This makes it possible to generate predicted value information that prioritizes the median of displacement vectors of adjacent points and also prioritizes candidates for predicted values that are close in distance, thereby improving encoding efficiency.
[0615] Furthermore, although an example of allocating predicted values with priority given to the median or mean value has been shown, the present invention is not necessarily limited to this.
[0616] 78 and 79 are explanatory diagrams showing examples of predicted value information of displacement vectors in this embodiment.
[0617] For example, the encoding device calculates statistical information of predicted values of prediction modes as shown in Figure 78. The statistical information may be, for example, the median, mean, variance, or standard deviation of neighboring points. Then, the encoding device can change the allocation of predicted values based on the calculated statistical information (see Figure 79).
[0618] Modified examples of setting the predicted value of the 3D point a to be coded in the coding frame will be described with reference to FIGS.
[0619] Fig. 80 is an explanatory diagram showing an example of points to be coded in this embodiment. Fig. 81 is an explanatory diagram showing an example of predicted value information of a displacement vector in this embodiment. Fig. 82 is an explanatory diagram showing an example of temporal dv in this embodiment. Fig. 83 is an explanatory diagram showing an example of predicted value information of a displacement vector in this embodiment.
[0620] The encoding device may set a predicted value of a 3D point a (see FIG. 80) to be encoded in a current frame to be encoded as predicted value information shown in FIG. 81. Specifically, the encoding device may set 0 (no prediction) as the predicted value for prediction mode 0, and may set the average value of the displacement vectors of adjacent points a, b, and c (dv0, dv1, and dv2, respectively) as the predicted value for prediction mode 1. Furthermore, the encoding device may set the displacement vectors of adjacent points a, b, and c (dv0, dv1, and dv2, respectively) as the predicted values for prediction modes 2, 3, and 4, respectively.
[0621] Note that the predicted values assigned to each prediction mode are not limited to these, and other predicted values may be assigned.
[0622] Furthermore, the encoding device may assign, for example, a displacement vector in a reference frame (see FIG. 82 ) different from the encoding target frame to the predicted value. Specifically, the encoding device may use, as the predicted value of the encoding target 3D point a, the displacement vector (hereinafter referred to as temporal dv) of the corresponding point a′ of the encoding target 3D point a in the reference frame that has already been coded or decoded. When the target object is moving at a constant motion, the value of the displacement vector of the encoding target 3D point tends to be relatively close to the value of the displacement vector of the corresponding point of the encoding target 3D point in the reference frame. Therefore, by adding the temporal dv as a predicted value to the prediction candidates, it is possible to improve encoding efficiency.
[0623] An example of prediction value information in the case where a prediction value of prediction mode 5 is added as temporal dv is shown in Fig. 83. Note that temporal dv may be added as a prediction value of another prediction mode (that is, any of prediction modes 0 to 4). Also, the prediction value of any prediction mode in the prediction value information shown in Fig. 81 may be changed to temporal dv.
[0624] The encoding device may calculate the temporal dv for each prediction unit (Displacement Vector Group, DVG) of the displacement vector in the reference frame, store the calculated temporal dv in memory, and use the temporal dv of the DVG to which the corresponding point a' belongs as the temporal dv of the corresponding point a'. This reduces the amount of memory required.
[0625] The encoding device may calculate the temporal dv of the DVG from the displacement vectors of the 3D points belonging to the DVG. For example, the average value of the displacement vectors of the 3D points belonging to the DVG may be used as the temporal dv of the DVG. This may reduce the amount of memory required to store the temporal dv, while adding the temporal dv to prediction candidates, thereby potentially improving encoding efficiency.
[0626] Alternatively, for example, the encoding device may calculate a global displacement vector (hereinafter referred to as a global dv) of the current frame to be encoded and add the global dv to prediction candidates to obtain a predicted value. The encoding device may calculate the global dv from, for example, the average value of the displacement vectors in the current frame to be encoded or the reference frame. The encoding device may also add the calculated global dv to the bitstream. This allows the decoding device to decode the global dv added by the encoding device as a prediction candidate from the bitstream and add the same global dv as the encoding device to the prediction candidates.
[0627] Alternatively, the encoding device may select at least two or more displacement vectors from the displacement vectors added to the prediction candidates as a new predicted value, and add an average value of the selected two or more displacement vectors to the prediction candidates as a new predicted value, which may improve encoding efficiency.
[0628] Furthermore, the encoding device may store one or more displacement vectors used in the past in a memory as new predicted values, and add at least one of the displacement vectors to the prediction candidates as a new predicted value. This may improve encoding efficiency. The encoding device may periodically or irregularly store the displacement vectors used for encoding or decoding in a memory (i.e., the memory storing the one or more displacement vectors used in the past), and delete old displacement vectors that have been stored for a certain period of time or more from the memory. In this way, by updating the displacement vectors stored in the memory, the encoding device can assign new displacement vectors to prediction candidates, which may improve encoding efficiency.
[0629] When the encoding device encodes the displacement vector of a three-dimensional point, a DVG, which is a prediction unit, may be provided according to the encoding or decoding order, and encoding or decoding may be performed for each DVG. For example, it is possible to define the number of three-dimensional points included in the DVG (DVG Size), and divide the three-dimensional points into multiple DVGs according to the encoding or decoding order, and then encode or decode them. Note that the encoding or decoding order of the displacement vector of the three-dimensional point may be any order. For example, it may be possible to generate an LoD and encode or decode in order for each LoD layer. Alternatively, it may be possible to encode or decode the displacement vector in the encoding or decoding order of the position information (vertex) of the three-dimensional point without generating an LoD. Alternatively, it may be possible to generate a Morton code using the position information of the three-dimensional point, and encode or decode in the order of the Morton code.
[0630] An example of a DVG definition is described below.
[0631] FIG. 84 is an explanatory diagram showing an example of a reference destination of a DVG in this embodiment.
[0632] In the example of the DVG reference destination shown in Fig. 84 (also referred to as the first example), 3D points within the same DVG are defined as not being able to be referenced. For example, 3D points within the same DVG may not be able to be added to adjacent points.
[0633] Also, a 3D point in a different DVG that has been encoded or decoded may be defined as being referable. For example, a 3D point in a different DVG that has been encoded or decoded may not be added to an adjacent point.
[0634] The size of the DVG may also be written in the header or the like (see FIG. 85). For example, if the size of the DVG (DVGSize) is 16, information "DVGSize=16" may be added to the header. Alternatively, DVGSize may be set to 2^n, and the value of n may be added to the header.
[0635] Also, 3D points within the same DVG can be encoded or decoded in parallel.
[0636] FIG. 85 is an explanatory diagram showing an example of syntax in this embodiment.
[0637] The example syntax shown in Figure 85 shows an example of the structure of information contained in a bitstream generated by an encoding device.
[0638] The syntax shown in Figure 85 includes a displacement vector_header, which includes a DVGSize.
[0639] DVGSize indicates the number of 3D points contained in the DVG.
[0640] FIG. 86 is an explanatory diagram showing an example of a reference destination of a DVG in this embodiment.
[0641] In the example of a DVG reference destination shown in Figure 86 (also referred to as the second example), encoded or decoded 3D points within the same DVG are defined as referable. 3D points that have not been encoded or decoded are defined as not referable. For example, encoded or decoded 3D points within the same DVG may be allowed to be added to adjacent points, and 3D points that have not been encoded or decoded may not be added.
[0642] Also, a 3D point in a different encoded or decoded DVG may be defined as referable, for example, a 3D point in a different encoded or decoded DVG may be allowed to be added to an adjacent point.
[0643] The size of the DVG may also be written in the header or the like (see FIG. 85). For example, if the size of the DVG (DVGSize) is 16, information "DVGSize=16" may be added to the header. Alternatively, DVGSize may be set to 2^n, and the value of n may be added to the header.
[0644] In this way, even for three-dimensional points within the same DVG, three-dimensional points that have already been coded and decoded can be referenced, thereby improving prediction accuracy and encoding efficiency.
[0645] FIG. 87 is an explanatory diagram showing an example of a reference destination of a DVG in this embodiment.
[0646] In the example of a DVG reference destination shown in Figure 87 (also referred to as the third example), encoded or decoded 3D points within the same DVG are defined as referable. 3D points that are not encoded or decoded are defined as non-referable. For example, encoded or decoded 3D points within the same DVG may be allowed to be added to adjacent points, and non-encoded or non-decoded 3D points may not be allowed to be added to adjacent points.
[0647] Also, 3D points in different DVGs are defined as not being able to be referenced, for example, 3D points in different DVGs may not be able to be added to adjacent points.
[0648] The size of the DVG may also be written in the header or the like (see FIG. 85). For example, if the size of the DVG (DVGSize) is 16, information "DVGSize=16" may be added to the header. Alternatively, DVGSize may be set to 2^n, and the value of n may be added to the header.
[0649] In this way, by prohibiting reference between DVGs, dependency between DVGs is eliminated, and multiple DVGs can be encoded or decoded in parallel.
[0650] Furthermore, by making it possible to refer to an encoded or decoded three-dimensional point using a three-dimensional point within the same DVG, it is possible to improve prediction accuracy and improve coding efficiency.
[0651] In the description with reference to FIG. 84, when encoding a displacement vector of a 3D point, an example was shown in which a DVG is provided according to the encoding or decoding order, and encoding or decoding is performed for each DVG. For example, an example was shown in which the number of 3D points included in a DVG (DVGSize) is defined, and the 3D points are divided into multiple DVGs according to the encoding or decoding order and encoded or decoded. Here, the prediction mode PredMode for encoding the displacement vector, or information disp_dvg_inter_mode indicating whether inter-prediction is applied, may be set for each DVG. In this case, 3D points included in the same DVG will share PredMode or disp_dvg_inter_mode, and the same value may be set. This can improve encoding efficiency by reducing the amount of code for PredMode or disp_dvg_inter_mode. Note that it is not necessarily required to set PredMode or disp_dvg_inter_mode for each DVG, but it is also possible to set PredMode or disp_dvg_inter_mode for each set of other three-dimensional points.
[0652] FIG. 88 is an explanatory diagram showing an example of a reference destination of a DVG in this embodiment.
[0653] In the example of a DVG reference destination shown in Figure 88 (also referred to as the fourth example), encoded or decoded 3D points within the same DVG are defined as referable. 3D points that have not been encoded or decoded are defined as not referable. For example, encoded or decoded 3D points within the same DVG may be allowed to be added to adjacent points, and 3D points that have not been encoded or decoded may not be added.
[0654] Also, a 3D point in a different encoded or decoded DVG is defined as being referable, for example, a 3D point in a different encoded or decoded DVG may be added to an adjacent point.
[0655] The encoding device may also add PredMode or disp_dvg_inter_mode to each DVG and predictively encode three-dimensional points within the DVG using the same PredMode or dips_dvg_inter_mode.
[0656] Furthermore, the encoding device may determine whether to add PredMode or disp_dvg_inter_mode for each DVG. For example, the encoding device may calculate PredMode or disp_dvg_inter_mode of the DVG to which the 3D point to be encoded belongs using a variance value of the displacement vectors of decoded 3D points in different DVGs. Furthermore, if the calculated variance value is equal to or greater than a threshold, PredMode or disp_dvg_inter_mode may be added to the DVG, and if not, PredMode or disp_dvg_inter_mode may not be added. If PredMode or disp_dvg_inter_mode is not added, it may be assumed that PredMode=0 or disp_dvg_inter_mode=0.
[0657] The size of the DVG may also be written in the header or the like (see FIG. 85). For example, if the size of the DVG (DVGSize) is 16, information "DVGSize=16" may be added to the header. Alternatively, DVGSize may be set to 2^n, and the value of n may be added to the header.
[0658] In this way, by allowing reference to three-dimensional points that have already been coded or decoded, even for three-dimensional points within the same DVG, prediction accuracy can be improved, and coding efficiency can be improved.
[0659] Furthermore, by adding PredMode or dips_dvg_inter_mode to each DVG, it is possible to reduce overhead and improve coding efficiency compared to adding PredMode or dips_dvg_inter_mode to each 3D point.
[0660] FIG. 89 is an explanatory diagram showing an example of syntax in this embodiment.
[0661] The example syntax shown in Figure 89 shows an example of the structure of information contained in a bitstream generated by an encoding device.
[0662] The syntax shown in Figure 89 includes a displacement vector_header, which includes a DVGSize.
[0663] DVGSize indicates the unit in which the displacement vector of a 3D point is predicted. PredMode or disp_dvg_inter_mode is added to each of the DVGSize 3D points, and 3D points in the same DVG are encoded and decoded using the same PredMode or disp_dvg_inter_mode.
[0664] When the displacement vectors are coded by dividing them into LoD layers, a different DVGSize may be set for each LoD layer. In this case, the DVGSize for each LoD layer may be added to the header. This allows the decoding device to correctly decode a bitstream generated by setting the DVGSize for each LoD layer.
[0665] For example, when a displacement vector is encoded using an LoD hierarchical layer, high-frequency components are collected in a lower LoD hierarchical layer by a lifting transform, and the values of the transform coefficients are likely to become smaller. Therefore, the lower the LoD hierarchical layer, the more likely it is that the accuracy of inter prediction will increase. Therefore, by increasing the value of DVGSize in a lower LoD hierarchical layer and sharing dips_dvg_inter_mode between many three-dimensional points, the amount of code required to encode disp_dvg_inter_mode can be reduced, potentially improving encoding efficiency. On the other hand, low-frequency components are collected in the above LoD hierarchical layer by a lifting transform, and the values of the transform coefficients are likely to become larger. Therefore, the lower the LoD hierarchical layer, the more likely it is that the accuracy of inter prediction will decrease. Therefore, by reducing the value of DVGSize in a lower LoD hierarchical layer and allowing detailed setting of whether or not to perform inter prediction, there is a possibility that encoding efficiency can be improved.
[0666] FIG. 90 is an explanatory diagram showing an example of syntax in this embodiment.
[0667] The example syntax shown in Figure 90 shows an example of the structure of information contained in a bitstream generated by an encoding device.
[0668] The syntax shown in Figure 90 includes displacement_vector_data, which may include PredMode, disp_dvg_inter_mode, dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k] for each of the 0th to NumLoDth layers of LoD (also referred to as the jth layer).
[0669] PredMode is information indicating a prediction mode for encoding or decoding a displacement vector of the i-th three-dimensional point of the j-th layer. PredMode takes a value from 0 to M-1 (M is the total number of prediction modes). If PredMode is not included in the bitstream (if the condition of the if statement "maxdiff>=Thfix[i] && NumPredMode[i]>1" is not satisfied), PredMode may be estimated to be 0. Note that the estimated value is not limited to 0, and any value from 0 to M-1 may be used. In addition, an estimated value when PredMode is not included in the bitstream may be added to a separate header or the like. Furthermore, PredMode may be binarized with a truncated unary code using the number of prediction modes to which the predicted value is assigned, and then arithmetically coded.
[0670] disp_dvg_inter_mode is information indicating whether the i-th displacement vector of the j-th LoD layer is encoded or decoded using inter prediction. A value of 1 indicates that inter prediction is applied, and a value of 0 indicates that inter prediction is not applied.
[0671] The dispd_is_zero[k], dispd_is_one[k], dispd_minus2[k], and dispd_sign[k] are the same as the information in FIG.
[0672] An example of the encoding process in this embodiment will be described below.
[0673] FIG. 91 is a flowchart showing an example of the encoding process in this embodiment.
[0674] The encoding process shown in FIG. 91 is a method for encoding displacement data of three-dimensional points, which is executed by an encoding device.
[0675] In step S9101, the encoding device generates a predicted value of the first displacement data by performing inter prediction using second displacement data at a different time from the first displacement data, which is the displacement data to be encoded.
[0676] In step S9102, the encoding device generates a prediction residual using the first displacement data and the generated prediction value.
[0677] In step S9103, the encoding device encodes the generated prediction residual.
[0678] This allows the encoding device to appropriately encode the displacement data of the 3D point by encoding the prediction residual generated by inter-prediction. By using inter-prediction, it is possible to reduce the amount of code when the difference between the first displacement data to be encoded and the second displacement data at a different time is relatively small, which may improve the encoding process. In this way, the above encoding method can contribute to improving the encoding process related to the displacement vector.
[0679] For example, when generating a predicted value of the first displacement data, it is possible to determine whether or not to perform inter-prediction when generating the predicted value of the first displacement data, and if it is determined that inter-prediction is to be performed, the predicted value of the first displacement data is generated by performing inter-prediction, and if it is determined that inter-prediction is not to be performed, the predicted value of the first displacement data is generated without performing inter-prediction.
[0680] As a result, when encoding the first displacement data to be encoded, the encoding device determines in advance whether to use inter-prediction to generate a predicted value of the displacement data, and can switch whether to use inter-prediction to generate a predicted value of the displacement data depending on the determination result.As a result, for example, if using inter-prediction can improve the encoding process, inter-prediction can be used in the encoding process, and if using inter-prediction cannot improve the encoding process (or deteriorates the encoding process), inter-prediction can be not used in the encoding process.This may further reduce the amount of code.In this way, the above encoding method can contribute to further improvement of encoding processes related to displacement vectors, etc.
[0681] For example, when determining whether to perform inter-prediction, a first sum, which is the sum of transformation coefficients for the first displacement data, and a second sum, which is the sum of prediction residuals when inter-prediction is applied to the transformation coefficients, may be calculated, and if it is determined that the second sum is smaller than the first sum, it may be determined that inter-prediction is to be performed, and if it is determined that the second sum is not smaller than the first sum, it may be determined that inter-prediction is not to be performed.
[0682] As a result, the encoding device can switch whether to use inter prediction to generate a predicted value of the displacement data by comparing the sum of transform coefficients when inter prediction is applied to the transform coefficients of the first displacement data to be encoded with the above transform coefficients (in other words, the sum of transform coefficients when inter prediction is not used). Specifically, it can determine to use inter prediction when it is determined that the sum of transform coefficients when inter prediction is applied is small. This may make it possible to reduce the amount of code by making a simpler decision. In this way, the above encoding method can contribute to improving encoding processes and the like related to displacement vectors.
[0683] For example, information indicating whether inter prediction was used to generate the prediction residual of the first displacement data may also be transmitted.
[0684] As a result, the encoding device can transmit information indicating whether inter prediction was used during encoding, thereby notifying a decoding device that receives and decodes the encoded displacement data of whether inter prediction was used during encoding. This can contribute to appropriate decoding of encoded data. Specifically, this can contribute to ensuring that data encoded using inter prediction is decoded using inter prediction, and data encoded without using inter prediction is decoded without using inter prediction. In this way, the above encoding method can contribute to improving encoding processes and the like related to displacement vectors.
[0685] For example, the three-dimensional points include a plurality of three-dimensional points, each of which belongs to one of a plurality of hierarchies, and when generating a predicted value of the first displacement data, for each of one or more hierarchies to which the plurality of three-dimensional points belong, it is determined whether or not to perform inter-prediction when generating a predicted value of the first displacement data for the three-dimensional points among the plurality of three-dimensional points that belong to that hierarchical level, and for a hierarchical level among the one or more hierarchies for which it is determined that inter-prediction should be performed, the predicted value of the first displacement data may be generated by performing inter-prediction, and for a hierarchical level among the one or more hierarchies for which it is determined that inter-prediction should not be performed, the predicted value of the first displacement data may be generated without performing inter-prediction.
[0686] As a result, when encoding the first displacement data to be encoded, the encoding device determines in advance whether to use inter-prediction to generate a predicted value of the displacement data for each layer to which the three-dimensional point belongs, and can switch whether to use inter-prediction to generate a predicted value of the displacement data for each layer to which the three-dimensional point belongs according to the determination result. As a result, for example, inter-prediction can be used in the encoding process of a layer where using inter-prediction can improve the encoding process, and inter-prediction can be avoided in the encoding process of a layer where using inter-prediction cannot improve (or deteriorates) the encoding process. This may further reduce the amount of code. In this way, the above encoding method can contribute to further improvement of encoding processes related to displacement vectors, etc.
[0687] For example, information indicating a layer in which inter prediction was used to generate the prediction residual of the first displacement data may be further transmitted.
[0688] As a result, the encoding device can transmit information indicating whether inter prediction was used for each layer to which the 3D points belong during encoding, thereby informing a decoding device that receives and decodes the encoded displacement data of whether inter prediction was used for each layer to which the 3D points belong during encoding. This can contribute to appropriate decoding of encoded data. Specifically, this can contribute to ensuring that data of a layer encoded using inter prediction is decoded using inter prediction, and that data of a layer encoded without using inter prediction is decoded without using inter prediction. In this way, the above encoding method can contribute to improving encoding processes and the like related to displacement vectors.
[0689] An example of the decoding process in this embodiment will be described below.
[0690] FIG. 92 is a flow diagram showing an example of the decoding process in this embodiment.
[0691] The decoding process shown in FIG. 92 is a method for decoding displacement data of three-dimensional points, which is executed by a decoding device.
[0692] In step S9201, the decoding device obtains a prediction residual by decoding the encoded data.
[0693] In step S9202, the decoding device generates a predicted value of the first displacement data by performing inter prediction using second displacement data at a time different from that of the first displacement data, which is the displacement data to be decoded.
[0694] In step S9203, the decoding device generates first displacement data using the prediction residual and the generated prediction value.
[0695] This allows the decoding device to properly decode the displacement data of the 3D point by decoding the prediction residual generated by inter prediction. By using inter prediction, when the difference between the first displacement data to be decoded and the second displacement data at a different time is relatively small, it is possible to reduce the amount of code, which may improve the decoding process. In this way, the above decoding method can contribute to improving the decoding process related to the displacement vector.
[0696] For example, when generating a predicted value of the first displacement data, it is possible to determine whether or not to perform inter-prediction when generating the predicted value of the first displacement data, and if it is determined that inter-prediction is to be performed, the predicted value of the first displacement data is generated by performing inter-prediction, and if it is determined that inter-prediction is not to be performed, the predicted value of the first displacement data is generated without performing inter-prediction.
[0697]
[0013] As a result, when decoding the first displacement data to be decoded, the decoding device determines in advance whether to use inter prediction to generate a predicted value of the displacement data, and can switch whether to use inter prediction to generate a predicted value of the displacement data depending on the determination result. This may further reduce the amount of code. In this way, the above decoding method can contribute to further improvement of decoding processes related to displacement vectors, etc.
[0698] For example, information indicating whether inter prediction was used to generate the prediction residual of the first displacement data may also be received.
[0699] As a result, the decoding device can know whether the encoding device used inter prediction during encoding by receiving information indicating whether inter prediction was used during encoding. If the encoding device used inter prediction during encoding, the decoding device can decode the displacement vector using inter prediction, and if the encoding device did not use inter prediction during encoding, the decoding device can decode the displacement vector without using inter prediction. As a result, the decoding device can properly decode the data encoded by the encoding device.
[0700] For example, the three-dimensional points include a plurality of three-dimensional points, each of which belongs to one of a plurality of hierarchies, and when generating a predicted value of the first displacement data, for each of one or more hierarchies to which the plurality of three-dimensional points belong, it is determined whether or not to perform inter-prediction when generating a predicted value of the first displacement data for the three-dimensional points among the plurality of three-dimensional points that belong to that hierarchical level, and for a hierarchical level among the one or more hierarchies for which it is determined that inter-prediction should be performed, the predicted value of the first displacement data may be generated by performing inter-prediction, and for a hierarchical level among the one or more hierarchies for which it is determined that inter-prediction should not be performed, the predicted value of the first displacement data may be generated without performing inter-prediction.
[0701] As a result, when decoding the first displacement data to be encoded, the decoding device determines in advance whether to use inter-prediction to generate a predicted value of the displacement data for each layer to which the three-dimensional point belongs, and can switch whether to use inter-prediction to generate a predicted value of the displacement data for each layer to which the three-dimensional point belongs depending on the determination result. This may further reduce the amount of code. In this way, the above decoding method can contribute to further improvement of encoding processes related to displacement vectors, etc.
[0702] For example, the information indicating the layer in which inter prediction was used to generate the prediction residual of the first displacement data may also be received.
[0703] As a result, the decoding device receives information indicating whether inter prediction was used for each layer to which the 3D points belong during encoding, thereby enabling the encoding device to know whether inter prediction was used for each layer to which the 3D points belong during encoding.The decoding device can then decode the displacement vector using inter prediction for data of a layer that the encoding device used inter prediction for during encoding, and can decode the displacement vector without using inter prediction for data of a layer that the encoding device did not use inter prediction for during encoding.This allows the decoding device to properly decode the data encoded by the encoding device.
[0704] An example of the encoding process in this embodiment will be described below.
[0705] FIG. 93 is an explanatory diagram showing an example of a frame structure in this embodiment.
[0706] An example of operation will be described below in which, when encoding or decoding a displacement vector by inter-prediction, the reference frame and the frame to be encoded or decoded have different frame structures.
[0707] Here, the difference in structure includes, for example, a difference in the LoD hierarchical structure between the reference frame and the frame to be encoded or decoded. Here, the difference in the LoD hierarchical structure includes, for example, a difference in the number of LoD hierarchies (in other words, the number of LoD hierarchies), or a difference in the number of 3D points belonging to the LoD hierarchies in one or more LoD hierarchies.
[0708] In this way, when the frame structure differs between the reference frame and the frame to be coded or decoded, inter prediction of the frame to be coded or decoded may be prohibited. Specifically, when the frame structure differs between the reference frame and the frame to be coded or decoded, inter prediction may be prohibited by limiting the value of disp_frame_inter_mode or disp_lod_inter_mode to 0 in a standard or the like. This makes it possible to prevent the application of inter prediction when there is no reference 3D point in the reference frame for the 3D point to be coded or decoded due to the frame structure differing between the reference frame and the frame to be coded or decoded, thereby stabilizing coding efficiency.
[0709] In addition, if the frame structures of the reference frame and the frame to be coded or decoded are different, and the decoding device receives a bitstream in which the value of disp_frame_inter_mode or disp_lod_inter_mode is set to 1, an error may be output as non-compliance with the standard.
[0710] Note that when the reference frame and the frame to be coded or decoded have different frame structures, if the decoding device receives a bitstream in which the value of disp_frame_inter_mode or disp_lod_inter_mode is set to 1, the decoding device may perform inter prediction with a prediction value of 0. Alternatively, the decoding device may not perform inter prediction regardless of the value of disp_frame_inter_mode or disp_lod_inter_mode.
[0711] A reference frame and a target frame are shown in Fig. 93. The target frame is a frame to be coded or decoded, and is coded or decoded by referring to the reference frame.
[0712] The reference frame has LoD 0 and LoD 1. The number of LoD layers of the reference frame, NumLoD_Ref, is 2.
[0713] The target frame has LoD0, LoD1, and LoD2. The number of LoD layers NumLoD_Cur of the target frame is 3. The target frame includes a displacement vector_header. The displacement vector_header includes disp_frame_inter_mode.
[0714] The LoD0 and LoD1 of the target frame are inter-predictable because the LoDs (i.e., LoD0 and LoD1) of the same layer exist in the reference frame. The LoD2 of the target frame is not inter-predictable because the LoD (i.e., LoD2) of the same layer does not exist in the reference frame.
[0715] The number of LoDs, NumLoD, may be calculated by adding 1 to the number of subdivision iterations from the base mesh.
[0716] For example, when the number of subdivisions is 1, LoD0 and LoD1 are generated, and NumLoD is calculated as 2.
[0717] For example, when the number of subdivisions is 2, LoD0, LoD1, and LoD2 are generated, and NumLoD is calculated as 3.
[0718] Instead of NumLoD, the encoding device may add the number of subdivisions to the header, and the decoding device may calculate NumLoD using the number of subdivisions decoded from the header. This may reduce the amount of bits by encoding a value that is 1 smaller than NumLoD.
[0719] In a frame to be encoded or decoded, when the number of LoD layers NumLoD_Cur of the frame to be encoded or decoded is greater than the number of LoD layers NumLoD_Ref of the reference frame, inter prediction may be prohibited by limiting the value of disp_frame_inter_mode to 0 in a standard, etc. This makes it possible to prevent the application of inter prediction when there is no reference 3D point in the reference frame that corresponds to the encoded or decoded 3D point due to differences in frame structure between the reference frame and the frame to be encoded or decoded, thereby stabilizing encoding efficiency.
[0720] In addition, if the number of LoD layers of the target frame, NumLoD_Cur, is greater than the number of LoD layers of the reference frame, NumLoD_Ref, and the decoding device receives a bitstream with disp_frame_inter_mode set to 1, it may determine that the bitstream does not comply with the standard and output an error.
[0721] Note that if the number of LoD layers NumLoD_Cur of the target frame is larger than the number of LoD layers NumLoD_Ref of the reference frame, and the decoding device receives a bitstream in which disp_frame_inter_mode is set to 1, the decoding device may set the prediction value to 0 and not perform inter prediction. Alternatively, the decoding device may not perform inter prediction regardless of the value of disp_frame_inter_mode.
[0722] In addition, when the number of layers NumLoD_Cur of the LoD of the target frame is greater than the number of layers NumLoD_Ref of the LoD of the reference frame, inter prediction may be applied to the LoD in which a reference 3D point exists in the reference frame, and inter prediction may not be applied to the LoD in which a reference 3D point does not exist in the reference frame. This allows inter prediction to be applied to at least the LoD in which a reference 3D point exists, thereby improving coding efficiency. In this case, the value of disp_frame_inter_mode may be set to 1. This allows the decoder side to apply the same process to correctly decode the bitstream.
[0723] Fig. 94 is an explanatory diagram showing an example of a frame structure in this embodiment. Fig. 95 is an explanatory diagram showing an example of a process for determining whether or not to apply inter prediction in this embodiment.
[0724] Prohibition of inter prediction will be described with reference to FIGS. 94 and 95.
[0725] When i≧NumLoD_Ref in the i-th LoD of the frame to be coded or decoded, inter prediction may be prohibited by limiting the value of disp_lod_inter_mode[i] to 0 according to standards or the like.
[0726] For example, when the number of LoD layers NumLoD_Ref of the reference frame shown in FIG. 94 is 2, disp_lod_inter_mode[0] and disp_lod_inter_mode[1] may be set to 1, and disp_lod_inter_mode[2] may be set to 0.
[0727] Then, for the LoD of the i-th frame to be encoded or decoded (where i is an integer that varies from 0 to a number one less than the number of LoDs of the target frame), if it is determined that disp_lod_inter_mode[i] is 1 and i is smaller than the number of LoDs of the reference frame, inter prediction is applied; otherwise, inter prediction may not be applied, in other words, the application of inter prediction may be prohibited (see Figure 95).
[0728] This prevents the application of inter-prediction when there is no reference three-dimensional point in the reference frame that corresponds to the three-dimensional point to be encoded or decoded due to differences in frame structure between the reference frame and the frame to be encoded or decoded, thereby stabilizing encoding efficiency.
[0729] In addition, if i≧NumLoD_Ref for the i-th LoD of the frame to be coded or decoded, and the decoding device receives a bitstream in which disp_lod_inter_mode[i] is set to 1, it may determine that the bitstream does not comply with the standard and output an error.
[0730] Note that if i≧NumLoD_Ref is satisfied for the i-th LoD of the frame to be coded or decoded, and the decoding device receives a bitstream in which disp_lod_inter_mode[i] is set to 1, the decoding device may perform inter prediction with the predicted value of the i-th LoD set to 0. Alternatively, the decoding device may not perform inter prediction for the i-th LoD that satisfies i≧NumLoD_Ref, regardless of the value of disp_lod_inter_mode[i].
[0731] In this case, disp_lod_inter_mode[i] may be used to notify different information.
[0732] In addition, when i≧NumLoD_Ref for the i-th LoD of the frame to be encoded or decoded, inter prediction may be applied to the LoD for which a reference 3D point exists in the reference frame, and inter prediction may not be applied to the LoD for which a reference 3D point does not exist in the reference frame.
[0733] This allows inter prediction to be applied to at least the LoDs where a reference 3D point exists, thereby improving coding efficiency. In this case, the value of disp_lod_inter_mode[i] corresponding to the LoD to which inter prediction is applied may be set to 1. This allows the decoding device to apply the same process and correctly decode the bitstream.
[0734] FIG. 96 is a flowchart showing a specific example of the encoding process according to this embodiment.
[0735] In the encoding process shown in Figure 96, when the number of points in LoD(i), which is the i-th LoD layer of the frame (also referred to as the processing frame) that is the target of the encoding or decoding process, and the reference frame are equal (that is, when the number of points in the LoD of the same layer of the processing frame and the reference frame are equal), encoding is possible using inter prediction, the unit of inter prediction of the displacement vector is determined for each LoD, and information disp_lod_inter_mode[i] indicating whether inter prediction has been applied for each LoD is added to the header. Here, i is an integer that varies from 0 to a number that is 1 less than the number of LoDs, and when i is 0, LoD(i) (that is, LoD(0)) corresponds to the base LoD layer.
[0736] Furthermore, if the number of points within LoD(i) of the processing frame and the reference frame are not equal, it is determined that inter prediction is not possible, and the output process of information disp_lod_inter_mode[i] indicating whether inter prediction has been applied can be skipped (in other words, disp_lod_inter_mode[i] does not need to be added to the bitstream). This reduces the amount of coding. However, without being limited to this, for example, disp_lod_inter_mode[i] may be set to 0 and added to the bitstream as information indicating whether inter prediction has been applied. This allows the decoding device to determine whether inter prediction has been applied by decoding disp_lod_inter_mode[i], and to correctly decode the bitstream.
[0737] In step S9601, the inter predictor 4805 performs start processing of loop A, which repeatedly executes the processing of steps S9602 to S9607, step S9611, and steps S9615 to S9616 (however, some processing is excluded depending on the determination results of steps S9602 and S9605), which will be described later. In loop A, attention is paid to each of one or more LoDs, and processing is executed for the focused LoD, and control is exercised so that processing is ultimately performed for all LoDs. In loop A, an integer i, which varies from 0 to a number one less than the number of LoDs, is used as an index. Note that the LoD of interest (i.e., LoD(i)) is also referred to as the LoD of interest or the LoD of interest(i). Loop A may also be referred to as an LoD loop.
[0738] In step S9602, the inter predictor 4805 determines whether the number of 3D points in the target LoD(i) between the processing frame and the reference frame is the same. If it is determined that the number of 3D points is the same (Yes in step S9602), the process proceeds to step S9603; if not (No in step S9602), the process proceeds to step S9611.
[0739] In step S9611, the inter predictor 4805 determines to encode the transform coefficients of the displacement vectors in the LoD(i) of interest without using inter prediction, and outputs the transform coefficients to the quantizer 4806. Note that the inter predictor 4805 may add information indicating that the transform coefficients of the displacement vectors in the LoD(i) of interest have been encoded without using inter prediction to the header. For example, by setting the information disp_lod_inter_mode(i), which is included in the header and indicates that the transform coefficients of the displacement vectors have been encoded using inter prediction, to 0, it is possible to indicate that the transform coefficients of the displacement vectors in the LoD(i) of interest have been encoded without using inter prediction.
[0740] In step S9603, the inter predictor 4805 calculates the sum (also referred to as sum_lod_nointer) of the transform coefficients of the displacement vector within the LoD(i) of interest.
[0741] In step S9604, the inter predictor 4805 calculates the sum (also referred to as sum_lod_inter) of prediction residuals within the LoD(i) of interest when inter prediction is applied to the transform coefficients of the displacement vector.
[0742] In step S9605, the inter predictor 4805 determines whether sum_lod_inter calculated in step S9604 is smaller than sum_lod_nointer calculated in step S9603. If it is determined that sum_lod_inter is smaller than sum_lod_nointer (Yes in step S9605), the process proceeds to step S9606; otherwise (No in step S9605), the process proceeds to step S9615.
[0743] In step S9606, the inter predictor 4805 determines to encode the transform coefficients of the displacement vectors in the current LoD(i) using inter prediction, and outputs the prediction residual to the quantizer 4806.
[0744] In step S9607, the inter predictor 4805 adds information indicating that the transform coefficients of the displacement vectors in the LoD(i) of interest have been coded using inter prediction to a header (e.g., a stream header). For example, by setting the information disp_lod_inter_mode(i), which is included in the header and indicates that the transform coefficients of the displacement vectors have been coded using inter prediction, to 1, it is possible to indicate that the transform coefficients of the displacement vectors in the LoD(i) of interest have been coded using inter prediction.
[0745] In step S9615, the inter predictor 4805 determines to encode the transform coefficients of the displacement vectors in the current LoD(i) without using inter prediction, and outputs the transform coefficients to the quantizer 4806.
[0746] In step S9616, the inter predictor 4805 adds information indicating that the transform coefficients of the displacement vectors in the LoD(i) of interest have been coded without using inter prediction to the header. For example, by setting the information disp_lod_inter_mode(i), which is included in the header and indicates that the transform coefficients of the displacement vectors have been coded using inter prediction, to 0, it is possible to indicate that the transform coefficients of the displacement vectors in the LoD(i) of interest have been coded without using inter prediction.
[0747] In step S9608, the inter predictor 4805 performs processing to end loop A. Specifically, the inter predictor 4805 determines whether or not the processing of steps S9602 to S9607, step S9611, and steps S9615 to S9616 (however, depending on the determination results of step S9602 and step S9605, some processing is excluded) has been performed for all LoDs, and if not, controls to perform processing focusing on the LoDs that have not yet been performed.
[0748] Note that the above encoding process is merely an example. In the above example, when the number of 3D points in the reference frame is the same and inter-prediction is possible, the encoding device compares the sums of prediction residuals when each of the multiple methods is used in all LoD layers to determine the method to be used for encoding. However, the encoding device may always perform inter-prediction and skip the process of adding information indicating that the transform coefficients of the displacement vector in the target LoD(i) have been encoded using inter-prediction to the header (specifically, the process of adding disp_lod_inter_mode(i) as 1 to the header, the above-mentioned step S9607). Furthermore, the encoding device may output disp_lod_inter_mode(i) as 0 when inter-prediction is not possible.
[0749] 96, the encoding device can perform inter prediction only when the number of points in LoD(i) of the processing frame and the reference frame is equal. Furthermore, the encoding device can turn off inter prediction when the reference frame cannot be referenced, such as when the subdivision method changes between the processing frame and the reference frame. This allows the encoding device to avoid degradation of encoding efficiency.
[0750] FIG. 97 is a flowchart showing a specific example of the encoding process according to this embodiment.
[0751] In the encoding process shown in Fig. 97, among the processes for each LoD, for LoD(0) where i = 0 (i.e., the LoD layer corresponding to the base mesh), if the number of 3D points in the LoD(0) of the processing frame and the reference frame is the same, inter prediction is performed. Also, for LoD(i) where i > 0, if the subdivision division method is different between the processing frame and the reference frame, or if the number of 3D points in LoD(i) is different, inter prediction is not performed. Under other conditions, inter prediction is performed.
[0752] In step S9701, the inter predictor 4805 performs start processing of loop A, which repeatedly executes the processing of steps S9702 to S9705, step S9711, and step S9715 (described below, however, some processing is excluded depending on the determination results of step S9702 and step S9711). Loop A focuses on each of one or more LoDs, executes processing for the focused LoD, and controls so that processing is ultimately performed for all LoDs. Loop A uses an integer i, which varies from 0 to a number one less than the number of LoDs, as an index. Note that the LoD of interest (i.e., LoD(i)) is also referred to as the LoD of interest or LoD(i). Loop A may also be referred to as an LoD loop.
[0753] In step S9702, the inter predictor 4805 determines whether i > 0. Since i > 0 corresponds to the LoD(i) of interest being a different LoD from the layer corresponding to the base mesh, it can also be said that the inter predictor 4805 determines whether the LoD(i) of interest is a different LoD from the layer corresponding to the base mesh. If it is determined that i > 0 is true (Yes in step S9702), the process proceeds to step S9703; if not (No in step S9702), the process proceeds to step S9711.
[0754] In step S9711, the inter predictor 4805 determines whether the number of 3D points within the LoD(0) of interest between the processing frame and the reference frame is the same. If it is determined that the number of 3D points within the LoD(0) of interest between the processing frame and the reference frame is the same (Yes in step S9711), proceed to step S9705; if not (No in step S9711), proceed to step S9715.
[0755] In step S9703, the inter predictor 4805 determines whether the sub-division methods of the processing frame and the reference frame are the same. If it is determined that the sub-division methods of the processing frame and the reference frame are the same (Yes in step S9703), the process proceeds to step S9704; if not (No in step S9703), the process proceeds to step S9715.
[0756] In step S9704, the inter predictor 4805 determines whether the number of 3D points in the target LoD(i) of the processing frame and the reference frame is the same. If it is determined that the number of 3D points in the target LoD(i) of the processing frame and the reference frame is the same (Yes in step S9704), the process proceeds to step S9705; if not (No in step S9704), the process proceeds to step S9715.
[0757] In step S9705, the inter predictor 4805 determines to encode the transform coefficients of the displacement vectors in the LoD(i) of interest using inter prediction, and outputs the prediction residual to the quantizer 4806. Note that the inter predictor 4805 may add information indicating that the transform coefficients of the displacement vectors in the LoD(i) of interest have been encoded using inter prediction to a header (e.g., a stream header). For example, by setting the information disp_lod_inter_mode(i), which is included in the header and indicates that the transform coefficients of the displacement vectors have been encoded using inter prediction, to 1, it is possible to indicate that the transform coefficients of the displacement vectors in the LoD(i) of interest have been encoded using inter prediction.
[0758] In step S9715, the inter predictor 4805 determines to encode the transform coefficients of the displacement vectors in the LoD(i) of interest without using inter prediction, and outputs the transform coefficients to the quantizer 4806. Note that the inter predictor 4805 may add information indicating that the transform coefficients of the displacement vectors in the LoD(i) of interest have been encoded without using inter prediction to the header. For example, by setting the information disp_lod_inter_mode(i), which is included in the header and indicates that the transform coefficients of the displacement vectors have been encoded using inter prediction, to 0, it is possible to indicate that the transform coefficients of the displacement vectors in the LoD(i) of interest have been encoded without using inter prediction.
[0759] In step S9706, the inter predictor 4805 performs processing to end loop A. Specifically, the inter predictor 4805 determines whether or not the processing of steps S9702 to S9705, step S9711, and step S9715 (however, some processing is excluded depending on the determination results of step S9702 and step S9711) has been performed for all LoDs, and if not, controls to perform processing focusing on the LoDs that have not yet been performed.
[0760] 96, the encoding device can be expected to avoid degradation of encoding efficiency by not performing inter prediction when the number of 3D points in the base mesh is different, when the subdivision division method differs between the processing frame and the reference frame at the LoD(i) layer (where i>0), or when the number of 3D points in LoD(i) is different. This is because in the above cases, there is a high possibility that coordinates will be misaligned between frames.
[0761] The division method for the subdivision may be determined based on the related parameters, the difference in the coded coordinates, or the number of 3D points of the base mesh.
[0762] Note that if inter prediction is not performed, information disp_lod_inter_mode[i] indicating whether inter prediction has been applied does not need to be added to the bitstream. This allows for a reduction in the amount of coding. Note that, without being limited to this, for example, the value of disp_lod_inter_mode[i] may be set to 0 and added to the bitstream. This allows the decoding device to determine whether inter prediction has been applied by decoding disp_lod_inter_mode[i], and to correctly decode the bitstream.
[0763] Note that the above encoding process is merely an example. In the above example, when the number of 3D points in the reference frame is the same and inter-prediction is possible, the encoding device compares the sums of prediction residuals when each of the multiple methods is used in all LoD layers between the multiple methods to determine the method to be used for encoding. However, the encoding device may always perform inter-prediction and skip the process of adding information indicating that the transform coefficients of the displacement vector in the target LoD(i) have been encoded using inter-prediction to the header (specifically, the process of adding disp_lod_inter_mode(i) as 1 to the header, the above-mentioned step S9705). Furthermore, the encoding device may output disp_lod_inter_mode(i) as 0 when inter-prediction is not possible.
[0764] FIG. 98 is a flowchart showing a specific example of the encoding process according to this embodiment.
[0765] In the encoding process shown in Figure 98, when the base mesh structures of the processing frame and the reference frame are the same, inter prediction processing is performed in the same manner as the series of processes shown in Figure 97, and when the base mesh structures of the processing frame and the reference frame are different, inter prediction processing is not performed. Note that the structure of the base mesh can be determined by parameters related to the structure of the base mesh or the number of three-dimensional points included in the base mesh, etc.
[0766] In step S9801, the inter predictor 4805 determines whether the LoD(0) structures of the processing frame and the reference frame are the same. If it is determined that the LoD(0) structures of the processing frame and the reference frame are the same (Yes in step S9801), the process proceeds to step S9802; if not (No in step S9801), the process proceeds to step S9811.
[0767] In step S9811, the inter predictor 4805 performs start processing of loop B, which repeatedly executes the processing of step S9812, which will be described later. In loop B, attention is focused on each of one or more LoDs, and processing is executed for the focused LoD, and control is exercised so that processing is ultimately performed for all LoDs. In loop B, an integer i, which varies from 0 to a number one less than the number of LoDs, is used as an index. Note that the LoD of interest (i.e., LoD(i)) is also referred to as the LoD of interest or the LoD of interest(i). Loop B may also be referred to as an LoD loop.
[0768] In step S9812, the inter predictor 4805 determines to encode the transform coefficients of the displacement vectors in the LoD(i) of interest without using inter prediction, and outputs the transform coefficients to the quantizer 4806. Note that the inter predictor 4805 may add information indicating that the transform coefficients of the displacement vectors in the LoD(i) of interest have been encoded without using inter prediction to the header. For example, by setting the information disp_lod_inter_mode(i), which is included in the header and indicates that the transform coefficients of the displacement vectors have been encoded using inter prediction, to 0, it is possible to indicate that the transform coefficients of the displacement vectors in the LoD(i) of interest have been encoded without using inter prediction.
[0769] In step S9813, the inter predictor 4805 performs processing to end loop B. Specifically, the inter predictor 4805 determines whether or not the processing of step S9812 has been executed for all LoDs, and if not, controls to execute processing focusing on the LoDs that have not yet been executed.
[0770] The processes of steps S9802 to S9807, step S9821, and step S9825 are similar to steps S9701 to S9706, step S9711, and step S9715 shown in FIG. 97, respectively.
[0771] By the process shown in Figure 98, the encoding device may be able to avoid degradation of encoding efficiency by not performing inter prediction. If the base mesh structures of the processing frame and the reference frame are different, there is a high possibility that the correlation between the frames is relatively small, and the effect of inter prediction cannot be expected. Therefore, the encoding device can avoid degradation of encoding efficiency by not performing inter prediction.
[0772] If inter prediction is not performed, information disp_lod_inter_mode[i] indicating whether inter prediction has been applied does not need to be added to the bitstream. This allows the amount of coding to be reduced. However, this is not limiting, and for example, the value of disp_lod_inter_mode[i] may be set to 0 and added to the bitstream. This allows the decoding device to determine whether inter prediction has been applied by decoding disp_lod_inter_mode[i], and can correctly decode the bitstream.
[0773] Note that the above encoding process is merely an example. In the above example, the encoding device performs inter prediction if the conditions are such that inter prediction is possible, and does not perform inter prediction if the conditions are such that inter prediction is not possible. However, the encoding device may also be configured to determine whether to perform inter prediction if the conditions are such that inter prediction is possible, determine which process to perform based on the determination result, and output disp_lod_inter_mode, as in the encoding process exemplified in Figure 49, Figure 52, or Figure 96.
[0774] FIG. 99 is a flow diagram showing a specific example of the encoding process according to this embodiment.
[0775] In the encoding process shown in Figure 99, for each 3D point P (i, j) in LoD (i) in the processing frame, the transform coefficient when encoding using inter prediction is performed and the prediction residual when encoding without using inter prediction are compared, and the process used to derive the smaller of the transform coefficient and the prediction residual is selected, and the selection information disp_point_inter_mode of the process at the 3D point P (i, j) is added to the header. Note that the 3D point P (i, j) represents the j-th 3D point in LoD (i).
[0776] Since the encoding device can select whether to perform inter-prediction on a 3D point basis, it is expected that encoding efficiency will improve even when, for example, the correlation between frames changes within LoD.
[0777] In step S9901, the inter predictor 4805 performs start processing of loop A, which repeatedly executes the processing of steps S9902 to S9908 and steps S9911 to S9912 (however, some processing is excluded depending on the determination result of step S9905), which will be described later. In loop A, attention is paid to each of one or more LoDs, and processing is executed for the focused LoD, and control is exercised so that processing is ultimately performed for all LoDs. In loop A, an integer i, which varies from 0 to a number one less than the number of LoDs, is used as an index. Note that the LoD of interest (i.e., LoD(i)) is also referred to as the LoD of interest or the LoD of interest(i). Loop A may also be referred to as an LoD loop.
[0778] In step S9902, the inter predictor 4805 performs start processing of loop B, which repeatedly executes the processing of steps S9903 to S9907 and steps S9911 to S9912 (described later, however, some processing is excluded depending on the determination result of step S9905). Loop B focuses on each of one or more 3D points included in the LoD of interest, executes processing for the 3D points of interest, and controls so that processing is ultimately performed for all 3D points included in the LoD of interest. Loop B uses an integer j, which varies from 0 to a number one less than the number of 3D points included in the LoD of interest (i.e., LoD(i)), as an index. The 3D point of interest is also referred to as the 3D point of interest or 3D point P(i,j). Loop B may also be referred to as a Point loop.
[0779] In step S9903, the inter predictor 4805 calculates the transformation coefficient (also called point_lod_nointer) of the displacement vector of the 3D point P(i,j).
[0780] In step S9904, the inter predictor 4805 calculates a prediction residual (also referred to as point_lod_inter) when inter prediction is applied to the transform coefficient of the displacement vector of the 3D point P(i, j).
[0781] In step S9905, the inter predictor 4805 determines whether point_lod_inter calculated in step S9904 is smaller than point_lod_nointer calculated in step S9903. If it is determined that point_lod_inter is smaller than point_lod_nointer (Yes in step S9905), the process proceeds to step S9906; otherwise (No in step S9905), the process proceeds to step S9911.
[0782] In step S9906, the inter predictor 4805 determines to encode the transform coefficients of the displacement vector of the 3D point P(i,j) using inter prediction, and outputs the prediction residual to the quantizer 4806.
[0783] In step S9907, the inter predictor 4805 adds information indicating that the transform coefficients of the displacement vector of the three-dimensional point P(i, j) have been coded using inter prediction to a header (e.g., the header of the stream). For example, by setting the information disp_point_inter_mode(i, j), which is included in the header and indicates that the transform coefficients of the displacement vector of the three-dimensional point P(i, j) have been coded using inter prediction, to 1, it is possible to indicate that the transform coefficients of the displacement vector of the three-dimensional point P(i, j) have been coded using inter prediction.
[0784] In step S9911, the inter predictor 4805 determines to encode the transform coefficients of the displacement vector of the 3D point P(i,j) without using inter prediction, and outputs the transform coefficients to the quantizer 4806.
[0785] In step S9912, the inter predictor 4805 adds information indicating that the transform coefficients of the displacement vector of the three-dimensional point P(i, j) have been coded without using inter prediction to the header. For example, by setting the information disp_point_inter_mode(i, j), which is included in the header and indicates that the transform coefficients of the displacement vector of the three-dimensional point P(i, j) have been coded using inter prediction, to 0, it is possible to indicate that the transform coefficients of the displacement vector of the three-dimensional point P(i, j) have been coded without using inter prediction.
[0786] In step S9908, the inter predictor 4805 performs processing to end loop B. Specifically, the inter predictor 4805 determines whether or not the processing of steps S9903 to S9907 and steps S9911 to S9912 (however, depending on the determination result of step S9905, some processing is excluded) has been performed for all 3D points, and if not, controls so that processing is performed with a focus on 3D points that have not yet been performed.
[0787] In step S9909, the inter predictor 4805 performs processing to end loop A. Specifically, the inter predictor 4805 determines whether or not the processing of steps S9902 to S9908 and steps S9911 to S9912 (however, depending on the determination result of step S9905, some processing is excluded) has been performed for all LoDs, and if not, controls to perform processing focusing on the LoDs that have not yet been performed.
[0788] It should be noted that the above encoding process is merely an example. In the above example, the encoding device performs processing at all points of the displacement vector, but for example, it may determine whether to perform inter-prediction for each LoD unit, and perform the above process for each 3D point only for the LoD for which inter-prediction is performed. Also, it may be combined with the encoding process exemplified in Figure 96, Figure 97 or Figure 98. Also, instead of LoD, it may be used a prediction unit DVG exemplified in Figure 84, Figure 86, Figure 87 or Figure 88.
[0789] FIG. 100 is a flow diagram showing a specific example of the encoding process according to this embodiment.
[0790] In the encoding process shown in Figure 100, only when a three-dimensional point P(i, j) exists within LoD(i) in the reference frame, the transform coefficient when encoding using inter-prediction is compared with the prediction residual when encoding without inter-prediction, the process used to derive the smaller of the transform coefficient and the prediction residual is selected, and selection information disp_point_inter_mode for the process at the three-dimensional point P(i, j) is added to the header.
[0791] Furthermore, the encoding device does not perform inter prediction if there is no 3D point P(i,j) within LoD(i) in the reference frame.
[0792] The encoding device can perform inter-prediction even if the number of three-dimensional points within LoD differs between the processing frame and the reference frame, and can avoid a deterioration in encoding efficiency by not performing inter-prediction if reference is not possible.
[0793] Furthermore, if there is no 3D point P'(i, j) within LoD(i) in the reference frame, the encoding device may set the selection information disp_point_inter_mode to 0 and add it to the header.
[0794] The processing in steps S10001 and S10002 is the same as the processing in steps S9901 and S9902 shown in FIG.
[0795] In step S10003, the inter predictor 4805 determines whether or not the 3D point P(i,j) exists in the reference frame. If it is determined that the 3D point P(i,j) exists in the reference frame (Yes in step S10003), the process proceeds to step S10004; if not (No in step S10003), the process proceeds to step S10021.
[0796] In step S10021, the inter predictor 4805 determines to encode the transform coefficients of the displacement vector of the three-dimensional point P(i,j) without using inter prediction, and outputs the transform coefficients to the quantizer 4806. Note that the inter predictor 4805 may add information indicating that the transform coefficients of the displacement vector of the three-dimensional point P(i,j) have been encoded without using inter prediction to the header. For example, by setting the information disp_point_inter_mode(i,j), which is included in the header and indicates that the transform coefficients of the displacemen...
Claims
1. A three-dimensional data encoding method, which generates a bitstream including data to be encoded and one or more flags for determining whether or not to perform a quantization process on the data to be encoded.
2. The three-dimensional data encoding method according to claim 1, wherein one or more of the flags are assigned to each of a plurality of levels into which the encoding target data is hierarchically divided.
3. The three-dimensional data encoding method according to claim 2, wherein the plurality of levels include a sequence level including the entire data to be encoded, a frame level into which the sequence level is divided in time, and a patch level into which the frame level is divided in space.
4. The three-dimensional data encoding method according to claim 2, wherein the one or more flags include a first flag indicating whether a quantization parameter is present or not.
5. The three-dimensional data encoding method according to claim 4, wherein the first flag assigned to each of the plurality of levels is set to 0 if the quantization process is skipped.
6. The three-dimensional data encoding method according to claim 2, wherein the one or more flags include a second flag indicating whether or not the quantization process is to be skipped.
7. The three-dimensional data encoding method according to claim 2, wherein whether or not to perform the quantization process is determined based on at least one of the flags assigned to each of the plurality of levels.
8. The three-dimensional data encoding method according to claim 7, wherein whether or not to perform the quantization process is determined based on a combination of the flags assigned to each of the plurality of levels.
9. The three-dimensional data encoding method according to claim 7, wherein it is determined to skip the quantization process when all of the flags assigned to the respective levels have the same value.
10. The three-dimensional data encoding method according to claim 1, wherein, if the quantization process is performed, the bitstream includes information signaling a quantization parameter.
11. The three-dimensional data encoding method according to claim 1, wherein the encoding target data includes information corresponding to each of a plurality of vertices in a three-dimensional space.
12. A three-dimensional data decoding method comprising: acquiring a bitstream; acquiring from the acquired bitstream data to be decoded and one or more flags for determining whether or not to perform inverse quantization processing on the data to be decoded; performing a decoding process to decode the data to be decoded; and determining, in the decoding process, whether or not to perform inverse quantization processing based on the one or more flags.
13. The three-dimensional data decoding method according to claim 12, wherein each of the one or more flags is assigned to each of a plurality of levels into which the decoding target data is hierarchically divided.
14. A three-dimensional data decoding method as described in claim 13, wherein the plurality of levels include a sequence level including the entire data to be decoded, a frame level into which the sequence level is divided in time, and a patch level into which the frame level is divided in space.
15. The three-dimensional data decoding method according to claim 13, wherein the one or more flags include a first flag indicating whether or not an inverse quantization parameter is present.
16. The three-dimensional data decoding method according to claim 15, wherein, if the inverse quantization process is skipped, the first flag assigned to each of the plurality of levels is set to 0.
17. The three-dimensional data decoding method according to claim 13, wherein the one or more flags include a second flag indicating whether or not the inverse quantization process is to be skipped.
18. The three-dimensional data decoding method according to claim 13, wherein whether or not to perform the inverse quantization process is determined based on at least one of the flags assigned to each of the plurality of levels.
19. The three-dimensional data decoding method according to claim 18, wherein whether or not to perform the inverse quantization process is determined based on a combination of the flags assigned to each of the plurality of levels.
20. The three-dimensional data decoding method according to claim 18, wherein it is determined to skip the inverse quantization process if all of the flags assigned to the respective levels have the same value.
21. The three-dimensional data decoding method according to claim 12, wherein the data to be decoded includes information corresponding to each of a plurality of vertices in a three-dimensional space.
22. A three-dimensional data encoding device comprising: a memory; and a circuit capable of accessing the memory, wherein the circuit, in operation, generates a bitstream including data to be encoded and one or more flags for determining whether or not to perform a quantization process on the data to be encoded.
23. A three-dimensional data decoding device comprising: a memory; and a circuit capable of accessing the memory, wherein the circuit, in operation, acquires a bitstream; acquires from the acquired bitstream data to be decoded and one or more flags for determining whether or not to perform inverse quantization processing on the data to be decoded; performs a decoding process to decode the data to be decoded; and determines, in the decoding process, whether or not to perform inverse quantization processing based on the one or more flags.
Citation Information
Patent Citations
Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
WO2021261516A1