Three-d data decoder and three-d data encoder
The 3D data decoding device addresses the inefficiencies in tile correspondence by using an atlas information decoding unit and attribute decoding unit to clarify tile mappings, resulting in improved decoding efficiency.
Patent Information
- Application Number
- JP2024033701
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-06
- Publication Date
- 2025-09-19
AI Technical Summary
The method of encoding and decoding attribute images using tiles in existing 3D data coding methods, such as MPEG-I's V3C and V-PCC, often fails to establish a clear correspondence between tiles in an atlas frame and tiles in an attribute image, leading to inefficiencies in the decoding process.
A 3D data decoding device is provided with an atlas information decoding unit that decodes parameters related to tile division of an atlas frame and an attribute decoding unit that decodes tile mapping parameters to clarify the correspondence between atlas tiles and video tiles in the attribute image, ensuring efficient decoding.
This approach allows for efficient decoding of 3D data by establishing a clear correspondence between tiles, enhancing the decoding process and improving overall performance.
Smart Images

Figure 2025135759000001_ABST
Abstract
Description
[Technical Field]
[0001] An embodiment of the present invention relates to a 3D data encoding device and a 3D data decoding device. [Background technology]
[0002] To efficiently transmit or record 3D data, there are 3D data encoding devices that convert the 3D data into 2D images, encode them using a video encoding method, and generate encoded data, and 3D data decoding devices that decode the 2D images from the encoded data and reconstruct the 3D data.
[0003] Specific examples of 3D data encoding methods include MPEG-I's ISO / IEC 23090-5 V3C (Volumetric Video-based Coding) and V-PCC (Video-based Point Cloud Compression). V3C can encode and decode point clouds consisting of point positions and attribute information. Furthermore, it can also be used to encode and decode multi-viewpoint video and mesh video using ISO / IEC 23090-12 (MPEG Immersive Video, MIV) and ISO / IEC 23090-29 (Video-based Dynamic Mesh Coding, V-DMC), which is currently being standardized. The latest draft document of the V-DMC method is disclosed in Non-Patent Document 1.
[0004] In these 3D data coding methods, the geometry and attributes that make up the 3D data are encoded and decoded as images using video coding methods such as H.265 / HEVC (High Efficiency Video Coding) and H.266 / VVC (Versatile Video Coding).
[0005] In the case of a point cloud, the geometry image is the depth to the projection plane, and the attribute image is the image of the attributes projected onto the projection plane.
[0006] 3D data (mesh) such as that described in Non-Patent Document 1 consists of a base mesh, mesh displacement, mesh displacement array, and texture mapping image. A vertex encoding method such as Draco can be used for the base mesh, the geometry image is a mesh displacement image obtained by two-dimensionalizing the mesh displacement, and the attribute image is a texture mapping image. These are encoded and decoded using a video encoding method such as HEVC or VVC, as described above. [Prior art documents] [Non-patent literature]
[0007] [Non-Patent Document 1] WD 5.0 of V-DMC (MDS23318_WG07_N00744), ISO / IEC JTC 1 / SC 29 / WG 7 N0744, October 2023 Summary of the Invention [Problem to be solved by the invention]
[0008] The method of encoding and decoding attribute images using tiles in Non-Patent Document 1 has a problem in that tiles of an atlas frame may not correspond to tiles of attribute images in video encoding.
[0009] The present invention aims to efficiently decode 3D data by clarifying the correspondence between tiles in an atlas frame and tiles in an attribute image when encoding and decoding 3D data using a video encoding method. [Means for solving the problem]
[0010] In order to solve the above problem, a 3D data decoding device according to one aspect of the present invention is provided. a 3D data decoding device for decoding atlas information, the device comprising: an atlas information decoding unit that decodes parameters related to tile division of an atlas frame from an atlas data stream whose Unit Type of the encoded data is V3C_AD; and an attribute decoding unit that decodes attribute images from an attribute data stream whose Unit Type of the encoded data is V3C_AVD, wherein the attribute decoding unit decodes tile mapping parameters that indicate the correspondence between atlas tiles that indicate the partition division of the atlas frame and video tiles that indicate the partition division of a video codec of the attribute image. [Effects of the Invention]
[0011] According to one aspect of the present invention, 3D data can be efficiently decoded by clarifying the correspondence between tiles of an atlas frame and tiles of an attribute image. [Brief explanation of the drawings]
[0012] [Figure 1] 1 is a schematic diagram showing the configuration of a 3D data transmission system according to the present embodiment. [Figure 2] FIG. 2 is a diagram showing a hierarchical structure of data in an encoded stream. [Figure 3] FIG. 2 is a functional block diagram showing a schematic configuration of a 3D data decoding device 31. [Figure 4] FIG. 2 is a functional block diagram showing the configuration of a base mesh decoding unit 303. [Figure 5] FIG. 10 is a functional block diagram showing the configuration of a mesh displacement decoding unit 305. [Figure 6] FIG. 2 is a functional block diagram showing the configuration of a mesh reconstruction unit 307. [Figure 7] 10 is an example of a syntax for transmitting coordinate transformation parameters and displacement mapping parameters at the sequence level (ASPS). [Figure 8]10 is an example of a syntax for transmitting coordinate transformation parameters and displacement mapping parameters at the picture / frame level (AFPS). [Figure 9] This is an example where the mesh displacement image gFrame is divided into segments (slices) equal to the number of dimensions of the geometry. [Figure 10] FIG. 10 is a diagram for explaining the operation of the mesh reconstruction unit 307. [Figure 11] 1 is a functional block diagram showing a schematic configuration of a 3D data encoding device 11. FIG. [Figure 12] FIG. 2 is a functional block diagram showing the configuration of a base mesh encoding unit 103. [Figure 13] FIG. 2 is a functional block diagram showing the configuration of a mesh displacement encoding unit 107. [Figure 14] FIG. 2 is a functional block diagram showing the configuration of a mesh separation unit 115. [Figure 15] 10 is a diagram for explaining the operation of the mesh separating unit 115. FIG. [Figure 16] 10 is an example of the syntax of parameters related to tiling of an atlas frame. [Figure 17] FIG. 10 is a diagram illustrating the relationship between video tiles and atlas tiles. [Figure 18] 10 is an example of the syntax of tile mapping parameters. [Figure 19] 10 is an example of the syntax of tile mapping parameters. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0014] FIG. 1 is a schematic diagram showing the configuration of a 3D data transmission system 1 according to this embodiment.
[0015] The 3D data transmission system 1 is a system that transmits an encoded stream obtained by encoding 3D data to be encoded, decodes the transmitted encoded stream, and displays the 3D data. The 3D data transmission system 1 includes a 3D data encoding device 11, a network 21, a 3D data decoding device 31, and a 3D data display device 41.
[0016] The 3D data T is input to the 3D data encoding device 11.
[0017] The network 21 transmits the encoded stream Te generated by the 3D data encoding device 11 to the 3D data decoding device 31. The network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination of these. The network 21 is not necessarily limited to a bidirectional communication network, and may be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. Furthermore, the network 21 may be replaced by a storage medium on which the encoded stream Te is recorded, such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blu-ray Disc: registered trademark).
[0018] The 3D data decoding device 31 decodes each of the coded streams Te transmitted by the network 21, and generates one or more decoded 3D data Td.
[0019] The 3D data display device 41 displays all or part of one or more pieces of decoded 3D data Td generated by the 3D data decoding device 31. The 3D data display device 41 is equipped with a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. Display forms include stationary, mobile, and HMD. Furthermore, if the 3D data decoding device 31 has high processing power, it displays high-quality images, and if it has only lower processing power, it displays images that do not require high processing power or display power.
[0020] <operator> The operators used in this specification are listed below.
[0021] >> is a right bit shift, << is a left bit shift, & is a bitwise AND, | is a bitwise OR, |= is the OR assignment operator, and || indicates logical sum.
[0022] x?y:z is a ternary operator that takes y if x is true (non-zero) and z if x is false (zero). y..z denotes the set of integers from y to z.
[0023] <Structure of the coded stream Te> Before proceeding to a detailed description of the 3D data encoding device 11 and the 3D data decoding device 31 according to this embodiment, the data structure of the encoded stream Te generated by the 3D data encoding device 11 and decoded by the 3D data decoding device 31 will be described.
[0024] 2 is a diagram showing the hierarchical structure of data in the coded stream Te. The coded stream Te has the data structure of either a V3C sample stream or a V3C unit stream. A V3C sample stream includes a sample stream header and a V3C unit. A V3C unit stream includes a V3C unit.
[0025] A V3C unit includes a V3C unit header and a V3C unit payload. The V3C unit header begins with a Unit Type, which is an ID that indicates the type of V3C unit, and takes values indicated by labels such as V3C_VPS, V3C_AD, V3C_AVD, V3C_GVD, and V3C_OVD.
[0026] If the Unit Type is V3C_VPS (Video Parameter Set), the V3C unit contains a V3C parameter set.
[0027] When the Unit Type is V3C_AD (Atlas Data), the V3C unit includes a VPS ID, an atlas ID, a sample stream NAL header, and multiple NAL units. The ID (Identification) takes an integer value of 0 or greater.
[0028] A NAL unit consists of NALUnitType, layerID, TemporalID, and RBSP (Raw byte sequence payload).
[0029] NAL units are identified by NALUnitType, and are classified as ASPS (Atlas Sequence Parameter Set), AAPS (Atlas Adaptation Parameter Set), ATL (Atlas Tile layer), SEI (Supplemental Enhancement Information), etc.
[0030] The ATL includes an ATL header and an ATL data unit, and the ATL data unit includes information such as the patch position and size, such as patch information data.
[0031] The SEI includes payloadType indicating the type of SEI, payloadSize indicating the size (number of bytes) of the SEI, and sei_payload of the SEI data.
[0032] If the Unit Type is V3C_AVD (Attribute Video Data), The unit includes the VPS ID, atlasID, attribute image ID attrIdx, partition ID partIdx, map ID mapIdx, a flag indicating whether it is auxiliary data or not, and a video stream. The video stream indicates data encoded using HEVC, VVC, etc. In V-DMC, it corresponds to a texture image.
[0033] If the Unit Type is V3C_GVD (Geometry Video Data), the V3C unit includes VPS ID, atlasID, mapIdx, auxFlag, and video stream. In V-DMC, it corresponds to mesh displacement.
[0034] If the Unit Type is V3C_OVD (Occumancy Video Data), the V3C unit includes a VPS ID, an atlas ID, and a video stream.
[0035] If the Unit Type is V3C_MD (Mesh data), the V3C unit contains the VPS ID, atlas ID, and mesh_payload. In V-DMC, it corresponds to the base mesh.
[0036] (Configuration of 3D data decoding device according to the first embodiment) 3 is a functional block diagram showing a schematic configuration of a 3D data decoding device 31 according to the first embodiment. The 3D data decoding device 31 is composed of a demultiplexing unit 301, an atlas information decoding unit 302, a base mesh decoding unit 303, a mesh displacement decoding unit 305, a mesh reconstruction unit 307, an attribute decoding unit 306, and a color space conversion unit 308. The 3D data decoding device 31 inputs encoded 3D data and outputs atlas information, meshes, and attribute images.
[0037] The demultiplexing unit 301 inputs encoded data multiplexed in a byte stream format, ISOBMFF (ISO Base Media File Format), etc., demultiplexes it, and outputs an atlas information encoded stream (V3C_AD Atlas Data stream, NAL unit), a base mesh encoded stream (V3C_MD mesh_payload), a mesh displacement encoded stream (V3C_GVD video stream), and an attribute video stream (V3C_AVD video stream).
[0038] The atlas information decoding unit 302 receives the atlas information coded stream output from the demultiplexing unit 301 and decodes the atlas information.
[0039] The base mesh decoding unit 303 decodes the base mesh coded stream coded by vertex coding (3D data compression coding method, for example, Draco), and outputs a base mesh. The base mesh will be described later.
[0040] The mesh displacement decoding unit 305 decodes a geometry video stream (mesh displacement coded stream) coded using VVC, HEVC, or the like, and outputs mesh displacement. The type of codec used for coding is determined by the ptl_profile_c parameter set obtained by decoding the V3C parameter set of the coded data. It can also be represented by the Four CC code (four-character code, 4CC code) indicated by gi_geometry_codec_id[atlasID] in the V3C parameter set. gi_geometry_codec_id[atlasID] indicates the index corresponding to the codec ID of the decoder used to decode the geometry video stream in the atlas ID. The set indicating the correspondence between the codec ID (ccm_codec_id) and its 4CC code (ccm_codec_4cc[ccm_codec_id]) can be transmitted in a separate codec mapping SEI (component_codec_mapping SEI). The codec may decode mesh displacement in units of subsets (segments) that further divide a frame. HEVC and VVC can divide frames into slices. Slices are coded in units of Coded Tree Units (CTUs). Note that subpictures or tiles may be used as subsets instead of slices. These subpictures, tiles, and slices can be decoded independently, so it is possible to decode only a part of a frame without decoding the entire frame. When using subpictures or tiles, slices are replaced with subpictures or tiles.
[0041] The mesh reconstructing unit 307 receives the base mesh and the mesh displacement and reconstructs the mesh in the 3D space.
[0042] The attribute decoding unit 306 decodes an attribute video stream encoded using VVC, HEVC, or the like, and outputs an attribute image in YCbCr format. The attribute image may be a texture image unwrapped along the UV axis (a texture-mapped image converted using the UV atlas method). The type of codec used for encoding is indicated by ptl_profile_codec_group_idc, which is obtained by decoding the V3C parameter set of the encoded data. It may also be indicated by the Four CC code indicated by ai_geometry_codec_id[atlasID] in the V3C parameter set. ai_geometry_codec_id[atlasID] indicates the index in the atlas ID that corresponds to the codec ID of the decoder used to decode the attribute video stream.
[0043] The attribute image may be decoded for each video tile that indicates partition division of video coding such as VVC, HEVC, etc. Also, the attribute image may be decoded for each atlas tile that indicates partition division of an atlas frame.
[0044] FIG. 16 shows an example of the syntax of parameters related to tile division of an atlas frame (atlas frame tile information). For example, afti_single_tile_in_atlas_frame_flag is a flag indicating whether the atlas frame is composed of only one tile. If this flag is equal to true (e.g., 1), the atlas frame is composed of only one tile. If this flag is equal to false (e.g., 0), the atlas frame is composed of a number of tiles greater than 1. afti_num_tiles_in_atlas_frame_minus1 indicates the number of tiles in the atlas frame minus 1. The atlas frame tile information is decoded by the atlas information decoding unit 302.
[0045] FIG. 17 is a diagram for explaining the relationship between video tiles and atlas tiles. FIG. 17(a) is an example of tile division (video tile) in video coding. Video tiles indicate the partition division of an attribute image in a video codec. Specifically, they indicate the partition division of a frame when transmitting an attribute frame in a video codec such as AVC, HEVC, or VVC. A video tile may be a slice in units of a macroblock in AVC / H.264, a slice in units of a CTU in HEVC / H.265, or a slice, tile, or subpicture in VVC / H.266. A video tile is a unit that can be decoded independently, at least in bitstream decoding. In other words, arithmetic coding such as CABAC is initialized at the beginning of each partition (also called a subset or segment) into which a frame is divided. The video tile partition division method involves decoding the parameter sets (Sequence Parameter Set (SPS), Picture Parameter Set (PPS)) of the video codec bitstream and the slice header. Atlas tiles are obtained by decoding the atlas frame tile information (atlas_frame_tile_information()) in the atlas information coded stream. The atlas tile partitioning method is obtained by decoding the atlas frame tile information (atlas_frame_tile_information()) in the atlas information coded stream. In the example of Figure 17, an image frame is composed of four video tiles N (N = 0..3). Figures 17(b), 17(c), and 17(d) are examples of tile division (atlas tiles) in an atlas frame. Figure 17(b) is an example where video tiles and atlas tiles match. Figure 17(c) is an example where an atlas tile contains multiple video tiles. For example, atlas tile 0 contains video tile 0 and video tile 1. Atlas tile 1 contains video tile 2 and video tile 3. Figure 17(d) is an example where an atlas tile contains a portion of a video tile. For example, atlas tile 0 and atlas tile 2 each contain a portion of video tile 0. Atlas tile 1 and atlas tile 3 each contain a portion of video tile 1. Attratile 4 and attratile 6 each contain a portion of video tile 2. Attratile 5 and attratile 7 each contain a portion of video tile 3.
[0046] The attribute decoding unit 306 or the atlas information decoding unit 302 may decode the tile mapping parameters (tile mapping information, tile mapping SEI) to obtain the correspondence between the atlas tiles and the video tiles.
[0047] (Example of tile mapping parameters) Figure 18 shows an example of the syntax of the tile mapping parameters. The semantics of each field are as follows:
[0048] aftm_signalled_video_tile_id_length_minus1: Indicates the number of bits required to express the syntax element aftm_video_tile_id minus 1.
[0049] aftm_num_video_tiles_in_atlas_tile_minus1[i]: Indicates the number of video tiles contained in the i-th (0..afti_num_tiles_in_atlas_frame_minus1) atlas tile minus 1.
[0050] aftm_video_tile_id[i][j]: The tile ID of the jth (0..aftm_num_video_tiles_in_atlas_tile_minus1[i]) video tile in the ith atlas tile (0..afti_num_tiles_in_atlas_frame_minus1). aftm_video_tile_id[i][j] is encoded using the binarization of u(v), and the length of aftm_video_tile_id[i][j] is aftm_signalled_video_tile_id_length_minus1+1 bits. u(n): An unsigned integer using n bits. For n=v, the number of bits is specified separately for u(v).
[0051] Note that afti_single_tile_in_atlas_frame_flag and afti_num_tiles_in_atlas_frame_minus1 refer to the values of the atlas frame tile information (FIG. 16) decoded by the atlas information decoding unit 302.
[0052] In the case of the tile configuration of FIG. 17(b), the values of each field are as follows: afti_single_tile_in_atlas_frame_flag = 0 afti_num_tiles_in_atlas_frame_minus1 = 3 aftm_signalled_video_tile_id_length_minus1 = 1 aftm_num_video_tiles_in_atlas_tile_minus1[0] = 0 aftm_num_video_tiles_in_atlas_tile_minus1[1] = 0 aftm_num_video_tiles_in_atlas_tile_minus1[2] = 0 aftm_num_video_tiles_in_atlas_tile_minus1[3] = 0 aftm_video_tile_id[0][0] = 0 aftm_video_tile_id[1][0] = 1 aftm_video_tile_id[2][0] = 2 aftm_video_tile_id[3][0] = 3 In the case of the tile configuration of FIG. 17(c), the values of each field are as follows: afti_single_tile_in_atlas_frame_flag = 0 afti_num_tiles_in_atlas_frame_minus1 = 1 aftm_signalled_video_tile_id_length_minus1 = 1 aftm_num_video_tiles_in_atlas_tile_minus1[0] = 1 aftm_num_video_tiles_in_atlas_tile_minus1[1] = 1 aftm_video_tile_id[0][0] = 0 aftm_video_tile_id[0][1] = 1 aftm_video_tile_id[1][0] = 2 aftm_video_tile_id[1][1] = 3 In the case of the tile configuration of FIG. 17(d), the values of each field are as follows: afti_single_tile_in_atlas_frame_flag = 0 afti_num_tiles_in_atlas_frame_minus1 = 7 aftm_signalled_video_tile_id_length_minus1 = 1 aftm_num_video_tiles_in_atlas_tile_minus1[0] = 0 aftm_num_video_tiles_in_atlas_tile_minus1[1] = 0 aftm_num_video_tiles_in_atlas_tile_minus1[2] = 0 aftm_num_video_tiles_in_atlas_tile_minus1[3] = 0 aftm_num_video_tiles_in_atlas_tile_minus1[4] = 0 aftm_num_video_tiles_in_atlas_tile_minus1[5] = 0 aftm_num_video_tiles_in_atlas_tile_minus1[6] = 0 aftm_num_video_tiles_in_atlas_tile_minus1[7] = 0 aftm_video_tile_id[0][0] = 0 aftm_video_tile_id[1][0] = 1 aftm_video_tile_id[2][0] = 0 aftm_video_tile_id[3][0] = 1 aftm_video_tile_id[4][0] = 2 aftm_video_tile_id[5][0] = 3 aftm_video_tile_id[6][0] = 2 aftm_video_tile_id[7][0] = 3 The attribute decoding unit 306 or the atlas information decoding unit 302 decodes the tile mapping parameters from the encoded data. As in the above example, the tile mapping parameters may be decoded from the SEI. The tile mapping parameters may be the video tile ID (aftm_video_tile_id). The attribute decoding unit 306 or the atlas information decoding unit 302 may decode, as the tile mapping parameters, aftm_video_tile_id[i][j], which is a string of IDs of video tiles corresponding to tile i in the range of i = 0..afti_num_tiles_in_atlas_frame_minus1 (j = 0..aftm_num_video_tiles_in_atlas_tile_minus1[i]) as in the syntax of FIG. 18. The attribute decoding unit 306 or the atlas information decoding unit 302 stores the list of IDs of the video tiles included in the atlas tile of TileId in a variable VideoTileIdInTile[TileId][]. That is, VideoTileIdInTile[i][j] = aftm_video_tile_id[i][j] where i=0..afti_num_tiles_in_atlas_frame_minus1, j = 0..aftm_num_video_tiles_in_atlas_tile_minus1[i].
[0053] (Another example of tile mapping parameters) Figure 19 shows another example of the syntax of the tile mapping parameters. The semantics of each field are as follows:
[0054] aftm_alignment_flag: Tiling and alignment in atlas frames This flag indicates whether the tile divisions (video tiles) in the video match. If this flag is true (e.g., 1), the atlas tiles and video tiles match. If this flag is false (e.g., 0), the atlas tiles and video tiles do not match.
[0055] aftm_aggregation_flag: A flag indicating whether the attra tile contains multiple video tiles. If this flag is equal to true (e.g., 1), the attra tile can contain multiple video tiles. That is, the attra tile and the video tile have a 1:N correspondence. If this flag is equal to false (e.g., 0), the attra tile does not contain multiple (more than 1) video tiles. That is, the attra tile and the video tile have a 1:1 correspondence. If this flag is not present, the value of this flag may be inferred to be false.
[0056] aftm_signalled_video_tile_id_length_minus1: Same as Figure 18.
[0057] aftm_video_tile_id[i]: Indicates the tile ID of the video tile included in the i-th (0..afti_num_tiles_in_atlas_frame_minus1) atlas tile.
[0058] aftm_num_video_tiles_in_atlas_tile_minus1[i]: Same as Figure 18.
[0059] aftm_video_tile_id[i][j]: Same as Figure 18.
[0060] In the case of the tile configuration of FIG. 17(b), the values of each field are as follows: afti_single_tile_in_atlas_frame_flag = 0 afti_num_tiles_in_atlas_frame_minus1 = 3 aftm_alignment_flag = 1 In the case of the tile configuration of FIG. 17(c), the values of each field are as follows: afti_single_tile_in_atlas_frame_flag = 0 afti_num_tiles_in_atlas_frame_minus1 = 1 aftm_alignment_flag = 0 aftm_aggregation_flag = 1 aftm_signalled_video_tile_id_length_minus1 = 1 aftm_num_video_tiles_in_atlas_tile_minus1[0] = 1 aftm_num_video_tiles_in_atlas_tile_minus1[1] = 1 aftm_video_tile_id[0][0] = 0 aftm_video_tile_id[0][1] = 1 aftm_video_tile_id[1][0] = 2 aftm_video_tile_id[1][1] = 3 In the case of the tile configuration of FIG. 17(d), the values of each field are as follows: afti_single_tile_in_atlas_frame_flag = 0 afti_num_tiles_in_atlas_frame_minus1 = 7 aftm_alignment_flag = 0 aftm_aggregation_flag = 0 aftm_signalled_video_tile_id_length_minus1 = 1 aftm_video_tile_id[0] = 0 aftm_video_tile_id[1] = 1 aftm_video_tile_id[2] = 0 aftm_video_tile_id[3] = 1 aftm_video_tile_id[4] = 2 aftm_video_tile_id[5] = 3 aftm_video_tile_id[6] = 2 aftm_video_tile_id[7] = 3 The attribute decoding unit 306 or the atlas information decoding unit 302 may decode aftm_video_tile_id[i], the ID of one video tile corresponding to tile i in the range of i = 0..afti_num_tiles_in_atlas_frame_minus1, as in the syntax of Fig. 19 as a tile mapping parameter. Alternatively, it may decode aftm_video_tile_id[i][j], which is a sequence of IDs of video tiles (j = 0..aftm_num_video_tiles_in_atlas_tile_minus1[i]) corresponding to tile i in the range of i = 0..afti_num_tiles_in_atlas_frame_minus1. The attribute decoding unit 306 or the atlas information decoding unit 302 stores a list of IDs of video tiles included in the atlas tile of TileId in a variable VideoTileIdInTile[TileId][]. That is, If aftm_aggregation_flag = false, VideoTileIdInTile[i] = aftm_video_tile_id[i] where i = 0..afti_num_tiles_in_atlas_frame_minus1. If aftm_aggregation_flag = true, VideoTileIdInTile[i][j] = aftm_video_tile_id[i][j] where i = 0..afti_num_tiles_in_atlas_frame_minus1, j = 0..aftm_num_video_tiles_in_atlas_tile_minus1[i]. If aftm_alignment_flag = true, VideoTileIdInTile[i] = i where i = 0.. afti_num_tiles_in_atlas_frame_minus1.
[0061] By decoding the tile mapping parameters, when a specific atlas tile is to be decoded, the corresponding video tile can be easily identified, allowing for efficient decoding of the 3D data.
[0062] By introducing aftm_alignment_flag, it is possible to reduce the amount of code required for tile mapping parameters when the atlas tile and video tile match.
[0063] By introducing the aftm_aggregation_flag, it is possible to reduce the amount of coding for the tile mapping parameters when the atlas tile does not contain multiple video tiles.
[0064] The color space conversion unit 308 converts the color space of the attribute image from YCbCr format to RGB format. Note that an attribute video stream coded in RGB format may be decoded and color space conversion may be omitted.
[0065] (Decoding the base mesh) FIG. 4 is a functional block diagram showing the configuration of the base mesh decoding unit 303. The base mesh decoding unit 303 is composed of a mesh decoding unit 3031, a motion information decoding unit 3032, a mesh motion compensation unit 3033, a reference mesh memory 3034, a switch 3035, and a switch 3036. The base mesh decoding unit 303 may also include a base mesh inverse quantization unit (not shown) before the output of the base mesh. When the base mesh to be decoded is coded (intra-coded) without reference to other base meshes (e.g., base meshes that have already been coded and decoded), the switches 3035 and 3036 are connected to the side that does not perform motion compensation. On the other hand, when the base mesh to be decoded is coded (inter-coded) with reference to other base meshes, the switches 3035 and 3036 are connected to the side that performs motion compensation. When motion compensation is performed, the target vertex coordinates are derived by referring to already decoded vertex coordinates and motion information.
[0066] The mesh decoding unit 3031 decodes the intra-coded base mesh coded stream and outputs the base mesh. As the coding method, Draco, Edge Breaker, etc. are used.
[0067] The motion information decoding unit 3032 decodes the inter-coded base mesh coded stream and outputs motion information for each vertex of a reference mesh (to be described later). Entropy coding such as arithmetic coding is used as the coding method.
[0068] The mesh motion compensation unit 3033 performs motion compensation on each vertex of the reference mesh input from the reference mesh memory 3034 based on the motion information, and outputs a motion-compensated mesh.
[0069] The reference mesh memory 3034 is a memory that holds the decoded mesh for reference in subsequent decoding processes.
[0070] (Mesh displacement decoding) 5 is a functional block diagram showing the configuration of the mesh displacement decoding unit 305. The mesh displacement decoding unit 305 is composed of a displacement unmapper 3052 (image unpacker, displacement decoder), an inverse quantizer 3053, an inverse transformer 3054, and a coordinate system converter 3055. The displacement unmapper may also be called a displacement mapper. As shown in the figure, the mesh displacement decoding unit 305 may further include a video decoder 3051, or the mesh displacement decoding unit 305 may not include the video decoder 3051 and the video decoder 31 may be used to decode the displacement image (displacement sequence). The mesh displacement decoding unit 305 may also not include the inverse quantizer 3053 and image quality control may be performed only by the video decoder 31.
[0071] The atlas information decoding unit 302 decodes, from the encoded data, coordinate system transformation information displacementCoordinateSystem (asps_vdmc_ext_displacement_coordinate_system, afps_vdmc_ext_displacement_coordinate_system) indicating the coordinate system. It may also decode mesh displacement slice division information (displacement slice division parameter, displacement segment parameter, displacement slice division flag, displacement segment flag). The slice division information may be displacementSliceFlag, which indicates whether to divide into segments, or a slice division type displacementSliceType (asps_vdmc_ext_displacement_slice_type, afps_vdmc_ext_displacement_slice_type), which indicates how to divide the segments. The slice division information may also include the component height dispCompHeight. The slice division information may also include the size ctuSize of the block for aligning the slice, or an index ctuSizeIdx indicating the ctuSize. The slice division type is a parameter indicating the type of slice division. Component height is a parameter that indicates the height of the image corresponding to each component (e.g., n, t, b) of the mesh displacement 3D vector. The components n (normal), t (tangent), and b (bitangent) are also called components.
[0072] The slice division information may be a syntax element displacementSliceFlag (displacementSliceEnabledFlag, displacementSliceUsedFlag) indicating that the displacement is divided into slices on a component basis (component basis). Here, "the displacement is divided into slices on a component basis" means that, when encoding the displacement using AVC, HEVC, VVC, or the like, at least the first component (normal) of normal, tangent, and bitangent is encoded in a slice different from the second and third components (tangent and bitangent). For example, displacementSliceFlag==1 may be a case in which the screen is divided into two, with slice 0 being normal and slice 1 being tangent and bitangent. DisplacementSliceFlag==0 is a case in which the screen is not divided into slices or does not explicitly indicate that it is divided into slices. Furthermore, for example, displacementSliceFlag==1 may be a case in which the screen is divided into three, with slice 0 being normal, slice 1 being tangent, and slice 2 being bitangent. It may also be defined that the screen is divided into slices at the boundaries of the displacement components. As will be described later, asps_vdmc_ext_displacement_slice_type may be used to further indicate whether the screen is divided into two slices or three slices. In this way, by indicating in advance that the screen has been divided into slices using an explicit syntax element, the 3D data decoder can determine whether it is a specific slice. By decoding only the entire video, scalable decoding can be achieved according to the decoding capability and power consumption. Note that although slice division is mentioned here, when supporting segments (decoding units) other than slices, such as HEVC or VVC, slices may be interpreted as tiles or subpictures. The same applies below.
[0073] The slice partition information may further include a flag indicating that multiple temporally consecutive frames are partitioned into rectangular regions of the same size and that inter-frame prediction is performed only within the corresponding rectangular regions. For example, tile partitioning using the HEVC Motion Constrained Tile Set (MCTS) may be used. Alternatively, VVC partitioning may be used, which restricts prediction other than from temporally consecutive subpictures. Referenceable subpictures are co-located subpictures with temporal prediction restrictions. Typically, slice partitioning provides independence in that prediction and filtering are not performed with respect to slices other than the slice in question. However, the subpicture is characterized by being independent and not dependent on the slice in the spatial direction or the temporal direction. The information regarding the HEVC tiles or VVC subpictures is referred to as spatiotemporally independent slice partitioning information.
[0074] Note that a separate gating flag may be provided, and each piece of coordinate system transformation information may be decoded only when the gating flag is 1. The gating flag may be, for example, afps_vdmc_ext_overriden_flag. Furthermore, a gating flag may also be provided for slice division information, and the slice division information may be decoded only when the gating flag is 1. The gating flag may be, for example, afps_vdmc_ext_displacement_slice_alignment_flag.
[0075] (Coordinate system) The following two types of coordinate systems are used for mesh displacement (3D vector). Cartesian coordinate system (canonical): A rectangular coordinate system commonly defined throughout the entire 3D space. (X, Y, Z) coordinate system. A rectangular coordinate system whose direction does not change at the same time (within the same frame, within the same tile). Local coordinate system (local): A Cartesian coordinate system defined for each region or vertex in 3D space. A Cartesian coordinate system whose direction can change at the same time (within the same frame, the same tile). Normal (D), tangent (U), bi-tangent (V) coordinate system. That is, it is a Cartesian coordinate system consisting of the first axis (D) indicated by the normal vector n_vec at a vertex (or the surface containing the vertex), and the second axis (U) and third axis (V) indicated by two tangent vectors t_vec and b_vec that are perpendicular to the normal vector n_vec. n_vec, t_vec, and b_vec are three-dimensional vectors. The (D, U, V) coordinate system may also be called the (n, t, b) coordinate system.
[0076] (Decoding and derivation of sequence-level control parameters) Here, the control parameters used in the mesh displacement decoding unit 305 will be explained.
[0077] FIG. 7 shows an example of syntax for transmitting coordinate system transformation parameters and displacement slice division parameters in sequence-level ASPS. ASPS (Atlas Sequence Parameter Set, or Atlas sequence mesh information) is one of the NAL units of atlas information and contains syntax elements that apply to coded atlas sequences. In ASPS, coordinate system transformation parameters and displacement slice division parameters are transmitted using the asps_vdmc_extension() syntax. The semantics of each field are as follows:
[0078] asps_vdmc_ext_displacement_coordinate_system: Coordinate system transformation information indicating the coordinate system of the mesh displacement. If the value is equal to a given first value (e.g., 0), it indicates a Cartesian coordinate system. If the value is equal to another second value (e.g., 1), it indicates a local coordinate system.
[0079] asps_vdmc_ext_displacement_slice_type: Slice of each component of mesh displacement Indicates the type of segmentation. The meaning of the value is as follows: 0:Reservation. 1: Assign the first component of the mesh displacement (e.g., normal, n) to the first slice of the geometry video stream, the second component (e.g., tangent, t) to the second slice of the geometry video stream, and the third component of the geometry video stream (e.g., bi-tangent, b) to the third slice. 2: Assign the first component (eg, n) of the mesh displacement to the first slice, and the second and third components (eg, t, b) to the second slice. 3: Reservation.
[0080] As mentioned above, the following may also be true: 0: Specifies that the mesh displacement is not sliced or is not sliced. 1: Assign the first component of the mesh displacement (e.g., normal) to the first slice of the geometry video stream, the second component (e.g., tangent) to the second slice of the geometry video stream, and the third component (e.g., bitangent) to the third slice of the geometry video stream. As mentioned above, the following may also be true: 0: Specifies that the mesh displacement is not sliced or is not sliced. 1: The first component (e.g., normal) of the mesh displacement is assigned to the first slice, and the second and third components (e.g., tangent, bitangent) are assigned to the second slice.
[0081] asps_vdmc_ext_displacement_component_height: Indicates the image height corresponding to each component of the mesh displacement (e.g., Y, U, V).
[0082] (Decoding and derivation of picture / frame level control parameters) Figure 8 shows an example of syntax for transmitting coordinate system transformation parameters and displacement slice division parameters in AFPS at the picture / frame level. AFPS (Atlas Frame Parameter Set or Atlas frame mesh information) is one of the NAL units of atlas information and contains syntax elements applied to coded atlas frames. In AFPS, coordinate system transformation parameters and displacement slice division parameters are transmitted using the afps_vdmc_extension() syntax. The semantics of each field are as follows:
[0083] afps_vdmc_ext_overriden_flag: A flag indicating whether to update the coordinate system of mesh displacement. If this flag is set to true, the coordinate system of mesh displacement is updated based on the value of afps_vdmc_ext_displacement_coordinate_system, which will be described later. If this flag is set to false, the coordinate system of mesh displacement is not updated.
[0084] afps_vdmc_ext_displacement_coordinate_system: Coordinate system transformation information indicating the coordinate system of the mesh displacement. If the value is equal to the first value (e.g., 0), it indicates a Cartesian coordinate system. If the value is equal to the second value (e.g., 1), it indicates a local coordinate system. If the syntax element is not present, the value is inferred as the value decoded by ASPS, and the default coordinate system is the coordinate system indicated by ASPS.
[0085] afps_vdmc_ext_displacement_slice_alignment_update_flag: A flag indicating whether to update the slice division information of the mesh displacement. If this flag is true (for example, 1), the slice division information of the mesh displacement is updated based on the values of asps_vdmc_ext_displacement_slice_type and asps_vdmc_ext_displacement_component_height, which will be described later. If this flag is false, If it is equal to e (for example, 0), the slice division information of the mesh displacement is not updated.
[0086] afps_vdmc_ext_displacement_slice_type: Indicates the type of slice division for each component of the mesh displacement. The meaning of the values is as described in the semantics of asps_vdmc_ext_displacement_slice_type. If afps_vdmc_ext_displacement_slice_type does not appear, afps_vdmc_ext_displacement_slice_type is set equal to asps_vdmc_ext_displacement_slice_type.
[0087] afps_vdmc_ext_displacement_component_height: indicates the image height corresponding to each component of the mesh displacement 3D vector (e.g., n, t, b). If afps_vdmc_ext_displacement_component_height does not appear, set afps_vdmc_ext_displacement_component_height equal to asps_vdmc_ext_displacement_component_height.
[0088] The mesh displacement decoding unit 305 may derive the coordinate system transformation information displacementCoordinateSystem as follows:
[0089] displacementCoordinateSystem = afps_vdmc_ext_displacement_coordinate_system Note that if afps_vdmc_ext_displacement_coordinate_system does not appear, then afps_vdmc_ext_displacement_coordinate_system shall be set equal to asps_vdmc_ext_displacement_coordinate_system.
[0090] (Derivation of displacement slice division parameters) The mesh displacement decoding unit 305 derives the displacement slice division parameters displacementSliceType and dispCompHeight as follows:
[0091] displacementSliceType = afps_vdmc_ext_displacement_slice_type dispCompHeight = afps_vdmc_ext_displacement_component_height That is, if a syntax element for a displacement slice division parameter appears in the AFPS, the value of the corresponding syntax element in the AFPS is used, and if not, the value of the corresponding syntax element in the ASPS is used.
[0092] Instead of displacementSliceType, a slice division flag displacementSliceFlag indicating whether or not a segment is used may be decoded from the Atlas frame mesh information.
[0093] (Operation of mesh displacement decoding unit) The video decoding unit 3051 decodes a geometry video stream (V3C_GVD video stream) encoded using VVC, HEVC, or the like, and outputs a decoded image (mesh displacement image, mesh displacement array) in which the (quantized) mesh displacement is used as a pixel value. The color components of the geometry are represented using DecGeoChromaFormat. The image may be in YCbCr4:2:0 format. The mesh displacement image may also be a converted mesh displacement image, or it may be the residual of the mesh displacement image.
[0094] The displacement unmapping unit 3052 generates a mesh displacement from a mesh displacement image. Specifically, it derives mesh displacement Qdisp[pos][compIdx], a one-dimensional signal in component compIdx units (=cIdx), from gFrame[compIdx][y][x] of the two-dimensional mesh displacement image according to the correspondence of coordinate positions. Note that gFrame may be an image array DecGeoFrames[mapIdx][frameIdx] or GeoFramesNF[mapIdx][compTimeIdx] decoded from the geometry video stream (V3C_GVD video stream). Here, the correspondence of coordinate positions may be a Z-order scan in units of blocks. NF stands for nominal format, and is the image after adjusting the image size, color sampling, etc. frameIdx and compTimeIdx are the construction time indexes. The name of the array of the mesh displaced image gFrame[compIdx][y][x] may be quantized displaced wavelet coefficients dispQuantCoeffFrame, and the order of the array indexes is not limited to gFrame[compIdx][y][x] but may be dispQuantCoeffFrame[x][y][compIdx] (the same applies below).
[0095] The displacement unmapper 3052 derives DisplacementDim according to the value of a flag asps_vdmc_ext_1d_displacement_flag that indicates whether or not to use the one-dimensional displacement decoded from the coded data.
[0096] DisplacementDim = (asps_vdmc_ext_1d_displacement_flag) ? 1 : 3 Here, asps_vdmc_ext_1d_displacement_flag=1 indicates that only one dimension of the three-dimensional displacement is transmitted. It indicates that the normal component, or x-component (first component) of the displacement is present in the (compressed) geometry image. If the 1D flag is true, the displacement unmapper 3052 infers that the remaining two components are 0. asps_vdmc_ext_1d_displacement_flag=0 indicates that all three components of the displacement are present in the (compressed) geometry image.
[0097] The displacement unmapping unit 3052 derives the number of blocks (blockCount) from the number of mesh displacements (number of points) verCoordCount, and derives the displacement height (dispCompHeight) from blockCount. Blocks are coefficient blocks of displacement, and blockSize is a variable indicating the size of the coefficient block of displacement. Width and height are variables indicating the width and height of the mesh displacement image. pixelsPerBlock = blockSize * blockSize widthInBlocks = width / blockSize shift = (1 << bitDepth) >> 1 blockCount = (verCoordCount + pixelsPerBlock - 1) / pixelsPerBlock heightInBlocks = (blockCount + widthInBlocks - 1) / widthInBlocks In one configuration, the displacement unmapper 3052 may use dispCompHeight, which is decoded from a syntax element (eg, asps_vdmc_ext_displacement_component_height) as described above.
[0098] In another configuration, the displacement unmapper 3052 may derive dispCompHeight using 1 / 3 of the height of the mesh-divided image if the mesh displacement is sliced (if displacementSliceType is a predetermined value, and displacementSliceFlag is true).
[0099] dispCompHeight = height / 3 In another configuration, if the mesh displacement is divided into slices (if displacementSliceType is a predetermined value, or if displacementSliceFlag is true), the displacement unmapper 3052 may derive the image height dispCompHeight corresponding to each component of the mesh displacement 3D vector by aligning it to a predetermined size according to the size of the Coded Tree Unit (CTU Size, ctuSize, videoBlockSize) of the codec used to encode the displacement. Note that this is not limited to the CTU size, and may be a predetermined block size of the mesh image. In this case, it is called videoBlockSize instead of ctuSize. dispCompHeight = heightInBlock * blockSize dispCompHeight = (dispCompHeight + ctuSize - 1) / ctuSize * ctuSize dispCompHeight = (dispCompHeight + ctuSize - 1) & ~(ctuSize - 1) Here, ~ is the bitwise negation operator, which inverts each bit. That is, it may be derived so as to be a constant multiple of a predetermined value ctuSize. Also, blockSize may be set to ctuSize. Also, regardless of the value of displacementSliceFlag, blockSize may be set to ctuSize. Here, the displacement demapping unit 3052 may derive ctuSize from the parameters SPS (Sequence Parameter Set) of the geometry video stream of the codec indicated by gi_geometry_codec_id[DecAtlasID] of V3C. from the SPS.
[0100] The displacement demapping unit 3052 may also decode ctuSize from the encoded data of the NAL unit of the atlas, for example, from the syntax elements of ASPS. Also, it may decode the value of ctuSizeIdx (videoBlockSizeIdx) and derive ctuSize from 16<<ctuSizeIdx or 32<<ctuSizeIdx, 64<<ctuSizeIdx.
[0101] The displacement demapping unit 3052 may use, for example, if gi_geometry_codec_id[DecAtlasID] is HEVC, 64, and if it is VVC, 128 as follows.
[0102] ctuSize = ptl_profile_codec_group_idc == 3 (VVC)? 128 : 64 Here, the value of ptl_profile_codec_group_idc indicates 0: AVC Progressive High, 1: HEVC Main 10, 2: HEVC Main 444, 3: VVC Main 10.
[0103] ctuSize = gi_geometry_codec_id[DecAtlasID] has a 4CC code indicating HEVC? 64 : 128 Here, the 4CC code strings indicating HEVC and VVC are "hev1" and "vvi1", respectively.
[0104] Alternatively, the displacement unmapper 3052 may use the larger of the maximum CTU size of 64 for HEVC and the maximum CTU size of 128 for VVC, as a fixed value, 128.
[0105] Figure 9 shows an example where the mesh displacement image gFrame is divided into segments (slices) equal to the number of dimensions of the geometry. In this example, it is divided into three slices, and the values of the first, second, and third components of QDisp are placed in the first component of gFrame. Here, W = width, H = height.
[0106] In one configuration, the displacement undapping unit 3052 derives a (quantized) mesh displacement array Qdisp from a mesh displaced image gFrame of 3 x height x width. In the following configuration, a 3D data decoding device includes a video decoding unit that decodes mesh displaced images decoded from a geometry video stream whose encoded data Unit Type is V3C_GVD, and a displacement undapping unit that derives mesh displacement QDisp[compIdx][pos] for each position pos and component compIdx from the mesh displaced image gFrame[compIdx][y][x] having the x, y position component compIdx. When the geometry image is 420, the displacement undapping unit derives the mesh displacement by deriving the Y coordinate of the geometry image from the product of a variable compIdx, which ranges from 0 to −1 indicating the number of dimensions DisplacementDim of the geometry, and the height dispCompHeight.
[0107] Here, in the case of a 420 image (DecGeoChromaFormat==1), Qdisp may be derived from gFrame[0][y][x], which is the first component (luminance (Y) image component) of gFrame, according to dispCompHeight. if (!asps_vdmc_ext_1d_displacement_flag) last = dispCompHeight * width - 1 else last = (width * height) - 1 for (v=0; v<verCoordCount; v++) { / / basic process Frame2QDisp v0 = asps_vdmc_ext_packing_method ? last - v : v blockIndex = v0 / pixelsPerBlock indexWithinBlock = v0 % pixelsPerBlock x0 = (blockIndex % widthInBlocks) * blockSize y0 = (blockIndex / widthInBlocks) * blockSize (x, y) = computeMorton2D(indexWithinBlock) x1 = x0 + x y1 = y0 + y for (compIdx=0; compIdx<DisplacementDim; compIdx++) { if (DecGeoChromaFormat==4:2:0) { Qdisp[v0][compIdx] = gFrame[0][compIdx * dispCompHeight + y1][x1] - shift } else { Qdisp[v0][compIdx] = gFrame[compIdx][y1][x1] - shift } } } Here, asps_vdmc_ext_packing_method=0 indicates that the displacement component samples are packed in ascending order. asps_vdmc_ext_packing_method=1 indicates that the displacement component samples are packed in descending order. computeMorton2D is a function for realizing Z-order scan and is defined below. x = extracOddBits(x) { x = x & 0x55555555 x = (x | (x >> 1)) & 0x33333333 x = (x | (x >> 2)) & 0x0F0F0F0F x = (x | (x >> 4)) & 0x00FF00FF x = (x | (x >> 8)) & 0x0000FFFF } (x, y) = computeMorton2D(i) { x = extracOddBits(i>>1) y = extracOddBits(i) } The inverse quantization unit 3053 performs inverse quantization based on the quantization scale value iscale and derives the mesh displacement Tdisp after transformation (e.g., wavelet transform). Tdisp may be in a Cartesian coordinate system or a local coordinate system. iscale is a value derived from the quantization parameter of each component of the mesh displacement image. Tdisp[0][pos] = (Qdisp[0][pos] * iscale[0] + iscaleOffset) >>iscaleShift Tdisp[1][pos] = (Qdisp[1][pos] * iscale[1] + iscaleOffset) >>iscaleShift Tdisp[2][pos] = (Qdisp[2][pos] * iscale[2] + iscaleOffset) >>iscaleShift Here, iscaleOffset = 1<<(iscaleShift-1). iscaleShift may be a predetermined constant, or may be a value decoded from encoded data at the sequence level, picture / frame level, tile / patch level, etc.
[0108] The inverse transform unit 3054 performs an inverse transform g (for example, an inverse wavelet transform) to derive a mesh displacement d. d[0][pos] = g(Tdisp[0][pos]) d[1][pos] = g(Tdisp[1][pos]) d[2][pos] = g(Tdisp[2][pos]) The coordinate system conversion unit 3055 converts the mesh displacement (coordinate system of the mesh displacement) into a Cartesian coordinate system based on the value of the coordinate system conversion information displacementCoordinateSystem. Specifically, when displacementCoordinateSystem = 1, the displacement in the local coordinate system is converted into a displacement in the Cartesian coordinate system. Here, d is a three-dimensional vector indicating the mesh displacement before the coordinate system conversion. disp is a three-dimensional vector indicating the mesh displacement after the coordinate system conversion, and is a Cartesian coordinate system. n_vec, t_vec, and b_vec are three-dimensional vectors (in the Cartesian coordinate system) corresponding to each axis of the local coordinate system of the target region or target vertex. if (displacementCoordinateSystem == 0) { disp = d } else if (displacementCoordinateSystem == 1){ disp = d[0] * n_vec + d[1] * t_vec + d[2] * b_vec } The derivation method shown above for vector multiplication can be expressed individually as scalars as follows: if (displacementCoordinateSystem == 0) { for (i = 0; i < 3; i++) { disp[i] = d[i] } } else if (displacementCoordinateSystem == 1){ for (i = 0; i < 3; i++) { disp[i] = d[0] * n_vec[i] + d[1] * t_vec[i] + d[2] * b_vec[i] } } Alternatively, the same variable name may be assigned before and after the transformation as disp=d, and the value of d may be updated by the coordinate transformation.
[0109] Alternatively, the following configuration may be used. if (displacementCoordinateSystem == 0) { disp = d } else if (displacementCoordinateSystem == 1){ disp = d[0] * n_vec + d[1] * t_vec + d[2] * b_vec } else if (displacementCoordinateSystem == 2){ disp = d[0] * n_vec2 + d[1] * t_vec2 + d[2] * b_vec2 } Here, n_vec2, t_vec2, and b_vec2 are three-dimensional vectors (in the Cartesian coordinate system) corresponding to the axes of the local coordinate system of the adjacent region.
[0110] The following configuration may also be used. if (displacementCoordinateSystem == 0) { disp = d } else if (displacementCoordinateSystem == 1){ disp = d[0] * n_vec3 + d[1] * t_vec3 + d[2] * b_vec3 } Here, n_vec3, t_vec3, and b_vec3 are three-dimensional vectors (in the Cartesian coordinate system) corresponding to the axes of the local coordinate system of the target region with reduced fluctuations. For example, the vectors in the coordinate system used for decoding are derived from the previous and current coordinate systems as follows:
[0111] n_vec3 = (w*n_vec3 + (WT-w)*n_vec)>>wShift t_vec3 = (w*t_vec3 + (WT-w)*t_vec)>>wShift b_vec3 = (w*b_vec3 + (WT-w)*b_vec)>>wShift Here, for example, wShift=2,3,4, WT=1< <wShift、w=1..WT-1。 For example, if w=3 and wShift=3, the following may be used.
[0112] n_vec3 = (3*n_vec3 + 5*n_vec)>>3 t_vec3 = (3*t_vec3 + 5*t_vec)>>3 b_vec3 = (3*b_vec3 + 5*b_vec)>>3 Alternatively, the selection may be made according to the value of the coordinate system transformation information displacementCoordinateSystem decoded from the coded data, as in the following configuration. if (displacementCoordinateSystem == 0) { disp = d } else if (displacementCoordinateSystem == 1){ disp = d[0] * n_vec + d[1] * t_vec + d[2] * b_vec } else if (displacementCoordinateSystem == 6){ disp = d[0] * n_vec3 + d[1] * t_vec3 + d[2] * b_vec3 } (Mesh reconstruction) 6 is a functional block diagram showing the configuration of the mesh reconstruction unit 307. The mesh reconstruction unit 307 is made up of a mesh division unit 3071 and a mesh deformation unit 3072.
[0113] The mesh dividing unit 3071 divides the base mesh output from the base mesh decoding unit 303 to generate divided meshes.
[0114] Figure 10(a) shows a part (triangle) of the base mesh, and the triangle is composed of vertices v1, v2, and v3. v1, v2, and v3 are three-dimensional vectors. The mesh division unit 3071 generates and outputs divided meshes by adding new vertices v12, v13, and v23 to the middle of each side of the triangle (Figure 10(b)). v12 = (v1 + v2) / 2 v13 = (v1 + v3) / 2 v23 = (v2 + v3) / 2 The following is also possible: v12 = (v1 + v2 + 1) >>1 v13 = (v1 + v3 + 1) >>1 v23 = (v2 + v3 + 1) >>1 The mesh deformation unit 3072 receives the division mesh and the mesh displacement, and outputs the mesh displacement d12, A deformed mesh is generated and output by adding d13 and d23 (FIG. 10(c)). The mesh displacement is the output of the mesh displacement decoding unit 305 (coordinate system conversion unit 3055). d12, d13, and d23 are mesh displacements corresponding to the vertices v12, v13, and v23 added by the mesh division unit 3071. v12' = v12 + d12 v13' = v13 + d13 v23' = v23 + d23 It should be noted that d12 = disp[0][], d23 = disp[1][], and d23 = disp[3][] may also be used.
[0115] (Configuration of 3D data encoding device according to the first embodiment) 11 is a functional block diagram showing a schematic configuration of a 3D data encoding device 11 according to the first embodiment. The 3D data encoding device 11 includes an atlas information encoding unit 101, a base mesh encoding unit 103, a base mesh decoding unit 104, a mesh displacement updating unit 106, a mesh displacement encoding unit 107, a mesh displacement decoding unit 108, a mesh reconstruction unit 109, an attribute updating unit 110, a padding unit 111, a color space conversion unit 112, an attribute encoding unit 113, a multiplexing unit 114, and a mesh separation unit 115. The 3D data encoding device 11 receives as input atlas information, a base mesh, a mesh displacement, a mesh, and an attribute image as 3D data, and outputs encoded data.
[0116] The atlas information encoding unit 101 encodes the atlas information and outputs an atlas information encoded stream.
[0117] The base mesh encoding unit 103 encodes the base mesh and outputs a base mesh encoded stream using a coding method such as Draco.
[0118] The base mesh decoding unit 104 is similar to the base mesh decoding unit 303, and therefore a description thereof will be omitted.
[0119] The mesh displacement update unit 106 updates the (original) base mesh and the decoded base mesh. Based on the mesh, adjust the mesh displacement and output the updated mesh displacement.
[0120] The mesh displacement encoding unit 107 encodes the updated mesh displacement and outputs a mesh displacement encoded stream, using a coding method such as VVC or HEVC.
[0121] The mesh displacement decoding unit 108 is similar to the mesh displacement decoding unit 305, and therefore a description thereof will be omitted.
[0122] The mesh reconstruction unit 109 is similar to the mesh reconstruction unit 307, and therefore a description thereof will be omitted.
[0123] The attribute update unit 110 inputs the (original) mesh, the reconstructed mesh output from the mesh reconstruction unit 109 (mesh deformation unit 3072), and the attribute image, updates the attribute image to match the position (coordinates) of the reconstructed mesh, and outputs the updated attribute image.
[0124] The padding unit 111 receives the attribute image and performs padding on areas where pixel values are empty.
[0125] The color space conversion unit 112 performs color space conversion from the RGB format to the YCbCr format.
[0126] The attribute encoding unit 113 encodes the attribute image in YCbCr format output from the color space conversion unit 112, and outputs an attribute video stream. As the encoding method, VVC, HEVC, or the like is used.
[0127] The attribute encoding unit 113 or the atlas information encoding unit 101 may encode tile mapping parameters (tile mapping information, tile mapping SEI) that indicate the correspondence between atlas tiles and video tiles.
[0128] The multiplexing unit 114 multiplexes the atlas information coded stream, base mesh coded stream, mesh displacement coded stream, and attribute video stream and outputs the result as coded data. As a multiplexing method, a byte stream format, ISOBMFF, etc. is used.
[0129] (Mesh separation unit operation) The mesh separator 115 generates a base mesh and a mesh displacement from the mesh.
[0130] 14 is a functional block diagram showing the configuration of the mesh separation unit 115. The mesh separation unit 115 is made up of a mesh thinning unit 1151, a mesh division unit 1152, and a mesh displacement derivation unit 1153.
[0131] The mesh thinning unit 1151 generates a base mesh by thinning out some of the vertices from the mesh.
[0132] Figure 15(a) shows a part of a mesh, which has vertices v1, v2, v3, v4, v5, 15B. The mesh thinning unit 1151 generates and outputs a base mesh by thinning out the vertices v4, v5, and v6 (FIG. 15B).
[0133] The mesh dividing unit 1152 divides the base mesh to generate divided meshes, similar to the mesh dividing unit 3071 (FIG. 15(c)). v4' = (v1 + v2) / 2 v5' = (v1 + v3) / 2 v6' = (v2 + v3) / 2 The mesh displacement derivation unit derives and outputs the displacements d4, d5, d6 of vertices v4, v5, v6 relative to vertices v4', v5', v6' as mesh displacements based on the mesh and the divided meshes (FIG. 15(d)). d4 = v4 - v4' d5 = v5 - v5' d6 = v6 - v6' (Base mesh encoding) 12 is a functional block diagram showing the configuration of the base mesh encoding unit 103. The base mesh encoding unit 103 is composed of a mesh encoding unit 1031, a mesh decoding unit 1032, a motion information encoding unit 1033, a motion information decoding unit 1034, a mesh motion compensation unit 1035, a reference mesh memory 1036, a switch 1037, and a switch 1038. The base mesh encoding unit 103 may also include a base mesh quantization unit (not shown) after inputting the base mesh. When encoding a base mesh without referring to other base meshes (e.g., an already encoded base mesh) (intra-coding), the switches 1037 and 1038 are connected to the side that does not perform motion compensation. When encoding a base mesh with reference to other base meshes (inter-coding), the switches 1037 and 1038 are connected to the side that performs motion compensation.
[0134] The mesh encoding unit 1031 has an intra-encoding function, intra-encodes the base mesh, and outputs a base mesh encoded stream. Draco or the like is used as the encoding method.
[0135] The mesh decoding unit 1032 is similar to the mesh decoding unit 3031, and therefore a description thereof will be omitted.
[0136] The motion information encoding unit 1033 has an inter-encoding function, performs inter-encoding on the base mesh, and outputs a base mesh encoded stream. The encoding method used is entropy encoding such as arithmetic encoding.
[0137] The motion information decoding unit 1034 is similar to the motion information decoding unit 3032, and therefore a description thereof will be omitted.
[0138] The mesh motion compensation unit 1035 is similar to the mesh motion compensation unit 3033, and therefore a description thereof will be omitted.
[0139] The reference mesh memory 1036 is similar to the reference mesh memory 3034, and therefore a description thereof will be omitted.
[0140] (Mesh displacement encoding) 13 is a functional block diagram showing the configuration of the mesh displacement encoding unit 107. The mesh displacement encoding unit 107 is made up of a coordinate system conversion unit 1071, a conversion unit 1072, a quantization unit 1073, and a displacement mapping unit 1074 (image packing unit, displacement encoding unit). As shown in the figure, the mesh displacement encoding unit 107 may further include a video encoding unit 1075. Alternatively, the video encoding unit 1075 may not be included in the mesh displacement encoding unit 107, and an external image encoding device may be used to encode the displaced image.
[0141] The coordinate system conversion unit 1071 converts the coordinate system of the mesh displacement from a Cartesian coordinate system to a coordinate system that encodes the displacement (for example, a local coordinate system) based on the value of the coordinate conversion information displacementCoordinateSystem. Here, disp is a three-dimensional vector indicating the mesh displacement before the coordinate system conversion, d is a three-dimensional vector indicating the mesh displacement after the coordinate system conversion, and n_vec, t_vec, and b_vec are three-dimensional vectors (in the Cartesian coordinate system) indicating each axis of the local coordinate system. if (displacementCoordinateSystem == 0) { d = disp } else if (displacementCoordinateSystem == 1){ d = (disp * n_vec, disp * t_vec, disp * b_vec) } The mesh displacement coding unit 107 may update the value of the displacementCoordinateSystem at the picture / frame level.
[0142] When encoding displacementCoordinateSystem at the sequence level, use the syntax of the configuration in Figure 7. For asps_vdmc_ext_displacement_coordinate_system, set 0 for the Cartesian coordinate system and 1 for the local coordinate system.
[0143] When changing the displacementCoordinateSystem at the picture / frame level, use the syntax in the configuration in Figure 8. For afps_vdmc_ext_overriden_flag, set 1 if you want to update the coordinate system, or 0 if you do not want to update the coordinate system. For afps_vdmc_ext_displacement_coordinate_system, set 0 if you want to use a Cartesian coordinate system, or 1 if you want to use a local coordinate system.
[0144] The transform unit 1072 performs a transform f (for example, a wavelet transform) to derive the transformed mesh displacement Tdisp, for pos=0..NumDisp-1, where NumDisp is the number of mesh vertices. Tdisp[0][pos] = f(d[0][pos]) Tdisp[1][pos] = f(d[1][pos]) Tdisp[2][pos] = f(d[2][pos]) The quantization unit 1073 performs quantization based on the quantization scale value "scale" derived from the quantization parameter of each component of the mesh displacement, and derives the mesh displacement Qdisp after quantization. Qdisp[0][pos] = Tdisp[0][pos] / scale[0] Qdisp[1][pos] = Tdisp[1][pos] / scale[1] Qdisp[2][pos] = Tdisp[2][pos] / scale[2] Alternatively, the scale value may be approximated by a power of 2 and Qdisp may be derived using the following formula: scale[i] = 1 << scale2[i] Qdisp[0][pos] = Tdisp[0][pos] >> scale2[0] Qdisp[1][pos] = Tdisp[1][pos] >> scale2[1] Qdisp[2][pos] = Tdisp[2][pos] >> scale2[2] The displacement mapping unit 1074 generates an image gFrame from the quantized mesh displacement Qdisp based on the value of the displacement mapping parameter displacementChromaLocationType.
[0145] The displacement mapping unit 1074 packs the first element of the (quantized) mesh displacement array, Qdisp[0], into the luma (Y) image component. For x=0..W-1, y=0..H-1, do the following:
[0146] For image width W and height H, (y=0..H-1, x=0..W-1), apply the following, where yc=y / 2 and xc=x / 2. H = dispCompHeight shift = (1 << bitDepth) >> 1 gFrame[0][ y][x] = Qdisp[0][pos] + shift gFrame[0][ H+y][x] = Qdisp[1][pos] + shift gFrame[0][2*H+y][x] = Qdisp[2][pos] + shift pos++ Do the following for xc=0..W / 2-1, yc=0..H / 2-1. gFrame[1][ yc][xc] = shift gFrame[1][H / 2+yc][xc] = shift gFrame[1][ H+yc][xc] = shift gFrame[2][ yc][xc] = shift gFrame[2][H / 2+yc][xc] = shift gFrame[2][ H+yc][xc] = shift Alternatively, you can do the following: for (compIdx = 0; compIdx < 3; compIdx++) { gFrame[0][compIdx*H+y][x] = Qdisp[compIdx][pos] + shift if (compIdx != 0) { gFrame[compIdx][compIdx*H / 2+yc][xc] = shift } } pos++ Alternatively, you can do the following: asps_vdmc_ext_subdivision_iteration_count = lodCount - 1 H = 0 for (i=0; i <lodCount; i++) { H[i] = dispCount[i] / W + 1 asps_vdmc_ext_displacement_component_height_lod[i] = H[i] H += H[i] } shift = (1 << bitDepth) >> 1 y0 = 0 pos0 = 0 for (i=0; i <lodCount; i++) { pos = 0 for (y=0; y <H[i]; y++) { for (x=0; x <W; x++) { if (pos == dispCount[i]) continue gFrame[0][ y0+y][x] = Qdisp[0][pos0+pos] + shift gFrame[0][ H+y0+y][x] = Qdisp[1][pos0+pos] + shift gFrame[0][2*H+y0+y][x] = Qdisp[2][pos0+pos] + shift pos++ } } y0 += H[i] pos0 += dispCount[i] } Alternatively, you can do the following: asps_vdmc_ext_subdivision_iteration_count = lodCount - 1 for (i=0; i <lodCount; i++) { H[i] = (dispCount[i] * 3) / W + 1 asps_vdmc_ext_displacement_component_height_lod[i] = H[i] } shift = (1 << bitDepth) >> 1 y0 = 0 pos0 = 0 for (i=0; i <lodCount; i++) { pos = 0 dim = 0 for (y=0; y <H[i]; y++) { for (x=0; x <W; x++) { if (pos == dispCount[i]) { pos = 0 dim++ } if (dim == 3) continue gFrame[0][y0+y][x] = Qdisp[dim][pos0+pos] + shift pos++ } } y0 += H[i] pos0 += dispCount[i] } The processing may be switched depending on DecGeoChromaFormat. That is, when DecGeoChromaFormat=1 (4:2:0), the above processing is performed, and when DecGeoChromaFormat=3 (4:4:4), the following processing is performed. gFrame[0][y][x] = Qdisp[0][pos] gFrame[1][y][x] = Qdisp[1][pos] gFrame[2][y][x] = Qdisp[2][pos] pos=pos+1 The mesh displacement encoding unit 107 may update the values of displacementSliceType and dispCompHeight at the picture / frame level.
[0147] The meanings of the displacementSliceType and dispCompHeight values are as explained in asps_vdmc_ext_displacement_slice_type and asps_vdmc_ext_displacement_component_height.
[0148] When updating displacementSliceType at the picture / frame level, use the syntax in the configuration in Figure 8. Set the value of displacementSliceType to afps_vdmc_ext_displacement_slice_type. Set the value of dispCompHeight to afps_vdmc_ext_displacement_component_height.
[0149] The video encoding unit 1075 encodes the image in YCbCr4:2:0 format including the (quantized) mesh displacement image, and outputs a mesh displacement encoded stream. The encoding method used is VVC, HEVC, or the like.
[0150] The video encoding unit 1075 may divide the mesh displaced image into slices for each dispCompHeight and encode the slices. Alternatively, the dispCompHeight may be aligned to a predetermined size according to the CTU size.
[0151] The video encoding unit 1075 may encode the mesh displacement by assigning the first component (e.g., D) to the first slice, the second component (e.g., U) to the second slice, and the third component (e.g., V) to the third slice (displacementSliceType=1).
[0152] Alternatively, the video encoding unit 1075 may encode the mesh displacement by allocating the first component to the first slice and the second and third components to the second slice (displacementSliceType=2).
[0153] As described above, by assigning a different slice to each component of the mesh displacement, the decoder can simplify processing by decoding only some of the components, and also achieve scalability by allowing the decoder to decode slices containing the second and third components of the mesh displacement as needed. Furthermore, even if an error is mixed into the coded data, the decoding device can decode only the slices (components) that are free of errors, thereby improving error resistance.
[0154] One embodiment of the present invention has been described in detail above with reference to the drawings, but the specific configuration is not limited to that described above, and various design changes and the like are possible within the scope that does not deviate from the gist of the present invention.
[0155] [Application example] The above-described 3D data encoding device 11 and 3D data decoding device 31 can be mounted on various devices that transmit, receive, record, and play back 3D data. The 3D data may be natural 3D data captured by a camera or the like, or artificial 3D data (including CG and GUI) generated by a computer or the like.
[0156] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. In other words, embodiments obtained by combining technical means modified appropriately within the scope of the claims are also included in the technical scope of the present invention. [Industrial Applicability]
[0157] The embodiments of the present invention can be suitably applied to a 3D data decoding device that decodes coded data in which 3D data has been coded, and a 3D data coding device that generates coded data in which 3D data has been coded, and can also be suitably applied to the data structure of coded data that is generated by the 3D data coding device and referenced by the 3D data decoding device. [Explanation of symbols]
[0158] 11 3D data encoding device 101 Atlas Information Encoding Unit 103 Base mesh coding unit 1031 Mesh coding unit 1032 Mesh Decoding Unit 1033 Motion information encoding unit 1034 Motion information decoding unit 1035 Mesh motion compensation unit 1036 reference mesh memory 1037 Switch 1038 Switch 104 Base mesh decoding unit 106 Mesh displacement update section 107 Mesh displacement coding unit 1071 Coordinate system conversion unit 1072 Conversion Unit 1073 Quantization section 1074 Displacement Mapping Section 1075 Video Encoding Unit 108 Mesh displacement decoding unit 109 Mesh reconstruction unit 110 Attribute Update Section 111 Padding section 112 Color space conversion unit 113 Attribute Encoding Unit 114 Multiplexer 115 mesh separation section 1151 Mesh thinning section 1152 Mesh division section 1153 Mesh displacement derivation part 21 Network 31 3D data decoding device 301 Demultiplexer 302 Atlas Information Decoding Unit 303 Base mesh decoding unit 3031 Mesh Decoding Unit 3032 Motion information decoding unit 3033 Mesh Motion Compensation Unit 3034 Reference Mesh Memory 3035 Switch 3036 Switch 305 Mesh displacement decoding unit 3051 Video Decoding Unit 3052 Displacement Unmapping Unit 3053 Inverse quantization section 3054 Inverse Conversion Unit 3055 Coordinate system conversion unit 307 Mesh Reconstruction Unit 306 Attribute Decoding Unit 3071 Mesh division section 3072 Mesh deformation part 308 Color Space Conversion Unit 41 3D data display device
Claims
1. A 3D data decoding device that decodes encoded data, an atlas information decoding unit that decodes parameters related to tile division of an atlas frame from an atlas data stream whose Unit Type of the encoded data is V3C_AD; an attribute decoding unit that decodes an attribute image from an attribute data stream in which the Unit Type of the encoded data is V3C_AVD, The 3D data decoding device is characterized in that the attribute decoding unit decodes tile mapping parameters that indicate the correspondence between atlas tiles that indicate the partition division of the atlas frame and video tiles that indicate the partition division of the video codec of the attribute image.
2. 2. The 3D data decoding device according to claim 1, wherein the tile mapping parameters include a flag indicating whether the atlas tile and the video tile match.
3. 3. The 3D data decoding device according to claim 1, wherein the tile mapping parameters include a flag indicating whether the atlas tile includes multiple of the video tiles.
4. 4. The 3D data decoding device according to claim 1, wherein the tile mapping parameters include the number of the video tiles included in the atlas tile.
5. 5. The 3D data decoding device according to claim 1, wherein the tile mapping parameters include IDs of the video tiles included in the atlas tiles.
6. A 3D data encoding device that encodes 3D data, an atlas information encoding unit that encodes parameters related to tile division of an atlas frame; an attribute encoding unit that encodes the attribute image; The 3D data encoding device is characterized in that the attribute encoding unit encodes tile mapping parameters indicating the correspondence between atlas tiles indicating the partition division of the atlas frame and video tiles indicating the partition division of the video codec of the attribute image.