3D data decoding device and 3D data encoding device
The 3D data decoding device addresses inefficiencies in existing methods by using an atlas information decoding unit to decode tile relationships, facilitating parallel processing and improving encoding and decoding efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SHARP KK
- Filing Date
- 2024-10-24
- Publication Date
- 2026-05-12
Smart Images

Figure 2026076618000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to a 3D data encoding device and a 3D data decoding device.
Background Art
[0002] In order to efficiently transmit or record 3D data, a 3D data encoding device that converts 3D data into a two-dimensional image and encodes it using a video image encoding method to generate encoded data, and a 3D data decoding device that decodes a two-dimensional image from the encoded data to reconstruct 3D data exist.
[0003] [[ID=Heights]] Specific 3D data encoding methods include, for example, ISO / IEC 23090-5 V3C (Volumetric Video-based Coding) and V-PCC (Video-based Point Cloud Compression) of MPEG-I. V3C can encode and decode a point cloud composed of point positions and attribute information. Further, ISO / IEC 23090-12 (MPEG Immersive Video, MIV) and ISO / IEC 23090-29 (Video-based Dynamic Mesh Coding, V-DMC) under standardization are also used for encoding and decoding multi-view video and mesh video. The V-DMC method is disclosed in the latest draft document of Non-Patent Document 1.
[0004] In these 3D data encoding methods, the geometry and attributes constituting the 3D data are encoded and decoded using a moving image encoding method such as H.265 / HEVC (High Efficiency Video Coding) or H.266 / VVC (Versatile Video Coding) as an image.
[0005] In the case of point clouds, the geometry image is the depth to the projection plane, and the attribute image is the image of the attributes projected onto the projection plane.
[0006] 3D data (mesh) like that in Non-Patent Document 1 consists of a base mesh, mesh displacement, and texture mapping image. The base mesh is encoded using vertex codes such as Draco. A coding method can be used. The encoding of mesh displacement is done by converting the mesh displacement into a 2D mesh. Besides encoding the displacement image using a video codec, there is also a method of direct encoding using arithmetic coding. The texture mapping image is encoded as an attribute image using a video codec. The HEVC and VVC video codecs mentioned above can be used. [Prior art documents] [Non-patent literature]
[0007] [Non-Patent Document 1] Study of technologies for Video-based mesh coding, ISO / IEC JTC 1 / SC 29 / WG 7 N0960, 2024-09-03 [Overview of the project] [Problems that the invention aims to solve]
[0008] In the 3D data encoding method described in Non-Patent Literature 1, the relationship between the tiles that divide the mesh submesh and point cloud, the mesh geometry, and attribute pictures, and the tiles of the video codec is unclear. Therefore, when decoding a specific submesh or specific tile, it is necessary to decode all the tiles of the video codec, which presents a problem.
[0009] The present invention aims to efficiently encode and decode non-redundant 3D data by decoding the correspondence between a specific tile and the codec tile to be decoded, thereby enabling the parallel decoding of the codec tile necessary to decode a particular tile, and shortening the time required to decode a specific tile. [Means for solving the problem]
[0010] To solve the above problems, a 3D data decoding device according to one aspect of the present invention is a 3D data decoding device that decodes mesh data or point cloud data, and comprises an atlas information decoding unit that decodes atlas information from encoded data in which the mesh data or point cloud data has been encoded, wherein the atlas information decoding unit decodes tile information and additional information indicating the relationship between tiles and codec tiles.
[0011] A 3D data decoding device for decoding mesh data or point cloud data is provided with an atlas information decoding unit that decodes atlas information from encoded data in which the mesh data or point cloud data has been encoded, and the atlas information decoding unit is characterized by decoding attribute tile information and additional information indicating the relationship between attribute tiles and codec tiles.
[0012] The above additional information is characterized by being a mapping flag that indicates the correspondence between tile information and codec tile information.
[0013] The above additional information is characterized by being an index of codec tiles that shows the correspondence between tile information and codec tile information.
[0014] The above additional information includes a flag indicating whether the tile division and the codec tile division are the same, and in the above atlas information decoding unit, each i-th tile is a displacement subbits If it contains only the region of the i-th codec tile of the trim, decode the first value, and then In the case of an external environment, the second value is decoded.
[0015] The additional information includes a flag indicating whether the division of the attribute tile and the division of the codec tile are the same. In the atlas information decoding unit, when each i-th attribute tile includes only the area of the i-th codec tile of the attribute video sub-bitstream, the first value is decoded, and in other cases, the second value is decoded. is characterized by
[0016] In order to solve the above problems, a 3D data encoding device according to an aspect of the present invention is a 3D data encoding device that decodes mesh data or point cloud data, and includes an atlas information encoding unit that encodes atlas information from the encoded data in which the mesh data or point cloud data is encoded. The atlas information encoding unit is characterized by encoding tile information and additional information indicating the relationship between the tile and the codec tile.
[0017] In a 3D data encoding device that decodes mesh data or point cloud data, an atlas information encoding unit that encodes atlas information from the encoded data in which the mesh data or point cloud data is encoded is provided. The atlas information encoding unit is characterized by encoding attribute tile information and additional information indicating the relationship between the attribute tile and the codec tile.
Effect of the Invention
[0018] According to one aspect of the present invention, the flexibility of tile definition can be enhanced, and 3D data can be encoded and decoded with high quality.
Brief Description of the Drawings
[0019] [Figure 1] It is a schematic diagram showing the configuration of a 3D data transmission system according to this embodiment. [Figure 2] It is a diagram showing the hierarchical structure of the data of the encoded stream. [Figure 3] It is a functional block diagram showing the schematic configuration of the 3D data decoding device 31. [Figure 4] It is a functional block diagram showing the configuration of the atlas information decoding unit 302. [Figure 5] It is a functional block diagram showing the configuration of the base mesh decoding unit 303. [Figure 6] It is a functional block diagram showing the configuration of the mesh displacement decoding unit 305. [Figure 7] It is a functional block diagram showing the configuration of the mesh reconstruction unit 307. [Figure 8] It is an example of the syntax of ASVE (ASPS Vdmc Extension), which is a sequence-level mesh data extension coding parameter set. [Figure 9] It is an example of the syntax of the extended coding parameter information in AFPS, which is a picture / frame-level parameter set. [Figure 10] It is an example of the syntax of the atlas style information in the atlas frame. [Figure 11] It is an example of the syntax of the attribute tile information in the atlas frame. [Figure 12] It is a diagram for explaining the operation of the mesh reconstruction unit 307. [Figure 13] It is a functional block diagram showing the schematic configuration of the 3D data encoding device 11. [Figure 14] It is a functional block diagram showing the configuration of the atlas information encoding unit 101. [Figure 15] It is a functional block diagram showing the configuration of the base mesh encoding unit 103. [Figure 16] It is a functional block diagram showing the configuration of the mesh displacement encoding unit 107. [Figure 17] It is a functional block diagram showing the configuration of the mesh separation unit 115. [Figure 18] It is a diagram for explaining the operation of the mesh separation unit 115. [Figure 19] It is a diagram for explaining the mapping flag showing the relationship between the tile and the codec tile. [Figure 20]This diagram illustrates the codec tile index, showing the relationship between tiles and codec tiles. [Figure 21] This diagram illustrates the mapping flags that show the relationship between attribute tiles and codec tiles included in attribute video data. [Figure 22] This diagram illustrates the codec tile index, which shows the relationship between attribute tiles and codec tiles included in attribute video data. [Figure 23] This is an example of the syntax for submesh tile information SEI. [Figure 24] This is another example of the syntax for submesh tile information SEI. [Figure 25] This is an example of the syntax for submesh attribute tile information SEI. [Figure 26] This is another example of the syntax for submesh attribute tile information SEI. [Figure 27] This is an example of alignment flag syntax when the number of divisions for each attribute video data is the same. [Modes for carrying out the invention]
[0020] Embodiments of the present invention will be described below with reference to the drawings.
[0021] Figure 1 is a schematic diagram showing the configuration of the 3D data transmission system 1 according to this embodiment.
[0022] The 3D data transmission system 1 is a system that transmits an encoded stream containing encoded 3D data to be encoded, decodes the transmitted encoded stream, and displays the 3D data. The 3D data transmission system 1 consists of a 3D data encoding device 11, a network 21, a 3D data decoding device 31, and a 3D data display device 41.
[0023] The 3D data encoding device 11 receives 3D data T as input.
[0024] Network 21 transmits the encoded stream Te generated by the 3D data encoding device 11 to the 3D data decoding device 31. Network 21 is the Internet, Wide Area Network (WAN), Local Area Network (LAN), or These are combinations. Network 21 is not necessarily limited to a bidirectional communication network; it may also be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. Network21 is a manufacturer of DVD (Digital Versatile Disc: registered trademark) and BD (Blu-ray Disc: registered trademark). This may be replaced by a storage medium that records an encoded stream Te such as the standard.
[0025] The 3D data decoding device 31 decodes each of the encoded streams Te transmitted by the network 21 and generates one or more decoded 3D data Td.
[0026] The 3D data display device 41 displays all or part of one or more decoded 3D data Td generated by the 3D data decoding device 31. The 3D data display device 41 includes, for example, a display device such as a liquid crystal display or an organic EL (electro-luminescence) display. Examples of display forms include stationary, mobile, and HMDs. Furthermore, the 3D data decoding device 31 is high If the system has sufficient processing power, it will display high-resolution images; if it has less processing power, it will display images that do not require high processing or display capabilities.
[0027] <operators> The operators used in this specification are listed below.
[0028] >> is a right bit shift, << is a left bit shift, & is a bitwise AND, and | is a bitwise OR. |= is the OR assignment operator, and || represents logical disjunction.
[0029] x?y:z is a ternary operator that takes the value y if x is true (non-zero) and z if x is false (0).
[0030] y..z represents a set of integers from y to z.
[0031] Log2(x) is a function that returns the base 2 logarithm of x.
[0032] Ceil(x) is a function that returns the smallest integer greater than or equal to x.
[0033] Floor(x) is a function that returns the largest integer less than or equal to x.
[0034] The Sign(x) function returns 1 if x is greater than 0, 0 if x is equal to 0, and -1 if x is less than 0.
[0035] Abs(x) is a function that returns the absolute value of x.
[0036] The Round(x) function returns an integer obtained by rounding x to the first decimal place.
[0037] Round( x ) = Sign( x ) * Floor( Abs( x ) + 0.5 ).
[0038] The operator ' / ' is integer division, truncating towards zero. For example, 7 / 4 is truncated to 1, and -7 / 4 is truncated to -1.
[0039] The division operator (÷) performs division without rounding or truncation.
[0040] <Structure of the coded stream Te> Prior to a detailed description of the 3D data encoding device 11 and the 3D data decoding device 31 according to this embodiment, the data structure of the encoded stream Te generated by the 3D data encoding device 11 and decoded by the 3D data decoding device 31 will be described. The 3D data may be ISO / IEC 23090-5 V3C (Volumetric Video-based Coding) and V-PCC (Video-based Point Cloud Compression) of MPEG-I, and ISO / IEC 23090-12 (MPEG Immersive Video, MIV) and ISO / IEC 23090-29 (Video-based Dynamic Mesh Coding, V-DMC) based on V3D.
[0041] Figure 2 shows the hierarchical structure of data in the encoded stream Te. A Te object has either a V3C sample stream or a V3C unit stream data structure. A V3C sample stream includes a sample stream header and a V3C unit. A V3C unit stream includes a V3C unit.
[0042] A V3C unit includes a V3C unit header and a V3C unit payload. The V3C unit header is the Unit Type, which is an ID indicating the type of V3C unit, such as V3C_VPS, V3C_AD, V3C_AVD, V3C_GVD, V3C_OVD. Which label will indicate the value?
[0043] If the Unit Type is V3C_VPS (Video Parameter Set), the V3C unit includes the V3C parameter set.
[0044] When the Unit Type is V3C_AD (Atlas Data), the V3C unit includes the VPS ID, atlasID, sample stream NAL header, and multiple NAL units. The atlasID is an ID (Identification). It can take a non-negative integer value.
[0045] A NAL unit includes NALUnitType, layerID, TemporalID, and RBSP (Raw byte sequence payload).
[0046] NAL units are identified by NALUnitType, ASPS (Atlas Sequence Parameter Set), and AAPS. This includes Atlas Adaptation Parameter Set (ATLAS), Atlas Tile layer (ATLAS), and Supplemental Enhancement Information (SEI).
[0047] An ATL file includes an ATL header and an ATL data unit, which contains information such as the location and size of a patch, including patch information data.
[0048] SEI includes payloadType, which indicates the type of SEI; payloadSize, which indicates the size (in bytes) of the SEI; and sei_payload, which contains the SEI data.
[0049] If the Unit Type is V3C_AVD (Attribute Video Data), then V3C The unit includes the VPS ID, atlasID, attrIdx (attribute image ID), partIdx (partition ID), mapIdx (map ID), auxFlag (flag indicating whether it is auxiliary data), and video stream. The video stream is data encoded in HEVC, VVC, etc. Attribute data In V-DMC, this corresponds to a texture image. attrIdx can be an integer between 0 and ai_attribute_count[RecAtlasID]-1, where ai_attribute_count is the syntax element of attribute_information and RecAtlasID is the ID of the target atlas (atlasID).
[0050] Here, ai_attribute_count[j] indicates the number of attributes associated with the atlas ID at index j. If none exist, the value of ai_attribute_count[j] is assumed to be 0.
[0051] When NalUnitType is V3C_GVD (Geometry Video Data), the V3C unit includes VPS ID, atlasID, mapIdx, auxFlag, and video stream. In V-DMC, the geometry data corresponds to mesh displacement.
[0052] If the Unit Type is V3C_OVD (Occupancy Video Data), the V3C unit includes the VPS ID, atlasID, and video stream.
[0053] If the Unit Type is V3C_MD (Mesh data), the V3C unit is the VPS ID, atl Includes asID and mesh_payload. V-DMC supports base mesh.
[0054] (Configuration of the 3D data decoding device according to the first embodiment) Figure 3 is a functional block diagram showing the schematic configuration of the 3D data decoding device 31 according to the first embodiment. The 3D data decoding device 31 consists of a demultiplexing unit 301, an atlas information decoding unit 302, a base mesh decoding unit 303, a mesh displacement decoding unit 305, a mesh reconstruction unit 307, an attribute decoding unit 306, and a color space conversion unit 308. The 3D data decoding device 31 receives encoded 3D data as input. It outputs additional information, atlas information, mesh, and attribute images.
[0055] The demultiplexing unit 301 receives encoded data multiplexed in a byte stream format, ISOBMFF (ISO Base Media File Format), etc., demultiplexes it, and outputs an atlas information encoded stream (V3C_AD's Atlas Data stream, NALunit), a base mesh encoded stream (V3C_MD's mesh_payload), a mesh displacement encoded stream (geometry video stream, V3C_GVD's geometry video stream), and an attribute video stream (V3C_AVD's attribute video stream). The geometry video stream and attribute video stream are 2D. This is a video, and HEVC and VVC video codecs may be used for encoding and decoding. Furthermore, the codecs for the following mesh displacement stream and attribute video stream may be video codecs.
[0056] The Atlas information decoding unit 302 receives the Atlas information encoded stream output from the demultiplexing unit 301 and decodes the Atlas information.
[0057] The base mesh decoding unit 303 decodes the base mesh encoding stream encoded with vertex coding (3D data compression encoding scheme, e.g., Draco) and outputs the base mesh. The base mesh will be described later. The codec type of the base mesh may be obtained by decoding the syntax elements bmsps_intra_mesh_codec_id and bmsps_inter_mesh_codec_id.
[0058] The mesh displacement decoding unit 305 decodes the mesh displacement coding stream and calculates the mesh displacement. Output. The type of codec used for encoding is determined by the V3C parameter set of the encoded data. This is indicated by the ptl_profile_codec_group_idc obtained by the command. Alternatively, it may be indicated by the Four CC code (4-character code, 4CC code) indicated by gi_geometry_codec_id[atlasID] in the V3C parameter set. gi_geometry_codec_id[atlasID] is the geometry video story in the atlas ID. This shows the index corresponding to the codec ID of the decoder used to decrypt the data. Also, even if you decrypt the syntax element dsps_codec_id which indicates the type of codec from the parameter set... Good. The set showing the correspondence between the codec ID (ccm_codec_id) and its 4CC code (ccm_codec_4cc[ccm_codec_id]) may be transmitted using a separate codec mapping SEI (component_codec_mapping SEI).
[0059] The mesh reconstruction unit 307 takes the base mesh and mesh displacement as input and reconstructs the mesh in 3D space. Reconstruct the layout.
[0060] The attribute decoding unit 306 decodes the attribute video stream encoded with VVC, HEVC, etc., and outputs an attribute image. The attribute image is a texture image unfolded along the UV axis (a texture mapping image converted using the UV atlas method) with YCbCr formatting. It may also be in a format. The type of codec used for encoding is indicated by ptl_profile_codec_group_idc, which is obtained by decoding the V3C parameter set of the encoded data. Alternatively, it may be indicated by the Four CC code indicated by ai_attribute_codec_id[atlasID] in the V3C parameter set. ai_attribute_codec_id[atlasID] indicates the index in the atlas ID that corresponds to the codec ID of the decoder used to decode the attribute video stream.
[0061] The color space conversion unit 308 converts the attribute image from YCbCr format to RGB format. Perform a color space conversion. Note that the attribute video stream encoded in RGB format... A configuration that decodes the image and omits color space conversion is also acceptable.
[0062] (Decryption of Atlas information) Figure 4 is a functional block diagram showing the configuration of the atlas information decoding unit 302. The atlas information decoding unit 302 includes a parameter decoding unit 3021, a tile information decoding unit 3022, an extended information decoding unit 3023, and additional information It consists of a signal decoding unit 3024.
[0063] (Decoding and derivation of coding parameters) The parameter decoding unit 3021 decodes the encoded parameters from the Atlas information encoded stream. The encoded parameters include the sequence-level parameter set ASPS (Atlas Sequence Parameter Set) and the picture / frame-level parameter set AFPS (Atlas Frame Parameter Set).
[0064] Figure 8 shows an example of the syntax for ASVE (ASPS Vdmc Extension), a sequence-level mesh data extended coding parameter set. The semantics of each element of the syntax The details are as follows:
[0065] asve_subdivision_iteration_count: Indicates the number of iterations for mesh subdivision.
[0066] asve_1d_displacement_flag: This flag indicates whether the mesh displacement is one-dimensional or not. A value of true indicates that the mesh displacement is one-dimensional. A value of false indicates that the mesh displacement is three-dimensional.
[0067] asve_num_attribute_video: Indicates the number of attributes signaled through the video subbitstream. The value of asve_num_attribute_video is ai_attribute_count[j] The V3C bitstream compliance requirement is that the value is equal to [this value].
[0068] (Decoding and derivation of tile encoding parameters) This section describes the atlas style information (tile selection information, tile division information), which is an encoding parameter that defines the tile to be decoded from the encoded data by the tile information decoding unit 3022. V3C The standard allows defining a common screen division (partition division) for the atlas frame, occupancy frame, geometry frame, and attribute frame as a tile. Since the information representing the atlas, occupancy, geometry, and attributes is the atlas, the tiles defined here may be called atlas styles, and the tile information may be called atlas style information. Furthermore, if a tile dedicated to a specific component, such as an attribute, is defined, it is called an attribute tile. Note that the unit of a tile is a rectangle, and (common to atlas style information and attribute tile information) the tile definition may include the number of columns, the number of rows, the width of the tiles in a given column, and the height of the tiles in a given row that make up the screen.
[0069] Figure 10 shows an example of the syntax for tile information in the Atlas Frame Parameter Set (AFPS), which is a picture / frame-level parameter set. The semantics of each syntax element are as follows. When dividing the atlas frame into tiles, the occupancy frame, geometry frame, and attribute frame are also divided into tiles in the same way. However, the attribute frame can be divided into independent attribute tiles as described later.
[0070] afti_single_tile_in_atlas_frame_flag: Tiles in each atlas frame that references AFPS A flag indicating whether only one tile exists. If the value is true, it means that there is only one tile in each atlas frame referencing the AFPS. If the value is false, it means that there are multiple (greater than 1) tiles in each atlas frame referencing the AFPS.
[0071] afti_single_partition_per_tile_flag: Each tile references AFPS A flag indicating whether or not there is only one tile partition. If the value is true, it indicates that each tile referencing AFPS contains only one tile partition, and if the value is false, AFPS Indicates that each tile referenced contains multiple (greater than 1) tile partitions. If none exist, the value of afti_single_partition_per_tile_flag is presumed to be equal to 1.
[0072] afti_num_tiles_in_atlas_frame_minus1: Specifies the number of tiles in each atlas frame referencing AFPS. The value of afti_num_tiles_in_atlas_frame_minus1 must be in the range of 0 to NumPartitionsInAtlasFrame-1. If it does not exist and afti_single_partition_per_tile_flag is equal to 1, the value of afti_num_tiles_in_atlas_frame_minus1 is presumed to be equal to NumPartitionsInAtlasFrame-1.
[0073] afti_signalled_tile_id_flag: Indicates whether the tile ID of each tile is signaled or not. A flag. If the flag is equal to 1, the tile ID of each tile is signaled. Flag If it is equal to 0, do not signal the tile ID.
[0074] afti_signalled_tile_id_length_minus1:afti_signalled_tile_id_length_minus1+1 is the syntax element afti_tile_id[i] (if it exists) of the tile header and syntax Specifies the number of bits used to represent the element ath_id. The value of afti_signalled_tile_id_length_minus1 must be in the range of 0 to 15.
[0075] afti_tile_id[i]: Specifies the tile ID of the i-th tile. If it does not exist, the value of afti_tile_id[i] is assumed to be equal to i for each i in the range from 0 to afti_num_tiles_in_atlas_frame_minus1. afti_tile_id[i] is equal to afti_tile_id[j] for all i != j. The requirement for bitstream conformance is that the bitstreams are not equal (they must not be equal). The 3D data decoding device 31 decodes a bitstream that satisfies the conformance requirement (the same applies hereafter).
[0076] When the tile information decoding unit 3022 decodes and encodes afti_single_tile_in_atlas_frame_flag and afti_single_partition_per_tile_flag, it decodes the syntax element afti_num_tiles_in_atlas_frame_minus2, which indicates the number of tiles minus 2 (the number of tiles minus 2). It may be encoded. Alternatively, the value of afti_single_tile_in_atlas_frame_flag is false. If the value of afti_single_partition_per_tile_flag is false, then the syntax element afti_num_tiles_in_atlas_frame_minus2, which indicates the number of tiles referenced minus 2, is decoded and coded. You may use numbering. The following semantics may also be used.
[0077] afti_num_tiles_in_atlas_frame_minus2: Specifies the number of tiles in each atlas frame that references the Atlas Frame Parameter Set AFPS. The value of afti_num_tiles_in_atlas_frame_minus1 must be in the range of 0 to NumPartitionsInAtlasFrame-2. If it does not exist and afti_single_partition_per_tile_flag is equal to 1, the value of afti_num_tiles_in_atlas_frame_minus2 is presumed to be equal to NumPartitionsInAtlasFrame-2.
[0078] In this configuration, if there is only one tile being referenced, afti_single_tile_in_atlas_frame_flag Since it can be represented by this, the syntax element indicating the number of tiles minus 2 is decoded and the symbol By using a numbering system, the overhead of the coding amount is reduced.
[0079] (Decoding and derivation of extended coding parameters) The extended coding parameters decoded from the coded data by the extended information decoding unit 3023 will be described below.
[0080] Figure 9 shows the extended encoding in AFPS, which is a picture / frame level parameter set. This is an example of the syntax for parameter information.
[0081] afve_overriden_flag: This flag indicates whether or not to update the coordinate system of the mesh displacement. If this flag is equal to true, the coordinate system of the mesh displacement will be updated based on the value of mdu_displacement_coordinate_system described below. If this flag is equal to false, the coordinate system of the mesh displacement will be updated. The system will not be updated.
[0082] afve_subdivision_iteration_count: Indicates the number of mesh subdivision iterations.
[0083] (Decoding and derivation of attribute tile-level coding parameters) This section explains the "attribute-level tile definition (attribute tile information)," which is an attribute tile-level encoding parameter decoded from the encoded data by the extended information decoding unit 3023. In the V3C standard, the partitioning common to each data is called an attribute tile. Although defined as such, there may be partitioning that applies only to attribute data. Such tiles are called "attribute tiles," and their definitions are called "attribute tile information." The syntax elements and parameters of attribute tile information are basically the same as those of attribute tile information, but differ in that their application is limited to attributes, and they are encoded and decoded on an attribute-by-attribute basis (for each attrIdx).
[0084] Figure 11 shows an example of the syntax for attribute tile information coding parameters in AFPS, which is a picture / frame-level parameter set.
[0085] afati_single_tile_in_atlas_frame_flag[attrIdx]: If afati_single_tile_in_atlas_frame_flag[attrIdx] is equal to 1, it specifies that there is only one tile signaled for the attribute in the attribute video data unit with index attrIdx. If afati_single_tile_in_atlas_frame_flag[attrIdx] is equal to 0, it specifies that there are two or more tiles signaled for the attribute in the attribute video data unit with index attrIdx.
[0086] If afati_uniform_partition_spacing_flag[attrIdx] is equal to 1, it specifies that the atlas tiling for the signaled attribute in the attribute video data unit with index attrIdx will use a method that uniformly divides the column and row boundaries across the attribute atlas frame (attribute frame). Information corresponding to these boundaries is signaled using the syntax elements afati_partition_cols_width_minus1[attrIdx] and afati_partition_rows_height_minus1[attrIdx], respectively. If afati_uniform_partition_spacing_flag[attrIdx] is equal to 0, it specifies that the atlas tiling for signaled attributes in the attribute video data unit with index attrIdx will use a method that may result in column and row boundaries that are not uniformly divided across atlas frames. In this case, these boundaries are defined by the syntax elements afati_num_partition_columns_minus1[attrIdx] and afati_num_partition_rows_minus1[attrIdx], as well as the syntax elements afati_partition_column_width_minus1[attrIdx][i] and afati_partition_row_height_minus1[attrIdx][i]. Signaling is performed using afati_ti_uniform_partition_spacing_fl. If it does not exist, afati_ti_uniform_partition_spacing_fl The value of ag[attrIdx] is presumed to be equal to 1.
[0087] The value obtained by adding 1 to afati_partition_cols_width_minus1[attrIdx] is equal to afati_uniform_partition_spacing_flag[attrIdx] = 1. At that time, the attribute video data unit with index attrIdx is extracted, excluding the attribute tile partition column at the right edge of the attribute atlas frame in units of 64 samples. Specifies the width of the tile partition column (the width of the tile column) for the knit attribute. The value of afati_partition_cols_width_minus1[attrIdx] should be in the range of 0 to asve_attribute_frame_width[attrIdx] / 64 - 1. If it does not exist, the value of afati_partition_cols_width_minus1[attrIdx] is assumed to be equal to asve_attribute_frame_width[attrIdx] / 64 - 1. It can be done.
[0088] The value obtained by adding 1 to afati_partition_rows_height_minus1[attrIdx] is equivalent to afati_uniform_partition_spacing_flag[attrIdx] being 1. When this happens, the attribute video data of index attrIdx is obtained, excluding the attribute tile partition row at the bottom of the 64-sample unit attribute atlas frame. Specifies the height of the unit attribute tile partition row. The value of afati_partition_rows_height_minus1[attrIdx] should be in the range of 0 to asve_attribute_frame_height[attrIdx] / 64 - 1. If it does not exist, the value of afati_partition_rows_height_minus1[attrIdx] is It is estimated to be equal to asve_attribute_frame_height[attrIdx] / 64 - 1.
[0089] The value of afati_num_partition_columns_minus1[attrIdx] plus 1 is equivalent to afati_uniform_partition_spacing_flag[attrIdx] being 0. In the case of a problem, the number of attribute tile partition columns of attribute video data with index attrIdx used to divide the attribute frame is Specify the value of afati_num_partition_columns_minus1[attrIdx], which ranges from 0 to asve_attribute_frame_width[attrIdx] / 64 - 1. If afati_single_tile_in_atlas_frame_flag[attrIdx] is 1, the value of afati_num_partition_columns_minus1[attrIdx] is assumed to be equal to 0.
[0090] The value obtained by adding 1 to afati_num_partition_rows_minus1[attrIdx] is equal to afati_uniform_partition_spacing_flag[attrIdx] if afati_uniform_partition_spacing_flag[attrIdx] is equal to 0. This refers to the number of attribute tile partition rows of attribute video data with index attrIdx, which are used to divide the attribute atlas frame. The value of afati_num_partition_rows_minus1[attrIdx] is set to a range from 0 to asve_attribute_frame_height[attrIdx] / 64 - 1. If afati_single_tile_in_atlas_frame_flag[attrIdx] is 1, the value of afati_num_partition_rows_minus1[attrIdx] is assumed to be equal to 0. It will be done.
[0091] The value obtained by adding 1 to afati_partition_column_width_minus1[attrIdx][i] specifies the width of the i-th attribute tile partition column of the attribute video data with index attrIdx in units of 64 samples.
[0092] The value obtained by adding 1 to afati_partition_row_height_minus1[attrIdx][i] specifies the height in 64 samples of the i-th attribute tile partition row of the attribute video data with index attrIdx.
[0093] afati_single_partition_per_tile_flag[attrIdx][ i ]: If afati_single_partition_per_tile_flag[attrIdx] is equal to 1, the attribute video with index attrIdx Each attribute tile of an attribute shown in data units is one tile partition. Specifies that it includes a partition. If afati_single_partition_per_tile_flag[attrIdx] is equal to 0, the attribute video data unit attribute with index attrIdx Specifies that an attribute tile may contain multiple attribute tile partitions. If none exist, the value of afati_single_partition_per_tile_flag[attrIdx] is assumed to be equal to 1.
[0094] afati_num_tiles_in_atlas_frame_minus1[attrIdx]: The value of afati_num_tiles_in_atlas_frame_minus1[attrIdx] plus 1 specifies the number of attribute tiles in each attribute atlas frame for the attribute signaled by the attribute video data unit at index attrIdx. The value of afati_num_tiles_in_atlas_frame_minus1[attrIdx] is in the range of 0 to NumPartitionsInAtlasFrameAtt[attrIdx]-1. If afati_num_tiles_in_atlas_frame_minus1[attrIdx] does not exist and afati_single_partition_per_tile_flag[attrIdx] is equal to 1, then the value of afati_num_tiles_in_atlas_frame_minus1[attrIdx] is presumed to be equal to NumPartitionsInAtlasFrameAtt[attrIdx] - 1. Here, the variable NumPartitionsInAtlasFrameAtt[attrIdx] is set to be equal to NumPartitionColumnsAtt[attrIdx] * NumPartitionRowsAtt[attrIdx]. If afati_single_tile_in_atlas_frame_flag[attrIdx] is equal to 0, then NumPartitionsInAtlasFrameAtt[attrIdx] must be greater than 1.
[0095] afati_top_left_partition_idx[attrIdx][i]: afati_top_left_partition_idx[attrIdx][i] specifies the partition index of the attribute tile partition located in the top-left corner of the i-th tile of the attribute video data with index attrIdx. The value of afati_top_left_partition_idx[attrIdx][i] is in the range of 0 to NumPartitionsInAtlasFrameAtt[attrIdx] - 1. The length of a syntax element is Ceil( Log2( NumPartitionsInAtlasFrameAtt[attrIdx] ) ) bits.
[0096] afati_bottom_right_partition_column_offset[attrIdx][i]: afati_bottom_right_partition_column_offset[attrIdx][i] is located at the bottom right corner of the i-th attribute tile. Attribute tiles for video data that have an index attrIdx to place. The column position of the partition and afati_bottom_right_partition_column_offset[attrIdx][ i Specifies the offset between the column position of the attribute tile partition and the partition index equal to ]. afati_single_partition_per_tile_flag[attrIdx] When is equal to 1, the value of afati_bottom_right_partition_column_offset[attrIdx][i] is presumed to be equal to 0.
[0097] afati_bottom_right_partition_row_offset[attrIdx][ i ]: afati_bottom_right_partition_row_offset[attrIdx][ i ] is located at the bottom right corner of the i-th attribute tile. Attribute tile part of attribute video data with index attridx Specifies the offset between the row position of the partition and the row position of the attribute tile partition that has a partition index equal to afati_top_left_partition_idx[attrIdx][i]. When afati_single_partition_per_tile_flag[attrIdx] is equal to 1, the value of afati_bottom_right_partition_row_offset[attrIdx][i] is presumed to be equal to 0.
[0098] afati_signalled_tile_id_flag[attrIdx]: If afati_signalled_tile_id_flag[attrIdx] is equal to 1, each attribute video data attribute with index attrIdx Specifies that the attribute tile ID of the root tile should be signaled. If afati_signalled_tile_id_flag[attrIdx] is equal to 0, it specifies that the attribute tile ID should not be signaled.
[0099] afati_signalled_tile_id_length_minus1[attrIdx]: The value of afati_signalled_tile_id_length_minus1[attrIdx] plus 1 specifies the number of bits used to represent the syntax element afati_tile_id[attrIdx][i], if it exists. The value of afati_signalled_tile_id_length_minus1[attrIdx] is in the range of 0 to 15. If it does not exist, the value of afati_signalled_tile_id_length_minus1[attrIdx] is estimated to be equal to Ceil( Log2( afati_num_tiles_in_atlas_frame_minus1[attrIdx] + 1 ) ) - 1.
[0100] afati_tile_id[attrIdx][i]: Specifies the attribute tile ID of the i-th attribute tile in the attribute video data with index attridx. If it does not exist... In total, the value of afati_tile_id[attrIdx][i] is presumed to be equal to i for each i in the range from 0 to afati_num_tiles_in_atlas_frame_minus1[attrIdx]. A bitstream conformance requirement is that afati_tile_id[attrIdx][i] is not equal to afati_tile_id[attrIdx][j] for all i != j. The variable FirstTileIDAtt[attrIdx] is calculated as follows: It can be done.
[0101] FirstTileIDAtt[attrIdx] = afati_tile_id[attrIdx]
[0000] for(i = 1; i < afati_num_tiles_in_atlas_frame_minus1[attrIdx] + 1; i++) FirstTileIDAtt[attrIdx] = Min( FirstTileID[attrIdx], afati_tile_id[attrIdx][ i ] ) The arrays TileIDToIndexAtt[attrIdx] and TileIndexToIDAtt[attrIdx] provide forward and reverse mappings, respectively, between the ID associated with each attribute tile and the ordered index of how each attribute tile was specified in the attribute tile information of the atlas frame.
[0102] (Structure of attribute tile information syntax) An attribute video stream frame (attribute frame) can be divided into one or more partitions, and attribute tiles can be constructed from these units (partitions). Typical examples include the following: - The attribute frame is not split, and the entire attribute frame is used as a single attribute tile (afati_single_tile_in_atlas_frame_flag[attrIdx]==1). - Divide the attribute frame into multiple partitions and use one partition as one attribute tile (afati_single_tile_in_atlas_frame_flag[attrIdx]==0 and afati_single_partition_per_tile_flag[attrIdx]==1). - Divide the attribute frame into multiple partitions, and use one or more consecutive partitions in the horizontal and vertical directions as a single attribute tile (afati_single_tile_in_atlas_frame_flag[attrIdx]==0 and afati_single_partition_per_tile_flag[attrIdx]==0).
[0103] Furthermore, an attribute frame can be divided into tile partitions (hereinafter also called partitions) of NumPartitionColumns * NumPartitionRows, and when dividing, you can choose to divide the frame at equal intervals or at specified units. NumPartitionColumns and NumPartitionRows are the partitions in the horizontal and vertical directions, respectively. It is a percentage.
[0104] Note that tiles are not limited to attribute frames; they can also be attributes, geometries, displacements, or meshes. In other words, the following syntax elements and their bitstream conformance conditions can also be used for attribute, geometry, displacement, and mesh tiles.
[0105] Figure 11 shows the atlas_frame_attribute_tile_information() tile information for V-DMC. This is a diagram of taxes.
[0106] The extended information decoding unit 3023 of the atlas information decoding unit 302 decodes the syntax element afati_single_tile_in_atlas_frame_flag[attrIdx]. afati_single_tile_in_atlas_frame_flag[attrIdx] is a binary flag that indicates whether or not the attribute frame consists of a single tile, and the value that indicates the attribute frame consists of a single tile (for example) 1) Or, a value indicating that the attribute frame consists of multiple tiles (e.g.) For example, it has 0). If the value of afati_single_tile_in_atlas_frame_flag[attrIdx] is a value indicating multiple tiles, the extended information decoding unit 3023 decodes the syntax element afati_uniform_partition_spacing_flag[attrIdx]. Here, afati_uniform_partition_spacing_flag[attrIdx] is a binary flag indicating whether or not to divide the attribute frame into equally spaced partitions. A lag is a value that indicates dividing the attribute frame into equally spaced partitions (example). For example, 1) or a value (e.g., 0) that indicates the attribute frame will not be divided into equally spaced partitions.
[0107] The extended information decoding unit 3023 decodes parameters indicating the position and size of the tiles.
[0108] If afati_uniform_partition_spacing_flagAtt[attrIdx] is a value of 1, then the extended information The information decoding unit 3023 decodes the syntax elements afati_partition_cols_width_minus1[attrIdx], afati_partition_cols_width_minus1[attrIdx], which indicate the width (column width) and height (row height) of each partition, excluding the rightmost column (rightmost column) and the bottommost row (bottommost row). For each i=0..NumPartitionColumnsAtt[attrIdx]-1, j=0..NumPartitionRowsAtt[attrIdx] Then, the x, y coordinates, width, and height of the top left of each partition are calculated as follows: PartitionPosXAtt[attrIdx][i], PartitionPosYAtt[attrIdx][j], PartitionWidthAtt[attrIdx][i], PartitionHeightAtt[attrIdx][j].
[0109] widthPartition = ( afati_partition_cols_width_minus1[attrIdx] + 1 ) * 64 NumPartitionColumnsAtt[attIdx] = asve_attribute_frame_width[attrIdx] / widthPartition PartitionPosXAtt[attrIdx]
[0000] = 0 PartitionWidthAtt[attrIdx]
[0000] = widthPartition for( i = 1; i < NumPartitionColumnsAtt[attrIdx] - 1; i++ ) { PartitionPosXAtt[attrIdx][ i ] = PartitionPosXAtt[attrIdx][ i - 1 ] + PartitionWidthAtt[attrIdx][ i - 1 ] PartitionWidthAtt[attrIdx][ i ] = widthPartition } partitionHeightAtt[attrIdx] = (afati_partition_rows_height_minus1[attrIdx] + 1) * 64 NumPartitionRowsAtt[attrIdx] = asve_attribute_frame_height[attrIdx] / partitionHeight PartitionPosYAtt[attrIdx]
[0000] = 0 PartitionHeightAtt[attrIdx]
[0000] = heightPartition for( j = 1; j < NumPartitionRowsAtt[attrIdx] - 1; j++ ) { PartitionPosYAtt[attrIdx][ j ] = PartitionPosYAtt[attrIdx] [ j - 1 ] + PartitionHeightAtt[attrIdx][ j - 1 ] PartitionHeightAtt[attrIdx][ j ] = heightPartition } If afati_uniform_partition_spacing_flagAtt[attrIdx] is a value of 0, then the extended information The information decoding unit 3023 uses syntax elements afati_num_partition_columns_minus1[attrIdx], afati_num_partition_rows_ to indicate the number of horizontal and vertical columns of the Atlas style partition. Decrypt minus1[attrIdx].
[0110] For each i=0..NumPartitionColumnsAtt[attrIdx]-1, j=0..NumPartitionRowsAtt[attrIdx] In this case, the x, y coordinates, width, and height of the top left of each partition are calculated as follows: PartitionPosXAtt[attrIdx][i], PartitionPosYAtt[attrIdx][j], PartitionWidthAtt[attrIdx][i], PartitionHeightAtt[attrIdx][j].
[0111] NumPartitionColumnsAtt[attrIdx] = afati_num_partition_columns_minus1[attrIdx] + 1 PartitionPosXAtt[attrIdx]
[0000] = 0 partitionWidthAtt[attrIdx]
[0000] = ( afati_partition_column_width_minus1[attrIdx]
[0000] + 1 ) * 64 for( i = 1; i < NumPartitionColumnsAtt[attrIdx] - 1; i++ ) { PartitionPosXAtt[attrIdx][ i ] = PartitionPosXAtt[attrIdx][ i - 1 ] + PartitionWidthAtt[attrIdx][ i - 1 ] PartitionWidthAtt[attrIdx][ i ] = ( afati_partition_column_width_minus1[attrIdx][ i ] + 1 ) * 64 } NumPartitionRowsAtt[attrIdx] = afati_num_partition_rows_minus1[attrIdx] + 1 PartitionPosYAtt[attrIdx]
[0000] = 0 PartitionHeightAtt[attrIdx]
[0000] = ( afati_partition_row_height_minus1[attrIdx]
[0000] + 1 ) * 64 for( j = 1; j < NumPartitionRowsAtt[attrIdx] - 1; j++ ) { PartitionPosYAtt[attrIdx][ j ] = PartitionPosYAtt[attrIdx][ j - 1 ] + PartitionHeightAtt[attrIdx][ j - 1 ] PartitionHeightAtt[attrIdx][ j ] = ( afati_partition_row_height_minus1[attrIdx][ j ] + 1 ) * 64 } Also, if the number of partitions in the horizontal and vertical directions is two or more, the rightmost and bottommost partitions The PartitionPosXAtt[attrIdx][i], PartitionPosYAtt[attrIdx][j], PartitionWidthAtt[attrIdx][i], and PartitionHeightAtt[attrIdx][j], which represent the x, y coordinates, width, and height of the top-left of the partition, are decoded as follows.
[0112] PartitionPosXAtt[attrIdx][ NumPartitionColumnsAtt[attrIdx - 1 ] = PartitionPosXAtt[attrIdx][ NumPartitionColumnsAtt[attrIdx - 2 ] + PartitionWidthAtt[attrIdx][ NumPartitionColumnsAtt[attrIdx] - 2 ] PartitionWidthAtt[attrIdx][ NumPartitionColumnsAtt[attrIdx] - 1 ] = asve_attribute_frame_width[attrIdx] - PartitionPosXAtt[attrIdx][ NumPartitionColumnsAtt[attrIdx] - 1 ] PartitionPosYAtt[attrIdx][ NumPartitionRowsAtt[attrIdx - 1 ] = PartitionPosYAtt[attrIdx][ NumPartitionRowsAtt[attrIdx - 2 ] + partitionHeightAtt[attrIdx][ NumPartitionRowsAtt[attrIdx - 2 ] PartitionHeight[ NumPartitionRowsAtt[attrIdx - 1 ] = asve_attribute_frame_height[attrIdx] - PartitionPosYAtt[attrIdx][ NumPartitionRows - 1 ] Here, the width and height of each partition are set to multiples of 64, but they are not limited to 64; you can replace 64 with 32, 128, or 256.
[0113] The extended information decoding unit 3023 decodes the syntax element afati_single_partition_per_tile_flag[attrIdx]. Here, afati_single_partition_per_tile_flag[attrIdx] is each tile A flag indicating whether a tile consists of only a single partition, where a value (e.g., 1) indicates that each tile consists of only a single partition, or each tile consists of multiple partitions. It has a value (e.g., 0) that indicates it consists of a single partition. If afati_single_partition_per_tile_flag[attrIdx] is a value that indicates multiple partitions, the extended information decoding unit 3023 decodes the syntax element afati_num_tiles_in_atlas_frame_minus1[attrIdx] and selects one The following process is performed to decode the tile parameters from the above partitions. Here, afati_num_tiles_in_atlas_frame_minus1 is the number of tiles that make up the attribute frame. That is the case.
[0114] The extended information decoding unit 3023 processes each i=0..afati_num_tiles_in_atlas_frame_minus1[attrIdx] Therefore, we decode the syntax elements afati_top_left_partition_idxAtt[attrIdx][i], afati_bottom_right_partition_column_offset[attrIdx][i], and afati_bottom_right_partition_row_offset[attrIdx][i]. Here, afati_top_left_partition_idx[attrIdx][i] is the index of the partition where the top-left corner (point) of the i-th tile is located, and afati_bottom_right_partition_column_offset[attrIdx][i] is the index of the partition where the top-left corner (point) of the i-th tile is located. afati_bottom_right_partition_row_offset[attrIdx][i] is the horizontal offset amount of the bottom right edge of the i-th tile relative to the top left edge of the row, and is the vertical offset amount of the bottom right edge of the i-th tile relative to the top left edge of the i-th tile.
[0115] Based on the decoded syntax above, the top left horizontal, height, and bottom right of each tile i The horizontal and vertical partition indices topLeftColumnAtt[attrIdx][i], topLeftRowAtt[attrIdx][i], bottomRightColumnAtt[attrIdx][i], and bottomRowAtt[attrIdx][i] are calculated as follows.
[0116] topLeftColumnAtt[attrIdx][ i ] = afati_top_left_partition_idxAtt[attrIdx][ i ] % NumPartitionColumnsAtt[attrIdx] topLeftRowAtt[attrIdx][ i ] = afati_top_left_partition_idxAtt[attrIdx][ i ] / NumPartitionColumnsAtt[attrIdx] bottomRightColumnAtt[attrIdx][ i ] = topLeftColumnAtt[attrIdx][ i ] + afati_bottom_right_partition_column_offsetAtt[attrIdx][ i ] bottomRightRowAtt[attrIdx][ i ] = topLeftRowAtt[attrIdx][ i ] + afati_bottom_right_partition_row_offsetAtt[attrIdx][ i ] Here, bottomRightColumnAtt[attrIdx][i] and bottomRightRowAtt[attrIdx][i] are They may also be less than or equal to (asve_attribute_frame_width[attrIdx] + 63) / 64 - 1 and (asve_attribute_frame_height[attrIdx] + 63) / 64 - 1, respectively.
[0117] In a 3D data decoding device 31 that decodes mesh data or point cloud data, the syntax elements indicating the position of attribute tiles are decoded, and the column of the top-left partition of the tile, topLeftColumnAtt[attrIdx], the row of the top-left partition, topLeftRowAtt[attrIdx], and The column bottomRightColumnAtt[attrIdx] of the lower right partition of the rivute tile, the lower right partition The 3D data decoding device 31 has means for decoding the bottomRightRowAtt[attrIdx] row of the attribute tile, and the partition column of the i-th attribute tile (topLeftColumnAtt[attrIdx][i] and bottomRightColumnAtt[attrIdx][i]), the partition column of the j-th attribute tile (topLeftColumnAtt[attrIdx][j]), and the partition of the i-th attribute tile The row of the tition (topLeftRowAtt[attrIdx][i], bottomRightRowAtt[attrIdx][i]), and j The row of the partition of the nth attribute tile (topLeftRowAtt[attrIdx][j]) The 3D data decoding device 31 may decode a bitstream that satisfies specific bitstream conformance conditions.
[0118] (Decoding of additional information) The additional information decoding unit 3024 decodes the additional information (SEI) of submesh tile information SEI and submesh attribute tile information SEI from the encoded data (additional information encoded stream). Furthermore, post-processing such as zippering to remove cracks in the mesh is performed. The SEI of the (Ri) or the SEI that extracts mesh displacement for each LoD may be decoded. The extended information decoding unit 3023 and the extended information encoding unit 1011 decode and encode the syntax elements of the submesh tile information SEI and submesh attribute tile information SEI, respectively.
[0119] (The relationship between Atlas tile and Codec tile) The 3D data decoding device 31, which decodes mesh data or point cloud data, decodes the atlas frame (displacement frame gFrame) and attribute frame aFrame on a tile-by-tile basis. The atlas frame (displacement frame gFrame) and attribute frame aFrame are HEVC, VVC These are encoded using codecs such as HEVC and VVC. The frame is partitioned into rectangular regions. A tile is a unit that can be decoded independently (in parallel) within that frame. To distinguish the defined rectangular area of a tile from attribute tiles and other similar tiles, it is called a "codec tile (independently decodeable unit, video tile)." The definition of this tile is called "codec tile information." If the displacement subbitstream is encoded in ISO / IEC 23008-2 (HEVC), the codec tile is defined as a tile as defined in ISO / IEC 23008-2. If the displacement subbitstream is encoded in ISO / IEC 23090-3 (VVC), the tile is defined as a tile as defined in ISO / IEC 23090-3. Defined as "ru".
[0120] Figure 23 shows an example of the syntax for submesh tile information SEI. This is an example of SEI syntax for decoding mapping flags, which are mapping information indicating the relationship between tiles and codec tiles. The submesh tile information SEI consists of smtm_num_submeshes_minus1, smtm_num_tiles_minus1, smtm_tile_idx[i], and smtm_mapping_flag[i][j]. It includes syntax elements. The submesh tile information SEI may further include SEI cancellation and persistence flags such as smtm_cancel_mapping_flag and smtm_persistance_mapping_flag, smtm_signalled_submesh_id_flag and smtm_signalled_submesh_id_length_minus1. The semantics of each syntax element are as follows. The semantics of the same syntax names are defined the same way in other embodiments (and so on).
[0121] smtm_cancel_mapping_flag: smtm_cancel_mapping_flag is a flag that indicates whether to cancel the persistence of SEI messages for the mapping information submesh_tiles_mapping. If smtm_cancel_mapping_flag is a first value (e.g., 1), the SEI message will be earlier in the output order. This indicates that the persistence of the mapping information SEI message is being canceled. If smtm_cancel_mapping_flag is a second value (e.g., 0), it indicates that the mapping information will continue.
[0122] Here, aFrmA is the current atlas frame associated with the mapping information SEI message. If smtm_persistance_mapping_flag is equal to the first value (e.g., 1), sub Specifies that mesh tile mapping may be used for the current image and all subsequent images of the current layer in the output order. This persists until one of the following conditions is true: • A new CAS (Casual Air Service) is starting. The bitstream ends. - In the current layer, if an atlas frame aFrmB to which a mapping information message is applied is output, and AtlasFrmOrderCnt(aFrmB) is greater than AtlasFrmOrderCnt(aFrmA). Here, AtlasFrmOrderCnt(aFrmB) and AtlasFrmOrderCnt(aFrmA) are the atlas frame order count values for aFrmB and aFrmA, respectively, and the atlas frame order count of aFrmB is This is the value immediately after the code process was invoked.
[0123] smtm_persistance_mapping_flag: If smtm_persistance_mapping_flag is equal to 1, it indicates that submesh tile mapping will persist. If it is equal to 0, it indicates that submesh tile mapping will persist. This indicates that the illuminating mapping is only valid for the current frame.
[0124] smtm_signalled_submesh_id_flag: A flag that indicates whether to transmit the submesh ID. If this flag is true, smtm_signalled_submesh_id_length_minus1 and smtm_submesh_id[i] are decoded and encoded; if false, they are not decoded or encoded.
[0125] smtm_signalled_submesh_id_length_minus1: Syntax element of smtm_submesh_id[i] This indicates the length in bits (bit length).
[0126] smtm_submesh_id[i]: smtm_submesh_id[i] specifies the ID of the i-th submesh. As a bitstream conformance requirement, smtm_submesh_id[i] is not equal to smtm_submesh_id[k] for all i != k. The length is smtm_signalled_submesh_id_length_minus1+1 bits.
[0127] smtm_num_submeshes_minus1: The value obtained by adding 1 to smtm_num_submeshes_minus1 is stored in CAS. Indicates the number of submeshes associated with the tile. smtm_tile_idx[i]: smtm_tile_idx[i] indicates the index of the tile contained in the i-th submesh.
[0128] smtm_num_tiles_minus1: The value of smtm_num_tiles_minus1 plus 1 indicates the number of tiles in the frame. Also, the value of smtm_num_tiles_minus1 must be within the range of 0 to afti_num_tiles_in_atlas_frame_minus1 (including both ends).
[0129] smtm_num_codec_tiles_minus1: The value obtained by adding 1 to smtm_num_codec_tiles_minus1 indicates the number of independently decodeable units, i.e., codec tiles, in the subbitstream.
[0130] smtm_mapping_flag[i][j]: smtm_mapping_flag[i][j] is a tile and code. This flag maps the relationships between codecs. If this flag is true (e.g., 1), the i-th tile contains the j-th codec tile. If this flag is false (e.g., 1) If the value is equal to 0, then the i-th tile does not contain the j-th codec tile.
[0131] Figure 19 is a diagram illustrating the values of each flag in the case of smtm_mapping_flag
[0000]
j
[0132] Figure 19 shows an example of an image frame consisting of four tiles and sixteen codec tiles. The four tiles and sixteen codec tiles are numbered (indexed) in the order of the raster scan. Let's assume that a code is assigned to the 0th tile. In the example in Figure 19, the 0th tile has codecs 0, 1, 4, and 5. It contains il. Each number listed on the tile / codec tile in Figure 19 is smtm_m This indicates the value of `smtm_mapping_flag
[0000] [j]`, and in this example, if the 0th tile contains a codec tile, the value of the flag `smtm_mapping_flag
[0000] [j]` is 1, and the 0th tile If the codec tile is not included, flag smtm_mapping_flag
[0000] [ j ] The value of 0 is decoded as the value of . In the example in Figure 19, the value of the flag smtm_mapping_flag
[0000] [j] can be shown as follows.
[0133] smtm_mapping_flag
[0000]
[0000] = 1 smtm_mapping_flag
[0000]
[0001] = 1 smtm_mapping_flag
[0000]
[0002] = 0 smtm_mapping_flag
[0000]
[0003] = 0 smtm_mapping_flag
[0000]
[0004] = 1 smtm_mapping_flag
[0000]
[0005] = 1 smtm_mapping_flag
[0000]
[0006] = 0 smtm_mapping_flag
[0000]
[0007] = 0 smtm_mapping_flag
[0000]
[0008] = 0 smtm_mapping_flag
[0000]
[0009] = 0 smtm_mapping_flag
[0000]
[0010] = 0 smtm_mapping_flag
[0000]
[0011] = 0 smtm_mapping_flag
[0000]
[0012] = 0 smtm_mapping_flag
[0000]
[0013] = 0 smtm_mapping_flag
[0000]
[0014] = 0 smtm_mapping_flag
[0000]
[0015] = 0 According to the above configuration, the additional information decoding unit 3024 decodes tile i from the encoded data into a codec The mapping flag smtm_mapping_flag[i][j], which indicates whether tile j is included, is decoded. This decodes the correspondence between a specific tile i and the codec tile j that needs to be decoded in order to decode that tile, allowing the necessary codec tiles to be decoded for a given tile to be decoded in parallel.
[0134] Alternatively, the configuration could involve decoding an index that shows the relationship between tiles and codec tiles.
[0135] Figure 24 shows another example of the syntax for submesh tile information SEI. Tiles and Coordinates This is an example of SEI syntax for decoding an index that shows the relationship between tictiles. The tics are as follows: The submesh tile information SEI includes the syntax elements smtm_num_submeshes_minus1, smtm_num_tiles_minus1, and smtm_codec_tile_idx[i][j]. The submesh tile information SEI may also include SEI cancellation and persistence flags such as smtm_cancel_mapping_flag and smtm_persistance_mapping_flag, smtm_signalled_submesh_id_flag, smtm_signalled_submesh_id_length_minus1, and smtm_tile_idx[i].
[0136] smtm_num_codec_tiles_in_tile_minus1[i]: The value of smtm_num_codec_tiles_in_tile_minus1[i] plus 1 is the code of the i-th tile included in the displacement subbitstream. This indicates the number of ctiles.
[0137] smtm_codec_tile_idx[i][j]: smtm_codec_tile_idx[i][j] is the codec tile of the current frame that corresponds to the j-th codec tile contained in the i-th tile. This indicates the index. If smtm_codec_tile_idx[i][j] does not exist for the i-th tile, then j will be in the range from 0 to smtm_num_codec_tiles_in_tile_minus1[i] (including both endpoints), which is inferred to be equal to i.
[0138] smtm_num_codec_tiles_in_tile_minus[ i ] = i.
[0139] The additional information decoding unit 3024 may further derive variables CodecTileIdxInSubmesh[i][j] that indicate the relationship between the submesh and the codec tile. The variables CodecTileIdxInSubmesh[i][j] represent the codec tile of the displacement subbitstream contained in the i-th submesh. This shows the index. i ranges from 0 to smtm_submesh_minus1, and j ranges from 0 to smtm_num_codec_tiles_in_tile_minus1. CodecIdxInSubmesh[i][j] is smtm_codec_tile_idx[smtm_tile_idx[i]][j] based on smtm_tile_idx[i].
[0140] CodecIdxInSubmesh[ i ][ j ] = smtm_codec_tile_idx[ smtm_tile_idx[ i ] ][ j ].
[0141] Figure 20 shows an example of an image frame consisting of four tiles and sixteen codec tiles. The four tiles and sixteen codec tiles are numbered (indexed) in the order of the raster scan. Let's assume that a codec is assigned to the 0th tile. In the example in Figure 20, the 0th tile has codecs 0, 1, 4, and 5. The tile is included. Each number listed on the tile / codec tile in Figure 20 represents the value of smtm_codec_tile_idx[i][j]. The values of the index smtm_codec_tile_idx[i][j] in the example in Figure 20 can be shown as follows:
[0142] smtm_codec_tiles_idx
[0000]
[0000] = 0 smtm_codec_tiles_idx
[0000]
[0001] = 1 smtm_codec_tiles_idx
[0000]
[0002] = 4 smtm_codec_tiles_idx
[0000]
[0003] = 5 smtm_codec_tiles_idx
[0001]
[0000] = 2 smtm_codec_tiles_idx
[0001]
[0001] = 3 smtm_codec_tiles_idx
[0001]
[0002] = 6 smtm_codec_tiles_idx
[0001]
[0003] = 7 smtm_codec_tiles_idx
[0002]
[0000] = 8 smtm_codec_tiles_idx
[0002]
[0001] = 9 smtm_codec_tiles_idx
[0002]
[0002] = 12 smtm_codec_tiles_idx
[0002]
[0003] = 13 smtm_codec_tiles_idx
[0003]
[0000] = 10 smtm_codec_tiles_idx
[0003]
[0001] = 11 smtm_codec_tiles_idx
[0003]
[0002] = 14 smtm_codec_tiles_idx
[0003]
[0003] = 15 According to the above configuration, the additional information decoding unit 3024 decodes tile i from the encoded data into a codec Decode the index smtm_codec_tiles_idx[i][j] which indicates whether tile j is included. This allows for the decoding of the codec tile index corresponding to a specific tile, enabling the decoding of the necessary codec tiles in parallel to decode a particular tile.
[0143] In another configuration, the displacement subbitstream is coded in the submesh tile information SEI. Decode the alignment flag smtm_codec_tile_alignment_flag, which indicates whether the tile and codec tile partitions are the same. If smtm_codec_tile_alignment_flag has a first value of true (e.g., 1), it signals that the tile and codec tile partitions are the same; if it has a second value of false (e.g., 0), it indicates that the tile and codec tile partitions may be different.
[0144] Here is an example of the syntax for submesh tile information SEI including alignment flags: It is as follows:
[0145] The additional information decoding unit 3024, when the alignment flag is the second value false, outputs smtm_num_tiles_minus1, smtm_num_codec_tiles_in_tiles_minus1[i], and smtm_codec_tile_idx[i][j]. Decode it.
[0146] submesh_tiles_mapping( payloadSize ) { smtm_cancel_mapping_flag if( !smtm_cancel_mapping_flag ) { smtm_persistance_flag smtm_num_submeshes_minus1 smtm_signalled_submesh_id_flag if( smtm_signalled_submesh_id_flag ) smtm_signalled_submesh_id_length_minus1 for( i = 0; i < smtm_num_submeshes_minus1 + 1; i++ ) { smtm_submesh_id[ i ] SubmeshIdxToID[ i ] = smtm_submesh_id[ i ] } } for( i = 0; i < smtm_num_submeshes_minus1 + 1; i++) { smtm_tile_idx[ i ] } smtm_codec_tile_alignment_flag if( !smtm_codec_tile_alignment_flag ) { smtm_num_tiles_minus1 for( i = 0; i < smtm_num_tiles_minus1 + 1; i++) { smtm_num_codec_tiles_in_tiles_minus1[ i ] for( j = 0; j < smtm_num_codec_tiles_in_tiles_minus1[ i ] + 1; j++) smtm_codec_tile_idx[ i ][ j ] } } } } Here, the semantics of the syntax element smtm_codec_tile_alignment_flag are as follows:
[0147] smtm_codec_tile_alignment_flag: smtm_codec_tile_alignment_flag is the frame's Indicates whether the tile division is the same as the codec tile division of the displacement subbitstream. A first value (e.g., 1) indicates that the tile and codec tile divisions are the same, while a second value (e.g., 0) indicates that the tile and codec tile divisions may be different.
[0148] According to the above configuration, the additional information decoding unit 3024 decodes the alignment flag smtm_codec_tile_alignment_flag, which indicates whether the tile division of the frame is the same as the codec tile division of the displacement subbitstream, that is, whether the i-th tile contains only the region of the j-th codec tile. In this case, the i-th tile contains only the j-th codec tile when the indices are equal (i == j).
[0149] Therefore, when the tile and codec tile divisions are the same, the syntax component smtm_codec_tile_idx, which shows the relationship between the frame tile division and the video stream's codec tile division, can be derived without encoding, thus reducing the overhead of the coding amount. It is effective.
[0150] Figure 25 shows an example of the syntax for submesh attribute tile information SEI. SEI decodes the mapping flag that shows the relationship between the rivet tile and the codec tile. This is an example of syntax. The semantics of each syntax element are as follows:
[0151] smatm_cancel_mapping_flag: smatm_cancel_mapping_flag is a flag that indicates whether or not to cancel the persistence of SEI messages for the mapping information submesh_attribute_tiles_mapping. Therefore, if smatm_cancel_mapping_flag is the first value (for example, 1), the SEI message is output. This indicates canceling the persistence of previous mapping information SEI messages for the force sequence. If smatm_cancel_mapping_flag is a second value (e.g., 0), it indicates that the mapping information will continue. .
[0152] Here, aFrmA is the current atlas frame associated with the mapping information SEI message. If smatm_persistance_mapping_flag is equal to the first value (e.g., 1), it specifies that submesh tile mapping may be used for the frames of the current image and all subsequent images of the current layer in the output order. This persists until one of the following conditions is true: • A new CAS (Casual Air Service) is starting. The bitstream ends. - In the current layer, if an atlas frame aFrmB to which a mapping information message is applied is output, and AtlasFrmOrderCnt(aFrmB) is greater than AtlasFrmOrderCnt(aFrmA). Here, AtlasFrmOrderCnt(aFrmB) and AtlasFrmOrderCnt(aFrmA) are the atlas frame order count values for aFrmB and aFrmA, respectively, and the atlas frame order count of aFrmB is This is the value immediately after the code process was invoked.
[0153] smatm_persistance_mapping_flag: If smtm_persistance_mapping_flag is equal to 1 , indicates that the submesh tile mapping persists. If equal to 0, the submesh This indicates that tile mapping is only valid for the current frame.
[0154] smatm_num_attribute_video_minus1: The value of smatm_num_attribute_video_minus1 plus 1 indicates the number of attributes signaled in the attribute video subbitstream. V3C bitstream As a requirement for the conformance of the atlas, the value of smatm_num_attribute_video is the ID of the current atlas. For some value of j, it must be equal to the value of ai_attribute_count[j].
[0155] Here, the attribute video data looped through with the number of attributes smatm_num_attribute_video_minus1+1 includes attributes such as Color, Material ID, Transparency, Reflectance, Normals, Texture coordinate, and Facegroup ID.
[0156] smatm_attribute_tile_idx[i][j]: smatm_attribute_tile_idx[i][j] is the index of the attribute tile of the j-th attribute video data that contains the i-th submesh. It indicates a problem.
[0157] smatm_codec_tiles_minus1: The value of smatm_codec_tiles_minus1 plus 1 indicates the number of independently decodeable units, i.e., codec tiles, in the subbitstream.
[0158] smatm_num_attribute_tiles_minus[i]: The value of smatm_num_attribute_tiles_minus1 plus 1 indicates the number of attribute tiles included in the i-th attribute video data.
[0159] smatm_mapping_flag[ i ][ j ][ k ] : smatm_mapping_flag[ i ][ j ][ k ] is, This flag maps the relationship between the attribute tile and the codec tile. If this flag is equal to true (e.g., 1), the j-th attribute tile of the i-th attribute video data... The codec tile contains the kth codec tile. If this flag is equal to false (e.g., 0) In total, the i-th attribute video data's j-th attribute tile corresponds to the k-th codec tile. The character "ル" is not included.
[0160] The additional information decoding unit 3024 may further derive a variable AttrCodecTileIdxInSubmesh[i][j][k] that indicates the relationship between the submesh and the codec tile. The variable AttrCodecTileIdxInSubmesh[i][j][k] represents the attribute video subbitstory contained in the i-th submesh. This shows the index of the codec tile for the submesh. i ranges from 0 to smtm_submesh_minus1. , where j ranges from 0 to smatm_num_attribute_video_minus1 and k ranges from 0 to smatm_num_codec_tiles_in_tiles_minus1[i][j], and AttrCodecIdxInSubmesh[i][j][k] is Based on smatm_attribute_tile_idx[i][j], the result is smatm_codec_tile_idx[i][smatm_attribute_tile_idx[i][j][k].
[0161] AttrCodecIdxInSubmesh[ i ][ j ][ k ] = smatm_codec_tile_idx[ i ][ smatm_attribute_tile_idx[ i ][ j ] ][ k ].
[0162] Figure 21 shows an example of smatm_tile_mapping_flag[i][j][k], specifically smatm_mapping_flag[0 This is a diagram illustrating the values of each flag in the case of
[0000] [k]. A codec tile indicates the partitioning of the attribute video subbitstream. A codec tile may be a macroblock slice in AVC / H.264, a slice or tile in HEVC / H.265, or a slice, tile or subpicture in VVC / H.266. A codec tile (also called a subset or segment) is each partition that divides the frame and initializes an arithmetic code such as CABAC at the beginning so that it can be decoded in parallel. The segmentation method is derived by decoding the video codec's bitstream parameter sets (Sequence Parameter Set (SPS), Picture Parameter Set (PPS)) and slice header. do.
[0163] Figure 21 shows an example of an image frame consisting of 4 tiles and 16 codec tiles for attribute 0 video data. The 4 tiles and 16 codec tiles contain rascals. Assume that numbers (indexes) are assigned in order. In the example in Figure 21, the 0th attribute is video data. The 0th tile of the array contains the 0th, 1st, 4th, and 5th codec tiles. Figure 21 shows the array. Each number listed in the attribute tile / codec tile indicates the value of smatm_mapping_flag
[0000]
[0000] [k]. In this example, if the 0th attribute tile contains a codec tile, the value of the flag smatm_mapping_flag
[0000]
[0000] [k] is 1, and if the 0th attribute tile does not contain a codec tile, the flag is false. The value of smatm_mapping_flag
[0000]
[0000] [k] is decoded to be 0. In the example in Figure 21, the value of the flag smatm_mapping_flag
[0000]
[0000] [k] can be shown as follows.
[0164] smatm_mapping_flag
[0000]
[0000]
[0000] = 1 smatm_mapping_flag
[0000]
[0000]
[0001] = 1 smatm_mapping_flag
[0000]
[0000]
[0002] = 0 smatm_mapping_flag
[0000]
[0000]
[0003] = 0 smatm_mapping_flag
[0000]
[0000]
[0004] = 1 smatm_mapping_flag
[0000]
[0000]
[0005] = 1 smatm_mapping_flag
[0000]
[0000]
[0006] = 0 smatm_mapping_flag
[0000]
[0000]
[0007] = 0 smatm_mapping_flag
[0000]
[0000]
[0008] = 0 smatm_mapping_flag
[0000]
[0000]
[0009] = 0 smatm_mapping_flag
[0000]
[0000]
[0010] = 0 smatm_mapping_flag
[0000]
[0000]
[0011] = 0 smatm_mapping_flag
[0000]
[0000]
[0012] = 0 smatm_mapping_flag
[0000]
[0000]
[0013] = 0 smatm_mapping_flag
[0000]
[0000]
[0014] = 0 smatm_mapping_flag
[0000]
[0000]
[0015] = 0 According to the above configuration, the mapping flag smatm_mapping_flag[i][j][k] is decoded. To decode a specific attribute tile, the codec tiles to be decoded for decoding the corresponding relationship with the codec tiles can be decoded in parallel for decoding the attribute tiles of a certain attribute video.
[0165] In another configuration, it may also be a configuration that decodes an index indicating the relationship between the attribute tile and the codec tile.
[0166] FIG. 26 is an example of another syntax of the sub-mesh attribute tile information SEI. The syntax of the SEI that decodes an index indicating the relationship between the attribute tile and the codec tile is an example. The semantics are as follows.
[0167] smatm_num_codec_tiles_in_tile_minus1[ i ][ j ]: The value obtained by adding 1 to smatm_num_codec_tiles_in_tile_minus1[ i ] indicates the number of codec tiles of the j-th attribute tile of the i-th attribute video data. shows.
[0168] smatm_codec_tile_idx[ i ][ j ][ k ]: smatm_codec_tile_idx[ i ][ j ][ k ] indicates the index of the current frame corresponding to the k-th codec tile included in the j-th attribute tile of the i-th attribute video data.
[0169] FIG. 22 shows an example of an image frame composed of 4 tiles of the 0-th attribute video data and 16 codec tiles. It is assumed that numbers (indices) are assigned to the 4 tiles and 16 codec tiles in raster scan order. In the example of FIG. 22, the 0-th attribute tile of the 0-th attribute video contains the 0th, 1st, 4th, and 5th codec tiles. tile. Each number listed in the attribute tile / codec tile in Figure 22 represents the value of smatm_codec_tiles_idx[i][j][k]. The values of the index smatm_codec_tiles_idx[i][j][k] in the example in Figure 22 can be expressed as follows:
[0170] smatm_codec_tiles_idx
[0000]
[0000]
[0000] = 0 smatm_codec_tiles_idx
[0000]
[0000]
[0001] = 1 smatm_codec_tiles_idx
[0000]
[0000]
[0002] = 4 smatm_codec_tiles_idx
[0000]
[0000]
[0003] = 5 smatm_codec_tiles_idx
[0000]
[0001]
[0000] = 2 smatm_codec_tiles_idx
[0000]
[0001]
[0001] = 3 smatm_codec_tiles_idx
[0000]
[0001]
[0002] = 6 smatm_codec_tiles_idx
[0000]
[0001]
[0003] = 7 smatm_codec_tiles_idx
[0000]
[0002]
[0000] = 8 smatm_codec_tiles_idx
[0000]
[0002]
[0001] = 9 smatm_codec_tiles_idx
[0000]
[0002]
[0002] = 12 smatm_codec_tiles_idx
[0000]
[0002]
[0003] = 13 smatm_codec_tiles_idx
[0000]
[0003]
[0000] = 10 smatm_codec_tiles_idx
[0000]
[0003]
[0001] = 11 smatm_codec_tiles_idx
[0000]
[0003]
[0002] = 14 smatm_codec_tiles_idx
[0000]
[0003]
[0003] = 15 According to the above configuration, the index smatm_codec_tiles_idx[i][j][k] is decoded. This corresponds to the codec tile of a specific attribute video. Since the index is known, the codec tiles necessary to decode a particular attribute tile can be decoded in parallel.
[0171] (Configuration of alignment flags) In another configuration, the attribute of the video data of index i is as follows: Decode the alignment flag smatm_codec_tile_alignment_flag[i] which indicates whether the tile division is the same as the codec tile division of the attribute stream. The first value (for example) If the first value is (i), it indicates that the attribute tile and codec tile divisions are the same in the video data at index i. If the second value (e.g., 0) is (i), it indicates that the attribute tile and codec tile divisions may be different, and a subsequent single tile such as smatm_codec_tile_idx[i][j][k] that shows the relationship between the attribute tile and the codec tile is not shown. Further decryption of the tax.
[0172] submesh_attribute_tiles_mapping( payloadSize ) { smatm_cancel_mapping_flag if( !smatm_cancel_mapping_flag ) { smatm_persistance_flag smtm_num_submeshes_minus1 smtm_signalled_submesh_id_length_minus1 if( smtm_signallied_submesh_id_length_minus1 ) { smtm_signalled_submesh_id_length_minus1 for( i = 0; i < smtm_num_submeshes_minus1 + 1; i++ ) { smtm_submesh_id[ i ] SubmeshIdxToID[ i ] = smtm_submesh_id[ i ] } } smatm_num_attribute_video_minus1 for( i = 0; i < smtm_num_submeshes_minus1 + 1; i++ ) { for( j = 0; j < smatm_num_attribute_video_minus1 + 1; j++ ) smatm_attribute_tile_idx[ i ][ j ] } for( i = 0; i < smatm_attribute_video_minus1 + 1; i++ ) { smatm_codec_tile_alignment_flag[ i ] if( !smatm_codec_tile_alignment_flag[ i ] ) { smatm_num_attribute_tiles_minus1[ i ] for( j = 0; j < smatm_num_attribute_tiles_minus1[i] + 1; j++) { smatm_num_codec_tiles_in_tiles_minus1[ i ][ j ] for (k = 0; j < smatm_num_codec_tiles_in_tiles_minus1[i][j] + 1; k++) smatm_codec_tile_idx[i][j][k] } } } } } Here, the semantics of the syntax element smatm_codec_tile_alignment_flag[i] are as follows.
[0173] smatm_codec_tile_alignment_flag[i]: smatm_codec_tile_alignment_flag[i] indicates whether the attribute tile division of the attribute video data of a certain index i is the same as the codec tile division of the attribute video sub bitstream. In the case of the first value (e.g., 1), it indicates that the division of the attribute tile and the codec tile is the same. In the case of the second value (e.g., 0), it indicates that the division of the attribute tile and the codec tile may be different.
[0174] According to the above configuration, the flag smatm_codec_tiles_alignment_flag[i] decodes an alignment flag indicating whether the attribute tile division of the attribute video [[ID=,30]]data is the same as the codec tile division of the attribute video sub-bitstream. Only when the tile division of the frame is different from the codec tile division of the displacement sub-bitstream when the division of the tile and the codec tile is the same, the syntax indicating the relationship between the attribute tile and the codec tile is encoded. Thereby, it has the effect of reducing the overhead of the coding amount.
[0175] (Another configuration of the alignment flag) In an alternative configuration, as shown in Figure 27, the smatm_codec_tile_alignment_flag, which indicates whether all attribute tile divisions are the same as the codec tile divisions of the attribute video subbitstream, may be decoded if the attribute tile divisions of all attribute video data are the same as the codec tile divisions of the attribute video subbitstream. Here, the semantics of the syntax element smatm_codec_tile_alignment_flag are as follows: It will look like this. smatm_codec_tile_alignment_flag: smatm_codec_tile_alignment_flag is all attributes Indicates whether the bit tile division is the same as the codec tile division of the attribute video subbitstream. A first value (e.g., 1) indicates that the attribute tile and codec tile divisions are the same, while a second value (e.g., 0) indicates that the attribute tile and codec tile divisions may be different.
[0176] (Decoding the base mesh) Figure 5 is a functional block diagram showing the configuration of the base mesh decoding unit 303. The base mesh decoding unit 303 consists of a mesh decoding unit 3031, a motion information decoding unit 3032, a mesh motion compensation unit 3033, and a reference It consists of a light mesh memory 3034, switches 3035 and 3036, and a skip decoding unit 3037. The base mesh decoding unit 303 processes a base mesh (not shown) before outputting the base mesh. The configuration may also include an inverse quantization unit. Switches 3035 and 3036 are connected to the mesh decoding unit 3031 if the base mesh to be decoded is encoded without referencing other base meshes (e.g., base meshes that have already been encoded and decoded). If, instead, the base mesh to be decoded is encoded by referencing other base meshes (inter-coding), they are connected to the motion compensation side. When motion compensation is performed, the target vertex coordinates are decoded by referencing already decoded vertex coordinates and motion information. If, instead, the base mesh to be decoded is skipped and another base mesh is encoded as the target (skip encoding), they are connected to the skip decoding unit 3037.
[0177] Each base mesh consists of one or more submeshes. If a submesh exists, the tile header of the atlas data subbitstream requires an ID to find the submesh corresponding to the tile. Here, a submesh is a 3D A subset of a mesh is defined by specifying a part of the model, and is a mesh created by dividing a mesh into multiple parts. Meshes are used to finely control a part of a 3D model. By dividing the mesh into subsets, it is possible to define specific ranges of meshes individually. Each sub-mesh has its own vertex coordinates, normal vectors, texture coordinates, etc., and can be manipulated and edited individually. The mesh of a given frame is called a mesh frame.
[0178] The mesh decoding unit 3031 decodes the intra-encoded base mesh encoded stream. The system outputs the base mesh (base mesh vertex positions, base mesh vertex position vector). Encoding methods such as Draco and edgebreaker are used.
[0179] The motion information decoding unit 3032 decodes the intercoded base mesh coded stream and outputs motion information (mesh motion information, mesh motion vector) for each vertex of the reference mesh described later. Entropy coding such as arithmetic coding is used as the coding method.
[0180] The mesh motion compensation unit 3033 performs motion compensation on each vertex of the reference mesh input from the reference mesh memory 3034 based on motion information, and outputs the motion-compensated mesh.
[0181] The reference mesh memory 3034 is a memory that holds the decoded mesh for reference in subsequent decoding processes.
[0182] (Decoding of mesh displacement) Figure 6 is a functional block diagram showing the configuration of the mesh displacement decoding unit 305. The mesh displacement decoding unit 305 consists of a CABAC decoding unit (arithmetic decoding unit 3051, multi-level conversion unit 3052, context selection unit 3056, context initialization unit 3057), an inverse quantization unit 3053, an inverse transformation unit 3054, and a coordinate system transformation unit 3055.
[0183] (Coordinate system) The coordinate systems used for mesh displacement (3D vectors) are the following two types of coordinate systems.
[0184] Cartesian coordinate system (canonical): A Cartesian coordinate system defined commonly across the entire 3D space. (X,Y,Z) coordinate system. A Cartesian coordinate system where direction does not change at the same time (same frame, same tile).
[0185] Local coordinate system: A set of orthogonal coordinates defined for each region or vertex in 3D space. A coordinate system. A Cartesian coordinate system in which the direction can change at the same time (same frame, same tile). A coordinate system with normal (D), tangent (U), and bi-tangent (V) axes. That is, a certain vertex (including a certain vertex) The first axis (D) is defined by the normal vector n_vec on the plane, and the second axis (U) and third axis (V) are defined by two tangent vectors t_vec and b_vec that are orthogonal to the normal vector n_vec. This is a Cartesian coordinate system. n_vec, t_vec, and b_vec are 3-dimensional vectors. (D, U, V) coordinates. The system may also be called the (n, t, b) coordinate system.
[0186] (Decoding and derivation of sequence-level control parameters) Here, the sequence level control of the decoding from the encoded data in the mesh displacement decoding unit 305 Let's explain the parameters.
[0187] Figure 8 shows an example of the syntax for ASVE (ASPS Vdmc Extension), a sequence-level mesh data extended coding parameter set. ASVE is one of the NAL units of atlas information and contains syntax elements to be applied to the atlas information coded stream. The semantics of each syntax element are as follows:
[0188] asve_subdivision_iteration_count: Indicates the number of iterations for mesh subdivision.
[0189] asve_displacement_coordinate_system: Coordinate system transformation information indicating the coordinate system of mesh displacement. If the value is equal to a predetermined first value (e.g., 0), it indicates a Cartesian coordinate system. If the value is another second value If the value is equal to (for example, 1), it indicates the local coordinate system.
[0190] asve_1d_displacement_flag: This flag indicates whether the mesh displacement is one-dimensional or not. A value of true indicates that the mesh displacement is one-dimensional. A value of false indicates that the mesh displacement is three-dimensional.
[0191] (Decoding and derivation of picture / frame-level control parameters) Figure 9 shows the extended encoding in AFPS, which is a picture / frame level parameter set. This is an example of the syntax for parameter information. AFPS is one of the NAL units of atlas information. This includes syntax elements to be applied to the Atlas Information Coded Stream. The semantics of each syntax element are as follows: AFPS includes atlas_frame_mesh_information().
[0192] afve_overriden_flag: This flag indicates whether or not to update the coordinate system of the mesh displacement. If this flag is equal to true, the value of afve_displacement_coordinate_system described below will be used. Update the coordinate system of the mesh displacement. If this flag is equal to false, update the coordinate system of the mesh displacement. The standard system will not be updated.
[0193] afve_subdivision_iteration_count: Indicates the number of mesh subdivision iterations.
[0194] afve_displacement_coordinate_system: Coordinate system transformation information indicating the coordinate system of mesh displacement. If the value is equal to the first value (e.g., 0), it indicates the Cartesian coordinate system. If the value is equal to the second value (e.g., 1), it indicates the local coordinate system. If no syntax elements appear, the default coordinate system is assumed to be the coordinate system indicated by ASPS, assuming the value is the value decoded by ASPS.
[0195] (Operation of the mesh displacement decoding unit) The arithmetic decoding unit 3051 decodes the mesh displacement coding stream, which has been arithmetically coded according to the value (context) representing the random variable, and outputs a binary signal. The binary signal may be an alpha code or a k-th order exponential Golomb code. An exponential Golomb code consists of a prefix and a suffix code. The prefix is an exponentially increasing value, and the suffix is its remainder. When encoding and decoding the variable rem with an exponential Golomb code, the prefix and suffix of the exponential Golomb code are also called the prefix and suffix of rem.
[0196] The multi-leveling unit 3052 decodes the binary signal into a multi-level signal, which is a quantized mesh displacement Qdisp.
[0197] The context selection unit 3056 (context memory) has memory for holding contexts, derives a context used for arithmetic decoding of mesh displacements according to the state, and updates the value as necessary.
[0198] The context initialization unit 3057 initializes the context (the probability of a binary signal occurring).
[0199] (Derivation process of mesh displacement) The mesh displacement decoding unit 305 decodes the syntax elements dismu_nz_subBlock, dismu_coeff_abs_level_gt0, dismu_coeff_abs_level_gt1, dismu_coeff_abs_level_gt2, dismu_coeff_abs_level_gt3, dismu_coeff_abs_level_rem, and dismu_coeff_sign through the following process. Then, decode the mesh displacement Qdisp.
[0200] The inverse quantization unit 3053 performs inverse quantization based on the quantization scale value iscale and decodes the mesh displacement Tdisp after transformation (e.g., wavelet transform). Tdisp may be in a Cartesian coordinate system or a local coordinate system. iscale is a value derived from the quantization parameters of each component of the mesh displacement image. The inverse quantization is shown by subMeshID (= displSubMeshID). It can also be done at the bushmesh level. Tdisp[subMeshID][0][] = (Qdisp[subMeshID][0][] * iscale[0] + iscaleOffset) >> iscaleShift Tdisp[subMeshID][1][] = (Qdisp[subMeshID][1][] * iscale[1] + iscaleOffset) >> iscaleShift Tdisp[subMeshID][2][] = (Qdisp[subMeshID][2][] * iscale[2] + iscaleOffset) >> iscaleShift Here, iscaleOffset = 1 << (iscaleShift - 1). iscaleShift may be a predetermined constant, or it may be encoded at the sequence level, picture / frame level, submesh level indicated by subMeshID (= displSubMeshID), tile / patch level, etc., and the encoded data You may also use the decoded value.
[0201] The inverse transform unit 3054 performs an inverse transform g (for example, an inverse wavelet transform) and decodes the mesh displacement d. d[0][] = g(Tdisp[subMeshID][0][]) d[1][] = g(Tdisp[subMeshID][1][]) d[2][] = g(Tdisp[subMeshID][2][]) The coordinate system conversion unit 3055 converts the mesh displacement (coordinate system of the mesh displacement) to the Cartesian coordinate system based on the value of the coordinate system conversion information displacementCoordinateSystem. Specifically, when displacementCoordinateSystem == 1, it converts the displacement in the local coordinate system to the displacement in the Cartesian coordinate system. Here, d is a 3D vector indicating the mesh displacement before coordinate system conversion. disp is a 3D vector indicating the mesh displacement after coordinate system conversion, and it is in the Cartesian coordinate system. n_vec, t_vec, b_vec are 3D vectors (in the Cartesian coordinate system) corresponding to each axis of the local coordinate system of the target region or target vertex. Here, d is a 3D vector indicating the mesh displacement before coordinate system conversion. disp is a 3D vector indicating the mesh displacement after coordinate system conversion, and it is in the Cartesian coordinate system. n_vec, t_vec, b_vec are 3D vectors (in the Cartesian coordinate system) corresponding to each axis of the local coordinate system of the target region or target vertex. if (displacementCoordinateSystem == 0) { disp = d } else if (displacementCoordinateSystem == 1){ disp = d[0] * n_vec3 + d[1] * t_vec3 + d[2] * b_vec3 } Here, n_vec3, t_vec3, b_vec3 are 3D vectors (in the Cartesian coordinate system) corresponding to each axis of the local coordinate system of the target region with fluctuations suppressed. For example, the vectors of the coordinate system used for decoding are decoded from the previous coordinate system and the current coordinate system as follows. For example, the vectors of the coordinate system used for decoding are decoded from the previous coordinate system and the current coordinate system as follows.
[0202] n_vec3 = (w*n_vec3 + (WT - w)*n_vec)>>wShift t_vec3 = (w*t_vec3 + (WT - w)*t_vec)>>wShift b_vec3 = (w*b_vec3 + (WT - w)*b_vec)>>wShift Here, for example, wShift = 2, 3, 4, WT = 1<<wShift, w = 1..WT - 1. For example, when w = 3 and wShift = 3, the coordinate system vectors are decoded as follows.
[0203] n_vec3 = (3*n_vec3 + 5*n_vec)>>3 t_vec3 = (3*t_vec3 + 5*t_vec)>>3 b_vec3 = (3*b_vec3 + 5*b_vec)>>3 (Mesh reconstruction) Figure 7 is a functional block diagram showing the configuration of the mesh reconstruction unit 307. The mesh reconstruction unit 307 consists of a mesh division unit 3071 and a mesh deformation unit 3072.
[0204] The mesh division unit 3071 divides the base mesh output from the base mesh decoding unit 303. Divide and generate a subdivided mesh.
[0205] Figure 12(a) shows a portion of the base mesh (a triangle), where the triangle has vertices v1, v It consists of 2 and v3. v1, v2, and v3 are 3D vectors. The mesh division unit 3071 generates a divided mesh by adding new vertices v12, v13, and v23 in the middle of each side of the triangle. Then, output (Figure 12(b)). v12 = (v1 + v2) / 2 v13 = (v1 + v3) / 2 v23 = (v2 + v3) / 2 The following is also acceptable. v12 = (v1 + v2 + 1) >> 1 v13 = (v1 + v3 + 1) >> 1 v23 = (v2 + v3 + 1) >> 1 The mesh deformation unit 3072 receives the divided mesh and the mesh displacement, and the mesh displacement d12, A deformed mesh is generated and output by adding d13 and d23 (Figure 12(c)). The mesh displacement is the output of the mesh displacement decoding unit 305 (coordinate system transformation unit 3055). d12, d13, and d23 are the mesh displacements corresponding to the vertices v12, v13, and v23 added by the mesh division unit 3071. v12' = v12 + d12 v13' = v13 + d13 v23' = v23 + d23 Note that d12 = disp[0][], d23 = disp[1][], and d23 = disp[3][] are also acceptable.
[0206] (Structure of tile information syntax) An Atlas frame can be divided into one or more partition units, and tiles can be constructed from these units (partitions). Typical examples include the following: - The entire atlas frame is used as a single tile without dividing it (afti_single_tile_in_atlas_frame_flag==1). The atlas frame is divided into multiple partitions, and each partition is used as one tile (afti_single_tile_in_atlas_frame_flag==0 and afti_single_partition_per_tile_flag==1). The atlas frame is divided into multiple partitions, and one or more partitions that are consecutive horizontally and vertically are used as a single tile (afti_single_tile_in_atlas_frame_flag==0 and afti_single_partition_per_tile_flag==0).
[0207] The atlas frame can be divided into tile partitions (hereinafter also called partitions) of NumPartitionColumns * NumPartitionRows, and when dividing, you can choose to divide the frame at equal intervals or at specified units. NumPartitionColumns and NumPartitionRows are the number of partition divisions in the horizontal and vertical directions, respectively. be.
[0208] Note that tiles are not limited to atlas frames; they can also be attributes, geometries, displacements, or meshes. In other words, the following syntax elements and their bitstream conformance conditions can also be used for attribute, geometry, displacement, and mesh tiles.
[0209] Figure 10 shows the syntax for tile information. Tile information may also be provided using the atlas_frame_tile_information() function defined in the ISO / IEC 23090-5 V3C standard.
[0210] The tile information decoding unit 3022 decodes the syntax element afti_single_tile_in_atlas_frame_flag. afti_single_tile_in_atlas_frame_flag is a binary flag that indicates whether or not the atlas frame consists of a single tile. A value (e.g., 1) indicating that it is, or that the atlas frame is composed of multiple tiles. It has a value (e.g., 0) that indicates it is multiple. If the value indicates a number of tiles, the tile information decoding unit 3022 decodes the syntax element afti_uniform_partition_spacing_flag. Here, afti_uniform_partition_spacing_flag is a binary flag that indicates whether or not to divide the atlas frame into equally spaced partitions. A value (e.g., 1) that indicates dividing the frame into equally spaced partitions, or It has a value (e.g., 0) that indicates dividing the last frame into partitions with different intervals.
[0211] The tile information decoding unit 3022 decodes parameters indicating the position and size of the tiles.
[0212] If afti_uniform_partition_spacing_flag is valued as 1, the tile information decoding unit 3022 decodes the syntax elements afti_partition_cols_width_minus1 and afti_partition_cols_width_minus1, which indicate the width (column width) and height (row height) of each partition except the rightmost column and the bottommost row. For each i=0..NumPartitionColumns-1, j=0..NumPartitionRows Next, we determine PartitionPosX[i], PartitionPosY[j], PartitionWidth[i], and PartitionHeight[j], which represent the x, y coordinates, width, and height of the top left corner of each partition, as follows.
[0213] partitionWidth = ( afti_partition_cols_width_minus1 + 1 ) * 64 NumPartitionColumns = asps_frame_width / partitionWidth PartitionPosX
[0000] = 0 PartitionWidth
[0000] = partitionWidth for( i = 1; i < NumPartitionColumns - 1; i++ ) { PartitionPosX[ i ] = PartitionPosX[ i - 1 ] + PartitionWidth[ i - 1 ] PartitionWidth[ i ] = partitionWidth } partitionHeight = (afti_partition_rows_height_minus1 + 1) * 64 NumPartitionRows = asps_frame_height / partitionHeight PartitionPosY
[0000] = 0 PartitionHeight
[0000] = partitionHeight for( j = 1; j < NumPartitionRows - 1; j++ ) { PartitionPosY[ j ] = PartitionPosY[ j - 1 ] + PartitionHeight[ j - 1 ] PartitionHeight[ j ] = partitionHeight } If afti_uniform_partition_spacing_flag is a value of 0, the tile information decoding unit 3022 decodes the syntax elements afti_num_partition_columns_minus1 and afti_num_partition_rows_minus1, which indicate the number of tile partitions in the horizontal and vertical directions.
[0214] In each i=0..NumPartitionColumns-1, j=0..NumPartitionRows, each partition The x, y coordinates, width, and height of the top-left corner, known as PartitionPosX[i], PartitionPosY[j], PartitionWidth[i], and PartitionHeight[j], are calculated as follows.
[0215] NumPartitionColumns = afti_num_partition_columns_minus1 + 1 PartitionPosX
[0000] = 0 partitionWidth
[0000] = ( afti_partition_column_width_minus1
[0000] + 1 ) * 64 for( i = 1; i < NumPartitionColumns - 1; i++ ) { PartitionPosX[ i ] = PartitionPosX[ i - 1 ] + PartitionWidth[ i - 1 ] PartitionWidth[ i ] = ( afti_partition_column_width_minus1[ i ] + 1 ) * 64 } NumPartitionRows = afti_num_partition_rows_minus1 + 1 PartitionPosY
[0000] = 0 PartitionHeight
[0000] = ( afti_partition_row_height_minus1
[0000] + 1 ) * 64 for( j = 1; j < NumPartitionRows - 1; j++ ) { PartitionPosY[ j ] = PartitionPosY[ j - 1 ] + PartitionHeight[ j - 1 ] PartitionHeight[ j ] = ( afti_partition_row_height_minus1[ j ] + 1 ) * 64 } Also, if the number of partitions in the horizontal and vertical directions is two or more, the rightmost column and the bottommost column are used. The PartitionPosX[i], PartitionPosY[j], PartitionWidth[i], and PartitionHeight[j] values, which represent the x, y coordinates, width, and height of the top-left partition of each row, are decoded as follows.
[0216] PartitionPosX[ NumPartitionColumns - 1 ] = PartitionPosX[ NumPartitionColumns - 2 ] + PartitionWidth[ NumPartitionColumns - 2 ] PartitionWidth[ NumPartitionColumns - 1 ] = asps_frame_width - PartitionPosX[ NumPartitionColumns - 1 ] PartitionPosY[ NumPartitionRows - 1 ] = PartitionPosY[ NumPartitionRows - 2 ] + partitionHeight[ NumPartitionRows - 2 ] PartitionHeight[ NumPartitionRows - 1 ] = asps_frame_height - PartitionPosY[ NumPartitionRows - 1 ] Here, the width and height of each partition are set to multiples of 64, but they are not limited to 64; you can replace 64 with 32, 128, or 256.
[0217] The tile information decoding unit 3022 processes the syntax element afti_single_partition_per_tile_flag Decode. Here, afti_single_partition_per_tile_flag means that each tile is a single partition A flag indicating whether a tile consists of only partitions; a value (e.g., 1) indicates that each tile consists of only a single partition, or each tile consists of multiple partitions. It has a value (e.g., 0) that indicates that there are multiple partitions. If afti_single_partition_per_tile_flag is a value that indicates multiple partitions, the tile information decoding unit 3022 decodes the syntax element afti_num_tiles_in_atlas_frame_minus1 and selects the number of tiles from one or more partitions. The following process is performed to decode the lattice. Here, afti_num_tiles_in_atlas_frame_minus1 is the number of tiles that make up the atlas frame.
[0218] The tile information decoding unit 3022 performs the following for each i=0..afti_num_tiles_in_atlas_frame_minus1: Decode the syntax elements afti_top_left_partition_idx[i], afti_bottom_right_partition_column_offset[i], and afti_bottom_right_partition_row_offset[i]. afti_top_left_partition_idx[i] is the top-left corner (point) of the i-th tile. The index of the partition to which it is located, afti_bottom_right_partition_column_offset[ `afti_bottom_right_partition_row_offset[i]` is the horizontal offset amount of the bottom-right edge of the i-th tile relative to the top-left edge of the i-th tile, and `afti_bottom_right_partition_row_offset[i]` is the vertical offset amount of the bottom-right edge of the i-th tile relative to the top-left edge of the i-th tile.
[0219] Based on the decoded syntax above, the top left horizontal, height, and bottom right of each tile i The indices of the horizontal and vertical partitions, topLeftColumn[i], topLeftRow[i], bottomRightColumn[i], and bottomRightRow[i], are calculated as follows.
[0220] topLeftColumn[ i ] = afti_top_left_partition_idx[ i ] % NumPartitionColumns topLeftRow[ i ] = afti_top_left_partition_idx[ i ] / NumPartitionColumns bottomRightColumn[ i ] = topLeftColumn[ i ] + afti_bottom_right_partition_column_offset[ i ] bottomRightRow[ i ] = topLeftRow[ i ] + afti_bottom_right_partition_row_offset[ i ] Here, bottomRightColumn[i] and bottomRightRow[i] may be less than or equal to (asps_frame_width + 63) / 64 - 1 and (asps_frame_height + 63) / 64 - 1, respectively.
[0221] In a 3D data decoding device 31 that decodes mesh data or point cloud data, the syntax element indicating the position of the tile is decoded, and the column of the top-left partition of the tile, topLeftColumn, is decoded. The system has means for decoding the row of the top-left partition (topLeftRow), the column of the bottom-right partition of the tile (bottomRightColumn), and the row of the bottom-right partition (bottomRightRow), and for 3D data decoding. Device 31 has a column of partitions for the i-th tile (topLeftColumn[i] and bottomRightColumn[i]), a column of partitions for the j-th tile (topLeftColumn[j]), and the i-th tile The partition row (topLeftRow[i], bottomRightRow[i]) and the partition of the j-th tile For a row of the tytion (topLeftRow[j]), a specific bitstream conformance clause The 3D data decoding device 31 may decode a bitstream that satisfies the following bitstream conformance conditions.
[0222] (Configuration of the 3D data encoding device according to the first embodiment) Figure 13 is a functional block diagram showing the schematic configuration of the 3D data encoding device 11 according to the first embodiment. The 3D data encoding device 11 comprises an atlas information encoding unit 101, a base mesh encoding unit 103, a base mesh decoding unit 104, a mesh displacement update unit 106, a mesh displacement encoding unit 107, and The 3D data encoding device 11 consists of a mesh displacement decoding unit 108, a mesh reconstruction unit 109, an attribute update unit 110, a padding unit 111, a color space conversion unit 112, an attribute encoding unit 113, a multiplexing unit 114, and a mesh separation unit 115. The 3D data encoding device 11 takes additional information, atlas information, a base mesh, mesh displacement, a mesh, and an attribute image as input as 3D data and outputs encoded data.
[0223] The Atlas information coding unit 101 codes the Atlas information and generates an Atlas information coding stream. Outputs "mu".
[0224] The base mesh coding unit 103 encodes the base mesh and the base mesh coding unit Outputs a stream. The encoding scheme used is Draco, among others.
[0225] The base mesh decoding unit 104 is the same as the base mesh decoding unit 303, so its explanation is omitted.
[0226] The mesh displacement update unit 106 uses the (original) base mesh and the decoded base mesh. Based on the mesh, the mesh displacement is adjusted and the updated mesh displacement is output.
[0227] The mesh displacement coding unit 107 encodes the updated mesh displacement and generates a mesh displacement code. Outputs a stream. Encoding schemes such as VVC and HEVC are used.
[0228] The mesh displacement decoding unit 108 is the same as the mesh displacement decoding unit 305, so its description is omitted.
[0229] The mesh reconstruction unit 109 is the same as the mesh reconstruction unit 307, so its description is omitted.
[0230] The attribute update unit 110 receives the (original) mesh, the reconstructed mesh output from the mesh reconstruction unit 109 (mesh deformation unit 3072), and the attribute image as input, updates the attribute image to match the position (coordinates) of the reconstructed mesh, and outputs the updated attribute image.
[0231] The padding unit 111 receives an attribute image and pads areas where the pixel value is empty. Perform the sizing process.
[0232] The color space conversion unit 112 performs a color space conversion from RGB format to YCbCr format.
[0233] The attribute encoding unit 113 converts the YCbCr format output from the color space conversion unit 112. It encodes attribute images and outputs attribute video streams. Encoding methods such as VVC and HEVC are used.
[0234] The multiplexing unit 114 includes an Atlas information coded stream, a base mesh coded stream, The mesh displacement encoded stream and attribute video stream are multiplexed and output as encoded data. Multiplexing methods include byte stream format and ISOBMFF. Use this.
[0235] (Operation of the mesh separation unit) The mesh separation unit 115 generates a base mesh and mesh displacement from the mesh.
[0236] Figure 17 is a functional block diagram showing the configuration of the mesh separation unit 115. The mesh separation unit 115 consists of a mesh thinning unit 1151, a mesh division unit 1152, and a mesh displacement extraction unit 1153.
[0237] The mesh thinning unit 1151 generates a base mesh by thinning out some vertices from the mesh.
[0238] Figure 18(a) shows a portion of the mesh, where the mesh consists of vertices v1, v2, v3, v4, v5, It consists of v6. v1, v2, v3, v4, v5, and v6 are each 3D vectors. The mesh thinning unit 1151 generates and outputs a base mesh by thinning vertices v4, v5, and v6 (Figure 18(b)).
[0239] The mesh division unit 1152, like the mesh division unit 3071, divides the base mesh and generates a divided mesh (Figure 18(c)).
[0240] v4' = (v1 + v2) / 2 v5' = (v1 + v3) / 2 v6' = (v2 + v3) / 2 The mesh displacement derivation unit calculates the displacement for vertices v4', v5', and v6' based on the mesh and the subdivided mesh. The displacements d4, d5, and d6 of vertices v4, v5, and v6 are derived as mesh displacements and output (Figure 18(d)). .
[0241] d4 = v4 - v4' d5 = v5 - v5' d6 = v6 - v6' (Encoding of Atlas information) Figure 14 is a functional block diagram showing the configuration of the Atlas information coding unit 101. The encoding unit 101 consists of an extended information encoding unit 1011, a tile information encoding unit 1012, a parameter encoding unit 1013, and an additional information encoding unit 1014.
[0242] The extended information encoding unit 1011 encodes extended encoding parameters related to the mesh data.
[0243] The tile information encoding unit 1012 encodes the number of tiles and tile IDs referenced at the picture / frame level.
[0244] The parameter coding unit 1013 encodes coding parameters related to the 3D data.
[0245] The additional information encoding unit 1014 encodes additional information such as SEI. , a loop of the number of tiles in a frame at smtm_num_tiles_minus1 and a displacement subbit Loop through the number of codec tiles in the stream, which is smtm_num_codec_tiles_minus1. Encodes mapping information that shows the relationship between a given tile and a codec tile.
[0246] Furthermore, the additional information encoding unit 1014 loops through the number of attribute video data smatm_attribute_video_minus1, the number of attribute tiles contained in that data smatm_num_tiles_minus1[i], and the codec tiles contained in a certain attribute video subbitstream. Encodes mapping information that shows the relationship between a given tile and a codec tile, looping through the number of tiles smatm_num_codec_tiles_minus1.
[0247] In another configuration, the division of a tile / attribute tile is the same as the division of a codec tile. The above additional information decoding unit 1014 includes a flag smtm_codec_tile_alignment_flag / smatm_codec_tile_alignment_flag / smatm_codec_tile_alignment_flag[i] that indicates whether they are the same. In this, the division method and displacement / attribute video subbit of the i-th tile / attribute tile If the codec tile segmentation method of the stream includes only the region, the first value (e.g., 1) Otherwise, a second value (e.g., 0) may be encoded.
[0248] (Encoding of the base mesh) Figure 15 is a functional block diagram showing the configuration of the base mesh coding unit 103. The mesh coding unit 103 consists of a mesh coding unit 1031, a mesh decoding unit 1032, a motion information coding unit 1033, a motion information decoding unit 1034, a mesh motion compensation unit 1035, a reference mesh memory 1036, switches 1037 and 1038, and a skip coding unit 1039. The base mesh coding unit 103 may also include a base mesh quantization unit (not shown) after the input of the base mesh. Switches 1037 and 1038 are connected to the side that does not perform motion compensation or skipping when coding a base mesh without referencing another base mesh (e.g., an already coded base mesh) (intra coding). Instead, they are connected to the side that performs motion compensation when coding a base mesh by referencing another base mesh (inter coding). Instead, they are connected to the side that performs skip coding when skipping the coding of the base mesh (skip coding).
[0249] The mesh coding unit 1031 has an intra coding function, intra-codes the base mesh, and outputs a base mesh coded stream. The coding method used is Draco, etc. Yes, they are.
[0250] The mesh decoding unit 1032 is the same as the mesh decoding unit 3031, so its description is omitted.
[0251] The motion information coding unit 1033 has an intercoding function, intercodes the base mesh, and outputs a base mesh coded stream. Entropy coding such as arithmetic coding is used as the coding method.
[0252] The motion information decoding unit 1034 is the same as the motion information decoding unit 3032, so its explanation is omitted.
[0253] The mesh motion compensation unit 1035 is the same as the mesh motion compensation unit 3033, so its explanation is omitted.
[0254] The reference mesh memory 1036 is the same as the reference mesh memory 3034, so its description is omitted.
[0255] (Encoding of mesh displacement) Figure 16 is a functional block diagram showing the configuration of the mesh displacement coding unit 107. The encoding unit 107 consists of a coordinate system transformation unit 1071, a transformation unit 1072, a quantization unit 1073, and a displacement mapping unit 1074. It consists of an image packing unit and a displacement coding unit. As shown in the figure, the mesh displacement coding unit 107 may further include a video coding unit 1075. Alternatively, the video coding unit 1075 is The mesh displacement coding unit 107 does not include the encoding of the displacement image, and an external image coding device is used. It is also acceptable to use this configuration.
[0256] The coordinate system transformation unit 1071 transforms the coordinate system of the mesh displacement from the Cartesian coordinate system to the coordinate system that encodes the displacement (e.g., the local coordinate system) based on the value of the coordinate transformation information displacementCoordinateSystem. Here, disp is a 3D vector representing the mesh displacement before the coordinate system transformation, d is a 3D vector representing the mesh displacement after the coordinate system transformation, and n_vec, t_vec, and b_vec are 3D vectors (in the Cartesian coordinate system) representing each axis of the local coordinate system.
[0257] if (displacementCoordinateSystem == 0) { d = disp } else if (displacementCoordinateSystem == 1){ d = (disp * n_vec, disp * t_vec, disp * b_vec) } The mesh displacement coding unit 107 converts the value of displacementCoordinateSystem into picture / frame It's okay to update it at the M level.
[0258] When encoding the displacementCoordinateSystem at the sequence level, the configuration shown in Figure 9 is Use the `intax`. `asve_displacement_coordinate_system` contains the field of the Cartesian coordinate system. Set 0 to 1 for the local coordinate system.
[0259] To change the displacementCoordinateSystem at the picture / frame level, use the syntax shown in Figure 12. Set afve_overriden_flag to 1 if the coordinate system is to be updated, or 0 if the coordinate system is not to be updated. Set afve_displacement_coordinate_system to 0 for the Cartesian coordinate system, or 1 for the local coordinate system.
[0260] The transformation unit 1072 performs a transformation f (for example, a wavelet transform) and decodes the transformed mesh displacement Tdisp. For pos=0..NumDisp-1, the following is performed. Here, NumDisp is the mesh apex. The number of points.
[0261] dispCoeffArray[v][d] = f(d[d][v]) The quantization unit 1073 performs quantization based on the quantization scale value scale, which is decoded from the quantization parameters of each component of the mesh displacement, and decodes the quantized mesh displacement dispQuantCoeffArray.
[0262] Vcount0 = 0 for( i = 0; i < subdivisionIterationCount; i++ ) { vcount1 = levelOfDetailCounts[ i ] for( v = vcount0; v < vcount1; v++ ) { for( d = 0; d < DisplacementDim; d++ ) { dispQuantCoeffArray[v][d] = dispCoeffArray[ v ][ d ] / iscale[ i ][ d ] } } vcount0 = vcount1 } Alternatively, the scale value can be approximated by a power of 2, and dispQuantCoeffArray can be derived using the following formula. .
[0263] Vcount0 = 0 for( i = 0; i < subdivisionIterationCount; i++ ) { scale[i] = 1 << scale2[i] vcount1 = levelOfDetailCounts[ i ] for( v = vcount0; v < vcount1; v++ ) { for( d = 0; d < DisplacementDim; d++ ) { dispQuantCoeffArray[v][d] = dispCoeffArray[v][d] >> scale2[i][d] } } vcount0 = vcount1 } The displacement mapping unit 1074 generates an image dispQuantCoeffFrame from the quantized mesh displacement dispQuantCoeffArray based on the value of the displacement mapping parameter displacementChromaLocationType.
[0264] The displacement mapping unit 1074 performs the first generation of the (quantized) mesh displacement array as follows: The minutes dispQuantCoeffArray[v][0] may be mapped to the luminance (Y) image component. For an image with width W and height H, apply the following to (y=0..H-1, x=0..W-1):
[0265] H = origHeight shift = (1 << bitDepth) >> 1 dispQuantCoeffFrame[x][ y][0] = dispQuantCoeffArray[v][0] + shift dispQuantCoeffFrame[x][ H+y][0] = dispQuantCoeffArray[v][1] + shift dispQuantCoeffFrame[x][2*H+y][0] = dispQuantCoeffArray[v][2] + shift v++ dispQuantCoeffFrame[x / 2][ y / 2][1] = shift dispQuantCoeffFrame[x / 2][H / 2+y / 2][1] = shift dispQuantCoeffFrame[x / 2][ H+y / 2][1] = shift dispQuantCoeffFrame[x / 2][ y / 2][2] = shift dispQuantCoeffFrame[x / 2][H / 2+y / 2][2] = shift dispQuantCoeffFrame[x / 2][ H+y / 2][2] = shift Alternatively, the displacement mapping unit 1074 may encode the mesh displacement for each submesh.
[0266] Furthermore, the processing may be switched depending on the DecGeoChromaFormat. That is, if DecGeoChromaFormat=1 (4:2:0), the above processing is performed, and if DecGeoChromaFormat=3 (4:4:4), the following processing is performed.
[0267] dispQuantCoeffFrame[x][y][d] = dispQuantCoeffArray[v][0] dispQuantCoeffFrame[x][y][d] = dispQuantCoeffArray[v][1] dispQuantCoeffFrame[x][y][d] = dispQuantCoeffArray[v][2] v++ The mesh displacement encoding unit 107 may update the origHeight and origWidth values at the picture / frame level.
[0268] The video encoding unit 1075 encodes an image in YCbCr4:2:0 format, including a (quantized) mesh displacement image, and outputs a mesh displacement encoded stream. The encoding scheme used may include VVC or HEVC.
[0269] The video encoding unit 1075 divides the mesh displacement image into slices for each origHeight and encodes them. Alternatively, the origHeight may be aligned to a predetermined size according to the CTU size. .
[0270] [Application Examples] The 3D data encoding device 11 and the 3D data decoding device 31 described above can be installed and used in various devices that transmit, receive, record, and reproduce 3D data. The 3D data may be natural 3D data captured by a camera or the like, or artificial 3D data (including CG and GUI) generated by a computer or the like.
[0271] The embodiments of the present invention are not limited to those described above, and various modifications are possible within the scope of the claims. That is, embodiments obtained by combining technical means that have been appropriately modified within the scope of the claims are also included in the technical scope of the present invention. [Industrial applicability]
[0272] Embodiments of the present invention can be suitably applied to a 3D data decoding device that decodes encoded data obtained by encoding 3D data, and a 3D data encoding device that generates encoded data obtained by encoding 3D data. Furthermore, it can be suitably applied to the data structure of encoded data generated by the 3D data encoding device and referenced by the 3D data decoding device. [Explanation of Symbols]
[0273] 11 3D Data Encoding Device 101 Atlas Information Coding Unit 1011 Extended Information Encoding Unit 1012 Tile Information Encoding Unit 1013 Parameter coding section 1014 Additional Information Decoding Unit 103 Base Mesh Coding Unit 1031 Mesh coding section 1032 Mesh Decoding Unit 1033 Motion Information Encoding Unit 1034 Motion Information Decoding Unit 1035 Mesh motion compensation unit 1036 Reference Mesh Memory 1037 Switch 1038 Switch 1039 Skip coding section 104 Base Mesh Decoding Unit 106 Mesh displacement update section 107 Mesh Displacement Coding Unit 1071 Coordinate System Transformation Unit 1072 Conversion Unit 1073 Quantization section 1074 Displacement Unmapping Section 1075 Image Encoding Unit 108 Mesh displacement decoding unit 109 Mesh Reconstruction Section 110 Attribute Update Section 111 Padding section 112 Color Space Conversion Unit 113 Attribute Encoding Section 114 Multiplexer 115 Mesh separation section 1151 Mesh thinning section 1152 Mesh division section 1153 Mesh displacement derivation section 21 Network 31 3D Data Decoder 301 Demultiplexer 302 Atlas Information Decoding Unit 3021 Parameter Decoding Unit 3022 Tile Information Decoding Unit 3023 Extended Information Decoding Unit 3024 Additional Information Decoding Unit 303 Base Mesh Decoding Unit 3031 Mesh Decoding Unit 3032 Motion Information Decoding Unit 3033 Mesh motion compensation unit 3034 Reference Mesh Memory 3035 Switch 3036 Switch 3037 Skip Decoding Unit 305 Mesh displacement decoding unit 3051 Video Decoding Unit 3052 Displacement Unmapping Section 3053 Inverse quantization section 3054 Inverse Transformer 3055 Coordinate System Transformation Unit 307 Mesh Reconstruction Section 306 Attribute Decoding Unit 3071 Mesh division section 3072 Mesh deformation area 308 Color Space Conversion Unit 41 3D Data Display Device
Claims
1. A 3D data decoding device for decoding mesh data or point cloud data, comprising an atlas information decoding unit that decodes atlas information from encoded data in which the mesh data or point cloud data has been encoded, wherein the atlas information decoding unit decodes tile information and additional information indicating the relationship between tiles and codec tiles.
2. A 3D data decoding device for decoding mesh data or point cloud data, comprising an atlas information decoding unit for decoding atlas information from encoded data in which mesh data or point cloud data has been encoded, wherein the atlas information decoding unit decodes attribute tile information and additional information indicating the relationship between attribute tiles and codec tiles.
3. The 3D data decoding device according to claim 1 or 2, characterized in that the above additional information is a mapping flag indicating the correspondence between tile information and codec tile information.
4. The 3D data decoding device according to claim 1 or 2, characterized in that the above additional information is an index of codec tiles indicating the correspondence between tile information and codec tile information.
5. The above additional information includes a flag indicating whether the tile division and the codec tile division are the same, and in the above atlas information decoding unit, each i-th tile is a displacement subbits If it contains only the region of the i-th codec tile of the trim, decode the first value, and then The 3D data decoding device according to claim 1, 3, or 4, characterized in that, in the case of an external value, it decodes a second value.
6. The above additional information includes a flag indicating whether the division of the attribute tile and the division of the codec tile are the same, and in the above atlas information decoding unit, each i-th attribute The codec tile only applies to the region of the i-th codec tile of the attribute video subbitstream. A 3D data decoding device according to claim 2, 3, or 4, characterized in that it decodes the first value if it contains, and otherwise decodes the second value.
7. A 3D data encoding device for decoding mesh data or point cloud data, comprising an atlas information encoding unit that encodes atlas information from encoded data obtained by encoding mesh data or point cloud data, wherein the atlas information encoding unit encodes tile information and additional information indicating the relationship between tiles and codec tiles.
8. A 3D data encoding device for decoding mesh data or point cloud data, comprising an atlas information encoding unit for encoding atlas information from encoded data obtained by encoding mesh data or point cloud data, wherein the atlas information encoding unit encodes attribute tile information and additional information indicating the relationship between attribute tiles and codec tiles.