3D data decoder and 3D data encoder
The 3D data decoding and encoding devices improve encoding efficiency by decoding and encoding mesh displacement in sub-mesh units, addressing inefficiencies in existing methods and achieving high-quality 3D data processing.
Patent Information
- Application Number
- JP2024004374
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-29
AI Technical Summary
Existing 3D data encoding methods, such as those described in Non-Patent Document 1, struggle with encoding and decoding mesh displacement in units smaller than a frame (sub-mesh units), leading to inefficiencies in encoding granularity.
A 3D data decoding device and encoding device that utilize a sub-mesh decoding unit, a base mesh decoding unit, and a mesh displacement decoding unit to decode and encode 3D data in sub-mesh units, improving the encoding efficiency of mesh displacement.
Enhances the encoding efficiency of mesh displacement and allows for high-quality encoding and decoding of 3D data.
Smart Images

Figure 2025110508000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to a 3D data encoding device and a 3D data decoding device.
Background Art
[0002] In order to efficiently transmit or record 3D data, a 3D data encoding device that converts 3D data into a 2D image and encodes it using a video image encoding method to generate encoded data, and a 3D data decoding device that decodes a 2D image from the encoded data to reconstruct 3D data exist.
[0003] Specific 3D data encoding methods include, for example, ISO / IEC 23090-5 V3C (Volumetric Video-based Coding) and V-PCC (Video-based Point Cloud Compression) of MPEG-I. V3C can encode and decode a point cloud composed of point positions and attribute information. Further, ISO / IEC 23090-12 (MPEG Immersive Video, MIV) and ISO / IEC 23090-29 (Video-based Dynamic Mesh Coding, V-DMC) under standardization are also used for encoding and decoding multi-view video and mesh video. The V-DMC method is disclosed in the latest draft document of Non-Patent Document 1.
[0004]
[0005] In these 3D data encoding methods, the geometry and attributes constituting the 3D data are encoded and decoded using a moving image encoding method such as H.265 / HEVC (High Efficiency Video Coding) or H.266 / VVC (Versatile Video Coding) as an image.
[0005] In the case of a point group, the geometry image is the depth on the projection plane, and the attribute image is the image obtained by projecting the attributes onto the projection plane.
[0006] 3D data (mesh) such as in Non-Patent Document 1 is composed of a base mesh, mesh displacement, and texture mapping image. The encoding of the base mesh can use a vertex encoding method such as Draco. For the encoding of mesh displacement, in addition to the method of encoding the mesh displacement image obtained by two-dimensioning the mesh displacement using a video codec, there is a method of directly encoding by arithmetic coding. The texture mapping image is encoded as an attribute image using a video codec. As the video codec, the above-mentioned HEVC or VVC can be used.
Prior Art Documents
Non-Patent Documents
[0007]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0008] In the 3D data encoding method in Non-Patent Document 1, the mesh displacement (mesh displacement array, mesh displacement image) constituting the 3D data (mesh) can be encoded and decoded using an arithmetic coding method. Although the mesh displacement is arithmetic-coded to be encoded and decoded in frame units, there is a problem that it cannot be encoded and decoded in units smaller than a frame (sub-mesh units).
[0009] An object of the present invention is to improve the encoding granularity of mesh displacement and encode and decode 3D data with high efficiency in the encoding and decoding of 3D data using an arithmetic coding method.
Means for Solving the Problem
[0010] In order to solve the above problems, a 3D data decoding device according to an aspect of the present invention is a 3D data decoding device that decodes mesh data or point cloud data. The 3D data decoding device includes a sub-mesh decoding unit that decodes sub-mesh information from encoded data in which the mesh data or point cloud data is encoded, a base mesh decoding unit that decodes a base mesh from the encoded data and the sub-mesh information, a mesh displacement decoding unit that decodes a mesh displacement from the encoded data and the sub-mesh information, and a mesh reconstruction unit that decodes a mesh from the decoded base mesh and the mesh displacement. In the mesh displacement decoding unit, the mesh displacement is decoded from the encoded data using the sub-mesh information decoded by the sub-mesh decoding unit.
[0011] In order to solve the above problems, a 3D data encoding device according to an aspect of the present invention is a 3D data encoding device that encodes mesh data or point cloud data. The 3D data encoding device includes a sub-mesh encoding unit that encodes sub-mesh information, a base mesh encoding unit that encodes a base mesh using the sub-mesh information, and a mesh displacement encoding unit that encodes a mesh displacement using the sub-mesh information. In the mesh displacement encoding unit, the mesh displacement is encoded using the sub-mesh information encoded by the sub-mesh encoding unit.
Advantages of the Invention
[0012] According to an aspect of the present invention, the encoding efficiency of the mesh displacement can be improved, and 3D data can be encoded and decoded with high quality.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Embodiments for Carrying Out the Invention
[0014] Hereinafter, embodiments of the present invention will be described with reference to the drawings.
[0015] FIG. 1 is a schematic diagram showing the configuration of a 3D data transmission system 1 according to this embodiment.
[0016] The 3D data transmission system 1 is a system that transmits an encoded stream obtained by encoding 3D data to be encoded, decodes the transmitted encoded stream, and displays the 3D data. The 3D data transmission system 1 includes a 3D data encoding device 11, a network 21, a 3D data decoding device 31, and a 3D data display device 41.
[0017] 3D data T is input to the 3D data encoding device 11.
[0018] The network 21 transmits the encoded stream Te generated by the 3D data encoding device 11 to the 3D data decoding device 31. The network 21 may be the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network 21 is not necessarily limited to a two-way communication network, and may be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. Further, the network 21 may be replaced by a storage medium that records the encoded stream Te such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blu-ray Disc: registered trademark).
[0019] The 3D data decoding device 31 decodes each of the encoded streams Te transmitted by the network 21 and generates one or more decoded 3D data Td that have been decoded.
[0020] The 3D data display device 41 displays all or part of the one or more decoded 3D data Td generated by the 3D data decoding device 31. The 3D data display device 41 includes, for example, a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. Examples of the form of the display include a stationary type, a mobile type, and an HMD. Further, when the 3D data decoding device 31 has a high processing capacity, it displays an image with high image quality, and when it has only a lower processing capacity, it displays an image that does not require a high processing capacity or display capacity.
[0021] <Operator> The operators used in this specification are described below.
[0022] >> is a right bit shift, << is a left bit shift, & is a bitwise AND, and | is a bitwise OR , |= is an OR assignment operator, and || indicates a logical OR.
[0023] x?y:z is a ternary operator that takes y when x is true (non-zero) and z when x is false (0). y..z indicates a set of integers from y to z.
[0024] <Structure of the Encoded Stream Te> Prior to the detailed description of the 3D data encoding device 11 and the 3D data decoding device 31 according to this embodiment, it is generated by the 3D data encoding device 11 and decoded by the 3D data decoding device 31 The data structure of the encoded stream Te will be described.
[0025] FIG. 2 is a diagram showing the hierarchical structure of the data in the encoded stream Te. The encoded stream - Te has either the data structure of a V3C sample stream or a V3C unit stream. The V3C sample stream includes a sample stream header and V3C units. The V3C unit stream includes V3C units.
[0026] A V3C unit includes a V3C unit header and a V3C unit payload. The V3C unit header is the Unit Type which is an ID indicating the type of the V3C unit, and takes values indicated by labels such as V3C_VPS, V3C_AD, V3C_AVD, V3C_GVD, V3C_OVD and so on.
[0027] When the Unit Type is V3C_VPS (Video Parameter Set), the V3C unit includes a V3C parameter set.
[0028] When the Unit Type is V3C_AD (Atlas Data, Atlas data), the V3C unit includes a VPS ID, an atlasID, a sample stream nal header, and multiple NAL units. The atlasID is an ID (Identification) and takes an integer value of 0 or greater. It is present and takes an integer value of 0 or greater.
[0029] The NAL unit includes a NALUnitType, a layerID, a TemporalID, and an RBSP (Raw byte sequence payload).
[0030] The NAL unit is identified by the NALUnitType and includes an ASPS (Atlas Sequence Parameter Set), an AAPS (Atlas Adaptation Parameter Set), an ATL (Atlas Tile layer), an SEI (Supplemental Enhancement Information), etc.
[0031] The ATL includes an ATL header and an ATL data unit, and the ATL data unit includes information such as the position and size of a patch, such as patch information data.
[0032] The SEI includes a payloadType indicating the type of SEI, a payloadSize indicating the size (number of bytes) of the SEI, and a sei_payload of the SEI data.
[0033] When the Unit Type is V3C_AVD (Attribute Video Data, attribute data), the V3C The unit includes a VPS ID, an atlasID, an attrIdx which is the ID of the attribute image, a partIdx which is the partition ID, a mapIdx which is the map ID, a flag auxFlag indicating whether it is Auxiliary data, and a video stream. The video stream is data encoded with HEVC, VVC, etc. The attribute data corresponds to the texture image in V-DMC.
[0034] When the NalUnitType is V3C_GVD (Geometory Video Data, geometry data), the V3C unit includes a VPS ID, an atlasID, a mapIdx, an auxFlag, and a video stream. The geometry data corresponds to the mesh displacement in V-DMC.
[0035] When the Unit Type is V3C_OVD (Occupancy Video Data, occupancy data), the V3C unit includes a VPS ID, an atlasID, and a video stream.
[0036] When the Unit Type is V3C_MD (Mesh data, mesh data), the V3C unit includes a VPS ID, an atlasID, and a mesh_payload. It corresponds to the base mesh in V-DMC.
[0037] (Configuration of the 3D data decoding device according to the first embodiment) FIG. 3 is a functional block diagram showing the schematic configuration of the 3D data decoding device 31 according to the first embodiment. The 3D data decoding device 31 includes a demultiplexing unit 301, a sub-mesh decoding unit 309, an atlas information decoding unit 302, a base mesh decoding unit 303, a mesh displacement decoding unit 305, a mesh reconstruction unit 307, an attribute decoding unit 306, and a color space conversion unit 308. The 3D data decoding device 31 inputs the encoded data of the 3D data and outputs the atlas information, the mesh, and the attribute image.
[0038] The inverse multiplexing unit 301 inputs the encoded data multiplexed in a byte stream format, ISOBMFF (ISO Base Media File Format), etc., and performs inverse multiplexing to obtain an atlas information encoded stream (V3C_AD's Atlas Data stream, NAL unit), a base mesh encoded stream (V3C_MD's mesh_payload), a mesh displacement encoded stream (V3C_GVD's video stream), and an attribute video stream (V3C_AVD's video stream) and outputs them.
[0039] The sub-mesh decoding unit 309 inputs the atlas information sub-mesh encoded stream output from the inverse multiplexing unit 301 and decodes the sub-mesh information.
[0040] The atlas information decoding unit 302 inputs the atlas information encoded stream output from the sub-mesh decoding unit 309 and decodes the atlas information.
[0041] The atlas information decoding unit 302 in FIG. 3 decodes coordinate system conversion information displacementCoordinateSystem (asps_vdmc_ext_displacement_coordinate_system, afps_vdmc_ext_displacement_coordinate_system) indicating the coordinate system from the encoded data. Additionally, a gating flag may be provided separately, and each coordinate system conversion information may be decoded only when the gating flag is 1. The gating flag is, for example, afve_displacement_coordinate_system_enable_flag.
[0042] The base mesh decoding unit 303 decodes the base mesh encoded stream encoded by vertex encoding (3D data compression encoding method, e.g., Draco) and outputs the base mesh. The base mesh will be described later.
[0043] The mesh displacement decoding unit 305 decodes the mesh displacement coded stream and outputs the mesh displacement. Output.
[0044] The mesh reconstruction unit 307 inputs the base mesh and the mesh displacement and reconstructs the mesh in the 3D space. Reconstruct.
[0045] The attribute decoding unit 306 decodes the attribute video stream encoded by VVC, HEVC, etc. and outputs the attribute image. The attribute image may be a texture image (texture mapping image converted by the UV atlas method) expanded on the UV axis and in the YCbCr format. The type of codec used for encoding is indicated by the ptl_profile_codec_group_idc obtained by decoding the V3C parameter set of the encoded data. It may also be indicated by the Four CC code indicated by ai_geometry_codec_id[atlasID] of the V3C parameter set. ai_geometry_codec_id[atlasID] indicates the index corresponding to the codec ID of the decoder used for decoding the attribute video stream at the atlas ID. format. The type of codec used for encoding is indicated by the ptl_profile_codec_group_idc obtained by decoding the V3C parameter set of the encoded data. It may also be indicated by the Four CC code indicated by ai_geometry_codec_id[atlasID] of the V3C parameter set. ai_geometry_codec_id[atlasID] indicates the index corresponding to the codec ID of the decoder used for decoding the attribute video stream at the atlas ID.
[0046] The color space conversion unit 308 performs color space conversion on the attribute image from the YCbCr format to the RGB format. Note that a configuration may be adopted in which the attribute video stream encoded in the RGB format is decoded and the color space conversion is omitted. format to the RGB format. Note that a configuration may be adopted in which the attribute video stream encoded in the RGB format is decoded and the color space conversion is omitted. eam is decoded and the color space conversion is omitted.
[0047] (Decoding of the base mesh) FIG. 4 is a functional block diagram showing the configuration of the base mesh decoding unit 303. The base mesh decoding unit 303 includes a mesh decoding unit 3031, a motion information decoding unit 3032, a mesh motion compensation unit 3033, and a reference The base mesh decoding unit 303 is configured with a reference mesh memory 3034, a switch 3035, a switch 3036, and a skip decoding unit 3037. The base mesh decoding unit 303 performs a base mesh decoding (not shown) before outputting the base mesh. The switches 3035 and 3036 may be configured to include a mesh dequantization unit. If the base mesh to be decoded has been coded (intra-coded) without reference to other base meshes (for example, base meshes that have already been coded and decoded), the switches 3035 and 3036 are connected to the mesh decoding unit 3031. If the base mesh to be decoded has been coded (inter-coded) with reference to other base meshes, the switches are connected to the side that performs motion compensation. When motion compensation is performed, the target vertex coordinates are derived by referencing already decoded vertex coordinates and motion information. If the base mesh to be decoded has been skipped and another base mesh has been coded (skip coded), the switches are connected to the skip decoding unit 3037.
[0048] Each base mesh is made up of one or more sub-meshes. If a mesh exists, the tile header in the atlas data sub-bitstream needs the ID to find the sub-mesh corresponding to the tile. A mesh is a subset of a mesh defined by specifying a part of the original model, and is a mesh created by dividing a mesh into multiple parts. By dividing the mesh into subsets, specific areas of the mesh can be defined individually. Each sub-mesh has its own vertex coordinates, normal vectors, texture coordinates, etc., and can be manipulated and edited individually. The mesh of a frame is called a mesh frame.
[0049] The mesh decoding unit 3031 decodes the intra-coded base mesh coded stream and outputs the base mesh (base mesh vertex positions, base mesh vertex position vectors). As the coding method, Draco, Edge Breaker, etc. are used.
[0050] The motion information decoding unit 3032 decodes the inter-coded base mesh coded stream and outputs motion information (mesh motion information, mesh motion vector) for each vertex of the reference mesh described later. As the coding method, entropy coding such as arithmetic coding is used.
[0051] The mesh motion compensation unit 3033 performs motion compensation on each vertex of the reference mesh input from the reference mesh memory 3034 based on the motion information and outputs the motion-compensated mesh.
[0052] The reference mesh memory 3034 is a memory that holds the decoded mesh for reference in subsequent decoding processes.
[0053] (Decoding of Mesh Displacement) FIG. 5 is a functional block diagram showing the configuration of the mesh displacement decoding unit 305. The mesh displacement decoding unit 305 is composed of a CABAC decoding unit (arithmetic decoding unit 3051, quantization unit 3052, context selection unit 3056, context initialization unit 3057), an inverse quantization unit 3053, an inverse transformation unit 3054, and a coordinate system transformation unit 3055.
[0054] (Context-Adaptive Binary Arithmetic Coding) The arithmetic decoding unit 3051, quantization unit 3052, context selection unit 3056, and context initialization unit 3057 use a context called Context-Adaptive Binary Arithmetic Coding (CABAC). CABAC uses a decoding method that was previously used. In CABAC, a binary string consisting of 0 and 1 is coded and decoded bit by bit using a state variable called a context (CABAC state). All CABAC states are initialized at the beginning of the segment. The CABAC decoding unit decodes each bit of the binary string (Bin String) corresponding to the syntax element. When a context is used, a context index ctxInc is derived for each bit of the syntax element, the bit is decoded using the context, and the CABAC state of the context is updated. Bits that do not use a context are coded with equal probability. The context is decoded at the rate (EP, bypass), and the index ctxIdx specifying the context and the update of the specified context are omitted. The context is the probability (state) of CABAC. It is a variable (memory area) to hold the data, and is identified by the value of ctxIdx (0, 1, 2, …). Also, the case where 0 and 1 are always equally probable, that is, 0.5, 0.5, is called EP (Equal Probability) or is called bypass. In this case, there is no need to maintain a state for a specific syntax element, so no context is used. Also, the probability is fixed at 0.5, so it is a static state that does not need to be updated. A context may be used. In this sense, it may be called static rather than bypass. An integer value such as 128 may be used to indicate a probability of 0.5.
[0055] The process of decoding one bit without using the context (bypassing) is as follows: A code may also be used. rangeTimesProb = IvlRange >> 1 binVal = ( rangeTimesProb <= ( IvlCode - IvlLow ) ) if (binVal == 0) IvlRange = rangeTimesProb else { IvlLow += rangeTimesProb IvlRange -= rangeTimesProb } In addition, the process of decoding 1 bit using the context may use the following pseudo-code . Here, prob0 is a variable indicating the probability of the context. rangeTimesProb = IvlRange * prob0 >> 16 binVal = ( rangeTimesProb <= ( IvlCode - IvlLow ) ) if (binVal == 0) IvlRange = rangeTimesProb else { IvlLow += rangeTimesProb IvlRange -= rangeTimesProb } (Coordinate system) The coordinate systems for mesh displacement (3D vector) use the following two types of coordinate systems. Cartesian coordinate system (canonical): An orthogonal coordinate system commonly defined throughout 3D space. (X, Y, Z) coordinate system. An orthogonal coordinate system whose direction does not change at the same time (within the same frame, within the same tile). Local coordinate system (local): An orthogonal coordinate system defined for each region or each vertex in 3D space . An orthogonal coordinate system whose direction can change at the same time (within the same frame, within the same tile). A coordinate system with axes of normal (D), tangent (U), and bi-tangent (V). That is, at a certain vertex (a surface including a certain vertex) the first axis (D) indicated by the normal vector n_vec, and the second axis (U) and the third axis (V) indicated by two tangent vectors t_vec, b_vec orthogonal to the normal vector n_vec which consists of an orthogonal coordinate system. n_vec, t_vec, b_vec are 3D vectors. The (D, U, V) coordinate system may also be referred to as the (n, t, b) coordinate system.
[0056] (Decoding and Derivation of Sequence-Level Control Parameters) Here, the control parameters of the sequence level decoded from the encoded data by the mesh displacement decoding unit 305 will be described.
[0057] Fig. 7 is an example of the syntax of an ASPS (Atlas Sequence Parameter Set), which is a parameter set at the sequence level. ASPS is one of the NAL units of atlas information and includes syntax elements applied to the atlas information encoding stream. The semantics of each field are as follows.
[0058] asve_subdivision_iteration_count: Indicates the number of mesh subdivision iterations.
[0059] asve_displacement_coordinate_system: Coordinate system conversion information indicating the coordinate system of the mesh displacement. When the value is equal to a predetermined first value (e.g., 0), it indicates a Cartesian coordinate system. When the value is equal to another second value (e.g., 1), it indicates a local coordinate system.
[0060] asve_1d_displacement_flag: A flag indicating whether the mesh displacement is one-dimensional. When the value is true, it indicates that the mesh displacement is one-dimensional. When the value is false, it indicates that the mesh displacement is three-dimensional.
[0061] (Decoding and Derivation of Picture / Frame-Level Control Parameters) Figure 8 shows an example of the syntax of an AFPS (Atlas Frame Parameter Set), which is a picture / frame-level parameter set. AFPS is one of the NAL units of atlas information and contains syntax elements to be applied to the atlas information coding stream. The semantics of each field are as follows. AFPS includes atlas_frame_mesh_information().
[0062] afve_overriden_flag: A flag indicating whether to update the coordinate system of the mesh displacement. When this flag is equal to true, the coordinate system of the mesh displacement is updated based on the value of afve_displacement_coordinate_system described below. When this flag is equal to false, the coordinate system of the mesh displacement is not updated.
[0063] afve_subdivision_iteration_count: Indicates the number of mesh subdivision iterations.
[0064] afve_displacement_coordinate_system: Coordinate system conversion information indicating the coordinate system of the mesh displacement . When the value is equal to the first value (e.g., 0), it indicates a Cartesian coordinate system. When the value is equal to the second value (e.g., 1), it indicates a local coordinate system. If the syntax element does not appear, the value is estimated as the value decoded by ASPS, and the default coordinate system is the coordinate system indicated by ASPS.
[0065] (Decoding and Derivation of Mesh-Level Control Parameters) Figures 21, 22, 23, and 24 are examples of the syntax structure of atlas_frame_mesh_information() that transmits base mesh and sub-mesh information of mesh displacement via AFPS. In the example of the syntax structure in Figure 21, regardless of the number of sub-meshes referenced, the number of sub-mesh IDs is encoded and decoded. atlas_frame_mesh_information() may contain any of the following syntax elements. The semantics of each field are as follows.
[0066] afmi_use_single_mesh_flag: In each atlas frame referring to AFPS, a flag indicating whether there is only one sub-mesh referenced by the mesh patch. If the value is true, it indicates that there is only one sub-mesh referenced. If the value is false, it indicates that there may be multiple sub-meshes referenced.
[0067] afmi_submesh_alignment_flag: A flag indicating whether the sub-mesh of the base mesh and the sub-mesh of the mesh displacement correspond. If the value is true, it indicates that the sub-mesh of the base mesh and the sub-mesh of the mesh displacement correspond. If the value is false, it indicates that the sub-mesh of the base mesh and the sub-mesh of the mesh displacement do not correspond. Indicates that the sub - meshes of the [[unknown character, perhaps "シュ"]] displacement may not correspond. Here, the correspondence between the sub - meshes of the base mesh and the sub - meshes of the mesh displacement means that the vertices of the base mesh with the same ID and the corresponding vertices of the mesh displacement exist in the same region. From the sub - meshes of the base mesh and the sub - meshes of the mesh displacement, a set of meshes (sub - meshes) can be decoded. Also, the correspondence between the sub - meshes of the base mesh and the sub - meshes of the mesh displacement means that the number of sub - meshes is equal. Therefore, when defining the sub - mesh information of the mesh displacement in AFPS, the value of afmi_num_submesh_minus1 may be used instead of afmi_num_displ_submesh_minus1 without encoding and decoding afmi_num_displ_submesh_minus1. Also, the vertices of the base mesh with a certain sub - mesh ID are referenced by the mesh displacement with the same sub - mesh ID. Or, the mesh displacement with a certain sub - mesh ID may decode the vertices by referring only to the vertices of the base mesh with the same sub - mesh ID. Also, there may be a condition of the bit - stream that the mesh displacement with a certain sub - mesh ID refers only to the vertices of the base mesh with the same sub - mesh ID. The mesh displacement with a certain sub - mesh ID refers to it. Or, the mesh displacement with a certain sub - mesh ID may decode the vertices by referring only to the vertices of the base mesh with the same sub - mesh ID. Also, there may be a condition of the bit - stream that the mesh displacement with a certain sub - mesh ID refers only to the vertices of the base mesh with the same sub - mesh ID.
[0068] afmi_num_submeshes_minus1: A parameter indicating the number of sub - meshes referred to by the mesh patch. The number of sub - meshes is afmi_num_submeshes_minus1 + 1. A parameter indicating the number of sub - meshes referred to by the mesh patch. The number of sub - meshes is afmi_num_submeshes_minus1 + 1.
[0069] afmi_num_displ_submesh_minus1: A parameter indicating the number of displacement sub - meshes referred to by the mesh patch. The number of sub - meshes is afmi_num_displ_submesh_minus1 + 1. A parameter indicating the number of displacement sub - meshes referred to by the mesh patch. The number of sub - meshes is afmi_num_displ_submesh_minus1 + 1. If it does not exist, the value of afmi_num_displ_submesh_minus1 is presumed to be equal to afmi_num_submeshes_minus1.
[0070] afmi_signalled_submesh_id_flag: A flag indicating whether the submesh ID referenced by the mesh patch is signaled. If the value is true, it indicates that the submesh ID is signaled. If the value is false, it indicates that the submesh ID is not signaled. Indicates.
[0071] afmi_signalled_submesh_id_length_minus1: In the patch data unit with index patchIdx, when there exists a syntax element afmi_submesh_id[i] within the atlas tile where the tile ID is equal to tileID, it is a parameter indicating the number of bits of the syntax element pdu_submesh_id / bmpdu_submesh_id[tileID][patchIdx]. The value of afmi_signalled_submesh_id_length_minus1 must be within the range of 0 to 15. If it does not exist, its value is assumed to be equal to the following formula. Ceil( Log2( afmi_num_submeshes_minus1 + 1 ) ) Is inferred to be equal to the following formula.
[0072] Ceil( Log2( afmi_num_submeshes_minus1 + 1 ) ) afmi_submesh_id[i]: It is a parameter indicating the submesh ID of the i-th submesh. When the value of afmi_signalled_submesh_id_flag is true, the length of the syntax element of afmi_submesh_id[i] is afmi_signalled_submesh_id_length_minus1 bits. When the value of afmi_signalled_submesh_id_flag is true, the length of the syntax element of afmi_submesh_id[i] is afmi_signalled_submesh_id_length_minus1 bits.
[0073] afmi_signalled_displ_submesh_id_flag: A flag indicating whether the submesh ID referenced by the mesh displacement patch is signaled. If the value is true, it indicates that the submesh ID is signaled. If the value is false, it indicates that the submesh ID is signaled. Indicates non-existence. When the value of afmi_use_single_mesh_flag is true, the value of afmi_signalled_displ_submesh_id_flag may always be set to false to reduce the overhead of the signed quantity.
[0074] afmi_displ_submesh_id: A parameter indicating the submesh ID of the i-th displacement submesh. The number of bits of the length of the syntax element of afmi_displ_submesh_id is derived by the following formula as follows.
[0075] Ceil(Log2(afmi_num_displ_submeshes_minus1 + 1)) bits The atlas information decoding unit 302 (submesh decoding unit 309) decodes the frame mesh information from the encoded data of the AFPS of the atlas information. For example, it decodes afmi_use_single_mesh_flag, afmi_num_submeshes_minus1, afmi_num_displ_submeshes_minus1, afmi_signalled_submesh_id_flag, afmi_signalled_displ_submesh_id_flag, afmi_signalled_submesh_id_length_minus1, afmi_signalled_displ_submesh_id_length_minus1, afmi_submesh_id, afmi_displ_submesh_id. Further, afmi_submesh_alignment_flag may be decoded. Also, when afmi_submesh_alignment_flag is false, it may be configured to decode and encode the submesh information afmi_signalled_displ_submesh_id_length_minus1, afmi_submesh_id, afmi_displ_submesh_id including the mesh displacement. That is, afmi_submesh_alignment_ When the flag is true, it may be configured not to include the sub-mesh information afmi_signalled_displ_submesh_id_length_minus1, afmi_submesh_id, and afmi_displ_submesh_id of the mesh displacement and not to perform decoding and encoding. The atlas information encoding unit 101 (sub-mesh encoding unit 116) encodes the frame mesh information into the encoded data of the AFPS of the atlas information.
[0076] When the afmi_signalled_submesh_id_flag is true, the sub-mesh decoder 309 decodes the afmi_submesh_id[ i ] in the range of i = 0.. afmi_num_submeshes_minus1 by the number of sub-meshes (afmi_num_submeshes_minus1 + 1), and derives the arrays SubMeshIDToIndex and SubMeshIndextoID for i = 0.. afmi_num_submeshes_minus1 as follows.
[0077] SubMeshIDToIndex[ afmi_submesh_id[ i ] ] = i SubMeshIndextoID[ i ] = afmi_submesh_id[ i ] Note that as shown in FIG. 25, in the syntax configurations of FIGS. 21, 22, 23, and 24, the afmi_signalled_submesh_id_flag, if (afmi_signalled_submesh_id_flag), and the subsequent {} parts may be made to exist only when if (!afmi_use_single_mesh_flag). Or when the number of sub-meshes (NumSubmeshes = afmi_num_submeshes_minus1 + 1) is greater than 1 the above parts may be made to exist.
[0078] In this case, when afmi_use_single_mesh_flag is false (or when the number of sub - meshes (NumSubmeshes = afmi_num_submeshes_minus1 + 1) is greater than 1) and afmi_signalled_submesh_id_flag is true, the sub - mesh decoder 309 may decode afmi_submesh_id[i] for i = 0..afmi_num_submeshes_minus1 by afmi_signalled_submesh_id_length_minus1 and the number of sub - meshes - 1 (afmi_num_submeshes_minus1). Also, when afmi_use_single_mesh_flag is true (or when the number of sub - meshes (NumSubmeshes = afmi_num_submeshes_minus1 + 1) is 1), SubMeshIDToIndex[0] = 0, SubMeshIndexToID[0] = 0, and the ID may always be 0. or when the number of sub - meshes (NumSubmeshes = afmi_num_submeshes_minus1 + 1) is greater than 1 and afmi_signalled_submesh_id_flag is true, the sub - mesh decoder 309 may decode afmi_submesh_id[i] for i = 0..afmi_num_submeshes_minus1 by afmi_signalled_submesh_id_length_minus1 and the number of sub - meshes - 1 (afmi_num_submeshes_minus1). Also, when afmi_use_single_mesh_flag is true (or when the number of sub - meshes (NumSubmeshes = afmi_num_submeshes_minus1 + 1) is 1), SubMeshIDToIndex[0] = 0, SubMeshIndexToID[0] = 0, and the ID may always be 0.
[0079] Also, in the syntax configurations of FIGS. 21, 22, 23, and 24 instead of the example of the syntax configuration of FIG. 25, when the value of afmi_use_single_mesh_flag is true, the value of afmi_signalled_submesh_id_flag may always be false, SubMeshIDToIndex[0] = 0, SubMeshIndexToID[0] = 0, and the ID may always be 0.
[0080] When afmi_signalled_submesh_id_flag is true, the sub - mesh decoder 309 does not decode afmi_submesh_id[i], and for i = 0..afmi_num_submeshes_minus1, the following array is derived as follows. is derived.
[0081] SubMeshIDToIndex[i] = i SubMeshIndextoID[ i ] = i When afmi_submesh_alignment_flag is true, the sub-mesh decoding unit 309 may derive an array as follows for i = 0..afmi_num_displ_submeshes_minus1.
[0082] DisplSubMeshIDToIndex[ i ] = SubMeshIDToIndex[ i ] DisplSubMeshIndextoID[ i ] = SubMeshIndextoID[ i ] Also, when afmi_submesh_alignment_flag is true, instead of the arrays DisplSubMeshIDToIndex and DisplSubMeshIndextoID of mesh displacements, SubMeshIDToIndex and SubMeshIndextoID common to the base mesh and mesh displacements may be used.
[0083] When afmi_signalled_displ_submesh_id_flag is true, the sub-mesh decoding unit 309 decodes afmi_displ_submesh_id[ i ] in the range of i = 0.. afmi_num_displ_submeshes_minus1 for the number of sub-meshes (afmi_num_displ_submeshes_minus1 + 1), and derives the arrays DisplSubMeshIDTo Index and DisplSubMeshIndextoID as follows for i = 0..afmi_num_displ_submeshes_minus1.
[0084] DisplSubMeshIDToIndex[ afmi_displ_submesh_id[ i ] ] = i DisplSubMeshIndextoID[ i ] = afmi_displ_submesh_id[ i ] When afmi_signalled_displ_submesh_id_flag is false, the sub-mesh decoding unit 309 does not decode afmi_displ_submesh_id[ i ] for i = 0.. afmi_num_displ_submeshes_minus1 and derives the array as follows. as follows
[0085] DisplSubMeshIDToIndex[ i ] = i DisplSubMeshIndextoID[ i ] = i In this configuration, it is possible to decode and display in units of sub-meshes even in the case of mesh displacement. In the above configuration, when afmi_submesh_alignment_flag is false, the sub-mesh decoding unit 309 decodes the sub-mesh information (afmi_signalled_displ_submesh_id_flag, afmi_signalled_displ_submesh_id_length_minus1, afmi_displ_submesh_id[ i ]) regarding the mesh displacement. When afmi_submesh_alignment_flag is true, it does not decode the sub-mesh information regarding the mesh displacement. information (afmi_signalled_displ_submesh_id_flag, afmi_signalled_displ_submesh_id_length_minus1, afmi_displ_submesh_id[ i ]) meshes, and when afmi_submesh_alignment_flag is true, it does not decode the sub-mesh information regarding the mesh displacement. meshes.
[0086] Also, in a configuration where afmi_submesh_alignment_flag is decoded and encoded, since it is possible to determine whether the sub-meshes of the base mesh and the mesh displacement are the same or different by looking at this flag, it is possible to easily decode and display in units of sub-meshes. Also, since the encoding of the array can be omitted, the amount of code can be reduced. Also, in a configuration where afmi_submesh_alignment_flag is decoded and encoded, since it is possible to determine whether the sub-meshes of the base mesh and the mesh displacement are the same or different by looking at this flag, it is possible to easily decode and display in units of sub-meshes. Also, since the encoding of the array can be omitted, the amount of code can be reduced.
[0087] (Another configuration) In the configuration shown in FIG. 22, when the value of afmi_submesh_alignment_flag is false, the sub-mesh decoder 309 further decodes afmi_num_displ_submeshes_minus1 as information on the sub-mesh regarding the mesh displacement. When the value of afmi_submesh_alignment_flag is true, the sub-mesh decoder 309 does not decode afmi_num_displ_submeshes_minus1 as information on the sub-mesh regarding the mesh displacement. None.
[0088] In this configuration, when the value of afmi_submesh_alignment_flag is true, since the sub-mesh decoder 309 does not decode afmi_num_displ_submeshes_minus1 as information on the sub-mesh regarding the mesh displacement, it has the effect of reducing the overhead of the code amount. None.
[0089] In the configuration shown in FIG. 23, when the value of afmi_use_single_flag is false, the sub-mesh decoder 309 decodes afmi_submesh_alignment_flag. When the value of afmi_use_single_flag is true, the sub-mesh decoder 309 does not decode afmi_submesh_alignment_flag.
[0090] In this configuration, when the value of afmi_use_single_flag is true, since the sub-mesh decoder 309 does not decode afmi_submesh_alignment_flag, it has the effect of reducing the overhead of the code amount.
[0091] Alternatively, the sub-mesh decoder 309 may always derive the following arrangement without decoding and encoding afmi_submesh_alignment_flag, not in the syntax structure examples of FIGS. 21, 22, 23, and 24. For example, it may always derive the following arrangement without decoding and encoding afmi_submesh_alignment_flag. None.
[0092] DisplSubMeshIDToIndex[ i ] = SubMeshIDToIndex[ i ] DisplSubMeshIndextoID[ i ] = SubMeshIndextoID[ i ] That is, the configuration may be such that the sub-meshes of the base mesh and the mesh displacement are always corresponding. In this case, the sub-mesh information decoded by the base mesh is also used for the decoding of the mesh displacement.
[0093] In this configuration, there is no degree of freedom to make the sub-meshes of the base mesh and the mesh displacement different. However, since it is known in advance on the decoding side that the mesh division is always the same, the effect is that the decoding and display in units of sub-meshes are facilitated.
[0094] In the configuration shown in FIG. 25, the sub-mesh decoding unit 309 decodes and encodes the sub-mesh information of the mesh displacement that does not always depend on the sub-meshes of the base mesh without decoding and encoding the afmi_submesh_alignment_flag.
[0095] In this configuration, since there is no dependency on the sub-meshes of the base mesh and the mesh displacement, mesh reconstruction cannot be performed in units of sub-meshes. However, since the base mesh and the mesh displacement can always be decoded and encoded independently, the effect is that it becomes easy to decode and encode each sub-mesh in parallel.
[0096] Also, when the sub-mesh decoding unit 309 decodes and encodes the afmi_use_single_mesh_flag instead of the syntax structures of FIGS. 21, 22, 23, and 24, the number of sub-meshes is not, The syntax elements afmi_num_submeshes_minus2 and afmi_num_displ_submeshes_minus2, which indicate nas2 (the value obtained by subtracting 2 from the number of sub-meshes), may be decoded and encoded. Alternatively, the syntax elements afmi_num_submeshes_minus2 and afmi_num_displ_submeshes_minus2, which indicate the number of sub-meshes minus 2 of the referenced sub-meshes, may be decoded and encoded only when the value of afmi_use_single_mesh_flag is false. The semantics may use the following examples.
[0097] afmi_num_submeshes_minus2: A parameter indicating the number of sub-meshes referenced by the mesh patch. The number of sub-meshes is afmi_num_submeshes_minus2 + 2. It is a parameter that indicates. The number of sub-meshes is afmi_num_submeshes_minus2 + 2.
[0098] afmi_num_displ_submeshes_minus2: A parameter indicating the number of displacement sub-meshes referenced by the mesh patch. The number of sub-meshes is afmi_num_displ_submeshes_minus2 + 2. If afmi_num_displ_submesh_minus2 does not exist, the value of afmi_num_displ_submesh_minus2 is presumed to be equal to afmi_num_submeshes_minus2. It is a parameter that indicates the number of sub-meshes. The number of sub-meshes is afmi_num_displ_submeshes_minus2 + 2. If afmi_num_displ_submesh_minus2 does not exist, the value of afmi_num_displ_submesh_minus2 is presumed to be equal to afmi_num_submeshes_minus2.
[0099] In this configuration, since the case where there is one referenced sub-mesh can be represented by afmi_use_single_mesh_flag, decoding and encoding the syntax element indicating the number of sub-meshes minus 2 has the effect of reducing the overhead of the coded amount. By doing so, it has the effect of reducing the overhead of the coded amount.
[0100] Also, the sub-mesh decoding unit 309 is in the examples of the syntax structures in FIGS. 21, 22, 23, and 24. Instead, when the value of afmi_submesh_alignment_flag is false, or when always decoding and encoding sub-mesh information of mesh displacement that does not depend on the sub-mesh information of the base mesh, the sub-mesh information of the mesh displacement may be decoded and encoded using the example of the syntax structure of mesh displacement in FIG. 17 described later, rather than atlas_frame_mesh_information().
[0101] (Syntax structure of mesh displacement) In the 3D data encoding method in Non-Patent Document 1, mesh displacement (displacement data) was encoded and decoded at the frame level, but there was a problem that encoding and decoding could not be performed using sub-mesh information indicating the unit for dividing a frame into a plurality of meshes. That is, the base mesh encoding unit 103 and the base me ssh decoding unit 303 that perform encoding and decoding independently for each sub-mesh and the mesh displacement encoding unit 107 and the mesh displacement decoding unit 305 that perform encoding and decoding for each frame can only process mesh reconstruction at the frame level, and there was a problem that it could not be processed at the sub-mesh level level.
[0102] As shown in the syntax structure described later, in this example, it has the NAL unit type of FIG. 20. That is, mesh displacement is decoded at the sub-mesh level. That is, mesh displacement is decoded at the sub-mesh level.
[0103] FIG. 15 is an example of the syntax of a configuration for transmitting mesh displacement parameters by a sequence-level DSPS. DSPS (Displacement Sequence Parameter Set) is one of the NAL units of mesh displacement and includes syntax elements applied to the mesh displacement encoding stream. The semantics of each field are as follows. That is, mesh displacement is decoded at the sub-mesh level.
[0104] dsps_sequence_parameter_set_id: Indicates the identifier of the mesh displacement sequence parameter set for reference by other syntax elements.
[0105] dsps_single_dimension_flag: A flag indicating whether the mesh displacement is one-dimensional. . If the value is true, it indicates that the mesh displacement is one-dimensional. If the value is false, it indicates that the mesh displacement is three-dimensional.
[0106] dsps_lod_count: Indicates the number of levels of detail (LoD) of the mesh displacement. Note that dsps_lod_count_minus1 (the number of LoD minus 1) may be encoded and decoded. In that case, dsps_lod_count = dsps_lod_count_minus1 + 1 is used.
[0107] Figure 16 is an example of the syntax of a configuration transmitted in a DFPS (Displacement Frame Parameter Set), which is a picture / frame-level parameter set of mesh displacement parameters. . DFPS is one of the NAL units of mesh displacement and contains syntax elements applied to the mesh displacement coding stream. The semantics of each field are as follows.
[0108] dfps_displ_sequence_parameter_set_id: Indicates the value of dsps_sequence_parameter_set_id of the active mesh displacement sequence parameter set.
[0109] dfps_displ_frame_parameter_set_id: Indicates the identifier of the mesh displacement frame parameter set for reference by other syntax elements.
[0110] The mesh displacement decoding unit 305 and the mesh displacement encoding unit 107 decode the dfps_displ_sequence_parameter_set_id and dfps_output_flag_present_flag from the encoded data of the DFPS, and encode the encoded data of the DFPS.
[0111] Figure 17 is an example of the syntax structure of displ_sub_mesh_information() transmitted by the DFPS. The semantics of each field are as follows. displ_sub_mesh_information() is sub-mesh information indicating a unit that divides a frame into a plurality of meshes. The sub-mesh information may include the number of sub-meshes, the ID of the sub-mesh, and the code length of the sub-mesh ID.
[0112] dsi_use_single_mesh_flag: When dsi_use_single_mesh_flag is equal to 1, it indicates that only one sub-mesh exists in each mesh frame referring to the DFPS. When dsi_use_single_mesh_flag is equal to 0, it indicates that there may be a plurality of sub-meshes in each mesh frame referring to the DFPS. dsi_num_submeshes_minus1: It indicates the number of sub-meshes - 1 in each mesh frame referring to the DFPS. When dsi_use_single_mesh_flag is equal to 1, it is presumed that the number of dsi_num_submeshes_minus1 + 1
[0113] is equal to 1. is equal to 1.
[0114] dsi_signalled_submesh_id_flag: When dsi_signalled_submesh_id_flag is equal to 1, Indicates that the sub-mesh ID of each mesh frame is signaled. When dsi_signalled_submesh_id_flag is equal to 0, it indicates that the sub-mesh ID is not signaled.
[0115] dsi_submesh_id: dsi_submesh_id[i] specifies the i-th sub-mesh ID. The length (number of bits) of the dsi_submesh_id[i] syntax element is derived by the following formula.
[0116] Ceil( Log2( dsi_num_displ_submeshes_minus1 + 1 ) ) bits The mesh displacement decoding unit 305 decodes the mesh displacement sub-mesh information from the encoded data of the DFPS. For example, it decodes dsi_use_single_mesh_flag, dsi_num_submeshes_minus1, dsi_signalled_submesh_id_flag, dsi_signalled_submesh_id_length_minus1, and dsi_submesh_id. The mesh displacement encoding unit 107 encodes the mesh displacement sub-mesh information into the encoded data of the DFPS.
[0117] When dsi_signalled_submesh_id_flag is true, the mesh displacement decoding unit 305 decodes dsi_submesh_id[ i ] for the range of i = 0.. dsi_num_submeshes_minus1 by the number of sub-meshes (dsi_num_submeshes_minus1 + 1), and derives the arrays DisplSubMeshIDToIndex and DisplSubMeshIndextoID for i = 0.. dsi_num_submeshes_minus1 as follows.
[0118] DisplSubMeshIDToIndex[ dsi_submesh_id[ i ] ] = i DisplSubMeshIndextoID[ i ] = dsi_submesh_id[ i ] When dsi_signalled_submesh_id_flag is false, the mesh displacement decoding unit 305 does not decode dsi_submesh_id[ i ], and derives an array as follows for i = 0.. dsi_num_submeshes_minus1.
[0119] DisplSubMeshIDToIndex[ i ] = i DisplSubMeshIndextoID[ i ] = i Also, instead of the example of the syntax structure in FIG. 17, when decoding and encoding dsi_use_single_mesh_flag, a syntax element dsi_num_displ_submeshes_minus2 indicating the number of submeshes minus 2 (the value obtained by subtracting 2 from the number of submeshes) may be decoded and encoded. Alternatively, only when the value of dsi_use_single_mesh_flag is false, the number of submeshes minus 2, the syntax element dsi_num_displ_submeshes_minus2, may be decoded and encoded. The semantics may use the following example.
[0120] dsi_num_displ_submeshes_minus2: A parameter indicating the number of submeshes in each mesh frame referring to DFPS. The number of submeshes is dsi_num_displ_submeshes_minus2 + 2. When dsi_use_single_mesh_flag is equal to 1, the number is presumed to be equal to 1.
[0121] In this configuration, the case where there is one referenced submesh is determined by dsi_use_single_mesh_flag Since it can be represented, a syntax element indicating the number of sub - meshes minus 2 is decoded, and by encoding, the effect of reducing the overhead of the code amount is achieved.
[0122] Although not shown, the mesh displacement decoding unit 305 and the mesh displacement encoding unit 107 may decode and encode a displacement layer displ_layer_rbsp( ) including a displacement header displ_header( ), displacement data displ_data_unit( displ_id ), and rbsp_trailing_bits( ) from the encoded data. The displacement data displ_data_unit( displ_id ) may have a syntax structure described later with respect to FIG. 18. The mesh displacement decoding unit 305 and the mesh displacement encoding unit 107 may decode and encode the following syntax from the displacement header displ_header( ).
[0123] The mesh displacement decoding unit 305 and the mesh displacement encoding unit 107 may decode and encode the following syntax from the displacement header displ_header( ). dh_frame_parameter_set_id: Indicates the ID of the parameter set.
[0124] dh_frame_parameter_set_id: Indicates the ID of the parameter set.
[0125] displ_submesh_id: Indicates the ID of the sub - mesh of the mesh displacement.
[0126] dh_type: Is the encoding type of the mesh displacement. It indicates either intra - encoding (I_DISPLACEMENT) or inter - encoding (P_DISPLACEMENT) type. A variable st may be used as a variable indicating the sub - mesh type. Also, st = dh_type may be used. dh_type: Is the encoding type of the mesh displacement. It indicates either intra - encoding (I_DISPLACEMENT) or inter - encoding (P_DISPLACEMENT) type. A variable st may be used as a variable indicating the sub - mesh type. Also, st = dh_type may be used.
[0127] dh_output_flag: Flag indicating whether to output.
[0128] dh_frm_order_cnt_lsb: Is the value of the least significant bit (LSB) of Picture Order Cnd (POC).
[0129] Figures 18 and 19 are examples of the syntax structure of mesh displacement. The semantics are as follows. Mesh displacement is a column of sub-mesh ID (subMeshID), position level, and values (coefficients) of the k component (k component), and is represented by the array Qdisp[subMeshID][level][k]. Displacement is a three-dimensional signal in the Cartesian coordinate system (xyz) or the local coordinate system (ntb), and each component of the three-dimensional displacement is called a component. The displacement Qdisp here is also called a coefficient because it is the value after being transformed by discrete wavelet transform, lifting transform, DCT transform, etc. Component The variable k takes values of 0, 1, and 2. The variable name is not limited to k, and it can also be dim or other variable names. The order of the indices of QDisp may be reversed, that is, Qdisp[subMeshID][k][level] may be used instead of Qdisp[subMeshID][level][k]. As in the example of the syntax structure in Figure 18(a), regardless of the encoding type of mesh displacement, the syntax structure ddu_intra_sub_mesh_unit(displSubmeshID) in Figure 18(b) may be used. Or, depending on the encoding type, ddu_intra_sub_mesh_unit(displSubmeshID) and ddu_inter_sub_mesh_unit(displSubmeshID) may be used as follows. displ_data_unit(displSubmeshID) { if( dh_type == I_DISPLACEMENT) { ddu_intra_sub_mesh_unit(displSubmeshID) } else if( df_type == P_DISPLACEMENT ) { ddu_inter_sub_mesh_unit(displSubmeshID) } } Here, ddu_inter_sub_mesh_unit(displSubmeshID) may use a syntax structure in which the mesh displacement to be decoded performs motion compensation with reference to other mesh displacements. The mesh displacement decoding unit 305 decodes the syntax shown in FIGS. 18 and 19 in units of submeshes indicated by displSubMeshID. For example, as shown below, the number of displacement coefficients, the absolute value of the displacement coefficients, the sign of the displacement coefficients, etc. may be decoded in units of displSubMeshID. Note that displSubMeshID may use the value displ_submesh_id of the displacement header displ_header(). In another form
[0130] The mesh displacement decoding unit 305 decodes the syntax shown in FIGS. 18 and 19 in units of submeshes indicated by displSubMeshID. For example, as shown below, the number of displacement coefficients, the absolute value of the displacement coefficients, the sign of the displacement coefficients, etc. may be decoded in units of displSubMeshID. Note that displSubMeshID may use the value displ_submesh_id of the displacement header displ_header(). In another form The mesh displacement decoding unit 305 decodes the syntax shown in FIGS. 18 and 19 in units of submeshes indicated by displSubMeshID. For example, as shown below, the number of displacement coefficients, the absolute value of the displacement coefficients, the sign of the displacement coefficients, etc. may be decoded in units of displSubMeshID. Note that displSubMeshID may use the value displ_submesh_id of the displacement header displ_header(). In another form as, displSubMeshID may use the value obtained by decoding the syntax element dsi_submesh_id[i] included in displ_sub_mesh_information(). as, displSubMeshID may use the value obtained by decoding the syntax element dsi_submesh_id[i] included in displ_sub_mesh_information().
[0131] displSubMeshID = dsi_submesh_id[i] Also, when decoding a plurality of submesh displacements indicated by the index i, ID (= displSubMeshID) may be derived from the index i using an array that has already been decoded or derived and used. Also, when decoding a plurality of submesh displacements indicated by the index i, ID (= displSubMeshID) may be derived from the index i using an array that has already been decoded or derived and used.
[0132] displSubMeshID = DisplSubMeshIndextoID[i] Also, for the index i, the mesh displacement decoding unit 305 and the mesh displacement encoding unit 107 may perform loop processing regarding i to decode and encode the displacement header displ_header() and the displacement data displ_data_unit( displ_id).
[0133] Alternatively, the mesh displacement decoding unit 305 and the mesh displacement encoding unit 107 may decode and encode one displacement header displ_header( ), and then perform a loop process for i to decode and encode the displacement data displ_data_unit( displ_id ). The displacement data uses the values decoded and encoded in one displ_header( ) in common.
[0134] The syntax elements in Figure 19 have the following meanings:
[0135] dismu_vertex_count_lod[displSubMeshID][i]: Coordinates included in the division (LoD) level i (variable vertCount[i] = dismu_vertex_count_lod[i], which indicates the number of vertices in block i. A number vertCount may be derived.
[0136] dismu_coeff_abs_level_gt0[displSubmeshID][k][v]: Submesh of displSubmeshID Indicates whether the absolute value of the non-zero mesh displacement coefficient of the vertex with index v in the component with index k is greater than 0. If it is greater, it is 1; otherwise, it is 0.
[0137] dismu_coeff_abs_level_gt1[displSubmeshID][k][v]: Submesh of displSubmeshID In this case, the absolute value of the non-zero mesh displacement coefficient of the vertex with index v of the component with index k is greater than 1. If it is greater, it is set to 1; otherwise, it is set to 0. If no box is present, it is assumed to be 0.
[0138] dismu_coeff_abs_level_gt2[displSubmeshID][k][v]: Submesh of displSubmeshID indicates whether the absolute value of the non-zero mesh displacement coefficient of the vertex with index v of the component with index k is greater than 2 in it. If it is greater, it is 1; otherwise, it is 0. If this syntax does not exist, it is presumed to be 0.
[0139] dismu_coeff_abs_level_gt3[displSubmeshID][k][v]: In the submesh of displSubmeshID indicates whether the absolute value of the non-zero mesh displacement coefficient of the vertex with index v of the component with index k is greater than 3. If it is greater, it is 1; otherwise, it is 0. If this syntax does not exist, it is presumed to be 0.
[0140] dismu_coeff_sign[displSubmeshID][k][v]: Indicates whether the non-zero mesh displacement coefficient of the vertex with index v of the component with index k is a positive number in the submesh of displSubmeshID. If it is a positive number, it is 1; otherwise (if it is a negative number), it is 0. If this syntax does not exist, it is presumed to be 1.
[0141] dismu_coeff_abs_level_rem[displSubmeshID][k][v]: In the submesh of displSubmeshID is the value obtained by subtracting 4 from the absolute value of the non-zero mesh displacement coefficient of the vertex with index v of the component with index k. If this syntax does not exist, it is presumed to be 0.
[0142] The mesh displacement decoding unit 305 decodes dismu_nz_subBlock for each sub-block of the mesh displacement. When dismu_nz_subBlock[displSubmeshID][k][block] is 1, the subsequent syntax elements at the component with index k and the sub-block level with index block are decoded.
[0143] The mesh displacement decoding unit 305 decodes dismu_coeff_abs_level_gt0 for each sub-block of the mesh displacement, and when dismu_coeff_abs_level_gt0 is a predetermined value (e.g., other than 0), it decodes the subsequent dismu_coeff_sign and dismu_coeff_abs_level_gt1.
[0144] The mesh displacement decoding unit 305, when dismu_coeff_abs_level_gt1 is a predetermined value (e.g., other than 0) decodes the subsequent dismu_coeff_abs_level_gt2.
[0145] The mesh displacement decoding unit 305, when dismu_coeff_abs_level_gt2 is a predetermined value (e.g., other than 0) decodes the subsequent dismu_coeff_abs_level_gt3.
[0146] The mesh displacement decoding unit 305, when dismu_coeff_abs_level_gt3 is a predetermined value (e.g., other than 0) decodes the subsequent dismu_coeff_abs_level_rem.
[0147] (Operation of the mesh displacement decoding unit) The arithmetic decoding unit 3051 decodes the mesh displacement encoded stream that has been arithmetically encoded according to a value (context) indicating a random variable, and outputs a binary signal. The binary signal may be an alpha code or a k-th order Exp-Golomb code. The Exp-Golomb code is composed of a prefix code and a suffix code. The prefix is a value that increases exponentially, and the suffix is the remainder. When encoding and decoding the variable rem using the Exp-Golomb code, the prefix and suffix of the Exp-Golomb code are also referred to as the prefix and suffix of rem.
[0148] The multi-valued conversion unit 3052 decodes the quantized mesh displacement Qdisp, which is a multi-valued signal, from a binary signal.
[0149] The context selection unit 3056 (context memory) has a memory for holding contexts, derives a context used for arithmetic decoding of the mesh displacement according to the state, and updates values as necessary. In the arithmetic decoding of each coefficient of the mesh displacement, different context arrays may be used according to the sub-mesh type dh_type (e.g., 0: intra-sub-mesh, 1: inter-sub-mesh), the level of detail lod (level of detail) of the mesh division, and the component dim of the mesh displacement vector. The context includes a variable indicating the generation probability of the binary signal. Depending on the level of detail lod (level of detail) of the mesh division and the component dim of the mesh displacement vector the following different context arrays may be used. The context includes a variable indicating the generation probability of the binary signal. is included. ctxCodedSubBlock[numST][numLOD][numDim] ctxCoeffGtN[numST][numLOD][MAX_GTN + 1][numDim] ctxCoeffRemPrefix[numST][numLOD][numDim][numPrefixBin] Note that the static context with a fixed probability without context update is ctxStatic. The decoding of the syntax element indicated by ctxStatic may be decoded without using a context. decode(ctxStatic) may be decode_bypass(), using a process dedicated to bypass.
[0150] Here, numST is the number of types of sub-mesh types, and it may be numST = 2. numPrefixBin is the number of bins using a context in the prefix, and it may be numPrefixBin = 2. numLOD is the maximum number of levels of detail of the mesh division, and it may be numLOD = 4. numDim is the number of dimensions of the mesh displacement vector, and it may be numDim = 3. The maximum value MAX_GTN of the threshold for the coefficient is 3.
[0151] ctxCodedSubBlock[numST][numLoD][numDim] is an array of contexts used for decoding the syntax element dismu_nz_subBlock. The arithmetic decoder 3051 uses the value of ctxCodedSubBlock[st][lod][dim] to decode dismu_nz_subBlock in the sub-mesh type st, detail level lod, and the dimension dim of the mesh displacement vector.
[0152] ctxCoeffGtN[numST][numLoD][MAX_GTN + 1][numDim] is an array of contexts used for decoding the syntax element dismu_coeff_abs_level_gtN (where N is replaced by 0, 1, 2, MAX_GTN). The arithmetic decoder 3051 uses the value of ctxCoeffGtN[st][lod][N][dim] to decode dismu_coeff_abs_level_gtN in the sub-mesh type st, detail level lod, and the dimension dim of the mesh displacement vector. The arithmetic decoder 3051 uses bypass to decode dismu_coeff_sign in the sub-mesh type st, detail level lod, and the
[0153] dimension dim of the mesh displacement vector. The arithmetic decoder 3051 uses bypass to decode dismu_coeff_sign in the sub-mesh type st, detail level lod, and the
[0154] ctxCoeffRemPrefix[numST][numLoD][numDim] is an array of contexts used for decoding the syntax element dismu_coeff_abs_level_rem. The arithmetic decoder 3051 uses the value of ctxCoeffRemPrefix[st][lod][dim] to decode dismu_coeff_abs_level_rem in the sub-mesh type st, detail level lod, and the mesh displacement vector in dimension dim. st may use dh_type decoded from the encoded data (the same applies hereinafter).
[0155] The context initialization unit 3057 initializes the context (the generation probability of a binary signal). The context may be initialized for each sub-mesh, or the context may be initialized for every plurality of sub-meshes. When initializing the context for each sub-mesh, since there is no context dependency between sub-meshes, random access to any sub-mesh can be easily performed. When initializing the context for every plurality of sub-meshes, since the initialization frequency is low, the coding efficiency can be improved as compared with the case of initializing for each sub-mesh.
[0156] (Derivation process of mesh displacement) The mesh displacement decoder 305 decodes the syntax elements dismu_nz_subBlock, dismu_coeff_abs_level_gt0, dismu_coeff_abs_level_gt1, dismu_coeff_abs_level_gt2, dismu_coeff_abs_level_gt3, dismu_coeff_abs_level_rem, and dismu_coeff_sign through the following process and derives the mesh displacement Qdisp. Here, regarding the sub-mesh of the base mesh decoder 303 and the sub-mesh of the mesh displacement decoder 305, constraints may be imposed so as to have a corresponding relationship using the syntax element afmi_submesh_alignment_flag decoded from atlas_frame_mesh_information() shown in FIGS. 21, 22, 23, and 24. Here, the mesh displacement decoder 305 decodes dismu_nz_subBlock in units of sub-blocks of SubBlockSize. When dismu_nz_subBlock is a predetermined value, the mesh displacement coefficients within the sub-block are decoded. The decoding may be performed in units of sub-meshes indicated by subMeshID (= displSubMeshID). st may use dh_type decoded from the encoded data. It may be performed. for (k = 0; k < numDim; k++) { / / dimension (component) loop for (b = 0; b<numLOD; b++) { / / Level of Detail loop, block loop numSubBlocks = dispCount[b] / subBlockSize + 1 for (s = 0; s < numSubBlocks; s++) { / / subblock loop / / decode dismu_nz_subBlock dismu_nz_subBlock [k][b][s] = decode(ctxCodedSubBlock[st][b][k]) if (dismu_nz_subBlock [k][b][s]) { for (v = 0; v < subBlockSize; v++) { / / coefficient loop within subblock value = 0 / / decode dismu_coeff_abs_level_gt0 dismu_coeff_abs_level_gt0[k][b][s][v] = decode(ctxCoeffGtN[st][b][0][k]) if (dismu_coeff_abs_level_gt0[k][b][s][v]) { value++ / / decode dismu_coeff_sign dismu_coeff_sign[k][b][s][v] = decode(ctxStatic) / / decode dismu_coeff_abs_level_gt1 dismu_coeff_abs_level_gt1[k][b][s][v] = decode(ctxCoeffGtN[st][b][1][k]) if (dismu_coeff_abs_level_gt1[k][b][s][v]) { value++ / / decode dismu_coeff_abs_level_gt2 dismu_coeff_abs_level_gt2[k][b][s][v] = decode(ctxCoeffGtN[st][b][2][k]) if (dismu_coeff_abs_level_gt2[k][b][s][v]) { value++ / / decode dismu_coeff_abs_level_gt3 dismu_coeff_abs_level_gt3[k][b][s][v] = decode(ctxCoeffGtN[st][b][3][k]) if (dismu_coeff_abs_level_gt3[k][b][s][v]) { / / decode dismu_coeff_abs_level_rem dismu_coeff_abs_level_rem[k][b][s][v] = decodeExpGolomb(ctxCoeffRemPrefix[st][b][k]) value += (1 + dismu_coeff_abs_level_rem) } } } if (dismu_coeff_sign[k][b][s][v]) { value = -value } } Qdisp[displSubMeshID][dispOffset + s * subBlockSize + v][k] = value } } } dispOffset += dispCount[b] } } Here, decode(ctx) is a function that decodes a 1-bit value with the corresponding context ctx as an argument, and decodeExpGolomb(ctxPrefix, ctxSuffix) is a function that decodes a value binary-ized with a k-th Golomb code (e.g., k = 0) Here, decode(ctx) is a function that decodes a 1-bit value with the corresponding context ctx as an argument, and decodeExpGolomb(ctxPrefix, ctxSuffix) is a function that decodes a value binary-ized with a k-th Golomb code (e.g., k = 0). ctxPrefix[n] is used as the context of the prefix bin position n, and ctxSuffix[m] is used as the context of the suffix bin position m. When not using the context for the suffix (using bypass), it can simply be described as decodeExpGolomb(ctxPrefix). value++ is an operation that increments the variable value by 1, and value += 1, value = value + 1. subBlockSize is the size of the sub-block. for indicates a loop. subBlockSize may use a value that is a power of 2 from 16 to 4096. For example, it may be 128 or 256. dispCount[b] is the number of mesh displacements at the detailed level b.
[0157] The mesh displacement decoder 305 may derive the value of the mesh displacement value from dismu_coeff_abs_level_gt0, dismu_coeff_abs_level_gt1, dismu_coeff_abs_level_gt2, dismu_coeff_abs_level_gt3, dismu_coeff_abs_level_rem, dismu_coeff_level_sign as follows, instead of the method of the pseudo-code described above. The value is stored in QDisp. The decoding may be performed in units of sub-meshes indicated by subMeshID (= displSubMeshID).
[0158] absCoeff = dismu_coeff_abs_level_gt0 + dismu_coeff_abs_level_gt1 + dismu_coeff_abs_level_gt2 + dismu_coeff_abs_level_gt3 + dismu_coeff_abs_level_rem value = absCoeff * (1 - 2 * dismu_coeff_sign) Alternatively, the mesh displacement decoding unit 305 may decode the syntax elements dismu_nz_subBlock, dismu_coeff_abs_level_gtN, dismu_coeff_abs_level_rem, dismu_coeff_sign by the following process and derive the mesh displacement Qdisp. for (k = 0; k < numDim; k++) { / / dimension (component) loop dispOffset = 0 for (b = 0; b <numLOD; b++) { / / Level of Detail loop, block loop numSubBlocks = dispCount[b] / subBlockSize + 1 for (s = 0; s < numSubBlocks; s++) { / / subblock loop / / decode dismu_nz_subBlock dismu_nz_subBlock [k][b][s] = decode(ctxCodedSubBlock[st][b][k]) if (dismu_nz_subBlock [k][b][s]) { for (v = 0; v < subBlockSize; v++) { / / coefficient loop within subblock value = 0 / / decode dismu_coeff_abs_level_gt0 dismu_coeff_abs_level_gt0[k][b][s][v] = decode(ctxCoeffGtN[st][b][0][k]) if (dismu_coeff_abs_level_gt0[k][b][s][v]) { / / decode dismu_coeff_sign dismu_coeff_sign[k][b][s][v] = decode(ctxStatic) N = 1 maxGtN = 3 while (N <= maxGtN) { value++ / / decode dismu_coeff_abs_level_gtN (N=1..maxGtN) dismu_coeff_abs_level_gtN[k][b][s][v] = decode(ctxCoeffGtN[st][b][N][k]) if (!dismu_coeff_abs_level_gtN[k][b][s][v]) break N++ } if (dismu_coeff_abs_level_gtN[k][b][s][v]) { / / decode dismu_coeff_abs_level_rem dismu_coeff_abs_level_rem[k][b][s][v] = decodeExpGolomb(ctxCoeffRemPrefix[st][b][k]) value += (1 + dismu_coeff_abs_level_rem[k][b][s][v]) } if (dismu_coeff_sign[k][b][s][v]) { value = -value } } Qdisp[displSubMeshID][dispOffset + s * subBlockSize + v][k] = value } } } dispOffset += dispCount[b] } } The break in the pseudo code is meant to skip the subsequent operations and exit from the nearest loop. That is what it means.
[0159] Note that maxGtN is not limited to 3. For example, when maxGtN = 2, a configuration for encoding / decoding the syntax elements dismu_coeff_abs_level_gt0, dismu_coeff_abs_level_gt1, dismu_coeff_abs_level_gt2 may be used. Or when maxGtN = 4, a configuration for encoding / decoding the syntax elements dismu_coeff_abs_level_gt0, dismu_coeff_abs_level_gt1, dismu_coeff_abs_level_gt2, dismu_coeff_abs_level_gt3, dismu_coeff_abs_level_gt4 may be used. _abs_level_gt0, dismu_coeff_abs_level_gt1, dismu_coeff_abs_level_gt2 may be used. Or when maxGtN = 4, a configuration for encoding / decoding the syntax elements dismu_coeff_abs_level_gt0, dismu_coeff_abs_level_gt1, dismu_coeff_abs_level_gt2, dismu_coeff_abs_level_gt3, dismu_coeff_abs_level_gt4 may be used. dismu_coeff_abs_level_gt1, dismu_coeff_abs_level_gt2, dismu_coeff_abs_level_gt3, dismu_coeff_abs_level_gt4 may be used.
[0160] The inverse quantization unit 3053 performs inverse quantization based on the quantization scale value iscale to derive the mesh displacement Tdisp after conversion (e.g., wavelet transform). Tdisp may be in the Cartesian coordinate system or the local coordinate system. iscale is a value derived from the quantization parameters of each component of the mesh displacement image. The inverse quantization may be performed in units of the sub - mesh indicated by subMeshID (= displSubMeshID). It may be performed in units of the sub - mesh. Tdisp[subMeshID][0][] = (Qdisp[subMeshID][0][] * iscale[0] + iscaleOffset) >> iscaleShift Tdisp[subMeshID][1][] = (Qdisp[subMeshID][1][] * iscale[1] + iscaleOffset) >> iscaleShift Tdisp[subMeshID][2][] = (Qdisp[subMeshID][2][] * iscale[2] + iscaleOffset) >> iscaleShift Here, iscaleOffset = 1<<(iscaleShift-1). iscaleShift may be a predetermined constant, or may be coded at the sequence level, picture / frame level, submesh level indicated by subMeshID (= displSubMeshID), tile / patch level, etc., and then decoded from the coded data. Alternatively, the value decoded from the
[0161] The inverse transform unit 3054 performs an inverse transform g (for example, an inverse wavelet transform) to derive a mesh displacement d. d[0][] = g(Tdisp[subMeshID][0][]) d[1][] = g(Tdisp[subMeshID][1][]) d[2][] = g(Tdisp[subMeshID][2][]) The coordinate system conversion unit 3055 converts the mesh displacement (coordinate system of the mesh displacement) into the Cartesian coordinate system based on the value of the coordinate system conversion information displacementCoordinateSystem. Specifically, when displacementCoordinateSystem==1, the displacement in the local coordinate system is converted into the displacement in the Cartesian coordinate system. Here, d is a 3D vector indicating the mesh displacement before the coordinate system transformation. disp is a 3D vector indicating the mesh displacement after the coordinate system transformation, and is in the Cartesian coordinate system. n_vec, t_vec, b_vec are 3D vectors (in the Cartesian coordinate system) corresponding to each axis of the local coordinate system of the target region or target vertex. if (displacementCoordinateSystem == 0) { disp = d } else if (displacementCoordinateSystem == 1){ disp = d[0] * n_vec + d[1] * t_vec + d[2] * b_vec } The derivation method represented by the above vector multiplication is individually expressed by scalars as follows. if (displacementCoordinateSystem == 0) { for (i = 0; i < 3; i++) {disp[i] = d[i]} } else if (displacementCoordinateSystem == 1){ for (i = 0; i < 3; i++) {disp[i] = d[0] * n_vec[i] + d[1] * t_vec[i] + d[2] * b_vec[i]} } Note that the same variable name may be assigned before and after the conversion as disp = d, and the value of d is updated by the coordinate conversion to be.
[0162] Alternatively, the following configuration may also be used. if (displacementCoordinateSystem == 0) { disp = d } else if (displacementCoordinateSystem == 1){ disp = d[0] * n_vec + d[1] * t_vec + d[2] * b_vec } else if (displacementCoordinateSystem == 2){ disp = d[0] * n_vec2 + d[1] * t_vec2 + d[2] * b_vec2 } Here, n_vec2, t_vec2, and b_vec2 are 3D vectors (in the Cartesian coordinate system) corresponding to the respective axes of the local coordinate system of the adjacent region.
[0163] Also, the following configuration may be used. if (displacementCoordinateSystem == 0) { disp = d } else if (displacementCoordinateSystem == 1){ disp = d[0] * n_vec3 + d[1] * t_vec3 + d[2] * b_vec3 } Here, n_vec3, t_vec3, and b_vec3 are 3D vectors (in the Cartesian coordinate system) corresponding to the respective axes of the local coordinate system of the target region with fluctuations suppressed. For example, the vectors of the coordinate system used for decoding from the previous coordinate system and the current coordinate system are derived as follows n_vec3 = (w*n_vec3 + (WT - w)*n_vec)>>wShift t_vec3 = (w*t_vec3 + (WT - w)*t_vec)>>wShift b_vec3 = (w*b_vec3 + (WT - w)*b_vec)>>wShift Here, for example, wShift = 2, 3, 4, WT = 1<<wShift, and w = 1..WT - 1. For example, when w = 3 and wShift = 3, n_vec3 = (3*n_vec3 + 5*n_vec)>>3 t_vec3 = (3*t_vec3 + 5*t_vec)>>3 b_vec3 = (3*b_vec3 + 5*b_vec)>>3 Also, a configuration may be used where it can be selected according to the value of the coordinate system conversion information displacementCoordinateSystem decoded from the encoded data as in the following configuration. if (displacementCoordinateSystem == 0) { disp = d } else if (displacementCoordinateSystem == 1){ disp = d[0] * n_vec + d[1] * t_vec + d[2] * b_vec } else if (displacementCoordinateSystem == 6){ disp = d[0] * n_vec3 + d[1] * t_vec3 + d[2] * b_vec3 } (Mesh Reconstruction) FIG. 6 is a functional block diagram showing the configuration of the mesh reconstruction unit 307. The mesh reconstruction unit 307 is composed of a mesh division unit 3071 and a mesh deformation unit 3072.
[0164] The mesh division unit 3071 divides the base mesh output from the base mesh decoding unit 303 and generates a divided mesh.
[0165] FIG. 9(a) shows a part (triangle) of the base mesh. The triangle is composed of vertices v1, v2, and v3. v1, v2, and v3 are three-dimensional vectors. The mesh division unit 3071 generates a divided mesh by adding new vertices v12, v13, and v23 in the middle of each side of the triangle and outputs it (FIG. 9(b)). v12 = (v1 + v2) / 2 v13 = (v1 + v3) / 2 v13 = (v1 + v3) / 2 v23 = (v2 + v3) / 2 Or the following may also be used. v12 = (v1 + v2 + 1) >> 1 v13 = (v1 + v3 + 1) >> 1 v23 = (v2 + v3 + 1) >> 1 The mesh deformation unit 3072 inputs the divided mesh and the mesh displacement, and generates and outputs a deformed mesh by adding the mesh displacements d12, d13, and d23 (Fig. 9(c)). The mesh displacement is the output of the mesh displacement decoding unit 305 (coordinate system conversion unit 3055). d12, d13, and d23 are the mesh displacements corresponding to the respective vertices v12, v13, and v23 added by the mesh division unit 3071. v12' = v12 + d12 v13' = v13 + d13 v23' = v23 + d23 Note that d12 = disp[0][], d23 = disp[1][], and d23 = disp[3][] may be used.
[0166] (Configuration of 3D data encoding device according to the first embodiment) Fig. 10 is a functional block diagram showing the schematic configuration of the 3D data encoding device 11 according to the first embodiment. The 3D data encoding device 11 includes an atlas information encoding unit 101, a base mesh encoding unit 103, a base mesh decoding unit 104, a mesh displacement update unit 106, a mesh displacement encoding unit 107, a me sh displacement decoding unit 108, a mesh reconstruction unit 109, an attribute update unit 110, a padding unit 111, a color space conversion unit 112, an attribute encoding unit 113, a sub-mesh encoding unit 116, a multiplexing unit 114, and a mesh separation unit 115. The 3D data encoding device 11 inputs atlas information, a base mesh, a mesh displacement, a mesh, and an attribute image as 3D data, and outputs encoded data.
[0167] The atlas information encoding unit 101 encodes the atlas information and outputs an atlas information encoding stream .
[0168] The base mesh encoding unit 103 encodes the base mesh and outputs a base mesh encoding s Output the team. Use Draco or the like as the encoding method.
[0169] Since the base mesh decoding unit 104 is the same as the base mesh decoding unit 303, the description thereof is omitted.
[0170] The mesh displacement update unit 106 adjusts the mesh displacement based on the (original) base mesh and the decoded base mesh, and outputs the updated mesh displacement.
[0171] The mesh displacement encoding unit 107 encodes the updated mesh displacement and outputs a mesh displacement code stream.
[0172] Since the mesh displacement decoding unit 108 is the same as the mesh displacement decoding unit 305, the description thereof is omitted.
[0173] Since the mesh reconstruction unit 109 is the same as the mesh reconstruction unit 307, the description thereof is omitted.
[0174] The attribute update unit 110 inputs the (original) mesh, the reconstructed mesh output from the mesh reconstruction unit 109 (mesh deformation unit 3072), and the attribute image, updates the attribute image to match the position (coordinates) of the reconstructed mesh, and outputs the updated attribute image.
[0175] The padding unit 111 inputs the attribute image and performs padding processing on the area where the pixel value is empty.
[0176] The color space conversion unit 112 performs color space conversion from the RGB format to the YCbCr format.
[0177] The attribute encoding unit 113 encodes the YCbCr format output from the color space conversion unit 112 Encode the attribute image and output an attribute video stream. As the encoding method, use VVC, HEVC, etc.
[0178] The sub-mesh encoding unit 116 encodes the sub-mesh information of the atlas information encoding stream. Perform encoding.
[0179] The multiplexing unit 114 multiplexes the atlas information sub-mesh encoding stream, the base mesh encoding stream, the mesh displacement encoding stream, and the attribute video stream, and outputs them as encoded data. As the multiplexing method, use the byte stream format, ISOBMFF, etc. Perform multiplexing and output as encoded data. As the multiplexing method, use the byte stream format, ISOBMFF, etc.
[0180] (Operation of the mesh separation unit) The mesh separation unit 115 generates a base mesh and a mesh displacement from the mesh.
[0181] FIG. 13 is a functional block diagram showing the configuration of the mesh separation unit 115. The mesh separation unit 115 includes a mesh thinning unit 1151, a mesh division unit 1152, and a mesh displacement derivation unit 1153.
[0182] The mesh thinning unit 1151 generates a base mesh by thinning out some vertices from the mesh and outputs it.
[0183] FIG. 14(a) shows a part of the mesh, and the mesh is composed of vertices v1, v2, v3, v4, v5, v6. v1, v2, v3, v4, v5, v6 are each three-dimensional vectors. The mesh thinning unit 1151 generates and outputs a base mesh by thinning out vertices v4, v5, v6 (FIG. 14(b)).
[0184] The mesh division unit 1152 divides the base mesh in the same manner as the mesh division unit 3071 to generate divided meshes (FIG. 14(c)). v4' = (v1 + v2) / 2 v5' = (v1 + v3) / 2 v6' = (v2 + v3) / 2 Based on the mesh and the divided meshes, the mesh displacement derivation unit derives and outputs the displacements d4, d5, and d6 of the vertices v4, v5, and v6 corresponding to the vertices v4', v5', and v6 as mesh displacements (Fig. 14(d)) 。 d4 = v4 - v4' d5 = v5 - v5' d6 = v6 - v6' (Encoding of the base mesh) Fig. 11 is a functional block diagram showing the configuration of the base mesh encoding unit 103. The base mesh encoding unit 103 includes a mesh encoding unit 1031, a mesh decoding unit 1032, a motion information encoding unit 1033, a motion information decoding unit 1034, a mesh motion compensation unit 1035, a reference mesh memory 1036, a switch 1037, and a switch 1038. The base mesh encoding unit 103 may have a configuration including a base mesh quantization unit (not shown) after input of the base mesh. The switches 1037 and 3038 are connected to the side where no motion compensation is performed when encoding the base mesh without referring to another base mesh (e.g., a previously encoded base mesh) (intra encoding). Otherwise, they are connected to the side where motion compensation is performed when encoding the base mesh with reference to another base mesh (inter encoding).
[0185] The mesh encoding unit 1031 has an intra encoding function, intra-encodes the base mesh, and outputs a base mesh encoding stream. As the encoding method, Draco or the like is used
[0186] The mesh decoding unit 1032 is the same as the mesh decoding unit 3031, and thus the description is omitted
[0187] The motion information encoding unit 1033 has an inter-encoding function, inter-encodes the base mesh, and outputs a base mesh encoded stream. As the encoding method, entropy encoding such as arithmetic encoding is used.
[0188] Since the motion information decoding unit 1034 is the same as the motion information decoding unit 3032, the description thereof is omitted.
[0189] Since the mesh motion compensation unit 1035 is the same as the mesh motion compensation unit 3033, the description thereof is omitted.
[0190] Since the reference mesh memory 1036 is the same as the reference mesh memory 3034, the description thereof is omitted.
[0191] (Encoding of mesh displacement) FIG. 12 is a functional block diagram showing the configuration of the mesh displacement encoding unit 107. Mesh displacement The encoding unit 107 is composed of a coordinate system conversion unit 1071, a conversion unit 1072, a quantization unit 1073, a binarization unit 1074, an arithmetic encoding unit 1075, a context selection unit 1076, and a context initialization unit 1077.
[0192] The coordinate system conversion unit 1071 converts the coordinate system of the mesh displacement from the Cartesian coordinate system to a coordinate system (for example, a local coordinate system) for encoding the displacement based on the value of the coordinate system conversion information displacementCoordinateSystem. Here, disp is a three-dimensional vector indicating the mesh displacement before coordinate system conversion, d is a three-dimensional vector indicating the mesh displacement after coordinate system conversion, and n_vec, t_vec, b_vec are three-dimensional vectors (in the Cartesian coordinate system) indicating the respective axes of the local coordinate system. if (displacementCoordinateSystem == 0) { d = disp } else if (displacementCoordinateSystem == 1){ d = (disp * n_vec, disp * t_vec, disp * b_vec) } The mesh displacement encoding unit 107 may update the value of displacementCoordinateSystem at the sequence level or at the picture / frame level. The initial value is 0 indicating the Cartesian coordinate system.
[0193] When updating displacementCoordinateSystem at the sequence level, use the syntax of the configuration in FIG. 7 Set 0 for the Cartesian coordinate system and 1 for the local coordinate system in asps_vdmc_ext_displacement_coordinate_system.
[0194] When changing displacementCoordinateSystem at the picture / frame level, use the syntax of the configuration in FIG. 8 Set 1 when updating the coordinate system and 0 when not updating the coordinate system in afps_vdmc_ext_displacement_coordinate_system_enable_flag. Set 0 for the Cartesian coordinate system and 1 for the local coordinate system in afps_vdmc_ext_displacement_coordinate_system.
[0195] The conversion unit 1072 performs a conversion f (e.g., wavelet conversion) to derive the mesh displacement Tdisp after conversion. Tdisp[displSubMeshID][0][] = f(d[displSubMeshID][0][]) Tdisp[displSubMeshID][1][] = f(d[displSubMeshID][1][]) Tdisp[displSubMeshID][2][] = f(d[displSubMeshID][2][]) The quantization unit 1073 performs quantization based on the quantization scale value scale derived from the quantization parameters of each component of the mesh displacement, and derives the quantized mesh displacement Qdisp. Qdisp[displSubMeshID][0][] = Tdisp[displSubMeshID][0][] / scale[0] Qdisp[displSubMeshID][1][] = Tdisp[displSubMeshID][1][] / scale[1] Qdisp[displSubMeshID][2][] = Tdisp[displSubMeshID][2][] / scale[2] Alternatively, the Qdisp may be derived by the following formula by approximating the scale value with a power of 2. scale[i] = 1 << scale2[i] Qdisp[displSubMeshID][0][] = Tdisp[displSubMeshID][0][] >> scale2[0] Qdisp[displSubMeshID][1][] = Tdisp[displSubMeshID][1][] >> scale2[1] Qdisp[displSubMeshID][2][] = Tdisp[displSubMeshID][2][] >> scale2[2] The binary quantization unit 1074 encodes the quantized mesh displacement Qdisp, which is a multi-valued signal, into a binary signal. The binary signal may be a k-th order exponential Golomb code.
[0196] The arithmetic coding unit 1075 arithmetically encodes the binary signal and outputs a mesh displacement encoded stream.
[0197] The context selection unit 1076 is the same as the context selection unit 3056, so the description is omitted.
[0198] Let ctxStatic be a static context with a fixed probability and no context update. Encoding of the syntax elements indicated by ctxStatic may be performed without using a context. encode(ctxStatic) may be implemented as encode_bypass() and use bypass-only processing.
[0199] The context initialization unit 1077 is the same as the context initialization unit 3057, and thus its description is omitted. Here, an example of using a context is described, but some syntax elements may be bypass-encoded without using a context. In a configuration for bypass encoding, it is effective to reduce the context memory and processing amount. For example, the syntax element dismu_coeff_abs_level_rem may be bypass-encoded without using a context. By bypass-encoding these syntax elements, it is possible to reduce the context memory and processing amount while maintaining the encoding efficiency.
[0200] The mesh displacement encoding unit 107 encodes the mesh displacement Qdisp by the following processing. for (k = 0; k < numDim; k++) { / / dimension (component) loop if (!lastSig) continue dispOffset = 0 for (b = 0; b <numLOD; b++) { / / Level of Detail loop, block loop numBlocks = dispCount[b] / subBlockSize + 1 for (s = 0; s < numBlocks; s++) { / / subblock loop / / encode dismu_nz_subBlock encode(dismu_nz_subBlock[k][b][s], ctxCodedSubBlock[st][b][k]) for (v = 0; v < subBlockSize; v++) { / / coefficient loop within subblock / / encode dismu_coeff_abs_level_gt0 d = Qdisp[dispOffset + s * subBlockSize + v][k] encode(d != 0, ctxCoeffGtN[st][b][0][k]) if (!d) continue / / encode dismu_coeff_sign encode(d < 0, ctxStatic) d = abs(d) - 1 / / encode dismu_coeff_abs_level_gt1 encode(d != 0, ctxCoeffGtN[st][b][1][k]) if (!d) continue d = abs(d) - 1 / / encode dismu_coeff_abs_level_gt2 encode(d != 0, ctxCoeffGtN[st][b][2][k]) if (!d) continue d = abs(d) - 1 / / encode dismu_coeff_abs_level_gt3 encode(d != 0, ctxCoeffGtN[st][b][3][k]) if (!d) continue / / encode dismu_coeff_abs_level_rem encodeExpGolomb(--d, ctxCoeffRemPrefix[st][b][k])} } dispOffset += dispCount[b] } } In the suspect code, continue means skipping the subsequent operations and jumping to the beginning of the loop (the next iteration). Here, encode() and encodeExpGolomb() are functions that take a value and the corresponding context as arguments and each arithmetically encode a 1-bit value and the binary sequence of a k-th Golomb code. dispCount[b] is , the number of mesh displacements at the detailed level b. lastSig is a flag indicating whether the current coefficient is the last non-zero coefficient within the sub-block in the scan order. lastSig = 0 indicates that the current coefficient is not the last non-zero coefficient within the sub-block in the scan order. lastSig = 1 indicates that the current coefficient is the last non-zero coefficient within the sub-block in the scan order. Or, the mesh displacement Qdisp may be encoded by the following process.
[0201] Or, the mesh displacement Qdisp may be encoded by the following process. for (k = 0; k < numDim; k++) { / / dimension (component) loop if (!lastSig) continue dispOffset = 0 for (b = 0; b <numLOD; b++) { / / Level of Detail loop, block loop numBlocks = dispCount[b] / subBlockSize + 1 for (s = 0; s < numBlocks; s++) { / / subblock loop / / encode dismu_nz_subBlock encode(dismu_nz_subBlock[k][b][s], ctxCodedSubBlock[st][b][k]) for (v = 0; v < subBlockSize; v++) { / / coefficient loop within subblock / / encode dismu_coeff_abs_level_gt0 d = Qdisp[dispOffset + s * subBlockSize + v][k] encode(d != 0, ctxCoeffGtN[st][b][0][k]) if (!d) continue / / encode dismu_coeff_sign encode(d < 0, ctxStatic) N = 1 maxGtN = 3 while (N <= maxGtN) { d = abs(d) - 1 / / encode dismu_coeff_abs_level_gtN (N=1..maxGtN) encode(d != 0, ctxCoeffGtN[st][b][N][k]) if (!d) break N++ } if (d) { / / encode dismu_coeff_abs_level_rem encodeExpGolomb(--d, ctxCoeffRemPrefix[st][b][k]) } } } dispOffset += dispCount[b] } } Note that maxGtN is not limited to 3. For example, when maxGtN = 2, it may be configured to encode the syntax elements dismu_coeff_abs_level_gt0, dismu_coeff_abs_level_gt1, dismu_coeff_abs_level_gt2. Or when maxGtN = 4, it may be configured to encode the syntax elements dismu_coeff_abs_level_gt0, dismu_coeff_abs_level_gt1, dismu_coeff_abs_level_gt2, dismu_coeff_abs_level_gt3, dismu_coeff_abs_level_gt4.
[0202] As described above, an embodiment of the present invention has been described in detail with reference to the drawings. However, the specific configuration is not limited to the above, and various design changes and the like can be made without departing from the gist of the present invention.
[0203] 〔Application Example〕 The above-described 3D data encoding device 11 and 3D data decoding device 31 can be mounted and used in various devices that perform transmission, reception, recording, and reproduction of 3D data. Note that the 3D data may be natural 3D data captured by a camera or the like, or may be artificial 3D data (including CG and GUI) generated by a computer or the like.
[0204] The embodiments of the present invention are not limited to the above-described embodiments, and various changes are possible within the scope shown in the claims. That is, embodiments obtained by appropriately combining technical means changed within the scope shown in the claims are also included in the technical scope of the present invention.
Industrial Applicability
[0205] Embodiments of the present invention can be suitably applied to a 3D data decoding device that decodes encoded data in which 3D data is encoded, and a 3D data encoding device that generates encoded data in which 3D data is encoded. Further, it can be suitably applied to the data structure of the encoded data generated by the 3D data encoding device and referred to by the 3D data decoding device.
Explanation of Signs
[0206] 11 3D data encoding device 101 Atlas information encoding unit 103 Base mesh encoding unit 1031 Mesh encoding unit 1032 Mesh decoding unit 1033 Motion information encoding unit 1034 Motion information decoding unit 1035 Mesh motion compensation unit 1036 Reference mesh memory 1037 Switch 1038 Switch 1039 Skip encoding 104 Base mesh decoding unit 106 Mesh displacement update unit 107 Mesh displacement encoding unit 1071 Coordinate system conversion unit 1072 Conversion unit 1073 Quantization unit 1074 Binarization unit 1075 Arithmetic encoding unit 1076 Context selection unit 1077 Context initialization unit 108 Mesh displacement decoding unit 109 Mesh reconstruction unit 110 Attribute update unit 111 Padding unit 112 Color space conversion unit 113 Attribute encoding unit 114 Multiplexing unit 115 Mesh separation unit 1151 Mesh thinning unit 1152 Mesh Division Unit 1153 Mesh Displacement Derivation Unit 116 Sub-Mesh Encoding Unit 21 Network 31 3D Data Decoder 301 Demultiplexing Unit 302 Atlas Information Decoding Unit 303 Base Mesh Decoding Unit 3031 Mesh Decoding Unit 3032 Motion Information Decoding Unit 3033 Mesh Motion Compensation Unit 3034 Reference Mesh Memory 3035 Switch 3036 Switch 3037 Skip Decoding Unit 305 Mesh Displacement Decoding Unit 3051 Arithmetic Decoding Unit 3052 Multi-Valuing Unit 3053 Inverse Quantization Unit 3054 Inverse Transformation Unit 3055 Coordinate System Transformation Unit 3056 Context Selection Unit 3057 Context Initialization Unit 307 Mesh Reconstruction Unit 306 Attribute Decoding Unit 3071 Mesh Division Unit 3072 Mesh Deformation Unit 308 Color Space Transformation Unit 309 Sub-Mesh Decoding Unit 41 3D Data Display Device
Claims
1. A 3D data decoding device for decoding mesh data or point cloud data, comprising: a submesh decoding unit for decoding submesh information from encoded data in which the mesh data or point cloud data is encoded; a base mesh decoding unit for decoding a base mesh from the encoded data and the submesh information; a mesh displacement decoding unit for decoding mesh displacement from the encoded data and the submesh information; and a mesh reconstruction unit for decoding a mesh from the decoded base mesh and mesh displacement, wherein the mesh displacement decoding unit decodes the mesh displacement from the encoded data using the submesh information decoded by the submesh decoding unit.
2. 2. The 3D data decoding device according to claim 1, wherein the submesh decoding unit includes a flag indicating whether a submesh of the base mesh corresponds to a submesh of the mesh displacement.
3. The 3D data decoding device according to claim 1, characterized in that the submesh decoding unit decodes syntax elements equal to the number of submeshes minus 2, depending on a flag indicating whether there are multiple submeshes.
4. The 3D data decoding device according to claim 1, characterized in that the mesh displacement decoding unit decodes the sub-mesh information of the mesh displacement in the mesh displacement decoding unit when the sub-mesh information of the mesh displacement is not decoded in the sub-mesh decoding unit.
5. A 3D data encoding device that encodes mesh data or point cloud data comprises a submesh encoding unit that encodes submesh information, a base mesh encoding unit that encodes a base mesh using the submesh information, and a mesh displacement encoding unit that encodes mesh displacement using the submesh information, wherein the mesh displacement encoding unit encodes the mesh displacement using the submesh information encoded by the submesh encoding unit.
6. 6. The 3D data encoding device according to claim 5, wherein the submesh encoding unit includes a flag indicating whether a submesh of the base mesh corresponds to a submesh of the mesh displacement.
7. The 3D data encoding device according to claim 5, wherein the sub-mesh encoding unit encodes a syntax element of a number of sub-meshes minus two according to a flag indicating whether there are a plurality of sub-meshes.
8. The 3D data encoding device according to claim 5, wherein the mesh displacement encoding unit encodes sub-mesh information of the mesh displacement when the sub-mesh encoding unit does not encode the sub-mesh information of the mesh displacement.