Mesh decoding device, mesh decoding method, and program

WO2026196722A1PCT designated stage Publication Date: 2026-09-24KDDI CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/044304
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-21
Filing Date
2025-12-18
Publication Date
2026-09-24

Smart Images

  • Figure JP2025044304_24092026_PF_FP_ABST
    Figure JP2025044304_24092026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention calculates an optimal value that falls within a prescribed bit depth for a base mesh and an AC displacement amount. A mesh decoding device 200 according to the present invention comprises a base mesh decoding unit 202 that decodes a base mesh bit stream to generate and output a base mesh. The base mesh decoding unit 202 clips coordinates of a decoded vertex according to the bit depth such that the coordinates of the vertex fall within the prescribed bit depth.
Need to check novelty before this filing date? Find Prior Art

Description

Mesh decoding apparatus, mesh decoding method and program

[0001] The present invention relates to a mesh decoding apparatus, a mesh decoding method and a program.

[0002] Non-Patent Document 4 utilizes the framework of Non-Patent Document 5 to decode a mesh by dividing it into a rough base mesh and detailed displacement amounts. For the base mesh, the base mesh decoding unit decodes information such as vertex coordinates, Connectivity, UV coordinates (a type of Attribute), and the like, and then reconstructs a mesh from the base mesh and the displacement amount by adding the atlas decoded by the atlas data decoding unit. A technology for reconstructing a mesh from a base mesh and a displacement amount by adding an atlas decoded by an atlas data decoding unit is disclosed.

[0003] Further, in the decoding of the above-described base mesh, any one of decoding by an intra frame (I frame), decoding by an inter frame (P frame), and decoding by a skip frame is used.

[0004] Further, there are two types of the above-described displacement amount decoding methods: a method using a video codec and a method using arithmetic decoding. In the present specification, these are referred to as video displacement amounts and AC displacement amounts, respectively. Further, the description of decoding the displacement amount may be omitted.

[0005] It should be noted that Non-Patent Document 4 discloses that, in addition to meshes, video data called texture, which is a type of Attribute, is decoded by a video decoding unit.

[0006] Khaled Mammou, Jungsun Kim, Alexis M Tourapis, Dimitri Podborski, and Krasimir Kolarov, “[V-CG] Apple's Dynamic Mesh Coding CfP Response,” April 2022, ISO / IEC JTC 1 / SC 29 / WG 7 m59281Google Draco, accessed May 26, 2022 [Online], https: / / google.github.io / dracoJean-Eudes Marvie, Olivier Mocquard, “[V-DMC][EE4.4-related] An efficient EdgeBreaker implementation,” April 2023, ISO / IEC JTC 1 / SC 29 / WG 7 m63344”Technologies for V-DMC,” ISO / IEC JTC 1 / SC 29 / WG7 N01099, Jan. 2025“Information technology - Coded Representation of Immersive Media - Part 5: Visual Volumetric Video-based Coding (V3C) and Video-based Point Cloud Compression (V-PCC),” ISO / IEC JTC 1 / SC 29 / WG 7, ISO / IEC 23090-5:2021(2E)

[0007] However, Non-Patent Document 4 states that when decoding the basic mesh or AC displacement using an interframe (P-frame), adding the corresponding value from the reference frame to calculate the value of that frame may exceed the specified bit depth. In that case, there was a problem that converting to the specified bit depth by bit shifting would not result in the correct value.

[0008] For example, at a vertex of a basic mesh, the motion vector (MV) is (1,1,1), and the coordinates of the corresponding vertex in the reference frame are (7,7,7). Therefore, the vertex coordinates in that frame would be (8,8,8). However, if the default bit depth is 3, bit shifting will cause the vertex coordinates to become (4,4,4).

[0009] Therefore, the present invention has been made in view of the above-mentioned problems, and aims to provide a mesh decoding device, a mesh decoding method, and a program that can be expected to have the effect of calculating an optimal value that fits within a specified bit depth for the basic mesh and AC displacement amount.

[0010] The first feature of the present invention is a mesh decoding device comprising a basic mesh decoding unit that decodes a basic mesh bitstream to generate and output a basic mesh, wherein the basic mesh decoding unit clips the coordinates of the decoded vertices by the bit depth so that the coordinates of the vertices fit within a specified bit depth.

[0011] A second feature of the present invention is a mesh decoding device comprising a displacement decoding unit that decodes a displacement bitstream to generate and output a displacement amount, wherein the displacement decoding unit clips the decoded coefficient sequence by the bit depth so that the sequence fits within a predetermined bit depth.

[0012] A third feature of the present invention is a mesh decoding method comprising the steps of decoding a basic mesh bitstream to generate and output a basic mesh, wherein the gist of the method is to clip the coordinates of the decoded vertices using a specified bit depth so that the coordinates of the vertices fit within a specified bit depth.

[0013] A fourth feature of the present invention is a program that causes a computer to function as a mesh decoding device, wherein the mesh decoding device comprises a basic mesh decoding unit that decodes a basic mesh bitstream to generate and output a basic mesh, and the basic mesh decoding unit clips the coordinates of the decoded vertices by the bit depth so that the coordinates of the vertices are contained within a specified bit depth.

[0014] According to the present invention, it is possible to provide a mesh decoding device, a mesh decoding method, and a program that can be expected to have the effect of calculating optimal values ​​that fit within a specified bit depth for the basic mesh and AC displacement amount.

[0015] Figure 1 shows an example of the configuration of a mesh processing system 1 according to one embodiment. Figure 2 shows an example of the functional blocks of a mesh decoding device 200 according to one embodiment. Figure 3 shows an example of a basic mesh and a subdivided mesh. Figure 4 shows an example of the functional blocks of the basic mesh decoding unit 202 of the mesh decoding device 200 according to one embodiment. Figure 5 shows an example of the functional blocks of the intra decoding unit 202B of the basic mesh decoding unit 202 of the mesh decoding device 200 according to one embodiment. Figure 6 shows an example of the correspondence between the vertices of the basic mesh of a P frame and the vertices of the basic mesh of an I frame. Figure 7 shows an example of ASPS. Figure 8 shows an example of AFPS. Figure 9 shows an example of ATH. Figure 10 shows an example of MMDU. Figure 11 shows an example of IMDU. Figure 12 shows an example of the configuration of a basic mesh bitstream. Figure 13 shows an example of BMSPS. Figure 14 shows an example of BMFPS. Figure 15 shows an example of BMSH. Figure 16 is a diagram showing an example of the functional block of the displacement amount decoding unit 206 of the mesh decoding device 200 according to one embodiment. Figure 17 is a diagram illustrating an example of AC displacement. Figure 18 is a diagram showing an example of DSPS. Figure 19 is a diagram showing an example of DFPS. Figure 20 is a diagram showing an example of DH. Figure 21 is a diagram showing an example of DH. Figure 22 is a diagram showing an example of the functional block of the inter-decoding unit 202E of the basic mesh decoding unit 202 of the mesh decoding device 200 according to one embodiment. Figure 23 is a diagram illustrating an example of how the motion vector prediction unit 202E3 of the inter-decoding unit 202E of the basic mesh decoding unit 202E of the mesh decoding device 200 according to one embodiment calculates the MVP of the vertices to be decoded. Figure 24 is a diagram illustrating an example of a basic mesh. Figure 25 is a diagram illustrating an example of a basic mesh. Figure 26 is a diagram illustrating the mesh buffer unit 202C of the basic mesh decoding unit 202 of the mesh decoding device 200 according to one embodiment. Figure 27 is a flowchart showing an example of the processing of the mesh buffer section 202C of the basic mesh decoding section 202 of a mesh decoding device 200 according to one embodiment.Figure 28 shows an example of a NAL header. Figure 29 shows an example of a BMSPS.

[0016] Embodiments of the present invention will be described below with reference to the drawings. Note that the components in the following embodiments can be replaced with existing components as appropriate, and various variations are possible, including combinations with other existing components. Therefore, the description of the following embodiments does not limit the content of the invention as described in the claims.

[0017] <First Embodiment> The mesh processing system according to this embodiment will be described below with reference to Figures 1 to 29.

[0018] Figure 1 shows an example of the configuration of the mesh processing system 1 according to this embodiment. As shown in Figure 1, the mesh processing system 1 includes a mesh encoding device 100 and a mesh decoding device 200.

[0019] Figure 2 shows an example of the functional blocks of the mesh decoding device 200 according to this embodiment.

[0020] As shown in Figure 2, the mesh decoding device 200 includes a multiplexing unit 201, a basic mesh decoding unit 202, a subdivision unit 203, a mesh decoding unit 204, a patch integration unit 205, a displacement decoding unit 206, a video decoding unit 207, and an atlas data decoding unit 208.

[0021] Here, the basic mesh decoding unit 202, the subdivision unit 203, the mesh reconstruction unit 204, and the displacement decoding unit 206 are configured to process the mesh in patch units, and the results of these processes may then be integrated in the patch integration unit 205.

[0022] In the example shown in Figure 3, the mesh is divided into patch 1, which consists of basic surfaces 1 and 2, and patch 2, which consists of basic surfaces 3 and 4.

[0023] The multiplexing unit 201 is configured to separate the multiplexed bitstream into a basic mesh bitstream, a displacement bitstream, a texture bitstream, and an atlas bitstream.

[0024] The subdivision unit 203 is configured to generate and output subdivided vertices and their connection information from the basic mesh decoded by the basic mesh decoding unit 202 using a subdivision method indicated by the control information (first control information and second control information). The basic mesh consists of one submesh or multiple submeshes.

[0025] Here, the base mesh, the added subdivided vertices, and their connection information are collectively referred to as the "subdivided mesh." Similarly, the submesh, the added subdivided vertices, and their connection information are collectively referred to as the "subdivided submesh."

[0026] The mesh decoding unit 204 is configured to generate and output a decoded mesh using control information, a subdivided mesh, the subdivided vertex normals, and displacement amounts.

[0027] The displacement amount decoding unit 206 is configured to decode the displacement amount bitstream based on control information, generate a displacement amount, and output it.

[0028] The displacement amount decoding unit 206 may decode and output a level value image from the displacement amount bitstream using a video codec, or it may decode and output a sequence of level values ​​from the displacement amount bitstream using arithmetic decoding.

[0029] The video decoding unit 207 is configured to decode and output textures using a video codec.

[0030] The Atlas data decoding unit 208 is configured to decode the Atlas bitstream and output control information. This control signal may be used as metadata by the basic mesh decoding unit 202, the subdivision unit 203, the mesh decoding unit 204, the displacement decoding unit 206, and the video decoding unit 207.

[0031] The atlas data decoding unit 208 initializes a variable called PredictorIdx at the beginning of decoding, setting the initial value of PredictorIdx to zero.

[0032] However, PredictorIdx indicates the index of the reference frame of the patch immediately preceding the current patch, and the index of the reference frame of the current patch is used as the predicted value.

[0033] According to this embodiment, it is expected that the PredictorIdx value will be correctly obtained when decoding the atlas data.

[0034] <Basic Mesh Decoding Unit 202> The basic mesh decoding unit 202 is configured to decode the basic mesh bitstream, generate a basic mesh, and output it.

[0035] Here, the basic mesh consists of multiple vertices in three-dimensional space and edges that connect these multiple vertices.

[0036] The basic mesh decoding unit 202 may be configured to decode the basic mesh bitstream using, for example, Draco as shown in Non-Patent Document 2 or the technology described in Non-Patent Document 3.

[0037] As shown in Figure 4, the basic mesh decoding unit 202 comprises a separation unit 202A, an intra decoding unit 202B, a mesh buffer unit 202C, a connection information decoding unit 202D, and an inter decoding unit 202E.

[0038] (Separation unit 202A) The separation unit 202A is configured to classify the basic mesh bitstream into I-frame bitstreams, P-frame bitstreams, and skip frame bitstreams, and to extract data from sub-mesh data units accordingly.

[0039] Specifically, the data for a submesh data unit having the submesh ID "submeshID" is extracted from the basic mesh submesh data unit bmesh_submesh_unit(submeshID) which has the submesh ID "submeshID".

[0040] (Intra decoding unit 202B) The intra decoding unit 202B is configured to decode the coordinates and connection information of vertices of an I-frame from the bitstream of the I-frame, for example, by using Draco described in Non-Patent Document 2 or the technology described in Non-Patent Document 3.

[0041] FIG. 5 is a diagram showing an example of functional blocks of the intra decoding unit 202B.

[0042] As shown in FIG. 6, the intra decoding unit 202B includes an arbitrary intra decoding unit 202B1 and an aligning unit 202B2.

[0043] The arbitrary intra decoding unit 202B1 is configured to decode the coordinates and connection information of unordered vertices of an I-frame from the bitstream of the I-frame by using any method including Draco described in Non-Patent Document 2 or the technology described in Non-Patent Document 3.

[0044] The aligning unit 202B2 is configured to output vertices by rearranging the unordered vertices in a predetermined order.

[0045] As the predetermined order, for example, Morton code order may be used, or raster scan order may be used.

[0046] In addition, the aligning unit 202B2 may collect overlapping vertices, which are a plurality of vertices with matching coordinates in the decoded base mesh, into a single vertex, and then rearrange them in a predetermined order.

[0047] (Mesh buffer unit 202C) The mesh buffer unit 202C is configured to store the coordinates and connection information of vertices of an I-frame decoded by the intra decoding unit 202B. Here, a specific buffer may be provided that stores pairs of indices A(k) and B(k) of vertices existing as overlapping vertices in a predetermined order.

[0048] (Connection information decoding unit 202D) The connection information decoding unit 202D is configured to convert the connection information of an I-frame or a reference frame extracted from the mesh buffer unit 202C into connection information of a P-frame.

[0049] (Inter-decoding unit 202E) The inter-decoding unit 202E is configured to decode the coordinates of the vertices of the P-frame by adding the coordinates of the vertices of the reference frame taken from the mesh buffer unit 202C and the motion vector decoded from the bitstream of the P-frame.

[0050] Furthermore, the inter-decoding unit 202E can adjust the vertex indices of the P-frame using pairs of vertex indices A(k) and B(k) of vertices that exist as duplicate vertices stored in the specific buffer.

[0051] In this embodiment, as shown in Figure 6, a correspondence exists between the vertices of the basic mesh of the P-frame and the vertices of the basic mesh of the reference frame (I-frame or P-frame). Here, the motion vector decoded by the inter-decoding unit 202E is the difference vector between the coordinates of the vertices of the basic mesh of the P-frame and the coordinates of the vertices of the basic mesh of the I-frame.

[0052] (Configuration of Atlas Bitstream) The Atlas bitstream may include ASPS (Atlas Sequence Parameter Set), AFPS (Atlas Frame Parameter Set), and ATS (Atlas Tile Header), which are sets of control information related to Atlas decoding.

[0053] The following describes an example of the configuration of an Atlas bitstream, referring to Figures 7 to 11.

[0054] In Figures 7 to 11, u(n) means an n-bit code, and ue(v) means an unsigned variable-length zero-order exponential Golomb code.

[0055] As shown in Figure 7, the ASPS may include a control signal, asps_atlas_sequence_parameter_set_id, that indicates its own ASPS ID.

[0056] The control signal asps_atlas_sequence_parameter_set_id is a type of APSP ID and is decoded by ue(v). However, the range of the control signal asps_atlas_sequence_parameter_set_id is limited to 0 to 15.

[0057] Furthermore, as shown in Figure 8, the AFPS may include a control signal afps_atlas_sequence_parameter_set_id that indicates the ASPS ID it references, and a control signal afps_atlas_frame_parameter_set_id that indicates its own AFPS ID.

[0058] Here, the control signal afps_atlas_sequence_parameter_set_id is a type of ASPS ID and is decoded by ue(v). However, the range of the control signal afps_atlas_sequence_parameter_set_id is limited to 0 to 15.

[0059] Furthermore, the control signal afps_atlas_frame_parameter_set_id is a type of AFPS ID and is decoded by ue(v). However, the range of the control signal afps_atlas_frame_parameter_set_id is limited to 0 to 63.

[0060] Furthermore, as shown in Figure 9, ATH may include a control signal ath_atlas_frame_parameter_set_id that indicates the AFPS ID it is referencing.

[0061] Here, the control signal ath_atlas_frame_parameter_set_id is a type of AFPS ID and is decoded by ue(v). However, the range of the control signal ath_atlas_frame_parameter_set_id is limited to 0 to 63.

[0062] Furthermore, Non-Patent Document 4 states that patch modes such as intra, inter, merge, and skip can be set for the data units of Atlas.

[0063] The merge patch data unit MMDU may include a control signal mmdu_ref_index[tileID][patchIdx] indicating the index of the reference frame, as shown in Figure 10.

[0064] The control signal mmdu_ref_index[tileID][patchIdx] is signaled only when NumRefIdxActive is greater than 1.

[0065] In other words, if NumRefIdxActive is 1, the control signal mmdu_ref_index[tileID][patchIdx] is not signaled, and the value of the control signal mmdu_ref_index[tileID][patchIdx] is set to zero.

[0066] Here, NumRefIdxActive is the maximum number of reference frames that the current atlas frame can reference. This can be calculated using control signals signaled to AFPS or ATH.

[0067] Furthermore, the interpatch data unit IMDU may include a control signal imdu_ref_index[tileID][patchIdx] indicating the index of the reference frame, as shown in Figure 11.

[0068] The control signal imdu_ref_index[tileID][patchIdx] is signaled only when NumRefIdxActive is greater than 1.

[0069] In other words, if NumRefIdxActive is 1, the control signal imdu_ref_index[tileID][patchIdx] is not signaled, and the value of the control signal imdu_ref_index[tileID][patchIdx] is set to zero.

[0070] Here, NumRefIdxActive is the maximum number of reference frames that the current atlas frame can reference.

[0071] This embodiment is expected to have the effect of optimizing the sub-bitstream of the atlas.

[0072] (Basic Mesh Bitstream Configuration) An example of the basic mesh bitstream configuration will be described below with reference to Figures 12 to 15.

[0073] Figure 12 shows an example of the configuration of a basic mesh bitstream.

[0074] As shown in Figure 12, the basic mesh bitstream may include BMSPS (Basemesh Sequence Parameter Set), BMFPS (Basemesh Frame Parameter Set), BMSH (Basemesh Submesh Header), and BMSDU (Basemesh Submesh Data Unit), which are sets of control information related to the decoding of the basic mesh.

[0075] Furthermore, as shown in Figure 13, the BMSPS may include a control signal bmsps_sequence_parameter_set_id that indicates its own BMSPS ID.

[0076] Here, the control signal bmsps_sequence_parameter_set_id is a type of BMSPS ID.

[0077] As shown in Non-Patent Document 4, when the control signal bmsps_sequence_parameter_set_id is decoded by u(4), it does not match the corresponding control signal asps_atlas_sequence_parameter_set_id of Atlas.

[0078] In this embodiment, the control signal bmsps_sequence_parameter_set_id is decoded by ue(v). Furthermore, the range of the control signal bmsps_sequence_parameter_set_id is limited to 0 to 15.

[0079] Furthermore, as shown in Figure 14, the BMFPS may include a control signal bmfps_sequence_parameter_set_id that indicates the BMSPS ID it references, and a control signal bmfps_frame_parameter_set_id that indicates its own BMFPS ID.

[0080] Here, the control signal bmfps_sequence_parameter_set_id is a type of BMSPS ID.

[0081] As shown in Non-Patent Document 4, when the control signal bmfps_sequence_parameter_set_id is decoded by u(4), it does not match the corresponding control signal afps_atlas_sequence_parameter_set_id of Atlas.

[0082] In this embodiment, the control signal bmfps_sequence_parameter_set_id is decoded using ue(v). Furthermore, the range of the control signal bmfps_sequence_parameter_set_id is limited to 0 to 15.

[0083] The control signal bmfps_frame_parameter_set_id is a type of BMFPS ID.

[0084] As shown in Non-Patent Document 4, when the control signal bmfps_frame_parameter_set_id is decoded by u(4) and the range limit of the control signal bmfps_frame_parameter_set_id is 0 to 15, the control signal bmfps_frame_parameter_set_id does not match the corresponding control signal afps_atlas_frame_parameter_set_id of Atlas.

[0085] In this embodiment, the control signal bmfps_frame_parameter_set_id is decoded by ue(v). Furthermore, the range of the control signal bmfps_frame_parameter_set_id is limited to 0 to 63.

[0086] However, as an example of a change, the control signal bmfps_frame_parameter_set_id may be changed to u(6) instead of ue(v) for the decoding method, while only the range limitation is matched with Atlas.

[0087] Furthermore, as shown in Figure 15, BMSH may include a control signal bmsh_basemesh_frame_parameter_set_id that indicates the BMFPS ID it is referencing.

[0088] Here, the control signal bmsh_basemesh_frame_parameter_set_id is a type of BMFPS ID.

[0089] As shown in Non-Patent Document 4, when the control signal bmsh_basemesh_frame_parameter_set_id is decoded by u(4) and the range limit of the control signal bmsh_basemesh_frame_parameter_set_id is 0 to 15, the control signal bmsh_basemesh_frame_parameter_set_id does not match the corresponding control signal ath_atlas_frame_parameter_set_id of Atlas.

[0090] In this embodiment, the control signal bmsh_basemesh_frame_parameter_set_id is decoded by ue(v). Furthermore, the range of the control signal bmsh_basemesh_frame_parameter_set_id is limited to 0 to 63.

[0091] According to this embodiment, by matching the decoding method and range of the basic mesh control signals with those of the atlas, consistency with the number of FPS IDs in the atlas is maintained, and it is expected that the degree of freedom and performance will be maximized without limiting the capabilities of the atlas.

[0092] (Displacement Decoding Unit 206) Figure 16 shows an example of the configuration of the displacement decoding unit 206. As shown in Figure 16, the displacement decoding unit 206 includes an arithmetic decoding unit 206A, an inverse quantization unit 206B, a frame buffer 206C, an interpretation unit 206D, and an adder 206E.

[0093] The arithmetic decoding unit 206A is configured to decode a sequence of level values ​​for each vertex from the displacement bitstream by arithmetic decoding and output it.

[0094] The inverse quantization unit 206B is configured to generate and output a coefficient sequence by inverse quantizing a level value sequence.

[0095] The frame buffer 206C is configured to acquire and store a sequence of coefficients from the inverse quantization unit 206B or the adder 206E. The frame buffer 206C is configured to output a sequence of coefficients in the reference frame of the displacement amount according to control information (not shown).

[0096] The interpretation unit 206D is configured to read a reference frame of the displacement amount indicated by the reference list of displacement amounts from the frame buffer 206C, perform interpretation using the coefficient sequence, generate a prediction coefficient sequence, and output it.

[0097] The interpretation unit 206D may directly refer to the coefficients of the corresponding frequencies in the displacement reference frame to determine the prediction coefficients for each frequency in the frame to be decoded.

[0098] The adder 206E is configured to obtain a sequence of predicted coefficients from the interpretation unit 206D, obtain a sequence of coefficients (actually a sequence of predicted residuals) from the inverse quantization unit 206B, and add them together to generate a sequence of coefficients (displacement).

[0099] The coefficient sequence generated by the adder 206E is output to the frame buffer 206C.

[0100] Furthermore, in order to fit the decoded coefficient sequence into a specified bit depth, the calculated coefficient sequence is clipped (Clip3) by the bit depth. Non-patent document 5 defines such clipping (Clip3) as follows.

[0101] Furthermore, this clipping process is performed for all coefficients. Specifically, for example, this clipping process may be carried out using the following procedure.

[0102] maxValue = 1 << dsps_displacement_3d_bit_depth_minus1 - 1 for( i = 0,vcount0=0; i <= subdivisionCount; i++ ) { vcount1 = levelOfDetailVertexCount[ i ] for( v = vcount0; v < vcount1; v++ ) { for( d = 0; d < MaxDimension; d++ ) { dispCoeffArray[ v ][ d ] = Clip3( -maxValue - 1, maxValue, dispCoeffArray[ v ][ d ])}} vcount0 = vcount1} However, dsps_displacement_3d_bit_depth_minus1 is the value obtained by subtracting 1 from the specified bit depth, subdivisionCount is the number of subdivisions, levelOfDetailVertexCount[i] is the sum of the number of coefficients from LOD=1 to i, and MaxDimension is 1 or 3.

[0103] In this embodiment, the basic mesh before subdivision is a group of vertices with LOD = 0, and when subdivision is performed once, the generated group of vertices becomes one LOD. Therefore, the number of LODs is the number of subdivisions + 1.

[0104] The displacement bitstream may include a control signal that represents the sum of the number of coefficients from LOD = 0 to i.

[0105] Alternatively, the displacement bitstream may include a control signal indicating the number of coefficients for each LOD. In that case, levelOfDetailVertexCount[i] adds the control signals from LOD=0 to i.

[0106] For example, as shown in Figure 17, if the predicted residual is (1,1,1) and the corresponding predicted value of the reference frame is (7,7,7) for a certain coefficient of AC displacement (MaxDimension = 3), then the coefficient for that frame becomes (8,8,8). If the specified bit depth is 4, then the maxValue mentioned above is 7 and the minValue is -8. Due to the clipping (Clip3) mentioned above, the coefficient sequence becomes (7,7,7).

[0107] On the other hand, if only bit shifting is performed without clipping, the result after bit shifting becomes (1,1,1) as shown in Figure 17, which is not the optimal value.

[0108] According to this embodiment, it is expected that the optimal value for AC displacement that falls within a specified bit depth can be calculated.

[0109] The displacement decoding unit 206 may also output the composition time for each displacement frame.

[0110] The configuration time variable DecDispCompTime may be calculated using the following method.

[0111] If the AC displacement bitstream is encapsulated using the sample stream format defined in Appendix D of Non-Patent Literature 4, the displacement decoding unit 206 can derive such configuration time (DecDispCompTime) using, if available, the frame order count of the displacement frame (DisplFrmOrderCntVal) or the timing information of the displacement frame.

[0112] If the AC displacement bitstream is encapsulated in an ISO media file format, such configuration time (DecDispCompTime) is expressed according to ISO / IEC DIS 23090-10.

[0113] Otherwise, such configuration time (DecDispCompTime) is provided through external means.

[0114] According to this embodiment, when using AC displacement, the displacement decoding unit 206 can be expected to calculate the composition time of the displacement frame, which is an expected benefit.

[0115] (Configuration of the displacement bitstream using arithmetic coding) Non-patent document 4 states that the displacement decoding unit 206 has two methods: one that decodes the displacement using video coding and another that decodes the displacement using arithmetic coding.

[0116] The following describes an example of the configuration of the displacement bitstream for arithmetic coding, with reference to Figures 18 to 21.

[0117] The arithmetic coding displacement bitstream may include a set of control information related to the decoding of the displacement, such as a DSPS (Displacement Sequence Parameter Set), a DFPS (Displacement Frame Parameter Set), or a DH (Displacement Header).

[0118] Furthermore, as shown in Figure 18, the DSPS may include a control signal dsps_sequence_parameter_set_id that indicates its own DSPS ID.

[0119] Here, the control signal dsps_sequence_parameter_set_id is a type of DSPS ID.

[0120] As shown in Non-Patent Document 4, when the control signal dsps_sequence_parameter_set_id is decoded by u(4), it does not match the corresponding control signal asps_atlas_sequence_parameter_set_id of Atlas.

[0121] In this embodiment, the control signal dsps_sequence_parameter_set_id is decoded by ue(v). Furthermore, the range of the control signal dsps_sequence_parameter_set_id is limited to 0 to 15.

[0122] Furthermore, as shown in Figure 19, the DFPS may include a control signal dfps_displ_sequence_parameter_set_id indicating the DSPS ID it references, and a control signal dfps_displ_frame_parameter_set_id indicating its own DFPS ID.

[0123] Here, the control signal dfps_displ_sequence_parameter_set_id is a type of DSPS ID.

[0124] As shown in Non-Patent Document 4, when the control signal dfps_displ_sequence_parameter_set_id is decoded by u(4), it does not match the corresponding control signal afps_atlas_sequence_parameter_set_id of Atlas.

[0125] In this embodiment, the control signal dfps_displ_sequence_parameter_set_id is decoded by ue(v). Furthermore, the range of the control signal dfps_displ_sequence_parameter_set_id is limited to 0 to 15.

[0126] The control signal dfps_displ_frame_parameter_set_id is a type of DFPS ID.

[0127] As shown in Non-Patent Document 4, when the control signal dfps_displ_frame_parameter_set_id is decoded by u(4) and the range limit of the control signal dfps_displ_frame_parameter_set_id is 0 to 15, the control signal dfps_displ_frame_parameter_set_id does not match the corresponding control signal afps_atlas_frame_parameter_set_id of Atlas.

[0128] In this embodiment, the control signal dfps_displ_frame_parameter_set_id is decoded by ue(v). Furthermore, the range limit of the control signal dfps_displ_frame_parameter_set_id is set to 0 to 63.

[0129] However, as an example of a change, the control signal dfps_displ_frame_parameter_set_id may be changed to u(6) instead of ue(v) for the decoding method, while only the range limitation is matched with Atlas.

[0130] Furthermore, as shown in Figure 20, DH may include a control signal dh_frame_parameter_set_id that indicates the DFPS ID it is referencing.

[0131] Here, the control signal dh_frame_parameter_set_id is a type of DFPS ID.

[0132] As shown in Non-Patent Document 4, when the control signal dh_frame_parameter_set_id is decoded by u(4) and the range limit of the control signal dh_frame_parameter_set_id is 0 to 15, the control signal dh_frame_parameter_set_id does not match the corresponding control signal ath_atlas_frame_parameter_set_id of Atlas.

[0133] In this embodiment, the control signal dh_frame_parameter_set_id is decoded by ue(v). Furthermore, the range of the control signal dh_frame_parameter_set_id is limited to 0 to 63.

[0134] However, as an example of a change, the control signal dh_frame_parameter_set_id may be modified to match the range limit of Atlas, while the decoding method may be changed from ue(v) to u(6).

[0135] According to this embodiment, by matching the decoding method and range of the displacement control signal using arithmetic coding with that of Atlas, it is expected that consistency with the number of FPS IDs in Atlas will be maintained, and the degree of freedom and performance will be maximized without limiting the capabilities of Atlas.

[0136] Furthermore, DH may include a control signal dh_num_subblock_lod_minus1[i] indicating the number of subblocks minus 1 in each LOD, as shown in Figure 21. Alternatively, the displacement data unit DDU may include a control signal indicating the number of subblocks minus 1 in each LOD. The number of subblocks in each LOD is calculated by adding 1 to such a control signal.

[0137] According to this embodiment, it is expected that the number of subblocks for each LOD can be prohibited if it is less than the minimum value of 1.

[0138] Note that DH may include a control signal dh_id indicating the ID of the sub-displacement. The decoding method for the control signal dh_id may be ue(v). The range limit for the control signal dh_id is 0 to (2NumSubdispls + di_signalled_subdispl_id_delta_length - 1).

[0139] However, NumSubdispls is the number of sub-displacements, and di_signalled_subdispl_id_delta_length is the control signal decoded from the displacement bitstream.

[0140] (Inter-decoding unit 202E) Figure 22 shows an example of the functional blocks of the inter-decoding unit 202E.

[0141] As shown in Figure 22, the inter-decoding unit 202E includes a motion vector residual decoding unit 202E1, a motion vector buffer unit 202E2, a motion vector prediction unit 202E3, a motion vector calculation unit 202E4, and an adder 202E5.

[0142] The motion vector residual decoding unit 202E1 is configured to generate an MVR (Motion Vector Residual) from the bitstream of the P-frame.

[0143] Here, MVR is the motion vector residual that shows the difference between MV (Motion Vector) and MVP (Motion Vector Prediction). MV is the difference vector (motion vector) between the coordinates of the vertex in the corresponding I-frame and the coordinates of the vertex in the P-frame. MVP is the predicted value of the MV of the target vertex (predicted value of the motion vector) using MV.

[0144] The motion vector buffer unit 202E2 is configured to sequentially save the MV output by the motion vector calculation unit 202E4.

[0145] The motion vector prediction unit 202E3 is configured to obtain decoded MVs from the motion vector buffer unit 202E2 for vertices connected to the vertex to be decoded, and to output the MVP of the vertex to be decoded using all or part of the obtained decoded MVs, as shown in Figure 23.

[0146] The motion vector calculation unit 202E4 is configured to add the MVR generated by the motion vector residual decoding unit 202E1 and the MVP output from the motion vector prediction unit 202E3, and output the MV of the vertex to be decoded.

[0147] The adder 202E5 is configured to add the coordinates of the vertices to be decoded, obtained from the decoded basic mesh of the corresponding reference frame (I-frame or P-frame), to the motion vector MV output from the motion vector calculation unit 202E3, and output the coordinates of the vertices to be decoded.

[0148] Furthermore, in order to fit the decoded vertex coordinates into a specified bit depth, the calculated vertex coordinates are clipped (Clip3) by the bit depth. Non-patent document 5 defines such clipping (Clip3) as follows.

[0149] This clipping operation is performed on the coordinates of all vertices. Specifically, for example, this clipping operation may be performed using the following procedure.

[0150] maxValue = 1 << ( bmsps_geometry_3d_bit_depth_minus1 + 1) - 1 for( v = 0; v < DecInterSubmesh.verCoordCount; v++ ) { for( k = 0; k < 3; k++ ) { DecInterMotionCodec.verCoordsArray[ v ][ k ] = Clip3( 0, maxValue, DecInterMotionCodec.verCoordsArray[ v ][ k ])}} where bmsps_geometry_3d_bit_depth_minus1 is the value obtained by subtracting 1 from the specified bit depth, and DecInterSubmesh.verCoordCount is the number of vertices in the base mesh.

[0151] For example, as shown in Figure 24, if the MV of a vertex in the basic mesh is (1,1,1) and the coordinates of the corresponding vertex in the reference frame are (7,7,7), then the vertex coordinates of that frame will be (8,8,8). If the specified bit depth is 3, then the maxValue mentioned above is 7. Due to the clipping (Clip3) mentioned above, the vertex coordinates become (7,7,7).

[0152] As shown in Figure 25, if the MV of a vertex in the basic mesh is (-1,-1,-1) and the coordinates of the corresponding vertex in the reference frame are (0,0,0), then the vertex coordinates of that frame will be (-1,-1,-1). If the specified bit depth is 3, then the aforementioned clipping (Clip3) will result in such vertex coordinates becoming (0,0,0).

[0153] On the other hand, if only bit shifting is performed without clipping, the result after bit shifting becomes (4,4,4) in the example shown in Figure 24, and (7,7,7) in the example shown in Figure 25, which are not optimal values.

[0154] According to this embodiment, it is expected that the optimal value that fits within a specified bit depth can be calculated in the basic mesh.

[0155] (Mesh buffer unit 202C) The mesh buffer unit 202C is configured to store one or more reference decoding base meshes in a predetermined order.

[0156] The basic mesh contains metadata such as frame numbers and sub-mesh numbers, as well as the coordinates of at least each vertex and the index of that vertex, and is stored in the mesh buffer unit 202C in a predetermined order determined by the reference frame list.

[0157] Here, as shown in Figure 26, the reference frame list (ref_list0) is a list of information that identifies all the reference decoded basic meshes stored in the mesh buffer unit 202C.

[0158] The reference frame list may be determined by the control signals decoded from the bitstream, as shown in Figure 26, or it may be naturally calculated from the decoding order of the frames.

[0159] The control signal decoded from the bitstream may be expressed as a relative distance to the frame being decoded, or as an absolute value of the frame index.

[0160] Furthermore, control signals may be used to utilize short-term or long-term reference frames.

[0161] For example, using a short-term reference frame, the absolute value (abs_delta_mfoc_st) and its sign (sign_flag) of the difference in display order (Display Order) between the frame to be decoded (cur) and the reference frame (ref) can be decoded from the bitstream, and the display order (Display Order) of the reference frame can be specified by the following formula. If (sign_flag) { Display Order (ref) = Display Order (cur) + abs_delta_mfoc_st} else { Display Order (ref) = Display Order (cur) - abs_delta_mfoc_st} Also, if a method is used that is naturally calculated from the decoding order of frames, for example, in the reference frame list, when no control signals are present, the frames may be arranged sequentially in fixed numbers starting from the immediately preceding decoded frame. That is, the reference frame list may be {0, -1, -2, ..., -(N-1)}.

[0162] Basically, the reference frame list does not change from frame to frame except under special circumstances (for example, when a Re-ordering instruction is received).

[0163] Furthermore, when there are multiple submeshes, the information for each submesh must be saved in an item in the reference frame list (for example, ref_list0[0]).

[0164] Non-patent document 4 states that when saving information about each submesh to an item in the reference frame list, the information is saved in a buffer using the submesh ID.

[0165] However, because the submesh ID can be freely set, it can sometimes result in a huge waste of buffer space in the reference frame.

[0166] For example, consider a basic mesh with two submeshes, where the ID of the first submesh is 10000 and the ID of the second submesh is 20000.

[0167] If we save the submesh using the submesh ID, the buffer size becomes 20,000 times the size of the submesh.

[0168] In this embodiment, the mesh buffer unit 202C stores the information of each submesh in an item in the reference frame list in index order to avoid waste.

[0169] In other words, the mesh buffer unit 202C stores only two sub-mesh information items in the reference frame list, in the order of the first and second sub-meshes.

[0170] Then, the mesh buffer unit 202C uses an index instead of a submesh ID to obtain reference submesh information from the reference frame list.

[0171] Therefore, the mesh buffer unit 202C must calculate an index from the submesh ID. As shown in Figure 27 and below, this mapping relationship BasemeshSubmeshIDToIndex is calculated.

[0172] for( i = 0; i < NumBmeshSubMeshes; i++ ) { BasemeshSubmeshIDToIndex[ bmsi_submesh_id[ i ] ] = i BasemeshSubmeshIndexToID[ i ] = bmsi_submesh_id[ i ]} Here, bmsi_submesh_id[i] is a control signal parsed from the subbitstream of the base mesh, and indicates the ID of the i-th submesh.

[0173] According to this embodiment, a reference frame list is constructed using an index instead of a submesh ID, which is expected to significantly reduce memory usage.

[0174] The mesh buffer section 202C may be updated as follows.

[0175] When the basic mesh is decoded, the mesh buffer unit 202C, in the case of I-frames and P-frames, deletes one or more existing reference frames in a predetermined order determined by the reference frame list, inserts one or more basic meshes including the basic mesh of the decoded frame, or creates and inserts one basic mesh from multiple basic meshes, thereby adjusting the order of the reference frames.

[0176] Such deletion operations may be performed only when the mesh buffer unit 202C is full. The number of basic meshes that can be stored in the mesh buffer unit 202C is predetermined. In this embodiment, the mesh buffer unit 202C is defined as being full when the number of basic meshes is reached.

[0177] In the creation process described above, the coordinates of the vertices corresponding to the decoded frame's base mesh and the existing base mesh stored in the mesh buffer section 202C may be weighted and averaged to form a single base mesh.

[0178] The weights used in such a weighted average may be predetermined, calculated using frame indices, or decoded from control signals.

[0179] Furthermore, when the mesh buffer unit 202C receives a control signal indicating a Re-ordering instruction based on a control signal decoded from the bitstream, it updates the reference frame list and adjusts the order of the reference frames according to a predetermined order determined by the updated reference frame list (ref_list0).

[0180] Furthermore, when a submesh exists as defined in Non-Patent Document 4 above, all submeshes will either be given the same control signal (smh_mesh_frm_order_cnt_lsb) or the control signal (smh_mesh_frm_order_cnt_lsb) will be applied to all submeshes.

[0181] The value indicated by such a control signal (smh_mesh_frm_order_cnt_lsb) may be the difference from the display order of the frames to be decoded, or it may be the order within a predetermined frame set MaxMeshFrmOrderCntLsb.

[0182] Furthermore, if the decoding order and the display order are different, and the decoded basic meshes are arranged in the decoding order, the basic mesh decoding unit 202 may rearrange the decoded basic meshes to the display order.

[0183] Furthermore, in order to achieve Temporary scalability, control signals are defined for each frame to indicate whether or not to decode the basic mesh, displacement amount, and texture, and these are decoded from the bitstream accordingly.

[0184] Furthermore, the Temporal_IDs of the atlas and the base mesh may be matched within the same frame. Similarly, the Temporal_IDs of the atlas and the textures may be matched within the same frame. Finally, the Temporal_IDs of the atlas and the displacement amount may be matched within the same frame.

[0185] This configuration is expected to have the effect of avoiding situations where frames cannot be decoded and unnecessary data is not included.

[0186] Furthermore, it is desirable that the interval between adjacent frames with the same Temporal_ID remains constant.

[0187] For adjacent frames with the same Temporal_ID, the POC is the closest.

[0188] As mentioned above, by keeping the frame interval constant, it is possible to maintain a constant frame rate when displaying the decoded frames.

[0189] Furthermore, the decoding order of atlases and base meshes with the same display order may be matched. The decoding order of atlases and displacement values ​​with the same display order may also be matched. Additionally, the decoding order of atlases and textures with the same display order may be matched.

[0190] Alternatively, the random access points of atlases and basic meshes with the same display order may be matched. Furthermore, the random access points of atlases and displacements with the same display order may be matched. Also, the random access points of atlases and textures with the same display order may be matched. Note that random access points are defined in Non-Patent Document 4 or Non-Patent Document 5.

[0191] With this configuration, it is expected that the mesh can be reconstructed without waiting for the decoding of the basic mesh, displacement, and texture to be completed.

[0192] Furthermore, frames with a Temporary_ID higher than the control signal Temporary_ID of the frame to be decoded will not be used as reference frames for that frame.

[0193] This is expected to have the effect of eliminating the possibility of reference frames being discarded.

[0194] The atlas, base mesh, displacement, and texture bitstreams are encapsulated by a Network Abstraction Layer (NAL) unit. The NAL unit may have an NAL header as shown in Figure 28.

[0195] The TID, defined as the last three bits in the NAL header, is Temporary_ID plus 1. The TID ranges from 1 to 7, and zero is prohibited. Therefore, Temporary_ID is calculated by subtracting 1 from the TID.

[0196] The LayerID / R6, defined as the six bits immediately preceding the TID in the NAL header, specifies the identifier of the layer to which the NAL unit belongs.

[0197] The LayerID / R6 value must be within the range of 0 to 62. The value 63 may be designated by ISO / IEC in the future.

[0198] Aside from determining the amount of data in the bitstream's decode unit, the mesh decoder 200 ignores all data following the value 63 in the NAL unit, and a mesh decoder 200 conforming to a specified profile ignores all NAL units where the LayerID-R6 value is not 0 (i.e., removes and discards them from the bitstream).

[0199] The LayerID / R6 value 63 can be used in future extensions to indicate an extended layer identifier.

[0200] Furthermore, in the AC displacement bitstream of Non-Patent Document 4, the LayerID / R6 value must be the same for all DCL NAL units in the encoded displacement frame.

[0201] The LayerID / R6 value of the encoded displacement frame is set to the LayerID / R6 value of the DCL NAL unit of that encoded displacement frame.

[0202] For example, it is determined whether the LayerID / R6 values ​​are the same for all DCL NAl units in the encoded displacement frame.

[0203] If they are the same, the processing of the displacement decoding unit 206 is executed. On the other hand, if they are not the same, the processing of the displacement decoding unit 206 is not executed, and the displacement decoding unit 206 may output empty data or stop.

[0204] Furthermore, in the AC displacement bitstream of Non-Patent Document 4, if NALType is equal to DNAL_DEOB, then LayerID / R6 must be 0.

[0205] According to this embodiment, it is expected that all sub-displacement amounts will be decoded when decoding the AC displacement amount, thereby enabling the decoding of the complete frame.

[0206] Furthermore, if a submesh defined in Non-Patent Document 4 exists, all submeshes shall be given the same TID, or the TID shall be applied to all submeshes.

[0207] The following describes an example of achieving Temporary scalability using the aforementioned Temporary_ID.

[0208] Regarding the atlas, Non-Patent Document 5 can be used, and for displacement and texture, the video encoding methods HEVC and VVC can be used, so the basic mesh will be described below.

[0209] As shown in Figure 29, the BMSPS of the basic mesh bitstream may include a control signal bmsps_max_sub_layers_minus1 in u(3) that indicates the maximum number of Temporal sublayers.

[0210] Furthermore, each Temporal sublayer may include a control signal bmsps_max_dec_mesh_frame_buffing_minus1 indicating the buffer size of the largest basic mesh, and a control signal bmsps_max_num_reorder_frames indicating the difference from the display order of the largest decoded frame.

[0211] The LayerID / R6 values ​​of all BMCL NAL units in the encoded base mesh frame must be the same. The LayerID / R6 value of the encoded base mesh frame is the LayerID / R6 value of the BMCL NAL unit in the encoded base mesh frame.

[0212] If NALType is equal to NAL_EOB, then the value of LayerID / R6 must be equal to 0.

[0213] If NALType falls within the range from NAL_BLA_W_LP to NAL_RSV_BMCL_29 as defined in Non-Patent Document 4, that is, if it belongs to an IRAP-coded basic mesh frame, then Temporal_ID must be 0.

[0214] If NALType is equal to NAL_TSA_R or NAL_TSA_N, then Temporal_ID must not be equal to 0.

[0215] If NALType is equal to 0, and NALType is equal to NAL_STSA_R or NAL_STSA_N, then Temporal_ID must not be equal to 0.

[0216] The value of Temporal_ID must be the same for all BMCL NAL units within the access unit.

[0217] The value of the Temporary_ID of the coded basic mesh frame or access unit is the value of the Temporary_ID of the BMCL NAL unit of the coded basic mesh frame or access unit.

[0218] The Temporal_ID value of the sublayer representation is the maximum value of the Temporal_IDs of all BMCL NAL units within the sublayer representation.

[0219] The Temporal_ID value of a non-BMCL NAL unit is restricted as follows: - If NALType is equal to NAL_BMSPS, the Temporal_ID must be 0, and the Temporal_ID of the access unit containing the NAL unit must be 0. - Otherwise, if NALType is equal to NAL_EOS or NAL_EOB, the Temporal_ID must be 0. - Otherwise, if NALType is equal to NAL_AUD or NAL_FD, the Temporal_ID must be equal to the Temporal_ID of the access unit containing the NALL unit. - Otherwise, the Temporal_ID must be greater than or equal to the Temporal_ID of the access unit containing the NAL unit.

[0220] If the NAL unit is not BMCL, the Temporal_ID value will be equal to the minimum Temporal_ID value of all access units to which the non-BMCL NAL unit applies.

[0221] If NALType is equal to NAL_BMFPS, then Temporal_ID can be greater than or equal to the Temporal_ID of the included access unit, since the entire set of basic mesh frame parameters (BMFPS) is included at the beginning of the bitstream where the Temporal_ID of the first encoded basic mesh frame is 0.

[0222] The skip decoding unit 202F will refer to the specified tIDTarget and discard any NAL units whose Temporal_ID is higher than tIDTarget without decoding them.

[0223] Here, tIDTarget may be specified by a predetermined value, or it may be specified by the network conditions or the terminal capabilities of the mesh decoding device 200.

[0224] For example, a lower tIDTarget is specified for wireless connections than for wired connections. Also, a lower tIDTarget is specified when the network conditions are poor. Furthermore, a lower tIDTarget is specified when decoding is performed by a low-spec mesh decoding device 200.

[0225] However, a requirement for bitstream compliance is that at least one NAL unit with a Temporal_ID not higher than tIDTarget must exist in the bitstream.

[0226] The following describes an example of a modification that enables Temporary scalability using the aforementioned Temporary_ID.

[0227] The bitstreams of the atlas, base mesh, displacement, and texture are encapsulated by a Network Abstraction Layer (NAL) unit. The NAL unit may have an NAL header as shown in Figure 28.

[0228] The TID, defined as the last three bits in the NAL header, is Temporary_ID plus 1. The TID ranges from 1 to 7, and zero is prohibited.

[0229] The LayerID / R6, defined as the six bits immediately preceding the TID in the NAL header, specifies the identifier of the layer to which the NAL unit belongs.

[0230] The LayerID / R6 value must be within the range of 0 to 62. The value 63 may be designated by ISO / IEC in the future.

[0231] Aside from determining the amount of data in the bitstream's decode unit, the mesh decoder 200 ignores all data following the value 63 in the NAL unit, and a mesh decoder 200 conforming to a specified profile ignores all NAL units where the LayerID-R6 value is not 0 (i.e., removes and discards them from the bitstream).

[0232] The LayerID / R6 value 63 can be used in future extensions to indicate an extended layer identifier.

[0233] Furthermore, if there are submeshes of the basic mesh as defined in Non-Patent Document 4, all submeshes shall be given the same TID, or the TID shall be applied to all submeshes.

[0234] The bitstreams for the atlas, base mesh, displacement, and texture may each have their own independently set TID. For example, for the atlas, the TID is fixed to zero using Annex A of Non-Patent Document 5. Thus, the base mesh, displacement, and texture each have their own independently set TID.

[0235] In other words, depending on the content, at least one of the base mesh, displacement, and texture may have its Temporal_ID fixed to zero. In that case, LD settings can also be used. An example is shown in Table 1.

[0236]

[0237] Even if each bitstream is set independently, the displacement amount and texture will utilize the HEVC or VVC video encoding scheme, while the basic mesh will utilize the embodiment described above.

[0238] If set independently for each bitstream, the system will refer to the tIDTarget specified in each bitstream and discard NAL units whose TID is higher than tIDTarget without decoding them.

[0239] Here, tIDTarget may be specified by a predetermined value, or it may be specified by the network conditions or the terminal capabilities of the mesh decoding device 200.

[0240] For example, a lower tIDTarget is specified for wireless connections than for wired connections. Also, a lower tIDTarget is specified when the network conditions are poor. Furthermore, a lower tIDTarget is specified when decoding is performed by a low-spec mesh decoding device 200.

[0241] However, a requirement for bitstream compatibility is that at least one NAL unit must exist in the bitstream whose TID is not higher than tIDTarget.

[0242] On the other hand, if each bitstream is set independently, then if even one element—the base mesh, displacement, or texture—is discarded in a particular frame, the others will also be discarded.

[0243] Alternatively, if each bitstream is set independently, if the base mesh is discarded in a particular frame, the displacement and textures will also be discarded, and the Reconstruction process will not be performed. However, if the displacement is discarded, all displacement values ​​will be set to zero, and the Reconstruction process will be performed. Similarly, if the texture is discarded, all texture values ​​will be set to zero, and the Reconstruction process will be performed.

[0244] Note that the number of submeshes may differ in each frame (intra frame, interframe, and skip frame).

[0245] In such cases, the intra decoding unit 202B, the inter decoding unit 202E, and the skip decoding unit 202F assign a unique submesh ID to each submesh in each frame.

[0246] Furthermore, the intra-decoding unit 202B and the inter-decoding unit 202E may assign different SubmeshIDs (submesh IDs) to corresponding submeshes between frames.

[0247] However, the interdecoding unit 202E is limited to referencing only submeshes that have the same SubmeshID within the reference frame.

[0248] Alternatively, the inter-decoding unit 202E may only reference submeshes that have the same number of vertices within the reference frame.

[0249] Alternatively, the intra-decoding unit 202B and the inter-decoding unit 202E may refer to a submesh specified in the reference frame.

[0250] In such cases, if there are multiple submeshes in the reference frame, the inter-decoding unit 202E may decode a control signal from the bitstream of the current submesh that specifies the SubmeshID of the referenceable submesh.

[0251] On the other hand, if there is only one submesh in the reference frame, the inter-decoding unit 202E may treat that submesh as a referenceable submesh.

[0252] However, if the above-mentioned control signals are not present, the inter-decoding unit 202E or the skip decoding unit 202F sets the SubmeshID of the referable submesh to the same SubmeshID as the submesh in the current frame.

[0253] Furthermore, the inter-decoding unit 202E may decode a control signal from the bitstream indicating whether or not the above-mentioned control signal exists.

[0254] The inter-decoding unit 202E may also decode a control signal that selects the method for determining the above-mentioned referable submesh.

[0255] The subdivision section 203 and the displacement decoding section 206 may conform to Non-Patent Document 4.

[0256] To summarize, the mesh decoding device 200 according to this embodiment includes a displacement amount decoding unit 206 that decodes a displacement amount bitstream to generate and output a displacement amount, and the displacement amount decoding unit 206 determines whether the LayerID / R6 values ​​are the same for all DCL NAl units of the displacement amount frame included in the displacement amount bitstream.

[0257] If the determination results are the same, the processing of the displacement amount decoding unit 206 is executed. If the determination results are not the same, the processing of the displacement amount decoding unit 206 is not executed, and the displacement amount decoding unit 206 may output empty data or stop.

[0258] Furthermore, the mesh decoding device 200 according to this embodiment includes a basic mesh decoding unit 202 that decodes a basic mesh bitstream to generate and output a basic mesh, and the basic mesh decoding unit 202 clips the coordinates of the decoded vertices by a specified bit depth so that the coordinates of the vertices fit within that bit depth.

[0259] Furthermore, the mesh decoding device 200 according to this embodiment includes a displacement amount decoding unit 206 that decodes a displacement amount bitstream to generate and output a displacement amount, and the displacement amount decoding unit 206 clips the decoded coefficient sequence by a predetermined bit depth so that the sequence fits within that bit depth.

[0260] Here, the basic mesh decoding unit 202 may perform the above-mentioned clipping on the coordinates of all vertices, and set the maximum and minimum values ​​of the vertex coordinates to specific values.

[0261] Furthermore, the displacement decoding unit 206 may perform the clipping described above for all coefficients and set the maximum and minimum values ​​of the coefficients to specific values.

[0262] Furthermore, the mesh decoding device 200 according to this embodiment includes an atlas data decoding unit 208 that decodes an atlas bitstream to generate and output an atlas containing control information, and the atlas data decoding unit 208 can set the patch mode to intra, inter, merge, and skip for the patch data units of the atlas described above.

[0263] Here, a patch data unit set to merge may not signal mmdu_ref_index[tileID][patchIdx], which is a control signal indicating the index of such a reference frame, if NumRefIdxActive, a variable indicating the maximum number of reference frames that the current atlas frame can reference, is 1, and the value of mmdu_ref_index[tileID][patchIdx] is set to zero, and may include mmdu_ref_index[tileID][patchIdx] if NumRefIdxActive is 1 or greater.

[0264] Furthermore, if NumRefIdxActive, a variable indicating the maximum number of reference frames that the current atlas frame can reference, is 1, the patch data unit configured on the interface may not signal imdu_ref_index[tileID][patchIdx], a control signal indicating the index of such reference frame, and set the value of imdu_ref_index[tileID][patchIdx] to zero. If NumRefIdxActive is 1 or greater, it may include imdu_ref_index[tileID][patchIdx].

[0265] Furthermore, the atlas data decoding unit 208 may, at the beginning of the decoding described above, initialize PredictorIdx, a variable that indicates the index of the reference frame of the patch immediately preceding the current patch and uses the index of the reference frame of the current patch as a predicted value, and may set the initial value of PredictorIdx to zero.

[0266] Furthermore, the mesh decoding device 200 according to this embodiment includes a displacement amount decoding unit 206 that decodes a displacement amount bitstream to generate and output displacement amounts. The displacement amount bitstream includes a control signal for each sub-displacement amount that indicates the number of sub-blocks in each LOD minus 1. The displacement amount decoding unit 206 calculates the number of sub-blocks in each LOD by adding 1 to the control signal included in the displacement amount bitstream.

[0267] Here, if the displacement bitstream is encapsulated using the sample stream format defined in Appendix D of Non-Patent Literature 4, the displacement decoding unit 206 derives the configuration time for each displacement frame using the frame order count of the displacement frame and the timing information of the displacement frame; if the displacement bitstream is encapsulated in the ISO media file format, the displacement decoding unit 206 derives the configuration time for each displacement frame according to ISO / IEC DIS 23090-10; otherwise, the displacement decoding unit 206 may obtain the configuration time for each displacement frame through external means.

[0268] Furthermore, in the displacement bitstream, the TID defined as the last three bits in the NAL header is Temporal_ID + 1, the range of TID is 1 to 7, and Temporal_ID may be calculated by subtracting 1 from TID.

[0269] The mesh coding device 100 and mesh decoding device 200 described above may be implemented as programs that cause a computer to execute each function (each process).

[0270] Furthermore, according to this embodiment, for example, it is possible to achieve an overall improvement in service quality in video communication, thereby contributing to Goal 9 of the United Nations-led Sustainable Development Goals (SDGs), "Build resilient infrastructure, promote sustainable industrialization and expand innovation."

[0271] 1...Mesh processing system 100...Mesh coding device 200...Mesh decoding device 201...Multiplex separation unit 202...Basic mesh decoding unit 202A...Separation unit 202B...Intra decoding unit 202B1...Arbitrary intra decoding unit 202B2...Alignment unit 202C...Mesh buffer unit 202D...Connection information decoding unit 202E...Inter decoding unit 202E1...Motion vector residue decoding unit 202E2...Motion vector buffer unit 202E3...Motion vector prediction unit 202E4...Motion vector calculation unit 202E5...Adder 203...Subdivision unit 204...Mesh decoding unit 205...Patch integration unit 206...Displacement amount decoding unit 206A...Arithmetic decoding unit 206B...Inverse quantization unit 206C...Frame buffer 206D...Inter prediction unit 206E...Adder 207...Video decoding unit 208...Atlas data decoding unit

Claims

1. A mesh decoding device comprising a basic mesh decoding unit that decodes a basic mesh bitstream to generate and output a basic mesh, wherein the basic mesh decoding unit clips the coordinates of the decoded vertices by the bit depth so that the coordinates of the vertices fit within a specified bit depth.

2. A mesh decoding device comprising a displacement amount decoding unit that decodes a displacement amount bitstream to generate and output a displacement amount, wherein the displacement amount decoding unit clips the decoded coefficient sequence by the bit depth so that the coefficient sequence fits within a predetermined bit depth.

3. The mesh decoding device according to claim 1, characterized in that the basic mesh decoding unit performs the clipping on the coordinates of all vertices and sets the maximum and minimum values ​​of the vertex coordinates to specific values.

4. The mesh decoding device according to claim 2, characterized in that the displacement decoding unit performs the clipping for all coefficients and sets the maximum and minimum values ​​of the coefficients to specific values.

5. A mesh decoding method comprising the step of decoding a basic mesh bitstream to generate and output a basic mesh, wherein in the step, the coordinates of the decoded vertices are clipped by a specified bit depth so that the coordinates of the vertices are contained within a specified bit depth.

6. A program that causes a computer to function as a mesh decoding device, wherein the mesh decoding device comprises a basic mesh decoding unit that decodes a basic mesh bitstream to generate and output a basic mesh, and the basic mesh decoding unit is characterized in that it clips the coordinates of the decoded vertices by the bit depth so that the coordinates of the vertices are contained within a specified bit depth.