Mesh decoding device, mesh decoding method, and program

The mesh decoding apparatus and method address the issue of specifying sub-meshes in inter-prediction coding by using intra and inter decoding units with varying sub-mesh counts, enhancing decoding accuracy and efficiency.

WO2025146751A1PCT designated stage expired Publication Date: 2025-07-10KDDI CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/041297
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-06
Filing Date
2024-11-21
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing mesh decoding technologies lack the ability to specify which sub-mesh to refer to among reference frames in inter-prediction coding, especially when information is absent.

Method used

A mesh decoding apparatus and method that includes an intra decoding unit for decoding vertex coordinates and connection information from an intra frame, and an inter decoding unit that uses a motion vector to specify the sub-mesh by adding decoded coordinates from an inter frame, with the number of sub-meshes varying for each frame type.

Benefits of technology

Enables accurate and efficient mesh decoding by specifying the sub-mesh to refer to, even in the absence of explicit information, thereby improving coding efficiency and reducing computational and memory requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024041297_10072025_PF_FP_ABST
    Figure JP2024041297_10072025_PF_FP_ABST
Patent Text Reader

Abstract

The purpose of the present invention is to specify which submesh to refer to. In a mesh decoding device 200 according to the present invention, each intra-frame has a different number of submeshes, and each inter-frame has a different number of submeshes.
Need to check novelty before this filing date? Find Prior Art

Description

Mesh decoding device, mesh decoding method and program

[0001] The present invention relates to a mesh decoding device, a mesh decoding method, and a program.

[0002] Non-Patent Document 1 or Non-Patent Document 4 discloses a technique for encoding a mesh using Non-Patent Document 2 or 3 according to the framework of Non-Patent Document 5.

[0003] Khaled Mammou, Jungsun Kim, Alexis M Tourapis, Dimitri Podborski, and Krasimir Kolarov, “[V-CG] Apple's Dynamic Mesh Coding CfP Response,” April 2022, ISO / IEC JTC 1 / SC 29 / WG 7 m59281.Google Draco, accessed May 26, 2022 [Online], https: / / google.github.io / dracoJean-Eudes Marvie, Olivier Mocquard, “[V-DMC][EE4.4-related] An efficient EdgeBreaker implementation,” April 2023, ISO / IEC JTC 1 / SC 29 / WG 7 m63344. “WD 5.0 ​​of V-DMC,” Oct. 2023, ISO / IEC JTC 1 / SC 29 / WG 7 N00744. “Information technology - Coded Representation of Immersive Media - Part 5: Visual Volumetric Video-based Coding (V3C) and Video-based Point Cloud Compression (V-PCC),” ISO / IEC JTC 1 / SC 29 / WG 7, ISO / IEC 23090-5:2021(2E).

[0004] However, in the prior art, there is a problem that there is no information on which sub-mesh to refer to in the reference frame of the basic mesh of the current frame in inter-prediction coding. Therefore, the present invention has been made in consideration of the above-mentioned problem, and aims to provide a mesh decoding device, a mesh decoding method and a program that can specify which sub-mesh to refer to by introducing a control signal indicating which sub-mesh to refer to in the reference frame of the current frame in inter-prediction coding.

[0005] Another object of the present invention is to provide a mesh decoding device, a mesh decoding method, and a program that can identify which sub-mesh to reference in inter-prediction coding even if there is no information regarding which sub-mesh to reference in the reference frame of the current frame.

[0006] A first feature of the present invention is a mesh decoding device comprising an intra decoding unit that decodes coordinates and connection information of vertices in an intra frame from a bit stream of the intra frame, and an inter decoding unit that decodes coordinates of the vertex to be decoded by adding a motion vector decoded from a bit stream of an inter frame to coordinates of a vertex corresponding to the vertex to be decoded in a reference frame, wherein the number of sub-meshes is different for each of the intra frames and the number of sub-meshes is different for each of the inter frames.

[0007] A second feature of the present invention is a mesh decoding method comprising a step A of decoding coordinates and connection information of vertices in an intra-frame from a bit stream of the intra-frame, and a step B of decoding coordinates of the vertex to be decoded by adding a motion vector decoded from a bit stream of an inter-frame to the coordinates of a vertex corresponding to the vertex to be decoded in a reference frame, wherein in steps A and B, the number of sub-meshes is different for each of the intra-frames and the number of sub-meshes is different for each of the inter-frames.

[0008] A third feature of the present invention is a program that causes a computer to function as a mesh decoding device, the mesh decoding device comprising an intra decoding unit that decodes coordinates and connection information of vertices in an intra frame from a bit stream of the intra frame, and an inter decoding unit that decodes coordinates of the vertex to be decoded by adding a motion vector decoded from a bit stream of an inter frame to coordinates of a vertex corresponding to the vertex to be decoded in a reference frame, wherein the number of sub-meshes is different for each of the intra frames and the number of sub-meshes is different for each of the inter frames.

[0009] According to the present invention, in inter-prediction coding, a mesh decoding device, a mesh decoding method and a program can be provided that can identify which sub-mesh to refer to by introducing a control signal indicating which sub-mesh to refer to in the reference frame of the current frame.

[0010] In addition, according to the present invention, it is possible to provide a mesh decoding device, a mesh decoding method and a program that can identify which sub-mesh to refer to in inter-prediction coding even if there is no information regarding which sub-mesh to refer to in the reference frame of the current frame.

[0011] FIG. 1 is a diagram showing an example of the configuration of a mesh processing system 1 according to an embodiment. FIG. 2 is a diagram showing an example of functional blocks of a mesh decoding device 200 according to an embodiment. FIG. 3A is a diagram showing an example of a base mesh and a subdivision mesh. FIG. 3B is a diagram showing an example of a base mesh and a subdivision mesh. FIG. 4 is a diagram showing an example of functional blocks of a basic mesh decoding unit 202 of a mesh decoding device 200 according to an embodiment. FIG. 5 is a diagram showing an example of functional blocks of an intra decoding unit 202B of the basic mesh decoding unit 202 of the mesh decoding device 200 according to an embodiment. FIG. 6 is a diagram showing an example of a correspondence relationship between vertices of a basic mesh of a P frame and vertices of a basic mesh of an I frame. FIG. 7 is a diagram showing an example of functional blocks of an inter decoding unit 202E of the basic mesh decoding unit 202 of the mesh decoding device 200 according to an embodiment. FIG. 8 is a diagram for explaining an example of a method for calculating the MVP of a vertex to be decoded by the motion vector prediction unit 202E3 of the inter decoding unit 202E of the basic mesh decoding unit 202 of the mesh decoding device 200 according to an embodiment. FIG. 9 shows a flowchart illustrating an example of the operation of the motion vector prediction unit 202E3 of the inter decoding unit 202E of the basic mesh decoding unit 202 of the mesh decoding device 200 according to an embodiment. FIG. 10A shows an example of the decoding order for a mesh. FIG. 10B shows an example of a list of vertices around a vertex to be decoded. FIG. 11 shows an example of statistical data indicating the relationship between the number of decoded motion vectors and the number of vertices around the vertex to be decoded. FIG. 12 is a diagram for explaining an example of the worst case. FIG. 13 is a diagram for explaining a second modification of the inter decoding unit 202E of the basic mesh decoding unit 202 of the mesh decoding device 200 according to an embodiment. FIG. 14 is a diagram for explaining a second modification of the inter decoding unit 202E of the basic mesh decoding unit 202 of the mesh decoding device 200 according to an embodiment. FIG. 15 is a diagram for explaining a third modification of the inter decoding unit 202E of the basic mesh decoding unit 202 of the mesh decoding device 200 according to an embodiment. FIG. 16 is a diagram showing a modification of the functional blocks of Modification 1 of the basic mesh decoding unit 202 of the mesh decoding device 200 according to an embodiment.Fig. 17 is a diagram illustrating a first modification of the basic mesh decoding unit 202 of the mesh decoding device 200 according to an embodiment. Fig. 18 is a diagram illustrating a mesh buffer unit 202C of the basic mesh decoding unit 202 of the mesh decoding device 200 according to an embodiment. Fig. 19 is a diagram illustrating the mesh buffer unit 202C of the basic mesh decoding unit 202 of the mesh decoding device 200 according to an embodiment. Fig. 20 is a diagram illustrating a modification of the basic mesh decoding unit 202 of the mesh decoding device 200 according to a second modification. Fig. 21 is a diagram illustrating a modification of the basic mesh decoding unit 202 of the mesh decoding device 200 according to the second modification. Fig. 22 is a diagram illustrating an example of a NAL header. Fig. 23 is a diagram illustrating an example of a case where the number of submeshes varies between frames. Fig. 24 is a diagram illustrating an example of a case where corresponding submeshes have different SubmeshIDs between frames.

[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the components in the following embodiments can be appropriately replaced with existing components, etc., and various variations, including combinations with other existing components, are possible. Therefore, the description of the following embodiments does not limit the content of the invention described in the claims.

[0013] First Embodiment A mesh processing system according to this embodiment will be described below with reference to FIGS.

[0014] 1 is a diagram showing an example of the configuration of a mesh processing system 1 according to this embodiment. As shown in FIG. 1, the mesh processing system 1 includes a mesh encoding device 100 and a mesh decoding device 200.

[0015] FIG. 2 is a diagram showing an example of functional blocks of a mesh decoding device 200 according to this embodiment.

[0016] As shown in Figure 2, the mesh decoding device 200 has a demultiplexing unit 201, a basic mesh decoding unit 202, a subdivision unit 203, a mesh decoding unit 204, a patch integration unit 205, a displacement amount decoding unit 206, a video decoding unit 207, and an atlas data decoding unit 208.

[0017] Here, the basic mesh decoding unit 202, the subdivision unit 203, the mesh decoding unit 204 and the displacement amount decoding unit 206 are configured to perform processing in units of patches into which the mesh is divided, and the results of these processes may then be integrated by the patch integration unit 205.

[0018] In the example of FIG. 3A, the mesh is divided into patch 1 consisting of basic faces 1 and 2 and patch 2 consisting of basic faces 3 and 4.

[0019] The demultiplexing unit 201 is configured to separate the multiplexed bit stream into a base mesh bit stream, a displacement bit stream, a texture bit stream, and an atlas bit stream.

[0020] The atlas data decoder 208 is configured to decode the atlas bitstream and output control information, which may be used as metadata by the base mesh decoder 202, the subdivision unit 203, the mesh decoder 204, the displacement decoder 206, and the video decoder 207.

[0021] <Basic Mesh Decoding Unit 202> The basic mesh decoding unit 202 is configured to decode the basic mesh bitstream, generate a basic mesh, and output it.

[0022] Here, the basic mesh is made up of a plurality of vertices in a three-dimensional space and edges connecting these vertices.

[0023] As shown in FIG. 3A, the basic mesh is formed by combining basic faces each represented by three vertices.

[0024] The base mesh decoder 202 may be configured to decode the base mesh bitstream using, for example, the techniques described in Draco in Non-Patent Document 2 or in Non-Patent Document 3.

[0025] The basic mesh decoding unit 202 may also be configured to generate "subdivision_method_id" (described later) as control information for controlling the type of subdivision method.

[0026] As shown in FIG. 4, the basic mesh decoding unit 202 includes a separating unit 202A, an intra decoding unit 202B, a mesh buffer unit 202C, a connection information decoding unit 202D, and an inter decoding unit 202E.

[0027] The separator 202A is configured to classify the basic mesh bitstream into a bitstream of I frames and a bitstream of P frames.

[0028] (Intra decoding unit 202B) The intra decoding unit 202B is configured to decode the coordinates and connection information of the vertices of the I frame from the bit stream of the I frame, for example, by using Draco shown in Non-Patent Document 2 or the technology described in Non-Patent Document 3.

[0029] FIG. 5 is a diagram showing an example of functional blocks of the intra decoder 202B.

[0030] As shown in FIG. 5, the intra decoding unit 202B includes an arbitrary intra decoding unit 202B1 and an alignment unit 202B2.

[0031] The optional intra decoder 202B1 is configured to decode the coordinates and connectivity information of the unordered vertices of the I-frame from the bitstream of the I-frame using any method including Draco shown in Non-Patent Document 2 or the technology described in Non-Patent Document 3.

[0032] The sorting unit 202B2 is configured to output the vertices by sorting the unordered vertices into a predetermined order.

[0033] The predetermined order may be, for example, a Morton code order or a raster scan order.

[0034] Furthermore, the alignment unit 202B2 may combine overlapping vertices, which are multiple vertices with the same coordinates in the decoded basic mesh, into a single vertex, and then rearrange them in a predetermined order.

[0035] The mesh buffer unit 202C is configured to store the coordinates and connection information of the vertices of the I frame decoded by the intra decoding unit 202B. Here, a specific buffer may be provided to store pairs of vertex indices A(k) and B(k) of overlapping vertices in a predetermined order.

[0036] The connection information decoding unit 202D is configured to convert the connection information of the I frame or reference frame extracted from the mesh buffer unit 202C into connection information of the P frame.

[0037] The inter-decoding unit 202E is configured to decode the coordinates of the vertices of the P frame by adding the coordinates of the vertices of the reference frame retrieved from the mesh buffer unit 202C and the motion vectors decoded from the bitstream of the P frame.

[0038] Furthermore, the inter-decoding unit 202E can adjust the index of the vertex of the P frame using the pair of vertex indices A(k) and B(k) that exist as overlapping vertices stored in the specific buffer.

[0039] Here, all or part of the indexes are decoded from the bitstream. The decoding method may be arithmetic coding. As a result, it is expected that there is no limit to the maximum value of the index to be decoded using arithmetic coding.

[0040] For example, arithmetic coding called ue(v) may be used, where ue(v) represents leftmost bit-first unsigned integer zeroth-order exponential-Golomb coding (Exp-Golomb).

[0041] Specifically, the parsing process for the syntax element of ue(v) begins from the current position in the bitstream, reading the bits containing the first non-zero bit, and counting the number of leading bits equal to 0. This process is specified as follows: leadingZeroBits=-1 for(b=0;!b;leadingZeroBits++ b=read_bits(l)) Next, the variable codeNum is assigned as follows:

[0042] codeNum = 2 leadingZeroBits - 1 + read_bits (leadingZeroBits) where the value returned by read_bits (leadingZeroBits) is interpreted as a binary representation of an unsigned integer with the most significant bit written first. Also, the value of ue(v) is equal to the value of codeNum.

[0043] Table 1 shows the structure of an Exp-Golomb code, separating the bit string into "prefix" bits and "suffix" bits.

[0044] Here, the "prefix" bits are the bits that are parsed as specified in the calculation of leadingZeroBits and are represented as 0 or 1 in the Bit String column of Table 1.

[0045] The "suffix" bits are the bits that are parsed in the calculation of codeNum, and are denoted by x in Table 1. i where i ranges from 0 to leadingZeroBits-1. Each x i is equal to either 0 or 1.

[0046] Table 2 shows how to explicitly assign bit strings to values ​​of codeNum, where the value of ue(v) is equal to the value of codeNum.

[0047]

[0048] In this embodiment, there is a correspondence between the vertices of the base mesh of the P frame and the vertices of the base mesh of the reference frame (I frame or P frame), as shown in Figure 6. Here, the motion vector decoded by the inter decoding unit 202E is a difference vector between the coordinates of the vertices of the base mesh of the P frame and the coordinates of the vertices of the base mesh of the I frame.

[0049] The inter-decoding unit 202E may decode the number of vertices of the current frame or the current submesh from the bitstream.

[0050] Here, when the inter-decoding unit 202E decodes the number of vertices of the current frame or current submesh from the above-mentioned bitstream, the compatibility requirement of such a bitstream is that the number of vertices of the decoded current frame or current submesh must be equal to the number of vertices of the reference frame or reference submesh.

[0051] In addition, when decoding the number of vertices of the current frame or current submesh from the above-mentioned bitstream, if the number of vertices of the decoded current frame or current submesh differs from the number of vertices of the reference frame or reference submesh, the inter-decoding unit 202E is configured to preferentially use the number of vertices of the current frame or current submesh decoded from the bitstream.

[0052] Furthermore, the inter decoding unit 202E may use the number of vertices of the reference frame or reference submesh as the number of vertices of the current frame or current submesh.

[0053] In such a case, if the number of vertices in the current frame or current submesh is greater than the number of vertices in the reference frame or reference submesh, the inter-decoding unit 202E may add dummy vertices and connection information for the dummy vertices to the reference frame or reference submesh.

[0054] Here, the inter decoding unit 202E may set the coordinates of such dummy vertices to fixed values ​​(for example, (0,0,0)), or may copy them from a predetermined vertex (for example, the last vertex of the reference frame).

[0055] Furthermore, the inter-decoding unit 202E may copy as much of the connectivity information of the vertices in the reference frame or reference submesh as necessary from the beginning as the connectivity information of the dummy vertices.

[0056] According to this configuration, it is expected that the decoding operation of the current frame or the current submesh can be guaranteed.

[0057] Note that since a base mesh has at least one face and such face has at least three or more vertices, the control signal indicating the number of vertices per frame or per submesh is limited to include at least three or more vertices.

[0058] For example, the control signals pdu_vertex_count_minus_1[titleID][patchIdx] defined in section 8.3.7.3 of Non-Patent Document 4, the control signals sismu_inter_vertex_count[subMeshID] defined in section H.8.1.3.8, and the control signal mesh_vertex_count defined in section I.8.3.7 are modified as shown in Table 3 and restricted to include three or more vertices.

[0059] Here, the control signal pdu_vertex_count_minus_1[titleID][patchIdx] is a control signal that specifies the number of vertices in a patch having a patch ID equal to the patch ID specified by [patchIdx] in an atlas tile having a tile ID equal to the tile ID specified by [titleID].

[0060] The control signal sismu_inter_vertex_count[subMeshID] is a control signal that specifies the number of vertices in the submesh having the same SubmeshID as the SubmeshID specified by [subMeshID].

[0061] The control signal mesh_vertex_count is a control signal that specifies the number of vertices in the decoded mesh.

[0062]

[0063] This configuration is expected to have the effect of preventing the situation where the system operates with meaningless data such as vertices being 1 or 2.

[0064] (Inter Decoding Unit 202E) FIG. 7 is a diagram showing an example of functional blocks of the inter decoding unit 202E.

[0065] As shown in FIG. 7, the inter decoding unit 202E includes a motion vector residual decoding unit 202E1, a motion vector buffer unit 202E2, a motion vector prediction unit 202E3, a motion vector calculation unit 202E4, and an adder 202E5.

[0066] The motion vector residual decoding unit 202E1 is configured to generate a motion vector residual (MVR) from a P frame bitstream.

[0067] Here, MVR is a motion vector residual indicating the difference between MV (Motion Vector) and MVP (Motion Vector Prediction). MV is a difference vector (motion vector) between the coordinates of a vertex of the corresponding I frame and the coordinates of a vertex of the corresponding P frame. MVP is a value (motion vector prediction value) predicted by using MV of the MV of the target vertex.

[0068] The motion vector buffer unit 202E2 is configured to sequentially store the motion vectors output by the motion vector calculation unit 202E4.

[0069] The motion vector prediction unit 202E3 is configured to obtain decoded MVs from the motion vector buffer unit 202E2 for vertices connected to the vertex to be decoded, and to output the MVP of the vertex to be decoded using all or part of the obtained decoded MVs, as shown in Figure 8.

[0070] The motion vector calculation unit 202E4 is configured to add the MVR generated by the motion vector residual decoding unit 202E1 and the MVP output from the motion vector prediction unit 202E3, and output the MV of the vertex to be decoded.

[0071] The adder 202E5 is configured to add the coordinates of the vertex corresponding to the vertex to be decoded, obtained from the decoded basic mesh of the corresponding reference frame (I frame or P frame), to the motion vector MV output from the motion vector calculation unit 202E3, and output the coordinates of the vertex to be decoded.

[0072] Each component of the inter decoding unit 202E will be described in detail below.

[0073] 9 is a flowchart showing an example of the operation of the motion vector prediction unit 202E3. Hereinafter, the operation of the motion vector prediction unit 202E3 will be referred to as the "average prediction method."

[0074] As shown in FIG. 9, in step S1001, the motion vector prediction unit 202E3 sets MVP and N to 0.

[0075] In step S1002, the motion vector prediction unit 202E3 obtains a set of MVs of vertices surrounding the vertex to be decoded from the motion vector buffer unit 202E2, identifies vertices for which subsequent processing has not been completed, and transitions to No. If subsequent processing has been completed for all vertices, it transitions to Yes.

[0076] In step S1003, if the MV of the vertex to be processed has not been decoded, the motion vector prediction unit 202E3 transitions to No, and if the MV of the vertex to be processed has been decoded, the motion vector prediction unit 202E3 transitions to Yes.

[0077] In step S1004, the motion vector prediction unit 202E3 adds MV to MVP and adds 1 to N.

[0078] In step S1005, if N is greater than 0, the motion vector prediction unit 202E3 outputs the result of dividing MVP by N, and if N is 0, it outputs 0, and ends the process.

[0079] That is, the motion vector prediction unit 202E3 is configured to output the MVP to be decoded by averaging the decoded motion vectors of the vertices around the vertex to be decoded.

[0080] The motion vector prediction unit 202E3 may be configured to set the MVP to 0 if the set of decoded motion vectors is an empty set.

[0081] The motion vector calculation unit 202E4 may be configured to calculate the MV of the vertex to be decoded from the MVP output by the motion vector prediction unit 202E3 and the MVR generated by the motion vector residual decoding unit 202E1 using equation (1).

[0082] MV(k)=MVP(k)+MVR(k) (1) where k is the index of the vertex, and MV, MVR, and MVP are vectors having x, y, and z components.

[0083] According to this configuration, since only the MVR is coded instead of the MV using the MVP, it is expected that the coding efficiency will be improved.

[0084] The adder 202E5 is configured to calculate the coordinates of a vertex by adding the MV of the vertex calculated by the motion vector calculation unit 202E4 to the coordinates of the vertex in the reference frame corresponding to the vertex, and to leave the connectivity information (Connectivity) as it is in the reference frame.

[0085] Specifically, the adder 202E5 calculates the coordinate v' of the kth vertex using equation (2). i (k) may be calculated.

[0086] v' i (k)=v′ j (k) + MV(k) ... (2) where v' i (k) is the coordinate of the kth vertex to be decoded in the frame to be decoded, and v' j (k) is the coordinate of the decoded k-th vertex of the reference frame, and MV(k) is the k-th MV of the frame to be decoded, where k=1, 2, . . . , K.

[0087] Furthermore, the connection information of the frame to be decoded is made the same as the connection information of the reference frame.

[0088] It should be noted that the motion vector prediction unit 202E3 calculates the MVP using the decoded MV, and therefore the order of decoding affects the MVP.

[0089] The decoding order is the order in which the vertices of the base mesh in the reference frame are decoded. Generally, if a decoding method is used that uses a fixed repetition pattern to increase the number of base faces one by one from the starting edge, the order of the vertices of the decoded base mesh is determined during the decoding process.

[0090] For example, the motion vector prediction unit 202E3 may use an Edgebreaker to determine the order in which vertices are decoded in the base mesh of the reference frame.

[0091] According to this configuration, since MVs from a reference frame are coded instead of vertex coordinates, it is expected that coding efficiency will be improved.

[0092] (Modification 1 of the inter decoding unit 202E) Hereinafter, a modification 1 of the inter decoding unit 202E will be described.

[0093] The motion vector prediction unit 202E3 of the inter-decoding unit 202E uses the "average prediction method" of averaging the decoded motion vectors of the vertices around the vertex to be decoded, and calculates the MVP using all or only some of the decoded motion vectors of the vertices around the vertex to be decoded so as not to exceed a predetermined maximum number of uses.

[0094] The predetermined maximum number of uses is decoded from the bitstream as a control signal.

[0095] Furthermore, if the number of decoded motion vectors of vertices around the vertex to be decoded exceeds the maximum number of uses, the motion vector prediction unit 202E3 picks up up to the maximum number of uses according to a certain rule.

[0096] For example, the motion vector prediction unit 202E3 may select the first or last vertex in decoding order as a rule.

[0097] The decoding order for a mesh such as that shown in FIG. 10A is as follows: D →v C →vA →v B is.

[0098] FIG. 10B shows the maximum number of decoded adjacent vertices set to 3. A ~v D This is a list of vertices around the vertex to be decoded, which is used when calculating the MVP of the vertex.

[0099] According to this configuration, by determining the maximum number of adjacent vertices, it is possible to expect the effect of reducing the amount of calculation and memory while maintaining or slightly decreasing the coding efficiency.

[0100] However, in order to achieve the above-mentioned effect, it is necessary to set an appropriate maximum number of adjacent vertices in the mesh coding device 100 and write it into the bitstream as a related control signal.

[0101] Therefore, since the range that can be set as the above-mentioned maximum number of adjacent vertices determines the amount of memory to be prepared in the mesh decoding device 200, the maximum number of adjacent vertices is coded / decoded so that it is less than or equal to a predetermined maximum value as a reasonable constraint on the maximum number of adjacent vertices.

[0102] In this way, by defining reasonable constraints on the maximum number of adjacent vertices, it is expected that the design of the mesh decoding device 200 will be made easier.

[0103] Generally, the average number of adjacent vertices in a closed 2-manifold triangle mesh is about 6, but statistically, the maximum number of adjacent vertices is often 7 to 8. As shown in Fig. 11 , the number of decoded motion vectors (vertical axis) changes dynamically depending on the number of vertices around the vertex to be decoded (horizontal axis).

[0104] Therefore, it is desirable to narrow the range that can be set as the maximum number of adjacent vertices described above.

[0105] For example, as shown in Figure 11, by including "3", which is the number of vertices around the vertex to be decoded that has the largest number of decoded motion vectors statistically, within the range that can be set as the maximum number of adjacent vertices in the above-mentioned control signal, or by setting a certain percentage (e.g., 50% or 120%) of the average statistical number of adjacent vertices or a value that is not greater than a natural number that can cover up to N bits (e.g., 3 bits), as the upper limit (maximum value) of the range that can be set as the maximum number of adjacent vertices in the above-mentioned control signal, it is possible to achieve the effect of reducing the amount of calculation and memory required.

[0106] On the other hand, if the range that can be set as the maximum number of adjacent vertices is set to a large value, for example, the worst case of 256 or 8 bits, it may not be possible to achieve the effect of reducing not only the amount of memory required but also the amount of calculation.

[0107] 12 shows an example of the worst case, where when n≧256, the number of decoded adjacent vertices exceeds 256. In FIG. 12, the number of decoded adjacent vertices for vertex n+1 is n.

[0108] If the upper limit of the maximum number of adjacent vertices were set to 256, the mesh decoding device 200 would require not only a huge amount of memory but also a huge amount of calculation, as shown in Fig. 10B. Therefore, the upper limit (maximum value) of the range that can be set as the above-mentioned maximum number of adjacent vertices may be set to 8.

[0109] Furthermore, the range that can be set as the maximum number of adjacent vertices in the above-mentioned control signal may be a definite value, or may be calculated from other control signals or data.

[0110] For example, Level 1 may define the range that can be set as the maximum number of adjacent vertices in the control signal.

[0111] Alternatively, the upper limit of the range that can be set as the maximum number of adjacent vertices in the control signal may be calculated from the number of vertices of the basic mesh using the following equation (3).

[0112] Upper limit of the range that can be set as the maximum number of adjacent vertices in the control signal = log2 (number of vertices in the basic mesh) ... Equation (3) With this configuration, the range that can be set as the maximum number of adjacent vertices can be determined appropriately, and even in the worst case, it is expected to reliably reduce both the amount of calculation and the amount of memory.

[0113] The inter decoding unit 202E decodes the coordinates of the vertex to be decoded by adding the MV decoded from the bit stream of the inter frame and the coordinates of the vertex corresponding to the vertex to be decoded in the reference frame.

[0114] Furthermore, the motion vector prediction unit 202E3 calculates a predicted value of the motion vector of the vertex to be decoded by averaging all or some of the motion vectors of decoded vertices adjacent to the vertex to be decoded, with reference to the adjacent vertex list. Here, the adjacent vertex list is a list of vertices adjacent to each vertex.

[0115] The motion vector prediction unit 202E3 can reuse the adjacent vertex list in the reference frame as the adjacent vertex list in the current frame. Here, it is assumed that the adjacent vertex list includes decoded vertices picked up up to the maximum number of uses (the set maximum number of adjacent vertices).

[0116] Here, the motion vector prediction unit 202E3 can reuse the adjacent vertex list in the reference frame if the following conditions (reuse conditions) are met: each vertex in the reference frame has a one-to-one correspondence with each vertex in the current frame, the decoding order of each vertex in the reference frame is the same as the decoding order of each vertex in the current frame, and an adjacent vertex list containing decoded vertices picked up up to the maximum number of uses in the reference frame has already been saved.

[0117] For example, when a reference frame is decoded, if the reference frame is a P frame or an S frame, there will be an adjacent vertex list containing the decoded vertices picked up to the maximum number of uses.

[0118] Therefore, when a reference frame is decoded, if the reference frame is a P frame or an S frame, the motion vector prediction unit 292E3 stores the decoded reference frame in a reference frame buffer and also stores the adjacent vertex list in the reference frame in a specific buffer.

[0119] Then, when the current frame is decoded, if the type of the reference frame is a P frame or an S frame, the motion vector prediction unit 292E3 reuses the adjacent vertex list in the reference frame stored in a specific buffer as the adjacent vertex list in the current frame.

[0120] However, the particular buffer may store an adjacent vertex list containing decoded vertices picked up to the maximum number of uses in one or more frames.

[0121] Therefore, when adjacent vertex lists for multiple frames are stored in a specific buffer, the motion vector prediction unit 202E3 reuses the adjacent vertex list in the frame corresponding to the reference frame (a frame that includes the same frame index as the reference frame) as the adjacent vertex list for the current frame.

[0122] The motion vector prediction unit 202E3 also applies the operation on the reference frame buffer to the specific buffer. Here, the operation on the reference frame buffer is, for example, the marking process described in Chapter 9.2.4.4 of Non-Patent Document 4 or Non-Patent Document 5.

[0123] According to this configuration, it is expected that the amount of calculation required to create an adjacent vertex list including decoded vertices picked up up to the maximum number of uses in the current frame can be reduced.

[0124] However, in such a case, a special buffer is required to store the adjacent vertex list for one or more frames. Here, in order to minimize the size of the special buffer, further conditions may be added to the reuse conditions mentioned above.

[0125] For example, the motion vector prediction unit 202E3 can reuse the adjacent vertex list in the reference frame as the adjacent vertex list in the current frame if, in addition to the above-mentioned reuse conditions, the condition that the reference frame is the frame immediately preceding the current frame in decoding order is met.

[0126] In this case, the motion vector prediction unit 202E3 may store only the adjacent vertex list for the frame immediately preceding the current frame in decoding order in the specific buffer. Furthermore, the motion vector prediction unit 202E3 does not apply operations on the reference frame buffer to the specific buffer. This configuration is expected to have the effect of reducing the size of the specific buffer.

[0127] (Modification 2 of Inter Decoding Unit 202E) Hereinafter, modification 2 of the inter decoding unit 202E will be described with reference to FIG.

[0128] The motion vector calculation unit 202E4 of the inter decoding unit 202E has a mode 1 and a mode 0.

[0129] In mode 1, the motion vector calculation unit 202E4 adds the MVR generated by the motion vector residual decoding unit 202E1 and the MVP output from the motion vector prediction unit 202E3, and outputs the MV of the vertex to be decoded (see A in Figure 13).

[0130] On the other hand, in mode 0, the motion vector calculation unit 202E4 outputs the MVR generated by the motion vector residual decoding unit 202E1 as the MV of the vertex to be decoded (see B in FIG. 13).

[0131] The operation of the motion vector calculation unit 202E4 in mode 0 corresponds to the operation of setting the MVP output from the motion vector prediction unit 202E3 to zero.

[0132] Furthermore, the motion vector calculation unit 202E4 may set the MV modes of N (N≧1) consecutive vertices in decoding order to the same mode.

[0133] The motion vector calculation unit 202E4 groups the above-mentioned N vertices into one group. The size of this group (group size) N is equal to or greater than 1. The motion vector calculation unit 202E4 decodes a control signal (group size shown in FIG. 13 ) from the bitstream that allows the calculation of this group size.

[0134] However, if the number of vertices remaining in the last group is smaller than the group size, the motion vector calculation unit 202E4 puts all of the remaining vertices into the group.

[0135] In this way, when N consecutive vertices are set to the same mode, the amount of code for the mode can be reduced, and therefore, the effect of improving coding efficiency can be expected.

[0136] Here, the greater the number of consecutive vertices with the same mode, the greater the effect of reducing the amount of code for that mode. Therefore, it is necessary to set an appropriate group size in the mesh encoding device 100 and decode it as a control signal from the bitstream in the mesh decoding device 200.

[0137] Therefore, it is desirable that the settable range of such a control signal is no smaller than the number of consecutive vertices that actually have the same mode.

[0138] For example, if a mode in which all vertices are almost the same is selected, the group size may be set to the total number of vertices.

[0139] Table 4 shows an example where the number of vertices for which mode 0 is selected is 80% or more, and an example where the number of vertices for which mode 1 is selected is 90% or more.

[0140] Therefore, the settable range of the control signal is set to cover from 1 to a predetermined maximum value, which is equal to or greater than the total number of vertices of the basic mesh.

[0141]

[0142] If the control signal (group size) is a natural number and is set to be equal to or greater than the total number of vertices, the absolute value is large, resulting in a large amount of code.

[0143] Therefore, the above control signal may be logarithmic. Specifically, the control signal may be log2_group_size, and the group size may be calculated by the following equation (4).

[0144] group size=2 log2_group_size (4) If there is only one group in the frame, that group is the last group. In other words, if group size is greater than the number of vertices, all vertices are put into the group.

[0145] Furthermore, the settable range of the above control signal may be a definite value, or may be calculated from other control signals or data.

[0146] For example, Level 1 may define the range that can be set in such a control signal.

[0147] Alternatively, the settable range of such a control signal may be calculated from the number of vertices of the basic mesh.

[0148] For example, the range that can be set in such a control signal may be the smallest natural number that is a power of two that can cover the number of vertices of the basic mesh.

[0149] Furthermore, the settable range of the above-mentioned control signal may be narrowed, and a predetermined flag (Mode flag) of another control signal may be introduced as shown in Fig. 14. In such a case, as shown in Fig. 14, if the predetermined flag is TRUE (Mode flag = 1), the motion vector calculation unit 202E4 groups all vertices into one group (i.e., the number of all vertices is the group size), and if the flag is FALSE, the group size calculated from the above-mentioned control signal remains unchanged.

[0150] The control signal may be set for each sequence or for each frame. When the control signal is set for each sequence, the group size of all frames is the same.

[0151] According to this configuration, by determining the range within which the group size can be set appropriately, it is possible to deal with any situation, reliably reduce the amount of code in the mode, and improve the coding efficiency.

[0152] (Third Modification of Inter Decoding Unit 202E) In a further modification of the above-described inter decoding unit 202E, the following functional blocks are added before implementing the above-described inter decoding unit 202E.

[0153] Specifically, as shown in FIG. 15, the inter decoding unit 202E includes, in addition to the configuration shown in FIG. 8, an overlapping vertex searching unit 202E6, an mv_signaled_flag obtaining unit (flag obtaining unit) 202E7, and a motion vector obtaining unit 202E8.

[0154] Here, derived_my_present_flag (first flag) is included at the beginning of the P frame bitstream and has at least two values, 0 or 1.

[0155] Furthermore, mv_signaled_flag (second flag) is included in the bitstream of a P frame when derived_my_present_flag indicates No, and has two values, 0 or 1, for each vertex.

[0156] If the derived_my_present_flag indicates No (for example, if the derived_my_present_flag is 0), the mv_signaled_flag acquisition unit 202E7 decodes the motion vectors of all vertices from the P frame bitstream, and sets the value of my_signaled_flag to 1 without decoding the my_signaled_flag of all vertices from the P frame bitstream.

[0157] If derived_my_present_flag indicates Yes (for example, if derived_my_present_flag is 1), the mv_signaled_flag acquisition unit 202E7 performs different processing on each vertex of the P frame. The mv_signaled_flag acquisition unit 202E7 may use mv_signaled_flag to determine the processing method for each vertex.

[0158] In addition, when derived_my_present_flag indicates Yes and the mv_signaled_flag of a certain vertex indicates Yes, the mv_signaled_flag acquisition unit 202E7 does not perform processing in the motion vector acquisition unit 202E8 for the motion vector of that vertex, but performs processing similar to that of the inter decoding unit 202E shown in Figure 7 or its modified example.

[0159] Furthermore, if derived_my_present_flag indicates Yes and mv_signaled_flag of a certain vertex indicates No, the mv_signaled_flag acquisition unit 202E7 performs processing on the motion vector of that vertex in the motion vector acquisition unit 202E8 to acquire the motion vector of that vertex.

[0160] For example, if derived_my_present_flag indicates Yes, the mv_signaled_flag acquisition unit 202E7 decodes mv_signaled_flag for each vertex from the bitstream of the P frame.

[0161] If derived_my_present_flag indicates Yes and if mv_signaled_flag of a certain vertex indicates Yes, the mv_signaled_flag acquisition unit 202E7 sets the prediction mode (MV mode) of that vertex to 2.

[0162] On the other hand, when derived_my_present_flag indicates Yes and mv_signaled_flag of a certain vertex indicates No, the mv_signaled_flag acquisition unit 202E7 sets the prediction mode of the certain vertex to a value other than 2.

[0163] Furthermore, if derived_my_present_flag indicates No, the mv_signaled_flag acquisition unit 202E7 does not decode the mv_signaled_flag of all vertices of the P frame from the bitstream, but sets the value to 1, and sets the MV mode of the vertex to a value other than 2.

[0164] The duplicated vertex search unit 202E6 is configured to search for the index of a vertex (hereinafter referred to as a duplicated vertex) with matching coordinates from the geometric information of the basic mesh of the decoded reference frame and store it in a buffer (not shown).

[0165] Specifically, the input to the overlapping vertex search unit 202E6 is the index (in decoding order) and position coordinates of each vertex of the base mesh of the decoded reference frame.

[0166] The output of the duplicate vertex search unit 202E6 is a list that stores the index (vindex1) of each vertex if there is a duplicate vertex associated with the index (vindex0) of that vertex, and stores the index (vindex0) of that vertex itself or a specific value (e.g., -1) that is not used in the index of each vertex if there is no duplicate vertex. Here, this list is stored in the buffer repVert in the order of index0.

[0167] Furthermore, since the vertex of vindex1 is decoded before vindex0, the relationship vindex0>vindex1 holds.

[0168] The duplicated vertex search unit 202E6 determines whether or not a duplicated vertex exists for each vertex (index: vindex0) of the basic mesh of the reference frame from the first vertex (index: 0) of the basic mesh of the decoded reference frame to the immediately previous vertex (index: vindex0-1), and if it determines that such a duplicated vertex exists, it outputs the index of the duplicated vertex using at least one of the following three methods.

[0169] (Method 1) The duplicated vertex search unit 202E6 sequentially searches for duplicated vertices with matching coordinates as follows: When a duplicated vertex exists, vref is vindex1, and when a duplicated vertex does not exist, vRef is −1.

[0170] vRef =firstVertexIndexDuplicated(vindex0) where firstVertexIndexDuplicated(v){ for( i = 0; i <v; i++){ if(referenceSubmeshVertexPositions[ i ] == referenceSubmeshVertexPositions[ v ]) { return i } } return -1 }

[0171] (Method 2) The duplicated vertex search unit 202E6 searches for duplicated vertices with matching coordinates using a binary search. For example, the duplicated vertex search unit 202E6 may use the find function of the associative array class map.

[0172] (Method 3) The duplicated vertex search unit 202E6 searches for duplicated vertices with matching coordinates using a hash table. For example, the duplicated vertex search unit 202E6 may use the find function of the hash associative array class unordered_map.

[0173] As a method for finding duplicate vertices in the base mesh of the reference frame, a method may be used in which, for vertices where duplicate vertices exist, the index of the duplicate vertex is decoded from a special signal instead of the position coordinate.

[0174] When MVmode is 2 (when derived_my_present_flag indicates Yes and mv_signaled_flag of the vertex indicates No), a duplicate vertex of the vertex exists, so the motion vector acquisition unit 202E8 is configured to acquire from the motion vector buffer unit 202E2 the motion vector of the vertex having the index (vindex1) of the duplicate vertex related to the index (vindex0) of the vertex output from the duplicate vertex search unit 202E6, and set the motion vector of the vertex as the motion vector of the vertex.

[0175] That is, it is the output of the duplicate vertex index (vindex1) 202E6 and is not decoded from the bitstream.

[0176] Here, if MVmode is other than 2 (if derived_my_present_flag indicates No, or if derived_my_present_flag indicates Yes and the mv_signaled_flag of the vertex indicates Yes), processing similar to that of the inter-decoding unit 202E shown in Figure 7 or a modified example thereof is performed instead of the motion vector acquisition unit 202E8.

[0177] According to this configuration, it is possible to expect an effect of reducing the decoding calculation and the amount of code for the motion vector for the vertices where overlapping vertices exist.

[0178] In a further modification of the inter decoding unit 202E described above, the overlapping vertex searching unit 202E6 searches for overlapping vertices only among the vertices for which mv_signaled_flag is set to No, rather than searching for overlapping vertices among all the vertices of the base mesh of the reference frame.

[0179] However, the input to the overlapping vertex search unit 202E6 includes mv_signaled_flag in addition to the index (in decoding order) and position coordinates of each vertex of the base mesh of the decoded reference frame.

[0180] According to this modification, overlapping vertices are searched for only among vertices that have overlapping vertices, rather than among all vertices, so that it is possible to expect an effect of reducing the amount of calculation required to decode motion vectors.

[0181] In a further modification of the inter decoding unit 202E described above, the mv_signaled_flag acquisition unit 202E7 decodes the mv_signaled_flag in two stages.

[0182] The mv_signaled_flag acquisition unit 202E7 groups N vertices together, decodes the mv_group_signaled_flag (third flag) from the P frame bitstream for each group, sets the mv_signaled_flag of all vertices in the group where mv_group_signaled_flag is 1 to 1, and decodes the mv_group_signaled_flag from the P frame bitstream for each vertex in the group where mv_group_signaled_flag is 0.

[0183] According to this modification, mv_group_signaled_flag is decoded in two stages, which is expected to reduce the amount of decoding calculations for motion vectors and the amount of code.

[0184] In a further modified example of the above-mentioned inter-decoding unit 202E, the mv_signaled_flag acquisition unit 202E7 decodes mv_signaled_flag from the P frame bitstream for vertices that have overlapping vertices, rather than for each vertex, and does not decode mv_signaled_flag from the P frame bitstream for vertices that do not have overlapping vertices, but sets mv_signaled_flag to 1.

[0185] The inter decoding unit 202E decodes, from the bitstream, a control signal indicating the number of vertices having overlapping vertices before mv_signaled_flag.

[0186] This control signal makes it possible to decode the mv_signaled_flag without performing the processing of the overlapping vertex searching unit 202E6.

[0187] Furthermore, as a requirement for the compatibility of the bit stream, the number of vertices having overlapping vertices output by the overlapping vertex search unit 202E6 must match the number of vertices indicated by the control signal.

[0188] Also, the duplicate vertex search unit 202E6 stores all duplicate vertex indices in a separate list. For example, the duplicate vertex search unit 202E6 may store all duplicate vertex indices in the duplicated_vertex_list as shown in the following (modified example of Method 1).

[0189] (Modified example of Method 1) The duplicate vertex search unit 202E6 sequentially searches for duplicate vertices whose coordinates match as shown below. When a duplicate vertex exists, vRef is vindex1, and when no duplicate vertex exists, vRef is -1.

[0190] vRef =firstVertexIndexDuplicated(vindex0,duplicated_vertex_list) where firstVertexIndexDuplicated(v,duplicated_vertex_list){ for( i = 0; i<v; i++){ if(referenceSubmeshVertexPositions[ i ] == referenceSubmeshVertexPositions[ v ]) { duplicated_vertex_list.push_back(v) return i } } return -1 } According to this modified example, since the mv_signalled_flag is provided only for vertices having duplicate vertices instead of all vertices, the decoding calculation of the motion vector and the effect of reducing the amount of code can be expected.

[0191] In a further modified example of the above-mentioned inter-decoding unit 202E, the mv_signalled_flag acquisition unit 202E7 may set the mv_signalled_flag of vertices having no duplicate vertices to 0.

[0192] Here, when the mv_signaled_flag of a vertex that does not have an overlapping vertex is 0, there are no overlapping vertices in the basic mesh of the reference frame, but there are vertices with the same motion vector, so the mv_signaled_flag acquisition unit 202E7 decodes the index of the vertex that has the same motion vector as the vertex in question from the interframe bitstream and acquires the motion vector of the vertex in question.

[0193] Specifically, when MVmode is 2 (when derived_my_present_flag indicates Yes and mv_signaled_flag of the vertex indicates No), if a duplicate vertex of the vertex exists, the motion vector acquisition unit 202E8 is configured to acquire from the motion vector buffer unit 202E2 the motion vector of a vertex having the index (vindex1) of the duplicate vertex related to the index (vindex0) of the vertex output by the duplicate vertex search unit 202E6, and set the motion vector of the vertex as the motion vector of the vertex; however, in this modified example, if a duplicate vertex of the vertex does not exist, the motion vector acquisition unit 202E8 is configured to decode from the bitstream the index (vindex1) of a vertex having the same motion vector as the vertex, acquire the motion vector of the vertex having the index (vindex1), and set the motion vector of the vertex as the motion vector of the vertex.

[0194] According to this modification, even for a vertex that does not have an overlapping vertex, if mv_signaled_flag indicates No, the motion vector is obtained from another vertex, so that it is possible to expect reductions in the motion vector decoding calculations and the amount of code.

[0195] Note that the above-mentioned modifications may or may not be used simultaneously. If they cannot be used simultaneously, a control signal indicating which modification to use is provided, and the control signal is decoded from the bitstream to determine which modification to use. However, the control signal may be an extension of an existing control signal.

[0196] In a further modification of the inter decoding unit 202E described above, if the reference frame is an inter frame, the overlapping vertex searching unit 202E6 reuses the results obtained in the reference frame in the frame to be decoded.

[0197] However, if the reference frame is an intraframe, the overlapping vertex searching unit 202E6 may have information about overlapping vertices and reuse such information.

[0198] Specifically, first, the duplicated vertex search unit 202E6 decodes from the bitstream a control signal indicating whether or not to reuse the result obtained in the reference frame (information about duplicated vertices, including the number of vertices that have duplicated vertices) in the frame to be decoded. However, the duplicated vertex search unit 202E6 may use an existing control signal as is or after extending it.

[0199] Second, when the control signal is Yes, if the reference frame is an interframe, the overlapping vertex search unit 202E6 reuses the results obtained in the reference frame for the frame to be decoded. However, if the reference frame is an intraframe, the overlapping vertex search unit 202E6 may assume that it has information about overlapping vertices and reuse such information.

[0200] Specifically, the output (the result described above) of the duplicate vertex search unit 202E6 for the reference frame is a list that stores the index (vindex1) of a duplicate vertex related to the index (vindex0) of each vertex if such a duplicate vertex exists, and stores the index (vindex0) of the vertex itself or a specific value (e.g., -1) when the index cannot be used if such a duplicate vertex does not exist.

[0201] Here, this list is stored in the buffer repVert in the order of vindex0.

[0202] In the following example, if there is no overlapping vertex, the index of the vertex (vindex0) itself is saved. tRef Display as.

[0203] repVert tRef (vindex0) = vRef vRef = vindex1 if vindex0 and vindex1 are overlapping vertices vRef = vindex0 if vindex0 and vindex1 are not overlapping vertices In the decoding target frame t, repVert tRef is reused, the duplicated vertex searching unit 202E6 does not need to search for duplicated vertices for each vertex (index: vindex0) of the base mesh of the reference frame of the decoding target frame, and tRef You can just use it as is.

[0204] repVert t (vindex0)=vindex0 if repVert tRef (vindex0)=vindex0 repVert t (vindex0)=repVert tRef (vindex0) if repVert tRef (vindex0)! = vindex0 According to this modification, overlapping vertices are not searched for, so that the effect of reducing the motion vector decoding calculation can be expected.

[0205] In a further modification of the inter decoding unit 202E described above, if the reference frame is an inter frame, an mv_signaled_flag acquisition unit (flag acquisition unit) 202E7 reuses the mv_signaled_flag acquired in the reference frame for the frame to be decoded.

[0206] However, if the reference frame is an intraframe, the mv_signaled_flag acquisition unit 202E7 may assume that the reference frame has an mv_signaled_flag and reuse the mv_signaled_flag.

[0207] Specifically, when derived_my_present_flag indicates Yes, the mv_signaled_flag acquisition unit 202E7 does not decode mv_signaled_flag for each vertex from the bitstream of the P frame, but decodes the difference from the mv_signaled_flag of the reference frame.

[0208] Alternatively, when there is no difference in all mv_signaled_flag, a control signal may be provided. In this case, the mv_signaled_flag acquisition unit 202E7 decodes the control signal, and if the control signal is a specific value (for example, TRUE), it sets the mv_signaled_flag of the reference frame as the mv_signaled_flag of the target frame as is, and if the control signal is a specific value (for example, FALSE), it further decodes the difference from the mv_signaled_flag of the reference frame and calculates the mv_signaled_flag of the target frame.

[0209] According to this modification, it is expected that the amount of code for mv_signaled_flag can be reduced.

[0210] In some cases, the above-mentioned modifications may or may not be used simultaneously. If they cannot be used simultaneously, a control signal indicating which modification to use is provided, and the control signal is decoded from the bitstream to determine which modification to use.

[0211] (Modification 1 of Basic Mesh Decoding Unit 202) Hereinafter, modification 1 of the basic mesh decoding unit 202 will be described with reference to Figs.

[0212] As shown in FIG. 16, the basic mesh decoding unit 202 according to the present modified example 1 includes a separating unit 202A, an intra decoding unit 202B, a mesh buffer unit 202C, an inter decoding unit 202E, and a skip decoding unit 202F.

[0213] The skip decoding unit 202F is configured to decode the basic mesh of the frame to be decoded by directly using the decoding basic mesh of the designated reference frame.

[0214] In this embodiment, the frame may be either a mesh or a submesh.

[0215] For example, as shown in FIG. 17, "P_SUBMESH" in smh_type may correspond to a P frame, "I_SUBMESH" in smh_type may correspond to an I frame, and "SKIP_SUBMESH" in smh_type may correspond to an S frame.

[0216] (Skip decoding unit 202F) The skip decoding unit 202F is configured to extract the decoded base mesh (reference decoded base mesh) of the specified reference frame from the mesh buffer unit 202C, and decode the coordinates of the vertices and the indexes of the vertices of the extracted reference decoded base mesh as they are, to decode the coordinates of the vertices and the indexes of the vertices of the base mesh of the frame to be decoded.

[0217] Here, the mesh buffer unit 202C has at least one reference frame and is configured to store at least one decoded basic mesh for each reference frame.

[0218] The skip decoding unit 202F may identify the designated reference decoding base mesh using a control signal decoded from the bitstream or a predetermined rule.

[0219] For example, the predetermined rule may be to extract the first reference frame in the reference frame list from the mesh buffer unit 202C, or to extract the reference frame whose frame index is closest to the frame to be decoded.

[0220] In this embodiment, a frame in which the coordinates of the vertices of the base mesh are decoded using the coordinates of the vertices of the decoded base mesh for reference and the indexes of those vertices as they are is called an "S frame."

[0221] According to this configuration, the skip decoding unit 202F does not need a motion vector, and therefore it is possible to expect a significant reduction in the amount of code and the amount of calculation.

[0222] (Mesh Buffer Unit 202C) The mesh buffer unit 202C is configured to store one or more reference decoded basic meshes in a predetermined order.

[0223] Such a basic mesh has metadata such as a frame number and a submesh number, as well as at least the coordinates of each vertex and the index of that vertex, and is stored in the mesh buffer unit 202C in a predetermined order determined by the reference frame list.

[0224] As shown in FIG. 18, the reference frame list (ref_list0) is a list of information that identifies all reference decoding basic meshes stored in the mesh buffer unit 202C.

[0225] The reference frame list may be determined by control signals decoded from the bitstream, as shown in FIG. 18, or may be calculated naturally from the decoding order of the frames.

[0226] The control signal decoded from the bitstream may be represented by a relative distance from the frame to be decoded, or may be represented by a frame index that is an absolute value.

[0227] Additionally, the control signal may use a short-term or long-term frame of reference.

[0228] For example, when a short-term reference frame is used, the absolute value (abs_delta_mfoc_st) of the difference in display order (Display Order) between the current frame (cur) and the reference frame (ref) and its sign (sign_flag) may be decoded from the bitstream, and the display order (Display Order) of the reference frame may be specified by the following formula: If(sign_flag) { Display Order(ref) = Display Order(cur) + abs_delta_mfoc_st} else { Display Order(ref) = Display Order(cur) - abs_delta_mfoc_st} Furthermore, when a method of calculating naturally from the decoding order of frames is used, for example, when no control signal is present in the reference frame list, frames may be arranged in order of a certain number of frames starting from the most recently decoded frame. That is, the reference frame list may be set to {0, -1, -2, ..., -(N-1)}.

[0229] Basically, the reference frame list does not change for each frame except in special circumstances (for example, when a re-ordering instruction is received).

[0230] The mesh buffer unit 202C may be updated as follows.

[0231] When a basic mesh is decoded, in the case of an I frame or a P frame, the mesh buffer unit 202C deletes one or more existing reference frames in a predetermined order determined by the reference frame list, and inserts one or more basic meshes including the basic mesh of the decoded frame, or creates and inserts one basic mesh from multiple basic meshes, thereby adjusting the order of the reference frames.

[0232] This deletion operation may be performed only when the mesh buffer unit 202C is full. The number of basic meshes that can be stored in the mesh buffer unit 202C is predetermined. In this embodiment, the mesh buffer unit 202C is defined as being full when this number of basic meshes is reached. In the above-mentioned creation operation, a single basic mesh may be created by performing a weighted average of the coordinates of the vertices corresponding to the basic meshes of the decoded frame and the existing basic meshes stored in the mesh buffer unit 202C.

[0233] The weights used in such a weighted average may be predetermined, calculated using the frame index, or decoded from the control signal.

[0234] However, the mesh buffer unit 202C may or may not perform such an update when it is an S frame.

[0235] In addition, when the mesh buffer unit 202C receives a control signal indicating a re-ordering instruction from a control signal decoded from the bitstream, it updates the reference frame list as shown in Figure 19 and adjusts the order of the reference frames according to the specified order determined by the updated reference frame list (ref_list0).

[0236] (Inter decoding unit 202E) The inter decoding unit 202E is configured to decode the coordinates of the vertices of the P frame by adding the coordinates of the vertices of the reference frame retrieved from the mesh buffer unit 202C and the motion vectors decoded from the bit stream of the P frame.

[0237] Furthermore, the inter-decoding unit 202E can adjust the indexes of the vertices of the P frame using the pair of indexes A(k) and B(k) of the vertices existing as overlapping vertices stored in the specific buffer. All or part of these indexes are decoded from the bitstream. This decoding method may be arithmetic coding. This configuration is expected to have the effect of eliminating the limit on the maximum value of the index to be decoded using arithmetic coding. For example, arithmetic coding called ue(v) may be used.

[0238] (Modification 2 of Basic Mesh Decoding Unit 202) Hereinafter, modification 2 of the basic mesh decoding unit 202 will be described with reference to FIG.

[0239] The skip decoding unit 202F will be described below, but it may also be applied to the inter decoding unit 202E.

[0240] As shown in FIG. 20, in the skip decoding unit 202F, the decoding order and the display order are different in order to allow reference to subsequent frames.

[0241] Here, the display order is the same as the input order when encoding, and is the same as the output order when decoding.

[0242] On the other hand, the decoding order is the same as the output order when encoding, and is the same as the input order when decoding.

[0243] The reference frame may be calculated by taking a weighted average of the subsequent frame and one or more other frames.

[0244] However, when referring to multiple frames including subsequent frames, MR_SUBMESH (MR frame or B frame) is defined as a new frame type (smh_type) in FIG. 14, and MR_SUBMESH is decoded from the bitstream.

[0245] Furthermore, the other frame may be a decoded frame immediately before the target frame, as shown in FIG.

[0246] Such weights may be calculated using the frame interval between the target frame and the subsequent frame and the frame interval between the target frame and another frame, or may be determined in advance.

[0247] The basic mesh decoding unit 202 decodes the control signal (smh_mesh_from_order_cnt_lsb) from the bitstream and decodes the order of the outputs.

[0248] When sub-meshes as defined in the above-mentioned non-patent document 4 exist, all sub-meshes are set to the same control signal (smh_mesh_from_order_cnt_lsb) or the control signal (smh_mesh_from_order_cnt_lsb) is applied to all sub-meshes.

[0249] The value indicated by this control signal (smh_mesh_from_order_cnt_lsb) may be a difference from the display order of the frame to be decoded, or may be an order within a predetermined frame group MaxMeshFrmOrderCntLsb.

[0250] In addition, when the decoding order (Decode Order) and the display order (Display Order) are different, if the decoded basic meshes are arranged in the decoding order (Decode Order), the basic mesh decoding unit 202 may rearrange the decoded basic meshes in the display order (Display Order).

[0251] In addition, in an S frame that can reference a subsequent frame, two mesh buffers 202C may be provided, or when only one mesh buffer unit 202 is provided, there will be at least one reference frame, including a reference frame whose display order is later than the frame to be decoded.

[0252] The skip decoding unit 202F specifies a reference frame by receiving a control signal decoded from the bitstream, a predetermined rule, or a re-ordering instruction.

[0253] Specifically, the skip decoding unit 202F specifies a reference frame in the reference frame list using the control signal.

[0254] Alternatively, the skip decoding unit 202F specifies the first reference frame in the reference frame list.

[0255] Alternatively, upon receiving the re-ordering instruction, the skip decoding unit 202F updates the reference frame list and the reference frame order in the mesh buffer unit 202C, and specifies the first reference frame in the reference frame list.

[0256] In this embodiment, even if an S frame is not decoded, it does not affect the decoding of other frames. Therefore, if some or all of the S frames are not decoded, temporal scalability can be achieved.

[0257] Furthermore, the base mesh decoding unit 202 may decode the base mesh of the S frame by integrating multiple reference frames in response to a control signal.

[0258] For example, the basic mesh decoding unit 202 may be configured to average the coordinates of corresponding vertices in the basic meshes of the two previous and next reference frames, and use the average coordinates and vertex indexes as they are to decode the coordinates of the vertices and the indexes of the vertices of the basic mesh of the frame to be decoded.

[0259] According to this configuration, the skip decoding unit 202F or the inter decoding unit 202E can generate a high-quality base mesh without requiring a motion vector, which is expected to improve the quality of the decoded mesh.Furthermore, it is expected to achieve temporal scalability.

[0260] However, in order to achieve temporal scalability, control signals indicating whether or not to decode the basic mesh, the displacement, and the texture are defined for each frame, and each is decoded from the bitstream.

[0261] Furthermore, the Temporal_ID of the atlas and the base mesh may be matched within the same frame. Furthermore, the Temporal_ID of the atlas and the texture may be matched within the same frame. Furthermore, the Temporal_ID of the atlas and the displacement amount may be matched within the same frame.

[0262] According to this configuration, it is expected that it will be possible to avoid frames that cannot be decoded and unnecessary data.

[0263] It is desirable that the interval between adjacent frames having the same Temporal_ID is constant.

[0264] Adjacent frames with the same Temporal_ID have the closest POC.

[0265] By keeping the frame intervals constant as described above, it is expected that a constant frame rate can be maintained when displaying decoded frames.

[0266] Furthermore, the decoding orders of atlases and base meshes having the same display order may be matched, the decoding orders of atlases and displacements having the same display order may be matched, and the decoding orders of atlases and textures having the same display order may be matched.

[0267] Alternatively, the random access points of an atlas and a base mesh that have the same display order may be matched. Alternatively, the random access points of an atlas and a displacement that have the same display order may be matched. Alternatively, the random access points of an atlas and a texture that have the same display order may be matched. Note that random access points are defined in Non-Patent Document 4 or Non-Patent Document 5.

[0268] According to this configuration, when decoding the basic mesh, the displacement amount, and the texture, it is possible to expect the effect that the mesh can be reproduced without waiting for each other to be decoded.

[0269] Furthermore, a frame having a Temporal_ID higher than the control signal Temporal_ID of the current frame to be decoded is not used as a reference frame for the current frame to be decoded.

[0270] This is expected to have the effect of eliminating the possibility of reference frames being discarded.

[0271] An example of realizing temporal scalability using the above-mentioned Temporal_ID will be described below.

[0272] The atlas, base mesh, displacement, and texture bitstreams are encapsulated by Network Abstraction Layer (NAL) units, which may have a NAL header as shown in Figure 22.

[0273] The TID, defined as the last 3 bits in the NAL header, is the Temporal_ID plus 1. The range of TID is 1 to 7, with zero being prohibited.

[0274] LayerID / R6, defined as the 6 bits immediately preceding the TID in the NAL header, specifies the identifier of the layer to which the NAL unit belongs.

[0275] The value of LayerID / R6 must be in the range of 0 to 62. The value 63 may be specified by ISO / IEC in the future.

[0276] For purposes other than determining the amount of data in a decoded unit of the bitstream, the mesh decoding device 200 ignores all data following the value 63 in a NAL unit, and a mesh decoding device 200 conforming to the specified profile will ignore (i.e., remove from the bitstream and discard) all NAL units whose LayerID-R6 value is not 0.

[0277] The value 63 of LayerID / R6 can be used to indicate an enhancement layer identifier in future enhancements.

[0278] When sub-meshes defined in Non-Patent Document 4 exist, all sub-meshes are assigned the same TID or the TID is applied to all sub-meshes.

[0279] Non-Patent Document 5 can be used for the atlas, and HEVC or VVC, a video coding method, can be used for the displacement and texture, so the following will explain the basic mesh.

[0280] The LayerID / R6 values ​​of all BMCL NAL units of a coded basic mesh frame must be the same. The LayerID / R6 value of a coded basic mesh frame is the LayerID / R6 value of the BMCL NAL unit of the coded basic mesh frame.

[0281] If NALType is equal to NAL_EOB, the value of LayerID / R6 must be equal to 0.

[0282] If the NALType is in the range from NAL_BLA_W_LP to NAL_RSV_BMCL_29 defined in Non-Patent Document 4, i.e., if it belongs to an IRAP coded basic mesh frame, Temporal_ID must be 0.

[0283] If NALType is equal to NAL_TSA_R or NAL_TSA_N, Temporal_ID must not be equal to 0.

[0284] If NALType is equal to 0 and NALType is equal to NAL_STSA_R or NAL_STSA_N, Temporal_ID must not be equal to 0.

[0285] The value of Temporal_ID must be the same for all BMCL NAL units within an access unit.

[0286] The value of Temporal_ID of a coded basic mesh frame or access unit is the value of Temporal_ID of the BMCL NAL unit of the coded basic mesh frame or access unit.

[0287] The value of Temporal_ID of a sub-layer representation is the maximum value of Temporal_ID of all BMCL NAL units within the sub-layer representation.

[0288] The value of Temporal_ID for non-BMCL NAL units is restricted as follows: - if NALType is equal to NAL_BMSPS, Temporal_ID shall be 0 and the Temporal_ID of the access unit containing the NAL unit shall be 0. - otherwise, if NALType is equal to NAL_EOS or NAL_EOB, Temporal_ID shall be 0. - otherwise, if NALType is equal to NAL_AUD or NALLFDD, Temporal_ID shall be equal to the Temporal_ID of the access unit containing the NAL unit. - otherwise, Temporal_ID shall be greater than or equal to the Temporal_ID of the access unit containing the NAL unit.

[0289] Note that if the NAL unit is not BMCL, the value of Temporal_ID is equal to the minimum value of the Temporal_ID values ​​of all access units to which the non-BMCL NAL unit applies.

[0290] When NALType is equal to NAL_BMFPS, Temporal_ID can be greater than or equal to the Temporal_ID of the included access unit, since all basic mesh frame parameter sets (BMFPS) are included at the beginning of the bitstream, where the Temporal_ID of the first coded basic mesh frame is 0.

[0291] The skip decoding unit 202F refers to the specified tIDTarget and discards, without decoding, any NAL unit whose Temporal_ID is higher than tIDTarget.

[0292] Here, tIDTarget may be specified by a predetermined value, or may be specified based on the network status or the terminal capability of the mesh decoding device 200 .

[0293] For example, a lower tIDTarget is specified in the wireless case than in the wired case. Also, a lower tIDTarget is specified when the network condition is poor. Also, a lower tIDTarget is specified when decoding is performed by a mesh decoding device 200 with low specifications.

[0294] However, the requirement for bitstream conformance is that there must be at least one NAL unit in the bitstream whose Temporal_ID is not higher than tIDTarget.

[0295] As shown in FIG. 23, the number of sub-meshes may differ for each frame (intraframe, interframe, and skip frame).

[0296] In this case, the intra decoder 202B, the inter decoder 202E, and the skip decoder 202F assign a unique submesh ID to each submesh in each frame.

[0297] Furthermore, as shown in FIG. 24, the intra decoder 202B, the inter decoder 202E, and the skip decoder 202F may assign different Submesh IDs to corresponding submeshes between frames.

[0298] However, the inter decoding unit 202E or the skip decoding unit 202F can only refer to submeshes that have the same SubmeshID in the reference frame.

[0299] Alternatively, the inter decoding unit 202E or the skip decoding unit 202F can refer to only submeshes that have the same number of vertices in the reference frame.

[0300] Alternatively, the intra decoder 202B and the inter decoder 202E can refer to a submesh specified in a reference frame.

[0301] In such a case, if there are multiple submeshes in the reference frame, the inter decoding unit 202E or the skip decoding unit 202F may decode a control signal specifying the Submesh ID of a referenceable submesh from the bit stream of the current submesh.

[0302] On the other hand, when there is only one submesh in the reference frame, the inter decoding unit 202E or the skip decoding unit 202F may treat this submesh as a referenceable submesh.

[0303] However, if the above-mentioned control signal does not exist, the inter decoding unit 202E or the skip decoding unit 202F sets the submesh ID of the referable submesh to the same submesh ID as the submesh in the current frame.

[0304] Additionally, the inter decoding unit 202E or the skip decoding unit 202F may decode, from the bitstream, a control signal indicating whether or not the above-mentioned control signal is present.

[0305] The inter decoding unit 202E or the skip decoding unit 202F may decode a control signal for selecting the method for determining the above-mentioned referenceable sub-meshes.

[0306] The subdivision unit 203 and the displacement amount decoding unit 206 may be configured in accordance with Non-Patent Document 4.

[0307] According to the present invention, the amount of calculation can be reduced by reusing the reference frame itself without searching for adjacent vertices that have already been decoded.

[0308] Furthermore, according to this embodiment, in inter-prediction coding, the base mesh of the current frame can be decoded even if the number of vertices of the base mesh of the current frame differs from the number of vertices of the reference frame or reference sub-mesh.

[0309] Furthermore, according to this embodiment, in inter-prediction coding, it is possible to avoid a situation in which the number of vertices of the base mesh of the current frame differs from the number of vertices of the reference frame or reference sub-mesh.

[0310] Furthermore, according to this embodiment, in inter-prediction coding, it is possible to identify which sub-mesh to refer to by introducing a control signal indicating which sub-mesh to refer to in the reference frame of the current frame.

[0311] Furthermore, according to this embodiment, in inter-prediction coding, even if there is no information in the reference frame of the current frame as to which sub-mesh to refer to, it is possible to identify which sub-mesh to refer to.

[0312] Furthermore, according to this embodiment, it is possible to ensure that the base mesh has at least one face.

[0313] Furthermore, according to this embodiment, the Temporal_scalability function can be realized.

[0314] Furthermore, according to this embodiment, the encoding efficiency of the mesh can be improved.

[0315] The mesh encoding device 100 and mesh decoding device 200 described above may be realized as a program that causes a computer to execute each function (each step).

[0316] According to this embodiment, for example, it is possible to improve the overall service quality in video communication, which makes it possible to contribute to Goal 9 of the Sustainable Development Goals (SDGs) led by the United Nations, which is to "Develop resilient infrastructure, promote sustainable industrialization and foster innovation."

[0317] 1...Mesh processing system 100...Mesh encoding device 200...Mesh decoding device 201...Demultiplexing unit 202...Basic mesh decoding unit 202A...Demultiplexing unit 202B...Intra decoding unit 202B1...Arbitrary intra decoding unit 202B2...Alignment unit 202C...Mesh buffer unit 202D...Connection information decoding unit 202E...Inter decoding unit 202E1...Motion vector residual decoding unit 202E2...Motion vector buffer unit 202E3...Motion vector prediction unit 202E4...Motion vector calculation unit 202E5...Adder 202E6...Overlapping vertex search unit 202E7...mv_signaled_flag acquisition unit 202E8...Motion vector acquisition unit 202F...Skip decoding unit 203...Subdivision unit 204...Mesh decoding unit 205...Patch integration unit 206: Displacement amount decoding unit 207: Video decoding unit 208: Atlas data decoding unit

Claims

1. A mesh decoding device, comprising: an intra decoder that decodes vertex coordinates and connection information in an intra frame from a bit stream of the intra frame; and an inter decoder that decodes the coordinates of a vertex to be decoded by adding a motion vector decoded from a bit stream of an inter frame and the coordinates of a vertex corresponding to the vertex to be decoded in a reference frame, wherein the number of sub-meshes is different for each of the intra frames, and the number of sub-meshes is different for each of the inter frames.

2. Each of the intra decoder and the inter decoder assigns a non-overlapping sub-mesh ID to each of the sub-meshes, and assigns different sub-mesh IDs to corresponding sub-meshes between the intra frames and between the inter frames. The mesh decoding device according to claim 1, characterized in that.

3. The inter decoder according to claim 1, characterized in that it can only refer to sub-meshes having the same sub-net ID in the reference frame.

4. The inter decoder according to claim 1, characterized in that it can only refer to sub-meshes having the same number of vertices in the reference frame.

5. The inter decoder according to claim 1, characterized in that it can refer to a specified sub-mesh in the reference frame.

6. The inter decoder according to claim 5, characterized in that when there are a plurality of sub-meshes in the reference frame, it decodes a control signal for specifying the sub-mesh ID of the sub-mesh that can be referred to from the bit stream of the current sub-mesh. Mesh decoding device.

7. The inter decoder according to claim 5, characterized in that when there is only one sub-mesh in the reference frame, the sub-mesh is used as a sub-mesh that can be referred to. Mesh decoding device.

8. The inter decoder according to claim 6, characterized in that when the control signal does not exist, the sub-mesh ID of the sub-mesh that can be referred to is set to the same SubmeshID as the sub-mesh in the current frame. Mesh decoding device.

9. A mesh decoding method, comprising: step A of decoding vertex coordinates and connection information in an intra-frame from a bit stream of the intra-frame; and step B of decoding the coordinates of a vertex to be decoded by adding a motion vector decoded from a bit stream of an inter-frame and the coordinates of a vertex corresponding to the vertex to be decoded in a reference frame, wherein in step A and step B, the number of sub-meshes is different for each of the intra-frames, and the number of sub-meshes is different for each of the inter-frames.

10. A program for causing a computer to function as a mesh decoding device, wherein the mesh decoding device includes an intra-decoding unit that decodes vertex coordinates and connection information in an intra-frame from a bit stream of the intra-frame, and an inter-decoding unit that decodes the coordinates of a vertex to be decoded by adding a motion vector decoded from a bit stream of an inter-frame and the coordinates of a vertex corresponding to the vertex to be decoded in a reference frame, and the number of sub-meshes is different for each of the intra-frames, and the number of sub-meshes is different for each of the inter-frames.

Citation Information

Patent Citations

  • Image / video-based mesh compression

    US20230290008A1

  • Video based mesh compression

    WO2022074515A1