Mesh decoding apparatus, mesh decoding method, and program
The mesh decoding device and method address the issue of ensuring a basic mesh has at least one face by using intra and inter decoding units to manage vertex coordinates and connection information, enhancing decoding efficiency and accuracy.
Patent Information
- Application Number
- JP2024000910
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-06
- Publication Date
- 2025-07-17
AI Technical Summary
Existing mesh decoding technologies fail to guarantee that the basic mesh has at least one face, leading to inefficiencies and potential errors in the decoding process.
A mesh decoding device and method that includes an intra decoding unit to decode vertices in an intra frame and an inter decoding unit to combine motion vectors with reference frame coordinates, ensuring that both intra and inter frames have at least one face by using techniques like Draco and arithmetic coding to manage vertex indices and connection information.
Ensures that the basic mesh has at least one face, improving decoding efficiency and preventing errors by maintaining the integrity of the mesh structure during the decoding process.
Smart Images

Figure 2025107109000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a mesh decoding device, a mesh decoding method, and a program.
Background Art
[0002] Non-Patent Document 1 or Non-Patent Document 4 discloses a technique for encoding a mesh using Non-Patent Document 2 or 3 according to the framework of Non-Patent Document 5.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
Non-Patent Document 5
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the prior art, there is a problem that it cannot be guaranteed that the basic mesh has at least one face. Therefore, the present invention has been made in view of the above problems, and an object thereof is to provide a mesh decoding device, a mesh decoding method, and a program capable of guaranteeing that the above-described basic mesh has at least one face.
Means for Solving the Problems
[0005] A first feature of the present invention is a mesh decoding device including an intra decoding unit that decodes the coordinates and connection information of vertices in the intra frame from a bit stream of the intra frame, and an inter decoding unit that decodes the coordinates of the vertices of the vertex to be decoded by adding the motion vector decoded from the bit stream of the inter frame and the coordinates of the vertex corresponding to the vertex to be decoded in the reference frame. The gist is that the basic meshes in the intra frame and the inter frame have at least one face.
[0006] The second feature of the present invention is a mesh decoding method, which includes step A of decoding the coordinates and connection information of vertices in the intra-frame from the bitstream of the intra-frame, and step B of decoding the coordinates of the vertex to be decoded by adding the motion vector decoded from the bitstream of the inter-frame and the coordinates of the vertex corresponding to the vertex to be decoded in the reference frame. In step A and step B, the basic mesh in the intra-frame and the inter-frame has at least one surface as the gist.
[0007] The third feature of the present invention is a program for causing a computer to function as a mesh decoding device. The mesh decoding device includes an intra-decoding unit that decodes the coordinates and connection information of vertices in the intra-frame from the bitstream of the intra-frame, and an inter-decoding unit that decodes the coordinates of the vertex to be decoded by adding the motion vector decoded from the bitstream of the inter-frame and the coordinates of the vertex corresponding to the vertex to be decoded in the reference frame. The basic mesh in the intra-frame and the inter-frame has at least one surface as the gist.
Effect of the Invention
[0008] According to the present invention, it is possible to provide a mesh decoding device, a mesh decoding method, and a program that can guarantee that the basic mesh has at least one surface.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10A
Figure 10B
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the components in the following embodiments can be replaced with existing components as appropriate, and various variations including combinations with other existing components are possible. Therefore, the description of the following embodiments does not limit the content of the invention described in the claims.
[0011] <First Embodiment> Hereinafter, with reference to FIGS. 1 to 24, the mesh processing system according to this embodiment will be described.
[0012] FIG. 1 is a diagram showing an example of the configuration of a mesh processing system 1 according to this embodiment. As shown in FIG. 1, the mesh processing system 1 includes a mesh encoding device 100 and a mesh decoding device 200.
[0013] FIG. 2 is a diagram showing an example of the functional blocks of the mesh decoding device 200 according to this embodiment.
[0014] As shown in FIG. 2, the mesh decoding device 200 includes a multiplex separation unit 201, a basic mesh decoding unit 202, a sub-division unit 203, a mesh decoding unit 204, a patch integration unit 205, a displacement amount decoding unit 206, a video decoding unit 207, and an atlas data decoding unit 208.
[0015] Here, the basic mesh decoding unit 202, the sub-division unit 203, the mesh decoding unit 204, and the displacement amount decoding unit 206 are configured to perform processing in units of patches obtained by dividing a mesh, and thereafter, the patch integration unit 205 may be configured to integrate the processing results thereof.
[0016] In the example of FIG. 3A, the mesh is divided into patch 1 composed of basic surfaces 1 and 2 and patch 2 composed of basic surfaces 3 and 4.
[0017] The multiplex separation unit 201 is configured to separate the multiplexed bit stream into a basic mesh bit stream, a displacement amount bit stream, a texture bit stream, and an atlas bit stream.
[0018] The atlas data decoding unit 208 is configured to decode the atlas bit stream and output control information. Such a control signal may be used as metadata in the basic mesh decoding unit 202, the subdivision unit 203, the mesh decoding unit 204, the displacement amount decoding unit 206, and the video decoding unit 207.
[0019] <Basic mesh decoding unit 202> The basic mesh decoding unit 202 is configured to decode the basic mesh bit stream, generate a basic mesh, and output it.
[0020] Here, the basic mesh is composed of a plurality of vertices in a three-dimensional space and edges connecting the plurality of vertices.
[0021] As shown in FIG. 3A, the basic mesh is configured by combining basic faces represented by three vertices.
[0022] The basic mesh decoding unit 202 may be configured to decode the basic mesh bit stream using, for example, Draco shown in Non-Patent Document 2 or the technique described in Non-Patent Document 3.
[0023] Further, the basic mesh decoding unit 202 may be configured to generate "subdivision_method_id" described later as control information for controlling the type of subdivision method.
[0024] As shown in FIG. 4, the basic mesh decoding unit 202 includes a separation unit 202A, an intra decoding unit 202B, a mesh buffer unit 202C, a connection information decoding unit 202D, and an inter decoding unit 202E.
[0025] The separation unit 202A is configured to classify the basic mesh bit stream into the bit stream of an I frame and the bit stream of a P frame.
[0026] (Intra decoding unit 202B) The intra decoding unit 202B is configured to decode the coordinates and connection information of the vertices of the I frame from the bit stream of the I frame, for example, using Draco shown in Non-Patent Document 2 or the technology described in Non-Patent Document 3.
[0027] FIG. 5 is a diagram showing an example of the functional block of the intra decoding unit 202B.
[0028] As shown in FIG. 5, the intra decoding unit 202B includes an arbitrary intra decoding unit 202B1 and an alignment unit 202B2.
[0029] The arbitrary intra decoding unit 202B1 is configured to decode the coordinates and connection information of the unordered vertices of the I frame from the bit stream of the I frame using an arbitrary method including Draco shown in Non-Patent Document 2 or the technology described in Non-Patent Document 3.
[0030] The alignment unit 202B2 is configured to output vertices by rearranging the unordered vertices in a predetermined order.
[0031] As the predetermined order, for example, the Morton code order may be used, or the raster scan order may be used.
[0032] Further, the alignment unit 202B2 may group duplicate vertices, which are a plurality of vertices having the same coordinates in the decoded basic mesh, into a single vertex and then rearrange them in a predetermined order.
[0033] The mesh buffer unit 202C is configured to store the vertex coordinates and connection information of the I-frame decoded by the intra decoding unit 202B. Here, a specific buffer may be provided for storing pairs of vertex indices A(k) and B(k) that exist as duplicate vertices in a predetermined order.
[0034] The connection information decoding unit 202D is configured to convert the connection information of the I-frame or reference frame retrieved from the mesh buffer unit 202C into the connection information of the P-frame.
[0035] The inter decoding unit 202E is configured to decode the vertex coordinates of the P-frame by adding the vertex coordinates of the reference frame retrieved from the mesh buffer unit 202C and the motion vectors decoded from the bit stream of the P-frame.
[0036] Furthermore, the inter decoding unit 202E can adjust the vertex indices of the P-frame based on the pairs of vertex indices A(k) and B(k) that exist as duplicate vertices stored in such a specific buffer.
[0037] Here, all or part of the above-mentioned indices are decoded from the bit stream. Such a decoding method may be arithmetic coding. As a result, it is possible to expect the effect that there is no limit to the maximum value of the index to be decoded using arithmetic coding.
[0038] For example, arithmetic coding such as ue(v) may be used. ue(v) indicates an unsigned integer zero-order exponential Golomb coding (Exp-Golomb) with the leftmost bit first.
[0039] Specifically, the parsing process of the syntax element of ue(v) starts from the current position in the bit stream, reads the bits including the first non-zero bit, and starts by counting the number of leading bits equal to 0. This process is specified as follows. leadingZeroBits=-1 for(b = 0;!b; leadingZeroBits++) b = read_bits(l) Next, the variable codeNum is assigned as follows.
[0040] codeNum = 2 leadingZeroBits -1 + read_bits(leadingZeroBits) However, the value returned by read_bits(leadingZeroBits) is interpreted as the binary representation of an unsigned integer with the most significant bit written first. Also, the value of ue(v) is equal to the value of codeNum.
[0041] Table 1 shows the structure of the Exp-Golomb code by separating the bit sequence into "prefix" bits and "suffix" bits.
[0042]
Table 1
[0043] Here, the "prefix" bits are the bits parsed as specified in the calculation of leadingZeroBits and are shown as 0 or 1 in the columns of the bit sequence in Table 1.
[0044] The "suffix" bits are the bits parsed in the calculation of codeNum and are shown as x i in Table 1. i ranges from 0 to leadingZeroBits - 1. Each x i is equal to either 0 or 1.
[0045] Table 2 shows how to explicitly assign the bit sequence to the value of codeNum. Here, the value of ue(v) is equal to the value of codeNum.
[0046]
Table 2
[0047] In this embodiment, as shown in FIG. 6, there is a correspondence relationship between the vertices of the basic mesh of the P frame and the vertices of the basic mesh of the reference frame (I frame or P frame). Here, the motion vector decoded by the inter-decoding unit 202E is the difference vector between the coordinates of the vertices of the basic mesh of the P frame and the coordinates of the vertices of the basic mesh of the I frame.
[0048] Note that the inter-decoding unit 202E may decode the number of vertices of the current frame or the current sub-mesh from the bit stream.
[0049] Here, when the inter-decoding unit 202E decodes the number of vertices of the current frame or the current sub-mesh from the above-mentioned bit stream, the requirement for the conformity of such a bit stream is that the number of vertices of the decoded current frame or current sub-mesh must be equal to the number of vertices of the reference frame or the reference sub-mesh.
[0050] Note that when the inter-decoding unit 202E decodes the number of vertices of the current frame or the current sub-mesh from the above-mentioned bit stream, if the number of vertices of the decoded current frame or current sub-mesh is different from the number of vertices of the reference frame or the reference sub-mesh, it is configured to preferentially use the number of vertices of the current frame or the current sub-mesh decoded from the bit stream.
[0051] Also, the inter-decoding unit 202E may use the number of vertices of the reference frame or the reference sub-mesh as the number of vertices of the current frame or the current sub-mesh as it is.
[0052] In such a case, when the number of vertices of the current frame or the current sub-mesh is larger than the number of vertices of the reference frame or the reference sub-mesh, the inter-decoding unit 202E may add dummy vertices and connection information of the dummy vertices to the reference frame or the reference sub-mesh.
[0053] Here, the inter-decoding unit 202E may use fixed values (for example, (0, 0, 0)) for the coordinates of such dummy vertices, or may copy them from a predetermined vertex (for example, the last vertex of the reference frame).
[0054] Also, the inter-decoding unit 202E may copy the necessary portion from the beginning of the connection information of the vertices in the reference frame or the reference sub-mesh as the connection information of the dummy vertices.
[0055] According to such a configuration, an effect of ensuring the decoding operation of the current frame or the current sub-mesh can be expected.
[0056] Note that since the basic mesh has at least one face and such a face has at least three or more vertices, the control signal indicating the number of vertices for each frame or each sub-mesh is limited to include at least three or more vertices.
[0057] For example, the control signal pdu_vertex_count_minus_1[titleID][patchIdx] defined in Section 8.3.7.3 of Non-Reference Document 4, the control signal sismu_inter_vertex_count[subMeshID] defined in Section H.8.1.3.8, and the control signal mesh_vertex_count defined in Section I.8.3.7 are changed as shown in Table 3 and limited to include three or more vertices.
[0058] Here, the control signal pdu_vertex_count_minus_1[titleID][patchIdx] is a control signal that specifies the number of vertices in a patch having a patch ID equal to the patch ID specified by [patchIdx] within an atlas tile having a tile ID equal to the tile ID specified by [titleID].
[0059] The control signal sismu_inter_vertex_count[subMeshID] is a control signal that specifies the number of vertices in the submesh having the SubmeshID equal to the SubmeshID specified by [subMeshID].
[0060] The control signal mesh_vertex_count is a control signal that specifies the number of vertices in the decoded mesh.
[0061]
Table 3
[0062] According to such a configuration, an effect can be expected that it is possible to prevent a situation where vertices operate with data having no meaning such as 1 or 2.
[0063] (Inter-decoding unit 202E) FIG. 7 is a diagram showing an example of the functional blocks of the inter-decoding unit 202E.
[0064] As shown in FIG. 7, the inter-decoding unit 202E includes a motion vector residual decoding unit 202E1, a motion vector buffer unit 202E2, a motion vector prediction unit 202E3, a motion vector calculation unit 202E4, and an adder 202E5.
[0065] The motion vector residual decoding unit 202E1 is configured to generate an MVR (Motion Vector Residual) from the bit stream of the P frame.
[0066] Here, the MVR is a motion vector residual indicating the difference between an MV (Motion Vector) and an MVP (Motion Vector Prediction). The MV is a difference vector (motion vector) between the coordinates of the vertex of the corresponding I frame and the coordinates of the vertex of the P frame. The MVP is a predicted value of the MV of the target vertex (predicted value of the motion vector) using the MV.
[0067] The motion vector buffer unit 202E2 is configured to sequentially store the MVs output by the motion vector calculation unit 202E4.
[0068] The motion vector prediction unit 202E3 acquires the decoded MVs from the motion vector buffer unit 202E2 for the vertices connected to the vertex to be decoded, and as shown in FIG. 8, outputs the MVP of the vertex to be decoded using all or part of the acquired decoded MVs.
[0069] The motion vector calculation unit 202E4 adds the MVR generated by the motion vector residual decoding unit 202E1 and the MVP output from the motion vector prediction unit 202E3, and is configured to output the MV of the vertex to be decoded.
[0070] The adder 202E5 adds the coordinates of the vertex corresponding to the vertex to be decoded obtained from the decoded basic mesh of the reference frame (I frame or P frame) having a correspondence relationship and the motion vector MV output from the motion vector calculation unit 202E3, and is configured to output the coordinates of the vertex to be decoded.
[0071] Hereinafter, the details of each unit of the inter-decoding unit 202E will be described.
[0072] FIG. 9 shows a flowchart showing an example of the operation of the motion vector prediction unit 202E3. Hereinafter, the operation of the motion vector prediction unit 202E3 is referred to as the "average prediction method".
[0073] As shown in FIG. 9, in step S1001, the motion vector prediction unit 202E3 sets 0 for MVP and N.
[0074] In step S1002, the motion vector prediction unit 202E3 acquires a set of MVs of the vertices around the vertex to be decoded from the motion vector buffer unit 202E2, identifies the vertices for which the subsequent processing has not been completed, and transitions to No. If the subsequent processing has been completed for all vertices, it transitions to Yes.
[0075] In step S1003, if the MV of the vertex to be processed is not decoded, the motion vector prediction unit 202E3 transitions to No, and if the MV of the vertex to be processed is decoded, it transitions to Yes.
[0076] In step S1004, the motion vector prediction unit 202E3 adds the MV to the MVP and adds 1 to N.
[0077] In step S1005, if N is greater than 0, the motion vector prediction unit 202E3 outputs the result of dividing the MVP by N, and if N is 0, it outputs 0 and ends the process.
[0078] That is, the motion vector prediction unit 202E3 is configured to output the MVP to be decoded by averaging the decoded motion vectors of the vertices around the vertex to be decoded.
[0079] Note that the motion vector prediction unit 202E3 may be configured to set the MVP to 0 when the set of such decoded motion vectors is an empty set.
[0080] The motion vector calculation unit 202E4 may be configured to calculate the MV of the vertex to be decoded from the MVP output by the motion vector prediction unit 202E3 and the MVR generated by the motion vector residual decoding unit 202E1 according to equation (1).
[0081] MV(k) = MVP(k) + MVR(k) … (1) Here, k is the index of the vertex. MV, MVR, and MVP are vectors having x, y, and z components.
[0082] According to such a configuration, since only the MVR is encoded instead of the MV using the MVP, an effect of improving the encoding efficiency can be expected.
[0083] The adder 202E5 is configured to calculate the coordinates of a vertex by adding the MV of the vertex calculated by the motion vector calculation unit 202E4 and the coordinates of the vertex of the reference frame corresponding to such vertex, and to keep the connection information as that of the reference frame.
[0084] Specifically, the adder 202E5 may be configured to calculate the coordinates v’ i (k) of the k-th vertex using Equation (2).
[0085] v’ i (k) = v’ j (k) + MV(k) … (2) Here, v’ i (k) is the coordinates of the k-th vertex to be decoded in the frame to be decoded, v’ j (k) is the coordinates of the decoded k-th vertex of the reference frame, MV(k) is the k-th MV of the frame to be decoded, and k = 1, 2…, K.
[0086] Also, the connection information of the frame to be decoded is made the same as the connection information of the reference frame.
[0087] Note that since the motion vector prediction unit 202E3 calculates the MVP using the decoded MV, the decoding order affects the MVP.
[0088] Such decoding order is set to the decoding order of the vertices of the basic mesh of the reference frame. Generally, if a decoding method that increases the basic surface one by one from the starting edge using a certain repeating pattern is used, the order of the decoded vertices of the basic mesh is determined in the decoding process.
[0089] For example, the motion vector prediction unit 202E3 may determine the decoding order of the vertices using Edgebreaker in the basic mesh of the reference frame.
[0090] According to such a configuration, since the MV from the reference frame is encoded instead of the coordinates of the vertex, an effect of improving the encoding efficiency can be expected.
[0091] (Modification Example 1 of Inter Decoder 202E) Hereinafter, Modification Example 1 of Inter Decoder 202E will be described.
[0092] In the "average prediction method" where the motion vector prediction unit 202E3 of the inter decoder 202E averages the decoded motion vectors of the vertices around the vertex to be decoded, in order not to exceed the predetermined maximum number of uses, all or only a part of the decoded motion vectors of the vertices around the vertex to be decoded are used to calculate the MVP.
[0093] Note that the predetermined maximum number of uses is decoded from the bit stream as a control signal.
[0094] In addition, when the number of decoded motion vectors of the vertices around the vertex to be decoded exceeds the maximum number of uses, the motion vector prediction unit 202E3 picks up to the maximum number of uses according to a certain rule.
[0095] For example, as such a rule, the motion vector prediction unit 202E3 selects the first or last vertex in the decoding order.
[0096] The decoding order for the mesh as shown in FIG. 10A is, as indicated by the arrow, vertex v D →v C →v A →v B That is.
[0097] FIG. 10B is a list of vertices around the vertex to be decoded that are used when calculating the MVP of each vertex v A ~v D when the maximum value of the number of decoded adjacent vertices is set to 3.
[0098] According to such a configuration, by determining the maximum number of adjacent vertices, it is possible to expect the effect of reducing the amount of calculation and the amount of memory while maintaining or reducing the coding efficiency.
[0099] However, in order to exhibit the above-described effect, it is necessary to set an appropriate maximum number of adjacent vertices in the mesh encoding device 100 and write it into the bit stream as a related control signal.
[0100] Therefore, the range that can be set as the above-described maximum number of adjacent vertices is encoded / decoded so that the maximum number of adjacent vertices is less than or equal to a predetermined maximum value as a reasonable constraint regarding the maximum number of adjacent vertices in order to determine the amount of memory prepared in the mesh decoding device 200.
[0101] In this way, it is possible to expect the effect of facilitating the design of the mesh decoding device 200 by defining a reasonable constraint regarding the maximum number of adjacent vertices.
[0102] Generally, the average number of adjacent vertices in a Closed 2-manifold triangle mesh is about 6. Statistically, however, the maximum number of adjacent vertices is often 7 to 8. As shown in FIG. 11, the number of decoded motion vectors (vertical axis) changes dynamically according to the number of vertices (horizontal axis) around the vertex to be decoded.
[0103] Therefore, it is desirable to narrow down the range that can be set as the above-described maximum number of adjacent vertices.
[0104] For example, as shown in FIG. 11, statistically, the number "3" of vertices around the vertex to be decoded having the largest number of decoded motion vectors is included within the range that can be set as the maximum number of adjacent vertices in the above-described control signal, or a certain ratio (for example, 50% or 120%) of the average of the statistical number of adjacent vertices or a value not greater than a natural number that can cover up to N bits (for example, 3 bits) is set as the upper limit (maximum value) of the range that can be set as the maximum number of adjacent vertices in the above-described control signal, whereby the effect of reducing the amount of calculation and the amount of memory can be exhibited.
[0105] On the other hand, if the range that can be set as the maximum number of adjacent vertices is set to a large value, for example, 256 or 8 bits in the worst case, it may not be possible to reduce not only the memory amount but also the computational amount.
[0106] FIG. 12 shows an example of the worst case, where when n ≧ 256, the number of decoded adjacent vertices exceeds 256. In FIG. 12, the number of decoded adjacent vertices at vertex n + 1 is n.
[0107] If the upper limit of the maximum number of adjacent vertices is set to 256, the mesh decoding device 200 requires not only a huge amount of memory but also a huge amount of computational amount as shown in FIG. 10B. Therefore, the upper limit (maximum value) of the range that can be set as the above-mentioned maximum number of adjacent vertices may be 8.
[0108] Furthermore, the range that can be set as the maximum number of adjacent vertices in the above control signal may be a clear value or may be calculated from other control signals and data.
[0109] For example, Level1 may define the range that can be set as the maximum number of adjacent vertices in the control signal.
[0110] Alternatively, the upper limit of the range that can be set as the maximum number of adjacent vertices in the control signal may be calculated from the number of vertices of the basic mesh by the following formula (3).
[0111] Upper limit of the range that can be set as the maximum number of adjacent vertices in the control signal = log2 (number of vertices of the basic mesh) … Formula (3) According to such a configuration, the range that can be appropriately set as the maximum number of adjacent vertices can be determined, and the effect of surely reducing both the computational amount and the memory amount can be expected even in the worst case.
[0112] The inter-decoding unit 202E decodes the coordinates of the vertex to be decoded by adding the MV decoded from the bitstream of the inter-frame and the coordinates of the vertex corresponding to the vertex to be decoded in the reference frame.
[0113] Further, the motion vector prediction unit 202E3 calculates a predicted value of the motion vector of the vertex to be decoded by referring to the adjacent vertex list and averaging all or part of the motion vectors of the decoded vertices adjacent to the vertex to be decoded. Here, the adjacent vertex list is a list of vertices adjacent to each vertex.
[0114] The motion vector prediction unit 202E3 can reuse the adjacent vertex list in the reference frame as the adjacent vertex list in the current frame. Here, it is assumed that such an adjacent vertex list includes the decoded vertices picked up up to the maximum usage number (the set maximum number of adjacent vertices).
[0115] Here, when the condition (reuse condition) that each vertex in the reference frame and each vertex in the current frame have a one-to-one correspondence, the decoding order of each vertex in the reference frame and the decoding order of each vertex in the current frame are the same, and an adjacent vertex list including the decoded vertices picked up up to the maximum usage number in the reference frame is already saved is satisfied, the adjacent vertex list in the reference frame can be reused.
[0116] For example, when the reference frame is decoded, if the reference frame is a P frame or an S frame, there exists an adjacent vertex list including the decoded vertices picked up up to the maximum usage number.
[0117] Therefore, when the reference frame is decoded, if the reference frame is a P frame or an S frame, the motion vector prediction unit 292E3 saves the decoded reference frame in the reference frame buffer and saves the adjacent vertex list in the reference frame in a specific buffer.
[0118] When the current frame is decoded, if the type of the reference frame is a P frame or an S frame, the motion vector prediction unit 292E3 reuses the adjacent vertex list in the reference frame stored in the specific buffer as the adjacent vertex list in the current frame.
[0119] However, the specific buffer may store an adjacent vertex list including decoded vertices picked up up to the maximum usage number in one or more frames.
[0120] Therefore, when the adjacent vertex lists in a plurality of frames are stored in the specific buffer, the motion vector prediction unit 202E3 reuses the adjacent vertex list in the frame corresponding to the reference frame (the frame including the same frame index as the reference frame) as the adjacent vertex list in the current frame.
[0121] The motion vector prediction unit 202E3 applies the operations on the reference frame buffer to the specific buffer as well. Here, the operations on the reference frame buffer are, for example, the marking process described in Chapter 9.2.4.4 of Non-Patent Document 4 or Non-Patent Document 5.
[0122] According to such a configuration, an effect of reducing the computational complexity for creating an adjacent vertex list including decoded vertices picked up up to the maximum usage number in the current frame can be expected.
[0123] However, in such a case, a specific buffer for storing the adjacent vertex lists in one or more frames is required. Here, in order to minimize the size of the specific buffer, additional conditions may be added to the above reuse conditions.
[0124] For example, in addition to the above reuse conditions, when the condition that the reference frame is the frame immediately before the current frame in decoding order is satisfied, the motion vector prediction unit 202E3 can reuse the adjacent vertex list in the reference frame as the adjacent vertex list in the current frame.
[0125] In such a case, the motion vector prediction unit 202E3 may save only the adjacent vertex list in the frame immediately before the current frame in decoding order in the specific buffer. Further, the motion vector prediction unit 202E3 does not apply operations on the reference frame buffer to the specific buffer. According to such a configuration, an effect of reducing the size of the specific buffer can be expected.
[0126] (Modification Example 2 of Inter Decoder 202E) Hereinafter, with reference to FIG. 13, a modification example 2 of the inter decoder 202E will be described.
[0127] The motion vector calculation unit 202E4 of the inter decoder 202E has mode 1 and mode 0.
[0128] In mode 1, the motion vector calculation unit 202E4 adds the MVR generated by the motion vector residual decoding unit 202E1 and the MVP output from the motion vector prediction unit 202E3, and outputs the MV of the vertex to be decoded (see A in FIG. 13).
[0129] On the other hand, in mode 0, the motion vector calculation unit 202E4 outputs the MVR generated by the motion vector residual decoding unit 202E1 as the MV of the vertex to be decoded (see B in FIG. 13).
[0130] Note that the operation of the motion vector calculation unit 202E4 in mode 0 corresponds to an operation of setting the MVP output from the motion vector prediction unit 202E3 to zero.
[0131] Furthermore, the motion vector calculation unit 202E4 may make the modes of the MVs of N (N≧1) consecutive vertices in decoding order the same.
[0132] The motion vector calculation unit 202E4 groups the above-mentioned N vertices into one group. The size of such a group (group size) N is 1 or more. The motion vector calculation unit 202E4 decodes a control signal (group size shown in FIG. 13) that can calculate such a group size from the bit stream.
[0133] However, when the number of vertices remaining in the last group is smaller than the group size, the motion vector calculation unit 202E4 will put all the remaining vertices into the group.
[0134] In this way, when N consecutive vertices are set to the same mode, the amount of code for the mode can be reduced, so an effect of improving the coding efficiency can be expected.
[0135] Here, the larger the number of consecutive vertices having the same mode, the greater the effect of reducing the amount of code for the mode. Therefore, it is necessary to set an appropriate group size in the mesh encoding device 100 and decode it from the bit stream as a control signal in the mesh decoding device 200.
[0136] Therefore, it is desirable that the range that can be set in such a control signal is not smaller than the actual number of consecutive vertices having the same mode.
[0137] For example, when almost all vertices select the same mode, the group size may be set to the total number of vertices.
[0138] Table 4 shows examples when the number of vertices selecting mode 0 is 80% or more, or when the number of vertices selecting mode 1 is 90% or more.
[0139] Therefore, make sure that the range that can be set in the above-mentioned control signal can cover from 1 to a predetermined maximum value. Such a maximum value is set to be equal to or greater than the total number of vertices of the basic mesh.
[0140]
Table 4
[0141] Incidentally, when the above control signal (group size) is a natural number and is set to be equal to or greater than the total number of vertices, since the absolute value is large, the amount of sign becomes large.
[0142] Therefore, it is also possible to use the logarithm of the above control signal. Specifically, the group size may be calculated by the following equation (4) with the control signal as log2_group_size.
[0143] group size = 2log2_group_size … Equation (4) Here, if there is only one group in the frame, that group is set as the last group. That is, if the group size is larger than the number of vertices, all vertices are put into the group.
[0144] Furthermore, the range that can be set in the above control signal may be a definite value or may be calculated from other control signals or data.
[0145] For example, the range that can be set in such a control signal may be defined by Level1.
[0146] Alternatively, the range that can be set in such a control signal may be calculated from the number of vertices of the basic mesh.
[0147] For example, regarding the range that can be set in such a control signal, it may be set to the smallest natural number that is a power of 2 and can cover the number of vertices of the basic mesh.
[0148] Furthermore, regarding the configurable range in the above control signal, after setting it to a small range, as shown in FIG. 14, a predetermined flag (Mode flag) of another control signal may be introduced. In such a case, as shown in FIG. 14, if such a predetermined flag is TRUE (Mode flag = 1), the motion vector calculation unit 202E4 groups all the vertices into one group (that is, sets the number of all vertices as the group size), and if it is FALSE, it keeps the group size calculated from the above control signal.
[0149] Note that the above control signal may be set for each sequence or for each frame. When the above control signal is set for each sequence, the group size of all frames is the same.
[0150] According to such a configuration, by appropriately determining the configurable range of the group size, it is possible to handle all situations, surely reduce the amount of mode bits, and expect the effect of improving the coding efficiency.
[0151] (Modification Example 3 of the Inter Decoder 202E) In a further modification example of the above inter decoder 202E, before implementing the above inter decoder 202E, it is configured to add the following functional blocks.
[0152] Specifically, as shown in FIG. 15, in addition to the configuration shown in FIG. 8, the inter decoder 202E includes a duplicate vertex search unit 202E6, an mv_signalled_flag acquisition unit (flag acquisition unit) 202E7, and a motion vector acquisition unit 202E8.
[0153] Here, the derived_my_present_flag (first flag) is included at the beginning of the bitstream of the P frame and has at least two values of 0 or 1.
[0154] Also, the mv_signalled_flag (second flag) is included in the bitstream of the P frame and has two values, 0 or 1, for each vertex when derived_my_present_flag indicates No.
[0155] When derived_my_present_flag indicates No (for example, when derived_my_present_flag is 0), the mv_signalled_flag acquisition unit 202E7 decodes the motion vectors of all vertices from the bitstream of the P frame and sets the value of my_signalled_flag to 1 without decoding my_signalled_flag of all vertices from the bitstream of the P frame.
[0156] When derived_my_present_flag indicates Yes (for example, when derived_my_present_flag is 1), the mv_signalled_flag acquisition unit 202E7 performs different processing for each vertex of the P frame. The mv_signalled_flag acquisition unit 202E7 may determine the processing method for each vertex using mv_signalled_flag.
[0157] Also, when derived_my_present_flag indicates Yes and the mv_signalled_flag of a certain vertex indicates Yes, the mv_signalled_flag acquisition unit 202E7 does not perform the processing in the motion vector acquisition unit 202E8 for the motion vector of such a vertex, and performs the same processing as the inter-decoding unit 202E shown in FIG. 7 or a modified example thereof.
[0158] Also, when derived_my_present_flag indicates Yes and the mv_signalled_flag of a certain vertex indicates No, the mv_signalled_flag acquisition unit 202E7 performs the processing in the motion vector acquisition unit 202E8 for the motion vector of such a vertex and acquires the motion vector of the vertex.
[0159] For example, when the mv_signalled_flag acquisition unit 202E7 determines that derived_my_present_flag indicates Yes, it decodes mv_signalled_flag for each vertex from the bit stream of the P frame.
[0160] When derived_my_present_flag indicates Yes and the mv_signalled_flag of a certain vertex indicates Yes, the mv_signalled_flag acquisition unit 202E7 sets the prediction mode (MV mode) of such vertex to 2.
[0161] On the other hand, when derived_my_present_flag indicates Yes and the mv_signalled_flag of a certain vertex indicates No, the mv_signalled_flag acquisition unit 202E7 sets the prediction mode of such vertex to a value other than 2.
[0162] Also, when derived_my_present_flag indicates No, the mv_signalled_flag acquisition unit 202E7 does not decode mv_signalled_flag of all vertices of the P frame from the bit stream, sets its value to 1, and sets the MV mode of the vertex to a value other than 2.
[0163] The duplicate vertex search unit 202E6 is configured to search for the indices of vertices with matching coordinates (hereinafter referred to as duplicate vertices) from the geometric information of the basic mesh of the decoded reference frame and store them in a buffer (not shown).
[0164] Specifically, the input to the duplicate vertex search unit 202E6 is the index (in decoding order) and position coordinates of each vertex of the basic mesh of the decoded reference frame.
[0165] In addition, the output of the duplicate vertex search unit 202E6 is a list that stores the index (vindex1) of such a duplicate vertex if there is a duplicate vertex related to the index (vindex0) of each vertex, and stores the index (vindex0) itself of the vertex or a specific value (e.g., "-1") not used as the index of each vertex if there is no such duplicate vertex. Here, such a list is stored in the buffer repVert in the order of index0.
[0166] Also, since the vertex of vindex1 is decoded before the vertex of vindex0, the relationship is vindex0 > vindex1.
[0167] The duplicate vertex search unit 202E6 determines whether there is a duplicate vertex related to the immediately preceding vertex (index: vindex0 - 1) from the first vertex (index: 0) of the decoded reference frame's basic mesh for each vertex (index: vindex0) of the reference frame's basic mesh. If it is determined that there is such a duplicate vertex, the index of such a duplicate vertex is output by at least one of the following three methods.
[0168] (Method 1) The duplicate vertex search unit 202E6 sequentially searches for duplicate vertices with matching coordinates as follows. When there is a duplicate vertex, vref is vindex1, and when there is no duplicate vertex, vRef is -1.
[0169] vRef = firstVertexIndexDuplicated(vindex0) where firstVertexIndexDuplicated(v) { for( i = 0; i < v; i++) { if(referenceSubmeshVertexPositions[ i ] == referenceSubmeshVertexPositions[ v ]) { return i } } return -1 } (Method 2) The duplicate vertex search unit 202E6 searches for duplicate vertices with matching coordinates using binary search. For example, the duplicate vertex search unit 202E6 may utilize the find function of the associative array class map. (Method 3) The duplicate vertex search unit 202E6 searches for duplicate vertices with matching coordinates using a hash table. For example, the duplicate vertex search unit 202E6 may utilize the find function of the hash associative array class unordered_map.
[0170] Note that as a method of finding duplicate vertices in the basic mesh of the reference frame, for vertices where duplicate vertices exist, a method of decoding the index of the duplicate vertex instead of the position coordinates from a special signal may be used.
[0171] When MVmode is 2 (when derived_my_present_flag indicates Yes and mv_signalled_flag of the vertex indicates No), since there are duplicate vertices for the vertex, the motion vector acquisition unit 202E8 acquires, from the motion vector buffer unit 202E2, the motion vector of the vertex having the index (vindex1) of the duplicate vertex corresponding to the index (vindex0) of the vertex output from the duplicate vertex search unit 202E6, and is configured such that the motion vector of such vertex is used as the motion vector of the vertex.
[0172] That is, it is the output of the index (vindex1) of the duplicate vertex 202E6 and is not decoded from the bitstream.
[0173] Here, when MVmode is other than 2 (when derived_my_present_flag indicates No, or when derived_my_present_flag indicates Yes and the mv_signalled_flag of the vertex indicates Yes), instead of the motion vector acquisition unit 202E8, the same processing as the inter-decoding unit 202E shown in FIG. 7 or its modified example is performed.
[0174] According to such a configuration, it is possible to expect the effect of reducing the decoding calculation of the motion vector and the amount of code for the vertex having overlapping vertices.
[0175] In a further modified example of the above-described inter-decoding unit 202E, the overlapping vertex search unit 202E6 does not search for overlapping vertices from all vertices of the basic mesh of the reference frame, but searches for overlapping vertices only from the vertices where mv_signalled_flag is No.
[0176] However, the input to the overlapping vertex search unit 202E6 includes the mv_signalled_flag in addition to the index (decoding order) and position coordinates of each vertex of the basic mesh of the decoded reference frame.
[0177] According to this modified example, since overlapping vertices are searched only for vertices having overlapping vertices instead of all vertices, it is possible to expect the effect of reducing the decoding calculation of the motion vector.
[0178] In a further modified example of the above-described inter-decoding unit 202E, the mv_signalled_flag acquisition unit 202E7 decodes the mv_signalled_flag in two steps.
[0179] The mv_signalled_flag acquisition unit 202E7 groups N vertices into one group, and for each group, decodes the mv_group_signalled_flag (the third flag) from the bit stream of the P frame. For all vertices in the group where mv_group_signalled_flag is 1, it sets mv_signalled_flag to 1, and for each vertex in the group where mv_group_signalled_flag is 0, it decodes mv_group_signalled_flag from the bit stream of the P frame.
[0180] According to this modification example, since the mv_group_signalled_flag is decoded in two stages, an effect of reducing the decoding calculation of the motion vector and the coding amount can be expected.
[0181] In a further modification example of the above-mentioned inter-decoding unit 202E, the mv_signalled_flag acquisition unit 202E7 decodes mv_signalled_flag from the bit stream of the P frame for vertices with duplicate vertices instead of for each vertex, and for vertices without duplicate vertices, it does not decode mv_signalled_flag from the bit stream of the P frame and sets mv_signalled_flag to 1.
[0182] Note that the inter-decoding unit 202E decodes a control signal indicating the number of vertices with duplicate vertices from the bit stream before mv_signalled_flag.
[0183] With such a control signal, it is possible to decode mv_signalled_flag without performing the process of the duplicate vertex search unit 202E6.
[0184] Also, as a requirement for the compatibility of the bit stream, it is specified that the number of vertices with duplicate vertices output by the duplicate vertex search unit 202E6 matches the number of vertices indicated by the control signal.
[0185] In addition, the duplicate vertex search unit 202E6 stores all duplicate vertex indices in another list. For example, the duplicate vertex search unit 202E6 may store all duplicate vertex indices in the duplicated_vertex_list as shown in the following (modified example of Method 1).
[0186] (Modified example of Method 1) The duplicate vertex search unit 202E6 sequentially searches for duplicate vertices with matching coordinates as shown below. When a duplicate vertex exists, vRef is vindex1, and when no duplicate vertex exists, vRef is -1.
[0187] vRef = firstVertexIndexDuplicated(vindex0, duplicated_vertex_list) where firstVertexIndexDuplicated(v, duplicated_vertex_list) { for (i = 0; i < v; i++) { if (referenceSubmeshVertexPositions[i] == referenceSubmeshVertexPositions[v]) { duplicated_vertex_list.push_back(v) return i } } return -1 } According to this modified example, since the mv_signalled_flag is provided only for vertices having duplicate vertices instead of all vertices, a decoding calculation of motion vectors and a reduction effect of the coded amount can be expected.
[0188] In a further modified example of the above-mentioned inter-decoding unit 202E, the mv_signalled_flag acquisition unit 202E7 may set the mv_signalled_flag of vertices having no duplicate vertices to 0.
[0189] Here, when the mv_signalled_flag of a vertex without duplicate vertices is 0, in the basic mesh of the reference frame, although there are no duplicate vertices, there are vertices with the same motion vector. Therefore, the index of the vertex with the same motion vector as the vertex is decoded from the bitstream of the inter-frame, and the motion vector of the vertex is obtained.
[0190] Specifically, when MVmode is 2 (when derived_my_present_flag indicates Yes and the mv_signalled_flag of the vertex indicates No), when there are duplicate vertices of such a vertex, from the motion vector buffer unit 202E2, the motion vector of the vertex having the index (vindex1) of the duplicate vertex corresponding to the index (vindex0) of the vertex output by the duplicate vertex search unit 202E6 is obtained, and the motion vector of such a vertex is configured to be the motion vector of the vertex. However, in this modification example, when there are no duplicate vertices of the vertex, the index (vindex1) of the vertex having the same motion vector as the vertex is decoded from the bitstream, the motion vector of the vertex having such an index (vindex1) is obtained, and the motion vector of such a vertex is configured to be the motion vector of the vertex.
[0191] According to this modification example, even for a vertex without duplicate vertices, when the mv_signalled_flag indicates No, since the motion vector is obtained from another vertex, an effect of reducing the decoding calculation of the motion vector and the amount of code can be expected.
[0192] Note that there are cases where the above modification examples can be used simultaneously and cases where they cannot. In the case where they cannot be used simultaneously, a control signal indicating which modification example to use is provided, the control signal is decoded from the bitstream, and a determination is made as to which modification example to use. However, such a control signal may be an extension of an existing control signal.
[0193] In a further modification example of the above-mentioned inter-decoding unit 202E, if the reference frame is an inter-frame, the duplicate vertex search unit 202E6 re-uses the results obtained in the reference frame in the frame to be decoded.
[0194] However, if the reference frame is an intra-frame, the duplicate vertex search unit 202E6 may re-use such information on the assumption that it has information on duplicate vertices.
[0195] Specifically, first, the duplicate vertex search unit 202E6 decodes a control signal from the bit stream indicating whether to re-use the results obtained in the reference frame (information on duplicate vertices including the number of vertices having duplicate vertices) in the frame to be decoded. However, for such a control signal, the duplicate vertex search unit 202E6 may use an existing control signal as it is or after extension.
[0196] Second, when such a control signal is Yes, if the reference frame is an inter-frame, the duplicate vertex search unit 202E6 re-uses the results obtained in the reference frame in the frame to be decoded. However, if the reference frame is an intra-frame, the duplicate vertex search unit 202E6 may re-use such information on the assumption that it has information on duplicate vertices.
[0197] Specifically, the output of the duplicate vertex search unit 202E6 related to the reference frame (the above-mentioned results) is a list that stores the index (vindex1) of such a duplicate vertex if a duplicate vertex exists for the index (vindex0) of each vertex, and stores the index (vindex0) of the vertex itself or a specific value (e.g., -1) for which the index cannot be used if such a duplicate vertex does not exist.
[0198] Here, such a list is stored in the buffer repVert in the order of vindex0.
[0199] In the following example, when there are no duplicate vertices, the index (vindex0) of the vertex itself is saved. Also, to clarify the index tRef of the frame, it is displayed in the buffer repVert tRef as follows.
[0200] repVert tRef (vindex0) = vRef vRef = vindex1 if vindex0 and vindex1 are duplicate vertices vRef = vindex0 if vindex0 and vindex1 are not duplicate vertices In the frame t to be decoded, when repVert tRef is reused, the duplicate vertex search unit 202E6 does not need to search for duplicate vertices for each vertex (index: vindex0) of the basic mesh of the reference frame of the frame to be decoded, and repVert tRef can be used as it is.
[0201] repVert t (vindex0) = vindex0 if repVert tRef (vindex0) = vindex0 repVert t (vindex0) = repVert tRef (vindex0) if repVert tRef (vindex0)!= vindex0 According to this modification example, since duplicate vertices are not searched for, an effect of reducing the decoding calculation of motion vectors can be expected.
[0202] In a further modification example of the above-mentioned inter-decoding unit 202E, if the reference frame is an inter-frame, the mv_signalled_flag acquisition unit (flag acquisition unit) 202E7 reuses the mv_signalled_flag acquired in the reference frame in the frame to be decoded.
[0203] However, if the reference frame is an intra-frame, the mv_signalled_flag acquisition unit 202E7 may reuse such mv_signalled_flag on the assumption that it has the mv_signalled_flag.
[0204] Specifically, when derived_my_present_flag indicates Yes, the mv_signalled_flag acquisition unit 202E7 decodes, for each vertex, not the mv_signalled_flag from the P-frame bit stream but the difference from the mv_signalled_flag of the reference frame.
[0205] Also, a control signal may be provided when there is no difference in all mv_signalled_flags. In that case, the mv_signalled_flag acquisition unit 202E7 decodes such control signal, and if such control signal is a specific value (for example, TRUE), uses the mv_signalled_flag of the reference frame as it is as the mv_signalled_flag of the target frame, and if such control signal is a specific value (for example, FALSE), further decodes the difference from the mv_signalled_flag of the reference frame to calculate the mv_signalled_flag of the target frame.
[0206] According to this modification example, an effect of reducing the amount of sign of mv_signalled_flag can be expected.
[0207] Note that there are cases where the above modification examples can be used simultaneously and cases where they cannot. In the case where they cannot be used simultaneously, a control signal indicating which modification example to use is provided, the control signal is decoded from the bit stream, and a determination is made as to which modification example to use.
[0208] (Modification Example 1 of the Basic Mesh Decoding Unit 202) Hereinafter, with reference to FIGS. 16 and 17, Modification Example 1 of the basic mesh decoding unit 202 will be described.
[0209] As shown in FIG. 16, the basic mesh decoding unit 202 according to the first modification example includes a separation unit 202A, an intra decoding unit 202B, a mesh buffer unit 202C, an inter decoding unit 202E, and a skip decoding unit 202F.
[0210] The skip decoding unit 202F is configured to decode the basic mesh of the frame to be decoded by directly using the decoded basic mesh of the specified reference frame.
[0211] In this embodiment, the frame may be either a mesh or a submesh.
[0212] For example, as shown in FIG. 17, "P_SUBMESH" in smh_type may correspond to a P frame, "I_SUBMESH" in smh_type may correspond to an I frame, and "SKIP_SUBMESH" in smh_type may correspond to an S frame.
[0213] (Skip decoding unit 202F) The skip decoding unit 202F is configured to extract the decoded basic mesh (reference decoded basic mesh) of the specified reference frame from the mesh buffer unit 202C, and use the coordinates of the vertices of the extracted reference decoded basic mesh and the indices of the vertices as they are to decode the coordinates of the vertices of the basic mesh of the frame to be decoded and the indices of the vertices.
[0214] Here, the mesh buffer unit 202C has at least one reference frame and is configured to store at least one decoded basic mesh for each reference frame.
[0215] The skip decoding unit 202F may identify the specified reference decoded basic mesh by using a control signal decoded from the bitstream or a predetermined rule.
[0216] For example, such a predetermined rule may be to extract the first reference frame from the reference frame list from the mesh buffer unit 202C, or to extract the reference frame with the closest frame index to the frame to be decoded, etc.
[0217] In this embodiment, a frame that decodes the coordinates of the vertices of the basic mesh by directly using the coordinates of the vertices of the reference decoding basic mesh and the indices of the vertices is called an "S frame".
[0218] According to such a configuration, the motion vector can be made unnecessary in the skip decoding unit 202F, so that a significant reduction effect in the amount of code and a significant reduction effect in the amount of calculation can be expected.
[0219] (Mesh buffer unit 202C) The mesh buffer unit 202C is configured to store one or more reference decoding basic meshes in a predetermined order.
[0220] Note that such a basic mesh has metadata such as a frame number and a sub-mesh number, and at least the coordinates of each vertex and the index of the vertex, and is stored in the mesh buffer unit 202C in a predetermined order determined by the reference frame list.
[0221] Here, as shown in FIG. 18, such a reference frame list (ref_list0) is a list of information for specifying all the reference decoding basic meshes stored in the mesh buffer unit 202C.
[0222] As shown in FIG. 18, the reference frame list may be determined by a control signal decoded from the bitstream, or may be calculated naturally from the decoding order of the frames.
[0223] Note that the control signal decoded from the bitstream may be indicated by the relative distance from the frame to be decoded, or may be the frame index with an absolute value.
[0224] Furthermore, either a short-term reference frame or a long-term reference frame may be used according to a control signal.
[0225] For example, when using a short-term reference frame, the absolute value (abs_delta_mfoc_st) and sign (sign_flag) of the difference in the display order between the current frame (cur) and the reference frame (ref) may be decoded from the bit stream, and the display order of the reference frame may be specified by the following formula. If(sign_flag){ Display Order(ref)=Display Order(cur)+abs_delta_mfoc_st }else{ Display Order(ref)=Display Order(cur)-abs_delta_mfoc_st } Also, when a method calculated naturally from the decoding order of frames is used, for example, in the reference frame list, when there is no control signal, they may be arranged sequentially by a certain number of frames from the previously decoded frame. That is, the reference frame list may be {0, -1, -2,…, -(N−1)}.
[0226] Basically, the reference frame list does not change for each frame except in special circumstances (for example, when receiving a re-ordering instruction).
[0227] The mesh buffer unit 202C may be updated as follows.
[0228] When the basic mesh is decoded, in the case of an I-frame and a P-frame, the mesh buffer unit 202C deletes one or more existing reference frames in a predetermined order determined by the reference frame list, and includes one or more basic meshes including the basic mesh of the decoded frame, or creates and includes one basic mesh from a plurality of basic meshes to adjust the order of the reference frames.
[0229] Such deletion operation may be performed only when the mesh buffer unit 202C is full. Note that the number of basic meshes that can be stored in the mesh buffer unit 202C is determined in advance. Here, in the present embodiment, when the number of such basic meshes is reached, it is defined that the mesh buffer unit 202C is full. In the above creation operation, the vertex coordinates corresponding to the basic meshes of the decoded frame and the existing basic meshes stored in the mesh buffer unit 202C may be weighted and averaged to form one basic mesh.
[0230] The weights used in such weighted averaging may be determined in advance, may be calculated using the frame index, or may be decoded from the control signal.
[0231] However, when the mesh buffer unit 202C is an S frame, such update may or may not be performed.
[0232] Note that when the mesh buffer unit 202C receives a control signal indicating a Re-ordering instruction from the control signal decoded from the bit stream, as shown in FIG. 19, it updates the reference frame list and adjusts the order of the reference frames in a predetermined order determined by the updated reference frame list (ref_list0).
[0233] (Inter-decoding unit 202E) The inter-decoding unit 202E is configured to decode the vertex coordinates of the P frame by adding the vertex coordinates of the reference frame taken out from the mesh buffer unit 202C and the motion vector decoded from the bit stream of the P frame.
[0234] Furthermore, the inter-decoding unit 202E can adjust the index of the vertices of the P frame by means of the pair of indices A(k) and B(k) of the vertices existing as duplicate vertices stored in such a specific buffer. All or part of such indices are decoded from the bit stream. Such a decoding method may be arithmetic coding. According to such a configuration, it is possible to expect the effect that there is no limit to the maximum value of the index to be decoded using arithmetic coding. For example, arithmetic coding such as ue(v) may be used.
[0235] (Modification Example 2 of the Basic Mesh Decoding Unit 202) Hereinafter, with reference to FIG. 20, a modification example 2 of the basic mesh decoding unit 202 will be described.
[0236] Hereinafter, the skip decoding unit 202F will be described, but it may also be applied to the inter-decoding unit 202E.
[0237] As shown in FIG. 20, in the skip decoding unit 202F, in order to enable reference to subsequent frames, the decoding order and the display order are different.
[0238] Here, the display order is the same as the order of the input during encoding and the same as the order of the output during decoding.
[0239] On the other hand, the decoding order is the same as the order of the output during encoding and the same as the order of the input during decoding.
[0240] Note that such a reference frame may be calculated by weighted averaging of a subsequent frame and one or more other frames.
[0241] However, when referring to a plurality of frames including a subsequent frame, as a new frame type (smh_type) in FIG. 14, MR_SUBMESH (MR frame or B frame) is defined, and MR_SUBMESH is decoded from the bit stream.
[0242] Also, as shown in FIG. 21, such other frame may be the decoded frame immediately before the target frame.
[0243] Such weight may be calculated using the frame interval between the target frame and the subsequent frame and the frame interval between the target frame and the other frame, or may be determined in advance.
[0244] The basic mesh decoder 202 decodes a control signal (smh_mesh_frm_order_cnt_lsb) from the bit stream and decodes the order of such output.
[0245] When there is a sub-mesh defined in the above-mentioned Non-Patent Document 4, all sub-meshes will have the same control signal (smh_mesh_frm_order_cnt_lsb) or the control signal (smh_mesh_frm_order_cnt_lsb) will be applied to all sub-meshes.
[0246] The value indicated by such control signal (smh_mesh_frm_order_cnt_lsb) may be the difference from the display order of the frame to be decoded, or may be the order in a pre-determined frame group MaxMeshFrmOrderCntLsb.
[0247] When the decoding order (Decode Order) and the display order (Display Order) are different, if the decoded basic meshes are arranged in the decoding order (Decode Order), the basic mesh decoder 202 may rearrange the decoded basic meshes in the display order (Display Order).
[0248] In the S frame where the subsequent frame can be referenced, two mesh buffers 202C may be provided. When only one mesh buffer unit 202 is provided, there is at least one reference frame including the reference frame whose display order is after the frame to be decoded.
[0249] The skip decoding unit 202F specifies a reference frame by receiving a control signal decoded from a bit stream, a predetermined rule, or a Re-ordering instruction.
[0250] Specifically, the skip decoding unit 202F specifies a reference frame in a reference frame list by such a control signal.
[0251] Alternatively, the skip decoding unit 202F specifies the first reference frame in the reference frame list.
[0252] Alternatively, the skip decoding unit 202F receives a Re-ordering instruction, updates the reference frame list and the reference frame order of the mesh buffer unit 202C, and specifies the first reference frame in such a reference frame list.
[0253] Note that in this embodiment, even if the S frame is not decoded, it does not affect the decoding of other frames. Therefore, when part or all of the S frame is not decoded, Temporal scalability can be realized.
[0254] Furthermore, the basic mesh decoding unit 202 may decode the basic mesh of the S frame by integrating a plurality of reference frames according to a control signal.
[0255] For example, the basic mesh decoding unit 202 may be configured to decode the coordinates and indices of the vertices of the basic mesh of the frame to be decoded by averaging the coordinates of the corresponding vertices in the basic meshes of the two reference frames before and after and using the average coordinates and the indices of the vertices as they are.
[0256] According to such a configuration, it is possible to obtain a high-quality basic mesh while eliminating the need for motion vectors in the skip decoding unit 202F or the inter decoding unit 202E. Therefore, an effect of improving the quality of the decoded mesh can be expected. Furthermore, an effect of realizing Temporal scalability can be expected.
[0257] However, in order to achieve temporal scalability, control signals indicating whether to decode the base mesh, displacement amount, and texture in each frame are defined respectively, and they are decoded from the bitstream respectively.
[0258] Also, within the same frame, the Temporal_IDs of the atlas and the base mesh may be made to match. Also, within the same frame, the Temporal_IDs of the atlas and the texture may be made to match. Also, within the same frame, the Temporal_IDs of the atlas and the displacement amount may be made to match.
[0259] According to such a configuration, an effect of being able to avoid the situation where a frame cannot be decoded and wasteful data can be expected.
[0260] Note that it is desirable that the interval between adjacent frames having the same Temporal_ID is constant.
[0261] Adjacent frames having the same Temporal_ID have the closest POC.
[0262] By making the interval between frames constant as described above, an effect of maintaining a constant frame rate when displaying the decoded frames can be expected.
[0263] Furthermore, the decoding order of the atlas and the base mesh having the same display order may be made to match. Also, the decoding order of the atlas and the displacement amount having the same display order may be made to match. Also, the decoding order of the atlas and the texture having the same display order may be made to match.
[0264] Alternatively, the random access points of the atlas and the base mesh having the same display order may be made to coincide. Also, the random access points of the atlas and the displacement amount having the same display order may be made to coincide. Further, the random access points of the atlas and the texture having the same display order may be made to coincide. Note that the random access point is defined in Non-Patent Document 4 or Non-Patent Document 5.
[0265] According to such a configuration, when decoding the base mesh, the displacement amount, and the texture, respectively, an effect that the mesh can be reproduced without waiting for each other's decoding can be expected.
[0266] Furthermore, a frame having a Temporal_ID higher than the control signal Temporal_ID of the frame to be decoded is not used as a reference frame for such a frame to be decoded.
[0267] As a result, an effect that the possibility of discarding the reference frame can be eliminated can be expected.
[0268] Hereinafter, an example of realizing Temporal scalability using the above-described Temporal_ID will be described.
[0269] The bitstreams of the atlas, the base mesh, the displacement amount, and the texture are encapsulated by a Network Abstraction Layer (NAL) unit. The NAL unit may have an NAL header as shown in FIG. 22.
[0270] The TID defined as the last 3 bits in the NAL header is Temporal_ID + 1. The range of the TID is from 1 to 7, and zero is prohibited.
[0271] The LayerID / R6 defined as the 6 bits immediately before the TID in the NAL header specifies the identifier of the layer to which the NAL unit belongs.
[0272] The value of LayerID / R6 must be within the range of 0 to 62. The value 63 may be specified by ISO / IEC in the future.
[0273] For purposes other than determining the data volume of the bitstream decoding unit, the mesh decoder 200 ignores all data following the value 63 in the NAL unit, and the mesh decoder 200 compliant with the specified profile ignores all NAL units where the value of LayerID-R6 is not 0 (i.e., removes and discards them from the bitstream).
[0274] The value 63 of LayerID / R6 can be used to indicate an extended layer identifier in future extensions.
[0275] When there are sub-meshes defined in Non-Patent Document 4, all sub-meshes will have the same TID or the TID will be applied to all sub-meshes.
[0276] For atlases, Non-Patent Document 5 can be used, and for displacement amounts and textures, HEVC or VVC of video coding formats can be used. Therefore, the following will describe the basic mesh.
[0277] The values of LayerID / R6 for all BMCL NAL units of the encoded basic mesh frame must be the same. The value of LayerID / R6 of the encoded basic mesh frame is the value of LayerID / R6 of the BMCL NAL unit of the encoded basic mesh frame.
[0278] When NALType is equal to NAL_EOB, the value of LayerID / R6 must be equal to 0.
[0279] When the NALType is in the range from NAL_BLA_W_LP to NAL_RSV_BMCL_29 defined in Non-Patent Document 4, that is, when it belongs to an IRAP-coded basic mesh frame, the Temporal_ID must be 0.
[0280] When the NALType is equal to NAL_TSA_R or NAL_TSA_N, the Temporal_ID must not be equal to 0.
[0281] When the NALType is equal to 0 and the NALType is equal to NAL_STSA_R or NAL_STSA_N, the Temporal_ID must not be equal to 0.
[0282] The value of the Temporal_ID must be the same for all BMCL NAL units within an access unit.
[0283] The value of the Temporal_ID of a coded basic mesh frame or an access unit is the value of the Temporal_ID of the BMCL NAL unit of the coded basic mesh frame or the access unit.
[0284] The value of the Temporal_ID of a sublayer representation is the maximum value of the Temporal_IDs of all BMCL NAL units within the sublayer representation.
[0285] The value of the Temporal_ID of a non-BMCL NAL unit is restricted as follows. - When the NALType is equal to NAL_BMSPS, the Temporal_ID must be 0, and the Temporal_ID of the access unit containing the NAL unit must be 0. - In other cases, when the NALType is equal to NAL_EOS or NAL_EOB, the Temporal_ID must be 0. - In other cases, when the NALType is equal to NAL_AUD or NALLFDD, the Temporal_ID must be equal to the Temporal_ID of the access unit containing the NALL unit. - Otherwise, the Temporal_ID must be greater than or equal to the Temporal_ID of the access unit containing the NAL unit.
[0286] Note that when the NAL unit is not BMCL, the value of the Temporal_ID is equal to the minimum value of the Temporal_ID values of all access units to which the non - BMCL NAL unit is applied.
[0287] When the NALType is equal to NAL_BMFPS, the Temporal_ID can be greater than or equal to the Temporal_ID of the containing access unit because all basic mesh frame parameter sets (BMFPS) are included at the beginning of the bitstream where the Temporal_ID of the first encoded basic mesh frame is 0.
[0288] Note that the skip decoding unit 202F refers to the specified tIDTarget and discards NAL units with a Temporal_ID higher than tIDTarget without decoding them.
[0289] Here, tIDTarget may be specified by a pre - determined value, or may be specified according to the network situation or the terminal capabilities of the mesh decoding device 200.
[0290] For example, a lower tIDTarget is specified for the wireless case than for the wired case. Also, a lower tIDTarget is specified when the network situation is poor. Also, a lower tIDTarget is specified when decoding is performed by a low - spec mesh decoding device 200.
[0291] However, as a requirement for the compatibility of the bitstream, it is stipulated that there must be at least one NAL unit whose Temporal_ID is not higher than tIDTarget in the bitstream.
[0292] Note that, as shown in FIG. 23, the number of sub-meshes may be different for each frame (intra-frame, inter-frame, and skip frame).
[0293] In such a case, the intra decoder 202B, the inter decoder 202E, and the skip decoder 202F assign non-overlapping sub-mesh IDs to each of the sub-meshes in each frame.
[0294] Also, as shown in FIG. 24, the intra decoder 202B, the inter decoder 202E, and the skip decoder 202F may assign different SubmeshIDs (sub-mesh IDs) to corresponding sub-meshes between frames.
[0295] However, it is assumed that the inter decoder 202E or the skip decoder 202F can only refer to sub-meshes having the same SubmeshID in the reference frames.
[0296] Alternatively, it is assumed that the inter decoder 202E or the skip decoder 202F can only refer to sub-meshes having the same number of vertices in the reference frames.
[0297] Alternatively, it is assumed that the intra decoder 202B and the inter decoder 202E can refer to the specified sub-mesh in the reference frames.
[0298] In such a case, when there are multiple sub-meshes in the reference frame, the inter decoder 202E or the skip decoder 202F may decode a control signal that designates the SubmeshID of the referable sub-mesh from the bitstream of the current sub-mesh.
[0299] On the other hand, if there is only one sub-mesh in the reference frame, the inter-decoding unit 202E or the skip-decoding unit 202F may use such a sub-mesh as a referenceable sub-mesh.
[0300] However, if the above control signal does not exist, the inter-decoding unit 202E or the skip-decoding unit 202F sets the SubmeshID of the referenceable sub-mesh to be the same as the SubmeshID of the sub-mesh in the current frame.
[0301] In addition, the inter-decoding unit 202E or the skip-decoding unit 202F may decode a control signal from the bit stream indicating whether the above control signal exists.
[0302] Note that the inter-decoding unit 202E or the skip-decoding unit 202F may also decode a control signal for selecting a method for determining the above referenceable sub-mesh.
[0303] The subdivision unit 203 and the displacement amount decoding unit 206 may follow Non-Patent Document 4.
[0304] According to the present invention, the amount of calculation can be reduced by reusing the reference frame itself without searching for decoded adjacent vertices.
[0305] In addition, according to the present embodiment, in inter-prediction coding, even when the number of vertices of the basic mesh of the current frame is different from the number of vertices of the reference frame or the reference sub-mesh, the basic mesh of the current frame can be decoded.
[0306] In addition, according to the present embodiment, in inter-prediction coding, a situation where the number of vertices of the basic mesh of the current frame is different from the number of vertices of the reference frame or the reference sub-mesh can be avoided.
[0307] Also, according to the present embodiment, in inter-prediction coding, by introducing a control signal indicating which sub-mesh to refer to among the reference frames of the current frame, it is possible to specify which sub-mesh to refer to.
[0308] Also, according to the present embodiment, in inter-prediction coding, even if there is no information indicating which sub-mesh to refer to among the reference frames of the current frame, it is possible to specify which sub-mesh to refer to.
[0309] Also, according to the present embodiment, it is possible to ensure that the basic mesh has at least one face.
[0310] Also, according to the present embodiment, the Temporal_scalability function can be realized.
[0311] Furthermore, according to the present embodiment, the coding efficiency of the mesh can be improved.
[0312] The above-described mesh encoding device 100 and mesh decoding device 200 may be realized by a program that causes a computer to execute each function (each process).
Industrial Applicability
[0313] Note that according to the present embodiment, for example, since it is possible to realize an improvement in the overall service quality in video communication, it is possible to contribute to Goal 9 of the Sustainable Development Goals (SDGs) led by the United Nations, "Build resilient infrastructure, promote sustainable industrialization, and foster innovation."
Explanation of Signs
[0314] 1... Mesh processing system 100... Mesh encoding device 200... Mesh decoding device 201... Multiplex separation unit 202... Basic mesh decoding unit 202A… Separation unit 202B… Intra decoding unit 202B1… Arbitrary intra decoding unit 202B2… Sorting unit 202C… Mesh buffer unit 202D… Connection information decoding unit 202E… Inter decoding unit 202E1… Motion vector residual decoding unit 202E2… Motion vector buffer unit 202E3… Motion vector prediction unit 202E4… Motion vector calculation unit 202E5… Adder 202E6… Duplicate vertex search unit 202E7… mv_signalled_flag acquisition unit 202E8… Motion vector acquisition unit 202F… Skip decoding unit 203… Subdivision unit 204… Mesh decoding unit 205… Patch integration unit 206… Displacement amount decoding unit 207… Video decoding unit 208… Atlas data decoding unit
Claims
1. A mesh decoding device, comprising: an intra decoder that decodes vertex coordinates and connection information in the intra frame from a bit stream of the intra frame; an inter decoder that decodes the coordinates of a vertex to be decoded by adding a motion vector decoded from a bit stream of an inter frame and the coordinates of a vertex corresponding to the vertex to be decoded in a reference frame; The mesh decoding device, wherein the basic meshes in the intra frame and the inter frame each have at least one face.
2. The mesh decoding device according to claim 1, wherein a control signal indicating the number of vertices per frame or per sub-mesh is limited to include at least three or more vertices.
3. The mesh decoding device according to claim 2, wherein the control signal is a control signal for designating the number of vertices in a patch within an atlas style.
4. The mesh decoding device according to claim 2, wherein the control signal is a control signal for designating the number of vertices in a sub-mesh.
5. The mesh decoding device according to claim 2, wherein the control signal is a control signal for designating the number of vertices in a decoded mesh.
6. A mesh decoding method, comprising: Step A of decoding vertex coordinates and connection information in the intra frame from a bit stream of the intra frame; Step B of decoding the coordinates of a vertex to be decoded by adding a motion vector decoded from a bit stream of an inter frame and the coordinates of a vertex corresponding to the vertex to be decoded in a reference frame; The mesh decoding method, wherein in step A and step B, the basic meshes in the intra frame and the inter frame each have at least one face.
7. A program for causing a computer to function as a mesh decoding device, wherein the mesh decoding device comprises an intra decoder that decodes vertex coordinates and connection information in the intra frame from a bit stream of the intra frame; and an inter decoder that decodes the coordinates of a vertex to be decoded by adding a motion vector decoded from a bit stream of an inter frame and the coordinates of a vertex corresponding to the vertex to be decoded in a reference frame. A program characterized in that the basic meshes in the intra-frame and the inter-frame have at least one surface.