Mesh decoding device, mesh decoding method, and program
The mesh decoding device and method address the inefficiency of searching for decoded adjacent vertices by reusing a reference frame's vertex list for motion vector prediction, thereby reducing computational load and improving decoding efficiency.
Patent Information
- Application Number
- JP2024000907
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-06
- Publication Date
- 2025-07-17
AI Technical Summary
Existing mesh decoding technologies require a heavy computational load due to the need to search for decoded adjacent vertices for motion vector prediction, which is inefficient and resource-intensive.
A mesh decoding device and method that calculates motion vector prediction by averaging motion vectors of adjacent vertices using an adjacent vertex list reused from a reference frame, reducing the need to search for decoded adjacent vertices.
This approach significantly reduces computational load and resource requirements by reusing the reference frame's adjacent vertex list, enhancing decoding efficiency.
Smart Images

Figure 2025107106000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a mesh decoding device, a mesh decoding method, and a program.
Background Art
[0002] Non-Patent Document 1 or Non-Patent Document 4 discloses a technique for encoding a mesh using Non-Patent Document 2 or 3 according to the framework of Non-Patent Document 5.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
Non-Patent Document 5
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the prior art, there is a problem that a heavy computational load is required to search for the decoded adjacent vertices for calculating the motion vector prediction value in the inter-prediction coding every time. Therefore, the present invention has been made in view of the above problems, and an object thereof is to provide a mesh decoding device, a mesh decoding method, and a program capable of reducing the computational load by reusing the reference frame itself without searching for the above-mentioned decoded adjacent vertices.
Means for Solving the Problems
[0005] A first feature of the present invention is a mesh decoding device, comprising an inter-decoding unit that decodes the coordinates of a vertex to be decoded by adding a motion vector decoded from a bitstream of an inter-frame and the coordinates of a vertex corresponding to the vertex to be decoded in a reference frame, wherein the inter-decoding unit comprises a motion vector prediction unit that calculates a prediction value of the motion vector of the vertex to be decoded by averaging all or part of the motion vectors of the decoded adjacent vertices adjacent to the vertex to be decoded with reference to an adjacent vertex list that is a list of decoded vertices adjacent to each vertex, and the gist of the motion vector prediction unit is to reuse the adjacent vertex list in the reference frame as the adjacent vertex list in the current frame.
[0006] The second feature of the present invention is a mesh decoding method, which includes a step of decoding the coordinates of a vertex to be decoded by adding a motion vector decoded from a bit stream of an inter-frame and the coordinates of a vertex corresponding to the vertex to be decoded in a reference frame. In this step, by referring to an adjacent vertex list that is a list of decoded vertices adjacent to each vertex, and averaging all or part of the motion vectors of the decoded vertices adjacent to the vertex to be decoded, a predicted value of the motion vector of the vertex to be decoded is calculated. The gist is to reuse the adjacent vertex list in the reference frame as the adjacent vertex list in the current frame.
[0007] The third feature of the present invention is a program for causing a computer to function as a mesh decoding device. The mesh decoding device includes an inter-decoding unit that decodes the coordinates of a vertex to be decoded by adding a motion vector decoded from a bit stream of an inter-frame and the coordinates of a vertex corresponding to the vertex to be decoded in a reference frame. The inter-decoding unit includes a motion vector prediction unit that calculates a predicted value of the motion vector of the vertex to be decoded by referring to an adjacent vertex list that is a list of decoded vertices adjacent to each vertex, and averaging all or part of the motion vectors of the decoded vertices adjacent to the vertex to be decoded. The gist of the motion vector prediction unit is to reuse the adjacent vertex list in the reference frame as the adjacent vertex list in the current frame.
Advantages of the Invention
[0008] According to the present invention, it is possible to provide a mesh decoding device, a mesh decoding method, and a program that can reduce the amount of calculation by reusing the reference frame itself without searching for decoded adjacent vertices.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10A
Figure 10B
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Mode for Carrying Out the Invention
[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the components in the following embodiments can be replaced with existing components as appropriate, and various variations including combinations with other existing components are possible. Therefore, the description of the following embodiments does not limit the content of the invention described in the claims.
[0011] <First Embodiment> Hereinafter, with reference to FIGS. 1 to 24, a mesh processing system according to the present embodiment will be described.
[0012] FIG. 1 is a diagram showing an example of the configuration of a mesh processing system 1 according to the present embodiment. As shown in FIG. 1, the mesh processing system 1 includes a mesh encoding device 100 and a mesh decoding device 200.
[0013] FIG. 2 is a diagram showing an example of the functional blocks of the mesh decoding device 200 according to the present embodiment.
[0014] As shown in FIG. 2, the mesh decoding device 200 includes a multiplex separation unit 201, a basic mesh decoding unit 202, a subdivision unit 203, a mesh decoding unit 204, a patch integration unit 205, a displacement amount decoding unit 206, a video decoding unit 207, and an atlas data decoding unit 208.
[0015] Here, the basic mesh decoding unit 202, the subdivision unit 203, the mesh decoding unit 204, and the displacement amount decoding unit 206 are configured to perform processing in units of patches obtained by dividing a mesh, and thereafter, the patch integration unit 205 may be configured to integrate the processing results thereof.
[0016] In the example of FIG. 3A, the mesh is divided into a patch 1 composed of basic surfaces 1 and 2 and a patch 2 composed of basic surfaces 3 and 4.
[0017] The multiplex separation unit 201 is configured to separate the multiplexed bit stream into a basic mesh bit stream, a displacement amount bit stream, a texture bit stream, and an atlas bit stream.
[0018] The atlas data decoding unit 208 is configured to decode the atlas bit stream and output control information. Such control signals may be used as metadata in the basic mesh decoding unit 202, the subdivision unit 203, the mesh decoding unit 204, the displacement amount decoding unit 206, and the video decoding unit 207.
[0019] <Basic mesh decoding unit 202> The basic mesh decoding unit 202 is configured to decode the basic mesh bit stream, generate a basic mesh, and output it.
[0020] Here, the basic mesh is composed of a plurality of vertices in a three-dimensional space and edges connecting the plurality of vertices.
[0021] As shown in FIG. 3A, the basic mesh is configured by combining basic faces represented by three vertices.
[0022] The basic mesh decoding unit 202 may be configured to decode the basic mesh bit stream using, for example, Draco shown in Non-Patent Document 2 or the technology described in Non-Patent Document 3.
[0023] Further, the basic mesh decoding unit 202 may be configured to generate "subdivision_method_id" described later as control information for controlling the type of the subdivision method.
[0024] As shown in FIG. 4, the basic mesh decoding unit 202 includes a separation unit 202A, an intra decoding unit 202B, a mesh buffer unit 202C, a connection information decoding unit 202D, and an inter decoding unit 202E.
[0025] The separation unit 202A is configured to classify the basic mesh bit stream into the bit stream of an I frame and the bit stream of a P frame.
[0026] (Intra decoding unit 202B) The intra decoding unit 202B is configured to decode the coordinates and connection information of the vertices of the I frame from the bit stream of the I frame, for example, using Draco shown in Non-Patent Document 2 or the technology described in Non-Patent Document 3.
[0027] FIG. 5 is a diagram showing an example of the functional block of the intra decoding unit 202B.
[0028] As shown in FIG. 5, the intra decoding unit 202B includes an arbitrary intra decoding unit 202B1 and an alignment unit 202B2.
[0029] The arbitrary intra decoding unit 202B1 is configured to decode the coordinates and connection information of the unordered vertices of the I frame from the bit stream of the I frame using an arbitrary method including Draco shown in Non-Patent Document 2 or the technology described in Non-Patent Document 3.
[0030] The alignment unit 202B2 is configured to output vertices by rearranging the unordered vertices in a predetermined order.
[0031] As the predetermined order, for example, the Morton code order or the raster scan order may be used.
[0032] Further, the alignment unit 202B2 may group duplicate vertices, which are a plurality of vertices having the same coordinates in the decoded basic mesh, into a single vertex and then rearrange them in a predetermined order.
[0033] The mesh buffer section 202C is configured to store the vertex coordinates and connection information of the I-frame decoded by the intra decoding section 202B. Here, a specific buffer may be provided for storing pairs of vertex indices A(k) and B(k) that exist as duplicate vertices in a predetermined order.
[0034] The connection information decoding section 202D is configured to convert the connection information of the I-frame or reference frame retrieved from the mesh buffer section 202C into the connection information of the P-frame.
[0035] The inter decoding section 202E is configured to decode the vertex coordinates of the P-frame by adding the vertex coordinates of the reference frame retrieved from the mesh buffer section 202C and the motion vector decoded from the bit stream of the P-frame.
[0036] Furthermore, the inter decoding section 202E can adjust the vertex index of the P-frame based on the pairs of vertex indices A(k) and B(k) that exist as duplicate vertices stored in such a specific buffer.
[0037] Here, all or part of the above-mentioned indices are decoded from the bit stream. Such a decoding method may be arithmetic coding. As a result, it is possible to expect the effect that there is no limit to the maximum value of the index to be decoded using arithmetic coding.
[0038] For example, arithmetic coding such as ue(v) may be used. ue(v) indicates leftmost bit-first unsigned integer zero-order exponential Golomb coding (Exp-Golomb).
[0039] Specifically, the parsing process of the syntax element of ue(v) starts from the current position in the bit stream, reads the bit including the first non-zero bit, and counts the number of leading bits equal to 0. This process is specified as follows. leadingZeroBits=-1 for (b = 0;!b; leadingZeroBits++) b = read_bits(l) Next, the variable codeNum is assigned as follows.
[0040] codeNum = 2 leadingZeroBits -1 + read_bits(leadingZeroBits) However, the value returned by read_bits(leadingZeroBits) is interpreted as the binary representation of an unsigned integer with the most significant bit written first. Also, the value of ue(v) is equal to the value of codeNum.
[0041] Table 1 shows the structure of the Exp-Golomb code by separating the bit sequence into "prefix" bits and "suffix" bits.
[0042]
Table 1
[0043] Here, the "prefix" bits are the bits that are parsed as specified in the calculation of leadingZeroBits and are shown as 0 or 1 in the columns of the bit sequence in Table 1.
[0044] The "suffix" bits are the bits that are parsed in the calculation of codeNum and are shown as x i in Table 1. i is in the range from 0 to leadingZeroBits - 1. Each x i is equal to either 0 or 1.
[0045] Table 2 shows a way to explicitly assign the bit sequence to the value of codeNum. Here, the value of ue(v) is equal to the value of codeNum.
[0046]
Table 2
[0047] In this embodiment, as shown in FIG. 6, there is a correspondence between the vertices of the basic mesh of the P frame and the vertices of the basic mesh of the reference frame (I frame or P frame). Here, the motion vector decoded by the inter-decoding unit 202E is the difference vector between the coordinates of the vertices of the basic mesh of the P frame and the coordinates of the vertices of the basic mesh of the I frame.
[0048] Note that the inter-decoding unit 202E may decode the number of vertices of the current frame or the current sub-mesh from the bit stream.
[0049] Here, when the inter-decoding unit 202E decodes the number of vertices of the current frame or the current sub-mesh from the above-mentioned bit stream, the requirement for the conformity of such a bit stream is that the number of vertices of the decoded current frame or current sub-mesh must be equal to the number of vertices of the reference frame or the reference sub-mesh.
[0050] Note that when the inter-decoding unit 202E decodes the number of vertices of the current frame or the current sub-mesh from the above-mentioned bit stream, if the number of vertices of the decoded current frame or current sub-mesh is different from the number of vertices of the reference frame or the reference sub-mesh, it is configured to preferentially use the number of vertices of the current frame or the current sub-mesh decoded from the bit stream.
[0051] Also, the inter-decoding unit 202E may use the number of vertices of the reference frame or the reference sub-mesh as the number of vertices of the current frame or the current sub-mesh as it is.
[0052] In such a case, when the number of vertices of the current frame or the current sub-mesh is larger than the number of vertices of the reference frame or the reference sub-mesh, the inter-decoding unit 202E may add dummy vertices and connection information of the dummy vertices to the reference frame or the reference sub-mesh.
[0053] Here, the inter-decoding unit 202E may set the coordinates of such dummy vertices to fixed values (for example, (0, 0, 0)), or may copy them from a predetermined vertex (for example, the last vertex of the reference frame).
[0054] Also, the inter-decoding unit 202E may copy the necessary portion from the beginning of the connection information of the vertices in the reference frame or the reference sub-mesh as the connection information of the dummy vertices.
[0055] According to such a configuration, an effect of ensuring the decoding operation of the current frame or the current sub-mesh can be expected.
[0056] Note that since the basic mesh has at least one face and such a face has at least three or more vertices, the control signal indicating the number of vertices for each frame or each sub-mesh is restricted to include at least three or more vertices.
[0057] For example, the control signal pdu_vertex_count_minus_1[titleID][patchIdx] defined in Section 8.3.7.3 of Non - Reference Document 4, the control signal sismu_inter_vertex_count[subMeshID] defined in Section H.8.1.3.8, and the control signal mesh_vertex_count defined in Section I.8.3.7 are changed as shown in Table 3 and restricted to include three or more vertices.
[0058] Here, the control signal pdu_vertex_count_minus_1[titleID][patchIdx] is a control signal that specifies the number of vertices in a patch having a patch ID equal to the patch ID specified by [patchIdx] within an atlas tile having a tile ID equal to the tile ID specified by [titleID].
[0059] The control signal sismu_inter_vertex_count[subMeshID] is a control signal that specifies the number of vertices in the submesh having the SubmeshID equal to the SubmeshID specified by [subMeshID].
[0060] The control signal mesh_vertex_count is a control signal that specifies the number of vertices in the decoded mesh.
[0061]
Table 3
[0062] According to such a configuration, an effect can be expected that it is possible to prevent a situation where vertices operate with data having no meaning such as 1 or 2.
[0063] (Inter-decoding unit 202E) FIG. 7 is a diagram showing an example of the functional blocks of the inter-decoding unit 202E.
[0064] As shown in FIG. 7, the inter-decoding unit 202E includes a motion vector residual decoding unit 202E1, a motion vector buffer unit 202E2, a motion vector prediction unit 202E3, a motion vector calculation unit 202E4, and an adder 202E5.
[0065] The motion vector residual decoding unit 202E1 is configured to generate an MVR (Motion Vector Residual) from the bit stream of the P frame.
[0066] Here, the MVR is a motion vector residual indicating the difference between an MV (Motion Vector) and an MVP (Motion Vector Prediction). The MV is a difference vector (motion vector) between the coordinates of the vertex of the corresponding I frame and the coordinates of the vertex of the P frame. The MVP is a predicted value of the MV of the target vertex (predicted value of the motion vector) using the MV.
[0067] The motion vector buffer unit 202E2 is configured to sequentially store the MVs output by the motion vector calculation unit 202E4.
[0068] The motion vector prediction unit 202E3 acquires the decoded MVs from the motion vector buffer unit 202E2 for the vertices connected to the vertex to be decoded, and as shown in FIG. 8, uses all or part of the acquired decoded MVs to output the MVP of the vertex to be decoded.
[0069] The motion vector calculation unit 202E4 adds the MVR generated by the motion vector residual decoding unit 202E1 and the MVP output from the motion vector prediction unit 202E3, and is configured to output the MV of the vertex to be decoded.
[0070] The adder 202E5 adds the coordinates of the vertex corresponding to the vertex to be decoded obtained from the decoded basic mesh of the reference frame (I-frame or P-frame) having a corresponding relationship and the motion vector MV output from the motion vector calculation unit 202E3, and is configured to output the coordinates of the vertex to be decoded.
[0071] Hereinafter, the details of each part of the inter-decoding unit 202E will be described.
[0072] FIG. 9 shows a flowchart showing an example of the operation of the motion vector prediction unit 202E3. Hereinafter, the operation of the motion vector prediction unit 202E3 is referred to as the "average prediction method".
[0073] As shown in FIG. 9, in step S1001, the motion vector prediction unit 202E3 sets 0 for MVP and N.
[0074] In step S1002, the motion vector prediction unit 202E3 acquires a set of MVs of the vertices around the vertex to be decoded from the motion vector buffer unit 202E2, identifies the vertices for which the subsequent processing has not been completed, and transitions to No. When the subsequent processing has been completed for all vertices, it transitions to Yes.
[0075] In step S1003, if the MV of the vertex to be processed is not decoded, the motion vector prediction unit 202E3 transitions to No, and if the MV of the vertex to be processed is decoded, it transitions to Yes.
[0076] In step S1004, the motion vector prediction unit 202E3 adds the MV to the MVP and adds 1 to N.
[0077] In step S1005, if N is greater than 0, the motion vector prediction unit 202E3 outputs the result of dividing the MVP by N, and if N is 0, it outputs 0 and ends the process.
[0078] That is, the motion vector prediction unit 202E3 is configured to output the MVP to be decoded by averaging the decoded motion vectors of the vertices around the vertex to be decoded.
[0079] Note that the motion vector prediction unit 202E3 may be configured to set the MVP to 0 when the set of such decoded motion vectors is an empty set.
[0080] The motion vector calculation unit 202E4 may be configured to calculate the MV of the vertex to be decoded from the MVP output by the motion vector prediction unit 202E3 and the MVR generated by the motion vector residual decoding unit 202E1 according to equation (1).
[0081] MV(k) = MVP(k) + MVR(k) … (1) Here, k is the index of the vertex. MV, MVR, and MVP are vectors having x, y, and z components.
[0082] According to such a configuration, since only the MVR is encoded instead of the MV using the MVP, an effect of improving the encoding efficiency can be expected.
[0083] The adder 202E5 is configured to calculate the coordinates of such a vertex by adding the MV of the vertex calculated by the motion vector calculation unit 202E4 and the coordinates of the vertex of the reference frame corresponding to such a vertex, and to keep the connectivity information as the reference frame.
[0084] Specifically, the adder 202E5 may be configured to calculate the coordinates v’ i (k) of the k-th vertex using Equation (2).
[0085] v’ i (k)=v’ j (k)+MV(k) … (2) Here, v’ i (k) is the coordinate of the k-th vertex to be decoded in the frame to be decoded, v’ j (k) is the coordinate of the decoded k-th vertex of the reference frame, MV(k) is the k-th MV of the frame to be decoded, and k = 1, 2…, K.
[0086] Also, the connectivity information of the frame to be decoded is made the same as the connectivity information of the reference frame.
[0087] Note that since the motion vector prediction unit 202E3 calculates the MVP using the decoded MV, the decoding order affects the MVP.
[0088] Such a decoding order is set to the decoding order of the vertices of the basic mesh of the reference frame. Generally, if a decoding method that increases the basic surface one by one from the starting edge using a certain repeating pattern is used, the order of the decoded vertices of the basic mesh is determined in the decoding process.
[0089] For example, the motion vector prediction unit 202E3 may determine the decoding order of the vertices using Edgebreaker in the basic mesh of the reference frame.
[0090] According to such a configuration, since the MV from the reference frame is encoded instead of the coordinates of the vertex, an effect of improving the encoding efficiency can be expected.
[0091] (Modification Example 1 of Inter Decoder 202E) Hereinafter, Modification Example 1 of Inter Decoder 202E will be described.
[0092] In the "average prediction method" where the motion vector prediction unit 202E3 of the inter decoder 202E averages the decoded motion vectors of the vertices around the vertex to be decoded, in order not to exceed the previously determined maximum number of usages, all or only a part of the decoded motion vectors of the vertices around the vertex to be decoded are used to calculate the MVP.
[0093] Note that the previously determined maximum number of usages is decoded from the bitstream as a control signal.
[0094] Also, when the number of decoded motion vectors of the vertices around the vertex to be decoded exceeds the maximum number of usages, the motion vector prediction unit 202E3 picks up to the maximum number of usages according to a certain rule.
[0095] For example, as such a rule, the motion vector prediction unit 202E3 selects the first or last vertex in the decoding order.
[0096] The decoding order for the mesh as shown in FIG. 10A is, as indicated by the arrow, vertex v D →v C →v A →v B That is.
[0097] FIG. 10B is a list of the vertices around the vertex to be decoded that are used when calculating the MVP of each vertex v A ~v D when the maximum value of the number of decoded adjacent vertices is set to 3.
[0098] According to such a configuration, by determining the maximum number of adjacent vertices, it is possible to expect the effect of reducing the amount of calculation and the amount of memory while maintaining or reducing the coding efficiency.
[0099] However, in order to exert the above-described effect, it is necessary to set an appropriate maximum number of adjacent vertices in the mesh encoding device 100 and write it into the bit stream as a related control signal.
[0100] Therefore, the range that can be set as the above-described maximum number of adjacent vertices is encoded / decoded so that the maximum number of adjacent vertices is less than or equal to a predetermined maximum value as a reasonable constraint regarding the maximum number of adjacent vertices in order to determine the amount of memory prepared in the mesh decoding device 200.
[0101] Thus, it is possible to expect the effect of facilitating the design of the mesh decoding device 200 by defining a reasonable constraint regarding the maximum number of adjacent vertices.
[0102] Generally, the average number of adjacent vertices in a Closed 2-manifold triangle mesh is about 6. Statistically, however, the maximum number of adjacent vertices is often 7 to 8. As shown in FIG. 11, the number of decoded motion vectors (vertical axis) changes dynamically according to the number of vertices (horizontal axis) around the vertex to be decoded.
[0103] Therefore, it is desirable to narrow down the range that can be set as the above-described maximum number of adjacent vertices.
[0104] For example, as shown in FIG. 11, statistically, the number "3", which is the number of vertices around the vertex to be decoded having the largest number of decoded motion vectors, is included within the range that can be set as the maximum number of adjacent vertices in the above-described control signal, or a certain ratio (for example, 50% or 120%) of the statistical average number of adjacent vertices or a value not greater than a natural number that can cover up to N bits (for example, 3 bits) is set as the upper limit (maximum value) of the range that can be set as the maximum number of adjacent vertices in the above-described control signal, whereby the effect of reducing the amount of calculation and the amount of memory can be exerted.
[0105] On the other hand, if the range that can be set as the maximum number of adjacent vertices is set to a large value, for example, 256 or 8 bits in the worst case, it may not be possible to reduce not only the memory amount but also the calculation amount.
[0106] FIG. 12 shows an example of the worst case, where when n ≧ 256, the number of decoded adjacent vertices exceeds 256. In FIG. 12, the number of decoded adjacent vertices at vertex n + 1 is n.
[0107] If the upper limit of the maximum number of adjacent vertices is set to 256, the mesh decoding device 200 requires not only a huge amount of memory but also a huge amount of calculation, as shown in FIG. 10B. Therefore, the upper limit (maximum value) of the range that can be set as the above-mentioned maximum number of adjacent vertices may be 8.
[0108] Furthermore, the range that can be set as the maximum number of adjacent vertices in the above control signal may be a clear value or may be calculated from other control signals and data.
[0109] For example, Level1 may define the range that can be set as the maximum number of adjacent vertices in the control signal.
[0110] Alternatively, the upper limit of the range that can be set as the maximum number of adjacent vertices in the control signal may be calculated from the number of vertices of the basic mesh by the following formula (3).
[0111] Upper limit of the range that can be set as the maximum number of adjacent vertices in the control signal = log2 (number of vertices of the basic mesh) … Formula (3) According to such a configuration, the range that can be appropriately set as the maximum number of adjacent vertices can be determined, and the effect of surely reducing both the calculation amount and the memory amount can be expected even in the worst case.
[0112] The inter-decoding unit 202E decodes the coordinates of the vertex to be decoded by adding the MV decoded from the bit stream of the inter-frame and the coordinates of the vertex corresponding to the vertex to be decoded in the reference frame.
[0113] Also, the motion vector prediction unit 202E3 calculates the predicted value of the motion vector of the vertex to be decoded by referring to the adjacent vertex list and averaging all or part of the motion vectors of the decoded vertices adjacent to the vertex to be decoded. Here, the adjacent vertex list is a list of vertices adjacent to each vertex.
[0114] The motion vector prediction unit 202E3 can reuse the adjacent vertex list in the reference frame as the adjacent vertex list in the current frame. Here, it is assumed that such an adjacent vertex list includes the decoded vertices picked up up to the maximum number of uses (the set maximum number of adjacent vertices).
[0115] Here, when the condition (reuse condition) that each vertex in the reference frame and each vertex in the current frame have a one-to-one correspondence, the decoding order of each vertex in the reference frame and the decoding order of each vertex in the current frame are the same, and an adjacent vertex list including the decoded vertices picked up up to the maximum number of uses in the reference frame is already stored is satisfied, the adjacent vertex list in the reference frame can be reused.
[0116] For example, when the reference frame is decoded, if the reference frame is a P frame or an S frame, there is an adjacent vertex list including the decoded vertices picked up up to the maximum number of uses.
[0117] Therefore, when the reference frame is decoded, if the reference frame is a P frame or an S frame, the motion vector prediction unit 292E3 stores the decoded reference frame in the reference frame buffer and stores the adjacent vertex list in the reference frame in a specific buffer.
[0118] When the current frame is decoded, if the type of the reference frame is a P frame or an S frame, the motion vector prediction unit 292E3 re-uses the adjacent vertex list in the reference frame stored in the specific buffer as the adjacent vertex list in the current frame.
[0119] However, the specific buffer may store an adjacent vertex list including decoded vertices picked up up to the maximum number of uses in one or more frames.
[0120] Therefore, when the specific buffer stores adjacent vertex lists in a plurality of frames, the motion vector prediction unit 202E3 re-uses the adjacent vertex list in the frame corresponding to the reference frame (the frame including the same frame index as the reference frame) as the adjacent vertex list in the current frame.
[0121] The motion vector prediction unit 202E3 applies operations on the reference frame buffer also to the specific buffer. Here, the operations on the reference frame buffer are, for example, the marking process described in Chapter 9.2.4.4 of Non-Patent Document 4 or Non-Patent Document 5.
[0122] According to such a configuration, an effect of reducing the computational amount for creating an adjacent vertex list including decoded vertices picked up up to the maximum number of uses in the current frame can be expected.
[0123] However, in such a case, a specific buffer for storing adjacent vertex lists in one or more frames is required. Here, in order to minimize the size of the specific buffer, further conditions may be added to the above re-use conditions.
[0124] For example, in addition to the above reuse conditions, when the condition that the reference frame is the frame immediately before the current frame in decoding order is satisfied, the motion vector prediction unit 202E3 can reuse the adjacent vertex list in the reference frame as the adjacent vertex list in the current frame.
[0125] In such a case, the motion vector prediction unit 202E3 may save only the adjacent vertex list in the frame immediately before the current frame in decoding order in the specific buffer. Also, the motion vector prediction unit 202E3 does not apply the operations on the reference frame buffer to the specific buffer. According to such a configuration, an effect of reducing the size of the specific buffer can be expected.
[0126] (Modification Example 2 of Inter Decoder 202E) Hereinafter, with reference to FIG. 13, a modification example 2 of the inter decoder 202E will be described.
[0127] The motion vector calculation unit 202E4 of the inter decoder 202E has mode 1 and mode 0.
[0128] In mode 1, the motion vector calculation unit 202E4 adds the MVR generated by the motion vector residual decoding unit 202E1 and the MVP output from the motion vector prediction unit 202E3, and outputs the MV of the vertex to be decoded (see A in FIG. 13).
[0129] On the other hand, in mode 0, the motion vector calculation unit 202E4 outputs the MVR generated by the motion vector residual decoding unit 202E1 as the MV of the vertex to be decoded (see B in FIG. 13).
[0130] Note that the operation of the motion vector calculation unit 202E4 in mode 0 corresponds to the operation of setting the MVP output from the motion vector prediction unit 202E3 to zero.
[0131] Furthermore, the motion vector calculation unit 202E4 may make the modes of the MVs of N (N≧1) consecutive vertices in decoding order the same.
[0132] The motion vector calculation unit 202E4 groups the above-mentioned N vertices into one group. The size of such a group (group size) N is 1 or more. The motion vector calculation unit 202E4 decodes a control signal (group size shown in FIG. 13) that can calculate such a group size from the bit stream.
[0133] However, when the number of vertices remaining in the last group is smaller than the group size, the motion vector calculation unit 202E4 puts all the remaining vertices into the group.
[0134] In this way, when N consecutive vertices are set to the same mode, the amount of code for the mode can be reduced, so an effect of improving the coding efficiency can be expected.
[0135] Here, the larger the number of consecutive vertices having the same mode, the greater the effect of reducing the amount of code for the mode. Therefore, it is necessary to set an appropriate group size in the mesh encoding device 100 and decode it from the bit stream as a control signal in the mesh decoding device 200.
[0136] Therefore, it is desirable that the range that can be set in such a control signal is not smaller than the actual number of consecutive vertices having the same mode.
[0137] For example, when almost all vertices select the same mode, the group size may be set to the total number of vertices.
[0138] Table 4 shows examples when the number of vertices selecting mode 0 is 80% or more, or when the number of vertices selecting mode 1 is 90% or more.
[0139] Therefore, the range that can be set in the above-mentioned control signal is made to cover from 1 to a predetermined maximum value. Such a maximum value is set to be equal to or greater than the total number of vertices of the basic mesh.
[0140]
Table 4
[0141] In addition, when the above control signal (group size) is a natural number, if it is set to be greater than or equal to the total number of vertices, since the absolute value is large, the amount of sign will be large.
[0142] Therefore, it is also possible to use the logarithm for the above control signal. Specifically, with the control signal as log2_group_size, the group size may be calculated by the following formula (4).
[0143] group size = 2log2_group_size … Formula (4) Here, if there is only one group in the frame, that group is made the last group. That is, if the group size is larger than the number of vertices, all vertices are put into the group.
[0144] Furthermore, the range that can be set in the above control signal may be a clear value or may be calculated from other control signals and data.
[0145] For example, by Level1, the range that can be set in such a control signal may be defined.
[0146] Alternatively, the range that can be set in such a control signal may be calculated from the number of vertices of the basic mesh.
[0147] For example, regarding the range that can be set in such a control signal, it may be set to the smallest natural number that is a power of 2 covering the number of vertices of the basic mesh.
[0148] Furthermore, regarding the configurable range in the above control signal, after setting it to a small range, as shown in FIG. 14, a predetermined flag (Mode flag) of another control signal may be introduced. In such a case, as shown in FIG. 14, if such a predetermined flag is TRUE (Mode flag = 1), the motion vector calculation unit 202E4 makes all vertices into one group (that is, uses the number of all vertices as the group size), and if it is FALSE, it keeps the group size calculated from the above control signal.
[0149] Note that the above control signal may be set for each sequence or for each frame. When the above control signal is set for each sequence, the group size of all frames is the same.
[0150] According to such a configuration, by appropriately determining the configurable range of the group size, it is possible to handle all situations, surely reduce the amount of mode codes, and expect the effect of improving the coding efficiency.
[0151] (Modification Example 3 of the Inter Decoder 202E) In a further modification example of the above inter decoder 202E, before implementing the above inter decoder 202E, it is configured to add the following functional blocks.
[0152] Specifically, as shown in FIG. 15, in addition to the configuration shown in FIG. 8, the inter decoder 202E includes a duplicate vertex search unit 202E6, an mv_signalled_flag acquisition unit (flag acquisition unit) 202E7, and a motion vector acquisition unit 202E8.
[0153] Here, the derived_my_present_flag (first flag) is included at the beginning of the P-frame bitstream and has at least two values of 0 or 1.
[0154] Also, mv_signalled_flag (the second flag) is included in the bitstream of the P frame and has a binary value of 0 or 1 for each vertex when derived_my_present_flag indicates No.
[0155] When derived_my_present_flag indicates No (for example, when derived_my_present_flag is 0), the mv_signalled_flag acquisition unit 202E7 decodes the motion vectors of all vertices from the bitstream of the P frame, and sets the value of my_signalled_flag to 1 without decoding my_signalled_flag of all vertices from the bitstream of the P frame.
[0156] When derived_my_present_flag indicates Yes (for example, when derived_my_present_flag is 1), the mv_signalled_flag acquisition unit 202E7 performs different processing for each vertex in the P frame. The mv_signalled_flag acquisition unit 202E7 may determine the processing method for each vertex using mv_signalled_flag.
[0157] Also, when derived_my_present_flag indicates Yes and the mv_signalled_flag of a certain vertex indicates Yes, the mv_signalled_flag acquisition unit 202E7 does not perform the processing in the motion vector acquisition unit 202E8 for the motion vector of such a vertex, and performs the same processing as the inter-decoding unit 202E shown in FIG. 7 or a modified example thereof.
[0158] Also, when derived_my_present_flag indicates Yes and the mv_signalled_flag of a certain vertex indicates No, the mv_signalled_flag acquisition unit 202E7 performs the processing in the motion vector acquisition unit 202E8 for the motion vector of such a vertex, and acquires the motion vector of the vertex.
[0159] For example, when the mv_signalled_flag acquisition unit 202E7 determines that derived_my_present_flag indicates Yes, it decodes mv_signalled_flag for each vertex from the bit stream of the P frame.
[0160] When derived_my_present_flag indicates Yes and the mv_signalled_flag of a certain vertex indicates Yes, the mv_signalled_flag acquisition unit 202E7 sets the prediction mode (MV mode) of such vertex to 2.
[0161] On the other hand, when derived_my_present_flag indicates Yes and the mv_signalled_flag of a certain vertex indicates No, the mv_signalled_flag acquisition unit 202E7 sets the prediction mode of such vertex to a value other than 2.
[0162] Also, when derived_my_present_flag indicates No, the mv_signalled_flag acquisition unit 202E7 does not decode the mv_signalled_flag of all vertices of the P frame from the bit stream, sets its value to 1, and sets the MV mode of the said vertex to a value other than 2.
[0163] The duplicate vertex search unit 202E6 is configured to search for the indices of vertices with matching coordinates (hereinafter referred to as duplicate vertices) from the geometric information of the basic mesh of the decoded reference frame and store them in a buffer (not shown).
[0164] Specifically, the input to the duplicate vertex search unit 202E6 is the index (in decoding order) and position coordinates of each vertex of the basic mesh of the decoded reference frame.
[0165] In addition, the output of the duplicate vertex search unit 202E6 is a list that stores the index (vindex1) of the duplicate vertex if it exists for the index (vindex0) of each vertex, and stores the index (vindex0) of the vertex itself or a specific value (e.g., "-1") not used as the index of each vertex if such a duplicate vertex does not exist. Here, such a list is stored in the buffer repVert in the order of index0.
[0166] Also, since the vertex of vindex1 is decoded before the vertex of vindex0, the relationship vindex0 > vindex1 holds.
[0167] The duplicate vertex search unit 202E6 determines whether there is a duplicate vertex related to the vertex (index: vindex0) of the basic mesh of the reference frame from the first vertex (index: 0) to the immediately preceding vertex (index: vindex0 - 1) of the decoded basic mesh of the reference frame for each vertex (index: vindex0) of the basic mesh of the reference frame. If it is determined that such a duplicate vertex exists, the index of such a duplicate vertex is output by any of at least the following three methods.
[0168] (Method 1) The duplicate vertex search unit 202E6 sequentially searches for duplicate vertices with matching coordinates as follows. When a duplicate vertex exists, vref is vindex1, and when no duplicate vertex exists, vRef is -1.
[0169] vRef = firstVertexIndexDuplicated(vindex0) where firstVertexIndexDuplicated(v) { for( i = 0; i < v; i++) { if(referenceSubmeshVertexPositions[ i ] == referenceSubmeshVertexPositions[ v ]) { return i } } return -1 } (Method 2) The duplicate vertex search unit 202E6 searches for duplicate vertices with matching coordinates using binary search. For example, the duplicate vertex search unit 202E6 may utilize the find function of the associative array class map. (Method 3) The duplicate vertex search unit 202E6 searches for duplicate vertices with matching coordinates using a hash table. For example, the duplicate vertex search unit 202E6 may utilize the find function of the hash associative array class unordered_map.
[0170] Note that as a method for finding duplicate vertices in the basic mesh of the reference frame, for vertices where duplicate vertices exist, a method of decoding the index of the duplicate vertex instead of the position coordinates from a special signal may be used.
[0171] When MVmode is 2 (when derived_my_present_flag indicates Yes and mv_signalled_flag of the vertex indicates No), since there exists a duplicate vertex for the vertex, the motion vector acquisition unit 202E8 acquires, from the motion vector buffer unit 202E2, the motion vector of the vertex having the index (vindex1) of the duplicate vertex corresponding to the index (vindex0) of the vertex output from the duplicate vertex search unit 202E6, and is configured such that the motion vector of such vertex is used as the motion vector of the vertex.
[0172] That is, it is the output of the index (vindex1) of the duplicate vertex 202E6 and is not decoded from the bitstream.
[0173] Here, when MVmode is other than 2 (when derived_my_present_flag indicates No, or when derived_my_present_flag indicates Yes and the mv_signalled_flag of the vertex indicates Yes), instead of the motion vector acquisition unit 202E8, the same processing as the inter-decoding unit 202E shown in FIG. 7 or a modified example thereof is performed.
[0174] According to such a configuration, a decoding calculation of a motion vector and a reduction effect of a coded amount can be expected for a vertex having overlapping vertices.
[0175] In a further modified example of the above-described inter-decoding unit 202E, the overlapping vertex search unit 202E6 does not search for overlapping vertices from all vertices of the basic mesh of the reference frame, but searches for overlapping vertices only from vertices where mv_signalled_flag is No.
[0176] However, the input to the overlapping vertex search unit 202E6 includes mv_signalled_flag in addition to the index (decoding order) and position coordinates of each vertex of the basic mesh of the decoded reference frame.
[0177] According to this modified example, since overlapping vertices are searched only for vertices having overlapping vertices instead of all vertices, a reduction effect of the decoding calculation of the motion vector can be expected.
[0178] In a further modified example of the above-described inter-decoding unit 202E, the mv_signalled_flag acquisition unit 202E7 decodes the mv_signalled_flag in two stages.
[0179] The mv_signalled_flag acquisition unit 202E7 groups N vertices into one group, and for each group, it decodes the mv_group_signalled_flag (the third flag) from the bit stream of the P frame. For all vertices in the group where mv_group_signalled_flag is 1, it sets mv_signalled_flag to 1, and for each vertex in the group where mv_group_signalled_flag is 0, it decodes mv_group_signalled_flag from the bit stream of the P frame.
[0180] According to this modification example, since the mv_group_signalled_flag is decoded in two steps, an effect of reducing the decoding calculation of the motion vector and the coding amount can be expected.
[0181] In a further modification example of the above-mentioned inter-decoding unit 202E, the mv_signalled_flag acquisition unit 202E7 decodes the mv_signalled_flag from the bit stream of the P frame for vertices having duplicate vertices instead of for each vertex. For vertices without duplicate vertices, it does not decode the mv_signalled_flag from the bit stream of the P frame and sets mv_signalled_flag to 1.
[0182] Note that the inter-decoding unit 202E decodes a control signal indicating the number of vertices having duplicate vertices from the bit stream before the mv_signalled_flag.
[0183] Such a control signal enables the decoding of the mv_signalled_flag without performing the processing of the duplicate vertex search unit 202E6.
[0184] Also, as a requirement for the compatibility of the bit stream, it is stipulated that the number of vertices having duplicate vertices output by the duplicate vertex search unit 202E6 coincides with the number of vertices indicated by the control signal.
[0185] In addition, the duplicate vertex search unit 202E6 stores all duplicate vertex indices in a separate list. For example, the duplicate vertex search unit 202E6 may store all duplicate vertex indices in the duplicated_vertex_list as shown in the following (modified example of Method 1).
[0186] (Modified example of Method 1) The duplicate vertex search unit 202E6 sequentially searches for duplicate vertices with matching coordinates as shown below. When a duplicate vertex exists, vRef is vindex1, and when no duplicate vertex exists, vRef is -1.
[0187] vRef = firstVertexIndexDuplicated(vindex0, duplicated_vertex_list) where firstVertexIndexDuplicated(v, duplicated_vertex_list) { for (i = 0; i < v; i++) { if (referenceSubmeshVertexPositions[i] == referenceSubmeshVertexPositions[v]) { duplicated_vertex_list.push_back(v) return i } } return -1 } According to this modified example, since the mv_signalled_flag is provided only for vertices having duplicate vertices instead of all vertices, the decoding calculation of the motion vector and the effect of reducing the amount of code can be expected.
[0188] In a further modified example of the above-mentioned inter-decoding unit 202E, the mv_signalled_flag acquisition unit 202E7 may set the mv_signalled_flag of vertices having no duplicate vertices to 0.
[0189] Here, when the mv_signalled_flag of a vertex without duplicate vertices is 0, in the basic mesh of the reference frame, although there are no duplicate vertices, there are vertices with the same motion vector. Therefore, the index of the vertex with the same motion vector as the vertex is decoded from the bitstream of the inter-frame, and the motion vector of the vertex is obtained.
[0190] Specifically, when MVmode is 2 (when derived_my_present_flag indicates Yes and the mv_signalled_flag of the vertex indicates No), when there are duplicate vertices of such a vertex, the motion vector of the vertex having the index (vindex1) of the duplicate vertex corresponding to the index (vindex0) of the vertex output by the duplicate vertex search unit 202E6 is obtained from the motion vector buffer unit 202E2, and the motion vector of such a vertex is configured to be the motion vector of the vertex. However, in this modified example, when there are no duplicate vertices of the vertex, the index (vindex1) of the vertex having the same motion vector as the vertex is decoded from the bitstream, the motion vector of the vertex having such an index (vindex1) is obtained, and the motion vector of such a vertex is configured to be the motion vector of the vertex.
[0191] According to this modified example, even for a vertex without duplicate vertices, when the mv_signalled_flag indicates No, since the motion vector is obtained from another vertex, the decoding calculation of the motion vector and the effect of reducing the coding amount can be expected.
[0192] Note that there are cases where the above-described modified examples can be used simultaneously and cases where they cannot. When they cannot be used simultaneously, a control signal indicating which modified example to use is provided, the control signal is decoded from the bitstream, and a determination is made as to which modified example to use. However, such a control signal may be an extension of an existing control signal.
[0193] In a further modification example of the above-described inter-decoding unit 202E, if the reference frame is an inter-frame, the duplicate vertex search unit 202E6 reuses the result obtained in the reference frame in the frame to be decoded.
[0194] However, if the reference frame is an intra-frame, the duplicate vertex search unit 202E6 may reuse such information on the assumption that it has information on duplicate vertices.
[0195] Specifically, first, the duplicate vertex search unit 202E6 decodes a control signal from the bit stream indicating whether to reuse the result obtained in the reference frame (information on duplicate vertices including the number of vertices having duplicate vertices) in the frame to be decoded. However, for such a control signal, the duplicate vertex search unit 202E6 may use the existing control signal as it is or after extension.
[0196] Second, when such a control signal is Yes, if the reference frame is an inter-frame, the duplicate vertex search unit 202E6 reuses the result obtained in the reference frame in the frame to be decoded. However, if the reference frame is an intra-frame, the duplicate vertex search unit 202E6 may reuse such information on the assumption that it has information on duplicate vertices.
[0197] Specifically, the output of the duplicate vertex search unit 202E6 related to the reference frame (the above result) is a list that stores the index (vindex1) of such a duplicate vertex if a duplicate vertex related to the index (vindex0) of each vertex exists, and stores the index (vindex0) of the vertex itself or a specific value for which the index cannot be used (for example, -1) if such a duplicate vertex does not exist.
[0198] Here, such a list is stored in the buffer repVert in the order of vindex0.
[0199] In the following example, when there are no duplicate vertices, the index of the vertex (vindex0) itself is saved. Also, to explicitly indicate the index tRef of the frame, it is displayed in the buffer repVert tRef as follows.
[0200] repVert tRef (vindex0) = vRef vRef = vindex1 if vindex0 and vindex1 are duplicate vertices vRef = vindex0 if vindex0 and vindex1 are not duplicate vertices In the frame t to be decoded, when repVert tRef is reused, the duplicate vertex search unit 202E6 does not need to search for duplicate vertices for each vertex (index: vindex0) of the basic mesh of the reference frame of the frame to be decoded, and repVert tRef can be used as it is.
[0201] repVert t (vindex0) = vindex0 if repVert tRef (vindex0) = vindex0 repVert t (vindex0) = repVert tRef (vindex0) if repVert tRef (vindex0)!= vindex0 According to this modified example, since duplicate vertices are not searched for, a reduction effect in the decoding calculation of motion vectors can be expected.
[0202] In a further modified example of the above-mentioned inter-decoding unit 202E, if the reference frame is an inter-frame, the mv_signalled_flag acquisition unit (flag acquisition unit) 202E7 reuses the mv_signalled_flag acquired in the reference frame in the frame to be decoded.
[0203] However, if the reference frame is an intra-frame, the mv_signalled_flag acquisition unit 202E7 may reuse such mv_signalled_flag on the assumption that it has the mv_signalled_flag.
[0204] Specifically, when derived_my_present_flag indicates Yes, the mv_signalled_flag acquisition unit 202E7 decodes, for each vertex, not the mv_signalled_flag from the P-frame bit stream but the difference from the mv_signalled_flag of the reference frame.
[0205] Also, a control signal may be provided when there is no difference in all mv_signalled_flags. In that case, the mv_signalled_flag acquisition unit 202E7 decodes such control signal, and if such control signal is a specific value (for example, TRUE), uses the mv_signalled_flag of the reference frame as it is as the mv_signalled_flag of the target frame, and if such control signal is a specific value (for example, FALSE), further decodes the difference from the mv_signalled_flag of the reference frame to calculate the mv_signalled_flag of the target frame.
[0206] According to this modification example, an effect of reducing the coded amount of mv_signalled_flag can be expected.
[0207] Note that there are cases where the above modification examples can be used simultaneously and cases where they cannot. In the case where they cannot be used simultaneously, a control signal indicating which modification example to use is provided, the control signal is decoded from the bit stream, and a determination is made as to which modification example to use.
[0208] (Modification Example 1 of the Basic Mesh Decoder 202) Hereinafter, with reference to FIGS. 16 and 17, modification example 1 of the basic mesh decoder 202 will be described.
[0209] As shown in FIG. 16, the basic mesh decoding unit 202 according to the first modification example includes a separation unit 202A, an intra decoding unit 202B, a mesh buffer unit 202C, an inter decoding unit 202E, and a skip decoding unit 202F.
[0210] The skip decoding unit 202F is configured to decode the basic mesh of the frame to be decoded by directly using the decoded basic mesh of the specified reference frame.
[0211] In this embodiment, the frame may be either a mesh or a submesh.
[0212] For example, as shown in FIG. 17, "P_SUBMESH" in smh_type may correspond to a P frame, "I_SUBMESH" in smh_type may correspond to an I frame, and "SKIP_SUBMESH" in smh_type may correspond to an S frame.
[0213] (Skip decoding unit 202F) The skip decoding unit 202F is configured to extract the decoded basic mesh (reference decoded basic mesh) of the specified reference frame from the mesh buffer unit 202C, and use the coordinates of the vertices of the extracted reference decoded basic mesh and the indices of the vertices as they are to decode the coordinates of the vertices of the basic mesh of the frame to be decoded and the indices of the vertices.
[0214] Here, the mesh buffer unit 202C has at least one reference frame and is configured to store at least one decoded basic mesh for each reference frame.
[0215] The skip decoding unit 202F may identify the specified reference decoded basic mesh by using a control signal decoded from the bitstream or a predetermined rule.
[0216] For example, such a predetermined rule may be to extract the first reference frame from the reference frame list from the mesh buffer unit 202C, or to extract the reference frame with the closest frame index to the frame to be decoded, etc.
[0217] In this embodiment, a frame that decodes the coordinates of the vertices of the basic mesh by directly using the coordinates of the vertices of the reference decoding basic mesh and the indices of the vertices is called an "S frame".
[0218] According to such a configuration, since the motion vector can be made unnecessary in the skip decoding unit 202F, a significant reduction effect in the amount of code and a significant reduction effect in the amount of calculation can be expected.
[0219] (Mesh buffer unit 202C) The mesh buffer unit 202C is configured to store one or more reference decoding basic meshes in a predetermined order.
[0220] Note that such a basic mesh has metadata such as a frame number and a sub-mesh number, and at least the coordinates of each vertex and the index of the vertex, and is stored in the mesh buffer unit 202C in a predetermined order determined by the reference frame list.
[0221] Here, as shown in FIG. 18, such a reference frame list (ref_list0) is a list of information for specifying all the reference decoding basic meshes stored in the mesh buffer unit 202C.
[0222] As shown in FIG. 18, the reference frame list may be determined by a control signal decoded from the bit stream, or may be calculated naturally from the decoding order of the frames.
[0223] Note that the control signal decoded from the bit stream may be indicated by the relative distance from the frame to be decoded, or may be the frame index with an absolute value.
[0224] Furthermore, a short-term reference frame or a long-term reference frame may be used according to a control signal.
[0225] For example, when using a short-term reference frame, the absolute value (abs_delta_mfoc_st) and its sign (sign_flag) of the difference in the display order between the current frame (cur) and the reference frame (ref) may be decoded from the bit stream, and the display order of the reference frame may be specified by the following formula. If(sign_flag){ Display Order(ref)=Display Order(cur)+abs_delta_mfoc_st }else{ Display Order(ref)=Display Order(cur)-abs_delta_mfoc_st } Also, when a method calculated from the natural decoding order of frames is used, for example, in the reference frame list, when there is no control signal, they may be arranged sequentially by a certain number of frames from the previously decoded frame. That is, the reference frame list may be {0, -1, -2,…, -(N−1)}.
[0226] Basically, the reference frame list does not change for each frame except in special circumstances (for example, when receiving a Re-ordering instruction).
[0227] The mesh buffer unit 202C may be updated as follows.
[0228] When the basic mesh is decoded, in the case of an I-frame and a P-frame, the mesh buffer unit 202C deletes one or more existing reference frames in a predetermined order determined by the reference frame list, and includes one or more basic meshes including the basic mesh of the decoded frame, or creates and includes one basic mesh from a plurality of basic meshes to adjust the order of the reference frames.
[0229] Such deletion operation may be performed only when the mesh buffer unit 202C is full. Note that the number of basic meshes that can be stored in the mesh buffer unit 202C is determined in advance. Here, in the present embodiment, when the number of such basic meshes is reached, it is defined that the mesh buffer unit 202C is full. In the above creation operation, the vertex coordinates corresponding to the basic mesh of the decoded frame and the existing basic meshes stored in the mesh buffer unit 202C may be weighted and averaged to form one basic mesh.
[0230] The weights used in such weighted averaging may be determined in advance, may be calculated using the frame index, or may be decoded from the control signal.
[0231] However, when the mesh buffer unit 202C is an S frame, such update may or may not be performed.
[0232] Note that when the mesh buffer unit 202C receives a control signal indicating a re-ordering instruction from the control signal decoded from the bitstream, as shown in FIG. 19, it updates the reference frame list and adjusts the order of the reference frames in the predetermined order determined by the updated reference frame list (ref_list0).
[0233] (Inter-decoding unit 202E) The inter-decoding unit 202E is configured to decode the vertex coordinates of the P frame by adding the vertex coordinates of the reference frame taken out from the mesh buffer unit 202C and the motion vector decoded from the bitstream of the P frame.
[0234] Furthermore, the inter-decoding unit 202E can adjust the index of the vertices in the P-frame by using the pair of indices A(k) and B(k) of the vertices that exist as duplicate vertices stored in such a specific buffer. All or part of such indices are decoded from the bitstream. Such a decoding method may be arithmetic coding. According to such a configuration, an effect that there is no limit to the maximum value of the index to be decoded can be expected by using arithmetic coding. For example, arithmetic coding such as ue(v) may be used.
[0235] (Modification Example 2 of the Basic Mesh Decoding Unit 202) Hereinafter, with reference to FIG. 20, a modification example 2 of the basic mesh decoding unit 202 will be described.
[0236] Hereinafter, the skip decoding unit 202F will be described, but it may also be applied to the inter-decoding unit 202E.
[0237] As shown in FIG. 20, in the skip decoding unit 202F, in order to enable reference to subsequent frames, the decoding order and the display order are different.
[0238] Here, the display order is the same as the order of the input during encoding and the same as the order of the output during decoding.
[0239] On the other hand, the decoding order is the same as the order of the output during encoding and the same as the order of the input during decoding.
[0240] Note that such a reference frame may be calculated by weighted averaging of one or more other frames and a subsequent frame.
[0241] However, when referring to a plurality of frames including a subsequent frame, as a new frame type (smh_type) in FIG. 14, MR_SUBMESH (MR frame or B frame) is defined, and MR_SUBMESH is decoded from the bitstream.
[0242] Also, as shown in FIG. 21, such another frame may be the decoded frame immediately before the target frame.
[0243] Such a weight may be calculated using the frame interval between the target frame and the subsequent frame and the frame interval between the target frame and the other frame, or may be determined in advance.
[0244] The basic mesh decoder 202 decodes a control signal (smh_mesh_frm_order_cnt_lsb) from the bit stream and decodes the order of such output.
[0245] When there are sub-meshes defined in Non-Patent Document 4 described above, all sub-meshes are set to the same control signal (smh_mesh_frm_order_cnt_lsb) or the control signal (smh_mesh_frm_order_cnt_lsb) is applied to all sub-meshes.
[0246] The value indicated by such a control signal (smh_mesh_frm_order_cnt_lsb) may be the difference from the display order of the frame to be decoded, or may be the order in a previously determined frame group MaxMeshFrmOrderCntLsb.
[0247] When the decoding order (Decode Order) and the display order (Display Order) are different, if the decoded basic meshes are arranged in the decoding order (Decode Order), the basic mesh decoder 202 may rearrange the decoded basic meshes in the display order (Display Order).
[0248] In the S frame where the subsequent frame can be referred to, two mesh buffers 202C may be provided, or when only one mesh buffer unit 202 is provided, there is at least one reference frame including a reference frame whose display order is later than that of the frame to be decoded.
[0249] The skip decoding unit 202F designates a reference frame by receiving a control signal decoded from a bit stream, a predetermined rule, or a Re-ordering instruction.
[0250] Specifically, the skip decoding unit 202F designates a reference frame in the reference frame list by such a control signal.
[0251] Alternatively, the skip decoding unit 202F designates the first reference frame in the reference frame list.
[0252] Alternatively, the skip decoding unit 202F receives a Re-ordering instruction, updates the reference frame list and the reference frame order of the mesh buffer unit 202C, and designates the first reference frame in such a reference frame list.
[0253] Note that in this embodiment, not decoding the S frame does not affect the decoding of other frames. Therefore, when part or all of the S frame is not decoded, Temporal scalability can be realized.
[0254] Furthermore, the basic mesh decoding unit 202 may decode the basic mesh of the S frame by integrating a plurality of reference frames according to a control signal.
[0255] For example, the basic mesh decoding unit 202 may be configured to decode the coordinates and indices of the vertices of the basic mesh of the frame to be decoded by averaging the coordinates of the corresponding vertices in the basic meshes of the two reference frames before and after, and directly using the average coordinates and the indices of the vertices.
[0256] According to such a configuration, it is possible to obtain a high-quality basic mesh while eliminating the need for motion vectors in the skip decoding unit 202F or the inter decoding unit 202E. Therefore, an effect of improving the quality of the decoded mesh can be expected. Furthermore, an effect of realizing Temporal scalability can be expected.
[0257] However, in order to achieve temporal scalability, control signals indicating whether to decode the base mesh, displacement amount, and texture in each frame are defined respectively, and they are decoded from the bitstream respectively.
[0258] Also, within the same frame, the Temporal_IDs of the atlas and the base mesh may be made to match. Also, within the same frame, the Temporal_IDs of the atlas and the texture may be made to match. Also, within the same frame, the Temporal_IDs of the atlas and the displacement amount may be made to match.
[0259] According to such a configuration, an effect of being able to avoid the situation where a frame cannot be decoded or wasteful data can be expected.
[0260] Note that it is desirable that the interval between adjacent frames having the same Temporal_ID be constant.
[0261] Adjacent frames having the same Temporal_ID have the closest POC.
[0262] By making the interval between frames constant as described above, an effect of maintaining a certain frame rate when displaying the decoded frames can be expected.
[0263] Furthermore, the decoding order of the atlas and the base mesh having the same display order may be made to match. Also, the decoding order of the atlas and the displacement amount having the same display order may be made to match. Also, the decoding order of the atlas and the texture having the same display order may be made to match.
[0264] Alternatively, the random access points of the atlas and the basic mesh having the same display order may be made to coincide. Also, the random access points of the atlas and the displacement amount having the same display order may be made to coincide. Also, the random access points of the atlas and the texture having the same display order may be made to coincide. Note that the random access point is defined in Non-Patent Document 4 or Non-Patent Document 5.
[0265] According to such a configuration, when decoding the basic mesh, the displacement amount, and the texture, respectively, an effect that the mesh can be reproduced without waiting for each other's decoding can be expected.
[0266] Furthermore, a frame having a Temporal_ID higher than the control signal Temporal_ID of the frame to be decoded is not used as a reference frame for such a frame to be decoded.
[0267] As a result, an effect that the possibility of discarding the reference frame can be eliminated can be expected.
[0268] An example of realizing Temporal scalability using the above-described Temporal_ID will be described below.
[0269] The bitstreams of the atlas, the basic mesh, the displacement amount, and the texture are encapsulated by a Network Abstraction Layer (NAL) unit. The NAL unit may have an NAL header as shown in FIG. 22.
[0270] The TID defined as the last 3 bits in the NAL header is Temporal_ID + 1. The range of the TID is from 1 to 7, and zero is prohibited.
[0271] The LayerID / R6 defined as the 6 bits immediately before the TID in the NAL header specifies the identifier of the layer to which the NAL unit belongs.
[0272] The value of LayerID / R6 must be within the range of 0 to 62. The value 63 may be specified by ISO / IEC in the future.
[0273] For purposes other than determining the data volume of the bitstream decoding unit, the mesh decoder 200 ignores all data following the value 63 in the NAL unit, and the mesh decoder 200 compliant with the specified profile ignores all NAL units where the value of LayerID-R6 is not 0 (i.e., removes and discards them from the bitstream).
[0274] The value 63 of LayerID / R6 can be used to indicate an extended layer identifier in future extensions.
[0275] When there are sub-meshes defined in Non-Patent Document 4, all sub-meshes will have the same TID or the TID will be applied to all sub-meshes.
[0276] For atlases, Non-Patent Document 5 can be used, and for displacement and texture, HEVC or VVC of the video coding method can be used. Therefore, the following will describe the basic mesh.
[0277] The values of LayerID / R6 of all BMCL NAL units of the encoded basic mesh frame must be the same. The value of LayerID / R6 of the encoded basic mesh frame is the value of LayerID / R6 of the BMCL NAL unit of the encoded basic mesh frame.
[0278] When NALType is equal to NAL_EOB, the value of LayerID / R6 must be equal to 0.
[0279] When the NALType is in the range from NAL_BLA_W_LP to NAL_RSV_BMCL_29 defined in Non-Patent Document 4, that is, when it belongs to an IRAP-coded basic mesh frame, the Temporal_ID must be 0.
[0280] When the NALType is equal to NAL_TSA_R or NAL_TSA_N, the Temporal_ID must not be equal to 0.
[0281] When the NALType is equal to 0 and the NALType is equal to NAL_STSA_R or NAL_STSA_N, the Temporal_ID must not be equal to 0.
[0282] The value of the Temporal_ID must be the same for all BMCL NAL units within an access unit.
[0283] The value of the Temporal_ID of a coded basic mesh frame or an access unit is the value of the Temporal_ID of the BMCL NAL unit of the coded basic mesh frame or the access unit.
[0284] The value of the Temporal_ID of a sublayer representation is the maximum value of the Temporal_IDs of all BMCL NAL units within the sublayer representation.
[0285] The value of the Temporal_ID of a non-BMCL NAL unit is restricted as follows. - When the NALType is equal to NAL_BMSPS, the Temporal_ID must be 0, and the Temporal_ID of the access unit containing the NAL unit must be 0. - In other cases, when the NALType is equal to NAL_EOS or NAL_EOB, the Temporal_ID must be 0. - In other cases, when the NALType is equal to NAL_AUD or NALLFDD, the Temporal_ID shall be equal to the Temporal_ID of the access unit containing the NALL unit. - In other cases, the Temporal_ID shall be greater than or equal to the Temporal_ID of the access unit containing the NAL unit.
[0286] Note that when the NAL unit is not BMCL, the value of the Temporal_ID shall be equal to the minimum value of the Temporal_ID values of all access units to which the non - BMCL NAL unit is applied.
[0287] When the NALType is equal to NAL_BMFPS, the Temporal_ID can be greater than or equal to the Temporal_ID of the access unit included because all basic mesh frame parameter sets (BMFPS) are included at the beginning of the bitstream where the Temporal_ID of the first encoded basic mesh frame is 0.
[0288] Note that the skip decoding unit 202F refers to the specified tIDTarget and discards NAL units with a Temporal_ID higher than tIDTarget without decoding them.
[0289] Here, tIDTarget may be specified by a pre - determined value, or may be specified according to the network situation and the terminal capabilities of the mesh decoding device 200.
[0290] For example, a lower tIDTarget is specified for the wireless case than for the wired case. Also, a lower tIDTarget is specified when the network situation is poor. Also, a lower tIDTarget is specified when decoding is performed by a low - spec mesh decoding device 200.
[0291] However, as a requirement for the compatibility of the bitstream, it is stipulated that there must be at least one NAL unit whose Temporal_ID is not higher than tIDTarget in the bitstream.
[0292] Note that, as shown in FIG. 23, the number of sub-meshes may be different for each frame (intra-frame, inter-frame, and skip frame).
[0293] In such a case, the intra decoder unit 202B, the inter decoder unit 202E, and the skip decoder unit 202F assign non-overlapping sub-mesh IDs to each of the sub-meshes in each frame.
[0294] Also, as shown in FIG. 24, the intra decoder unit 202B, the inter decoder unit 202E, and the skip decoder unit 202F may assign different SubmeshIDs (sub-mesh IDs) to corresponding sub-meshes between frames.
[0295] However, it is assumed that the inter decoder unit 202E or the skip decoder unit 202F can only refer to sub-meshes having the same SubmeshID among the reference frames.
[0296] Alternatively, it is assumed that the inter decoder unit 202E or the skip decoder unit 202F can only refer to sub-meshes having the same number of vertices among the reference frames.
[0297] Alternatively, it is assumed that the intra decoder unit 202B and the inter decoder unit 202E can refer to the specified sub-meshes among the reference frames.
[0298] In such a case, when there are multiple sub-meshes in the reference frame, the inter decoder unit 202E or the skip decoder unit 202F may decode a control signal that specifies the SubmeshID of the referable sub-mesh from the bitstream of the current sub-mesh.
[0299] On the other hand, when there is only one sub-mesh in the reference frame, the inter-decoding unit 202E or the skip decoding unit 202F may use such a sub-mesh as a referenceable sub-mesh.
[0300] However, when the above control signal does not exist, the inter-decoding unit 202E or the skip decoding unit 202F sets the SubmeshID of the referenceable sub-mesh to be the same as the SubmeshID of the sub-mesh in the current frame.
[0301] In addition, the inter-decoding unit 202E or the skip decoding unit 202F may decode a control signal indicating whether the above control signal exists from the bit stream.
[0302] Note that the inter-decoding unit 202E or the skip decoding unit 202F may also decode a control signal for selecting a method for determining the above referenceable sub-mesh.
[0303] The subdivision unit 203 and the displacement amount decoding unit 206 may follow Non-Patent Document 4.
[0304] According to the present invention, the amount of calculation can be reduced by reusing the reference frame itself without searching for decoded adjacent vertices.
[0305] In addition, according to the present embodiment, in inter-prediction coding, even when the number of vertices of the basic mesh of the current frame is different from the number of vertices of the reference frame or the reference sub-mesh, the basic mesh of the current frame can be decoded.
[0306] In addition, according to the present embodiment, in inter-prediction coding, it is possible to avoid a situation where the number of vertices of the basic mesh of the current frame is different from the number of vertices of the reference frame or the reference sub-mesh.
[0307] Also, according to the present embodiment, in inter-prediction coding, by introducing a control signal indicating which sub-mesh to reference among the reference frames of the current frame, it is possible to specify which sub-mesh to reference.
[0308] Furthermore, according to the present embodiment, in inter-prediction coding, even if there is no information indicating which sub-mesh to reference among the reference frames of the current frame, it is possible to specify which sub-mesh to reference.
[0309] Also, according to the present embodiment, it is possible to ensure that the basic mesh has at least one face.
[0310] Also, according to the present embodiment, the Temporal_scalability function can be realized.
[0311] Furthermore, according to the present embodiment, the coding efficiency of the mesh can be improved.
[0312] The above-described mesh encoding device 100 and mesh decoding device 200 may be realized by a program that causes a computer to execute each function (each process).
Industrial Applicability
[0313] Note that according to the present embodiment, for example, since it is possible to realize an improvement in overall service quality in video communication, it is possible to contribute to Goal 9 of the Sustainable Development Goals (SDGs) led by the United Nations, which is to "build resilient infrastructure, promote sustainable industrialization, and foster innovation."
Explanation of Signs
[0314] 1…Mesh processing system 100…Mesh encoding device 200…Mesh decoding device 201…Multiplex separation unit 202…Basic mesh decoding unit 202A… Separation unit 202B… Intra decoding unit 202B1… Arbitrary intra decoding unit 202B2… Sorting unit 202C… Mesh buffer unit 202D… Connection information decoding unit 202E… Inter decoding unit 202E1… Motion vector residual decoding unit 202E2… Motion vector buffer unit 202E3… Motion vector prediction unit 202E4… Motion vector calculation unit 202E5… Adder 202E6… Duplicate vertex search unit 202E7… mv_signalled_flag acquisition unit 202E8… Motion vector acquisition unit 202F… Skip decoding unit 203… Subdivision unit 204… Mesh decoding unit 205… Patch integration unit 206… Displacement amount decoding unit 207… Video decoding unit 208… Atlas data decoding unit
Claims
1. A mesh decoding device, comprising: an inter-decoding unit that decodes the coordinates of a vertex to be decoded by adding a motion vector decoded from a bitstream of an inter-frame and the coordinates of a vertex corresponding to the vertex to be decoded in a reference frame; the inter-decoding unit includes a motion vector prediction unit that calculates a predicted value of the motion vector of the vertex to be decoded by averaging all or part of the motion vectors of the decoded vertices adjacent to the vertex to be decoded with reference to an adjacent vertex list that is a list of decoded vertices adjacent to each vertex; the motion vector prediction unit is characterized in that, as the adjacent vertex list in the current frame, the adjacent vertex list in the reference frame is reused.
2. The mesh decoding device according to claim 1, wherein the motion vector prediction unit has a one-to-one correspondence between each vertex in the reference frame and each vertex in the current frame, the decoding order of each vertex in the reference frame and the decoding order of each vertex in the current frame are the same, and when the condition that the adjacent vertex list including the decoded vertices picked up up to the maximum utilization number in the reference frame has already been saved is satisfied, the adjacent vertex list in the reference frame is reused as the adjacent vertex list in the current frame.
3. The motion vector prediction unit: when the reference frame is decoded, if the reference frame is a P frame or an S frame, saves the adjacent vertex list in the reference frame in a specific buffer; when the current frame is decoded, if the type of the reference frame is a P frame or an S frame, reuses the adjacent vertex list saved in the specific buffer as the adjacent vertex list in the current frame.
4. The mesh decoding device according to claim 3, wherein the motion vector prediction unit reuses the adjacent vertex list in the frame corresponding to the reference frame as the adjacent vertex list in the current frame when the adjacent vertex list including the decoded vertices picked up up to the maximum utilization number in a plurality of frames is saved in the specific buffer.
5. The motion vector prediction unit applies an operation on a reference frame buffer that stores the decoded reference frame to an operation on the specific buffer, in the mesh decoding apparatus according to claim 3.
6. In addition to the above conditions, when the condition that the reference frame is the frame immediately preceding the current frame in decoding order is satisfied, the motion vector prediction unit reuses the adjacent vertex list in the reference frame as the adjacent vertex list in the current frame, in the mesh decoding apparatus according to claim 2.
7. When the reference frame is decoded, if the reference frame is a P frame or an S frame, the motion vector prediction unit stores only the adjacent vertex list in the immediately preceding frame in the specific buffer, in the mesh decoding apparatus according to claim 6.
8. A mesh decoding method, comprising a step of decoding the coordinates of a vertex to be decoded by adding a motion vector decoded from a bit stream of an inter-frame and the coordinates of a vertex corresponding to the vertex to be decoded in a reference frame, in the step, by referring to an adjacent vertex list which is a list of decoded vertices adjacent to each vertex, and averaging all or part of the motion vectors of the decoded vertices adjacent to the vertex to be decoded, calculating a predicted value of the motion vector of the vertex to be decoded, characterized in that the adjacent vertex list in the reference frame is reused as the adjacent vertex list in the current frame, in the mesh decoding method.
9. A program for causing a computer to function as a mesh decoding apparatus, wherein the mesh decoding apparatus comprises an inter-decoding unit that decodes the coordinates of a vertex to be decoded by adding a motion vector decoded from a bit stream of an inter-frame and the coordinates of a vertex corresponding to the vertex to be decoded in a reference frame, the inter-decoding unit comprises a motion vector prediction unit that calculates a predicted value of the motion vector of the vertex to be decoded by referring to an adjacent vertex list which is a list of decoded vertices adjacent to each vertex, and averaging all or part of the motion vectors of the decoded vertices adjacent to the vertex to be decoded, The motion vector prediction unit reuses the adjacent vertex list in the reference frame as the adjacent vertex list in the current frame, characterized by the program.