Information processing device and method
The method addresses the challenge of maintaining displacement vector accuracy and bit depth in 3D mesh encoding by using a higher bit depth displacement coefficient group to generate information for a lower bit depth group, effectively encoding this information into a displacement map.
Patent Information
- Application Number
- PCT/JP2024/041647
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-11-25
- Publication Date
- 2025-06-19
AI Technical Summary
Existing methods for encoding 3D meshes, such as V-DMC, face challenges in suppressing the increase in bit depth of displacement maps while maintaining the accuracy of displacement vectors, which is crucial for maintaining the quality of restored meshes.
The proposed solution involves using a first displacement coefficient group with a higher bit depth to generate displacement information for a second displacement coefficient group with a lower bit depth, and then encoding this information into a displacement map. This approach includes a displacement information generation unit, a packing unit, and an encoding unit to manage the bit depth and accuracy of displacement vectors.
This method effectively suppresses the increase in bit depth of the displacement map while maintaining the accuracy of displacement vectors, thereby ensuring the quality of the restored mesh.
Smart Images

Figure JP2024041647_19062025_PF_FP_ABST
Abstract
Description
Information processing device and method
[0001] The present disclosure relates to an information processing device and method, and more particularly to an information processing device and method that can suppress an increase in the bit depth of a disparity map while suppressing a decrease in the accuracy of a disparity vector.
[0002] Conventionally, V-DMC (Video-based Dynamic Mesh Coding) has been used as a method for encoding meshes, which are 3D data that represent the three-dimensional structure of an object using vertices and connections (see, for example, Non-Patent Document 1). In V-DMC, an original mesh is represented by a decimated base mesh and a displacement vector. The displacement vector is encoded by storing a displacement coefficient corresponding to the displacement vector in a displacement map.
[0003] Generally, the bit depth of the displacement map is set to 10 bits, and the displacement map is encoded and decoded using a 10-bit encoder / decoder. However, an 8-bit encoder / decoder is more versatile and less expensive. Furthermore, the smaller the bit depth of the displacement map, the more efficient the coding.
[0004] Khaled Mammou, Jungsun Kim, Alexis Tourapis, Dimitri Podborski, Krasimir Kolarov, "[V-CG] Apple's Dynamic Mesh Coding CfP Response", ISO / IEC JTC 1 / SC 29 / WG 7 m59281, April 2022
[0005] However, reducing the bit depth of the displacement map may result in an insufficient bit depth for the displacement coefficients, which may reduce the accuracy of the displacement coefficients (i.e., the displacement vectors). Because the accuracy of the displacement vectors has a significant impact on the quality of the reconstructed mesh, it is desirable to minimize the reduction in accuracy as much as possible.
[0006] The present disclosure has been made in view of such circumstances, and makes it possible to suppress an increase in the bit depth of a displacement map while suppressing a decrease in the accuracy of a displacement vector.
[0007] an encoding unit configured to encode a displacement video having the displacement map as a frame image; and a displacement information generation unit configured to generate, for a base mesh, displacement information relating to a second displacement coefficient set including second displacement coefficients of a second bit depth lower than the first bit depth, using a first displacement coefficient set including first displacement coefficients of a first bit depth; a packing unit configured to generate a displacement map of the second bit depth in which the second displacement coefficient set is stored as a pixel value set; and an encoding unit configured to encode a displacement video having the displacement map as a frame image. The base mesh is a mesh with lower resolution than an original mesh to be encoded, which is composed of vertices and connections that represent a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh; the first displacement coefficients are coefficients corresponding to displacement vectors that indicate displacements of vertices or division points of the base mesh; the division points are vertices generated by subdividing the base mesh; the first displacement coefficient set is a coefficient set corresponding to a displacement vector set corresponding to the base mesh; and the displacement information is information regarding the displacement indicated by the displacement vector set corresponding to the base mesh.
[0008] an information processing method according to one aspect of the present technology; using a first set of displacement coefficients including first displacement coefficients of a first bit depth, to generate, for a base mesh, displacement information regarding a second set of displacement coefficients including second displacement coefficients of a second bit depth lower than the first bit depth; generating a displacement map of the second bit depth in which the second set of displacement coefficients is stored as a set of pixel values; and encoding a displacement video using the displacement map as frame images; the base mesh is a mesh with lower resolution than an original mesh to be encoded, the original mesh being composed of vertices and connections that represent a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh; the first displacement coefficients are coefficients corresponding to displacement vectors that indicate displacements of vertices or division points of the base mesh; the division points are vertices generated by subdividing the base mesh; the first set of displacement coefficients are a coefficient set corresponding to a set of displacement vectors corresponding to the base mesh; and the displacement information is information regarding the displacement indicated by the set of displacement vectors corresponding to the base mesh.
[0009] According to another aspect of the present technology, there is provided an information processing device including: a decoding unit configured to generate a displacement video having frame images of a displacement map that stores, as a pixel value group, a first displacement coefficient group including first displacement coefficients of a first bit depth by decoding encoded data; and a displacement coefficient generation unit configured to generate, using displacement information regarding the first displacement coefficient group extracted from the displacement map, a second displacement coefficient group including second displacement coefficients of a second bit depth higher than the first bit depth, wherein the second displacement coefficients are coefficients corresponding to displacement vectors that indicate displacements of vertices or division points of a base mesh, the base mesh being a mesh with lower resolution than the original mesh to be encoded and consisting of vertices and connections that represent a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh, the original mesh being composed of vertices and connections that represent the three-dimensional structure of the object, and the division points are vertices generated by subdividing the base mesh, the displacement information is information regarding the displacement indicated by the displacement vector group corresponding to the base mesh, and the second displacement coefficient group is a coefficient group corresponding to the displacement vector group corresponding to the base mesh.
[0010] Another aspect of the present technology is an information processing method for decoding encoded data to generate a displacement video having frame images of a displacement map that stores, as a pixel value group, a first displacement coefficient group including first displacement coefficients of a first bit depth, and generates, using displacement information related to the first displacement coefficient group extracted from the displacement map, a second displacement coefficient group including second displacement coefficients of a second bit depth higher than the first bit depth, the second displacement coefficients being coefficients corresponding to displacement vectors that indicate displacements of vertices or division points of a base mesh, the base mesh being a mesh with lower resolution than the original mesh to be encoded that is composed of vertices and connections that represent a three-dimensional structure of an object and that is generated by thinning out vertices from the original mesh, the division points being vertices generated by subdividing the base mesh, the displacement information being information related to the displacement indicated by a displacement vector group corresponding to the base mesh, and the second displacement coefficient group being a coefficient group that corresponds to the displacement vector group corresponding to the base mesh.
[0011] In an information processing device and method according to one aspect of the present technology, a first group of displacement coefficients including first displacement coefficients of a first bit depth is used to generate displacement information for a base mesh regarding a second group of displacement coefficients including second displacement coefficients of a second bit depth lower than the first bit depth, a displacement map of the second bit depth is generated in which the second group of displacement coefficients is stored as a group of pixel values, and a displacement video using the displacement map as frame images is encoded.
[0012] In an information processing device and method according to another aspect of the present technology, encoded data is decoded to generate a displacement video having a displacement map as a frame image, the displacement map storing a first group of displacement coefficients including first displacement coefficients of a first bit depth as a group of pixel values, and displacement information regarding the first group of displacement coefficients extracted from the displacement map is used to generate a second group of displacement coefficients including second displacement coefficients of a second bit depth higher than the first bit depth.
[0013] 1 is a diagram illustrating a mesh. FIG. 1 is a diagram illustrating V-DMC. FIG. 1 is a diagram illustrating an example of the relationship between dQP and the number of vertices where the bit depth of the displacement coefficient exceeds the bit depth of the displacement map. FIG. 1 is a diagram illustrating an example of a displacement map. FIG. 2 is a diagram illustrating an example of a method for transmitting displacement information. FIG. 3 is a diagram illustrating an example of applying divided displacement coefficients as displacement information. FIG. 4 is a diagram illustrating an example of an array of displacement coefficients after division. FIG. 5 is a diagram illustrating an example of dislocation coefficient arrangement. FIG. 6 is a diagram illustrating an example of applying excess information as displacement information. FIG. 7 is a diagram illustrating an example of transmitting a signed excess amount. FIG. 8 is a diagram illustrating an example of transmitting excess information for each code. FIG. 9 is a diagram illustrating an example of transmitting an unsigned excess amount. FIG. 10 is a diagram illustrating an example of transmitting an index difference value. FIG. 11 is a diagram illustrating an example of applying a shifted displacement coefficient as displacement information. FIG. 12 is a diagram illustrating an example of shifting displacement coefficients. FIG. 13 is a diagram illustrating an example of displacement coefficients. FIG. 14 is a block diagram illustrating an example of the main configuration of an encoding device. FIG. 15 is a flowchart illustrating an example of the flow of an encoding process. FIG. 16 is a flowchart illustrating an example of the flow of an encoding process. FIG. 17 is a block diagram illustrating an example of the main configuration of a decoding device. FIG. 18 is a flowchart illustrating an example of the flow of a decoding process. Fig. 1 is a flowchart illustrating an example of the flow of a decoding process. Fig. 2 is a flowchart illustrating an example of the flow of a decoding process. Fig. 3 is a block diagram illustrating an example of the main configuration of a computer.
[0014] Hereinafter, modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described. The description will be made in the following order: 1. Literature etc. supporting technical content and technical terminology 2. Transmission of displacement coefficients 3. Transmission of displacement information 4. First embodiment (encoding device) 5. Second embodiment (decoding device) 6. Supplementary notes
[0015] <1. Literature, etc. supporting technical content and technical terminology> The scope of disclosure of the present technology includes not only the content described in the embodiments, but also the content described in the following non-patent documents, etc. that were publicly known at the time of filing, and the content of other documents referenced in the following non-patent documents.
[0016] Non-patent document 1: (mentioned above)
[0017] In other words, the contents of the above-mentioned non-patent documents and the contents of other documents referenced in the above-mentioned non-patent documents are also used as the basis for determining the support requirements.
[0018] <2. Transmission of Displacement Coefficients> <V-DMC> Conventionally, 3D data representing the three-dimensional structure of a three-dimensional structure (object with a three-dimensional shape) has been available as a mesh, which represents the three-dimensional shape of the object surface by forming polygons with vertices and connections (also called edges).
[0019] As shown in the upper left of Figure 1, in a mesh, vertices 11 and connections 12 connecting these vertices 11 form polygonal planes (polygons). These polygons (also called faces) represent the surface of a three-dimensional object, i.e., the three-dimensional shape of the object. A texture 13 can be applied to each face of this mesh.
[0020] Mesh data is composed of information such as that shown in the lower part of Figure 1. Vertex information 14, shown first from the left in the lower part of Figure 1, is information indicating the three-dimensional position (three-dimensional coordinates (X, Y, Z)) of each vertex 11 that constitutes the mesh. Connection information 15, shown second from the left in the lower part of Figure 1, is information indicating each connection (edge) 12 that constitutes the mesh. A texture image 16, shown third from the left in the lower part of Figure 1, is map information for the texture 13 that is applied to each face. A UV map 17, shown fourth from the left in the lower part of Figure 1, is information indicating the correspondence between the vertices 11 and the texture 13. The UV map 17 indicates the coordinates (UV coordinates) of each vertex 11 in the texture image 16.
[0021] As an example of such a mesh coding method, there is V-DMC (Video-based Dynamic Mesh Coding) as disclosed in Non-Patent Document 1.
[0022] In V-DMC, the mesh to be encoded (referred to in this specification as the original mesh) is represented as a base mesh that is less fine (i.e., coarser) than the original mesh, and displacement vectors of the division points obtained by subdividing the base mesh, and the base mesh and displacement vectors are then encoded.
[0023] For example, assume that there is an original mesh as shown in the top row of Figure 2. The original mesh is a mesh composed of vertices and connections that represent the three-dimensional structure of an object, and is the target of encoding. For example, the original mesh is generated from a captured image of an object in real space (by camera capture). In Figure 2, black dots represent vertices, and lines connecting the black dots represent connections (edges). As described above, a mesh essentially forms polygons using vertices and edges, but for convenience of explanation, it is described here as a group of vertices connected linearly (in series).
[0024] By simplifying the original mesh, a coarse (low-resolution) mesh like the one shown in the second row from the top of Figure 2 is formed. This is called the base mesh. One simplification method is to thin out some of the vertices (decimate). In other words, the base mesh is a mesh with lower resolution than the original mesh, generated by thinning out vertices from the original mesh (i.e., simplifying the original mesh).
[0025] By subdividing each polygon of this base mesh, vertices and edges are added, as shown in the third row from the top of Figure 2. The degree of subdivision is arbitrary. That is, the number of vertices and edges added is arbitrary. For example, this subdivision can add vertices equal to the number of vertices thinned out from the original mesh. That is, subdivision can be used to maintain the same number of vertices as the original mesh. In this specification, these added vertices are also referred to as division points. This subdivision can also be repeated recursively. For example, in a technique called midpoint, the process of adding vertices to the midpoints of edges (subdivision) is repeated recursively. In other words, recursive subdivision increases the number of vertices and improves the resolution of the mesh. In this way, it is possible to perform subdivision up to any desired level of resolution (i.e., control the resolution of the subdivided mesh). In other words, the subdivided mesh can be layered according to its level of resolution. In other words, this can be considered a layering of the subdivision process and the vertices (division points) and edges obtained by the subdivision process.
[0026] However, the connections of the base mesh have been updated when the vertices of the original mesh are thinned out. Therefore, the division points obtained by subdivision are formed on these updated connections (edges). As a result, the shape of the subdivided base mesh differs from that of the original mesh. More specifically, as shown in the bottom part of Figure 2, the positions of the division points (on the dotted line) differ from those of the original mesh.
[0027] In other words, by moving the positions of the vertices of the subdivided base mesh (the vertices or division points of the base mesh) closer to the vertex positions of the original mesh, the difference in shape between the subdivided base mesh and the original mesh can be reduced. In this specification, such movement of the vertices of the subdivided base mesh (the vertices or division points of the base mesh) is also referred to as displacement. Furthermore, the amount and direction of this displacement, expressed as a vector, is also referred to as a displacement vector. Ideally, by displacing each vertex of the subdivided base mesh, the shape of the subdivided base mesh can be made to match the shape of the original mesh. In other words, the original mesh can be expressed as a base mesh and a displacement vector.
[0028] In V-DMC, such base meshes and displacement vectors are coded instead of the original mesh (geometry). By coding the base meshes and displacement vectors in this way, it is possible to code with a reduced number of polygons (i.e., the number of vertices and edges) compared to coding the original mesh, which generally reduces the amount of code for the same quality. In other words, it is possible to improve coding efficiency.
[0029] During decoding, as described above, a mesh is restored (generated) by subdividing a base mesh and applying a displacement vector to each vertex of the subdivided base mesh to displace it. In this specification, this restored mesh is also referred to as a restored mesh. Ideally, a restored mesh equivalent to the original mesh can be generated. Note that although the shape of the polygon (face) may be any polygonal shape, the following description will be given assuming that the polygon is triangular. Therefore, in the following, a polygon (face) will also be referred to as a triangle.
[0030] <Encoding and Decoding of V-DMC Data> In the case of V-DMC, mesh data consists of a base mesh, displacement vectors, attributes, and atlas information. This data group is also referred to as V-DMC data. The base mesh consists of information indicating vertices and connections, and is encoded using an existing mesh encoding method such as Draco.
[0031] Displacement vectors are converted into displacement coefficients using a predetermined method. These displacement coefficients are arranged as pixel values in a two-dimensional area (also called a displacement map). This arrangement (mapping) of displacement coefficients is also called packing. A moving image (also called a displacement video) with the displacement map as its frame image is coded using a coding method for 2D moving images. In other words, the displacement coefficients are scalar values corresponding to the displacement vectors. A displacement map is map information (also called image data) that stores the displacement coefficients as pixel values. A displacement video is moving image data with the displacement map as its frame image.
[0032] An attribute is information other than geometry that is applied to a mesh (geometry), which is 3D data. For example, an attribute may include a texture applied to a face of the mesh (geometry). The attribute (e.g., texture) is divided into multiple subregions, each of which is projected in a predetermined projection direction, and the projected images (patches) are placed in a two-dimensional region (also called an attribute map). A video (also called an attribute video) using the attribute map as frame images is encoded using a 2D video encoding method. In other words, the attribute map is map information (also called image data) that stores the patches (projected textures) as pixel values. The attribute video is video data using the attribute map as frame images.
[0033] Atlas information is information used when reconstructing a mesh. For example, atlas information may include correspondence between the base mesh and a displacement map or attribute map (such as a UV map), quantized values of displacement vectors, etc. This atlas information is encoded using a predetermined encoding method.
[0034] The coded data (bitstream) of each data is decoded by a decoding method corresponding to the coding method. In other words, by decoding the coded data (bitstream), various information such as base meshes, displacement vectors, attributes, and atlas information is restored (generated).
[0035] <Encoding / Decoding of Displacement Coefficients> Generally, the bit depth of the displacement map is 10 bits, and the displacement video is encoded and decoded using an encoder / decoder that supports a YUV component bit depth of 10 bits. However, encoders / decoders that support a YUV component bit depth of 8 bits are more versatile and cheaper than those that support a YUV component bit depth of 10 bits. In addition, the lower the bit depth of the displacement map, the more efficient the coding.
[0036] However, reducing the bit depth of the displacement map may result in an insufficient bit depth for the displacement coefficients to be stored, which may reduce the precision of the displacement coefficients (i.e., displacement vectors).Since the precision of the displacement vectors has a significant impact on the quality of the reconstructed mesh, it is desirable to minimize the reduction in precision as much as possible.
[0037] For example, if the displacement coefficients could be stored in the displacement map by reducing the bits that exceed the bit depth of the displacement map, the accuracy of the displacement coefficients (i.e., displacement vectors) would be reduced, and there was a risk that the reconstructed mesh would fail at a practical level.
[0038] Displacement coefficients are derived, for example, by wavelet transforming and quantizing displacement vectors. Lowering the quantization parameter (also referred to as displacement QP or dQP) improves the accuracy of displacement coefficients, but increases their values. That is, as shown in the table of FIG. 3 , lowering dQP increases the number of displacement coefficients whose bit depth exceeds the displacement map (i.e., displacement coefficients whose bit depth is higher than the displacement map). The table in FIG. 3 shows examples of the number of vertices whose displacement coefficient bit depth exceeds the displacement map in two sequences (mesh of sequence A and mesh of sequence B) for each dQP value. For example, when dQP=14, there are no vertices whose displacement coefficient bit depth exceeds the displacement map in either the mesh of sequence A or the mesh of sequence B. In contrast, when dQP=10, there are six vertices whose displacement coefficient bit depth exceeds the displacement map in the mesh of sequence A and six vertices in the mesh of sequence B. In contrast, when dQP=6, there are 72 vertices in the mesh of sequence A and 48 vertices in the mesh of sequence B where the bit depth of the displacement coefficient exceeds the displacement map.
[0039] In other words, a higher displacement QP allows for a better fit of the bit depth, but reduces the precision of the displacement coefficients (i.e., displacement vectors).
[0040] The displacement coefficients are arranged so as to be grouped by layer (also referred to as LoD (Level Of Detail)), as in the displacement map 41 shown in FIG. 4 . That is, the displacement coefficients of each layer are arranged in consecutive pixel groups on the displacement map 41 to form an area. In the example of FIG. 4 , the displacement coefficients of LoD1 are stored in the LoD1 area on the displacement map 41. Similarly, the displacement coefficients of LoD2 are stored in the LoD2 area on the displacement map 41. The displacement coefficients of LoD3 are stored in the LoD3 area on the displacement map 41.
[0041] <3. Transmission of Displacement Information> <Transmission of Displacement Information> Therefore, as shown in the top row of the table in Fig. 5 , displacement information including a set of displacement coefficients with a bit depth of the displacement map corresponding to a set of displacement coefficients including displacement coefficients with a bit depth higher than that of the displacement map is transmitted.
[0042] Hereinafter, an information processing device that encodes 3D data including a base mesh and a displacement vector will also be referred to as a first information processing device. For example, the first information processing device may include a displacement information generation unit configured to generate, for the base mesh, displacement information related to a second displacement coefficient set including displacement coefficients at a second bit depth lower than the first bit depth using a first displacement coefficient set including displacement coefficients at the first bit depth; a packing unit configured to generate a displacement map at the second bit depth in which the second displacement coefficient set is stored as a pixel value set; and an encoding unit configured to encode a displacement video using the displacement map as a frame image. In the present disclosure, displacement coefficients included in the first displacement coefficient set may be referred to as first displacement coefficients. Also, displacement coefficients included in the second displacement coefficient set may be referred to as second displacement coefficients.
[0043] Also, in the first information processing device, a first displacement coefficient group including displacement coefficients of a first bit depth is used to generate displacement information for a base mesh regarding a second displacement coefficient group including displacement coefficients of a second bit depth lower than the first bit depth, a displacement map of the second bit depth in which the second displacement coefficient group is stored as a pixel value group, and a displacement video using the displacement map as a frame image is encoded.
[0044] The base mesh is a mesh with lower resolution than the original mesh to be coded, which is composed of vertices and connections that represent the three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh. The displacement coefficients are coefficients corresponding to displacement vectors that indicate the displacement of vertices or division points of the base mesh. The division points are vertices generated by subdividing the base mesh. The first displacement coefficient group is a group of coefficients corresponding to the displacement vector group corresponding to the base mesh. The displacement information is information regarding the displacement indicated by the displacement vector group corresponding to the base mesh.
[0045] By doing so, the first information processing apparatus can suppress an increase in the bit depth of the disparity map while suppressing a decrease in the accuracy of the disparity vector.
[0046] In the following description, an information processing device that decodes a bitstream and generates 3D data including a base mesh and a displacement vector is also referred to as a “second information processing device.” The second information processing device includes: a decoding unit configured to generate a displacement video having frame images as a displacement map in which a first displacement coefficient set including displacement coefficients at a first bit depth is stored as a pixel value set by decoding encoded data; and a displacement coefficient generation unit configured to generate a second displacement coefficient set including displacement coefficients at a second bit depth higher than the first bit depth, using displacement information related to the first displacement coefficient set extracted from the displacement map.
[0047] In addition, in a second information processing device, by decoding the encoded data, a displacement video is generated in which a displacement map storing a first group of displacement coefficients including displacement coefficients of a first bit depth as a group of pixel values is used as a frame image, and a second group of displacement coefficients including displacement coefficients of a second bit depth higher than the first bit depth is generated using displacement information regarding the first group of displacement coefficients extracted from the displacement map.
[0048] The displacement coefficients are coefficients corresponding to displacement vectors indicating the displacement of the vertices or division points of the base mesh. The base mesh is a mesh with lower resolution than the original mesh to be coded, which is composed of vertices and connections that represent the three-dimensional structure of the object, and is generated by thinning out vertices from the original mesh. The division points are vertices generated by subdividing the base mesh. The displacement information is information regarding the displacement indicated by the displacement vector group corresponding to the base mesh. The second displacement coefficient group is a coefficient group corresponding to the displacement vector group corresponding to the base mesh.
[0049] By doing so, the second information processing device can suppress an increase in the bit depth of the disparity map while suppressing a decrease in the accuracy of the disparity vector.
[0050] <Method 1> For example, as shown in the second row from the top of the table in Figure 5, the displacement information may be transmitted as a displacement map in which one displacement coefficient is divided into multiple displacement coefficients and stored in multiple pixels (Method 1).
[0051] For example, in the above-mentioned first information processing device, the displacement information generation unit may be configured to divide each displacement coefficient (first displacement coefficient) of the first displacement coefficient group into a plurality of displacement coefficients (second displacement coefficients) of a second bit depth, and generate, as displacement information, a second displacement coefficient group composed of the displacement coefficients after the division.
[0052] In the second information processing device described above, the first displacement coefficient group may be configured by first displacement coefficients obtained by dividing the second displacement coefficients into a plurality of first displacement coefficients, and the displacement coefficient generation unit may be configured to generate the second displacement coefficient group by combining corresponding first displacement coefficients of the first displacement coefficient group extracted from the displacement video.
[0053] In other words, multiple pixels of the displacement map are used to store displacement coefficients, so that displacement coefficients with a bit depth higher than the displacement map can be stored in the displacement map without reducing the amount of information.
[0054] For example, assume that the bit depth of the displacement coefficient 101 shown in FIG. 6 is 10 bits, and the bit depth of the displacement map that stores the displacement coefficient 101 is 8 bits. The displacement coefficient 101 is divided into displacement coefficients 102-1 and 102-2, each with a bit depth of 8 bits. The divided displacement coefficients 102-1 and 102-2 are then stored in the displacement map. That is, in this case, two pixels (pixel0, pixel1) of the displacement map are used to store the displacement coefficient 101. Simply comparing the bit lengths, the displacement coefficients 102-1 and 102-2 are each 8 bits, for a total of 16 bits. Therefore, the 10-bit displacement coefficient 101 can be divided into the displacement coefficients 102-1 and 102-2 without reducing the amount of information.
[0055] During decoding, the displacement coefficient 101 can be restored by combining the displacement coefficients 102-1 and 102-2 extracted from the displacement map. As described above, the amount of information is not reduced, so ideally, the displacement coefficient 101 can be restored without degradation.
[0056] Here, when there is no need to distinguish between the displacement coefficients 102-1 and 102-2 after division, they are also referred to as displacement coefficients 102. Note that the displacement coefficient 101 may be divided into any number of displacement coefficients 102. One displacement coefficient 101 may be divided into three displacement coefficients 102, four displacement coefficients 102, or five or more displacement coefficients 102.
[0057] The bit depths of the displacement coefficients 101 before division and the displacement coefficients 102 after division may be any, but it is desirable that they be in a relationship that allows division without reducing the amount of information of the displacement coefficients 101 before division. Also, in order to reduce the bit depth of the encoder / decoder, it is desirable that the bit depth of the displacement coefficients 102 after division (i.e., the bit depth of the displacement map) be lower than the maximum bit depth of the displacement coefficients 101 before division.
[0058] Note that a group of displacement coefficients generated from one displacement coefficient before division can be considered a set, and each displacement coefficient after division can be considered an element. In this specification, this element is also referred to as a post-division element. Each displacement coefficient to be divided is converted into a set of the same configuration as each other by the division. For example, displacement coefficient 101 in FIG. 6 is divided into two types of post-division elements: a post-division element stored in pixel0 and a post-division element stored in pixel1. Other displacement coefficients to be divided are similarly divided into the two types of post-division elements. The same applies when dividing into three or more parts. Each displacement coefficient to be divided is divided into sets of the same configuration as each other (the same number of types of post-division elements).
[0059] In this specification, "division of displacement coefficients" means generating multiple displacement coefficients from one displacement coefficient. "Combining displacement coefficients" means generating one displacement coefficient from multiple displacement coefficients. In other words, division and composition are inverse processes. Note that while it is desirable for this division and composition to be lossless processing, errors may be included as long as they do not cause practical problems.
[0060] Any method of division and combination may be used. For example, values may be divided and combined. Bit strings may be divided and combined. A predetermined calculation may be performed on values or bit strings. The division may be equal or biased. For example, a displacement coefficient of "50" may be divided into two displacement coefficients of "25". A displacement coefficient of "50" may be divided into a displacement coefficient of "20" and a displacement coefficient of "30". The division may be between the upper and lower bit strings of the displacement coefficient. The division may be between odd and even bits of the displacement coefficient. The division may also be biased in terms of the number of bits, such as the upper 6 bits and the lower 4 bits. Division may also be performed using division, or the remainder of the division may be added to one side. The division may also be performed by multiplying a weighting coefficient. Although the division into two parts has been described as an example, the same applies to division into three or more parts.
[0061] An example of this division (encoding process) and synthesis (decoding process) is shown as a pseudo program in box 103 in Figure 6. First, the division (encoding process) will be explained. "displacements" on the first line indicates the displacement coefficient before division. "1 << 7" indicates that the value "1" is shifted left by 7 bits. That is, on the first line (dispValue = displacements + (1 << 7)), the value obtained by adding 128 to the displacement coefficient "displacements" is stored in the variable dispValue.
[0062] In the second line (pixel0 = dispValue / 2 + dispValue % 2), the quotient (dispValue / 2) obtained by dividing the variable dispValue by the value "2" is added to the remainder (dispValue % 2), and the resulting value is stored as the displacement coefficient "pixel0" after division. In the third line (pixel1 = dispValue / 2), the quotient (dispValue / 2) obtained by dividing the variable dispValue by the value "2" is stored as the displacement coefficient "pixel1" after division.
[0063] In other words, in this example, the displacement coefficient is divided into two equal parts, and the remainder is added to one of them. By dividing in this way, the values of the displacement coefficients after division (the values of each divided element) become approximately the same, which makes it possible to suppress an increase in the difference between pixel values in the displacement map and suppress a decrease in coding efficiency.
[0064] Next, the synthesis (decoding process) will be described. In the first line (dispValue = pixel0 + pixel1), the value obtained by adding the displacement coefficients "pixel0" and "pixel1" is stored in the variable dispValue. This value is the value obtained by adding 128 to the displacement coefficient before division. Therefore, in the second line (Displacements = dispValue - (1 << 7)), the displacement coefficient "displacements" is derived by subtracting 128 (1 << 7) from the value of the variable dispValue. In this way, the value of the displacement coefficient before division can be easily derived from the displacement coefficient after division.
[0065] Another example of division (encoding process) and synthesis (decoding process) is shown as a pseudo program within box 104 in FIG. 6 . In the division (encoding process), the quotient (dispValue / α) obtained by dividing the variable dispValue (which is derived from the pre-division displacement coefficient "displacements" in the same manner as in the example of box 103) by the variable α is stored as the post-division displacement coefficient "pixel0" in the first line (pixel0 = dispValue / α). Furthermore, the difference between the variable dispValue and the post-division displacement coefficient "pixel0" is stored as the post-division displacement coefficient "pixel1" in the second line (pixel1 = dispValue - pixel0). In the synthesis (decoding process), the addition result of the post-division displacement coefficient "pixel0" and the post-division displacement coefficient "pixel1" is stored as the variable dispValue in the first line (dispValue = α*pixel0 + pixel1). As a result, the displacement coefficient "displacements" is derived in the same manner as in the example of box 103. In this way, the value of the displacement coefficient before division can be easily derived from the displacement coefficient after division.
[0066] Of course, the calculation example is arbitrary and is not limited to the example of FIG.
[0067] This allows the displacement coefficients to be stored and transmitted in a displacement map with a lower bit depth than the displacement coefficients without reducing the amount of information contained in the displacement coefficients, thereby preventing the increase in bit depth of the displacement map while preventing a reduction in the accuracy of the displacement vectors.
[0068] <Method 1-1> When Method 1 is applied, as shown in the third row from the top of the table in Figure 5, the displacement coefficients after division may be arranged together for each post-division element in the displacement map (Method 1-1).
[0069] For example, in the first information processing device described above, the displacement information generation unit may be configured to arrange the divided displacement coefficients together for each divided element. More specifically, the displacement information generation unit may be configured to divide the first displacement coefficients into first divided coefficients and second divided coefficients as second displacement coefficients, and arrange the second coefficient sequence representing a group of the second divided coefficients consecutively with the first coefficient sequence representing a group of the first divided coefficients. Here, the first divided coefficients and the second divided coefficients each correspond to one of the divided elements. The packing unit may then generate a displacement map in which the divided displacement coefficients are arranged to form pixel regions for each divided element according to the arrangement. More specifically, the packing unit may be configured to generate a displacement map in which the second pixel region representing the second coefficient sequence is arranged consecutively with the first pixel region representing the first coefficient sequence according to the arrangement of the first coefficient sequence and the second coefficient sequence. In the second information processing device described above, the displacement map may include a first pixel region in which a first coefficient sequence representing a group of first division coefficients is arranged, and a second pixel region in which a second coefficient sequence representing a group of second division coefficients is arranged, and the displacement coefficient generation unit may be configured to generate the second displacement coefficient group by combining the corresponding first division coefficients and second division coefficients extracted from the displacement video.
[0070] As described with reference to Fig. 4, when the displacement coefficients are stored in the displacement map, they are arranged together for each layer (each LoD). At this time, the displacement coefficients are arranged so that displacement coefficients of the same LoD are consecutive, as in the displacement coefficient sequence 111-1 in Fig. 7, and are arranged in the displacement map in that order.
[0071] When dividing the displacement coefficients as described above, the divided displacement coefficients may be arranged together for each division element and placed in the displacement map in the order of the arrangement. For example, suppose the displacement coefficients of LoD1 are divided into two. In this case, as in the displacement coefficient sequence 111-2 of FIG. 7 , the displacement coefficient group (LoD1(pixel0)) of the division element (pixel0) and the displacement coefficient group (LoD1(pixel1)) of the division element (pixel1) may be arranged together (so that the displacement coefficients of the same division element are consecutive), and placed in the displacement map in the order of the arrangement. More specifically, a second division coefficient group representing the division element pixel1 is placed consecutively to a first division coefficient group (coefficient sequence) representing the division element pixel0. Note that, in the present disclosure, the division coefficient corresponding to pixel0 may be larger than the division coefficient corresponding to pixel1. In other words, the first division coefficient / division element pixel0 may be a reference division element in which the remainder when the first displacement coefficient is divided is stored. When the displacement coefficient is divided into three or more parts, the second division coefficient may correspond to a plurality of division elements, such as pixel1, pixel2, etc. In this way, displacement coefficients of the same division elements having the same LoD can be arranged in the same pixel region, thereby further suppressing a decrease in coding efficiency.
[0072] <Method 1-2> When Method 1 is applied, as shown in the fourth row from the top of the table in Figure 5, the displacement coefficients after division may be arranged together for each pre-division displacement coefficient in the displacement map (Method 1-2).
[0073] For example, in the first information processing device described above, the displacement information generation unit may be configured to arrange the divided displacement coefficients together for each pre-division displacement coefficient. More specifically, the displacement information generation unit may be configured to divide a first displacement coefficient into a first division coefficient and a second division coefficient, respectively, to generate a set of the first division coefficient and the second division coefficient as the second displacement coefficient, and to arrange the set of the first division coefficient and the second division coefficient consecutively according to the arrangement of the first displacement coefficient before division. The packing unit may then be configured to generate a displacement map in which a group of pixels in which a plurality of divided displacement coefficients corresponding to one pre-division displacement coefficient are consecutively arranged according to the arrangement / arrangement. More specifically, the packing unit may be configured to generate a displacement map in which a group of pixels representing the arrangement of the set of the first division coefficient and the second division coefficient is arranged according to the arrangement of the set of the first division coefficient and the second division coefficient. In the second information processing device, the displacement map may include pixel groups representing the arrangement of sets of first and second division coefficients obtained by dividing the second displacement coefficient, and the displacement coefficient generation unit may be configured to generate the second displacement coefficient set by combining the corresponding first and second division coefficients extracted from the displacement video.
[0074] That is, when dividing displacement coefficients as described above, the divided displacement coefficients may be arranged together for each displacement coefficient before division, and then arranged in the displacement map in the same order. For example, assume that the displacement coefficients of LoD1 are divided into two. In this case, as shown in the displacement coefficient sequence 111-3 of FIG. 7, the displacement coefficients of the divided element (pixel0) and the displacement coefficient of the divided element (pixel0) generated from the same displacement coefficients of LoD1 may be arranged consecutively, and then arranged in the displacement map in the same order. That is, in this example, the displacement coefficients of the divided element (pixel0) and the displacement coefficient of the divided element (pixel0) are arranged alternately. This allows displacement coefficients of the same LoD to be arranged in the same pixel region. Therefore, a decrease in coding efficiency can be further suppressed.
[0075] <Method 1-3> When Method 1 is applied, the displacement coefficients after division may be arranged in reverse order in the displacement map, as shown in the fifth row from the top of the table in FIG. 5 (Method 1-3).
[0076] For example, in the first information processing device described above, the displacement information generation unit may be configured to arrange the divided displacement coefficients in a predetermined order, i.e., to arrange the second displacement coefficients as a coefficient sequence in a predetermined order. The divided displacement coefficients may then be arranged in the displacement map in the order of their arrangement, starting from the bottom right pixel of the displacement map in a reverse scanning order. Furthermore, in the second information processing device described above, the divided displacement coefficients may be arranged in the displacement map in a predetermined order, starting from the bottom right pixel of the displacement map in a reverse scanning order. That is, the displacement map may include a coefficient sequence in which the second displacement coefficients are arranged in a predetermined order, starting from the bottom right pixel of the displacement map in a reverse scanning order.
[0077] Typically, displacement coefficients are arranged starting from the upper left pixel of the displacement map, as in the example of FIG. 4 , but this arrangement pattern may be reversed. That is, displacement coefficients may be arranged starting from the lower right pixel of the displacement map in the reverse scanning order. In this case, the arrangement order of the displacement coefficient sequence is arbitrary. For example, if displacement coefficients are arranged like the displacement coefficient sequence 111-1 of FIG. 7 , the displacement coefficient sequence 111-1 may be arranged starting from the lower right pixel in the reverse scanning order, as in the displacement map 121 of FIG. 8 . Furthermore, if displacement coefficients are arranged like the displacement coefficient sequence 111-2 of FIG. 7 , the displacement coefficient sequence 111-2 may be arranged starting from the lower right pixel in the reverse scanning order, as in the displacement map 122 of FIG. 8 . Furthermore, if displacement coefficients are arranged like the displacement coefficient sequence 111-3 of FIG. 7 , the displacement coefficient sequence 111-3 may be arranged starting from the lower right pixel in the reverse scanning order, as in the displacement map 123 of FIG. 8 .
[0078] Since image codecs encode from the top left, packing flat high LoD images from the top left improves encoding efficiency. In other words, by arranging the displacement coefficients in the displacement map as described above, it is possible to further suppress a decrease in encoding efficiency. It is also possible to select whether to arrange the displacement coefficients from the top left or the bottom right.
[0079] <Method 2> For example, as shown in the sixth row from the top of the table in FIG. 5, excess information indicating a displacement coefficient that exceeds the bit depth of the displacement map may be transmitted as the displacement information (Method 2).
[0080] For example, in the first information processing device described above, the displacement information generation unit may be configured to generate, as displacement information, excess information indicating displacement coefficients included in the first displacement coefficient set that exceed the second bit depth, and to generate a second displacement coefficient set by replacing the displacement coefficients in the first displacement coefficient set that exceed the second bit depth with values of the second bit depth, and the encoding unit may be configured to encode the excess information and encode the displacement video in which the second displacement coefficient set is stored as a pixel value set.
[0081] In the second information processing device, the decoding unit may be configured to generate a displaced video and excess information indicating displaced coefficients included in the second displaced coefficient set that exceed the first bit depth by decoding the encoded data, and the displaced coefficient generation unit may be configured to generate the second displaced coefficient set using the first displaced coefficient set and the excess information extracted from the displaced video.
[0082] That is, displacement coefficients that exceed the bit depth of the displacement map may be transmitted separately (as excess information) without being stored in the displacement map.
[0083] By transmitting excess information in this manner, displacement coefficients can be transmitted using a displacement map with a lower bit depth than the displacement coefficients without reducing the amount of information contained therein, thereby preventing the increase in bit depth of the displacement map while preventing the accuracy of the displacement vectors from being reduced.
[0084] The excess information may be in any format as long as it can indicate displacement coefficients that exceed the bit depth of the disparity map. For example, information indicating displacement coefficients that exceed the bit depth of the disparity map may be compiled into a list, and the list may be transmitted as the excess information.
[0085] <Method 2-1> When Method 2 is applied, as shown in the seventh row from the top of the table in FIG. 5, the identification information of the target displacement coefficient and a list of the excess amount may be transmitted as the excess information (Method 2-1).
[0086] For example, in the first information processing device described above, the excess information may be information indicating identification information of displacement coefficients that exceed the second bit depth and the amount of excess. Also, in the second information processing device described above, the excess information may be information indicating identification information of displacement coefficients that exceed the first bit depth and the amount of excess.
[0087] That is, a target displacement coefficient (a displacement coefficient that exceeds the bit depth of the displacement map) may be represented by its identification information and an excess amount. Here, excess means exceeding the bit depth of the displacement map (i.e., the range of values that can be represented in the displacement map). Excess can be in both a positive direction (values that are too large) and a negative direction (values that are too small). The excess amount indicates the value that exceeds the range.
[0088] For example, suppose displacement coefficient sequence 131 in Fig. 9 is an array of displacement coefficients with a bit depth of 10 bits. The 123rd element of displacement coefficient sequence 131 is a displacement coefficient with a value of "-138." Also, suppose the 285th element is a displacement coefficient with a value of "207." These displacement coefficient values exceed the value range (-128 to 128) for an 8-bit bit depth.
[0089] Therefore, in this case, a displacement coefficient sequence 132 with a bit depth of 8 bits and excess information 133 indicating these displacement coefficients are generated and transmitted. At this time, instead of the above-mentioned displacement coefficients, any value expressible in 8 bits (any value between −128 and 128) is stored (replaced with a dummy value) in the 123rd and 285th elements of the displacement coefficient sequence 132. In this way, a displacement coefficient sequence 132 with a bit depth of 8 bits can be generated. During decoding, the values of the 123rd and 285th elements of the displacement coefficient sequence 131 are restored based on the excess information 133, and the values of the 123rd and 285th elements of the displacement coefficient sequence 132 are replaced with the restored values. In this way, the displacement coefficient sequence 131 can be restored.
[0090] 9, the excess information 133 is a list of identification information (Idx) of displacement coefficients and their excess amounts (values). From this list, the value of the 123rd element, "-138," and the value of the 285th element, "207," can be easily derived. In other words, based on the excess information, displacement coefficients with higher bit depths can be more easily derived than from the displacement map.
[0091] The displacement coefficient sequence 132 is stored in a displacement map and transmitted. The excess information 133 may be transmitted as data separate from the displacement map. For example, the excess information 133 may be transmitted as atlas information. The excess information 133 may be coded using a predetermined method, or may not be coded at all. The excess information 133 may be stored in a displacement map and transmitted.
[0092] The content (structure) of the excess information is arbitrary.
[0093] <Method 2-1-1> When Method 2-1 is applied, the excess amount included in the excess information (list) may include a sign, as shown in the eighth row from the top of the table in FIG. 5 (Method 2-1-1).
[0094] For example, in the first information processing device and the second information processing device described above, the excess amount may be information including a positive or negative sign. In the present disclosure, the positive or negative sign may be simply referred to as a sign.
[0095] In other words, in this method, when the value of the displacement coefficient exceeds (goes above) the upper limit of the range of values that can be represented by the displacement map (the range based on the bit depth), the excess is expressed as a positive value, and when it exceeds (goes below) the lower limit of that range, the excess is expressed as a negative value.
[0096] An example of the encoding process in this case is shown in box 141-1 in Figure 10. Note that the bit depth of the displacement map is 8 bits, the value range is "255", and the displacement coefficient shift amount (shift) is "128". Also, the identification information of the vertex corresponding to the displacement coefficient (dispValue) to be processed is set to vertexIdx. The displacement coefficient (dispValue) is shifted (+128) and the value is stored in the variable dshift (dshift = dispValue + shift).
[0097] If this variable dshift is negative, the displacement coefficient exceeds in the negative direction. Therefore, the identification information (vertexIdx) of the vertex to be processed and the excess amount (value) are added to the list (excess information). The excess amount (value) is stored as the value of the variable dshift. In other words, a negative value is stored. The variable dshift is replaced with a dummy value of 8 bits or less (dshift - shift < 0). This dummy value may be any value. For example, it may be set to suppress a decrease in coding efficiency.
[0098] On the other hand, if the variable dshift is greater than the variable range, the displacement coefficient exceeds in the positive direction. Therefore, the identification information (vertexIdx) of the vertex to be processed and the excess amount (value) are added to the list (excess information). The excess amount (value) is stored as the value of (dshift - range). In other words, a positive value is stored. The variable dshift is replaced with a dummy value of 8 bits or less (dshift - shift > 0). This dummy value may be any value. For example, it may be set to suppress a decrease in coding efficiency.
[0099] An example of the decoding process in this case is shown in box 141-2 in Figure 10. Note that the bit depth of the displacement map is 8 bits, the value range is "255", and the displacement coefficient shift amount is "128". Also, the identification information of the vertex corresponding to the displacement coefficient (dispValue) to be processed is set to vertexIdx.
[0100] If the identification information (vertexIdx) of the vertex to be processed exists in the list (if (posOfDispInALoD == vertexId)), and its excess amount (value) is negative (if (value < 0)), the excess amount (value) is stored in the variable dshift. If the excess amount (value) is positive, (range + value) is stored in the variable dshift. The displacement coefficient is derived by shifting the variable dshift inversely (-128) (dshift -= shift).
[0101] In this way, it is easier to derive displacement factors for higher bit depths based on excess information than the displacement map.
[0102] <Method 2-1-2> When Method 2-1 is applied, a list of excess amounts by positive and negative signs may be transmitted as shown in the ninth row from the top of the table in FIG. 5 (Method 2-1-2).
[0103] For example, in the first information processing device and the second information processing device described above, the excess information may be information for each positive and negative sign of the excess amount.
[0104] In other words, with this method, the list is divided into two parts: when the displacement coefficient value exceeds (is above) the upper limit of the range of values that can be represented by the displacement map (the range based on the bit depth), and when it exceeds (is below) the lower limit of that range.
[0105] An example of the encoding process in this case is shown in box 142-1 in Figure 11. Note that the bit depth of the displacement map is 8 bits, the value range is "255", and the displacement coefficient shift amount (shift) is "128". Also, the identification information of the vertex corresponding to the displacement coefficient (dispValue) to be processed is set to vertexIdx. The displacement coefficient (dispValue) is shifted (+128) and the value is stored in the variable dshift (dshift = dispValue + shift).
[0106] If the variable dshift is negative, the displacement coefficient exceeds the negative direction. Therefore, the identification information (vertexIdx) of the vertex to be processed and the excess amount (value) are added to the list for the negative direction (lsitMinus). The excess amount (value) is set to (-1 * dshift). The variable dshift is replaced with a dummy value of 8 bits or less (dshift - shift < 0). This dummy value may be any value. For example, it may be set to suppress a decrease in coding efficiency.
[0107] On the other hand, if the variable dshift is greater than the variable range, the displacement coefficient exceeds the positive direction. Therefore, the identification information (vertexIdx) of the vertex to be processed and the excess amount (value) are added to the list for the positive direction (listPlus). The excess amount (value) is set to the value of (dshift - range). The variable dshift is replaced with a dummy value of 8 bits or less (dshift - shift > 0). This dummy value may be any value. For example, it may be set to suppress a decrease in coding efficiency.
[0108] An example of the decoding process in this case is shown in box 142-2 in Figure 11. Note that the bit depth of the displacement map is 8 bits, the value range is "255", and the displacement coefficient shift amount is "128". Also, the identification information of the vertex corresponding to the displacement coefficient (dispValue) to be processed is set to vertexIdx.
[0109] If the identification information (vertexIdx) of the vertex to be processed is in the negative direction list (listMinus), (-1 * value) is stored in the variable dshift. Also, if the identification information (vertexIdx) of the vertex to be processed is in the positive direction list (listPlus), (range + value) is stored in the variable dshift. The displacement coefficient is derived by shifting the variable dshift inversely (-128) (dshift -= shift).
[0110] In this way, it is easier to derive displacement factors for higher bit depths based on excess information than the displacement map.
[0111] <Method 2-1-3> When Method 2-1 is applied, the excess amount does not need to include a positive or negative sign, as shown in the tenth row from the top of the table in FIG. 5 (Method 2-1-3).
[0112] For example, in the first information processing device and the second information processing device described above, the excess amount may be information that does not include a positive or negative sign.
[0113] In this method, whether the value of the displacement coefficient exceeds (is above) the upper limit of the range of values representable by the displacement map (a range based on the bit depth) or exceeds (is below) the lower limit of that range is identified by dshift - shift < 0 or dshift - shift > 0.
[0114] An example of the encoding process in this case is shown in box 143-1 in Figure 12. Note that the bit depth of the displacement map is 8 bits, the value range is "255", and the displacement coefficient shift amount (shift) is "128". Also, the identification information of the vertex corresponding to the displacement coefficient (dispValue) to be processed is set to vertexIdx. The displacement coefficient (dispValue) is shifted (+128) and the value is stored in the variable dshift (dshift = dispValue + shift).
[0115] If the variable dshift is negative, the displacement coefficient exceeds the negative direction. Therefore, the identification information (vertexIdx) of the vertex to be processed and the excess amount (value) are added to the list (excess information). The excess amount (value) is set to a value of (-1 * dshift). The variable dshift is replaced with a dummy value of 8 bits or less (where dshift - shift < 0). This dummy value may be any value. For example, it may be set to suppress a decrease in coding efficiency.
[0116] On the other hand, if the variable dshift is greater than the variable range, the displacement coefficient exceeds the positive direction. Therefore, the identification information (vertexIdx) of the vertex to be processed and the excess amount (value) are added to the list (excess information). The excess amount (value) is set to the value of (dshift - range). The variable dshift is replaced with a dummy value of 8 bits or less (where dshift - shift > 0). This dummy value may be any value. For example, it may be set to suppress a decrease in coding efficiency.
[0117] An example of the decoding process in this case is shown in box 143-2 in Figure 12. Note that the bit depth of the displacement map is 8 bits, the value range is "255", and the displacement coefficient shift amount is "128". Also, the identification information of the vertex corresponding to the displacement coefficient (dispValue) to be processed is set to vertexIdx.
[0118] If the identification information (vertexIdx) of the vertex to be processed is in the list, if dshift - shift < 0, it is determined to be an excess in the negative direction, and (-1 * value) is stored in the variable dshift. Also, if dshift - shift < 0 is not true, it is determined to be an excess in the positive direction, and (range + value) is stored in the variable dshift. The displacement coefficient is derived by shifting the variable dshift inversely (-128) (dshift -= shift).
[0119] In this way, it is easier to derive displacement factors for higher bit depths based on excess information than the displacement map.
[0120] <Method 2-1-4> When Method 2-1 is applied, a list of difference values and excess amounts of identification information may be transmitted as shown in the eleventh row from the top of the table in FIG. 5 (Method 2-1-4).
[0121] For example, in the first information processing device described above, the excess information may include, as information indicating the identification information, a difference value between identification information of displacement coefficients that exceed the second bit depth. Also, in the second information processing device described above, the excess information may include, as information indicating the identification information, a difference value between identification information of displacement coefficients that exceed the first bit depth.
[0122] That is, in this method, the identification information (Idx) is transmitted as a differential value. Other than that, it is the same as in method 2-1-1. When the value of the displacement coefficient exceeds the upper limit of the range of values that can be represented by the displacement map (the range based on the bit depth), the excess amount is expressed as a positive value, and when it exceeds the lower limit of that range, the excess amount is expressed as a negative value.
[0123] An example of the encoding process in this case is shown in box 144-1 in Figure 13. Note that the bit depth of the displacement map is 8 bits, the value range is "255", and the displacement coefficient shift amount (shift) is "128". Also, the identification information of the vertex corresponding to the displacement coefficient (dispValue) to be processed is set to vertexIdx. Also, the identification information of the vertex (displacement coefficient) immediately before the one that exceeds the displacement map range is set to prevVertexIdx, and its initial value is set to "0". The value obtained by shifting the displacement coefficient (dispValue) (+128) is stored in the variable dshift (dshift = dispValue + shift).
[0124] If this variable dshift is negative, the displacement coefficient exceeds in the negative direction. Therefore, the difference value (vertexIdx - prevVertexIdx) of the identification information of the vertex to be processed and the excess amount (value) are added to the list (excess information). The excess amount (value) is stored as the value of the variable dshift. In other words, a negative value is stored. The variable dshift is replaced with a dummy value of 8 bits or less (dshift - shift < 0). This dummy value may be any value. For example, it may be set to suppress a decrease in coding efficiency.
[0125] On the other hand, if the variable dshift is greater than the variable range, the displacement coefficient exceeds in the positive direction. Therefore, the difference value (vertexIdx - prevVertexIdx) of the identification information of the vertex to be processed and the excess amount (value) are added to the list (excess information). The excess amount (value) is stored as the value of (dshift - range). In other words, a positive value is stored. The variable dshift is replaced with a dummy value of 8 bits or less (dshift - shift > 0). This dummy value may be any value. For example, it may be set to suppress a decrease in coding efficiency.
[0126] An example of the decoding process in this case is shown in box 144-2 in Figure 13. Note that the bit depth of the displacement map is 8 bits, the value range is "255", and the displacement coefficient shift amount (shift) is "128". Also, the identification information of the vertex corresponding to the displacement coefficient (dispValue) to be processed is set to vertexIdx. The identification information of the vertex (displacement coefficient) immediately before the one that exceeds the displacement map range is set to prevVertexIdx, and its initial value is set to "0".
[0127] If the identification information of the vertex to be processed (vertexIdx + prevVertexIdx) is present in the list, and its excess amount (value) is negative (if (value < 0)), the excess amount (value) is stored in the variable dshift. If the excess amount (value) is positive, (range + value) is stored in the variable dshift. The displacement coefficient is derived by inversely shifting the variable dshift (-128) (dshift -= shift).
[0128] By doing so, it is possible to suppress a decrease in coding efficiency.
[0129] <Method 3> For example, as shown in the 12th row from the top of the table in FIG. 5, a displacement map storing a group of displacement coefficients shifted by an offset may be transmitted as the displacement information (Method 3).
[0130] For example, in the first information processing device described above, the displacement information generating unit may be configured to generate a second set of displacement coefficients as displacement information by shifting the first set of displacement coefficients by an amount corresponding to the offset.
[0131] In the second information processing device, the first set of displacement coefficients may be obtained by shifting the second set of displacement coefficients by an offset, and the displacement coefficient generation unit may be configured to generate the second set of displacement coefficients by inversely shifting the first set of displacement coefficients extracted from the displacement video by the offset.
[0132] For example, suppose that the displacement coefficients are biased toward the upper end of the displacement map range (0 to 256), greatly exceeding the upper limit of 256, as shown in graph 151-1 in Fig. 14. In such a case, by subtracting the offset indicated by the double-headed arrow from each displacement coefficient, the excess amount can be reduced, as shown in graph 151-2. In the example of Fig. 14, the excess is 7 bits in graph 151-1, whereas the excess is reduced to 4 bits in graph 151-2.
[0133] In this way, when the displacement coefficients are biased relative to the range of the displacement map, the range of the displacement map can be more effectively utilized by shifting the displacement coefficients by an offset amount. Therefore, the displacement coefficients can be transmitted using a displacement map with a lower bit depth than the displacement coefficients without reducing the amount of information contained in the displacement coefficients. Therefore, the increase in the bit depth of the displacement map can be suppressed while suppressing a reduction in the accuracy of the displacement vectors.
[0134] An example of the encoding process for this method is shown in box 152-1 of Figure 15. As shown in this example, the displacement factor is shifted by the offset (dshift = dispValue + shift - offset). An example of the decoding process is shown in box 152-2 of Figure 15. As shown in this example, the displacement factor is shifted back by the offset (dshift -= shift + offset).
[0135] In this way, the shift of the displacement coefficient can be easily realized.
[0136] <Method 3-1> When Method 3 is applied, as shown in the thirteenth row from the top of the table in FIG. 5, it is not necessary to transmit the offset (Method 3-1).
[0137] For example, in the first information processing device and the second information processing device described above, the offset may be a predetermined value.
[0138] That is, the encoder and decoder may shift or reverse shift the data by a predetermined offset amount. In this case, since the shift amount is known, there is no need to transmit the offset value. Therefore, it is possible to suppress a decrease in coding efficiency.
[0139] <Method 3-2> When Method 3 is applied, an offset may be transmitted as shown in the 14th row from the top of the table in FIG. 5 (Method 3-2).
[0140] For example, in the first information processing device described above, the displacement information generation unit may be configured to set an offset and shift the first displacement coefficient set by the set offset to generate a second displacement coefficient set, and the encoding unit may be configured to encode the offset and encode the displacement video in which the second displacement coefficient set is stored as a pixel value set.
[0141] In the second information processing device, the decoding unit may be configured to generate a displaced video and an offset by decoding the encoded data, and the displacement coefficient generation unit may be configured to generate a second set of displacement coefficients by reverse-shifting a first set of displacement coefficients extracted from the displaced video by the generated offset.
[0142] That is, the encoder sets an offset value to shift, and the decoder uses the offset value to perform the inverse shift. In this way, any value can be applied as the offset, so a more appropriate offset value can be set depending on, for example, the displacement coefficient. Therefore, the bit depth of the displacement map can be used more effectively in a wider variety of cases.
[0143] <Method 3-3> When Method 3 is applied, an independent offset may be applied for each LoD, as shown in the 15th row from the top of the table in FIG. 5 (Method 3-3).
[0144] For example, in the first information processing device described above, the offset may have an independent value for each subdivision level of the base mesh, and the displacement information generating unit may be configured to generate the second set of displacement coefficients by shifting the first displacement coefficient to be processed in the first set of displacement coefficients by the offset value corresponding to the level.
[0145] In the second information processing device, the offset may have an independent value for each subdivision level of the base mesh, and the displacement coefficient generation unit may be configured to generate the second set of displacement coefficients by shifting the first displacement coefficients to be processed in the first set of displacement coefficients extracted from the displacement video back by the offset value corresponding to the level.
[0146] By doing so, an offset according to the characteristics of each LoD can be applied, and the bit depth of the disparity map can be used more effectively.
[0147] <Method 3-4> When Method 3 is applied, an independent offset may be applied to each positive and negative sign of the displacement coefficient, as shown in the 16th row from the top of the table in FIG. 5 (Method 3-4).
[0148] For example, in the first information processing device described above, the offset may have an independent value for each positive or negative sign of the displacement coefficient (e.g., the first transform coefficient), and the displacement information generating unit may be configured to generate the second displacement coefficient set by shifting the displacement coefficient to be processed in the first displacement coefficient set by an offset value corresponding to the sign of the displacement coefficient.
[0149] In the second information processing device described above, the offset may have an independent value for each positive or negative sign of the displacement coefficient (e.g., the second transform coefficient), and the displacement coefficient generation unit may be configured to generate the second displacement coefficient set by reverse-shifting a displacement coefficient (first transform coefficient) to be processed in the first displacement coefficient set extracted from the displacement video by an offset whose value corresponds to the sign of the second transform coefficient corresponding to the first displacement coefficient.
[0150] For example, there may be cases where the bias differs between positive displacement coefficients and negative displacement coefficients. By allowing the positive shift amount (offset value) and the negative shift amount (offset value) to be set independently of each other, such cases can be handled appropriately. In other words, an offset according to the characteristics of each code of the displacement coefficient can be applied, thereby making more effective use of the bit depth of the displacement map.
[0151] <Method 4> For example, as shown in the 17th row from the top of the table in FIG. 5, control information may be transmitted (Method 4).
[0152] For example, in the first information processing device described above, the displacement information generation unit may be configured to generate control information for controlling processing related to the displacement information, and the encoding unit may be configured to encode the generated control information.
[0153] In the second information processing device, the decoding unit may be configured to generate control information for controlling processing related to the displacement information by decoding the encoded data, and the displacement coefficient generation unit may be configured to generate a second set of displacement coefficients based on the generated control information.
[0154] In this case, the method for generating the displacement information is arbitrary. Any of the above-described methods 1 to 3, or a combination of these methods, may be applied, or the displacement information may be generated by another method.
[0155] In this way, the transmission of the above-mentioned displacement information can be controlled. Therefore, the displacement information can be transmitted according to the situation. Therefore, it is possible to suppress an increase in the bit depth of the displacement map while suppressing a decrease in the accuracy of the displacement vector. Furthermore, since the decoder only needs to control the processing of the displacement information based on the control information, it is possible to suppress an increase in the load of the decoding processing.
[0156] <Method 4-1> When Method 4 is applied, execution control information may be transmitted as shown in the 18th row from the top of the table in FIG. 5 (Method 4-1).
[0157] For example, in the first information processing device described above, the control information may include execution control information indicating whether to generate displacement information. Also, in the second information processing device described above, the control information may include execution control information indicating whether to apply the displacement information. Also, the displacement coefficient generation unit may be configured to generate a second displacement coefficient group using the displacement information when the execution control information indicates that the displacement information is to be applied.
[0158] For example, an execution control flag may be transmitted to indicate whether or not to apply the displacement information and transmit displacement coefficients with a bit depth higher than the displacement map. The decoding side can control whether or not to apply the displacement information based on the execution control flag. Also, an enable flag may be transmitted to indicate whether or not to apply the displacement information and transmit displacement coefficients with a bit depth higher than the displacement map.
[0159] In this way, it is possible to control the application of the disparity information depending on the situation, and therefore it is possible to suppress an increase in the bit depth of the disparity map while suppressing a decrease in the accuracy of the disparity vector. Furthermore, since the decoder only needs to control the execution of processing on the disparity information based on the control information, it is possible to suppress an increase in the load of the decoding process.
[0160] <Method 4-1-1> When Method 4-1 is applied, LoD designation information may be transmitted as shown in the 19th row from the top of the table in FIG. 5 (Method 4-1-1).
[0161] For example, in the first information processing device described above, the execution control information may include information specifying a subdivision level of the base mesh. Also, in the second information processing device described above, the execution control information may include information specifying a subdivision level of the base mesh. The displacement coefficient generation unit may be configured to generate a second set of displacement coefficients for the level specified by the execution control information using the displacement information.
[0162] Graph 161 in Figure 16 shows the distribution of displacement coefficient values. In this graph 161, dotted line A to dotted line B represent the displacement coefficient values for LoD0, dotted line B to dotted line C represent the displacement coefficient values for LoD1, dotted line C to dotted line D represent the displacement coefficient values for LoD2, and dotted line D to dotted line E represent the displacement coefficient values for LoD3. As shown in this graph, the displacement coefficient values can vary significantly for each LoD. In general, the displacement coefficient values for LoD1 tend to be significantly exceeded.
[0163] Therefore, for example, displacement information may be applied only to displacement coefficients of LoD1, and displacement coefficients with a higher bit depth than the displacement map may be transmitted. For example, execution control for each LoD may be performed by transmitting LoD designation information that designates the LoD to which the displacement information is to be applied as execution control information. In this way, for example, displacement information may be applied only to some LoDs. Since the application of displacement information can be controlled for each LoD in this way, it is possible to suppress an increase in the bit depth of the displacement map while suppressing a decrease in the accuracy of the displacement vector. Furthermore, since the decoder only needs to control the execution of processing for the displacement information based on the control information, it is possible to suppress an increase in the load of the decoding process.
[0164] <Method 4-1-2> When Method 4-1 is applied, channel designation information may be transmitted as shown in the 20th row from the top of the table in FIG. 5 (Method 4-1-2).
[0165] For example, in the first information processing device described above, the execution control information may include information specifying a channel of a component. Also, in the second information processing device described above, the execution control information may include information specifying a channel of a component. The displacement coefficient generation unit may be configured to generate a second displacement coefficient group for the channel specified by the execution control information using the displacement information.
[0166] In other words, by transmitting such information, it is possible to apply displacement information only to some channels of components, such as luminance (Y) and chrominance (Cb, Cr). In other words, since it is possible to control the application of displacement information for each component channel, it is possible to suppress an increase in the bit depth of the displacement map while suppressing a decrease in the accuracy of the displacement vector. Furthermore, since the decoder only needs to control the execution of processing on the displacement information based on the control information, it is possible to suppress an increase in the load of the decoding process.
[0167] <Method 4-1-3> When Method 4-1 is applied, method designation information may be transmitted as shown in the bottom row of the table in FIG. 5 (Method 4-1-3).
[0168] For example, in the first information processing device described above, the execution control information may include information specifying a method for generating the displacement information. Also, in the second information processing device described above, the execution control information may include information specifying a method for generating the displacement information. The displacement coefficient generation unit may be configured to apply the generation method specified by the execution control information and generate a second displacement coefficient group using the displacement information.
[0169] For example, it may be possible to specify a method for generating displacement information, such as whether to apply method 1, method 2, or method 3. It may also be possible to specify detailed conditions, such as what parameters to use. By transmitting the method specification information in this way, it is possible to control the method for generating displacement information, and therefore it is possible to generate displacement information in a manner that is more suited to the situation. Therefore, it is possible to suppress an increase in the bit depth of the displacement map while suppressing a decrease in the accuracy of the displacement vector. Furthermore, since the decoder only needs to perform processing in accordance with the control information, it is possible to easily and correctly perform processing, and it is possible to suppress an increase in the load of the decoding process.
[0170] <Combination> Each of the above-described methods may be applied in combination with any other method as long as no contradiction occurs. Three or more methods may be applied in combination. For example, any two or more of Method 1, Method 2, Method 3, and Method 4 may be applied in combination. Furthermore, the methods that can be combined are not limited to those shown in the table of FIG. 5 as "Method," but may include all of the elements described above in <3. Transmission of Displacement Information>. Furthermore, each of the above-described methods may be applied in combination with methods other than those described above.
[0171] In this specification, a description of a higher-level method also applies to lower-level methods that belong to that higher-level method, unless a contradiction arises. For example, when it is described that "Method 1 may be applied," it is possible to apply Methods 1-1, 1-2, and 1-3. When it is described that "Method 2 may be applied," it is possible to apply Method 2-1, and it is also possible to apply Methods 2-1-1, 2-1-2, 2-1-3, and 2-1-4.
[0172] 4. First embodiment Encoding device The present technology can be applied to an encoding device that encodes meshes (V-DMC data). Fig. 17 is a block diagram showing an example of the configuration of an encoding device that is one aspect of an information processing device to which the present technology is applied. The encoding device 300 shown in Fig. 17 is a device that encodes meshes (V-DMC data).
[0173] Fig. 17 shows the main processing units, data flows, etc., but is not limited to all that is shown in Fig. 17. In other words, in encoding device 300, there may be processing units that are not shown as blocks in Fig. 17, and there may be processing and data flows that are not shown as arrows, etc. in Fig. 17.
[0174] The encoding device 300 encodes meshes using a method basically similar to the V-DMC described in the above-mentioned non-patent document, except that the present technology is applied.
[0175] 17 , the encoding device 300 includes a displacement vector processing unit 311, a displacement information generation unit 312, a packing unit 313, a mesh reconstruction unit 314, an attribute map conversion unit 315, an encoding unit 316, and a multiplexing unit 317. The encoding unit 316 includes an atlas information encoding unit 331, a base mesh encoding unit 332, an excess information encoding unit 333, an offset encoding unit 334, a control information encoding unit 335, a displacement video encoding unit 336, and an attribute video encoding unit 337.
[0176] The displacement vector processing unit 311 performs processing related to displacement vectors. For example, the displacement vector processing unit 311 may acquire a base mesh and a displacement vector supplied to the encoding device 300. The displacement vector processing unit 311 may acquire encoded data of the base mesh generated by the base mesh encoding unit 332. The displacement vector processing unit 311 may decode the encoded data of the base mesh, compare the base mesh before and after encoding to determine the encoding distortion, and correct the displacement vector according to the encoding distortion. The displacement vector processing unit 311 may generate a displacement coefficient from the corrected displacement vector. For example, the displacement vector processing unit 311 may derive the displacement coefficient by performing a wavelet transform on the displacement vector and quantizing it. The displacement vector processing unit 311 may supply the generated displacement coefficient to the displacement information generating unit 312. The displacement vector processing unit 311 may supply the decoded base mesh to the mesh reconstruction unit 314.
[0177] The displacement information generation unit 312 executes processing related to the generation of displacement information. For example, the displacement information generation unit 312 may acquire a displacement coefficient supplied from the displacement vector processing unit 311. The displacement information generation unit 312 may generate displacement information using the displacement coefficient. In this case, the displacement information generation unit 312 may apply the present technology described above in <3. Transmission of Displacement Information>. For example, the displacement information generation unit 312 may generate displacement information by applying any of the methods described above in <3. Transmission of Displacement Information>.
[0178] For example, the displacement information generation unit 312 may apply Method 1 to generate a group of displacement coefficients after division as displacement information. In this case, the displacement information generation unit 312 may supply the generated group of displacement coefficients after division to the packing unit 313.
[0179] Alternatively, the displacement information generation unit 312 may apply Method 2 to generate a displacement coefficient set and excess information as displacement information. In this case, the displacement information generation unit 312 may supply the generated displacement coefficient set to the packing unit 313. Alternatively, the displacement information generation unit 312 may supply the generated excess information to the excess information encoding unit 333.
[0180] Alternatively, the displacement information generation unit 312 may apply Method 3 to generate a set of displacement coefficients and an offset as displacement information. In this case, the displacement information generation unit 312 may supply the generated set of displacement coefficients to the packing unit 313. Alternatively, the displacement information generation unit 312 may supply the generated offset to the offset encoding unit 334.
[0181] Alternatively, the displacement information generation unit 312 may apply Method 4 to generate a displacement coefficient set and control information as displacement information. In this case, the displacement information generation unit 312 may supply the generated displacement coefficient set to the packing unit 313. Alternatively, the displacement information generation unit 312 may supply the generated control information to the control information encoding unit 335.
[0182] The packing unit 313 executes processing related to packing. For example, the packing unit 313 may acquire a set of displacement coefficients supplied from the displacement information generation unit 312. The packing unit 313 may perform packing and generate a displacement map in which the acquired set of displacement coefficients is stored. The packing unit 313 may supply the displacement map to the displacement video encoding unit 336. The packing unit 313 may supply the acquired set of displacement coefficients to the mesh reconstruction unit 314.
[0183] The mesh reconstruction unit 314 performs processing related to mesh reconstruction. For example, the mesh reconstruction unit 314 may acquire a base mesh supplied from the displacement vector processing unit 311. The mesh reconstruction unit 314 may also acquire a set of displacement coefficients supplied from the packing unit 313. The mesh reconstruction unit 314 may derive a displacement vector by inverse quantizing the set of displacement coefficients and performing an inverse wavelet transform on it. The mesh reconstruction unit 314 may reconstruct a mesh using the acquired base mesh or the derived displacement vector. The mesh reconstruction unit 314 may supply the reconstructed mesh to the attribute map conversion unit 315.
[0184] The attribute map conversion unit 315 performs processing related to attribute map conversion. For example, the attribute map conversion unit 315 may acquire a reconstructed mesh supplied from the mesh reconstruction unit 314. Alternatively, the attribute map conversion unit 315 may acquire an original mesh and an attribute map input to the encoding device 300. The attribute map conversion unit 315 may update the acquired attribute map based on other acquired information. The attribute map conversion unit 315 may supply the updated attribute map to the attribute video encoding unit 337.
[0185] The encoding unit 316 performs processing related to encoding of various data. In this case, the encoding unit 316 may apply the present technology described above in <3. Transmission of Displacement Information>. For example, the encoding unit 316 may apply any of the methods described above in <3. Transmission of Displacement Information> to encode various data including displacement information.
[0186] The atlas information encoding unit 331 performs processing related to encoding of atlas information. For example, the atlas information encoding unit 331 may acquire atlas information input to the encoding device 300. The atlas information encoding unit 331 may also encode the acquired atlas information using a predetermined encoding method to generate encoded data of the atlas information. The atlas information encoding unit 331 may supply the generated encoded data of the atlas information to the multiplexing unit 317.
[0187] The base mesh encoding unit 332 performs processing related to encoding of the base mesh. For example, the base mesh encoding unit 332 may acquire a base mesh input to the encoding device 300. The base mesh encoding unit 332 may also quantize the acquired base mesh and encode it using a predetermined encoding method (e.g., Draco) to generate encoded data of the base mesh. The base mesh encoding unit 332 may supply the generated encoded data of the base mesh to the multiplexing unit 317.
[0188] The excess information encoding unit 333 performs processing related to encoding of excess information. For example, the excess information encoding unit 333 may acquire excess information supplied from the displacement information generating unit 312. The excess information encoding unit 333 may also encode the acquired excess information using a predetermined encoding method to generate encoded data of the excess information. The excess information encoding unit 333 may also supply the generated encoded data of the excess information to the multiplexing unit 317. Note that if excess information is not supplied or if the excess information is not encoded, the excess information encoding unit 333 may be omitted.
[0189] The offset coding unit 334 performs processing related to coding of the offset. For example, the offset coding unit 334 may acquire an offset supplied from the displacement information generation unit 312. Furthermore, the offset coding unit 334 may encode the acquired offset using a predetermined encoding method to generate coded data of the offset. The offset coding unit 334 may supply the generated coded data of the offset to the multiplexing unit 317. Note that if no offset is supplied or if the offset is not coded, the offset coding unit 334 may be omitted.
[0190] The control information encoder 335 performs processing related to encoding of control information. For example, the control information encoder 335 may acquire control information supplied from the displacement information generator 312. The control information encoder 335 may also encode the acquired control information using a predetermined encoding method to generate encoded data of the control information. The control information encoder 335 may also supply the generated encoded data of the control information to the multiplexer 317. Note that if no control information is supplied or if the control information is not encoded, the control information encoder 335 may be omitted.
[0191] The displacement video encoding unit 336 performs processing related to encoding of the displacement video. The displacement video is a moving image in which a displacement map, which is a two-dimensional area in which displacement vectors are packed, is used as a frame image. For example, the displacement video encoding unit 336 may acquire a displacement map supplied from the packing unit 313. The displacement video encoding unit 336 may generate a displacement video in which the displacement map is used as a frame image. Furthermore, the displacement video encoding unit 336 may encode the generated displacement video using a predetermined encoding method for 2D moving images to generate encoded data of the displacement video. The displacement video encoding unit 336 may supply the generated encoded data of the displacement video to the multiplexing unit 317.
[0192] The attribute video encoding unit 337 performs processing related to encoding of the attribute video. For example, the attribute video encoding unit 337 may acquire an attribute map supplied from the attribute map conversion unit 315. The attribute video encoding unit 337 may also generate attribute video using the acquired attribute map as frame images. The attribute video encoding unit 337 may also encode the generated attribute video using a predetermined encoding method for 2D moving images to generate encoded data of the attribute video. The attribute video encoding unit 337 may also supply the generated encoded data of the attribute video to the multiplexing unit 317.
[0193] The multiplexing unit 317 performs processing related to multiplexing of encoded data (substreams). For example, the multiplexing unit 317 may acquire encoded data of atlas information supplied from the atlas information encoding unit 331. The multiplexing unit 317 may acquire encoded data of base meshes supplied from the base mesh encoding unit 332. The multiplexing unit 317 may acquire encoded data of excess information supplied from the excess information encoding unit 333. The multiplexing unit 317 may acquire encoded data of offsets supplied from the offset encoding unit 334. The multiplexing unit 317 may acquire encoded data of control information supplied from the control information encoding unit 335. The multiplexing unit 317 may acquire encoded data of displacement video supplied from the displacement video encoding unit 336. The multiplexing unit 317 may also acquire encoded data of attribute video supplied from the attribute video encoding unit 337. The multiplexing unit 317 may multiplex these encoded data as substreams to generate a V-DMC bitstream. Therefore, multiplexing unit 317 can also be considered a bitstream generation unit. Multiplexing unit 317 may output the generated V-DMC bitstream to the outside of encoding device 300. For example, multiplexing unit 317 may supply the V-DMC bitstream to decoding device 400, which will be described later. Therefore, multiplexing unit 317 can also be considered a supplier (provider) of the V-DMC bitstream.
[0194] In the encoding device 300 configured as above, the various methods (the present technology) described above in <3. Transmission of displacement information> may be applied.
[0195] In this way, the encoding device 300 can achieve the same effect as that described above in <3. Transmission of Displacement Information>. That is, the encoding device 300 can suppress an increase in the bit depth of the displacement map while suppressing a decrease in the accuracy of the displacement vector.
[0196] <Flow of Encoding Process> An example of the flow of the encoding process executed by the encoding device 300 will be described with reference to the flowchart in Fig. 18. Note that the flowchart in Fig. 18 shows an example of the flow of the encoding process when method 1 of the present technology is applied.
[0197] When the encoding process starts, the atlas information encoding unit 331 of the encoding device 300 encodes the atlas information in step S301. The base mesh encoding unit 332 encodes the base mesh in step S302. The displacement vector processing unit 311 corrects the displacement vector and derives a displacement coefficient in step S303.
[0198] In step S304, the disparity information generating unit 312 divides each displacement coefficient of the displacement coefficient group including a displacement coefficient with a high bit depth into a plurality of displacement coefficients with a low bit depth, and sets the resulting displacement coefficients as displacement information.
[0199] In step S305, the packing unit 313 stores the divided displacement coefficients in a displacement map. In step S306, the displacement video encoding unit 336 encodes a displacement video using the displacement map as a frame image. In step S307, the mesh reconstruction unit 314 reconstructs a mesh. In step S308, the attribute map conversion unit 315 updates the attribute map. In step S309, the attribute video encoding unit 337 encodes an attribute video using the attribute map as a frame image.
[0200] In step S310, the multiplexing unit 317 multiplexes the coded data generated as described above to generate a V-DMC bitstream.
[0201] When the process of step S310 is completed, the encoding process ends.
[0202] By performing each process as described above, the encoding device 300 can apply Method 1. Therefore, the encoding device 300 can suppress an increase in the bit depth of the displacement map while suppressing a decrease in the accuracy of the displacement vector.
[0203] <Flow of Encoding Process> Another example of the flow of the encoding process executed by the encoding device 300 will be described with reference to the flowchart of Fig. 19. Note that the flowchart of Fig. 19 illustrates an example of the flow of the encoding process when Method 2 of the present technology is applied.
[0204] When the encoding process is started, the processes from step S321 to step S323 are executed in the same manner as the processes from step S301 to step S303 (FIG. 18).
[0205] In step S324, the displacement information generation unit 312 generates excess information indicating displacement coefficients with a high bit depth. In step S325, the excess information encoding unit 333 encodes the excess information. In step S326, the displacement information generation unit 312 generates a group of displacement coefficients with a low bit depth.
[0206] The processes from step S327 to step S332 are executed in the same manner as the processes from step S305 to step S310 (FIG. 18).
[0207] When the process of step S332 ends, the encoding process ends.
[0208] By performing each process as described above, the encoding device 300 can apply Method 2. Therefore, the encoding device 300 can suppress an increase in the bit depth of the displacement map while suppressing a decrease in the accuracy of the displacement vector.
[0209] <Flow of Encoding Process> Another example of the flow of the encoding process executed by the encoding device 300 will be described with reference to the flowchart of Fig. 20. Note that the flowchart of Fig. 20 illustrates an example of the flow of the encoding process when Method 3 of the present technology is applied.
[0210] When the encoding process is started, the processes from step S341 to step S343 are executed in the same manner as the processes from step S301 to step S303 (FIG. 18).
[0211] In step S344, the displacement information generation unit 312 sets an offset. In step S345, the displacement information generation unit 312 shifts each displacement coefficient of the displacement coefficient set including the high bit depth displacement coefficient by the offset. In step S346, the offset encoding unit 334 encodes the offset.
[0212] The processes from step S347 to step S352 are executed in the same manner as the processes from step S305 to step S310 (FIG. 18).
[0213] When the process of step S352 is completed, the encoding process ends.
[0214] By performing each process as described above, the encoding device 300 can apply Method 3. Therefore, the encoding device 300 can suppress an increase in the bit depth of the displacement map while suppressing a decrease in the accuracy of the displacement vector.
[0215] <Flow of Encoding Process> Another example of the flow of the encoding process executed by the encoding device 300 will be described with reference to the flowchart of Fig. 21. Note that the flowchart of Fig. 21 illustrates an example of the flow of the encoding process when Method 4 of the present technology is applied.
[0216] When the encoding process is started, the processes from step S361 to step S363 are executed in the same manner as the processes from step S301 to step S303 (FIG. 18).
[0217] In step S364, the displacement information generation unit 312 generates control information. In step S365, the control information encoding unit 335 encodes the control information. In step S366, the displacement information generation unit 312 generates displacement information based on the control information. In step S367, the encoding unit 316 (e.g., the displacement video encoding unit 336) encodes the displacement information.
[0218] The processes from step S368 to step S373 are executed in the same manner as the processes from step S305 to step S310 (FIG. 18).
[0219] When the process of step S373 is completed, the encoding process ends.
[0220] By performing each process as described above, the encoding device 300 can apply Method 4. Therefore, the encoding device 300 can suppress an increase in the bit depth of the displacement map while suppressing a decrease in the accuracy of the displacement vector.
[0221] 22 is a block diagram showing an example of the configuration of a decoding device, which is one aspect of an information processing device to which the present technology is applied. The decoding device 400 shown in Fig. 22 is a device that decodes, for example, encoded data of meshes generated in the encoding device 300 (V-DMC bitstream generated by the multiplexing unit 317).
[0222] Fig. 22 shows the main processing units, data flows, etc., but does not necessarily include all of them. In other words, in the decoding device 400, there may be processing units that are not shown as blocks in Fig. 22, and there may be processing and data flows that are not shown as arrows, etc. in Fig. 22.
[0223] The decoding device 400 decodes coded data of meshes that have been coded using a method basically similar to the V-DMC described in the above-mentioned non-patent document, except that the present technology is applied.
[0224] 22 , the decoding device 400 includes a demultiplexing unit 411, a decoding unit 412, a base mesh reconstruction unit 413, a subdivision unit 414, an unpacking unit 415, a displacement coefficient generation unit 416, a displacement vector generation unit 417, a displacement vector application unit 418, and a display processing unit 419. The decoding unit 412 includes an atlas information decoding unit 431, a base mesh decoding unit 432, a displacement video decoding unit 433, an excess information decoding unit 434, an offset decoding unit 435, a control information decoding unit 436, and an attribute video decoding unit 437.
[0225] The demultiplexing unit 411 performs demultiplexing processing. For example, the demultiplexing unit 411 may acquire a V-DMC bitstream to be decoded and supplied to the decoding device 400. The demultiplexing unit 411 may demultiplex the acquired V-DMC bitstream to extract coded data of atlas information, coded data of base meshes, coded data of displacement video, coded data of excess information, coded data of offsets, coded data of control information, and coded data of attribute video. Thus, the demultiplexing unit 411 can also be considered an acquisition unit for the V-DMC bitstream or various information contained in the V-DMC bitstream. The demultiplexing unit 411 may supply the coded data of the extracted atlas information to the atlas information decoding unit 431. The demultiplexing unit 411 may also supply the coded data of the extracted base meshes to the base mesh decoding unit 432. The demultiplexing unit 411 may also supply the coded data of the extracted displacement video to the displacement video decoding unit 433. The demultiplexing unit 411 may also supply the extracted coded data of excess information to the excess information decoding unit 434. The demultiplexing unit 411 may also supply the extracted coded data of offset to the offset decoding unit 435. The demultiplexing unit 411 may also supply the extracted coded data of control information to the control information decoding unit 436. The demultiplexing unit 411 may also supply the extracted coded data of attribute video to the attribute video decoding unit 437.
[0226] The decoding unit 412 executes processing related to decoding of various coded data. In this case, the decoding unit 412 may apply the present technology described above in <3. Transmission of Displacement Information>. For example, the decoding unit 412 may decode various coded data including displacement information by applying any of the methods described above in <3. Transmission of Displacement Information>. Note that if Method 2 is not applied, the excess information decoding unit 434 may be omitted. Also, if Method 3 is not applied, the offset decoding unit 435 may be omitted. Also, if Method 4 is not applied, the control information decoding unit 436 may be omitted.
[0227] The atlas information decoding unit 431 performs processing related to decoding of encoded data of atlas information. For example, the atlas information decoding unit 431 may acquire encoded data of atlas information supplied from the demultiplexing unit 411. The atlas information decoding unit 431 may also decode the acquired encoded data of atlas information using a predetermined decoding method to generate (restore) atlas information. The atlas information decoding unit 431 may supply the generated atlas information to the base mesh reconstruction unit 413 and the display processing unit 419.
[0228] The base mesh decoding unit 432 performs processing related to decoding of the coded data of the base mesh. For example, the base mesh decoding unit 432 may acquire the coded data of the base mesh supplied from the demultiplexing unit 411. The base mesh decoding unit 432 may also decode the acquired coded data of the base mesh using a predetermined decoding method (e.g., Draco) to generate (restore) information about the base mesh (e.g., a vertex list, a triangle list, etc.). The base mesh decoding unit 432 may supply the generated information about the base mesh to the base mesh reconstruction unit 413.
[0229] The displacement video decoding unit 433 performs processing related to decoding of encoded data of the displacement video. For example, the displacement video decoding unit 433 may acquire encoded data of the displacement video supplied from the demultiplexing unit 411. Furthermore, the displacement video decoding unit 433 may decode the acquired encoded data of the displacement video using a predetermined decoding method for 2D moving images to generate (restore) the displacement video. The displacement video decoding unit 433 may supply a disparity map, which is frame images of the generated displacement video, to the unpacking unit 415.
[0230] The excess information decoding unit 434 performs processing related to decoding of the coded data of excess information. For example, the excess information decoding unit 434 may acquire the coded data of excess information supplied from the demultiplexing unit 411. The excess information decoding unit 434 may also decode the acquired coded data of excess information using a predetermined decoding method to generate (restore) the excess information. The excess information decoding unit 434 may also supply the generated excess information to the displacement coefficient generating unit 416. Note that if coded data of excess information is not supplied, the excess information decoding unit 434 may be omitted.
[0231] The offset decoding unit 435 performs processing related to decoding of the coded data of the offset. For example, the offset decoding unit 435 may acquire the coded data of the offset supplied from the demultiplexing unit 411. The offset decoding unit 435 may also decode the acquired coded data of the offset using a predetermined decoding method to generate (restore) the offset. The offset decoding unit 435 may also supply the generated offset to the displacement coefficient generation unit 416. Note that if coded data of the offset is not supplied, the offset decoding unit 435 may be omitted.
[0232] The control information decoder 436 performs processing related to decoding of coded data of control information. For example, the control information decoder 436 may acquire coded data of control information supplied from the demultiplexer 411. The control information decoder 436 may also decode the acquired coded data of control information using a predetermined decoding method to generate (restore) control information. The control information decoder 436 may also supply the generated control information to the displacement coefficient generator 416. Note that if coded data of control information is not supplied, the control information decoder 436 may be omitted.
[0233] The attribute video decoding unit 437 performs processing related to decoding of the attribute video. For example, the attribute video decoding unit 437 may acquire coded data of the attribute video supplied from the demultiplexing unit 411. The attribute video decoding unit 437 may also decode the acquired coded data of the attribute video using a predetermined decoding method for 2D moving images to generate (restore) the attribute video. The attribute video decoding unit 437 may supply an attribute map, which is a frame image of the generated attribute video, to the display processing unit 419.
[0234] The base mesh reconstructing unit 413 executes processing related to the reconstruction of the base mesh. For example, the base mesh reconstructing unit 413 may acquire information about the base mesh supplied from the base mesh decoding unit 432. The base mesh reconstructing unit 413 may also acquire atlas information supplied from the atlas information decoding unit 431. The base mesh reconstructing unit 413 may reconstruct the base mesh using the acquired atlas information and information about the base mesh. The base mesh reconstructing unit 413 may also supply the generated base mesh to the subdivision unit 414.
[0235] The subdivision unit 414 performs processing related to subdivision of triangles of a base mesh. For example, the subdivision unit 414 may obtain a base mesh provided from the base mesh reconstruction unit 413. The subdivision unit 414 may subdivide triangles of the base mesh to generate division points. The subdivision unit 414 may provide the subdivided base mesh to the displacement vector application unit 418.
[0236] The unpacking unit 415 performs processing related to unpacking. For example, the unpacking unit 415 may obtain a displacement map supplied from the displacement video decoding unit 433. The unpacking unit 415 may also perform unpacking and extract displacement coefficients (displacement information) from the displacement map. The unpacking unit 415 may supply the displacement coefficients (displacement information) obtained in this manner to the displacement coefficient generation unit 416.
[0237] The displacement coefficient generation unit 416 executes processing related to the generation of displacement coefficients. In this case, the displacement coefficient generation unit 416 may apply the present technology described above in <3. Transmission of Displacement Information>. For example, the displacement coefficient generation unit 416 may apply any of the methods described above in <3. Transmission of Displacement Information> to generate displacement coefficients from displacement information.
[0238] For example, when applying method 1, the displacement coefficient generation unit 416 may obtain the set of displacement coefficients after division supplied from the unpacking unit 415. The displacement coefficient generation unit 416 may combine the set of displacement coefficients after division to restore (generate) the set of displacement coefficients before division (a set of displacement coefficients including displacement coefficients with a bit depth higher than the bit depth of the displacement map). The displacement coefficient generation unit 416 may supply the generated set of displacement coefficients to the displacement vector generation unit 417.
[0239] For example, when applying Method 2, the displacement coefficient generation unit 416 may obtain a set of displacement coefficients supplied from the unpacking unit 415. The displacement coefficient generation unit 416 may also obtain excess information supplied from the excess information decoding unit 434. The displacement coefficient generation unit 416 may use the displacement coefficients and excess information to restore (generate) a set of displacement coefficients including displacement coefficients with a bit depth higher than the bit depth of the displacement map. The displacement coefficient generation unit 416 may also supply the generated set of displacement coefficients to the displacement vector generation unit 417.
[0240] For example, when applying Method 3, the displacement coefficient generation unit 416 may obtain a set of displacement coefficients supplied from the unpacking unit 415. Furthermore, the displacement coefficient generation unit 416 may obtain an offset supplied from the offset decoding unit 435. The displacement coefficient generation unit 416 may de-shift the displacement coefficients by the offset amount. The displacement coefficient generation unit 416 may supply the de-shifted set of displacement coefficients to the displacement vector generation unit 417.
[0241] For example, when applying method 4, the displacement coefficient generation unit 416 may acquire a set of displacement coefficients (displacement information) supplied from the unpacking unit 415. Furthermore, the displacement coefficient generation unit 416 may acquire control information supplied from the control information decoding unit 436. Based on the control information, the displacement coefficient generation unit 416 may restore (generate) a set of displacement coefficients from the displacement information, including displacement coefficients with a bit depth higher than the bit depth of the displacement map. The displacement coefficient generation unit 416 may supply the generated set of displacement coefficients to the displacement vector generation unit 417.
[0242] The displacement vector generation unit 417 executes processing related to the generation of displacement vectors. For example, the displacement vector generation unit 417 acquires a group of displacement coefficients supplied from the displacement coefficient generation unit 416. The displacement vector generation unit 417 generates a group of displacement vectors using the group of displacement coefficients. For example, the displacement vector generation unit 417 generates displacement vectors by inverse quantizing the acquired displacement coefficients and performing an inverse wavelet transform. The displacement vector generation unit 417 supplies the generated group of displacement vectors to the displacement vector application unit 418.
[0243] The displacement vector application unit 418 performs processing related to application of displacement vectors to the subdivided base mesh. For example, the displacement vector application unit 418 may acquire the subdivided base mesh supplied from the subdivision unit 414. Alternatively, the displacement vector application unit 418 may acquire a group of displacement vectors supplied from the displacement vector generation unit 417. Alternatively, the displacement vector application unit 418 may apply displacement vectors to division points of the subdivided base mesh. In other words, the displacement vector application unit 418 may generate a reconstructed mesh. The displacement vector application unit 418 may supply the generated reconstructed mesh to the display processing unit 419.
[0244] The display processing unit 419 performs processing related to mesh display. For example, the display processing unit 419 may acquire atlas information supplied from the atlas information decoding unit 431. The display processing unit 419 may also acquire a reconstructed mesh supplied from the displacement vector application unit 418. The display processing unit 419 may also acquire an attribute map supplied from the attribute video decoding unit 437. The display processing unit 419 may use the acquired atlas information to extract a texture from the attribute map and apply the texture to a face of the decoded mesh corresponding to the extracted texture. In other words, the display processing unit 419 may attach the texture to the face. The display processing unit 419 may render the reconstructed mesh to which the texture has been applied and generate a display image for displaying the reconstructed mesh to which the texture has been applied. The display processing unit 419 may then supply the generated display image to an external device outside the decoding device 400, causing the display image to be displayed by another device or the like.
[0245] In the decoding device 400 configured as above, the various methods (the present technology) described above in <3. Transmission of displacement information> may be applied.
[0246] By doing so, the decoding device 400 can obtain the same effect as that described above in <3. Transmission of Displacement Information>. That is, the decoding device 400 can suppress an increase in the bit depth of the displacement map while suppressing a decrease in the accuracy of the displacement vector.
[0247] <Flow of Decoding Process> An example of the flow of the decoding process executed by the decoding device 400 will be described with reference to the flowchart in Fig. 23. Note that the flowchart in Fig. 23 shows an example of the flow of the decoding process when method 1 of the present technology is applied.
[0248] When the decoding process starts, the demultiplexing unit 411 of the decoding device 400 demultiplexes the V-DMC bitstream in step S401. In step S402, the atlas information decoding unit 431 decodes the coded data of the atlas information. In step S403, the base mesh decoding unit 432 decodes the coded data of the base mesh. In step S404, the base mesh reconstruction unit 413 reconstructs the base mesh using the atlas information. In step S405, the subdivision unit 414 subdivides the reconstructed base mesh. In step S406, the displacement video decoding unit 433 decodes the coded data of the displacement video. In step S407, the unpacking unit 415 extracts displacement coefficients from the displacement map.
[0249] In step S408, the displacement coefficient generator 416 combines a plurality of low-bit-depth displacement coefficients, and in step S409, the displacement vector generator 417 generates a displacement vector using the combined displacement coefficients.
[0250] In step S410, the displacement vector application unit 418 applies displacement vectors to the division points of the subdivided base mesh. In step S411, the attribute video decoding unit 437 decodes the encoded data of the attribute video. In step S412, the display processing unit 419 applies attributes to the decoded mesh using the atlas information to generate a display image.
[0251] When the process of step S412 is completed, the decoding process ends.
[0252] By performing each process as described above, the decoding device 400 can apply Method 1. Therefore, the decoding device 400 can suppress an increase in the bit depth of the disparity map while suppressing a decrease in the accuracy of the disparity vector.
[0253] <Flow of Decoding Process> Another example of the flow of the encoding process executed by the decoding device 400 will be described with reference to the flowchart of Fig. 24. Note that the flowchart of Fig. 24 illustrates an example of the flow of the decoding process when Method 2 of the present technology is applied.
[0254] When the decoding process is started, the processes from step S421 to step S427 are executed in the same manner as the processes from step S401 to step S407 (FIG. 23).
[0255] In step S428, the excess information decoder 434 decodes the excess information. In step S429, the displacement coefficient generator 416 generates a displacement coefficient set that may include a high-bit-depth displacement coefficient using the low-bit-depth displacement coefficient set and the excess information. In step S430, the displacement vector generator 417 generates a displacement vector using the generated displacement coefficients.
[0256] The processes from step S431 to step S433 are executed in the same manner as the processes from step S410 to step S412 (FIG. 23).
[0257] When the process of step S433 is completed, the decoding process ends.
[0258] By performing each process as described above, the decoding device 400 can apply Method 2. Therefore, the decoding device 400 can suppress an increase in the bit depth of the disparity map while suppressing a decrease in the accuracy of the disparity vector.
[0259] <Decoding Process Flow> Another example of the flow of the encoding process executed by the decoding device 400 will be described with reference to the flowchart of Fig. 25. Note that the flowchart of Fig. 25 illustrates an example of the flow of the decoding process when Method 3 of the present technology is applied.
[0260] When the decoding process is started, the processes from step S441 to step S447 are executed in the same manner as the processes from step S401 to step S407 (FIG. 23).
[0261] In step S448, the offset decoding unit 435 decodes the offset. In step S449, the displacement coefficient generation unit 416 de-shifts each displacement coefficient by the offset. In step S450, the displacement vector generation unit 417 generates a displacement vector using the de-shifted displacement coefficients.
[0262] The processes from step S451 to step S453 are executed in the same manner as the processes from step S410 to step S412 (FIG. 23).
[0263] When the process of step S453 is completed, the decoding process ends.
[0264] By performing each process as described above, the decoding device 400 can apply Method 3. Therefore, the decoding device 400 can suppress an increase in the bit depth of the disparity map while suppressing a decrease in the accuracy of the disparity vector.
[0265] <Flow of Decoding Process> Another example of the flow of the encoding process executed by the decoding device 400 will be described with reference to the flowchart of Fig. 26. Note that the flowchart of Fig. 26 illustrates an example of the flow of the decoding process when Method 4 of the present technology is applied.
[0266] When the decoding process is started, the processes from step S461 to step S467 are executed in the same manner as the processes from step S401 to step S407 (FIG. 23).
[0267] In step S468, the control information decoder 436 decodes the control information. In step S469, the decoder 412 (e.g., the displacement video decoder 433) decodes the displacement information according to the control information. In step S470, the displacement coefficient generator 416 generates a displacement coefficient set that may include a high-bit-depth displacement coefficient using the low-bit-depth displacement coefficient set and the displacement information according to the control information. In step S471, the displacement vector generator 417 generates a displacement vector using the displacement coefficients.
[0268] The processes from step S472 to step S474 are executed in the same manner as the processes from step S410 to step S412 (FIG. 23).
[0269] When the process of step S474 ends, the decoding process ends.
[0270] By performing each process as described above, the decoding device 400 can apply Method 4. Therefore, the decoding device 400 can suppress an increase in the bit depth of the displacement map while suppressing a decrease in the accuracy of the displacement vector.
[0271] <6. Supplementary Notes> <Polygon Shape> In the above, the polygon shape has been described as being triangular, but this shape is just one example. The polygon shape may be any polygonal shape. <Standards> In the above, V-DMC has been used as an example of an encoding method to which the present technology is applied, but the present technology is not limited to this example and can be applied to any encoding method that encodes a base mesh, a displacement vector, an attribute map including texture, and atlas information, or information equivalent thereto.
[0272] <Computer> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs that make up the software are installed on a computer. Here, the term "computer" includes computers built into dedicated hardware, and general-purpose personal computers, etc., that can execute various functions by installing various programs.
[0273] FIG. 27 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.
[0274] In a computer 900 shown in FIG. 27, a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, and a RAM (Random Access Memory) 903 are interconnected via a bus 904.
[0275] An input / output interface 910 is also connected to the bus 904. To the input / output interface 910, an input unit 911, an output unit 912, a storage unit 913, a communication unit 914, and a drive 915 are connected.
[0276] The input unit 911 includes, for example, a keyboard, a mouse, a microphone, a touch panel, and an input terminal. The output unit 912 includes, for example, a display, a speaker, and an output terminal. The storage unit 913 includes, for example, a hard disk, a RAM disk, and a non-volatile memory. The communication unit 914 includes, for example, a network interface. The drive 915 drives removable media 921 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0277] In a computer configured as described above, the CPU 901 performs the above-described series of processes by, for example, loading a program stored in the storage unit 913 into the RAM 903 via the input / output interface 910 and the bus 904 and executing the program. The RAM 903 also stores data necessary for the CPU 901 to execute various processes as appropriate.
[0278] The program executed by the computer can be applied by recording it on, for example, a removable medium 921 such as a package medium. In this case, the program can be installed in the storage unit 913 via the input / output interface 910 by inserting the removable medium 921 into the drive 915.
[0279] This program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, digital satellite broadcasting, etc. In this case, the program can be received by the communication unit 914 and installed in the storage unit 913.
[0280] Alternatively, this program can be installed in advance in the ROM 902 or the storage unit 913 .
[0281] <Application of the Present Technology> The present technology can be applied to any configuration. For example, the present technology can be applied to various electronic devices.
[0282] Furthermore, for example, the present technology can also be implemented as part of an apparatus, such as a processor (e.g., a video processor) as a system LSI (Large Scale Integration), a module using multiple processors (e.g., a video module), a unit using multiple modules (e.g., a video unit), or a set in which other functions are added to a unit (e.g., a video set).
[0283] Furthermore, for example, the present technology can also be applied to a network system configured with multiple devices. For example, the present technology may be implemented as cloud computing in which multiple devices share and collaborate on processing via a network. For example, the present technology may be implemented in a cloud service that provides image (video)-related services to any terminal, such as a computer, an AV (Audio Visual) device, a portable information processing terminal, or an IoT (Internet of Things) device.
[0284] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0285] <Fields and uses to which this technology can be applied> Systems, devices, processing units, etc. to which this technology is applied can be used in any field, for example, transportation, medical care, crime prevention, agriculture, livestock farming, mining, beauty, factories, home appliances, weather, nature monitoring, etc. In addition, the uses thereof are also arbitrary.
[0286] <Others> In this specification, a "flag" refers to information for identifying multiple states, and includes not only information used to identify two states, true (1) or false (0), but also information capable of identifying three or more states. Therefore, the value that this "flag" can take may be, for example, two values, 1 / 0, or three or more values. That is, the number of bits constituting this "flag" is arbitrary, and may be one bit or multiple bits. Furthermore, identification information (including flags) can be included not only in a bitstream, but also in a bitstream that includes differential information of the identification information relative to certain reference information. Therefore, in this specification, "flag" and "identification information" encompass not only the information itself, but also differential information relative to the reference information.
[0287] Furthermore, various types of information (e.g., metadata) related to the coded data (bitstream) may be transmitted or recorded in any form as long as they are associated with the coded data. Here, the term "associate" means, for example, that one piece of data can be used (linked) when processing the other piece of data. That is, the associated pieces of data may be combined into one piece of data or may be separate pieces of data. For example, information associated with coded data (image) may be transmitted over a transmission path separate from that of the coded data (image). Furthermore, for example, information associated with coded data (image) may be recorded on a recording medium separate from that of the coded data (image) (or on a different recording area of the same recording medium). Note that this "association" may refer not to the entire data, but to only part of the data. For example, an image and information corresponding to that image may be associated with each other in any unit, such as multiple frames, one frame, or a portion of a frame.
[0288] In this specification, terms such as "composite," "multiplex," "add," "integrate," "include," "store," "embed," "insert," and the like refer to combining multiple items into one, such as combining encoded data and metadata into one piece of data, and refer to one method of "associating" as described above.
[0289] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.
[0290] For example, a configuration described as one device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, configurations described above as multiple devices (or processing units) may be combined and configured as one device (or processing unit). Of course, configurations other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, as long as the configuration and operation of the entire system are substantially the same, part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).
[0291] Furthermore, for example, the above-described program may be executed in any device, as long as the device has the necessary functions (functional blocks, etc.) and is able to obtain the necessary information.
[0292] Also, for example, each step of a single flowchart may be executed by a single device, or may be shared and executed by multiple devices. Furthermore, when a single step includes multiple processes, the multiple processes may be executed by a single device, or may be shared and executed by multiple devices. In other words, multiple processes included in a single step can be executed as multiple step processes. Conversely, processes described as multiple steps can be executed collectively as a single step.
[0293] For example, the steps of a program executed by a computer may be executed in chronological order in the order described herein, or may be executed in parallel or individually at the required timing, such as when a call is made. In other words, as long as no contradiction occurs, the steps may be executed in an order different from the order described above. Furthermore, the steps of this program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.
[0294] Furthermore, for example, multiple technologies related to the present technology can be implemented independently and independently, as long as no contradiction occurs. Of course, any multiple technologies can also be implemented in combination. For example, part or all of the present technology described in any embodiment can be implemented in combination with part or all of the present technology described in another embodiment. Furthermore, part or all of any of the above-described present technologies can be implemented in combination with other technologies not described above.
[0295] Note that the present technology can also be configured as follows: (1) An information processing device including: a displacement information generation unit configured to generate, for a base mesh, displacement information related to a second displacement coefficient set including second displacement coefficients of a second bit depth lower than the first bit depth, using a first displacement coefficient set including first displacement coefficients of a first bit depth; a packing unit configured to generate a displacement map of the second bit depth in which the second displacement coefficient set is stored as a pixel value set; and an encoding unit configured to encode a displacement video using the displacement map as a frame image, wherein the base mesh is a mesh with lower resolution than an original mesh to be encoded, the original mesh being composed of vertices and connections that represent a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh, the original mesh being composed of vertices and connections that represent the three-dimensional structure of the object; the displacement coefficients are coefficients corresponding to displacement vectors that indicate displacements of vertices or division points of the base mesh; the division points are vertices generated by subdividing the base mesh; the first displacement coefficient set is a coefficient set corresponding to a displacement vector set corresponding to the base mesh; and the displacement information is information related to the displacement indicated by the displacement vector set corresponding to the base mesh. (2) The information processing device according to (1), wherein the displacement information generation unit is configured to generate the second displacement coefficient group as the displacement information by dividing each of the first displacement coefficients into a plurality of the second displacement coefficients. (3) The information processing device according to (2), wherein the displacement information generation unit is configured to divide the first displacement coefficients into first divided coefficients and second divided coefficients as the second displacement coefficients, and arrange a second coefficient sequence representing a group of the second divided coefficients contiguous to a first coefficient sequence representing a group of the first divided coefficients, and the packing unit is configured to generate the displacement map in which a second pixel region representing the second coefficient sequence is arranged contiguous to a first pixel region representing the first coefficient sequence in accordance with the arrangement of the first coefficient sequence and the second coefficient sequence.(4) The information processing device according to any of (2) to (4), wherein the displacement information generation unit is configured to generate a set of the first and second division coefficients as the second displacement coefficients by dividing the first displacement coefficients into first and second division coefficients, and to arrange the set of first and second division coefficients consecutively according to an arrangement of the first displacement coefficients before division, and the packing unit is configured to generate the displacement map in which a group of pixels representing the arrangement of the set of first and second division coefficients is arranged according to the arrangement of the set of first and second division coefficients. (5) The information processing device according to any of (2) to (4), wherein the displacement information generation unit is configured to arrange the second displacement coefficients as a coefficient sequence in a predetermined order, and the packing unit is configured to generate the displacement map in which the second displacement coefficients are arranged in reverse scanning order from the bottom right pixel of the displacement map in the order of arrangement of the coefficient sequence. (6) The information processing device according to any one of (1) to (5), wherein the displacement information generation unit is configured to generate, as the displacement information, excess information indicating the displacement coefficients included in the first displacement coefficient set that exceed the second bit depth, and generate the second displacement coefficient set by replacing the displacement coefficients in the first displacement coefficient set that exceed the second bit depth with values of the second bit depth, and the encoding unit is configured to encode the excess information and encode the displacement video in which the second displacement coefficient set is stored as a pixel value set. (7) The information processing device according to (6), wherein the excess information is information indicating identification information of the displacement coefficients that exceed the second bit depth and an excess amount. (8) The information processing device according to (7), wherein the excess amount is information including a positive or negative sign. (9) The information processing device according to (7) or (8), wherein the excess information is information for each positive or negative sign of the excess amount. (10) The information processing device according to any one of (7) to (9), wherein the excess amount is information that does not include a positive or negative sign.(11) The information processing device according to any one of (7) to (10), wherein the excess information includes, as information indicating the identification information, a difference value between the identification information of the displacement coefficients that exceed the second bit depth. (12) The information processing device according to any one of (1) to (11), wherein the displacement information generation unit is configured to generate the second set of displacement coefficients as the displacement information by shifting the first set of displacement coefficients by an offset. (13) The information processing device according to (12), wherein the offset is a predetermined value. (14) The information processing device according to (12) or (13), wherein the displacement information generation unit is configured to set the offset and shift the first set of displacement coefficients by the set offset to generate the second set of displacement coefficients, and the encoding unit is configured to encode the offset and encode the displacement video in which the second set of displacement coefficients is stored as a set of pixel values. (15) The information processing device according to any of (12) to (14), wherein the offset has an independent value for each subdivision level of the base mesh, and the displacement information generation unit is configured to generate the second set of displacement coefficients by shifting the first displacement coefficient to be processed in the first displacement coefficient group by the offset, the value of which corresponds to the level. (16) The information processing device according to any of (12) to (15), wherein the offset has an independent value for each positive or negative sign of the first displacement coefficient, and the displacement information generation unit is configured to generate the second set of displacement coefficients by shifting the first displacement coefficient to be processed in the first displacement coefficient group by the offset, the value of which corresponds to the sign of the first displacement coefficient. (17) The information processing device according to any of (1) to (16), wherein the displacement information generation unit is configured to generate control information for controlling processing related to the displacement information, and the encoding unit is configured to encode the generated control information. (18) The information processing device according to (17), wherein the control information includes execution control information indicating whether the displacement information is to be generated.(19) The information processing device according to (18), wherein the execution control information includes information specifying a subdivision level of the base mesh. (20) The information processing device according to (18) or (19), wherein the execution control information includes information specifying a channel of a component. (21) The information processing device according to any one of (18) to (20), wherein the execution control information includes information specifying a method of generating the displacement information. (22) An information processing method, comprising: using a first group of displacement coefficients including first displacement coefficients of a first bit depth, to generate displacement information for a base mesh regarding a second group of displacement coefficients including second displacement coefficients of a second bit depth lower than the first bit depth; generating a displacement map of the second bit depth in which the second group of displacement coefficients is stored as a group of pixel values; and encoding a displacement video using the displacement map as a frame image; wherein the base mesh is a mesh with lower resolution than an original mesh to be encoded, the original mesh being composed of vertices and connections that represent a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh; the first displacement coefficients are coefficients corresponding to displacement vectors that indicate displacements of vertices or division points of the base mesh; the division points are vertices generated by subdividing the base mesh; the first group of displacement coefficients are a group of coefficients corresponding to a group of displacement vectors corresponding to the base mesh; and the displacement information is information regarding the displacement indicated by the group of displacement vectors corresponding to the base mesh.
[0296] (31) An information processing device comprising: a decoding unit configured to generate a displacement video having frame images as a displacement map storing, as pixel value groups, a first group of displacement coefficients including first displacement coefficients of a first bit depth by decoding encoded data; and a displacement coefficient generation unit configured to generate, by using displacement information on the first group of displacement coefficients extracted from the displacement map, a second group of displacement coefficients including second displacement coefficients of a second bit depth higher than the first bit depth, wherein the displacement coefficients are coefficients corresponding to displacement vectors indicating displacement of vertices or division points of a base mesh, the base mesh is a mesh with lower resolution than an original mesh to be encoded that is composed of vertices and connections that express a three-dimensional structure of an object and is generated by thinning out vertices from the original mesh, the original mesh being generated by thinning out vertices from the original mesh, the original mesh being composed of vertices and connections that express the three-dimensional structure of the object, the division points are vertices generated by subdividing the base mesh, the displacement information is information on the displacement indicated by a group of displacement vectors corresponding to the base mesh, and the second group of displacement coefficients is a group of coefficients corresponding to the group of displacement vectors corresponding to the base mesh. (32) The information processing device according to (31), wherein the first displacement coefficient group is constituted by the first displacement coefficients obtained by dividing the second displacement coefficient into a plurality of parts, and the displacement coefficient generation unit is configured to generate the second displacement coefficient group by combining the first displacement coefficients of the first displacement coefficient group extracted from the displacement video that correspond to each other. (33) The information processing device according to (32), wherein the displacement map includes a first pixel region in which a second coefficient sequence representing a group of second divided coefficients is arranged, adjacent to a first pixel region in which a first coefficient sequence representing a group of first divided coefficients is arranged, and the displacement coefficient generation unit is configured to generate the second displacement coefficient group by combining the first divided coefficients and the second divided coefficients that correspond to each other and are extracted from the displacement video.(34) The information processing device according to (32) or (33), wherein the displacement map includes a group of pixels representing an arrangement of sets of first and second division coefficients obtained by dividing the second displacement coefficient, and the displacement coefficient generation unit is configured to generate the second displacement coefficient set by combining the first and second division coefficients corresponding to each other and extracted from the displacement video. (35) The information processing device according to any of (32) to (34), wherein the displacement map includes a coefficient sequence in which the second displacement coefficients are arranged in a predetermined order, and the coefficient sequence is arranged in a reverse scanning order from the bottom right pixel of the displacement map. (36) The information processing device according to any one of (31) to (35), wherein the decoding unit is configured to generate the displacement video and excess information indicating displacement coefficients included in the second displacement coefficient set that exceed the first bit depth by decoding the encoded data, and the displacement coefficient generation unit is configured to generate the second displacement coefficient set using the first displacement coefficient set and the excess information extracted from the displacement video. (37) The information processing device according to (36), wherein the excess information is information indicating identification information of the displacement coefficients that exceed the first bit depth and an excess amount. (38) The information processing device according to (37), wherein the excess amount is information including a positive or negative sign. (39) The information processing device according to (37) or (38), wherein the excess information is information for each positive or negative sign of the excess amount. (40) The information processing device according to any one of (37) to (39), wherein the excess amount is information without including a positive or negative sign. (41) The information processing device according to any one of (37) to (40), wherein the excess information includes, as information indicating the identification information, a difference value between the identification information of the displacement coefficients that exceed the first bit depth. (42) The information processing device according to any one of (31) to (41), wherein the first set of displacement coefficients is obtained by shifting the second set of displacement coefficients by an offset, and the displacement coefficient generation unit is configured to generate the second set of displacement coefficients by shifting the first set of displacement coefficients extracted from the displacement video inversely by the offset.(43) The information processing device according to (42), wherein the offset is a predetermined value. (44) The information processing device according to (42) or (43), wherein the decoding unit is configured to generate the displacement video and the offset by decoding the encoded data, and the displacement coefficient generation unit is configured to generate the second set of displacement coefficients by decrementing the first set of displacement coefficients extracted from the displacement video by the generated offset. (45) The information processing device according to any of (42) to (44), wherein the offset has an independent value for each level of subdivision of the base mesh, and the displacement coefficient generation unit is configured to generate the second set of displacement coefficients by decrementing the first displacement coefficient to be processed of the first set of displacement coefficients extracted from the displacement video by the offset whose value corresponds to the level. (46) The information processing device according to any one of (42) to (45), wherein the offset has an independent value for each positive or negative sign of the second displacement coefficient, and the displacement coefficient generation unit is configured to generate the second set of displacement coefficients by reversely shifting the first displacement coefficient to be processed of the first set of displacement coefficients extracted from the displacement video by the offset, the value of which corresponds to the sign of the second displacement coefficient. (47) The information processing device according to any one of (31) to (46), wherein the decoding unit is configured to generate control information for controlling processing related to the displacement information by decoding the encoded data, and the displacement coefficient generation unit is configured to generate the second set of displacement coefficients based on the generated control information. (48) The information processing device according to (47), wherein the control information includes execution control information indicating whether to apply the displacement information, and the displacement coefficient generation unit is configured to generate the second set of displacement coefficients using the displacement information when the execution control information indicates that the displacement information is to be applied.(49) The information processing device according to (48), wherein the execution control information includes information specifying a subdivision level of the base mesh, and the displacement coefficient generation unit is configured to generate the second set of displacement coefficients for the level specified by the execution control information using the displacement information. (50) The information processing device according to (48) or (49), wherein the execution control information includes information specifying a channel of a component, and the displacement coefficient generation unit is configured to generate the second set of displacement coefficients for the channel specified by the execution control information using the displacement information. (51) The information processing device according to any of (48) to (50), wherein the execution control information includes information specifying a generation method of the displacement information, and the displacement coefficient generation unit is configured to apply the generation method specified by the execution control information and generate the second set of displacement coefficients using the displacement information. (52) An information processing method, comprising: decoding encoded data to generate a displacement video having frame images as a displacement map storing, as pixel value groups, a first group of displacement coefficients including first displacement coefficients of a first bit depth; generating, using displacement information regarding the first group of displacement coefficients extracted from the displacement map, a second group of displacement coefficients including second displacement coefficients of a second bit depth higher than the first bit depth; the displacement coefficients are coefficients corresponding to displacement vectors indicating displacement of vertices or division points of a base mesh; the base mesh is a mesh with lower resolution than an original mesh to be encoded, which is composed of vertices and connections that represent a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh; the division points are vertices generated by subdividing the base mesh; the displacement information is information regarding the displacement indicated by a group of displacement vectors corresponding to the base mesh; and the second group of displacement coefficients is a group of coefficients corresponding to the group of displacement vectors corresponding to the base mesh.
[0297] 300 Encoding device, 311 Displacement vector processing unit, 312 Displacement information generation unit, 313 Packing unit, 314 Mesh reconstruction unit, 315 Attribute map conversion unit, 316 Encoding unit, 317 Multiplexing unit, 331 Atlas information encoding unit, 332 Base mesh encoding unit, 333 Excess information encoding unit, 334 Offset encoding unit, 335 Control information encoding unit, 336 Displacement video encoding unit, 337 Attribute video encoding unit, 400 Decoding device, 411 Demultiplexing unit, 412 Decoding unit, 413 Base mesh reconstruction unit, 414 Subdivision unit, 415 Unpacking unit, 416 Displacement coefficient setting unit, 417 Displacement vector generation unit, 418 Displacement vector application unit, 419 Display processing unit, 431 Atlas information decoding unit 432 Base mesh decoding unit 433 Displacement video decoding unit 434 Excess information decoding unit 435 Offset decoding unit 436 Control information decoding unit 437 Attribute video decoding unit 900 Computer
Claims
1. An information processing device comprising: a displacement information generation unit configured to generate, for a base mesh, displacement information regarding a second displacement coefficient group including second displacement coefficients of a second bit depth lower than the first bit depth, using a first displacement coefficient group including first displacement coefficients of a first bit depth; a packing unit configured to generate a displacement map of the second bit depth in which the second displacement coefficient group is stored as a pixel value group; and an encoding unit configured to encode a displacement video having the displacement map as a frame image, wherein the base mesh is a mesh with lower resolution than an original mesh to be encoded, which is composed of vertices and connections that represent a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh, the original mesh being composed of vertices and connections that represent a three-dimensional structure of the object; the first displacement coefficients are coefficients corresponding to displacement vectors that indicate the displacement of vertices or division points of the base mesh; the division points are vertices generated by subdividing the base mesh; the first displacement coefficient group is a coefficient group corresponding to a displacement vector group corresponding to the base mesh; and the displacement information is information regarding the displacement indicated by the displacement vector group corresponding to the base mesh.
2. The information processing device according to claim 1, wherein the displacement information generating unit is configured to generate the second displacement coefficient group as the displacement information by dividing each of the first displacement coefficients into a plurality of the second displacement coefficients.
3. The information processing device according to claim 2, wherein the displacement information generation unit is configured to divide the first displacement coefficient into a first divided coefficient and a second divided coefficient as the second displacement coefficient, and to arrange a second coefficient sequence representing a group of the second divided coefficients adjacent to a first coefficient sequence representing a group of the first divided coefficients, and the packing unit is configured to generate the displacement map in which a second pixel region representing the second coefficient sequence is arranged adjacent to a first pixel region representing the first coefficient sequence in accordance with the arrangement of the first coefficient sequence and the second coefficient sequence.
4. The information processing device according to claim 2, wherein the displacement information generation unit is configured to generate a set of the first and second division coefficients as the second displacement coefficient by dividing the first displacement coefficient into a first division coefficient and a second division coefficient, respectively, and to arrange the first and second division coefficients consecutively in accordance with the arrangement of the first displacement coefficient before division; and the packing unit is configured to generate the displacement map in which a group of pixels representing the arrangement of the first and second division coefficients is arranged in accordance with the arrangement of the first and second division coefficients.
5. The information processing device according to claim 2, wherein the displacement information generation unit is configured to arrange the second displacement coefficients as a coefficient sequence in a predetermined order, and the packing unit is configured to generate the displacement map in which the second displacement coefficients are arranged in the order of arrangement of the coefficient sequence, starting from the bottom right pixel of the displacement map in the reverse direction of scanning order.
6. The information processing device of claim 1, wherein the displacement information generation unit is configured to generate, as the displacement information, excess information indicating displacement coefficients included in the first displacement coefficient group that exceed the second bit depth, and to generate the second displacement coefficient group by replacing the displacement coefficients in the first displacement coefficient group that exceed the second bit depth with values of the second bit depth, and the encoding unit is configured to encode the excess information and encode the displacement video in which the second displacement coefficient group is stored as a pixel value group.
7. The information processing device according to claim 6, wherein the excess information is information indicating identification information of the displacement coefficient that exceeds the second bit depth and an amount of excess.
8. The information processing device according to claim 7, wherein the excess amount is information including a positive or negative sign.
9. The information processing device according to claim 7, wherein the excess information is information for each positive and negative sign of the excess amount.
10. The information processing device according to claim 7, wherein the excess amount is information that does not include a positive or negative sign.
11. The information processing device according to claim 7, wherein the excess information includes, as information indicating the identification information, a difference value between the identification information of the displacement coefficients that exceed the second bit depth.
12. The information processing device according to claim 1, wherein the displacement information generating section is configured to generate the second set of displacement coefficients as the displacement information by shifting the first set of displacement coefficients by an offset amount.
13. The information processing device according to claim 12, wherein the offset is a predetermined value.
14. The information processing device according to claim 12, wherein the displacement information generation unit is configured to generate the second group of displacement coefficients by setting the offset and shifting the first group of displacement coefficients by the set offset, and the encoding unit is configured to encode the offset and encode the displacement video in which the second group of displacement coefficients is stored as a group of pixel values.
15. The information processing device according to claim 12, wherein the offset has an independent value for each level of subdivision of the base mesh, and the displacement information generating unit is configured to generate the second group of displacement coefficients by shifting the first displacement coefficient to be processed in the first group of displacement coefficients by an amount equal to the offset, the value of which corresponds to the level.
16. The information processing device according to claim 12, wherein the offset has an independent value for each positive or negative sign of the first displacement coefficient, and the displacement information generation unit is configured to generate the second displacement coefficient group by shifting the first displacement coefficient to be processed in the first displacement coefficient group by an amount of the offset whose value corresponds to the sign of the first displacement coefficient.
17. The information processing device according to claim 1, wherein the displacement information generation unit is configured to generate control information for controlling processing related to the displacement information, and the encoding unit is configured to encode the generated control information.
18. An information processing method comprising: generating, for a base mesh, displacement information relating to a second displacement coefficient group including second displacement coefficients of a second bit depth lower than the first bit depth, using a first displacement coefficient group including first displacement coefficients of a first bit depth; generating a displacement map of the second bit depth in which the second displacement coefficient group is stored as a pixel value group; and encoding a displacement video having the displacement map as a frame image, wherein the base mesh is a mesh with lower resolution than an original mesh to be encoded, the original mesh being composed of vertices and connections expressing a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh; the first displacement coefficients are coefficients corresponding to displacement vectors indicating displacements of vertices or division points of the base mesh; the division points are vertices generated by subdividing the base mesh; the first displacement coefficient group is a coefficient group corresponding to a displacement vector group corresponding to the base mesh; and the displacement information is information regarding the displacement indicated by the displacement vector group corresponding to the base mesh.
19. An information processing device comprising: a decoding unit configured to generate a displacement video having a displacement map, the frame images of which are a displacement map storing a first displacement coefficient group including first displacement coefficients of a first bit depth as a group of pixel values by decoding encoded data; and a displacement coefficient generation unit configured to generate a second displacement coefficient group including second displacement coefficients of a second bit depth higher than the first bit depth using displacement information regarding the first displacement coefficient group extracted from the displacement map, wherein the second displacement coefficients are coefficients corresponding to displacement vectors indicating the displacement of vertices or division points of a base mesh, the base mesh is a mesh with lower resolution than an original mesh to be encoded, the original mesh being composed of vertices and connections expressing a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh, the original mesh being composed of vertices and connections expressing a three-dimensional structure of an object, the division points are vertices generated by subdividing the base mesh, the displacement information is information regarding the displacement indicated by a displacement vector group corresponding to the base mesh, and the second displacement coefficient group is a coefficient group corresponding to the displacement vector group corresponding to the base mesh.
20. An information processing method comprising: generating a displacement video having a displacement map, as frame images, storing a first displacement coefficient group including first displacement coefficients of a first bit depth as a group of pixel values by decoding encoded data; generating a second displacement coefficient group including second displacement coefficients of a second bit depth higher than the first bit depth using displacement information regarding the first displacement coefficient group extracted from the displacement map; the second displacement coefficients are coefficients corresponding to displacement vectors indicating the displacement of vertices or division points of a base mesh; the base mesh is a mesh with lower resolution than an original mesh to be encoded, the original mesh being composed of vertices and connections that express a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh; the division points are vertices generated by subdividing the base mesh; the displacement information is information regarding the displacement indicated by a displacement vector group corresponding to the base mesh; and the second displacement coefficient group is a coefficient group corresponding to the displacement vector group corresponding to the base mesh.
Citation Information
Patent Citations
Apparatus and method for displacement mesh compression
JP2021149942A
Encoding device, decoding device, encoding method, and decoding method
WO2024214745A1