Information processing device and method, and bitstream
By differentially encoding UV vertex positions between frames, the encoding efficiency of non-tracked meshes is improved, addressing the inefficiencies in existing 3D data encoding methods and maintaining image quality.
Patent Information
- Application Number
- PCT/JP2024/045358
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-10
- Filing Date
- 2024-12-23
- Publication Date
- 2025-07-17
AI Technical Summary
Existing 3D data encoding methods, such as V-DMC, face reduced encoding efficiency when encoding non-tracked meshes due to the inability to apply inter-encoding, leading to potential reductions in quality and efficiency.
Differentially encode UV vertex positions between frames to enable inter-encoding, generating a bitstream that includes encoded data of a base mesh, displacement video, attribute video, and atlas information, using methods like Draco for encoding and decoding.
Enhances encoding efficiency by allowing inter-encoding of non-tracked meshes, maintaining quality and reducing the risk of reduced image quality in reconstructed meshes.
Smart Images

Figure JP2024045358_17072025_PF_FP_ABST
Abstract
Description
Information processing device and method, and bitstream
[0001] The present disclosure relates to an information processing device and method, and a bitstream, and in particular to an information processing device and method, and a bitstream that can suppress a decrease in encoding efficiency of 3D data including a base mesh and a displacement vector.
[0002] Conventionally, V-DMC (Video-based Dynamic Mesh Coding) has been used as a method for encoding meshes, which are 3D data that represent the three-dimensional structure of an object using vertices and connections (see, for example, Non-Patent Document 1). In V-DMC, inter-coding can be applied when the mesh to be encoded is a tracked mesh. In a tracked mesh, the topology of the geometry vertices in the sequence matches the keyframe, and the topology and positions of the UV vertices match the keyframe.
[0003] Khaled Mammou, Jungsun Kim, Alexis Tourapis, Dimitri Podborski, Krasimir Kolarov, "[V-CG] Apple's Dynamic Mesh Coding CfP Response", ISO / IEC JTC 1 / SC 29 / WG 7 m59281, April 2022
[0004] However, in other words, if the mesh to be coded is not a tracked mesh, inter-coding cannot be applied. Therefore, the number of frames to which inter-coding can be applied is reduced, and there is a risk that coding efficiency will decrease.
[0005] The present disclosure has been made in view of such circumstances, and makes it possible to suppress a decrease in the coding efficiency of 3D data including a base mesh and a displacement vector.
[0006] An information processing device according to one aspect of the present technology includes an atlas information encoding unit that differentially encodes UV vertex positions between a current frame and a key frame, and generates encoded data of atlas information including the differentially encoded UV vertex positions; and a bitstream generation unit that generates a bitstream including encoded data of a base mesh, encoded data of a displacement video, encoded data of an attribute video, and encoded data of the atlas information, wherein the base mesh is a mesh with lower resolution than an original mesh to be encoded, the original mesh being composed of vertices and connections that represent a three-dimensional structure of an object, and the displacement video is a moving image having frame images of a displacement map, which is a two-dimensional area packed with displacement vectors indicating the displacement of division points, which are vertices generated by subdividing the base mesh; the attribute video is a moving image having frame images of an attribute map, which is a two-dimensional area packed with projected images of the texture of the original mesh; the atlas information is information used for mesh reconstruction, including information indicating the correspondence between the base mesh, the displacement map, and the attribute map; and the UV vertex positions are position information indicating the positions of the vertices of the base mesh in the attribute map.
[0007] An information processing method according to one aspect of the present technology includes: differentially encoding UV vertex positions between a current frame and a key frame; generating encoded data of atlas information including the differentially encoded UV vertex positions; and generating a bitstream including encoded data of a base mesh, encoded data of a displacement video, encoded data of an attribute video, and encoded data of the atlas information; the base mesh is a mesh with lower resolution than an original mesh to be encoded, which is composed of vertices and connections that represent a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh; and the displacement video is generated by subdividing the base mesh. the attribute video is a moving image whose frame images are a displacement map, which is a two-dimensional area into which displacement vectors indicating the displacement of division points, which are vertices generated by converting the original mesh into a pixel value, are packed; the attribute video is a moving image whose frame images are an attribute map, which is a two-dimensional area into which projected images of the texture of the original mesh are packed; the atlas information is information used for reconstructing a mesh, including information indicating the correspondence between the base mesh, the displacement map, and the attribute map; and the UV vertex positions are position information indicating the positions of the vertices of the base mesh in the attribute map.
[0008] An information processing device according to another aspect of the present technology includes a demultiplexing unit that demultiplexes a bitstream and extracts coded data of a base mesh, coded data of a displacement video, coded data of an attribute video, and coded data of atlas information, and an atlas information decoding unit that differentially decodes UV vertex positions that are differentially coded between a current frame and a key frame and are included in the extracted coded data of the atlas information, wherein the base mesh is a mesh with lower resolution than an original mesh to be coded, the original mesh being composed of vertices and connections that represent a three-dimensional structure of an object, and the displacement video is generated by thinning out vertices from the original mesh, and the atlas information decoding unit is configured to decode the UV vertex positions. The information processing device is a moving image having frame images of a displacement map, which is a two-dimensional area packed with displacement vectors indicating the displacement of division points, which are vertices generated by subdividing a base mesh; the attribute video is a moving image having frame images of an attribute map, which is a two-dimensional area packed with projected images of the texture of the original mesh; the atlas information is information used to reconstruct a mesh, including information indicating the correspondence between the base mesh, the displacement map, and the attribute map; and the UV vertex positions are position information indicating the positions of the vertices of the base mesh in the attribute map.
[0009] An information processing method according to another aspect of the present technology includes demultiplexing a bitstream, extracting coded data of a base mesh, coded data of a displacement video, coded data of an attribute video, and coded data of atlas information, and differentially decoding UV vertex positions that are differentially coded between a current frame and a key frame and are included in the extracted coded data of the atlas information, the base mesh being a mesh with lower resolution than an original mesh to be coded, the original mesh being composed of vertices and connections that represent a three-dimensional structure of an object, and the displacement video being generated by thinning out vertices from the original mesh, the original mesh being a mesh with lower resolution than the original mesh, the original mesh being a mesh to be coded, the original mesh being composed of vertices and connections that represent a three-dimensional structure of an object, and the displacement video being generated by subdividing the base mesh. the attribute video is a moving image whose frame images are a displacement map, which is a two-dimensional area into which displacement vectors indicating the displacement of division points, which are vertices generated by converting the original mesh into a pixel value, are packed; the attribute video is a moving image whose frame images are an attribute map, which is a two-dimensional area into which projected images of the texture of the original mesh are packed; the atlas information is information used for reconstructing a mesh, including information indicating the correspondence between the base mesh, the displacement map, and the attribute map; and the UV vertex positions are position information indicating the positions of the vertices of the base mesh in the attribute map.
[0010] According to yet another aspect of the present technology, there is provided an information processing device comprising: a base mesh encoding unit that, for a current frame in which a geometry vertex topology matches a key frame and at least one of a UV vertex position and a UV vertex topology changes with respect to the key frame, differentially encodes geometry vertex positions relative to the key frame to generate coded data of a base mesh; an atlas information encoding unit that differentially encodes at least one of the UV vertex positions and the UV vertex identifiers with respect to the current frame to generate coded data of atlas information; and a bitstream generation unit that generates a bitstream including the coded data of the base mesh, coded data of a displacement video, coded data of an attribute video, and coded data of the atlas information, wherein the base mesh is a mesh with lower resolution than an original mesh to be coded, the original mesh being composed of vertices and connections that represent a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh; The displacement video is a moving image whose frame images are displacement maps, which are two-dimensional areas packed with displacement vectors indicating the displacement of division points, which are vertices generated by subdividing the base mesh; the attribute video is a moving image whose frame images are attribute maps, which are two-dimensional areas packed with projected images of the texture of the original mesh; the atlas information is information used to reconstruct a mesh, including information indicating the correspondence between the base mesh, the displacement map, and the attribute map; the geometry vertex positions indicate the positions of the vertices of the base mesh in a three-dimensional area; the geometry vertex topology indicates how the vertices of the base mesh are connected in the three-dimensional area; the UV vertex positions indicate the positions of the vertices of the base mesh in the attribute map; and the UV vertex topology indicates how the vertices of the base mesh are connected in the attribute map.
[0011] An information processing method according to yet another aspect of the present technology includes: for a current frame in which a geometry vertex topology matches a key frame and at least one of a UV vertex position and a UV vertex topology changes with respect to the key frame, differentially encoding geometry vertex positions relative to the key frame to generate coded data of a base mesh; differentially encoding at least one of the UV vertex positions and UV vertex identifiers for the current frame to generate coded data of atlas information; and generating a bitstream including coded data of the base mesh, coded data of a displacement video, coded data of an attribute video, and coded data of the atlas information; the base mesh is a mesh with lower resolution than the original mesh to be coded, which is formed by thinning vertices from the original mesh to be coded and which is composed of vertices and connections that represent a three-dimensional structure of an object; and the displacement video is generated by subdividing the base mesh. the attribute video is a moving image having frame images of an attribute map, which is a two-dimensional area into which projection images of the texture of the original mesh are packed; the atlas information is information used for reconstructing a mesh, including information indicating a correspondence between the base mesh and the displacement map and the attribute map; the geometry vertex positions indicate positions of the vertices of the base mesh in a three-dimensional area; the geometry vertex topology indicates how the vertices of the base mesh are connected in the three-dimensional area; the UV vertex positions indicate positions of the vertices of the base mesh in the attribute map; and the UV vertex topology indicates how the vertices of the base mesh are connected in the attribute map.
[0012] A bitstream according to yet another aspect of the present technology includes: encoded data of a base mesh in which geometry vertex positions are differentially encoded relative to a key frame for a frame in which geometry vertex topology matches a key frame and at least one of UV vertex positions and UV vertex topology changes relative to the key frame; and encoded data of atlas information in which at least one of the UV vertex positions and UV vertex identifiers are differentially encoded relative to the key frame, wherein the base mesh is a mesh with lower resolution than an original mesh to be encoded that is configured by vertices and connections that represent a three-dimensional structure of an object and that is generated by thinning out vertices from the original mesh, and the atlas information includes the base mesh, a displacement map, and an attribute map. the displacement map is a two-dimensional region packed with displacement vectors indicating the displacement of division points, which are vertices generated by subdividing the base mesh; the attribute map is a two-dimensional region packed with a projected image of the texture of the original mesh; the geometry vertex positions indicate the positions of the vertices of the base mesh in a three-dimensional region; the geometry vertex topology indicates how the vertices of the base mesh are connected in the three-dimensional region; the UV vertex positions indicate the positions of the vertices of the base mesh in the attribute map; and the UV vertex topology is a bit stream indicating how the vertices of the base mesh are connected in the attribute map.
[0013] In an information processing device and method according to one aspect of the present technology, UV vertex positions are differentially encoded between a current frame and a key frame, encoded data of atlas information including the differentially encoded UV vertex positions is generated, and a bitstream including encoded data of a base mesh, encoded data of a displacement video, encoded data of an attribute video, and encoded data of the atlas information is generated.
[0014] In an information processing device and method according to another aspect of the present technology, the bitstream is demultiplexed, and encoded data of the base mesh, encoded data of the displacement video, encoded data of the attribute video, and encoded data of the atlas information are extracted, and the UV vertex positions differentially encoded between the current frame and the key frame contained in the extracted encoded data of the atlas information are differentially decoded.
[0015] In an information processing device and method according to yet another aspect of the present technology, for a current frame in which the geometry vertex topology matches a key frame and at least one of the UV vertex positions and UV vertex topology changes relative to the key frame, the geometry vertex positions are differentially encoded relative to the key frame to generate encoded data for the base mesh, and for that current frame, at least one of the UV vertex positions and UV vertex identifiers are differentially encoded to generate encoded data for the atlas information, and a bitstream is generated that includes the encoded data for the base mesh, the encoded data for the displacement video, the encoded data for the attribute video, and the encoded data for the atlas information.
[0016] A bitstream according to yet another aspect of the present technology includes encoded data of a base mesh in which the geometry vertex positions are differentially encoded relative to the key frame for frames in which the geometry vertex topology matches the key frame and at least one of the UV vertex positions and UV vertex topology changes relative to the key frame, and encoded data of atlas information in which at least one of the UV vertex positions and UV vertex identifiers are differentially encoded relative to the key frame.
[0017] FIG. 1 is a diagram illustrating a mesh. FIG. 1 is a diagram illustrating V-DMC. FIG. 1 is a diagram illustrating a topology. FIG. 1 is a diagram illustrating an example of a tracked mesh. FIG. 1 is a diagram illustrating an example of a non-tracked mesh. FIG. 1 is a diagram illustrating an example of a non-tracked mesh. FIG. 1 is a diagram illustrating an example of information transmitted in inter-coding. FIG. 1 is a diagram illustrating an example of an encoding result. FIG. 1 is a diagram illustrating an example of a method for differential encoding of UV vertices. FIG. 1 is a diagram illustrating differential encoding of UV vertices. FIG. 1 is a diagram illustrating differential encoding of UV vertices. FIG. 1 is a diagram illustrating an example of syntax for differential encoding of UV vertex positions. FIG. 1 is a diagram illustrating an example of differential decoding of UV vertex positions. FIG. 1 is a diagram illustrating an example of syntax for differential encoding of UV vertex identifiers. FIG. 1 is a diagram illustrating an example of differential decoding of UV vertex identifiers. FIG. 1 is a diagram illustrating an example of encoding result. A block diagram illustrating an example of the main configuration of an encoding device. A block diagram illustrating an example of the main configuration of a V-DMC encoding unit. A flowchart illustrating an example of the flow of encoding processing. A flowchart illustrating an example of the flow of V-DMC encoding processing. A block diagram illustrating an example of the main configuration of a decoding device. A flowchart illustrating an example of the flow of decoding processing. A block diagram illustrating an example of the main configuration of a computer.
[0018] Below, modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described. The description will be made in the following order: 1. Literature etc. supporting technical content and technical terminology 2. Application of inter coding in V-DMC 3. Inter coding of UV vertices 4. First embodiment (encoding device) 5. Second embodiment (decoding device) 6. Supplementary notes
[0019] <1. Literature, etc. supporting technical content and technical terminology> The scope of disclosure of the present technology includes not only the content described in the embodiments, but also the content described in the following non-patent documents, etc. that were publicly known at the time of filing, and the content of other documents referenced in the following non-patent documents.
[0020] Non-patent document 1: (mentioned above)
[0021] In other words, the contents of the above-mentioned non-patent documents and the contents of other documents referenced in the above-mentioned non-patent documents are also used as the basis for determining the support requirements.
[0022] <2. Application of inter-coding in V-DMC> <V-DMC> Conventionally, 3D data representing the three-dimensional structure of a three-dimensional structure (object with a three-dimensional shape) has been available as a mesh, which represents the three-dimensional shape of the object surface by forming polygons with vertices and connections (also called edges).
[0023] As shown in the upper left of Figure 1, in a mesh, vertices 11 and connections 12 connecting these vertices 11 form polygonal planes (polygons). These polygons (also called faces) represent the surface of a three-dimensional object, i.e., the three-dimensional shape of the object. A texture 13 can be applied to each face of this mesh.
[0024] Mesh data is composed of information such as that shown in the lower part of Figure 1. Vertex information 14, shown first from the left in the lower part of Figure 1, is information indicating the three-dimensional position (three-dimensional coordinates (X, Y, Z)) of each vertex 11 that constitutes the mesh. Connection information 15, shown second from the left in the lower part of Figure 1, is information indicating each connection (edge) 12 that constitutes the mesh. A texture image 16, shown third from the left in the lower part of Figure 1, is map information for the texture 13 that is applied to each face. A UV map 17, shown fourth from the left in the lower part of Figure 1, is information indicating the correspondence between the vertices 11 and the texture 13. The UV map 17 indicates the coordinates (UV coordinates) of each vertex 11 in the texture image 16.
[0025] As an example of such a mesh coding method, there is V-DMC (Video-based Dynamic Mesh Coding) as disclosed in Non-Patent Document 1.
[0026] In V-DMC, the mesh to be encoded (referred to in this specification as the original mesh) is represented as a base mesh that is less fine (i.e., coarser) than the original mesh, and displacement vectors of the division points obtained by subdividing the base mesh, and the base mesh and displacement vectors are then encoded.
[0027] For example, assume that there is an original mesh as shown in the top row of Figure 2. The original mesh is a mesh composed of vertices and connections that represent the three-dimensional structure of an object, and is the target of encoding. For example, the original mesh is generated from a captured image of an object in real space (by camera capture). In Figure 2, black dots represent vertices, and lines connecting the black dots represent connections (edges). As described above, a mesh essentially forms polygons using vertices and edges, but for convenience of explanation, it is described here as a group of vertices connected linearly (in series).
[0028] By simplifying the original mesh, a coarse (low-resolution) mesh like the one shown in the second row from the top of Figure 2 is formed. This is called the base mesh. One simplification method is to thin out some of the vertices (decimate). In other words, the base mesh is a mesh with lower resolution than the original mesh, generated by thinning out vertices from the original mesh (i.e., simplifying the original mesh).
[0029] By subdividing each polygon of this base mesh, vertices and edges are added, as shown in the third row from the top of Figure 2. The degree of subdivision is arbitrary. That is, the number of vertices and edges added is arbitrary. For example, this subdivision can add vertices equal to the number of vertices thinned out from the original mesh. That is, subdivision can be used to maintain the same number of vertices as the original mesh. In this specification, these added vertices are also referred to as division points. This subdivision can also be repeated recursively. For example, in a technique called midpoint, the process of adding vertices to the midpoints of edges (subdivision) is repeated recursively. In other words, recursive subdivision increases the number of vertices and improves the resolution of the mesh. In this way, it is possible to perform subdivision up to any desired level of resolution (i.e., control the resolution of the subdivided mesh). In other words, the subdivided mesh can be layered according to its level of resolution. In other words, this can be considered a layering of the subdivision process and the vertices (division points) and edges obtained by the subdivision process.
[0030] However, the connections of the base mesh have been updated when the vertices of the original mesh are thinned out. Therefore, the division points obtained by subdivision are formed on these updated connections (edges). As a result, the shape of the subdivided base mesh differs from that of the original mesh. More specifically, as shown in the bottom part of Figure 2, the positions of the division points (on the dotted line) differ from those of the original mesh.
[0031] In other words, by moving the positions of the vertices of the subdivided base mesh (the vertices or division points of the base mesh) closer to the vertex positions of the original mesh, the difference in shape between the subdivided base mesh and the original mesh can be reduced. In this specification, such movement of the vertices of the subdivided base mesh (the vertices or division points of the base mesh) is also referred to as displacement. Furthermore, the amount and direction of this displacement, expressed as a vector, is also referred to as a displacement vector. Ideally, by displacing each vertex of the subdivided base mesh, the shape of the subdivided base mesh can be made to match the shape of the original mesh. In other words, the original mesh can be expressed as a base mesh and a displacement vector.
[0032] In V-DMC, such base meshes and displacement vectors are coded instead of the original mesh (geometry). By coding the base meshes and displacement vectors in this way, it is possible to code with a reduced number of polygons (i.e., the number of vertices and edges) compared to coding the original mesh, which generally reduces the amount of code for the same quality. In other words, it is possible to improve coding efficiency.
[0033] During decoding, as described above, a mesh is restored (generated) by subdividing a base mesh and applying a displacement vector to each vertex of the subdivided base mesh to displace it. In this specification, this restored mesh is also referred to as a restored mesh. Ideally, a restored mesh equivalent to the original mesh can be generated. Note that although the shape of the polygon (face) may be any polygonal shape, the following description will be given assuming that the polygon is triangular. Therefore, in the following, a polygon (face) will also be referred to as a triangle.
[0034] <Encoding and Decoding of V-DMC Data> In the case of V-DMC, mesh data consists of a base mesh, displacement vectors, attributes, and atlas information. This data group is also referred to as V-DMC data. The base mesh consists of information indicating vertices and connections, and is encoded using an existing mesh encoding method such as Draco.
[0035] Displacement vectors are converted into displacement coefficients using a predetermined method. These displacement coefficients are arranged as pixel values in a two-dimensional area (also called a displacement map). This arrangement (mapping) of displacement coefficients is also called packing. A moving image (also called a displacement video) with the displacement map as its frame image is coded using a coding method for 2D moving images. In other words, the displacement coefficients are scalar values corresponding to the displacement vectors. A displacement map is map information (also called image data) that stores the displacement coefficients as pixel values. A displacement video is moving image data with the displacement map as its frame image.
[0036] An attribute is non-geometry information applied to a mesh (geometry), which is 3D data. For example, an attribute may include a texture applied to a face of the mesh (geometry). The attribute (e.g., texture) is divided into multiple subregions, each of which is projected in a predetermined projection direction, and the projected images (patches) are arranged in a two-dimensional region (also called an attribute map). In other words, attribute patches are packed into the attribute map. A video (also called attribute video) using the attribute map as frame images is encoded using a 2D video encoding method. In other words, the attribute map is map information (also called image data) that stores the patches (projected textures) as pixel values. Attribute video is video data using the attribute map as frame images.
[0037] Atlas information is information used when reconstructing a mesh. For example, atlas information may include correspondence between the base mesh and a displacement map or attribute map (such as a UV map), quantized values of displacement vectors, etc. This atlas information is encoded using a predetermined encoding method.
[0038] The coded data (bitstream) of each data is decoded by a decoding method corresponding to the coding method. In other words, by decoding the coded data (bitstream), various information such as base meshes, displacement vectors, attributes, and atlas information is restored (generated).
[0039] <Application of Inter-Coding> V-DMC can encode mesh data that changes over time, such as video. In other words, it can encode a sequence of multiple meshes that represent the three-dimensional structure of a three-dimensional structure at different times. In this specification, each mesh at a time in such a mesh sequence is referred to as a frame, just as in the case of video.
[0040] In V-DMC, in addition to intra-coding, which independently codes each frame of such a mesh sequence, it may be possible to apply inter-coding (also called differential coding), which utilizes the correlation between frames.
[0041] For example, the base mesh includes information indicating the three-dimensional position (three-dimensional coordinates (X, Y, Z)) of each vertex. Note that, hereinafter, such a vertex in a three-dimensional region is also referred to as a geometry vertex. The position of the geometry vertex is also referred to as a geometry vertex position. In other words, the geometry vertex position indicates the position of a vertex of the base mesh in a three-dimensional region. The coordinates of the geometry vertex are also referred to as geometry coordinates. In other words, the base mesh includes information indicating the geometry vertex position (geometry coordinates).
[0042] The atlas information also includes information indicating the two-dimensional position (two-dimensional coordinates (u, v)) of each vertex in a texture image (a two-dimensional area where a texture is placed). Note that, hereinafter, vertices placed in such a two-dimensional area (also referred to as a texture image, attribute map, etc.) are also referred to as UV vertices. The positions of these UV vertices are also referred to as UV vertex positions. In other words, the UV vertex positions indicate the positions of the vertices of the base mesh in the two-dimensional area (also referred to as a texture image, attribute map, etc.). The coordinates of these UV vertices are also referred to as UV coordinates. In other words, the atlas information includes information (UV map) indicating the UV vertex positions (UV coordinates).
[0043] The atlas information also includes information indicating the identifiers of the vertices (geometry vertices or UV vertices) that make up each face. Note that, hereinafter, such information is also referred to as triangle information.
[0044] In some cases, inter-coding utilizing the correlation between frames can be applied to this information. Generally, as with video, inter-coding provides higher coding efficiency than intra-coding. In other words, in V-DMC, applying such inter-coding can generally be expected to improve coding efficiency compared to applying intra-coding.
[0045] However, to apply inter-coding, the mesh to be coded must be a tracked mesh, which is a mesh whose geometry vertex topology and UV vertex topology and positions match those of the keyframes in the sequence.
[0046] Topology refers to the configuration of faces, i.e., how each vertex is connected. For example, the mesh shown in FIG. 3A has five vertices forming three faces (triangles). The mesh shown in FIG. 3B has seven vertices forming three faces (triangles). The mesh shown in FIG. 3C has nine vertices forming three faces (triangles). In other words, the meshes in each example have the same number of faces, but the vertices that make up each face are different. In other words, the way the vertices are connected is different. Therefore, the topologies of these meshes do not match. Hereinafter, the topology of geometry vertices will also be referred to as geometry vertex topology. The topology of UV vertices will also be referred to as UV vertex topology. In other words, the geometry vertex topology refers to how the vertices of a base mesh are connected in a three-dimensional domain, and the UV vertex topology refers to how the vertices of a base mesh are connected in a two-dimensional domain (also referred to as a texture image, attribute map, etc.).
[0047] FIG. 4 shows an example of changes in geometry vertices and UV vertices in a sequence. In FIG. 4, the time direction progresses from left to right as indicated by the arrows. FIG. 4 also shows the appearance of a mesh of geometry vertices and a mesh of UV vertices in two frames. As shown in the upper part of the figure, the geometry vertex positions change between the two frames, but the geometry vertex topology remains consistent (the way each geometry vertex is connected does not change in the time direction). In contrast, as shown in the lower part of the figure, both the UV vertex positions and the UV vertex topology remain consistent between the two frames (the position and connection of each UV vertex does not change in the time direction). Therefore, the mesh in the example of FIG. 4 is a tracked mesh, and inter-coding can be applied.
[0048] FIG. 5 shows an example of changes in geometry vertices and UV vertices in a sequence. In FIG. 5, the time direction progresses from left to right as indicated by the arrows. FIG. 5 also shows the appearance of a mesh of geometry vertices and a mesh of UV vertices in two frames. As shown in the upper part of the figure, the geometry vertex topology changes between the two frames (the way each geometry vertex is connected changes in the time direction). Therefore, the mesh in the example of FIG. 5 is not a tracked mesh. Therefore, inter-coding cannot be applied to this mesh (intra-coding is applied).
[0049] FIG. 6 shows an example of changes in geometry vertices and UV vertices in a sequence. In FIG. 6, the time direction progresses from left to right as indicated by the arrows. FIG. 6 also shows the appearance of a mesh of geometry vertices and a mesh of UV vertices in two frames. As shown at the bottom of the figure, the UV vertex position (the position of UV vertex [0]) changes between the two frames. Therefore, the mesh in the example of FIG. 6 is not a tracked mesh. Therefore, inter-coding cannot be applied to this mesh (intra-coding is applied).
[0050] FIG. 7 shows an example of changes in geometry vertices and UV vertices in a sequence. In FIG. 7, the time direction progresses from left to right as indicated by the arrows. FIG. 7 also shows the appearance of a mesh of geometry vertices and a mesh of UV vertices in two frames. As shown at the bottom of the figure, the geometry vertex topology changes between the two frames (the way each geometry vertex is connected changes in the time direction). Therefore, the mesh in the example of FIG. 7 is not a tracked mesh. Therefore, inter-coding cannot be applied to this mesh (intra-coding is applied).
[0051] FIG. 8 is a diagram illustrating an example of information transmitted in conventional inter-coding. As described above, in V-DMC, the geometry vertex positions, UV vertex positions, and triangle information of each frame are encoded and transmitted. Let v be the geometry vertex position in a key frame, vt be the UV vertex position, and f be the triangle information. If v' is the geometry vertex position in the current frame (decode frame), in conventional inter-coding, the difference (v'-v) in the geometry vertex positions between the key frame and the current frame is encoded and transmitted. In the case of tracked meshes, as described above, the geometry vertex topology, UV vertex positions, and UV vertex topology do not change, so the UV vertex positions vt of the key frame are applied to the UV vertex positions of the current frame. The triangle information f of the key frame is applied to the triangle information of the current frame.
[0052] In other words, if any one of the geometry vertex topology, UV vertex positions, and UV vertex topology changes, the UV vertex positions vt of the key frame and the triangle information f of the key frame cannot be applied to the current frame, making inter-coding impossible, which could result in reduced coding efficiency.
[0053] 9 is a diagram showing an example of an encoding result. The upper part of the figure shows a sequence of meshes to be encoded, and the lower part shows an example of the encoding result (applied encoding method). As shown by the arrow, the time direction progresses from left to right in the figure.
[0054] As shown in the upper part of Figure 9, in the frame with number (fid) 1220, the geometry vertex topology does not change, but the UV vertex topology does. Therefore, inter-coding cannot be applied to the frame with number (fid) 1220, and intra-coding is applied instead. This could result in a decrease in coding efficiency.
[0055] For example, when an object (mesh) is deformed in a three-dimensional area, the face may be deformed, and the previous texture image may no longer be optimal. In such cases, the quality of the texture may be reduced, which may result in a reduction in the quality of the reconstructed mesh. Therefore, it is possible to reduce the reduction in the quality of the texture by adjusting the pixel distribution in the texture image by changing the UV vertex positions and UV vertex topology to match the new shape of the object (mesh), thereby reducing the reduction in the quality of the reconstructed mesh.
[0056] However, if the UV vertex positions or UV vertex topology are changed, inter-coding may not be applicable, as in the example of Fig. 9, and coding efficiency may be reduced. In other words, if the coding target is a tracked mesh so that inter-coding can be applied, the quality of the reconstructed mesh may be reduced.
[0057] <3. Inter-Coding of UV Vertices> <Method 1> Therefore, as shown in the top row of the table in FIG. 10 , UV vertex positions are differentially coded and transmitted (Method 1). In the case of conventional inter-coding, as shown in the example of FIG. 11A, differential coding is performed only on geometry vertex positions. In contrast, in the present technology, differential coding is applied to UV vertex positions (UV coordinates), as shown in the example of FIG. 11B. For example, as shown in FIG. 11B, when the shape of a mesh is reduced in a texture image, the UV vertex topology remains consistent, but the UV vertex positions change. Conventional methods could not apply inter prediction to such cases, but by applying differential coding to UV vertex positions (UV coordinates), inter-coding can be applied even in such cases.
[0058] Hereinafter, an information processing device that encodes 3D data including a base mesh and a displacement vector is also referred to as a first information processing device. For example, this first information processing device includes an atlas information encoding unit that differentially encodes UV vertex positions between a current frame and a key frame and generates encoded data of atlas information including the differentially encoded UV vertex positions, and a bitstream generation unit that generates a bitstream including the encoded data of the base mesh, the encoded data of the displacement video, the encoded data of the attribute video, and the encoded data of the atlas information.
[0059] In addition, in the first information processing device, UV vertex positions are differentially encoded between the current frame and the key frame, encoded data of atlas information including the differentially encoded UV vertex positions is generated, and a bitstream including encoded data of the base mesh, encoded data of the displacement video, encoded data of the attribute video, and encoded data of the atlas information is generated.
[0060] A base mesh is a mesh with lower resolution than an original mesh to be coded, which is composed of vertices and connections that represent the three-dimensional structure of an object. It is generated by thinning out vertices from the original mesh. A displacement video is a moving image whose frame images are displacement maps, which are two-dimensional regions packed with displacement vectors indicating the displacement of division points, which are vertices generated by subdividing a base mesh. An attribute video is a moving image whose frame images are attribute maps, which are two-dimensional regions packed with projected images of the texture of the original mesh. Atlas information is information used for mesh reconstruction, including information indicating the correspondence between the base mesh and the displacement map and attribute map. UV vertex positions are positional information indicating the positions of the vertices of the base mesh in the attribute map.
[0061] Here, differential coding (inter coding) of UV vertex positions is a technique for deriving differences between UV vertex positions and coding the differences. In the case of intra coding, the UV vertex positions of the current frame are coded. Therefore, inter coding generally results in higher coding efficiency than intra coding. In other words, by doing as described above, the first information processing device can suppress a decrease in coding efficiency of 3D data including a base mesh and a displacement vector.
[0062] In the following description, an information processing device that decodes the bitstream and generates 3D data including a base mesh and a displacement vector is also referred to as a "second information processing device." This second information processing device includes a demultiplexing unit that demultiplexes the bitstream and extracts coded data of the base mesh, coded data of the displacement video, coded data of the attribute video, and coded data of the atlas information, and an atlas information decoding unit that differentially decodes UV vertex positions that are differentially coded between the current frame and the key frame and that are included in the coded data of the extracted atlas information.
[0063] In addition, in the second information processing device, the bit stream is demultiplexed, and the encoded data of the base mesh, the encoded data of the displacement video, the encoded data of the attribute video, and the encoded data of the atlas information are extracted, and the UV vertex positions differentially encoded between the current frame and the key frame, which are included in the extracted encoded data of the atlas information, are differentially decoded.
[0064] A base mesh is a mesh with lower resolution than an original mesh to be coded, which is composed of vertices and connections that represent the three-dimensional structure of an object. It is generated by thinning out vertices from the original mesh. A displacement video is a moving image whose frame images are displacement maps, which are two-dimensional regions packed with displacement vectors indicating the displacement of division points, which are vertices generated by subdividing a base mesh. An attribute video is a moving image whose frame images are attribute maps, which are two-dimensional regions packed with projected images of the texture of the original mesh. Atlas information is information used for mesh reconstruction, including information indicating the correspondence between the base mesh and the displacement map and attribute map. UV vertex positions are positional information indicating the positions of the vertices of the base mesh in the attribute map.
[0065] Here, differential decoding (inter decoding) of UV vertex positions is a decoding method corresponding to the above-mentioned differential encoding. That is, it is a technique for decoding inter-coded encoded data to derive differences between UV vertex positions, and then adding the UV vertex positions of a key frame (already-encoded UV vertex positions) to the differences to derive UV vertex positions of the current frame. By applying inter decoding, inter-coded encoded data can be correctly decoded. In contrast, by applying intra decoding, intra-coded encoded data can be correctly decoded. In general, inter coding improves coding efficiency compared to intra coding. In other words, by doing as described above, the second information processing device can suppress a decrease in coding efficiency of 3D data including a base mesh and a displacement vector.
[0066] <UV Vertices from which Differences are Derive> Such differences in UV vertex positions are derived between corresponding UV vertices between the current frame and the key frame. For example, the UV vertices may be arranged in a predetermined order, and the UV coordinates may be differentially encoded between the UV vertices in the same order in the key frame. This makes it easy to encode and decode differences between corresponding UV vertices.
[0067] For example, in the first information processing device described above, the atlas information encoding unit may differentially encode UV vertex positions between vertices of the current frame and vertices of the key frame that are in the same order in a predetermined sorting order of UV vertex positions. Also, in the second information processing device described above, the atlas information decoding unit may differentially decode UV vertex positions between vertices of the current frame and vertices of the key frame that are in the same order in a predetermined sorting order of UV vertex positions.
[0068] <Key Frame> Note that a key frame may be any frame that can be referenced from the current frame. For example, a key frame may be a frame that is processed temporally before the current frame. For example, a key frame may be the frame immediately before the current frame.
[0069] <Geometry Vertices> Note that the geometry vertex topology between the current frame and the key frame may match that of the key frame. For example, in a first information processing device, when the geometry vertex topology between the current frame and the key frame matches that of the key frame, the atlas information encoding unit may differentially encode UV vertex positions between the current frame and the key frame, and generate encoded data of atlas information including the differentially encoded UV vertex positions. Furthermore, in a second information processing device, when the geometry vertex topology between the current frame and the key frame matches that of the key frame, the atlas information decoding unit may differentially decode the UV vertex positions differentially encoded between the current frame and the key frame, which are included in the encoded data of the atlas information.
[0070] Furthermore, geometry vertex positions may also be differentially encoded and differentially decoded. For example, the first information processing device may further include a base mesh encoding unit that differentially encodes geometry vertex positions indicating the positions of base mesh vertices in a three-dimensional region. For example, when the geometry vertex topology between the current frame and the key frame matches that of the key frame, the base mesh encoding unit may differentially encode the geometry vertex positions indicating the positions of base mesh vertices in a three-dimensional region. By doing so, the first information processing device can suppress a decrease in encoding efficiency of 3D data including the base mesh and displacement vectors.
[0071] For example, the second information processing device may further include a base mesh decoding unit that differentially decodes differentially encoded geometry vertex positions indicating the positions of the vertices of the base mesh in a three-dimensional region. For example, when the geometry vertex topology between the current frame and the key frame matches that of the key frame, the base mesh decoding unit may differentially decode the differentially encoded geometry vertex positions indicating the positions of the vertices of the base mesh in a three-dimensional region. By doing so, the second information processing device can suppress a decrease in the encoding efficiency of 3D data including the base mesh and the displacement vectors.
[0072] <Method 1-1> When Method 1 is applied, as shown in the second row from the top of the table in FIG. 10 , if derived_uv_present_flag is true, UV vertex positions may be differentially encoded and transmitted between the current frame and the key frame (Method 1-1). derived_uv_present_flag is flag information that controls whether or not to differentially encode UV vertex positions. If derived_uv_present_flag is true, differential encoding of UV vertex positions may be applied. If derived_uv_present_flag is false, application of differential encoding of UV vertex positions is prohibited. Hereinafter, derived_uv_present_flag is also referred to as UV vertex position encoding control information or first UV vertex position encoding control information. For example, a first information processing device may set derived_uv_present_flag and apply differential encoding to UV vertex positions according to derived_uv_present_flag. Alternatively, the first information processing device may encode derived_uv_present_flag (transmit it to the decoding side).
[0073] For example, the first information processing apparatus may further include an encoding control unit that generates first UV vertex position encoding control information that controls whether to differentially encode UV vertex positions. The atlas information encoding unit may then differentially encode the UV vertex positions in accordance with the first UV vertex position encoding control information, encode atlas information including the first UV vertex position encoding control information, and generate encoded data for the atlas information. In this manner, the first information processing apparatus can easily control differential encoding of UV vertex positions.
[0074] For example, derived_uv_present_flag may be set to true only when the number of faces of the mesh in the current frame is the same as that in the key frame (for an area with the same number of faces). For example, in a first information processing device, the encoding control unit may set the first UV vertex position encoding control information for an area where the number of faces of the mesh in the current frame and that in the key frame are the same to a value indicating that the UV vertex positions can be differentially encoded. This makes it possible to more reliably execute differential encoding and differential decoding correctly.
[0075] Furthermore, derived_uv_present_flag may be set to true only when the UV vertex positions of the current frame differ from those of the key frame (for areas with different UV vertex positions). For example, in the first information processing device, when the UV vertex positions of the current frame and the key frame differ from each other, the encoding control unit may set the first UV vertex position encoding control information to a value indicating that the UV vertex positions can be differentially encoded. This makes it possible to more reliably execute differential encoding and differential decoding correctly.
[0076] Note that if derived_uv_present_flag is false, the UV vertex positions of the current frame may be intra-coded. For example, in the first information processing device, if the first UV vertex position encoding control information is a value indicating that the UV vertex positions are not to be differentially encoded, the atlas information encoding unit may intra-code the UV vertex positions of the current frame.
[0077] <Method 1-1-1> When Method 1-1 is applied, as shown in the third row from the top of the table in FIG. 10 , if derived_uv_mode is set to delta, UV vertex positions may be differentially encoded and transmitted between the current frame and the key frame (Method 1-1-1). derived_uv_mode is control information that controls the encoding method for UV vertex positions. This derived_uv_mode can be set to delta or no_delta. delta is a value indicating differential encoding. That is, when derived_uv_mode is set to delta, differential encoding is applied as the encoding method for UV vertex positions. Also, no_delta is a value that does not indicate differential encoding. That is, when derived_uv_mode is set to no_delta, differential encoding is not applied as the encoding method for UV vertex positions (intra-encoding is applied). Hereinafter, derived_uv_mode is also referred to as UV vertex position encoding control information or second UV vertex position encoding control information. For example, a first information processing device may set this derived_uv_mode and apply differential encoding to UV vertex positions according to the derived_uv_mode. Alternatively, the first information processing device may encode the derived_uv_mode (transmit it to the decoding side). Note that the derived_uv_mode can be used in combination with the derived_uv_present_flag.
[0078] For example, in the first information processing apparatus, the encoding control unit may further generate second UV vertex position encoding control information that controls the encoding method of UV vertex positions. Then, when the first UV vertex position encoding control information is set to a value indicating that UV vertex positions can be differentially encoded and the second UV vertex position encoding control information is set to a value indicating differential encoding, the atlas information encoding unit may differentially encode the UV vertex positions between the current frame and the key frame. This allows the first information processing apparatus to easily control the differential encoding of UV vertex positions.
[0079] For example, in the second information processing device, when UV vertex position encoding control information that controls the encoding method of UV vertex positions is set to a value indicating differential encoding, the atlas information decoding unit may differentially decode the UV vertex positions between the current frame and the key frame. In this way, the second information processing device can easily control the differential decoding of the UV vertex positions.
[0080] Also, for example, derived_uv_mode may be set to delta only when the number of UV vertices in the current frame is the same as that in the key frame. As shown in the table of FIG. 12, if the number of UV vertices in the current frame and the key frame is the same (Tn), only the Tn differences need to be transmitted. In other words, in this case, differential encoding can be applied or not. That is, the value of derived_uv_mode may be set to delta or no_delta. On the other hand, if the number of UV vertices in the current frame and the key frame is different (Tn2 and Tn1), not all differences can be derived. That is, differential encoding cannot be applied in this case. That is, the value of derived_uv_mode can only be set to no_delta. Then, the Tn2 UV vertex positions of the current frame are intra-coded and transmitted.
[0081] For example, in the first information processing device, the encoding control unit may be prohibited from setting the second UV vertex position encoding control information to a "value indicating differential encoding" when the number of vertices between the current frame and the key frame is different. By doing so, it is possible to more reliably execute differential encoding and differential decoding correctly.
[0082] <Method 1-1-2> When Method 1-1 is applied, as shown in the fourth row from the top of the table in Figure 10, if derived_uv_mode is no_delta, the UV vertex positions of the current frame may be intra-coded and transmitted (Method 1-1-2). For example, in a first information processing device, an atlas information encoding unit may intra-code the UV vertex positions of the current frame if the second UV vertex position encoding control information has a value that does not indicate differential encoding. Also, in a second information processing device, an atlas information decoding unit may intra-decode the UV vertex positions of the current frame if the UV vertex position encoding control information has a value that does not indicate differential encoding.
[0083] As shown in the table in FIG. 12, derived_uv_mode can be set to no_delta whether the number of UV vertices in the current frame is the same as or different from that in the key frame.
[0084] <Syntax> An example of syntax for control using the UV vertex position encoding control information described above is shown in Figure 13. As shown in Figure 13, in this example, when derived_uv_present_flag is set to true, derived_uv_mode is set, and if derived_uv_mode is delta, differential encoding is applied to the UV vertex positions. If derived_uv_mode is no_delta, intra encoding is applied to the UV vertex positions.
[0085] <Decoding Process> An example of a pseudo program showing the decoding process when control is performed using the UV vertex position encoding control information described above is shown in Fig. 14. As shown in Fig. 14, in this example, when derived_uv_mode is delta, differential decoding is applied to the UV vertex positions, and when derived_uv_mode is no_delta, intra decoding is applied to the UV vertex positions.
[0086] <Method 1-2> When Method 1 is applied, as shown in the fifth row from the top of the table in FIG. 10 , UV vertex identifiers (also referred to as UV vertex Idx) may be differentially encoded and transmitted (Method 1-2). For example, as shown in C of FIG. 11 , differential encoding may be applied to UV vertex positions (UV coordinates) and UV vertex topology (i.e., UV vertex identifiers). For example, as shown in C of FIG. 11 , separating a triangle in a texture image changes not only the UV vertex positions but also the UV vertex topology. Conventional methods could not apply inter prediction to such cases, but by applying differential encoding to the UV vertex positions (UV coordinates) and UV vertex identifiers (UV vertex topology), inter prediction can be applied even in such cases.
[0087] For example, in the first information processing device, the atlas information encoding unit may differentially encode UV vertex identifiers that identify vertices of a base mesh in an attribute map, and generate encoded data of atlas information that includes the differentially encoded UV vertex identifiers. This allows the first information processing device to apply inter-coding in a wider variety of cases. Therefore, the first information processing device can suppress a decrease in the encoding efficiency of 3D data that includes a base mesh and a displacement vector.
[0088] In addition, in the second information processing device, the atlas information decoding unit may differentially decode UV vertex identifiers that identify the vertices of the base mesh in the attribute map, which have been differentially encoded. By doing so, the second information processing device can apply inter-decoding in a wider variety of cases. Therefore, the second information processing device can suppress a decrease in the encoding efficiency of 3D data including the base mesh and the displacement vector.
[0089] <Method 1-2-1> When Method 1-2 is applied, as shown in the sixth row from the top of the table in FIG. 10 , if derived_face_present_flag is true, UV vertices Idx may be differentially encoded and transmitted. derived_face_present_flag is flag information that controls whether or not to differentially encode UV vertex identifiers. If derived_face_present_flag is true, differential encoding of UV vertex identifiers may be applied. If derived_face_present_flag is false, application of differential encoding of UV vertex identifiers is prohibited. Hereinafter, derived_face_present_flag is also referred to as UV vertex identifier encoding control information or first UV vertex identifier encoding control information. For example, the first information processing device may set derived_face_present_flag and apply differential encoding to UV vertex identifiers according to derived_face_present_flag. Furthermore, the first information processing device may encode derived_face_present_flag (transmit it to the decoding side).
[0090] For example, the first information processing apparatus may further include an encoding control unit that generates first UV vertex identifier encoding control information that controls whether to differentially encode UV vertex identifiers. The atlas information encoding unit may then differentially encode the UV vertex identifiers in accordance with the first UV vertex identifier encoding control information, encode the atlas information including the first UV vertex identifier encoding control information, and generate encoded data for the atlas information. In this manner, the first information processing apparatus can easily control the differential encoding of UV vertex identifiers.
[0091] For example, derived_face_present_flag may be set to true only when the number of faces of the mesh in the current frame is the same as that in the key frame (for an area where the number of faces is the same). For example, in a first information processing device, the encoding control unit may set the first UV vertex identifier encoding control information for an area where the number of faces of the mesh in the current frame and that in the key frame are the same to a value indicating that the UV vertex identifiers can be differentially encoded. This makes it possible to more reliably execute differential encoding and differential decoding correctly.
[0092] Furthermore, derived_face_present_flag may be set to true only when the UV vertex topology of the current frame is different from that of the key frame (for an area with a different UV vertex topology). For example, in the first information processing device, when the UV vertex identifiers differ between the current frame and the key frame, the encoding control unit may set the first UV vertex identifier encoding control information to a value indicating that the UV vertex identifiers can be differentially encoded. This makes it possible to more reliably execute differential encoding and differential decoding correctly.
[0093] In addition, if derived_face_present_flag is false, the UV vertex identifiers of the current frame may be intra-coded. For example, in the first information processing device, if the first UV vertex identifier encoding control information is a value indicating that the UV vertex identifiers are not differentially encoded, the atlas information encoding unit may intra-code the UV vertex identifiers of the current frame.
[0094] <Method 1-2-1-1> When Method 1-2-1 is applied, as shown in the seventh row from the top of the table in FIG. 10, if derived_face_mode is delta, UV vertex Idx may be differentially encoded and transmitted between the current frame and the key frame (Method 1-2-1-1). derived_face_mode is control information that controls the encoding method of UV vertex identifiers. As shown in the table in FIG. 12, derived_face_mode may be set to delta, delta_in_face, or no_delta. delta is a value that indicates differential encoding between the key frame. In other words, when derived_face_mode is delta, differential encoding between the key frame is applied as the encoding method of UV vertex identifiers. Furthermore, delta_in_face is a value that indicates differential encoding within the current frame. In other words, when derived_face_mode is delta_in_face, differential encoding within the current frame is applied as the encoding method of UV vertex identifiers. Furthermore, no_delta is a value that does not indicate differential encoding. In other words, when derived_face_mode is no_delta, differential encoding is not applied as the encoding method for UV vertex identifiers (intra encoding is applied). Hereinafter, derived_face_mode is also referred to as UV vertex identifier encoding control information or second UV vertex identifier encoding control information. For example, the first information processing device may set this derived_face_mode and apply differential encoding to UV vertex identifiers according to the derived_face_mode. Furthermore, the first information processing device may encode the derived_face_mode (transmit it to the decoding side). Note that the derived_face_mode can be used in conjunction with the derived_face_present_flag.
[0095] For example, in the first information processing apparatus, the encoding control unit may further generate second UV vertex identifier encoding control information that controls the encoding method of UV vertex identifiers. Then, when the first UV vertex identifier encoding control information is set to a value indicating that UV vertex identifiers may be differentially encoded and the second UV vertex identifier encoding control information is set to a value indicating differential encoding between the current frame and the key frame, the atlas information encoding unit may differentially encode the UV vertex identifiers between the current frame and the key frame. This allows the first information processing apparatus to easily control differential encoding of UV vertex identifiers.
[0096] For example, in the second information processing device, when UV vertex identifier encoding control information that controls the encoding method of UV vertex identifiers is set to a value indicating differential encoding between the current frame and the key frame, the atlas information decoding unit may differentially decode the UV vertex identifiers between the current frame and the key frame. In this way, the second information processing device can easily control the differential decoding of the UV vertex identifiers.
[0097] <UV Identifiers for Deriving Differences> Such UV vertex identifier differences are derived between corresponding UV vertices in the current frame and the key frame. For example, the UV vertices may be arranged in a predetermined order, and identifiers may be differentially encoded between UV vertices in the same order in the key frame. This makes it easy to encode and decode differences between corresponding UV vertices.
[0098] For example, in a first information processing device, the atlas information encoding unit may differentially encode UV vertex identifiers between vertices of a current frame and vertices of a key frame that are in the same order in a predetermined sort order of UV vertex identifiers, and in a second information processing device, the atlas information decoding unit may differentially decode UV vertex identifiers between vertices of a current frame and vertices of a key frame that are in the same order in a predetermined sort order of UV vertex identifiers.
[0099] <Method 1-2-1-2> When Method 1-2-1 is applied, as shown in the eighth row from the top of the table in FIG. 10 , if derived_face_mode is delta_in_face, UV vertex Idx between vertices in the current frame may be differentially encoded and transmitted (Method 1-2-1-2). For example, in a first information processing device, if the atlas information encoding unit sets the first UV vertex identifier encoding control information to a value indicating that UV vertex identifiers can be differentially encoded and the second UV vertex identifier encoding control information to a value indicating differential encoding within the current frame, the atlas information encoding unit may differentially encode UV vertex identifiers within the current frame. In this way, the first information processing device can easily control the differential encoding of UV vertex identifiers.
[0100] For example, in the second information processing device, when the UV vertex identifier encoding control information is set to a value indicating differential encoding within the current frame, the atlas information decoding unit may differentially decode the UV vertex identifiers within the current frame. In this way, the second information processing device can easily control the differential decoding of the UV vertex identifiers.
[0101] <UV Identifiers for Deriving Differences> Such UV vertex identifier differences may be derived between any UV vertices in the current frame. For example, the UV vertices may be arranged in a predetermined order, and identifiers may be differentially encoded between UV vertices in the same order in the key frame. This makes it easy to encode and decode differences between corresponding UV vertices.
[0102] For example, in a first information processing device, the atlas information encoding unit may differentially encode UV vertex identifiers between a vertex of a current processing target and a vertex of a previous processing target in a predetermined sort order of UV vertex identifiers in a current frame. Also, in a second information processing device, the atlas information decoding unit may differentially decode UV vertex identifiers between a vertex of a current processing target and a vertex of a previous processing target in a predetermined sort order of UV vertex identifiers in a current frame. Note that only the UV vertex at the top of this predetermined sort order may encode its identifier instead of the difference.
[0103] <Method 1-2-1-3> When Method 1-2-1 is applied, as shown in the bottom row of the table in FIG. 10 , if derived_uv_mode is no_delta, the UV vertex Idx of the current frame may be intra-coded and transmitted (Method 1-2-1-3). For example, in a first information processing device, an atlas information encoding unit may intra-code the UV vertex identifiers of the current frame when the second UV vertex identifier encoding control information is a value that does not indicate differential encoding. Also, in a second information processing device, an atlas information decoding unit may intra-decode the UV vertex identifiers when the UV vertex identifier encoding control information is a value that does not indicate differential encoding.
[0104] <Syntax> Fig. 15 shows an example of syntax when control is performed using the UV vertex identifier encoding control information described above. As shown in Fig. 15, in this example, when derived_face_present_flag is set to true, derived_face_mode is set, and when derived_face_mode is delta, differential encoding between the current frame and the key frame is applied to the UV vertex identifiers. When derived_face_mode is delta_in_face, differential encoding within the current frame is applied to the UV vertex identifiers. When derived_face_mode is no_delta, intra encoding is applied to the UV vertex identifiers.
[0105] <Decoding Process> An example of a pseudo program showing the state of the decoding process when control is performed using the UV vertex identifier encoding control information described above is shown in Figure 16. As shown in this example, when derived_face_mode is delta, differential decoding between the current frame and the key frame is applied to the UV vertex identifiers, when derived_face_mode is delta_in_face, differential decoding within the current frame is applied to the UV vertex identifiers, and when derived_face_mode is no_delta, intra decoding is applied to the UV vertex identifiers.
[0106] <Encoding Result> Fig. 17 is a diagram showing an example of an encoding result. The upper part of the figure shows a sequence of meshes to be encoded, and the lower part of the figure shows an example of the encoding result (applied encoding method). As shown by the arrow, the time direction progresses from left to right in the figure.
[0107] As shown in the upper part of Figure 17, in the frame with number (fid) 1220, the geometry vertex topology remains unchanged, but the UV vertex topology has changed. Therefore, as shown in the coding result A shown in the lower part of Figure 17, with the conventional method, inter coding could not be applied to the frame with number (fid) 1220, and intra coding was applied instead. In contrast, by applying the above-described present technology, it is possible to apply inter coding to the frame with number (fid) 1220, as shown in the coding result B shown in the lower part of Figure 17. Therefore, it is possible to suppress a decrease in coding efficiency.
[0108] That is, as in the encoding result B in this example, the bitstream may include, for frames in which the geometry vertex topology matches the keyframe and at least one of the UV vertex positions and the UV vertex topology changes relative to the keyframe, encoded data of the base mesh in which the geometry vertex positions are differentially encoded relative to the keyframe, and encoded data of the atlas information in which at least one of the UV vertex positions and the UV vertex identifiers are differentially encoded relative to the keyframe. By configuring the bitstream in this way, it is possible to suppress a decrease in encoding efficiency as described above.
[0109] In other words, the first information processing device may generate a bitstream configured as described above. That is, the first information processing device may include: a base mesh encoding unit that, for a current frame in which the geometry vertex topology matches the key frame and at least one of the UV vertex positions and the UV vertex topology changes relative to the key frame, differentially encodes the geometry vertex positions relative to the key frame to generate coded data of a base mesh, an atlas information encoding unit that differentially encodes at least one of the UV vertex positions and the UV vertex identifiers for the current frame to generate coded data of atlas information, and a bitstream generation unit that generates a bitstream including the coded data of the base mesh, the coded data of the displacement video, the coded data of the attribute video, and the coded data of the atlas information.
[0110] In addition, in the first information processing device, for a current frame in which the geometry vertex topology matches the key frame and at least one of the UV vertex positions and the UV vertex topology changes relative to the key frame, the geometry vertex positions are differentially encoded between the key frame to generate encoded data for the base mesh, and for the current frame, at least one of the UV vertex positions and the UV vertex identifiers are differentially encoded to generate encoded data for the atlas information, and a bitstream including the encoded data for the base mesh, the encoded data for the displacement video, the encoded data for the attribute video, and the encoded data for the atlas information may be generated.
[0111] A base mesh is a mesh with lower resolution than an original mesh to be coded, which is composed of vertices and connections that represent the three-dimensional structure of an object. It is generated by thinning out vertices from the original mesh. A displacement video is a moving image whose frame images are displacement maps, which are two-dimensional regions packed with displacement vectors indicating the displacement of division points, which are vertices generated by subdividing a base mesh. An attribute video is a moving image whose frame images are attribute maps, which are two-dimensional regions packed with projected images of the texture of the original mesh. Atlas information is information used for mesh reconstruction, including information indicating the correspondence between the base mesh, the displacement map, and the attribute map. Geometry vertex positions indicate the positions of the vertices of the base mesh in the three-dimensional region. Geometry vertex topology indicates how the vertices of the base mesh are connected in the three-dimensional region. UV vertex positions indicate the positions of the vertices of the base mesh in the attribute map. UV vertex topology indicates how the vertices of the base mesh are connected in the attribute map.
[0112] By doing so, the first information processing device can suppress a decrease in the coding efficiency of 3D data including the base mesh and the displacement vector, as described above.
[0113] <Combination> Each of the above-described methods may be applied in combination with any other method as long as no contradiction occurs. Three or more methods may be applied in combination. Furthermore, the techniques that can be combined are not limited to those shown in the table of FIG. 10 as "methods," but may include all of the elements described above in <3. Inter-coding of UV vertices>. Furthermore, each of the above-described methods may be applied in combination with methods other than those described above.
[0114] In this specification, a description of a higher-level method also applies to lower-level methods that belong to that higher-level method, unless a contradiction arises. For example, when it is described that "Method 1 may be applied," it is possible to apply Methods 1-1 and 1-2. Furthermore, when it is described that "Method 1 may be applied," it is possible to apply Methods 1-1 and 1-2, or Methods 1-1-1 and 1-1-2, or Method 1-2-1, or Methods 1-2-1-1, 1-2-1-2, and 1-2-1-3.
[0115] 4. First embodiment Encoding device The present technology can be applied to an encoding device that encodes a mesh. Fig. 18 is a block diagram showing an example of the configuration of an encoding device that is one aspect of an information processing device to which the present technology is applied. The encoding device 300 shown in Fig. 18 is a device that encodes a mesh.
[0116] Fig. 18 shows the main processing units, data flows, etc., but is not limited to all that is shown in Fig. 18. In other words, in encoding device 300, there may be processing units that are not shown as blocks in Fig. 18, and there may be processing and data flows that are not shown as arrows, etc. in Fig. 18.
[0117] The encoding device 300 encodes a mesh using a method essentially similar to the V-DMC method described in the aforementioned non-patent document, except that the present technology is applied. For example, the encoding device 300 acquires an original mesh to be encoded and an attribute map including a texture corresponding to the original mesh. Note that the original mesh includes not only information about the mesh's geometry but also information indicating the correspondence with the attribute map (e.g., a UV list). The encoding device 300 encodes the original mesh and attribute map using the V-DMC method, generates a V-DMC bitstream, and outputs it.
[0118] 18, the encoding device 300 includes a preprocessing unit 311 and an encoding unit 312. The preprocessing unit 311 performs preprocessing before encoding. As shown in FIG. 18, the preprocessing unit 311 includes a mesh decimation unit 321, an atlas information generation unit 322, an encoding control unit 323, and a displacement vector generation unit 324.
[0119] The mesh decimation unit 321 performs processing related to mesh decimation. For example, the mesh decimation unit 321 may obtain an original mesh to be input to the encoding device 300. The mesh decimation unit 321 may also perform decimation processing (thinning out vertices) on the original mesh to generate a base mesh. The mesh decimation unit 321 may supply the generated base mesh together with the original mesh to the atlas information generation unit 322.
[0120] The atlas information generation unit 322 performs processing related to the generation of atlas information corresponding to the base mesh. For example, the atlas information generation unit 322 may acquire the base mesh or original mesh supplied from the mesh decimation unit 321. Alternatively, the atlas information generation unit 322 may generate the atlas information by UV unwrapping the base mesh, for example. The atlas information generation unit 322 may supply the generated atlas information to the displacement vector generation unit 324 together with the base mesh, etc.
[0121] The encoding control unit 323 executes processing related to controlling the encoding of the base mesh (particularly, information about geometry vertex positions) and atlas information (particularly, information about UV vertex positions and UV vertex topology). For example, the encoding control unit 323 may acquire the base mesh and atlas information processed by the atlas information generation unit 322 and control the encoding method thereof. For example, the encoding control unit 323 may control whether inter-encoding or intra-encoding is applied to the base mesh (particularly, information about geometry vertex positions) and the atlas information (particularly, information about UV vertex positions and UV vertex topology). For example, the encoding control unit 323 may set the encoding method, generate control information (e.g., derived_uv_present_flag, derived_uv_mode, derived_face_present_flag, derived_face_mode, etc.), and supply it to the atlas information generation unit 322 as atlas information.
[0122] The displacement vector generation unit 324 performs processing related to the generation of displacement vectors. For example, the displacement vector generation unit 324 may acquire a base mesh, atlas information, etc. supplied from the atlas information generation unit 322. Alternatively, the displacement vector generation unit 324 may acquire an original mesh input to the encoding device 300. Alternatively, the displacement vector generation unit 324 may use the acquired information to generate displacement vectors that displace division points generated by subdividing the base mesh. The displacement vector generation unit 324 may supply the generated displacement vectors to the encoding unit 312 together with the base mesh (including vertex information), atlas information (including control information, UV maps, triangle information, etc.), etc.
[0123] The encoding unit 312 performs processing related to encoding of V-DMC data. For example, the encoding unit 312 may acquire an original mesh input to the encoding device 300. The encoding unit 312 may also acquire a base mesh (including vertex information), a displacement vector, atlas information (including control information, UV maps, triangle information, etc.) supplied from the displacement vector generation unit 324. The encoding unit 312 may also acquire an attribute map input to the encoding device 300. The encoding unit 312 may use this information to encode the atlas information, base mesh, displacement vector, and attribute map, respectively, and generate the respective encoded data. The encoding unit 312 may also multiplex these encoded data as substreams to generate a single bitstream. This bitstream is also referred to as a V-DMC bitstream. The encoding unit 312 may output the generated V-DMC bitstream to the outside of the encoding device 300.
[0124] <Encoding Unit> Fig. 19 is a block diagram showing an example of the main configuration of the encoding unit 312. Note that Fig. 19 shows the main processing units, data flows, etc., and does not necessarily show everything. In other words, the encoding unit 312 may include processing units that are not shown as blocks in Fig. 19, or processes or data flows that are not shown as arrows, etc. in Fig. 19.
[0125] As shown in FIG. 19, the encoding unit 312 has an atlas information encoding unit 351, a base mesh encoding unit 352, a displacement vector correction unit 353, a displacement video encoding unit 354, a mesh reconstruction unit 355, an attribute map conversion unit 356, an attribute video encoding unit 357, and a multiplexing unit 358.
[0126] The atlas information encoding unit 351 performs processing related to encoding of atlas information. For example, the atlas information encoding unit 351 may acquire atlas information supplied from the displacement vector generation unit 324. This atlas information may include, for example, control information, UV maps, triangle information, and the like. The atlas information encoding unit 351 may also encode the acquired atlas information using a predetermined encoding method to generate encoded data of the atlas information. The atlas information encoding unit 351 may also supply the generated encoded data of the atlas information to the multiplexing unit 358.
[0127] The base mesh encoding unit 352 performs processing related to encoding of the base mesh. For example, the base mesh encoding unit 352 may acquire the base mesh supplied from the displacement vector generation unit 324. The base mesh encoding unit 352 may quantize the acquired base mesh and encode it using a predetermined encoding method (e.g., Draco) to generate encoded data of the base mesh. The base mesh encoding unit 352 may encode the base mesh according to control information included in the atlas information. The base mesh encoding unit 352 may supply the generated encoded data of the base mesh to the displacement vector correction unit 353. The base mesh encoding unit 352 may supply the generated encoded data of the base mesh to the multiplexing unit 358.
[0128] The displacement vector correction unit 353 performs processing related to the correction of the displacement vector. For example, the displacement vector correction unit 353 may acquire a base mesh and a displacement vector supplied from the displacement vector generation unit 324. Alternatively, the displacement vector correction unit 353 may acquire encoded data of the base mesh supplied from the base mesh encoding unit 352. The displacement vector correction unit 353 may correct the displacement vector based on this information. For example, the displacement vector correction unit 353 may decode the acquired encoded data of the base mesh, compare the base mesh before and after encoding to determine encoding distortion of the base mesh, and correct the displacement vector in accordance with the encoding distortion. The displacement vector correction unit 353 may supply the corrected displacement vector to the displacement video encoding unit 354. Alternatively, the displacement vector correction unit 353 may dequantize the decoded base mesh and supply it to the mesh reconstruction unit 355.
[0129] The displacement video encoding unit 354 performs processing related to encoding of the displacement video. The displacement video is a moving image having a displacement map, which is a two-dimensional area in which displacement vectors are packed, as frame images. For example, the displacement video encoding unit 354 may acquire displacement vectors supplied from the displacement vector correction unit 353. The displacement video encoding unit 354 may also generate a displacement map by wavelet transforming the displacement vectors, quantizing them, and packing them into a two-dimensional area. The displacement video encoding unit 354 may also generate a displacement video having the displacement map as frame images. The displacement video encoding unit 354 may also encode the generated displacement video using a predetermined encoding method for 2D moving images to generate encoded data of the displacement video. The displacement video encoding unit 354 may also supply the generated encoded data of the displacement video to the multiplexing unit 358. The displacement video encoding unit 354 may also decode the generated encoded data of the displacement video, unpack the displacement vectors from the displacement map, and dequantize the displacement vectors. The displacement video encoder 354 may provide the dequantized displacement vectors to the mesh reconstructor 355 .
[0130] The mesh reconstruction unit 355 performs processing related to mesh reconstruction. For example, the mesh reconstruction unit 355 may acquire a base mesh supplied from the displacement vector correction unit 353. The mesh reconstruction unit 355 may also acquire a displacement vector supplied from the displacement video encoding unit 354. The mesh reconstruction unit 355 may reconstruct a mesh using the acquired displacement vectors. The mesh reconstruction unit 355 may supply the reconstructed mesh to the attribute map conversion unit 356.
[0131] The attribute map conversion unit 356 performs processing related to attribute map conversion. For example, the attribute map conversion unit 356 may acquire a reconstructed mesh supplied from the mesh reconstruction unit 355. Alternatively, the attribute map conversion unit 356 may acquire atlas information supplied from the displacement vector generation unit 324. Alternatively, the attribute map conversion unit 356 may acquire an original mesh and an attribute map input to the encoding device 300. The attribute map conversion unit 356 may convert the acquired attribute map based on other acquired information. For example, the attribute map conversion unit 356 may convert the attribute map based on the atlas information, the original mesh, or the like so that it corresponds to the reconstructed mesh. In other words, the attribute map conversion unit 356 can be said to generate a converted attribute map. Therefore, the attribute map conversion unit 356 can also be said to be an attribute map generation unit. The attribute map conversion unit 356 may supply the converted attribute map to the attribute video encoding unit 357 .
[0132] The attribute video encoding unit 357 performs processing related to encoding of the attribute video. For example, the attribute video encoding unit 357 may acquire an attribute map supplied from the attribute map conversion unit 356. The attribute video encoding unit 357 may also generate attribute video using the acquired attribute map as frame images. The attribute video encoding unit 357 may also encode the generated attribute video using a predetermined encoding method for 2D moving images to generate encoded data of the attribute video. The attribute video encoding unit 357 may also supply the generated encoded data of the attribute video to the multiplexing unit 358.
[0133] The multiplexing unit 358 performs processing related to multiplexing of encoded data (substreams). For example, the multiplexing unit 358 may acquire encoded data of atlas information supplied from the atlas information encoding unit 351. Alternatively, the multiplexing unit 358 may acquire encoded data of base meshes supplied from the base mesh encoding unit 352. Alternatively, the multiplexing unit 358 may acquire encoded data of displacement video supplied from the displacement video encoding unit 354. Alternatively, the multiplexing unit 358 may acquire encoded data of attribute video supplied from the attribute video encoding unit 357. The multiplexing unit 358 may multiplex these encoded data as substreams to generate a V-DMC bitstream. Therefore, the multiplexing unit 358 can also be referred to as a bitstream generation unit. The multiplexing unit 358 may output the generated V-DMC bitstream to an external device from the encoding device 300. For example, the multiplexing unit 358 may supply the V-DMC bitstream to a decoding device 400, which will be described later. Therefore, the multiplexing unit 358 can also be said to be a supply unit (providing unit) of the V-DMC bitstream.
[0134] In the encoding device 300 configured as above, the various methods (the present technology) described above in <3. Inter-encoding of UV vertices> may be applied.
[0135] For example, in the encoding device 300 (first information processing device), the atlas information encoding unit 351 may differentially encode UV vertex positions between a current frame and a key frame, and generate encoded data of atlas information including the differentially encoded UV vertex positions. The multiplexing unit 358 may generate a bitstream including encoded data of a base mesh, encoded data of a displacement video, encoded data of an attribute video, and encoded data of atlas information. Furthermore, the atlas information encoding unit 351 may differentially encode UV vertex positions between vertices of the current frame and vertices of a key frame that are in the same order in a predetermined sorting order of UV vertex positions. The key frame may be the frame immediately preceding the current frame. The base mesh encoding unit 352 may differentially encode geometry vertex positions that indicate the positions of the vertices of the base mesh in a three-dimensional domain.
[0136] The encoding control unit 323 may also generate first UV vertex position encoding control information that controls whether to differentially encode UV vertex positions. The atlas information encoding unit 351 may differentially encode UV vertex positions according to the first UV vertex position encoding control information, encode atlas information including the first UV vertex position encoding control information, and generate encoded data for the atlas information. The encoding control unit 323 may set the first UV vertex position encoding control information for an area where the number of mesh faces between the current frame and the key frame is the same to a value indicating that UV vertex positions can be differentially encoded. If the UV vertex positions between the current frame and the key frame are different, the encoding control unit 323 may set the first UV vertex position encoding control information to a value indicating that UV vertex positions can be differentially encoded. If the first UV vertex position encoding control information is a value indicating that UV vertex positions should not be differentially encoded, the atlas information encoding unit 351 may intra-code the UV vertex positions of the current frame.
[0137] The encoding control unit 323 may further generate second UV vertex position encoding control information that controls the encoding method of the UV vertex positions. When the first UV vertex position encoding control information is set to a value indicating that the UV vertex positions can be differentially encoded and the second UV vertex position encoding control information is set to a value indicating differential encoding, the atlas information encoding unit 351 may differentially encode the UV vertex positions between the current frame and the key frame. When the number of vertices differs between the current frame and the key frame, the encoding control unit 323 may prohibit setting the second UV vertex position encoding control information to a value indicating differential encoding.
[0138] Furthermore, if the second UV vertex position encoding control information is a value that does not indicate differential encoding, the atlas information encoding unit 351 may intra-encode the UV vertex positions of the current frame.
[0139] In addition, the atlas information encoding unit 351 may differentially encode UV vertex identifiers that identify the vertices of the base mesh in the attribute map, and generate encoded data of the atlas information that includes the differentially encoded UV vertex identifiers.
[0140] The encoding control unit 323 may also generate first UV vertex identifier encoding control information that controls whether to differentially encode UV vertex identifiers. The atlas information encoding unit 351 may differentially encode UV vertex identifiers according to the first UV vertex identifier encoding control information, encode atlas information including the first UV vertex identifier encoding control information, and generate encoded data for the atlas information. The encoding control unit 323 may set the first UV vertex identifier encoding control information for a region where the number of mesh faces between the current frame and the key frame is the same to a value indicating that UV vertex identifiers can be differentially encoded. If the UV vertex identifiers between the current frame and the key frame are different, the encoding control unit 323 may set the first UV vertex identifier encoding control information to a value indicating that UV vertex identifiers can be differentially encoded. If the first UV vertex identifier encoding control information is a value indicating that UV vertex identifiers are not differentially encoded, the atlas information encoding unit may intra-code the UV vertex identifiers of the current frame.
[0141] The encoding control unit 323 may further generate second UV vertex identifier encoding control information that controls the encoding method of UV vertex identifiers. When the first UV vertex identifier encoding control information is set to a value indicating that UV vertex identifiers may be differentially encoded and the second UV vertex identifier encoding control information is set to a value indicating differential encoding between the current frame and the key frame, the atlas information encoding unit 351 may differentially encode UV vertex identifiers between the current frame and the key frame. The atlas information encoding unit 351 may differentially encode UV vertex identifiers between vertices of the current frame and vertices of the key frame that are in the same order as each other in a predetermined sort order of UV vertex identifiers.
[0142] Furthermore, when the first UV vertex identifier encoding control information is set to a value indicating that UV vertex identifiers can be differentially encoded and the second UV vertex identifier encoding control information is set to a value indicating differential encoding within the current frame, the atlas information encoding unit 351 may differentially encode UV vertex identifiers within the current frame. The atlas information encoding unit 351 may differentially encode UV vertex identifiers between the vertex of the current processing target and the vertex of the previous processing target in a predetermined sort order of UV vertex identifiers within the current frame.
[0143] Furthermore, if the second UV vertex identifier encoding control information is a value that does not indicate differential encoding, the atlas information encoding unit 351 may intra-encode the UV vertex identifiers of the current frame.
[0144] Furthermore, in the encoding device 300 (first information processing device), the base mesh encoding unit 352 may differentially encode the geometry vertex positions of a current frame, in which the geometry vertex topology matches the key frame and at least one of the UV vertex positions and the UV vertex topology changes relative to the key frame, to generate encoded data of the base mesh. The atlas information encoding unit 351 may differentially encode at least one of the UV vertex positions and the UV vertex identifiers of the current frame to generate encoded data of the atlas information. The multiplexing unit 358 may generate a bitstream including the encoded data of the base mesh, the encoded data of the displacement video, the encoded data of the attribute video, and the encoded data of the atlas information.
[0145] In addition, the bitstream generated by the encoding device 300 (first information processing device) may include encoded data of a base mesh in which the geometry vertex positions are differentially encoded with respect to the key frame for frames in which the geometry vertex topology matches the key frame and at least one of the UV vertex positions and UV vertex topology changes relative to the key frame, and encoded data of atlas information in which at least one of the UV vertex positions and UV vertex identifiers are differentially encoded with respect to the key frame.
[0146] By doing so, the encoding device 300 can obtain the same effect as that described above in <3. Inter-coding of UV vertices>. In other words, the encoding device 300 can suppress a decrease in the encoding efficiency of 3D data including a base mesh and a displacement vector.
[0147] <Flow of Encoding Process> An example of the flow of the encoding process executed by the encoding device 300 will be described with reference to the flowchart of FIG.
[0148] When the encoding process starts, in step S301, the mesh decimation unit 321 of the encoding device 300 decimates the original mesh to be encoded to generate a base mesh.
[0149] In step S302, the atlas information generating unit 322 generates atlas information for the base mesh.
[0150] In step S303, the encoding control unit 323 sets an encoding method for the geometry (i.e., the base mesh). For example, the encoding control unit 323 sets whether to apply inter-coding or intra-coding to the geometry and generates control information for that. In step S304, the encoding control unit 323 sets an encoding method for the UV vertex positions. For example, the encoding control unit 323 sets whether to apply inter-coding or intra-coding to the UV vertex positions and generates control information for that. In step S305, the encoding control unit 323 sets an encoding method for the UV vertex topology (UV vertex identifiers). For example, the encoding control unit 323 sets whether to apply inter-coding or intra-coding to the UV vertex topology (UV vertex identifiers) and generates control information for that.
[0151] In step S306, the displacement vector generating unit 324 generates a displacement vector.
[0152] In step S307, the encoding unit 312 performs V-DMC encoding processing on the V-DMC data (atlas information, base mesh, displacement vector, and attribute map) generated as described above, and generates a V-DMC bitstream.
[0153] In step S308, the encoding unit 312 determines whether all frames have been processed. If unprocessed frames remain, the process returns to step S301, and the subsequent processes are executed. That is, the processes from step S301 to step S308 are executed for each frame in the mesh sequence. Then, if it is determined in step S308 that all frames have been processed (that is, if there are no more frames to be processed), the encoding process ends.
[0154] <Flow of V-DMC Encoding Process> Next, an example of the flow of the V-DMC encoding process executed in step S307 of FIG. 20 will be described with reference to the flowchart of FIG.
[0155] When the V-DMC encoding process starts, in step S321, the atlas information encoding unit 351 encodes control information (e.g., derived_uv_present_flag, derived_uv_mode, derived_face_present_flag, derived_face_mode, etc.) as atlas information. In step S322, the atlas information encoding unit 351 encodes UV vertex positions as atlas information in accordance with the control information. In step S323, the atlas information encoding unit 351 encodes UV vertex topology (UV vertex identifiers) as atlas information in accordance with the control information. In step S324, the atlas information encoding unit 351 encodes other atlas information. For example, the atlas information encoding unit 351 may differentially encode UV vertex positions between the current frame and the key frame, and generate encoded data of atlas information including the differentially encoded UV vertex positions. In addition, the atlas information encoding unit 351 may differentially encode at least one of the UV vertex positions and UV vertex identifiers for a current frame in which the geometry vertex topology matches the key frame and at least one of the UV vertex positions and UV vertex topology changes relative to the key frame, thereby generating encoded data for the atlas information.
[0156] In step S325, the base mesh encoding unit 352 encodes the base mesh in accordance with the control information. For example, for a current frame in which the geometry vertex topology matches that of a key frame and at least one of the UV vertex positions and the UV vertex topology changes relative to the key frame, the base mesh encoding unit 352 may differentially encode the geometry vertex positions relative to the key frame to generate encoded data for the base mesh.
[0157] In step S326, the displacement vector correction unit 353 corrects the displacement vector.
[0158] In step S327, the displaced video encoding unit 354 encodes a displaced video in which the displaced map in which the corrected displaced vectors are packed is used as a frame image.
[0159] In step S328, the mesh reconstructing unit 355 reconstructs the mesh.
[0160] In step S329, the attribute map conversion unit 356 converts the attribute map.
[0161] In step S330, the attribute video encoding unit 357 encodes the attribute video using the attribute map as a frame image.
[0162] In step S331, the multiplexing unit 358 multiplexes the coded data of the atlas information, the coded data of the bitstream, the coded data of the displacement video, and the coded data of the attribute video to generate a V-DMC bitstream.
[0163] When the process of step S331 ends, the encoding process of FIG. 21 ends, and the process returns to the process of FIG.
[0164] By performing the above-described processes, the encoding device 300 can achieve the same effect as described above in <3. Inter-coding of UV vertices>. In other words, the encoding device 300 can suppress a decrease in the encoding efficiency of 3D data including a base mesh and a displacement vector.
[0165] 22 is a block diagram showing an example of the configuration of a decoding device, which is one aspect of an information processing device to which the present technology is applied. The decoding device 400 shown in Fig. 22 is a device that decodes, for example, encoded data of meshes generated in the encoding device 300 (V-DMC bitstream generated by the multiplexing unit 358).
[0166] Fig. 22 shows the main processing units, data flows, etc., but does not necessarily include all of them. In other words, in the decoding device 400, there may be processing units that are not shown as blocks in Fig. 22, and there may be processing and data flows that are not shown as arrows, etc. in Fig. 22.
[0167] The decoding device 400 decodes coded data of a mesh that has been coded using a method essentially similar to the V-DMC method described in the aforementioned non-patent document, except that the present technology is applied. For example, the decoding device 400 obtains a V-DMC bitstream. This V-DMC bitstream may be generated by, for example, the coding device 300. The decoding device 400 decodes the V-DMC bitstream, reconstructs a mesh (also referred to as a decoded mesh), applies a texture to the decoded mesh, generates a display image for displaying the decoded mesh, and outputs the display image to an external device.
[0168] As shown in FIG. 22, the decoding device 400 has a demultiplexing unit 411, an atlas information decoding unit 412, a base mesh decoding unit 413, a base mesh reconstruction unit 414, a subdivision unit 415, a displacement video decoding unit 416, an unpacking unit 417, a displacement vector application unit 418, an attribute video decoding unit 419, and a display processing unit 420.
[0169] The demultiplexing unit 411 performs demultiplexing processing. For example, the demultiplexing unit 411 may acquire a V-DMC bitstream to be decoded and supplied to the decoding device 400. The demultiplexing unit 411 may also demultiplex the acquired V-DMC bitstream to extract coded data of atlas information, coded data of base meshes, coded data of displacement video, and coded data of attribute video. Therefore, the demultiplexing unit 411 can also be considered an acquisition unit for the V-DMC bitstream or various information contained in the V-DMC bitstream. The demultiplexing unit 411 may supply the extracted coded data of atlas information to the atlas information decoding unit 412. The demultiplexing unit 411 may also supply the extracted coded data of base meshes to the base mesh decoding unit 413. The demultiplexing unit 411 may also supply the extracted coded data of displacement video to the displacement video decoding unit 416. Furthermore, the demultiplexing unit 411 may supply the extracted coded data of the attribute video to the attribute video decoding unit 419 .
[0170] The atlas information decoding unit 412 performs processing related to decoding of coded data of atlas information. For example, the atlas information decoding unit 412 may acquire coded data of atlas information supplied from the demultiplexing unit 411. The atlas information decoding unit 412 may also decode the acquired coded data of atlas information to generate (restore) atlas information. For example, the atlas information decoding unit 412 may decode coded data of control information (e.g., derived_uv_present_flag, derived_uv_mode, derived_face_present_flag, derived_face_mode, etc.). The atlas information decoding unit 412 may also decode coded data of UV vertex positions according to the control information. The atlas information decoding unit 412 may also decode coded data of UV vertex identifiers according to the control information. The atlas information decoding unit 412 may supply the generated atlas information to the base mesh reconstruction unit 414 and the display processing unit 420. The atlas information decoding unit 412 may also supply the control information to the base mesh decoding unit 413.
[0171] The base mesh decoding unit 413 performs processing related to decoding of the coded data of the base mesh. For example, the base mesh decoding unit 413 may acquire the coded data of the base mesh supplied from the demultiplexing unit 411. The base mesh decoding unit 413 may also decode the acquired coded data of the base mesh using a predetermined decoding method (e.g., Draco) to generate (restore) information about the base mesh (e.g., a vertex list, a triangle list, etc.). For example, the base mesh decoding unit 413 may decode coded data of geometry vertex positions in accordance with control information supplied from the atlas information decoding unit 412. The base mesh decoding unit 413 may supply the generated information about the base mesh to the base mesh reconstruction unit 414.
[0172] The base mesh reconstruction unit 414 executes processing related to the reconstruction of the base mesh. For example, the base mesh reconstruction unit 414 may acquire information about the base mesh supplied from the base mesh decoding unit 413. Alternatively, the base mesh reconstruction unit 414 may acquire atlas information supplied from the atlas information decoding unit 412. The base mesh reconstruction unit 414 may reconstruct the base mesh using the acquired atlas information and information about the base mesh. The base mesh reconstruction unit 414 may supply the generated base mesh to the subdivision unit 415.
[0173] The subdivision unit 415 performs processing related to subdivision of triangles of a base mesh. For example, the subdivision unit 415 may obtain a base mesh provided from the base mesh reconstruction unit 414. The subdivision unit 415 may subdivide triangles of the base mesh to generate division points. The subdivision unit 415 may provide the subdivided base mesh to the displacement vector application unit 418.
[0174] The displacement video decoding unit 416 performs processing related to decoding of encoded data of the displacement video. For example, the displacement video decoding unit 416 may acquire encoded data of the displacement video supplied from the demultiplexing unit 411. Furthermore, the displacement video decoding unit 416 may decode the acquired encoded data of the displacement video using a predetermined decoding method for 2D moving images to generate (restore) the displacement video. The displacement video decoding unit 416 may supply the generated displacement video to the unpacking unit 417.
[0175] The unpacking unit 417 performs processing related to unpacking of the displacement vector. For example, the unpacking unit 417 may acquire the displacement video supplied from the displacement video decoding unit 416. The unpacking unit 417 may also unpack the displacement vector from a displacement map, which is a frame image of the displacement video. The unpacking unit 417 may supply the displacement vector obtained in this way to the displacement vector application unit 418.
[0176] The displacement vector application unit 418 performs processing related to application of a displacement vector to a subdivided base mesh. For example, the displacement vector application unit 418 may obtain the subdivided base mesh supplied from the subdivision unit 415. Alternatively, the displacement vector application unit 418 may obtain a displacement vector supplied from the unpacking unit 417. Alternatively, the displacement vector application unit 418 may apply a displacement vector to a division point of the subdivided base mesh. In other words, the displacement vector application unit 418 may generate a decoded mesh. The displacement vector application unit 418 may supply the generated decoded mesh to the display processing unit 420.
[0177] The attribute video decoding unit 419 performs processing related to decoding of the attribute video. For example, the attribute video decoding unit 419 may acquire coded data of the attribute video supplied from the demultiplexing unit 411. The attribute video decoding unit 419 may also decode the acquired coded data of the attribute video using a predetermined decoding method for 2D moving images to generate (restore) the attribute video. The attribute video decoding unit 419 may supply an attribute map, which is a frame image of the generated attribute video, to the display processing unit 420.
[0178] The display processing unit 420 performs processing related to mesh display. For example, the display processing unit 420 may acquire atlas information supplied from the atlas information decoding unit 412. The display processing unit 420 may also acquire a decoded mesh supplied from the displacement vector application unit 418. The display processing unit 420 may also acquire an attribute map supplied from the attribute video decoding unit 419. The display processing unit 420 may use the acquired atlas information to extract a texture from the attribute map and apply it to a face of the decoded mesh corresponding to the texture. In other words, the display processing unit 420 may attach the texture to the face. The display processing unit 420 may render the decoded mesh to which the texture has been applied and generate a display image for displaying the decoded mesh to which the texture has been applied. The display processing unit 420 may then supply the generated display image to an external device outside the decoding device 400, causing the display image to be displayed by another device or the like.
[0179] In the decoding device 400 configured as above, the various methods (the present technology) described above in <3. Inter-coding of UV vertices> may be applied.
[0180] For example, in the decoding device 400 (second information processing device), the demultiplexing unit 411 may demultiplex the bitstream and extract coded data of the base mesh, coded data of the displacement video, coded data of the attribute video, and coded data of the atlas information. The atlas information decoding unit 412 may differentially decode UV vertex positions differentially coded between the current frame and the key frame, which are included in the extracted coded data of the atlas information. The atlas information decoding unit 412 may differentially decode UV vertex positions between the vertices of the current frame and the vertices of the key frame that are in the same order in a predetermined sorting order of the UV vertex positions. The key frame may be the frame immediately preceding the current frame. The base mesh decoding unit 413 may differentially decode geometry vertex positions that indicate the positions of the vertices of the base mesh in a three-dimensional region, which have been differentially coded.
[0181] In addition, if the UV vertex position encoding control information that controls the encoding method of the UV vertex positions is set to a value indicating differential encoding, the atlas information decoding unit 412 may differentially decode the UV vertex positions between the current frame and the key frame.
[0182] Furthermore, if the UV vertex position encoding control information is a value that does not indicate differential encoding, the atlas information decoding unit 412 may intra-decode the UV vertex positions of the current frame.
[0183] Furthermore, the atlas information decoding unit 412 may differentially decode the differentially encoded UV vertex identifiers that identify the vertices of the base mesh in the attribute map.
[0184] Furthermore, when UV vertex identifier encoding control information that controls the encoding method of UV vertex identifiers is set to a value indicating differential encoding between the current frame and the key frame, the atlas information decoding unit 412 may differentially decode UV vertex identifiers between the current frame and the key frame.The atlas information decoding unit 412 may differentially decode UV vertex identifiers between vertices of the current frame and vertices of the key frame that are in the same order in a predetermined sorting order of UV vertex identifiers.
[0185] Furthermore, when the UV vertex identifier encoding control information is set to a value indicating differential encoding within the current frame, the atlas information decoding unit 412 may differentially decode UV vertex identifiers within the current frame. The atlas information decoding unit 412 may differentially decode UV vertex identifiers between the vertex of the current processing target and the vertex of the previous processing target in a predetermined sort order of UV vertex identifiers within the current frame.
[0186] Furthermore, if the UV vertex identifier encoding control information is a value that does not indicate differential encoding, the atlas information decoding unit 412 may intra-decode the UV vertex identifiers.
[0187] By doing so, the decoding device 400 can obtain the same effect as that described above in <3. Inter-coding of UV vertices>. In other words, the decoding device 400 can suppress a decrease in the coding efficiency of 3D data including a base mesh and a displacement vector.
[0188] <Flow of Decoding Process> An example of the flow of the decoding process executed by the decoding device 400 will be described with reference to the flowchart of FIG.
[0189] When the decoding process starts, the demultiplexing unit 411 of the decoding device 400 demultiplexes the V-DMC bitstream in step S401.
[0190] In step S402, the atlas information decoding unit 412 decodes coded data of control information (e.g., derived_uv_present_flag, derived_uv_mode, derived_face_present_flag, derived_face_mode, etc.) as atlas information to generate (restore) the control information. In step S403, the atlas information decoding unit 412 decodes coded data of UV vertex positions as atlas information in accordance with the control information to generate (restore) UV vertex positions. In step S404, the atlas information decoding unit 412 decodes coded data of UV vertex topology (UV vertex identifiers) as atlas information in accordance with the control information to generate (restore) UV vertex topology (UV vertex identifiers). In step S405, the atlas information decoding unit 412 decodes coded data of other atlas information to generate (restore) other atlas information.
[0191] In step S406, the base mesh decoding unit 413 decodes the coded data of the base mesh in accordance with the control information, and generates (restores) information about the base mesh.
[0192] In step S407, the base mesh reconstructing unit 414 reconstructs the base mesh using the atlas information and information related to the base mesh.
[0193] In step S408, the subdivision unit 415 subdivides the reconstructed base mesh to generate division points.
[0194] In step S409, the displaced video decoding unit 416 decodes the coded data of the displaced video to generate (restore) the displaced video.
[0195] In step S410, the unpacking unit 417 unpacks the displacement map, which is the frame images of the displacement video, and extracts the displacement vectors.
[0196] In step S411, the displacement vector application unit 418 applies the unpacked displacement vectors to the division points of the subdivided base mesh to generate a decoded mesh.
[0197] In step S412, the attribute video decoding unit 419 decodes the coded data of the attribute video to generate (restore) the attribute video.
[0198] In step S413, the display processing unit 420 uses the atlas information to apply a texture extracted from an attribute map, which is a frame image of the attribute video, to the face of the generated decoded mesh, and renders the decoded mesh to generate a display image.
[0199] When the process of step S413 ends, the decoding process ends.
[0200] By performing the above-described processes, the decoding device 400 can achieve the same effect as that described above in <3. Inter-coding of UV vertices>. In other words, the decoding device 400 can suppress a decrease in the coding efficiency of 3D data including a base mesh and a displacement vector.
[0201] 6. Supplementary Notes Polygon Shape In the above description, the polygon shape is a triangle, but this shape is just an example. The polygon shape may be any polygonal shape.
[0202] <Scope of Application of the Present Technology> In the above, V-DMC has been used as an example of an encoding method to which the present technology is applied, but the present technology is not limited to this example. In other words, the standard and specifications of the encoding method to which the present technology is applied may be any as long as they do not contradict the above description. The present technology can be applied to any encoding method as long as it encodes a base mesh, a displacement vector, an attribute map including texture, atlas information, or information equivalent thereto. The same applies to decoding. Furthermore, the specifications and names of the 3D data (mesh data or V-DMC data) to be encoded may also be any as long as they do not contradict the above description. For example, mesh data to be encoded may be generated by converting other 3D data such as a point cloud.
[0203] <Computer> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs that make up the software are installed on a computer. Here, the term "computer" includes computers built into dedicated hardware, and general-purpose personal computers, etc., that can execute various functions by installing various programs.
[0204] FIG. 24 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.
[0205] In a computer 900 shown in FIG. 24, a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, and a RAM (Random Access Memory) 903 are interconnected via a bus 904.
[0206] An input / output interface 910 is also connected to the bus 904. To the input / output interface 910, an input unit 911, an output unit 912, a storage unit 913, a communication unit 914, and a drive 915 are connected.
[0207] The input unit 911 includes, for example, a keyboard, a mouse, a microphone, a touch panel, and an input terminal. The output unit 912 includes, for example, a display, a speaker, and an output terminal. The storage unit 913 includes, for example, a hard disk, a RAM disk, and a non-volatile memory. The communication unit 914 includes, for example, a network interface. The drive 915 drives removable media 921 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0208] In a computer configured as described above, the CPU 901 performs the above-described series of processes by, for example, loading a program stored in the storage unit 913 into the RAM 903 via the input / output interface 910 and the bus 904 and executing the program. The RAM 903 also stores data necessary for the CPU 901 to execute various processes as appropriate.
[0209] The program executed by the computer can be applied by recording it on, for example, a removable medium 921 such as a package medium. In this case, the program can be installed in the storage unit 913 via the input / output interface 910 by inserting the removable medium 921 into the drive 915.
[0210] This program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, digital satellite broadcasting, etc. In this case, the program can be received by the communication unit 914 and installed in the storage unit 913.
[0211] Alternatively, this program can be installed in advance in the ROM 902 or the storage unit 913 .
[0212] <Application of the Present Technology> The present technology can be applied to any configuration. For example, the present technology can be applied to various electronic devices.
[0213] Furthermore, for example, the present technology can also be implemented as part of an apparatus, such as a processor (e.g., a video processor) as a system LSI (Large Scale Integration), a module using multiple processors (e.g., a video module), a unit using multiple modules (e.g., a video unit), or a set in which other functions are added to a unit (e.g., a video set).
[0214] Furthermore, for example, the present technology can also be applied to a network system configured with multiple devices. For example, the present technology may be implemented as cloud computing in which multiple devices share and collaborate on processing via a network. For example, the present technology may be implemented in a cloud service that provides image (video)-related services to any terminal, such as a computer, an AV (Audio Visual) device, a portable information processing terminal, or an IoT (Internet of Things) device.
[0215] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0216] <Fields and uses to which this technology can be applied> Systems, devices, processing units, etc. to which this technology is applied can be used in any field, for example, transportation, medical care, crime prevention, agriculture, livestock farming, mining, beauty, factories, home appliances, weather, nature monitoring, etc. In addition, the uses thereof are also arbitrary.
[0217] <Others> In this specification, a "flag" refers to information for identifying multiple states, and includes not only information used to identify two states, true (1) or false (0), but also information capable of identifying three or more states. Therefore, the value that this "flag" can take may be, for example, two values, 1 / 0, or three or more values. That is, the number of bits constituting this "flag" is arbitrary, and may be one bit or multiple bits. Furthermore, identification information (including flags) can be included not only in a bitstream, but also in a bitstream that includes differential information of the identification information relative to certain reference information. Therefore, in this specification, "flag" and "identification information" encompass not only the information itself, but also differential information relative to the reference information.
[0218] Furthermore, various types of information (e.g., metadata) related to the coded data (bitstream) may be transmitted or recorded in any form as long as they are associated with the coded data. Here, the term "associate" means, for example, that one piece of data can be used (linked) when processing the other piece of data. That is, the associated pieces of data may be combined into one piece of data or may be separate pieces of data. For example, information associated with coded data (image) may be transmitted over a transmission path separate from that of the coded data (image). Furthermore, for example, information associated with coded data (image) may be recorded on a recording medium separate from that of the coded data (image) (or on a different recording area of the same recording medium). Note that this "association" may refer not to the entire data, but to only part of the data. For example, an image and information corresponding to that image may be associated with each other in any unit, such as multiple frames, one frame, or a portion of a frame.
[0219] In this specification, terms such as "composite," "multiplex," "add," "integrate," "include," "store," "embed," "insert," and the like refer to combining multiple items into one, such as combining encoded data and metadata into one piece of data, and refer to one method of "associating" as described above.
[0220] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.
[0221] For example, a configuration described as one device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, configurations described above as multiple devices (or processing units) may be combined and configured as one device (or processing unit). Of course, configurations other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, as long as the configuration and operation of the entire system are substantially the same, part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).
[0222] Furthermore, for example, the above-described program may be executed in any device, as long as the device has the necessary functions (functional blocks, etc.) and is able to obtain the necessary information.
[0223] Also, for example, each step of a single flowchart may be executed by a single device, or may be shared and executed by multiple devices. Furthermore, when a single step includes multiple processes, the multiple processes may be executed by a single device, or may be shared and executed by multiple devices. In other words, multiple processes included in a single step can be executed as multiple step processes. Conversely, processes described as multiple steps can be executed collectively as a single step.
[0224] For example, the steps of a program executed by a computer may be executed in chronological order in the order described herein, or may be executed in parallel or individually at the required timing, such as when a call is made. In other words, as long as no contradiction occurs, the steps may be executed in an order different from the order described above. Furthermore, the steps of this program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.
[0225] Furthermore, for example, multiple technologies related to the present technology can be implemented independently and independently, as long as no contradiction occurs. Of course, any multiple technologies can also be implemented in combination. For example, part or all of the present technology described in any embodiment can be implemented in combination with part or all of the present technology described in another embodiment. Furthermore, part or all of any of the above-described present technologies can be implemented in combination with other technologies not described above.
[0226] The present technology can also be configured as follows: (1) An atlas information encoding unit that differentially encodes UV vertex positions between a current frame and a key frame, and generates encoded data of atlas information including the differentially encoded UV vertex positions, and a bitstream generating unit that generates a bitstream including encoded data of a base mesh, encoded data of a displacement video, encoded data of an attribute video, and encoded data of the atlas information, wherein the base mesh is a mesh with lower resolution than an original mesh to be encoded, which is composed of vertices and connections that represent a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh, the displacement video is a video having frame images of a displacement map that is a two-dimensional region packed with displacement vectors that indicate the displacement of division points that are vertices generated by subdividing the base mesh, and the attribute video is a video having frame images of an attribute map that is a two-dimensional region packed with projected images of the texture of the original mesh, and the atlas information is information used for mesh reconstruction, including information indicating a correspondence between the base mesh, the displacement map, and the attribute map, The information processing device according to (1), wherein the UV vertex positions are position information indicating positions of vertices of the base mesh in the attribute map. (2) The information processing device according to (1), wherein the atlas information encoding unit is configured to differentially encode the UV vertex positions between vertices of the current frame and vertices of the key frame that are in the same order in a predetermined sorting order of the UV vertex positions. (3) The information processing device according to (2), wherein the key frame is the frame immediately preceding the current frame. (4) The information processing device according to (3), further comprising a base mesh encoding unit that differentially encodes geometry vertex positions that indicate positions of the vertices of the base mesh in a three-dimensional region.(5) The information processing device according to any one of (1) to (4), further comprising an encoding control unit that generates first UV vertex position encoding control information that controls whether to differentially encode the UV vertex positions, wherein the atlas information encoding unit is configured to differentially encode the UV vertex positions in accordance with the first UV vertex position encoding control information, encode the atlas information including the first UV vertex position encoding control information, and generate encoded data of the atlas information. (6) The information processing device according to (5), wherein the encoding control unit is configured to set the first UV vertex position encoding control information for an area where the number of mesh faces between the current frame and the key frame is the same to a value indicating that the UV vertex positions can be differentially encoded. (7) The information processing device according to (6), wherein the encoding control unit is configured to set the first UV vertex position encoding control information to a value indicating that the UV vertex positions can be differentially encoded when the UV vertex positions between the current frame and the key frame are different from each other. (8) The information processing device according to any one of (5) to (7), wherein the atlas information encoding unit is configured to intra-code the UV vertex positions of the current frame when the first UV vertex position encoding control information is a value indicating not to differentially encode the UV vertex positions. (9) The information processing device according to any one of (5) to (8), wherein the encoding control unit further generates second UV vertex position encoding control information for controlling an encoding method of the UV vertex positions, and when the first UV vertex position encoding control information is set to a value indicating that the UV vertex positions may be differentially encoded and the second UV vertex position encoding control information is set to a value indicating differential encoding, the atlas information encoding unit is configured to differentially encode the UV vertex positions between the current frame and the key frame. (10) The information processing device according to any one of (9), wherein the encoding control unit is configured to prohibit setting the second UV vertex position encoding control information to the value indicating differential encoding when the number of vertices between the current frame and the key frame differs from each other.(11) The information processing device according to (9) or (10), wherein the atlas information encoding unit is configured to intra-code the UV vertex positions of the current frame when the second UV vertex position encoding control information is a value that does not indicate the differential encoding. (12) The information processing device according to (1) to (11), wherein the atlas information encoding unit is configured to differentially encode UV vertex identifiers that identify vertices of the base mesh in the attribute map, and generate encoded data of the atlas information including the differentially encoded UV vertex identifiers. (13) The information processing device according to (12), further comprising an encoding control unit that generates first UV vertex identifier encoding control information that controls whether to differentially encode the UV vertex identifiers, wherein the atlas information encoding unit is configured to differentially encode the UV vertex identifiers in accordance with the first UV vertex identifier encoding control information, encode the atlas information including the first UV vertex identifier encoding control information, and generate encoded data of the atlas information. (14) The information processing device according to (13), wherein the encoding control unit is configured to set the first UV vertex identifier encoding control information for an area where the number of mesh faces between the current frame and the key frame is the same to a value indicating that the UV vertex identifiers can be differentially encoded. (15) The information processing device according to (14), wherein the encoding control unit is configured to set the first UV vertex identifier encoding control information to a value indicating that the UV vertex identifiers can be differentially encoded when the UV vertex identifiers are different between the current frame and the key frame. (16) The information processing device according to any of (13) to (15), wherein the atlas information encoding unit is configured to intra-code the UV vertex identifiers of the current frame when the first UV vertex identifier encoding control information is a value indicating that the UV vertex identifiers will not be differentially encoded.(17) The information processing device according to any one of (13) to (16), wherein the encoding control unit further generates second UV vertex identifier encoding control information that controls an encoding method of the UV vertex identifiers, and the atlas information encoding unit is configured to differentially encode the UV vertex identifiers between the current frame and the key frame when the first UV vertex identifier encoding control information is set to a value indicating that the UV vertex identifiers can be differentially encoded and the second UV vertex identifier encoding control information is set to a value indicating differential encoding between the current frame and the key frame. (18) The information processing device according to (17), wherein the atlas information encoding unit is configured to differentially encode the UV vertex identifiers between vertices of the current frame and vertices of the key frame that are in the same order as each other in a predetermined sorting order of the UV vertex identifiers. (19) The information processing device according to (17) or (18), wherein the atlas information encoding unit is configured to differentially encode the UV vertex identifiers within the current frame when the first UV vertex identifier encoding control information is set to a value indicating that the UV vertex identifiers may be differentially encoded and the second UV vertex identifier encoding control information is set to a value indicating differential encoding within the current frame. (20) The information processing device according to (19), wherein the atlas information encoding unit is configured to differentially encode the UV vertex identifiers between a vertex of a current processing target and a vertex of a previous processing target in a predetermined sort order of the UV vertex identifiers within the current frame. (21) The information processing device according to (17) to (20), wherein the atlas information encoding unit is configured to intra-code the UV vertex identifiers of the current frame when the second UV vertex identifier encoding control information is a value not indicating the differential encoding.(22) Differentially encode UV vertex positions between a current frame and a key frame, and generate encoded data of atlas information including the differentially encoded UV vertex positions; and generate a bitstream including encoded data of a base mesh, encoded data of a displacement video, encoded data of an attribute video, and encoded data of the atlas information, wherein the base mesh is a mesh with lower resolution than an original mesh to be encoded, which is composed of vertices and connections that represent a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh, the displacement video is a moving image having frame images of a displacement map, which is a two-dimensional area packed with displacement vectors that indicate the displacement of division points, which are vertices generated by subdividing the base mesh; and the attribute video is a moving image having frame images of an attribute map, which is a two-dimensional area packed with projected images of the texture of the original mesh; the atlas information is information used for mesh reconstruction, including information indicating a correspondence between the base mesh, the displacement map, and the attribute map; and the UV vertex positions are position information indicating the positions of the vertices of the base mesh in the attribute map. Information processing methods.
[0227] (31) A video encoding system comprising: a demultiplexing unit that demultiplexes a bitstream and extracts coded data of a base mesh, coded data of a displacement video, coded data of an attribute video, and coded data of atlas information; and an atlas information decoding unit that differentially decodes UV vertex positions that are differentially coded between a current frame and a key frame and are included in the extracted coded data of the atlas information, wherein the base mesh is a mesh with lower resolution than an original mesh to be encoded, which is composed of vertices and connections that represent a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh; the displacement video is a video having frame images of a displacement map that is a two-dimensional area packed with displacement vectors that indicate the displacement of division points that are vertices generated by subdividing the base mesh; the attribute video is a video having frame images of an attribute map that is a two-dimensional area packed with projected images of the texture of the original mesh; and the atlas information is information used for mesh reconstruction, including information indicating a correspondence between the base mesh, the displacement map, and the attribute map. The information processing device according to (31), wherein the UV vertex positions are position information indicating positions of vertices of the base mesh in the attribute map. (32) The information processing device according to (31), wherein the atlas information decoding unit is configured to differentially decode the UV vertex positions between vertices of the current frame and vertices of the key frame that are in the same order in a predetermined sorting order of the UV vertex positions. (33) The information processing device according to (32), wherein the key frame is the frame immediately preceding the current frame. (34) The information processing device according to (33), further comprising a base mesh decoding unit that differentially decodes geometry vertex positions that indicate positions of the vertices of the base mesh in a three-dimensional region.(35) The information processing device according to any one of (31) to (34), wherein the atlas information decoding unit is configured to differentially decode the UV vertex positions between the current frame and the key frame when UV vertex position encoding control information controlling a method for encoding the UV vertex positions is set to a value indicating differential encoding. (36) The information processing device according to (35), wherein the atlas information decoding unit is configured to intra-decode the UV vertex positions of the current frame when the UV vertex position encoding control information is a value not indicating the differential encoding. (37) The information processing device according to any one of (31) to (36), wherein the atlas information decoding unit is configured to differentially decode UV vertex identifiers that identify vertices of the base mesh in the attribute map that have been differentially encoded. (38) The information processing device according to (37), wherein the atlas information decoding unit is configured to differentially decode the UV vertex identifiers between the current frame and the key frame when UV vertex identifier encoding control information controlling a method for encoding the UV vertex identifiers is set to a value indicating differential encoding between the current frame and the key frame. (39) The information processing device according to (38), wherein the atlas information decoding unit is configured to differentially decode the UV vertex identifiers between vertices of the current frame and vertices of the key frame that are in the same order in a predetermined sorting order of the UV vertex identifiers. (40) The information processing device according to (38) or (39), wherein the atlas information decoding unit is configured to differentially decode the UV vertex identifiers within the current frame when the UV vertex identifier encoding control information is set to a value indicating differential encoding within the current frame. (41) The information processing device according to (40), wherein the atlas information decoding unit is configured to differentially decode the UV vertex identifiers between a vertex of a current processing target and a vertex of a previous processing target in a predetermined sorting order of the UV vertex identifiers within the current frame. (42) The information processing device according to any of (38) to (41), wherein the atlas information decoding unit is configured to intra-decode the UV vertex identifiers when the UV vertex identifier encoding control information is a value that does not indicate differential encoding.(43) Demultiplexing a bitstream to extract coded data of a base mesh, coded data of a displacement video, coded data of an attribute video, and coded data of atlas information; and differentially decoding UV vertex positions differentially coded between a current frame and a key frame, which are included in the extracted coded data of the atlas information; wherein the base mesh is a mesh with lower resolution than an original mesh to be coded, which is composed of vertices and connections that represent a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh; the displacement video is a video having frame images of a displacement map, which is a two-dimensional area packed with displacement vectors indicating the displacement of division points, which are vertices generated by subdividing the base mesh; the attribute video is a video having frame images of an attribute map, which is a two-dimensional area packed with projected images of the texture of the original mesh; the atlas information is information used for mesh reconstruction, including information indicating a correspondence between the base mesh, the displacement map, and the attribute map; and the UV vertex positions are position information indicating the positions of the vertices of the base mesh in the attribute map. Information processing methods.
[0228] (51) A video encoding system comprising: a base mesh encoding unit that, for a current frame in which a geometry vertex topology matches a key frame and at least one of a UV vertex position and a UV vertex topology changes with respect to the key frame, differentially encodes geometry vertex positions relative to the key frame to generate coded data of a base mesh; an atlas information encoding unit that differentially encodes at least one of the UV vertex positions and UV vertex identifiers for the current frame to generate coded data of atlas information; and a bitstream generating unit that generates a bitstream including the coded data of the base mesh, coded data of a displacement video, coded data of an attribute video, and coded data of the atlas information, wherein the base mesh is a mesh with lower resolution than an original mesh to be encoded, the original mesh being composed of vertices and connections that represent a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh, the original mesh being composed of vertices and connections that represent the three-dimensional structure of the object; and the displacement video is a video in which frame images are displacement maps that are two-dimensional regions in which displacement vectors indicating displacements of division points, which are vertices generated by subdividing the base mesh, are packed, The attribute video is a moving image whose frame images are an attribute map, which is a two-dimensional area in which projected images of the texture of the original mesh are packed; the atlas information is information used to reconstruct a mesh, including information indicating the correspondence between the base mesh, the displacement map, and the attribute map; the geometry vertex positions indicate the positions of the vertices of the base mesh in a three-dimensional area; the geometry vertex topology indicates how the vertices of the base mesh are connected in the three-dimensional area; the UV vertex positions indicate the positions of the vertices of the base mesh in the attribute map; and the UV vertex topology indicates how the vertices of the base mesh are connected in the attribute map.(52) For a current frame in which a geometry vertex topology matches a key frame and at least one of a UV vertex position and a UV vertex topology changes with respect to the key frame, differentially encode the geometry vertex positions relative to the key frame to generate coded data of a base mesh; for the current frame, differentially encode at least one of the UV vertex positions and UV vertex identifiers to generate coded data of atlas information; and generate a bitstream including the coded data of the base mesh, the coded data of a displacement video, the coded data of an attribute video, and the coded data of the atlas information; the base mesh is a mesh with lower resolution than an original mesh to be coded, which is composed of vertices and connections that represent a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh; the displacement video is a moving image in which frame images are a displacement map that is a two-dimensional area in which displacement vectors indicating displacements of division points that are vertices generated by subdividing the base mesh are packed; and the attribute video is a moving image in which frame images are an attribute map that is a two-dimensional area in which projected images of the texture of the original mesh are packed; An information processing method, wherein the atlas information is information used to reconstruct a mesh, including information indicating the correspondence between the base mesh and the displacement map and the attribute map; the geometry vertex positions indicate the positions of the vertices of the base mesh in a three-dimensional region; the geometry vertex topology indicates how the vertices of the base mesh are connected in the three-dimensional region; the UV vertex positions indicate the positions of the vertices of the base mesh in the attribute map; and the UV vertex topology indicates how the vertices of the base mesh are connected in the attribute map.(53) A data set includes: encoded data of a base mesh, for a frame in which a geometry vertex topology matches a key frame and at least one of a UV vertex position and a UV vertex topology changes relative to the key frame, wherein the encoded data of the base mesh is differentially encoded relative to the key frame; and encoded data of atlas information, for which at least one of the UV vertex position and the UV vertex identifier is differentially encoded relative to the key frame; wherein the base mesh is a mesh with lower resolution than the original mesh to be encoded, which is composed of vertices and connections that represent a three-dimensional structure of an object, and is generated by thinning out vertices from the original mesh; the atlas information is information used for mesh reconstruction, including information indicating a correspondence between the base mesh, and a displacement map and an attribute map; the displacement map is a two-dimensional region in which displacement vectors indicating displacements of division points, which are vertices generated by subdividing the base mesh, are packed; and the attribute map is a two-dimensional region in which a projected image of a texture of the original mesh is packed; and the geometry vertex positions indicate positions of the vertices of the base mesh in a three-dimensional region. The geometry vertex topology indicates how the vertices of the base mesh are connected in the three-dimensional domain, the UV vertex positions indicate the positions of the vertices of the base mesh in the attribute map, and the UV vertex topology indicates how the vertices of the base mesh are connected in the attribute map.
[0229] 300 Encoding device, 311 Preprocessing unit, 312 V-DMC encoding unit, 321 Mesh decimation unit, 322 Atlas information generation unit, 323 Encoding control unit, 324 Displacement vector generation unit, 351 Atlas information encoding unit, 352 Base mesh encoding unit, 353 Displacement vector correction unit, 354 Displacement video encoding unit, 355 Mesh reconstruction unit, 356 Attribute map conversion unit, 357 Attribute video encoding unit, 358 Multiplexing unit, 400 Decoding device, 411 Demultiplexing unit, 412 Atlas information decoding unit, 413 Base mesh decoding unit, 414 Base mesh reconstruction unit, 415 Subdivision unit, 416 Displacement video decoding unit, 417 Unpacking unit, 418 Displacement vector application unit, 419 Attribute video decoding unit, 420 display processing unit, 900 computer
Claims
1. An atlas information encoding unit that differentially encodes UV vertex positions between a current frame and a key frame and generates encoded data of atlas information including the differentially encoded UV vertex positions; and a bitstream generation unit that generates a bitstream including encoded data of a base mesh, encoded data of a displacement video, encoded data of an attribute video, and the encoded data of the atlas information. The base mesh is a mesh with a lower level of detail than the original mesh, which is generated by decimating vertices from the original mesh to be encoded, which is composed of vertices and connections representing the three-dimensional structure of an object. The displacement video is a moving image with a frame image being a displacement map, which is a two-dimensional area in which displacement vectors indicating displacements of split points, which are vertices generated by subdividing the base mesh, are packed. The attribute video is a moving image with a frame image being an attribute map, which is a two-dimensional area in which a projected image of the texture of the original mesh is packed. The atlas information is information used for mesh reconstruction, including information indicating a correspondence relationship between the base mesh, the displacement map, and the attribute map. The UV vertex position is position information indicating the position of a vertex of the base mesh in the attribute map. An information processing apparatus.
2. The information processing apparatus according to claim 1, wherein the atlas information encoding unit is configured to differentially encode the UV vertex positions between the vertices of the current frame and the vertices of the key frame that are in the same order as each other in a predetermined alignment order of the UV vertex positions.
3. The information processing apparatus according to claim 1, further comprising an encoding control unit that generates first UV vertex position encoding control information for controlling whether to differentially encode the UV vertex positions. The atlas information encoding unit is configured to differentially encode the UV vertex positions according to the first UV vertex position encoding control information, encode the atlas information including the first UV vertex position encoding control information, and generate the encoded data of the atlas information.
4. The encoding control unit is configured to set the first UV vertex position encoding control information for a region where the number of faces of the mesh is the same between the current frame and the key frame to a value indicating that the UV vertex position can be differentially encoded. The information processing apparatus according to claim 3.
5. The encoding control unit is configured to set the first UV vertex position encoding control information to a value indicating that the UV vertex position can be differentially encoded when the UV vertex positions are different between the current frame and the key frame. The information processing apparatus according to claim 4.
6. The encoding control unit further generates second UV vertex position encoding control information for controlling an encoding method of the UV vertex position. The atlas information encoding unit is configured to differentially encode the UV vertex position between the current frame and the key frame when the first UV vertex position encoding control information is set to a value indicating that the UV vertex position can be differentially encoded and the second UV vertex position encoding control information is set to a value indicating differential encoding. The information processing apparatus according to claim 3.
7. The encoding control unit is configured to prohibit setting the second UV vertex position encoding control information to a value indicating the differential encoding when the number of vertices is different between the current frame and the key frame. The information processing apparatus according to claim 6.
8. The atlas information encoding unit is configured to intra-encode the UV vertex position of the current frame when the second UV vertex position encoding control information does not indicate the differential encoding. The information processing apparatus according to claim 6.
9. The atlas information encoding unit is configured to differentially encode a UV vertex identifier that identifies a vertex of the base mesh in the attribute map and generate encoded data of the atlas information including the differentially encoded UV vertex identifier. The information processing apparatus according to claim 1.
10. The information processing apparatus according to claim 9, further comprising an encoding control unit that generates first UV vertex identifier encoding control information for controlling whether to differentially encode the UV vertex identifier, wherein the atlas information encoding unit differentially encodes the UV vertex identifier according to the first UV vertex identifier encoding control information, encodes the atlas information including the first UV vertex identifier encoding control information, and generates encoded data of the atlas information.
11. The information processing apparatus according to claim 10, wherein the encoding control unit is configured to set the first UV vertex identifier encoding control information for a region where the number of faces of the mesh is the same between the current frame and the key frame to a value indicating that the UV vertex identifier can be differentially encoded.
12. The information processing apparatus according to claim 11, wherein the encoding control unit is configured to set the first UV vertex identifier encoding control information to a value indicating that the UV vertex identifier can be differentially encoded when the UV vertex identifiers are different between the current frame and the key frame.
13. The information processing apparatus according to claim 10, wherein the encoding control unit further generates second UV vertex identifier encoding control information for controlling an encoding method of the UV vertex identifier, and the atlas information encoding unit differentially encodes the UV vertex identifier between the current frame and the key frame when the first UV vertex identifier encoding control information is set to a value indicating that the UV vertex identifier can be differentially encoded and the second UV vertex identifier encoding control information is set to a value indicating differential encoding between the current frame and the key frame.
14. The information processing apparatus according to claim 13, wherein the atlas information encoding unit is configured to differentially encode the UV vertex identifier between the vertex of the current frame and the vertex of the key frame that are in the same order in a predetermined alignment order of the UV vertex identifiers.
15. The atlas information encoding unit is configured to differentially encode the UV vertex identifier within the current frame when the first UV vertex identifier encoding control information is set to a value indicating that the UV vertex identifier can be differentially encoded and the second UV vertex identifier encoding control information is set to a value indicating differential encoding within the current frame. The information processing apparatus according to claim 13.
16. The atlas information encoding unit is configured to differentially encode the UV vertex identifier between the current processing target vertex and the vertex of the previous processing target in a predetermined alignment order of the UV vertex identifiers within the current frame. The information processing apparatus according to claim 15.
17. The atlas information encoding unit is configured to intra-encode the UV vertex identifier of the current frame when the second UV vertex identifier encoding control information is a value that does not indicate differential encoding. The information processing apparatus according to claim 13.
18. Differentially encode the UV vertex positions between the current frame and the key frame, generate encoded data of atlas information including the differentially encoded UV vertex positions, generate a bitstream including the encoded data of the base mesh, the encoded data of the displacement video, the encoded data of the attribute video, and the encoded data of the atlas information. The base mesh is a mesh with a lower level of detail than the original mesh, which is generated by decimating vertices from the original mesh of the encoding target composed of vertices and connections representing the three-dimensional structure of the object. The displacement video is a moving image with a displacement map, which is a two-dimensional area in which displacement vectors indicating the displacements of the split points, which are the vertices generated by subdividing the base mesh, are packed, as the frame image. The attribute video is a moving image with an attribute map, which is a two-dimensional area in which the projected image of the texture of the original mesh is packed, as the frame image. The atlas information is information used for mesh reconstruction, including information indicating the correspondence between the base mesh, the displacement map, and the attribute map. The UV vertex position is position information indicating the position of the vertex of the base mesh in the attribute map. Information processing method.
19. A demultiplexing unit that demultiplexes a bitstream and extracts encoded data of a base mesh, encoded data of displacement video, encoded data of attribute video, and encoded data of atlas information; and an atlas information decoding unit that differentially decodes UV vertex positions that are differentially encoded between a current frame and a key frame, included in the extracted encoded data of the atlas information. The base mesh is a mesh with a lower level of detail than the original mesh, generated by thinning out vertices from the original mesh to be encoded, which is composed of vertices and connections representing the three-dimensional structure of an object. The displacement video is a moving image having, as a frame image, a displacement map that is a two-dimensional region in which displacement vectors indicating displacements of division points, which are vertices generated by subdividing the base mesh, are packed. The attribute video is a moving image having, as a frame image, an attribute map that is a two-dimensional region in which a projection image of the texture of the original mesh is packed. The atlas information is information used for mesh reconstruction, including information indicating a correspondence relationship between the base mesh, the displacement map, and the attribute map. The UV vertex position is position information indicating the position of a vertex of the base mesh in the attribute map. Information processing apparatus.
20. Demultiplex a bitstream to extract encoded data of a base mesh, encoded data of displacement video, encoded data of attribute video, and encoded data of atlas information; perform differential decoding on UV vertex positions that are differentially encoded between the current frame and a key frame and are included in the extracted encoded data of the atlas information; the base mesh is a mesh with lower fineness than the original mesh, which is generated by thinning out vertices from the original mesh to be encoded, the original mesh being composed of vertices and connections representing the three-dimensional structure of an object; the displacement video is a moving image having, as a frame image, a displacement map that is a two-dimensional area in which displacement vectors indicating displacements of split points, which are vertices generated by subdividing the base mesh, are packed; the attribute video is a moving image having, as a frame image, an attribute map that is a two-dimensional area in which a projected image of the texture of the original mesh is packed; the atlas information is information used for mesh reconstruction, including information indicating the correspondence between the base mesh and the displacement map and the attribute map; the UV vertex position is position information indicating the position of the vertex of the base mesh in the attribute map. Information processing method.
21. For a current frame in which the geometry vertex topology coincides with the key frame and at least one of the UV vertex position and the UV vertex topology changes with respect to the key frame, a base mesh encoding unit that differentially encodes the geometry vertex positions between the current frame and the key frame to generate encoded data of the base mesh; an atlas information encoding unit that differentially encodes at least one of the UV vertex position and the UV vertex identifier for the current frame to generate encoded data of the atlas information; and a bitstream generation unit that generates a bitstream including the encoded data of the base mesh, the encoded data of the displacement video, the encoded data of the attribute video, and the encoded data of the atlas information, wherein the base mesh is a mesh with a lower level of detail than the original mesh, which is generated by decimating vertices from the original mesh to be encoded, which is composed of vertices and connections representing the three-dimensional structure of the object; the displacement video is a moving image with a displacement map, which is a two-dimensional area in which displacement vectors indicating displacements of split points, which are vertices generated by subdividing the base mesh, are packed, as a frame image; the attribute video is a moving image with an attribute map, which is a two-dimensional area in which a projected image of the texture of the original mesh is packed, as a frame image; the atlas information is information used for reconstructing the mesh, including information indicating the correspondence between the base mesh, the displacement map, and the attribute map; the geometry vertex position indicates the position of the vertices of the base mesh in a three-dimensional area; the geometry vertex topology indicates the way in which the vertices of the base mesh in the three-dimensional area are connected; the UV vertex position indicates the position of the vertices of the base mesh in the attribute map; and the UV vertex topology indicates the way in which the vertices of the base mesh in the attribute map are connected. Information processing apparatus.
22. For a current frame in which the geometry vertex topology matches the key frame and at least one of the UV vertex position and the UV vertex topology changes with respect to the key frame, the geometry vertex position is differentially encoded with respect to the key frame to generate encoded data of the base mesh, For the current frame, at least one of the UV vertex position and the UV vertex identifier is differentially encoded to generate encoded data of the atlas information, A bitstream including the encoded data of the base mesh, the encoded data of the displacement video, the encoded data of the attribute video, and the encoded data of the atlas information is generated, The base mesh is a mesh with a lower level of detail than the original mesh, which is generated by decimating vertices from the original mesh to be encoded, which is composed of vertices and connections representing the three-dimensional structure of the object, The displacement video is a moving image having, as a frame image, a displacement map which is a two-dimensional area in which displacement vectors indicating displacements of division points, which are vertices generated by subdividing the base mesh, are packed, The attribute video is a moving image having, as a frame image, an attribute map which is a two-dimensional area in which a projection image of the texture of the original mesh is packed, The atlas information is information used for mesh reconstruction, including information indicating the correspondence between the base mesh and the displacement map and the attribute map, The geometry vertex position indicates the position of the vertex of the base mesh in the three-dimensional area, The geometry vertex topology indicates the way in which the vertices of the base mesh are connected in the three-dimensional area, The UV vertex position indicates the position of the vertex of the base mesh in the attribute map, The UV vertex topology indicates the way in which the vertices of the base mesh are connected in the attribute map Information processing method.
23. Encoded data of a base mesh in which the geometry vertex positions are differentially encoded with respect to the key frame for a frame in which the geometry vertex topology coincides with the key frame and at least one of the UV vertex positions and the UV vertex topology changes with respect to the key frame, and encoded data of atlas information in which at least one of the UV vertex positions and the UV vertex identifiers is differentially encoded with respect to the key frame, wherein the base mesh is a mesh of lower fineness than the original mesh, generated by decimating vertices from the original mesh to be encoded, which is composed of vertices and connections representing the three-dimensional structure of the object, the atlas information is information used for mesh reconstruction, including information indicating the correspondence between the base mesh, the displacement map, and the attribute map, the displacement map is a two-dimensional region in which displacement vectors indicating the displacements of the split points, which are the vertices generated by subdividing the base mesh, are packed, the attribute map is a two-dimensional region in which the projected image of the texture of the original mesh is packed, the geometry vertex position indicates the position of the vertex of the base mesh in the three-dimensional region, the geometry vertex topology indicates the way in which the vertices of the base mesh in the three-dimensional region are connected, the UV vertex position indicates the position of the vertex of the base mesh in the attribute map, and the UV vertex topology indicates the way in which the vertices of the base mesh in the attribute map are connected Bitstream.
Citation Information
Patent Citations
Atlas sampling based mesh compression with charts of general topology
WO2023192027A1
Information processing device and method
WO2024157930A1