Information processing apparatus and method

By selecting the appropriate encoding method in the encoding unit, the problem of decreasing coding efficiency caused by the reduction of the number of sub-grid segmentation is solved, and efficient encoding is achieved in the case of reducing the number of faces.

CN119999192APending Publication Date: 2025-05-13SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071185.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-19
Filing Date
2023-10-03
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When the sub-grid is segmented, the number of faces of the coding unit decreases, resulting in a decrease in the use effect of reference results of adjacent faces, which may reduce the coding efficiency.

Method used

The first method or the second method is selected as the encoding method of the basic grid. In the first mode, the vertex information and the connection information are converted into secondary information indicating the relationship between adjacent faces of the base grid and encoded. In the second mode, the vertex information and connection information are directly encoded without conversion.

Benefits of technology

By selecting an appropriate encoding method, the reduction in encoding efficiency can be suppressed, and even if the number of faces is reduced, a high encoding efficiency can be maintained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119999192A_ABST
    Figure CN119999192A_ABST
Patent Text Reader

Abstract

The present disclosure pertains to an information processing device and method that make it possible to suppress a decrease in encoding efficiency. The first mode or the second mode is selected as a coding mode of the base grid to be intra-coded. In a case where the first mode is selected as an encoding mode, vertex information and connection information of the base grid are converted into secondary information indicating a relationship between adjacent faces of the base grid, and then the secondary information is encoded. When the second mode is selected as the encoding mode, the vertex information and the connection information are encoded without converting the vertex information and the connection information into secondary information. The present disclosure can be applied to, for example, an information processing apparatus, an electronic device, an information processing method, or a program.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus and method, and more particularly, to an information processing apparatus and method capable of reducing a decrease in encoding efficiency. Background Art

[0002] Conventionally, as a method for encoding a mesh of 3D data representing a three-dimensional structure of an object by connecting vertices, there has been video-based dynamic mesh coding (V-DMC) (for example, see non-patent document 1). In V-DMC, a mesh to be encoded is represented by a rough base mesh and displacement vectors of division points obtained by subdividing the base mesh, and the base mesh and the displacement vectors are encoded. The displacement vectors are stored (packed) in a two-dimensional image and encoded as a moving image (displacement video) having the two-dimensional image as a frame.

[0003] As a coding method of the base mesh, there are intra-frame coding that performs coding independently for each frame and inter-frame coding that performs coding using correlation between frames. As intra-frame coding, for example, a coding method that refers to adjacent faces such as Draco has been proposed. In such a coding method, vertex information indicating the positions of vertices constituting the base mesh and connection information indicating connections between vertices constituting the base mesh are converted into "secondary information indicating the relationship between adjacent faces" and encoded.

[0004] Meanwhile, in recent years, a method has been proposed to divide a frame into a plurality of sub-grids (patches) and encode them independently of each other (for example, see Non-Patent Document 2 and Non-Patent Document 3). By dividing into sub-grids, an encoding method more suitable for the shape characteristics of the grid can be selected.

[0005] Citation List

[0006] Non-patent literature

[0007] Non-Patent Literature 1: Khaled Mammou, Jungsun Kim, Alexis Tourapis, DimitriPodborski, Krasimir Kolarov, “[V-CG] Apple's Dynamic Mesh Coding CfP Response”, ISO / IEC JTC 1 / SC 29 / WG 7m59281, April 2022

[0008] Non-patent literature 2: Alexis Tourapis, Jungsun Kim, Dimitri Podborski, Khaled Mammou, "Base mesh data substream format for VDMC", ISO / IEC JTC 1 / SC 29 / WG7m60362, July 2022

[0009] Non-patent literature 3: Jungsun Kim, Alexis Tourapis, Dimitri Podborski, Khaled Mammou, David, “VDMC support in the V3C framework”, ISO / IEC JTC 1 / SC 29 / WG7m60363, July 2022. Summary of the invention

[0010] Problems to be solved by the present invention

[0011] However, when the sub-grid is divided, the number of faces of a coding unit (a data unit that can be encoded independently of other coding units) is reduced. In the encoding method of referring to adjacent faces as described above, when the number of faces in the coding unit is reduced, the use effect of the reference result of the adjacent faces is reduced, and the encoding efficiency may be reduced.

[0012] The present disclosure has been made in view of such circumstances, and an object of the present disclosure is to make it possible to reduce the reduction in encoding efficiency.

[0013] Solution to the problem

[0014] An information processing device according to one aspect of the present technology includes: an encoding method determination unit that selects a first method or a second method as an encoding method of a base mesh to be subjected to intra-frame encoding; a first encoding unit that, when the first method is selected as the encoding method, converts vertex information indicating positions of vertices constituting the base mesh and connection information indicating connections between vertices constituting the base mesh into secondary information indicating relationships between adjacent faces of the base mesh and encodes the secondary information; and a second encoding unit that, when the second method is selected as the encoding method, encodes the vertex information and the connection information without converting them into secondary information. The base mesh is a mesh generated by removing vertices from an original mesh to be encoded that is composed of vertices and connections representing a three-dimensional structure of an object, and is coarser than the original mesh.

[0015] An information processing method according to one aspect of the present technology is an information processing method, including: selecting a first mode or a second mode as an encoding mode of a base mesh to be subjected to intra-frame encoding; in a case where the first mode is selected as the encoding mode, converting vertex information indicating positions of vertices constituting the base mesh and connection information indicating connections between vertices constituting the base mesh into secondary information indicating relationships between adjacent faces of the base mesh and encoding the secondary information; and in a case where the second mode is selected as the encoding mode, encoding the vertex information and the connection information without converting them into secondary information. The base mesh is a mesh generated by removing vertices from an original mesh to be encoded that is composed of vertices and connections representing a three-dimensional structure of an object, and is coarser than the original mesh.

[0016] An information processing device according to another aspect of the present technology includes: a first decoding unit, which, when the encoding mode control information indicates the first mode as the encoding mode of the intra-encoded base mesh, decodes the encoded data of the intra-encoded base mesh, generates secondary information indicating the relationship between adjacent faces of the base mesh, and converts the generated secondary information into vertex information and connection information of the base mesh; and a second decoding unit, which, when the encoding mode control information indicates the second mode as the encoding mode, decodes the encoded data of the base mesh, and generates vertex information and connection information that are not converted into secondary information. The base mesh is a mesh generated by removing vertices from an original mesh to be encoded that is composed of vertices and connections representing a three-dimensional structure of an object, and is coarser than the original mesh. Vertex information is information indicating the positions of vertices constituting the base mesh. Connection information is information indicating connections between vertices constituting the base mesh. The first mode is an encoding mode that converts vertex information and connection information into secondary information and encodes the secondary information. The second mode is an encoding mode that encodes vertex information and connection information without converting them into secondary information.

[0017] According to another aspect of the present technology, an information processing method includes: in a case where the encoding mode control information indicates a first mode as an encoding mode of an intra-encoded base mesh, decoding the encoded data of the intra-encoded base mesh, generating secondary information indicating a relationship between adjacent faces of the base mesh, and converting the generated secondary information into vertex information and connection information of the base mesh; and in a case where the encoding mode control information indicates a second mode as an encoding mode, decoding the encoded data of the base mesh, and generating vertex information and connection information that are not converted into secondary information. The base mesh is a mesh generated by removing vertices from an original mesh to be encoded that is composed of vertices and connections representing a three-dimensional structure of an object, and is coarser than the original mesh. Vertex information is information indicating the positions of vertices constituting the base mesh. Connection information is information indicating connections between vertices constituting the base mesh. The first mode is an encoding mode that converts vertex information and connection information into secondary information and encodes the secondary information. The second mode is an encoding mode that encodes vertex information and connection information without converting them into secondary information.

[0018] In an information processing device and method according to one aspect of the present technology, a first mode or a second mode is selected as an encoding mode of a base mesh to be intra-encoded, and when the first mode is selected as the encoding mode, vertex information indicating the positions of vertices constituting the base mesh and connection information indicating connections between vertices constituting the base mesh are encoded by converting them into secondary information indicating a relationship between adjacent faces of the base mesh, and when the second mode is selected as the encoding mode, the vertex information and the connection information are encoded without converting them into secondary information.

[0019] In an information processing device and method according to another aspect of the present technology, in a case where the encoding mode control information indicates a first mode as an encoding mode of a base mesh that has been intra-encoded, the encoded data of the base mesh is decoded to generate secondary information indicating a relationship between adjacent faces in the base mesh, and the generated secondary information is converted into vertex information and connection information of the base mesh. In a case where the encoding mode control information indicates a second mode as an encoding mode, the encoded data of the base mesh is decoded to generate vertex information and connection information that are not converted into secondary information. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a diagram for explaining a grid.

[0021] Figure 2 It is a diagram for explaining V-DMC.

[0022] Figure 3It is a diagram for explaining displacement vectors.

[0023] Figure 4 A diagram for explaining a displacement video.

[0024] Figure 5 This is a diagram for explaining an example of the intra-frame encoding method.

[0025] Figure 6 This is a diagram for explaining an example of the intra-frame encoding method.

[0026] Figure 7 This is a diagram for explaining an example of the intra-frame encoding method.

[0027] Figure 8 This is a diagram for explaining sub-gridding.

[0028] Fig. 9 This is a diagram for explaining sub-gridding.

[0029] Fig.10 This is a diagram for explaining sub-gridding.

[0030] Fig.11 This is a diagram for explaining sub-gridding.

[0031] Fig.12 This is a diagram for explaining comparison of bit amounts.

[0032] Fig.13 is a diagram showing a comparative example of bit amounts.

[0033] Fig.14 is a diagram showing an example of an encoding method.

[0034] Fig.15 This is a diagram showing an overview of the encoding method.

[0035] Fig.16 is a diagram showing an example of an original patch.

[0036] Fig.17 is a diagram showing an example of an original patch.

[0037] Fig.18 is a diagram showing an example of storing original patches.

[0038] Fig.19 is a diagram showing a comparative example of bit rates.

[0039] Fig. 20 is a block diagram showing a main configuration example of an encoding device.

[0040] Fig.21 is a block diagram showing a main configuration example of an intra unit.

[0041] Fig. 22 is a flowchart for explaining an example of the flow of encoding processing.

[0042] Fig.23 : is a flowchart for explaining an example of the flow of intra-frame encoding processing.

[0043] Fig.24 is a block diagram showing a main configuration example of a decoding device.

[0044] Fig.25 is a block diagram showing an example of a main configuration of an intra decoding unit.

[0045] Fig.26 is a flowchart for explaining an example of the flow of decoding processing.

[0046] Fig. 27 : is a flowchart for explaining an example of the flow of intra-frame decoding processing.

[0047] Fig.28 is a block diagram showing a main configuration example of a computer. DETAILED DESCRIPTION

[0048] Modes for carrying out the present disclosure (hereinafter referred to as embodiments) are described below. Note that the description will be made in the following order.

[0049] 1. Documents supporting technical content and technical terminology, etc.

[0050] 2. Intra-frame coding of base grid

[0051] 3. Encoding without conversion into secondary information

[0052] 4. Implementation Method

[0053] 5. Additional Notes

[0054] <1. Documents supporting technical content and technical terminology, etc.>

[0055] In addition to the contents disclosed in the embodiments, the scope disclosed in the present technology also includes contents described in the following non-patent documents and the like known at the time of filing, and contents of other documents cited in the following non-patent documents and the like.

[0056] Non-Patent Document 1: (described above)

[0057] That is, the contents described in the above non-patent documents, the contents of other documents cited in the above non-patent documents, etc. are also the basis for determining the support requirements.

[0058] <2. Intra-frame coding of base grid>

[0059] <v-dmc>

[0060] Conventionally, as 3D data representing a three-dimensional structure of a vertical structure (an object having a three-dimensional shape), there is known a mesh that represents the three-dimensional shape of an object surface by forming polygons by vertices and connections.

[0061] like Figure 1 As shown in the upper left part of FIG. 1 , in the mesh, vertices 11 and connections 12 connecting the vertices 11 form polygonal planes (polygons). The surface of an object having a three-dimensional structure, that is, the three-dimensional shape of the object is represented by a polygon (also called a face). Note that a texture 13 can be applied to each face of the mesh.

[0062] Grid data includes e.g. Figure 1 The information shown in the lower part. Figure 1 The vertex information 14 shown first from the left in the lower portion of FIG. 1 is information indicating the three-dimensional position (three-dimensional coordinates (X, Y, Z)) of each vertex 11 constituting the mesh. Figure 1 The connection information 15 shown second from the left in the lower portion is information indicating each connection 12 constituting the mesh. Figure 1 The texture image 16 shown third from the left in the lower portion is mapping information of the texture 13 attached to each face. Figure 1 The UV map 17 shown fourth from the left in the lower portion is information indicating the correspondence relationship between the vertices 11 and the texture 13. In the UV map 17, the coordinates (UV coordinates) of each vertex 11 in the texture image 16 are shown.

[0063] For example, as a method for encoding such a mesh, there is video-based dynamic mesh coding (V-DMC) as disclosed in Non-Patent Document 1.

[0064] In V-DMC, a mesh to be encoded is represented by a coarse base mesh and displacement vectors of division points obtained by subdividing the base mesh, and the base mesh and the displacement vectors are encoded.

[0065] For example, suppose there is a Figure 2 The original mesh is shown in the top part of Figure 2 In FIG. 1 , black dots indicate vertices and lines connecting the black dots indicate connections. As described above, a mesh is initially formed into polygons by connecting vertices, but here, for convenience of description, a mesh is described as a linearly (serially) connected group of vertices.

[0066] By removing some vertices from the original mesh, we get Figure 2 The coarse mesh shown in the second row from the top in . This is used as the base mesh.

[0067] By subdividing each grid of this base grid, such as Figure 2 vertices are added as shown in the third row from the top in FIG. Here, it is assumed that by this subdivision, vertices are added by the number obtained by subtracting vertices from the middle of the original mesh. Therefore, a mesh having the same number of vertices as the original mesh is obtained. In this specification, the added vertices are also referred to as split points.

[0068] However, since the connections are updated when the vertices of the original mesh are thinned out, and the split points are formed on the updated connections, the shape of the subdivided base mesh is different from that of the original mesh. More specifically, Figure 2 As shown in the bottom part of , the position of the segmentation point (on the dotted line) is different from the original mesh. In this specification, the difference between the position of the segmentation point and the position of the vertex of the original mesh is called a displacement vector.

[0069] For example, assuming that the original mesh 21 and the subdivided base mesh 22 are Figure 3 In addition, it is assumed that there are vertices 23 and 24 of the subdivided base mesh 22. In this case, Figure 3 As shown, the position of vertex 23 and the position of vertex 23' of original mesh 21 corresponding to vertex 23 are different from each other. This difference is represented as a displacement vector (displacement vector 25). Similarly, the difference between the position of vertex 24 and the position of vertex 24' of original mesh 21 corresponding to vertex 24 is represented as displacement vector 26. In this way, a displacement vector is set for each vertex of the subdivided base mesh.

[0070] In the encoder, since the original mesh is known, the base mesh can be generated, and such a displacement vector can be further obtained. The decoder can generate (restore) the original mesh (the mesh corresponding to the original mesh) by subdividing the base mesh and applying the displacement vector to each vertex.

[0071] As described above, by reducing the number of vertices of the original mesh and encoding the original mesh as a base mesh, the amount of encoding can be reduced. In addition, the displacement vector is stored in a two-dimensional image and encoded using 2D encoding, thereby improving the encoding efficiency. In this specification, storing data in a two-dimensional image is also referred to as packing.

[0072] Note that V-DMC supports scalable decoding. For example, Figure 4 As shown in FIG. 1 , the displacement vector is layered for each fineness (number of divisions of the base grid), and is packed into a two-dimensional image 31 as data for each layer. Figure 4 In the example of FIG. 3 , LoD0 packed in the two-dimensional image 31 indicates data of displacement vectors of vertices in the uppermost layer (lowest level of detail) among displacement vectors of vertices layered at each level of detail. Similarly, LoD1 indicates data of displacement vectors of vertices in a layer above LoD0. LoD2 indicates data of displacement vectors of vertices in a layer above LoD1. In this way, the displacement vectors are divided for each layer (put together as data for each layer) and packed.

[0073] Note that the displacement vectors can be packed as transform coefficients in a two-dimensional image by a coefficient transformation such as a wavelet transform. In addition, the displacement vectors can be quantized. For example, the displacement vectors can be converted into transform coefficients by a wavelet transform, the transform coefficients can be quantized, and the quantized transform coefficients (quantized coefficients) can be packed.

[0074] In addition, the three-dimensional shape of the object can change in the time direction. In this specification, the change in the time direction is also referred to as "dynamic". Therefore, the grid (i.e., the base grid and the displacement vector) is also dynamic. Therefore, the displacement vector is encoded as a motion image with a two-dimensional image as a frame. In this specification, the motion image is also referred to as displacement video.

[0075] <Base mesh encoding>

[0076] As a coding method of the base mesh, there are intra-frame coding that performs coding independently for each frame and inter-frame coding that performs coding using correlation between frames. For example, as intra-frame coding, a coding method that refers to adjacent faces, such as Draco, has been proposed. In such a coding method, vertex information indicating the positions of vertices constituting the base mesh and connection information indicating connections between vertices constituting the base mesh are converted into "secondary information indicating the relationship between adjacent faces" and encoded.

[0077] For example, in the case of Draco, the connection information is classified and symbolized for each face according to the pattern of the presence or absence of detection of adjacent faces in the search order. Then, the symbols are encoded. For example, Figure 5 As shown by the gray line in , the search order is set to a spiral shape, each face (the triangle in the figure) is detected in the search order, and clustering (classification) and symbolization are performed based on the presence or absence of detection results of adjacent faces.

[0078] For example, Figure 6 As shown, clustering is performed into five types of symbols, namely, C, L, R, S, and E. For example, in the case where the processing target vertex v is an unvisited vertex and there are no adjacent faces on the left and right sides of the processing target face, the processing target face is clustered into the symbol C. In addition, in the case where the processing target vertex v is a visited vertex and the face adjacent to the left side of the processing target face is a visited vertex, the processing target face is clustered into the symbol L. In addition, in the case where the processing target vertex v is a visited vertex and the face adjacent to the right side of the processing target face is a visited vertex, the processing target face is clustered into the symbol R. In addition, in the case where the processing target vertex v is a visited vertex and the faces adjacent to the left and right sides of the processing target face are unvisited vertices, the processing target face is clustered into the symbol S. In addition, in the case where the processing target vertex v is a visited vertex and the faces adjacent to the left and right sides of the processing target face are visited vertices, the processing target face is clustered into the symbol E. The symbols in which the faces are converted in this manner are encoded as connection information.

[0079] As described above, each symbol is set according to the detection situation of the adjacent surface. That is, each symbol is set according to the arrangement of the surface (the positional relationship between the processing target surface and its adjacent surface). In other words, the connection information is converted into "secondary information indicating the relationship between adjacent surfaces" and encoded.

[0080] In addition, for vertex information, such as Figure 7 As shown, the vertex position (gray point) is predicted by using the parallelogram prediction of the encoded adjacent faces. Then, the prediction residual is encoded as the vertex information of the processing target vertex v. The prediction residual depends on the position of the prediction point, that is, the parallelogram prediction, that is, the shape correlation between the processing target face and its adjacent faces. In other words, the vertex information is converted into "secondary information indicating the relationship between adjacent faces" and encoded.

[0081] In this way, in the encoding method of referring to adjacent faces, the encoding efficiency is improved by using the reference result, that is, the relationship between adjacent faces.

[0082] Meanwhile, in recent years, for example, as disclosed in Non-Patent Document 2 and Non-Patent Document 3, a method has been proposed in which one frame is divided into a plurality of sub-grids (patches) and encoded independently of each other. Figure 8 As shown, the grid 41 is divided into subgrids 42 and 43, and each subgrid is encoded independently. In this way, an encoding method can be selected for each subgrid. Therefore, an encoding method that is more suitable for the shape characteristics of the grid can be selected.

[0083] For example, suppose there is a Fig. 9 The base mesh 50 representing a human face is shown on the left side of FIG. Assume that the base mesh 50 includes a head 51 having a relatively simple three-dimensional shape and a hair portion 52 having a relatively complex three-dimensional shape. In general, if the correlation between frames is high enough, inter-frame coding tends to have higher coding efficiency than intra-frame coding. Since the three-dimensional shape of the head 51 is relatively simple, the correlation between frames tends to be high, and the head is suitable for inter-frame coding. On the other hand, since the three-dimensional shape of the hair portion 52 is relatively complex, the correlation between frames tends to be low, and the hair portion is not suitable for inter-frame coding. Therefore, when encoding the base mesh 50, the correlation between frames is reduced due to the influence of the hair portion 52, intra-frame coding is easily applied, and the coding efficiency may be reduced. In addition, even if inter-frame coding is applied, it will not last for a long time, and inter-frame coding and intra-frame coding are switched every few frames, which may reduce the coding efficiency.

[0084] In such a case, Fig. 9 As shown on the right side of , by sub-gridding the head 51 and the hair portion 52 and encoding them independently, inter-frame encoding can be stably selected for the head 51 and a decrease in encoding efficiency can be suppressed.

[0085] However, when the sub-grids are divided, the number of faces of a coding unit (a data unit that can be encoded independently of other data units) is reduced. In the encoding method of referring to adjacent faces as described above, when the number of faces in the coding unit is reduced, the use effect of the reference result of the adjacent faces is reduced, and the encoding efficiency may be reduced.

[0086] For example, in Figure 5 The grid shown in Fig.10 The thick line shown in indicates the case of division into two sub-meshes, and the faces on both sides of the thick line cannot be referenced. Therefore, the relationship between the faces adjacent to each other in this part cannot be used, and the encoding efficiency may be reduced.

[0087] In addition, subgridding increases the amount of independently encoded data. Fig.11 As shown, as the number of headers increases, there is a possibility that information serving as overhead increases and encoding efficiency decreases.

[0088] <Comparison of bit amounts>

[0089] For example, the bit amounts in the case of encoding using a coding method that references adjacent planes, such as Draco, are compared in three cases. In the first case, Fig.12 As shown on the left side of , one facet 61 is encoded at a time. In the second case, as Fig.12 As shown in the center of FIG. , a patch 61 is divided into three patches A to C, and these patches are collectively (merged) encoded at one time. In the third case, as Fig.12 As shown on the right side of , one patch 61 is divided into three patches A to C, and patch 61-1 (patch A), patch 61-2 (patch B), and patch 61-3 (patch C) are independently encoded.

[0090] Fig.13 The comparison results of the output bits of each case are shown in . As shown in the table, in case 1 (original), the output bits are 6496 bits. In addition, in case 2 (merged), the output bits are 7424 bits. In addition, in case 3 (plane A, plane B, plane C), the total number of output bits is 8824. It is obvious from the comparison between the output bits of case 1 and case 2 that by dividing into sub-grids, the use effect of the reference results of adjacent planes is reduced, and the amount of output bits is increased. In addition, it is obvious from the comparison between the output bits of case 2 and case 3 that by encoding each sub-grid independently of each other, the number of headers is increased and the amount of output bits is increased. As described above, there is a possibility of reducing the coding efficiency by sub-gridding.

[0091] <3. Encoding without conversion into secondary information>

[0092] <Method 1>

[0093] Therefore, when the base mesh is intra-coded, the vertex information and connection information of the base mesh are coded by a coding method different from the coding method of the reference adjacent faces. Fig.14 As shown in the top row of the table in , the vertex information and connection information of the base mesh are encoded without being converted into "secondary information indicating the relationship between adjacent faces" (method 1).

[0094] By doing so, it is possible to suppress a decrease in encoding efficiency due to a decrease in the use effect of the reference result of the adjacent surface through sub-gridding. That is, even when the number of surfaces to be encoded is reduced, a decrease in encoding efficiency can be suppressed.

[0095] <Method 1-1>

[0096] When applying this method 1, for example, Fig.14 As shown in the second row from the top of the table in , an encoding method can be selected and encoding can be performed by the selected encoding method (method 1-1). That is, the encoding method of the base mesh to be intra-encoded can be determined (the encoding method can be variable), and an encoding method for encoding the vertex information and connection information of the base mesh without converting the vertex information and connection information into "secondary information indicating the relationship between adjacent faces" can be applied as the encoding method.

[0097] For example, encoding method candidates may be prepared in advance, the encoding method candidates including encoding methods for encoding vertex information and connection information of a base mesh without converting them into "secondary information indicating the relationship between adjacent faces", and the encoding method to be applied to the base mesh for intra-frame encoding may be selected (determined) from the candidates. For example, an encoding method for encoding vertex information and connection information by converting them into "secondary information indicating the relationship between adjacent faces" (hereinafter also referred to as the first method), and an encoding method for encoding vertex information and connection information of the base mesh without converting them into secondary information (hereinafter also referred to as the second method) may be prepared as candidates, and any one of these encoding methods may be applied to the base mesh to be intra-frame encoded. In this specification, "selecting an encoding method" is also referred to as "determining an encoding method".

[0098] For example, an information processing device (also referred to as a first information processing device) includes: an encoding method determination unit, which selects a first method or a second method as an encoding method of a base mesh for intra-frame encoding; a first encoding unit, which, when the first method is selected as the encoding method, encodes vertex information and connection information by converting vertex information indicating the positions of vertices constituting the base mesh and connection information indicating connections between vertices constituting the base mesh into secondary information indicating a relationship between adjacent faces of the base mesh; and a second encoding unit, which, when the second method is selected as the encoding method, encodes the vertex information and the connection information without converting the vertex information and the connection information into secondary information.

[0099] In addition, in an information processing method performed by a first information processing device, the first method or the second method is selected as an encoding method for a basic mesh to be intra-encoded, and when the first method is selected as the encoding method, vertex information indicating the positions of vertices constituting the basic mesh and connection information indicating connections between vertices constituting the basic network are encoded by converting the vertex information indicating the positions of vertices constituting the basic mesh and connection information indicating connections between vertices constituting the basic network into secondary information indicating a relationship between adjacent faces of the basic mesh, and when the second method is selected as the encoding method, the vertex information and the connection information are encoded without converting them into secondary information.

[0100] For example, an information processing device (also referred to as a second information processing device) includes: a first decoding unit, which, when the encoding mode control information indicates the first mode as the encoding mode of the base mesh encoded within the frame, decodes the encoded data of the base mesh, generates secondary information indicating the relationship between adjacent faces of the base mesh, and converts the generated secondary information into vertex information and connection information of the base mesh; and a second decoding unit, which, when the encoding mode control information indicates the second mode as the encoding mode, decodes the encoded data of the base mesh, and generates vertex information and connection information that have not been converted into secondary information.

[0101] Furthermore, in the information processing method executed by the second information processing device, in a case where the encoding mode control information indicates the first mode as the encoding mode of the base mesh that has been intra-encoded, the encoded data of the base mesh is decoded to generate secondary information indicating a relationship between adjacent faces in the base mesh, and the generated secondary information is converted into vertex information and connection information of the base mesh. In a case where the encoding mode control information indicates the second mode as the encoding mode, the encoded data of the base mesh is decoded to generate vertex information and connection information that are not converted into secondary information.

[0102] Note that the base mesh is a mesh generated by removing vertices from an original mesh to be encoded consisting of vertices and connections representing the three-dimensional structure of an object, and is coarser than the original mesh.

[0103] Furthermore, vertex information is information indicating the positions of vertices constituting the base mesh. Connection information is information indicating connections between vertices constituting the base mesh. The first method is a coding method for encoding vertex information and connection information by converting the vertex information and connection information into the above-mentioned secondary information. The second method is a coding method for encoding vertex information and connection information without converting the vertex information and connection information into the above-mentioned secondary information.

[0104] Applying the second method does not always improve the coding efficiency compared to applying the first method. Therefore, by making it possible to select the coding method in this way, it is possible to suppress the reduction in coding efficiency in more various situations.

[0105] <Encoding using correlations between vertices>

[0106] Note that in this second manner, the vertex information and connection information of the base mesh can be encoded and decoded in any manner as long as the vertex information and connection information are not converted into "secondary information indicating the relationship between adjacent faces".

[0107] For example, in the first information processing device described above, the second encoding unit may encode the vertex information using the correlation between the vertices of the base mesh. For example, the second encoding unit may perform differential encoding on the vertex information. In addition, the second encoding unit may perform arithmetic encoding on the vertex information.

[0108] Similarly, in the above-mentioned second information processing device, the second decoding unit can decode the encoded data of the vertex information encoded using the correlation between the vertices of the base mesh by a decoding method corresponding to the encoding method. For example, the second decoding unit can perform differential decoding on the encoded data of the vertex information after differential encoding. In addition, the second decoding unit can perform arithmetic decoding on the encoded data of the vertex information after arithmetic encoding.

[0109] Furthermore, in the above-mentioned first information processing apparatus, the second encoding unit may perform fixed bit encoding on the vertex information. Similarly, in the above-mentioned second information processing apparatus, the second decoding unit may perform fixed bit decoding on the vertex information that has been subjected to fixed bit encoding.

[0110] <Method 1-1-1>

[0111] In the case where the encoding scheme is variable as described above, it is only necessary to apply the same encoding scheme to the encoder and the decoder.

[0112] For example, Fig.14 As shown in the third row from the top of the table in , the encoding mode control information indicating the encoding mode selected in the encoder can be transmitted from the encoder to the decoder (method 1-1-1). For example, in the encoder, the encoding mode control information indicating the applied (selected) encoding mode can be generated and sent. In addition, as Fig.15 As shown, the decoder can select to apply the first method or the second method for each sub-grid (face patch) according to the transmitted coding method control information, and can decode the sub-grid. Then, the sub-grids (face patches) decoded by the corresponding method can be combined to reconstruct the basic grid. For example, the coding method control information can be stored in a header and sent. In addition, the coding method control information can be encoded and sent (as encoded data).

[0113] For example, in the first information processing device, the second encoding unit may further store encoding mode control information indicating the applied (selected) encoding mode in the header, and encode the encoding mode control information. In addition, the first information processing device may further include a third encoding unit, which stores the encoding mode control information indicating the applied (selected) encoding mode in the header, and encodes the encoding mode control information.

[0114] Similarly, in the above-mentioned second information processing device, the second decoding unit can also decode the encoded data of the encoding mode control information stored in the header. In addition, the above-mentioned second information processing device may further include a third decoding unit, which decodes the encoded data of the encoding mode control information stored in the header. Then, the first decoding unit and the second decoding unit can perform encoding based on the encoding mode control information. That is, when the first mode is specified in the encoding mode control information, the first decoding unit can decode the encoded data of the base mesh, and when the second mode is specified in the encoding mode control information, the second decoding unit can decode the encoded data of the base mesh. As described above, the first decoding unit decodes the encoded data of the base mesh, generates secondary information indicating the relationship between adjacent faces of the base mesh, and converts the secondary information into vertex information and connection information of the base mesh. In addition, the second decoding unit decodes the encoded data of the base mesh and generates vertex information and connection information that are not converted into secondary information.

[0115] As described above, by transmitting encoding scheme control information indicating an encoding scheme selected by an encoder from an encoder to a decoder, the decoder can more easily perform decoding by applying the same encoding scheme as that applied by the encoder at the time of encoding.

[0116] <Encoding mode control information>

[0117] The encoding mode control information may be any information as long as the information indicates the encoding mode applied in the encoder. For example, the encoding mode control information may include an index indicating the applied encoding mode. For example, in the case of applying the first mode, an index (e.g., 0) indicating the first mode may be transmitted as the encoding mode control information, and in the case of applying the second mode, an index (e.g., 1) indicating the second mode may be transmitted as the encoding mode control information.

[0118] Furthermore, flag information indicating whether the second mode (or the first mode) is applied may be transmitted as the encoding mode control information.

[0119] <Method 1-2>

[0120] As described above, when the vertex information and the connection information of the base mesh are encoded by an encoding method that does not convert the vertex information and the connection information of the base mesh into "secondary information indicating the relationship between adjacent faces" for encoding, the base mesh can be converted into information for transmission (RAW face patch). Fig.14 As shown in the fourth row from the top of the table, a primitive patch can be generated from the base mesh, and the primitive patch can be encoded (method 1-2). The primitive patch is information for transmitting the base mesh, and although it has a data format different from that of a general base mesh, it is information indicating vertex information and connection information, and is not information indicating a relationship between adjacent faces of the base mesh.

[0121] For example, Fig.15 As shown, Draco may be prepared as a first method, and an encoding method for encoding original patches may be prepared as a second method, and any one of the first method and the second method may be applied to encoding and decoding the base mesh.

[0122] For example, the first information processing device may further include an original patch generation unit that converts the base mesh and generates an original patch as information for transmission. Then, in the first information processing device, when the second method is selected as the encoding method of the base mesh for intra-frame encoding, the second encoding unit may encode the original patch without converting the original patch into secondary information indicating the relationship between adjacent faces of the base mesh.

[0123] Furthermore, in the second information processing device, the second decoding unit may decode the encoded data of the base mesh and generate an original patch as information of the transmission format of the base mesh. Then, the second information processing device may further include a reconstruction unit that reconstructs the base mesh using the original patch.

[0124] <Original patch>

[0125] Next, the data of the original patch will be described. The data of the original patch includes mesh information and original patch information, the mesh information includes vertex information and connection information of the base mesh, and the original patch information includes meta information related to the base mesh. The mesh information may also include UV coordinates (information indicating the correspondence between vertices and textures). In addition, for example, the mesh information may also include attribute information such as normal vectors. The original patch information may include information indicating the number of original patches. In addition, the original patch information may include information indicating the number of faces in each original patch. In addition, the original patch information may include information indicating the number of vertices in each original patch. In addition, the original patch information may include information indicating the number of UV coordinates (also referred to as the number of UVs) in each original patch. In addition, the original patch information may include the number of attribute information (e.g., the number of normal vectors) in each original patch.

[0126] <Method 1-2-1>

[0127] The data format of the original patch (grid information and original patch information) can be in any format. Fig.14 As shown in the fifth row from the top of the table, in the mesh information of the original patch, the overlap information can be omitted, and each element (eg, vertex information) can be indicated by the minimum number (hereinafter also referred to as the first format) (Method 1-2-1).

[0128] For example, suppose there is a Fig.16 As shown in FIG. 11 , two original patches of patch A and patch B are provided. As shown in block 111, patch A has a vertex number of 8 (v: 8), a UV number of 8 (vt: 8), and a face number of 4 (f: 4). As shown in block 112, it is assumed that patch B has a vertex number of 4 (v: 4), a UV number of 4 (vt: 4), and a face number of 2 (f: 2). Note that in block 113, connection information (vN), UV coordinates (vtN), and connection information (fN) (for each face) of patch A are shown.

[0129] For such a patch, when the overlap information is omitted in the mesh information and each element is indicated by the minimum number, for example, the original patch information and the mesh information are indicated as follows Fig.17 As in block 121 in FIG. Fig.17 In FIG. 1 , only the mesh information for patch A is shown.

[0130] That is, in this case, the original patch information indicates that the number of original patches is 2, the number of faces in each original patch is 4 and 2, the number of vertices in each original patch is 8 and 4, and the number of UVs in each original patch is 8 and 4. In addition, in the mesh information of patch A, the three-dimensional coordinates (xyz) of each vertex are indicated as vertex information (v1 to v8). In addition, in the mesh information of patch A, the coordinates (uv) in the texture map of each vertex are indicated as UV coordinates (vt1 to vt8). In addition, in the mesh information of patch A, the edge (connection) of each face is indicated as connection information. In addition, the mesh information of patch A may also include attribute information such as normal vectors (normal lines).

[0131] <Method 1-2-2>

[0132] In addition, if Fig.14 As shown in the sixth row from the top of the table in FIG. 1 , the mesh information of the original patch may have a format (hereinafter also referred to as the second format) indicating the vertex coordinates of each face, etc., for each connection (method 1-2-2). Fig.16 In the patch A and patch B, for example, the original patch information and mesh information can be as follows Fig.17 As indicated by box 122 in FIG.

[0133] In this case, the original patch information indicates that the number of original patches is 2, and the number of faces in each original patch is 4 and 2. The number of vertices and UVs in each original patch can be omitted. In addition, in the mesh information of patch A, the three-dimensional coordinates (xyz) of each vertex are indicated as vertex information for each connection. In addition, in the mesh information of patch A, the coordinates (uv) in the texture map of each vertex are indicated as UV coordinates (vt1 to vt8) for each connection. In this case, the arrangement order of the vertices in the vertex information (and the arrangement order of the vertices in the UV coordinates) indicates the connection. That is, the arrangement order corresponds to the connection information. Therefore, the explicit indication of the connection information as in box 121 can be omitted.

[0134] <Method 1-2-3>

[0135] The coded data of the original patch can be stored at any position in the bitstream. Fig.14 As shown in the seventh row from the top of the table, all original facets may be stored in the header and encoded (method 1-2-3). For example, in the above-mentioned first information processing device, the second encoding unit may store the vertex information in the header and encode the vertex information. In addition, in the above-mentioned second information processing device, the second decoding unit may decode the encoded data of the vertex information stored in the header to generate the vertex information.

[0136] <Method 1-2-4>

[0137] In addition, for example, Fig.14 As shown at the bottom of the table in , vertex information can be stored in the displacement video, and other information can be stored in the header and encoded (method 1-2-4). Fig.18 As shown, vertex information can be packed together with displacement vectors in a two-dimensional image 131 (frame image of displacement video). In this case, the three-dimensional coordinates (x, y, z) of each vertex are stored as vertex information. That is, the coordinate value of one component (x component, y component or z component) of one vertex is stored as a pixel value. That is, the absolute information indicating the coordinates is stored in the frame image of the displacement video together with relative information such as displacement vectors.

[0138] Note that in the case where the displacement video includes three components storing the corresponding three components x, y and z of the displacement vector, the corresponding components of the vertex information (three-dimensional coordinates) can be stored in different components. In addition, in the case where the displacement video includes one component, all components of the vertex information can be stored in one component. In this case as well, the coordinate value of one component of one vertex is stored as one pixel value. That is, three pixels are used to store the coordinate value (x component, y component and z component) of one vertex.

[0139] For example, in the above-mentioned first information processing device, the second encoding unit may store the vertex information in a displacement video having a 2D image in which the displacement vector is stored as a frame, and encode the vertex information. Note that the displacement vector is the position difference between the vertex of the subdivided base mesh and the vertex of the original mesh. In addition, the first information processing device may further include a third encoding unit that stores encoding mode control information indicating the applied (selected) encoding mode in a header and encodes the encoding mode control information.

[0140] In addition, in the above-mentioned second information processing device, the second decoding unit can decode the encoded data of the displacement video having the 2D image in which the displacement vector is stored as a frame, and generate vertex information stored in the displacement video. Note that the displacement vector is the position difference between the vertex of the subdivided base mesh and the vertex of the original mesh. In addition, the second information processing device may further include a third decoding unit that decodes the encoded data of the encoding mode control information stored in the header.

[0141] <Comparison of bit amounts>

[0142] Fig.19 The table shown in shows a comparison result of the bit rate (Mbps) of the encoded data between the case where 11 slices are encoded using the first method and the case where the slices are encoded using the second method.

[0143] As shown in the table, in patches with a relatively small number of target faces (e.g., patches 3 to 10), the bit rate of the second method (original patch) is lower than the bit rate of the first method (Draco). Therefore, by applying the second method as in method 1, a decrease in encoding efficiency can be suppressed. In particular, when performing sub-gridding in which there is a high possibility that the number of faces becomes small, by applying the second method as in method 1, a decrease in encoding efficiency can be suppressed.

[0144] On the contrary, in a patch with a large number of faces (e.g., patches 0 to 2), the bit rate of the first mode (Draco) is lower than the bit rate of the second mode (original patch). That is, the bit rate of the second mode is not always lower than the bit rate of the first mode. Therefore, by making the encoding mode selectable as in method 1-1, it is possible to suppress a decrease in encoding efficiency in more cases.

[0145] Any method may be used to select (determine) the encoding method. For example, the selection (determination) may be performed based on the number of faces, or the selection (determination) may be performed by comparing the encoding costs. Note that when selecting the encoding method from pre-prepared options, the number of candidates may be any number. In addition, any encoding method may be included in the options. In addition, the encoding method may be selected (determined) in any data unit. For example, the encoding method may be selected (determined) for each frame, the encoding method may be selected (determined) for each base grid, or the encoding method may be selected (determined) for each subgrid (patch).

[0146] <Combination>

[0147] The above-described various methods (method 1, method 1-1, method 1-1-1, method 1-2, method 1-2-1, method 1-2-2, method 1-2-3, method 1-2-4, and other methods described above) may be appropriately combined and applied.

[0148] <4. Implementation Method>

[0149] <Encoding device>

[0150] The present technology can be applied to an encoding device for encoding a grid. Fig. 20 : is a block diagram showing a configuration example of an encoding device as one mode of an information processing device to which the present technology is applied. Fig. 20 The encoding device 200 shown in FIG. 1 is a device that encodes a lattice. The encoding device 200 encodes a lattice by a method basically similar to the V-DMC described in Non-Patent Document 1.

[0151] In this case, the encoding device 200 encodes the grid by applying the method 1, method 1-1, method 1-1-1, and method 1-2 described above in <3. Encoding without conversion into secondary information>. In addition, the encoding device 200 may apply one or more methods from method 1-2-1 to method 1-2-4. Therefore, the encoding device 200 may also be referred to as a first information processing device.

[0152] Note that Fig. 20 The main processing units, data flows, etc. are shown in FIG. Fig. 20 These are not necessarily all. That is, in the encoding device 200, there may be Fig. 20 A processing unit not shown as a box in the diagram, or there may be Fig. 20 Not shown are processes or data flows as arrows or the like.

[0153] like Fig. 20 As shown, the encoding device 200 includes a pre-processing unit 211, an encoding mode determination unit 212, an intra-frame encoding unit 213, an inter-frame encoding unit 214 and a bit stream generation unit 215.

[0154] The preprocessing unit 211 acquires a mesh, and generates a base mesh and a displacement vector for each patch corresponding to the mesh. The preprocessing unit 211 supplies the base mesh generated for each patch to the encoding method determination unit 212.

[0155] The encoding method determination unit 212 determines the encoding method of the base grid of each patch. For example, the encoding method determination unit 212 determines whether to apply intra-frame encoding or inter-frame encoding, and in the case of determining to apply intra-frame encoding, selects whether to apply the first method or the second method, and generates encoding method control information indicating the selected encoding method (first method or second method). The encoding method determination unit 212 supplies the base grid (of the patch) to be intra-frame encoded and the displacement vector corresponding to the base grid to the intra-frame encoding unit 213 together with the encoding method control information. In addition, the encoding method determination unit 212 supplies the base grid (of the patch) to be inter-frame encoded and the displacement vector corresponding to the base grid to the inter-frame encoding unit 214.

[0156] Note that the encoding method determination unit 212 controls (each processing unit of) the intra-frame encoding unit 213 to encode the base grid by the determined encoding method. For example, in the case where the first method is selected, the encoding method determination unit 212 controls the intra-frame encoding unit 213 to encode the base grid by the first method. In addition, in the case where the second method is selected, the encoding method determination unit 212 controls the intra-frame encoding unit 213 to encode the base grid by the second method.

[0157] The intra-frame encoding unit 213 encodes the supplied base grid by the method specified by the encoding mode control information. In addition, the intra-frame encoding unit 213 stores the supplied displacement vector in the displacement video and encodes the displacement vector. The intra-frame encoding unit 213 supplies the generated encoded data to the bit stream generation unit 215.

[0158] The inter-frame encoding unit 214 performs inter-frame encoding on the supplied base grid. In addition, the inter-frame encoding unit 214 stores the supplied displacement vector in the displacement video and encodes the displacement vector. The inter-frame encoding unit 214 supplies the generated encoded data to the bit stream generation unit 215.

[0159] The bitstream generation unit 215 collects the encoded data of each slice supplied from the intra encoding unit 213 and the inter encoding unit 214 and generates a bitstream. The bitstream generation unit 215 outputs the generated bitstream to the outside of the encoding device 200. The bitstream is supplied to the decoder via any transmission medium or storage medium or both, for example.

[0160] <Intra-frame coding unit>

[0161] Fig.21 It is shown Fig. 20 213 in the intra-frame encoding unit 213. Fig.21 The main processing units, data flows, etc. are shown in FIG. Fig.21 These are not necessarily all. That is, in the intra-frame coding unit 213, there may be Fig.21 Processing parts not depicted as boxes in the Fig.21 No process or data flow is depicted as arrows or the like.

[0162] like Fig.21 As shown, the intra-frame coding unit 213 includes a basic grid coding unit 251, an original patch generation unit 252, a displacement vector correction unit 253, a packing unit 254, a displacement video coding unit 255, a grid reconstruction unit 256, an attribute map correction unit 257, an attribute video coding unit 258, a header coding unit 259 and a combination unit 260. Each processing unit operates under the control of the coding mode determination unit 212.

[0163] The base grid encoding unit 251 acquires the base grid and performs intra-frame encoding by the first method. That is, the base grid encoding unit 251 converts the base grid into "secondary information indicating the relationship between adjacent faces", and encodes the converted information to generate encoded data of the base grid. In the case where the encoding method determination unit 212 selects the first method, the base grid encoding unit 251 supplies the generated encoded data of the base grid to the combination unit 260. That is, the base grid encoding unit 251 can also be referred to as a first encoding unit.

[0164] In addition, the base mesh encoding unit 251 can decode the encoded data to generate (restore) the base mesh. The generated (restored) base mesh includes encoding distortion. The base mesh encoding unit 251 supplies the base mesh to the displacement vector correction unit 253 and the mesh reconstruction unit 256.

[0165] When the encoding mode determination unit 212 selects the second mode, the original patch generation unit 252 obtains the base mesh and generates an original patch (mesh information and original patch information) corresponding to the base mesh. Fig.17 As described above, the original patch generation unit 252 may generate an original patch in the first format or an original patch in the second format.

[0166] Note that the vertex information may be stored in the header or may be stored in the displacement video. In the case where the vertex information is stored in the header, the original patch generation unit 252 supplies all data of the generated original patch to the header encoding unit 259. In addition, in the case where the vertex information is stored in the displacement video, the original patch generation unit 252 supplies the vertex information to the displacement vector correction unit 253, and supplies other data to the header encoding unit 259.

[0167] The displacement vector correction unit 253 acquires the base mesh including the encoding distortion supplied from the base mesh encoding unit 251. In addition, in the case where the encoding method determination unit 212 selects the second method, the displacement vector correction unit 253 acquires the vertex information supplied from the original patch generation unit 252. In addition, the displacement vector correction unit 253 acquires the displacement vector supplied from the encoding method determination unit 212.

[0168] The displacement vector correction unit 253 subdivides the acquired base mesh and corrects the displacement vector using the subdivided base mesh. Note that this correction may be omitted. In addition, in the case where vertex information is supplied from the original patch generation unit 252, the displacement vector correction unit 253 may correct the vertex information using the subdivided base mesh. The displacement vector correction unit 253 supplies the displacement vector and vertex information to the packing unit 254.

[0169] The packing unit 254 obtains the displacement vector supplied from the displacement vector correction unit 253, and packs the displacement vector into a two-dimensional image (also referred to as a frame image) of the current frame. At this time, the packing unit 254 may perform coefficient transformation (e.g., wavelet transformation) on the displacement vector and pack the transformation coefficient into the two-dimensional image. In addition, the packing unit 254 may quantize the displacement vector (or transformation coefficient) and pack the quantized coefficient into the two-dimensional image. The packing unit 254 supplies the two-dimensional image in which the displacement vector or information corresponding thereto is packed to the displacement video encoding unit 255.

[0170] In addition, in the case where the encoding method determination unit 212 selects the second method, the packing unit 254 may obtain the vertex information supplied from the displacement vector correction unit 253 and pack it in the two-dimensional image (also referred to as the frame image) of the current frame. In this case, the packing unit 254 supplies the two-dimensional image in which the displacement vector or the information corresponding thereto and the vertex information or the information corresponding thereto are packed to the displacement video encoding unit 255.

[0171] The displacement video encoding unit 255 obtains the two-dimensional image supplied from the packing unit 254. The displacement video encoding unit 255 uses the two-dimensional image as a frame image, and encodes (intra-encodes) the image into a motion image (displacement video) using a 2D codec. That is, in the case where vertex information is included in the frame image of the displacement video as a pixel value, the displacement video encoding unit 255 does not convert the vertex information into "secondary information indicating the relationship between adjacent faces of the base mesh" but encodes the vertex information. Therefore, it can also be said that the displacement video encoding unit 255 encodes the vertex information in a second manner. Therefore, the displacement video encoding unit 255 can also be referred to as a second encoding unit. The displacement video encoding unit 255 supplies the encoded data of the generated displacement video to the mesh reconstruction unit 256 and the combination unit 260.

[0172] The mesh reconstruction unit 256 decodes the encoded data of the displacement video supplied from the displacement video encoding unit 255, and obtains a displacement vector by performing unpacking or the like from the two-dimensional image. In addition, the mesh reconstruction unit 256 subdivides the base mesh (including encoding distortion) supplied from the base mesh encoding unit 251, and reconstructs the mesh by applying the obtained displacement vector. The mesh includes encoding distortion. The mesh reconstruction unit 256 supplies the reconstructed mesh to the attribute map correction unit 257.

[0173] The property map correction unit 257 acquires a property map such as texture, and corrects the property map using the mesh (including encoding distortion) supplied from the mesh reconstruction unit 256. The property map correction unit 257 supplies the corrected property map to the property video encoding unit 258. Note that this correction may be omitted.

[0174] The attribute video encoding unit 258 uses the attribute map supplied from the attribute map correction unit 257 as a frame image, and encodes the image into a moving image (attribute video). The attribute video encoding unit 258 supplies the generated encoded data of the attribute video to the combining unit 260.

[0175] The header encoding unit 259 encodes the header. For example, in the case where the encoding method determination unit 212 selects the second method, the header encoding unit 259 obtains the original patch supplied from the original patch generation unit 252, stores the original patch in the header, and encodes the original patch by the second method. For example, the header encoding unit 259 can encode the original patch using the correlation between the vertices of the base mesh. For example, the header encoding unit 259 can perform differential encoding on the original patch. In addition, the header encoding unit 259 can perform arithmetic encoding on the original patch. In addition, the header encoding unit 259 can perform fixed bit encoding on the original patch. That is, the header encoding unit 259 does not convert the base mesh into "secondary information indicating the relationship between adjacent faces" and encodes the base mesh. Therefore, the header encoding unit 259 can also be referred to as a second encoding unit. Note that the original patch may or may not include vertex information.

[0176] In addition, the header encoding unit 259 obtains the encoding method control information supplied from the encoding method determination unit 212, stores the information in the header, and encodes the information. Therefore, for example, in the case where the vertex information is stored in the displacement video, the header encoding unit 259 can also be called a third encoding unit. The header encoding unit 259 supplies the encoded data of the generated header to the combining unit 260.

[0177] The combining unit 260 combines (multiplexes) the supplied encoded data of the header, the encoded data of the base grid, the encoded data of the displacement video, and the encoded data of the attribute video. The combining unit 260 supplies the combined encoded data to the bit stream generating unit 215.

[0178] With such a configuration, the encoding device 200 can encode the base grid as described above in <3. Encoding without conversion into secondary information>, so that a decrease in encoding efficiency can be suppressed.

[0179] <Encoding Process Flow>

[0180] Will refer to Fig. 22 The flowchart describes an example of the flow of the encoding process performed by the encoding device 200.

[0181] When the encoding process starts, in step S201 , the pre-processing unit 211 generates a base mesh and a displacement vector for each patch.

[0182] In step S202, the encoding method determination unit 212 determines the encoding method of the base mesh of each patch generated in step S201. For example, the encoding method determination unit 212 selects the first method or the second method as the encoding method of the base mesh of each patch generated in step S201. Then, the encoding method determination unit 212 generates encoding method control information indicating the selected encoding method.

[0183] In step S203 , the intra encoding unit 213 performs intra encoding processing on the patch to be subjected to intra encoding, and encodes the base mesh and the displacement vector.

[0184] In step S204, the inter-coding unit 214 performs an inter-coding process on the patch to be inter-coded, and encodes the base mesh and the displacement vector.

[0185] In step S205 , the bit stream generation unit 215 combines (multiplexes) the encoded data obtained by the process of step S203 and the encoded data obtained by the process of step S204 to generate a bit stream.

[0186] When the process of step S205 ends, the encoding process ends.

[0187] <Flow of Intra-frame Coding Processing>

[0188] Next, we will refer to Fig.23 The flowchart in the description is in Fig. 22 An example of the flow of intra-frame encoding processing performed in step S203 in .

[0189] When the intra-coding process starts, in step S251, the base grid encoding unit 251 determines whether to encode the original patch. In the case where it is determined not to encode the original patch, the process proceeds to step S252.

[0190] In step S252, the base mesh encoding unit 251 encodes the base mesh in the first manner. That is, the base mesh encoding unit 251 converts the base mesh into “secondary information indicating the relationship between adjacent faces of the base mesh” to encode the converted information.

[0191] When the process of step S252 ends, the process proceeds to step S254. In addition, when it is determined in step S251 that the original patch is to be encoded, the process proceeds to step S253.

[0192] In step S253, the original patch generation unit 252 converts the base mesh and generates an original patch as information for transmission. When the process of step S253 ends, the process proceeds to step S254.

[0193] In step S254, the displacement vector correction unit 253 appropriately corrects the displacement vector using the base grid including the encoding distortion.

[0194] In step S255, the packing unit 254 packs the appropriately corrected displacement vector into the frame image of the displacement video.

[0195] In step S256 , the displacement video encoding unit 255 encodes the displacement video.

[0196] In step S257 , the mesh reconstruction unit 256 reconstructs a mesh using the base mesh including the coding distortion and the displacement vector.

[0197] In step S258 , the attribute map correction unit 257 uses the reconstructed mesh to appropriately correct the attribute map.

[0198] In step S259, the attribute video encoding unit 258 encodes the appropriately corrected attribute map.

[0199] In step S260, the header encoding unit 259 encodes the header. Note that, in the case where it is determined in step S251 to encode the original patch, the header encoding unit 259 stores the original patch generated in step S253 in the header and encodes the original patch. In addition, the header encoding unit 259 stores the encoding mode control information generated by the process of step S202 in the header and encodes it.

[0200] In step S261 , the combining unit 260 multiplexes the encoded data generated in each process to generate a bit stream.

[0201] In step S262, the base grid encoding unit 251 determines whether all patches have been processed. In the case where it is determined that there are unprocessed patches, the process returns to step S251 and subsequent processing is performed. In addition, in the case where it is determined in step S262 that all patches have been processed, the intra-frame encoding process ends, and the process returns to step S263. Fig. 22 .

[0202] By performing each process in this manner, the encoding device 200 can encode the base grid as described above in <3. Encoding without conversion into secondary information>, so that a decrease in encoding efficiency can be suppressed.

[0203] <Decoding device>

[0204] The present technology can be applied to a decoding device that decodes coded data of a grid. Fig.24 : is a block diagram showing an example of the configuration of a decoding device as one aspect of an information processing device to which the present technology is applied. Fig.24 The decoding device 300 shown in FIG. 1 is a device that decodes the encoded data of the mesh. The decoding device 300 decodes the encoded data of the mesh by a method basically similar to that of the V-DMC described in Non-Patent Document 1.

[0205] At this time, the decoding device 300 decodes the coded data of the grid by applying the method 1, method 1-1, method 1-1-1, and method 1-2 described above in <3. Encoding without conversion into secondary information>. In addition, the decoding device 300 can apply one or more of the methods 1-2-1 to 1-2-4 described above. Therefore, the decoding device 300 can also be called a second information processing device.

[0206] That is, the decoding device 300 is Fig. 20 The decoding device corresponds to the encoding device 200 in the embodiment, and can decode the bit stream generated by the encoding device 200 to reconstruct the grid.

[0207] Note that although Fig.24 The main elements such as processing units and data flow are shown, but Fig.24 Those elements shown in do not necessarily include all elements. That is, in the decoding device 300, there may be Fig.24 A processing unit is not shown as a box, or may exist Fig.24 Not shown are processes or data flows as arrows or the like.

[0208] like Fig.24 As shown, the decoding device 300 includes a demultiplexing unit 311, a header decoding unit 312, an intra-frame decoding unit 313, an inter-frame decoding unit 314 and a combining unit 315.

[0209] A bit stream generated by an encoding device (for example, the encoding device 200 ) that encodes a trellis by the V-DMC method is supplied to the decoding device 300 .

[0210] The demultiplexing unit 311 demultiplexes the bitstream and extracts each coded data included in the bitstream. For example, the demultiplexing unit 311 extracts the coded data of the header from the bitstream and supplies the coded data to the header decoding unit 312. In addition, the demultiplexing unit 311 supplies the coded data of the patch in which the base grid is intra-coded to the intra-decoding unit 313 based on the control of the header decoding unit 312, and supplies the coded data of the patch in which the base grid is inter-coded to the inter-decoding unit 314.

[0211] The header decoding unit 312 decodes the encoded data of the header supplied from the demultiplexing unit 311, and appropriately supplies necessary information to the demultiplexing unit 311, the intra decoding unit 313, and the inter decoding unit 314. In addition, the header decoding unit 312 controls the operations of the demultiplexing unit 311, the intra decoding unit 313, and the inter decoding unit 314.

[0212] In addition, the header decoding unit 312 decodes the coded data of the coding mode control information and generates (restores) the coding mode control information. Therefore, the header decoding unit 312 can also be called a third decoding unit. The header decoding unit 312 supplies the coding mode control information to the intra-frame decoding unit 313.

[0213] In addition, in the case where the encoding mode control information specifies the second mode, the header decoding unit 312 obtains the encoded data of the original patch stored in the header, decodes the encoded data by the second mode, and generates (restores) the original patch. That is, the header decoding unit 312 generates a base mesh that is not converted into secondary information indicating the relationship between adjacent faces of the base mesh. Therefore, the header decoding unit 312 can also be referred to as a second decoding unit. The encoded data can be encoded using the correlation between the vertices of the base mesh. For example, the header decoding unit 312 can perform differential decoding on the encoded data of the original patch that has been differentially encoded. In addition, the header decoding unit 312 can perform arithmetic decoding on the encoded data of the original patch after arithmetic encoding. In addition, the header decoding unit 312 can perform fixed bit decoding on the encoded data of the original patch that has been subjected to fixed bit encoding. The header decoding unit 312 supplies the generated (restored) original patch to the intra-frame decoding unit 313.

[0214] Based on the control of the header decoding unit 312, the intra decoding unit 313 acquires the encoded data of the intra-encoded patch from the demultiplexing unit 311 and performs intra decoding. At this time, the intra decoding unit 313 applies the encoding method specified by the encoding method control information supplied from the header decoding unit 312, and decodes the encoded data. The intra decoding unit 313 supplies the mesh and attribute map of the patch generated (restored) by decoding to the combining unit 315.

[0215] Based on the control of the header decoding unit 312 , the inter-frame decoding unit 314 obtains the encoded data of the inter-frame encoded patch from the demultiplexing unit 311 , performs inter-frame decoding, generates (restores) the mesh and attribute map of the patch, and supplies the mesh and attribute map to the combining unit 315 .

[0216] The combining unit 315 combines the meshes and attribute maps of the corresponding patches supplied from the intra decoding unit 313 and the inter decoding unit 314 to generate and output the mesh and attribute map of the entire object.

[0217] <Intra-frame decoding unit>

[0218] Fig.25 It is shown Fig.24 A block diagram of a main configuration example of the intra-frame decoding unit 313 in FIG. Fig.25 The main processing units, data flows, etc. are shown in FIG. Fig.25 Those shown in are not necessarily all. That is, in the intra-frame decoding unit 313, there may be Fig.25 Processing that is not depicted as a box, or may exist Fig.25 No process or data flow is depicted as arrows or the like.

[0219] like Fig.25 As shown, the intra decoding unit 313 includes a displacement video decoding unit 351, a depacketizing unit 352, a base grid decoding unit 353, a reconstruction unit 354, a subdivision unit 355, a displacement vector application unit 356, and an attribute video decoding unit 357. Each processing unit operates under the control of the header decoding unit 312.

[0220] The displacement video decoding unit 351 decodes the encoded data of the displacement video supplied from the demultiplexing unit 311, and generates (restores) the displacement video. Note that the original patch (vertex information) can be stored in the displacement video. In this case, it can also be said that the displacement video decoding unit 351 decodes the encoded data of the displacement video, and generates vertex information stored in the displacement video. Therefore, the displacement video decoding unit 351 can also be referred to as a second decoding unit. The displacement video decoding unit 351 supplies the generated displacement video (the current frame) to the unpacking unit 352.

[0221] The unpacking unit 352 unpacks (extracts) a displacement vector from the current frame of the displacement video supplied from the displacement video decoding unit 351. At this time, the unpacking unit 352 may unpack the transform coefficients and perform coefficient transform (e.g., wavelet transform) on the transform coefficients to obtain the displacement vector. In addition, the unpacking unit 352 may unpack the quantization coefficients and inversely quantize the quantization coefficients to obtain the displacement vector. In addition, the unpacking unit 352 may unpack the quantization coefficients, inversely quantize the quantization coefficients to obtain the transform coefficients, and perform coefficient transform (e.g., wavelet transform) on the transform coefficients to obtain the displacement vector. The unpacking unit 352 supplies the unpacked displacement vector to the displacement vector application unit 356.

[0222] Note that, in the case where the original patch (vertex information) is stored in the displacement video, the unpacking unit 352 also unpacks the vertex information, and supplies the vertex information to the reconstruction unit 354 .

[0223] The base mesh decoding unit 353 acquires the encoding mode control information supplied from the header decoding unit 312. In the case where the encoding mode control information specifies the first mode, the base mesh decoding unit 353 decodes the encoded data of the base mesh supplied from the demultiplexing unit 311 by the first mode, and generates (restores) the base mesh. That is, the base mesh decoding unit 353 decodes the encoded data of the base mesh, generates secondary information indicating the relationship between adjacent faces of the base mesh, and converts the generated secondary information into vertex information and connection information of the base mesh. Therefore, the base mesh decoding unit 353 can also be referred to as a first decoding unit. The base mesh decoding unit 353 supplies the generated base mesh to the subdivision unit 355.

[0224] The reconstruction unit 354 obtains the encoding mode control information supplied from the header decoding unit 312. In the case where the encoding mode control information specifies the second mode, the reconstruction unit 354 obtains the original patch (the original patch included in the header) supplied from the header decoding unit 312. Note that, as described above, the vertex information constituting the original patch can be stored in the header, or can be stored in the displacement video. In the case where the vertex information is stored in the displacement video, the reconstruction unit 354 also obtains the vertex information supplied from the unpacking unit 352. The reconstruction unit 354 uses the original patch obtained in this manner to reconstruct the base mesh. That is, the reconstruction unit 354 converts the original patch as information for transmission into the base mesh. Note that the data format of the original patch can be the first format, the second format or other formats. The reconstruction unit 354 supplies the reconstructed base mesh to the subdivision unit 355.

[0225] The subdivision unit 355 subdivides the base mesh supplied from the base mesh decoding unit 353, and supplies the subdivided base mesh to the displacement vector application unit 356. Furthermore, the subdivision unit 355 subdivides the base mesh supplied from the reconstruction unit 354, and supplies the subdivided base mesh to the displacement vector application unit 356.

[0226] The displacement vector application unit 356 applies the displacement vector supplied from the unpacking unit 352 to the vertices of the subdivided base mesh supplied from the subdividing unit 355 to reconstruct the mesh. In this specification, the reconstructed mesh is also referred to as a decoded mesh. That is, it can also be said that the displacement vector application unit 356 generates a decoded mesh. The displacement vector application unit 356 supplies the generated decoded mesh to the combining unit 315.

[0227] The attribute video decoding unit 357 decodes the encoded data of the attribute video supplied from the demultiplexing unit 311 and generates (the current frame of) the attribute video. The attribute video decoding unit 357 supplies the generated current frame of the attribute video, ie, the attribute map corresponding to the decoded mesh, to the combining unit 315.

[0228] With such a configuration, the decoding device 300 can decode the encoded data of the base grid as described above in <3. Encoding without conversion into secondary information>, so that a decrease in encoding efficiency can be suppressed.

[0229] <Decoding Process Flow>

[0230] Will refer to Fig.26 The flowchart in describes an example of the flow of the decoding process performed by the decoding device 300.

[0231] When the decoding process starts, in step S301 , the demultiplexing unit 311 demultiplexes a bit stream and extracts various encoded data.

[0232] In step S302, the header decoding unit 312 decodes the encoded data of the header extracted in step S301, and generates (restores) the information stored in the header. For example, the header decoding unit 312 decodes the encoded data of the encoding mode control information, and generates (restores) the encoding mode control information. In addition, in the case where the encoding mode control information specifies the second mode, the header decoding unit 312 obtains the encoded data of the original face patch stored in the header, decodes the encoded data by the second mode, and generates (restores) the original face patch. That is, the header decoding unit 312 generates a base mesh that is not converted into secondary information indicating the relationship between adjacent faces of the base mesh.

[0233] In step S303 , the intra decoding unit 313 performs an intra decoding process, and performs intra decoding on the encoded data of the intra-encoded slice.

[0234] In step S304 , the inter decoding unit 314 performs an inter decoding process, and performs inter decoding on the encoded data of the inter-encoded slice.

[0235] In step S305 , the combining unit 315 combines the 3D data of each patch generated by the processing of steps S303 and S304 .

[0236] When the process of step S305 ends, the decoding process ends.

[0237] <Flow of Intra-frame Decoding Process>

[0238] Next, we will refer to Fig. 27 The flowchart in the description is in Fig.26 An example of the flow of the intra-frame decoding process performed in step S303 in .

[0239] When the intra-frame decoding process starts, in step S351, the displacement video decoding unit 351 decodes the encoded data of the displacement video and generates (restores) the displacement video. Note that the original patch (vertex information) can be stored in the displacement video. In this case, it can also be said that the displacement video decoding unit 351 decodes the encoded data of the displacement video and generates the vertex information stored in the displacement video.

[0240] In step S352, the unpacking unit 352 unpacks (extracts) the displacement vector from the current frame of the displacement video. At this time, the unpacking unit 352 can unpack the transform coefficients and perform coefficient transformation (e.g., wavelet transform) on the transform coefficients to obtain the displacement vector. In addition, the unpacking unit 352 can unpack the quantization coefficients and inverse quantize the quantization coefficients to obtain the displacement vector. In addition, the unpacking unit 352 can unpack the quantization coefficients, inverse quantize the quantization coefficients to obtain the transform coefficients, and perform coefficient transformation (e.g., wavelet transform) on the transform coefficients to obtain the displacement vector. Note that in the case where the original face (vertex information) is stored in the displacement video, the unpacking unit 352 also unpacks the vertex information.

[0241] In step S353, the base mesh decoding unit 353 determines whether to decode the original patch. In the case where the encoding method control information specifies the first method, it is determined not to decode the original patch, and the process proceeds to step S354.

[0242] In step S354, the base mesh decoding unit 353 decodes the encoded data of the base mesh extracted from the bit stream by the demultiplexing in step S301 by the first method, and generates (restores) the base mesh. That is, the base mesh decoding unit 353 decodes the encoded data of the base mesh, generates secondary information indicating the relationship between adjacent faces of the base mesh, and converts the generated secondary information into vertex information and connection information of the base mesh.

[0243] When the process of step S354 ends, the process proceeds to step S356. Furthermore, in step S353, in the case where the encoding method control information specifies the second method, it is determined that the original face is decoded, and the process proceeds to step S355.

[0244] In step S355, the reconstruction unit 354 obtains the original patch (the original patch stored and sent in the header) generated (restored) by the processing of step S302. Note that the vertex information constituting the original patch can be stored in the header or can be stored in the displacement video. In the case where the vertex information is stored in the displacement video, the reconstruction unit 354 also obtains the unpacked vertex information from the displacement video. The reconstruction unit 354 uses the original patch obtained in this way to reconstruct the basic mesh. That is, the reconstruction unit 354 converts the original patch as information for transmission into the basic mesh. Note that the data format of the original patch can be the first format, the second format or other format. When the processing of step S355 is completed, the processing proceeds to step S356.

[0245] In step S356 , the subdivision unit 355 subdivides the base mesh obtained in step S354 or step S355 .

[0246] In step S357 , the displacement vector application unit 356 applies the displacement vectors unpacked from the displacement video to the vertices of the subdivided base mesh, and reconstructs the mesh (generates a decoded mesh).

[0247] In step S358 , the attribute video decoding unit 357 decodes the encoded data of the attribute video extracted from the bit stream by the demultiplexing in step S301 , and generates (the current frame of) the attribute video.

[0248] In step S359, the base grid decoding unit 353 determines whether all the patches have been processed. In the case where it is determined that there are unprocessed patches, the process returns to step S351 and the subsequent process is performed. In addition, in the case where it is determined in step S359 that all the patches have been processed, the intra-frame decoding process ends, and the process returns to step S352. Fig.26 .

[0249] By performing each process in this manner, the decoding device 300 can decode the encoded data of the base grid as described above in <3. Encoding without conversion into secondary information>, so that a decrease in encoding efficiency can be suppressed.

[0250] <5. Additional Notes>

[0251] <Computer>

[0252] The above series of processing can be performed by hardware or software. In the case of performing a series of processing by software, the program included in the software is installed on the computer. Here, examples of computers include, for example, computers built into dedicated hardware, general-purpose personal computers that can perform various functions by installing various programs, etc.

[0253] Fig.28 : is a block diagram showing a hardware configuration example of a computer that executes the series of processes described above according to a program.

[0254] exist Fig.28 In a computer 900 shown in , a central processing unit (CPU) 901 , a read only memory (ROM) 902 , and a random access memory (RAM) 903 are connected to each other via a bus 904 .

[0255] The bus 904 is also connected to an input / output interface 910. To the input / output interface 910, an input unit 911, an output unit 912, a storage unit 913, a communication unit 914, and a drive 915 are connected.

[0256] The input unit 911 includes, for example, a keyboard, a mouse, a microphone, a touch panel, an input terminal, etc. The output unit 912 includes, for example, a display, a speaker, an output terminal, etc. The storage unit 913 includes, for example, a hard disk, a RAM disk, a nonvolatile memory, etc. The communication unit 914 includes, for example, a network interface. The drive 915 drives a removable medium 921, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0257] In the computer configured as described above, for example, the series of processes described above are performed by the CPU 901 loading and executing the program recorded in the storage unit 913 into the RAM 903 via the input / output interface 910 and the bus 904. The RAM 903 also appropriately stores data necessary for the CPU 901 to perform various processes.

[0258] The program executed by the computer can be applied by being recorded on the removable medium 921 as, for example, a package medium, etc. In this case, by attaching the removable medium 921 to the drive 915 , the program can be installed in the storage unit 913 via the input / output interface 910 .

[0259] In addition, the program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting. In this case, the program can be received by the communication unit 914 and installed in the storage unit 913.

[0260] Furthermore, the program may be installed in the ROM 902 and the storage unit 913 in advance.

[0261] <Targets of application of this technology>

[0262] The present technology can be applied to any configuration. For example, the present technology can be applied to various electronic devices.

[0263] In addition, for example, the present technology can also be implemented as a partial configuration of a device, such as a processor as a system large-scale integrated circuit (LSI) or the like (e.g., a video processor), a module using multiple processors or the like (e.g., a video module), a unit using multiple modules or the like (e.g., a video unit), or a collection obtained by further adding other functions to the unit (e.g., a video collection).

[0264] In addition, for example, the present technology can also be applied to a network system including a plurality of devices. For example, the present technology can be implemented as cloud computing that is collaboratively shared and processed by a plurality of devices through a network. For example, the present technology can be implemented in a cloud service that provides image (moving image) related services to any terminal such as a computer, an audio-visual (AV) device, a portable information processing terminal, or an Internet of Things (IoT) device.

[0265] Note that in this specification, a system means a collection of multiple components (devices, modules (parts), etc.), and it is not important whether all the components are in the same housing. Therefore, multiple devices stored in different housings and connected via a network and a device in which multiple modules are stored in one housing are both systems.

[0266] <Fields and applications of this technology>

[0267] The system, device, processing unit, etc. applying this technology can be used in any field such as transportation, medical care, crime prevention, agriculture, animal husbandry, mining, beauty care, factory, household appliances, weather and nature monitoring. In addition, its application is also arbitrary.

[0268] <Others>

[0269] Note that in this specification, a "flag" is information for identifying multiple states, and includes not only information for identifying two states of true (1) and false (0), but also information capable of identifying three or more states. Therefore, the value that a "flag" can take can be binary or ternary or more, such as 1 / 0. That is, the number of bits forming the "flag" is any number, and can be one bit or multiple bits. In addition, it is assumed that the identification information (including the flag) includes not only its identification information in the bit stream, but also the difference information of the identification information relative to a certain reference information in the bit stream. Therefore, in this specification, "flag" and "identification information" include not only its information, but also the difference information relative to the reference information.

[0270] In addition, various information (e.g., metadata) related to the coded data (bitstream) can be sent or recorded in any form as long as the information is associated with the coded data. Here, the term "association" means, for example, that when processing one data, other data is allowed to be used (linked). That is, data associated with each other can be collected as one data or can be separate data. For example, information associated with coded data (image) can be sent on a transmission path different from the transmission path of the coded data (image). In addition, for example, information associated with coded data (image) can be recorded in a recording medium different from the recording medium of the coded data (image) (or another recording area of ​​the same recording medium). Note that the "association" may not be the entire data, but a part of the data. For example, an image and information corresponding to the image can be associated with each other in any unit such as multiple frames, one frame, or a part within a frame.

[0271] Note that in this specification, for example, terms such as "combine", "multiplex", "add", "integrate", "include", "store", "put in", "introduce" and "insert" mean, for example, combining multiple objects into one, for example, combining encoded data and metadata into one data, and mean a method of the above-mentioned "association".

[0272] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications may be made without departing from the scope of the present technology.

[0273] For example, a configuration described as one device (or processing unit) may be divided and configured as a plurality of devices (or processing units). Conversely, a configuration described as a plurality of devices (or processing units) may be collectively configured as one device (or processing unit). Furthermore, it goes without saying that a configuration other than the above configuration may be added to the configuration of each device (or each processing unit). Furthermore, in the case where the configuration and operation as the entire system are substantially the same, a portion of the configuration of a certain device (or processing unit) may be included in the configuration of another device (or another processing unit).

[0274] In addition, for example, the above-mentioned program can be executed in an arbitrary device. In this case, it is only necessary that the device has necessary functions (functional blocks, etc.) and obtains necessary information.

[0275] In addition, for example, each step in one flowchart may be performed by one device, or may be performed by being shared by a plurality of devices. In addition, when a plurality of processes are included in one step, the plurality of processes may be performed by one device, or may be shared and performed by a plurality of devices. In other words, the plurality of processes included in one step may be performed as a plurality of steps. Conversely, the processes described as a plurality of steps may also be performed together as one step.

[0276] In addition, for example, in a program executed by a computer, the processing of the steps describing the program may be performed in a time series order in the order described in this specification, or may be performed in parallel or individually at a desired timing such as when a call is made. That is, as long as there is no contradiction, the processing of each step may be performed in an order different from the above order. In addition, the processing of the steps describing the program may be performed in parallel with the processing of another program, or may be performed in combination with the processing of another program.

[0277] In addition, for example, multiple technologies related to the present technology can be independently implemented as a single entity as long as there is no contradiction. It goes without saying that any multiple technologies in the present technology can be implemented in combination. For example, part or all of the present technology described in any embodiment can be implemented in combination with part or all of the present technology described in other embodiments. In addition, part or all of any of the above-mentioned present technologies can be implemented together with another technology not described above.

[0278] Note that the present technology may also have the following configurations.

[0279] (1) An information processing device comprising:

[0280] an encoding method determination unit that selects the first method or the second method as an encoding method of a base grid to be subjected to intra-frame encoding;

[0281] a first encoding unit that, when the first mode is selected as the encoding mode, converts vertex information indicating positions of vertices constituting a base mesh and connection information indicating connections between vertices constituting the base mesh into secondary information indicating a relationship between adjacent faces of the base mesh and encodes the secondary information; and

[0282] A second encoding unit, which, when the second mode is selected as the encoding mode, encodes the vertex information and the connection information without converting the vertex information and the connection information into secondary information, wherein:

[0283] The base mesh is a mesh generated by removing vertices from an original mesh to be encoded consisting of vertices and connections representing a three-dimensional structure of an object, and is coarser than the original mesh.

[0284] (2) The information processing device according to (1), wherein:

[0285] The second encoding unit encodes vertex information using correlations between vertices of the base mesh.

[0286] (3) The information processing device according to (2), wherein:

[0287] The second encoding unit performs differential encoding on the vertex information.

[0288] (4) The information processing device according to (2), wherein:

[0289] The second encoding unit performs arithmetic encoding on the vertex information.

[0290] (5) The information processing device according to (1), wherein:

[0291] The second encoding unit performs fixed-bit encoding on the vertex information.

[0292] (6) The information processing device according to any one of (1) to (5), wherein:

[0293] The second encoding unit stores the vertex information in the header and encodes the vertex information.

[0294] (7) The information processing device according to (6), wherein:

[0295] The second encoding unit further stores encoding mode control information indicating the selected encoding mode in the header and encodes the encoding mode control information.

[0296] (8) The information processing device according to any one of (1) to (4), wherein:

[0297] The second encoding unit stores the vertex information in a displacement video having a 2D image in which the displacement vector is stored as a frame and encodes the vertex information, and

[0298] The displacement vector is the difference between the position of the vertices of the subdivided base mesh and the vertices of the original mesh.

[0299] (9) The information processing device according to (8), further comprising:

[0300] A third encoding unit stores encoding mode control information indicating the selected encoding mode in a header and encodes the encoding mode control information.

[0301] (10) The information processing device according to any one of (1) to (9), further comprising:

[0302] The original patch generation unit converts the basic mesh and generates the original patch as the transmission information, wherein:

[0303] When the second mode is selected as the encoding mode, the second encoding unit encodes the original facet without converting the original facet into secondary information.

[0304] (11) The information processing device according to (10), wherein:

[0305] The original patch includes mesh information and original patch information. The mesh information includes vertex information and connection information. The original patch information includes meta information related to the basic mesh.

[0306] (12) The information processing device according to (11), wherein:

[0307] Overlapping information in vertex information is omitted and vertices are indicated by the minimum number.

[0308] (13) The information processing device according to (11), wherein:

[0309] The vertex information indicates a vertex of each connection in the connection.

[0310] (14) The information processing device according to any one of (11) to (13), wherein:

[0311] The original patch information includes information indicating the number of original patches and information indicating the number of faces in the original patch.

[0312] (15) An information processing method comprising:

[0313] selecting a first mode or a second mode as an encoding mode of a base grid to be subjected to intra-frame encoding;

[0314] In a case where the first mode is selected as the encoding mode, converting vertex information indicating positions of vertices constituting a base mesh and connection information indicating connections between vertices constituting the base mesh into secondary information indicating relationships between adjacent faces of the base mesh and encoding the secondary information; and

[0315] When the second mode is selected as the encoding mode, the vertex information and the connection information are encoded without converting them into secondary information, wherein:

[0316] The base mesh is a mesh generated by removing vertices from an original mesh to be encoded consisting of vertices and connections representing a three-dimensional structure of an object, and is coarser than the original mesh.

[0317] (21) An information processing device comprising:

[0318] a first decoding unit that, when the encoding mode control information indicates the first mode as the encoding mode of the intra-coded base mesh, decodes the encoded data of the intra-coded base mesh, generates secondary information indicating a relationship between adjacent faces of the base mesh, and converts the generated secondary information into vertex information and connection information of the base mesh; and

[0319] A second decoding unit, which decodes the encoded data of the base mesh and generates vertex information and connection information that are not converted into secondary information when the encoding mode control information indicates the second mode as the encoding mode, wherein

[0320] The base mesh is a mesh generated by removing vertices from an original mesh to be encoded consisting of vertices and connections representing the three-dimensional structure of an object, and is coarser than the original mesh.

[0321] Vertex information is information indicating the positions of vertices constituting the base mesh.

[0322] The connection information is information indicating the connection between the vertices constituting the base mesh.

[0323] The first method is an encoding method of converting vertex information and connection information into secondary information and encoding the secondary information, and

[0324] The second method is a coding method of coding the vertex information and the connection information without converting them into secondary information.

[0325] (22) The information processing device according to (21), wherein:

[0326] The second decoding unit decodes the encoded data of the vertex information encoded using the correlation between the vertices of the base mesh.

[0327] (23) The information processing device according to (22), wherein:

[0328] The second decoding unit decodes the encoded data of the vertex information that has been differentially encoded.

[0329] (24) The information processing device according to (22), wherein:

[0330] The second decoding unit decodes the encoded data of the vertex information that has been arithmetically encoded.

[0331] (25) The information processing device according to (21), wherein:

[0332] The second decoding unit decodes the vertex information that has been subjected to the fixed-bit encoding.

[0333] (26) The information processing device according to any one of (21) to (25), wherein:

[0334] The second decoding unit decodes the encoded data of the vertex information stored in the header, and generates the vertex information.

[0335] (27) The information processing device according to (26), wherein:

[0336] The second decoding unit further decodes the encoded data of the encoding mode control information stored in the header.

[0337] (28) The information processing device according to any one of (21) to (23), wherein:

[0338] The second decoding unit decodes the encoded data of the displacement video having the 2D image in which the displacement vector is stored as a frame, and generates vertex information stored in the displacement video, and

[0339] The displacement vector is the difference between the position of the vertices of the subdivided base mesh and the vertices of the original mesh.

[0340] (29) The information processing device according to (28), further comprising:

[0341] A third decoding unit decodes the encoded data of the encoding mode control information stored in the header.

[0342] (30) The information processing device according to any one of (21) to (29), wherein:

[0343] The second decoding unit decodes the encoded data of the base grid and generates an original patch as information in a transmission form of the base grid, and

[0344] The information processing device further includes a reconstruction unit for reconstructing a base mesh using the original patch.

[0345] (31) The information processing device according to (30), wherein:

[0346] The original patch includes mesh information and original patch information. The mesh information includes vertex information and connection information. The original patch information includes meta information related to the basic mesh.

[0347] (32) The information processing device according to (31), wherein:

[0348] Overlapping information in vertex information is omitted and vertices are indicated by the minimum number.

[0349] (33) The information processing device according to (31), wherein:

[0350] The vertex information indicates a vertex of each connection in the connection.

[0351] (34) The information processing device according to any one of (31) to (33), wherein:

[0352] The original patch information includes information indicating the number of original patches and information indicating the number of faces in the original patch.

[0353] (35) An information processing method comprising:

[0354] In a case where the encoding mode control information indicates the first mode as the encoding mode of the intra-coded base mesh, decoding the encoded data of the intra-coded base mesh, generating secondary information indicating a relationship between adjacent faces of the base mesh, and converting the generated secondary information into vertex information and connection information of the base mesh; and

[0355] In a case where the encoding mode control information indicates the second mode as the encoding mode, the encoded data of the base mesh is decoded, and vertex information and connection information that are not converted into secondary information are generated, wherein

[0356] The base mesh is a mesh generated by removing vertices from an original mesh to be encoded consisting of vertices and connections representing the three-dimensional structure of an object, and is coarser than the original mesh.

[0357] Vertex information is information indicating the positions of vertices constituting the base mesh.

[0358] The connection information is information indicating the connection between the vertices constituting the base mesh.

[0359] The first method is a coding method of converting vertex information and connection information into secondary information and encoding the secondary information, and

[0360] The second method is a coding method of coding the vertex information and the connection information without converting them into secondary information.

[0361] Reference numerals list

[0362] 200 Encoding Device

[0363] 211 Pre-processing unit

[0364] 212 Coding mode determination unit

[0365] 213 Intra-frame coding unit

[0366] 214 inter-frame coding units

[0367] 215 bit stream generation unit

[0368] 251 basic grid coding units

[0369] 252 original patch generation unit

[0370] 253 displacement vector correction unit

[0371] 254 Packing Units

[0372] 255 bit shift video encoding unit

[0373] 256 grid reconstruction units

[0374] 257 attribute mapping correction unit

[0375] 258 attribute video coding unit

[0376] 259 header encoding unit

[0377] 260 combination units

[0378] 300 decoding device

[0379] 311 Demultiplexing Unit

[0380] 312 Header decoding unit

[0381] 313 Intraframe decoding unit

[0382] 314 Interframe decoding unit

[0383] 315 combination unit

[0384] 351 displacement video decoding unit

[0385] 352 Unpacking Unit

[0386] 353 basic grid decoding unit

[0387] 354 Reconstruction Unit

[0388] 355 subdivision units

[0389] 356 Displacement Vector Application Unit

[0390] 357 attribute video decoding unit

[0391] 900 Computer

Claims

1. An information processing device, comprising: an encoding method determination unit that selects the first method or the second method as an encoding method of a base grid to be subjected to intra-frame encoding; a first encoding unit that, when the first mode is selected as the encoding mode, converts vertex information indicating positions of vertices constituting the base mesh and connection information indicating connections between vertices constituting the base mesh into secondary information indicating a relationship between adjacent faces of the base mesh and encodes the secondary information; as well as A second encoding unit, which, when the second mode is selected as the encoding mode, encodes the vertex information and the connection information without converting the vertex information and the connection information into the secondary information, wherein: The base mesh is a mesh generated by removing vertices from an original mesh to be encoded consisting of vertices and connections representing a three-dimensional structure of an object, and is coarser than the original mesh.

2. The information processing device according to claim 1, wherein: The second encoding unit encodes the vertex information using correlation between vertices of the base mesh.

3. The information processing device according to claim 1, wherein: The second encoding unit stores the vertex information in a header and encodes the vertex information.

4. The information processing device according to claim 3, wherein: The second encoding unit further stores encoding mode control information indicating the selected encoding mode in the header and encodes the encoding mode control information.

5. The information processing device according to claim 1, wherein: The second encoding unit stores the vertex information in a displacement video having a 2D image in which a displacement vector is stored as a frame and encodes the vertex information, and The displacement vectors are the differences in position between the vertices of the subdivided base mesh and the vertices of the original mesh.

6. The information processing device according to claim 5, further comprising: A third encoding unit stores encoding mode control information indicating the selected encoding mode in a header and encodes the encoding mode control information.

7. The information processing device according to claim 1, further comprising: An original patch generation unit converts the basic mesh and generates an original patch as transmission information, wherein: When the second mode is selected as the encoding mode, the second encoding unit encodes the original surface patch without converting the original surface patch into the secondary information.

8. The information processing device according to claim 7, wherein: The original patch includes mesh information and original patch information, the mesh information includes the vertex information and the connection information, and the original patch information includes meta information related to the basic mesh.

9. The information processing device according to claim 8, wherein: The original patch information includes information indicating the number of the original patches and information indicating the number of faces in the original patch.

10. An information processing method, comprising: selecting a first mode or a second mode as an encoding mode of a base grid to be subjected to intra-frame encoding; In a case where the first mode is selected as the encoding mode, converting vertex information indicating positions of vertices constituting the base mesh and connection information indicating connections between vertices constituting the base mesh into secondary information indicating relationships between adjacent faces of the base mesh and encoding the secondary information; as well as When the second mode is selected as the encoding mode, the vertex information and the connection information are encoded without converting them into the secondary information, wherein: The base mesh is a mesh generated by removing vertices from an original mesh to be encoded consisting of vertices and connections representing a three-dimensional structure of an object, and is coarser than the original mesh.

11. An information processing device, comprising: a first decoding unit that, when the encoding mode control information indicates a first mode as an encoding mode of the intra-coded base mesh, decodes the encoded data of the intra-coded base mesh, generates secondary information indicating a relationship between adjacent faces of the base mesh, and converts the generated secondary information into vertex information and connection information of the base mesh; as well as A second decoding unit, which decodes the encoded data of the base mesh and generates the vertex information and the connection information which are not converted into the secondary information when the encoding mode control information indicates the second mode as the encoding mode, wherein The base mesh is a mesh generated by removing vertices from an original mesh to be encoded consisting of vertices and connections representing a three-dimensional structure of an object, and is coarser than the original mesh. The vertex information is information indicating the positions of vertices constituting the base mesh, The connection information is information indicating connections between vertices constituting the base mesh, The first method is a coding method of converting the vertex information and the connection information into the secondary information and encoding the secondary information, and The second method is a coding method of coding the vertex information and the connection information without converting the vertex information and the connection information into the secondary information.

12. The information processing device according to claim 11, wherein: The second decoding unit decodes the encoded data of the vertex information encoded using correlation between vertices of the base mesh.

13. The information processing device according to claim 11, wherein: The second decoding unit decodes the encoded data of the vertex information stored in the header and generates the vertex information.

14. The information processing device according to claim 13, wherein: The second decoding unit further decodes the encoded data of the encoding mode control information stored in the header.

15. The information processing device according to claim 11, wherein: The second decoding unit decodes the encoded data of the displacement video having the 2D image in which the displacement vector is stored as a frame, and generates the vertex information stored in the displacement video, and The displacement vectors are the differences in position between the vertices of the subdivided base mesh and the vertices of the original mesh.

16. The information processing device according to claim 15, further comprising: A third decoding unit decodes the encoded data of the encoding mode control information stored in the header.

17. The information processing device according to claim 11, wherein: The second decoding unit decodes the encoded data of the base grid and generates an original patch as information in a transmission format of the base grid, and The information processing device further includes a reconstruction unit for reconstructing the basic mesh using the original patch.

18. The information processing device according to claim 17, wherein: The original patch includes mesh information and original patch information, the mesh information includes the vertex information and the connection information, and the original patch information includes meta information related to the basic mesh.

19. The information processing device according to claim 18, wherein: The original patch information includes information indicating the number of the original patches and information indicating the number of faces in the original patch.

20. An information processing method, comprising: In a case where the encoding mode control information indicates the first mode as the encoding mode of the intra-coded base mesh, decoding the encoded data of the intra-coded base mesh, generating secondary information indicating the relationship between adjacent faces of the base mesh, and converting the generated secondary information into vertex information and connection information of the base mesh; as well as When the encoding mode control information indicates the second mode as the encoding mode, the encoded data of the base mesh is decoded, and the vertex information and the connection information that are not converted into the secondary information are generated, wherein: The base mesh is a mesh generated by removing vertices from an original mesh to be encoded consisting of vertices and connections representing a three-dimensional structure of an object, and is coarser than the original mesh. The vertex information is information indicating the positions of vertices constituting the base mesh, The connection information is information indicating connections between vertices constituting the base mesh, The first method is a coding method of converting the vertex information and the connection information into the secondary information and encoding the secondary information, and The second method is a coding method of coding the vertex information and the connection information without converting the vertex information and the connection information into the secondary information.