Information processing device and method

By generating and storing scalable decoding information in a metadata area of a content file, the method addresses the increased load of the playback process by allowing selective extraction and decoding of G-PCC content slices, enhancing efficiency in processing large-scale point cloud data.

JP7746995B2Active Publication Date: 2025-10-01SONY GROUP CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022547574
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-08
Filing Date
2021-09-06
Publication Date
2025-10-01
Estimated Expiration
2041-09-06

AI Technical Summary

Technical Problem

The existing method for transmitting G-PCC content does not support the slice structure proposed in Non-Patent Document 4, requiring the entire content to be transmitted and parsed for decoding, which increases the load of the playback process.

Method used

An information processing device and method that generates scalable decoding information based on depth information and dependency relationships between slices in G-PCC content, storing this information in a metadata area of a content file to facilitate efficient extraction and decoding of specific slices.

Benefits of technology

This approach allows for reduced decoder processing by enabling selective transmission and decoding of necessary slices, thereby suppressing the increase in playback processing load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007746995000001
    Figure 0007746995000001
  • Figure 0007746995000002
    Figure 0007746995000002
  • Figure 0007746995000003
    Figure 0007746995000003
Patent Text Reader

Abstract

The present disclosure relates to an information processing device and a method which make it possible to suppress an increase in a reproduction processing load. Scalable decoding information is generated on the basis of a depth of a slice and a dependent relationship between slices in G-PCC content, and stored in a metadata region of a content file that stores the G-PCC content. Moreover, on the basis of the scalable decoding information stored in the metadata region of said content file, an arbitrary slice of said G-PCC content is extracted from said content file and decoded. The present disclosure may be applied, for example, to an information processing device, an information processing method, or the like.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing device and method, and more particularly to an information processing device and method that can suppress an increase in the load of playback processing. [Background technology]

[0002] Conventionally, various methods have been proposed as video encoding technologies, such as HEVC (High Efficiency Video Coding). A technology for transmitting such encoded video is ISOBMFF (International Organization for Standardization Base Media File Format), which is a file container specification of MPEG-4 (Moving Picture Experts Group-4), an international standard technology for video compression (see, for example, Non-Patent Document 1).

[0003] Furthermore, in encoding technologies such as HEVC, encoded data can be hierarchically structured based on, for example, resolution, etc., making it possible to support scalable decoding. A file format has also been proposed in which encoded data is divided into tracks for each layer, and stored, allowing selective transmission of only the encoded data of a desired layer (see, for example, Non-Patent Document 2).

[0004] Meanwhile, as a method for encoding point clouds that represent three-dimensional objects as a collection of points, an encoding technology called G-PCC (Geometry-based Point Cloud Compression) is currently being standardized in MPEG-I Part 9 (ISO / IEC 23090-9), which separates and encodes point cloud data into geometry, which indicates position information of the points, and attributes, which indicate attribute information of the points (see, for example, Non-Patent Document 3). In this G-PCC, it has been proposed to form a slice structure that can be decoded independently (see Non-Patent Document 4).

[0005] In addition, as a technique for transmitting G-PCC content obtained by applying this G-PCC to encode a point cloud, storing the content in the above-mentioned ISOBMFF has been proposed (see, for example, Non-Patent Document 5). [Prior art documents] [Non-patent literature]

[0006] [Non-Patent Document 1] "Information technology - Coding of audio-visual objects - Part 12:ISO base media file format", ISO / IEC 14496-12, 2015-02-20 [Non-patent document 2] "Information technology - Coding of audio-visual objects - Part 15: Carriage of network abstraction layer (NAL) unit structured video in the ISO base media file format", ISO / IEC FDIS 14496-15:2014(E), 2014-01-13 [Non-patent document 3] "G-PCC Future Enhancements", ISO / IEC JTC 1 / SC 29 / WG 11 N19328, 2020-6-26 [Non-patent document 4] David Flynn, Khaled Mammou, "G-PCC: A hierarchical geometry slice structure", ISO / IEC JCTC1 / SC29 / WG11 MPEG / m54677, April 2020, Online [Non-patent document 5] Sejin Oh, Ryohei Takahashi, Youngkwon Lim, "Text of ISO / IEC CD 23090-18 Carriage of Geometry-based Point Cloud Compression Data", ISO / IEC JTC 1 / SC 29 / WG 11 N19442, 2020-07-30 Summary of the Invention [Problem to be solved by the invention]

[0007] However, the method described in Non-Patent Document 5 does not support the slice structure described in Non-Patent Document 4. Therefore, in order to decode some slices of the G-PCC content, the entire G-PCC content must be transmitted and parsed (analyzed), which may increase the load of the playback process.

[0008] The present disclosure has been made in consideration of such circumstances, and aims to make it possible to suppress an increase in the load of the regeneration process. [Means for solving the problem]

[0009] An information processing device according to one aspect of the present technology is an information processing device that includes a scalable decoding information generation unit that generates scalable decoding information regarding scalable decoding of Geometry-based Point Cloud Compression (G-PCC) content based on depth information indicating the quality hierarchy level of each slice in the G-PCC content, which includes a first slice and a second slice, and a dependency relationship between the first slice and the second slice in the G-PCC content, and a content file generation unit that generates a content file that stores the G-PCC content and stores the scalable decoding information in a metadata area of ​​the content file.

[0010] An information processing method according to one aspect of the present technology is an information processing method that generates scalable decoding information related to scalable decoding of Geometry-based Point Cloud Compression (G-PCC) content based on depth information indicating the quality hierarchy level of each slice in the G-PCC content, including a first slice and a second slice, and a dependency relationship between the first slice and the second slice in the G-PCC content, generates a content file that stores the G-PCC content, and stores the scalable decoding information in a metadata area of ​​the content file.

[0011] Another aspect of the present technology is an information processing device that includes an extraction unit that extracts any slice of Geometry-based Point Cloud Compression (G-PCC) content from a content file that stores G-PCC content including a first slice and a second slice based on scalable decoding information stored in a metadata area of ​​the content file, and a decoding unit that decodes the slice of the G-PCC content extracted by the extraction unit, wherein the scalable decoding information is information regarding scalable decoding of the G-PCC content and is information generated based on depth information indicating the quality hierarchy level of the slice in the G-PCC content and the dependency between the first slice and the second slice in the G-PCC content.

[0012] Another aspect of the information processing method of the present technology is an information processing method that extracts any slice of Geometry-based Point Cloud Compression (G-PCC) content from a content file storing G-PCC content including a first slice and a second slice based on scalable decoding information stored in a metadata area of ​​the content file, and decodes the extracted slice of the G-PCC content, wherein the scalable decoding information is information regarding scalable decoding of the G-PCC content, and is information generated based on depth information indicating the quality hierarchy level of the slice in the G-PCC content and the dependency between the first slice and the second slice in the G-PCC content.

[0013] In an information processing device and method according to one aspect of the present technology, scalable decoding information for scalable decoding of Geometry-based Point Cloud Compression (G-PCC) content including a first slice and a second slice is generated based on depth information indicating the quality hierarchy level of each slice in the G-PCC content and the dependency between the first slice and the second slice in the G-PCC content, a content file storing the G-PCC content is generated, and the scalable decoding information is stored in a metadata area of ​​the content file.

[0014] In another aspect of the information processing device and method of the present technology, based on scalable decoding information stored in a metadata area of ​​a content file storing G-PCC (Geometry-based Point Cloud Compression) content including a first slice and a second slice, an arbitrary slice of the G-PCC content is extracted from the content file, and the extracted slice of the G-PCC content is decoded. [Brief explanation of the drawings]

[0015] [Figure 1]FIG. 1 is a diagram illustrating an L-HEVC file format. [Figure 2] FIG. 10 is a diagram illustrating a slice structure of geometry. [Figure 3] FIG. 1 is a diagram illustrating the structure of a bitstream. [Figure 4] FIG. 2 is a diagram illustrating the structure of a content file. [Figure 5] FIG. 1 is a diagram illustrating a method for transmitting G-PCC content. [Figure 6] FIG. 1 is a diagram illustrating a method for transmitting G-PCC content. [Figure 7] FIG. 10 is a diagram illustrating slice configuration information. [Figure 8] FIG. 10 is a diagram showing an example of a subsample information box. [Figure 9] FIG. 10 is a diagram illustrating slice dependency relationship information. [Figure 10] 10A and 10B are diagrams illustrating attribute geometry slice identification information. [Figure 11] FIG. 10 is a diagram illustrating an attribute slice to which non-scalable coding is applied. [Figure 12] FIG. 10 is a diagram illustrating non-scalable coding attribute geometry slice identification information. [Figure 13] FIG. 10 is a diagram illustrating slice dependency relationship information. [Figure 14] FIG. 10 is a diagram illustrating geometry slice depth information. [Figure 15] FIG. 10 is a diagram illustrating track configuration information. [Figure 16] FIG. 10 is a diagram illustrating track depth information. [Figure 17] FIG. 10 is a diagram illustrating track dependency relationship information. [Figure 18] FIG. 10 is a diagram illustrating an example of the configuration of a Matryoshka media container. [Figure 19] FIG. 2 is a block diagram illustrating an example of the main configuration of a file generation device. [Figure 20] 10 is a flowchart illustrating an example of the flow of a file generation process. [Figure 21] FIG. 2 is a block diagram illustrating an example of the main configuration of a decoding device. [Figure 22] FIG. 2 is a block diagram showing an example of the main configuration of a playback processing unit. [Figure 23] 10 is a flowchart showing an example of the flow of a playback process. [Figure 24] FIG. 10 is a diagram illustrating a control file of a G-PCC content. [Figure 25] FIG. 10 is a diagram illustrating adaptation set configuration information. [Figure 26] FIG. 10 is a diagram illustrating an example of a description of an MPD. [Figure 27] FIG. 2 is a block diagram illustrating an example of the main configuration of a file generation device. [Figure 28] 10 is a flowchart illustrating an example of the flow of a file generation process. [Figure 29] FIG. 2 is a block diagram illustrating an example of the main configuration of a decoding device. [Figure 30] 10 is a flowchart showing an example of the flow of a playback process. [Figure 31] FIG. 1 is a block diagram illustrating an example of the main configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION

[0016] Hereinafter, modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described in the following order. 1. Transmission of G-PCC content with slice structure 2. Transmission of scalable decoding information by content file 3. First embodiment (file generation device, playback device) 4. Transmission of scalable decoding information via control file 5. Second embodiment (file generation device, playback device) 6. Supplementary Notes

[0017] <1. Transmission of G-PCC content with slice structure> <References supporting technical content and technical terminology> The scope of disclosure of the present technology includes not only the contents described in the embodiments but also the contents described in the following non-patent documents that were publicly known at the time of filing, as well as the contents of other documents referenced in the following non-patent documents.

[0018] Non-patent document 1: (mentioned above) Non-patent document 2: (mentioned above) Non-patent document 3: (mentioned above) Non-patent document 4: (mentioned above) Non-patent document 5: (mentioned above) Non-patent document 6: https: / / www.matroska.org / index.html

[0019] In other words, the contents of the above-mentioned non-patent documents and the contents of other documents referenced in the above-mentioned non-patent documents are also used as the basis for determining the support requirements.

[0020] <hevc> Conventionally, various methods have been proposed as video encoding technologies, such as High Efficiency Video Coding (HEVC). As a technology for transmitting such encoded video, there is the International Organization for Standardization Base Media File Format (ISOBMFF), which is a file container specification of the international standard technology for video compression, "MPEG-4 (Moving Picture Experts Group-4)," as described in Non-Patent Document 1.

[0021] Furthermore, in encoding technologies such as HEVC, encoded data can be hierarchically structured based on, for example, resolution, etc., making it possible to support scalable decoding. As described in Non-Patent Document 2, for example, a file format (L-HEVC file format) has also been proposed in which encoded data is divided into tracks for each layer and stored, and only encoded data of a desired layer can be selectively transmitted.

[0022] In the L-HEVC file format, a video bitstream has a hierarchical structure, and one or more of its layers can be stored in individual tracks of ISOBMFF. As shown in FIG. 1, in ISOBMFF, layer information included in each sample is stored in a layer information sample group. In L-HEVC, which encodes a bitstream so that it has a hierarchical structure, inter prediction is applied, so the hierarchical structure does not change frequently. Therefore, it is preferable to store layer information in a sample group. This sample group can also be used as information for selecting a track.

[0023] <Point Cloud> By the way, as 3D data representing a three-dimensional object (also referred to as a 3D object), there is a point cloud that represents a 3D object as a set of points.

[0024] For example, in the case of a point cloud, a 3D object that is a three-dimensional structure is represented as a set of a large number of points. A point cloud is composed of the position information (also referred to as geometry) of each point and the attribute information (also referred to as attribute). The attribute can contain any information. For example, color information, reflectance information, normal information, etc. of each point may be included in the attribute. In this way, the point cloud has a relatively simple data structure and can represent the three-dimensional shape of a 3D object with sufficient accuracy by using a sufficiently large number of points.

[0025] <Overview of G-PCC> In Non-Patent Document 3, an encoding technique called Geometry-based Point Cloud Compression (G-PCC) that encodes this point cloud separately into geometry and attributes was disclosed. G-PCC is in the process of being standardized in MPEG-I Part 9 (ISO / IEC 23090-9).

[0026] For the encoding of geometry, tree-structured encoding as shown in FIG. 2 is applied. In tree-structured encoding, first, the geometry is tree-structured. The three-dimensional space is recursively divided, and the geometry is quantized in each division region at each stage, thereby forming a tree structure of geometry as shown in FIG. 2. In the case of FIG. 2, a tree structure hierarchically divided into depths (also referred to as LOD) 0 to 6 shown in the vertical direction in the figure is formed.

[0027] As the number of divisions increases, the number of division regions increases and each division region becomes smaller, resulting in a 3D object being represented by a larger number of points with higher precision geometry. In other words, the deeper the depth (the lower the layer), the higher the resolution of the point cloud. The geometry of each layer is then encoded according to this tree structure. For example, the difference between the geometry of each layer and the geometry of the layer immediately above it is calculated, and this difference is encoded. While each layer can be encoded independently, encoding the difference in this way can improve encoding efficiency.

[0028] Since the difference with the higher layer is encoded, to obtain geometry at a desired depth, it is sufficient to decode the encoded data of that depth and the layer higher than that depth. For example, in the case of FIG. 2, by decoding the encoded data of depths 0 to 2, geometry at a resolution of depth 2 can be obtained. In other words, decoding depths 3 to 6 is not necessary. This decoding method, which can restore (generate) information at a desired layer by decoding only a portion of the encoded data, is called scalable decoding. In other words, by encoding in a tree structure as described above, geometry can be decoded in a scalable manner according to resolution.

[0029] For example, if scalable decoding is not supported, all encoded data must be decoded to restore the geometry of the lowest resolution (i.e., the highest resolution), and then a desired resolution lower than that must be generated. In contrast, if the encoded data is compatible with scalable decoding, the desired resolution can be restored by decoding only the necessary portion of the encoded data as described above. In other words, an increase in the load of the playback process can be suppressed.

[0030] 2, a binary tree is shown as an example of the tree structure of this geometry, but any tree structure may be used, such as an octree or a kd tree.

[0031] The coded data (bit stream) generated by coding the geometry as described above is also called a geometry bit stream.

[0032] Furthermore, methods such as Predicting Weight Lifting, Region Adaptive Hierarchical Transform (RAHT), or Fix Weight Lifting are applied to compress attributes. The coded data (bitstream) generated by coding attributes is also called an attribute bitstream.

[0033] Furthermore, a bitstream that combines a geometry bitstream and an attribute bitstream is also called a G-PCC bitstream or G-PCC content.

[0034] <Slice> Non-Patent Document 4 proposes forming a slice structure in the bitstream for this G-PCC. A slice is a unit for dividing geometry and attribute data. In this specification, a geometry slice is also referred to as a geometry slice. An attribute slice is also referred to as an attribute slice.

[0035] Slices are divided into units of depths in the tree structure of geometry. That is, a slice is composed of data for a single depth or multiple consecutive depths. For example, in FIG. 2, the data areas divided by bold lines indicate slices. In the figure, circled numbers indicate the identification information of each slice. For example, data for depths 0 to 3 form slice #1. Slices can also be divided by position (area) in three-dimensional space. For example, in FIG. 2, data for depths 4 and 5 are divided into areas A to D, forming four slices: slice #2, slice #3, slice #4, and slice #5. Similarly, data for depth 6 is divided into areas A to D, forming four slices: slice #6, slice #7, slice #8, and slice #9.

[0036] Geometry and attribute data are coded for each slice. In other words, the geometry bitstream and attribute bitstream can be decoded for each slice. However, as mentioned above, since geometry codes the difference with the upper layer, there are two types of slices: independent slices that can be decoded independently, and dependent slices that require other slices for decoding.

[0037] For example, in the case of the tree structure of FIG. 2, to obtain the geometry of slice #1 (e.g., the geometry of depth 3) by decoding, it is sufficient to decode the geometry of slice #1. On the other hand, to obtain the geometry of slice #2 (e.g., the geometry of region A at depth 4) by decoding, it is necessary to decode the geometries of slice #1 and slice #2. Similarly, to decode slices #3 to #5, it is necessary to decode slice #1. Furthermore, to decode slice #6, it is necessary to decode slices #1 and #2. To decode slice #7, it is necessary to decode slices #1 and #3. To decode slice #8, it is necessary to decode slices #1 and #4. To decode slice #9, it is necessary to decode slices #1 and #5.

[0038] That is, slice #1 is an independent slice, and slices #2 to #9 are dependent slices.

[0039] Fig. 3 is a diagram illustrating an example of the main configuration of a bitstream having such a slice structure. Bitstream 30 shown in gray in Fig. 3 is a bitstream of a point cloud in which geometry is tree-structured and sliced ​​as shown in Fig. 2 and then encoded. Fig. 3 illustrates only a portion of it.

[0040] As shown in FIG. 3, the bitstream 30 has a sequence parameter set (SPS), a geometry parameter set (GPS), an attribute parameter set (APS), and a tile inventory. The sequence parameter set is a parameter set related to the entire sequence. The geometry parameter set is a parameter set related to geometry. The geometry parameter set may differ for each geometry slice. The attribute parameter set is a parameter set related to attributes. The attribute parameter set may differ for each attribute slice. The tile inventory stores position information of tiles. The number of tiles and their position information are variable for each frame.

[0041] Following this data, as indicated by the dotted arrow 31, the bit stream of the point cloud is arranged for each sample. A sample is a point cloud at a certain time and corresponds to a frame of a video. Within each sample, the bit stream is arranged for each slice. Within each slice, the geometry bit stream is arranged first, followed by the attribute bit stream.

[0042] In Figure 3, each square represents a data unit. The data unit of "geom_slice#1" is a data unit that stores the geometry bitstream of slice #1. The data unit of "attr slice(s)" that follows this data unit is a data unit that stores the attribute bitstream of slice #1. The same slice identification information (slice_id) is assigned to the geometry and attribute data units that make up the same slice. Note that attributes can also be divided into slices for each parameter, and multiple attribute slices can be stored in one data unit. In other words, these data units correspond to slice #1, which is an independent slice, and store data from depth 0 to depth 3 in the tree structure of Figure 2.

[0043] In this specification, a data unit that stores a geometry bitstream is also referred to as a geometry data unit, and a data unit that stores an attribute bitstream is also referred to as an attribute data unit.

[0044] The data unit of "geom_slice#2" is a geometry data unit that stores the geometry bitstream of slice #2. The data unit of "attr slice(s)" that follows this geometry data unit is an attribute data unit that stores the attribute bitstream of slice #2. In other words, these data units are data units corresponding to slice #2, which is a dependent slice, and store data for area A at depths 4 and 5 in the tree structure of Figure 2. Slice #2 is dependent on slice #1, as indicated by arrow 32.

[0045] The data unit of "geom_slice#6" is a geometry data unit that stores the geometry bitstream of slice #6. The data unit of "attr slice(s)" that follows this geometry data unit is an attribute data unit that stores the attribute bitstream of slice #6. In other words, these data units are data units corresponding to slice #6, which is a dependent slice, and store data of region A at depth 6 in the tree structure of Figure 2. Slice #6 is directly dependent on slice #2, as indicated by arrow 33. In other words, slice #6 is also indirectly dependent on slice #1.

[0046] The data unit of "geom_slice#3" is a geometry data unit that stores the geometry bitstream of slice #3. The data unit of "attr slice(s)" that follows this geometry data unit is a data unit that stores the attribute bitstream of slice #3. In other words, these data units are data units that correspond to slice #3, which is a dependent slice, and store data of area B at depths 4 and 5 in the tree structure of Figure 2. Slice #3 is dependent on slice #1, as indicated by arrow 34.

[0047] The data unit "geom_slice#7" is a geometry data unit that stores the geometry bitstream of slice#7. The data unit "attr slice(s)" that follows this geometry data unit is a data unit that stores the attribute bitstream of slice#7. In other words, these data units are data units that correspond to slice#7, which is a dependent slice, and store data of region B at depth 6 in the tree structure of Figure 2. Slice#7 is directly dependent on slice#3, as indicated by arrow 35. In other words, slice#7 is also indirectly dependent on slice#1.

[0048] Although not shown in the figure, data units (geometry data units and attribute data units) corresponding to slices #4, #8, #5, and #9 are similarly arranged thereafter. Slices and tiles are linked by tile identification information (tile_id) stored in the geometry data units.

[0049] If such a slice structure is not formed in a bitstream, a decoder needs to parse (analyze) the bitstream to determine which depth data corresponds to which part of the bitstream in order to perform scalable decoding. In contrast, if the bitstream forms the slice structure described above, the decoder can easily select data to be decoded in slice units.

[0050] For example, to obtain data for slice #1 in the tree structure of Figure 2, the decoder simply decodes the bitstream stored in the data unit corresponding to slice #1 in Figure 3. To obtain data for slice #2, the decoder simply decodes the bitstream stored in the data unit corresponding to slice #1 and the bitstream stored in the data unit corresponding to slice #2. To obtain data for slice #6, the decoder simply decodes the bitstream stored in the data unit corresponding to slice #1, the bitstream stored in the data unit corresponding to slice #2, and the bitstream stored in the data unit corresponding to slice #6.

[0051] Similarly, to obtain data for slice #3, the decoder simply decodes the bitstream stored in the data unit corresponding to slice #1 and the bitstream stored in the data unit corresponding to slice #3. To obtain data for slice #7, the decoder simply decodes the bitstream stored in the data unit corresponding to slice #1, the bitstream stored in the data unit corresponding to slice #3, and the bitstream stored in the data unit corresponding to slice #7.

[0052] To obtain data for slice #4, the decoder simply decodes the bitstream stored in the data unit corresponding to slice #1 and the bitstream stored in the data unit corresponding to slice #4. To obtain data for slice #8, the decoder simply decodes the bitstream stored in the data unit corresponding to slice #1, the bitstream stored in the data unit corresponding to slice #4, and the bitstream stored in the data unit corresponding to slice #8.

[0053] To obtain data for slice #5, the decoder simply decodes the bitstream stored in the data unit corresponding to slice #1 and the bitstream stored in the data unit corresponding to slice #5. To obtain data for slice #9, the decoder simply decodes the bitstream stored in the data unit corresponding to slice #1, the bitstream stored in the data unit corresponding to slice #5, and the bitstream stored in the data unit corresponding to slice #9.

[0054] Therefore, the decoder can more easily perform scalable decoding.

[0055] In this specification, an independent slice of geometry is also referred to as an independent geometry slice. A dependent slice of geometry is also referred to as a dependent geometry slice. Further, an independent slice of an attribute is also referred to as an independent attribute slice. A dependent slice of an attribute is also referred to as a dependent attribute slice.

[0056] <Storage of G-PCC Content in ISOBMFF> Non-Patent Document 5 discloses a method of storing G-PCC content (G-PCC bitstream) in ISOBMFF for the purpose of improving the efficiency of playback processing from local storage of G-PCC content (G-PCC bitstream) and network distribution. This method is currently in the process of being standardized in MPEG-I Part 18 (ISO / IEC 23090-18).

[0057] FIG. 4 is a diagram showing an example of the file structure in that case. In this specification, the content stored with G-PCC content in ISOBMFF is also referred to as a content file.

[0058] The sequence parameter set is stored in the GPCCDecoderConfigurationRecord in the metadata area of the content file. The GPCCDecoderConfigurationRecord may further include a geometry parameter set, an attribute parameter set, and a tile inventory according to the sample entry type.

[0059] A sample in the Media data box (Media) includes a geometry slice corresponding to a 1-point cloud frame and an attribute slice. Further, it may include a geometry parameter set, an attribute parameter set, and a tile inventory according to the sample entry type.

[0060] <Storing G-PCC content with slice structure in ISOBMFF> For example, if the bitstream could be divided into tracks in slice units and stored as in the L-HEVC file format described in Non-Patent Document 2, so-called progressive download decoding would be possible when distributing G-PCC content, in which only a portion corresponding to the depth to be played back is transmitted and decoded. This would prevent increases in the amount of data transmitted and the amount of data to be decoded. For example, this would prevent increases in display delays when distributing large-scale point clouds.

[0061] However, Non-Patent Document 5 does not disclose storing G-PCC content having such a slice structure in ISOBMFF. Therefore, even when decoding only a portion of the slices of the G-PCC content, the entire G-PCC content must be transmitted, potentially increasing the amount of data to be transmitted. Furthermore, even if tracks are used to divide the bitstream, information necessary for scalable decoding, such as which depth data is contained in which track, is only present in the bitstream, so the decoder must parse the bitstream. Furthermore, because G-PCC applies intra-coding, it is common for the depth hierarchical structure to change for each sample (frame). Therefore, the decoder must parse the entire bitstream to obtain the information necessary for scalable decoding. This potentially increases the load of the playback process.

[0062] <2. Transmission of scalable decoding information using content files> Therefore, as shown in the top row of the table in Fig. 5, scalable decoding information of G-PCC content having a slice structure stored in the metadata area of ​​the content file is transmitted (Method 1). Note that in this specification, scalable decoding information is information related to (used for) scalable decoding of G-PCC content having a slice structure. Furthermore, this scalable decoding information is set based on the depth of each slice of G-PCC content having a slice structure and the dependency relationships between slices.

[0063] For example, an information processing device may include a scalable decoding information generation unit that generates scalable decoding information related to scalable decoding of G-PCC content based on depth information indicating the quality hierarchy level of the geometry contained in each slice in the G-PCC content including a first slice and a second slice and the dependency relationship between the first slice and the second slice in the G-PCC content, and a content file generation unit that generates a content file that stores the G-PCC content and stores the scalable decoding information in the metadata area of ​​the content file.

[0064] For example, in an information processing method, scalable decoding information regarding scalable decoding of G-PCC content including a first slice and a second slice is generated based on depth information indicating the quality hierarchy level of the geometry contained in each slice in the G-PCC content and the dependency between the first slice and the second slice in the G-PCC content, a content file for storing the G-PCC content is generated, and the scalable decoding information is stored in a metadata area of ​​the content file.

[0065] Furthermore, for example, an information processing device may include an extraction unit that extracts any slice of G-PCC content from a content file storing G-PCC content including a first slice and a second slice, based on scalable decoding information stored in a metadata area of ​​the content file, and a decoding unit that decodes the slice of G-PCC content extracted by the extraction unit. Note that the scalable decoding information is information related to scalable decoding of G-PCC content, and is information generated based on depth information indicating the quality hierarchical level of geometry included in the slice of the G-PCC content and the dependency relationship between the first slice and the second slice in the G-PCC content.

[0066] For example, in an information processing method, an arbitrary slice of G-PCC content including a first slice and a second slice is extracted from a content file storing the G-PCC content based on scalable decoding information stored in a metadata area of ​​the content file, and the extracted slice of the G-PCC content is decoded. Note that the scalable decoding information is information related to scalable decoding of the G-PCC content, and is generated based on depth information indicating the quality hierarchical level of geometry included in the slice of the G-PCC content and the dependency relationship between the first slice and the second slice in the G-PCC content.

[0067] In this way, the decoder can extract and decode slices necessary to reproduce a point cloud of a desired depth or region based on the scalable decoding information in the metadata area of ​​the content file, and generate presentation information. This reduces unnecessary decoder processing (such as transmission of unnecessary information and parsing of bitstreams). Therefore, it is possible to suppress an increase in the load of the reproduction processing.

[0068] One use case for G-PCC content is the encoding of large-scale point cloud data, such as point cloud map data or virtual assets in film production (digitalized versions of actual film sets). Such large-scale point clouds have a very large data volume overall, and playing back the entire data is not realistic in terms of processing load and processing delay. Therefore, scalable decoding is desired, which allows only a portion of the data to be played back by limiting the playback area or reducing the resolution.

[0069] When a G-PCC file is stored in a content file and transmitted, scalable decoding information can be stored in the metadata area of ​​the content file, and by making it compatible with the above-mentioned scalable decoding, it becomes possible to transmit only the necessary information or to more easily decode only the necessary information. As mentioned above, the larger the data, the more effectively it is possible to suppress the increase in the load of the playback process, and thus obtain greater benefits.

[0070] A content file may have a structure in which geometry and attributes are stored in one track (also called a single track encapsulation structure), or a structure in which geometry and attributes are stored in different tracks (also called a multi-track encapsulation structure). Although the following description uses a single track as an example, the present technology can be applied to the case of multi-tracks in the same way as the case of a single track. Note that even in the case of a single track, there may be multiple tracks (there may be multiple tracks containing geometry and attributes).

[0071] In the following description, it is assumed that the SPS, GPS, APS, and tile inventory are stored in the GPCCDecoderConfigurationRecord, and that only geometry slices and attribute slices are stored in the sample, as described with reference to Figures 3 and 4. However, some or all of the SPS, GPS, APS, and tile inventory may be stored in the sample.

[0072] <2-1. Slice configuration information for each sample> As shown in the second row from the top of the table in Fig. 5, the scalable decoding information may include slice configuration information for each sample (method 1-1). The slice configuration information is information about the configuration of slices within a sample. That is, the decoder can obtain the slice configuration information for each sample from the metadata area of ​​the content file. Therefore, the decoder can understand the slice configuration for each sample without parsing the bitstream (i.e., more easily).

[0073] For example, as shown in the third row from the top of the table in FIG. 5, this slice configuration information may be stored in codec specific parameters in a subsample information box (SubSampleInformationBox) in the metadata area of ​​the content file (Method 1-1-1). For example, a content file generation unit of an encoder may set a subsample for each slice and store the slice configuration information in codec specific parameters in the subsample information box in the metadata area of ​​the content file. Furthermore, an extraction unit of a decoder may extract any slice of the G-PCC content from the content file based on the slice configuration information stored in the codec specific parameters in the subsample information box in the metadata area of ​​the content file for the subsample set for each slice.

[0074] Because G-PCC applies intra-coding, it is common for the depth hierarchical structure to change for each sample (frame). Therefore, if slice configuration information is stored in sample groups, as in the L-HEVC file format, each sample must be associated with a sample group that contains different information, which can unnecessarily increase the file size by the amount of grouping information. Furthermore, to determine the slice configuration of a desired sample, the decoder must check all sample groups, which can increase the playback processing load.

[0075] As described above, by setting a subsample for each slice and storing slice configuration information in the codec specific parameters of the subsample information box in the metadata area of ​​the content file, it is possible to suppress an increase in file size. Furthermore, since the decoder only needs to check the codec specific parameters of the desired subsample information box, it is possible to more easily check the slice configuration information. In other words, it is possible to suppress an increase in the load of the playback process.

[0076] <2-1-1. Slice dependency information> For example, as shown in the fourth row from the top of the table in FIG. 5, this slice configuration information may include slice dependency information (method 1-1-2). In this specification, slice dependency information is information indicating the dependency relationship between slices or slice groups. For example, slice dependency information indicates the dependency relationship between a first slice and a second slice included in a G-PCC content. Also, in this specification, a slice group refers to multiple slices that correspond to the same depth.

[0077] For example, bitstream 100 shown in Fig. 7 shows a portion of G-PCC content similar to bitstream 30 in Fig. 3. Bitstream 101 shown in gray shows a portion of the structure within a sample of bitstream 100. Each data unit of this bitstream 101 has inter-slice dependencies similar to those in Fig. 3, as indicated by arrows 111 to 114.

[0078] For example, the geometry data unit of "geom_slice#2" and the attribute data unit of "attr slice(s)" that follows are data units corresponding to slice #2, and are subordinate to the data unit of slice #1 (the geometry data unit of "geom_slice#1" and the attribute data unit of "attr slice(s)" that follows), as shown by arrow 111. In other words, slice #2 is subordinate to slice #1.

[0079] Furthermore, the geometry data unit of "geom_slice#6" and the attribute data unit of "attr slice(s)" that follows it are data units corresponding to slice #6, and as shown by arrow 112, they are directly dependent on the data unit of slice #2 (the geometry data unit of "geom_slice#2" and the attribute data unit of "attr slice(s)" that follows it). In other words, slice #6 is also indirectly dependent on slice #1.

[0080] Furthermore, the geometry data unit of "geom_slice#3" and the attribute data unit of "attr slice(s)" that follows it are data units corresponding to slice #3, and are subordinate to the data unit of slice #1 (the geometry data unit of "geom_slice#1" and the attribute data unit of "attr slice(s)" that follows it), as shown by arrow 113. In other words, slice #3 is subordinate to slice #1.

[0081] Furthermore, the geometry data unit of "geom_slice#7" and the attribute data unit of "attr slice(s)" that follows it are data units corresponding to slice #7, and as shown by arrow 114, they are directly dependent on the data unit of slice #3 (the geometry data unit of "geom_slice#3" and the attribute data unit of "attr slice(s)" that follows it). In other words, slice #7 is also indirectly dependent on slice #1.

[0082] That is, the information indicated by the arrows (e.g., arrows 111 to 114) between slices surrounded by a dotted frame 121 is slice dependency information. This slice dependency information is stored in the metadata area of ​​the content file. In this way, the decoder can obtain this slice dependency information from the metadata area of ​​the content file. Therefore, the decoder can understand the slice dependency information without parsing the bitstream (i.e., more easily).

[0083] For example, as shown in the fifth row from the top of the table in Fig. 5, this slice dependency information may be stored in the codex specific parameters of the subsample information box in the metadata area of ​​the content file (method 1-1-2-1). For example, a content file generation unit of an encoder may set a subsample for each slice and store the slice dependency information in the codex specific parameters of the subsample information box in the metadata area of ​​the content file. Furthermore, an extraction unit of a decoder may extract any slice of the G-PCC content from the content file based on the slice dependency information stored in the codex specific parameters of the subsample information box in the metadata area of ​​the content file for the subsample set for each slice.

[0084] Because G-PCC applies intra-coding, it is common for the depth hierarchical structure to change for each sample (frame). Therefore, if slice dependency information is stored in sample groups, as in the L-HEVC file format, each sample must be associated with a sample group that contains different information, which can unnecessarily increase the file size by the amount of grouping information. Furthermore, since the decoder must check all sample groups to determine the slice structure of a desired sample, this can increase the playback processing load.

[0085] As described above, by setting a subsample for each slice and storing slice dependency information in the codec specific parameters of the subsample information box in the metadata area of ​​the content file, it is possible to suppress an increase in file size. Also, since the decoder only needs to check the codec specific parameters of the desired subsample information box, it is possible to more easily check the slice dependency. In other words, it is possible to suppress an increase in the load of the playback process.

[0086] For example, as shown in the sixth row from the top of the table in FIG. 5, this slice dependency relationship information may include reference-source geometry slice identification information and referenced-geometry slice identification information (method 1-1-2-2). In this specification, the reference-source geometry slice identification information is the identification information of the geometry slice to which the information corresponds. In other words, the reference-source geometry slice identification information is the identification information of the geometry slice that serves as the reference source (start point of the arrow in the dotted-line frame 121 in FIG. 7) in the dependency relationship between the above-mentioned slices or slice groups. Also, in this specification, the referenced-geometry slice identification information is the identification information of another geometry slice referenced by the geometry slice to which the information corresponds. In other words, the referenced-geometry slice identification information is the identification information of the geometry slice that serves as the reference destination (end point of the arrow in the dotted-line frame 121 in FIG. 7) in the dependency relationship between the above-mentioned slices or slice groups.

[0087] In this way, the decoder can obtain the referenced geometry slice identification information and the referenced geometry slice identification information from the metadata area of ​​the content file, and thus can grasp the referenced geometry slice identification information without parsing the bitstream (i.e., more easily).

[0088] For example, as shown in the seventh row from the top of the table in Fig. 5, if the slice corresponding to this slice dependency relationship information (i.e., the slice corresponding to the reference geometry slice identification information) is an independent geometry slice, the reference geometry slice identification information and the referenced geometry slice identification information may be the same (method 1-1-2-2-1). In other words, since an independent geometry slice can be decoded without requiring other slices, the reference may be set to the independent geometry slice itself. Furthermore, for an independent geometry slice, storage of the referenced geometry slice identification information may be omitted.

[0089] For example, as shown in the eighth row from the top of the table in Fig. 5, the reference-source geometry slice identification information and the reference-destination geometry slice identification information may be stored in the codex specific parameters of the subsample information box in the metadata area of ​​the content file (method 1-1-2-2-1). For example, a content file generation unit of an encoder may set a subsample for each slice and store the reference-source geometry slice identification information and the reference-destination geometry slice identification information in the codex specific parameters of the subsample information box in the metadata area. Furthermore, an extraction unit of a decoder may extract any slice of the G-PCC content from the content file based on the reference-source geometry slice identification information and the reference-destination geometry slice identification information stored in the codex specific parameters of the subsample information box in the metadata area of ​​the subsample set for each slice.

[0090] An example of the syntax of the subsample information box is shown in Figure 8. In this subsample information box, flags are set, as shown in the underlined line 131. Also, codec specific parameters are provided, as shown in the underlined line 132. The codec specific parameters store subsample information determined for each encoding codec.

[0091] An example of the syntax of Codex Specific Parameters is shown in Figure 9. As shown in the dotted line frame 133, the Codex Specific Parameters stores reference source geometry slice identification information (geom_slice_id) and reference destination geometry slice identification information (ref_geom_slice_id). The slice identification information (slice_id) of the geometry slice corresponding to this information is set in geom_slice_id. The slice identification information (slice_id) of another geometry slice referenced by that geometry slice is set in ref_geom_slice_id.

[0092] In this way, by setting a subsample for each slice and storing the reference source geometry slice identification information and the reference destination geometry slice identification information in the codec specific parameters of the subsample information box in the metadata area of ​​the content file, it is possible to suppress an increase in file size. Also, since the decoder only needs to check the codec specific parameters of the desired subsample information box, it is possible to more easily confirm the reference relationship between geometry slices (reference source and reference destination geometry slices). In other words, it is possible to suppress an increase in the load of the playback process.

[0093] In the example of Fig. 7, the geometry data unit and the attribute data unit immediately following it constitute one slice, so the reference relationship from the attribute slice to the geometry slice is obvious. Therefore, in the example of Fig. 9, the reference relationship from the attribute slice to the geometry slice is omitted.

[0094] <Attribute geometry slice identification information> It is also possible to explicitly indicate the reference relationship from the attribute slice to the geometry slice. For example, as shown in the ninth row from the top of the table in FIG. 5, the slice dependency relationship information may include attribute geometry slice identification information (method 1-1-2-3). The attribute geometry slice identification information is identification information of the geometry slice referenced by the attribute slice. By explicitly indicating the reference relationship from the attribute slice to the geometry slice in this way, it is possible to eliminate restrictions on the positional relationship between the geometry data unit and the attribute data unit.

[0095] For example, as shown in the tenth row from the top of the table in Fig. 5, this attribute geometry slice identification information may be stored in the codex specific parameters of the subsample information box in the metadata area of ​​the content file (method 1-1-2-3-1). For example, a content file generation unit of an encoder may set a subsample for each slice and store the attribute geometry slice identification information in the codex specific parameters of the subsample information box in the metadata area. Furthermore, an extraction unit of a decoder may extract any slice of the G-PCC content from the content file based on the attribute geometry slice identification information stored in the codex specific parameters of the subsample information box in the metadata area of ​​the subsample set for each slice.

[0096] An example of the syntax of the Codex Specific Parameters in this case is shown in Figure 10. The Codex Specific Parameters stores attribute geometry slice identification information (ref_attr_geom_slice_id) as shown in the underlined line 134. The ref_attr_geom_slice_id is set to the slice identification information (slice_id) of the geometry slice referenced by the attribute slice to which the information corresponds.

[0097] In this way, by setting a subsample for each slice and storing attribute geometry slice identification information in the codec specific parameters of the subsample information box in the metadata area of ​​the content file, it is possible to suppress an increase in file size. Also, since the decoder only needs to check the codec specific parameters of the desired subsample information box, it can more easily check the geometry slice referenced by the attribute. In other words, it is possible to suppress an increase in the load of the playback process.

[0098] <Non-scalable coding> Non-scalable coding may be applied to attributes. Non-scalable coding is a coding method that does not support scalable decoding. For example, when non-scalable coding is applied to attributes, an attribute slice is configured so that it can be decoded independently of other attribute slices (i.e., without referring to other attribute slices). Therefore, data (depth) overlaps between attribute data units.

[0099] An example of the configuration of a bitstream when non-scalable coding is applied to attributes in this way is shown in Fig. 11. A gray bitstream 141 shown in Fig. 11 shows part of the configuration of G-PCC content (G-PCC bitstream). In this bitstream 141, attribute data unit 151 stores attribute slices corresponding to geometries of depth 0 to depth 3. Attribute data unit 152 stores attribute slices corresponding to geometries of depth 0 to depth 5. Attribute data unit 153 stores attribute slices corresponding to geometries of depth 0 to depth 6.

[0100] Therefore, for example, by decoding attribute data unit 153, it is possible to obtain attributes corresponding to geometries at depths 0 to 6. In other words, to obtain attributes corresponding to geometry at depth 6, it is sufficient to decode attribute data unit 153 (there is no need to decode other attribute data units).

[0101] Similarly, by decoding attribute data unit 152, attributes corresponding to geometries at depths 0 to 5 can be obtained. In other words, to obtain attributes corresponding to geometries at depths 4 or 5, it is sufficient to decode attribute data unit 152 (there is no need to decode other attribute data units).

[0102] Similarly, by decoding the attribute data unit 151, it is possible to obtain attributes corresponding to geometries at depths 0 to 3. In other words, to obtain attributes corresponding to any of the geometries at depths 0 to 3, it is sufficient to decode the attribute data unit 151 (there is no need to decode other attribute data units).

[0103] In such a case, multiple geometry data units may be candidates for reference for a single attribute data unit, such as attribute data unit 152 and attribute data unit 153. Therefore, when non-scalable coding is applied to an attribute, the attribute geometry slice identification information is set to the slice identification information of the geometry slice corresponding to the lowest layer (highest depth) of the depths corresponding to that attribute slice (attribute data unit). The reference relationship between geometry slices (reference relationship with higher-level geometry slices) is indicated by ref_geom_slice_id.

[0104] In this specification, the attribute geometry slice identification information when non-scalable coding is applied to the attribute is also referred to as non-scalable coding attribute geometry slice identification information.

[0105] In other words, as shown in the 11th row from the top of the table in Figure 5, the slice dependency information may include non-scalable coding attribute geometry slice identification information, which is identification information of the geometry slice referenced by the attribute slice to which non-scalable coding is applied (method 1-1-2-4).

[0106] For example, as shown in the twelfth row from the top of the table in Fig. 5, this non-scalable coding attribute geometry slice identification information may include identification information of the geometry slice or geometry slice group including the geometry with the maximum depth information among the geometry slices or geometry slice groups referenced by the attribute slice corresponding to that information (Method 1-1-2-4-1). Note that in this specification, a geometry slice group refers to multiple geometry slices that correspond to the same depth.

[0107] In this way, even when non-scalable coding is applied to attributes, restrictions on the positional relationship between geometry data units and attribute data units can be eliminated.

[0108] For example, as shown in the thirteenth row from the top of the table in Fig. 5, non-scalable coding attribute geometry slice identification information may be stored in the codex specific parameters of the subsample information box in the metadata area of ​​the content file (method 1-1-2-4-2). For example, a content file generation unit of an encoder may set a subsample for each slice and store non-scalable coding attribute geometry slice identification information in the codex specific parameters of the subsample information box in the metadata area of ​​the content file. Furthermore, an extraction unit of a decoder may extract any slice of the G-PCC content from the content file based on the non-scalable coding attribute geometry slice identification information stored in the codex specific parameters of the subsample information box in the metadata area of ​​the content file for the subsample set for each slice.

[0109] In this way, by setting a subsample for each slice and storing non-scalable coding attribute geometry slice identification information in the codec specific parameters of the subsample information box in the metadata area of ​​the content file, it is possible to suppress an increase in file size. Furthermore, since the decoder only needs to check the codec specific parameters of the desired subsample information box, even if non-scalable coding is applied to the attribute, it is possible to more easily check the geometry slice referenced by the attribute. In other words, it is possible to suppress an increase in the load of the playback process.

[0110] Furthermore, as shown in the 14th row from the top of the table in Fig. 5, the slice dependency information may include a non-scalable coding flag (method 1-1-2-4-3). The non-scalable coding flag is flag information indicating whether or not non-scalable coding has been applied to an attribute slice. By storing such information, a decoder can easily determine whether or not non-scalable coding has been applied.

[0111] 5, a non-scalable encoding flag may be stored in the codex specific parameters of the subsample information box in the metadata area of ​​the content file (Method 1-1-2-4-3-1). For example, a content file generation unit of an encoder may set a subsample for each slice and store a non-scalable encoding flag in the codex specific parameters of the subsample information box in the metadata area of ​​the content file. Furthermore, an extraction unit of a decoder may extract an arbitrary slice of the G-PCC content from the content file based on the non-scalable encoding flag stored in the codex specific parameters of the subsample information box in the metadata area of ​​the content file for the subsample set for each slice.

[0112] An example of the syntax of the codex specific parameters in this case is shown in Fig. 12. As shown in a dotted frame 161, the codex specific parameters store a non-scalable coding flag (non_scalable_flag) and non-scalable coding attribute geometry slice identification information (ref_attr_geom_slice_id).

[0113] If the attribute slice is scalably coded, the value of non_scalable_flag is set to 0 (false). If the attribute slice is non-scalable coded, the value of non_scalable_flag is set to 1 (true).

[0114] If the value of non_scalable_flag is 1 (true), ref_attr_geom_slice_id is set as non-scalable coding attribute geometry slice identification information. That is, in this case, ref_attr_geom_slice_id is set to the identification information (slice_id) of the geometry slice or geometry slice group including the maximum depth among the geometry slices or geometry slice groups referenced by the attribute slice corresponding to that information.

[0115] If the value of non_scalable_flag is 0 (false), ref_attr_geom_slice_id is set as attribute geometry slice identification information. That is, in this case, ref_attr_geom_slice_id is set to the slice identification information (slice_id) of the geometry slice referenced by the attribute slice to which the information corresponds.

[0116] <Change payload type> In addition, in the Codex Specific Parameters, as shown in the bottom row of the table in Figure 5, the payload type (PayloadType) of the slice dependency information of an independent geometry slice and the payload type of the slice dependency information of a dependent geometry slice may be set to different values.

[0117] 13 is a diagram showing an example of the syntax of the Codex Specific Parameters in this case. In this example, as shown in the dotted-line box 162, the payload type of the slice dependency information of the independent geometry slice is set to "2," and the payload type of the slice dependency information of the dependent geometry slice is set to "9." This allows the decoder to more easily identify the type of geometry slice based on the payload type.

[0118] 13, the payload type values ​​are described as "2" and "9", but these values ​​are merely examples. The payload type values ​​of the independent geometry slice and the dependent geometry slice may be different from each other, and are not limited to this example.

[0119] <2-1-2. Geometry slice depth information> As shown in the top row of the table in Fig. 6, the slice configuration information may include geometry slice depth information (method 1-1-3). In this specification, geometry slice depth information refers to information about depth information of geometry included in the geometry slice or geometry slice group corresponding to the information. The slice configuration information may include both slice dependency relationship information (method 1-1-2) and geometry slice depth information.

[0120] For example, in Figure 7, the geometry data unit of "geom_slice#1" corresponds to slice #1. Slice #1 includes geometries at depths 0 to 3. The geometry data unit of "geom_slice#2" corresponds to slice #2. Slice #2 includes geometries at depths 4 and 5. The geometry data unit of "geom_slice#6" corresponds to slice #6. Slice #6 includes geometry at depth 6. The geometry data unit of "geom_slice#3" corresponds to slice #3. Slice #3 includes geometries at depths 4 and 5. The geometry data unit of "geom_slice#7" corresponds to slice #7. Slice #7 includes geometry at depth 6.

[0121] That is, the depth information of each slice shown within the dotted frame 122 is geometry slice depth information. This geometry slice depth information is stored in the metadata area of ​​the content file. In this way, a decoder can obtain this geometry slice depth information from the metadata area of ​​the content file. Therefore, the decoder can understand the geometry slice depth information without parsing the bitstream (i.e., more easily).

[0122] Note that a geometry slice may contain geometry at multiple depths.

[0123] Therefore, as shown in the second row from the top of the table in FIG. 6, the geometry slice depth information may include minimum depth information (method 1-1-3-1). In this specification, minimum depth information refers to information indicating the minimum value of depth information in the geometry slice or geometry slice group corresponding to that information. In this way, the decoder can obtain this minimum depth information from the metadata area of ​​the content file. Therefore, the decoder can determine the minimum depth value included in the geometry slice without parsing the bitstream (i.e., more easily).

[0124] Furthermore, as shown in the third row from the top of the table in FIG. 6, the geometry slice depth information may include maximum depth information (method 1-1-3-2). In this specification, maximum depth information refers to information indicating the maximum value of depth information in the geometry slice or geometry slice group corresponding to that information. In this way, the decoder can obtain this maximum depth information from the metadata area of ​​the content file. Therefore, the decoder can determine the maximum depth value included in the geometry slice without parsing the bitstream (i.e., more easily).

[0125] Of course, the geometry slice depth information may include both minimum and maximum depth information, allowing a decoder to determine the depth range contained in the geometry slice without parsing the bitstream (i.e., more easily).

[0126] As shown in the fourth row from the top of the table in Fig. 6, this geometry slice depth information may be stored in the codex specific parameters of the subsample information box in the metadata area of ​​the content file (method 1-1-3-3). For example, a content file generation unit of an encoder may set a subsample for each slice or slice group and store the geometry slice depth information in the codex specific parameters of the subsample information box in the metadata area of ​​the content file. Alternatively, an extraction unit of a decoder may extract any slice of the G-PCC content from the content file based on the geometry slice depth information stored in the codex specific parameters of the subsample information box in the metadata area of ​​the content file for the subsample set for each slice or slice group.

[0127] An example of the syntax of the codex specific parameters in this case is shown in Fig. 14. In the codex specific parameters shown in Fig. 14, minimum depth information (min_depth) and maximum depth information (max_depth) are set, as shown in a dotted line frame 171.

[0128] In this way, by setting a subsample for each slice or slice group and storing geometry slice depth information in the codec specific parameters of the subsample information box in the metadata area of ​​the content file, it is possible to suppress an increase in file size. Also, since the decoder only needs to check the codec specific parameters of the desired subsample information box, it is possible to more easily confirm the depth included in the geometry slice. In other words, it is possible to suppress an increase in the load of the playback process.

[0129] Note that when multiple geometry slices with the same depth range are consecutive, they are treated as one subsample in a geometry slice group as described above. Geometry slice depth information is then generated for each subsample. That is, the geometry slice depth information for each slice in the geometry slice group is combined into one piece of geometry slice depth information. Therefore, compared to generating geometry slice depth information for each slice in the geometry slice group, an increase in the number of subsample entries can be suppressed. Therefore, an increase in bit cost (i.e., the amount of data in the bitstream) can be suppressed.

[0130] It should be noted that when a geometry slice or a group of geometry slices includes only one depth, the minimum depth information (min_depth) and the maximum depth information (max_depth) are set to the same value.

[0131] Furthermore, as shown in the fifth row from the top of the table in Figure 6, the flags of the subsample information box of Codex Specific Parameters that stores this geometry slice depth information and the flags of the subsample information box of Codex Specific Parameters that stores the slice dependency information may be set to different values ​​(Method 1-1-3-3-1).

[0132] That is, the slice configuration information may further include slice dependency relationship information indicating the dependency relationship between the first slice and the second slice, in addition to the geometry slice depth information.The content file generation unit of the encoder may set the flags of the subsample information box storing the geometry slice depth information to a value different from the flags of the subsample information box storing the slice dependency information.Furthermore, the flags of the subsample information box storing the geometry slice depth information and the flags of the subsample information box storing the slice dependency information may be set to different values.

[0133] For example, as shown in Fig. 14, the flags of the subsample information box storing the slice dependency information is set to "0," whereas the flags of the subsample information box storing the geometry slice depth information is set to "2." Of course, these values ​​are merely examples, and the value of flags is not limited to this example (it is arbitrary).

[0134] In this way, the geometry slice depth information and the slice dependency information can be stored in different subsample information boxes having different flag values, and these can be used in combination.

[0135] <2-2. Track configuration information> As shown in the sixth row from the top of the table in FIG. 6, the scalable decoding information may include track configuration information (method 1-2). In this specification, track configuration information is information about the configuration of tracks in a content file that store G-PCC content in slice units. That is, a decoder can obtain the configuration information of tracks included in a content file from the metadata area of ​​the content file. Therefore, the decoder can understand the track configuration of the content file without parsing the bitstream (i.e., more easily). Note that the scalable decoding information may include both slice configuration information for each sample (method 1-1) and track configuration information.

[0136] <2-2-1. Track depth information> As shown in the seventh row from the top of the table in Fig. 6, the track configuration information may include track depth information (Method 1-2-1). In this specification, track depth information refers to information regarding depth information of the geometry of all slices included in the track corresponding to that information.

[0137] 15 is a diagram showing an example of how bitstreams are stored. It is assumed that a content file (ISOBMFF) has track 1 indicated by a rectangle 181, track 2 indicated by a rectangle 182, and track 3 indicated by a rectangle 183. It is assumed that track 1 stores geometry bitstreams and attribute bitstreams for depths 0 to 3 (depth=0 to 3). It is assumed that track 2 stores geometry bitstreams and attribute bitstreams for depths 4 and 5 (depth=4 to 5). It is assumed that track 3 stores geometry bitstreams and attribute bitstreams for depth 6 (depth=6).

[0138] In this case, each data unit (slice) of the bitstream 101 is stored in each track as indicated by the dotted arrows in Figure 15. That is, the geometry and attributes of slice #1 are stored in track 1. The geometry and attributes of slice #2 are stored in track 2. The geometry and attributes of slice #3 are stored in track 2. The geometry and attributes of slice #6 are stored in track 3. The geometry and attributes of slice #7 are stored in track 7.

[0139] In this way, the track depth information indicates which depth data is stored in each track. The track depth information is set for each track. In other words, the track depth information indicates which depth data is stored in the track corresponding to that information. As described above, data of multiple depths can be stored in a track. Also, data of multiple samples can be stored in a track. The slice configurations of each sample do not have to be identical to each other. In other words, the depth included in each track can vary for each sample.

[0140] As shown in the eighth row from the top of the table in Fig. 6, the track depth information may include track minimum depth information indicating the minimum value of the depth information in the track corresponding to that information (Method 1-2-1-1). In other words, the track minimum depth information indicates the minimum value of the depth information among all samples included in the track corresponding to that information.

[0141] As shown in the ninth row from the top of the table in FIG. 6, the track depth information may also include track maximum depth information indicating the maximum value of depth information in the track corresponding to that information (Method 1-2-1-2). In other words, the track maximum depth information indicates the maximum value of depth information among all samples included in the track corresponding to that information. Note that the track depth information may also include both track minimum depth information (Method 1-2-1-1) and track maximum depth information.

[0142] Furthermore, as shown in the tenth row from the top of the table in FIG. 6, the track depth information may include a match flag (method 1-2-1-3). In this specification, this match flag is flag information indicating whether sample minimum depth information, which is the minimum value of depth information for each sample included in the track corresponding to the information, matches the track minimum depth information, and whether sample maximum depth information, which is the maximum value of depth information for each sample included in the track corresponding to the information, matches the track maximum depth information. In other words, the match flag is flag information indicating whether the minimum and maximum values ​​of depth information are common to all samples in the track. Note that the track depth information may include all of this match flag, track minimum depth information (method 1-2-1-1), and track maximum depth information (method 1-2-1-2).

[0143] Alternatively, as shown in the eleventh row from the top of the table in Fig. 6, track depth information may be stored in a depth information box (DepthInfoBox) of a sample entry (SampleEntry) corresponding to each track in the metadata area (Method 1-2-1-4). For example, a content file generation unit of an encoder may store track depth information in a depth information box of a sample entry in the metadata area. Alternatively, an extraction unit of a decoder may extract any slice of the G-PCC content from the content file based on the track depth information stored in the depth information box of the sample entry in the metadata area.

[0144] Fig. 16 is a diagram illustrating an example of the syntax of a depth information box (DepthInfoBox). As shown in Fig. 16, in the depth information box, track minimum depth information (track_min_depth), track maximum depth information (track_max_depth), and a match flag (fixed_depth) are set.

[0145] If the match flag is "0" (false), the sample minimum depth information and sample maximum depth information of all samples in the track take values ​​within the range of the track minimum depth information (track_min_depth) to the track maximum depth information (track_max_depth). In other words, in this case, the minimum and / or maximum depth values ​​of each sample can vary from sample to sample.

[0146] If the match flag is "1" (true), it indicates that the sample minimum depth information of each sample in the track matches the track minimum depth information (track_min_depth), and the sample maximum depth information of each sample in the track matches the track maximum depth information (track_min_depth). In other words, in this case, the minimum and maximum depth values ​​of each sample are common values ​​for all samples.

[0147] In addition, if a track contains data of only one depth, such as track 3, the value of the match flag (fixed_depth) is set to "1", and the track minimum depth information (track_min_depth) and the track maximum depth information (track_max_depth) have the same value (track_min_depth=track_max_depth).

[0148] As described above, by storing track depth information in the depth information box of the sample entry corresponding to each track, the decoder can more easily determine (without parsing the bitstream) which depth data is stored in each track based on that information. In other words, it is possible to suppress an increase in the load of the playback process.

[0149] For example, if fixed_depth=1 and the desired LoD is within the range of track_min_depth to track_max_depth, the desired LoD can be reliably obtained by processing this track and the tracks it references. In other words, this track depth information can be useful information when the client selects a track.

[0150] Note that instead of depth, the maximum and minimum LoD values ​​obtained by processing this track and the reference track (if any) may be signaled.

[0151] A non-scalable coding attribute flag (non_scalable_attribute_flag) may also be added to indicate whether a track contains attribute slices to which non-scalable coding has been applied. If this non-scalable coding attribute flag (non_scalable_attribute_flag) is "1" (true), it indicates that the track contains attribute slices to which non-scalable coding has been applied. If this non-scalable coding attribute flag (non_scalable_attribute_flag) is "0" (false), it indicates that the track does not contain attribute slices to which non-scalable coding has been applied.

[0152] <2-2-2. Track dependency information> As shown in the twelfth row from the top of the table in Fig. 6, the track configuration information may include track dependency information (method 1-2-2). In this specification, track dependency information is information indicating the dependency relationship between tracks (for example, the dependency relationship between a first track and a second track). Note that the track configuration information may include both track depth information (method 1-2-1) and track dependency information.

[0153] For example, in FIG. 15, the tracks (track 1 to track 3) have inter-track dependency relationships as indicated by arrows 184 and 185.

[0154] For example, track 2 is dependent on track 1, as indicated by arrow 184. That is, to decode the bitstream of track 2 and restore the geometry and attributes of depth 4 or depth 5, it is necessary to also decode the bitstream of track 1, which corresponds to depths 0 to 3.

[0155] Similarly, track 3 directly depends on track 2, as indicated by arrow 185. That is, track 3 indirectly depends on track 1, as indicated by arrow 186. The track dependency information indicates such inter-track dependency relationships.

[0156] As shown in the thirteenth row from the top of the table in FIG. 6, this track dependency information may also include dependency information indicating other tracks containing slices necessary for decoding the dependent slices included in the track corresponding to the information (Method 1-2-2-1). The dependency information, as track dependency information on the start side of the above-mentioned arrows (arrows 184 to 186), indicates the dependency destination, i.e., the track on the end side of the above-mentioned arrows. By storing such dependency information in the metadata area of ​​the content file, the decoder can more easily identify other tracks necessary for decoding the track being processed based on the dependency information without parsing the bitstream. This can suppress an increase in the load of the playback process.

[0157] As shown in the 14th row from the top of the table in Fig. 6, this dependency information may indicate all other tracks that contain slices necessary for decoding the dependency slices included in the track corresponding to that information (method 1-2-2-1-1). For example, in the case of Fig. 15, the dependency information of track 3 may include both the information corresponding to arrow 185 and the information corresponding to arrow 186.

[0158] 6, the dependency information may indicate other tracks that include slices referenced by the information (method 1-2-2-1-2). In other words, the dependency information may indicate only slices to which the slice corresponding to the information directly depends. For example, in the case of FIG. 15, the dependency information of track 3 may include only information corresponding to arrow 185.

[0159] As shown in the bottom row of the table in FIG. 6, the track dependency information may be stored as a track reference in the metadata area.

[0160] For example, the dependent information as shown in Fig. 15 may be stored as a track reference. In this case, the reference type of the track reference may be, for example, depd (reference_type = 'depd').

[0161] Alternatively, a track reference may be used to link only the track containing the slice that is directly referenced when decoding the dependent slice contained in that track.

[0162] The Sample Entry 4CC of a track that contains an independent slice and is independently decodable may be specified as 'gpc1', and the Sample Entry 4CC of a track that does not contain an independent slice and is not independently decodable may be specified as 'lgp1'.

[0163] Furthermore, as shown in the 16th row from the top of the table in Figure 6, the track dependency information may include independent information indicating other tracks that contain dependent slices that require independent slices included in the track corresponding to that information when decoding (method 1-2-2-2).

[0164] For example, in Figure 15, an independent slice in track 1 is referenced when decoding a dependent slice in track 2, as indicated by arrow 184. Similarly, an independent slice in track 1 is referenced when decoding a dependent slice in track 3, as indicated by arrow 186.

[0165] 17, as indicated by arrow 191, an independent slice stored in track 1 is used to decode a dependent slice in track 2. Also, as indicated by arrow 192, an independent slice stored in track 1 is used to decode a dependent slice in track 3. Independent information is information that indicates such a dependent relationship. In other words, independent information is reverse lookup information for dependent information.

[0166] Note that track 3 is indirectly dependent on track 1. Therefore, the dependent relationship indicated by arrow 192 may or may not be included in the independent information. That is, as shown in the 17th row from the top of the table in FIG. 6, the independent information may indicate other tracks including other slices that reference the independent slice during decoding. Furthermore, the track dependency information may include both dependent information and independent information.

[0167] As described above, the track dependency information may be stored as a track reference in the metadata area. For example, the independent information shown in FIG. 17 may be stored as a track reference. In this case, the reference type of the track reference may be, for example, indd (reference_type='indd'). By using this list of track references, the order of slice arrangement when reconstructing a G-PCC bitstream from the slices stored in each track can be stored in the metadata area.

[0168] <2-3. Matryoshka Media Container> Although the above describes an example in which ISOBMFF is used as the file format, the file in which the G-PCC bitstream is stored may be any format other than ISOBMFF. For example, the G-PCC content may be stored in a Matroska Media Container. An example of the main configuration of a Matroska Media Container is shown in Fig. 18.

[0169] In this case, for example, the tile management information (tile identification information) may be stored as a newly defined element under the track entry element. Also, when the tile management information (tile identification information) is stored in timed metadata, the timed metadata may be stored in a track entry different from the track entry in which the G-PCC content is stored.

[0170] 3. First Embodiment <3-1. File generation device> An encoding device will be described. The present technology (each method) described above can be applied to any device. Fig. 19 is a block diagram showing an example of the configuration of a file generation device, which is one aspect of an information processing device to which the present technology is applied. The file generation device 300 shown in Fig. 19 is a device that applies G-PCC to encode point cloud data and stores the G-PCC content (G-PCC bitstream) generated by the encoding in a content file (ISOBMFF).

[0171] In this case, the file generation device 300 applies the present technology described above in Chapter <2. Transmission of scalable decoding information by content file>. That is, the file generation device 300 generates scalable decoding information based on the depth of slices in the G-PCC content and the dependency relationships between the slices, generates a content file for storing the G-PCC content, and stores the generated scalable decoding information in the metadata area of ​​the generated content file.

[0172] Note that Fig. 19 shows the main processing units, data flows, etc., and does not necessarily include everything shown in Fig. 19. In other words, in file generation device 300, there may be processing units that are not shown as blocks in Fig. 19, and there may be processing or data flows that are not shown as arrows, etc. in Fig. 19.

[0173] 19, the file generation device 300 includes an extraction unit 311, an encoding unit 312, a bitstream generation unit 313, a scalable decoding information generation unit 314, and a file generation unit 315. The encoding unit 312 also includes a geometry encoding unit 321, an attribute encoding unit 322, and a metadata generation unit 323.

[0174] The extraction unit 311 extracts geometry data and attribute data from the point cloud data input to the file generation device 300. The extraction unit 311 supplies the extracted geometry data to a geometry encoding unit 321 of the encoding unit 312. The extraction unit 311 also supplies the extracted attribute data to an attribute encoding unit 322 of the encoding unit 312.

[0175] The encoding unit 312 encodes the point cloud data. The geometry encoding unit 321 encodes the geometry data supplied from the extraction unit 311 to generate a geometry bitstream. The geometry encoding unit 321 supplies the generated geometry bitstream to the metadata generation unit 323. The geometry encoding unit 321 also supplies the generated geometry bitstream to the attribute encoding unit 322.

[0176] The attribute encoding unit 322 encodes the attribute data supplied from the extraction unit 311 to generate an attribute bitstream. The attribute encoding unit 322 supplies the generated attribute bitstream to the metadata generation unit 323.

[0177] The metadata generation unit 323 generates metadata by referring to the supplied geometry bitstream and attribute bitstream, and supplies the generated metadata to the bitstream generation unit 313 together with the geometry bitstream and attribute bitstream.

[0178] The bitstream generation unit 313 multiplexes the supplied geometry bitstream, attribute bitstream, and metadata to generate G-PCC content (G-PCC bitstream). The bitstream generation unit 313 supplies the generated G-PCC content to the scalable decoding information generation unit 314.

[0179] The scalable decoding information generation unit 314 acquires the G-PCC content including the first slice and the second slice supplied from the bitstream generation unit 313. The scalable decoding information generation unit 314 applies the present technology described above in Chapter <2. Transmission of Scalable Decoding Information by Content File> to generate scalable decoding information for scalable decoding of the G-PCC content based on depth information indicating the quality hierarchical level of the geometry included in each slice in the G-PCC content and the dependency relationship between the first slice and the second slice in the G-PCC content. The scalable decoding information generation unit 314 supplies the generated scalable decoding information to the file generation unit 315 together with the G-PCC content.

[0180] The file generation unit 315 applies the present technology described above in <2. Transmission of scalable decoding information by content file> to generate a content file that stores the G-PCC content supplied from the scalable decoding information generation unit 314, and stores the scalable decoding information in a metadata area of ​​the generated content file. The file generation unit 315 outputs the content file generated as described above to the outside of the file generation device 300.

[0181] The scalable decoding information generation unit 314 may generate scalable decoding information including slice configuration information for each sample. The file generation unit 315 may set a subsample for each slice and store the slice configuration information in the codec specific parameters of the subsample information box in the metadata area of ​​the content file.

[0182] The scalable decoding information generator 314 may generate slice configuration information including slice dependency information. The file generator 315 may set a subsample for each slice and store the slice dependency information in the codec specific parameters of the subsample information box in the metadata area of ​​the content file.

[0183] The scalable decoding information generator 314 may generate slice dependency relationship information including reference geometry slice identification information and referenced geometry slice identification information. If a slice corresponding to the slice dependency relationship information (i.e., a slice corresponding to the reference geometry slice identification information) is an independent geometry slice, the reference geometry slice identification information and the referenced geometry slice identification information may be the same.

[0184] The file generation unit 315 may set a subsample for each slice, and store the reference source geometry slice identification information and the reference destination geometry slice identification information in the codec specific parameters of the subsample information box in the metadata area.

[0185] The scalable decoding information generator 314 may generate slice dependency information including attribute geometry slice identification information. The file generator 315 may set a subsample for each slice and store the attribute geometry slice identification information in the codec specific parameters of the subsample information box in the metadata area.

[0186] The scalable decoding information generation unit 314 may generate slice dependency information including non-scalable encoding attribute geometry slice identification information, which is identification information of a geometry slice referenced by an attribute slice to which non-scalable encoding is applied.

[0187] The scalable decoding information generation unit 314 may generate non-scalable encoding attribute geometry slice identification information including identification information of a geometry slice or a group of geometry slices including a geometry with the maximum depth information among geometry slices or groups of geometry slices referenced by an attribute slice corresponding to the information. The file generation unit 315 may set a subsample for each slice and store the non-scalable encoding attribute geometry slice identification information in the codec specific parameters of the subsample information box in the metadata area of ​​the content file.

[0188] Alternatively, the scalable decoding information generation unit 314 may generate slice dependency information including a non-scalable encoding flag. Alternatively, the file generation unit 315 may set a subsample for each slice and store the non-scalable encoding flag in the codec specific parameters of the subsample information box in the metadata area of ​​the content file.

[0189] Note that the file generation unit 315 may set the payload type (PayloadType) of the slice dependency information of the independent geometry slice and the payload type of the slice dependency information of the dependent geometry slice to different values ​​in the Codex specific parameters.

[0190] The scalable decoding information generation unit 314 may generate slice configuration information including geometry slice depth information. In this case, the scalable decoding information generation unit 314 may generate geometry slice depth information including minimum depth information. Alternatively, the scalable decoding information generation unit 314 may generate geometry slice depth information including maximum depth information. Then, the file generation unit 315 may set a subsample for each slice or slice group and store this geometry slice depth information in the codec specific parameters of the subsample information box in the metadata area of ​​the content file.

[0191] The scalable decoding information generation unit 314 may generate slice configuration information that further includes slice dependency relationship information indicating the dependency relationship between the first slice and the second slice in addition to the geometry slice depth information.The file generation unit 315 may then set the flags of the subsample information box that stores the geometry slice depth information to a value different from the flags of the subsample information box that stores the slice dependency relationship information.

[0192] The scalable decoding information generator 314 may generate scalable decoding information including track configuration information. The scalable decoding information generator 314 may generate track configuration information including track depth information.

[0193] The scalable decoding information generator 314 may generate track depth information including track minimum depth information indicating the minimum value of depth information in the track corresponding to the information. The scalable decoding information generator 314 may also generate track depth information including track maximum depth information indicating the maximum value of depth information in the track corresponding to the information. Furthermore, the scalable decoding information generator 314 may generate track depth information including a match flag. The file generator 315 may then store such track depth information in a depth information box of a sample entry in the metadata area.

[0194] The scalable decoding information generator 314 may generate track configuration information including track dependency information. Alternatively, the scalable decoding information generator 314 may generate track dependency information including dependency information indicating other tracks that include slices necessary for decoding a dependent slice included in a track corresponding to the information. Alternatively, the scalable decoding information generator 314 may generate dependency information that indicates all other tracks that include slices necessary for decoding a dependent slice included in a track corresponding to the information. Alternatively, the scalable decoding information generator 314 may generate dependency information that indicates other tracks that include a slice referenced from the information.

[0195] The scalable decoding information generator 314 may generate track dependency information including independent information indicating other tracks including dependent slices that require an independent slice included in a track corresponding to the information for decoding. The scalable decoding information generator 314 may generate independent information indicating other tracks including other slices that reference the independent slice for decoding. Alternatively, the scalable decoding information generator 314 may generate independent information in which the track dependency information includes both dependent information and independent information.

[0196] Then, the file generator 315 may store the track dependency information as a track reference in the metadata area.

[0197] By doing so, as described above in the chapter <2. Transmission of scalable decoding information by content file>, it is possible to suppress an increase in the load of the playback process.

[0198] <File generation process flow> An example of the flow of the file generation process executed by this file generation device 300 will be described with reference to the flowchart of FIG.

[0199] When the file generation process starts, the extraction unit 311 of the file generation device 300 extracts geometry and attributes from the point cloud in step S301.

[0200] In step S302, the encoding unit 312 encodes the geometry and attributes extracted in step S301 to generate a geometry bitstream and an attribute bitstream, and further generates metadata for the encoded geometry and attribute bitstream.

[0201] In step S303, the bitstream generation unit 313 multiplexes the geometry bitstream, attribute bitstream, and metadata generated in step S302 to generate a G-PCC bitstream (G-PCC content).

[0202] In step S304, the scalable decoding information generation unit 314 applies the present technology described above in chapter <2. Transmission of scalable decoding information via content file>, and generates scalable decoding information regarding scalable decoding of the G-PCC content based on the depth of the slices in the G-PCC content generated in step S303 and the dependency relationships between slices in the G-PCC content.

[0203] In step S305, the file generation unit 315 generates other information and generates a content file (e.g., ISOBMFF) that stores the G-PCC content generated in step S303. Then, the file generation unit 315 applies the present technology described above in chapter <2. Transmission of scalable decoding information by content file>, and stores the scalable decoding information generated in step S304 in the metadata area of ​​the generated content file.

[0204] In step S306, the file generation unit 315 outputs the generated content file (a content file storing scalable decoding information) to the outside of the file generation device 300. For example, the file generation unit 315 transmits the content file to another device (e.g., a playback device) via a network or the like. Alternatively, for example, the file generation unit 315 supplies the content file to a storage medium external to the file generation device 300 for storage. In this case, the content file is supplied to the playback device or the like via the storage medium.

[0205] When the process of step S306 is completed, the file generation process ends.

[0206] As described above, in the file generation process, the file generation device 300 applies the present technology described in Chapter <2. Transmission of scalable decoding information by content file> and stores the scalable decoding information in the metadata area of ​​the content file. In this way, it is possible to reduce the processing (decoding, etc.) of unnecessary information, and to suppress an increase in the load of the playback process.

[0207] <3-2. Playback device> Fig. 21 is a block diagram showing an example of the configuration of a playback device, which is one aspect of an information processing device to which the present technology is applied. The playback device 400 shown in Fig. 21 is a device that decodes a G-PCC file, constructs a point cloud, and performs rendering to generate presentation information. In this case, the playback device 400 applies the present technology described above in Chapter <2. Transmission of scalable decoded information using a content file>, extracts slices necessary for playing back a desired depth in the point cloud from the content file generated by the file generation device 300, decodes them, and plays them back.

[0208] Note that Fig. 21 shows the main processing units, data flows, etc., and does not necessarily show everything. In other words, playback device 400 may have processing units that are not shown as blocks in Fig. 21, or may have processing or data flows that are not shown as arrows, etc. in Fig. 21.

[0209] 21, the playback device 400 has a control unit 401, a file acquisition unit 411, a playback processing unit 412, and a presentation processing unit 413. The playback processing unit 412 has a file processing unit 421, a decoding unit 422, and a presentation information generation unit 423.

[0210] The control unit 401 controls each processing unit in the playback device 400. The file acquisition unit 411 acquires a content file that stores the point cloud to be played back, and supplies it to (the file processing unit 421 of) the playback processing unit 412. The playback processing unit 412 performs processing related to the playback of the point cloud stored in the supplied content file.

[0211] The file processing unit 421 of the playback processing unit 412 acquires the content file supplied from the file acquisition unit 411 and extracts a bitstream from the content file. At that time, the file processing unit 421 applies the present technology described above in Chapter <2. Transmission of scalable decoding information via content file> to extract only the bitstream of the slices necessary for playback of the desired depth. The file processing unit 421 supplies the extracted bitstream to the decoding unit 422.

[0212] The decoding unit 422 decodes the bitstream supplied from the file processing unit 421 and generates geometry and attribute data. The decoding unit 422 supplies the generated geometry and attribute data to the presentation information generation unit 423. The presentation information generation unit 423 constructs a point cloud using the supplied geometry and attribute data and generates presentation information, which is information for presenting (e.g., displaying) the point cloud. For example, the presentation information generation unit 423 performs rendering using the point cloud and generates, as presentation information, a display image of the point cloud viewed from a predetermined viewpoint. The presentation information generation unit 423 supplies the presentation information generated in this manner to the presentation processing unit 413.

[0213] The presentation processing unit 413 performs processing to present the supplied presentation information. For example, the presentation processing unit 413 supplies the presentation information to a display device or the like external to the playback device 400, and causes the presentation information to be presented.

[0214] <Playback Processing Unit> Fig. 22 is a block diagram showing an example of the main configuration of the playback processing unit 412. As shown in Fig. 22, the file processing unit 421 has a bitstream extraction unit 431. The decoding unit 422 has a geometry decoding unit 441 and an attribute decoding unit 442. The presentation information generation unit 423 has a point cloud construction unit 451 and a presentation processing unit 452.

[0215] The bitstream extraction unit 431 extracts a bitstream from the content file supplied from the file acquisition unit 411. In doing so, the bitstream extraction unit 431 applies the present technology described above in Chapter <2. Transmission of Scalable Decoding Information via Content File> to extract only the bitstream of the slices required for playback of the desired depth. That is, the bitstream extraction unit 431 extracts any slice of the G-PCC content from the content file storing the G-PCC content based on scalable decoding information stored in the metadata area of ​​the content file. Note that this scalable decoding information is information related to the scalable decoding of the G-PCC content, and is generated based on depth information indicating the quality hierarchical levels of the geometry included in the slices in the G-PCC content and the dependency relationships between slices in the G-PCC content (e.g., between the first slice and the second slice).

[0216] The scalable decoding information may include slice configuration information for each sample. For example, the bitstream extraction unit 431 may determine the slice configuration within each sample based on the slice configuration information for each sample, and extract any slice of the G-PCC content from the content file based on the configuration. The bitstream extraction unit 431 may extract any slice of the G-PCC content from the content file based on slice configuration information stored in the codec specific parameters of the subsample information box in the metadata area of ​​the content file for the subsample set for each slice.

[0217] This slice configuration information may include slice dependency information. For example, the bitstream extraction unit 431 may understand the dependency relationships between slices or slice groups based on the slice dependency information and extract any slice of the G-PCC content from the content file based on the dependency relationships. The bitstream extraction unit 431 may extract any slice of the G-PCC content from the content file based on the slice dependency information stored in the codec specific parameters of the subsample information box in the metadata area of ​​the content file for the subsample set for each slice.

[0218] The slice dependency information may include reference-source geometry slice identification information and reference-destination geometry slice identification information. For example, the bitstream extraction unit 431 may determine the dependency relationship between slices or slice groups based on the reference-source geometry slice identification information and the reference-destination geometry slice identification information, and extract other geometry slices necessary for decoding a desired geometry slice of the G-PCC content from the content file based on the dependency relationship. If the slice corresponding to the slice dependency information (i.e., the slice corresponding to the reference-source geometry slice identification information) is an independent geometry slice, the reference-source geometry slice identification information and the reference-destination geometry slice identification information may be the same. The bitstream extraction unit 431 may extract any slice of the G-PCC content from the content file based on the reference-source geometry slice identification information and the reference-destination geometry slice identification information stored in the codec specific parameters of the subsample information box in the metadata area of ​​the subsample set for each slice.

[0219] The slice dependency information may include attribute geometry slice identification information. For example, the bitstream extraction unit 431 may identify a correspondence between an attribute slice and a geometry slice based on the attribute geometry slice identification information, and extract any slices (attribute slices and geometry slices) of the G-PCC content from the content file based on the identified correspondence. The bitstream extraction unit 431 may extract any slice of the G-PCC content from the content file based on the attribute geometry slice identification information stored in the codec specific parameters of the subsample information box in the metadata area of ​​the subsample set for each slice.

[0220] The slice dependency information may include non-scalable coding attribute geometry slice identification information, which is identification information of a geometry slice referenced by an attribute slice to which non-scalable coding has been applied. For example, the bitstream extraction unit 431 may determine a correspondence between an attribute slice to which non-scalable coding has been applied and a geometry slice based on the non-scalable coding attribute geometry slice identification information, and extract any slices of the G-PCC content (attribute slices and geometry slices to which non-scalable coding has been applied) from the content file based on the correspondence. This non-scalable coding attribute geometry slice identification information may include identification information of a geometry slice or a geometry slice group that includes the maximum depth among the geometry slices or geometry slice groups referenced by the attribute slice corresponding to the information. The bitstream extraction unit 431 may extract any slice of the G-PCC content from the content file based on the non-scalable coding attribute geometry slice identification information stored in the codec specific parameters of the subsample information box in the metadata area of ​​the content file for the subsample set for each slice.

[0221] The slice dependency information may include a non-scalable coding flag. For example, the bitstream extraction unit 431 may identify whether non-scalable coding has been applied based on the non-scalable coding flag, identify a correspondence between an attribute slice and a geometry slice based on the identification result, and extract any slices (attribute slices and geometry slices) of the G-PCC content from the content file based on the correspondence. The bitstream extraction unit 431 may extract any slices of the G-PCC content from the content file based on the non-scalable coding flag stored in the codec specific parameters of the subsample information box in the metadata area of ​​the content file for the subsample set for each slice.

[0222] In the Codex Specific Parameters, the payload type (PayloadType) of the slice dependency information of the independent geometry slice and the payload type of the slice dependency information of the dependent geometry slice may be different values. For example, the bitstream extraction unit 431 may identify whether the slice dependency information is of an independent geometry slice or of a dependent geometry slice based on the payload type, and analyze the slice dependency information based on the identification result.

[0223] The slice configuration information may include geometry slice depth information. For example, the bitstream extraction unit 431 may determine the depth of the geometry included in a geometry slice or a geometry slice group based on the geometry slice depth information, and extract any slice of the G-PCC content from the content file based on the depth information. The geometry slice depth information may include minimum depth information. The geometry slice depth information may include maximum depth information. The bitstream extraction unit 431 may extract any slice of the G-PCC content from the content file based on the geometry slice depth information stored in the codec specific parameters of the subsample information box in the metadata area of ​​the content file for the subsample set for each slice or slice group.

[0224] The slice configuration information may further include slice dependency relationship information indicating dependency relationships between slices in addition to the geometry slice depth information. The flags of the subsample information box storing the geometry slice depth information and the flags of the subsample information box storing the slice dependency relationship information may be set to different values. For example, the bitstream extraction unit 431 may identify whether the subsample information box stores the geometry slice depth information or the subsample information box stores the slice dependency relationship information based on the flags.

[0225] The scalable decoding information may include track configuration information. For example, the bitstream extraction unit 431 may determine the track configuration based on the track configuration information and extract any slice of the G-PCC content from the content file based on the track configuration. The track configuration information may include track depth information. For example, the bitstream extraction unit 431 may determine geometric depth information of all slices included in the track based on the track depth information, identify a track in which a desired slice is stored based on the depth information, and extract the desired slice from the track. The track depth information may include track minimum depth information indicating the minimum value of depth information in the track corresponding to the information. The track depth information may include track maximum depth information indicating the maximum value of depth information in the track corresponding to the information. The track depth information may include a match flag. For example, the bitstream extraction unit 431 may determine geometric depth information included in the track based on this information, identify a track in which a desired slice is stored based on the depth information, and extract the desired slice from the track. The bitstream extraction unit 431 may extract any slice of the G-PCC content from the content file based on the track depth information stored in the depth information box of the sample entry in the metadata area.

[0226] The track configuration information may include track dependency information. For example, the bitstream extraction unit 431 may understand track dependency relationships based on the track dependency information and extract any slice of the G-PCC content from the content file based on the dependency relationships. The track dependency information may include dependency information indicating other tracks containing slices necessary for decoding the dependent slices included in the track corresponding to the information. For example, the bitstream extraction unit 431 may identify other tracks containing slices necessary for decoding the dependent slices included in the track corresponding to the information based on the dependency information. The dependency information may indicate all other tracks containing slices necessary for decoding the dependent slices included in the track corresponding to the information. The dependency information may also indicate other tracks containing slices referenced by the information. The track dependency information may also include independent information indicating other tracks containing dependent slices required for decoding by the independent slices included in the track corresponding to the information. For example, the bitstream extraction unit 431 may identify other tracks containing dependent slices required for decoding by the independent slices included in the track corresponding to the information based on the independence information. The independent information may also indicate other tracks containing other slices that reference the independent slices when decoding. Additionally, the track dependency information may be stored as a track reference in the metadata area.

[0227] The bitstream extraction unit 431 supplies the extracted geometry bitstream to the geometry decoding unit 441. The bitstream extraction unit 431 also supplies the extracted attribute bitstream to the attribute decoding unit 442.

[0228] The geometry decoding unit 441 decodes the supplied geometry bitstream to generate geometry data. The geometry decoding unit 441 supplies the generated geometry data to the point cloud construction unit 451. The attribute decoding unit 442 decodes the supplied attribute bitstream to generate attribute data. The attribute decoding unit 442 supplies the generated attribute data to the point cloud construction unit 451.

[0229] The point cloud construction unit 451 constructs a point cloud using the supplied geometry and attribute data. That is, the point cloud construction unit 451 can construct a point cloud at a desired depth. The point cloud construction unit 451 supplies the constructed point cloud data to the presentation processing unit 452.

[0230] The presentation processing unit 452 generates presentation information using the supplied point cloud data. The presentation processing unit 452 supplies the generated presentation information to the presentation processing unit 413.

[0231] With this configuration, the playback device 400 can more easily extract, decode, construct, and present only desired tiles based on the tile management information (tile identification information) stored in the G-PCC file, without having to parse the entire bitstream, thereby reducing the load of playback processing.

[0232] <Recycling process flow> An example of the flow of the playback process executed by this playback device 400 will be described with reference to the flowchart of FIG.

[0233] When the playback process starts, the file acquisition unit 411 of the playback device 400 acquires a content file to be played back in step S401.

[0234] In step S402, the bitstream extraction unit 431 extracts any slice from the content file acquired in step S401. At this time, the bitstream extraction unit 431 applies the present technology described above in <2. Transmission of scalable decoding information by content file>, and extracts the slice based on the scalable decoding information stored in the metadata area of ​​the content file.

[0235] In step S403, the geometry decoding unit 441 of the decoding unit 422 decodes the geometry bitstream of the slice extracted in step S402 to generate a geometry of the desired depth. Also, the attribute decoding unit 442 decodes the attribute bitstream of the slice extracted in step S402 to generate an attribute corresponding to the geometry of the desired depth.

[0236] In step S404, the point cloud construction unit 451 constructs a point cloud using the geometry and attributes generated in step S403. That is, the point cloud construction unit 451 can construct a point cloud at a desired depth.

[0237] In step S405, the presentation processing unit 452 generates presentation information by performing rendering using the point cloud constructed in step S404, etc. In step S406, the presentation processing unit 413 supplies the presentation information to an external device outside the playback device 400, where it is presented.

[0238] When the process of step S406 ends, the playback process ends.

[0239] As described above, in the playback process, the playback device 400 applies the present technology described in Chapter <2. Transmission of scalable decoding information by content file> to extract and decode desired slices from a content file based on scalable decoding information stored in the metadata area of ​​the content file. In this way, it is possible to reduce the processing (decoding, etc.) of unnecessary information, and to suppress an increase in the load of the playback process.

[0240] <4. Transmission of scalable decoding information using control files> This technology can also be applied to, for example, MPEG-DASH (Moving Picture Experts Group phase - Dynamic Adaptive Streaming over HTTP). For example, in MPEG-DASH, scalable decoding information may be stored by extending the MPD (Media Presentation Description), which is a control file that stores control information related to bitstream distribution. For example, as the scalable decoding information, adaptation set configuration information, which is information related to the configuration of an adaptation set that describes information about the tracks of a content file, may be stored in the MPD.

[0241] In other words, as shown in the top row of the table in Figure 24, adaptation set configuration information based on the depth of each slice of G-PCC content having a slice structure and the dependencies between slices stored in the control file is transmitted (Method 2).

[0242] For example, an information processing device may include an adaptation set configuration information generation unit that generates adaptation set configuration information based on depth information indicating the quality hierarchical level of geometry included in each slice of G-PCC content including a first slice and a second slice and on the dependency relationship between the first slice and the second slice in the G-PCC content, and a control file generation unit that generates a control file for controlling playback of a content file that stores the G-PCC content and stores the adaptation set configuration information in the control file.The content file stores the G-PCC content in tracks on a slice-by-slice basis.The adaptation set configuration information is information about the configuration of an adaptation set that describes information about the tracks of the content file.

[0243] For example, in an information processing method, adaptation set configuration information is generated based on depth information indicating the quality hierarchical level of geometry included in each slice in G-PCC content including a first slice and a second slice, and on the dependency relationship between the first slice and the second slice in the G-PCC content, and a control file for controlling playback of a content file storing the G-PCC content is generated, and the adaptation set configuration information is stored in the control file.The content file then stores the G-PCC content in tracks on a slice-by-slice basis.The adaptation set configuration information is information about the configuration of an adaptation set that describes information about the tracks of the content file.

[0244] For example, an information processing device may include an analyzer that analyzes a control file that controls playback of a content file that stores G-PCC content including a first slice and a second slice in a track on a slice-by-slice basis, and identifies an adaptation set required to obtain an arbitrary slice of the G-PCC content based on adaptation set configuration information stored in the control file, an acquirer that acquires a track of the content file corresponding to the adaptation set identified by the analyzer, and a decoder that decodes the slice of the G-PCC content stored in the track acquired by the acquirer. The adaptation set configuration information is information about the configuration of the adaptation set that describes information about the track of the content file, and is information generated based on depth information indicating the quality hierarchical level of the geometry included in the slice of the G-PCC content and the dependency relationship between the first slice and the second slice in the G-PCC content.

[0245] For example, an information processing method analyzes a control file that controls playback of a content file that stores G-PCC content in tracks on a slice-by-slice basis, identifies an adaptation set required to obtain any slice of the G-PCC content based on adaptation set configuration information stored in the control file, acquires a track of the content file corresponding to the identified adaptation set, and decodes the slice of the G-PCC content stored in the acquired track. The adaptation set configuration information is information about the configuration of the adaptation set that describes information about the track of the content file, and is information generated based on depth information indicating the quality hierarchical level of the geometry included in the slice in the G-PCC content and the dependency between the first slice and the second slice in the G-PCC content.

[0246] In this way, the decoder can select a track storing slices necessary for reproducing a point cloud of a desired depth or region based on the adaptation set configuration information stored in the control file, and acquire the selected track. Therefore, it is possible to suppress the transmission of unnecessary data. This can suppress an increase in the load on the transmission path and communication processing. Furthermore, since the decoder can suppress an increase in the amount of data to be processed, it can suppress an increase in the load on the reproduction processing. This can suppress an increase in delays related to data transmission and reproduction processing.

[0247] As mentioned above in Chapter 2. Transmission of scalable decoding information using content files, applying this technology to particularly large-scale point clouds can further reduce the load on data transmission and playback processing, resulting in greater benefits.

[0248] <4-1. Adaptation Set Depth Information> As shown in the second row from the top of the table in Fig. 24, the adaptation set configuration information may include adaptation set depth information (method 2-1). The adaptation set depth information is information regarding the depth information of the geometry of all slices included in the track corresponding to the adaptation set corresponding to that information. This adaptation set depth information is the same as the track depth information described in <2-2-1. Track Depth Information>. In other words, the adaptation set depth information is the track depth information applied to the control file (MPD).

[0249] Based on this adaptation set depth information stored in the MPD, the decoder can easily (without parsing the bitstream) determine which slice of depth corresponds to which adaptation set (i.e., track), which allows the decoder to more easily (without parsing the bitstream) select a track to acquire.

[0250] In the case of an MPD, a new supplemental property or essential property may be defined, schemeIdUri = "urn:mpeg:mpegI:gpcc:2020:depth" may be set, and adaptation set depth information may be stored therein.

[0251] As shown in the third row from the top of the table in Fig. 24, the adaptation set depth information may include adaptation set minimum depth information (method 2-1-1). The adaptation set minimum depth information is information that indicates the minimum value of depth information in the track corresponding to the adaptation set corresponding to that information.

[0252] Fig. 25 is a diagram showing examples of parameters to be added to the MPD as adaptation set depth information. @minDepth shown in Fig. 25 is adaptation set minimum depth information, and is information similar to track_min_depth (track minimum depth information) of the track depth information described in <2-2-1. Track Depth Information>. In other words, @minDepth is track_min_depth applied to the control file (MPD).

[0253] As shown in the fourth row from the top of the table in Fig. 24, the adaptation set depth information may include adaptation set maximum depth information (method 2-1-2). The adaptation set maximum depth information is information that indicates the maximum value of depth information in the track corresponding to the adaptation set corresponding to that information.

[0254] 25 is adaptation set maximum depth information, and is the same information as track_max_depth (track maximum depth information) of the track depth information described in <2-2-1. Track Depth Information>. In other words, @maxDepth is track_max_depth applied to the control file (MPD).

[0255] As shown in the fifth row from the top of the table in FIG. 24, the adaptation set depth information may include a match flag (method 2-1-3). The match flag is flag information indicating whether the sample minimum depth information of each sample included in the track corresponding to the adaptation set corresponding to the information matches the adaptation set minimum depth information and whether the sample maximum depth information of each sample included in the track matches the adaptation set maximum depth information (flag information indicating whether the minimum depth and maximum depth are common to all samples). The sample minimum depth information indicates the minimum value of the depth information for the sample. The sample maximum depth information indicates the maximum value of the depth information for the sample.

[0256] 25 is a match flag for the adaptation set. When @fixedDepth is "0" (false), the sample minimum depth information and sample maximum depth information of all samples in the track corresponding to the adaptation set take values ​​within the range from the adaptation set minimum depth information (@minDepth) to the adaptation set maximum depth information (@maxDepth). In other words, in this case, the minimum or maximum depth value, or both, of each sample can change for each sample.

[0257] On the other hand, if @fixedDepth is "1" (true), the sample minimum depth information of each sample in the track matches the track minimum depth information (track_min_depth), and the sample maximum depth information of each sample in the track matches the track maximum depth information (track_min_depth). In other words, in this case, the minimum and maximum depth values ​​of each sample are common values ​​for all samples.

[0258] For example, if a track contains only one depth data, the value of the match flag (@fixedDepth) is set to "1", and the adaptation set minimum depth information (@minDepth) and the adaptation set maximum depth information (@maxDepth) have the same value (@minDepth = @maxDepth).

[0259] In the case of MPD, these parameters (@fixedDepth, @minDepth, @maxDepth) of the adaptation set depth information may be set in the above-mentioned Supplemental Property or Essential Property.

[0260] As described above, by storing this information as adaptation set depth information in the control file (MPD), the decoder can more easily (without parsing the bitstream) determine which depth data is stored in each track based on this information. Therefore, the decoder can more easily (without parsing the bitstream) select the track to acquire.

[0261] <4-2. Representation dependency information> As shown in the sixth row from the top of the table in Fig. 24, the adaptation set configuration information may include representation dependency information (method 2-2). Representation dependency information is information indicating a dependency relationship between representations (for example, a dependency relationship between a first representation and a second representation). In the MPD, tracks are managed by representations of adaptation sets. This representation dependency information is the same as the track dependency information described in <2-2-2. Track dependency information>. In other words, the representation dependency information is track dependency information applied to the control file (MPD).

[0262] Based on this representation dependency information stored in the MPD, a decoder can more easily determine the dependencies between tracks (without parsing the bitstream), and therefore more easily select which tracks to retrieve (without parsing the bitstream).

[0263] As shown in the seventh row from the top of the table in Figure 24, the representation dependency information may include dependent information indicating other representations required for decoding the representation corresponding to that information (method 2-2-1).

[0264] The dependency information is information that links a representation that does not include an independent slice (depth=0) to all representations that include slices necessary for decoding that representation.

[0265] By storing such dependency information in the control file (MPD), the decoder can more easily check other representations (i.e., tracks) required for decoding the representation (i.e., track) to be processed based on the dependency information without parsing the bitstream, thereby suppressing an increase in the load of the playback process.

[0266] Furthermore, as shown in the eighth row from the top of the table in Figure 24, this dependent information may also indicate all other representations necessary to decode the representation corresponding to that information (method 2-2-1-1).

[0267] 24, this subordinate information may indicate other representations referenced by the representation corresponding to that information (Method 2-2-1-2). In other words, this subordinate information may indicate the representations that the referencing representation directly references.

[0268] 24, the representation dependency information may include independent information indicating other representations that require a representation corresponding to the information during decoding (Method 2-2-2). In other words, the independent information is information that links a representation including an independent slice (depth=0) to a representation including a dependent slice that is decoded by referring to the slice. In other words, the independent information is reverse lookup information for the dependent information.

[0269] As shown in the 11th row from the top of the table in Figure 24, this independent information may indicate all other representations that require a representation corresponding to that information when decoding (method 2-2-2-1).

[0270] 24, the representation dependency information may be stored in the control file (MPD) as two parameters, that is, representation association identification information (Representation@associationId) and association type (associationType) (Method 2-2-3). In other words, the linking of dependent information and independent information may be performed using two parameters, that is, representation association identification information (Representation@associationId) and association type (associationType).

[0271] In the case of dependent information, the association type may be set to "depd" (associationType="depd"), and in the case of independent information, the association type may be set to "indd" (associationType="indd").

[0272] Note that flag information (@nonScalableAttributeFlag) indicating whether or not an attribute slice to which non-scalable coding is applied is included in the adaptation set may be stored in the control file (MPD).

[0273] <4-3. Description example> An example of MPD description is shown in Fig. 26. In the MPD 520 shown in Fig. 26, in the underlined line 521, an essential property is defined for the adaptation set (AdaptationSet id="1") whose identification information is "1", and schemeIdUri="urn:mpeg:mpegI:gpcc:2020:depth" is set. Then, adaptation set depth information such as @fixedDepth, @minDepth, and @maxDepth is set.

[0274] Also, in the MPD 520 shown in FIG. 26, in the underlined line 522, an essential property is defined for the adaptation set (AdaptationSet id="2") having identification information "2", and schemeIdUri="urn:mpeg:mpegI:gpcc:2020:depth" is set. Then, adaptation set depth information such as @fixedDepth, @minDepth, and @maxDepth is set. Furthermore, in the next line, representation association identification information (Representation@associationId) and representation dependency information such as association type (associationType) are set. Since this representation dependency information is subordinate information, the association type (associationType) is set to "depd".

[0275] Also, in the MPD 520 shown in FIG. 26, in the line underlined 523, an essential property is defined for the adaptation set (AdaptationSet id="3") having identification information "3", and schemeIdUri="urn:mpeg:mpegI:gpcc:2020:depth" is set. Then, adaptation set depth information such as @fixedDepth, @minDepth, and @maxDepth is set. Furthermore, in the next line, representation association identification information (Representation@associationId) and representation dependency information such as association type (associationType) are set. Since this representation dependency information is subordinate information, the association type (associationType) is set to "depd".

[0276] As described above, since the adaptation set depth information and the representation dependency information are stored in the MPD, the decoder can more easily identify the representation (i.e., track) including the desired slice based on the information without parsing the bitstream. Therefore, it is possible to suppress an increase in the load of the playback process.

[0277] Note that either or both of the technology described in this chapter (<4. Transmission of scalable decoding information by control file>), i.e., storing scalable decoding information (adaptation set configuration information) in the MPD, and the technology described in chapter <2. Transmission of scalable decoding information by content file>, i.e., storing scalable decoding information (at least slice configuration information for each sample) in the metadata area of ​​the content file, may be applied.

[0278] 5. Second Embodiment <5-1. File Generation Device> The present technology (each method) described above can be applied to any device. Fig. 27 is a block diagram showing an example of the configuration of a file generation device, which is one aspect of an information processing device to which the present technology is applied. Like file generation device 300, file generation device 600 shown in Fig. 27 is a device that applies G-PCC to encode point cloud data and stores the G-PCC content (G-PCC bitstream) generated by the encoding in a content file (ISOBMFF). However, file generation device 600 also generates an MPD corresponding to the content file.

[0279] In this case, the file generation device 600 may apply the present technology described above in chapters <2. Transmission of scalable decoding information by a content file> and <4. Transmission of scalable decoding information by a control file>. That is, the file generation device 600 may generate scalable decoding information based on the slice depths and inter-slice dependency relationships in the G-PCC content, generate a content file for storing the G-PCC content, and store at least slice configuration information for each sample of the generated scalable decoding information in the metadata area of ​​the generated content file. Furthermore, the file generation device 600 may store adaptation set configuration information of the generated scalable decoding information in the MPD.

[0280] Note that Fig. 27 shows the main processing units, data flows, etc., and is not necessarily all that is shown in Fig. 27. In other words, in file generation device 600, there may be processing units that are not shown as blocks in Fig. 27, and there may be processing and data flows that are not shown as arrows, etc. in Fig. 27.

[0281] 27, file generation device 600 has basically the same configuration as file generation device 300 (FIG. 19). However, file generation device 600 has a content file generation unit 615 and an MPD generation unit 616 instead of file generation unit 315.

[0282] In this case, the scalable decoding information generator 314 also applies the present technology described above in <2. Transmission of scalable decoding information by content file> to generate scalable decoding information (at least slice configuration information for each sample). That is, the scalable decoding information generator 314 generates scalable decoding information (at least slice configuration information for each sample) based on depth information indicating the quality hierarchical level of geometry included in each slice in the G-PCC content including the first slice and the second slice, and the dependency relationship between the first slice and the second slice in the G-PCC content. Also, in this case, the scalable decoding information generator 314 applies the present technology described above in Chapter <4. Transmission of scalable decoding information by control file> to generate adaptation set configuration information as the scalable decoding information instead of track configuration information. In other words, the scalable decoding information generation unit 314, as an adaptation set configuration information generation unit, generates adaptation set configuration information based on depth information indicating the quality hierarchy level of the geometry contained in each slice in the G-PCC content including the first slice and the second slice, and the dependency between the first slice and the second slice in the G-PCC content.

[0283] The scalable decoding information generation unit 314 supplies the generated scalable decoding information together with the G-PCC content to the content file generation unit 615. The scalable decoding information generation unit 314 also supplies the generated scalable decoding information to the MPD generation unit 616 together with the G-PCC content.

[0284] The content file generation unit 615 applies the present technology described above in <2. Transmission of scalable decoding information by content file> to generate a content file that stores the supplied G-PCC content in a track on a slice-by-slice basis, and stores the scalable decoding information (at least slice configuration information for each sample) in the metadata area of ​​the generated content file. The content file generation unit 615 outputs the content file generated as described above to the outside of the file generation device 300.

[0285] The MPD generation unit 616, as a control file generation unit, generates an MPD and stores information about the supplied G-PCC content in the MPD. The MPD generation unit 616 also applies the present technology described above in Chapter <4. Transmission of scalable decoding information using a control file>, and stores the supplied scalable decoding information (adaptation set configuration information) in the MPD. The MPD generation unit 616 outputs the MPD generated as described above to an external device (such as a content file distribution server) outside the file generation device 300.

[0286] This adaptation set configuration information may include adaptation set depth information. That is, the scalable decoding information generation unit 314 may generate adaptation set configuration information including adaptation set depth information, and the MPD generation unit 616 may store the adaptation set configuration information in an adaptation set of the MPD. Furthermore, the adaptation set depth information may include adaptation set minimum depth information. Furthermore, the adaptation set depth information may include adaptation set maximum depth information. Furthermore, the adaptation set depth information may include a match flag. That is, the scalable decoding information generation unit 314 may generate adaptation set depth information including this information, and the MPD generation unit 616 may store the adaptation set depth information in an adaptation set of the MPD. For example, the MPD generation unit 616 may newly define a supplemental property or an essential property and store the adaptation set depth information therein.

[0287] Furthermore, the adaptation set configuration information may include representation dependency information. That is, the scalable decoding information generation unit 314 may generate adaptation set configuration information including representation dependency information, and the MPD generation unit 616 may store the adaptation set configuration information in an adaptation set of the MPD. Furthermore, the representation dependency information may include dependency information indicating other representations required for decoding the representation corresponding to the information. Furthermore, this dependency information may indicate all other representations required for decoding the representation corresponding to the information. Furthermore, this dependency information may indicate other representations referenced by the representation corresponding to the information. That is, the scalable decoding information generation unit 314 may generate representation dependency information including such dependency information, and the MPD generation unit 616 may store the representation dependency information in an adaptation set of the MPD.

[0288] Furthermore, the representation dependency information may include independent information indicating other representations that require a representation corresponding to the information when decoding. This independent information may then indicate all other representations that require a representation corresponding to the information when decoding. That is, the scalable decoding information generation unit 314 may generate representation dependency information that includes such independent information, and the MPD generation unit 616 may store the representation dependency information in an adaptation set of the MPD.

[0289] The representation dependency information may be stored in the control file (MPD) as two parameters: representation association identification information (Representation@associationId) and an association type (associationType).

[0290] By doing so, as described above in chapters <2. Transmission of scalable decoding information by content file> and <4. Transmission of scalable decoding information by control file>, it is possible to suppress an increase in the load of playback processing.

[0291] <File generation process flow> An example of the flow of the file generation process executed by this file generation device 600 will be described with reference to the flowchart of FIG.

[0292] When the file generation process is started, the processes of steps S601 to S603 are executed in the same manner as the processes of steps S301 to S303 in the flowchart of the file generation process of FIG.

[0293] In step S604, the scalable decoding information generation unit 314 applies the present technology described in Section <2. Transmission of scalable decoding information by content file> to generate slice configuration information for each sample as the scalable decoding information. Also, the scalable decoding information generation unit 314 applies the present technology described in Section <4. Transmission of scalable decoding information by control file> to generate adaptation set configuration information as the scalable decoding information.

[0294] In step S605, the content file generation unit 615 applies the present technology described in Chapter <2. Transmission of scalable decoding information using a content file>. That is, the content file generation unit 615 generates a content file and stores the G-PCC content in slice units in a track of the content file. Then, the content file generation unit 615 stores slice configuration information for each sample in the metadata area of ​​the content file.

[0295] In step S606, the content file generation unit 615 outputs the generated content file (a content file storing scalable decoding information) to the outside of the file generation device 600. For example, the content file generation unit 615 transmits the content file to another device (e.g., a playback device) via a network or the like. Also, for example, the content file generation unit 615 supplies the content file to a storage medium external to the file generation device 600 for storage. In this case, the content file is supplied to the playback device or the like via the storage medium.

[0296] In step S607, the MPD generation unit 616 applies the present technology described in Chapter <4. Transmission of scalable decoding information using a control file>. That is, the MPD generation unit 616 generates an MPD corresponding to the content file generated in step S605, and stores the adaptation set configuration information generated in step S604 in the MPD.

[0297] In step S608, the MPD generation unit 616 outputs the MPD to the outside of the file generation device 600. For example, the MPD is provided to a content file distribution server or the like.

[0298] When the process of step S608 ends, the file generation process ends.

[0299] As described above, in the file generation process, the file generation device 600 applies the present technology described in chapters <2. Transmission of scalable decoding information by content file> and <4. Transmission of scalable decoding information by control file>, and stores the scalable decoding information in the metadata area of ​​the content file or in the MPD. In this way, it is possible to reduce the transmission and processing (decoding, etc.) of unnecessary information, and to suppress an increase in the load of data transmission and playback processing.

[0300] <5-2. Playback device> Fig. 29 is a block diagram showing an example of the configuration of a playback device, which is one aspect of an information processing device to which the present technology is applied. Similar to the playback device 400, the playback device 700 shown in Fig. 29 is a device that decodes a content file, constructs a point cloud, and performs rendering to generate presentation information. In this case, the playback device 700 may apply the present technology described above in chapters <2. Transmission of scalable decoded information by content file> and <4. Transmission of scalable decoded information by control file>.

[0301] Note that Fig. 29 shows the main processing units, data flows, etc., and does not necessarily include everything that is shown in Fig. 29. In other words, playback device 700 may have processing units that are not shown as blocks in Fig. 29, or processes or data flows that are not shown as arrows, etc. in Fig. 29.

[0302] 29, playback device 700 basically has the same configuration as playback device 400 (FIG. 21). However, playback device 700 has a file acquisition unit 711 and an MPD analysis unit 712 instead of file acquisition unit 411.

[0303] The file acquisition unit 711 acquires an MPD corresponding to a desired content file (a content file to be played back), and supplies the MPD to the MPD analysis unit 712. The file acquisition unit 711 also requests and acquires, from the supplier of the content file to be played back, the track requested by the MPD analysis unit 712 from among the tracks of the content file. The file acquisition unit 711 supplies the acquired track (the bitstream stored in that track) to the playback processing unit 412 (file processing unit 421).

[0304] When the MPD parser 712 acquires the MPD from the file acquirer 711, it analyzes the MPD and selects a desired track. In this case, the MPD parser 712 may apply the present technology described above in Chapter <4. Transmission of scalable decoding information using a control file>. That is, the MPD parser 712 identifies an adaptation set (i.e., a track) required to obtain an arbitrary slice of the G-PCC content based on the adaptation set configuration information stored in the adaptation set of the MPD. The MPD parser 712 requests the file acquirer 711 to acquire a track corresponding to the identified adaptation set.

[0305] This adaptation set configuration information may include adaptation set depth information. For example, the MPD analysis unit 712 may determine depth information of the geometry of all slices included in a track based on the adaptation set depth information stored in an adaptation set of the MPD, and identify an adaptation set (i.e., a track) required to obtain an arbitrary slice based on the depth information. The adaptation set depth information may also include adaptation set minimum depth information. The adaptation set depth information may also include adaptation set maximum depth information. Furthermore, the adaptation set depth information may include a match flag. For example, the MPD analysis unit 712 may determine depth information of the geometry of all slices included in a track based on the information stored in an adaptation set of the MPD, and identify an adaptation set (i.e., a track) required to obtain an arbitrary slice based on the depth information.

[0306] The adaptation set configuration information may also include representation dependency information. For example, the MPD analysis unit 712 may understand the dependency relationships between representations (i.e., tracks) based on the representation dependency information stored in the adaptation set of the MPD, and may identify a representation (i.e., track) required to obtain a given slice based on the dependency relationships. The representation dependency information may also include dependency information indicating other representations required to decode the representation corresponding to the information. Furthermore, this dependency information may indicate all other representations required to decode the representation corresponding to the information. Furthermore, this dependency information may indicate other representations referenced by the representation corresponding to the information. For example, the MPD analysis unit 712 may identify other representations (i.e., other tracks) required to decode the representation corresponding to the information based on the dependency information stored in the adaptation set of the MPD.

[0307] The representation dependency information may also include independent information indicating other representations that require a representation corresponding to the information for decoding. This independent information may then indicate all other representations that require a representation corresponding to the information for decoding. For example, the MPD parser 712 may identify other representations (i.e., other tracks) that require a representation corresponding to the information for decoding based on this dependency information stored in an adaptation set of the MPD.

[0308] The representation dependency information may be stored in the control file (MPD) as two parameters: representation association identification information (Representation@associationId) and an association type (associationType). That is, the MPD analysis unit 712 may refer to these parameters stored in the adaptation set of the MPD, understand the dependency relationships between representations (i.e., tracks) based on these parameters, and identify the representations (i.e., tracks) required to obtain any slice based on the dependency relationships.

[0309] By doing so, as described above in chapters <2. Transmission of scalable decoding information by content file> and <4. Transmission of scalable decoding information by control file>, it is possible to suppress an increase in the load of playback processing.

[0310] The decoding unit 422 applies the present technology described above in chapter <4. Transmission of scalable decoding information using a control file> to decode slices of the G-PCC content stored in the track supplied from the file acquisition unit 711.

[0311] <Recycling process flow> An example of the flow of playback processing executed by this playback device 700 will be described with reference to the flowchart of FIG.

[0312] When the playback process starts, in step S701, the file acquisition unit 711 of the playback device 700 acquires an MPD corresponding to the content file to be played back.

[0313] In step S702, the MPD analysis unit 712 identifies an adaptation set required to obtain G-PCC content of a desired depth, based on the adaptation set configuration information stored in the MPD.

[0314] In step S703, the file acquisition unit 711 acquires the encoded data stored in the track corresponding to the adaptation set identified in step S702 of the content file to be played back.

[0315] In step S704, the file processing unit 421 obtains from the content file the encoded data of the slices required to obtain G-PCC content of the desired depth from the acquired encoded data based on the slice configuration information for each sample.

[0316] The processes in steps S705 to S708 are executed in the same manner as the processes in steps S403 to S406 in the playback process of Fig. 23. When the process in step S708 ends, the playback process ends.

[0317] As described above, in the playback process, the playback device 700 applies the present technology described in chapters <2. Transmission of scalable decoding information by content file> and <4. Transmission of scalable decoding information by control file>, and acquires and decodes a desired track of a content file based on adaptation set configuration information stored in the MPD and slice configuration information for each sample stored in the metadata area of ​​the content file. In this way, it is possible to reduce unnecessary information processing (data transmission, decoding, etc.), and to suppress an increase in the load of the playback process.

[0318] <6. Notes> <Computer> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the term "computer" includes computers built into dedicated hardware, and general-purpose personal computers, etc., that can execute various functions by installing various programs.

[0319] FIG. 31 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.

[0320] In a computer 900 shown in FIG. 31, a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, and a RAM (Random Access Memory) 903 are interconnected via a bus 904.

[0321] An input / output interface 910 is also connected to the bus 904. To the input / output interface 910, an input unit 911, an output unit 912, a storage unit 913, a communication unit 914, and a drive 915 are connected.

[0322] The input unit 911 includes, for example, a keyboard, a mouse, a microphone, a touch panel, an input terminal, etc. The output unit 912 includes, for example, a display, a speaker, an output terminal, etc. The storage unit 913 includes, for example, a hard disk, a RAM disk, a non-volatile memory, etc. The communication unit 914 includes, for example, a network interface. The drive 915 drives removable media 921 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0323] In a computer configured as above, the CPU 901 performs the above-described series of processes by, for example, loading a program stored in the storage unit 913 into the RAM 903 via the input / output interface 910 and the bus 904 and executing the program. The RAM 903 also stores data necessary for the CPU 901 to execute various processes as appropriate.

[0324] The program executed by the computer can be applied by recording it on removable media 921 such as package media, for example. In this case, the program can be installed in storage unit 913 via input / output interface 910 by inserting removable media 921 into drive 915.

[0325] This program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, digital satellite broadcasting, etc. In this case, the program can be received by the communication unit 914 and installed in the storage unit 913.

[0326] Alternatively, this program can be installed in advance in the ROM 902 or the storage unit 913 .

[0327] <Applicable targets of this technology> The above description has been given using an example in which G-PCC content having a slice structure is stored in ISOBMFF, but the cases to which this technology can be applied are not limited to this example. This technology can be applied to any technology as long as it does not contradict the above-mentioned technology. For example, while point cloud data has been used as an example of the encoding target, 3D data of any standard can also be used as the encoding target. Furthermore, while G-PCC has been used as an example of an encoding / decoding method, any encoding / decoding method can be applied as long as it is capable of generating encoded data having a slice structure (a method that supports scalable decoding). Furthermore, while ISOBMFF has been used as an example of a file format for storing G-PCC content, any file format can be applied as long as it can store scalable decoding information. Furthermore, as long as it does not contradict the above-mentioned technology, some of the above-mentioned processes and specifications may be omitted or combined with technologies not described above.

[0328] Furthermore, the present technology can be applied to any configuration, for example, various electronic devices.

[0329] Furthermore, for example, the present technology can also be implemented as part of an apparatus, such as a processor as a system LSI (Large Scale Integration), a module using multiple processors, a unit using multiple modules, or a set in which other functions are added to a unit.

[0330] Furthermore, for example, the present technology can also be applied to a network system configured with multiple devices. For example, the present technology may be implemented as cloud computing in which multiple devices share and collaborate on processing via a network. For example, the present technology may be implemented in a cloud service that provides image (video)-related services to any terminal, such as a computer, AV (Audio Visual) equipment, a portable information processing terminal, or an IoT (Internet of Things) device.

[0331] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0332] <Fields and applications where this technology can be applied> Systems, devices, processing units, etc. to which the present technology is applied can be used in any field, such as transportation, medical care, crime prevention, agriculture, livestock farming, mining, beauty, factories, home appliances, weather, and nature monitoring. In addition, the applications thereof are also arbitrary.

[0333] For example, the present technology can be applied to systems and devices used to provide viewing content, etc. Furthermore, for example, the present technology can also be applied to systems and devices used for transportation, such as monitoring traffic conditions and controlling automatic driving. Furthermore, for example, the present technology can also be applied to systems and devices used for security. Furthermore, for example, the present technology can also be applied to systems and devices used for automatic control of machines, etc. Furthermore, for example, the present technology can also be applied to systems and devices used for agriculture and livestock farming. Furthermore, for example, the present technology can also be applied to systems and devices used to monitor natural conditions, such as volcanoes, forests, and oceans, and wildlife. Furthermore, for example, the present technology can also be applied to systems and devices used for sports.

[0334] <Other> In this specification, a "flag" refers to information for identifying multiple states, and includes not only information used to identify two states, true (1) or false (0), but also information capable of identifying three or more states. Therefore, the value that this "flag" can take may be, for example, two values, 1 / 0, or three or more values. In other words, the number of bits constituting this "flag" is arbitrary, and may be one bit or multiple bits. Furthermore, identification information (including flags) can be assumed not only to include the identification information in the bit stream, but also to include difference information of the identification information relative to certain reference information in the bit stream. Therefore, in this specification, "flag" and "identification information" include not only the information itself, but also difference information relative to the reference information.

[0335] Furthermore, various types of information (metadata, etc.) related to the encoded data (bitstream) may be transmitted or recorded in any form as long as they are associated with the encoded data. Here, the term "associate" means, for example, making it possible to use (link) one piece of data when processing the other piece of data. That is, data associated with each other may be combined into one piece of data or may be individual pieces of data. For example, second data associated with first data may be transmitted over a transmission path different from that of the first data. Also, for example, second data associated with first data may be recorded on a recording medium different from that of the first data (or on a different recording area of ​​the same recording medium). Note that this "association" may refer not to the entire data, but to a portion of the data. For example, 3D data and metadata corresponding to that 3D data may be associated with each other in any unit, such as multiple samples, one sample, or a portion of a sample.

[0336] In this specification, the terms "composite," "multiplex," "add," "integrate," "include," "store," "incorporate," "insert," and the like refer to combining multiple things into one, and are one method of "associating" as described above.

[0337] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.

[0338] For example, a configuration described as one device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, configurations described above as multiple devices (or processing units) may be combined and configured as one device (or processing unit). Of course, configurations other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, as long as the configuration and operation of the entire system are substantially the same, part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).

[0339] Furthermore, for example, the above-described program may be executed in any device, as long as the device has the necessary functions (functional blocks, etc.) and can obtain the necessary information.

[0340] Also, for example, each step of a single flowchart may be executed by one device, or may be shared and executed by multiple devices. Furthermore, when one step includes multiple processes, the multiple processes may be executed by one device, or may be shared and executed by multiple devices. In other words, multiple processes included in one step can be executed as multiple step processes. Conversely, processes described as multiple steps can be executed collectively as one step.

[0341] Furthermore, the program executed by the computer may have the following features. For example, the processing of the steps of writing the program may be executed in chronological order according to the order described in this specification. The processing of the steps of writing the program may also be executed in parallel. Furthermore, the processing of the steps of writing the program may be executed individually at the necessary timing, such as when called. In other words, as long as no contradiction occurs, the processing of each step may be executed in an order different from the order described above. Furthermore, the processing of the steps of writing the program may be executed in parallel with the processing of another program. Furthermore, the processing of the steps of writing the program may be executed in combination with the processing of another program.

[0342] Furthermore, for example, multiple technologies related to the present technology can be implemented independently and independently, as long as no contradiction occurs. Of course, any multiple technologies can also be implemented in combination. For example, part or all of the present technology described in any embodiment can be implemented in combination with part or all of the present technology described in another embodiment. Furthermore, part or all of any of the above-described present technologies can be implemented in combination with other technologies not described above.

[0343] The present technology can also be configured as follows. (1) a scalable decoding information generation unit that generates scalable decoding information for scalable decoding of a Geometry-based Point Cloud Compression (G-PCC) content including a first slice and a second slice, based on depth information indicating a quality hierarchical level of geometry included in each slice in the G-PCC content and a dependency relationship between the first slice and the second slice in the G-PCC content; a content file generation unit that generates a content file for storing the G-PCC content and stores the scalable decoding information in a metadata area of ​​the content file; An information processing device comprising: (2) The scalable decoding information includes slice configuration information regarding a configuration of the slice for each sample. The information processing device described in (1). (3) The content file generation unit sets a subsample for each slice and stores the slice configuration information in a codec specific parameter of a subsample information box in the metadata area. (2) An information processing device according to the present invention. (4) The slice configuration information includes slice dependency information indicating the dependency relationship between the first slice and the second slice. An information processing device according to (2) or (3). (5) The content file generation unit sets a subsample for each slice, and stores the slice dependency relationship information in a codec specific parameter of a subsample information box in the metadata area. (4) An information processing device according to the present invention. (6) The slice dependency information includes reference-source geometry slice identification information and reference-destination geometry slice identification information, the reference geometry slice identification information is identification information of a geometry slice that serves as a reference in the dependency relationship between the first slice and the second slice; The referenced geometry slice identification information is identification information of a geometry slice that is a reference in the dependency relationship between the first slice and the second slice. An information processing device according to (4) or (5). (7) When the geometry slice corresponding to the reference geometry slice identification information is an independent geometry slice that can be independently decoded, the reference geometry slice identification information and the referenced geometry slice identification information are identical. (6) An information processing device according to (6). (8) The content file generation unit sets a subsample for each slice, and stores the reference source geometry slice identification information and the reference destination geometry slice identification information in Codex Specific Parameters of a subsample information box in the metadata area. An information processing device according to (6) or (7). (9) The slice dependency relationship information includes attribute geometry slice identification information, which is identification information of a geometry slice referenced by an attribute slice. An information processing device according to any one of (4) to (8). (10) The content file generation unit sets a subsample for each slice, and stores the attribute geometry slice identification information in a codec specific parameter of a subsample information box in the metadata area. (9) An information processing device according to (9). (11) The slice dependency information includes non-scalable coding attribute geometry slice identification information, which is identification information of a geometry slice to which an attribute slice to which non-scalable coding is applied refers. An information processing device according to any one of (4) to (8). (12) The non-scalable coding attribute geometry slice identification information includes identification information of the geometry slice that includes the geometry with the maximum depth information among the geometry slices referenced by the attribute slice. (11) An information processing device according to (11). (13) The content file generation unit sets a subsample for each slice, and stores the non-scalable coding attribute geometry slice identification information in a codec specific parameter of a subsample information box in the metadata area. The information processing device according to (11) or (12). (14) The non-scalable coding attribute geometry slice identification information includes a non-scalable coding flag indicating whether non-scalable coding is applied to the attribute slice. An information processing device according to any one of (11) to (13). (15) The content file generation unit sets a subsample for each slice, and stores the non-scalable encoding flag in a codec specific parameter of a subsample information box in the metadata area. (14) An information processing device according to (14). (16) The payload type of the slice dependency information of an independently decodable independent geometry slice is different from the payload type of the slice dependency information of a dependent geometry slice that references another geometry slice during decoding. An information processing device according to any one of (10) to (15). (17) The slice configuration information includes geometry slice depth information related to the depth information of the geometry included in the geometry slice. An information processing device according to any one of (2) to (16). (18) The geometry slice depth information includes minimum depth information indicating a minimum value of the depth information in the geometry slice. (17) An information processing device according to (17). (19) The geometry slice depth information includes maximum depth information indicating a maximum value of the depth information in the geometry slice. The information processing device according to (17) or (18). (20) The content file generation unit sets a subsample for each slice, and stores the geometry slice depth information in a codec specific parameter of a subsample information box in the metadata area. An information processing device according to any one of (17) to (19). (21) The slice configuration information further includes slice dependency information indicating the dependency relationship between the first slice and the second slice, The content file generation unit sets flags of the subsample information box that stores the geometry slice depth information to a value different from flags of the subsample information box that stores the slice dependency relationship information. (20) An information processing device according to (20). (22) The scalable decoding information includes track configuration information regarding a configuration of a track in the content file that stores the G-PCC content in slice units. An information processing device according to any one of (1) to (21). (23) The track configuration information includes track depth information regarding the depth information of the geometry of all the slices included in the track corresponding to the track configuration information. (22) An information processing device according to (22). (24) The track depth information includes track minimum depth information indicating the minimum value of the depth information in the track. (23) An information processing device according to (23). (25) The track depth information includes track maximum depth information indicating the maximum value of the depth information in the track. The information processing device according to (23) or (24). (26) The track depth information includes a match flag indicating whether the minimum value of the depth information in each sample included in the track matches the minimum value of the depth information in the track, and whether the maximum value of the depth information in each sample included in the track matches the maximum value of the depth information in the track. An information processing device according to any one of (23) to (25). (27) The content file generation unit stores the track depth information in a depth information box of a sample entry in the metadata area. An information processing device according to any one of (23) to (26). (28) The track configuration information includes track dependency information indicating a dependency relationship between a first track and a second track. An information processing device according to any one of (22) to (27). (29) The track dependency information includes dependency information indicating other tracks including slices necessary for decoding the dependent slices included in the track. (28) An information processing device according to (28). (30) The dependency information indicates all the other tracks including the slices necessary for decoding the dependency slice. (29) An information processing device according to (29). (31) The dependency information indicates the other track including the slice referenced by the dependency slice. (29) An information processing device according to (29). (32) The track dependency information includes independent information indicating other tracks including other slices that require an independent slice included in the track during decoding. An information processing device according to any one of (28) to (31). (33) The independent information indicates the other track including the other slice that references the independent slice during decoding. (32) An information processing device according to (32). (34) The content file generation unit stores the track dependency information in the metadata area as a track reference. An information processing device according to any one of (28) to (33). (35) generating scalable decoding information for scalable decoding of a Geometry-based Point Cloud Compression (G-PCC) content including a first slice and a second slice, based on depth information indicating a quality hierarchy level of geometry included in each slice in the G-PCC content and a dependency relationship between the first slice and the second slice in the G-PCC content; A content file for storing the G-PCC content is generated, and the scalable decoding information is stored in a metadata area of ​​the content file. Information processing methods.

[0344] (51) An extracting unit that extracts an arbitrary slice of a Geometry-based Point Cloud Compression (G-PCC) content from a content file storing the G-PCC content including a first slice and a second slice, based on scalable decoding information stored in a metadata area of ​​the content file; a decoding unit that decodes the slice of the G-PCC content extracted by the extraction unit; Equipped with The scalable decoding information is information related to scalable decoding of the G-PCC content, and is information generated based on depth information indicating a quality hierarchy level of geometry included in the slice in the G-PCC content and a dependency relationship between the first slice and the second slice in the G-PCC content. Information processing device. (52) The scalable decoding information includes slice configuration information regarding a configuration of the slice for each sample. (51) An information processing device according to (51). (53) The extracting unit extracts any of the slices of the G-PCC content from the content file based on the slice configuration information stored in Codex Specific Parameters of a subsample information box in the metadata area of ​​the subsample set for each slice. (52) An information processing device according to (52). (54) The slice configuration information includes slice dependency information indicating the dependency relationship between the first slice and the second slice. The information processing device according to (52) or (53). (55) The extracting unit extracts any of the slices of the G-PCC content from the content file based on the slice dependency information stored in Codex Specific Parameters of a subsample information box in the metadata area of ​​the subsample set for each slice. (54) An information processing device according to (54). (56) The slice dependency information includes reference-source geometry slice identification information and reference-destination geometry slice identification information, the reference geometry slice identification information is identification information of a geometry slice that serves as a reference in the dependency relationship between the first slice and the second slice; The referenced geometry slice identification information is identification information of a geometry slice that is a reference in the dependency relationship between the first slice and the second slice. The information processing device according to (54) or (55). (57) When the geometry slice corresponding to the reference geometry slice identification information is an independent geometry slice that can be independently decoded, the reference geometry slice identification information and the referenced geometry slice identification information are identical. (56) An information processing device according to (56). (58) The extracting unit extracts any of the slices of the G-PCC content from the content file based on the reference source geometry slice identification information and the reference destination geometry slice identification information stored in the codec specific parameters of the subsample information box in the metadata area of ​​the subsample set for each slice. The information processing device according to (56) or (57). (59) The slice dependency relationship information includes attribute geometry slice identification information, which is identification information of a geometry slice to which an attribute slice refers. An information processing device according to any one of (54) to (58). (60) The extracting unit extracts any of the slices of the G-PCC content from the content file based on the attribute geometry slice identification information stored in the codec specific parameters of the subsample information box in the metadata area of ​​the subsample set for each slice. (59) An information processing device according to (59). (61) The slice dependency information includes non-scalable coding attribute geometry slice identification information, which is identification information of a geometry slice to which an attribute slice to which non-scalable coding is applied refers. An information processing device according to any one of (54) to (58). (62) The non-scalable coding attribute geometry slice identification information includes identification information of the geometry slice that includes the geometry with the maximum depth information among the geometry slices referenced by the attribute slice. (61) An information processing device according to (61). (63) The extracting unit extracts any one of the slices of the G-PCC content from the content file based on the non-scalable coding attribute geometry slice identification information stored in Codex Specific Parameters of a subsample information box in the metadata area of ​​the subsample set for each slice. The information processing device according to (61) or (62). (64) The non-scalable coding attribute geometry slice identification information includes a non-scalable coding flag indicating whether non-scalable coding is applied to the attribute slice. An information processing device according to any one of (61) to (63). (65) The extracting unit extracts any of the slices of the G-PCC content from the content file based on the non-scalable encoding flag stored in the codec specific parameters of the subsample information box in the metadata area of ​​the subsample set for each slice. (64) An information processing device according to (64). (66) The payload type of the slice dependency information of an independently decodable independent geometry slice is different from the payload type of the slice dependency information of a dependent geometry slice that references another geometry slice during decoding. An information processing device according to any one of (60) to (65). (67) The slice configuration information includes geometry slice depth information related to the depth information of the geometry included in the geometry slice. An information processing device according to any one of (52) to (66). (68) The geometry slice depth information includes minimum depth information indicating a minimum value of the depth information in the geometry slice. (67) An information processing device according to (67). (69) The geometry slice depth information includes maximum depth information indicating a maximum value of the depth information in the geometry slice. The information processing device according to (67) or (68). (70) The extraction unit extracts any of the slices of the G-PCC content from the content file based on the geometry slice depth information stored in the codec specific parameters of the subsample information box in the metadata area of ​​the subsample set for each slice. An information processing device according to any one of (67) to (69). (71) The slice configuration information further includes slice dependency information indicating the dependency relationship between the first slice and the second slice, The flags of the subsample information box storing the geometry slice depth information and the flags of the subsample information box storing the slice dependency relationship information are set to different values. (70) An information processing device according to (70). (72) The scalable decoding information includes track configuration information regarding a configuration of a track in the content file that stores the G-PCC content in slice units. An information processing device according to any one of (51) to (71). (73) The track configuration information includes track depth information regarding the depth information of the geometry of all the slices included in the track corresponding to the track configuration information. (72) An information processing device according to (72). (74) The track depth information includes track minimum depth information indicating the minimum value of the depth information in the track. (73) An information processing device according to (73). (75) The track depth information includes track maximum depth information indicating the maximum value of the depth information in the track. The information processing device according to (73) or (74). (76) The track depth information includes a match flag indicating whether the minimum value of the depth information in each sample included in the track matches the minimum value of the depth information in the track, and whether the maximum value of the depth information in each sample included in the track matches the maximum value of the depth information in the track. An information processing device according to any one of (73) to (75). (77) The extracting unit extracts any of the slices of the G-PCC content from the content file based on the track depth information stored in a depth information box of a sample entry in the metadata area. An information processing device according to any one of (73) to (76). (78) The track configuration information includes track dependency information indicating a dependency relationship between a first track and a second track. An information processing device according to any one of (72) to (77). (79) The track dependency information includes dependency information indicating other tracks including slices necessary for decoding the dependent slices included in the track. (78) An information processing device according to (78). (80) The dependency information indicates all the other tracks including the slices necessary for decoding the dependency slice. (79) An information processing device according to (79). (81) The dependency information indicates the other track including the slice referenced by the dependency slice. (79) An information processing device according to (79). (82) The track dependency information includes independence information indicating other tracks including other slices for which an independent slice included in the track is required for decoding. An information processing device according to any one of (78) to (81). (83) The independent information indicates the other track including the other slice that references the independent slice during decoding. (82) An information processing device according to (82). (84) The extracting unit extracts any of the slices of the G-PCC content from the content file based on the track dependency information stored as a track reference in the metadata area. An information processing device according to any one of (78) to (83). (85) Extracting an arbitrary slice of a Geometry-based Point Cloud Compression (G-PCC) content from a content file storing the G-PCC content including a first slice and a second slice, based on scalable decoding information stored in a metadata area of ​​the content file; Decrypting the slice of the extracted G-PCC content; The scalable decoding information is information related to scalable decoding of the G-PCC content, and is information generated based on depth information indicating a quality hierarchy level of geometry included in the slice in the G-PCC content and a dependency relationship between the first slice and the second slice in the G-PCC content. Information processing methods.

[0345] (101) An adaptation set configuration information generation unit that generates adaptation set configuration information based on depth information indicating a quality hierarchical level of geometry included in each slice in a G-PCC (Geometry-based Point Cloud Compression) content including a first slice and a second slice, and a dependency relationship between the first slice and the second slice in the G-PCC content; a control file generating unit that generates a control file for controlling playback of a content file that stores the G-PCC content, and stores the adaptation set configuration information in the control file; Equipped with the content file stores the G-PCC content in slice units in tracks; The adaptation set configuration information is information about the configuration of an adaptation set that describes information about the track of the content file. Information processing device. (102) The adaptation set configuration information includes adaptation set depth information regarding the depth information of the geometry of all the slices included in the track corresponding to the adaptation set. The information processing device according to (101). (103) The adaptation set depth information includes adaptation set minimum depth information indicating a minimum value of the depth information in the track. (102) An information processing device according to (102). (104) The adaptation set depth information includes adaptation set maximum depth information indicating a maximum value of the depth information in the track. The information processing device according to (102) or (103). (105) The adaptation set depth information includes a match flag indicating whether sample minimum depth information of each sample included in the track matches the adaptation set minimum depth information and whether sample maximum depth information of each sample included in the track matches the adaptation set maximum depth information; the sample minimum depth information indicates a minimum value of the depth information in the sample; the sample maximum depth information indicates a maximum value of the depth information in the sample; the adaptation set minimum depth information indicates a minimum value of the depth information in the track; The adaptation set maximum depth information indicates the maximum value of the depth information in the track. An information processing device according to any one of (102) to (104). (106) The adaptation set configuration information includes representation dependency information indicating a dependency relationship between a first representation and a second representation. The information processing device according to (101). (107) The representation dependency information includes dependency information indicating other representations required for decoding the representation corresponding to the representation dependency information. (106) An information processing device according to (106). (108) The dependency information indicates all the other representations necessary for decoding the representation corresponding to the dependency information. (107) An information processing device according to (107). (109) The dependent information indicates the other representation referenced from the representation corresponding to the dependent information. (107) An information processing device according to (107). (110) The representation dependency information includes independent information indicating other representations that require a representation corresponding to the representation dependency information when decoding. An information processing device according to any one of (106) to (109). (111) The independent information indicates all the other representations that require the representation corresponding to the independent information when decoding. (110) An information processing device according to (110). (112) The control file generation unit stores the representation dependency information in the control file as a representation association ID and an association type. An information processing device according to any one of (106) to (111). (113) Generate adaptation set configuration information based on depth information indicating a quality hierarchy level of geometry included in each slice in a G-PCC (Geometry-based Point Cloud Compression) content including a first slice and a second slice, and a dependency relationship between the first slice and the second slice in the G-PCC content; generating a control file for controlling playback of a content file storing the G-PCC content, and storing the adaptation set configuration information in the control file; the content file stores the G-PCC content in slice units in tracks; The adaptation set configuration information is information about the configuration of an adaptation set that describes information about the track of the content file. Information processing methods.

[0346] (151) An analysis unit that analyzes a control file that controls playback of a content file that stores G-PCC (Geometry-based Point Cloud Compression) content including a first slice and a second slice in a track on a slice-by-slice basis, and identifies an adaptation set required to obtain an arbitrary slice of the G-PCC content based on adaptation set configuration information stored in the control file; an acquisition unit that acquires the track of the content file corresponding to the adaptation set identified by the analysis unit; a decoding unit that decodes the slice of the G-PCC content stored in the track acquired by the acquisition unit; Equipped with The adaptation set configuration information is information about the configuration of the adaptation set that describes information about the track of the content file, and is information generated based on depth information that indicates a quality hierarchy level of geometry included in the slice in the G-PCC content and a dependency relationship between the first slice and the second slice in the G-PCC content. Information processing device. (152) The adaptation set configuration information includes adaptation set depth information regarding the depth information of the geometry of all the slices included in the track corresponding to the adaptation set. (151) An information processing device according to the present invention. (153) The adaptation set depth information includes adaptation set minimum depth information indicating a minimum value of the depth information in the track. (152) An information processing device according to the present invention. (154) The adaptation set depth information includes adaptation set maximum depth information indicating a maximum value of the depth information in the track. The information processing device according to (152) or (153). (155) The adaptation set depth information includes a match flag indicating whether sample minimum depth information of each sample included in the track matches the adaptation set minimum depth information and whether sample maximum depth information of each sample included in the track matches the adaptation set maximum depth information; the sample minimum depth information indicates a minimum value of the depth information in the sample; the sample maximum depth information indicates a maximum value of the depth information in the sample; the adaptation set minimum depth information indicates a minimum value of the depth information in the track; The adaptation set maximum depth information indicates the maximum value of the depth information in the track. An information processing device according to any one of (152) to (154). (156) The adaptation set configuration information includes representation dependency information indicating a dependency relationship between a first representation and a second representation. An information processing device according to any one of (151) to (155). (157) The representation dependency information includes dependency information indicating other representations required for decoding the representation corresponding to the representation dependency information. (156) An information processing device according to the present invention. (158) The dependency information indicates all the other representations necessary for decoding the representation corresponding to the dependency information. (157) An information processing device according to the present invention. (159) The dependent information indicates the other representation referenced from the representation corresponding to the dependent information. (157) An information processing device according to the present invention. (160) The representation dependency information includes independent information indicating other representations that require a representation corresponding to the representation dependency information when being decoded. An information processing device according to any one of (156) to (159). (161) The independent information indicates all the other representations that require the representation corresponding to the independent information during decoding. (160) An information processing device according to the present invention. (162) The analysis unit identifies the adaptation set based on the representation dependency information stored in the control file as a representation association ID and an association type. An information processing device according to any one of (156) to (161). (163) Analyzing a control file that controls playback of a content file storing G-PCC (Geometry-based Point Cloud Compression) content including a first slice and a second slice in a track on a slice-by-slice basis, and identifying an adaptation set required to obtain an arbitrary slice of the G-PCC content based on adaptation set configuration information stored in the control file; obtaining the track of the content file corresponding to the identified adaptation set; Decrypting the slice of the G-PCC content stored in the acquired track; The adaptation set configuration information is information about the configuration of the adaptation set that describes information about the track of the content file, and is information generated based on depth information that indicates a quality hierarchy level of geometry included in the slice in the G-PCC content and a dependency relationship between the first slice and the second slice in the G-PCC content. Information processing methods. [Explanation of symbols]

[0347] 300 file generation device, 311 extraction unit, 312 encoding unit, 313 bitstream generation unit, 314 scalable decoding information generation unit, 315 file generation unit, 321 geometry encoding unit, 322 attribute encoding unit, 323 metadata generation unit, 400 playback device, 401 control unit, 411 file acquisition unit, 412 playback processing unit, 413 presentation processing unit, 421 file processing unit, 422 decoding unit, 423 presentation information generation unit, 431 bitstream extraction unit, 441 geometry decoding unit, 442 attribute decoding unit, 451 point cloud construction unit, 452 presentation processing unit, 600 file generation device, 615 content file generation unit, 616 MPD generation unit, 700 playback device, 711 file acquisition unit, 712 MPD analysis unit< / hevc>

Claims

1. a scalable decoding information generation unit that generates scalable decoding information regarding scalable decoding of Geometry-based Point Cloud Compression (G-PCC) content based on depth information indicating a quality hierarchical level of each slice in the G-PCC content, including a first slice and a second slice, and a dependency relationship between the first slice and the second slice in the G-PCC content; a content file generation unit that generates a content file for storing the G-PCC content and stores the scalable decoding information in a metadata area of ​​the content file; An information processing device comprising:

2. The scalable decoding information includes slice configuration information regarding the configuration of the slice for each sample. The information processing device according to claim 1 .

3. The content file generation unit sets a subsample for each slice and stores the slice configuration information in Codex Specific Parameters of a Subsample Information Box in the metadata area. The information processing device according to claim 2 .

4. The slice configuration information includes slice dependency information indicating the dependency between the first slice and the second slice. The information processing device according to claim 2 .

5. the slice dependency information includes reference geometry slice identification information and referenced geometry slice identification information; the reference geometry slice identification information is identification information of a geometry slice that serves as a reference in the dependency relationship between the first slice and the second slice, The referenced geometry slice identification information is identification information of a geometry slice that is a reference in the dependency relationship between the first slice and the second slice. The information processing device according to claim 4 .

6. The slice configuration information includes geometry slice depth information related to the depth information of the geometry included in the geometry slice. The information processing device according to claim 2 .

7. The geometry slice depth information includes minimum depth information indicating a minimum value of the depth information in the geometry slice, and maximum depth information indicating a maximum value of the depth information in the geometry slice. The information processing device according to claim 6 .

8. The scalable decoding information includes track configuration information regarding the configuration of tracks in the content file that store the G-PCC content in slice units. The information processing device according to claim 1 .

9. The track configuration information includes track depth information regarding the depth information of the geometry of all the slices included in the track corresponding to the track configuration information. The information processing device according to claim 8 .

10. The track depth information includes track minimum depth information indicating the minimum value of the depth information in the track, and track maximum depth information indicating the maximum value of the depth information in the track. The information processing device according to claim 9 .

11. The content file generation unit stores the track depth information in a depth information box of a sample entry in the metadata area. The information processing device according to claim 9 .

12. The track configuration information includes track dependency information indicating a dependency relationship between a first track and a second track. The information processing device according to claim 8 .

13. The track dependency information includes dependency information indicating other tracks including slices necessary for decoding the dependent slices included in the track. The information processing device according to claim 12.

14. The track dependency information includes independence information indicating other tracks including other slices that require an independent slice included in the track during decoding. The information processing device according to claim 12.

15. The content file generation unit stores the track dependency information in the metadata area as a track reference. The information processing device according to claim 12.

16. generating scalable decoding information for scalable decoding of Geometry-based Point Cloud Compression (G-PCC) content including a first slice and a second slice, based on depth information indicating a quality hierarchical level of each slice in the G-PCC content and a dependency relationship between the first slice and the second slice in the G-PCC content; A content file for storing the G-PCC content is generated, and the scalable decoding information is stored in a metadata area of ​​the content file. Information processing methods.

17. an extracting unit that extracts an arbitrary slice of Geometry-based Point Cloud Compression (G-PCC) content from a content file storing G-PCC content including a first slice and a second slice, based on scalable decoding information stored in a metadata area of ​​the content file; a decoding unit that decodes the slice of the G-PCC content extracted by the extraction unit; Equipped with The scalable decoding information is information related to scalable decoding of the G-PCC content, and is information generated based on depth information indicating a quality hierarchical level of the slice in the G-PCC content and a dependency relationship between the first slice and the second slice in the G-PCC content. Information processing device.

18. The scalable decoding information includes slice configuration information regarding the configuration of the slice for each sample. The information processing device according to claim 17.

19. The scalable decoding information includes track configuration information regarding the configuration of tracks in the content file that store the G-PCC content in slice units. The information processing device according to claim 17.

20. extracting an arbitrary slice of Geometry-based Point Cloud Compression (G-PCC) content from a content file storing the G-PCC content including a first slice and a second slice, based on scalable decoding information stored in a metadata area of ​​the content file; Decrypting the slice of the extracted G-PCC content; The scalable decoding information is information related to scalable decoding of the G-PCC content, and is information generated based on depth information indicating a quality hierarchical level of the slice in the G-PCC content and a dependency relationship between the first slice and the second slice in the G-PCC content. Information processing methods.

Citation Information

Patent Citations

  • IEC14496-12,2015-02-20

  • Point cloud data transmission apparatus, point cloud data transmission method, point cloud data reception apparatus and point cloud data reception method

    US20210029187A1

  • An apparatus, a method and a computer program for video coding and decoding

    WO2020008106A1

  • Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device

    WO2020241723A1

  • Information processing device, information processing method, playback processing device, and playback processing method

    WO2021049333A1