Information processing device and method
By storing scalable decoding information in the metadata area of the G-PCC content file, the problem of increased playback processing load caused by the mismatch of the G-PCC content slice structure is solved, and more efficient decoding and playback processing is achieved.
Patent Information
- Application Number
- CN202180053432.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-08
- Filing Date
- 2021-09-06
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2041-09-06
AI Technical Summary
In the prior art, the slice structure of G-PCC content does not correspond, resulting in the need to transmit and parse the entire content during decoding, which increases the playback processing load.
Scalable decoding is supported by generating and storing scalable decoding information for Geometry-Based Point Cloud Compression (G-PCC) content, including depth information at each quality level of each slice and dependencies between slices, in the metadata area of the content file.
This reduces unnecessary processing of the decoder, suppresses the increase in the reproduction processing load, and improves decoding efficiency. In particular, when processing large-scale point cloud data, it is easier to extract and decode the necessary slices.
Smart Images

Figure CN116157838B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing apparatus and method, and more particularly to an information processing apparatus and method capable of suppressing an increase in a reproduction processing load. Background Art
[0002] In related art, various methods such as High Efficiency Video Coding (HEVC) have been proposed as moving picture coding technologies. As a technology for transmitting moving pictures encoded in this way, there is the International Organization for Standardization Base Media File Format (ISOBMFF), which is a file container specification of the international standard technology for moving picture compression "Moving Picture Experts Group-4 (MPEG-4)" (for example, see Non-Patent Document 1).
[0003] Furthermore, for example, in encoding technologies such as HEVC, encoded data is layered according to, for example, resolution, etc., to enable scalable decoding. Furthermore, a file format has been proposed in which encoded data is divided into tracks and stored on a layer basis, and only the encoded data of a desired layer can be selectively transmitted (for example, see Non-Patent Document 2).
[0004] Incidentally, as a method for encoding a point cloud representing a three-dimensional object as a set of points, a coding technique called Geometry-based Point Cloud Compression (G-PCC) is being standardized in MPEG-I Part 9 (ISO / IEC 23090-9), in which point cloud data is divided into a geometry indicating point position information and attributes indicating point attribute information (for example, see Non-Patent Document 3). In this G-PCC, the formation of an independently decodable slice structure has been proposed (see Non-Patent Document 4).
[0005] Furthermore, as a technique for transmitting G-PCC content obtained by encoding a point cloud applying G-PCC, storing the G-PCC content in the above-described ISOBMFF has been proposed (for example, see Non-Patent Document 5).
[0006] [Citation List]
[0007] [Non-patent literature]
[0008] Non-patent document 1: “Information technology-Coding of audio-visual objects-Part 12: ISO base media file format,” ISO / IEC 14496-12, February 20, 2015.
[0009] Non-patent document 2: "Information technology-Coding of audio-visual objects-Part15: Carriage of network abstraction layer (NAL) unit structured video in the ISO base media file format", ISO / IEC FDIS 14496-15:2014(E), 2014-01-13.
[0010] Non-patent document 3: “G-PCC Future Enhancements,” ISO / IEC JTC 1 / SC 29 / WG11N19328, June 26, 2020.
[0011] Non-Patent Literature 4: David Flynn, Khaled Mammou, “G-PCC: A hierarchical geometry slice structure,” ISO / IEC JCTC1 / SC29 / WG11 MPEG / m54677, April 2020, online.
[0012] Non-Patent Literature 5: Sejin Oh, Ryohei Takahashi, Youngkwon Lim, “Text of ISO / IEC CD 23090-18 Carriage of Geometry-based Point Cloud Compression Data,” ISO / IEC JTC 1 / SC 29 / WG 11N19442, July 30, 2020. Summary of the Invention
[0013] [Problems to be Solved by the Invention]
[0014] However, the method described in Non-Patent Document 5 does not correspond to the slice structure described in Non-Patent Document 4. Therefore, in order to decode some of the slices of the G-PCC content, it is necessary to transmit and parse (analyze) the entire G-PCC content, so there is a possibility of increasing the reproduction processing load.
[0015] The present invention has been made in view of such circumstances, and an object of the present invention is to suppress an increase in the reproduction processing load.
[0016] [Solution to the problem]
[0017] An information processing device according to one aspect of the present technology is an information processing device including: a scalable decoding information generation unit configured to generate scalable decoding information about scalable decoding of G-PCC content based on depth information indicating a quality hierarchy level of each slice in geometry-based point cloud compression (G-PCC) content including first and second slices and a dependency relationship between the first slice and the second slice in the G-PCC content; and a content file generation unit configured to generate a content file storing the G-PCC content and store the scalable decoding information in a metadata area of the content file.
[0018] An information processing method according to another aspect of the present technology is an information processing method comprising: generating scalable decoding information about scalable decoding of geometry-based point cloud compression (G-PCC) content based on depth information indicating a quality hierarchy level of each slice in G-PCC content including first and second slices and a dependency relationship between the first slice and the second slice in the G-PCC content; and generating a content file storing the G-PCC content and storing the scalable decoding information in a metadata area of the content file.
[0019] An information processing device according to another aspect of the present technology is an information processing device including: an extraction unit configured to extract an arbitrary slice of geometry-based point cloud compression (G-PCC) content from a content file based on scalable decoding information stored in a metadata area of a content file storing G-PCC content including first and second slices; and a decoding unit configured to decode the slice of the G-PCC content extracted by the extraction unit. The scalable decoding information is information about scalable decoding of the G-PCC content, and is information generated based on depth information indicating a quality hierarchy level of a slice in the G-PCC content and a dependency relationship between the first slice and the second slice in the G-PCC content.
[0020] An information processing method according to another aspect of the present technology is an information processing method including: extracting an arbitrary slice of geometry-based point cloud compression (G-PCC) content from a content file based on scalable decoding information stored in a metadata area of a content file storing G-PCC content including first and second slices; and decoding the extracted slice of the G-PCC content. The scalable decoding information is information about scalable decoding of the G-PCC content, and is information generated based on depth information indicating a quality hierarchy level of a slice in the G-PCC content and a dependency relationship between a first slice and a second slice in the G-PCC content.
[0021] In an information processing device and method according to one aspect of the present technology, scalable decoding information about scalable decoding of geometry-based point cloud compression (G-PCC) content including first and second slices is generated based on depth information indicating the quality hierarchy level of each slice in the G-PCC content and the dependency relationship between the first slice and the second slice in the G-PCC content; a content file storing the G-PCC content is generated, and the scalable decoding information is stored in a metadata area of the content file.
[0022] In an information processing device and method according to another aspect of the present technology, based on scalable decoding information stored in a metadata area of a content file storing geometry-based point cloud compression (G-PCC) content including first and second slices, an arbitrary slice of the G-PCC content is extracted from the content file; and the extracted slice of the G-PCC content is decoded. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a diagram showing the L-HEVC file format.
[0024] Figure 2 is a diagram showing the slice structure of a geometric body.
[0025] Figure 3 is a diagram showing the structure of a bit stream.
[0026] Figure 4 It is a diagram showing the structure of a content file.
[0027] Figure 5 is a diagram illustrating a method of transmitting G-PCC content.
[0028] Figure 6 is a diagram illustrating a method of transmitting G-PCC content.
[0029] Figure 7 It is a diagram showing slice composition information.
[0030] Figure 8 is a diagram showing an example of a subsample information box.
[0031] Figure 9 It is a diagram showing slice dependency information.
[0032] Figure 10 It is a diagram showing attribute geometry slice identification information.
[0033] Figure 11 is a diagram showing attribute slices to which non-scalable coding is applied.
[0034] Figure 12 is a diagram showing non-scalable coded attribute geometry slice identification information.
[0035] Figure 13 It is a diagram showing slice dependency information.
[0036] Figure 14 It is a diagram showing the depth information of the geometry slice.
[0037] Figure 15 It is a diagram showing track configuration information.
[0038] Figure 16 is a diagram showing track depth information.
[0039] Figure 17 It is a diagram showing track dependency information.
[0040] Figure 18 is a diagram showing a configuration example of a Matroska media container.
[0041] Figure 19 is a block diagram showing a main configuration example of a file generating device.
[0042] Figure 20 is a flowchart illustrating an example of the flow of file generation processing.
[0043] Figure 21 is a block diagram showing a main configuration example of a decoding device.
[0044] Figure 22 is a block diagram showing a main configuration example of a reproduction processing unit.
[0045] Figure 23 is a flowchart showing an example of the flow of the reproduction process.
[0046] Figure 24 This is a diagram showing a control file of the G-PCC content.
[0047] Figure 25 It is a diagram showing adaptation set composition information.
[0048] Figure 26 is a diagram showing a description example of MPD.
[0049] Figure 27 is a block diagram showing a main configuration example of a file generating device.
[0050] Figure 28 is a flowchart illustrating an example of the flow of file generation processing.
[0051] Figure 29 is a block diagram showing a main configuration example of a decoding device.
[0052] Figure 30is a flowchart showing an example of the flow of the reproduction process.
[0053] Figure 31 is a block diagram showing a main configuration example of a computer. DETAILED DESCRIPTION
[0054] Hereinafter, modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described. Note that the description will be made in the following order.
[0055] 1. Transmission of G-PCC content with a slice structure
[0056] 2. Transmission of scalable decoding information via content files
[0057] 3. First Embodiment (File Generation Device and Reproduction Device)
[0058] 4. Transmission of scalable decoding information by controlling files
[0059] 5. Second Embodiment (File Generation Device and Reproduction Device)
[0060] 6. Supplement
[0061] <1. Transmission of G-PCC Content with Slice Structure>
[0062] <Documents supporting technical content and technical terminology>
[0063] The scope disclosed in the present technology includes not only the contents described in the embodiments but also the contents described in the following non-patent documents and the like known at the time of filing, the contents of other documents cited in the following non-patent documents, and the like.
[0064] Non-Patent Document 1: (described above)
[0065] Non-Patent Document 2: (described above)
[0066] Non-Patent Document 3: (described above)
[0067] Non-Patent Document 4: (described above)
[0068] Non-Patent Document 5: (described above)
[0069] Non-Patent Literature 6: https: / / www.matroska.org / index.html
[0070] That is, the contents described in the above-mentioned non-patent documents, the contents of other documents cited in the above-mentioned non-patent documents, etc. are also the basis for determining the support requirements.
[0071] <hevc>
[0072] In the prior art, various methods such as High Efficiency Video Coding (HEVC) have been proposed as moving image coding techniques. As a technique for transmitting the coded moving images in this way, for example, there is the International Organization for Standardization Base Media File Format (ISOBMFF) described in Non-Patent Document 1, which is a file container specification of the international standard technology "Moving Picture Experts Group - 4 (MPEG-4)" for moving image compression.
[0073] In addition, for example, in coding techniques such as HEVC, the coded data is layered according to, for example, the resolution, etc., so as to be able to correspond to scalable decoding. Then, for example, as described in Non-Patent Document 2, a file format (L-HEVC file format) has been proposed in which the coded data is divided into tracks and stored on a layer basis, and only the coded data of the desired layer can be selectively transmitted.
[0074] In the L-HEVC file format, the bitstream of the moving image has a hierarchical structure, and single or multiple layers can be stored in each track of the ISOBMFF. As Figure 1 shown, in the ISOBMFF, the information about the layers included in each sample is stored in the sample group (layer information sample group). In L-HEVC where the coding is performed such that the bitstream has a hierarchical structure, since intra prediction is applied, the hierarchical structure does not change frequently. Therefore, it is preferable to store the information about the layers in the sample group. This sample group can also be used as information for track selection.
[0075] <Point cloud>
[0076] Incidentally, as 3D data representing a three-dimensional object (also referred to as a 3D object), there is a point cloud that represents the 3D object as a set of points.
[0077] For example, in the case of a point cloud, the 3D object as a three-dimensional structure is expressed as a set of a large number of points. The point cloud includes the position information (also referred to as geometry) and the attribute information (also referred to as attributes) of each point. The attributes can include any information. For example, the color information, reflectance information, normal information, etc. of each point can be included in the attributes. As described above, the point cloud has a relatively simple data structure and can express the three-dimensional shape of the 3D object with sufficient accuracy by using a sufficiently large number of points.
[0078] <Overview of G-PCC>
[0079] Non-Patent Document 3 discloses a coding technique called Geometry-based Point Cloud Compression (G-PCC), which is used to separately encode such point clouds into geometry and attributes. G-PCC is being standardized in MPEG-I Part 9 (ISO / IEC 23090-9).
[0080] like Figure 2 The tree structure coding shown in is applied to geometry coding. In tree structure coding, the geometry is first converted into a tree structure. The three-dimensional space is recursively divided, and the geometry is quantized in the divided area of each level, and thus forms a tree structure as shown in FIG. Figure 2 The tree structure of the geometry shown in . Figure 2 In this case, a tree structure is formed which is hierarchical into depth 0 to depth 6 (also referred to as LOD) shown in the vertical direction in the figure.
[0081] As the number of divisions increases, the number of divided areas increases, and each divided area becomes smaller. Therefore, the 3D object is represented by a larger number of points with higher precision geometry. That is, as the depth becomes deeper (as the depth becomes lower), the resolution of the point cloud becomes higher. The geometry of each layer is then encoded according to the tree structure. For example, the difference between the geometry of each layer and the geometry of the layer one level higher than the layer to which the geometry belongs is calculated. The difference is encoded. Each layer can be encoded separately. However, by encoding the difference in this way, the coding efficiency can be improved.
[0082] The difference with the upper layer is encoded. Therefore, in order to obtain the geometry at the desired depth, the depth and encoded data of the upper layer except for that depth can be decoded. For example, in Figure 2 In the case of , the geometry with a resolution of depth 2 is obtained by decoding the coded data of depth 0 to depth 2. In other words, decoding of depths 3 to 6 is unnecessary. A decoding method that can restore (generate) information about a desired layer by decoding only a portion of the coded data in this way is called scalable decoding. That is, by forming a tree structure and performing encoding as described above, the geometry can undergo scalable decoding with resolution.
[0083] For example, if the encoded data does not support scalable decoding, it is necessary to decode all the encoded data to restore the geometry at the lowest resolution (i.e., the highest resolution) and then generate the desired resolution at a resolution lower than that. In contrast, if the encoded data supports scalable decoding, the desired resolution can be restored by decoding only a portion of the necessary encoded data as described above. In other words, the increase in the reproduction processing load can be suppressed.
[0084] Note that in Figure 2 In FIG, a binary tree is shown as an example of a tree structure of the geometry, but any tree structure may be applied. For example, the tree may be an octree or a kd-tree.
[0085] The encoded data (bit stream) generated by encoding the geometry as described above is also called a geometry bit stream.
[0086] Furthermore, in compression of attributes, methods such as prediction weight boosting, region adaptive hierarchical transform (RAHT), or fixed weight boosting are applied. The encoded data (bitstream) generated by encoding the attributes is also called an attribute bitstream.
[0087] Furthermore, a bitstream in which a geometry bitstream and an attribute bitstream are combined into one stream is also referred to as a G-PCC bitstream or G-PCC content.
[0088] <Slice>
[0089] Non-patent document 4 proposes forming a slice structure in the bitstream of G-PCC. Slices are the basis for dividing data into geometry and attributes. In this specification, a slice of geometry is also referred to as a geometry slice. In addition, a slice of an attribute is also referred to as an attribute slice.
[0090] Slices are divided based on depth in the geometry tree. That is, a slice contains data of a single depth or multiple consecutive depths. For example, Figure 2 In the figure, the data area divided in the bold box indicates the slice. The numbers surrounded by circles in the figure indicate the identification information of each slice. For example, the data from depth 0 to depth 3 form slice #1. In addition, the slices can also be divided according to the position (area) in the three-dimensional space. For example, in Figure 2 In the example, the data at depth 4 and depth 5 are divided into regions A to D and form four slices: slice #2, slice #3, slice #4, and slice #5. Similarly, the data at depth 6 is divided into regions A to D and form four slices: slice #6, slice #7, slice #8, and slice #9.
[0091] Geometry or attribute data is encoded for each slice. That is, the geometry bitstream and attribute bitstream can be decoded for each slice. However, due to the difference in geometry coding from the upper layer, as described above, there are two types of slices: independent slices that can be decoded independently, and dependent slices that require additional slices for decoding.
[0092] For example, in Figure 2 In the case of a tree structure, the geometry of slice #1 can be decoded so that the geometry of slice #1 (for example, the geometry of depth 3) is obtained by decoding. On the other hand, in order to obtain the geometry of slice #2 (for example, the geometry of area A at depth 4) by decoding, it is necessary to decode the geometry of slice #1 and the geometry of slice #2. Similarly, in the decoding of slices #3 to #5, the decoding of slice #1 is also required. In addition, in the decoding of slice #6, the decoding of slice #1 and slice #2 is required. In the decoding of slice #7, the decoding of slice #1 and slice #3 is required. In the decoding of slice #8, the decoding of slice #1 and slice #4 is required. In the decoding of slice #9, the decoding of slice #1 and slice #5 is required.
[0093] That is, slice #1 is an independent slice, and slices #2 to #9 are dependent slices.
[0094] Figure 3 is a diagram showing a main configuration example of a bit stream having such a slice structure. Figure 3 The bit stream 30 indicated in gray in FIG is the bit stream of the point cloud, in which the geometry is processed into a tree structure and a slice structure and encoded, as shown in FIG. Figure 2 Same as in. Figure 3 A portion of the bitstream is shown.
[0095] like Figure 3 As shown in FIG, the bitstream 30 includes a sequence parameter set (SPS), a geometry parameter set (GPS), an attribute parameter set (APS), and a tile library. The sequence parameter set is a parameter set related to the entire sequence. The geometry parameter set is a parameter set related to the geometry. The geometry parameter set can be different based on the geometry slice. The attribute parameter set is a parameter set related to the attribute. The attribute parameter set can be different based on the attribute slice. The tile library stores the location information of the tiles. The number of tiles and the location information of the tiles are variable for each frame.
[0096] As indicated by dashed arrow 31, following the data, a point cloud bitstream is arranged for each sample. A sample is a point cloud at a certain time and corresponds to a frame of a moving image. Within each sample, a bitstream is arranged for each slice. Within each slice, the geometry bitstream and the attribute bitstream are arranged in this order.
[0097] exist Figure 3 In the data unit, each square indicates a data unit. The data unit of "geometry slice #1" is a data unit that stores the geometry bitstream of slice #1. The data unit of "attribute slice" following this data unit is a data unit that stores the attribute bitstream of slice #1. The same slice identification information (slice_id) is assigned to the data units of geometry and attributes included in the same slice. Note that the attribute can be divided into slices for each parameter, and multiple attribute slices can be stored in one data unit. That is, these data units are data units corresponding to slice #1 as an independent slice, and store Figure 2 The data from depth 0 to depth 3 in the tree structure.
[0098] In this specification, a data unit storing a bitstream of a geometry is also referred to as a geometry data unit. In addition, a data unit storing a bitstream of an attribute is also referred to as an attribute data unit.
[0099] The data unit of "Geometry Slice #2" is a geometry data unit that stores the geometry bitstream of Slice #2. The data unit of "Attribute Slice" following this geometry data unit is an attribute data unit that stores the attribute bitstream of Slice #2. That is, these data units are data units corresponding to Slice #2 as a subordinate slice, and store Figure 2 The data of region A at depth 4 and depth 5 of the tree structure of FIG. As indicated by arrow 32, slice #2 depends on slice #1.
[0100] The data unit of "Geometry Slice #6" is a geometry data unit that stores the geometry bitstream of slice #6. The data unit of "Attribute Slice" following this geometry data unit is an attribute data unit that stores the attribute bitstream of slice #6. That is, these data units are data units corresponding to slice #6 as a subordinate slice, and store Figure 2 The data of region A at depth 6 of the tree structure of . As indicated by arrow 33, slice #6 is directly subordinate to slice #2. In other words, slice #6 is also indirectly subordinate to slice #1.
[0101] The data unit of "Geometry Slice #3" is a geometry data unit that stores the geometry bitstream of Slice #3. The data unit of "Attribute Slice" following this geometry data unit is a data unit that stores the attribute bitstream of Slice #3. That is, these data units are data units corresponding to Slice #3 as a subordinate slice, and store Figure 2 The data of region B at depth 4 and depth 5 of the tree structure of FIG. As indicated by arrow 34, slice #3 is subordinate to slice #1.
[0102] The data unit of "Geometry Slice #7" is a geometry data unit that stores the geometry bitstream of slice #7. The data unit of "Attribute Slice" following this geometry data unit is a data unit that stores the attribute bitstream of slice #7. That is, these data units are data units corresponding to slice #7 as a subordinate slice, and store Figure 2 The data of region B at depth 6 of the tree structure of FIG. As indicated by arrow 35, slice #7 is directly subordinate to slice #3. In other words, slice #7 is also indirectly subordinate to slice #1.
[0103] Although not shown, the data units (geometry data units and attribute data units) corresponding to slices #4, #8, #5, and #9 are arranged in a similar manner. Note that slices and tiles are associated with each other using tile identification information (tile_id) stored in the geometry data units.
[0104] Without such a slice structure formed in the bitstream, in order to perform scalable decoding, the decoder needs to parse (analyze) the bitstream and determine which data depth corresponds to which part of the bitstream. On the other hand, as described above, by forming a bitstream with a slice structure, the decoder can easily select data to be decoded on a slice basis.
[0105] For example, in Figure 2 In the case of obtaining the data of slice #1 in the tree structure corresponding to Figure 3 In addition, in the case of obtaining data of slice #2, the decoder can decode the bit stream stored in the data unit corresponding to slice #1 and the bit stream stored in the data unit corresponding to slice #2. In the case of obtaining data of slice #6, the decoder can decode the bit stream stored in the data unit corresponding to slice #1, the bit stream stored in the data unit corresponding to slice #2, and the bit stream stored in the data unit corresponding to slice #6.
[0106] Similarly, when obtaining data of slice #3, the decoder can decode the bit stream stored in the data unit corresponding to slice #1 and the bit stream stored in the data unit corresponding to slice #3. When obtaining data of slice #7, the decoder can decode the bit stream stored in the data unit corresponding to slice #1, the bit stream stored in the data unit corresponding to slice #3, and the bit stream stored in the data unit corresponding to slice #7.
[0107] In the case of obtaining the data of slice #4, the decoder can decode the bitstream stored in the data unit corresponding to slice #1 and the bitstream stored in the data unit corresponding to slice #4. In the case of obtaining the data of slice #8, the decoder can decode the bitstream stored in the data unit corresponding to slice #1, the bitstream stored in the data unit corresponding to slice #4, and the bitstream stored in the data unit corresponding to slice #8.
[0108] In the case of obtaining the data of slice #5, the decoder can decode the bitstream stored in the data unit corresponding to slice #1 and the bitstream stored in the data unit corresponding to slice #5. In the case of obtaining the data of slice #9, the decoder can decode the bitstream stored in the data unit corresponding to slice #1, the bitstream stored in the data unit corresponding to slice #5, and the bitstream stored in the data unit corresponding to slice #9.
[0109] Therefore, the decoder can more easily perform scalable decoding.
[0110] Note that in this specification, an independent slice of a geometry is also referred to as an independent geometry slice. A dependent slice of a geometry is also referred to as a dependent geometry slice. In addition, an independent slice of an attribute is also referred to as an independent attribute slice. A dependent slice of an attribute is also referred to as a dependent attribute slice.
[0111] <Storage of G-PCC Content in ISOBMFF>
[0112] Non-patent document 5 discloses a method of storing G-PCC content (G-PCC bitstream) in ISOBMFF for the purpose of improving the reproduction processing and network distribution efficiency of G-PCC content from a local storage device. This method is standardized in MPEG-I Part 18 (ISO / IEC 23090-18).
[0113] Figure 4 It is a diagram showing an example of the file structure in this case. In this specification, the G-PCC content stored in ISOBMFF is also referred to as a content file. <000028A sample of a media data box (media) includes a geometry slice and an attribute slice corresponding to a point cloud frame. In addition, it may include a geometry parameter set, an attribute parameter set, and a tile library depending on the sample entry type.
[0116] <Storage of Slice-structured G-PCC Content in ISOBMFF>
[0117] For example, in the L-HEVC file format described in Non-Patent Document 2, when the bitstream can be divided into tracks based on slices and stored, so-called progressive download decoding is enabled, for example, when distributing G-PCC content, only a portion corresponding to the playback target depth is transmitted and decoded. As a result, increases in the amount of data transmitted and the amount of data to be decoded can be suppressed. For example, increases in display latency can be suppressed when delivering large-scale point clouds, etc.
[0118] However, in non-patent document 5, it is not disclosed that G-PCC content with such a slice structure is stored in ISOBMFF. Therefore, even in the case where some of the slices of the G-PCC content are decoded, the entire G-PCC content must be transmitted, and therefore there is a possibility of increasing the amount of data to be transmitted. In addition, even when the bitstream is divided using tracks, the information necessary for scalable decoding, such as which depth of data is included in which track, exists only in the bitstream. Therefore, a decoder is required to parse the bitstream. In addition, since intra-frame coding is applied to G-PCC, it is often possible that the hierarchical structure of the depth for each sample (frame) changes. Therefore, in order to obtain the information necessary for scalable decoding, a decoder is required to parse the entire bitstream. In this way, the reproduction processing load is likely to increase.
[0119] <2. Transmission of Scalable Decoding Information via Content Files>
[0120] Therefore, if Figure 5 As shown at the top of the table shown in , scalable decoding information for G-PCC content having a slice structure stored in the metadata area of the content file is transmitted (Method 1). Note that in this specification, scalable decoding information is information regarding scalable decoding of G-PCC content having a slice structure (for scalable decoding information). In addition, scalable decoding information is set based on the depth of each slice of the G-PCC content having a slice structure and a dependency relationship between slices.
[0121] For example, an information processing device includes: a scalable decoding information generation unit that generates scalable decoding information about scalable decoding of G-PCC content based on depth information indicating a quality hierarchy level of a geometry included in each slice of the G-PCC content including a first slice and a second slice, and a dependency relationship between the first slice and the second slice in the G-PCC content; and a content file generation unit that generates a content file in which the G-PCC content is stored and stores the scalable decoding information in a metadata area of the content file.
[0122] For example, according to the information processing method, based on depth information indicating the quality hierarchy level of a geometry included in each slice in the G-PCC content including the first and second slices and the dependency relationship between the first slice and the second slice in the G-PCC content, scalable decoding information about scalable decoding of the G-PCC content is generated, a content file storing the G-PCC content is generated, and the scalable decoding information is stored in a metadata area of the content file.
[0123] Furthermore, for example, the information processing apparatus includes an extraction unit that extracts any slice of G-PCC content from a content file storing G-PCC content including first and second slices based on scalable decoding information stored in a metadata area of the content file, and a decoding unit that decodes the slice of the G-PCC content extracted by the extraction unit. Note that the scalable decoding information is information regarding scalable decoding of the G-PCC content, and is information generated based on depth information indicating a quality hierarchy level of a geometry included in a slice in the G-PCC content and a dependency relationship between the first slice and the second slice in the G-PCC content.
[0124] For example, according to the information processing method, based on scalable decoding information stored in a metadata area of a content file in which G-PCC content including first and second slices is stored, any slice of G-PCC content is extracted from the content file, and the extracted slice of the G-PCC content is decoded. Note that the scalable decoding information is information regarding scalable decoding of the G-PCC content, and is information generated based on depth information indicating the quality hierarchy level of a geometry included in a slice in the G-PCC content and a dependency relationship between the first slice and the second slice in the G-PCC content.
[0125] In this way, the decoder can extract and decode the slices necessary to reproduce the point cloud of the desired depth or area based on the scalable decoding information in the metadata area of the content file, and can generate rendering information. Therefore, unnecessary processing of the decoder (transmission of unnecessary information, parsing of the bitstream, etc.) can be reduced. Therefore, the increase in the reproduction processing load can be suppressed.
[0126] As a use case for G-PCC content, large-scale point cloud data, such as point clouds and mapping data for virtual assets (real movie sets converted into digital data) in film production, is encoded. Such large-scale point clouds have a very large data volume as a whole, and reproducing the entire point cloud is unrealistic from the perspectives of processing load and processing latency. Therefore, it is preferable to perform scalable decoding that reproduces only a portion of the data by limiting the area to be reproduced or reducing the resolution.
[0127] When a G-PCC file is stored in a content file and transmitted, by storing scalable decoding information in the metadata area of the content file as described above and supporting scalable decoding, only the necessary information can be transmitted or only the necessary information can be decoded more easily. As described above, the increase in the reproduction processing load can be further suppressed as the data increases, and thus a better effect can be achieved.
[0128] In a content file, there are structures in which geometry and attributes are stored in one track (also called a single-track packaging structure) and structures in which geometry and attributes are stored in different tracks (also called a multi-track packaging structure). In the following description, an example of a single track will be used, but the present technology can be applied to cases of multiple tracks like the single track case. Note that even in the case of a single track, the number of tracks can be plural (there can be multiple tracks including geometry and attributes).
[0129] In addition, in the following description, as referenced Figure 3 and Figure 4 As described, it is assumed that SPS, GPS, APS, and tile inventory are stored in GPCCDecoderConfigurationRecord, and only geometry slices and attribute slices are stored in samples. However, some or all of SPS, GPS, APS, and tile inventory may be stored in samples.
[0130] <2-1. Slice composition information for each sample>
[0131] like Figure 5 As shown in the second row from the top of the table shown in , the scalable decoding information may include slice composition information for each sample (method 1-1). The slice composition information is information about the structure of the slices in the sample. That is, the decoder can obtain the structure information of the slices of each sample from the metadata area of the content file. Therefore, the decoder can determine the structure of the slices of each sample without parsing the bitstream (i.e., more easily).
[0132] For example, Figure 5 As shown in the third row from the top of the table shown in , the slice composition information can be stored in the codec-specific parameters of the SubSampleInformationBox in the metadata area of the content file (method 1-1-1). For example, the content file generation unit of the encoder can set subsamples for each slice and store the slice composition information in the codec-specific parameters of the Subsample Information Box in the metadata area of the content file. In addition, the extraction unit of the decoder can extract any slice of the G-PCC content from the content file based on the slice composition information stored in the codec-specific parameters of the Subsample Information Box in the metadata area of the content file of the subsample set for each slice.
[0133] Since intra-frame coding is applied to G-PCC, it is often possible that the hierarchical structure of the depth for each sample (frame) changes. Therefore, when slice composition information is stored in sample groups such as in the L-HEVC file format, it is necessary to associate the slice composition information with sample groups having different information for each sample. Therefore, the file size may be unnecessarily increased by a large amount of grouping information. In addition, in order to determine the construction of the slice of the desired sample, the decoder is required to check all sample groups. Therefore, the load of the reproduction process may increase.
[0134] As described above, by setting subsamples for each slice and storing slice configuration information in codec-specific parameters in a subsample information box within the metadata area of a content file, it is possible to suppress an increase in file size. Furthermore, since a decoder only needs to confirm the codec-specific parameters of a desired subsample information box, it is easier to confirm slice configuration information. In other words, it is possible to suppress an increase in the playback processing load.
[0135] <2-1-1. Slice dependency information>
[0136] For example, Figure 5 As shown in the fourth row from the top of the table shown in , the slice composition information may include slice dependency information (method 1-1-2). In this specification, slice dependency information is information indicating a dependency relationship between slices or slice groups. For example, slice dependency information indicates a dependency relationship between a first and a second slice included in the G-PCC content. In addition, in this specification, a slice group is a plurality of slices corresponding to the same depth.
[0137] For example, Figure 7 The bit stream 100 shown in FIG. Figure 3 The bitstream 101 indicated in grey indicates a partial configuration of samples of the bitstream 100. As indicated by arrows 111 to 114, each data unit of the bitstream 101 is similar to Figure 3 The case has dependencies between slices.
[0138] For example, the geometry data unit of "Geometry Slice #2" and the attribute data unit of its subsequent "Attribute Slice" are data units corresponding to Slice #2, and are subordinate to the data units of Slice #1 (the geometry data unit of "Geometry Slice #1" and the attribute data unit of its subsequent "Attribute Slice") as indicated by arrow 111. That is, Slice #2 is subordinate to Slice #1.
[0139] Furthermore, the geometry data unit of "Geometry Slice #6" and the attribute data unit of its subsequent "Attribute Slice" are data units corresponding to Slice #6, and are directly subordinate to the data units of Slice #2 (the geometry data unit of "Geometry Slice #2" and the attribute data unit of its subsequent "Attribute Slice") as indicated by arrow 112. That is, Slice #6 is directly subordinate to Slice #2. In other words, Slice #6 is also indirectly subordinate to Slice #1.
[0140] Furthermore, the geometry data unit of "Geometry Slice #3" and the attribute data unit of its subsequent "Attribute Slice" are data units corresponding to Slice #3, and are subordinate to the data units of Slice #1 (the geometry data unit of "Geometry Slice #1" and the attribute data unit of its subsequent "Attribute Slice") as indicated by arrow 113. That is, Slice #3 is subordinate to Slice #1.
[0141] Furthermore, the geometry data unit of "Geometry Slice #7" and the attribute data unit of its subsequent "Attribute Slice" are data units corresponding to slice #7 and are directly subordinate to the data units of slice #3 (the geometry data unit of "Geometry Slice #3" and the attribute data unit of its subsequent "Attribute Slice") as indicated by arrow 114. That is, slice #7 is directly subordinate to slice #3. In other words, slice #7 is also indirectly subordinate to slice #1.
[0142] That is, the information indicated by the arrows between the slices surrounded by the dotted box 121 (e.g., arrows 111 to 114) is slice dependency information. The slice dependency information is stored in the metadata area of the content file. In this way, the decoder can obtain the slice dependency information from the metadata area of the content file. Therefore, the decoder can determine the slice dependency information without parsing the bitstream (i.e., more easily).
[0143] For example, Figure 5 As shown in the fifth row from the top of the table shown in , the slice dependency information may be stored in the codec-specific parameters of the subsample information box in the metadata area of the content file (method 1-1-2-1). For example, the content file generation unit of the encoder may set subsamples for each slice and store the slice dependency information in the codec-specific parameters of the subsample information box in the metadata area of the content file. In addition, the extraction unit of the decoder may extract any slice of the G-PCC content from the content file based on the slice dependency information stored in the codec-specific parameters of the subsample information box in the metadata area of the content file of the subsample set for each slice.
[0144] Since intra-frame coding is applied to G-PCC, it is often possible that the hierarchy of the depth for each sample (frame) changes. Therefore, when slice dependency information is stored in sample groups such as in the L-HEVC file format, it is necessary to associate the slice dependency information with sample groups having different information for each sample. Therefore, it is possible to unnecessarily increase the file size by grouping information. In addition, in order to determine the configuration of the slice of the desired sample, the decoder is required to check all sample groups. Therefore, the load of the reproduction processing may increase.
[0145] As described above, by setting subsamples for each slice and storing slice dependency information in the codec-specific parameters of the subsample information box in the metadata area of the content file, it is possible to suppress an increase in file size. Furthermore, since the decoder only needs to confirm the codec-specific parameters of the desired subsample information box, slice dependency can be confirmed more easily. In other words, it is possible to suppress an increase in the load of the playback process.
[0146] For example, Figure 5 As shown in the sixth row from the top of the table shown in , the slice dependency information may include reference source geometry slice identification information and reference destination geometry slice identification information (method 1-1-2-2). In this specification, the reference source geometry slice identification information is identification information of the geometry slice corresponding to the information. That is, the reference source geometry slice identification information is used as a reference source in the dependency relationship between the above-mentioned slices or slice groups ( Figure 7 The reference destination geometric slice identification information is the identification information of the geometric slice of the image (the starting point of the arrow in the dotted box 121). In addition, in this specification, the reference destination geometric slice identification information is the identification information of another geometric slice referenced by the geometric slice corresponding to the information. That is, the reference destination geometric slice identification information is used as the reference destination ( Figure 7 The identification information of the geometric slice (the end point of the arrow in the dotted box 121).
[0147] In this way, the decoder can obtain the reference source geometry slice identification information and the reference destination geometry slice identification information from the metadata area of the content file. Therefore, the decoder can determine the reference source geometry slice identification information and the reference destination geometry slice identification information without parsing the bitstream (i.e., more easily).
[0148] For example, Figure 5 As shown in the seventh row from the top of the table shown in , when the slice corresponding to the slice dependency information (i.e., the slice corresponding to the reference source geometry slice identification information) is an independent geometry slice, the reference source geometry slice identification information and the reference destination geometry slice identification information can be the same (method 1-1-2-2-1). That is, since the independent geometry slice can be decoded without requiring another slice, the reference destination can be set to the independent geometry slice itself. In addition, for the independent geometry slice, the storage of the reference destination geometry slice identification information can be omitted.
[0149] For example, Figure 5 As shown in the eighth row from the top of the table shown in , the reference source geometry slice identification information and the reference destination geometry slice identification information can be stored in the codec-specific parameters of the subsample information box in the metadata area of the content file (method 1-1-2-2-1). For example, the content file generation unit of the encoder can set subsamples for each slice and store the reference source geometry slice identification information and the reference destination geometry slice identification information in the codec-specific parameters of the subsample information box in the metadata area. In addition, the extraction unit of the decoder can extract any slice of the G-PCC content from the content file based on the reference source geometry slice identification information and the reference destination geometry slice identification information stored in the codec-specific parameters of the subsample information box in the metadata area of the subsample set for each slice.
[0150] Figure 8 An example of the syntax of the subsample information box is shown in FIG. In the subsample information box, a flag is set as indicated by underline 131. In addition, as indicated by underline 132, codec specific parameters are provided. The codec specific parameters store information of subsamples determined for each encoding codec.
[0151] Figure 9 An example of the syntax of codec-specific parameters is shown in FIG. As indicated by the dotted frame 133, the codec-specific parameters store reference source geometry slice identification information (geom_slice_id) and reference destination geometry slice identification information (ref_geom_slice_id). The slice identification information (slice_id) of the geometry slice corresponding to the information is set to geom_slice_id. The slice identification information (slice_id) of another geometry slice referenced by the geometry slice is set to ref_geom_slice_id.
[0152] In this way, by setting subsamples for each slice and storing the reference source geometry slice identification information and the reference destination geometry slice identification information in the codec-specific parameters of the subsample information box in the metadata area of the content file, it is possible to suppress an increase in file size. Furthermore, since the decoder only needs to confirm the codec-specific parameters of the desired subsample information box, it is easier to confirm the reference relationship of the geometry slices (the geometry slices of the reference source and reference destination). In other words, it is possible to suppress an increase in the load of the reproduction process.
[0153] Note that in Figure 7 In the case of the example, since the geometry data unit and the attribute data unit immediately following it constitute a slice, the reference relationship from the attribute slice to the geometry slice is obvious. Figure 9 In the example, the reference relationship from the attribute slice to the geometry slice is omitted.
[0154] <attribute geometry slice identification information>
[0155] Note that you can clarify the reference relationship between the attribute slice and the geometry slice. For example, Figure 5 As shown in the ninth row from the top of the table shown in , the slice dependency information may include attribute geometry slice identification information (method 1-1-2-3). The attribute geometry slice identification information is identification information of the geometry slice referenced by the attribute slice. By specifying the reference relationship between the attribute slice and the geometry slice in this way, the constraint on the positional relationship between the geometry data unit and the attribute data unit can be eliminated.
[0156] For example, Figure 5 As shown in the tenth row from the top of the table shown in , the attribute geometry slice identification information can be stored in the codec-specific parameters of the subsample information box in the metadata area of the content file (method 1-1-2-3-1). For example, the content file generation unit of the encoder can set subsamples for each slice and store the attribute geometry slice identification information in the codec-specific parameters of the subsample information box in the metadata area. In addition, the extraction unit of the decoder can extract any slice of the G-PCC content from the content file based on the attribute geometry slice identification information stored in the codec-specific parameters of the subsample information box in the metadata area of the subsample set for each slice.
[0157] Figure 10 An example of the syntax of the codec-specific parameter in this case is shown in FIG. As indicated by underline 134, attribute geometry slice identification information (ref_attr_geom_slice_id) is stored in the codec-specific parameter. Slice identification information (slice_id) of the geometry slice referenced by the attribute slice corresponding to the information is set to ref_attr_geom_slice_id.
[0158] In this way, by setting a subsample for each slice and storing the attribute geometry slice identification information in the codec-specific parameters of the subsample information box in the metadata area of the content file, it is possible to suppress an increase in file size. Furthermore, since the decoder only needs to confirm the codec-specific parameters of the desired subsample information box, it is easier to confirm the geometry slice referenced from the attribute. In other words, it is possible to suppress an increase in the playback processing load.
[0159] <Non-scalable coding>
[0160] Non-scalable coding can be applied to attributes. Non-scalable coding is a coding scheme that is incompatible with scalable decoding. For example, when non-scalable coding is applied to an attribute, the attribute slice is set to be able to be decoded independently of other attribute slices (i.e., without reference to other attribute slices). Therefore, data (depth) overlaps between attribute data units.
[0161] Figure 11 An example of the configuration of a bit stream in the case where non-scalable coding is applied to attributes in this way is shown in FIG. Figure 11 The gray bitstream 141 shown in FIG. 1 shows a portion of the structure of the G-PCC content (G-PCC bitstream). In the case of the bitstream 141, the attribute data unit 151 stores attribute slices corresponding to the geometry of depth 0 to depth 3. The attribute data unit 152 stores attribute slices corresponding to the geometry of depth 0 to depth 5. The attribute data unit 153 stores attribute slices corresponding to the geometry of depth 0 to depth 6.
[0162] Thus, for example, attributes corresponding to geometry at depths 0 to 6 can be obtained by decoding attribute data unit 153. In other words, to obtain attributes corresponding to geometry at depth 6, attribute data unit 153 can be decoded (without decoding other attribute data units).
[0163] Similarly, by decoding attribute data unit 152, attributes corresponding to geometry at depths 0 to 5 can be obtained. In other words, to obtain attributes corresponding to geometry at depth 4 or depth 5, attribute data unit 152 can be decoded (without decoding other attribute data units).
[0164] Similarly, by decoding attribute data unit 151, attributes corresponding to geometries at depths 0 to 3 can be obtained. In other words, to obtain attributes corresponding to a geometry at depths 0 to 3, attribute data unit 151 can be decoded (without decoding other attribute data units).
[0165] In such a case, for example, multiple geometry data units can be candidates for reference destinations for a single attribute data unit (e.g., attribute data unit 152 or attribute data unit 153). Therefore, when non-scalable coding is applied to an attribute, slice identification information of a geometry slice corresponding to the lowest level (maximum depth) in the depth corresponding to the attribute slice (attribute data unit) is set in the attribute geometry slice identification information. The reference relationship between geometry slices (the reference relationship with the higher geometry slice) is indicated by ref_geom_slice_id.
[0166] In this specification, attribute geometry slice identification information in the case where non-scalable coding is applied to an attribute in this manner is also referred to as non-scalable coded attribute geometry slice identification information.
[0167] That is to say, if Figure 5 As shown in the eleventh row from the top of the table shown in , the slice dependency information may include non-scalable coded attribute geometry slice identification information as identification information of a geometry slice referenced by an attribute slice to which non-scalable coding is applied (method 1-1-2-4).
[0168] For example, Figure 5 The non-scalable coded attribute geometry slice identification information shown in the twelfth row from the top of the table shown in may include identification information of a geometry slice or a geometry slice group having geometry containing maximum depth information among the geometry slices or geometry slice groups referenced by the attribute slice corresponding to the information (method 1-1-2-4-1). Note that in this specification, a geometry slice group is a plurality of geometry slices corresponding to the same depth.
[0169] In this way, even in the case where non-scalable coding is applied to attributes, the constraints on the positional relationship between the geometry data unit and the attribute data unit can be eliminated.
[0170] For example, Figure 5 As shown in the thirteenth row at the top of the table shown in , the non-scalable coding attribute geometry slice identification information can be stored in the codec-specific parameters of the subsample information box in the metadata area of the content file (method 1-1-2-4-2). For example, the content file generation unit of the encoder can set subsamples for each slice and store the non-scalable coding attribute geometry slice identification information in the codec-specific parameters of the subsample information box in the metadata area of the content file. In addition, the extraction unit of the decoder can extract any slice of the G-PCC content from the content file based on the non-scalable coding attribute geometry slice identification information stored in the codec-specific parameters of the subsample information box in the metadata area of the content file of the subsample set for each slice.
[0171] In this way, by setting subsamples for each slice and storing the non-scalable coding attribute geometry slice identification information in the codec-specific parameters of the subsample information box in the metadata area of the content file, it is possible to suppress an increase in file size. Furthermore, the decoder only needs to confirm the codec-specific parameters of the desired subsample information box. Therefore, even when non-scalable coding is applied to an attribute, it is easier to confirm the geometry slice referenced from the attribute. In other words, it is possible to suppress an increase in the playback processing load.
[0172] In addition, if Figure 5 As shown in the fourteenth row from the top of the table shown in , the slice dependency information may include a non-scalable coding flag (method 1-1-2-4-3). The non-scalable coding flag is flag information indicating whether non-scalable coding is applied to the attribute slice. By storing such information, the decoder can easily determine whether non-scalable coding is applied.
[0173] In addition, if Figure 5 As shown in the fifteenth row from the top of the table shown in , the non-scalable coding flag can be stored in the codec-specific parameters of the subsample information box in the metadata area of the content file (method 1-1-2-4-3-1). For example, the content file generation unit of the encoder can set subsamples for each slice and store the non-scalable coding flag in the codec-specific parameters of the subsample information box in the metadata area of the content file. In addition, the extraction unit of the decoder can extract any slice of G-PCC content from the content file based on the non-scalable coding flag stored in the codec-specific parameters of the subsample information box in the metadata area of the content file of the subsample set for each slice.
[0174] Figure 12 An example of the syntax of the codec specific parameters in this case is shown in As indicated in a dotted box 161 , the codec specific parameters store a non-scalable coding flag (non_scalable_flag) and non-scalable coding attribute geometry slice identification information (ref_attr_geom_slice_id).
[0175] When the attribute slice is scalably coded, the value of non_scalable_flag is set to 0 (false). When the attribute slice is non-scalably coded, the value of non_scalable_flag is set to 1 (true).
[0176] When the value of non_scalable_flag is 1 (true), ref_attr_geom_slice_id is set to the non-scalable coding attribute geometry slice identification information. That is, in this case, in ref_attr_geom_slice_id, the identification information (slice_id) of the geometry slice or geometry slice group including the maximum depth among the geometry slice or geometry slice group referenced by the attribute slice corresponding to the information is set.
[0177] When the value of non_scalable_flag is 0 (false), ref_attr_geom_slice_id is set to attribute geometry slice identification information. That is, in this case, the slice identification information (slice_id) of the geometry slice referenced by the attribute slice corresponding to the information is set to ref_attr_geom_slice_id.
[0178] <Change Payload Type>
[0179] Note that in codec specific parameters, such as Figure 5 As shown in the bottom of the table shown in , the payload type (PayloadType) of the slice dependency information of the independent geometry slice and the payload type of the slice dependency information of the dependent geometry slice can be set to different values.
[0180] Figure 13 16 is a diagram showing an example of the syntax of codec-specific parameters in this case. In this example, as indicated in the dotted box 162, the payload type of the slice dependency information of the independent geometry slice is set to "2", and the payload type of the slice dependency information of the dependent geometry slice is set to "9". In this way, the decoder can more easily identify the type of geometry slice based on the payload type.
[0181] Note that in Figure 13 In the example of FIG, "2" and "9" have been described as examples of the value of the payload type, but these values are exemplary. The values of the payload type of the independent geometry slice and the dependent geometry slice can be different values from each other and are not limited to these examples.
[0182] <2-1-2. Geometry Slice Depth Information>
[0183] like Figure 6 As shown at the top of the table shown in , the slice composition information may include geometry slice depth information (method 1-1-3). In this specification, geometry slice depth information is information about depth information of the geometry included in the geometry slice or geometry slice group corresponding to the information. The slice composition information may include both slice dependency information (method 1-1-2) and geometry slice depth information.
[0184] For example, in Figure 7 , the geometry data unit of "Geometry Slice #1" corresponds to slice #1. Slice #1 includes geometry at depth 0 to depth 3. In addition, the geometry data unit of "Geometry Slice #2" corresponds to slice #2. Slice #2 includes geometry at depth 4 and depth 5. The geometry data unit of "Geometry Slice #6" corresponds to slice #6. Slice #6 includes geometry at depth 6. The geometry data unit of "Geometry Slice #3" corresponds to slice #3. Slice #3 includes geometry at depth 4 and depth 5. The geometry data unit of "Geometry Slice #7" corresponds to slice #7. Slice #7 includes geometry at depth 6.
[0185] That is, the depth information of each slice indicated in the dotted box 122 is the geometry slice depth information. The geometry slice depth information is stored in the metadata area of the content file. In this way, the decoder can obtain the geometry slice depth information from the metadata area of the content file. Therefore, the decoder can determine the geometry slice depth information without parsing the bitstream (i.e., more easily).
[0186] Note that geometry slices can include geometry at multiple depths.
[0187] Therefore, if Figure 6 As shown in the second row from the top of the table shown in , the geometry slice depth information may include minimum depth information (method 1-1-3-1). In this specification, the minimum depth information represents information indicating the minimum value of the depth information in the geometry slice or geometry slice group corresponding to the information. In this way, the decoder can obtain the minimum depth information from the metadata area of the content file. Therefore, the decoder can determine the minimum value of the depth included in the geometry slice without parsing the bitstream (i.e., more easily).
[0188] In addition, if Figure 6 As shown in the third row from the top of the table shown in , the geometry slice depth information may include maximum depth information (method 1-1-3-2). In this specification, the maximum depth information represents information indicating the maximum value of the depth information in the geometry slice or geometry slice group corresponding to the information. In this way, the decoder can obtain the maximum depth information from the metadata area of the content file. Therefore, the decoder can determine the maximum value of the depth included in the geometry slice without parsing the bitstream (i.e., more easily).
[0189] Of course, the geometry slice depth information may include both minimum depth information and maximum depth information. In this case, the decoder may determine the range of depths included in the geometry slice without parsing the bitstream (ie, more easily).
[0190] like Figure 6 As shown in the fourth row from the top of the table shown in , the geometric slice depth information can be stored in the codec-specific parameters of the subsample information box in the metadata area of the content file (method 1-1-3-3). For example, the content file generation unit of the encoder can set subsamples for each slice or slice group, and store the geometric slice depth information in the codec-specific parameters of the subsample information box in the metadata area of the content file. In addition, the extraction unit of the decoder can extract any slice of the G-PCC content from the content file based on the geometric slice depth information stored in the codec-specific parameters of the subsample information box in the metadata area of the content file of the subsample set for each slice or slice group.
[0191] Figure 14 An example of the syntax of codec specific parameters in this case is shown. As indicated in the dashed box 171, Figure 14 The minimum depth information (min_depth) and the maximum depth information (max_depth) are set in the codec specific parameters shown in .
[0192] In this way, by setting subsamples for each slice or slice group and storing the geometry slice depth information in the codec-specific parameters of the subsample information box in the metadata area of the content file, it is possible to suppress an increase in file size. Furthermore, since the decoder only needs to confirm the codec-specific parameters of the desired subsample information box, it is easier to confirm the depth included in the geometry slice. In other words, it is possible to suppress an increase in the reproduction processing load.
[0193] Note that, in the case where a plurality of geometry slices having the same depth range are continuous, as described above, the geometry slice is set as a subsample of the geometry slice group. Then, geometry slice depth information is generated for each subsample. That is, the geometry slice depth information of each slice of the geometry slice group is collected into one piece of geometry slice depth information. Therefore, compared with the case where geometry slice depth information is generated for each slice of the geometry slice group, the increase in subsample entries can be suppressed. Therefore, the increase in bit cost (i.e., the amount of data in the bitstream) can be suppressed.
[0194] Note that, in the case where a geometry slice or a geometry slice group includes only one depth, the minimum depth information (min_depth) and the maximum depth information (max_depth) are set to the same value.
[0195] In addition, if Figure 6 As shown in the fifth row from the top of the table shown in , the flag of the subsample information box of the codec-specific parameters storing the geometry slice depth information and the flag of the subsample information box of the codec-specific parameters storing the slice dependency information can be set to different values (method 1-1-3-3-1).
[0196] That is, the slice composition information may further include slice dependency information indicating a dependency relationship between the first slice and the second slice in addition to the geometry slice depth information. The content file generation unit of the encoder may then set a flag of the subsample information box storing the geometry slice depth information to a different value from a flag of the subsample information box storing the slice dependency information. Furthermore, the flag of the subsample information box storing the geometry slice depth information and the flag of the subsample information box storing the slice dependency information may be set to different values.
[0197] For example, Figure 14 As shown in FIG, although the flag of the subsample information box in which the slice dependency information is stored is set to “0”, the flag of the subsample information box in which the geometry slice depth information is stored is set to “2”. Of course, these values are exemplary, and the values of the flags are not limited to these examples (optional).
[0198] In this way, geometry slice depth information and slice dependency information can be stored in mutually different subsample information boxes having mutually different flag values, and these information can be used in combination.
[0199] <2-2. Track Configuration Information>
[0200] like Figure 6 As shown in the sixth row from the top of the table shown in , the scalable decoding information may include track composition information (method 1-2). In this specification, the track composition information is information about the construction of the track in which the slice-based G-PCC content is stored in the content file. That is, the decoder can obtain the construction information of the track included in the content file from the metadata area of the content file. Therefore, the decoder can determine the construction of the track of the content file without parsing the bitstream (i.e., more easily). Note that the scalable decoding information may include slice composition information (method 1-1) and track composition information for each sample.
[0201] <2-2-1. Track Depth Information>
[0202] like Figure 6 As shown in the seventh row from the top of the table shown in , the track configuration information may include track depth information (method 1-2-1). In this specification, the track depth information is information about the depth information of the geometry of all slices included in the track corresponding to the information.
[0203] Figure 15 1 is a diagram showing an example of a storage state of a bitstream. It is assumed that a content file (ISOBMFF) has track 1 (Track 1) indicated by rectangle 181, track 2 (Track 2) indicated by rectangle 182, and track 3 (Track 3) indicated by rectangle 183. It is assumed that the geometry bitstream and attribute bitstream of depth 0 to depth 3 (depth=0 to 3) are stored in track 1. It is assumed that the geometry bitstream and attribute bitstream of depth 4 and depth 5 (depth=4 to 5) are stored in track 2. It is assumed that the geometry bitstream and attribute bitstream of depth 6 (depth=6) are stored in track 3.
[0204] In this case, assuming that Figure 15 Each data unit (slice) of the bitstream 101, indicated by the dashed arrows in FIG, is stored in each track. That is, the geometry and attributes of slice #1 are stored in track 1. The geometry and attributes of slice #2 are stored in track 2. The geometry and attributes of slice #3 are stored in track 2. The geometry and attributes of slice #6 are stored in track 3. The geometry and attributes of slice #7 are stored in track 7.
[0205] In this way, the track depth information indicates which depth data is stored in each track. Track depth information is set for each track. In other words, the track depth information indicates which depth data is stored in the track corresponding to the information. As described above, data at multiple depths can be stored in a track. Furthermore, data for multiple samples can be stored in a track. The slice structure of each sample can be different. In other words, the depth included in each track can be changed for each sample.
[0206] like Figure 6 As shown in the eighth row from the top of the table shown in , the track depth information may include track minimum depth information indicating the minimum value of the depth information in the track corresponding to the information (method 1-2-1-1). That is, the track minimum depth information indicates the minimum value of the depth information among all samples included in the track corresponding to the information.
[0207] In addition, if Figure 6 As shown in the ninth row from the top of the table shown in , the track depth information may include track maximum depth information indicating the maximum value of the depth information in the track corresponding to the information (method 1-2-1-2). That is, the track maximum depth information indicates the maximum value of the depth information in all samples included in the track corresponding to the information. Note that the track depth information may include both the track minimum depth information (method 1-2-1-1) and the track maximum depth information.
[0208] In addition, if Figure 6 As shown in the tenth row from the top of the table shown in , the track depth information may include a match flag (method 1-2-1-3). In this specification, the match flag is flag information indicating whether the sample minimum depth information (which is the minimum value of the depth information in each sample included in the track corresponding to the information) matches the track minimum depth information and whether the sample maximum depth information (which is the maximum value of the depth information in each sample included in the track corresponding to the information) matches the track maximum depth information. That is, the match flag is flag information indicating whether the minimum value and the maximum value of the depth information are common to all samples in the track. Note that the track depth information may include all of the match flag, the track minimum depth information (method 1-2-1-1), and the track maximum depth information (method 1-2-1-2).
[0209] In addition, if Figure 6 As shown in the eleventh row from the top of the table shown in , the track depth information may be stored in the depth information box (DepthInfoBox) of the sample entry (SampleEntry) corresponding to each track in the metadata area (Method 1-2-1-4). For example, the content file generation unit of the encoder may store the track depth information in the depth information box of the sample entry in the metadata area. In addition, the extraction unit of the decoder may extract any slice of the G-PCC content from the content file based on the track depth information stored in the depth information box of the sample entry in the metadata area.
[0210] Figure 16 : is a diagram showing an example of the syntax of the depth information box (DepthInfoBox). Figure 16 As shown in , the track minimum depth information (track_min_depth), the track maximum depth information (track_max_depth) and the matching flag (fixed_depth) are set in the depth information box.
[0211] When the match flag is "0" (false), the minimum sample depth information and the maximum sample depth information of all samples in the track are set to values within the range of the track minimum depth information (track_min_depth) to the track maximum depth information (track_max_depth). That is, in this case, the minimum value, the maximum value, or both of the depth of each sample can be changed for each sample.
[0212] When the match flag is "1" (true), the flag indicates that the sample minimum depth information of each sample in the track matches the track minimum depth information (track_min_depth), and the sample maximum depth information of each sample in the track matches the track maximum depth information (track_min_depth). That is, in this case, the minimum and maximum values of the depth of each sample take the common values of all samples.
[0213] Note that, for example, in the case where the track includes data of only one depth as in Track 3, the value of the matching flag (fixed_depth) is set to "1", and the track minimum depth information (track_min_depth) and the track maximum depth information (track_max_depth) take the same value (track_min_depth=track_max_depth).
[0214] As described above, by storing track depth information in the depth information box of the sample entry corresponding to each track, the decoder can more easily determine which depth data is stored in each track based on this information (without parsing the bitstream). In other words, it is possible to suppress an increase in the playback processing load.
[0215] For example, when fixed_depth=1 and there is a desired LoD within the range of track_min_depth to track_max_depth, the desired LoD can be reliably obtained by processing the track and the track to be referenced. That is, track depth information can be useful information in the client's track selection process.
[0216] Instead of the depth, the maximum LoD value and the minimum LoD value obtained by processing the track and the track to be referenced (if any) may be signaled.
[0217] In addition, a non-scalable coding attribute flag (non_scalable_attribute_flag) can be added to clearly indicate whether attribute slices to which non-scalable coding is applied are included in the track. When the non-scalable coding attribute flag (non_scalable_attribute_flag) is "1" (true), this flag indicates that attribute slices to which non-scalable coding is applied are included in the track. In addition, when the non-scalable coding attribute flag (non_scalable_attribute_flag) is "0" (false), this flag indicates that attribute slices to which non-scalable coding is applied are not included in the track.
[0218] <2-2-2. Track Dependency Information>
[0219] like Figure 6 As shown in the twelfth row from the top of the table shown in , the track composition information may include track dependency information (method 1-2-2). In this specification, track dependency information is information indicating a dependency relationship between tracks (e.g., a dependency relationship between a first track and a second track). Note that the track composition information may include both track depth information (method 1-2-1) and track dependency information.
[0220] For example, in Figure 15 , each track (Track 1 to Track 3) has dependencies between tracks as indicated by arrows 184 and 185.
[0221] For example, as indicated by arrow 184, track 2 is dependent on track 1. That is, in order to decode the bitstream of track 2 and restore the geometry or attributes of depth 4 or depth 5, the bitstream of track 1 corresponding to depth 0 to depth 3 also needs to be decoded.
[0222] Similarly, as indicated by arrow 185, track 3 is directly subordinate to track 2. That is, as indicated by arrow 186, track 3 is indirectly subordinate to track 1. The track dependency information indicates such dependency between tracks.
[0223] like Figure 6 As shown in the thirteenth row from the top of the table shown in , the track dependency information may include subordinate information indicating another track, wherein the other track includes a slice necessary for decoding the subordinate slice included in the track corresponding to the information (method 1-2-2-1). The dependency information is information indicating the dependency destination (i.e., the track located on the end side of the above-mentioned arrow (arrows 184 to 186) according to the track dependency information located on the start side of the above-mentioned arrow). By storing the dependency information in the metadata area of the content file, the decoder can more easily confirm other tracks necessary for decoding the track to be processed based on the dependency information without parsing the bitstream. Therefore, the increase in the reproduction processing load can be suppressed.
[0224] Note that Figure 6 As shown in the fourteenth row from the top of the table shown in , the dependent information may indicate all other tracks including slices necessary for decoding the dependent slices included in the track corresponding to the information (method 1-2-2-1-1). Figure 15 In this case, the subordinate information of track 3 may include both the information corresponding to arrow 185 and the information corresponding to arrow 186.
[0225] In addition, if Figure 6 As shown in the fifteenth row from the top of the table shown in , the dependent information may indicate another track including the slice referenced by the dependent information (method 1-2-2-1-2). That is, the dependent information may indicate only the slices directly subordinate to the slice corresponding to the information. For example, in Figure 15 In this case, the subordinate information of track 3 may include only the information corresponding to arrow 185.
[0226] like Figure 6 As shown at the bottom of the table shown in , the track dependency information may be stored in the metadata area as a track reference.
[0227] For example, Figure 15 The dependency information shown in may be stored as a track reference. In this case, the reference type of the track reference may be, for example, depd (reference_type='depd').
[0228] Furthermore, a track reference can be used to associate only a track including slices directly referenced in decoding of dependent slices included in the track.
[0229] A sample entry 4CC of a track that includes independent slices and is decodable only by the track may be designated as “gpc1.” Also, a sample entry 4CC of a track that does not include independent slices and cannot be decodable only by the track may be designated as “lgp1.”
[0230] In addition, if Figure 6 As shown in the sixteenth row from the top of the table shown in , the track dependency information may include independent information indicating another track including dependent slices necessary for independent slices included in the track corresponding to the information in decoding (method 1-2-2-2).
[0231] For example, in Figure 15 , as indicated by arrow 184, when the dependent slice of track 2 is decoded, the independent slice of track 1 is referenced. Similarly, as indicated by arrow 186, when the dependent slice of track 3 is decoded, the independent slice of track 1 is referenced.
[0232] In other words, in Figure 17 In the example, as indicated by arrow 191, the independent slice stored in track 1 is used to decode the dependent slice of track 2. Furthermore, as indicated by arrow 192, the independent slice stored in track 1 is used to decode the dependent slice of track 3. Independent information is information indicating such a dependency relationship. In other words, independent information is reverse lookup information for dependent information.
[0233] Track 3 is indirectly subordinate to Track 1. Therefore, the subordinate relationship indicated by arrow 192 may or may not be included in the independent information. Figure 6 As shown in the seventeenth row from the top of the table shown in , the independent information may indicate another track including another slice that refers to the independent slice in decoding. In addition, the track dependency information may include both dependent information and independent information.
[0234] As mentioned above, track dependency information can be stored in the metadata area as track references. Figure 17 As shown in , independent information can be stored as a track reference. In this case, the reference type of the track reference can be, for example, indd (reference_type='indd'). By using this list of track references, the order of slice arrangement when reconstructing the G-PCC bitstream based on the slices stored in each track can be stored in the metadata area.
[0235] <2-3. Matroska Media Container>
[0236] Although an example of applying ISOBMFF as a file format has been described above, a file for storing a G-PCC bitstream is any file and may be a file other than ISOBMFF. For example, G-PCC content may be stored in a Matroska media container. Figure 18 The main configuration examples of the Matroska media container are shown in FIG.
[0237] In this case, for example, tile management information (tile identification information) may be stored as a newly defined element (element) under a track entry element (track entry element). In addition, when the tile management information (tile identification information) is stored in timed metadata, the timed metadata may be stored in a track entry different from the track entry in which the G-PCC content is stored.
[0238] <3. First embodiment>
[0239] <3-1. File Generation Device>
[0240] The encoding side device will be described. (Each method) of the present technology described above can be applied to any device. Figure 19 : is a block diagram showing an example of the configuration of a file generating apparatus as one type of information processing apparatus to which the present technology is applied. Figure 19 The file generation device 300 shown in FIG. 1 is a device that encodes point cloud data by applying G-PCC and stores G-PCC content (G-PCC bitstream) generated by the encoding in a content file (ISOBMFF).
[0241] At this time, the file generation device 300 applies the present technology described above in Section <2. Transmission of Scalable Decoding Information via Content File>. That is, the file generation device 300 generates scalable decoding information based on the slice depth and the dependency relationship between slices in the G-PCC content, generates a content file storing the G-PCC content, and stores the generated scalable decoding information in the metadata area of the generated content file.
[0242] Note that in Figure 19 In the figure, the main processing units, data flow, etc. are shown, and Figure 19 That is, in the file generating device 300, there may be Figure 19 A processing unit is not shown as a block, or may be present Figure 19 Not shown are processes or data flows as arrows or the like.
[0243] like Figure 19 As shown in FIG, the file generation device 300 includes an extraction unit 311, an encoding unit 312, a bitstream generation unit 313, a scalable decoding information generation unit 314, and a file generation unit 315. In addition, the encoding unit 312 includes a geometry encoding unit 321, an attribute encoding unit 322, and a metadata generation unit 323.
[0244] The extraction unit 311 extracts geometric body data and attribute data from the point cloud data input to the file generation device 300. The extraction unit 311 supplies the extracted geometric body data to the geometry encoding unit 321 of the encoding unit 312, and the extraction unit 311 supplies the extracted attribute data to the attribute encoding unit 322 of the encoding unit 312.
[0245] The encoding unit 312 encodes the point cloud data. The geometry encoding unit 321 encodes the geometry data supplied from the extraction unit 311 to generate a geometry bitstream. The geometry encoding unit 321 supplies the generated geometry bitstream to the metadata generation unit 323. In addition, the geometry encoding unit 321 also supplies the generated geometry bitstream to the attribute encoding unit 322.
[0246] The attribute encoding unit 322 encodes the data of the attribute supplied from the extraction unit 311 to generate an attribute bit stream. The attribute encoding unit 322 supplies the generated attribute bit stream to the metadata generation unit 323.
[0247] The metadata generation unit 323 generates metadata with reference to the supplied geometry bitstream and attribute bitstream.The metadata generation unit 323 supplies the generated metadata to the bitstream generation unit 313 together with the geometry bitstream and attribute bitstream.
[0248] The bitstream generation unit 313 multiplexes the supplied geometry bitstream, attribute bitstream, and metadata to generate G-PCC content (G-PCC bitstream). The bitstream generation unit 313 supplies the generated G-PCC content to the scalable decoding information generation unit 314.
[0249] The scalable decoding information generation unit 314 obtains the G-PCC content including the first slice and the second slice and supplied from the bitstream generation unit 313. The scalable decoding information generation unit 314 applies the present technology described above in Section <2. Transmission of scalable decoding information via content file> and generates scalable decoding information regarding scalable decoding of the G-PCC content based on depth information indicating the quality hierarchy level of the geometry included in each slice in the G-PCC content and the dependency relationship between the first slice and the second slice in the G-PCC content. The scalable decoding information generation unit 314 supplies the generated scalable decoding information to the file generation unit 315 together with the G-PCC content.
[0250] The file generation unit 315 applies the present technology described above in <2. Transmission of Scalable Decoding Information via Content File> to generate a content file storing the G-PCC content supplied from the scalable decoding information generation unit 314, and the scalable decoding information in the metadata area of the generated content file. The file generation unit 315 outputs the generated content file to the outside of the file generation device 300.
[0251] Note that the scalable decoding information generation unit 314 may generate scalable decoding information including slice configuration information for each sample. The file generation unit 315 may set subsamples for each slice and store the slice configuration information in codec-specific parameters of a subsample information box in the metadata area of the content file.
[0252] The scalable decoding information generation unit 314 may generate slice configuration information including slice dependency information. The file generation unit 315 may set a subsample for each slice and store the slice dependency information in a codec specific parameter of a subsample information box in the metadata area of the content file.
[0253] The scalable decoding information generation unit 314 may generate slice dependency information including reference source geometry slice identification information and reference destination geometry slice identification information. If the slice corresponding to the slice dependency information (i.e., the slice corresponding to the reference source geometry slice identification information) is an independent geometry slice, the reference source geometry slice identification information and the reference destination geometry slice identification information may be the same.
[0254] The file generation unit 315 may set a subsample for each slice and store the reference source geometry slice identification information and the reference destination geometry slice identification information in the codec specific parameters of the subsample information box in the metadata area.
[0255] The scalable decoding information generation unit 314 may generate slice dependency information including attribute geometry slice identification information. The file generation unit 315 may set subsamples for each slice and store the attribute geometry slice identification information in a codec specific parameter of a subsample information box in the metadata area.
[0256] The scalable decoding information generation unit 314 may generate slice dependency information including non-scalable coding attribute geometry slice identification information as identification information of a geometry slice referenced by an attribute slice to which non-scalable coding is applied.
[0257] The scalable decoding information generation unit 314 may generate non-scalable coded attribute geometry slice identification information, including identification information of a geometry slice or a geometry slice group having geometry containing maximum depth information among the geometry slices or geometry slice groups referenced by the attribute slice corresponding to the information. The file generation unit 315 may set a subsample for each slice and store the non-scalable coded attribute geometry slice identification information in a codec-specific parameter of a subsample information box in the metadata area of the content file.
[0258] In addition, the scalable decoding information generation unit 314 may generate slice dependency information including a non-scalable coding flag. In addition, the file generation unit 315 may set a subsample for each slice and store the non-scalable coding flag in the codec specific parameters of the subsample information box in the metadata area of the content file.
[0259] Note that the file generation unit 315 may set the payload type (PayloadType) of the slice dependency information of the independent geometry slice and the payload type of the slice dependency information of the dependent geometry slice to different values in the codec specific parameters.
[0260] The scalable decoding information generation unit 314 may generate slice composition information including geometry slice depth information. In this case, the scalable decoding information generation unit 314 may generate geometry slice depth information including minimum depth information. Furthermore, the scalable decoding information generation unit 314 may generate geometry slice depth information including maximum depth information. The file generation unit 315 may then set subsamples for each slice or slice group and store the geometry slice depth information in a codec-specific parameter of a subsample information box in the metadata area of the content file.
[0261] The scalable decoding information generation unit 314 may generate slice composition information further including slice dependency information, wherein the slice dependency information indicates a dependency relationship between a first slice and a second slice in addition to the geometry slice depth information. Then, the file generation unit 315 may set a flag of a subsample information box storing the geometry slice depth information to a different value from a flag of a subsample information box storing the slice dependency information.
[0262] The scalable decoding information generating unit 314 may generate scalable decoding information including track composition information. The scalable decoding information generating unit 314 may generate track composition information including track depth information.
[0263] The scalable decoding information generation unit 314 may generate track depth information including minimum track depth information, which indicates the minimum value of the depth information in the track corresponding to the information. Furthermore, the scalable decoding information generation unit 314 may generate track depth information including maximum track depth information, which indicates the maximum value of the depth information in the track corresponding to the information. Furthermore, the scalable decoding information generation unit 314 may generate track depth information including a match flag. The file generation unit 315 may then store such track depth information in the depth information box of the sample entry in the metadata area.
[0264] The scalable decoding information generation unit 314 may generate track composition information including track dependency information. Furthermore, the scalable decoding information generation unit 314 may generate track dependency information including dependency information indicating another track including slices required for decoding the dependent slices included in the track corresponding to the information. Furthermore, the scalable decoding information generation unit 314 may generate dependency information indicating all other tracks including slices required for decoding the dependent slices included in the track corresponding to the information. Furthermore, the scalable decoding information generation unit 314 may generate dependency information indicating another track including slices referenced by the information.
[0265] The scalable decoding information generation unit 314 may generate track dependency information including independent information indicating another track including a dependent slice for which an independent slice included in the track corresponding to the information is required for decoding. The scalable decoding information generation unit 314 may generate independent information indicating another track including another slice that references the independent slice during decoding. Furthermore, the scalable decoding information generation unit 314 may generate independent information in which the track dependency information includes both dependent information and independent information.
[0266] The file generation unit 315 may then store the track dependency information as a track reference in the metadata area.
[0267] In this way, as described above in the section <2. Transmission of scalable decoding information by content file>, an increase in the load of the reproduction process can be suppressed.
[0268] <Flow of File Generation Processing>
[0269] Will refer to Figure 20 An example of the flow of the file generation process executed by the file generation apparatus 300 is described with reference to a flowchart of FIG.
[0270] When the file generation process starts, the extraction unit 311 of the file generation device 300 extracts geometric bodies and attributes from the point cloud in step S301 .
[0271] In step S302, the encoding unit 312 encodes the geometry and attributes extracted in step S301 to generate a geometry bitstream and an attribute bitstream. The encoding unit 312 also generates metadata.
[0272] In step S303 , the bitstream generation unit 313 multiplexes the geometry bitstream, attribute bitstream, and metadata generated in step S302 to generate a G-PCC bitstream (G-PCC content).
[0273] In step S304, the scalable decoding information generation unit 314 applies the present technology described above in section <2. Transmission of scalable decoding information through content file> and generates scalable decoding information about the G-PCC content based on the depth of the slices in the G-PCC content generated in step S303 and the dependency between the slices in the G-PCC content.
[0274] In step S305, the file generation unit 315 generates other information and generates a content file (e.g., ISOBMFF) storing the G-PCC content generated in step S303. Then, the file generation unit 315 applies the present technology described above in section <2. Transmission of scalable decoding information via content file> and stores the scalable decoding information generated in step S304 in the metadata area of the generated content file.
[0275] In step S306, the file generation unit 315 outputs the generated content file (the content file storing the scalable decoding information) to the outside of the file generation device 300. For example, the file generation unit 315 transmits the content file to another device (e.g., a playback device) via a network or the like. Furthermore, for example, the file generation unit 315 supplies the content file and stores it in a storage medium external to the file generation device 300. In this case, the content file is supplied to the playback device or the like via the storage medium.
[0276] When the process of step S306 ends, the file generation process ends.
[0277] As described above, during the file generation process, the file generation device 300 applies the present technology described in Section 2. Transmission of Scalable Decoding Information via Content Files and stores the scalable decoding information in the metadata area of the content file. This reduces unnecessary information processing (decoding, etc.) and suppresses increases in the playback processing load.
[0278] <3-2. Playback Device>
[0279] Figure 21 : is a block diagram showing an example of the configuration of a reproduction device as one type of information processing device to which the present technology is applied. Figure 21 The reproduction device 400 shown in FIG is a device that decodes the G-PCC file, constructs a point cloud, and renders the point cloud to generate presentation information. At this time, the reproduction device 400 applies the present technology described above in Section <2. Transmission of Scalable Decoding Information via Content File> to extract, decode, and reproduce the slices necessary to reproduce the desired depth in the point cloud from the content file generated by the file generation device 300.
[0280] Note that in Figure 21 In the figure, the main processing units, data flow, etc. are shown, and Figure 21 That is, in the reproduction device 400, there may be Figure 21 A processing unit is not shown as a block, or may be present Figure 21 Not shown are processes or data flows as arrows or the like.
[0281] like Figure 21 As shown in FIG, the reproduction apparatus 400 includes a control unit 401, a file acquisition unit 411, a reproduction processing unit 412, and a presentation processing unit 413. The reproduction processing unit 412 includes a file processing unit 421, a decoding unit 422, and a presentation information generation unit 423.
[0282] The control unit 401 controls each processing unit in the reproduction device 400. The file acquisition unit 411 acquires a content file storing a point cloud to be reproduced and supplies the content file to the reproduction processing unit 412 (the file processing unit 421 thereof). The reproduction processing unit 412 performs processing related to the reproduction of the point cloud stored in the supplied content file.
[0283] The file processing unit 421 of the reproduction processing unit 412 obtains the content file supplied from the file acquisition unit 411 and extracts a bitstream from the content file. At this time, the file processing unit 421 applies the present technology described above in the section <2. Transmission of Scalable Decoding Information via Content File> and extracts only the bitstream of the slices necessary for reproduction at the desired depth. The file processing unit 421 supplies the extracted bitstream to the decoding unit 422.
[0284] The decoding unit 422 decodes the bitstream supplied from the file processing unit 421 to generate geometry and attribute data. The decoding unit 422 supplies the generated geometry and attribute data to the rendering information generation unit 423. The rendering information generation unit 423 constructs a point cloud using the supplied geometry and attribute data and generates rendering information as information for rendering (e.g., displaying) the point cloud. For example, the rendering information generation unit 423 performs rendering using the point cloud and generates a display image of the point cloud viewed from a predetermined viewpoint as rendering information. The rendering information generation unit 423 supplies the rendering information generated in this manner to the rendering processing unit 413.
[0285] The presentation processing unit 413 performs a process of presenting the supplied presentation information. For example, the presentation processing unit 413 supplies the presentation information to a display device or the like outside the reproduction device 400 to present the presentation information.
[0286] <Processing Unit>
[0287] Figure 22 4 is a block diagram showing a main configuration example of the reproduction processing unit 412. Figure 22 As shown in FIG, the file processing unit 421 includes a bit stream extraction unit 431. The decoding unit 422 includes a geometry decoding unit 441 and an attribute decoding unit 442. The rendering information generation unit 423 includes a point cloud construction unit 451 and a rendering processing unit 452.
[0288] The bitstream extraction unit 431 extracts a bitstream from the content file supplied by the file acquisition unit 411. At this time, the bitstream extraction unit 431 applies the present technology described above in the section <2. Transmission of scalable decoding information by content file>, and extracts only the bitstream of the slices necessary for reproduction at the desired depth. That is, the bitstream extraction unit 431 extracts any slice of the G-PCC content from the content file based on the scalable decoding information stored in the metadata area of the content file storing the G-PCC content. Note that this scalable decoding information is information about scalable decoding of the G-PCC content, and is information generated based on depth information indicating the quality hierarchy level of the geometry included in the slice in the G-PCC content and the dependency relationship between the slices in the G-PCC content (for example, between the first slice and the second slice).
[0289] Note that the scalable decoding information may include slice configuration information for each sample. For example, the bitstream extraction unit 431 may determine the configuration of slices in each sample based on the slice configuration information for each sample, and extract any slice of the G-PCC content from the content file based on the configuration. The bitstream extraction unit 431 may extract any slice of the G-PCC content from the content file based on the slice configuration information stored in the codec-specific parameters of the subsample information box in the metadata area of the content file for the subsample set for each slice.
[0290] The slice composition information may include slice dependency information. For example, the bitstream extraction unit 431 may determine the dependency between slices or slice groups based on the slice dependency information, and extract any slice of the G-PCC content from the content file based on the dependency. The bitstream extraction unit 431 may extract any slice of the G-PCC content from the content file based on the slice dependency information stored in the codec-specific parameters of the subsample information box in the metadata area of the content file for the subsample set for each slice.
[0291] The slice dependency information may include reference source geometry slice identification information and reference destination geometry slice identification information. For example, the bitstream extraction unit 431 may determine the dependency between slices or slice groups based on the reference source geometry slice identification information and the reference destination geometry slice identification information, and extract additional geometry slices required for decoding the desired geometry slice of the G-PCC content from the content file based on the dependency. In the case where the slice corresponding to the slice dependency information (i.e., the slice corresponding to the reference source geometry slice identification information) is an independent geometry slice, the above-mentioned reference source geometry slice identification information and the reference destination geometry slice identification information may be the same. The bitstream extraction unit 431 may extract any slice of the G-PCC content from the content file based on the reference source geometry slice identification information and the reference destination geometry slice identification information stored in the codec-specific parameters of the subsample information box in the metadata area of the subsample set for each slice.
[0292] The slice dependency information may include attribute geometry slice identification information. For example, the bitstream extraction unit 431 may identify the correspondence between the attribute slice and the geometry slice based on the attribute geometry slice identification information, and extract any slice (attribute slice and geometry slice) of the G-PCC content from the content file based on the identified correspondence. The bitstream extraction unit 431 may extract any slice of the G-PCC content from the content file based on the attribute geometry slice identification information stored in the codec-specific parameters of the subsample information box in the metadata area of the subsample set for each slice.
[0293] The slice dependency information may include non-scalable coded attribute geometry slice identification information, which is identification information of a geometry slice referenced by an attribute slice to which non-scalable coding is applied. For example, the bitstream extraction unit 431 may specify the correspondence between a geometry slice and an attribute slice to which non-scalable coding is applied based on the non-scalable coded attribute geometry slice identification information, and may extract any slice of the G-PCC content (geometry slices and attribute slices to which non-scalable coding is applied) from the content file based on the correspondence. The non-scalable coded attribute geometry slice identification information may include identification information of a geometry slice or a geometry slice group including a maximum depth in a geometry slice or a geometry slice group referenced by the attribute slice corresponding to the information. The bitstream extraction unit 431 may extract any slice of the G-PCC content from the content file based on the non-scalable coded attribute geometry slice identification information stored in the codec-specific parameters of the subsample information box in the metadata area of the content file of the subsample set for each slice.
[0294] The slice dependency information may include a non-scalable coding flag. For example, the bitstream extraction unit 431 may identify whether non-scalable coding is applied based on the non-scalable coding flag, identify the correspondence between attribute slices and geometry slices based on the identification result, and extract any slice (attribute slice and geometry slice) of the G-PCC content from the content file based on the correspondence. The bitstream extraction unit 431 may extract any slice of the G-PCC content from the content file based on the non-scalable coding flag stored in the codec-specific parameters of the subsample information box in the metadata area of the content file of the subsample set for each slice.
[0295] Note that in the codec specific parameters, the payload type (PayloadType) of the slice dependency information of the independent geometry slice and the payload type of the slice dependency information of the dependent geometry slice can have different values. For example, based on the payload type, the bitstream extraction unit 431 can identify whether it is the slice dependency information of the independent geometry slice or the slice dependency information of the dependent geometry slice, and analyze the slice dependency information based on the identification result.
[0296] The slice composition information may include geometry slice depth information. For example, the bitstream extraction unit 431 may determine the depth of the geometry included in the geometry slice or geometry slice group based on the geometry slice depth information, and may extract any slice of the G-PCC content from the content file based on the depth information. The geometry slice depth information may include minimum depth information. The geometry slice depth information may include maximum depth information. The bitstream extraction unit 431 may extract any slice of the G-PCC content from the content file based on the geometry slice depth information stored in the codec-specific parameters of the subsample information box in the metadata area of the content file of the subsample set for each slice or slice group.
[0297] The slice composition information may further include slice dependency information indicating dependencies between slices in addition to the geometry slice depth information. Then, the flag of the subsample information box storing the geometry slice depth information and the flag of the subsample information box storing the slice dependency information may be set to different values. For example, based on the flag, the bitstream extraction unit 431 may identify whether the box is a subsample information box storing the geometry slice depth information or a subsample information box storing the slice dependency information.
[0298] Scalable decoding information may include track composition information. For example, the bitstream extraction unit 431 may determine the track configuration based on the track composition information and extract any slice of G-PCC content from the content file based on the track configuration. The track composition information may also include track depth information. For example, the bitstream extraction unit 431 may determine the depth information of the geometry of all slices included in the track based on the track depth information, specify the track in which the desired slice is stored based on the depth information, and extract the desired slice from the track. The track depth information may include track minimum depth information indicating the minimum value of the depth information in the track corresponding to the information. The track depth information may also include track maximum depth information indicating the maximum value of the depth information in the track corresponding to the information. The track depth information may include a match flag. For example, the bitstream extraction unit 431 may determine the depth information of the geometry included in the track based on this information, specify the track in which the desired slice is stored based on the depth information, and extract the desired slice from the track. The bitstream extraction unit 431 may extract any slice of G-PCC content from the content file based on the track depth information stored in the depth information box of the sample entry in the metadata area.
[0299] Track composition information may include track dependency information. For example, the bitstream extraction unit 431 may determine track dependencies based on the track dependency information and extract any slices of G-PCC content from the content file based on the dependencies. The track dependency information may include dependency information indicating another track that includes slices necessary for decoding the dependent slices included in the track corresponding to the information. For example, based on the dependency information, the bitstream extraction unit 431 may specify another track that includes slices necessary for decoding the dependent slices included in the track corresponding to the information. The dependency information may indicate all other tracks that include slices necessary for decoding the dependent slices included in the track corresponding to the information. Furthermore, the dependency information may indicate another track that includes slices referenced by the information. Furthermore, the track dependency information may include independence information indicating another track that includes dependent slices necessary for decoding the independent slices included in the track corresponding to the information. For example, based on the independence information, the bitstream extraction unit 431 may specify another track that includes dependent slices necessary for decoding the independent slices included in the track corresponding to the information. The independence information may indicate another track including another slice that refers to the independent slice in decoding. In addition, track dependency information may be stored in the metadata area as a track reference.
[0300] The bitstream extraction unit 431 supplies the extracted geometry bitstream to the geometry decoding unit 441. Furthermore, the bitstream extraction unit 431 supplies the extracted attribute bitstream to the attribute decoding unit 442.
[0301] The geometry decoding unit 441 decodes the supplied geometry bitstream to generate geometry data. The geometry decoding unit 441 supplies the generated geometry data to the point cloud construction unit 451. The attribute decoding unit 442 decodes the supplied attribute bitstream to generate attribute data. The attribute decoding unit 442 supplies the generated attribute data to the point cloud construction unit 451.
[0302] The point cloud construction unit 451 constructs a point cloud using the supplied geometry and attribute data. That is, the point cloud construction unit 451 can construct a point cloud with a desired depth. The point cloud construction unit 451 supplies the constructed point cloud data to the rendering processing unit 452.
[0303] The rendering processing unit 452 generates rendering information using the supplied point cloud data, and supplies the generated rendering information to the rendering processing unit 413 .
[0304] With this configuration, the playback device 400 can more easily extract, decode, construct, and present only desired tiles based on the tile management information (tile identification information) stored in the G-PCC file without parsing the entire bitstream. Therefore, an increase in playback processing load can be suppressed.
[0305] <Flow of playback processing>
[0306] Will refer to Figure 23 An example of the flow of the reproduction process performed by the reproduction device 400 is described with reference to a flowchart of FIG.
[0307] When the reproduction process starts, the file acquisition unit 411 of the reproduction apparatus 400 acquires a content file to be reproduced in step S401 .
[0308] In step S402, the bitstream extraction unit 431 extracts any slices from the content file acquired in step S401. At this time, the bitstream extraction unit 431 applies the present technology described above in <2. Transmission of scalable decoding information via content file> and extracts the slices based on the scalable decoding information stored in the metadata area of the content file.
[0309] In step S403, the geometry decoding unit 441 of the decoding unit 422 decodes the geometry bitstream of the slice extracted in step S402 to generate the geometry of the desired depth. In addition, the attribute decoding unit 442 decodes the attribute bitstream of the slice extracted in step S402 and generates an attribute corresponding to the geometry of the desired depth.
[0310] In step S404, the point cloud construction unit 451 constructs a point cloud using the geometry and attributes generated in step S403. That is, the point cloud construction unit 451 can construct a point cloud with a desired depth.
[0311] In step S405 , the rendering processing unit 452 generates rendering information by performing rendering using the point cloud constructed in step S404 , etc. In step S406 , the rendering processing unit 413 supplies the rendering information and renders it to the outside of the reproduction device 400 .
[0312] When the process of step S406 ends, the reproduction process ends.
[0313] As described above, during the playback process, the playback device 400 applies the present technology described in Section <2. Transmission of Scalable Decoding Information via Content File> and extracts and decodes the desired slice from the content file based on the scalable decoding information stored in the metadata area of the content file. This reduces unnecessary information processing (decoding, etc.) and suppresses increases in the playback processing load.
[0314] <4. Transmission of Scalable Decoding Information via Control Files>
[0315] This technology can also be applied to, for example, Moving Picture Experts Group Dynamic Adaptive Streaming over HTTP (MPEG-DASH). For example, MPEG-DASH can store scalable decoding information by extending the Media Presentation Description (MPD), which is a control file that stores control information about the delivery of bitstreams. For example, as scalable decoding information, adaptation set configuration information, which is information about the configuration of adaptation sets that describe information about tracks of content files, can be stored in the MPD.
[0316] That is to say, if Figure 24 As described at the top of the table shown in , adaptation set composition information of the depth of each slice based on the G-PCC content having the dependency relationship between slices and the slice structure stored in the control file is transmitted (Method 2).
[0317] For example, the information processing device includes: an adaptation set composition information generation unit that generates adaptation set composition information based on depth information indicating the quality hierarchy level of a geometry included in each slice of G-PCC content including first and second slices and a dependency relationship between the first slice and the second slice in the G-PCC content; and a control file generation unit that generates a control file for controlling the reproduction of a content file storing the G-PCC content and stores the adaptation set composition information in the control file. The content file then stores the G-PCC content in tracks on a slice basis. In addition, the adaptation set composition information is information about the construction of an adaptation set that describes information about the track of the content file.
[0318] For example, the information processing method includes: generating adaptation set composition information based on depth information indicating the quality hierarchy level of a geometry included in each slice of G-PCC content including a first slice and a second slice, and a dependency relationship between the first slice and the second slice in the G-PCC content; generating a control file for controlling the reproduction of a content file storing the G-PCC content; and storing the adaptation set composition information in the control file. The content file then stores the G-PCC content in tracks based on the slices. Furthermore, the adaptation set composition information is information about the configuration of the adaptation set that describes information about the tracks of the content file.
[0319] For example, an information processing device includes: an analysis unit that analyzes a control file for controlling the reproduction of a content file storing G-PCC content including a first slice and a second slice in a track on a slice basis and specifies an adaptation set necessary for obtaining any slice of the G-PCC content based on adaptation set composition information stored in the control file; an acquisition unit that acquires a track of the content file corresponding to the adaptation set specified by the analysis unit; and a decoding unit that decodes the slice of the G-PCC content stored in the track acquired by the acquisition unit. Then, the adaptation set composition information is information about the configuration of the adaptation set that describes information about the track of the content file, and is information generated based on depth information indicating the quality hierarchy level of a geometry included in the slice in the G-PCC content and a dependency relationship between the first slice and the second slice in the G-PCC content.
[0320] For example, the information processing method includes: analyzing a control file for controlling the reproduction of a content file, the content file storing G-PCC content in tracks based on slices; specifying an adaptation set required to obtain any slice of the G-PCC content based on adaptation set composition information stored in the control file; acquiring a track corresponding to the specified adaptation set of the content file; and decoding a slice of the G-PCC content stored in the acquired track. Then, the adaptation set composition information is information about the configuration of the adaptation set that describes information about the track of the content file, and is information generated based on depth information indicating the quality hierarchy level of geometry included in the slice in the G-PCC content and a dependency relationship between a first slice and a second slice in the G-PCC content.
[0321] In this way, the decoder can select the track that stores the slices necessary to reproduce the point cloud of the desired depth or area based on the adaptation set composition information stored in the control file, and can then retrieve the track. This can suppress the transmission of unnecessary data. Consequently, it can suppress increases in the load on the transmission path and communication processing. Furthermore, since the decoder can suppress increases in the amount of data to be processed, it can also suppress increases in the load on the reproduction process. Consequently, it can suppress increases in delays associated with data transmission and reproduction processing.
[0322] As described above in the section "<2. Transmission of scalable decoding information through content files>", especially in the case of large-scale point clouds, by applying this technology, the increase in the load of data transmission and reproduction processing can be further suppressed and greater effects can be obtained.
[0323] <4-1. Adaptation Set Depth Information>
[0324] like Figure 24 As shown in the second row from the top of the table shown in , the adaptation set composition information may include adaptation set depth information (method 2-1). The adaptation set depth information is information about the depth information of the geometry of all slices included in the track corresponding to the adaptation set corresponding to the information. The adaptation set depth information is information similar to the track depth information described in <2-2-1. Track depth information>. That is, the adaptation set depth information is obtained by applying the track depth information to the control file (MPD).
[0325] The decoder can easily determine which slice at which depth corresponds to which adaptation set (i.e., track) based on the adaptation set depth information stored in the MPD (without parsing the bitstream). Therefore, the decoder can more easily select the track to obtain (without parsing the bitstream).
[0326] Note that in the case of MPD, a supplementary property or an essential property (Essential Property) may be newly defined, schemeIdUri="urn:mpeg:mpegI:gpcc:2020:depth" may be set, and adaptation set depth information may be stored therein.
[0327] like Figure 24 As shown in the third row from the top of the table shown in , the adaptation set depth information may include adaptation set minimum depth information (Method 2-1-1). The adaptation set minimum depth information is information indicating a minimum value of depth information in a track corresponding to the adaptation set corresponding to the information.
[0328] Figure 25 is a diagram showing an example of parameters to be added to MPD as adaptation set depth information. Figure 25 The @minDepth shown in is the adaptation set minimum depth information, and is information similar to track_min_depth (track minimum depth information) of the track depth information described in <2-2-1. Track depth information>. That is, @minDepth is obtained by applying track_min_depth to the control file (MPD).
[0329] like Figure 24 As shown in the fourth row from the top of the table shown in , the adaptation set depth information may include adaptation set maximum depth information (Method 2-1-2). The adaptation set maximum depth information is information indicating a maximum value of depth information in a track corresponding to the adaptation set corresponding to the information.
[0330] Figure 25 The @maxDepth shown in is the adaptation set maximum depth information, and is information similar to track_max_depth (track maximum depth information) of the track depth information described in <2-2-1. Track depth information>. That is, @maxDepth is obtained by applying track_max_depth to the control file (MPD).
[0331] like Figure 24 As shown in the fifth row from the top of the table shown in , the adaptation set depth information may include a matching flag (method 2-1-3). The matching flag is flag information indicating whether the sample minimum depth information of each sample included in the track corresponding to the adaptation set corresponding to the information matches the adaptation set minimum depth information, and whether the sample maximum depth information of each sample included in the track matches the adaptation set maximum depth information (flag information indicating whether the minimum depth and maximum depth are common to all samples). Note that the sample minimum depth information indicates the minimum value of the depth information in the sample. The sample maximum depth information indicates the maximum value of the depth information in the sample.
[0332] Figure 25 The @fixedDepth shown in the figure is a match flag for the adaptation set. When @fixedDepth is "0" (false), it indicates that the sample minimum depth information and sample maximum depth information of all samples in the track corresponding to the adaptation set take values within the range of the adaptation set minimum depth information (@minDepth) to the adaptation set maximum depth information (@maxDepth). In other words, in this case, the minimum or maximum value of the depth of each sample, or both, can be changed for each sample.
[0333] On the other hand, when @fixedDepth is "1" (true), it indicates that the sample minimum depth information of each sample in the track matches the track minimum depth information (track_min_depth), and the sample maximum depth information of each sample in the track matches the track maximum depth information (track_min_depth). That is, in this case, the minimum and maximum values of the depth of each sample take the common values in all samples.
[0334] For example, when the track includes data of only one depth, the value of the matching flag (@fixedDepth) is set to "1", and the adaptation set minimum depth information (@minDepth) and the adaptation set maximum depth information (@maxDepth) take the same value (@minDepth=@maxDepth).
[0335] In the case of MPD, these parameters (@fixedDepth, @minDepth, @maxDepth) of the adaptation set depth information may be set in the aforementioned supplementary attributes or essential properties.
[0336] As described above, by storing the adaptation set depth information in the control file (MPD), the decoder can more easily determine which depth data is stored in each track based on this information (without parsing the bitstream). Therefore, the decoder can more easily select the track to be acquired (without parsing the bitstream).
[0337] <4-2. Presenting Dependency Information>
[0338] like Figure 24 As shown in the sixth row from the top of the table shown in , the adaptation set composition information may include presentation dependency information (method 2-2). Presentation dependency information is information indicating the dependency between presentations (for example, the dependency between the first presentation and the second presentation). In the MPD, tracks are managed by presentations of the adaptation set. Then, the presentation dependency information is information similar to the track dependency information described in <2-2-2. Track dependency information>. That is, the presentation dependency information is obtained by applying the track dependency information to the control file (MPD).
[0339] The decoder can more easily determine the dependencies between tracks based on the presentation dependency information stored in the MPD (without parsing the bitstream). Therefore, the decoder can more easily select the track to obtain (without parsing the bitstream).
[0340] like Figure 24 As shown in the seventh row from the top of the table shown in , the presentation dependency information may include dependent information indicating an additional presentation necessary to decode the presentation corresponding to the information (method 2-2-1).
[0341] The dependency information is information to be associated from a presentation that does not include an independent slice (depth=0) with all presentations that include slices necessary for decoding the presentation.
[0342] By storing this dependency information in the control file (MPD), the decoder can more easily identify other presentations (i.e., tracks) required to decode the presentation (i.e., track) being processed based on the dependency information without parsing the bitstream. This can suppress increases in the playback load.
[0343] Note that Figure 24 As shown in the eighth row from the top of the table shown in , the dependent information can indicate all other presentations necessary to decode the presentation corresponding to the information (method 2-2-1-1).
[0344] In addition, if Figure 24 As shown in the ninth row from the top of the table shown in , the dependent information may indicate another presentation to be referenced by the presentation corresponding to the information (method 2-2-1-2). That is, the dependent information may indicate a presentation directly referenced by the presentation of the reference source.
[0345] In addition, if Figure 24 As shown in the tenth row from the top of the table shown in , the presentation dependency information may include independent information indicating that another presentation corresponding to the information is required for decoding (method 2-2-2). That is, the independent information is information associated with a presentation including an independent slice (depth = 0) and a presentation including a dependent slice for decoding. In other words, the independent information is reverse lookup information for the dependent information.
[0346] like Figure 24 As shown in the eleventh row from the top of the table shown in , the independent information can indicate all other presentations in which the presentation corresponding to the information is necessary in decoding (method 2-2-2-1).
[0347] In addition, if Figure 24 As shown at the bottom of the table shown in , the presentation dependency information can be stored in the control file (MPD) as two parameters of presentation association identification information (Representation@associationId) and association type (associationType) (Method 2-2-3). That is, the two parameters of presentation association identification information (Representation@associationId) and association type (associationType) can be used to perform association between dependent information and independent information.
[0348] In the case of dependent information, the association type may be set to "depd" (associationType="depd"). Also, in the case of independent information, the association type may be set to "indd" (associationType="indd").
[0349] Note that flag information (@nonScalableAttributeFlag) clearly indicating whether an adaptation set includes an attribute slice to which non-scalable coding is applied may be stored in a control file (MPD).
[0350] <4-3. Description Example>
[0351] Figure 26 An example of the description of MPD is shown in FIG. Figure 26 In the MPD 520 shown in FIG, in the underlined line 521, the basic property (EssentialProperty) is defined for the adaptation set (AdaptationSet id="1") whose identification information is "1" and schemeIdUri="urn:mpeg:mpegI:gpcc:2020:depth" is set. Then, the adaptation set depth information such as @fixedDepth, @minDepth and @maxDepth is set.
[0352] In addition, Figure 26 In the MPD 520 shown in FIG, in the line with the underline 522, the basic property (Essential Property) is defined for the adaptation set (AdaptationSet id="2") whose identification information is "2" and schemeIdUri="urn:mpeg:mpegI:gpcc:2020:depth" is set. Then, the adaptation set depth information such as @fixedDepth, @minDepth, and @maxDepth is set. In addition, in the next line, the presentation association identification information (Representation@associationId) and the presentation dependency information such as the association type (associationType) are set. Since the presentation dependency information is subordinate information, the association type (associationType) is set to "depd".
[0353] In addition, Figure 26 In the MPD 520 shown in FIG, in the line with the underline 523, the basic property (Essential Property) is defined for the adaptation set (AdaptationSet id="3") whose identification information is "3" and schemeIdUri="urn:mpeg:mpegI:gpcc:2020:depth" is set. Then, the adaptation set depth information such as @fixedDepth, @minDepth and @maxDepth is set. In addition, in the next line, the presentation association identification information (Representation@associationId) and the presentation dependency information such as the association type (associationType) are set. Since the presentation dependency information is subordinate information, the association type (associationType) is set to "depd".
[0354] As described above, since the adaptation set depth information and presentation dependency information are stored in the MPD, the decoder can more easily confirm the presentation (i.e., track) including the desired slice based on the information without parsing the bitstream. Therefore, the increase in the playback load can be suppressed.
[0355] Note that one or both of the present techniques described in (<4. Transmission of scalable decoding information via control file>) of this section - that is, storing scalable decoding information (adaptation set composition information) in the MPD, and the present technique described in <2. Transmission of scalable decoding information via content file> of this section - that is, storing scalable decoding information (at least slice composition information for each sample) in the metadata area of the content file can be applied.
[0356] <5. Second embodiment>
[0357] <5-1. File Generation Device>
[0358] The above-described (each method) of the present technology can be applied to any device. Figure 27 1 is a block diagram showing an example of the configuration of a file generating apparatus as one type of information processing apparatus to which the present technology is applied. As in the file generating apparatus 300, Figure 27 The file generation device 600 shown in FIG is a device that encodes point cloud data by applying G-PCC and stores the G-PCC content (G-PCC bitstream) generated by the encoding in a content file (ISOBMFF). However, the file generation device 600 also generates an MPD corresponding to the content file.
[0359] At this point, the file generation device 600 can apply the present technology described above in sections <2. Transmission of Scalable Decoding Information via Content Files> or <4. Transmission of Scalable Decoding Information via Control Files>. That is, the file generation device 600 can generate scalable decoding information based on the depth of slices in the G-PCC content and the dependencies between slices, generate a content file storing the G-PCC content, and store at least slice composition information for each sample in the generated scalable decoding information in the metadata area of the generated content file. Furthermore, the file generation device 600 can store the adaptation set composition information in the generated scalable decoding information in the MPD.
[0360] Note that in Figure 27 In the figure, the main processing units, data flow, etc. are shown, and Figure 27 That is, in the file generating device 600, there may be Figure 27 A processing unit is not shown as a block, or may be present Figure 27 Not shown are processes or data flows as arrows or the like.
[0361] like Figure 27 As shown in FIG, the file generating apparatus 600 has a configuration substantially similar to that of the file generating apparatus 300 ( Figure 19 ) configuration. However, instead of the file generating unit 315, the file generating device 600 includes a content file generating unit 615 and an MPD generating unit 616.
[0362] In this case as well, the scalable decoding information generation unit 314 applies the present technique described above in <2. Transmission of Scalable Decoding Information via Content File> to generate scalable decoding information (at least slice composition information for each sample). That is, the scalable decoding information generation unit 314 generates scalable decoding information (at least slice composition information for each sample) based on depth information indicating the quality hierarchy level of the geometry included in each slice of the G-PCC content including the first and second slices, and the dependency relationship between the first and second slices in the G-PCC content. Furthermore, in this case, the scalable decoding information generation unit 314 applies the present technique described above in <4. Transmission of Scalable Decoding Information via Control File> and generates adaptation set composition information as scalable decoding information rather than track composition information. That is, the scalable decoding information generation unit 314, as an adaptation set composition information generation unit, generates adaptation set composition information based on depth information indicating the quality hierarchy level of the geometry included in each slice of the G-PCC content including the first and second slices, and the dependency relationship between the first and second slices in the G-PCC content.
[0363] The scalable decoding information generating unit 314 supplies the generated scalable decoding information together with the G-PCC content to the content file generating unit 615. Furthermore, the scalable decoding information generating unit 314 supplies the generated scalable decoding information to the MPD generating unit 616 together with the G-PCC content.
[0364] The content file generation unit 615 applies the present technology described above in <2. Transmission of Scalable Decoding Information via Content File> to store the content file storing the supplied G-PCC content on a slice-by-slice basis in a track, and stores the scalable decoding information (at least slice configuration information for each sample) in the metadata area of the generated content file. The content file generation unit 615 outputs the content file generated as described above to the outside of the file generation device 300.
[0365] The MPD generation unit 616, as a control file generation unit, generates an MPD and stores information about the supplied G-PCC content in the MPD. Furthermore, the MPD generation unit 616 applies the present technology described above in Section <4. Transmission of Scalable Decoding Information via a Control File> and stores the supplied scalable decoding information (adaptation set configuration information) in the MPD. The MPD generation unit 616 outputs the MPD generated as described above to an external portion of the file generation device 300 (e.g., a content file distribution server, etc.).
[0366] The adaptation set composition information may include the adaptation set depth information. That is, the scalable decoding information generation unit 314 may generate the adaptation set composition information including the adaptation set depth information, and the MPD generation unit 616 may store the adaptation set composition information in the adaptation set of the MPD. In addition, the adaptation set depth information may include the adaptation set minimum depth information. In addition, the adaptation set depth information may include the adaptation set maximum depth information. In addition, the adaptation set depth information may include a matching flag. In other words, the scalable decoding information generation unit 314 may generate the adaptation set depth information including the information, and the MPD generation unit 616 may store the adaptation set depth information in the adaptation set of the MPD. For example, the MPD generation unit 616 may newly define supplementary attributes or basic attributes and store the adaptation set depth information therein.
[0367] In addition, the adaptation set composition information may include presentation dependency information. That is, the scalable decoding information generation unit 314 may generate adaptation set composition information including presentation dependency information, and the MPD generation unit 616 may store the adaptation set composition information in the adaptation set of the MPD. In addition, the presentation dependency information may include dependent information indicating additional presentations required for decoding the presentation corresponding to the information. In addition, the dependent information may indicate all other presentations required for decoding the presentation corresponding to the information. In addition, the dependent information may indicate additional presentations referenced from the presentation corresponding to the information. That is, the scalable decoding information generation unit 314 may generate presentation dependency information including such dependent information, and the MPD generation unit 616 may store the presentation dependency information in the adaptation set of the MPD.
[0368] Furthermore, the presentation dependency information may include independent information indicating that another presentation is required for decoding the presentation corresponding to the information. The independent information may then indicate that all other presentations corresponding to the presentation of the information are required for decoding. That is, the scalable decoding information generation unit 314 may generate presentation dependency information including such independent information, and the MPD generation unit 616 may store the presentation dependency information in an adaptation set of the MPD.
[0369] The presentation dependency information may be stored in the control file (MPD) as two parameters of presentation association identification information (Representation@associationId) and association type (associationType).
[0370] In this way, it is possible to suppress an increase in the load of the reproduction process as described above in the sections <2. Transmission of scalable decoding information by content file> and <4. Transmission of scalable decoding information by control file>.
[0371] <Flow of File Generation Processing>
[0372] Will refer to Figure 28 An example of the flow of the file generation process executed by the file generation device 600 is described with reference to a flowchart of FIG.
[0373] When the file generation process starts, each process of steps S601 to S603 is the same as Figure 20 Each process of steps S301 to S303 in the flowchart of the file generation process is similarly performed.
[0374] In step S604, the scalable decoding information generation unit 314 applies the present technology described in Section <2. Transmission of Scalable Decoding Information via Content File> to generate slice configuration information for each sample as scalable decoding information. Furthermore, the scalable decoding information generation unit 314 applies the present technology described in Section <4. Transmission of Scalable Decoding Information via Control File> to generate adaptation set configuration information as scalable decoding information.
[0375] In step S605, the content file generation unit 615 applies the present technology described in Section 2. Transmission of Scalable Decoding Information via Content Files. Specifically, the content file generation unit 615 generates a content file and stores the G-PCC content in a track of the content file on a slice-by-slice basis. The content file generation unit 615 then stores the slice configuration information for each sample in the metadata area of the content file.
[0376] In step S606, the content file generation unit 615 outputs the generated content file (the content file storing the scalable decoding information) to the outside of the file generation device 600. For example, the content file generation unit 615 transmits the content file to another device (e.g., a playback device) via a network or the like. Furthermore, for example, the content file generation unit 615 supplies the content file to a storage medium outside the file generation device 600 and stores the content file. In this case, the content file is supplied to the playback device or the like via the storage medium.
[0377] In step S607, the MPD generation unit 616 applies the present technology described in section <4. Transmission of scalable decoding information by control file>. That is, the MPD generation unit 616 generates an MPD corresponding to the content file generated in step S605 and stores the adaptation set configuration information generated in step S604 in the MPD.
[0378] In step S608, the MPD generation unit 616 outputs the MPD to the outside of the file generation device 600. For example, the MPD is provided to a content file distribution server or the like.
[0379] When the process of step S608 ends, the file generation process ends.
[0380] In this way, in the file generation process, the file generation device 600 applies the present technology described in Section <2. Transmission of Scalable Decoding Information via Content File> or <4. Transmission of Scalable Decoding Information via Control File> and stores the scalable decoding information in the metadata area of the content file or the MPD. In this way, the transmission and processing (decoding, etc.) of unnecessary information can be reduced, and the increase in the load of data transmission and reproduction processing can be suppressed.
[0381] <5-2. Playback Device>
[0382] Figure 29 400 is a block diagram showing an example of the configuration of a reproduction device as one type of information processing device to which the present technology is applied. Figure 29 The reproduction device 700 shown in FIG is a device that decodes a content file, constructs a point cloud, and renders the point cloud to generate presentation information. In this case, the reproduction device 700 can apply the present technology described above in Sections <2. Transmission of Scalable Decoding Information via Content Files> or <4. Transmission of Scalable Decoding Information via Control Files>.
[0383] Note that in Figure 29 In the figure, the main processing units, data flow, etc. are shown, and Figure 29 That is, in the reproduction device 700, there may be Figure 29 A processing unit is not shown as a block, or may be present Figure 29 Not shown are processes or data flows as arrows or the like.
[0384] like Figure 29 As shown in FIG, the reproduction apparatus 700 basically has the same functions as the reproduction apparatus 400 (see FIG. Figure 21 ) has the same configuration. However, instead of the file acquisition unit 411, the reproduction device 700 includes a file acquisition unit 711 and an MPD analysis unit 712.
[0385] The file acquisition unit 711 acquires the MPD corresponding to the desired content file (content file to be reproduced) and supplies the MPD to the MPD analysis unit 712. In addition, the file acquisition unit 711 requests and acquires the track requested from the MPD analysis unit 712 among the tracks of the content file from the supply source of the content file to be reproduced. The file acquisition unit 711 supplies the acquired track (the bit stream stored in the track) to the reproduction processing unit 412 (file processing unit 421).
[0386] After the MPD is acquired from the file acquisition unit 711, the MPD analysis unit 712 analyzes the MPD and selects the desired track. At this time, the MPD analysis unit 712 can apply the present technology described above in Section <4. Transmission of Scalable Decoding Information by Control File>. That is, the MPD analysis unit 712 specifies the adaptation set (i.e., track) required to obtain any slice of the G-PCC content based on the adaptation set configuration information stored in the adaptation set of the MPD. The MPD analysis unit 712 requests the file acquisition unit 711 to acquire the track corresponding to the specified adaptation set.
[0387] The adaptation set composition information may include adaptation set depth information. For example, the MPD analysis unit 712 may determine the depth information of the geometry of all slices included in the track based on the adaptation set depth information stored in the adaptation set of the MPD, and specify the adaptation set (i.e., track) required to obtain any slice based on the depth information. In addition, the adaptation set depth information may include adaptation set minimum depth information. In addition, the adaptation set depth information may include adaptation set maximum depth information. In addition, the adaptation set depth information may include a matching flag. For example, the MPD analysis unit 712 may determine the depth information of the geometry of all slices included in the track based on the information stored in the adaptation set of the MPD, and specify the adaptation set (i.e., track) required to obtain any slice based on the depth information.
[0388] In addition, the adaptation set composition information may include presentation dependency information. For example, the MPD analysis unit 712 may determine the dependency between presentations (i.e., tracks) based on the presentation dependency information stored in the adaptation set of the MPD, and specify the presentation (i.e., track) required to obtain any slice based on the dependency. The presentation dependency information may include dependent information indicating an additional presentation required to decode the presentation corresponding to the information. In addition, the dependent information may indicate all other presentations required to decode the presentation corresponding to the information. In addition, the dependent information may indicate an additional presentation referenced by the presentation corresponding to the information. For example, the MPD analysis unit 712 may specify an additional presentation (i.e., an additional track) required to decode the presentation corresponding to the information based on the dependent information stored in the adaptation set of the MPD.
[0389] In addition, the presentation dependency information may include independent information indicating that another presentation is required for decoding the presentation corresponding to the information. The independent information may then indicate that all other presentations corresponding to the presentation of the information are required for decoding. For example, the MPD parsing unit 712 may specify another presentation that is required for decoding the presentation corresponding to the information based on dependency information stored in the adaptation set of the MPD.
[0390] The presentation dependency information can be stored in the control file (MPD) as two parameters: presentation association identification information (Representation@associationId) and association type (associationType). That is, the MPD analysis unit 712 can refer to these parameters stored in the adaptation set of the MPD to determine the dependency between presentations (i.e., tracks) based on these parameters, and specify the presentation (i.e., track) required to obtain any slice based on the dependency.
[0391] In this way, it is possible to suppress an increase in the load of the reproduction process as described above in the sections <2. Transmission of scalable decoding information by content file> and <4. Transmission of scalable decoding information by control file>.
[0392] The decoding unit 422 applies the present technology described above in section <4. Transmission of scalable decoding information by control file> and decodes the slices of the G-PCC content stored in the track supplied from the file acquisition unit 711 .
[0393] <Flow of playback processing>
[0394] Will refer to Figure 30 An example of the flow of the reproduction process performed by the reproduction device 700 is described with reference to a flowchart of FIG.
[0395] When the reproduction process starts, the file acquisition unit 711 of the reproduction device 700 acquires the MPD corresponding to the content file to be reproduced in step S701.
[0396] In step S702, the MPD analyzing unit 712 specifies an adaptation set necessary to obtain G-PCC content having a desired depth based on the adaptation set configuration information stored in the MPD.
[0397] In step S703 , the file acquisition unit 711 acquires the encoded data of the content file to be reproduced, which is stored in the track corresponding to the adaptation set specified in step S702 .
[0398] In step S704 , the file processing unit 421 acquires, from the content file, the encoded data of the slices necessary to obtain the G-PCC content having the desired depth from the acquired encoded data, based on the slice configuration information of each sample.
[0399] The processing of steps S705 to S708 is as follows Figure 23 The reproduction process of steps S403 to S406 is executed. When the process of step S708 ends, the reproduction process ends.
[0400] As described above, in the reproduction process, the reproduction device 700 applies the present technology described in Section <2. Transmission of Scalable Decoding Information via Content File> or <4. Transmission of Scalable Decoding Information via Control File> and acquires and decodes the desired track of the content file based on the adaptation set configuration information stored in the MPD or the slice configuration information for each sample stored in the metadata area of the content file. In this way, the processing of unnecessary information (data transmission, decoding, etc.) can be reduced, and the increase in the reproduction processing load can be suppressed.
[0401] <6. Supplement>
[0402] <Computer>
[0403] The above series of processes can be performed by hardware or software. In the case of performing a series of processes by software, the software program is installed in a computer. Here, the computer is, for example, a computer incorporated into dedicated hardware, a general-purpose personal computer capable of performing various functions by installing various programs, etc.
[0404] Figure 31 : is a block diagram showing a configuration example of a computer that executes the above-described series of processes according to a program.
[0405] exist Figure 31 In a computer 900 shown in FIG, a central processing unit (CPU) 901 , a read only memory (ROM) 902 , and a random access memory (RAM) 903 are connected to one another via a bus 904 .
[0406] An input / output interface 910 is also connected to the bus 904. An input unit 911, an output unit 912, a storage unit 913, a communication unit 914, and a drive 915 are connected to the input / output interface 910.
[0407] The input unit 911 includes, for example, a keyboard, a mouse, a microphone, a touch panel, and input terminals. The output unit 912 includes, for example, a display, a speaker, and output terminals. The storage unit 913 includes, for example, a hard disk, a RAM disk, and a nonvolatile memory. The communication unit 914 includes, for example, a network interface. The drive 915 drives a removable medium 921 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0408] In the computer configured as described above, for example, the CPU 901 executes the above series of processes by loading a program stored in the storage unit 913 to the RAM 903 via the input / output interface 910 and the bus 904 and executing the program. The RAM 903 also appropriately stores data and the like necessary for the CPU 901 to execute various processes.
[0409] For example, the program executed by the computer can be applied by being recorded on the removable medium 921 as a package medium, etc. In this case, by attaching the removable medium 921 to the drive 915 , the program can be installed in the storage unit 913 via the input / output interface 910 .
[0410] In addition, the program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting. In this case, the program can be received by the communication unit 914 and can be installed in the storage unit 913.
[0411] In addition, the program may be installed in the ROM 902 or the storage unit 913 in advance.
[0412] <Targets for which this technology is applicable>
[0413] As described above, the case where G-PCC content with a slice structure is stored in ISOBMFF has been described as an example, but the case where the present technology can be applied is not limited to this example. The present technology can be applied to any technology as long as it is incompatible with the above-mentioned present technology. For example, although point cloud data has been described as the encoding target in the example, any standard 3D data can be set as the encoding target. In addition, although G-PCC has been described as an example of an encoding or decoding method, any encoding or decoding method can be applied as long as it is a method that can generate encoded data with a slice structure (a method corresponding to scalable decoding). In addition, ISOBMFF has been described as an example of a file format with which G-PCC content is stored, but any file format can be applied as long as scalable decoding information can be stored. In addition, as long as there is no incompatibility with the present technology, some of the above-mentioned processing and specifications may be omitted, or may be combined with technologies not described above.
[0414] Furthermore, the present technology can be applied to any configuration. For example, the present technology can be applied to various electronic devices.
[0415] In addition, for example, the present technology can also be implemented as a partial configuration of a device, such as a processor as a system large-scale integration (LSI), a module using multiple processors, a unit using multiple modules, or a collection in which other functions are added to the unit.
[0416] Furthermore, for example, the present technology can also be applied to a network system comprising multiple devices. For example, the present technology can be implemented as cloud computing in which multiple devices share and collaborate on processing via a network. For example, the present technology can be implemented in a cloud service that provides image (moving image) related services to any terminal such as a computer, audio-visual (AV) device, portable information processing terminal, or Internet of Things (IoT) device.
[0417] Note that in this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), and it does not matter whether all components are in the same housing. Therefore, multiple devices housed in different housings and connected via a network, as well as a single device in which multiple modules are housed in a single housing, are both systems.
[0418] <Fields and applications of this technology>
[0419] The systems, devices, processing units, etc. using this technology can be used in any field such as transportation, medical care, crime prevention, agriculture, animal husbandry, mining, beauty, factories, home appliances, weather and nature monitoring. In addition, its application is also arbitrary.
[0420] For example, the present technology can be applied to systems or devices that provide content for observation, etc. Furthermore, for example, the present technology can also be applied to systems and devices that provide for traffic, such as traffic condition monitoring and autonomous driving control. Furthermore, for example, the present technology can also be applied to systems or devices that provide for safety. Furthermore, for example, the present technology can also be applied to systems or devices that provide for automatic control of machinery, etc. Furthermore, for example, the present technology can also be applied to systems and devices used for agriculture and animal husbandry. Furthermore, the present technology can also be applied to systems and devices that monitor natural conditions such as volcanoes, forests and oceans, wildlife, etc. Furthermore, for example, the present technology can also be applied to systems and devices that provide for sports.
[0421] <Other>
[0422] Note that in this specification, a "flag" is information with which a plurality of states are identified, and includes not only information for identifying two states of true (1) and false (0), but also information with which three or more states can be identified. Therefore, the value that a "flag" can take can be, for example, binary 1 and 0 or ternary or more. That is, the number of bits included in the "flag" is any number, and can be one or more bits. In addition, since it is assumed that identification information (including a flag) includes not only identification information in a bit stream, but also different information of identification information relative to a certain reference information in the bit stream, in this specification, "flag" and "identification information" include not only information but also different information relative to the reference information.
[0423] In addition, various types of information (metadata, etc.) related to the encoded data (bitstream) can be transmitted or recorded in any form, as long as the information is associated with the encoded data. Here, the term "association" means that, for example, one data can be used (linked) when processing other data. That is, data associated with each other can be collected as one data or can be separate data. For example, second data associated with first data can be transmitted on a transmission path different from the transmission path of the first data. In addition, for example, the second data associated with the first data can be recorded on a recording medium different from the first data (or another recording area of the same recording medium). Note that this "association" can be a part of the data, not the entire data. For example, 3D data and metadata corresponding to 3D data can be associated with each other on any basis such as multiple samples, one sample, or a part of a sample.
[0424] Note that in this specification, terms such as "combine", "multiplex", "add", "integrate", "include", "store", "advance", "place", "insert", etc. mean putting multiple objects into one, and mean a method of the above-mentioned "association".
[0425] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications can be made without departing from the gist of the present technology.
[0426] For example, a configuration described as one device (or processing unit) may be divided and configured as a plurality of devices (or processing units). Conversely, the configuration described as a plurality of devices (or processing units) may be collectively configured as one device (or processing unit). Furthermore, configurations other than the above configurations may be added to the configuration of each device (or each processing unit). Furthermore, as long as the configuration and operation of the entire system are substantially the same, a portion of the configuration of a particular device (or processing unit) may be included in the configuration of another device (or another processing unit).
[0427] In addition, for example, the above-mentioned program can be executed in any device. In this case, it is sufficient that the device has necessary functions (functional blocks, etc.) and can obtain necessary information.
[0428] Furthermore, for example, each step of a flowchart may be executed by a single device, or may be shared and executed by multiple devices. Furthermore, when multiple processes are included in a single step, the multiple processes may be executed by a single device, or may be shared and executed by multiple devices. In other words, the multiple processes included in a single step may also be executed as a single step. Conversely, a process described as a single step may be collectively executed as a single step.
[0429] In addition, the program executed by the computer may have the following features. For example, the processing of the steps of the description program can be performed in chronological order according to the order described in this specification. In addition, the processing of the steps of the description program can be performed in parallel. In addition, the processing of the steps of the description program can be performed separately at necessary timing (for example, when called). That is, as long as there is no contradiction, the processing of each step can be performed in an order different from the above order. In addition, the processing of the steps of the description program can be performed in parallel with the processing of another program. In addition, the processing of the steps of the description program can be performed in combination with the processing of another program.
[0430] Furthermore, for example, multiple technologies related to the present technology can be independently implemented as a single technology, as long as there are no contradictions. Of course, multiple technologies can be implemented in combination. For example, some or all of the technologies described in an embodiment can be implemented in combination with some or all of the technologies described in other embodiments. In addition, some or all of the above technologies can be implemented in combination with other technologies not described above.
[0431] Note that the present technology can also have the following configurations.
[0432] (1) An information processing device comprising:
[0433] a scalable decoding information generating unit configured to generate scalable decoding information regarding scalable decoding of a Geometry-Based Point Cloud Compression (G-PCC) content based on depth information indicating a quality hierarchy level of a geometry included in each slice of the G-PCC content including a first slice and a second slice, and a dependency relationship between the first slice and the second slice in the G-PCC content; and
[0434] The content file generating unit is configured to generate a content file storing G-PCC content and store scalable decoding information in a metadata area of the content file.
[0435] (2) The information processing device according to (1), wherein the scalable decoding information includes slice configuration information about a configuration of a slice of each sample.
[0436] (3) The information processing device according to (2), wherein the content file generation unit sets a subsample for each slice and stores the slice configuration information in a codec specific parameter of a subsample information box of the metadata area.
[0437] (4) The information processing device according to (2) or (3), wherein the slice configuration information includes slice dependency information indicating a dependency relationship between the first slice and the second slice.
[0438] (5) The information processing device according to (4), wherein the content file generation unit sets a subsample for each slice and stores the slice dependency information in a codec specific parameter of a subsample information box of the metadata area.
[0439] (6) The information processing device according to (4) or (5),
[0440] The slice dependency information includes the reference source geometry slice identification information and the reference destination geometry slice identification information.
[0441] The reference source geometry slice identification information is identification information of a geometry slice used as a reference source in the dependency relationship between the first slice and the second slice, and
[0442] The reference destination geometry slice identification information is identification information of a geometry slice used as a reference destination in the dependency relationship between the first slice and the second slice.
[0443] (7) An information processing device according to (6), wherein, when the geometry slice corresponding to the reference source geometry slice identification information is an independent decodable independent geometry slice, the reference source geometry slice identification information and the reference destination geometry slice identification information are the same.
[0444] (8) An information processing device according to (6) or (7), wherein the content file generation unit sets a subsample for each of the slices, and stores the reference source geometry slice identification information and the reference destination geometry slice identification information in the codec-specific parameters of the subsample information box in the metadata area.
[0445] (9) The information processing device according to any one of (4) to (8), wherein the slice dependency information includes attribute geometry slice identification information as identification information of a geometry slice referenced by the attribute slice.
[0446] (10) The information processing device according to (9), wherein the content file generation unit sets a subsample for each slice and stores the attribute geometry slice identification information in a codec specific parameter of a subsample information box of the metadata area.
[0447] (11) An information processing device according to any one of (4) to (8), wherein the slice dependency information includes non-scalable coded attribute geometry slice identification information as identification information of a geometry slice referenced by an attribute slice to which non-scalable coding is applied.
[0448] (12) The information processing device according to (11), wherein the non-scalable coded attribute geometry slice identification information includes identification information of a geometry slice including a geometry having maximum depth information among geometry slices referenced by the attribute slice.
[0449] (13) An information processing device according to (11) or (12), wherein the content file generation unit sets a subsample for each slice and stores the non-scalable coding attribute geometry slice identification information in the codec-specific parameters of the subsample information box in the metadata area.
[0450] (14) The information processing device according to any one of (11) to (13), wherein the non-scalable coding attribute geometry slice identification information includes a non-scalable coding flag indicating whether non-scalable coding is applied to the attribute slice.
[0451] (15) The information processing device according to (14), wherein the content file generation unit sets a subsample for each of the slices and stores the non-scalable coding flag in a codec specific parameter of a subsample information box of the metadata area.
[0452] (16) An information processing device according to any one of (10) to (15), wherein the payload type of slice dependency information of an independently decodable independent geometry slice is different from the payload type of slice dependency information of a dependent geometry slice that references other geometry slices in decoding.
[0453] (17) The information processing device according to any one of (2) to (16), wherein the slice configuration information includes geometry slice depth information which is depth information about a geometry included in the geometry slice.
[0454] (18) The information processing device according to (17), wherein the geometry slice depth information includes minimum depth information indicating a minimum value of the depth information in the geometry slice.
[0455] (19) The information processing device according to (17) or (18), wherein the geometry slice depth information includes maximum depth information indicating a maximum value of the depth information in the geometry slice.
[0456] (20) An information processing device according to any one of (17) to (19), wherein the content file generation unit sets a subsample for each slice and stores the geometry slice depth information in a codec-specific parameter of a subsample information box in a metadata area.
[0457] (21) The information processing device according to (20),
[0458] The slice composition information further includes slice dependency information indicating a dependency relationship between the first slice and the second slice, and
[0459] The content file generation unit sets a flag of a subsample information box in which the geometry slice depth information is stored to a value different from a flag of a subsample information box in which the slice dependency information is stored.
[0460] (22) The information processing device according to any one of (1) to (21), wherein the scalable decoding information includes track configuration information about a configuration of tracks storing the G-PCC content in the content file on a slice basis.
[0461] (23) The information processing device according to (22), wherein the track configuration information includes track depth information of depth information about geometry of all slices included in the track corresponding to the track configuration information.
[0462] (24) The information processing device according to (23), wherein the track depth information includes track minimum depth information indicating a minimum value of the depth information in the track.
[0463] (25) According to the information processing device described in (23) or (24), the track depth information includes track maximum depth information indicating a maximum value of the depth information in the track.
[0464] (26) An information processing device according to any one of (23) to (25), wherein the track depth information includes a matching flag indicating whether a minimum value of the depth information in each sample included in the track matches a minimum value of the depth information in the track and whether a maximum value of the depth information in each sample included in the track matches a maximum value of the depth information in the track.
[0465] (27) The information processing device according to any one of (23) to (26), wherein the content file generation unit stores the track depth information in a depth information box of a sample entry in the metadata area.
[0466] (28) The information processing device according to any one of (22) to (27), wherein the track configuration information includes track dependency information indicating a dependency relationship between the first track and the second track.
[0467] (29) The information processing device according to (28), wherein the track dependency information includes dependent information indicating another track including a slice necessary for decoding the dependent slice included in the track.
[0468] (30) The information processing device according to (29), wherein the dependent information indicates all other tracks including slices necessary for decoding the dependent slice.
[0469] (31) The information processing device according to (29), wherein the dependent information indicates another track including a slice referenced from the dependent slice.
[0470] (32) An information processing device according to any one of (28) to (31), wherein the track dependency information includes independent information indicating another track including another slice necessary for decoding the independent slice included in the track.
[0471] (33) The information processing device according to (32), wherein the independent information indicates another track including another slice that refers to the independent slice in decoding.
[0472] (34) The information processing device according to any one of (28) to (33), wherein the content file generation unit stores the track dependency information as a track reference in the metadata area.
[0473] (35) An information processing method comprising:
[0474] generating scalable decoding information regarding scalable decoding of geometry-based point cloud compression (G-PCC) content based on depth information indicating a quality hierarchy level of geometry included in each slice of the G-PCC content including a first slice and a second slice, and a dependency relationship between the first slice and the second slice in the G-PCC content; and
[0475] A content file storing G-PCC content is generated and scalable decoding information is stored in a metadata area of the content file.
[0476] (51) An information processing device comprising:
[0477] an extraction unit configured to extract an arbitrary slice of Geometry-Based Point Cloud Compression (G-PCC) content from a content file based on scalable decoding information stored in a metadata area of a content file storing the G-PCC content including the first slice and the second slice; and
[0478] a decoding unit configured to decode the slices of the G-PCC content extracted by the extraction unit,
[0479] The scalable decoding information is information about scalable decoding of G-PCC content, and is information generated based on depth information indicating the quality hierarchy level of the geometry included in the slice in the G-PCC content and the dependency between the first slice and the second slice in the G-PCC content.
[0480] (52) The information processing device according to (51), wherein the scalable decoding information includes slice configuration information about a configuration of a slice for each sample.
[0481] (53) An information processing device according to (52), wherein the extraction unit extracts an arbitrary slice of the G-PCC content from the content file based on slice composition information stored in a codec specific parameter of a subsample information box in a metadata area of a subsample set for each slice.
[0482] (54) The information processing device according to (52) or (53), wherein the slice configuration information includes slice dependency information indicating a dependency relationship between the first slice and the second slice.
[0483] (55) An information processing device according to (54), wherein the extraction unit extracts an arbitrary slice of the G-PCC content from the content file based on the slice dependency information stored in the codec specific parameters of the subsample information box of the metadata area of the subsample set for each slice.
[0484] (56) The information processing device according to (54) or (55),
[0485] The slice dependency information includes the reference source geometry slice identification information and the reference destination geometry slice identification information, and
[0486] The reference source geometry slice identification information is identification information of a geometry slice used as a reference source in the dependency relationship between the first slice and the second slice, and
[0487] The reference destination geometry slice identification information is identification information of a geometry slice used as a reference destination in the dependency relationship between the first slice and the second slice.
[0488] (57) An information processing device according to (56), wherein, when the geometry slice corresponding to the reference source geometry slice identification information is an independent decodable independent geometry slice, the reference source geometry slice identification information and the reference destination geometry slice identification information are the same.
[0489] (58) An information processing device according to (56) or (57), wherein the extraction unit extracts an arbitrary slice of the G-PCC content from the content file based on the reference source geometry slice identification information and the reference destination geometry slice identification information stored in the codec-specific parameters of the subsample information box of the metadata area of the subsample set for each slice.
[0490] (59) An information processing device according to any one of (54) to (58), wherein the slice dependency information includes attribute geometry slice identification information as identification information of a geometry slice referenced by the attribute slice.
[0491] (60) An information processing device according to (59), wherein the extraction unit extracts an arbitrary slice of the G-PCC content from the content file based on the attribute geometry slice identification information stored in the codec specific parameters of the subsample information box of the metadata area of the subsample set for each slice.
[0492] (61) An information processing device according to any one of (54) to (58), wherein the slice dependency information includes non-scalable coded attribute geometry slice identification information as identification information of a geometry slice referenced by an attribute slice to which non-scalable coding is applied.
[0493] (62) The information processing device according to (61), wherein the non-scalable coded attribute geometry slice identification information includes identification information of a geometry slice including a geometry having maximum depth information among geometry slices referenced by the attribute slice.
[0494] (63) An information processing device according to (61) or (62), wherein the extraction unit extracts an arbitrary slice of the G-PCC content from the content file based on the non-scalable coding attribute geometry slice identification information stored in the codec-specific parameters of the subsample information box of the metadata area of the subsample set for each slice.
[0495] (64) The information processing device according to any one of (61) to (63), wherein the non-scalable coding attribute geometry slice identification information includes a non-scalable coding flag indicating whether non-scalable coding is applied to the attribute slice.
[0496] (65) An information processing device according to (64), wherein the extraction unit extracts an arbitrary slice of the G-PCC content from the content file based on a non-scalable coding flag stored in a codec-specific parameter of a subsample information box in a metadata area of a subsample set for each slice.
[0497] (66) An information processing device according to any one of (60) to (65), wherein the payload type of slice dependency information of an independently decodable independent geometry slice is different from the payload type of slice dependency information of a dependent geometry slice that references other geometry slices in decoding.
[0498] (67) The information processing device according to any one of (52) to (66), wherein the slice configuration information includes geometry slice depth information which is depth information about a geometry included in the geometry slice.
[0499] (68) The information processing device according to (67), wherein the geometry slice depth information includes minimum depth information indicating a minimum value of the depth information in the geometry slice.
[0500] (69) According to the information processing device described in (67) or (68), the geometry slice depth information includes maximum depth information indicating the maximum value of the depth information in the geometry slice.
[0501] (70) An information processing device according to any one of (67) to (69), wherein the extraction unit extracts an arbitrary slice of the G-PCC content from the content file based on the geometry slice depth information stored in the codec specific parameters of the subsample information box of the metadata area of the subsample set for each slice.
[0502] (71) The information processing device according to (70),
[0503] The slice composition information further includes slice dependency information indicating a dependency relationship between the first slice and the second slice, and
[0504] The flag of the subsample information box storing the geometry slice depth information and the flag of the subsample information box storing the slice dependency information are set to different values.
[0505] (72) The information processing device according to any one of (51) to (71), wherein the scalable decoding information includes track configuration information about a configuration of tracks storing the G-PCC content in the content file on a slice basis.
[0506] (73) The information processing device according to (72), wherein the track configuration information includes track depth information of depth information about geometry of all slices included in the track corresponding to the track configuration information.
[0507] (74) The information processing device according to (73), wherein the track depth information includes track minimum depth information indicating a minimum value of the depth information in the track.
[0508] (75) The information processing device according to (73) or (74), wherein the track depth information includes track maximum depth information indicating a maximum value of the depth information in the track.
[0509] (76) An information processing device according to any one of (73) to (75), wherein the track depth information includes a matching flag indicating whether a minimum value of the depth information in each sample included in the track matches a minimum value of the depth information in the track and whether a maximum value of the depth information in each sample included in the track matches a maximum value of the depth information in the track.
[0510] (77) An information processing device according to any one of (73) to (76), wherein the extraction unit extracts an arbitrary slice of the G-PCC content from the content file based on the track depth information stored in the depth information box of the sample entry of the metadata area.
[0511] (78) The information processing device according to any one of (72) to (77), wherein the track configuration information includes track dependency information indicating a dependency relationship between the first track and the second track.
[0512] (79) The information processing device according to (78), wherein the track dependency information includes dependent information indicating another track including a slice necessary for decoding the dependent slice included in the track.
[0513] (80) The information processing device according to (79), wherein the dependent information indicates all additional tracks including slices necessary for decoding the dependent slice.
[0514] (81) The information processing device according to (79), wherein the dependent information indicates other tracks including slices referenced from the dependent slice.
[0515] (82) An information processing device according to any one of (78) to (81), wherein the track dependency information includes independent information indicating an additional track including an additional slice necessary for decoding an independent slice included in the track
[0516] (83) The information processing device according to (82), wherein the independent information indicates another track including another slice that refers to the independent slice in decoding.
[0517] (84) An information processing device according to any one of (78) to (83), wherein the extraction unit extracts an arbitrary slice of the G-PCC content from the content file based on track dependency information stored as a track reference in the metadata area.
[0518] (85) An information processing method comprising:
[0519] extracting an arbitrary slice of Geometry-Based Point Cloud Compression (G-PCC) content from a content file based on scalable decoding information stored in a metadata area of the content file storing the G-PCC content including the first slice and the second slice; and
[0520] Decode the extracted slices of G-PCC content,
[0521] The scalable decoding information is information about scalable decoding of G-PCC content, and is information generated based on depth information indicating the quality hierarchy level of the geometry included in the slice in the G-PCC content and the dependency between the first slice and the second slice in the G-PCC content.
[0522] (101) An information processing device comprising:
[0523] an adaptation set composition information generating unit configured to generate adaptation set composition information based on depth information indicating a quality hierarchy level of a geometry included in each slice of Geometry-Based Point Cloud Compression (G-PCC) content including a first slice and a second slice, and a dependency relationship between the first slice and the second slice in the G-PCC content; and
[0524] A control file generating unit is configured to generate a control file for controlling reproduction of a content file storing G-PCC content and store adaptation set configuration information in the control file.
[0525] The content file stores G-PCC content in tracks based on slices, and
[0526] The adaptation set configuration information is information about the configuration of an adaptation set that describes information about a track of a content file.
[0527] (102) The information processing device according to (101), wherein the adaptation set configuration information includes adaptation set depth information regarding depth information of geometries of all slices included in a track corresponding to the adaptation set.
[0528] (103) The information processing device according to (102), wherein the adaptation set depth information includes adaptation set minimum depth information indicating a minimum value of the depth information in the track.
[0529] (104) The information processing device according to (102) or (103), wherein the adaptation set depth information includes adaptation set maximum depth information indicating a maximum value of depth information in the track.
[0530] (105) The information processing device according to any one of (102) to (104),
[0531] The adaptation set depth information includes a matching flag indicating whether the sample minimum depth information of each sample included in the track matches the adaptation set minimum depth information and whether the sample maximum depth information of each sample included in the track matches the adaptation set maximum depth information.
[0532] The minimum depth information of a sample indicates the minimum value of the depth information in the sample.
[0533] The maximum depth information of a sample indicates the maximum value of the depth information in the sample.
[0534] The adaptation set minimum depth information indicates the minimum value of the depth information in the track, and
[0535] The adaptation set maximum depth information indicates the maximum value of depth information in the track.
[0536] (106) The information processing apparatus according to (101), wherein the adaptation set configuration information includes presentation dependency information indicating a dependency relationship between the first presentation and the second presentation.
[0537] (107) The information processing apparatus according to (106), wherein the presentation dependency information includes dependent information indicating another presentation necessary for decoding the presentation corresponding to the presentation dependency information.
[0538] (108) The information processing device according to (107), wherein the dependent information indicates all other representations necessary for decoding the representation corresponding to the dependent information.
[0539] (109) The information processing apparatus according to (107), wherein the subordinate information indicates another presentation referred to from the presentation corresponding to the subordinate information.
[0540] (110) An information processing device according to any one of (106) to (109), wherein the presentation dependency information includes independent information indicating an additional presentation necessary in decoding for the presentation corresponding to the presentation dependency information.
[0541] (111) The information processing device according to (110), wherein the independent information indicates all other representations necessary in decoding for the representation corresponding to the independent information.
[0542] (112) The information processing device according to any one of (106) to (111), wherein the control file generation unit stores the presentation dependency relationship information as a presentation association ID and an association type in the control file.
[0543] (113) An information processing method comprising:
[0544] generating adaptation set composition information based on depth information indicating a quality hierarchy level of geometry included in each slice of Geometry-Based Point Cloud Compression (G-PCC) content including a first slice and a second slice, and a dependency relationship between the first slice and the second slice in the G-PCC content; and
[0545] Generate a control file for controlling the reproduction of a content file storing G-PCC content and store adaptation set configuration information in the control file,
[0546] The content file stores G-PCC content in tracks based on slices, and
[0547] The adaptation set configuration information is information about the configuration of an adaptation set that describes information about a track of a content file.
[0548] (151) An information processing device comprising:
[0549] an analyzing unit configured to analyze a control file that controls reproduction of a content file storing Geometry-Based Point Cloud Compression (G-PCC) content including a first slice and a second slice in a track on a slice basis and to specify an adaptation set necessary to obtain an arbitrary slice of the G-PCC content based on adaptation set composition information stored in the control file;
[0550] an acquisition unit configured to acquire a track of a content file corresponding to the adaptation set specified by the analysis unit; and
[0551] A decoding unit is configured to decode a slice of the G-PCC content stored in the track acquired by the acquisition unit.
[0552] The adaptation set composition information is information about the configuration of the adaptation set that describes information about the track of the content file, and is information generated based on depth information indicating the quality hierarchy level of the geometry included in the slice in the G-PCC content and the dependency relationship between the first slice and the second slice in the G-PCC content.
[0553] (152) The information processing device according to (151), wherein the adaptation set configuration information includes adaptation set depth information regarding depth information of geometries of all slices included in a track corresponding to the adaptation set.
[0554] (153) The information processing device according to (152), wherein the adaptation set depth information includes adaptation set minimum depth information indicating a minimum value of the depth information in the track.
[0555] (154) The information processing device according to (152) or (153), wherein the adaptation set depth information includes adaptation set maximum depth information indicating a maximum value of depth information in the track.
[0556] (155) The information processing device according to any one of (152) to (154),
[0557] The adaptation set depth information includes a matching flag indicating whether the sample minimum depth information of each sample included in the track matches the adaptation set minimum depth information and whether the sample maximum depth information of each sample included in the track matches the adaptation set maximum depth information.
[0558] The minimum depth information of a sample indicates the minimum value of the depth information in the sample.
[0559] The maximum depth information of a sample indicates the maximum value of the depth information in the sample.
[0560] The adaptation set minimum depth information indicates the minimum value of the depth information in the track, and
[0561] The adaptation set maximum depth information indicates the maximum value of depth information in the track.
[0562] (156) The information processing device according to any one of (151) to (155), wherein the adaptation set composition information includes presentation dependency information indicating a dependency relationship between the first presentation and the second presentation.
[0563] (157) The information processing apparatus according to (156), wherein the presentation dependency information includes dependent information indicating another presentation necessary for decoding the presentation corresponding to the presentation dependency information.
[0564] (158) The information processing device according to (157), wherein the dependent information indicates all other representations necessary for decoding the representation corresponding to the dependent information.
[0565] (159) The information processing apparatus according to (157), wherein the subordinate information indicates other presentations referenced from the presentation corresponding to the subordinate information.
[0566] (160) An information processing device according to any one of (156) to (159), wherein the presentation dependency information includes independent information indicating an additional presentation required in decoding for the presentation corresponding to the presentation dependency information.
[0567] (161) The information processing device according to (160), wherein the independent information indicates all other representations necessary in decoding for the representation corresponding to the independent information.
[0568] (162) The information processing device according to any one of (156) to (161), wherein the analysis unit specifies the adaptation set based on presentation dependency information stored in the control file as a presentation association ID and an association type.
[0569] (163) An information processing method comprising:
[0570] analyzing a control file for controlling reproduction of a content file storing geometry-based point cloud compression (G-PCC) content including a first slice and a second slice in a track on a slice basis, and specifying an adaptation set necessary to obtain any slice of the G-PCC content based on adaptation set composition information stored in the control file;
[0571] Obtaining a track of a content file corresponding to the identified adaptation set; and
[0572] Decode the slices of the G-PCC content stored in the acquired track,
[0573] The adaptation set composition information is information about the configuration of the adaptation set that describes information about the track of the content file, and is information generated based on depth information indicating the quality hierarchy level of the geometry included in the slice in the G-PCC content and the dependency relationship between the first slice and the second slice in the G-PCC content.
[0574] Reference Signs List
[0575] 300 File Generation Unit
[0576] 311 Extraction Unit
[0577] 312 coding units
[0578] 313 Bitstream Generation Unit
[0579] 314 Scalable decoding information generation unit
[0580] 315 File Generation Unit
[0581] 321 Geometry Encoding Unit
[0582] 322 attribute coding unit
[0583] 323 Metadata Generation Unit
[0584] 400 Reproduction Device
[0585] 401 Control Unit
[0586] 411 File Retrieval Unit
[0587] 412 Reproduction Processing Unit
[0588] 413 Rendering Processing Unit
[0589] 421 File Processing Unit
[0590] 422 decoding unit
[0591] 423 Presentation Information Generation Unit
[0592] 431 Bitstream Extraction Unit
[0593] 441 Geometry Decoding Unit
[0594] 442 Attribute Decoding Unit
[0595] 451 Point Cloud Construction Unit
[0596] 452 Rendering Processing Unit
[0597] 600 File Generation Unit
[0598] 615 Content file generation unit
[0599] 616 MPD generation unit
[0600] 700 Reproduction Device
[0601] 711 File Acquisition Unit
[0602] 712 MPD analysis unit.< / hevc>
Claims
1. An information processing device, comprising: a scalable decoding information generating unit configured to generate scalable decoding information regarding scalable decoding of a geometry-based point cloud compressed G-PCC content based on depth information indicating a quality hierarchy level of each slice in the G-PCC content including a first slice and a second slice, and a dependency relationship between the first slice and the second slice in the G-PCC content; as well as The content file generating unit is configured to generate a content file storing the G-PCC content and store the scalable decoding information in a metadata area of the content file.
2. The information processing device according to claim 1, wherein The scalable decoding information includes slice configuration information regarding the configuration of a slice of each sample.
3. The information processing device according to claim 2, wherein: The content file generation unit sets a subsample for each slice and stores the slice configuration information in a codec specific parameter of a subsample information box of the metadata area.
4. The information processing device according to claim 2, wherein: The slice composition information includes slice dependency information indicating the dependency relationship between the first slice and the second slice.
5. The information processing device according to claim 4, in, The slice dependency information includes reference source geometry slice identification information and reference destination geometry slice identification information. The reference source geometry slice identification information is identification information of a geometry slice used as a reference source in the dependency relationship between the first slice and the second slice, and The reference destination geometry slice identification information is identification information of a geometry slice used as a reference destination in the dependency relationship between the first slice and the second slice. The information processing apparatus according to claim 2 , wherein: The slice configuration information includes geometry slice depth information regarding depth information of the geometry included in the geometry slice.
7. The information processing apparatus according to claim 6, wherein: The geometry slice depth information includes minimum depth information indicating a minimum value of the depth information in the geometry slice and maximum depth information indicating a maximum value of the depth information in the geometry slice.
8. The information processing apparatus according to claim 1, wherein: The scalable decoding information includes track configuration information regarding the configuration of tracks storing the G-PCC content in the content file in units of slices.
9. The information processing apparatus according to claim 8, wherein: The track composition information includes track depth information regarding depth information of geometries of all slices included in the track corresponding to the track composition information.
10. The information processing apparatus according to claim 9, wherein: The track depth information includes track minimum depth information indicating a minimum value of the depth information in the track and track maximum depth information indicating a maximum value of the depth information in the track.
11. The information processing apparatus according to claim 9, wherein: The content file generation unit stores the track depth information in a depth information box of a sample entry in the metadata area.
12. The information processing apparatus according to claim 8, wherein: The track composition information includes track dependency information indicating a dependency relationship between a first track and a second track.
13. The information processing apparatus according to claim 12, wherein: The track dependency information includes dependency information indicating another track including slices necessary for decoding dependent slices included in the track.
14. The information processing apparatus according to claim 12, wherein: The track dependency information includes independent information indicating an additional track including an additional slice necessary in decoding of an independent slice included in the track.
15. The information processing apparatus according to claim 12, wherein: The content file generation unit stores the track dependency information in the metadata area as a track reference.
16. An information processing method, comprising: generating scalable decoding information regarding scalable decoding of geometry-based point cloud compressed G-PCC content based on depth information indicating a quality hierarchy level of each slice in the G-PCC content including a first slice and a second slice and a dependency relationship between the first slice and the second slice in the G-PCC content; as well as A content file storing the G-PCC content is generated and the scalable decoding information is stored in a metadata area of the content file.
17. An information processing device comprising: an extraction unit configured to extract an arbitrary slice of the geometry-based point cloud compressed G-PCC content from the content file based on scalable decoding information stored in a metadata area of the content file storing the G-PCC content including the first slice and the second slice; as well as a decoding unit configured to decode the slice of the G-PCC content extracted by the extraction unit, The scalable decoding information is information about scalable decoding of the G-PCC content, and is information generated based on depth information indicating the quality hierarchy level of the slice in the G-PCC content and the dependency relationship between the first slice and the second slice in the G-PCC content.
18. The information processing apparatus according to claim 17, wherein: The scalable decoding information includes slice configuration information regarding the configuration of a slice of each sample.
19. The information processing apparatus according to claim 17, wherein: The scalable decoding information includes track configuration information regarding the configuration of tracks storing the G-PCC content in the content file in units of slices.
20. An information processing method, comprising: extracting an arbitrary slice of the geometry-based point cloud compressed G-PCC content including the first slice and the second slice from the content file based on scalable decoding information stored in a metadata area of the content file; as well as decoding the extracted slice of the G-PCC content, The scalable decoding information is information about scalable decoding of the G-PCC content, and is information generated based on depth information indicating the quality hierarchy level of the slice in the G-PCC content and the dependency relationship between the first slice and the second slice in the G-PCC content.