Information processing device and method
The solution addresses the lack of spatial scalability in ISOBMFF by encoding 3D data layers as sub-bitstreams and storing scalability information in the system layer, enabling easy playback and optimal LoD selection for 3D data, enhancing media quality and bandwidth efficiency.
Patent Information
- Application Number
- JP2022530134
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-12
- Filing Date
- 2021-05-28
- Publication Date
- 2025-08-20
- Estimated Expiration
- 2041-05-28
AI Technical Summary
The ISOBMFF system does not support spatial scalability, making it difficult to store and utilize information about spatial scalability in the system layer, which is necessary for clients to construct 3D data at desired levels of detail (LoD) without cumbersome parsing tasks.
An information processing device and method that supports spatial scalability by encoding each layer of hierarchical 3D data as a sub-bitstream, generating a bitstream including a sub-bitstream, and storing spatial scalability information in the system layer of a file, allowing easy identification and selection of appropriate layers for decoding.
Enables clients to easily playback 3D data using spatial scalability, allowing selection of appropriate levels of detail (LoD) without complex parsing, optimizing bandwidth utilization and media quality under varying network conditions.
Smart Images

Figure 0007726209000001 
Figure 0007726209000002 
Figure 0007726209000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device and method, and more particularly to an information processing device and method that enable 3D data to be reproduced more easily by utilizing spatial scalability. [Background technology]
[0002] Conventionally, the Moving Picture Experts Group (MPEG) has been working on standardization of encoding and decoding of point clouds, which represent three-dimensional objects as a collection of points. A method has been proposed in which the geometry and attributes of the point cloud are projected onto a two-dimensional plane for each small region, the images (patches) projected onto the two-dimensional plane are placed within the frame images of a video, and the video is encoded using an encoding method for two-dimensional images (hereinafter also referred to as V-PCC (Video-based Point Cloud Compression)) (see, for example, Non-Patent Document 1).
[0003] There is also the International Organization for Standardization Base Media File Format (ISOBMFF), which is a file container specification of the international standard technology for video compression, Moving Picture Experts Group-4 (MPEG-4) (see, for example, Non-Patent Documents 2 and 3).
[0004] Then, with the aim of improving the efficiency of playback processing from local storage and network distribution of bitstreams coded with V-PCC (also called V3C bitstreams), methods of storing V3C bitstreams in ISOBMFF are being studied (see, for example, Non-Patent Document 4). Non-Patent Document 4 also discloses a partial access technology that decodes only a part of a point cloud object.
[0005] Furthermore, in MPEG-I Part 5 Visual Volumetric Video-based Coding (V3C) and Video-based Point Cloud Compression (V-PCC), a LoD patch mode was proposed that encodes the low-LoD (sparse) point clouds that make up a high-LoD (dense) point cloud independently, allowing the client to construct a low-LoD point cloud (see, for example, Non-Patent Document 5). [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] "V-PCC Future Enhancements (V3C + V-PCC)", ISO / IEC JTC 1 / SC 29 / WG 11 N19329, 2020-04-24 [Non-patent document 2] "Information technology - Coding of audio-visual objects - Part 12:ISO base media file format", ISO / IEC 14496-12, 2015-02-20 [Non-patent document 3] "Information technology - Coding of audio-visual objects - Part 15: Carriage of network abstraction layer (NAL) unit structured video in the ISO base media file format", ISO / IEC FDIS 14496-15:2014(E), ISO / IEC JTC 1 / SC 29 / WG 11,2014-01-13 [Non-patent document 4] "Text of ISO / IEC DIS 23090-10 Carriage of Visual Volumetric Video-based Coding Data", ISO / IEC JTC 1 / SC 29 / WG 11 N19285,2020-06-01 [Non-Patent Document 5] "Report on Scalability features in V-PCC", ISO / IEC JTC 1 / SC 29 / WG 11 N19156,2020-01-22 Summary of the Invention [Problem to be solved by the invention]
[0007] However, the ISOBMFF that stores the V3C bitstream described in Non-Patent Document 4 does not support spatial scalability, making it difficult to store information about spatial scalability in the system layer. Therefore, in order for a client to use spatial scalability to construct 3D data at a desired LoD, it is necessary to perform cumbersome tasks such as analyzing the V3C bitstream.
[0008] The present disclosure has been made in light of these circumstances, and aims to make it easier to play back 3D data using spatial scalability. [Means for solving the problem]
[0009] An information processing device according to one aspect of the present technology supports spatial scalability for controlling the resolution of reconstructed 3D data, and processes the 3D data having a hierarchical structure based on the resolution. encoding each layer of the hierarchical structure as a sub-bitstream, a coding unit for generating a bitstream including a sub-bitstream; and information including identification information of the layer corresponding to the sub-bitstream and information regarding the resolution of the 3D data obtained by reconstructing from the top layer of the hierarchical structure to the layer of the sub-bitstream.An information processing device comprising: a spatial scalability information generation unit that generates spatial scalability information; and a file generation unit that generates a file that stores the bitstream generated by the encoding unit and stores the spatial scalability information generated by the spatial scalability information generation unit in a system layer of the file.
[0010] An information processing method according to one aspect of the present technology supports spatial scalability for controlling the resolution of reconstructed 3D data, and processes the 3D data having a hierarchical structure based on the resolution. encoding each layer of the hierarchical structure as a sub-bitstream, generating a bitstream including a sub-bitstream, and determining the spatial scalability of the sub-bitstream; information including identification information of the layer corresponding to the sub-bitstream and information regarding the resolution of the 3D data obtained by reconstructing from the top layer of the hierarchical structure to the layer of the sub-bitstream. An information processing method for generating spatial scalability information, generating a file for storing the generated bitstream, and storing the generated spatial scalability information in a system layer of the file.
[0011] Another aspect of the present technology relates to a spatial scalability for controlling the resolution of reconstructed 3D data stored in a system layer of a file. information including identification information of the layer corresponding to the sub-bitstream coded for each layer of a hierarchical structure based on the resolution of the 3D data, and information about the resolution of the 3D data obtained by reconstructing from the top layer of the hierarchical structure to the layer of the sub-bitstream. Based on spatial scalability information, The layer to be decoded an extraction unit that extracts a sub-bitstream corresponding to the layer selected by the selection unit from the bitstream of the 3D data stored in the file; and a decoding unit that decodes the sub-bitstream extracted by the extraction unit.
[0012] Another aspect of the present technology is an information processing method that relates to spatial scalability, which controls the resolution of reconstructed 3D data stored in the system layer of a file. information including identification information of the layer corresponding to the sub-bitstream coded for each layer of a hierarchical structure based on the resolution of the 3D data, and information about the resolution of the 3D data obtained by reconstructing from the top layer of the hierarchical structure to the layer of the sub-bitstream. Based on spatial scalability information, The layer to be decodedthe information processing method, which selects a layer from the 3D data bitstream stored in the file, extracts a sub-bitstream corresponding to the selected layer from the 3D data bitstream stored in the file, and decodes the extracted sub-bitstream.
[0013] According to yet another aspect of the present technology, there is provided an information processing device that supports spatial scalability for controlling a resolution of reconstructed 3D data, and processes the 3D data having a hierarchical structure based on the resolution. encoding each layer of the hierarchical structure as a sub-bitstream, a coding unit for generating a bitstream including a sub-bitstream; and information including identification information of the layer corresponding to the sub-bitstream and information regarding the resolution of the 3D data obtained by reconstructing from the top layer of the hierarchical structure to the layer of the sub-bitstream. An information processing device comprising: a spatial scalability information generation unit that generates spatial scalability information; and a control file generation unit that generates a control file that stores the spatial scalability information generated by the spatial scalability information generation unit and control information regarding distribution of the bitstream generated by the encoding unit.
[0014] An information processing method according to yet another aspect of the present technology includes: a spatial scalability method for controlling a resolution of reconstructed 3D data; and a method for processing the 3D data having a hierarchical structure based on the resolution. encoding each layer of the hierarchical structure as a sub-bitstream, generating a bitstream including a sub-bitstream, and determining the spatial scalability of the sub-bitstream; information including identification information of the layer corresponding to the sub-bitstream and information regarding the resolution of the 3D data obtained by reconstructing from the top layer of the hierarchical structure to the layer of the sub-bitstream. An information processing method for generating spatial scalability information and generating a control file that stores the generated spatial scalability information and control information related to distribution of the generated bitstream.
[0015] An information processing device according to yet another aspect of the present technology supports spatial scalability for controlling the resolution of reconstructed 3D data, and includes a control file storing control information related to distribution of a bitstream in which the 3D data having a hierarchical structure based on the resolution is encoded. information including identification information of the layer corresponding to the sub-bitstream coded for each layer of the hierarchical structure, and information regarding the resolution of the 3D data obtained by reconstructing from the top layer of the hierarchical structure to the layer of the sub-bitstream.Based on spatial scalability information, The layer to be decoded an acquisition unit that acquires a sub-bitstream corresponding to the layer selected by the selection unit; and a decoding unit that decodes the sub-bitstream acquired by the acquisition unit.
[0016] An information processing method according to yet another aspect of the present technology relates to spatial scalability of a bitstream, wherein the spatial scalability controls the resolution of reconstructed 3D data, and the control information relating to distribution of the bitstream in which the 3D data having a hierarchical structure based on the resolution is stored in a control file. information including identification information of the layer corresponding to the sub-bitstream coded for each layer of the hierarchical structure, and information regarding the resolution of the 3D data obtained by reconstructing from the top layer of the hierarchical structure to the layer of the sub-bitstream. Based on spatial scalability information, The layer to be decoded a sub-bitstream corresponding to the selected layer, and decoding the sub-bitstream.
[0017] In an information processing device and method according to one aspect of the present technology, spatial scalability is supported for controlling the resolution of reconstructed 3D data, and 3D data having a hierarchical structure based on the resolution is generated. is coded as a sub-bitstream for each layer of the hierarchical structure, A bitstream including a sub-bitstream is generated, and spatial scalability of the sub-bitstream is calculated. information including identification information of a layer corresponding to the sub-bitstream and information regarding the resolution of 3D data obtained by reconstructing from the top layer of the hierarchical structure to the layer of the sub-bitstream. Spatial scalability information is generated, a file is generated that stores the generated bitstream, and the generated spatial scalability information is stored in a system layer of the file.
[0018] In another aspect of the present technology, an information processing device and method relate to spatial scalability, which controls the resolution of reconstructed 3D data stored in a system layer of a file. The information includes identification information of layers corresponding to sub-bitstreams coded for each layer of a hierarchical structure based on the resolution of the 3D data, and information on the resolution of the 3D data obtained by reconstructing the layers from the top layer of the hierarchical structure to the layer of the sub-bitstream. Based on spatial scalability information, Layer to decodeis selected, a sub-bitstream corresponding to the selected layer is extracted from the bitstream of the 3D data stored in the file, and the extracted sub-bitstream is decoded.
[0019] In an information processing device and method according to still another aspect of the present technology, spatial scalability for controlling the resolution of reconstructed 3D data is supported, and 3D data having a hierarchical structure based on the resolution is provided. Each layer of the hierarchical structure is coded as a sub-bitstream, and A bitstream including a sub-bitstream is generated, and spatial scalability of the sub-bitstream is calculated. information including identification information of a layer corresponding to the sub-bitstream and information regarding the resolution of 3D data obtained by reconstructing from the top layer of the hierarchical structure to the layer of the sub-bitstream. Spatial scalability information is generated, and a control file is generated that stores the generated spatial scalability information and control information regarding distribution of the generated bitstream.
[0020] In an information processing device and method according to still another aspect of the present technology, spatial scalability for controlling the resolution of reconstructed 3D data is supported, and control information relating to the distribution of a bitstream in which 3D data having a hierarchical structure based on the resolution is stored in a control file. information including identification information of layers corresponding to sub-bitstreams coded for each layer of the hierarchical structure, and information regarding the resolution of 3D data obtained by reconstructing from the top layer of the hierarchical structure to the layer of the sub-bitstream. Based on spatial scalability information, Layer to decode is selected, a sub-bitstream corresponding to the selected layer is obtained, and the obtained sub-bitstream is decoded. [Brief explanation of the drawings]
[0021] [Figure 1] FIG. 1 is a diagram illustrating an overview of a V-PCC. [Figure 2] FIG. 1 is a diagram illustrating an example of the main configuration of a V3C bitstream. [Figure 3] A diagram showing an example of the main structure of an atlas sub-bitstream. [Figure 4] A diagram showing an example of the structure of an ISOBMFF that stores a V3C bitstream. [Figure 5] FIG. 10 is a diagram illustrating an example of partial access information. [Figure 6] FIG. 10 is a diagram illustrating an example of a file structure when a 3D spatial region is static. [Figure 7] FIG. 10 is a diagram illustrating an example of a SpatialRegionGroupBox and a V3CSpatialRegionsBox. [Figure 8] FIG. 10 is a diagram illustrating an example of a file structure when a 3D spatial region changes dynamically. [Figure 9] FIG. 1 is a diagram illustrating spatial scalability. [Figure 10] FIG. 10 is a diagram illustrating an example of syntax related to spatial scalability. [Figure 11] FIG. 10 is a diagram illustrating an example of the configuration of data corresponding to spatial scalability. [Figure 12] FIG. 10 is a diagram illustrating an example of a file structure for storing spatial scalability information. [Figure 13] FIG. 10 is a diagram illustrating an example of syntax. [Figure 14] FIG. 10 is a diagram illustrating another example of syntax. [Figure 15] FIG. 10 is a diagram illustrating an example of the configuration of a Matryoshka media container. [Figure 16] FIG. 2 is a block diagram illustrating an example of the main configuration of a file generation device. [Figure 17] 10 is a flowchart illustrating an example of the flow of a file generation process. [Figure 18] FIG. 2 is a block diagram illustrating an example of the main configuration of a client device. [Figure 19] 10 is a flowchart illustrating an example of the flow of a client process. [Figure 20] FIG. 10 is a diagram illustrating an example of the configuration of an MPD that stores spatial scalability information. [Figure 21] FIG. 10 is a diagram illustrating an example of syntax. [Figure 22] FIG. 10 is a diagram illustrating another example of syntax. [Figure 23] FIG. 10 is a diagram illustrating an example of a description of an MPD. [Figure 24] FIG. 10 is a diagram illustrating an example of a description of an MPD. [Figure 25] FIG. 2 is a block diagram illustrating an example of the main configuration of a file generation device. [Figure 26] 10 is a flowchart illustrating an example of the flow of a file generation process. [Figure 27] FIG. 2 is a block diagram illustrating an example of the main configuration of a client device. [Figure 28] 10 is a flowchart illustrating an example of the flow of a client process. [Figure 29] FIG. 1 is a block diagram illustrating an example of the main configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION
[0022] Hereinafter, modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described in the following order. 1. Spatial scalability of V3C bitstream 2. First embodiment (file storing bitstream and spatial scalability information) 3. Second embodiment (control file for storing spatial scalability information) 4. Notes
[0023] <1. Spatial scalability of V3C bitstream> <References supporting technical content and technical terminology> The scope of disclosure of the present technology includes not only the contents described in the embodiments but also the contents described in the following non-patent documents that were publicly known at the time of filing, as well as the contents of other documents referenced in the following non-patent documents.
[0024] Non-patent document 1: (mentioned above) Non-patent document 2: (mentioned above) Non-patent document 3: (mentioned above) Non-patent document 4: (mentioned above) Non-Patent Document 5: (above) Non-Patent Document 6: https: / / www.matroska.org / index.html
[0025] That is, the content described in the above non-patent documents and the content of other documents referred to in the above non-patent documents, etc. also serve as a basis when judging the support requirements.
[0026] <Point cloud> Conventionally, 3D data such as a point cloud (Point cloud) representing a three-dimensional structure has existed based on point position information, attribute information, etc.
[0027] For example, in the case of a point cloud, a three-dimensional structure (a three-dimensional object) is represented as a set of a large number of points. A point cloud is composed of the position information (also referred to as geometry) and attribute information (also referred to as attribute) of each point. The attribute can include any information. For example, color information, reflectance information, normal information, etc. of each point may be included in the attribute. In this way, the point cloud has a relatively simple data structure and can represent any three-dimensional structure with sufficient accuracy by using a sufficiently large number of points.
[0028] <Overview of V-PCC> In V-PCC (Video-based Point Cloud Compression), the geometry and attributes of such a point cloud are projected onto a two-dimensional plane for each small region. In the present disclosure, these small regions may be referred to as partial regions. An image in which the geometry and attributes are projected onto a two-dimensional plane is also referred to as a projected image. Furthermore, the projected image for each small region (partial region) is referred to as a patch. For example, object 1 (3D data) in A of FIG. 1 is decomposed into patch 2 (2D data) as shown in B of FIG. 1. In the case of a geometry patch, each pixel value indicates the position information of a point. However, in this case, the position information of the point is expressed as position information (depth value) in the direction perpendicular to the projection plane (depth direction).
[0029] Then, each patch generated in this manner is arranged in a frame image (also referred to as a video frame) of a video sequence. A frame image in which geometry patches are arranged is also referred to as a geometry video frame. A frame image in which attribute patches are arranged is also referred to as an attribute video frame. For example, from object 1 in A of FIG. 1, a geometry video frame 11 in which geometry patches 3 are arranged as shown in C of FIG. 1, and an attribute video frame 12 in which attribute patches 4 are arranged as shown in D of FIG. 1 are generated. For example, each pixel value of geometry video frame 11 indicates the above-mentioned depth value.
[0030] These video frames are then encoded using a coding method for two-dimensional images, such as AVC (Advanced Video Coding) or HEVC (High Efficiency Video Coding). In other words, point cloud data, which is 3D data representing a three-dimensional structure, can be encoded using a codec for two-dimensional images.
[0031] An occupancy map can also be used. The occupancy map is map information that indicates the presence or absence of a projected image (patch) for each NxN pixel of a geometry video frame or an attribute video frame. For example, the occupancy map indicates areas (NxN pixels) in the geometry video frame or the attribute video frame where a patch exists with a value of "1" and areas (NxN pixels) in which a patch does not exist with a value of "0."
[0032] By referencing this occupancy map, the decoder can determine whether an area contains a patch, thereby suppressing the effects of noise caused by encoding and decoding and restoring 3D data more accurately. For example, even if depth values change due to encoding and decoding, the decoder can ignore depth values in areas where no patches exist by referencing the occupancy map. In other words, by referencing the occupancy map, the decoder can avoid processing the depth values as position information for 3D data.
[0033] For example, an occupancy map 13 as shown in Fig. 1E may be generated for the geometry video frame 11 and the attribute video frame 12. In the occupancy map 13, white areas indicate a value of "1" and black areas indicate a value of "0."
[0034] Such an occupancy map can be encoded as data (video frame) separate from the geometry video frame and the attribute video frame and transmitted to the decoding side. That is, the occupancy map can also be encoded using a coding method for two-dimensional images such as AVC or HEVC, just like the geometry video frame and the attribute video frame.
[0035] The coded data (bitstream) generated by coding a geometry video frame is also called a geometry video sub-bitstream. The coded data (bitstream) generated by coding an attribute video frame is also called an attribute video sub-bitstream. The coded data (bitstream) generated by coding an occupancy map is also called an occupancy map video sub-bitstream. Note that when there is no need to distinguish between the geometry video sub-bitstream, attribute video sub-bitstream, and occupancy map video sub-bitstream, they are all called video sub-bitstreams.
[0036] Furthermore, atlas information (atlas), which is information for reconstructing a point cloud (3D data) from the patches (2D data), is encoded and transmitted to the decoding side. The encoding method (and decoding method) of the atlas information is arbitrary. The encoded data (bitstream) generated by encoding the atlas information is also called the atlas sub-bitstream.
[0037] In the following, the point cloud (object) is assumed to be able to change in the time direction, like a moving image of a 2D image. That is, the geometry data and the attribute data have the concept of time direction and are data sampled at predetermined times like a moving image of a 2D image. Note that, like the video frames of a 2D image, the data at each sampling time is referred to as a frame. That is, the point cloud data (geometry data and attribute data) is assumed to be composed of a plurality of frames like a moving image of a 2D image. In the present disclosure, this frame of the point cloud is also referred to as a point cloud frame. In the case of V-PCC, even for such a point cloud of a moving image (a plurality of frames), by video-framing each point cloud frame to form a video sequence, it can be encoded with high efficiency using the encoding method of a moving image.
[0038] <Structure of V3C bitstream> The encoder multiplexes the encoded data of the geometry video frame, the attribute video frame, the occupancy map, and the atlas information as described above to generate one bitstream. This bitstream is also referred to as a V3C bitstream (V3C Bitstream).
[0039] FIG. 2 is a diagram showing a structural example of a V3C sample stream, which is one format of a V3C bitstream. As shown in FIG. 2, the V3C bitstream (V3C sample stream), which is an encoded stream of V-PCC, includes a plurality of V3C units (V3C unit).
[0040] A V3C unit includes a V3C unit header and a V3C unit payload. The V3C unit header includes information indicating the type of information stored in the V3C unit payload. Depending on the type stored in the V3C unit header, the V3C unit payload may store an attribute video sub-bitstream, a geometry video sub-bitstream, an occupancy video sub-bitstream, an atlas sub-bitstream, or the like.
[0041] <Atlas sub-bitstream structure> A in Fig. 3 is a diagram showing an example of the main structure of an atlas sub-bitstream. As shown in A in Fig. 3, an atlas sub-bitstream 31 is made up of a series of atlas NAL units 32. Each square in A in Fig. 3 represents an atlas NAL unit 32.
[0042] aud is the NAL unit for the access unit delimiter, atlas sps is the NAL unit for the atlas sequence parameter set, atlas fps is the NAL unit for the atlas frame parameter set, and atlas aps is the NAL unit for the atlas adaptation parameter set.
[0043] An atlas tile layer NAL unit is a NAL unit of an atlas tile layer. The atlas tile layer NAL unit has atlas tile information, which is information about an atlas tile. One atlas tile layer NAL unit has the information of one atlas tile. That is, the atlas tile layer NAL unit and the atlas tile correspond one-to-one.
[0044] The atlas fps stores the in-frame position information of an atlas tile, and the position information is linked to the atlas tile layer NAL unit via an id.
[0045] Atlas tiles are independently decodable from each other and have 2D-3D conversion information for patches in the corresponding rectangular region of a video sub-bitstream. The 2D-3D conversion information is information for converting a patch, which is 2D data, into a point cloud, which is 3D data. For example, an attribute video frame 12 shown in B of FIG. 3 is divided as shown by the dotted line, and a rectangular atlas tile 33 is formed.
[0046] The encoding of this atlas tile has constraints equivalent to those of a tile in HEVC. For example, it is configured not to depend on other atlas tiles in the same frame. Also, atlas frames with a reference relationship have the same atlas tile partitioning as each other. Furthermore, only the atlas tiles at the same position in the reference frame are referenced.
[0047] <Storage method to ISOBMFF> Non-Patent Document 4 specifies two types of methods for storing a V3C bitstream in ISOBMFF (International Organization for Standardization Base Media File Format): a multi-track structure and a single track structure.
[0048] The multi-track structure is a method of storing the geometry video sub-bitstream, attribute video sub-bitstream, occupancy video sub-bitstream, and atlas sub-bitstream in separate tracks. Since each video sub-bitstream is a conventional 2D video stream, it can be stored (managed) using the same method as for 2D. An example of a file structure when the multi-track structure is applied is shown in Figure 4.
[0049] The single track structure is a method of storing a V-PCC bitstream in one track, that is, in this case, the geometry video sub-bitstream, the attribute video sub-bitstream, the occupancy map video sub-bitstream, and the atlas sub-bitstream are stored in the same track.
[0050] <Partial Access> Non-Patent Document 4 defines partial access information for acquiring and decoding a portion of a point cloud object. For example, by using this partial access information, it becomes possible to control the acquisition of only the information of the display portion of a point cloud object during streaming distribution. This control can achieve the effect of effectively utilizing bandwidth and achieving high definition.
[0051] For example, suppose that a bounding box 51, which is a three-dimensional area that contains an object in a point cloud, is set for the object, as shown in A of Fig. 5. That is, in ISOBMFF, bounding box information (3DBoundingBoxStruct), which is information about the bounding box 51, is set, as shown in B of Fig. 5. In the bounding box information, the coordinates of the reference point (orgin) of the bounding box 51 are (0, 0, 0), and the size of the bounding box 51 is specified as (bb_dx, bb_dy, bb_dz).
[0052] By setting the partial access information, it is possible to set a 3D spatial region 52, which is an independently decodable partial region, within this bounding box 51, as shown in A of Fig. 5. That is, as shown in B of Fig. 5, 3D spatial region information (3dSpatialRegionStruct), which is information related to the 3D spatial region 52, is set as partial access information in ISOBMFF. In the 3D spatial region information, the region is specified by the coordinates (x, y, z) of its reference point and its size (cuboid_dx, cuboid_dy, cuboid_dz).
[0053] <File structure example> For example, it is assumed that the bitstream of object 61 in Fig. 6 is divided into three 3D spatial regions (3D spatial region 61A, 3D spatial region 61B, and 3D spatial region 61C) and stored in ISOBMFF. It is also assumed that a multi-track structure is applied and the 3D spatial region information is static (does not change in the time direction).
[0054] In this case, as shown on the right side of Fig. 6, the video sub-bitstreams are stored separately for each 3D spatial region (in different tracks). Then, tracks storing the geometry video sub-bitstream, attribute video sub-bitstream, and occupancy video sub-bitstream corresponding to the same 3D spatial region are grouped (indicated by the dotted line in Fig. 6). This group is also called a spatial region track group.
[0055] The video sub-bitstream of one 3D spatial region is stored in one or more spatial region track groups. In the example of Figure 6, three 3D spatial regions are configured, so three or more spatial region track groups are formed.
[0056] Each spatial region track group is assigned a track_group_id as track group identification information that identifies the spatial region track group. This track_group_id is stored in each track. In other words, tracks that belong to the same spatial region track group store the same track_group_id value. Therefore, it is possible to identify tracks that belong to a desired spatial region track group based on the value of this track_group_id.
[0057] In other words, the tracks storing the geometry video sub-bitstream, attribute video sub-bitstream, and occupancy video sub-bitstream corresponding to the same 3D spatial region each store the same track_group_id value. Therefore, based on the value of track_group_id, each video sub-bitstream corresponding to a desired 3D spatial region can be identified.
[0058] More specifically, tracks belonging to the same spatial region track group store spatial region group boxes (SpatialRegionGroupBox) with the same track_group_id, as shown in Fig. 7. The track_group_id is stored in a track group type box (TrackGroupTypeBox) that the spatial region group box inherits.
[0059] Note that atlas sub-bitstreams are stored in one V3C track regardless of the 3D spatial region. In other words, this one atlas sub-bitstream has 2D / 3D transformation information for patches of multiple 3D spatial regions. More specifically, as shown in Figure 7, a V3C spatial region box (V3CSpatialRegionsBox) is stored in the V3C track where the atlas sub-bitstream is stored, and each track_group_id is stored in that V3C spatial region box.
[0060] The atlas tile and the spatial region track group are linked by the NALUMapEntry sample group described in Non-Patent Document 3.
[0061] If the 3D spatial region information is dynamic (changes in the time direction), the 3D spatial region at each time can be represented using a timed metadata track, as shown in A of Fig. 8. That is, as shown in B of Fig. 8, a dynamic 3D spatial region sample entry (Dynamic3DSpatialRegionSampleEntry) and a dynamic spatial region sample (DynamicSpatialRegionSample) are stored in the ISOBMFF.
[0062] <Scalability of V-PCC Symbolization> In V-PCC symbolization, for example, as described above, by using the volumetric annotation SEI message family, region-based scalability can be realized, where only a partial point cloud at a specific 3D spatial position can be decoded and rendered.
[0063] Also, as described in Non-Patent Document 5, by using the LoD patch mode, spatial scalability can be realized, where only the point group of the point cloud corresponding to a specific LoD can be decoded and rendered.
[0064] LoD indicates the hierarchy when hierarchically classifying a point cloud object by point density. For example, the points of the point cloud are grouped (hierarchically classified) so that multiple hierarchies with different point densities (from a sparse point hierarchy to a dense point hierarchy) are formed, such as an octree using voxel quantization. Each hierarchy of such a hierarchical structure is also referred to as LoD.
[0065] The point cloud objects constructed at each LoD represent the same object, but their resolutions (number of points) are different from each other. That is, this hierarchical structure can also be said to be a hierarchical structure based on the resolution of the point cloud.
[0066] <LoD Patch Mode> In the LoD patch mode, the point cloud is encoded so that the client can decode the low LoD (sparse) point cloud that constitutes the high LoD (dense) point cloud alone and construct the low LoD point cloud.
[0067] In other words, by grouping each point as described above, the point cloud with the original density (dense point cloud) is divided into multiple sparse point clouds. The densities of these sparse point clouds may or may not be the same. Point clouds at each level of the hierarchical structure described above can be realized by using a single sparse point cloud or by combining multiple sparse point clouds. For example, the point cloud with the original density can be restored by combining all the sparse point clouds.
[0068] In LoD patch mode, this layering of point clouds can be performed for each patch. Then, the point spacing of the sparse point cloud patches can be scaled to a dense state (the point spacing of the original point cloud) and encoded. For example, as shown in A of Figure 9, downscaling the point spacing allows encoding as dense patches (small patches). This makes it possible to suppress the reduction in encoding efficiency due to layering.
[0069] For example, as shown in Figure 9B, sparse patches (large patches) can be restored by upscaling the point spacing by the same ratio as when encoding.
[0070] In this case, a scaling factor, which is information related to such scaling, is transmitted from the encoding side to the decoding side for each patch. That is, the scaling factor is stored in the V3C bitstream. FIG. 10 is a diagram showing an example of the syntax of this scaling factor. In FIG. 10, pdu_lod_scale_x_minus1[patchIndex] indicates the conversion ratio of downscaling in the x direction for each patch, and pdu_lod_scale_y[patchIndex] indicates the conversion ratio of downscaling in the y direction for each patch. During decoding, upscaling is performed using the conversion ratio indicated by these parameters (i.e., upscaling is performed based on the scaling factor), making it easy to upscale using the same conversion ratio as during encoding (downscaling).
[0071] As described above, in the LoD patch mode, a single point cloud is divided into multiple sparse point clouds that each represent the same object and then encoded. This division is performed for each patch. That is, as shown in FIG. 11, a patch is divided into multiple patches composed of points of different sample grids. Patch P0 shown on the left side of FIG. 11 represents the original patch (dense patch), and each circle represents a point that constitutes that patch. In FIG. 11, the points of this patch P0 are grouped into four types of sample grids to form four sparse patches. That is, white points, black points, gray points, and diagonal points are extracted from patch P0 and divided into different sparse patches. In this case, four sparse patches are formed, each with half the point density (double the point spacing) of the original patch P0 in both the x and y directions. By using a single sparse patch or combining multiple such sparse patches, spatial scalability (resolution scalability) can be achieved.
[0072] As described above, sparse patches are downscaled during encoding. In LoD patch mode, this division is performed on each of the original dense patches. The divided sparse patches are then grouped into atlas tiles for each sample grid when placed on a frame image. For example, in FIG. 11, a sparse patch consisting of points indicated by white circles is placed in "atlas tile 0," a sparse patch consisting of points indicated by black circles is placed in "atlas tile 1," a sparse patch consisting of points indicated by gray circles is placed in "atlas tile 2," and a sparse patch consisting of points indicated by diagonally shaded circles is placed in "atlas tile 3." By dividing the atlas tiles in this way, patches corresponding to the same sample grid can be decoded independently of each other. In other words, a point cloud can be constructed for each sample grid. This enables spatial scalability.
[0073] <No support for spatial scalability> As described above, MPEG-I Part 5 Visual Volumetric Video-based Coding (V3C) and Video-based Point Cloud Compression (V-PCC) encode in LoD patch mode, allowing the client to independently decode the low-LoD (sparse) point clouds that make up a high-LoD (dense) point cloud and construct a low-LoD point cloud.
[0074] By utilizing such spatial scalability, it becomes possible to obtain V-PCC content with an appropriate LoD depending on the network bandwidth limitations and fluctuations during V-PCC content distribution, the decoding and rendering performance of the client device, etc. Therefore, it is desirable for MPEG-I part 10 to support distribution using spatial scalability.
[0075] However, the ISOBMFF described in Non-Patent Document 4, which stores the V3C bitstream, does not support spatial scalability, and it is difficult to store information about spatial scalability in the system layer as information separate from the V3C bitstream. As a result, when distributing V-PCC content, the client is unable to identify the combination of point clouds that provide spatial scalability, and is unable to select a point cloud with an appropriate LoD according to the client environment.
[0076] In order for a client to use this spatial scalability to construct 3D data at the desired LoD, it was necessary to perform tedious tasks such as parsing the V3C bitstream (atlas sub-bitstream) down to the patch data unit (patch_data_unit).
[0077] 2. First Embodiment <Files that store bitstream and spatial scalability information> Therefore, information about spatial scalability is stored in a file (for example, ISOBMFF) that stores a V3C bitstream as information separate from the V3C bitstream (stored in the system layer).
[0078] For example, in an information processing method, a point cloud that represents a three-dimensional object as a collection of points is encoded as 2D data of the point cloud that corresponds to spatial scalability, a bitstream including sub-bitstreams in which the point cloud corresponding to one or more layers of the spatial scalability is encoded is generated, spatial scalability information regarding the spatial scalability of the sub-bitstream is generated, and a file that stores the generated bitstream and spatial scalability information is generated.
[0079] For example, an information processing device may include an encoding unit that encodes 2D data obtained by two-dimensionally converting a point cloud that represents a three-dimensional object as a collection of points and that corresponds to spatial scalability, and generates a bitstream including a sub-bitstream in which the point cloud corresponding to one or more layers of the spatial scalability is encoded; a spatial scalability information generation unit that generates spatial scalability information related to the spatial scalability of the sub-bitstream; and a file generation unit that generates a file that stores the bitstream generated by the encoding unit and the spatial scalability information generated by the spatial scalability information generation unit.
[0080] For example, as shown in Fig. 12, a spatial scalability infostruct is newly defined and stored in the VPCC spatial region box (VPCCspatialRegionsBox) of the sample entry (SampleEntry). Then, spatial scalability information is stored in the spatial scalability infostruct.
[0081] This allows spatial scalability information to be provided at the system layer to a client device that decodes a V3C bitstream, allowing the client device to more easily play back 3D data using spatial scalability without the need for complex tasks such as parsing the V3C bitstream.
[0082] For example, in an information processing method, a spatial scalability layer to be decoded is selected based on spatial scalability information regarding spatial scalability of a bitstream in which 2D data obtained by two-dimensionally converting a point cloud corresponding to spatial scalability stored in a file, the point cloud representing a three-dimensional object as a collection of points, is encoded; a sub-bitstream corresponding to the selected layer is extracted from the bitstream stored in the file; and the extracted sub-bitstream is decoded.
[0083] For example, an information processing device may include a selection unit that selects a spatial scalability layer to decode based on spatial scalability information regarding spatial scalability of a bitstream in which 2D data obtained by two-dimensionally converting a point cloud corresponding to spatial scalability stored in a file, the point cloud representing a three-dimensional object as a collection of points, an extraction unit that extracts a sub-bitstream corresponding to the layer selected by the selection unit from the bitstream stored in the file, and a decoding unit that decodes the sub-bitstream extracted by the extraction unit.
[0084] For example, as shown in Figure 12, in an ISOBMFF that stores the V3C bitstream described in non-patent document 4, the spatial scalability layer to be decoded is selected based on the spatial scalability information stored in the spatial scalability infostruct in the VPCC spatial region box of the sample entry.
[0085] In this way, the client device can identify a combination of point clouds that provides spatial scalability based on the spatial scalability information stored in the system layer. Therefore, for example, in V-PCC content distribution, the client device can select a point cloud with an appropriate LoD according to the client environment without requiring cumbersome operations such as analyzing the bitstream.
[0086] For example, a client device can control the acquisition of a portion of a point cloud close to the viewpoint at a high LoD and other portions farther away at a low LoD, thereby enabling the client device to more effectively utilize bandwidth even under limited network bandwidth conditions and provide users with a high-quality media experience.
[0087] This means that client devices can more easily play back 3D data using spatial scalability.
[0088] For example, as this spatial scalability information, base-enhancement grouping information that specifies the selection order (layer) of each group (sparse patch), such as which group (sparse patch) is the base layer and which group (sparse patch) is the enhancement layer, may be stored in the system layer.
[0089] This allows the client device to easily determine which groups are required to construct a point cloud with a desired LoD based on the spatial scalability information, and therefore the client device can more easily select a point cloud with an appropriate LoD.
[0090] Also, as shown in Fig. 12, the bitstreams of each layer (group) may be stored in different tracks (spatial region track groups) of the ISOBMFF. In the example of Fig. 12, the sparse patches made up of points indicated by white circles, the sparse patches made up of points indicated by black circles, the sparse patches made up of points indicated by gray circles, and the sparse patches made up of points indicated by diagonally shaded circles are stored in different spatial region track groups. In this way, the client device can select the bitstream of a desired layer (group) by selecting the track (spatial region track group) to decode. In other words, the client device can more easily obtain and decode the bitstream of a desired layer (group).
[0091] <Example of spatial scalability information> A in Fig. 13 is a diagram showing an example of the syntax of a VPCC spatial region box (VPCC spatial RegionsBox). In the example of A in Fig. 13, a spatial scalability infostruct (SpatialScalabilityInfoStruct()) is stored for each region in the VPCC spatial region box.
[0092] B of Fig. 13 is a diagram showing an example of the syntax of a spatial scalability info struct (SpatialScalabilityInfoStruct()). As shown in B of Fig. 13, layer identification information (layer_id) may be stored as spatial scalability information in this spatial scalability info struct. The layer identification information is identification information indicating a layer of ISOBMFF to which a sub-bitstream stored in a track group to which the spatial scalability info struct corresponds. For example, layer_id = 0 indicates a base layer, and layer_id = 1 to 255 indicates an enhancement layer.
[0093] That is, the file generation device may store this layer identification information (layer_id) in the spatial scalability infostruct, and the client device may select a sub-bitstream (track) based on this layer identification information. In this manner, the client device can determine the layer to which the sub-bitstream (sparse patch) stored in each track (spatial region track group) corresponds based on this layer identification information. Therefore, the client device can more easily select a point cloud that achieves high resolution in the order intended by the content creator. Therefore, the client device can more easily play back 3D data using spatial scalability.
[0094] In addition to this layer identification information, information (lod) regarding the resolution of the point cloud obtained by reconstructing the point cloud corresponding to each layer from the top layer of the spatial scalability to the layer indicated by the layer identification information may also be stored in the spatial scalability infostruct, as shown in B of Figure 13.
[0095] For example, when layer_id = 0, the information (lod) about the resolution of this point cloud indicates the LoD value of the base layer. Furthermore, when layer_id = not 0, the information (lod) about the resolution of this point cloud indicates the LoD value obtained by simultaneously displaying the point clouds of layers 0 to (layer_id-1). This LoD value may be a guideline value determined by the content creator. Alternatively, the information (lod) about the resolution of the point cloud may not be signaled, and the value of layer_id may signal information about the resolution of the point cloud. In other words, the information (lod) about the resolution of the point cloud may be included in the layer identification information (layer_id). For example, the value of layer_id may also indicate the resolution (lod value) of the point cloud corresponding to each layer from the highest layer of spatial scalability to the layer indicated by the layer identification information.
[0096] That is, the file generation device may store information about the resolution of the point cloud (LoD) in the spatial scalability infostruct, and the client device may select a sub-bitstream (track) based on the information about the resolution of the point cloud. This allows the client device to more easily determine which track (spatial region track group) to select to obtain the desired LoD. Therefore, the client device can more easily play back 3D data using spatial scalability.
[0097] Furthermore, as shown in B of Fig. 13, in addition to this layer identification information, spatial scalability identification information (spatial_scalability_id) that identifies spatial scalability may be stored in the spatial scalability infostruct. A group of regions (each loop in the for loop of num_region corresponds to one region) that have the same spatial scalability identification information (spatial_scalability_id) provides spatial scalability. In other words, combining multiple regions that have the same spatial scalability identification information (spatial_scalability_id) can result in a high-LoD point cloud.
[0098] That is, the file generation device may store this spatial scalability identification information (spatial_scalability_id) in the spatial scalability infostruct, and the client device may select a sub-bitstream (track) based on this spatial scalability identification information. In this way, the client device can more easily identify a group that provides spatial scalability. Therefore, the client device can more easily play back 3D data using spatial scalability.
[0099] As shown in A of Fig. 13, a spatial scalability flag (spatial_scalability_flag) may be stored in the VPCC spatial region box. The spatial scalability flag is flag information indicating whether or not a spatial scalability infostruct is stored. If the spatial scalability flag is true (for example, "1"), it indicates that a spatial scalability infostruct is stored. If the spatial scalability flag is false (for example, "0"), it indicates that a spatial scalability infostruct is not stored.
[0100] <Other examples of spatial scalability information> In addition, as shown in the example of A in Figure 14, in the VPCC spatial region box (VPCC spatial region box), for each region, the spatial scalability info struct (SpatialScalabilityInfoStruct()) and track group identification information (track_group_id) may be stored using a for loop for the number of layers.
[0101] In this case, the group stored by this for loop provides spatial scalability. In other words, the for loop groups together spatial scalability infostructs that provide the same spatial scalability. Therefore, in this case, there is no need to store spatial scalability identification information (spatial_scalability_id). In other words, spatial scalability can be identified without the need to store spatial scalability identification information.
[0102] An example of the syntax of the spatial scalability infostruct (SpatialScalabilityInfoStruct()) in this case is shown in B of Fig. 14. In the example of B of Fig. 14, the spatial scalability infostruct stores the above-mentioned layer identification information (layer_id) and information about the resolution of the point cloud (lod).
[0103] In this case, too, a spatial scalability flag (spatial_scalability_flag) may be stored in the VPCC spatial region box, as shown in A of FIG.
[0104] <Matryoshka Media Container> Although the above describes an example in which ISOBMFF is used as the file format, the file in which the V3C bitstream is stored may be any format other than ISOBMFF. For example, the V3C bitstream may be stored in a Matroska Media Container. An example of the main configuration of a Matroska Media Container is shown in Figure 15.
[0105] For example, spatial scalability information (or base and enhancement point cloud information) may be stored in an element under the Track Entry element of the track that stores the atlas sub-bitstream.
[0106] <File generation device> Fig. 16 is a block diagram showing an example of the configuration of a file generation device, which is one aspect of an information processing device to which the present technology is applied. The file generation device 300 shown in Fig. 16 is a device that applies V-PCC to encode point cloud data as video frames using an encoding method for two-dimensional images. The file generation device 300 also generates ISOBMFF and stores the V3C bitstream generated by the encoding.
[0107] At this time, the file generation device 300 applies the present technology described above in the present embodiment and stores information in the ISOBMFF so as to enable spatial scalability. That is, the file generation device 300 stores information related to spatial scalability in the ISOBMFF.
[0108] Note that Fig. 16 shows the main processing units, data flows, etc., and is not necessarily all that is shown in Fig. 16. In other words, in file generation device 300, there may be processing units that are not shown as blocks in Fig. 16, and there may be processing and data flows that are not shown as arrows, etc. in Fig. 16.
[0109] As shown in FIG. 16, the file generation device 300 includes a 3D2D conversion unit 301, a 2D encoding unit 302, a metadata generation unit 303, a PC stream generation unit 304, a file generation unit 305, and an output unit 306.
[0110] The 3D2D conversion unit 301 decomposes a point cloud, which is 3D data input to the file generation device 300, into patches and packs them. That is, the 3D2D conversion unit 301 generates a geometry video frame, an attribute video frame, and an occupancy video frame. In this case, as described with reference to, for example, FIGS. 11 and 12, the 3D2D conversion unit 301 divides the point cloud into multiple sparse point clouds and arranges each patch in a frame image so as to group the atlases into atlas tiles for each sample grid (for each patch that provides the same spatial scalability). The 3D2D conversion unit 301 also generates atlas information. The 3D2D conversion unit 301 supplies the generated geometry video frames, attribute video frames, occupancy video frames, atlas information, etc. to the 2D encoding unit 302.
[0111] The 2D encoding unit 302 performs encoding-related processing. For example, the 2D encoding unit 302 acquires geometry video frames, attribute video frames, occupancy video frames, atlas information, and the like supplied from the 3D2D conversion unit 301. The 2D encoding unit 302 encodes these frames and generates sub-bitstreams. For example, the 2D encoding unit 302 includes encoding units 311 to 314. The encoding unit 311 encodes geometry video frames and generates a geometry video sub-bitstream. The encoding unit 312 encodes attribute video frames and generates an attribute video sub-bitstream. The encoding unit 313 encodes occupancy video frames and generates an occupancy video sub-bitstream. The encoding unit 314 encodes atlas information and generates an atlas sub-bitstream.
[0112] In this case, the 2D encoding unit 302 applies the LoD patch mode, encodes each patch information of the sparse point cloud as an atlas tile, and generates one atlas sub-bitstream.The 2D encoding unit 302 also applies the LoD patch mode, encodes three images (geometry image, attribute image, and occupancy map) for each sparse point cloud, and generates a geometry video sub-bitstream, an attribute video sub-bitstream, and an occupancy video sub-bitstream.
[0113] The 2D encoding unit 302 supplies the generated sub-bitstreams to the metadata generation unit 303 and the PC stream generation unit 304. For example, the encoding unit 311 supplies the generated geometry video sub-bitstream to the metadata generation unit 303 and the PC stream generation unit 304. The encoding unit 312 supplies the generated attribute video sub-bitstream to the metadata generation unit 303 and the PC stream generation unit 304. The encoding unit 313 supplies the generated occupancy video sub-bitstream to the metadata generation unit 303 and the PC stream generation unit 304. The encoding unit 314 supplies the generated atlas sub-bitstream to the metadata generation unit 303 and the PC stream generation unit 304.
[0114] The metadata generation unit 303 performs processing related to the generation of metadata. For example, the metadata generation unit 303 acquires the video sub-bitstream and the atlas sub-bitstream supplied from the 2D encoding unit 302. The metadata generation unit 303 also generates metadata using this data.
[0115] For example, the metadata generation unit 303 generates, as metadata, spatial scalability information related to the spatial scalability of the acquired sub-bitstream. That is, the metadata generation unit 303 generates the spatial scalability information using any one of the various methods described with reference to Figures 12 to 14, etc., or by appropriately combining any two or more of the methods. Note that the metadata generation unit 303 may generate any metadata other than spatial scalability information.
[0116] After generating the metadata including the spatial scalability information in this way, the metadata generation unit 303 supplies the metadata to the file generation unit 305 .
[0117] The PC stream generation unit 304 performs processing related to the generation of a V3C bitstream. For example, the PC stream generation unit 304 acquires the video sub-bitstream and the atlas sub-bitstream supplied from the 2D encoding unit 302. The PC stream generation unit 304 also uses these to generate a V3C bitstream (a geometry video sub-bitstream, an attribute video sub-bitstream, an occupancy map video sub-bitstream, and an atlas sub-bitstream, or a combination of these), and supplies this to the file generation unit 305.
[0118] The file generation unit 305 performs processing related to file generation. For example, the file generation unit 305 acquires metadata including spatial scalability information supplied from the metadata generation unit 303. The file generation unit 305 also acquires a V3C bitstream supplied from the PC stream generation unit 304. The file generation unit 305 generates a file (e.g., ISOBMFF or Matryoshka Media Container) that stores the acquired metadata and V3C bitstream. In other words, the file generation unit 305 stores the spatial scalability information in a file separate from the V3C bitstream. In other words, the file generation unit 305 stores the spatial scalability information in a system layer.
[0119] At this time, the file generation unit 305 stores the spatial scalability information in a file using any one of the various methods described with reference to Figures 12 to 14, etc., or using an appropriate combination of any two or more of the methods. For example, the file generation unit 305 stores the spatial scalability information in the locations shown in the examples of Figures 12 to 14 in the file that stores the V3C bitstream.
[0120] The file generation unit 305 supplies the generated file to the output unit 306. The output unit 306 outputs the supplied file (a file including a V3C bitstream and spatial scalability information) to an external device (such as a distribution server) outside the file generation device 300.
[0121] As described above, the file generation device 300 applies the technology described above in this embodiment to generate a file (for example, ISOBMFF or Matryoshka Media Container) that stores a V3C bitstream and spatial scalability information.
[0122] This configuration allows spatial scalability information to be provided at the system layer to a client device that decodes a 3C bitstream, allowing the client device to more easily play back 3D data using spatial scalability without the need for complex tasks such as parsing the V3C bitstream.
[0123] Note that these processing units (the 3D2D conversion unit 301 to the output unit 306 and the encoding units 311 to 314) may have any configuration. For example, each processing unit may be configured with a logic circuit that realizes the above-mentioned processing. Furthermore, each processing unit may have, for example, a central processing unit (CPU), read-only memory (ROM), random access memory (RAM), etc., and may realize the above-mentioned processing by executing a program using these. Of course, each processing unit may have both of these configurations, and may realize some of the above-mentioned processing by a logic circuit and other by executing a program. The configurations of each processing unit may be independent of each other. For example, some processing units may realize some of the above-mentioned processing by a logic circuit, other processing units may realize the above-mentioned processing by executing a program, and still other processing units may realize the above-mentioned processing by both a logic circuit and by executing a program.
[0124] <File generation process flow> An example of the flow of the file generation process executed by this file generation device 300 will be described with reference to the flowchart of FIG.
[0125] When the file generation process starts, in step S301, the 3D2D conversion unit 301 of the file generation device 300 divides the point cloud into multiple sparse point clouds. In step S302, the 3D2D conversion unit 301 decomposes the point cloud into patches and generates patches of geometry and attributes. The 3D2D conversion unit 301 then packs the patches into a video frame. The 3D2D conversion unit 301 also generates an occupancy map and atlas information.
[0126] In step S303, the 2D encoding unit 302 applies the LoD patch mode, encodes each patch information of the sparse point cloud as an atlas tile, and generates one atlas sub-bitstream.
[0127] In step S304, the 2D encoding unit 302 encodes three images (geometry video frame, attribute video frame, and occupancy map video frame) for each sparse point cloud, respectively, to generate a geometry video sub-bitstream, an attribute video sub-bitstream, and an occupancy video sub-bitstream.
[0128] The PC stream generation unit 304 generates a V3C bitstream (point cloud stream) using the video sub-bitstream, atlas sub-bitstream, and the like.
[0129] In step S305, the metadata generation unit 303 generates metadata including spatial scalability information. That is, the metadata generation unit 303 generates the spatial scalability information by using any single method or an appropriate combination of any multiple methods among the various methods described with reference to Figures 12 to 14, etc. For example, the metadata generation unit 303 generates base enhancement point cloud information as the spatial scalability information.
[0130] In step S306, the file generation unit 305 generates a file, such as an ISOBMFF or Matryoshka media container, and stores the spatial scalability information and the V3C bitstream in the file. The file generation unit 305 stores the spatial scalability information in the file using any one of the various methods described with reference to Figures 12 to 14, or a combination of any two or more of the methods. For example, the file generation unit 305 stores the base enhancement point cloud information generated in step S305 in the file.
[0131] In step S307, the output unit 306 outputs the file generated in step S306, i.e., the file storing the V3C bitstream and spatial scalability information, to an external device (e.g., a distribution server) outside the file generation device 300. When the processing of step S307 ends, the file generation processing ends.
[0132] By performing these processes, spatial scalability information can be provided at the system layer to a client device that decodes a 3C bitstream, allowing the client device to more easily play back 3D data using spatial scalability without the need for complex tasks such as parsing the V3C bitstream.
[0133] <Client device> The present technology described in this embodiment can be applied not only to file generation devices but also to client devices. FIG. 18 is a block diagram showing an example of the configuration of a client device, which is one aspect of an information processing device to which the present technology is applied. The client device 400 shown in FIG. 18 is a device that applies V-PCC, acquires a V3C bitstream (a geometry video sub-bitstream, an attribute video sub-bitstream, an occupancy video sub-bitstream, and an atlas sub-bitstream, or a combination of these) encoded using a two-dimensional image encoding method with point cloud data as video frames from a file, decodes the V3C bitstream using a two-dimensional image decoding method, and generates (reconstructs) a point cloud. For example, the client device 400 can extract the V3C bitstream from the file generated by the file generation device 300, decode it, and generate a point cloud.
[0134] In this case, the client device 400 realizes spatial scalability by using any one of the various methods of the present technology described above in this embodiment, or by appropriately combining any two or more of the methods. That is, the client device 400 selects and decodes the bitstream (track) required to reconstruct a point cloud of a desired LoD based on the spatial scalability information stored in the file together with the V3C bitstream.
[0135] Note that Fig. 18 shows the main processing units, data flows, etc., and is not necessarily all that is shown in Fig. 18. In other words, in client device 400, there may be processing units that are not shown as blocks in Fig. 18, and there may be processing or data flows that are not shown as arrows, etc. in Fig. 18.
[0136] As shown in FIG. 18, the client device 400 includes a file processing unit 401, a 2D decoding unit 402, a display information generation unit 403, and a display unit 404.
[0137] The file processing unit 401 extracts a V3C bitstream (sub-bitstream) from a file input to the client device 400 and supplies it to the 2D decoding unit 402. At this time, the file processing unit 401 applies the present technology described in this embodiment and extracts a V3C bitstream (sub-bitstream) of a layer corresponding to a desired LoD, etc., based on spatial scalability information stored in the file. Then, the file processing unit 401 supplies the extracted V3C bitstream to the 2D decoding unit 402.
[0138] In other words, the file processing unit 401 excludes, from the decoding target, the V3C bitstreams of layers that are not required for constructing a point cloud of the desired LoD, based on the spatial scalability information.
[0139] The file processing unit 401 includes a file acquisition unit 411 , a file analysis unit 412 , and an extraction unit 413 .
[0140] The file acquisition unit 411 acquires a file input to the client device 400. As described above, this file stores a V3C bitstream and spatial scalability information. For example, this file is an ISOBMFF or Matryoshka media container. The file acquisition unit 411 supplies the acquired file to the file analysis unit 412.
[0141] The file analysis unit 412 acquires the file provided from the file acquisition unit 411. The file analysis unit 412 analyzes the acquired file. In doing so, the file analysis unit 412 analyzes the file using any single method or a suitable combination of any multiple methods among the various methods of the present technology described in this embodiment. For example, the file analysis unit 412 analyzes spatial scalability information stored in the file and selects sub-bitstreams to be decoded. For example, the file analysis unit 412 selects a combination of point clouds that provides spatial scalability according to the network environment and the processing capabilities of the client device 400 itself, based on the spatial scalability information (i.e., sub-bitstreams to be decoded). The file analysis unit 412 provides the analysis result, along with the file, to the extraction unit 413.
[0142] The extraction unit 413 extracts data to be decoded from the V3C bitstream stored in the file based on the analysis result by the file analysis unit 412. That is, the extraction unit 413 extracts the sub-bitstream selected by the file analysis unit 412. The extraction unit 413 supplies the extracted data to the 2D decoding unit 402.
[0143] The 2D decoding unit 402 performs decoding-related processing. For example, the 2D decoding unit 402 acquires a geometry video sub-bitstream, an attribute video sub-bitstream, an occupancy video sub-bitstream, an atlas sub-bitstream, and the like supplied from the file processing unit 401. The 2D decoding unit 402 decodes these to generate video frames and atlas information. For example, the 2D decoding unit 402 includes decoding units 421 to 424. The decoding unit 421 decodes the supplied geometry video sub-bitstream to generate geometry video frames (2D data). The decoding unit 422 decodes the attribute video sub-bitstream to generate attribute video frames (2D data). The decoding unit 423 decodes the occupancy video sub-bitstream to generate occupancy video frames (2D data). The decoding unit 424 decodes the atlas sub-bitstream to generate atlas information corresponding to the above-mentioned video frames.
[0144] The 2D decoding unit 402 supplies the generated bitstream to the display information generation unit 403. For example, the decoding unit 421 supplies the generated geometry video frame to the display information generation unit 403. The decoding unit 422 supplies the generated attribute video frame to the display information generation unit 403. The decoding unit 423 supplies the generated occupancy video frame to the display information generation unit 403. The decoding unit 424 supplies the generated atlas information to the display information generation unit 403.
[0145] The display information generation unit 403 performs processing related to the construction and rendering of a point cloud. For example, the display information generation unit 403 acquires video frames and atlas information supplied from the 2D decoding unit 402. Furthermore, the display information generation unit 403 generates a point cloud from patches packed in the acquired video frames based on the acquired atlas information. The display information generation unit 403 then renders the point cloud to generate a display image and supplies it to the display unit 404.
[0146] The display information generation unit 403 includes, for example, a 2D / 3D conversion unit 431 and a display processing unit 432.
[0147] The 2D3D conversion unit 431 converts, into a point cloud (3D data), patches (2D data) arranged in the video frames supplied from the 2D decoding unit 402. The 2D3D conversion unit 431 supplies the generated point cloud to the display processing unit 432.
[0148] The display processing unit 432 performs processing related to rendering. For example, the display processing unit 432 acquires a point cloud supplied from the 2D3D conversion unit 431. The display processing unit 432 also renders the acquired point cloud to generate an image for display. The display processing unit 432 supplies the generated image for display to the display unit 404.
[0149] The display unit 404 has a display device such as a monitor, and displays a display image. For example, the display unit 404 acquires a display image supplied from the display processing unit 432. The display unit 404 displays the display image on the display device and presents it to a user or the like.
[0150] With this configuration, the client device 400 can identify a combination of point clouds that provides spatial scalability based on the spatial scalability information stored in the system layer. Therefore, for example, in V-PCC content distribution, the client device can select a point cloud with an appropriate LoD according to the client environment without requiring complicated operations such as analyzing the bitstream. In other words, the client device can more easily play back 3D data using spatial scalability.
[0151] These processing units (file processing unit 401 to display unit 404, file acquisition unit 311 to extraction unit 413, decoding unit 421 to decoding unit 424, 2D3D conversion unit 431, and display processing unit 432) may have any configuration. For example, each processing unit may be configured with a logic circuit that realizes the above-mentioned processing. Furthermore, each processing unit may have, for example, a CPU, ROM, RAM, etc., and may execute a program using these to realize the above-mentioned processing. Of course, each processing unit may have both of these configurations, and may realize some of the above-mentioned processing using a logic circuit and other by executing a program. The configurations of each processing unit may be independent of each other. For example, some processing units may realize some of the above-mentioned processing using a logic circuit, other processing units may execute a program to realize the above-mentioned processing, and still other processing units may realize the above-mentioned processing using both a logic circuit and by executing a program.
[0152] <Client processing flow> An example of the flow of client processing executed by this client device 400 will be described with reference to the flowchart of FIG.
[0153] When client processing starts, in step S401, the file acquisition unit 411 of the client device 400 acquires a file to be provided to the client device 400. This file stores a V3C bitstream and spatial scalability information. For example, this file is an ISOBMFF or Matryoshka media container.
[0154] In step S402, the file analysis unit 412 selects a combination of point clouds that provides spatial scalability based on the network environment and the processing capabilities of the client device 400 itself, based on the spatial scalability information (e.g., base and enhancement point cloud information) stored in the file.
[0155] In step S403, the extraction unit 413 extracts, from the V3C bitstream stored in the file, the atlas sub-bitstream and the video sub-bitstreams corresponding to the multiple sparse point clouds selected in step S402.
[0156] In step S404, the 2D decoding unit 402 decodes the atlas sub-bitstream and video sub-bitstream extracted in step S403.
[0157] In step S405, the display information generation unit 403 constructs a point cloud based on the data obtained by decoding in step S403. That is, a point cloud of the desired LoD extracted from the file is constructed.
[0158] In step S406, the display information generation unit 403 renders the constructed point cloud and generates a display image.
[0159] In step S407, the display unit 404 displays the display image generated in step S406 on the display device.
[0160] When the process in step S407 is completed, the client process ends.
[0161] By performing the above processes, the client device 400 can identify a combination of point clouds that provides spatial scalability based on the spatial scalability information stored in the system layer. Therefore, for example, in V-PCC content distribution, the client device can select a point cloud with an appropriate LoD according to the client environment without requiring cumbersome tasks such as analyzing the bitstream. In other words, the client device can more easily play back 3D data using spatial scalability.
[0162] 3. Second Embodiment <Control file that stores spatial scalability information> This technology can also be applied to, for example, MPEG-DASH (Moving Picture Experts Group phase - Dynamic Adaptive Streaming over HTTP). For example, in MPEG-DASH, the MPD (Media Presentation Description), which is a control file that stores control information related to the distribution of bitstreams, may be extended to store spatial scalability information related to the spatial scalability of sub-bitstreams.
[0163] For example, in an information processing method, a point cloud that represents a three-dimensional object as a collection of points, where the point cloud corresponding to spatial scalability is two-dimensionalized, 2D data is encoded to generate a bitstream including sub-bitstreams in which the point cloud corresponding to one or more layers of the spatial scalability is encoded, spatial scalability information regarding the spatial scalability of the sub-bitstream is generated, and a control file is generated that stores the generated spatial scalability information and control information regarding distribution of the generated bitstream.
[0164] For example, an information processing device may include an encoding unit that encodes 2D data of a point cloud that represents a three-dimensional object as a collection of points, the point cloud corresponding to spatial scalability, and generates a bitstream including a sub-bitstream in which the point cloud corresponding to one or more layers of the spatial scalability is encoded; a spatial scalability information generation unit that generates spatial scalability information related to the spatial scalability of the sub-bitstream; and a control file generation unit that generates a control file that stores the spatial scalability information generated by the spatial scalability information generation unit and control information related to the distribution of the bitstream generated by the encoding unit.
[0165] For example, as shown in FIG. 20, the V3C3D region descriptor of the MPD may be extended to store spatial scalability information (for example, base and enhancement point cloud information).
[0166] By doing so, spatial scalability information can be provided in the system layer (MPD) to a client device that acquires a V3C bitstream to be decoded using this MPD. Therefore, the client device can more easily play back 3D data using spatial scalability without the need for complicated tasks such as analyzing the V3C bitstream.
[0167] For example, in an information processing method, a point cloud that represents a three-dimensional object as a collection of points is stored in a control file that stores control information regarding the distribution of a bitstream in which 2D data obtained by two-dimensionally converting the point cloud corresponding to spatial scalability is encoded.Based on spatial scalability information regarding the spatial scalability of the bitstream, a spatial scalability layer to be decoded is selected, a sub-bitstream corresponding to the selected layer is obtained, and the obtained sub-bitstream is decoded.
[0168] For example, an information processing device may include a selection unit that selects a spatial scalability layer to decode based on spatial scalability information regarding the spatial scalability of a bitstream stored in a control file that stores control information regarding the distribution of a bitstream in which 2D data obtained by two-dimensionally converting a point cloud that represents a three-dimensional object as a collection of points, an acquisition unit that acquires a sub-bitstream corresponding to the layer selected by the selection unit, and a decoding unit that decodes the sub-bitstream acquired by the acquisition unit.
[0169] For example, as shown in FIG. 20, the spatial scalability layer to be decoded may be selected based on spatial scalability information (e.g., base and enhancement point cloud information) stored in the V3C3D region descriptor of the MPD.
[0170] In this way, the client device can identify a combination of point clouds that provides spatial scalability based on the spatial scalability information stored in its system layer (MPD). Therefore, for example, in V-PCC content distribution, the client device can select a point cloud with an appropriate LoD according to the client environment without requiring complicated operations such as analyzing the bitstream.
[0171] For example, a client device can control the acquisition of a portion of a point cloud close to the viewpoint at a high LoD and other portions farther away at a low LoD, thereby enabling the client device to more effectively utilize bandwidth even under limited network bandwidth conditions and provide users with a high-quality media experience.
[0172] This means that client devices can more easily play back 3D data using spatial scalability.
[0173] For example, as this spatial scalability information, base-enhancement grouping information that specifies the selection order (layer) of each group (sparse patch), such as which group (sparse patch) is the base layer and which group (sparse patch) is the enhancement layer, may be stored in the system layer.
[0174] This allows the client device to easily determine which groups are required to construct a point cloud with a desired LoD based on the spatial scalability information, and therefore the client device can more easily select a point cloud with an appropriate LoD.
[0175] Also, as shown in Fig. 20, control information related to the distribution of the bitstream of each layer (group) may be stored in different adaptation sets of the MPD. In the example of Fig. 20, control information related to the sparse patch made up of points indicated by white circles, the sparse patch made up of points indicated by black circles, the sparse patch made up of points indicated by gray circles, and the sparse patch made up of points indicated by diagonal circles is stored in different adaptation sets. In this way, the client device can select the bitstream of a desired layer (group) by selecting an adaptation set to decode. In other words, the client device can more easily acquire and decode the bitstream of a desired layer (group).
[0176] <Example of spatial scalability information>
[0111] Fig. 21 is a diagram showing an example of the syntax of a V3C3D region descriptor. As shown in Fig. 21, layer identification information (layerId) may be stored as spatial scalability information in vpsr.spatialRegion.spatialScalabilityInfo of this V3C3D region descriptor. As in the case of ISOBMFF, the layer identification information is identification information indicating the layer corresponding to the sub-bitstream whose control information is stored in the adaptation set corresponding to that vpsr.spatialRegion.spatialScalabilityInfo. For example, layerId = 0 indicates the base layer, and layerId = 1 to 255 indicates the enhancement layer.
[0177] That is, the file generation device may store this layer identification information (layerId) in a V3C3D region descriptor of the MPD, and the client device may select a sub-bitstream (adaptation set) based on the layer identification information stored in the V3C3D region descriptor of the MPD. In this manner, the client device can determine, based on the layer identification information, the layer corresponding to the sub-bitstream (sparse patch) whose control information is stored in each adaptation set. Therefore, the client device can more easily select and acquire a point cloud that achieves high resolution in the order intended by the content creator. Therefore, the client device can more easily play back 3D data using spatial scalability.
[0178] In addition to this layer identification information, as shown in Figure 21, information (lod) regarding the resolution of the point cloud obtained by reconstructing the point cloud corresponding to each layer from the top layer of the spatial scalability to the layer indicated by the layer identification information may also be stored in vpsr.spatialRegion.spatialScalabilityInfo of this V3C3D region descriptor.
[0179] For example, when layerId = 0, the information (lod) about the resolution of this point cloud indicates the LoD value of the base layer. Also, for example, when layerId = 0, the information (lod) about the resolution of this point cloud indicates the LoD value obtained by simultaneously displaying it with point clouds of 0 to (layer_id-1). This LoD value may be a guideline value determined by the content creator. In this case, the information (lod) about the resolution of the point cloud may not be signaled, and the value of layer_id may signal the information about the resolution of the point cloud. In other words, the information (lod) about the resolution of the point cloud may be included in the layer identification information (layer_id). For example, the value of layer_id may also indicate the resolution (lod value) of the point cloud corresponding to each layer from the top layer of spatial scalability to the layer indicated by the layer identification information.
[0180] That is, the file generation device may store information about the resolution of the point cloud (LoD) in a V3C3D region descriptor of the MPD, and the client device may select a sub-bitstream (adaptation set) based on the information about the resolution of the point cloud stored in the V3C3D region descriptor of the MPD. This allows the client device to more easily determine which adaptation set to select to obtain the desired LoD. Therefore, the client device can more easily play back 3D data using spatial scalability.
[0181] Also, as shown in FIG. 21, in addition to this layer identification information, special scalability identification information (id) for identifying special scalability may be stored in vpsr.spatialRegion.spatialScalabilityInfo of this V3C3D region descriptor. A group of spatial regions (SpatialRegion) having the same special scalability identification information (id) with each other provides special scalability. That is, by combining a plurality of regions having the same special scalability identification information (id) with each other, a high LoD point cloud can be obtained.
[0182] That is, the file generation device stores this special scalability identification information (id) in the V3C3D region descriptor of the MPD, and the client device may select a sub-bitstream (adaptation set) based on the special scalability identification information (id) stored in the V3C3D region descriptor of this MPD. By doing so, the client device can more easily identify a group that provides special scalability. Therefore, the client device can more easily reproduce 3D data using special scalability.
[0183] <Another example of special scalability information> Note that, as in the example shown in FIG. 22, instead of signaling vpsr.spatialRegion.spatialScalabilityInfo@id, vpsr.spatialRegion.spatialScalabilityInfo and asIds for the number of layers may be signaled. At this time, a plurality of spatialScalabilityInfo under a specific vpsr.spatialRegion provides spatial scalability.
[0184] <MPD description example> Fig. 23 is a diagram showing an example of description of an MPD when the present technology is applied. Fig. 24 shows an example of description of the supplemental property shown in the fifth line from the top of Fig. 23.
[0185] In the example shown in Fig. 24, spatial scalability identification information (id), information on the resolution of the point cloud (lod), and layer identification information (layerId) are shown as v3c:spatialScalabilityInfo.
[0186] Therefore, a client device can obtain the bitstream necessary to construct a point cloud of a desired LoD based on this MPD, which means that the client device can more easily play back 3D data using spatial scalability.
[0187] <File generation device> Fig. 25 is a block diagram showing a main configuration example of a file generation device 300 in this case. That is, the file generation device 300 shown in Fig. 25 shows an example of the configuration of a file generation device, which is one aspect of an information processing device to which the present technology is applied. The file generation device 300 shown in Fig. 25 is a device that applies V-PCC to encode point cloud data as video frames using an encoding method for two-dimensional images. Furthermore, the file generation device 300 in this case generates an MPD that stores control information for controlling distribution of a V3C bitstream generated by the encoding.
[0188] At this time, the file generation device 300 applies the present technology described above in the present embodiment and stores information in the MPD so as to enable spatial scalability. That is, the file generation device 300 stores information related to spatial scalability in the MPD.
[0189] Note that Fig. 25 shows the main processing units, data flows, etc., and is not necessarily all that is shown in Fig. 25. In other words, in file generation device 300, there may be processing units that are not shown as blocks in Fig. 25, and there may be processing or data flows that are not shown as arrows, etc. in Fig. 25.
[0190] As shown in FIG. 25, file generation device 300 has an MPD generation unit 501 in addition to the configuration described with reference to FIG.
[0191] In this case, the metadata generation unit 303 generates metadata in the same manner as in the case of Fig. 16. For example, the metadata generation unit 303 generates, as metadata, spatial scalability information related to the spatial scalability of the acquired sub-bitstream. That is, the metadata generation unit 303 generates the spatial scalability information using any single method or an appropriate combination of multiple methods from among the various methods described with reference to Figs. 20 to 24, etc. Note that the metadata generation unit 303 may generate any metadata other than spatial scalability information.
[0192] After generating the metadata including the spatial scalability information in this way, the metadata generation unit 303 supplies the metadata to the MPD generation unit 501.
[0193] The MPD generation unit 501 acquires metadata including spatial scalability information supplied from the metadata generation unit 303. The MPD generation unit 501 generates an MPD that stores the acquired metadata. That is, the MPD generation unit 501 stores the spatial scalability information in the MPD. That is, the MPD generation unit 501 stores the spatial scalability information in the system layer.
[0194] At this time, the MPD generation unit 501 stores the spatial scalability information in the MPD using any one of the various methods described with reference to Figures 20 to 24, etc., or using an appropriate combination of any two or more methods. For example, as shown in Figure 24, the MPD generation unit 501 stores the spatial scalability information in v3c:spatialScalabilityInfo of the MPD.
[0195] The MPD generation unit 501 supplies the generated MPD to the output unit 306. The output unit 306 outputs the supplied MPD (MPD including spatial scalability information) to an external device of the file generation device 300 (for example, a distribution server, a client device, etc.).
[0196] As described above, the file generation device 300 applies the present technology described above in this embodiment to generate an MPD that stores spatial scalability information.
[0197] This configuration allows spatial scalability information to be provided at the system layer to a client device that decodes a V3C bitstream, enabling the client device to more easily play back 3D data using spatial scalability without the need for complex tasks such as parsing the V3C bitstream.
[0198] Note that these processing units (the 3D2D conversion unit 301 to the output unit 306, the MPD generation unit 501, and the encoding units 311 to 314) may have any configuration. For example, each processing unit may be configured with a logic circuit that realizes the above-mentioned processing. Furthermore, each processing unit may have, for example, a CPU, a ROM, a RAM, etc., and may realize the above-mentioned processing by executing a program using these. Of course, each processing unit may have both of these configurations, and may realize some of the above-mentioned processing by a logic circuit and other by executing a program. The configurations of each processing unit may be independent of each other. For example, some processing units may realize some of the above-mentioned processing by a logic circuit, other processing units may realize the above-mentioned processing by executing a program, and still other processing units may realize the above-mentioned processing by both a logic circuit and by executing a program.
[0199] <File generation process flow> An example of the flow of the file generation process executed by file generation device 300 in this case will be described with reference to the flowchart of FIG.
[0200] When the file generation process is started, the processes of steps S501 to S505 are executed in the same manner as the processes of steps S301 to S305 in FIG.
[0201] In step S506, the file generation unit 305 generates a file and stores the V3C bitstream (each sub-bitstream) in the file.
[0202] In step S507, the MPD generation unit 501 generates an MPD that stores spatial scalability information (for example, base and enhancement point cloud information). At this time, the MPD generation unit 501 stores the spatial scalability information in the MPD using any one of the various methods described with reference to Figures 20 to 24, etc., or using an appropriate combination of any two or more methods.
[0203] In step S508, the output unit 306 outputs the file generated in step S506 and the MPD storing the spatial scalability information generated in step S507 to an external device (such as a distribution server) outside the file generation device 300. When the processing in step S508 ends, the file generation processing ends.
[0204] By performing each process in this way, spatial scalability information can be provided at the system layer to a client device that decodes a V3C bitstream, allowing the client device to more easily play back 3D data using spatial scalability without the need for complex tasks such as parsing the V3C bitstream.
[0205] <Client device> The present technology described in this embodiment can be applied not only to file generation devices but also to client devices. FIG. 27 is a block diagram showing a main configuration example of a client device 400 in this case. That is, the client device 400 shown in FIG. 27 shows an example of the configuration of a client device, which is one aspect of an information processing device to which the present technology is applied. The client device 400 shown in FIG. 27 is a device that applies V-PCC and acquires a V3C bitstream (a geometry video sub-bitstream, an attribute video sub-bitstream, an occupancy video sub-bitstream, and an atlas sub-bitstream, or a combination of these) encoded using a two-dimensional image encoding method with point cloud data as video frames, based on an MPD, decodes the V3C bitstream using a two-dimensional image decoding method, and generates (reconstructs) a point cloud. For example, the client device 400 can acquire a V3C bitstream based on the MPD generated by the file generation device 300, decode it, and generate a point cloud.
[0206] In this case, the client device 400 realizes spatial scalability by using any one of the various methods of the present technology described above in this embodiment, or by using an appropriate combination of any two or more of the methods. That is, the client device 400 selects and acquires a bitstream (track) required to reconstruct a point cloud of a desired LoD, based on the spatial scalability information stored in the MPD.
[0207] Note that Fig. 27 shows the main processing units, data flows, etc., and is not limited to all that is shown in Fig. 27. In other words, in client device 400, there may be processing units that are not shown as blocks in Fig. 27, and there may be processing or data flows that are not shown as arrows, etc. in Fig. 27.
[0208] As shown in FIG. 27, the client device 400 includes an MPD analysis unit 601 in addition to the configuration shown in FIG.
[0209] The MPD analysis unit 601 analyzes the MPD acquired by the file acquisition unit 411, selects a bitstream to be decoded, and causes the file acquisition unit 411 to acquire the bitstream.
[0210] In this case, the MPD analysis unit 601 analyzes the MPD using any one of the various methods of the present technology described in this embodiment, or an appropriate combination of any two or more of the methods. For example, the MPD analysis unit 601 analyzes spatial scalability information stored in the MPD and selects sub-bitstreams to be decoded. For example, the MPD analysis unit 601 selects a combination of point clouds that provides spatial scalability according to the network environment and the processing capabilities of the client device 400 itself, based on the spatial scalability information (i.e., sub-bitstreams to be decoded). Based on the analysis result, the MPD analysis unit 601 controls the file acquisition unit 411 to acquire the selected bitstreams.
[0211] In this case, the file acquisition unit 411 acquires an MPD from a distribution server or the like, and supplies it to the MPD analysis unit 601. In addition, the file acquisition unit 411 is controlled by the MPD analysis unit 601, acquires a file including a bitstream selected by the MPD analysis unit 601 from the distribution server or the like, and supplies it to the file analysis unit 412.
[0212] The file analysis unit 412 analyzes the file, and the extraction unit 413 extracts the bitstream based on the analysis result and supplies it to the 2D decoding unit 402.
[0213] The 2D decoding unit 402 through the display unit 404 perform the same processing as in the case of FIG.
[0214] With this configuration, the client device 400 can identify a combination of point clouds that provides spatial scalability based on the spatial scalability information stored in the system layer. Therefore, for example, in V-PCC content distribution, the client device can select a point cloud with an appropriate LoD according to the client environment without requiring complicated operations such as analyzing the bitstream. In other words, the client device can more easily play back 3D data using spatial scalability.
[0215] Note that these processing units (file processing unit 401 to display unit 404, file acquisition unit 311 to extraction unit 413, decoding unit 421 to decoding unit 424, 2D3D conversion unit 431 and display processing unit 432, and MPD analysis unit 601) may have any configuration. For example, each processing unit may be configured with a logic circuit that realizes the above-mentioned processing. Furthermore, each processing unit may have, for example, a CPU, ROM, RAM, etc., and may realize the above-mentioned processing by executing a program using these. Of course, each processing unit may have both of these configurations, and may realize some of the above-mentioned processing by a logic circuit and other by executing a program. The configurations of each processing unit may be independent of each other. For example, some processing units may realize some of the above-mentioned processing by a logic circuit, other processing units may realize the above-mentioned processing by executing a program, and still other processing units may realize the above-mentioned processing by both a logic circuit and by executing a program.
[0216] <Client processing flow> An example of the flow of client processing executed by this client device 400 will be described with reference to the flowchart of FIG.
[0217] When the client process starts, the file acquisition unit 411 of the client device 400 acquires an MPD in step S601.
[0218] In step S602, the MPD analysis unit 601 selects a combination of point clouds that provides spatial scalability according to the network environment and client processing capabilities, based on the spatial scalability information (base and enhancement point cloud information) described in the MPD.
[0219] In step S603, the file acquisition unit 411 acquires a file storing the atlas sub-bitstream and the video sub-bitstreams corresponding to the multiple sparse point clouds selected in step S602.
[0220] In step S604, the extractor 413 extracts the tras sub-bitstream and the video sub-bitstream from the file.
[0221] The processes in steps S605 to S608 are executed in the same manner as the processes in steps S404 to S407 in FIG.
[0222] When the process of step S608 ends, the client process ends.
[0223] By performing the above processes, the client device 400 can identify a combination of point clouds that provides spatial scalability based on the spatial scalability information stored in the system layer. Therefore, for example, in V-PCC content distribution, the client device can select a point cloud with an appropriate LoD according to the client environment without requiring cumbersome tasks such as analyzing the bitstream. In other words, the client device can more easily play back 3D data using spatial scalability.
[0224] <4. Notes> <Computer> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the term "computer" includes computers built into dedicated hardware, and general-purpose personal computers, etc., that can execute various functions by installing various programs.
[0225] FIG. 29 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.
[0226] In a computer 900 shown in FIG. 29, a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, and a RAM (Random Access Memory) 903 are interconnected via a bus 904.
[0227] An input / output interface 910 is also connected to the bus 904. To the input / output interface 910, an input unit 911, an output unit 912, a storage unit 913, a communication unit 914, and a drive 915 are connected.
[0228] The input unit 911 includes, for example, a keyboard, a mouse, a microphone, a touch panel, an input terminal, etc. The output unit 912 includes, for example, a display, a speaker, an output terminal, etc. The storage unit 913 includes, for example, a hard disk, a RAM disk, a non-volatile memory, etc. The communication unit 914 includes, for example, a network interface. The drive 915 drives removable media 921 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0229] In a computer configured as above, the CPU 901 performs the above-described series of processes by, for example, loading a program stored in the storage unit 913 into the RAM 903 via the input / output interface 910 and the bus 904 and executing the program. The RAM 903 also stores data necessary for the CPU 901 to execute various processes as appropriate.
[0230] The program executed by the computer can be applied by recording it on removable media 921 such as package media, for example. In this case, the program can be installed in storage unit 913 via input / output interface 910 by inserting removable media 921 into drive 915.
[0231] This program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, digital satellite broadcasting, etc. In this case, the program can be received by the communication unit 914 and installed in the storage unit 913.
[0232] Alternatively, this program can be installed in advance in the ROM 902 or the storage unit 913 .
[0233] <Applicable targets of this technology> While the above describes the application of this technology to the encoding and decoding of point cloud data, this technology is not limited to these examples and can be applied to the encoding and decoding of 3D data of any standard. In other words, as long as it does not conflict with the above-described technology, various processes such as encoding and decoding methods and specifications of various data such as 3D data and metadata are arbitrary. Furthermore, as long as it does not conflict with the above-described technology, some of the above-described processes and specifications may be omitted.
[0234] Furthermore, the present technology can be applied to any configuration, for example, various electronic devices.
[0235] Furthermore, for example, the present technology can also be implemented as a part of an apparatus, such as a processor (e.g., a video processor) as a system LSI (Large Scale Integration), a module (e.g., a video module) using multiple processors, a unit (e.g., a video unit) using multiple modules, or a set in which other functions are added to a unit (e.g., a video set).
[0236] Furthermore, for example, the present technology can also be applied to a network system configured with multiple devices. For example, the present technology may be implemented as cloud computing in which multiple devices share and collaborate on processing via a network. For example, the present technology may be implemented in a cloud service that provides image (video)-related services to any terminal, such as a computer, AV (Audio Visual) equipment, a portable information processing terminal, or an IoT (Internet of Things) device.
[0237] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0238] <Fields and applications where this technology can be applied> Systems, devices, processing units, etc. to which the present technology is applied can be used in any field, such as transportation, medical care, crime prevention, agriculture, livestock farming, mining, beauty, factories, home appliances, weather, and nature monitoring. In addition, the applications thereof are also arbitrary.
[0239] For example, the present technology can be applied to systems and devices used to provide viewing content, etc. Furthermore, for example, the present technology can also be applied to systems and devices used for transportation, such as monitoring traffic conditions and controlling automatic driving. Furthermore, for example, the present technology can also be applied to systems and devices used for security. Furthermore, for example, the present technology can also be applied to systems and devices used for automatic control of machines, etc. Furthermore, for example, the present technology can also be applied to systems and devices used for agriculture and livestock farming. Furthermore, for example, the present technology can also be applied to systems and devices used to monitor natural conditions, such as volcanoes, forests, and oceans, and wildlife. Furthermore, for example, the present technology can also be applied to systems and devices used for sports.
[0240] <Other> In this specification, a "flag" refers to information for identifying multiple states, and includes not only information used to identify two states, true (1) or false (0), but also information capable of identifying three or more states. Therefore, the value that this "flag" can take may be, for example, two values, 1 / 0, or three or more values. In other words, the number of bits constituting this "flag" is arbitrary, and may be one bit or multiple bits. Furthermore, identification information (including flags) can be assumed not only to include the identification information in the bit stream, but also to include difference information of the identification information relative to certain reference information in the bit stream. Therefore, in this specification, "flag" and "identification information" include not only the information itself, but also difference information relative to the reference information.
[0241] Furthermore, various types of information (metadata, etc.) related to the coded data (bitstream) may be transmitted or recorded in any form as long as they are associated with the coded data. Here, the term "associate" means, for example, that one piece of data can be used (linked) when processing the other piece of data. In other words, data associated with each other may be combined into one piece of data or may be individual pieces of data. For example, information associated with coded data (image) may be transmitted over a transmission path separate from that of the coded data (image). Also, for example, information associated with coded data (image) may be recorded on a recording medium separate from that of the coded data (image) (or on a different recording area of the same recording medium). Note that this "association" may refer to only a portion of the data, rather than the entire data. For example, an image and information corresponding to that image may be associated with each other in any unit, such as multiple frames, one frame, or a portion of a frame.
[0242] In this specification, terms such as "composite," "multiplex," "add," "integrate," "include," "store," "embed," "insert," and the like refer to combining multiple items into one, such as combining encoded data and metadata into one piece of data, and refer to one method of "associating" as described above.
[0243] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.
[0244] For example, a configuration described as one device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, configurations described above as multiple devices (or processing units) may be combined and configured as one device (or processing unit). Of course, configurations other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, as long as the configuration and operation of the entire system are substantially the same, part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).
[0245] Furthermore, for example, the above-described program may be executed in any device, as long as the device has the necessary functions (functional blocks, etc.) and can obtain the necessary information.
[0246] Also, for example, each step of a single flowchart may be executed by one device, or may be shared and executed by multiple devices. Furthermore, when one step includes multiple processes, the multiple processes may be executed by one device, or may be shared and executed by multiple devices. In other words, multiple processes included in one step can be executed as multiple step processes. Conversely, processes described as multiple steps can be executed collectively as one step.
[0247] For example, the steps of a program executed by a computer may be executed in chronological order in the order described herein, or may be executed in parallel or individually at the required timing, such as when a call is made. In other words, as long as no contradiction occurs, the steps may be executed in an order different from the order described above. Furthermore, the steps of this program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.
[0248] Furthermore, for example, multiple technologies related to the present technology can be implemented independently and independently, as long as no contradiction occurs. Of course, any multiple technologies can also be implemented in combination. For example, part or all of the present technology described in any embodiment can be implemented in combination with part or all of the present technology described in another embodiment. Furthermore, part or all of any of the above-described present technologies can be implemented in combination with other technologies not described above.
[0249] The present technology can also be configured as follows. (1) A coding unit that encodes 2D data of a point cloud that represents a three-dimensional object as a set of points, the point cloud being two-dimensionalized and corresponding to spatial scalability, and generates a bitstream including sub-bitstreams in which the point cloud corresponding to one or more layers of the spatial scalability is encoded; a spatial scalability information generator that generates spatial scalability information related to the spatial scalability of the sub-bitstream; a file generation unit that generates a file that stores the bitstream generated by the encoding unit and the spatial scalability information generated by the spatial scalability information generation unit; An information processing device comprising: (2) The spatial scalability information includes layer identification information indicating the layer to which the sub-bitstream stored in the track group of the file corresponds. The information processing device described in (1). (3) The spatial scalability information further includes information regarding the resolution of the point cloud obtained by reconstructing the point cloud corresponding to each layer from the highest layer of the spatial scalability to the layer indicated by the layer identification information. (2) An information processing device according to the present invention. (4) The spatial scalability information further includes spatial scalability identification information that identifies the spatial scalability. (3) An information processing device according to the present invention. (5) A point cloud expressing a three-dimensional object as a set of points is encoded as 2D data of the point cloud corresponding to spatial scalability, and a bitstream including sub-bitstreams in which the point cloud corresponding to one or more layers of the spatial scalability is encoded is generated; generating spatial scalability information regarding the spatial scalability of the sub-bitstream; generating a file storing the generated bitstream and the spatial scalability information; Information processing methods.
[0250] (6) a selection unit that selects a layer of spatial scalability to be decoded based on spatial scalability information of a bitstream in which 2D data obtained by two-dimensionally converting the point cloud corresponding to spatial scalability into a point cloud stored in a file and representing a three-dimensional object as a set of points; an extracting unit that extracts, from the bitstream stored in the file, a sub-bitstream corresponding to the layer selected by the selecting unit; a decoding unit that decodes the sub-bitstream extracted by the extraction unit; An information processing device comprising: (7) The selection unit selects the layer of the spatial scalability to be decoded based on layer identification information included in the spatial scalability information and indicating the layer to which the sub-bitstream stored in a track group of the file corresponds. (6) An information processing device according to (6). (8) The selection unit further selects the layer of the spatial scalability to be decoded based on information on a resolution of the point cloud obtained by reconstructing the point cloud corresponding to each layer from the highest layer of the spatial scalability to the layer indicated by the layer identification information, the information being included in the spatial scalability information. (7) An information processing device according to (7). (9) The selection unit further selects the layer of the spatial scalability to be decoded based on spatial scalability identification information that is included in the spatial scalability information and identifies the spatial scalability. (8) An information processing device according to (8). (10) A point cloud is stored in a file and represents a three-dimensional object as a set of points. The point cloud corresponds to spatial scalability, and 2D data is two-dimensionally converted into 2D data. Based on the spatial scalability information of the bit stream, a layer of the spatial scalability to be decoded is selected. extracting a sub-bitstream corresponding to the selected layer from the bitstream stored in the file; Decoding the extracted sub-bitstream Information processing methods.
[0251] (11) A coding unit that encodes 2D data of a point cloud that represents a three-dimensional object as a set of points, the point cloud corresponding to spatial scalability, and generates a bitstream including sub-bitstreams in which the point cloud corresponding to one or more layers of the spatial scalability is encoded; a spatial scalability information generator that generates spatial scalability information related to the spatial scalability of the sub-bitstream; a control file generation unit that generates a control file that stores the spatial scalability information generated by the spatial scalability information generation unit and control information related to distribution of the bitstream generated by the encoding unit; An information processing device comprising: (12) The spatial scalability information includes layer identification information indicating the layer to which the sub-bitstream corresponding to the control information stored in the adaptation set of the control file corresponds. (11) An information processing device according to (11). (13) The spatial scalability information further includes information regarding the resolution of the point cloud obtained by reconstructing the point cloud corresponding to each layer from the highest layer of the spatial scalability to the layer indicated by the layer identification information. (12) An information processing device according to (12). (14) The spatial scalability information further includes spatial scalability identification information that identifies the spatial scalability. (13) An information processing device according to (13). (15) A point cloud expressing a three-dimensional object as a set of points is encoded as 2D data of the point cloud corresponding to spatial scalability, and a bitstream including sub-bitstreams in which the point cloud corresponding to one or more layers of the spatial scalability is encoded is generated; generating spatial scalability information regarding the spatial scalability of the sub-bitstream; generating a control file that stores the generated spatial scalability information and control information related to distribution of the generated bitstream; Information processing methods.
[0252] (16) A selection unit that selects a layer of spatial scalability to be decoded based on spatial scalability information about the spatial scalability of a bitstream that is stored in a control file that stores control information about distribution of a bitstream in which 2D data obtained by two-dimensionally converting the point cloud corresponding to spatial scalability is encoded; and an acquisition unit that acquires a sub-bitstream corresponding to the layer selected by the selection unit; a decoding unit that decodes the sub-bitstream acquired by the acquisition unit; An information processing device comprising: (17) The selection unit selects the layer of the spatial scalability to be decoded based on layer identification information included in the spatial scalability information and indicating the layer corresponding to the sub-bitstream for which the control information is stored in an adaptation set of the control file. (16) An information processing device according to (16). (18) The selection unit further selects the layer of the spatial scalability to be decoded based on information on a resolution of the point cloud obtained by reconstructing the point cloud corresponding to each layer from a top layer of the spatial scalability to the layer indicated by the layer identification information, the information being included in the spatial scalability information. (17) An information processing device according to (17). (19) The selection unit further selects the layer of the spatial scalability to be decoded based on spatial scalability identification information that is included in the spatial scalability information and identifies the spatial scalability. (18) An information processing device according to (18). (20) A point cloud expressing a three-dimensional object as a set of points, the point cloud corresponding to spatial scalability is two-dimensionalized, and a control file stores control information related to the distribution of a bit stream in which 2D data is encoded, and the spatial scalability layer to be decoded is selected based on spatial scalability information related to the spatial scalability of the bit stream; obtaining a sub-bitstream corresponding to the selected layer; Decoding the obtained sub-bitstream Information processing methods. [Explanation of symbols]
[0253] 300 file generation device, 301 3D2D conversion unit, 302 2D encoding unit, 303 metadata generation unit, 304 PC stream generation unit, 305 file generation unit, 306 output unit, 311 to 314 encoding units, 400 client device, 401 file processing unit, 402 2D decoding unit, 403 display information generation unit, 404 display unit, 411 file acquisition unit, 412 file analysis unit, 413 extraction unit, 421 to 424 decoding units, 431 2D3D conversion unit, 432 display processing unit, 501 MPD generation unit, 601 MPD analysis unit
Claims
1. an encoding unit that supports spatial scalability for controlling a resolution of the reconstructed 3D data, encodes the 3D data having a hierarchical structure based on the resolution as a sub-bitstream for each layer of the hierarchical structure, and generates a bitstream including the sub-bitstream; a spatial scalability information generating unit that generates information about the spatial scalability of the sub-bitstream, the information including identification information of the layer corresponding to the sub-bitstream and information about the resolution of the 3D data obtained by reconstructing layers from the highest layer of the hierarchical structure to the layer of the sub-bitstream; a file generation unit that generates a file for storing the bitstream generated by the encoding unit, and stores the spatial scalability information generated by the spatial scalability information generation unit in a system layer of the file; An information processing device comprising:
2. The spatial scalability information further includes spatial scalability identification information that identifies the spatial scalability. The information processing device according to claim 1 .
3. Corresponding to spatial scalability for controlling the resolution of the reconstructed 3D data, the 3D data having a hierarchical structure based on the resolution is encoded as a sub-bitstream for each layer of the hierarchical structure, and a bitstream including the sub-bitstream is generated; generating spatial scalability information about the spatial scalability of the sub-bitstream, the spatial scalability information including identification information of the layer corresponding to the sub-bitstream and information about the resolution of the 3D data obtained by reconstructing layers from the highest layer of the hierarchical structure to the layer of the sub-bitstream; generating a file for storing the generated bitstream, and storing the generated spatial scalability information in a system layer of the file; Information processing methods.
4. a selection unit that selects the layer to be decoded based on spatial scalability information that is stored in a system layer of a file and that controls a resolution of the reconstructed 3D data, the spatial scalability information including identification information of the layer corresponding to a sub-bitstream coded for each layer of a hierarchical structure based on the resolution of the 3D data, and information about the resolution of the 3D data obtained by reconstructing layers from the highest layer of the hierarchical structure to the layer of the sub-bitstream; an extracting unit that extracts the sub-bitstream corresponding to the layer selected by the selecting unit from the bitstream of the 3D data stored in the file; a decoding unit that decodes the sub-bitstream extracted by the extraction unit; An information processing device comprising:
5. The selection unit further selects the layer to be decoded based on spatial scalability identification information that identifies the spatial scalability and is included in the spatial scalability information. The information processing device according to claim 4 .
6. selecting the layer to be decoded based on spatial scalability information for controlling a resolution of the reconstructed 3D data, the spatial scalability information being stored in a system layer of the file, the spatial scalability information including identification information of the layer corresponding to a sub-bitstream coded for each layer of a hierarchical structure based on the resolution of the 3D data, and information on the resolution of the 3D data obtained by reconstructing layers from the highest layer of the hierarchical structure to the layer of the sub-bitstream; extracting a sub-bitstream corresponding to the selected layer from the bitstream of the 3D data stored in the file; Decoding the extracted sub-bitstream Information processing methods.
7. an encoding unit that supports spatial scalability for controlling a resolution of the reconstructed 3D data, encodes the 3D data having a hierarchical structure based on the resolution as a sub-bitstream for each layer of the hierarchical structure, and generates a bitstream including the sub-bitstream; a spatial scalability information generating unit that generates information about the spatial scalability of the sub-bitstream, the information including identification information of the layer corresponding to the sub-bitstream and information about the resolution of the 3D data obtained by reconstructing layers from the highest layer of the hierarchical structure to the layer of the sub-bitstream; a control file generation unit that generates a control file that stores the spatial scalability information generated by the spatial scalability information generation unit and control information related to distribution of the bitstream generated by the encoding unit; An information processing device comprising:
8. The spatial scalability information further includes spatial scalability identification information that identifies the spatial scalability. The information processing device according to claim 7 .
9. Supports spatial scalability to control the resolution of reconstructed 3D data. encoding the 3D data having a hierarchical structure based on the resolution as a sub-bitstream for each layer of the hierarchical structure, and generating a bitstream including the sub-bitstream; generating spatial scalability information about the spatial scalability of the sub-bitstream, the spatial scalability information including identification information of the layer corresponding to the sub-bitstream and information about the resolution of the 3D data obtained by reconstructing layers from the highest layer of the hierarchical structure to the layer of the sub-bitstream; generating a control file that stores the generated spatial scalability information and control information related to distribution of the generated bitstream; Information processing methods.
10. a selection unit that selects the layer to be decoded based on spatial scalability information related to the spatial scalability of the bitstream, the spatial scalability information corresponding to spatial scalability that controls the resolution of the reconstructed 3D data and stored in a control file that stores control information related to distribution of a bitstream in which the 3D data having a hierarchical structure based on the resolution is encoded, the spatial scalability information including identification information of the layer corresponding to a sub-bitstream encoded for each layer of the hierarchical structure and information related to the resolution of the 3D data obtained by reconstructing layers from the top layer of the hierarchical structure to the layer of the sub-bitstream; an acquisition unit that acquires a sub-bitstream corresponding to the layer selected by the selection unit; a decoding unit that decodes the sub-bitstream acquired by the acquisition unit; An information processing device comprising:
11. The selection unit further selects the layer to be decoded based on spatial scalability identification information that identifies the spatial scalability and is included in the spatial scalability information. The information processing device according to claim 10.
12. selecting the layer to be decoded based on information on the spatial scalability of the bitstream, the information corresponding to spatial scalability that controls the resolution of the reconstructed 3D data and stored in a control file in which control information related to distribution of a bitstream in which the 3D data having a hierarchical structure based on the resolution is stored, the spatial scalability information including identification information of the layer corresponding to a sub-bitstream coded for each layer of the hierarchical structure and information related to the resolution of the 3D data obtained by reconstructing layers from the top layer of the hierarchical structure to the layer of the sub-bitstream; obtaining a sub-bitstream corresponding to the selected layer; Decoding the obtained sub-bitstream Information processing methods.
Citation Information
Patent Citations
IEC14496-12,2015-02-20
Point cloud and mesh compression using image / video codecs
US20180268570A1
Point cloud compression using non-cubic projections and masks
US20190087978A1
Information processing device, information processing method, and program
WO2020008758A1