Method and apparatus for adaptive streaming of point clouds
Adaptive streaming of point clouds using video-based compression standards addresses the challenge of efficiently representing and compressing 3D point clouds, enabling high-quality streaming of static and dynamic scenes with interoperability and network adaptability.
Patent Information
- Application Number
- JP2024111684
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-01
- Filing Date
- 2024-07-11
- Publication Date
- 2025-11-06
- Estimated Expiration
- 2040-03-06
AI Technical Summary
Existing technologies face challenges in efficiently representing and compressing high-quality 3D point clouds for immersive media applications, particularly in terms of storage and transmission, especially for dynamic point clouds with moving objects, and there is a need for standardized lossy and lossless encoding methods.
Adaptive streaming of point clouds using video-based compression (V-PCC) standards, which include geometry-based and video-based compression methods, utilizing hybrid video coding systems and MPEG-DASH for efficient and interoperable storage and transmission.
Enables efficient and high-quality streaming of 3D point clouds, supporting both static and dynamic scenes, while maintaining interoperability and adaptability to varying network conditions.
Smart Images

Figure 0007765561000019 
Figure 0007765561000020 
Figure 0007765561000021
Abstract
Description
[Background technology]
[0001] background High-quality 3D point clouds have emerged in recent years as an advanced representation of immersive media. A point cloud consists of a set of points represented in 3D space using coordinates that indicate the location of each point along with one or more attributes associated with each point, such as color, transparency, acquisition time, laser reflectivity, or material properties. Data for creating a point cloud can be captured in several ways. For example, one technique for capturing point clouds uses multiple cameras and depth sensors. Light detection and ranging (LiDAR) laser scanners are also commonly used to capture point clouds. The number of points required to realistically reconstruct objects and scenes using point clouds is on the order of millions (or even billions). Therefore, efficient representation and compression are crucial for the storage and transmission of point cloud data.
[0002]
[0002] Recent advances in 3D point capture and rendering technology have led to novel applications in the fields of telepresence, virtual reality, and large-scale dynamic 3D maps. The 3D Graphics Subgroup of the ISO / IEC JTC1 / SC29 / WG11 Moving Picture Experts Group (MPEG) is currently working on the development of two 3D point cloud compression (PCC) standards: a geometry-based compression standard for static point clouds (point clouds of stationary objects) and a video-based compression standard for dynamic point clouds (point clouds of moving objects). The goal of these standards is to support efficient and interoperable storage and transmission of 3D point clouds. Among the requirements of these standards is to support lossy and / or lossless encoding of point cloud geometry coordinates and attributes.
[0003] BRIEF DESCRIPTION OF THE DRAWINGS A more detailed understanding may be had from the following detailed description, given by way of example in conjunction with the drawings attached hereto. The figures in such drawings, like the detailed description, are examples. Therefore, the figures and detailed description should not be considered as limiting, and other equally valid examples are possible and appropriate. Furthermore, like reference numerals in the figures indicate like elements. [Brief explanation of the drawings]
[0004] [Figure 1A] FIG. 4 is a block diagram illustrating an example video encoding and decoding system in which one or more embodiments may be implemented and / or practiced. [Figure 1B]
[0005] 1B is a block diagram illustrating an example video encoder unit for use with the video encoding and / or decoding system of FIG. 1A. [Figure 2]
[0006] FIG. 1 is a block diagram of a general block-based hybrid video coding system. [Figure 3]
[0007] FIG. 1 is a general block diagram of a block-based video decoder. [Figure 4]
[0008] The bitstream structure of video-based point cloud compression (V-PCC) is shown. [Figure 5]
[0009] 1 shows the MPD hierarchical data model. [Figure 6]
[0010] 10 illustrates an exemplary DASH configuration for grouping V-PCC components belonging to one point cloud within an MPEG-DASH MPD file. [Figure 7]
[0011] 1 is a flow diagram illustrating an exemplary decoder process for streaming point cloud content according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0005] Detailed Description Exemplary Systems in Which Embodiments May Be Implemented
[0012] 1A is a block diagram illustrating an example video encoding and decoding system 100 that may implement and / or practice one or more embodiments. System 100 may include a source device 112 that may transmit encoded video information to a destination device 114 via a communication channel 116.
[0006]
[0013] Source device 112 and / or destination device 114 can be any of a wide range of devices. In some representative embodiments, source device 112 and / or destination device 114 can include a wireless transmit and / or receive unit (WTRU), such as a wireless handset or any wireless device capable of communicating video information over communication channel 116, in which case communication channel 116 includes a wireless link. However, the methods, apparatuses, and systems described, disclosed, or otherwise explicitly, implicitly, and / or inherently provided (collectively "provided") herein are not necessarily limited to wireless applications or settings. For example, these techniques may apply to terrestrial television broadcast, cable television transmission, satellite television transmission, Internet video transmission, coded digital video encoded on a storage medium, and / or other situations. Communication channel 116 can include and / or be any combination of wireless or wired media suitable for transmitting coded video data.
[0007]
[0014] Source device 112 may include a video encoder unit 118, a transmit and / or receive (Tx / Rx) unit 120, and / or a Tx / Rx element 122. As shown, source device 112 may include a video source 124. Destination device 114 may include a Tx / RX element 126, a Tx / Rx unit 128, and / or a video decoder unit 130. As shown, destination device 114 may include a display device 132. Each of Tx / Rx units 120, 128 may be or include a transmitter, a receiver, or a combination of a transmitter and a receiver (e.g., a transceiver or a transmitter-receiver). Each of Tx / Rx elements 122, 126 may be, for example, an antenna. In accordance with this disclosure, video encoder unit 118 of source device 112 and / or video decoder unit 130 of destination device 114 may be configured and / or adapted (collectively "adapted") to apply the encoding techniques provided herein.
[0008]
[0015] Source device 112 and destination device 114 may include other elements / components or mechanisms. For example, source device 112 may be adapted to receive video data from an external video source. Destination device 114 may interface with an external display device (not shown) and / or include and / or use (e.g., integrated) display device 132. In some embodiments, the data stream generated by video encoder unit 118 may be communicated to other devices without modulating the data onto a carrier signal, such as by direct digital transfer, which may or may not modulate the data for transmission.
[0009]
[0016] The techniques provided herein may be performed by any digital video encoding and / or decoding device. Typically, the techniques provided herein are performed by separate video encoding and / or video decoding devices, although the techniques may also be performed by a video encoder / decoder combination, commonly referred to as a "CODEC." The techniques provided herein may also be performed by a video preprocessor, etc. Source device 112 and destination device 114 are merely examples of such encoding devices from which source device 112 may generate encoded video information (and / or receive and generate video data) for transmission to destination device 114. In some representative embodiments, source device 112 and destination device 114 may operate substantially symmetrically, such that each of devices 112, 114 may include both video encoding and decoding components and / or elements (collectively "elements"). Thus, system 100 may support either unidirectional and bidirectional video transmission between source device 112 and destination device 114 (e.g., for any video streaming, video playback, video broadcasting, video calling, and / or video conferencing, among others). In certain representative embodiments, source device 112 may be, for example, a video streaming server adapted to generate encoded video information (and / or receive and generate video data) for one or more destination devices, in which case the destination devices may communicate with source device 112 via wired and / or wireless communication systems.
[0010]
[0017] The external video source and / or video source 124 may be and / or may include a video capture device such as a video camera, a video archive containing previously captured video, and / or a video feed from a video content provider. In certain representative embodiments, the external video source and / or video source 124 may generate computer graphics-based data as the source video, or a combination of live video, archived video, and / or computer-generated video. In certain representative embodiments, if the video source 124 is a video camera, the source device 112 and the destination device 114 may be or implement a camera phone or video phone.
[0011]
[0018] Captured, pre-captured, computer-generated video, video feeds, and / or other types of video data (collectively “unencoded video”) may be encoded by video encoder unit 118 to form coded video information. Tx / Rx unit 120 may modulate the coded video information (e.g., form one or more modulated signals that carry the coded video information in accordance with a communication standard). Tx / Rx unit 120 may send the modulated signals to a transmitter for transmission. The transmitter may transmit the modulated signals to destination device 114 via Tx / Rx element 122.
[0012]
[0019] At destination device 114, Tx / Rx unit 128 may receive the modulated signal from over channel 116 via Tx / Rx element 126. Tx / Rx unit 128 may demodulate the modulated signal to obtain coded video information. Tx / RX unit 128 may pass the coded video information to video decoder unit 130.
[0013]
[0020] Video decoder unit 130 may decode the coded video information to obtain decoded video data. The coded video information may include syntax information defined by video encoder unit 118. This syntax information may include one or more elements (“syntax elements”), some or all of which may be useful in decoding the coded video information. The syntax elements may include, for example, characteristics of the coded video information. The syntax elements may also include characteristics of uncoded video used to form the coded video information and / or describe processing of the uncoded video.
[0014]
[0021] Video decoder unit 130 may output the decoded video data for subsequent storage and / or display on an external display (not shown). In certain representative embodiments, video decoder unit 130 may output the decoded video data to display device 132. Display device 132 may include any individual display device, multiple display devices, or combination of various display devices adapted to display the decoded video data to a user. Examples of such display devices include liquid crystal displays (LCDs), plasma displays, organic light-emitting diode (OLED) displays, and / or cathode ray tubes (CRTs), among others.
[0015]
[0022] The communication channel 116 may be any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines, or any combination of wireless and wired media. The communication channel 116 may be part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication channel 116 generally represents any suitable communication medium or collection of different communication media for transmitting video data from the source device 112 to the destination device 114, including any suitable combination of wired and / or wireless media. The communication channel 116 may include routers, switches, base stations, and / or any other equipment that may be useful in facilitating communication from the source device 112 to the destination device 114. Details of example communication systems that may facilitate such communication between the devices 112 and 114 are provided below with reference to FIGS. 15A-15E. Details of devices that may represent the source device 112 and the destination device 114 are likewise provided below.
[0016]
[0023] Video encoder unit 118 and video decoder unit 130 may operate according to one or more standards and / or specifications, such as, for example, MPEG-2, H.261, H.263, H.264, H.264 / AVC, and / or H.264 extended according to SVC extensions (“H.264 / SVC”), among others. Those skilled in the art will appreciate that the methods, apparatus, and / or systems described herein are applicable to other video encoders, decoders, and / or CODECs implemented according to (and / or compliant with) different standards, or proprietary video encoders, decoders, and / or CODECs, including future video encoders, decoders, and / or CODECs. The techniques described herein are not limited to any particular encoding standard.
[0017]
[0024] The relevant portions of the above-mentioned H.264 / AVC are available from the International Telecommunications Union as ITU-T Recommendation H.264, or more specifically, as “ITU-T Rec. H.264 and ISO / IEC 14496-10 (MPEG4-AVC), 'Advanced Video Coding for Generic Audiovisual Services,' v5, March, 2010;” which is incorporated herein by reference and may be referred to herein as the H.264 standard, the H.264 specification, the H.264 / AVC standard, and / or the specification. The techniques provided herein may be applicable to devices that conform to (e.g., generally conform to) the H.264 standard.
[0018]
[0025] 1A , each of video encoder unit 118 and video decoder unit 130 may include and / or be integrated with an audio encoder and / or audio decoder (as appropriate). Video encoder unit 118 and video decoder unit 130 may include a suitable MUX-DEMUX unit or other hardware and / or software that handles the encoding of both audio and video in a common data stream and / or separate data streams. Where applicable, the MUX-DEMUX unit may conform to other protocols, such as, for example, the ITU-T Recommendation H.223 Multiplexer Protocol and / or the User Datagram Protocol (UDP).
[0019]
[0026] One or more video encoder units 118 and / or video decoder units 130 may be included in one or more encoders and / or decoders, any of which may be integrated as part of a CODEC, and / or may be integrated with and / or associated with a respective camera, computer, mobile device, subscriber device, broadcast device, set-top box, and / or server, among others. The video encoder unit 118 and / or video decoder unit 130 may each be implemented as any of a wide variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. Either or both of the video encoder unit 118 and / or video decoder unit 130 may be implemented substantially in software, and the operations of the elements of the video encoder unit 118 and / or video decoder unit 130 may be performed by appropriate software instructions executed by one or more processors (not shown). Such embodiments may include off-chip components in addition to the processor, such as external storage (eg, in the form of non-volatile memory) and / or input / output interfaces, among others.
[0020]
[0027] In any embodiment in which the operations of elements of video encoder unit 118 and / or video decoder unit 130 may be performed by software instructions executed by one or more processors, the software instructions may be stored on a computer-readable medium including, for example, a magnetic disk, an optical disk, any other volatile (e.g., random access memory (“RAM”)), non-volatile (e.g., read-only memory (“ROM”)), and / or mass storage system readable by a CPU, among others. The computer-readable medium may reside exclusively on a processing system and / or may include cooperating or interconnected computer-readable media distributed across multiple interconnected processing systems, which may be local or remote to the processing system.
[0021]
[0028] 1B is a block diagram illustrating an example video encoder unit 118 for use with a video encoding and / or decoding system, such as system 100. Video encoder unit 118 may include a video encoder 133, an output buffer 134, and a system controller 136. Video encoder 133 (or one or more elements thereof) may be implemented in accordance with one or more standards and / or specifications, such as, for example, H.261, H.263, H.264, H.264 / AVC, the SVC extension of H.264 / AVC (H.264 / AVC Annex G), HEVC, and / or the scalable extension of HEVC (SHVC), among others. Those skilled in the art will appreciate that the methods, apparatus, and / or systems provided herein may also be applicable to other video encoders implemented in accordance with different standards and / or proprietary CODECs, including future CODECs.
[0022]
[0029] Video encoder 133 may receive a video signal provided from a video source, such as video source 124 and / or an external video source. This video signal may include unencoded video. Video encoder 133 may encode the unencoded video and provide an encoded (i.e., compressed) video bitstream (BS) at its output.
[0023]
[0030] The coded video bitstream BS may be provided to an output buffer 134. The output buffer 134 may buffer the coded video bitstream BS and provide such coded video bitstream BS as a buffered bitstream (BBS) for transmission over the communication channel 116.
[0024]
[0031] The buffered bitstream BBS output from the output buffer 134 may be sent to a storage device (not shown) for later viewing or transmission. In certain representative embodiments, the video encoder unit 118 may be configured for visual communication, which may transmit the buffered bitstream BBS over the communication channel 116 at a specified constant bitrate and / or variable bitrate (e.g., with delay (e.g., very slow or minimum delay)).
[0025]
[0032] The coded video bitstream BS, and in turn the buffering bitstream BBS, may carry bits of coded video information. The bits of the buffering bitstream BBS may be arranged as a stream of coded video frames. The coded video frames may be intra-coded frames (e.g., I-frames) or inter-coded frames (e.g., B-frames and / or P-frames). The stream of coded video frames may be arranged, for example, as a series of groups of pictures (GOPs), with the coded video frames of each GOP arranged in a specified order. Generally, each GOP may begin with an intra-coded frame (e.g., I-frame) followed by one or more inter-coded frames (e.g., P-frames and / or B-frames). Each GOP may contain only one intra-coded frame, but any GOP may contain multiple intra-coded frames. It is intended that B-frames may not be used in real-time, low-delay applications, for example, because bidirectional prediction may introduce additional coding delay compared to unidirectional prediction (P-frames). Additional and / or other frame types may be used, and the particular order of the encoded video frames may be changed, as will be understood by those skilled in the art.
[0026]
[0033] Each GOP may include syntax data ("GOP syntax data"), which may be located in the header of the GOP, in the headers of one or more frames of the GOP, and / or elsewhere. The GOP syntax data may indicate, quantify, classify, and / or describe the order of the coded video frames of each GOP. Each coded video frame may include syntax data ("coding frame syntax data"), which may indicate and / or describe the coding mode of each coded video frame.
[0027]
[0034] The system controller 136 may monitor various parameters and / or constraints associated with the channel 116, the computational capabilities of the video encoder unit 118, user demand, etc., and may establish target parameters to provide an associated quality of experience (QoE) appropriate to the specified constraints and / or conditions of the channel 116. One or more of the target parameters may be adjusted occasionally or periodically depending on the specified constraints and / or channel conditions. As an example, the QoE may be quantitatively evaluated using one or more metrics that evaluate video quality, including, for example, a metric commonly referred to as the relative perceptual quality of the encoded video sequence. For example, the relative perceptual quality of the encoded video sequence, measured using a peak signal-to-noise ratio (“PSNR”) metric, may be controlled by the bitrate (BR) of the encoded bitstream BS. One or more of the target parameters (e.g., including a quantization parameter (QP)) may be adjusted to maximize the relative perceptual quality within constraints associated with the bitrate of the encoded bitstream BS.
[0028]
[0035] FIG. 2 is a block diagram of a block-based hybrid video encoder 200 for use with a video encoding and / or decoding system such as system 100.
[0029]
[0036] Referring to FIG. 2, the block-based hybrid coding system 200 may include, among other things, a transform unit 204, a quantization unit 206, an entropy coding unit 208, an inverse quantization unit 210, an inverse transform unit 212, a first adder 216, a second adder 226, a spatial prediction unit 260, a motion prediction unit 262, a reference picture store 264, one or more filters 266 (e.g., loop filters), and / or a mode decision and encoder controller unit 280.
[0030]
[0037] The details of video encoder 200 are intended to be illustrative only, and real-world implementations may vary. A real-world implementation may, for example, include more, fewer, and / or different elements and / or may be arranged differently from the arrangement shown in FIG. 2. For example, while shown separately, some or all of the functionality of both transform unit 204 and quantization unit 206 may be highly integrated in some real-world implementations, such as implementations that use the core transforms of the H.264 standard. Similarly, inverse quantization unit 210 and inverse transform unit 212 may be highly integrated in some real-world implementations (e.g., H.264 or HEVC standard-compliant implementations), but are likewise shown separately for conceptual purposes.
[0031]
[0038] As described above, video encoder 200 may receive a video signal at its input 202. Video encoder 200 may generate coded video information from the received uncoded video and output the coded video information (e.g., any intra-frames or inter-frames) at its output 220 in the form of a coded video bitstream BS. Video encoder 200 may, for example, operate as a hybrid video encoder and utilize a block-based coding process to code the uncoded video. When performing such an coding process, video encoder 200 may operate on individual frames, pictures, and / or images (collectively "uncoded pictures") of the uncoded video.
[0032]
[0039] To facilitate the block-based encoding process, video encoder 200 may slice, partition, divide, and / or segment (collectively "partition") each uncoded picture received at its input 202 into multiple uncoded video blocks. For example, video encoder 200 may partition an uncoded picture into multiple uncoded video partitions (e.g., slices) and partition (e.g., subsequently partition) each of the uncoded video partitions into uncoded video blocks. Video encoder 200 may pass, feed, transmit, or provide the uncoded video blocks to spatial prediction unit 260, motion prediction unit 262, mode decision and coding controller unit 280, and / or first summer 216. As described in more detail below, the uncoded video blocks may be provided on a block-by-block basis.
[0033]
[0040] Spatial prediction unit 260 may receive uncoded video blocks and encode such video blocks in intra-mode. Intra-mode refers to any of several modes of spatial-based compression, and intra-mode encoding seeks to provide spatial-based compression of an uncoded picture. Spatial-based compression may result from reducing or eliminating spatial redundancy, if any, of video information within the uncoded picture. In forming the predictive blocks, spatial prediction unit 260 may perform spatial prediction (or "intra-prediction") of each uncoded video block with reference to one or more video blocks of the uncoded picture that have already been coded ("coded video blocks") and / or reconstructed ("reconstructed video blocks"). The coded video blocks and / or reconstructed video blocks may be neighboring, adjacent, or nearby (e.g., nearby) the uncoded video blocks.
[0034]
[0041] Motion prediction unit 262 may receive uncoded video blocks from input 202 and encode them in inter-mode. Inter-mode refers to any of several modes of temporal-based compression, including, for example, P-mode (unidirectional prediction) and / or B-mode (bidirectional prediction). Encoding in inter-mode seeks to provide temporal-based compression of the uncoded picture. The temporal-based compression may result from reducing or eliminating temporal redundancy, if any, of video information between the uncoded picture and one or more reference (e.g., adjacent) pictures. Motion / temporal prediction unit 262 may perform temporal prediction (or “inter-prediction”) of each uncoded video block relative to one or more video blocks of a reference picture (“reference video block”). The temporal prediction performed may be unidirectional prediction (e.g., for P-mode) and / or bidirectional prediction (e.g., for B-mode).
[0035]
[0042] For unidirectional prediction, the reference video blocks may be from one or more previously coded and / or previously reconstructed pictures. The coded and / or reconstructed picture(s) may be neighboring, adjacent, and / or nearby to the uncoded picture.
[0036]
[0043] For bidirectional prediction, the reference video blocks may be from one or more previously coded and / or previously reconstructed pictures, which may be neighboring, adjacent, and / or nearby to the uncoded picture.
[0037]
[0044] If multiple reference pictures are used (as may be the case in more recent video coding standards such as H.264 / AVC and / or HEVC), for each video block, its reference picture index may be sent to entropy coding unit 208 for subsequent output and / or transmission. The reference index may be used to identify which reference picture or pictures in reference picture store 264 the temporal prediction comes from.
[0038]
[0045] While typically highly integrated, the functions of motion / temporal prediction unit 262 related to motion estimation and motion compensation may be performed by separate entities or units (not shown). Motion estimation may be performed to estimate the motion of each uncoded video block relative to a reference picture video block and may include generating a motion vector for the uncoded video block. The motion vector may indicate the displacement of a predictive block relative to the uncoded video block being coded. This predictive block may be, for example, a reference picture video block that has been found to closely match, in terms of pixel difference, the uncoded video block being coded. The match may be determined by sum of absolute differences (SAD), sum of squared differences (SSD), and / or other difference metrics. Motion compensation may include fetching and / or generating a predictive block based on the motion vector identified by motion estimation.
[0039]
[0046] Motion prediction unit 262 may calculate a motion vector for the uncoded video block by comparing the uncoded video block to a reference video block from a reference picture stored in reference picture store 264. Motion prediction unit 262 may calculate values for fractional pixel locations of a reference picture included in reference picture store 264. In some cases, adder 226 or another unit of video encoder 200 may calculate fractional pixel locations of the reconstructed video block and store the reconstructed video block along with the calculated values of the fractional pixel locations in reference picture store 264. Motion prediction unit 262 may interpolate sub-integer pixels of a reference picture (e.g., of an I-frame, and / or a P-frame, and / or a B-frame).
[0040]
[0047] Motion prediction unit 262 may be configured to encode the motion vector relative to a selected motion predictor. The motion predictor selected by motion / temporal prediction unit 262 may be, for example, a vector equal to the average value of the motion vectors of already-encoded neighboring blocks. To encode the motion vector of an unencoded video block, motion / temporal prediction unit 262 may calculate a difference between the motion vector and the motion predictor to form a motion vector difference value.
[0041]
[0048] H.264 and HEVC refer to a set of potential reference frames as a "list." A set of reference pictures stored in reference picture store 264 may correspond to such a list of reference frames. Motion / temporal prediction unit 262 may compare reference video blocks of reference pictures from reference picture store 264 with uncoded video blocks (e.g., of P frames or B frames). If a reference picture in reference picture store 264 includes sub-integer pixel values, a motion vector calculated by motion / temporal prediction unit 262 may point to a sub-integer pixel location of the reference picture. Motion / temporal prediction unit 262 may send the calculated motion vector to entropy coding unit 208 and a motion compensation function of motion / temporal prediction unit 262. Motion prediction unit 262 (or its motion compensation function) may calculate an error value of a prediction block for the uncoded video block being coded. Motion prediction unit 262 may calculate prediction data based on the prediction block.
[0042]
[0049] The mode decision and encoder controller unit 280 may select one of the coding modes, intra mode or inter mode, based on, for example, a rate-distortion optimization method and / or error results generated for each mode.
[0043]
[0050] Video encoder 200 may form a residual block ("residual video block") by subtracting prediction data provided by motion prediction unit 262 from an unencoded video block being encoded. Adder 216 represents one or more elements that may perform this subtraction operation.
[0044]
[0051] Transform unit 204 may apply a transform to the residual video block to convert such residual video block from the pixel value domain to a transform domain, such as the frequency domain. The transform may be, for example, a discrete cosine transform (DCT), a transform provided herein, or any conceptually similar transform. Other examples of transforms include transforms defined in H.264 and / or HEVC, wavelet transforms, integer transforms, and / or sub-band transforms, among others. Application of the transform to the residual video block by transform unit 204 generates corresponding blocks of transform coefficients for the residual video block (“residual transform coefficients”). These residual transform coefficients may represent the magnitude of frequency components of the residual video block. Transform unit 204 may forward the residual transform coefficients to quantization unit 206.
[0045]
[0052] Quantization unit 206 may quantize the residual transform coefficients to further reduce the coding bit rate. For example, the quantization process may reduce the bit depth associated with some or all of the residual transform coefficients. In particular cases, quantization unit 206 may divide the values of the residual transform coefficients by a quantization level corresponding to a QP to form a block of quantized transform coefficients. The degree of quantization may be changed by adjusting the QP value. Quantization unit 206 may apply quantization using a desired number of quantization steps to represent the residual transform coefficients, and the number of steps used (or correspondingly, the value of the quantization level) may determine the number of coded video bits used to represent the residual video block. Quantization unit 206 may obtain the QP value from a rate controller (not shown). Following quantization, quantization unit 206 may provide the quantized transform coefficients to entropy coding unit 208 and inverse quantization unit 210.
[0046]
[0053] The entropy coding unit 208 may apply entropy coding to the quantized transform coefficients to form entropy-coded coefficients (i.e., a bitstream). The entropy coding unit 208 may form the entropy-coded coefficients using adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), and / or other entropy coding techniques. CABAC may require input of context information ("context"), as will be understood by those skilled in the art. This context may be based on neighboring video blocks, for example.
[0047]
[0054] Entropy encoding unit 208 may provide the entropy-coded coefficients, along with motion vectors and one or more reference picture indexes, in the form of a raw coded video bitstream to an internal bitstream format (not shown). This bitstream format may form coded video bitstream BS, which is provided to the output of buffer 134 (FIG. 1B), by appending additional information to the raw coded video bitstream, including headers and / or other information that enables video decoder unit 300 (FIG. 3), for example, to decode coded video blocks from the raw coded video bitstream. Following entropy encoding, coded video bitstream BS provided from entropy encoding unit 208 may be output, for example, to output buffer 134, transmitted, for example, to destination device 114 via channel 116, or archived for later transmission or retrieval.
[0048]
[0055] In certain representative embodiments, entropy coding unit 208 or another unit of video encoder 133, 200 may be configured to perform other coding functions in addition to entropy coding. For example, entropy coding unit 208 may be configured to identify a code block pattern (CBP) value for a video block. In certain representative embodiments, entropy coding unit 208 may perform run-length coding of quantized transform coefficients within a video block. As an example, entropy coding unit 208 may apply a zigzag scan or other scan pattern to place the quantized transform coefficients in a video block and encode lengths of zero for further compression. Entropy coding unit 208 may construct header information using syntax elements appropriate for transmission in the coded video bitstream BS.
[0049]
[0056] The inverse quantization unit 210 and the inverse transform unit 212 may apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain, for example, for later use as one of the reference video blocks (e.g., within one of the reference pictures in a reference picture list).
[0050]
[0057] Mode decision and encoder controller unit 280 may calculate a reference video block by applying the reconstructed residual video block to a predictive block of one of the reference pictures stored in reference picture store 264. Mode decision and encoder controller unit 280 may apply one or more interpolation filters to the reconstructed residual video block to calculate sub-integer pixel values (e.g., for half-pixel positions) for use in motion estimation.
[0051]
[0058] Adder 226 may add the reconstructed residual video block to the motion compensated prediction video block to generate a reconstructed video block that is stored in reference picture store 264. The reconstructed (pixel value domain) video block may be used by motion prediction unit 262 (or its motion estimation function and / or its motion compensation function) as one of the reference blocks for inter-coding an uncoded video block in a subsequent uncoded video.
[0052]
[0059] Filter 266 (e.g., loop filter) may include a deblocking filter. The deblocking filter is operable to remove visual artifacts that may be present in the reconstructed macroblocks. These artifacts may be introduced in the encoding process due to the use of different encoding modes, such as I-type, P-type, or B-type. The artifacts may be present, for example, at boundaries and / or edges of received video blocks, and the deblocking filter is operable to smooth the boundaries and / or edges of the video blocks to improve visual quality. The deblocking filter may filter the output of summer 226. Filter 266 may include other in-loop filters, such as a sample adaptive offset (SAO) filter supported by the HEVC standard.
[0053]
[0060] 3 is a block diagram illustrating an example video decoder 300 for use with a video decoder unit, such as video decoder unit 130 of FIG. 1A. The video decoder 300 may include an input 302, an entropy decoding unit 308, a motion compensation prediction unit 362, a spatial prediction unit 360, an inverse quantization unit 310, an inverse transform unit 312, a reference picture store 364, a filter 366, an adder 326, and an output 320. The video decoder 300 may generally perform a decoding process that is the inverse of the encoding process provided with respect to the video encoder 133, 200. This decoding process may be performed as described below.
[0054]
[0061] The motion compensated prediction unit 362 may generate prediction data based on the motion vector received from the entropy decoding unit 308. The motion vector may be coded with respect to a motion predictor of the video block corresponding to the coded motion vector. The motion compensated prediction unit 362 may determine the motion predictor, for example, as the median of motion vectors of neighboring blocks of the video block to be decoded. After determining the motion predictor, the motion compensated prediction unit 362 may extract a motion vector difference value from the coded video bitstream BS and decode the coded motion vector by adding the motion vector difference value to the motion predictor. The motion compensated prediction unit 362 may quantize the motion predictor to the same resolution as the coded motion vector. In certain representative embodiments, the motion compensated prediction unit 362 may use the same precision for some or all coded motion predictors. As another example, the motion compensated prediction unit 362 may be configured to use any of the above methods and determine which method to use by analyzing data included in a sequence parameter set, a slice parameter set, or a picture parameter set obtained from the coded video bitstream BS.
[0055]
[0062] After decoding the motion vector, motion compensated prediction unit 362 may extract the prediction video block identified by the motion vector from the reference picture in reference picture store 364. If the motion vector points to a fractional pixel location, such as a half pixel, motion compensated prediction unit 362 may interpolate values for the fractional pixel location. Motion compensated prediction unit 362 may use an adaptive interpolation filter or a fixed interpolation filter to interpolate these values. Motion compensated prediction unit 362 may obtain an indication of which of filters 366 to use, and in various representative embodiments, the coefficients of the filters 366, from the received coded video bitstream BS.
[0056]
[0063] The spatial prediction unit 360 may form a prediction video block from spatially neighboring blocks using an intra-prediction mode received in the coded video bitstream BS. The inverse quantization unit 310 may inverse quantize (e.g., inverse quantize the quantized block coefficients provided in the coded video bitstream BS and decoded by the entropy decoding unit 308). The inverse quantization process may include a conventional process, for example, as defined by H.264. The inverse quantization process may use a quantization parameter QP calculated by the video encoder 133, 200 for each video block to determine the degree of quantization and / or dequantization to apply.
[0057]
[0064] The inverse transform unit 312 may apply an inverse transform (e.g., an inverse of any transform described herein, an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients to generate a residual video block in the pixel domain. The motion compensated prediction unit 362 may generate a motion compensated block and perform interpolation based on an interpolation filter. An identifier of an interpolation filter to use in sub-pixel accurate motion estimation may be included in a syntax element of the video block. The motion compensated prediction unit 362 may calculate sub-integer pixel interpolated values of a reference block using an interpolation filter used by the video encoder 133, 200 during encoding of the video block. The motion compensated prediction unit 362 may identify the interpolation filter used by the video encoder 133, 200 according to the received syntax information and use the interpolation filter to generate the prediction block.
[0058]
[0065] The motion compensation prediction unit 262 may (1) use syntax information to identify the size of the video blocks used to encode one or more pictures of the encoded video sequence, (2) use partition information that describes how each video block of a frame of the encoded video sequence is partitioned, (3) use a mode (or mode information) that indicates how each partition is coded, (4) use one or more reference pictures for each inter-coded video block, and / or (5) use other information to decode the encoded video sequence.
[0059]
[0066] Adder 326 may sum the residual block with a corresponding prediction block generated by motion compensation prediction unit 362 or spatial prediction unit 360 to form a decoded video block. A loop filter 366 (e.g., a deblocking filter or SAO filter) may be applied to filter the decoded video block to remove blockiness artifacts and / or improve visual quality. The decoded video block may be stored in reference picture store 364, which may provide reference video blocks for subsequent motion compensation and may generate decoded video for presentation to a display device (not shown).
[0060] Point Cloud Compression
[0067] Figure 4 shows the structure of the bitstream for video-based point cloud compression (V-PCC). The generated video bitstream and metadata are multiplexed together to generate the final V-PCC bitstream.
[0061]
[0068] A V-PCC bitstream consists of a set of V-PCC units as shown in Figure 4. The syntax of a V-PCC unit defined in the latest version of the V-PCC Standard Community Draft (V-PCC CD) is given in Table 1, where each V-PCC unit has a V-PCC unit header and a V-PCC unit payload. The V-PCC unit header describes the V-PCC unit type (Table 2). V-PCC units with unit types 2, 3, and 4 are occupancy, geometry, and attribute data units, as defined in the community draft. These data units represent the three main components required for point cloud reconstruction. In addition to the V-PCC unit type, the V-PCC attribute unit header also specifies the attribute type and its index, allowing for support of multiple instances of the same attribute type.
[0062]
[0069] The payload of an occupancy, geometry, and attribute V-PCC unit (Table 3) corresponds to a video data unit (e.g., an HEVC NAL (Network Abstraction Layer) unit) that can be decoded by a video decoder specified in the corresponding occupancy, geometry, and attribute parameter set V-PCC unit.
[0063] [Table 1]
[0064] [Table 2]
[0065] [Table 3]
[0066] Dynamic Streaming over HTTP (DASH)
[0070] MPEG Dynamic Adaptive Streaming over HTTP (MPEG-DASH) is a universal delivery format that provides end users with the best possible video experience by dynamically adapting to changing network conditions.
[0067]
[0071] HTTP adaptive streaming, such as MPEG-DASH, requires that different bitrate alternatives for multimedia content be made available by the server. In addition, multimedia content may contain several media components (e.g., audio, video, text), each of which may have different characteristics. In MPEG-DASH, these characteristics are described by a media presentation description (MPD).
[0068]
[0072] Figure 5 shows the MPD hierarchical data model. The MPD describes a sequence of Periods in which the coded versions of a consistent set of media content components do not change. Each Period has a start time and a duration, and is composed of one or more Adaptation Sets.
[0069]
[0073] An Adaptation Set represents a set of coded versions of one or several media content components that share the same properties, such as language, media type, picture aspect ratio, role, accessibility, and rating properties. For example, an Adaptation Set may contain video components of the same multimedia content at different bit rates. Another Adaptation Set may contain audio components of the same multimedia content at different bit rates (e.g., lower quality stereo and higher quality surround sound). Each Adaptation Set typically contains multiple Representations.
[0070]
[0074] A Representation describes a distributable encoded version of one or several media components that may vary from other representations in bit rate, resolution, number of channels, or other characteristics. Each Representation consists of one or more Segments. Attributes of the Representation element, such as @id, @bandwidth, @qualityRanking, and @dependencyId, are used to specify properties of the associated Representation. A Representation may also contain sub-representations, which are parts of the Representation that describe and extract partial information from the Representation. A sub-representation may provide the ability to access lower-quality versions of the Representation in which it is contained.
[0071]
[0075] A Segment is the largest unit of data that can be retrieved using one HTTP request. Each segment has a URL, i.e., an addressable location on the server, that can be downloaded using an HTTP GET or an HTTP GET with a byte range.
[0072]
[0076] To use this data model, a DASH client parses the MPD XML and selects a collection of AdaptationSets appropriate for its environment based on the information provided in each AdaptationSet element. Within each AdaptationSet, the client selects one Representation, typically based on the value of the @bandwidth attribute, taking into account the client's decoding and rendering capabilities. The client downloads the initialization segment of the selected Representation and then accesses the content by requesting the entire Segment or a byte range of the Segment. Once the presentation has begun, the client continues to consume the media content by subsequently requesting Media Segments or portions of Media Segments and playing the content according to the media presentation timeline. The client may switch Representations taking into account updated information from its environment. The client should play the content continuously over the Period. Once the client has consumed the media contained in a Segment towards the end of the media announced in the Representation, the Media Presentation is ended, a new Period is started, or the MPD needs to be refetched.
[0073] Descriptors in DASH
[0077] MPEG-DASH introduces the concept of descriptors, which provide application-specific information about media content. Descriptor elements are all structured in the same way: they contain an @schemeIdUri attribute that provides a URI identifying the scheme, an optional @value attribute, and an optional @id attribute. The semantics of the elements are specific to the scheme used. The URI identifying the scheme can be a URN (Universal Resource Name) or a URL (Universal Resource Locator). The MPD does not provide any specific information on how to use these elements. It is up to applications using the DASH format to instantiate descriptor elements with the appropriate scheme information. A DASH application that uses one of these elements must first define a scheme identifier in the form of a URI, and then define the value space of the element when that scheme identifier is used. If structured data is needed, any extension elements or attributes may be defined in a separate namespace. Descriptors can appear at several levels within the MPD: -The presence of an element at the MPD level means that the element is a child of the MPD element. The presence of an element at the AdaptationSet level means that the element is a child element of the AdaptationSet element. The presence of an element at the representation level means that the element is a child element of the Representation element.
[0074] Preselection
[0078] In MPEG-DASH, a bundle is a set of media components that can be consumed together by one decoder instance. Each bundle contains decoder-specific information and contains the main media component that bootstraps the decoder. A PreSelection defines the subset of media components in a bundle that are expected to be consumed together.
[0075]
[0079] An AdaptationSet that contains a main media component is called a main AdaptationSet. The main media component is always included in any PreSelection associated with a bundle. In addition, each bundle may contain one or more partial AdaptationSets. A partial AdaptationSet may only be processed in combination with the main AdaptationSet.
[0076]
[0080] A Preselection may be defined through the PreSelection element, as defined in Table 4. The PreSelection selection is based on the attributes and elements contained in the PreSelection element.
[0077] [Table 4]
[0078] Adaptive Streaming of Point Clouds
[0081] While traditional multimedia applications such as video remain popular, there is significant interest in new media such as virtual reality (VR) and immersive 3D graphics. High-quality 3D point clouds have recently emerged as an advanced representation of immersive media, enabling new forms of interaction and communication with virtual worlds. The large amount of information required to represent such dynamic point clouds necessitates efficient encoding algorithms. MPEG's 3DG workgroup is currently working on developing a video-based point cloud compression standard, with the Community Draft (CD) version released at the MPEG#124 conference. The latest version of the CD defines a bitstream for compressed dynamic point clouds. In parallel, MPEG is also developing a system standard for transporting point cloud data.
[0079]
[0082] The above point cloud standards only address the encoding and storage aspects of point clouds. However, it is conceivable that practical point cloud applications will require streaming of point cloud data over a network. Such applications may perform live or on-demand streaming of point cloud content depending on how the content is generated. Furthermore, due to the large amount of information required to represent a point cloud, such applications need to support adaptive streaming techniques to avoid overloading the network and provide an optimal viewing experience at any given moment in relation to the network capacity at that moment.
[0080]
[0083] One strong candidate for adaptive delivery of point clouds is Dynamic Adaptive Streaming over HTTP (DASH). However, the current MPEG-DASH standard does not provide any signaling mechanism for point cloud media, including point cloud streams based on the MPEG V-PCC standard. Therefore, it is important to define new signaling elements that allow streaming clients to identify point cloud streams and their component substreams in media presentation descriptor (MPD) files. In addition, there is also a need to signal different kinds of metadata associated with point cloud components so that streaming clients can select the best version of the point cloud or its components that they can support.
[0081]
[0084] Unlike traditional media content, V-PCC media content is composed of several components, and some components have multiple layers. Each component (and / or layer) is encoded separately as a substream of the V-PCC bitstream. Some component substreams, such as geometry and occupancy maps (in addition to some attributes such as texture), are encoded using a traditional video encoder (e.g., H.264 / AVC or HEVC). However, these substreams, along with additional metadata, need to be decoded together to render the point cloud.
[0082]
[0085] Several XML elements and attributes are defined. These XML elements are defined in a separate namespace "urn:mpeg:mpegI:vpcc:2019". The namespace designator "vpcc:" is used to refer to this namespace in this document.
[0083] Signaling of V-PCC components in DASH MPD
[0086] Each V-PCC component and / or component layer can be represented in the DASH manifest (MPD) file as a separate AdaptationSet (hereinafter "component AdaptationSet"), with an additional AdaptationSet (hereinafter "main AdaptationSet") serving as the main access point for the V-PCC content. In another embodiment, one adaptation set per component per resolution is signaled.
[0084]
[0087] In an embodiment, the adaptation sets of a V-PCC stream, including all V-PCC component AdaptationSets, shall have the value of the @codecs attribute (e.g., as defined in the V-PCC) set to "vpc1", indicating that the MPD relates to cloud points. In another embodiment, only the main AdaptationSet has the @codecs attribute set to "vpc1", while the @codecs attribute of the AdaptationSet of the point cloud component (or each Representation, if the @codecs of the AdaptationSet element is not signaled) is set based on the respective codec used to encode that component. For video coding components, the value of @codecs shall be set to "resv.pccv.XXXX", where XXXX corresponds to the four-character code (4CC) of the video codec (e.g., avc1 or hvc1).
[0085]
[0088] To identify the type of V-PCC component within a component AdaptationSet (e.g., occupancy map, geometry, or attributes), the EssentialProperty descriptor may be used with a @schemeIdUri attribute equal to "urn:mpeg:mpegI:vpcc:2019:component". This descriptor is called the VPCCComponent descriptor.
[0086]
[0089] At the adaptation set level, one VPCCComponent descriptor may be signaled for each point cloud component that is presented in the adaptation set's Representation.
[0087]
[0090] In an embodiment, the @value attribute of the VPCCComponent descriptor shall not be present. The VPCCComponent descriptor may include elements and attributes as specified in Table 5.
[0088] [Table 5]
[0089] [Table 6]
[0090]
[0091] The data types of the various elements and attributes of the VPCCComponent descriptor may be defined in the following XML schema:
number
[0091]
[0092] In an embodiment, the main AdaptationSet shall contain either one initialization segment at the adaptation set level or multiple initialization segments at the representation level (one for each Representation). In an embodiment, the initialization segment shall contain a V-PCC sequence parameter set used to initialize the V-PCC decoder, as defined in the community draft. In the case of one initialization segment, the V-PCC sequence parameter sets of all Representations may be included in the initialization segment. If more than one Representation is signaled in the main AdaptationSet, the initialization segment for each Representation may contain the V-PCC sequence parameter set for that particular Representation. If the ISO Base Media File Format (ISOBMFF), as defined in the WD of ISO / IEC 23090-10, is used as the media container for V-PCC content, the initialization segment may also contain a meta box, as defined in ISO / IEC 14496-12. This meta box contains one or more VPCCGroupBox instances, as defined in the VPCC CD, that provide metadata information describing tracks and relationships between tracks at the file format level.
[0092]
[0093] In an embodiment, a media segment of a Representation of a main AdaptationSet contains one or more track fragments of a V-PCC track defined in the community draft, and a media segment of a Representation of a constituent AdaptationSet contains one or more track fragments of the corresponding constituent track at the file format level.
[0093]
[0094] In another embodiment, an additional attribute, referred to herein as the @videoCodec attribute, is defined in the VPCCComponent descriptor, whose value indicates the codec used to encode the corresponding point cloud component. This allows for supporting situations where more than one point cloud component is present in an AdaptationSet or Representation.
[0094]
[0095] In another embodiment, the Role descriptor element may be used in conjunction with newly defined values of the V-PCC components to indicate the role (e.g., geometry, occupancy map, or attribute) of the corresponding AdaptationSet or Representation. For example, the geometry, occupancy map, and attribute components may have the following corresponding values, respectively: vpcc-geometry, vpcc-occupancy, and vpcc-attribute. Additional EssentialProperty descriptor elements similar to those listed in Table 5 minus the component type attribute may be signaled at the adaptation set level to identify the component's layer and attribute type (if the component is a point cloud attribute).
[0095] V-PCC adaptive set grouping
[0096] A streaming client can identify the type of point cloud component in an AdaptationSet or Representation by checking the VPCCComponent descriptor in the corresponding element. However, a streaming client also needs to distinguish between the different point cloud streams present in the MPD file and identify each of their component streams.
[0096]
[0097] An EssentialProperty element with an @schemeIdUri attribute equal to "urn:mpeg:mpegI:vpcc:2019:vpc" may be introduced, which is referred to herein as a VPCC descriptor. At most one VPCC descriptor may exist at the adaptation set level of the main AdaptationSet of a point cloud. If more than one Representation is present in the main AdaptationSet, at most one VPCC descriptor may exist at the representation level (i.e., within each Representation element). Table 6 shows the attributes of a VPCC descriptor according to an embodiment.
[0097] [Table 7]
[0098]
[0098] If two or more versions of a point cloud are available (e.g., at different resolutions), each version may exist in a separate component AdaptationSet containing one Representation and a VPCC descriptor with the same value for the @pcId attribute. In another embodiment, different versions of a point cloud may be signaled as Representations in one (main) AdaptationSet. In such a case, a VPCC descriptor shall be present in each Representation, and the @pcId attribute may be signaled with the same value for all Representations in the main AdaptationSet, or may be omitted.
[0099]
[0099] In another embodiment, a PreSelection is signaled in the MPD with a value of the @preselectionComponents attribute containing the ID of the point cloud's main AdaptationSet followed by the IDs of the component AdaptationSets corresponding to the point cloud components. The @codecs attribute of the PreSelection shall be set to "vpc1", indicating that the PreSelection media is a video-based point cloud. The PreSelection may be signaled using a PreSelection element within a Period element, or a preselection descriptor at the adaptation set level (or at the representation level, if multiple versions / representations are available for the same point cloud). When a PreSelection element is used and two or more versions of the same point cloud are available, each version is signaled in a separate PreSelection element, and the first ID in the ID list of the @preselectionComponents attribute is the ID of the Representation of the corresponding point cloud version in the main AdaptationSet. Figure 6 shows an example DASH configuration for grouping V-PCC components belonging to one point cloud in an MPEG-DASH MPD file.
[0100]
[0100] Using the preselection descriptor, this grouping / association may be signaled as follows:
number
[0101]
[0101] In another embodiment, the main AdaptationSet of a point cloud or its Representation may list the identifiers of the constituent AdaptationSets and / or Representations using the @associationId attribute defined in ISO / IEC 23009-1, with the @associationType value set to 4CC of V-PCC (i.e., "vpc1").
[0102] In another embodiment, the main AdaptationSet of a point cloud or its Representation may list identifiers of the constituent AdaptationSets and / or Representations using the @dependencyId attribute defined in ISO / IEC 23009-1, since to reconstruct the point cloud, segments in the main AdaptationSet need to be decoded together with segments from the constituent AdaptationSets of the point cloud components, thus creating an inherent dependency.
[0103] Signaling Component Metadata
[0103] Geometry metadata and attribute metadata are typically used for rendering. They are signaled within the parameter set of the V-PCC bitstream. However, it may be necessary to signal these metadata elements in the MPD so that the streaming client can obtain information about these metadata elements as early as possible. In addition, the streaming client may choose between multiple versions of the point cloud with different geometry and attribute metadata values (e.g., based on whether the client supports the signaled values).
[0104] Geometry Metadata Signaling A SupplementalProperty element may be introduced with an @schemeIdUri attribute equal to "urn:mpeg:mpegI:vpcc:2019:geom_meta", which is referred to herein as a geometry metadata descriptor or geoMeta descriptor. At most one geomMeta descriptor may be present at an MPD level, in which case it applies to all point cloud geometry components signaled in the MPD, unless overridden by a geoMeta descriptor at a lower level as discussed below. At most one geomMeta descriptor may be present at the adaptation set level within the main AdaptationSet. At most one geomMeta descriptor may be present at the representation level within the main AdaptationSet. If a geomMeta descriptor is present at a particular level, the geomMeta descriptor overrides any geomMeta descriptor signaled at a higher level.
[0105]
[0105] In an embodiment, the @value attribute of the geomMeta descriptor shall not be present. In an embodiment, the geomMeta descriptor includes the elements and attributes specified in Table 7.
[0106] [Table 8]
[0107] [Table 9]
[0108] [Table 10]
[0109]
[0106] In an embodiment, the data types of the various elements and attributes of the geomMeta descriptor may be as defined in the following XML Schema:
number
number
[0110] Attribute Metadata Signaling A SupplementalProperty element may be introduced with an @schemeIdUri attribute equal to "urn:mpeg:mpegI:vpcc:2019:attr_meta", which is referred to herein as an attribute metadata descriptor or attrMeta descriptor. At most one attrMeta descriptor may be present at the adaptation set level within the main AdaptationSet. At most one attrMeta descriptor may be present at the representation level within the main AdaptationSet. If an attrMeta descriptor is present at the representation level, it overrides any attrMeta descriptor signaled at the adaptation set level of the AdaptationSet to which the Representation belongs.
[0111]
[0108] In an embodiment, there is no @value attribute of the attrMeta descriptor. In an embodiment, the attrMeta descriptor may include the elements and attributes specified in Table 8.
[0112] [Table 11]
[0113] [Table 12]
[0114] [Table 13]
[0115]
[0109] In an embodiment, the data types of the various elements and attributes of the attrMeta descriptor may be as defined in the following XML Schema:
number
[0116] Streaming of client behavior
[0110] The DASH client (decoder node) is guided by the information provided in the MPD. The following is an example client behavior for processing streaming point cloud content according to the signaling presented herein, assuming an embodiment in which the association of constituent AdaptationSets to the main point cloud AdaptationSet is signaled using a VPCC descriptor. Figure 7 is a flow diagram illustrating an example streaming client process according to an embodiment.
[0117]
[0111] The client first issues an HTTP request to download the MPD file from the content server at 711. The client then parses the MPD file to generate corresponding in-memory representations of the XML elements in the MPD file.
[0118] Next, to identify the available point cloud media content within the Period, the streaming client scans the AdaptationSet elements to find an AdaptationSet with a @codecs attribute set to "vpc1" and a VPCC descriptor element, at 713. The resulting subset is the main AdaptationSet for the point cloud content.
[0119]
[0113] Next, at 715, the streaming client identifies the number of unique point clouds by checking the VPCC descriptors of those AdaptationSets and groups AdaptationSets that have the same @pcId value in the VPCC descriptor as versions of the same content.
[0120] At 717, a group of AdaptationSets with @pcId values corresponding to the point cloud content the user wants to stream is identified. If the group contains more than one AdaptationSet, the streaming client selects the AdaptationSet with the supported version (e.g., video resolution). Otherwise, a single AdaptationSet from the group is chosen.
[0121] Next, at 719, the streaming client checks the VPCC descriptor of the selected AdaptationSet to identify the AdaptationSets of the point cloud components. These are identified from the values of the @occupancyId, @geometryId, and @attributeId attributes. If the geomMeta and / or attrMeta descriptors are present in the selected main AdaptationSet, the streaming client can identify whether it supports the signaled rendering configuration of the point cloud stream before downloading any segments. Otherwise, the client needs to extract this information from the initialization segment.
[0122]
[0116] Next, at 721, the client starts streaming the point cloud by downloading the initialization segment of the main AdaptationSet, which contains the parameter set required to initialize the V-PCC decoder.
[0123]
[0117] At 723, the initialization segment of the video encoding component stream is downloaded and cached in memory.
[0124]
[0118] At 725, the streaming client then begins downloading timed media segments from the main AdaptationSet and the constituent AdaptationSets in parallel via HTTP, and the downloaded segments are stored in an in-memory segment buffer.
[0125]
[0119] At 727, the timed media segments are removed from each buffer and concatenated with each initialization segment.
[0126]
[0120] Finally, in 729, the media container (eg ISOBMFF) is parsed to extract elementary stream information and structure the V-PCC bitstream according to the V-PCC standard, and the bitstream is then passed to a V-PCC decoder.
[0127] Although features and elements are described above in particular combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with the other features and elements. In addition, the methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of non-transitory computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, optical media such as magneto-optical media and CD-ROM disks, and digital versatile disks (DVDs). A processor associated with software may be used to implement a radio frequency transceiver for use in a WTRU 102, a UE, a terminal, a base station, an RNC, or any host computer.
[0128]
[0122] Furthermore, in the above-described embodiments, processing platforms, computing systems, controllers, and other devices including processors are described. These devices may include at least one central processing unit ("CPU") and memory. Those skilled in the art of computer programming will recognize that references to acts and symbolic representations of operations or instructions may be performed by various CPUs and memories. Such acts and operations or instructions may be referred to as being "executed," "computer-executed," or "CPU-executed."
[0129]
[0123] Those skilled in the art will appreciate that the symbolic representations of operations and calculations or instructions include the manipulation of electrical signals by a CPU. The electrical system may cause the transformation or reduction of electrical signals and the maintenance of data bits in memory locations in a memory system, thereby reconfiguring or otherwise altering the CPU's calculations and other processing of signals. The memory locations where the data bits are maintained are physical locations that have specific electrical, magnetic, optical, or organic properties that correspond to or represent the data bits. It should be understood that the exemplary embodiments are not limited to the above platforms or CPUs, and that other platforms and CPUs may also support the provided methods.
[0130]
[0124] Data bits may also be maintained on computer-readable media including magnetic disks, optical disks, and any other volatile (e.g., random access memory ("RAM")) or non-volatile (e.g., read-only memory ("ROM")) mass storage system that is CPU-readable. The computer-readable media may include cooperating or interconnected computer-readable media that reside solely on the processing system or that are distributed across multiple interconnected processing systems that may be local or remote to the processing system. It is understood that exemplary embodiments are not limited to the above memories, and that other platforms and memories may also support the described methods.
[0131] In an exemplary embodiment, any operations, processes, etc. described herein may be implemented as computer-readable instructions stored on a computer-readable medium. The computer-readable instructions may be executed by a processor of a mobile unit, a network element, and / or any other computing device.
[0132] There is little distinction between hardware and software implementations of aspects of the system. The use of hardware or software is generally (but not always, in that in certain circumstances the choice between hardware and software may be greater) a design choice representing a cost-efficiency tradeoff. There may be various means (e.g., hardware, software, and / or firmware) by which the processes, and / or systems, and / or other techniques described herein may be implemented, and the preferred means may vary with the context in which the processes, and / or systems, and / or other techniques are deployed. For example, if an implementer determines that speed and accuracy are paramount, the implementer may opt for a primarily hardware and / or firmware implementation. If flexibility is paramount, the implementer may opt for a primarily software implementation. Alternatively, the implementer may opt for some combination of hardware, software, and / or firmware.
[0133] The above detailed description has set forth various embodiments of devices and / or processes through the use of block diagrams, flowcharts, and / or examples. To the extent that such block diagrams, flowcharts, and / or examples include one or more functions and / or operations, those skilled in the art will appreciate that each function and / or operation within such block diagrams, flowcharts, or examples may, individually and / or collectively, be implemented by a wide variety of hardware, software, firmware, or substantially any combination thereof. By way of example, suitable processors include general-purpose processors, special-purpose processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors in association with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), an application-specific standard product (ASSP), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and / or a state machine.
[0134]
[0128] While features and elements are provided above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. The present disclosure should not be limited with respect to the specific embodiments described herein, which are intended as examples of various aspects. As will be apparent to those skilled in the art, many modifications and variations can be made without departing from the spirit and scope of the present disclosure. No element, operation, or instruction used in the description of the present application should be construed as critical or essential to the invention unless expressly provided as such. Functionally equivalent methods and apparatuses within the scope of the present disclosure, in addition to those listed herein, will be apparent to those skilled in the art from the above description. Such modifications and variations are intended to be within the scope of the appended claims. The present disclosure should be limited only by the appended claims and the full range of equivalents to which such claims are entitled. It is understood that the present disclosure is not limited to a particular method or system.
[0135] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the terms “station” and its abbreviation “STA,” and “user equipment” and its abbreviation “UE,” when referred to herein, may mean (i) a wireless transmit and / or receive unit (WTRU) such as described below, (ii) any of several embodiments of a WTRU such as described below, (iii) a wireless-enabled and / or wired-enabled (e.g., connectable) device configured with, among other things, some or all of the structure and functionality of a WTRU such as described below, (iii) a wireless-enabled and / or wired-enabled device configured with less than all of the structure and functionality of a WTRU such as described below, or (iv) the like.
[0136] In certain exemplary embodiments, some portions of the subject matter described herein may be implemented via an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), and / or other integrated format. However, those skilled in the art will recognize that some aspects of the embodiments disclosed herein may equally be implemented, in whole or in part, in an integrated circuit, as one or more computer programs running on one or more computers (e.g., as one or more programs running on one or more computer systems), as one or more programs running on one or more processors (e.g., as one or more programs running on one or more microprocessors), as firmware, or as substantially any combination thereof, and that designing circuitry and / or writing software and firmware code is well within the skill of those skilled in the art in light of this disclosure. Additionally, those skilled in the art will understand that the subject matter mechanisms described herein may be distributed as program products in a wide variety of forms, and that exemplary embodiments of the subject matter described herein apply regardless of the particular type of signal-bearing medium used in the actual implementation of the distribution. Examples of signal bearing media include, but are not limited to, recordable-type media such as floppy disks, hard disk drives, CDs, DVDs, digital tape, computer memory, and transmission-type media such as digital and / or analog communications media (e.g., fiber optic cables, wave guides, wired communications links, wireless communications links, etc.).
[0137]
[0131] The present description sometimes depicts different components contained within or connected to other different components. It should be understood that such illustrated architectures are merely examples, and that in fact, many other architectures that achieve the same functionality are possible. In a conceptual sense, any arrangement of components that achieve the same functionality is, in effect, "associated" such that the desired functionality can be achieved. Thus, any two components combined herein to achieve a particular function can be viewed as "associated" with each other such that the desired functionality is achieved, regardless of the architecture or intervening components. Similarly, any two components so associated can also be viewed as being "operably connected" or "operably coupled" to each other to achieve the desired functionality, and any two components so associateable can also be viewed as being "operably coupleable" to each other to achieve the desired functionality. Examples of operably coupleable components include, but are not limited to, physically matable and / or physically interacting components, wirelessly interacting and / or wirelessly interacting components, and / or logically interacting and / or logically interacting components.
[0138] With respect to the use of virtually any plural and / or singular term herein, those skilled in the art can convert from the plural to the singular and / or from the singular to the plural as appropriate to the situation and / or application. Various singular / plural permutations may be expressly set forth herein for clarity.
[0139]
[0133] In general, those skilled in the art will understand that the terms used in this specification, and particularly in the appended claims (e.g., the body of the appended claims), are generally intended as "open" terms (e.g., the term "comprising" should be interpreted as "including, but not limited to," the term "having" should be interpreted as "having at least," the term "comprising" should be interpreted as "including, but not limited to," etc.). Those skilled in the art will further understand that where a specific number of introduced claim items is intended, such intention will be expressly set forth in the claim; in the absence of such a statement, no such intention exists. For example, where only one item is intended, the term "one" or similar language may be used. As an aid to understanding, the following appended claims and / or description herein may include the use of the introductory phrases "at least one" and "one or more" to introduce claim items. However, the use of such phrases should not be construed as implying that introducing a claim statement with the indefinite article "a" or "an" limits any particular claim including such an introduced claim statement to embodiments including only one such statement, even if the same claim includes the introductory phrases "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" and / or "an" should be interpreted to mean "at least one" or "one or more"). The same applies to the use of definite articles used to introduce claim statements. Additionally, those skilled in the art will recognize that even when a specific number of introduced claim statements is explicitly recited, such recitation should be interpreted to mean at least the recited number (e.g., the bare recitation of "two statements" without other qualification means at least two statements or more than two statements).Furthermore, when phrases similar to "at least one of A, B, and C, etc." are used, such structures are generally intended in the sense that one of ordinary skill in the art would understand the phrase (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, systems having A only, B only, C only, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). When phrases similar to "at least one of A, B, or C, etc." are used, such structures are generally intended in the sense that one of ordinary skill in the art would understand the phrase (e.g., "a system having at least one of A, B, or C" includes, but is not limited to, systems having A only, B only, C only, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). Those skilled in the art will further appreciate that nearly all disjunctive words and / or disjunctive phrases presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibility of including one of those terms, either one of those terms, or both terms. For example, the phrase "A or B" is understood to include the possibilities of "A" or "B" or "A and B." Furthermore, the term "any of," when used herein following a list of multiple items and / or multiple categories of items, is intended to include "any of," "any combination of," "any plurality of," and / or "any combination of multiple of," the items and / or categories of items, individually or in combination with other items and / or other categories of items. Furthermore, when used herein, the term "set" or "group" is intended to include any number of items, including zero. Furthermore, when used herein, the term "number" is intended to include any number, including zero.
[0140]
[0134] Additionally, those skilled in the art will recognize that when features or aspects of the present disclosure are described in terms of a Markush group, the present disclosure is also thereby described in terms of any individual element or subgroup of elements of the Markush group.
[0141] As will be understood by those of skill in the art, for any and all purposes, including with respect to providing a written description, all ranges disclosed herein encompass any and all possible subranges and combinations of subranges. Any recited range can be readily recognized as fully descriptive and allowing for division of the range into at least 2, 3, 4, 5, 10, etc. divisions. As a non-limiting example, each range discussed herein can be readily divided into a lower third, a middle third, and an upper third, etc. As will also be understood by those of skill in the art, all terms such as "up to," "at least," "greater than," "less than," etc., are inclusive of the recited numbers and refer to ranges that can be further divided into subranges as described above. Finally, as will be understood by those of skill in the art, a range includes its individual elements. Thus, for example, a group having 1 to 3 cells refers to a group having 1, 2, or 3 cells. Similarly, a group having 1 to 5 cells refers to a group having 1, 2, 3, 4, or 5 cells, etc.
[0142]
[0136] Furthermore, the claims should not be read as limited to the order or elements provided unless stated to that effect. In addition, the use of the term "means for" in any claim is intended to invoke 35 U.S.C. 112(f) or means-plus-function claim format, and any claim without the term "means for" is not so intended.
[0143]
[0137] Although the invention is shown and described herein with reference to specific embodiments, it is not intended that the invention be limited to the details shown. Rather, various changes may be made in the details within the scope and breadth of equivalents of the claims without departing from the invention.
[0144]
[0138] Throughout this disclosure, those skilled in the art will understand that certain exemplary embodiments may be used in alternative exemplary embodiments or in combination with other exemplary embodiments.
[0145]
[0139] While features and elements are described above in particular combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with the other features and elements. In addition, the methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of non-transitory computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, optical media such as magneto-optical media and CD-ROM disks, and digital versatile disks (DVDs). A processor associated with software may be used to implement a radio frequency transceiver for use in a WRTU, UE, terminal, base station, RNC, or any host computer.
[0146]
[0140] Furthermore, in the above-described embodiments, processing platforms, computing systems, controllers, and other devices including processors are described. These devices may include at least one central processing unit ("CPU") and memory. Those skilled in the art of computer programming will recognize that references to acts and symbolic representations of operations or instructions may be performed by various CPUs and memories. Such acts and operations or instructions may be referred to as being "executed," "computer-executed," or "CPU-executed."
[0147]
[0141] Those skilled in the art will appreciate that the symbolic representations of operations and calculations or instructions include the manipulation of electrical signals by a CPU. The electrical system may cause the transformation or reduction of the electrical signals and the maintenance of data bits in memory locations in a memory system, thereby representing the data bits, which reconfigures or otherwise alters the CPU's operations and other processing of the signals. The memory locations in which the data bits are maintained are physical locations that have particular electrical, magnetic, optical, or organic properties that correspond to or represent the data bits.
[0148]
[0142] Data bits may also be maintained on computer-readable media including magnetic disks, optical disks, and any other volatile (e.g., random access memory ("RAM")) or non-volatile (e.g., read-only memory ("ROM")) mass storage system that is CPU-readable. The computer-readable media may include cooperating or interconnected computer-readable media that reside solely on the processing system or that are distributed across multiple interconnected processing systems that may be local or remote to the processing system. It is understood that exemplary embodiments are not limited to the above memories, and that other platforms and memories may also support the described methods.
[0149] No element, act, or instruction used in the description herein should be construed as critical or essential to the present invention unless expressly described as such. Additionally, the article "a" is intended to encompass one or more items. Where only one item is intended, the term "one" or similar language is used. Furthermore, the term "any of," followed by a list of multiple items and / or multiple categories of items, as used herein, is intended to include "any of," "any combination of," "any plurality of," and / or "any combination of multiple of," the items and / or categories of items, individually or in combination with other items and / or other categories of items. Furthermore, as used herein, the term "set" is intended to include any number of items, including zero. Furthermore, as used herein, the term "number" is intended to include any number, including zero.
[0150]
[0144] Furthermore, the claims should not be read as limited to the described order or elements unless expressly stated to that effect. In addition, use of the term "means for" in any claim is intended to invoke 35 U.S.C. 112(f), and any claim without the term "means for" is not so intended.
[0151]
[0145] By way of example, suitable processors include general purpose processors, special purpose processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application specific integrated circuits (ASICs), application specific standard products (ASSPs), field programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines.
[0152] Software and associated processors may be used to implement a radio transmit / receive unit (WRTU), a radio frequency transceiver for use in a user equipment (UE), a terminal, a base station, a mobility management entity (MME) or an evolved packet core (EPC), or any host computer. The WRTU may be used in conjunction with other components such as a camera, a video camera module, a videophone, a speakerphone, a vibration device, a speaker, a microphone, a television transceiver, a hands-free headset, a keyboard, a Bluetooth module, a frequency modulation (FM) radio unit, a near field communication (NFC) module, a liquid crystal display (LCD) display unit, an organic light emitting diode (OLED) display unit, a digital music player, a media player, a video game player module, an internet browser, and / or any wireless local area network (WLAN) or ultra-wideband (UWB) module.
[0153]
[0147] Although the present invention has been described with respect to a communications system, it is contemplated that the system may be implemented in software on a microprocessor / general purpose computer (not shown). In particular embodiments, one or more of the functions of the various components may be implemented in software controlling a general purpose computer.
[0154]
[0148] In addition, although the invention has been shown and described with reference to specific embodiments, it is not intended that the invention be limited to the details shown. Rather, various changes may be made in the details, within the scope and breadth of equivalents of the claims, without departing from the invention.
Claims
1. a decoding node, receiving information associated with streaming PC data (PCD) corresponding to a point cloud (PC) in a Dynamic Adaptive Streaming over Hypertext Transfer Protocol (HTTP) (DASH) Media Presentation Description (MPD); The DASH MPD includes at least: a main adaptation set (AS) of the PC, the main AS including at least a value of a codec attribute indicating that the main AS corresponds to video-based point cloud compressed (V-PCC) data, and information indicating an initialization segment including at least one V-PCC sequence parameter set of a representation of the PC; a plurality of component ASs, each of which corresponds to one of a plurality of V-PCC components, and each of which includes at least a V-PCC component descriptor identifying a type of the corresponding V-PCC component, the type including a geometry, occupancy, or attribute, and information indicating at least one property of the corresponding V-PCC component; and receiving the DASH MPD via a network; a decoding node comprising: a processor configured to perform
2. The decoding node of claim 1 , wherein the processor is further configured to parse the DASH MPD to generate a representation of the DASH MPD.
3. The decoding node of claim 1 , wherein the processor is further configured to identify available point cloud media content based on one or more of the values of the codec attributes or the initialization segment.
4. The decoding node of claim 1 , wherein the processor is further configured to determine a number of unique point clouds based on the V-PCC component descriptors.
5. The decoding node of claim 1 , wherein the processor is further configured to stream the PC by downloading the initialization segment of the main AS.
6. the processor receiving timed segments from the main AS and the plurality of constituent ASs via HTTP; storing the timed segments in a memory buffer; The decoding node of claim 5 , further configured to:
7. The decoding node of claim 6 , wherein the processor is further configured to concatenate the timed segment with a respective initialization segment.
8. The decoding node of claim 7 , wherein the processor is further configured to generate a V-PCC bitstream using the timed segments.
9. The decoding node of claim 8 , wherein the processor is further configured to decode the V-PCC bitstream.
10. 2. The decoding node of claim 1, wherein the International Organization for Standardization (ISO) Base Media File Format (ISOBMFF) is used as a media container for V-PCC content, and the initialization segment of the main AS includes a MetaBox containing one or more V-PCCGroupBox instances that provide metadata information associated with a V-PCC track.
11. The decoding node of claim 1 , wherein the initialization segment of the main AS includes information indicating one initialization segment at an adaptation level and V-PCC sequence parameter sets of multiple representations of the main AS.
12. The decoding node of claim 1, wherein the main AS includes information indicating an initialization segment for each of a plurality of representations of the PC, and each initialization segment corresponding to a representation of the PC includes a V-PCC sequence parameter set for that representation.
13. The decoding node of claim 1 , wherein the V-PCC component descriptor includes information indicating video codec attributes of a codec used to encode the corresponding PC component.
14. The decoding node of claim 1 , wherein the AS includes information indicating a role descriptor DASH element having a value indicating one of a geometry, an occupancy map, or an attribute of a corresponding V-PCC component.
15. The decoding node of claim 1 , wherein the decoding node is a DASH streaming client.
16. 1. A method of a decoding node streaming point cloud (PCD) data (PCD) corresponding to a point cloud (PC) including a plurality of video-based point cloud compressed (V-PCC) components that make up the PC over a network using Hypertext Transfer Protocol (HTTP), comprising: receiving information associated with streaming a PCD corresponding to a PC in a Dynamic Adaptive Streaming over HTTP (DASH) Media Presentation Description (MPD); The DASH MPD includes at least: a main adaptation set (AS) of the PC, the main AS including at least a value of a codec attribute indicating that the main AS corresponds to V-PCC data, and information indicating an initialization segment including at least one V-PCC sequence parameter set of a representation of the PC; a plurality of component ASs, each of which corresponds to one of a plurality of V-PCC components, and each of which includes at least a V-PCC component descriptor identifying a type of the corresponding V-PCC component, the type including a geometry, occupancy, or attribute, and information indicating at least one property of the corresponding V-PCC component; and receiving the DASH MPD via the network; A method comprising:
17. The method of claim 16 , further comprising parsing the DASH MPD to generate a representation of the DASH MPD.
18. The method of claim 16 , further comprising identifying available point cloud media content based on one or more of the values of the codec attributes or the initialization segments.
19. The method of claim 16 , further comprising determining a number of unique point clouds based on the V-PCC component descriptors.
20. The method of claim 16, further comprising streaming the PC by downloading the initialization segment of the main AS.
21. receiving timed segments from the main AS and the plurality of constituent ASs via HTTP; storing the timed segments in a memory buffer; 21. The method of claim 20, further comprising:
22. 22. The method of claim 21, further comprising concatenating the timed segment with each initialization segment.
23. 23. The method of claim 22, further comprising generating a V-PCC bitstream using the timed segments.
24. 24. The method of claim 23, further comprising decoding the V-PCC bitstream.
25. 17. The method of claim 16, wherein the International Organization for Standardization (ISO) Base Media File Format (ISOBMFF) is used as a media container for V-PCC content, and the initialization segment of the main AS includes a MetaBox containing one or more V-PCCGroupBox instances that provide metadata information associated with a V-PCC track.
26. The method of claim 16, wherein the initialization segment of the main AS includes information indicating one initialization segment at an adaptation level and includes V-PCC sequence parameter sets of multiple representations of the main AS.
27. 17. The method of claim 16, wherein the main AS includes information indicating an initialization segment for each of a plurality of representations of the PC, and each initialization segment corresponding to a representation of the PC includes a V-PCC sequence parameter set for that representation.
28. The method of claim 16, wherein the V-PCC component descriptor includes information indicating video codec attributes of a codec used to encode the corresponding PC component.
29. The method of claim 16, wherein the AS includes information indicating a role descriptor DASH element having a value indicating one of a geometry, an occupancy map, or an attribute of a corresponding V-PCC component.
Citation Information
Patent Citations
Image processing device and image processing method
WO2018150933A1