Decoding node and method for decoding node

TWI937946BActive Publication Date: 2026-09-01INTERDIGITAL VC HOLDINGS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
TW114126654
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-03-08
Filing Date
2020-03-06
Publication Date
2026-09-01
Estimated Expiration
2040-03-05

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently representing and compressing high-quality 3D point clouds for immersive media, particularly in supporting lossy and/or lossless encoding of point cloud geometric coordinates and attributes, which is necessary for applications in remote rendering, virtual reality, and large-scale dynamic 3D graphics.

Method used

A network node is configured to stream point cloud data using HTTP and includes video-based point cloud compression components, generating a Dynamic Adaptive Streaming (DASH) Media Presentation Description (MPD) with adaptability sets and component ASs, utilizing ISO Basic Media File Format (ISOBMFF) as the media container, and signaling separate master ASs for different versions of the point cloud data.

Benefits of technology

Enables efficient and interactive storage and transmission of 3D point clouds, supporting lossy and lossless encoding, and facilitating applications in remote rendering and virtual reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TB001908934_001
    Figure TWG2TB001908934_001
  • Figure TWG2TB001908934_002
    Figure TWG2TB001908934_002
  • Figure TWG2TB001908934_003
    Figure TWG2TB001908934_003
Patent Text Reader

Abstract

Methods, apparatus, and systems are intended to adaptively stream V-PCC (video-based point cloud compression) data using adaptive HTTP streaming protocols such as MPEG DASH. The method includes signaling the point cloud data in a DASH MPD, comprising: a primary adaptability set for the point cloud, the primary adaptability set including at least (1) a @codec attribute set to a unique value indicating that the corresponding adaptability set corresponds to V-PCC data, and (2) an initialization segment containing at least one set of V-PCC sequence parameters for the representation of the point cloud; and multiple component adaptability sets, each component adaptability set corresponding to one of the V-PCC components, and including at least (1) a VPCC component descriptor identifying the type of the corresponding V-PCC component, and (2) at least one property of the V-PCC component; and transmitting a DASH bitstream via the network.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Adaptive Video Streaming [Previous Technology]

[0002] High-quality 3D point clouds have recently emerged as an advanced representation for immersive media. A point cloud consists of a set of points represented in 3D space using coordinates indicating the location of each point and one or more attributes (such as color, transparency, acquisition time, laser reflectivity, or material properties associated with each point). Data used to create point clouds can be captured in various ways. For example, one technique for acquiring point clouds is using multiple cameras and depth sensors. Light detection and ranging (LiDAR) laser scanners are also commonly used to capture point clouds. The number of points required to realistically reconstruct objects and scenes using point clouds is on the order of millions (or even billions). Therefore, efficient representation and compression are necessary for storing and transmitting point cloud data.

[0003] Recent advancements in 3D point acquisition and rendering technologies have led to novel applications in remote rendering, virtual reality, and large-scale dynamic 3D graphics. The 3D Graphics Subgroup of the ISO / IEC JTC1 / SC29 / WG11 Moving Picture Experts Group (MPEG) is currently working on two 3D point cloud compression (PCC) standards: a geometry-based compression standard for static point clouds (point clouds for stationary objects), and a video-based compression standard for dynamic point clouds (point clouds for moving objects). The goal of these standards is to support efficient and interactive storage and transmission of 3D point clouds. One of the requirements of these standards is support for lossy and / or lossless encoding of point cloud geometric coordinates and attributes. [Summary of the Invention]

[0004] A network node may include circuitry comprising any one of a transmitter, a receiver, a processor, and memory. The network node may be configured to stream point cloud data (PCD) corresponding to a point cloud (PC) via a network using the HyperFile Transfer Protocol (HTTP), and may include a plurality of video-based point cloud compression (V-PCC) components comprising the PC. The network node may be configured to generate information associated with the PCD in a Dynamic Adaptive Streaming (DASH) Media Presentation Description (MPD) of the HTTP. The generated information in the DASH MPD may include information indicating at least a master adaptability set (AS) for the PC and a plurality of component ASs. The master AS may include a value indicating at least a codec attribute indicating that the master AS corresponds to V-PCC data, and information containing an initialization segment of at least one V-PCC sequence parameter set for the representation of the PC. Each of the plurality of component ASs may correspond to one of the V-PCC components and may contain a V-PCC component descriptor indicating at least the type of the corresponding V-PCC component, and information on at least one property of the corresponding V-PCC component. The type includes geometry, occupancy, or attributes. The network node may be configured to transmit the DASH MPD over the network. The network node may be configured to transmit the PCD over the network, with the International Organization for Standardization (ISO) Basic Media File Format (ISOBMFF) used as the media container for the V-PCC content. The initialization segment of the master AS may contain a meta-frame containing one or more V-PCC frame instances that provide relay data information associated with the V-PCC trajectory. The initialization segment of the master AS may contain information indicating a single initialization segment at the suitability level, and the V-PCC sequence parameter set for all representations of the master AS. The master AS may contain information indicating an initialization segment for each of the plurality of representations of the PC. Each initialization segment corresponding to the representation of the PC may contain a set of V-PCC sequence parameters for that representation. The V-PCC component descriptor may include information indicating video codec attributes used to encode the corresponding V-PCC component. Each of the plurality of component ASs may contain information indicating a role descriptor DASH element, which has a value indicating one of the geometry, occupancy map, or attribute of the corresponding component. The V-PCC component descriptor information may indicate any layer of the corresponding V-PCC component or the attribute type of the corresponding V-PCC component. The master AS may indicate V-PCC descriptors for indicators of the specific point cloud and component ASs corresponding to the master AS.The network node can also be configured to signal each version of the PC in each of the individual master ASs, provided that the V-PCC data used for the PC includes more than one version of the PC. Each individual master AS contains a single representation corresponding to that version of the PC and its own V-PCC descriptor. All individual master ASs corresponding to different versions of the PC can have the same value for the PC indicator (ID) attribute.

Implementation Method

[0006] An exemplary system in which embodiments may be implemented

[0007] FIG1A is a block diagram illustrating an exemplary video encoding and decoding system 100 in which one or more embodiments may be performed and implemented. System 100 may include a source device 112 that can transmit encoded video information to a destination device 114 via a communication channel 116.

[0008] The source device 112 and / or the destination device 114 can be any of the widest range of devices. In some representative embodiments, the source device 112 and / or the destination device 114 may include a wireless transmitting and / or receiving unit (WTRU), such as a wireless handheld device or any wireless device capable of transmitting video information via communication channel 116, in which case communication channel 116 includes a wireless link. However, the methods, apparatuses, and systems described, disclosed, or otherwise expressly, implicitly, and / or inherently provided (collectively, the “Provided”) herein are not necessarily limited to wireless applications or settings. For example, these techniques can be applied to wireless television broadcasting, cable television transmission, satellite television transmission, Internet video transmission, encoded digital video encoded onto storage media, and / or other scenarios. Communication channel 116 may include and / or may be any combination of wireless or wired media suitable for the transmission of encoded video data.

[0009] Source device 112 may include a video encoder unit 118, a transmit and / or receive (Tx / Rx) unit 120, and / or a Tx / Rx element 122. As shown, source device 112 may include a video source 124. Destination device 114 may include a Tx / Rx element 126, a Tx / Rx unit 128, and / or a video decoder unit 130. As shown, destination device 114 may include a display device 132. Each of the Tx / Rx units 120 and 128 may be or may include a transmitter, a receiver, or a combination of a transmitter and a receiver (e.g., a transceiver or a transmitter-receiver). Each of the Tx / Rx elements 122 and 126 may be, for example, an antenna. According to the present invention, the video encoder unit 118 of source device 112 and / or the video decoder unit 130 of destination device 114 may be configured and / or adapted (collectively, “adapted”) to apply the coding techniques provided herein.

[0010] The source device 112 and the destination device 114 may include other elements / components or arrangements. For example, the source device 112 may be adapted to receive video data from an external video source. The destination device 114 may have an interface with an external display device (not shown) and / or may include and / or use (e.g., integrate) the display device device 132. In some embodiments, the stream generated by the video encoder unit 118 may be transmitted to other devices (such as via direct digital forwarding) without modulating the data onto a carrier signal, and the other devices may or may not modulate the data for transmission.

[0011] The techniques provided herein can be implemented by any digital video encoding and / or decoding device. Although the techniques provided herein are typically implemented by individual video encoding and / or video decoding devices, they can also be implemented by a combined video encoder / decoder, commonly referred to as a “codec.” Source device 112 and destination device 114 are merely examples of such coding devices in which source device 112 can generate (and / or receive video data and generate) encoded video information for transmission to destination device 114. In some representative embodiments, source device 112 and destination device 114 can operate in a substantially symmetrical manner, such that each of devices 112, 114 may include video encoding and decoding components and / or elements (collectively, “elements”). Thus, system 100 can support either one-way or two-way video transmission between source device 112 and destination device 114 (e.g., for video streaming, video replay, video broadcasting, video telephony, and / or video conferencing, etc.). In some representative embodiments, the source device 112 may be, for example, a video streaming server adapted to generate (and / or receive video data and adapted to generate) coded video information for one or more destination devices, wherein the destination devices may communicate with the source device 112 via wired and / or wireless communication systems.

[0012] The external video source and / or video source 124 may be and / or include video capturing devices, such as a video camera, a video file containing previously captured video, and / or a video feed from a video content provider. In some representative embodiments, the external video source and / or video source 124 may generate computer graphics-based data as a combination of source video, or live video, archived video, and / or computer-generated video. In some representative embodiments, when the video source 124 is a video camera, the source device 112 and the destination device 114 may be or may embody a camera phone or video phone.

[0013] The captured, pre-captured, computer-generated video, video feed, and / or other types of video data (collectively referred to as "unencoded video") may be encoded by the video encoder unit 118 to form encoded video information. The Tx / Rx unit 120 may modulate the encoded video information (e.g., according to communication standards, to form one or more modulated signals carrying the encoded video information). The Tx / Rx unit 120 may transmit the modulated signal to its transmitter for transmission. The transmitter may send the modulated signal to the destination device 114 via the Tx / Rx element 122.

[0014] At the destination device 114, the Tx / Rx unit 128 can receive the modulated signal from the channel 116 via the Tx / Rx element 126. The Tx / Rx unit 128 can demodulate the modulated signal to obtain encoded video information. The Tx / Rx unit 128 can transmit the encoded video information to the video decoder unit 130.

[0015] The video decoder unit 130 can decode encoded video information to obtain decoded video data. The encoded video information may include syntax information defined by the video encoder unit 118. The syntax information may include one or more elements (“syntax elements”); some or all of them may be used to decode the encoded video information. The syntax elements may include, for example, characteristics of the encoded video information. The syntax elements may also include characteristics of the unencoded video used to form the encoded video information, and / or descriptions of the processing of the unencoded video.

[0016] The video decoder unit 130 can output decoded video data for later storage and / or display on an external display (not shown). In some representative embodiments, the video decoder unit 130 can output decoded video data to a display device 132. The display device 132 can be and / or may include any individual, multiple, or combination of a variety of display devices suitable for displaying decoded video data to a user. Examples of such display devices include liquid crystal displays (LCDs), plasma displays, organic light-emitting diode displays (OLEDs), and / or cathode ray tubes (CRTs).

[0017] Communication channel 116 can be any wireless or wired communication medium (such as radio frequency (RF) spectrum or one or more physical transmission lines), or any combination of wireless and wired media. Communication channel 116 may form part of a packet-based network, such as a local area network, a wide area network, or a global network, such as the Internet. Communication channel 116 generally represents any suitable communication medium or collection of different communication media for transmitting video data from source device 112 to destination device 114, including any suitable combination of wired and / or wireless media. Communication channel 116 may include routers, switches, base stations, and / or any other means that can be used to facilitate communication from source device 112 to destination device 114. Details of an exemplary communication system that can facilitate such communication between devices 112 and 114 are provided below with reference to Figures 1A and 1B. Details of means that can represent source device 112 and destination device 114 are also provided below.

[0018] The video encoder unit 118 and the video decoder unit 130 may operate according to one or more standards and / or specifications, such as MPEG-2, H.261, H.263, H.264, H.264 / AVC, and / or H.264 extended according to SVC extensions (“H.264 / SVC”). Those skilled in the art will understand that the methods, apparatus, and / or systems described herein are applicable to other video encoders, decoders, and / or codecs implemented according to (and / or conforming to) different standards, or to proprietary video encoders, decoders, and / or codecs (including future video encoders, decoders, and / or codecs). The techniques described herein are not limited to any particular coding standard.

[0019] The relevant portions of H.264 / AVC described above are available from the International Telecommunication Union as ITU-T Recommendation H.264, or more specifically, "ITU-T Rec.264 and ISO / IEC 14496-10 (MPEG4-AVC), March 2010 'Advanced Video Coding for General Audiovisual Services', version 5, which are incorporated herein by reference and may be referred to herein as the H.264 standard, H.264 specification, H.264 / AVC standard and / or specification. The techniques provided herein can be applied to devices that conform to (e.g., substantially conform to) the H.264 standard.

[0020] Although not shown in Figure 1A, each of the video encoder and video decoder units 118, 130 may include an audio encoder and / or an audio decoder (as applicable), and / or be integrated with an audio encoder and / or an audio decoder. The video encoder and video decoder units 118, 130 may include appropriate MUX-DEMUX units, or other hardware and / or software, to handle the encoding of both audio and video in a common stream and / or individual streams. If applicable, the MUX-DEMUX unit may comply with, for example, ITU-T Recommendation H.223 multiplexer protocol and / or other protocols such as User Datagram Protocol (UDP).

[0021] One or more video encoder and / or video decoder units 118, 130 may be included in one or more encoders and / or decoders; any of them may be integrated as part of a codec and may be integrated and / or combined with respective cameras, computers, mobile devices, subscriber devices, broadcasting devices, set-top boxes and / or servers, etc. The video encoder unit 118 and / or the video decoder unit 130 may be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. Either or both of the video encoder and video decoder units 118, 130 may be physically implemented in software, and the operation of the elements of the video encoder unit 118 and / or the video decoder unit 130 may be executed by appropriate software instructions executed by one or more processors (not shown). In addition to the processor, such embodiments may include off-chip components, such as external storage (e.g., in the form of non-volatile memory) and / or input / output interfaces, etc.

[0022] In any embodiment where the operation of the elements of the video encoder and / or video decoder units 118, 130 can be performed by software instructions executable by one or more processors, the software instructions may be stored on a computer-readable medium, including, for example, a magnetic disk, an optical disk, any other volatile (e.g., random access memory (“RAM”)), non-volatile (e.g., read-only memory (“ROM”)) and / or a large storage system readable by the CPU. The computer-readable medium may include cooperative or interconnected computer-readable media that may reside exclusively on the processing system and / or be distributed among multiple interconnected processing systems that may be located locally or remotely to the processing system.

[0023] Figure 1B is a block diagram illustrating an exemplary video encoder unit 118 for use with, for example, a video encoding and / or decoding system of system 100. The video encoder unit 118 may include a video encoder 133, an output buffer 134, and a system controller 136. The video encoder 133 (or one or more elements thereof) may be implemented according to one or more standards and / or specifications, such as, for example, H.261, H.263, H.264, H.264 / AVC, the SVC extension of H.264 / AVC (H.264 / AVC Appendix G), HEVC, and / or the Scalable Extension of HEVC (SHVC), etc. Those skilled in the art will understand that the methods, apparatus, and / or systems provided herein can be adapted to other video encoders implemented according to different standards and / or to proprietary codecs (including future codecs).

[0024] The video encoder 133 can receive video signals from a video source (e.g., video source 124 and / or an external video source). These video signals may include unencoded video. The video encoder 133 can encode the unencoded video and provide an encoded (i.e., compressed) video bitstream (BS) at its output.

[0025] The encoded video bitstream BS can be provided to the output buffer 134. The output buffer 134 can buffer the encoded video bitstream BS and can provide such encoded video bitstream BS as a buffered bitstream (BBS) for transmission via the communication channel 116.

[0026] The buffered bit stream BBS output from the output buffer 134 may be sent to a storage device (not shown) for later viewing or transmission. In some representative embodiments, the video encoder unit 118 may be configured for visual communication, wherein the buffered bit stream BBS may be transmitted via the communication channel 116 at a specified constant and / or variable bit rate (e.g., with delay (e.g., very low or minimal delay)).

[0027] The encoded video bitstream BS, and the subsequent buffered bitstream BBS, can carry bits of encoded video information. The bits of the buffered bitstream BBS can be arranged as a stream of encoded video frames. The encoded video frames can be intra-coded frames (e.g., I-frames) or inter-coded frames (e.g., B-frames and / or P-frames). The stream of encoded video frames can be arranged, for example, as a series of group of frames (GOPs), wherein the encoded video frames of each GOP are arranged in a specified order. Typically, each GOP can begin with an intra-coded frame (e.g., an I-frame), followed by one or more inter-coded frames (e.g., P-frames and / or B-frames). Each GOP can include only a single intra-coded frame; although any GOP can include multiple intra-coded frames. It is anticipated that B-frames may not be used for real-time, low-latency applications, as bidirectional prediction, for example, may result in additional coding latency compared to unidirectional prediction (P-frames). As those skilled in the art will understand, additional and / or other frame types may be used, and the specific ordering of encoded video frames may be modified.

[0028] Each GOP may include syntax data (“GOP syntax data”). The GOP syntax data may be placed in the header of the GOP, in the header of one or more frames of the GOP, and / or elsewhere. The GOP syntax data may indicate the order, quantity, or type, and / or describe the encoded video frames of the respective GOP. Each encoded video frame may contain syntax data (“encoded frame syntax data”). The encoded frame syntax data may indicate and / or describe the encoding pattern used to encode the respective video frames.

[0029] The system controller 136 can monitor various parameters and / or constraints associated with the computing power of the channel 116, the video encoder unit 118, user needs, etc., and can establish target parameters to provide an attendant quality of experience (QoE) suitable for specified constraints and / or channel 116 conditions. One or more target parameters can be adjusted from time to time or periodically according to specified constraints and / or channel conditions. As an example, QoE can be quantitatively evaluated using one or more metrics for evaluating video quality, including, for example, metrics commonly referred to as the relative perceived quality of the encoded video sequence. For example, the relative perceived quality of the encoded video sequence, measured using a peak signal-to-noise ratio (“PSNR”) metric, can be controlled by the bit rate (BR) of the encoded bitstream BS. One or more target parameters (including, for example, a quantization parameter (QP)) can be adjusted to maximize the relative perceived quality of the video within constraints associated with the bit rate of the encoded bitstream BS.

[0030] Figure 2 is a block diagram of a block-based hybrid video encoder 200 for use with a video encoding and / or decoding system such as system 100.

[0031] Referring to FIG2, the block-based hybrid coding system 200 may include a transform unit 204, a quantization unit 206, an entropy writing code unit 208, an inverse quantization unit 210, an inverse transform unit 212, a first adder 216, a second adder 226, a spatial prediction unit 260, a motion prediction unit 262, a reference image storage unit 264, a single or complex filter 266 (e.g., a loop filter) and / or a mode determination and encoder controller unit 280, etc.

[0032] The details of the video encoder 200 are for illustrative purposes only, and real-world implementations may differ. For example, real-world implementations may include more, fewer, and / or different components, and / or may be arranged differently from the arrangement shown in FIG. 2. For example, although shown separately, some or all of the functions of both the transform unit 204 and the quantization unit 206 may be highly integrated in some real-world implementations, such as implementations using the core transform of the H.264 standard. Similarly, the inverse quantization unit 210 and the inverse transform unit 212 may be highly integrated in some implementations of real-world implementations (e.g., H.264 or HEVC standard-compatible implementations), but are also described separately for conceptual purposes.

[0033] As described above, the video encoder 200 can receive video signals at its input 202. The video encoder 200 can generate encoded video information from the received unencoded video and output the encoded video information (e.g., any of in-frame or between-frame formats) from its output 220 in the form of an encoded video bitstream BS. The video encoder 200 can operate, for example, as a hybrid video encoder and employ a block-based coding process to encode the unencoded video. When performing this encoding process, the video encoder 200 can operate on individual frames, frames, and / or images (collectively, "unencoded frames") of the unencoded video.

[0034] To facilitate the block-based encoding process, the video encoder 200 may slice, partition, divide, and / or segment (collectively, “segment”) each unencoded frame received at its input 202 into a plurality of unencoded video blocks. For example, the video encoder 200 may segment an unencoded frame into a plurality of unencoded video segments (e.g., slices), and may (e.g., then) segment each of the unencoded video segments into unencoded video blocks. The video encoder 200 may pass, supply, send, or provide the unencoded video blocks to the spatial prediction unit 260, motion prediction unit 262, mode determination and encoder controller unit 280, and / or the first adder 216. As described in more detail below, unencoded video blocks may be provided on a block-by-block basis.

[0035] Spatial prediction unit 260 may receive uncoded video blocks and encode these video blocks in an intra-frame mode. An intra-frame mode refers to any of several modes of spatial compression, and the encoding in the intra-frame mode attempts to provide spatial compression of the uncoded image. If any spatial compression exists, it may result in reducing or eliminating spatial redundancy of video information within the uncoded image. When forming prediction blocks, spatial prediction unit 260 may perform spatial prediction (or "intra-frame prediction") for each uncoded video block relative to one or more encoded ("coded video block") and / or reconstructed ("reconstructed video block") video blocks of the uncoded image. The encoded and / or reconstructed video block may be an adjacent, neighboring, or close to (e.g., very close to) an uncoded video block.

[0036] The motion prediction unit 262 may receive uncoded video blocks from input 202 and encode them in an inter-frame mode. An inter-frame mode refers to any of several modes of time-based compression, including (for example) P mode (one-way prediction) and / or B mode (two-way prediction). Encoding in the inter-frame mode attempts to provide time-based compression of the uncoded image. If time-based compression is present, it can be used to reduce or remove temporal redundancy of video information between the uncoded image and one or more reference (e.g., adjacent) images. The motion / time prediction unit 262 may perform time prediction (or "inter-frame prediction") for each uncoded video block relative to one or more video blocks ("reference video blocks") of the reference images. The performed time prediction may be one-way prediction (e.g., for P mode) and / or two-way prediction (e.g., for B mode).

[0037] For unidirectional prediction, the reference video block may come from one or more previously encoded and / or reconstructed frames. The encoded and / or reconstructed frames (one or more) may be neighbors of the unencoded frames, adjacent to the unencoded frames, and / or close to the unencoded frames.

[0038] For bidirectional prediction, the reference video block may come from one or more previously encoded and / or reconstructed frames. The encoded and / or reconstructed frame may be an adjacency of the unencoded frame, adjacent to the unencoded frame, and / or close to the unencoded frame.

[0039] If multiple reference frames are used for each video block (as is the case for recent video coding standards such as H.264 / AVC and / or HEVC), then their reference frame indices can be sent to the entropy coding unit 208 for subsequent output and / or transmission. The reference index can be used to identify which reference frames(s) in the reference image storage 264 the timing prediction comes from.

[0040] Although generally highly integrated, the motion estimation and motion compensation functions of the motion / timing prediction unit 262 can be performed by separate entities or units (not shown). Motion estimation can be performed to estimate the motion of each uncoded video block relative to a reference frame video block, and may involve generating a motion vector for the uncoded video block. This motion vector may indicate the displacement of the predicted block relative to the uncoded video block being decoded. This predicted block is based on, for example, a reference frame video block whose pixel differences are found to closely match those of the uncoded video block being written. The match may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), and / or other difference metrics. Motion compensation may involve capturing and / or generating the predicted block based on the motion vector determined by the motion estimation.

[0041] The motion prediction unit 262 can calculate the motion vector of the uncoded video block by comparing it with a reference video block from a reference frame stored in the reference image storage 264. The motion prediction unit 262 can calculate the fractional pixel position value of the reference frame contained in the reference image storage 264. In some cases, the adder 226 of the video encoder 200 or another unit can calculate the fractional pixel position value of the reconstructed video block and store the reconstructed video block together with the calculated fractional pixel position value in the reference image storage 264. The motion prediction unit 262 can interpolate sub-integer pixels of the reference frame (e.g., I-frame and / or P-frame and / or B-frame).

[0042] The motion prediction unit 262 can be configured to encode motion vectors relative to a selected motion predictor. The motion predictor selected by the motion / timing prediction unit 262 can be, for example, a vector equal to the average of the motion vectors of already encoded neighboring blocks. In order to encode the motion vectors of unencoded video blocks, the motion / timing prediction unit 262 can calculate the difference between the motion vectors and the motion predictor to form a motion vector difference value.

[0043] H.264 and HEVC refer to the set of potential reference frames as a "list". The set of reference frames stored in the reference image storage 264 may correspond to this list of reference frames. The motion / time prediction unit 262 may compare reference video blocks from the reference frames in the reference image storage 264 with uncoded video blocks (e.g., P frames or B frames). When the reference frames in the reference image storage 264 contain values ​​of sub-integer pixels, the motion vector calculated by the motion / time prediction unit 262 may reference the sub-integer pixel positions of the reference frames. The motion / time prediction unit 262 may send the calculated motion vector to the entropy coding unit 208 and the motion compensation function of the motion / time prediction unit 262. The motion prediction unit 262 (or its motion compensation function module) may calculate the error value of the prediction block relative to the uncoded video block being coded. The motion prediction unit 262 may calculate prediction data based on the prediction block.

[0044] The mode determination and encoder controller unit 280 can select one of a code writing mode, an intra-frame mode, or an inter-frame mode. The mode determination and encoder controller unit 280 can perform this operation based on, for example, a rate distortion optimization method and / or the error results generated in each mode.

[0045] The video encoder 200 can form a residual block (“residual video block”) by subtracting prediction data provided by the motion prediction unit 262 from the uncoded video block being written. The adder 216 represents a single element or a complex element that can perform the subtraction operation.

[0046] Transform unit 204 may apply a transformation to the residual video block to convert the residual video block from the pixel value domain to the transform domain (e.g., the frequency domain). This transformation may be, for example, any of the transformations provided herein, a Discrete Cosine Transform (DCT), or a conceptually similar transformation. Other examples of transformations include those defined in H.264 and / or HEVC, wavelet transform, integer transform, and / or subband transform. The transformation applied by transform unit 204 to the residual video block produces a corresponding block of transform coefficients for the residual video block (“residual transform coefficients”). These residual transform coefficients may represent the magnitude of the frequency components of the residual video block. Transform unit 204 may forward the residual transform coefficients to quantization unit 206.

[0047] The quantization unit 206 can quantize the residual transform coefficients to further reduce the encoded bit rate. For example, the quantization process can reduce the bit depth associated with some or all of the residual transform coefficients. In some cases, the quantization unit 206 can divide the value of the residual transform coefficients by the quantization level corresponding to QP to form a quantized transform coefficient block. The degree of quantization can be modified by adjusting the QP value. The quantization unit 206 can apply quantization to represent the residual transform coefficients using the desired number of quantization steps; the number of steps used (or the corresponding quantization level value) determines the number of encoded video bits used to represent the residual video block. The quantization unit 206 can obtain the QP value from the rate controller (not shown). After quantization, the quantization unit 206 can provide the quantized transform coefficients to the entropy write coding unit 208 and the inverse quantization unit 210.

[0048] The entropy writing unit 208 may apply entropy writing to quantization transform coefficients to form entropy writing coefficients (i.e., bitstream). The entropy writing unit 208 may use adaptive variable-length writing code (CAVLC), context-adaptive binary arithmetic writing code (CABAC), and / or another entropy writing technique to form the entropy writing coefficients. As will be understood by those skilled in the art, CABAC may require input context information (“context”). For example, this context may be based on adjacent video blocks.

[0049] The entropy coding unit 208 can provide entropy coding coefficients, along with motion vectors and one or more complex reference frame indices, in the form of a raw encoded video bitstream to an internal bitstream format (not shown). This bitstream format can be configured by appending additional information, including headers and / or other information, to the raw encoded video bitstream so that, for example, the video decoder unit 300 (FIG. 3) can decode encoded video blocks from the raw encoded video bitstream to form an encoded video bitstream BS provided to the output buffer 134 (FIG. 1B). After entropy coding, the encoded video bitstream BS provided from the entropy coding unit 208 can be output to, for example, the output buffer 134, and can be transmitted via channel 116 to, for example, the destination device 114 or archived for later transmission or retrieval.

[0050] In some representative embodiments, the entropy writing unit 208 or another unit of the video encoder 133, 200 may be configured to perform other writing functions besides entropy writing. For example, the entropy writing unit 208 may be configured to determine the Code Block Pattern (CBP) value of the video block. In some representative embodiments, the entropy writing unit 208 may perform run length writing of the quantization transform coefficients in the video block. As an example, the entropy writing unit 208 may apply a zigzag scan or other scan pattern to arrange the quantization transform coefficients in the video block, and encode zero runs for further compression. The entropy writing unit 208 may construct header information using appropriate syntax elements for transmission in the encoded video bitstream BS.

[0051] The inverse quantization unit 210 and the inverse transform unit 212 may respectively apply inverse quantization and inverse transform to reconstruct the residual video block in the pixel domain, for example, for later use as one of the reference video blocks (e.g., in one of the reference frames in the reference frame list).

[0052] The mode determination and encoder controller unit 280 can calculate a reference video block by adding the reconstructed residual video block to the prediction block of one of the reference images stored in the reference image storage 264. The mode determination and encoder controller unit 280 can apply one or more interpolation filters to the reconstructed residual video block to calculate sub-integer pixel values ​​(e.g., for half-pixel positions) for motion estimation.

[0053] Adder 226 can add the reconstructed residual video block to the motion-compensated predicted video block to produce a reconstructed video block for storage in reference image storage 264. The reconstructed (pixel value domain) video block can be used by motion prediction unit 262 (or its motion estimation function and / or its motion compensation function) as one of the reference blocks for inter-frame coding of uncoded video blocks in subsequent uncoded video.

[0054] Filter 266 (e.g., a loop filter) may include a deblocking filter. The deblocking filter can operate to remove visual artifacts that may be present in the reconstructed macroblock. These artifacts may be introduced into the encoding process, for example, due to the use of different encoding modes such as Type I, Type P, or Type B. For example, artifacts may be present at the boundaries and / or edges of the received video block, and the deblocking filter can operate to smooth the boundaries and / or edges of the video block to improve visual quality. The deblocking filter may filter the output of adder 226. Filter 266 may include other in-loop filters, such as the Sample Adaptability Shift (SAO) filter supported by the HEVC standard.

[0055] FIG3 is a block diagram illustrating an example of a video decoder 300 used with a video decoder unit such as the video decoder unit 130 of FIG1A. The video decoder 300 may include an input 302, an entropy decoding unit 308, a motion compensation prediction unit 362, a spatial prediction unit 360, an inverse quantization unit 310, an inverse transform unit 312, a reference image storage unit 364, a filter 366, an adder 326, and an output 320. The video decoder 300 can perform a decoding process that is generally the inverse of the encoding process provided relative to the video encoders 133, 200. This decoding process can be performed as described below.

[0056] The motion compensation prediction unit 362 can generate prediction data based on motion vectors received from the entropy decoding unit 308. Motion vectors can be encoded relative to motion predictors used for video blocks corresponding to encoded motion vectors. The motion compensation prediction unit 362 can determine the motion predictor as, for example, the median of motion vectors of blocks adjacent to the video block to be decoded. After determining the motion predictor, the motion compensation prediction unit 362 can decode the encoded motion vector by extracting motion vector differences from the encoded video bitstream BS and adding the motion vector differences to the motion predictor. The motion compensation prediction unit 362 can quantize the motion predictor to the same resolution as the encoded motion vector. In some representative embodiments, the motion compensation prediction unit 362 can use the same precision for some or all encoded motion predictors. As another example, the motion compensation prediction unit 362 can be configured to use any of the above methods, and the method to be used is determined by analyzing data contained in the sequence parameter set, slice parameter set, or picture parameter set obtained from the encoded video bitstream BS.

[0057] After decoding the motion vector, the motion compensation prediction unit 362 can retrieve the predicted video block identified by the motion vector from the reference frame in the reference image storage 364. If the motion vector points to a fractional pixel position (e.g., half a pixel), then the motion compensation prediction unit 362 can interpolate the value of the fractional pixel position. The motion compensation prediction unit 362 can interpolate these values ​​using an adaptive interpolation filter or a fixed interpolation filter. The motion compensation prediction unit 362 can obtain the index of which filter 366 is used from the received encoded video bitstream BS, and in various representative embodiments, obtain the coefficients of the filter 366.

[0058] Spatial prediction unit 360 can use in-frame prediction modes received in the encoded video bitstream BS to form predicted video blocks from spatially adjacent blocks. Inverse quantization unit 310 can inversely quantize (e.g., dequantize) the quantized block coefficients provided in the encoded video bitstream BS and decoded by entropy decoding unit 308. The inverse quantization process may include conventional processes, such as those defined in H.264. The inverse quantization process may include using quantization parameters QP calculated by video encoders 133, 200 for each video block to determine the degree of quantization and / or inverse quantization to be applied.

[0059] The inverse transform unit 312 may apply an inverse transform (e.g., the inverse transform of any of the transforms provided herein, inverse DCT, inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients to generate a residual video block in the pixel domain. The motion compensation prediction unit 362 may generate a motion compensation block and may perform interpolation based on an interpolation filter. The identifier of the interpolation filter for motion estimation with sub-pixel precision may be included in the syntax elements of the video block. The motion compensation prediction unit 362 may use an interpolation filter, such as that used by the video encoders 133, 200 during encoding of the video block, to calculate the interpolated values ​​of the sub-integer pixels of the reference block. The motion compensation prediction unit 362 may determine the interpolation filter used by the video encoders 133, 200 based on the received syntax information and use the interpolation filter to generate the prediction block.

[0060] The motion compensation prediction unit 262 may use: (1) the syntax information to determine the size of the video block used to encode one or more frames of the encoded video sequence; (2) partitioning information describing how to partition each video block of the frame of the encoded video sequence; (3) a mode (or mode information) indicating how to encode each partition; (4) one or more reference frames for encoding the video block between each frame; and / or (5) other information for decoding the encoded video sequence.

[0061] Adder 326 can sum the residual block with the corresponding predicted block generated by motion compensation prediction unit 362 or spatial prediction unit 360 to form a decoded video block. A loop filter 366 (e.g., a deblocking filter or SAO filter) can be applied to filter the decoded video block to remove block artifacts and / or improve visual quality. The decoded video block can be stored in reference image storage 364, which provides a reference video block for subsequent motion compensation and can generate decoded video for presentation on a display device (not shown). Point cloud compression.

[0062] Figure 4 illustrates the structure of the bitstream used for video-based point cloud compression (V-PCC). The resulting video bitstream and relay data are multiplexed together to produce the final V-PCC bitstream.

[0063] A V-PCC bitstream consists of a set of V-PCC cells as shown in Figure 4. Table 1 shows the syntax of V-PCC cells defined in the latest version of the V-PCC standard community draft (V-PCC CD), where each V-PCC cell has a V-PCC cell header and a V-PCC cell payload. The V-PCC cell header describes the V-PCC cell type (Table 2). V-PCC cells with cell types 2, 3, and 4 are occupancy, geometry, and attribute data cells as defined in the community draft. These data cells represent the three main components required to reconstruct the point cloud. In addition to the V-PCC cell type, the V-PCC attribute cell header specifies the attribute type and its index, which allows for multiple instances of the same attribute type.

[0064] The payloads (Table 3) of the V-PCC cell for occupancy, geometry, and attributes correspond to video data cells (e.g., HEVC NAL (Network Abstraction Layer) cells), which can be decoded by the video decoder specified in the corresponding set of occupancy, geometry, and attribute parameters of the V-PCC cell. Table 1 V-PCC Cell Syntax vpcc_unit() { descriptor vpcc_unit_header() vpcc_unit_payload() } Table 2 V-PCC Unit Header Syntax vpcc_unit_header(){ descriptor vpcc_unit_type U(5) if( vpcc_unit_type == VPCC_AVD | | vpcc_unit_type == VPCC_GVD | | vpcc_unit_type == VPCC_OVD | | vpcc_unit_type == VPCC_PSD ) vpcc_sequence_parameter_set_id u(4) if( vpcc_unit_type = = VPCC_AVD ) { vpcc_attribute_index u(7) if( sps_multiple_layer_streams_present_flag ) { vpcc_layer_index u(4) pcm_separate_video_data( 11 ) } else pcm_separate_video_data( 15 ) } else if( vpcc_unit_type = = VPCC_GVD ) { if( sps_multiple_layer_streams_present_flag ) { vpcc_layer_index u(4) pcm_separate_video_data( 18 ) } else pcm_separate_video_data( 22 ) } else if( vpcc_unit_type = = VPCC_OVD | | vpcc_unit_type = = VPCC_PSD ) { vpcc_reserved_zero_23bits u(23) } else vpcc_reserved_zero_27bits u(27) } Table 3 V-PCC Unit Load Syntax vpcc_unit_payload() { descriptor if( vpcc_unit_type == VPCC_SPS ) sequence_parameter_set() else if( vpcc_unit_type == VPCC_PSD ) patch_sequence_data_unit() else if( vpcc_unit_type == VPCC_OVD | | vpcc_unit_type == VPCC_GVD | | vpcc_unit_type == VPCC_AVD) video_data_unit() } Through HTTP dynamic streaming (DASH)

[0065] MPEG Dynamic Adaptive Streaming (MPEG-DASH) via HTTP is a universal delivery format that provides end users with the best possible video experience by dynamically adapting to changing network conditions.

[0066] HTTP-compatible streaming, such as MPEG-DASH, requires various bitrate replacements of multimedia content available at the server. Furthermore, multimedia content can include several media components (e.g., audio, video, text), each of which can have different characteristics. In MPEG-DASH, these characteristics are described by a Media Presentation Description (MPD).

[0067] Figure 5 illustrates the MPD hierarchical data model. MPD describes a series of time periods in which the consistent set of encoded versions of media content components remains unchanged during the time period. Each time period has a start time and duration, and consists of one or more fitness sets (fitness sets).

[0068] An adaptation set represents a collection of encoded versions of one or more media content components that share the same properties, such as language, media type, aspect ratio, character, accessibility, and rating properties. For example, an adaptation set may contain different bitrates of video components of the same multimedia content. Another adaptation set may contain different bitrates of audio components of the same multimedia content (e.g., lower-quality immersive sound and higher-quality surround sound). Each adaptation set typically includes multiple representations.

[0069] A representation describes a deliverable coded version of one or more media components that differs from other representations in bit rate, resolution, number of channels, or other characteristics. Each representation consists of one or more segments. Attributes of representation elements, such as @id, @bandwidth, @quality ordering, and @dependencyId, are used to specify the properties of the associated representation. A representation may also include sub-representations as part of a representation to describe the representation and extract partial information from it. Sub-representations can provide the ability to access lower-quality versions of representations (in which they are contained).

[0070] A segment is the largest unit of data that can be retrieved with a single HTTP request. Each segment has a URL (i.e., an addressable location on a server) that can be downloaded using an HTTP GET or an HTTP GET with a byte range.

[0071] To use this data model, the DASH client parses the MPD XML file and selects a set of adaptation sets suitable for its environment based on the information provided in each adaptation set element. Within each adaptation set, the client typically selects a representation based on the value of the @bandwidth attribute, while also considering the client's decoding and rendering capabilities. The client downloads the initial segments of the selected representation and then accesses the content by requesting the entire segment or a byte range of a segment. Once rendering has begun, the client continues to consume media content by continuously requesting media segments or portions of media segments and playing the content according to the media rendering timeline. The client may switch representations based on updated information from its environment. The client should play content continuously across multiple time periods. Once the client is consuming media contained in a segment up to the end of the media announced in that representation, media rendering is terminated, a new time period begins, or the MPD needs to be retrieved again. Descriptors in DASH

[0072] MPEG-DASH introduces the concept of descriptors to provide application-specific information about media content. Descriptor elements are all structured in the same way, containing the `@schemaIdUri` attribute (providing a URI to identify the scheme), the optional `@value` attribute, and the optional `@id` attribute. The semantics of the element are specific to the scheme used. The URI identifying the scheme can be a URN (Universal Resource Name) or a URL (Universal Resource Locator). MPD does not provide any specific information on how to use these elements. This is determined by the application, which uses the DASH format to instantiate descriptor elements with appropriate scheme information. DASH applications using one of these elements must first define the scheme identifier in URI form, and then define the value space of the element when using the scheme identifier. If structured data is required, any extended elements or attributes can be defined in separate namespaces. Descriptors can appear at multiple levels within MPD: - The presence of an element at the MPD level means that the element is a child element of an MPD element. - The presence of an element at the fitness set level means that the element is a child element of a fitness set element. - The presence of an element at the hierarchy level means that the element is a child element of the element representing hierarchy. Preselection

[0073] In MPEG-DASH, a bundle is a set of media components that can be consumed jointly by a single decoder instance. Each bundle includes a main media component that contains decoder-specific information and bootstraps the decoder. Pre-selection defines a subset of the media components in the bundle that are expected to be consumed jointly.

[0074] The fitness set containing the main media component is called the main fitness set. The main media component is always included in any pre-selection associated with the bundle. In addition, each bundle may include one or more partial fitness sets. Partial fitness sets can only be processed in conjunction with the main fitness set.

[0075] Preselection can be defined using the preselection elements defined in Table 4. The selection of a preselection is based on the attributes and elements contained within the preselection elements. Table 4: Semantics of Preselection Elements Element or property name use describe Preliminary @id OD default=1 Specify the pre-selected ID. This will be unique within a given time period. @Pre-selected Components M Specify the included suitability set or the ID of the content component belonging to the preselection as a blank separator list in the processing order, where the first ID is the ID of the main media component. @language O The pre-selected language code is declared according to the syntax and semantics in IETF RFC5646. Accessibility 0 … N Specify information about the accessibility scheme. Role 0 … N Specify information regarding the character annotation scheme. Element or property name 0 … N Information on the specified rating scheme. Rating 0 … N Specify information regarding the rating scheme. viewpoint 0 … N Information regarding the viewpoint annotation scheme has been specified. Normal attribute elements - Specify common properties and elements (properties and elements from the basic type RepresentationBaseType). legend: For attributes: M = mandatory, O = optional, OD = optional with preset value, CM = conditionally mandatory. For elements: <minOccurs>..<maxOccurs> (N = unbounded) The element is bold; the attribute is non-bold and preceded by @. Adaptive streaming of point clouds

[0076] While traditional multimedia applications such as video remain popular, there is significant interest in new media such as VR and immersive 3D graphics. High-quality 3D point clouds have recently emerged as a high-level representation of immersive media, enabling new forms of interactive work and communication with the virtual world. The large amount of information required to represent such dynamic point clouds necessitates efficient coding algorithms. The MPEG 3DG Working Group is currently developing a standard for video-based point cloud compression, with a community draft (CD) version released at the MPEG #124 meeting. The latest version of the CD defines the bitstream for compressing dynamic point clouds. In parallel, MPEG is also developing a system standard for carrying point cloud data.

[0077] The aforementioned point cloud standard only addresses the issues of writing and storing point clouds. However, it is conceivable that practical point cloud applications will require streaming point cloud data over a network. Such applications can perform live or on-demand streaming of point cloud content depending on how the content is generated. Furthermore, due to the large amount of information required to represent point clouds, such applications need to support adaptive streaming technology to avoid overloading the network and to provide the best viewing experience for the network capacity at any given moment.

[0078] A strong candidate method for adaptive delivery of point clouds is Dynamic Adaptive Streaming (DASH) via HTTP. However, the current MPEG-DASH standard does not provide any communication mechanism for point cloud media, including point cloud streams based on the MPEG V-PCC standard. Therefore, it is important to define new communication elements that enable streaming clients to identify point cloud streams and their component substreams within a Media Presentation Descriptor (MPD) file. Additionally, it is necessary to signal the different types of relay data associated with the point cloud components so that streaming clients can select the best supported version (one or multiple) of the point cloud or its components.

[0079] Unlike traditional media content, V-PCC media content consists of multiple components, some of which have multiple layers. Each component (and / or layer) is encoded separately as a substream of the V-PCC bitstream. Some component substreams, such as geometry and occupancy maps (except for some attributes such as texture), are encoded using conventional video encoders (e.g., H.264 / AVC or HEVC). However, these substreams need to be decoded together with additional relay data in order to render the point cloud.

[0080] Defines a number of XML elements and attributes. These XML elements are defined in their respective namespaces "urn:mpeg:mpegI:vpcc:2019". The namespace identifier "vpcc:" is used to refer to that namespace in the file. The V-PCC component is notified by signaling in DASH MPD.

[0081] Each V-PCC component and / or component layer can be represented in a DASH manifest (MPD) file as a separate suitability set (hereinafter referred to as the "component suitability set"), which has an additional suitability set attached to the primary access point (hereinafter referred to as the "primary suitability set") used for V-PCC content. In another embodiment, one suitability set is signaled per resolution per component.

[0082] In one embodiment, the fitness set of a V-PCC stream, including all V-PCC component fitness sets, should have a value of the @codec attribute set to 'vpc1' (e.g., as defined for V-PCC), indicating that the MPD is attached to the cloud point. In another embodiment, only the primary fitness set has the @codec attribute set to 'vpc1', while the @codec attribute of the fitness set of the point cloud components (or, if not signaled to the @codec for fitness set elements) is set based on the separate codecs used to encode the components. In the case of video write-coded components, the value of @codec should be set to 'resv.pccv.XXXX', where XXXX corresponds to the four-character code (4CC) of the video codec (e.g., avc1 or hvc1).

[0083] To identify the type (e.g., occupancy map, geometry, or attribute) of a V-PCC component (one or multiple) in the component fitness set, an EssentialProperty descriptor and a @SchemeIdUri attribute equal to "urn:mpeg:mpegI:vpcc:2019:component" can be used. This descriptor is called the VPCC component descriptor.

[0084] At the fitness set level, a VPCC component descriptor can be signaled for each point cloud component present in the representation of the fitness set.

[0085] In one embodiment, the @ value attribute of the VPCC component descriptor should not exist. The VPCC component descriptor may include elements and attributes as specified in Table 5. Table 5 Elements and Attributes of VPCC Component Descriptors Elements and attributes used for VPCC component descriptors use Data type describe Quantity 0. . N vpcc: vpcc component type An element whose attributes specify information about one of the point cloud components present in the representation (one or multiple) of the adaptive set. Components @ Component Types M xs: string Indicates the type of point cloud components. The value 'geom' indicates a geometric structure component, 'occp' indicates an occupancy component, and 'attr' indicates an attribute component. Component @ Minimum Level Index O xs: integer Indicates the index of the first level of the component represented by the fitness set containing the VPCC component descriptor. If only one level exists in the representation of the fitness set, the minimum level index and the maximum level index should have the same value. Component @ Maximum Level Index CM xs: integer Indicates the index of the last level of the component represented by the fitness set of the VPCC component descriptors. It will only exist if the minimum level exists. If there is only one layer in the representation of the adaptive set, then the minimum layer index and the maximum layer index will have the same value. Component @ Attribute Type CM xs: unsigned byte Indicates the type of attribute (see Table 7.2 in V-PCC CD). Only values ​​between 0 and 15 are allowed (including end values). It only exists if the component is a point cloud attribute (i.e., the component type has the value 'attr'). Component @ Attribute Index CM xs: unsigned byte Indicates the index of the attribute. It should be a value between 0 and 127, inclusive. It only exists if the component is a point cloud attribute (i.e., the component type has the value "attr"). legend: For attributes: M = mandatory, O = optional, OD = optional with default value, CM = conditionally mandatory. For elements: <minOccurs>..<maxOccurs> (N = unbounded) The element is bold; the attribute is non-bold and preceded by @.

[0086] The data types of the various elements and attributes of the VPCC component descriptor can be as defined in the following XML outline. <?xml version="1.0" encoding="UTF-8"?> <xs:schema xmlns:xs="http: / / www.w3.org / 2001 / XMLSchema" targetNamespace="urn:mpeg:mpegI:vpcc:2019" xmlns:omaf="urn:mpeg:mpegI:vpcc:2019" elementFormDefault="qualified"> <xs:element name="component" type="vpcc:vpccComponentType" / > <xs:complexType name="vpccComponentType"> <xs:attribute name="component_type" type="xs:integer" use="required" / > <xs:attribute name="min_layer_index" type="xs:integer" use="optional" / > ​​<xs:attribute name="max_layer_index" type="xs:integer" / > <xs:attribute name="attribute_type" type="xs:unsignedByte" / > <xs:attribute name="attribute_index" type="xs:unsignedByte" / > < / xs:complexType> < / xs:schema>

[0087] In one embodiment, the primary adaptability set should be contained in a single initialization segment at the adaptability set level or multiple initialization segments at the representation level (one for each representation). In one embodiment, the initialization segment should contain a V-PCC sequence parameter set as defined in the community draft, which is used to initialize the V-PCC decoder. In the case of a single initialization segment, the V-PCC sequence parameter sets of all representations can be contained in the initialization segment. When more than one representation is signaled in the primary adaptability set, the initialization segment for each representation can contain the V-PCC sequence parameter set of that particular representation. When the ISO Basic Media File Format (ISOBMFF) is used as a media container for V-PCC content as defined in WD of ISO / IEC 23090-10, the initialization segment may also include a metabox as defined in ISO / IEC 14496-12. The metabox contains one or more instances of VPCC GroupBox (as defined in VPCC CD), which provide relay information about the tracks and the relationships between them at the file format level.

[0088] In one embodiment, the media segments used for the representation of the primary fitness set include one or more track segments of V-PCC tracks defined in the community draft. The media segments used for the representation of the secondary fitness set include one or more track segments of corresponding secondary tracks at the archive format level.

[0089] In another embodiment, an additional attribute, referred to herein as the @video codec attribute, is defined for the VPCC component descriptor, the value of which indicates the codec used to encode the corresponding point cloud component. This enables support for scenarios where more than one point cloud component exists in the suitability set or representation.

[0090] In another embodiment, the role descriptor element may be used with newly defined values ​​for V-PCC components to indicate the role of the corresponding fitness set or representation (e.g., geometry, occupancy map, or attribute). For example, geometry, occupancy map, and attribute components may have corresponding values: vpcc-geometry, vpcc-occupancy, and vpcc-attribute, respectively. Additional key property descriptor elements, similar to those described in Table 5 minus the component type attribute, can be signaled at the fitness set level to identify the layer and attribute type of the component (if the component is a point cloud attribute). V-PCC fitness sets are grouped.

[0091] The streaming client can identify the type of point cloud components in the fitness set or representation by examining the VPCC component descriptor within the corresponding element. However, the streaming client also needs to distinguish between different point cloud streams existing in the MPD file and identify their respective component streams.

[0092] A key property element having the attribute @SchemeIdUri equal to "urn:mpeg:mpegI:vpcc:2019:vpc" can be introduced and is referred to herein as a VPCC descriptor. At most one VPCC descriptor can exist at the fitness set level for the main fitness set of a point cloud. If more than one representation exists in the main fitness set, at most one VPCC descriptor can exist at the representation level (i.e., within each representation element). Table 6 shows the properties of a VPCC descriptor according to one embodiment. Table 6 Properties of VPCC Descriptors Attributes of VPCC descriptor use Data type describe vpcc:@pcId CM xs: string The ID used for the point cloud. This attribute should exist if multiple versions of the same point cloud are signaled in separate fitness sets. vpcc:@occupancy rate Id M String vector type The suitability set or representation ID of the point cloud occupancy map component. vpcc: Geometry ID M String vector type A list of spatial delimiters, corresponding to the values ​​of the @id attribute used for the fitness sets and / or representations of point cloud geometry components. vpcc: Attribute ID M String vector type A list of spatial separators, corresponding to the suitability set used for point cloud attribute components and / or the value of the @id attribute.

[0093] When more than one version of a point cloud is available (e.g., different resolutions), each version may exist in a separate component fitness set containing a single representation and a VPCC descriptor with the same value for the @pcId attribute. In another embodiment, different versions of the point cloud may be signaled as representations of a single (primary) fitness set. In this case, a VPCC descriptor will exist in each representation, and the @pcId attribute may be signaled with the same value for all representations in the primary fitness set, or it may be omitted.

[0094] In another embodiment, preselection is signaled in the MPD using the value of the @preselected component attribute, which includes the id of the primary fitness set for the point cloud (followed by the id of the component fitness set corresponding to the point cloud component). The preselected @codec attribute should be set to 'vpc1', indicating that the preselected media is a video-based point cloud. Preselection can be signaled using a preselection element within a time period element or a preselection descriptor at the fitness set level (or the representation level when multiple versions / representations are available for the same point cloud). When a preselection element is used and more than one version of the same point cloud is available, each version is signaled in the individual preselection element using the first id in the list of ids of the @preselection element attribute, where the first id is the representation id of the corresponding point cloud version in the primary fitness set. Figure 6 illustrates an exemplary DASH configuration for grouping V-PCC components of a single point cloud belonging to an MPEG-DASH MPD archive.

[0095] Using a preselected descriptor, the group / association can be signaled as follows: <Period> <AdaptationSet id="5" codecs="vpc1"> <SupplementalProperty schemeIdUri="urn:mpeg:dash:preselection:2016" value="Presel1,5 1 2 3 4" / > <Representation> ... < / Representation> < / AdaptationSet> < / Period>

[0096] In another embodiment, the primary fitness set or its representation (one or multiple) of the point cloud can be listed using the @AssociationId attribute, defined in ISO / IEC 23009-1, which lists the fitness set and / or representation identifier of the component with a @AssociationType value set to V-PCC (i.e., 'vpc1').

[0097] In another embodiment, the primary adaptability set or its representation (one or multiple) of the point cloud can be listed using the @dependencyId attribute defined in ISO / IEC 23009-1 to identify the adaptability set and / or representation of the component. This is because there is an inherent dependency, as segments in the primary adaptability set need to be decoded together with segments from the component adaptability sets of the point cloud components in order to reconstruct the point cloud. Component Relay Data

[0098] Geometric structure and attribute relay data are typically used for rendering. They are signaled in the parameter set of the V-PCC bitstream. However, it may be necessary to signal these relay data elements in the MPD so that the streaming client can obtain the information as early as possible. Additionally, the streaming client can choose between multiple versions of the point cloud with different geometric structure and attribute relay data values ​​(e.g., based on whether the client supports signaled values). Signaling geometric structure relay data.

[0099] A supplementary property element with the @schemaIdURri attribute equal to "urn:mpeg:mpegI:vpcc:2019:geom_meta" may be introduced and is referred to herein as a geometry relay data descriptor or geoMeta descriptor. At most, a geometa descriptor may exist at the MPD level, in which case it applies to all geometric components of the point cloud notified by signaling in the MPD, unless it is overridden by a lower-level geoMeta descriptor as described below. At most, a geometa descriptor may exist at the fitness set level of the master fitness set. At most, a geometa descriptor may exist at the representation level of the master fitness set. If a geometa descriptor exists at a certain level, it overrides any geometa descriptor notified by signaling at a higher level.

[0100] In one embodiment, the @ value attribute of the geomMeta descriptor will not exist. In one embodiment, the geomMeta descriptor includes the elements and attributes specified in Table 7. Table 7 Elements and attributes used for the geomMeta descriptor Elements and attributes used for geomMeta descriptors use Data type describe geom 0. .1 vpcc: Geometric Structure Relay Data Type A container element whose attributes and elements specify the geometric structure and relay information. geom@dotshape O xs: unsigned byte Indicates the shape of the geometric points used for rendering. Supported values ​​are in the range of 0 to 15 (inclusive). The corresponding shapes are obtained from Table 7-2 in the community draft. If not present, the default value should be 0. geom@dot size O xs: unsigned byte Indicates the size of the geometry points used for rendering. Supported values ​​are in the range of 1 to 65535 (inclusive). If not specified, the default value should be 1. geom.geomSmoothing 0. .1 vpcc: Geometry smoothing type Its properties provide information on the smoothness of geometry. geom. geom Smoothing@ Grid Size M xs: unsigned byte Specifies the grid size used for geometry smoothing. Allowed values ​​should be in the range of 2 to 128 (inclusive). If the geom.geomSmoothing element does not exist, the default grid size should be inferred to be 8. geom. geom Smoothing@critical value M xs: unsigned byte Smoothing threshold. If the geom.geom Smoothing element does not exist, the default threshold should be inferred to be 64. geom.geom scaling (scale) 0. .1 vpcc: Geometry scaling type The attribute provides information about the geometry scaling of the element. geom.geom scaling@x M xs: unsigned integer The scaling value along the X-axis. If the geom.geom Smoothing element does not exist, the default value should be inferred to be 1. geom.geom scaling@y M xs: unsigned integer The scaling value along the Y-axis. If the geom.geom Smoothing element does not exist, the default value should be inferred to be 1. geom.geom scaling@z M xs: unsigned integer The scaling value along the Z-axis. If the geom.geomSmoothing element does not exist, the default value should be inferred to be 1. geom.geom offset 0. .1 vpcc:geom.geom offset An element whose properties provide information about geometric offsets. geom.geom offset@x M xs: integer Offset along the X-axis. If the geom.geomSmoothing element does not exist, the default value should be inferred to be 0. geom.geom offset@y M xs: integer Offset along the Y-axis. If the geom.geomSmoothing element does not exist, the default value should be inferred to be 0. geom.geom offset@ z M xs: integer Offset along the Z-axis. If the geom.geomSmoothing element does not exist, the default value should be inferred to be 0. geom.geom rotation 0. .1 vpcc:geom.geom rotation type The attribute provides information about the rotation of the geometry of the element. geom.geom rotates @ x M xs: integer Rotation value along the X-axis by 2 -16 The unit is degrees. If the geom.geomSmoothing element does not exist, the default value should be inferred to be 0. geom.geom rotates @ y M xs: integer Rotation value along the Y-axis by 2 -16 The unit is degrees. If the geom.geomSmoothing element does not exist, the default value should be inferred to be 0. geom.geom rotates @ z M xs: integer Rotation value along the Z-axis by 2 -16 The unit is degrees. If the geom.geomSmoothing element does not exist, the default value should be inferred to be 0. legend: For attributes: M = mandatory, O = optional, OD = optional with default value, CM = conditionally mandatory. For elements: <minOccurs>..<maxOccurs> (N = unbounded) The element is bold; the attribute is non-bold and preceded by @.

[0101] In one embodiment, the data types of the various elements and attributes of the geomMeta descriptor can be as defined in the following XML schema. <?xml version="1.0" encoding="UTF-8"?> <xs:schema xmlns:xs="http: / / www.w3.org / 2001 / XMLSchema" targetNamespace="urn:mpeg:mpegI:vpcc:2019" xmlns:omaf="urn:mpeg:mpegI:vpcc:2019" elementFormDefault="qualified"> <xs:element name="geom" type="vpcc:geometryMetadataType" / > <xs:complexType name="geometryMetadataType"> <xs:attribute name="point_shape" type="xs:unsignedShort" use="optional" default="0" / > <xs:attribute name="point_size" type="xs:unsignedByte" use="optional" default="1" / > <xs:element name="geomSmoothing" type="vpcc:geometrySmoothingType" minOccurs="0" maxOccurs="1" / > <xs:element name="geomScale" type="vpcc:geometryScaleType" minOccurs="0" maxOccurs="1" / > <xs:element name="geomOffset" type="vpcc:geometryOffsetType" minOccurs="0" maxOccurs="1" / > <xs:element name="geomRotation" type="vpcc:geometryRotationType" minOccurs="0" maxOccurs="1" / > < / xs:complexType> <xs:complexType name="geometrySmoothingType"> <xs:attribute<xs:attribute name="grid_size" type="xs:unsignedByte" use="required" / > <xs:attribute name="threshold" type="xs:unsignedByte" use="required" / > < / xs:complexType> <xs:complexType name="geometryScaleType"> <xs:attribute name="x" type="xs:unsignedInt" use="required" / > <xs:attribute name="y" type="xs:unsignedInt" use="required" / > <xs:attribute name="z" type="xs:unsignedInt" use="required" / > < / xs:complexType> <xs:complexType name="geometryOffsetType"> <xs:attribute name="x" type="xs:int" use="required" / > <xs:attribute name="y" type="xs:int" use="required" / > <xs:attribute name="z" type="xs:int" use="required" / > < / xs:complexType> <xs:complexType name="geometryRotationType"> <xs:attribute name="x" type="xs:int" use="required" / > <xs:attribute name="y" type="xs:int" use="required" / > <xs:attribute name="z" type="xs:int" use="required" / > < / xs:complexType> < / xs:schema> Signal attribute metadata

[0102] A supplementary property element with the @schemaIdUri attribute equal to "urn:mpeg:mpegI:vpcc:2019:attr_meta" may be introduced, and is referred to herein as an attribute relay data descriptor or attrMeta descriptor. At most one attrMeta descriptor may exist at the primary fitness set level. At most one attrMeta descriptor may exist at the representation level within the primary fitness set. If an attrMeta descriptor exists at the representation level, it replaces any attrMeta descriptor signaled at the fitness set level that represents the fitness set to which it belongs.

[0103] In one embodiment, there is no @ value attribute for the attrMeta descriptor. In one embodiment, the attrMeta descriptor may include elements and attributes as specified in Table 8. Table 8 Elements and attributes for the attrMeta descriptor Elements and attributes used in attrMeta descriptors use Data type describe attm 0. .N vpcc: Attribute Relay Data Type A container element whose attributes and element-specified point cloud attributes are relay data information. attm@index M xs: unsigned byte Indicates the index of the attribute. It should be a value between 0 and 127 (inclusive). attm@Dimension M xs: unsigned byte The number of dimensions of point cloud attributes. attm.attrSmoothing 0. .1 vpcc: Attribute smoothing type The attribute provides smooth information for point cloud properties. attm.attrSmoothing@radius M xs: unsigned byte Detects the radius of neighboring elements for attribute smoothing. If the attm.attrSmoothing element does not exist, the default value will be inferred as 0. attm.attrSmoothing@Neighbor Counting M xs: unsigned byte The maximum number of adjacent points used for attribute smoothing. If the attm.attrSmoothing element does not exist, the default value should be inferred to be 0. attm.attrSmoothing@radius2bound M xs: unsigned byte The radius used for boundary point detection. If the attm.attrSmoothing element does not exist, the default value will be inferred as 0. attm.attrSmoothing@critical value M xs: unsigned byte Attribute smoothing threshold. If the unsigned tuple element xs: does not exist, the default value will be inferred as 0. attm.attrSmoothing@ Critical Local Entropy M xs: unsigned byte The local entropy threshold in the neighborhood of the boundary point. The value of this attribute should be in the range of 0 to 7 (inclusive). If the `attm.attrSmoothing` element does not exist, the default value will be inferred as 0. attm.attrScale 0. .1 vpcc: Attribute scaling type The attribute provides scaling information along each dimension of the point cloud attribute. attm.attrScale@value M xs: string A comma-separated string of scaling values ​​for each dimension of the point cloud attributes. attm.attrScaleOffset 0. .1 vpcc: Attribute offset type Its properties provide offset information along each dimension of the point cloud properties. attm.attrScale@value M xs: string A comma-separated string of offset values ​​for each dimension of the point cloud attributes. legend: For attributes: M = mandatory, O = optional, OD = optional with preset value, CM = conditionally mandatory. For elements: <minOccurs>..<maxOccurs> (N = unbounded) The element is bold; the attribute is non-bold and preceded by @.

[0104] In one embodiment, the data types of the various elements and attributes of the attrMeta descriptor may be as defined in the following XML outline. <?xml version="1.0" encoding="UTF-8"?> <xs:schema xmlns:xs="http: / / www.w3.org / 2001 / XMLSchema" targetNamespace="urn:mpeg:mpegI:vpcc:2019" xmlns:omaf="urn:mpeg:mpegI:vpcc:2019" elementFormDefault="qualified"> <xs:element name="attm" type="vpcc:attributeMetadataType" / > <xs:complexType name="attributeMetadataType"> <xs:attribute name="index" type="xs:unsignedByte" use="required" / > <xs:attribute name="num_dimensions" type="xs:unsignedByte" use="required" / > <xs:element name="attrSmoothing" type="vpcc:attributeSmoothingType" minOccurs="0" maxOccurs="1" / > <xs:element name="attrScale" type="vpcc:attributeScaleType" minOccurs="0" maxOccurs="1" / > <xs:element name="attrOffset" type="vpcc:attributeOffsetType" minOccurs="0" maxOccurs="1" / > < / xs:complexType> <xs:complexType name="attributeSmoothingType"> <xs:attribute name="radius" type="xs:unsignedByte" use="required" / > <xs:attribute name="neighbour_count" type="xs:unsignedByte"use="required" / > <xs:attribute name="radius2_boundary" type="xs:unsignedByte" use="required" / > <xs:attribute name="threshold" type="xs:unsignedByte" use="required" / > <xs:attribute name="threshold_local_entropy" type="xs:unsignedByte" use="required" / > < / xs:complexType> <xs:complexType name="attributeScaleType"> <xs:attribute name="values" type="xs:string" use="required" / > < / xs:complexType> <xs:complexType name="attributeOffsetType"> <xs:attribute name="values" type="xs:string" use="required" / > < / xs:complexType> < / xs:schema> Streaming client behavior

[0105] The DASH client (decoder node) is guided by information provided in the MPD. The following is an example client behavior for processing streaming point cloud content according to the communications presented in this specification, assuming an implementation using a VPCC descriptor to signal the association between the component fitness set and the main point cloud fitness set. Figure 7 is a flowchart illustrating the example streaming client process according to the embodiment.

[0106] At 711, the client first sends an HTTP request and downloads the MPD file from the content server. The client then parses the MPD file to generate a memory representation of the corresponding XML elements in the MPD file.

[0107] Next, at 713, in order to identify the available point cloud media content for a given time period, the streaming client scans the fitness set elements to find the fitness set with the @ encoder / decoder property set to 'vpc1' and the VPCC descriptor element. The resulting subset is a set of primary fitness sets used for the point cloud content.

[0108] Next, at 715, the streaming client identifies the number of unique point clouds by examining the VPCC descriptors of those fitness sets, and groups fitness sets with the same @pcId value in their VPCC descriptors into versions with the same content.

[0109] At 717, identify a set of fitness sets having @pcId values ​​corresponding to the point cloud content that the user expects to stream. If the set contains more than one fitness set, the streaming client selects the fitness set with the supported version (e.g., video resolution). Otherwise, select the only fitness set in the set.

[0110] Next, at 719, the streaming client checks the VPCC descriptors of the selected suitability set to identify the suitability set of the point cloud components. These are identified from the values ​​of the @OccupancyId, @GeometryId, and @AttributeId attributes. If the geomMeta and / or attrMeta descriptors exist in the selected primary suitability set, the streaming client can identify whether it supports the rendering configuration for signal notifications for point cloud streaming before downloading any segment. Otherwise, the client needs to retrieve this information from the initialization segment.

[0111] Next, at 721, the client starts streaming the point cloud by downloading the initialization segment for the primary adaptability set, which contains the parameter set required to initialize the V-PCC decoder.

[0112] At 723, the initialization segment for the video write code component stream is downloaded and cached in memory.

[0113] At 725, the streaming client then begins to download time-aligned media segments in parallel from the primary and secondary fitness sets via HTTP, and the downloaded segments are stored in the in-memory segment buffer.

[0114] At 727, the time-aligned media segment is removed from its respective buffer and concatenated with the respective initialized segment sequence.

[0115] Finally, at 729, the media container (e.g., ISOBMFF) is parsed to extract basic stream information and the V-PCC bitstream is structured according to the V-PCC standard, and then the bitstream is passed to the V-PCC decoder.

[0116] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in a computer program, software, or firmware incorporated into a computer-readable medium for execution by a computer or processor. Examples of non-transitory computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), temporary registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable discs, magneto-optical media, and optical media such as CD-ROM discs and digital multifunction discs (DVDs). The processor associated with the software can be used to implement a radio frequency transceiver used in a WTRU 102, UE, terminal, base station, RNC, or any host computer.

[0117] Furthermore, in the above embodiments, note the processing platform, computing system, controller, and other devices including a processor. These devices may include at least one central processing unit (“CPU”) and memory. According to the practice of those skilled in the art of computer programming, references to symbolic representations of actions and operations or instructions can be performed by various CPUs and memory. Such actions and operations or instructions may be referred to as “executed,” “computer-executed,” or “CPU-executed.”

[0118] Those skilled in the art will understand that the actions and symbols representing operations or instructions include manipulation of electrical signals by the CPU. Electrical systems represent data bits, which can cause the transformation or restoration of electrical signals and the maintenance of data bits at their storage locations in the memory system, thereby reconfiguring or otherwise altering the operation of the CPU and other signal processing. The memory location maintaining the data bits is a physical location having specific electrical, magnetic, optical, or organic properties corresponding to or representing the data bits. It should be understood that exemplary embodiments are not limited to the platforms or CPUs described above, and other platforms and CPUs may support the provided methods.

[0119] Data bits may also be stored on a computer-readable medium, including magnetic disks, optical disks, and any other CPU-readable volatile (e.g., random access memory (“RAM”)) or non-volatile (e.g., read-only memory (“ROM”)) mass storage system. The computer-readable medium may include cooperative or interconnected computer-readable media that reside exclusively on a processing system or are distributed among multiple interconnected processing systems that may be located locally or remotely on the processing system. It should be understood that representative embodiments are not limited to the memory described above, and other platforms and memories may support the described methods.

[0120] In illustrative embodiments, any operations, processes, etc., described herein may be implemented as computer-readable instructions stored on a computer-readable medium. These computer-readable instructions may be executed by a processor of an action unit, network element, and / or any other computing device.

[0121] There is little difference between the hardware and software implementations of various aspects of the system. The use of hardware or software is often (but not always, as the choice between hardware and software may become important in some cases) a design choice representing a trade-off between cost and efficiency. Various vehicles can exist by which the processes and / or systems and / or other technologies (e.g., hardware, software, and / or firmware) described herein can be influenced, and the preferred vehicle can vary depending on the context of the deployment process and / or system and / or other technologies. For example, if the implementer determines that speed and accuracy are of paramount importance, the implementer may choose a vehicle that is primarily hardware and / or firmware. If flexibility is of paramount importance, the implementer may choose a primarily software implementation. Alternatively, the implementer may choose some combination of hardware, software, and / or firmware.

[0122] The foregoing detailed description has illustrated various embodiments of the apparatus and / or process using block diagrams, flowcharts, and / or examples. Where such block diagrams, flowcharts, and / or examples contain one or more functions and / or operations, those skilled in the art will understand that each function and / or operation within such block diagrams, flowcharts, or examples can be implemented individually and / or collectively by a wide variety of hardware, software, firmware, or virtually any combination thereof. For example, suitable processors include general-purpose processors, special-purpose processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application-specific integrated circuits (ASICs), application-specific standard products (ASSPs); field-programmable gate array (FPGA) circuits, any other type of integrated circuit (IC) and / or state machine.

[0123] Although features and elements have been provided above in specific combinations, those skilled in the art will understand that each feature or element may be used alone or in any combination with other features and elements. This disclosure should not be limited to the specific embodiments described in this application, which are intended to illustrate various aspects. Many modifications and variations can be made without departing from the spirit and scope of the invention, as will be apparent to those skilled in the art. Unless expressly provided so, elements, actions, or instructions used in the description of this application should not be construed as critical or necessary to the invention. In addition to those listed herein, functionally equivalent methods and apparatus within the scope of this disclosure will be apparent to those skilled in the art based on the foregoing description. These modifications and variations are intended to fall within the scope of the appended claims. This disclosure is limited only by the terminology of the appended claims and the full scope of the equivalents granted by those claims. It should be understood that this disclosure is not limited to specific methods or systems.

[0124] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, when the terms “base station” and its abbreviation “STA” and “user equipment” and its abbreviation “UE” are referred to herein, they may mean (i) a wireless transmitting and / or receiving unit (WTRU), such as those described below; (ii) any of a plurality of embodiments of a WTRU, such as those described below; (iii) a device having wireless and / or wired capabilities (e.g., wired) configured with some or all of the structures and functions of a particular WTRU, such as those described below; (iii) a device having wireless and / or wired capabilities configured to have fewer of the structures and functions of a WTRU, such as those described below; or (iv) the like.

[0125] In some representative embodiments, certain portions of the subject matter described herein may be implemented by application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and / or other integrated formats. However, those skilled in the art will recognize that some aspects of the embodiments disclosed herein may be implemented, in whole or in part, equivalently in integrated circuits as one or more computer programs (e.g., one or more programs running on one or more computer systems), one or more programs running on one or more processors (e.g., one or more programs running on one or more microprocessors), firmware, or virtually any combination thereof, and that it is appropriate for those skilled in the art to design circuits and / or write code for software and / or firmware within the scope of their skills based on this disclosure. Furthermore, those skilled in the art will understand that the mechanisms of the subject matter described herein can be distributed as various forms of program products, and that the illustrative embodiments of the subject matter described herein apply regardless of the particular type of signal-bearing medium used for the actual performance of the distribution. Examples of signal-carrying media include, but are not limited to, the following: recordable media (e.g., floppy disks, hard disk drives, CDs, DVDs, digital magnetic tapes, computer memory, etc.), and transmission media (e.g., digital and / or analog communication media (e.g., fiber optic cables, waveguides, wired communication links, wireless communication links, etc.)).

[0126] The topics described herein sometimes illustrate different components contained within or connected to different other components. It should be understood that the architectures described in this way are merely examples, and many other architectures that can implement the same functionality can actually be implemented. Conceptually, any arrangement of elements that implement the same functionality is effectively “associated” to implement the desired functionality. Therefore, any two components combined in this document to implement a particular functionality can be considered “associated” with each other to implement the desired functionality, regardless of the architecture or intermediate components. Similarly, any two components that are so associated can also be considered “operably connected” or “operably coupled” to each other to implement the desired functionality, and any two components that can be so associated can also be considered “operably coupled” to each other to implement the desired functionality. Specific examples of operably coupled components include, but are not limited to, physically pairable and / or physically interactive components and / or wirelessly interactive and / or logically interactive and / or logically interactive components.

[0127] Regarding the use of virtually any plural and / or singular terms in this document, those skilled in the art may convert plural to singular and / or singular to plural as needed by the context and / or application. For clarity, various singular / plural substitutions may be explicitly described herein.

[0128] Those skilled in the art will understand that, in general, the terms used herein and especially in the appended claims (e.g., the body of the appended claims) are generally intended to be “open-ended” terms (e.g., the term “comprising” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “having at least,” the term “including” should be interpreted as “including but not limited to,” etc.). Those skilled in the art will also understand that if the intent is a specific number of features described in the introduced claims, such intent will be explicitly stated in the claims, and without such a statement, such intent does not exist. For example, the term “single” or similar language may be used where only one feature is desired. To aid understanding, the appended claims and / or the description herein may contain the use of introductory phrases such as “at least one” and “one or a plurality” to introduce claims. However, the use of such phrases should not be construed as implying that a patent application scope statement introduced by the indefinite article “a” or “one” limits any particular patent application scope containing such an introduced patent application scope statement to containing only one embodiment of such a statement, even when the same patent application scope includes the introductory phrases “one or plural” or “at least one” and indefinite articles such as “a” or “one” (e.g., “a” and / or “one” should be interpreted as meaning “at least one” or “one or plural”). The same applies to the use of definite articles to introduce patent application scope statements. Furthermore, even if a specific number of introduced patent application scope statements is explicitly stated, those skilled in the art will recognize that such a statement should be interpreted as meaning at least the number stated (e.g., in the absence of other modifiers, the unmodified statement of “two statements” means at least two statements, or two or more statements). Furthermore, in instances where the convention of “at least one of A, B, and C” is used, such a construction is generally intended to be understood by a person skilled in the art in the sense of the convention (e.g., “a system having at least one of A, B, and C” will include, but is not limited to, systems having only A, only B, only C, A and B, A and C, B and C, and / or systems having A, B, and C). In instances where the convention of “at least one of A, B, or C” is used, such a construction is generally intended to be understood by a person skilled in the art in the sense of the convention (e.g., “a system having at least one of A, B, or C” will include, but is not limited to, systems having only A, only B, only C, A and B, A and C, B and C, and / or systems having A, B, and C). A person skilled in the art will also understand that any transitional conjunctions and / or phrases that actually present two or more alternative terms, whether in the specification, the claims, or the drawings, should be understood to imply the possibility of including one, any one, or both of these terms.For example, the phrase “A or B” will be understood to include the possibility of “A” or “B” or “A and B”. Furthermore, as used herein, the term “any” preceding a list of plural features and / or plural feature categories is intended to include “any,” “any combination,” “any plural,” and / or “any plural combination” of features and / or feature categories, individually or in combination with other features and / or feature categories. Additionally, as used herein, the term “set” or “group” is intended to include any number (including zero) of features. Furthermore, as used herein, the term “number” is intended to include any number (including zero).

[0129] Furthermore, where features or aspects of this disclosure are described in accordance with the Markush Group, those skilled in the art will recognize that this disclosure is also described in accordance with any individual member or subgroup of the Markush Group.

[0130] As those skilled in the art will understand, for any and all purposes, such as in providing a written description, all scopes disclosed herein also encompass any and all possible subscopes and combinations thereof. Any listed scope can be readily considered sufficiently descriptive and such that the same scope can be decomposed into at least two, three, four, five, ten, etc., equal parts. As a non-limiting example, each scope discussed herein can be readily decomposed into a lower third (1 / 3), a middle third (1.5 / 3), and an upper third (2 / 3), etc. Those skilled in the art will also understand that all language such as “up to,” “at least,” “greater than,” “less than,” etc., includes the listed numbers and refers to a scope that can subsequently be decomposed into subscopes as described above. Finally, as those skilled in the art will understand, a scope includes each individual member. Thus, for example, a group having 1-3 cells means a group having 1, 2, or 3 cells. Similarly, a group having 1-5 cells means a group having 1, 2, 3, 4, or 5 cells, and so on.

[0131] Furthermore, the scope of the patent application should not be construed as limited to the order or elements provided, unless stated as such an effect. Additionally, the use of the term “apparatus for…” in any patent application scope is intended to invoke 35 U.SC §(f) or the means-function terminology format, and no patent application scope without the term “apparatus for…” is thus excluded.

[0132] Although the invention has been described and illustrated herein with reference to specific embodiments, the invention is not intended to be limited to the details shown. Rather, various modifications to the details may be made within the equivalents of the claims and without departing from the invention.

[0133] Throughout the disclosure, those skilled in the art will understand that certain representative embodiments may be used alternatively or in combination with other representative embodiments.

[0134] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in a computer program, software, or firmware incorporated into a computer-readable medium for execution by a computer or processor. Examples of non-transitory computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), temporary registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable discs, magneto-optical media, and optical media such as CD-ROM discs and digital multifunction discs (DVDs). The processor associated with the software can be used to implement a radio frequency transceiver used in a WRTU, UE, terminal, base station, RNC, or any host computer.

[0135] Furthermore, in the above embodiments, note the processing platform, computing system, controller, and other devices including a processor. These devices may include at least one central processing unit (“CPU”) and memory. According to the practice of those skilled in the art of computer programming, references to symbolic representations of actions and operations or instructions can be performed by various CPUs and memory. Such actions and operations or instructions may be referred to as “executed,” “computer-executed,” or “CPU-executed.”

[0136] Those skilled in the art will understand that actions and symbols representing operations or instructions include manipulation of electrical signals by the CPU. Electrical systems represent data bits, which can lead to the transformation or restoration of electrical signals and the maintenance of data bits at their storage locations in the memory system, thereby reconfiguring or otherwise altering the operation of the CPU and other signal processing. The storage location maintaining the data bits is a physical location having specific electrical, magnetic, optical, or organic properties corresponding to or representing the data bits.

[0137] Data bits may also be stored on a computer-readable medium, including magnetic disks, optical disks, and any other CPU-readable volatile (e.g., random access memory (“RAM”)) or non-volatile (e.g., read-only memory (“ROM”)) mass storage system. The computer-readable medium may include cooperative or interconnected computer-readable media that reside exclusively on a processing system or are distributed among multiple interconnected processing systems that may be located locally or remotely on the processing system. It should be understood that representative embodiments are not limited to the memory described above, and other platforms and memories may support the described methods.

[0138] Unless explicitly described herein, elements, actions, or instructions used in the description of this application should not be construed as critical or necessary to the invention. Additionally, as used herein, the article “a” is intended to include one or more features. Where only one feature is indicated, the term “a” or similar language is used. Furthermore, as used herein, the term “any” preceding a list of multiple features and / or multiple feature categories is intended to include “any,” “any combination,” “any multiple,” and / or “any multiple combination” of features and / or feature categories, individually or in combination with other features and / or other feature categories. Furthermore, as used herein, the term “set (combination)” is intended to include any number of “sets (combinations)” of features. Furthermore, as used herein, the term “number” is intended to include any number (including zero).

[0139] Furthermore, the scope of the patent application should not be construed as limited to the described order or elements unless stated as such an effect. Additionally, the use of the term “apparatus” in any patent application scope is intended to invoke 35 USC §112(f), and any patent application scope without the word “apparatus” is not so intended.

[0140] For example, suitable processors include general-purpose processors, special-purpose processors, conventional processors, digital signal processors (DSPs), complex microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application-specific integrated circuits (ASICs), application-specific standard products (ASSPs); field-programmable gate array (FPGA) circuits, any other type of integrated circuit (IC) and / or state machine.

[0141] The processor associated with the software can be used to implement radio frequency transceivers in wireless transmit receiver units (WRTUs), user equipment (UEs), terminals, base stations, mobility management entities (MMEs) or evolved packet cores (EPCs), or any host computer. The WRTU can be used in conjunction with modules implemented in hardware and / or software, including software-defined radio (SDR) and other components such as cameras, video phone modules, videophones, speakerphones, vibration devices, speakers, microphones, TV transceivers, hands-free headsets, keypads, Bluetooth® modules, FM radio units, near field communication (NFC) modules, liquid crystal display (LCD) units, organic light-emitting diode (OLED) display units, digital music players, media players, video game console modules, internet browsers, and / or any wireless local area network (WLAN) or ultra-wideband (UWB) modules.

[0142] Although the present invention has been described in accordance with a communication system, it is contemplated that the system can be implemented in software on a microprocessor / general-purpose computer (not shown). In some embodiments, one or more of the functions of the various elements can be implemented in software that controls the general-purpose computer.

[0143] Furthermore, although the invention has been shown and described herein with reference to specific embodiments, the invention is not limited to the details shown. Rather, various modifications to the details may be made within the equivalents of the claims and without departing from the invention. [Simplified Explanation of the Diagram]

[0005] A more detailed understanding can be obtained from the following detailed description, given by way of example in conjunction with the accompanying drawings. As with the detailed description, the figures in these drawings are exemplary. Therefore, the drawings and detailed description should not be considered limiting, and other equivalent examples are possible and feasible. Furthermore, the same reference numerals in the figures denote the same elements, wherein: Figure 1A is a block diagram illustrating an exemplary video encoding and decoding system in which one or more embodiments may be performed and implemented; Figure 1B is a block diagram illustrating an exemplary video encoder unit for use with the video encoding and / or decoding system of Figure 1A; Figure 2 is a block diagram of a general block-based hybrid video encoding system; Figure 3 is a general block diagram of a block-based video decoder; Figure 4 illustrates the structure of a bitstream for video-based point cloud compression (V-PCC); Figure 5 illustrates an MPD hierarchical data model; Figure 6 illustrates an exemplary DASH configuration for grouping V-PCC components belonging to a single point cloud within an MPEG-DASH MPD archive; and Figure 7 is a flowchart illustrating an exemplary decoder process for streaming point cloud content according to an embodiment.

Claims

1. A decoding node, comprising: A processor is configured to: receive information associated with a streaming PC data (PCD) corresponding to a point cloud (PC) in a Dynamic Adaptive Streaming (DASH) Media Presentation Description (MPD) over Hyper-File Transfer Protocol (HTTP), wherein the DASH MPD indicates at least: a master adaptability set (AS) for the PC, wherein the master AS includes information indicating at least: a value for a codec attribute indicating that the master AS corresponds to video-based point cloud compression (V-PCC) data, and an initialization segment including at least one set of V-PCC sequence parameters representing the PC, and a plurality of component ASs, wherein each of the plurality of component ASs corresponds to one of the plurality of V-PCC components and includes information indicating at least: a V-PCC component descriptor identifying a type of the corresponding V-PCC component, and at least one property of the corresponding V-PCC component, wherein the type includes geometry, occupancy, or attribute; and receive the DASH MPD via the network.

2. The decoding node as described in claim 1, wherein the processor is further configured to parse the DASH MPD to generate a representation of the DASH MPD.

3. The decoding node as described in claim 1, wherein the processor is further configured to identify an available point cloud media content based on the value used for the codec attribute or one or more of the initialization segment.

4. The decoding node as described in request item 1, wherein the processor is further configured to determine a unique number of point clouds based on the V-PCC component descriptor.

5. The decoding node as described in claim 1, wherein the processor is further configured to stream the PC by downloading the initialization segment for the master AS.

6. The decoding node as described in claim 5, wherein the processor is further configured to: receive time-aligned segments from the main AS and the plurality of component ASs via HTTP; and store the time-aligned segments in a memory buffer.

7. The decoding node as described in claim 6, wherein the processor is further configured to sequentially connect the time-aligned segments with the respective initialization segments.

8. The decoding node as described in claim 7, wherein the processor is further configured to generate a V-PCC bitstream using the time-aligned segments.

9. The decoding node as described in claim 8, wherein the processor is further configured to decode the V-PCC bitstream.

10. The decoding node as described in claim 1, wherein an International Organization for Standardization (ISO) Basic Media File Format (ISOBMFF) is used as a media container for a V-PCC content, and wherein the initialization segment of the main AS includes a meta-frame containing one or more V-PCC frame instances that provide relay data information associated with a V-PCC trajectory.

11. The decoding node as claimed in claim 1, wherein the initialization segment of the master AS includes information indicating a single initialization segment at an adaptability level and a set of V-PCC sequence parameters for a plurality of representations of the master AS.

12. The decoding node as claimed in claim 1, wherein the master AS includes an initialization segment for each of a plurality of representations of the PC, and wherein each initialization segment corresponding to a representation of the PC includes a V-PCC sequence parameter set for that representation.

13. The decoding node as claimed in claim 1, wherein the V-PCC component descriptor includes information indicating a video codec attribute of a codec used to encode a corresponding PC component.

14. The decoding node as described in claim 1, wherein the main AS includes information indicating a role descriptor DASH element having a value indicating a geometry, occupancy map, or attribute of a corresponding V-PCC component.

15. The decoding node as described in request item 1, wherein the decoding node is a DASH streaming user terminal.

16. A method for a decoding node, the decoding node using a HyperFile Transfer Protocol (HTTP) to stream point cloud data (PCD) corresponding to a point cloud (PC) via a network, and including a plurality of video-based point cloud compression (V-PCC) components comprising a PC, the method comprising: Receive information associated with a streaming PC data (PCD) corresponding to a point cloud (PC) in a Dynamic Adaptive Streaming (DASH) Media Presentation Description (MPD) over Hyper-File Transfer Protocol (HTTP), wherein the DASH MPD indicates at least: a master adaptability set (AS) for the PC, wherein the master AS includes information indicating at least: a value for a codec attribute indicating that the master AS corresponds to video-based point cloud compression (V-PCC) data, and an initialization segment including at least one set of V-PCC sequence parameters representing the PC, and a plurality of component ASs, wherein each of the plurality of component ASs corresponds to one of the plurality of V-PCC components and includes information indicating at least: a V-PCC component descriptor identifying a type of the corresponding V-PCC component, and at least one property of the corresponding V-PCC component, wherein the type includes geometry, occupancy, or attribute; and receive the DASH MPD via the network.

17. The method of claim 16 further includes parsing the DASH MPD to generate a representation of the DASH MPD.

18. The method of claim 16 further includes identifying an available point cloud media content based on the value used for the codec attribute or one or more of the initialization segment.

19. The method of claim 16 further includes determining a unique number of point clouds based on the V-PCC component descriptor.

20. The method as described in claim 16 further includes streaming the PC by downloading the initialization segment for the master AS.

21. The method as described in claim 20, further comprising: Receive time-aligned segments from the main AS and the plurality of component ASs via HTTP; And the time-aligned segments are stored in a memory buffer.

22. The method as described in claim 21 further includes sequentially connecting the time-aligned segments and the respective initialization segments.

23. The method of claim 22 further includes using the time-aligned segments to generate a V-PCC bitstream.

24. The method as described in claim 23 further includes decoding the V-PCC bitstream.

25. The method of claim 16, wherein an International Organization for Standardization (ISO) Basic Media File Format (ISOBMFF) is used as a media container for a V-PCC content, and wherein the initialization segment of the master AS includes a meta-frame containing one or more V-PCC frame instances that provide relay data information associated with a V-PCC track.

26. The method of claim 16, wherein the initialization segment of the master AS includes information indicating a single initialization segment at an adaptability level and a set of V-PCC sequence parameters for a plurality of representations of the master AS.

27. The method of claim 16, wherein the master AS includes an initialization segment for each of a plurality of representations of the PC, and wherein each initialization segment corresponding to a representation of the PC includes a V-PCC sequence parameter set for that representation.

28. The method of claim 16, wherein the V-PCC component descriptor includes information indicating a video codec attribute of a codec used to encode a corresponding PC component.

29. The method of claim 16, wherein the main AS includes information indicating a role descriptor DASH element having a value indicating a geometry, occupancy map, or attribute of a corresponding V-PCC component.

Citation Information

Patent Citations

  • Just-in-time dereferencing of remote elements in dynamic adaptive streaming over hypertext transfer protocol

    WO2015009723A1

  • Processing continuous multi-period content

    WO2015148519A1