Signaling of block-based region-adaptive hierarchical transform (RAHT) in a point cloud codec
The block-based RAHT transform partitions the octree into subtrees for independent attribute encoding and decoding, addressing the complexity and latency issues in point cloud codecs, enabling faster and more efficient real-time processing.
Patent Information
- Application Number
- PCT/EP2025/050640
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-15
- Filing Date
- 2025-01-13
- Publication Date
- 2025-07-24
AI Technical Summary
Existing point cloud codecs face challenges in achieving low complexity, latency, and memory footprint, particularly in applications requiring real-time encoding and decoding, due to the sequential compression of geometry and attributes, with the attribute encoding process not being parallelizable.
Implementing a block-based region-adaptive hierarchical transform (RAHT) that partitions the octree into non-overlapping subtrees, allowing independent encoding and decoding of attributes within each subtree, and signaling this information in the bitstream, enabling parallel processing.
This approach reduces encoding complexity and latency while maintaining encoding efficiency, allowing for faster and more efficient real-time processing of point cloud attributes.
Smart Images

Figure EP2025050640_24072025_PF_FP_ABST
Abstract
Description
SIGNALING OF BLOCK-BASED REGION-ADAPTIVE HIERARCHICAL TRANSFORM (RAHT) IN A POINT CLOUD CODECCROSS REFERENCE
[0001] This application claims the benefit of European Patent Application No. 24305091.1 , filed 15 January 2024, the entire disclosure of which is incorporated herein by reference.BACKGROUND
[0002] Advances in 3D capturing and rendering technologies are enabling new applications and services in the fields of autonomous driving, cultural heritage archival, immersive telepresence and virtual / augmented reality. Point clouds have arisen as one of the main 3D scene representations for such applications. A point cloud frame consists of a set of 3D points, each point being represented with its 3D position and possibly one or more attributes such as color, transparency, reflectance, and the like.
[0003] A standardization activity for point cloud compression is carried out by the ISO / IEC JTC1 / SC29 / WG7 “MPEG 3D Graphics and Haptics Coding” group, as described for example in D. Graziosi et al., “An overview of ongoing point cloud compression standardization activities: video-based (V-PCC) and geometry-based (G-PCC),” APSIPA Transactions on Signal and Information Processing, Volume 9, Vol. 9, No. 1 , e13, 2020. The first edition of the Geometrybased Point Cloud Compression standard (G-PCC) standard, part 9 of the ISO / IEC 23090 series on the coded representation of immersive media was published as ISO / IEC 23090-9:2023, Information technology — Coded representation of immersive media — Part 9: Geometry-based point cloud compression.
[0004] Within the framework of G-PCC second edition in construction, the compression of dense dynamic point clouds with a geometry-based approach - that is in the 3D space domain, without 3D-to-2D round trip to leverage existing 2D video codecs - has been identified as a separate target, and a preliminary specification is being drafted, as described in “Technology under consideration for Solid G-PCC,” ISO / IEC JTC1 / SC29 / WG7, 144th MPEG meeting, Hannover, Tech. Rep. N00759, October 2023. The current G-PCC encoder for dense dynamic point clouds, so called GeS-TM, is described in “Test model for geometry-based solid point cloud GeS TM v4.0,” ISO / IEC JTC1 / SC29 / WG7, 144th MPEG meeting, Hannover, Tech. Rep. N00750, October 2023. It combines the following tools:- Pruned occupancy tree (octree) and triangle soups for geometry coding.- Region-Adaptive Hierarchical Transform (RAHT) for color attribute coding.- Motion compensated inter-frame prediction.- Context-adaptive arithmetic coding.
[0005] In a conventional point cloud codec, the geometry is first compressed, then input attributes are transferred onto the reconstructed geometry (after decompression) and compressed. Which means that at decoder side, the reconstructed geometry is available when decoding the attributes.
[0006] Region-Adaptive Hierarchical Transform used for encoding and decoding of point cloud attributes (such as color) is described in, for example, R. L. de Queiroz and P. A. Chou, “Compression of 3d point clouds using a region-adaptive hierarchical transform,” IEEE Transactions on Image Processing, vol. 25, no. 8, pp. 3947-3956, Aug. 2016, which is incorporated herein by reference in its entirety.SUMMARY
[0007] A point cloud encoding method according to an example embodiment comprises: obtaining data representing a geometry of a point cloud as a plurality of voxels arranged in an octree; obtaining data representing at least one attribute for each of a plurality of the voxels; partitioning the octree into a plurality of subtrees; and for each respective subtree, independently from other subtrees, encoding the attributes of the voxels in the respective subtree using a region-adaptive hierarchical transform.
[0008] Some embodiments further include signaling in a bitstream information indicating that the attributes of the respective subtrees are encoded independently.
[0009] In some embodiments, each of the subtrees is an attribute coding unit (ACU). In some embodiments, each ACU corresponds to a geometry coding unit (GCU).
[0010] In some embodiments, each of the subtrees has a respective subtree root, and all of the subtree roots are at a same level of the octree. Some such embodiments further include signaling the level of the subtree roots in the bitstream.
[0011] Some embodiments further include: signaling in a bitstream information indicating that the attributes of the respective subtrees are encoded independently; and in response to signaling in the bitstream information indicating that the attributes of the respective subtrees are encoded independently, signaling the level of the subtree roots in the bitstream.
[0012] In some embodiments, the subtrees are non-overlapping. In some embodiments, the subtrees span all occupied voxels of the octree.
[0013] In some embodiments, for at least two of the subtrees, the encoding of attributes is performed in parallel.
[0014] In some embodiments, the attribute is a color parameter.
[0015] In some embodiments, independently encoding the attributes of a subtree using a region- adaptive hierarchical transform comprises: obtaining a plurality of transform coefficients based on the attributes in the subtree; and entropy coding the transform coefficients in a bitstream.
[0016] In some embodiments, the plurality of transform coefficients for a subtree are obtained independently from attributes of voxels in a different subtree. In some embodiments, the plurality of transform coefficients for a subtree are obtained independently from attributes of voxels in any other subtree.
[0017] A point cloud decoding method according to some embodiments comprises: obtaining data representing a geometry of a point cloud as a plurality of voxels arranged in an octree; for each of a plurality of subtrees in the octree, obtaining data representing coefficients in a region-adaptive hierarchical transform of attributes in the subtree; and for each respective subtree, independently from other subtrees, decoding the attributes of the voxels in the respective subtree using the region-adaptive hierarchical transform.
[0018] Some embodiments further include reading from a bitstream information indicating that the attributes of the respective subtrees are encoded independently.
[0019] In some embodiments, each of the subtrees is an attribute coding unit (ACU).
[0020] In some embodiments, each of the subtrees has a respective subtree root, and wherein all of the subtree roots are at a same level of the octree.
[0021] Some embodiments further include reading from the bitstream information indicating the level of the subtree roots.
[0022] Some embodiments further include: reading from the bitstream information indicating that the attributes of the respective subtrees are encoded independently; and in response to the information indicating that the attributes of the respective subtrees are encoded independently, reading from the bitstream information indicating the level of the subtree roots in the bitstream.
[0023] In some embodiments, the subtrees are non-overlapping. In some embodiments, the subtrees span all occupied voxels of the octree.
[0024] In some embodiments, for at least two of the subtrees, the decoding of attributes is performed in parallel.
[0025] In some embodiments, independently decoding the attributes of a subtree using the region-adaptive hierarchical transform comprises: entropy decoding a plurality of transform coefficients from the bitstream; and obtaining the attributes in the subtree based on the plurality of transform coefficients.
[0026] In some embodiments, the attributes for the voxels in a respective subtree are obtained independently from transform coefficients associated with different subtree.
[0027] An apparatus according to some embodiments comprises one or more processors, the apparatus being configured to perform any of the methods described herein.
[0028] An apparatus according to some embodiments comprises at least one processor and a computer-readable medium storing instructions for performing any of the methods described herein.
[0029] A computer-readable medium according to some embodiments store instructions for performing any of the methods described herein.
[0030] A computer-readable medium according to some embodiments stores a mesh encoded according to any of the methods described herein.
[0031] A signal according to some embodiments conveys a mesh encoded according to any of the methods described herein.BRIEF DESCRIPTION OF THE DRAWINGS
[0032] FIG. 1 is a functional block diagram of a point cloud encoder and decoder.
[0033] FIG. 2 schematically illustrates an RAHT transform tree.
[0034] FIG. 3 schematically illustrates a set of subtrees used in a block-based region-adaptive hierarchical transform according to an example embodiment. In this example, each subtree is an attribute coding unit (ACU).
[0035] FIG. 4 is a flow diagram of a point cloud encoding method according to some embodiments.
[0036] FIG. 5 is a flow diagram of a point cloud decoding method according to some embodiments.
[0037] FIG. 6 is a functional block diagram of an apparatus implementing a point cloud encoder and / or decoder according to example embodiments.DETAILED DESCRIPTION
[0038] In an example point cloud codec, geometry and attributes are sequentially compressed, as illustrated in FIG. 1 where the attributes are the colors in this example. At a point cloud encoder, a geometry encoder receives an input point cloud 102 with attributes and encodes the geometry at 104. The encoding of the geometry may be accomplished, for example, by dividing the space using an octree and signaling information indicating, for each node in the octree, whether the node is occupied by any point of the input point cloud. The encoded geometry is provided to a geometry decoder 106 that generates a reconstructed geometry (which, in the case of lossyencoding, may differ in some respects from the geometry of the input point cloud). An attribute transfer process 108 is applied to transfer the attributes (such as colors) of points of the input point cloud to points of the reconstructed point cloud (e.g. using interpolation filters or other techniques). Attributes of the points of the reconstructed point cloud are encoded at 110, e.g. using a region-adaptive hierarchical transform. Data representing the attributes and data representing the geometry is multiplexed at 112 into a bitstream 114 that may be transmitted in real time or stored for later processing or use as desired. A point cloud decoder receives the bitstream, and at 116 it demultiplexes the data representing the geometry and the data representing the attributes. A geometry decoder 118 reconstructs the geometry from the data and provides the reconstructed geometry to an attribute decoder 120. The attribute decoder uses the data representing the attributes to apply the attributes to the points of the reconstructed geometry, resulting in a decoded point cloud 122.
[0039] It is desirable for a point cloud codec to have a relatively low complexity, latency, and memory footprint. This is especially true when targeting applications such as teleconferencing, for which real-time encoding and decoding is required. For that purpose, a localization of the coding processes within the point cloud frame is useful in order to enable a parallelizable fast encoding / decoding scheme.
[0040] Currently, the geometry encoding part of GeS-TM is already compatible with a parallelized implementation, with the frame being divided in coding units that can be independently encoded / decoded. This is not the case for attribute encoding, however. In general, a region- adaptive hierarchical transform (RAHT) is used to encoding the entire frame’s attributes together.
[0041] In some example embodiments, the RAHT transform is locally applied on 3D-blocks corresponding to an independently coded geometry coding unit (GCU), or more generally on smaller subsets or larger supersets of GCUs. Such a technique may be referred to herein as a block-based RAHT attribute encoding method and may be used as an alternative to the original RAHT. Such a technique may have lower encoding efficiency (as not taking full advantage of the spatial correlations of attributes over the entire frame), but it is expected to have higher encoding speed and a lower memory footprint.
[0042] In some embodiments, the use of a block-based RAHT attribute encoding method may be signaled in a bitstream from an encoded to the decoder, together with its related parameters.
[0043] FIG. 2 provide an illustration of a non-block-based RAHT, where the hierarchical transform is applied up to Nthlevel, i.e. up to the root of the tree. For the sake of illustration, a binary tree is represented, but the RAHT transform tree is an octree where each node is divided into 2 in the x, y and z directions yielding up to 8 possible child nodes, depending on their occupancy.
[0044] FIG. 3 illustrates an alternative block-based RAHT according to an example embodiment. In the example of FIG. 3, the transform is performed up to a certain level. The dashed linesseparate different attribute coding units (ACU). They are processed independently, hence leveraging a parallel encoding feature.
[0045] In some embodiments, the stopping level corresponds to a geometry coding unit (GCU). In this case, each attribute coding unit has the same size as a geometry coding unit, and there are as many ACUs as GCUs. In other embodiments, the stopping level may be above or below the size of a GCU.
[0046] Once a block is transformed, it can be entropy coded and transmitted without waiting for the transformation of other blocks.
[0047] In an example embodiment, the attribute bitstream includes a concatenation of each independently transformed block.
[0048] Example embodiments further include systems and methods for signaling information associated with the localized block-based RAHT transform.
[0049] In some embodiments, the information associated with the localized block-based RAHT transform is signaled in an attribute parameter set data unit. In some embodiments, first information is signaled indicating whether a block-based region adaptive hierarchical transform is used. In cases where the first information indicates that block-based region adaptive hierarchical transform is used, second information is signaled indicating a level associated with the blockbased RAHT transform.
[0050] Some embodiments to support the block-based variant of the RAHT transform are implemented through a modification of the Solid G-PCC codec, which is described in “Technology under consideration for Solid G-PCC,” ISO / IEC JTC1 / SC29 / WG7, 144th MPEG meeting, Hannover, Tech. Rep. N00759, October 2023. In one such embodiment, the information associated with the block-based RAHT transform is provided in an Attribute Parameter Set (APS). The APS specifies properties related to attribute coding at the point cloud sequence level. In an example, a new coding type B-RAHT is introduced for the block-based RAHT, in addition to the existing RAHT and raw ones.
[0051] According to an example embodiment, the attribute parameter set data unit syntax (as described in section 7.3.2.6 of the Solid G-PCC codec) is modified as follows. Modified lines are indicated with a dagger (t).
[0052] According to an example embodiment, the semantics of the attribute parameter set data unit may be as follows:• attr_coding_type specifies the attribute coding method. Valid values are specified in the table below. Other values are reserved for future use by ISO / IEC. Decoders conforming to this version of this document shall ignore (remove from the bitstream and discard) attribute data units coded with reserved values of attr_coding_type.Table — Interpretation of attr_coding_type• block_raht_root_level specifies the root level of block-based local RAHT transform in the global RAHT transform tree. The value of block_raht_root_level is selected to be less than RahtRootLvI.
[0053] Example embodiments allow for a parallel encoding of point cloud attributes for faster encoding / decoding, with a smaller memory footprint, still using the RAHT transform.
[0054] An example point cloud encoding method according to some embodiments is illustrated in FIG. 4. In the example method, an encoder at 402 obtains data representing a geometry of a point cloud as a plurality of voxels arranged in an octree. At 404 the encoder also obtains data representing at least one attribute for each of a plurality of the voxels. At 406, the encoder partitions the octree into a plurality of subtrees. The partitioning may be performed, for example, by identifying a root node for each of a plurality of subtrees. Each subtree may be a separate attribute coding unit (ACU) as shown in FIG. 3. For each respective subtree, the attributes of the voxels in the respective subtree are encoded (e.g. at 408) using a region-adaptive hierarchical transform. The attributes of each subtree may be encoded independently from other subtrees. The encoded attribute information of the subtrees may be signaled in a bitstream at 410.
[0055] In some embodiments, information is further signaled in the bitstream indicating that the attributes of the respective subtrees are encoded independently. For example, a parameter such as attr_coding_type may take different values for different encoding types, with one value indicating the use of conventional RAHT and a different value indicating the use of block-based RAHT in which attributes of the respective subtrees are encoded independently.
[0056] In some embodiments, each of the subtrees has a respective subtree root node. All of the leaf nodes (which may be voxels) in the subtree may descend from the same subtree root node.
[0057] In some such embodiments, all of the subtree root nodes are at a same level of the octree. For example, in the tree of FIG. 3, the root nodes of all of the subtrees are at Level 2. In some embodiments, information indicating the level of the root nodes is signaled in the bitstream, for example using a parameter such as block_raht_root_level. In some embodiments, the level of the root nodes is signaled only if a parameter such as attr_coding_type indicates the use of blockbased RAHT.
[0058] In some embodiments, as shown in FIG. 3, each ACU is a subtree that does not overlap (e.g. does not share any nodes with) any other subtree.
[0059] In some embodiments, the subtrees for which attributes are encoded span all occupied voxels of the octree. Unoccupied voxels may or may not belong to any subtree for which attributes are encoded.
[0060] In some embodiments, the encoding of attributes for different subtrees (e.g. for at least two of the subtrees) is performed in parallel. For example, the encoding of attributes of onesubtree may begin even before encoding of attributes of a second one of the subtrees is completed.
[0061] In some embodiments, the encoding the attributes of a subtree using a region-adaptive hierarchical transform includes obtaining a plurality of transform coefficients based on the attributes in the subtree, and entropy coding the transform coefficients in a bitstream. The plurality of transform coefficients for one subtree may be obtained independently from attributes of voxels in a different subtree, allowing for parallel processing.
[0062] An example point cloud decoding method according to some embodiments is illustrated in FIG. 5. At 502, a decoder obtains data representing a geometry of a point cloud as a plurality of voxels arranged in an octree. The octree includes a plurality of subtrees. Each subtree may be a separate attribute coding unit (ACU) as shown in FIG. 3. For each of a plurality of subtrees in the octree, the decoder obtains at 504 data representing coefficients in a region-adaptive hierarchical transform of attributes in the subtree. For each respective subtree, the decoder decodes the attributes of the voxels in the respective subtree using the region-adaptive hierarchical transform, e.g. at 506. The decoding of the attributes of the voxels in a subtree may be performed independently from the attribute decoding in other subtrees.
[0063] In some embodiments, the decoder reads from the bitstream information indicating that the attributes of the respective subtrees are encoded independently. For example, a parameter such as attr_coding_type may take different values for different encoding types, with one value indicating the use of conventional RAHT and a different value indicating the use of block-based RAHT in which attributes of the respective subtrees are encoded independently.
[0064] In some embodiments, each of the subtrees has a respective subtree root node. All of the leaf nodes (which may be voxels) in the subtree may descend from the same subtree root node.
[0065] In some such embodiments, all of the subtree root nodes are at a same level of the octree. For example, in the tree of FIG. 3, the root nodes of all of the subtrees are at Level 2. In some embodiments, information indicating the level of the root nodes is signaled in the bitstream, for example using a parameter such as block_raht_root_level, and is read by the decoder. In some embodiments, the level of the root nodes is read by the decoder only if a parameter such as attr_coding_type indicates the use of block-based RAHT.
[0066] In some embodiments, the decoding of attributes for different subtrees (e.g. for at least two of the subtrees) is performed in parallel. For example, the decoding of attributes of one subtree may begin even before decoding of attributes of a second one of the subtrees is completed.
[0067] In some embodiments, the decoding of attributes of a subtree using a region-adaptive hierarchical transform includes entropy decoding from a bitstream a plurality of transformcoefficients based on the attributes in the subtree. An inverse RAHT transform may then be performed on the coefficients to obtain the attributes associated with the voxels in the subtree.
[0068] In some embodiments, the attributes for one subtree may be decoded independently from transform coefficients of a different subtree, allowing for parallel processing.
[0069] In some embodiments, the decoding of attributes for different subtrees (e.g. for at least two of the subtrees) is performed in parallel. For example, the decoding of attributes of one subtree may begin even before decoding of attributes of a second one of the subtrees is completed.
[0070] In some embodiments, the decoding the attributes of a subtree using a region-adaptive hierarchical transform includes entropy decoding a plurality of transform coefficients from the bitstream and obtaining the attributes in the subtree based on the plurality of transform coefficients, e.g. using an inverse transform.
[0071] An apparatus according to some embodiments comprises one or more processors, the apparatus being configured to perform any of the methods described herein.
[0072] An apparatus according to some embodiments comprises at least one processor and a computer-readable medium storing instructions for performing any of the methods described herein.
[0073] A computer-readable medium according to some embodiments store instructions for performing any of the methods described herein.
[0074] A computer-readable medium according to some embodiments stores a point cloud encoded according to any of the methods described herein.
[0075] A signal according to some embodiments conveys a point cloud encoded according to any of the methods described herein.Example system hardware.
[0076] Example embodiments of encoders and / or decoders (collectively coders) configured to implement embodiments described herein may be implemented using systems such as the system of FIG. 6. FIG. 6 is a block diagram of an example of a system in which various aspects and embodiments are implemented. System 1000 can be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this document. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 1000, singly or in combination, can be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoderelements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, the system 1000 is communicatively coupled to one or more other systems, or other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports. In various embodiments, the system 1000 is configured to implement one or more of the aspects described in this document.
[0077] The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this document. Processor 1010 can include embedded memory, input output interface, and various other circuitries as known in the art. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device, and / or a non-volatile memory device). System 1000 includes a storage device 1040, which can include non-volatile memory and / or volatile memory, including, but not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), flash, magnetic disk drive, and / or optical disk drive. The storage device 1040 can include an internal storage device, an attached storage device (including detachable and non-detachable storage devices), and / or a network accessible storage device, as non-limiting examples.
[0078] System 1000 includes an encoder / decoder module 1030 configured, for example, to process data to provide an encoded point cloud or decoded point cloud, and the encoder / decoder module 1030 can include its own processor and memory. The encoder / decoder module 1030 represents module(s) that can be included in a device to perform the encoding and / or decoding functions. As is known, a device can include one or both of the encoding and decoding modules. Additionally, encoder / decoder module 1030 can be implemented as a separate element of system 1000 or can be incorporated within processor 1010 as a combination of hardware and software as known to those skilled in the art.
[0079] Program code to be loaded onto processor 1010 or encoder / decoder 1030 to perform the various aspects described in this document can be stored in storage device 1040 and subsequently loaded onto memory 1020 for execution by processor 1010. In accordance with various embodiments, one or more of processor 1010, memory 1020, storage device 1040, and encoder / decoder module 1030 can store one or more of various items during the performance of the processes described in this document. Such stored items can include, but are not limited to, the input point cloud, the decoded point cloud or portions of the decoded point cloud, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0080] In some embodiments, memory inside of the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and to provide working memory for processing that is needed during encoding or decoding. In other embodiments, however, a memory external to theprocessing device (for example, the processing device can be either the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory can be the memory 1020 and / or the storage device 1040, for example, a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of, for example, a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for coding and decoding operations, such as for MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also referred to as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or WC (Versatile Video Coding, a new standard being developed by JVET, the Joint Video Experts Team).
[0081] The input to the elements of system 1000 can be provided through various input devices as indicated in block 1130. Such input devices include, but are not limited to, (i) a radio frequency (RF) portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Component (COMP) input terminal (or a set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples include composite video.
[0082] In various embodiments, the input devices of block 1130 have associated respective input processing elements as known in the art. For example, the RF portion can be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) downconverting the selected signal, (iii) bandlimiting again to a narrower band of frequencies to select (for example) a signal frequency band which can be referred to as a channel in certain embodiments, (iv) demodulating the downconverted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion can include a tuner that performs various of these functions, including, for example, downconverting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Adding elements can include inserting elements in between existing elements, such as, for example, insertingamplifiers and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna.
[0083] Additionally, the USB and / or HDMI terminals can include respective interface processors for connecting system 1000 to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed-Solomon error correction, can be implemented, for example, within a separate input processing IC or within processor 1010 as necessary. Similarly, aspects of USB or HDMI interface processing can be implemented within separate interface ICs or within processor 1010 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 1010, and encoder / decoder 1030 operating in combination with the memory and storage elements to process the datastream as necessary for presentation on an output device.
[0084] Various elements of system 1000 can be provided within an integrated housing, Within the integrated housing, the various elements can be interconnected and transmit data therebetween using suitable connection arrangement 1140, for example, an internal bus as known in the art, including the Inter-IC (I2C) bus, wiring, and printed circuit boards.
[0085] The system 1000 includes communication interface 1050 that enables communication with other devices via communication channel 1060. The communication interface 1050 can include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 1060. The communication interface 1050 can include, but is not limited to, a modem or network card and the communication channel 1060 can be implemented, for example, within a wired and / or a wireless medium.
[0086] Data is streamed, or otherwise provided, to the system 1000, in various embodiments, using a wireless network such as a Wi-Fi network, for example IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal of these embodiments is received over the communications channel 1060 and the communications interface 1050 which are adapted for Wi-Fi communications. The communications channel 1060 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 1000 using a set-top box that delivers the data over the HDMI connection of the input block 1130. Still other embodiments provide streamed data to the system 1000 using the RF connection of the input block 1130. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, for example a cellular network or a Bluetooth network.
[0087] The system 1000 can provide an output signal to various output devices, including a display 1100, speakers 1110, and other peripheral devices 1120. The display 1100 of various embodiments includes one or more of, for example, a touchscreen display, an organic lightemitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 can be for a television, a tablet, a laptop, a cell phone (mobile phone), or other device. The display 1100 can also be integrated with other components (for example, as in a smart phone), or separate (for example, an external monitor for a laptop). The other peripheral devices 1120 include, in various examples of embodiments, one or more of a stand-alone digital video disc (or digital versatile disc) (DVR, for both terms), a disk player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 1120 that provide a function based on the output of the system 1000. For example, a disk player performs the function of playing the output of the system 1000.
[0088] In various embodiments, control signals are communicated between the system 1000 and the display 1100, speakers 1110, or other peripheral devices 1120 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communications protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, the output devices can be connected to system 1000 using the communications channel 1060 via the communications interface 1050. The display 1100 and speakers 1110 can be integrated in a single unit with the other components of system 1000 in an electronic device such as, for example, a television. In various embodiments, the display interface 1070 includes a display driver, such as, for example, a timing controller (T Con) chip.
[0089] The display 1100 and speaker 1110 can alternatively be separate from one or more of the other components, for example, if the RF portion of input 1130 is part of a separate set-top box. In various embodiments in which the display 1100 and speakers 1110 are external components, the output signal can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0090] The embodiments can be carried out by computer software implemented by the processor 1010 or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. The memory 1020 can be of any type appropriate to the technical environment and can be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. The processor 1010 can be of any type appropriate to the technical environment, and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non-limiting examples.Additional embodiments.
[0091] A point cloud encoding method according to some embodiments comprises: obtaining data representing a geometry of a point cloud as a plurality of voxels arranged in an octree; obtaining data representing an attribute value for each of a plurality of the voxels; partitioning the octree into a plurality of subtrees; and for each respective subtree, independently from other subtrees, encoding the attribute values of the voxels in the respective subtree using a region-adaptive hierarchical transform.
[0092] A point cloud encoding apparatus according to some embodiments comprises one or more processors configured to perform at least: obtaining data representing a geometry of a point cloud as a plurality of voxels arranged in an octree; obtaining data representing an attribute value for each of a plurality of the voxels; partitioning the octree into a plurality of subtrees; and for each respective subtree, independently from other subtrees, encoding the attribute values of the voxels in the respective subtree using a region-adaptive hierarchical transform.
[0093] Some embodiments further include: signaling in a bitstream information indicating that the attribute values of the voxels in the respective subtrees are encoded independently from other subtrees.
[0094] In some embodiments, each of the subtrees is an attribute coding unit (ACU).
[0095] In some embodiments, each of the subtrees has a respective subtree root, and all of the subtree roots are at a same level of the octree.
[0096] Some embodiments further comprise signaling the level of the subtree roots in the bitstream.
[0097] Some embodiments further comprise signaling in a bitstream information indicating that the attribute values of the voxels in the respective subtrees are encoded independently; and in response to signaling in the bitstream information indicating that the attribute values of the voxels in the respective subtrees are encoded independently, signaling the level of the subtree roots in the bitstream.
[0098] In some embodiments, the subtrees are non-overlapping.
[0099] In some embodiments, the subtrees span all occupied voxels of the octree.
[0100] In some embodiments, for at least a first one of the subtrees and a second one of the subtrees, the encoding of the attribute values of the voxels in the first subtree is performed in parallel with the encoding of the attribute values of the voxels in the second subtree.
[0101] In some embodiments, the attribute values represent a color parameter.
[0102] In some embodiments, independently encoding the attributes values of the voxels in a subtree using a region-adaptive hierarchical transform comprises: obtaining a plurality oftransform coefficients based on the attribute values of the voxels in the subtree; and entropy coding the transform coefficients in a bitstream.
[0103] In some embodiments, for each of the subtrees, the plurality of transform coefficients for the respective subtree are obtained independently of the attribute values of the voxels in different subtrees.
[0104] A point cloud decoding method in some embodiments comprises: obtaining data representing a geometry of a point cloud as a plurality of voxels arranged in an octree; for each of a plurality of subtrees in the octree, obtaining data representing coefficients in a region-adaptive hierarchical transform of attribute values in the subtree; and for each respective subtree, independently from other subtrees, decoding attribute values of the voxels in the respective subtree using the region-adaptive hierarchical transform.
[0105] Some embodiments further comprise reading from a bitstream information indicating that the attribute values of the voxels in the respective subtrees are encoded independently from other subtrees.
[0106] In some embodiments, each of the subtrees has a respective subtree root, and all of the subtree roots are at a same level of the octree.
[0107] Some embodiments further comprise reading from the bitstream information indicating the level of the subtree roots.
[0108] Some embodiments further include: reading from the bitstream information indicating that the attribute values of the voxels in the respective subtrees are encoded independently from other subtrees; and in response to the information indicating that the attribute values of the voxels in the respective subtrees are encoded independently, reading from the bitstream information indicating the level of the subtree roots in the bitstream.
[0109] In some embodiments, the subtrees are non-overlapping.
[0110] In some embodiments, the subtrees span all occupied voxels of the octree.
[0111] In some embodiments, for at least two of the subtrees, the decoding of attributes is performed in parallel.
[0112] In some embodiments, the attribute is a color parameter.
[0113] In some embodiments, independently decoding the attribute values of the voxels in a subtree using the region-adaptive hierarchical transform comprises: entropy decoding a plurality of transform coefficients from the bitstream; and obtaining the attribute values of the voxels in the subtree based on the plurality of transform coefficients.
[0114] In some embodiments, the attribute values of the voxels in a respective subtree are obtained independently from transform coefficients associated with different subtrees.
[0115] Some embodiments include at least one processor and a computer-readable medium storing instructions for performing any of the methods described herein.
[0116] Some embodiments include a computer-readable medium storing instructions for performing any of the methods described herein.
[0117] A computer-readable medium according to some embodiments stores a point cloud encoded according to any of the methods described herein.
[0118] A signal according to some embodiments conveys a point cloud encoded according to any of the methods described herein.
[0119] This disclosure describes a variety of aspects, including tools, features, embodiments, models, approaches, etc. Many of these aspects are described with specificity and, at least to show the individual characteristics, are often described in a manner that may sound limiting. However, this is for purposes of clarity in description, and does not limit the disclosure or scope of those aspects. Indeed, all of the different aspects can be combined and interchanged to provide further aspects. Moreover, the aspects can be combined and interchanged with aspects described in earlier filings as well.
[0120] The aspects described and contemplated in this disclosure can be implemented in many different forms. While some embodiments are illustrated specifically, other embodiments are contemplated, and the discussion of particular embodiments does not limit the breadth of the implementations. At least one of the aspects generally relates to point cloud encoding and decoding, and at least one other aspect generally relates to transmitting a bitstream generated or encoded. These and other aspects can be implemented as a method, an apparatus, a computer readable storage medium having stored thereon instructions for encoding or decoding point cloud data according to any of the methods described, and / or a computer readable storage medium having stored thereon a bitstream generated according to any of the methods described.
[0121] Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., such as, for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.
[0122] Various numeric values may be used in the present disclosure, for example. The specific values are for example purposes and the aspects described are not limited to these specific values.
[0123] Embodiments described herein may be carried out by computer software implemented by a processor or other hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. The processor can be of any type appropriate to the technical environment and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non-limiting examples.
[0124] When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of a corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of a corresponding method / process.
[0125] The implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed can also be implemented in other forms (for example, an apparatus or program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate communication of information between endusers.
[0126] Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this disclosure are not necessarily all referring to the same embodiment.
[0127] Additionally, this disclosure may refer to “determining” various pieces of information. Determining the information can include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
[0128] Further, this disclosure may refer to “accessing” various pieces of information. Accessing the information can include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0129] Additionally, this disclosure may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information can include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations such as, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0130] It is to be appreciated that the use of any of the following 7”, “and / or”, and “at least one of’, for example, in the cases of “A / B”, “A and / or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended for as many items as are listed.
[0131] Also, as used herein, the word “signal” refers to, among other things, indicating something to a corresponding decoder. For example, in certain embodiments the encoder signals a particular one of a plurality of parameters for region-based filter parameter selection for de-artifact filtering. In this way, in an embodiment the same parameter is used at both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways. For example, one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments. While the preceding relates to the verb form of the word “signal”, the word “signal” can also be used herein as a noun.
[0132] Implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal can be formatted to carry the bitstream of a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting can include, for example, encodinga data stream and modulating a carrier with the encoded data stream. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
[0133] We describe a number of embodiments. Features of these embodiments can be provided alone or in any combination, across various claim categories and types. Further, embodiments can include one or more of the following features, devices, or aspects, alone or in any combination, across various claim categories and types:
[0134] A bitstream or signal that includes one or more of the described syntax elements, or variations thereof.
[0135] A bitstream or signal that includes syntax conveying information generated according to any of the embodiments described.
[0136] Creating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the described syntax elements, or variations thereof.
[0137] Creating and / or transmitting and / or receiving and / or decoding according to any of the embodiments described.
[0138] A method, process, apparatus, medium storing instructions, medium storing data, or signal according to any of the embodiments described.
[0139] Note that various hardware elements of one or more of the described embodiments are referred to as “modules” that carry out (i.e., perform, execute, and the like) various functions that are described herein in connection with the respective modules. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memory devices) deemed suitable for a given implementation. Each described module may also include instructions executable for carrying out the one or more functions described as being carried out by the respective module, and it is noted that those instructions could take the form of or include hardware (i.e., hardwired) instructions, firmware instructions, software instructions, and / or the like, and may be stored in any suitable non-transitory computer-readable medium or media, such as commonly referred to as RAM, ROM, etc.
[0140] Although features and elements are described above in particular combinations, each feature or element can be used alone or in any combination with the other features and elements. In addition, the methods described herein may be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, a readonly memory (ROM), a random access memory (RAM), a register, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magnetooptical media, and optical media such as CD-ROM disks, and digital versatile disks (DVDs). A processor in association with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
CLAIMS1. A point cloud encoding method comprising: obtaining data representing a geometry of a point cloud as a plurality of voxels arranged in an octree; obtaining data representing an attribute value for each of a plurality of the voxels; partitioning the octree into a plurality of subtrees; and for each respective subtree, independently from other subtrees, encoding the attribute values of the voxels in the respective subtree using a region-adaptive hierarchical transform.
2. A point cloud encoding apparatus comprising one or more processors configured to perform at least: obtaining data representing a geometry of a point cloud as a plurality of voxels arranged in an octree; obtaining data representing an attribute value for each of a plurality of the voxels; partitioning the octree into a plurality of subtrees; and for each respective subtree, independently from other subtrees, encoding the attribute values of the voxels in the respective subtree using a region-adaptive hierarchical transform.
3. The method of claim 1 or the apparatus of claim 2, further comprising: signaling in a bitstream information indicating that the attribute values of the voxels in the respective subtrees are encoded independently from other subtrees.
4. The method of claim 1 , or claim 3 as it depends from claim 1 , or the apparatus of claim 2, or claim 3 as it depends from claim 2, wherein each of the subtrees has a respective subtree root, and wherein all of the subtree roots are at a same level of the octree.
5. The method of claim 4 as it depends from claim 1 , or the apparatus of claim 4 as it depends from claim 2, further comprising signaling the level of the subtree roots in the bitstream.
6. The method of claim 4 as it depends from claim 1 , or the apparatus of claim 4 as it depends from claim 2, further comprising: signaling in a bitstream information indicating that the attribute values of the voxels in the respective subtrees are encoded independently; and in response to signaling in the bitstream information indicating that the attribute values of the voxels in the respective subtrees are encoded independently, signaling the level of the subtree roots in the bitstream.
7. The method of claim 1 , or any of claims 3-6 as they depend from claim 1 , or the apparatus of claim 2, or any of claims 3-6 as they depend from claim 2, wherein for at least a first one of the subtrees and a second one of the subtrees, the encoding of the attribute values of the voxels in the first subtree is performed in parallel with the encoding of the attribute values of the voxels in the second subtree.
8. The method of claim 1 , or any of claims 3-7 as they depend from claim 1 , or the apparatus of claim 2, or any of claims 3-7 as they depend from claim 2, wherein independently encoding the attributes values of the voxels in a subtree using a region-adaptive hierarchical transform comprises: obtaining a plurality of transform coefficients based on the attribute values of the voxels in the subtree; and entropy coding the transform coefficients in a bitstream.
9. A point cloud decoding method comprising: obtaining data representing a geometry of a point cloud as a plurality of voxels arranged in an octree; for each of a plurality of subtrees in the octree, obtaining data representing coefficients in a region-adaptive hierarchical transform of attribute values in the subtree; and for each respective subtree, independently from other subtrees, decoding attribute values of the voxels in the respective subtree using the region-adaptive hierarchical transform.
10. A point cloud decoding apparatus comprising one or more processors configured to perform at least: obtaining data representing a geometry of a point cloud as a plurality of voxels arranged in an octree; for each of a plurality of subtrees in the octree, obtaining data representing coefficients in a region-adaptive hierarchical transform of attribute values in the subtree; and for each respective subtree, independently from other subtrees, decoding attribute values of the voxels in the respective subtree using the region-adaptive hierarchical transform.11 . The method of claim 9 or the apparatus of claim 10, further comprising: reading from a bitstream information indicating that the attribute values of the voxels in the respective subtrees are encoded independently from other subtrees.
12. The method of claim 9, or claim 11 as it depends from claim 9, or the apparatus of claim 10, or claim 11 as it depends from claim 10, wherein each of the subtrees has a respective subtree root, and wherein all of the subtree roots are at a same level of the octree.
13. The method of claim 12 as it depends from claim 9, or the apparatus of claim 12 as it depends from claim 10, further comprising: reading from the bitstream information indicating that the attribute values of the voxels in the respective subtrees are encoded independently from other subtrees; and in response to the information indicating that the attribute values of the voxels in the respective subtrees are encoded independently, reading from the bitstream information indicating the level of the subtree roots in the bitstream.
14. The method of claim 9, or any of claims 11-13 as they depend from claim 9, or the apparatus of claim 10, or any of claims 11-13 as they depend from claim 10, wherein the subtrees are non-overlapping and span all occupied voxels of the octree.
15. The method of claim 9, or claims 11-14 as they depend from claim 9, or the apparatus of claim 10, or any of claims 11-14 as they depend from claim 10, wherein independently decoding the attribute values of the voxels in a subtree using the region-adaptive hierarchical transform comprises: entropy decoding a plurality of transform coefficients from the bitstream; and obtaining the attribute values of the voxels in the subtree based on the plurality of transform coefficients; wherein the attribute values of the voxels in a respective subtree are obtained independently from transform coefficients associated with different subtrees.
Citation Information
Patent Citations
Local coding of point cloud attributes
WO2025019418A1