Method and apparatus for encoding point cloud coefficients
By decomposing the transformation coefficients of point cloud data into set-index values and symbol-index values, and performing entropy coding and bypass encoding, the problem of poor coding complexity and compression efficiency in the prior art is solved, and more efficient point cloud data compression is achieved.
Patent Information
- Application Number
- CN202180002845.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-03
- Filing Date
- 2021-01-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-01-07
AI Technical Summary
The existing point cloud data encoding methods are still not satisfactory in terms of complexity/memory and compression efficiency, and are difficult to meet the needs of real-time communication and efficient compression.
Efficient compression of point cloud data is achieved by decomposing the transform coefficients associated with point cloud data into set-index values and symbol-index values, and entropy encoding and bypass encoding.
This method improves the encoding of the transform coefficients of attributes in G-PCC through the encoding of alphabet-division and alphabet-division information, and improves the balance of compression efficiency and coding complexity.
Smart Images

Figure CN114026789B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 958,839, filed on January 9, 2020, and U.S. Patent Application No. 17 / 110,691, filed on December 3, 2020, both of which are hereby incorporated by reference in their entireties. Technical Field
[0003] This application relates to graph - based point cloud compression (G - PCC), and more particularly, to methods and apparatus for encoding point cloud coefficients. Background Art
[0004] Advanced three - dimensional (3D) representations of the world enable more immersive forms of interaction and communication and also enable machines to understand, interpret, and navigate our world. 3D point clouds emerge as an enabling representation of such information. Many use cases associated with point cloud data have been identified, and corresponding requirements for point cloud representation and compression have been established. For example, point clouds can be used for object detection and localization in autonomous driving. Point clouds can also be used for mapping in geographic information systems (GIS) and for visualizing and archiving cultural heritage objects and collections in cultural heritage.
[0005] A point cloud is a collection of points in 3D space, each point having associated attributes such as color, material properties, etc. Point clouds can be used to reconstruct an object or scene as a combination of such points. These points can be captured using multiple cameras, depth sensors, or lidar sensors in various setups and can consist of thousands up to billions of points in order to realistically represent the reconstructed scene.
[0006] Compression techniques are needed to reduce the amount of data used to represent point clouds. Thus, techniques for lossy compression of point clouds are needed for real - time communication and six - degrees - of - freedom (6DoF) virtual reality. In addition, techniques for lossless point cloud compression are sought in the context of dynamic mapping for applications such as autonomous driving and cultural heritage. The Moving Picture Experts Group (MPEG) has started working on standards for compression that address geometric structures and attributes (such as color and reflectance, scalable / progressive coding, encoding of point cloud sequences captured over time, and random access to subsets of point clouds). However, current point cloud data encoding methods are still not satisfactory in terms of complexity / memory and compression efficiency. Summary of the Invention
[0007] This application provides a method and apparatus for encoding point cloud coefficients, which can improve the encoding of point cloud coefficients.
[0008] According to an embodiment, a method for point cloud coefficient coding is performed by at least one processor and includes decomposing transform coefficients associated with point cloud data into set-index values and sign-index values, where the sign-index values specify the positions of the transform coefficients within the sets. The decomposed transform coefficients can be partitioned into one or more sets based on the set-index values and the sign-index values. Entropy coding can be performed on the set-index values of the partitioned transform coefficients, and bypass coding can be performed on the sign-index values of the partitioned transform coefficients. The point cloud data can be compressed based on the entropy-coded set-index values and the bypass-coded sign-index values.
[0009] According to an embodiment, an apparatus for point cloud coefficient coding includes: at least one memory configured to store computer program code, and at least one processor configured to access the at least one memory and operate according to the computer program code. The computer program code includes code configured to cause the at least one processor to perform a method that can include decomposing transform coefficients associated with point cloud data into set-index values and sign-index values, where the sign-index values specify the positions of the transform coefficients within the sets. The decomposed transform coefficients can be partitioned into one or more sets based on the set-index values and the sign-index values. Entropy coding can be performed on the set-index values of the partitioned transform coefficients, and bypass coding can be performed on the sign-index values of the partitioned transform coefficients. The point cloud data can be compressed based on the entropy-coded set-index values and the bypass-coded sign-index values.
[0010] According to an embodiment, a non-transitory computer-readable storage medium stores instructions that cause at least one processor to decompose transform coefficients associated with point cloud data into set-index values and sign-index values, where the sign-index values specify the positions of the transform coefficients within the sets. The decomposed transform coefficients can be partitioned into one or more sets based on the set-index values and the sign-index values. Entropy coding can be performed on the set-index values of the partitioned transform coefficients, and bypass coding can be performed on the sign-index values of the partitioned transform coefficients. The point cloud data can be compressed based on the entropy-coded set-index values and the bypass-coded sign-index values.
[0011] In an embodiment of the present application, the transform coefficients associated with the point cloud data are decomposed into set-index values and sign-index values, and the sign-index values specify the positions of the transform coefficients in the set. The decomposed transform coefficients can be partitioned into one or more sets based on the set-index values and the sign-index values. The set-index values of the partitioned transform coefficients can be entropy encoded, and the sign-index values of the partitioned transform coefficients can be bypass encoded. The point cloud data can be compressed based on the entropy-encoded set-index values and the bypass-encoded sign-index values. In this way, the encoding of the transform coefficients from lifting, predictive transforms, and RAHT can be performed by frequency-sorted lookup table indexing, cache indexing, and direct encoding of sign values. In practice, these may require multiple lookup tables and caches with many (usually 32 to possibly up to 256) entries to cover a byte of code words. These lookup tables and caches may additionally need to be updated periodically, and the frequency may imply different trade-offs in terms of computational requirements and encoding efficiency. Therefore, it may be advantageous to improve the encoding of the transform coefficients of the attributes in G-PCC in terms of complexity / memory and compression efficiency trade-offs by alphabet partitioning and encoding of alphabet-partitioning information. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1A is a diagram showing a method for generating LoD in G-PCC.
[0013] Figure 1B is a diagram of an architecture for P / U lifting in G-PCC.
[0014] Figure 2 is a block diagram of a communication system according to an embodiment.
[0015] Figure 3 is a diagram showing the placement of a G-PCC compressor and a G-PCC decompressor in an environment according to an embodiment.
[0016] Figure 4 is a functional block diagram of a G-PCC compressor according to an embodiment.
[0017] Figure 5 is a functional block diagram of a G-PCC decompressor according to an embodiment.
[0018] Figure 6 is a flowchart showing a method for encoding point cloud coefficients according to an embodiment.
[0019] Figure 7 is a block diagram of an apparatus for encoding point cloud coefficients according to an embodiment.
[0020] Figure 8 is a diagram of a computer system suitable for implementing an embodiment. Detailed implementation mode
[0021] Figure 1A It is a diagram showing a method for generating levels of detail (LoD) in geometry-based point cloud compression (G-PCC).
[0022] Refer to Figure 1A , in the current G-PCC attribute coding, the LoD (i.e., group) of each 3D point (e.g., P0 to P9) is generated based on the distance of each 3D point, and then, by applying prediction in the LoD-based order 110 of the 3D points instead of the original order 105 of the 3D points, the attribute values of the 3D points in each LoD are encoded. For example, the attribute value of 3D point P2 is predicted by calculating the distance-based weighted average of 3D points P0, P5, and P4 that are encoded or decoded before 3D point P2.
[0023] The current anchoring method in G-PCC is carried out as follows.
[0024] First, the variability of the neighborhood of the 3D points is calculated to check the degree of difference in the neighboring values, and if the variability is below the threshold, then by predicting the attribute value (a i ) i∈0...k-1 , the calculation of the distance-based weighted average prediction is carried out using a linear interpolation process based on the distance of the nearest neighbors of the current point i. Let be the set of k nearest neighbors of the current point i, let be their decoded / reconstructed attribute values, and let be their distances from the current point i. Then, the predicted attribute value is given by the following formula:[[]]
[0025]
[0026] Note that when encoding the attributes, the geometric positions of all point clouds are already available. In addition, both the neighboring points and their reconstructed attribute values are available as a k-dimensional tree structure at both the encoder and the decoder, and this k-dimensional tree structure is used to facilitate the nearest neighbor search for each point in the same way.
[0027] Next, if the variability is higher than the threshold, perform rate-distortion optimization (RDO) prediction value selection. Based on the results of neighboring point search during LoD generation, create multiple prediction value candidates or candidate predictions. For example, when encoding the attribute value of 3D point P2 using prediction pairs, set the weighted average of the distances from 3D point P2 to 3D points P0, P5, and P4 to the prediction value index equal to 0. Then, set the distance from 3D point P2 to the nearest neighbor point P4 to the prediction value index equal to 1. Additionally, as shown in Table 1 below, set the distances from 3D point P2 to the next nearest neighbor points P5 and P0 to the prediction value indexes equal to 2 and 3, respectively.
[0028] Table 1: Samples of Prediction Value Candidates for Attribute Encoding
[0029] Predicted value index Predicted value 0 Average value 1 P4 (First nearest point) 2 P5 (Second nearest point) 3 P0 (Third nearest point)
[0030] After creating the prediction value candidates, select the best prediction value by applying rate-distortion optimization processing, then map the selected prediction value index to a truncated unary (TU) code, and perform arithmetic coding on the binary value (bin) of the truncated unary (TU) code. Note that shorter TU codes will be assigned to smaller prediction value indexes in Table 1.
[0031] Limit the maximum number of prediction value candidates MaxNumCand and encode it into the attribute header. In the current implementation, set the maximum number of prediction value candidates MaxNumCand to be equal to the number of nearest neighbors in the prediction + 1 (numberOfNearestNeighborsInPrediction + 1) and use it to encode and decode the prediction value index through truncated unary binarization.
[0032] The lifting transform for attribute encoding in G-PCC is built on top of the above-mentioned prediction transform. The main difference between the prediction scheme and the lifting scheme is the introduction of an update operator.
[0033] Figure 1BIt is a diagram of the architecture for P / U (prediction / update) enhancement in G-PCC. For the prediction step and update step during enhancement, it is necessary to split the signal into two sets with high correlation during the decomposition at each stage. In the enhancement scheme in G-PCC, the split is performed by leveraging the LoD structure, where such high correlation is expected between levels, and each level is constructed through nearest neighbor search to organize the non-uniform point cloud into structured data. The P / U decomposition step at level N generates the detail signal D(N - 1) and the approximation signal A(N - 1), and the detail signal D(N - 1) and the approximation signal A(N - 1) are further decomposed into D(N - 2) and A(N - 2). This step is repeatedly applied until the base approximation signal A(1) is obtained.
[0034] Therefore, in the enhancement scheme, instead of encoding the input attribute signal itself composed of LOD(N), …, LOD(1), ultimately D(N - 1), D(N - 2), …, D(1), A(1) are encoded. Note that the application of an effective P / U step typically results in sparse subband “coefficients” in D(N - 1), …, D(1), thus providing the advantage of transform coding gain.
[0035] Currently, the above distance-based weighted average prediction for predictive transform is used as the anchoring method in G-PCC for the prediction step during enhancement.
[0036] In the prediction and enhancement for attribute coding in G-PCC, the availability of neighboring attribute samples is crucial for compression efficiency because more neighboring attribute samples can provide better prediction. In the case where there are not enough neighboring attribute samples to perform prediction based on them, the compression efficiency may be affected.
[0037] Another type of transform for attribute coding in G-PCC can be the Region Adaptive Hierarchial Transform (RAHT). RAHT and its inverse can be performed for the hierarchy defined by the Morton code of voxel positions. The Morton code of the d-bit non-negative integer coordinates x, y, and z can be the 3d-bit non-negative integer obtained by interleaving the bits of x, y, and z. The Morton code M = morton(x, y, z) of the non-negative d-bit integer coordinates
[0038]
[0039] (where x l , y l , z l ∈{0, 1} can be the bits of x, y, and z from l = 1 (higher order) to l = d (lower order)) is the non-negative 3d-bit integer
[0040]
[0041] where m l′ ∈ {0, 1} can be the bits of M from l′ = 1 (higher order) to l′ = 3d (lower order).
[0042] can represent the l′-bit prefix of M. m can be such a prefix. The block at level l′ can be defined by the prefix m as the set of all points (x, y, z) where m = prefix l′ (morton(x, y, z)). If two blocks at level l′ have the same (l′-1)-bit prefix, they can be sibling blocks. The union of two sibling blocks at level i′ can be the block at level (l′-1) called their parent block.
[0043] Sequence A n , n = 1,..., N and its inverse's region adaptive Haar transform can include a base case and a recursive function. For the base case, A n can be the attribute of a point, and T n can be its transform, where T n = A n . For the recursive function, there can be two sibling blocks and their parent block. and can be the attributes of the points (x n , y n , z n ) listed in Morton increasing order in the sibling blocks, and and can be their respective transforms. Similarly, can be the attributes of all the points (x n , y n , z n ) in their parent block listed in Morton increasing order, and can be its transform. Then,[[]]
[0044]
[0045]
[0046] and
[0047]
[0048] where
[0049] and
[0050] The transformation of the parent block can be a concatenation of two sibling blocks, except that the first (DC) component of the transformations of the two sibling blocks can be replaced by their weighted sum and difference, and the inverse transformations of the two sibling blocks can be copied from the first part and the last part of the transformation of the parent block, except that the DC components of the transformations of the two sibling blocks can be replaced by their weighted difference and sum
[0051]
[0052] and
[0053]
[0054] To efficiently encode the transformed attribute coefficients, an adaptive lookup table (A-LUT) that can follow N (e.g., 32) most frequent coefficient symbols and a cache that can follow the last M (e.g., 16) different coefficient symbols observed can be used. The A-LUT can be initialized with N symbols provided by the user or computed offline based on the statistics of point clouds of a similar class. The cache can be initialized with M symbols provided by the user or computed offline based on the statistics of point clouds of a similar class. When a symbol S is encoded, binary information indicating whether S is in the A-LUT can be encoded. If S is in the A-LUT, the index of S in the A-LUT can be encoded by using a binary arithmetic encoder. The occurrence count of symbol S in the A-LUT can be incremented by 1. If S is not in the A-LUT, binary information indicating whether S is in the cache can be encoded. If S is in the cache, the binary representation of its index can be encoded by using a binary arithmetic encoder. If S is not in the cache, the binary representation of S can be encoded by using a binary arithmetic encoder. Symbol S can be added to the cache, and the oldest symbol in the cache can be removed.
[0055] The embodiments described herein provide methods and apparatuses for point cloud coefficient encoding. Specifically, the encoding of the transform coefficients from lifting, prediction transforms, and RAHT can be performed by frequency-sorted lookup table indexing, cache indexing, and direct encoding of the symbol values. In practice, these may require multiple lookup tables and caches with many (typically 32 to possibly up to 256) entries to cover a byte of codewords. These lookup tables and caches may additionally need to be updated periodically, the frequency of which may imply different trade-offs in terms of computational requirements and encoding efficiency. Therefore, it may be advantageous to improve the encoding of the transform coefficients of the attributes in G-PCC in terms of complexity / memory and compression efficiency trade-off by alphabet-partitioning and encoding of the alphabet-partitioning information.
[0056] Figure 2 FIG.
[0056] is a block diagram of a communication system 200 according to an embodiment. The communication system 200 may include at least two terminals 210 and 220 interconnected via a network 250. For a one-way transmission of data, a first terminal 210 may encode point cloud data at a local location for transmission via the network 250 to a second terminal 220. The second terminal 220 may receive the encoded point cloud data of the first terminal 210 from the network 250, decode the encoded point cloud data, and display the decoded point cloud data. One-way data transmission may be common in media service applications and the like.
[0057] Figure 2 Also shown is a second pair of terminals 230 and 240 provided to support a two-way transmission of encoded point cloud data that may occur, for example, during a video conference. For a two-way transmission of data, each terminal 230 or 240 may encode point cloud data captured at a local location for transmission via the network 250 to the other terminal. Each terminal 230 or 240 may also receive the encoded point cloud data transmitted by the other terminal, may decode the encoded point cloud data, and may display the decoded point cloud data on a local display device.
[0058] In Figure 2 FIG. Figure 2 , the terminals 210 to 240 may be shown as servers, personal computers, and smart phones, but the principles of the embodiment are not limited thereto. The embodiment is applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network 250 represents any number of networks that transfer encoded point cloud data between the terminals 210 to 240, including, for example, wired and / or wireless communication networks. The communication network 250 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless otherwise stated hereinafter, the architecture and topology of the network 250 may be unimportant for the operation of the embodiment.
[0059] Figure 3 FIG. is a diagram of the arrangement of a G-PCC compressor 303 and a G-PCC decompressor 310 in an environment according to an embodiment. The disclosed subject matter may be equally applicable to other point cloud-enabled applications, including, for example, video conferencing, digital television, storing compressed point cloud data on digital media including CDs, DVDs, memory sticks, and the like.
[0060] The streaming system 300 can include a capture subsystem 313, which can include a point cloud source 301 such as a digital camera. The point cloud source 301 creates uncompressed point cloud data 302, for example. The point cloud data 302 with a high data volume can be processed by a G-PCC compressor 303 coupled to the point cloud source 301. The G-PCC compressor 303 can include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded point cloud data 304 with a low data volume can be stored on the streaming server 305 for future use. One or more streaming clients 306 and 308 can access the streaming server 305 to retrieve copies 307 and 309 of the encoded point cloud data 304. The client 306 can include a G-PCC decompressor 310, which decodes the incoming copy 307 of the encoded point cloud data and creates outgoing point cloud data 311 that can be presented on a display 312 or other rendering device (not depicted). In some streaming systems, the encoded point cloud data 304, 307, and 309 can be encoded according to video coding / compression standards. Examples of these standards include those developed by MPEG for G-PCC.
[0061] Figure 4 is a functional block diagram of the G-PCC compressor 303 according to an embodiment.
[0062] As Figure 4 shown, the G-PCC compressor 303 includes a quantizer 405, a point removal module 410, an octree encoder 415, an attribute transfer module 420, a LoD generator 425, a prediction module 430, a quantizer 435, and an arithmetic encoder 440.
[0063] The quantizer 405 receives the positions of the points in the input point cloud. The positions can be (x, y, z) coordinates. The quantizer 405 also quantizes the received positions using, for example, a scaling algorithm and / or a shifting algorithm.
[0064] The point removal module 410 receives the quantized positions from the quantizer 405 and removes or filters duplicate positions from the received quantized positions.
[0065] The octree encoder 415 receives the filtered positions from the point removal module 410 and encodes the received filtered positions into occupancy symbols of an octree representing the input point cloud using an octree coding algorithm. The bounding box of the input point cloud corresponding to the octree can be any 3D shape, such as a cube.
[0066] The octree encoder 415 also reorders the received filtered positions based on the encoding of the filtered positions.
[0067] The attribute transfer module 420 receives the attributes of the points in the input point cloud. The attribute may include, for example, the color or RGB value and / or reflectivity of each point. The attribute transfer module 420 also receives the reordered positions from the octree encoder 415.
[0068] The attribute transfer module 420 also updates the received attributes based on the received reordered positions. For example, the attribute transfer module 420 may perform one or more of the preprocessing algorithms on the received attributes, and the preprocessing algorithms include: for example, weighting and averaging the received attributes; and interpolating additional attributes from the received attributes. The attribute transfer module 420 also transfers the updated attributes to the prediction module 430.
[0069] The LoD generator 425 receives the reordered positions from the octree encoder 415 and obtains the LoD of each point among the points corresponding to the received reordered positions. Each LoD can be considered as a set of points and can be obtained based on the distance of each point in the points. For example, as Figure 1A shown, the points P0, P5, P4, and P2 may be in the LoD LOD0, the points P0, P5, P4, P2, P1, P6, and P3 may be in the LoD LOD1, and the points P0, P5, P4, P2, P1, P6, P3, P9, P8, and P7 may be in the LoD LOD2.
[0070] The prediction module 430 receives the transferred attributes from the attribute transfer module 420 and receives the LoD of each point among the obtained points from the LoD generator 425. The prediction module 430 respectively obtains the prediction residuals (values) of the received attributes by applying the prediction algorithm to the received attributes in order based on the LoD of each point among the received points. The prediction algorithm may include any one of various prediction algorithms such as interpolation, weighted average calculation, nearest neighbor algorithm, and RDO.
[0071] For example, as Figure 1A shown, before the prediction residuals of the attributes of the points P1, P6, P3, P9, P8, and P7 respectively included in the received LoD LOD1 and LOD2, the prediction residuals of the attributes of the points P0, P5, P4, and P2 included in the received LoD LOD0 may be obtained first. The prediction residual of the attribute of the received point P2 may be obtained by calculating the distance based on the weighted average of the points P0, P5, and P4.
[0072] The quantizer 435 receives the obtained prediction residuals from the prediction module 430 and quantizes the received prediction residuals using, for example, a scaling algorithm and / or a shifting algorithm.
[0073] The arithmetic encoder 440 receives occupancy symbols from the octree encoder 415 and the quantized prediction residuals from the quantizer 435. The arithmetic encoder 440 performs arithmetic coding on the received occupancy symbols and the quantized prediction residuals to obtain a compressed bitstream. The arithmetic coding may include any one of various entropy coding algorithms, such as context-adaptive binary arithmetic coding.
[0074] Figure 5 is a functional block diagram of the G-PCC decompressor 310 according to an embodiment.
[0075] As Figure 5 shown, the G-PCC decompressor 310 includes an arithmetic decoder 505, an octree decoder 510, an inverse quantizer 515, a LoD generator 520, an inverse quantizer 525, and an inverse prediction module 530.
[0076] The arithmetic decoder 505 receives the compressed bitstream from the G-PCC compressor 303 and performs arithmetic decoding on the received compressed bitstream to obtain occupancy symbols and the quantized prediction residuals. The arithmetic decoding may include any one of various entropy decoding algorithms, such as context-adaptive binary arithmetic decoding.
[0077] The octree decoder 510 receives the obtained occupancy symbols from the arithmetic decoder 505 and decodes the received occupancy symbols into quantized positions using an octree decoding algorithm.
[0078] The inverse quantizer 515 receives the quantized positions from the octree decoder 510 and inverse quantizes the received quantized positions using, for example, a scaling algorithm and / or a shifting algorithm to obtain the reconstructed positions of the points in the input point cloud.
[0079] The LoD generator 520 receives the quantized positions from the octree decoder 510 and obtains the LoD of each of the points corresponding to the received quantized positions.
[0080] The inverse quantizer 525 receives the obtained quantized prediction residuals and inverse quantizes the received quantized prediction residuals using, for example, a scaling algorithm and / or a shifting algorithm to obtain the reconstructed prediction residuals.
[0081] The inverse prediction module 530 receives the obtained reconstructed prediction residuals from the inverse quantizer 525 and receives the LoD of each of the obtained points from the LoD generator 520. The inverse prediction module 530 obtains the reconstruction attributes of the received reconstructed prediction residuals respectively by applying a prediction algorithm to the received reconstructed prediction residuals in order based on the LoD of each of the received points. The prediction algorithm may include any one of various prediction algorithms such as interpolation, weighted average calculation, nearest neighbor algorithm, and RDO. The reconstructed attributes belong to the points in the input point cloud.
[0082] A method and apparatus for point cloud coefficient coding will now be described in detail. Such a method and apparatus may be implemented in the above-mentioned G-PCC compressor 303, i.e., the prediction module 430. The method and apparatus may also be implemented in the G-PCC decompressor 310, i.e., the inverse prediction module 530.
[0083] Alphabet - partition of transformation coefficients
[0084] The transformed coefficients or their 8-bit portions may be encoded by using a lookup table (e.g., the above-mentioned A-LUT) or bypass coding with 256 symbols. The 8-bit coefficient values may be decomposed into a set-index and a symbol-index within the set, and the symbol-index may specify the exact position of the coefficient value within the set. For example, the index value may correspond to a position within the lookup table or within the cache. As described in Table 2 below, 256 possible coefficient values may be grouped into N sets.
[0085] Table 2. Example of alphabet-partitioning of coefficient values
[0086]
[0087] In one or more embodiments, offline training may be performed to design the partitioning of coefficient values for a given number of partitions (N). The alphabet-partition boundary values may be signaled explicitly. Alternatively, an index may be signaled to indicate a particular alphabet-partition with associated boundary values in view of multiple alphabet-partition types shared between the encoder and the decoder. It can be understood that the partitioning may be designed such that more frequent symbols belong to sets with lower indices and smaller sizes, and vice versa, to improve coding efficiency.
[0088] In one or more embodiments, a cache or a frequency-sorted LUT may be used to follow the frequencies of coefficient values in descending order. When forming the alphabet-partitioning, lower set-indices may be assigned to more frequent coefficient values by using the indices in the cache or LUT instead of the coefficient values themselves, and vice versa. This process may be performed on-the-fly at both the encoder and the decoder.
[0089] Encoding of alphabet - partition information
[0090] The resulting set-index can be entropy-coded in various ways, and when the symbol distribution within the expected set is fairly uniform, the accompanying symbol-index can simply be bypass-coded.
[0091] In one or more embodiments, the resulting set-index is encoded by multi-symbol arithmetic coding or other types of context-based binary arithmetic coding. Different alphabet-partitions can be used to better utilize the different characteristics of the coefficients.
[0092] In one or more embodiments, different alphabet-partitions can be used for different levels of detail (LOD) layers of the lifting / prediction coefficients, because as a result of the lifting / prediction decomposition, higher LOD layers can have smaller coefficients.
[0093] In one or more embodiments, different alphabet-partitions can be used for different quantization parameters (QP), because higher QP tends to produce smaller quantized coefficients, and smaller quantized coefficients tend to produce higher QP.
[0094] In one or more embodiments, different alphabet-partitions can be used for different layers of the granular scalability of SNR scalable coding, because the enhancement layer (i.e., the layer added to refine the reconstructed signal to a smaller QP level) may be noisier or have a random nature in terms of the correlation between coefficients.
[0095] In one or more embodiments, in the case of SNR scalable coding, different alphabet-partitions can be used according to the value or a function of the value of the reconstructed samples from the corresponding positions in the lower quantization level layer. For example, it is possible that regions with zero or very small reconstructed values in the lower layer may have different coefficient characteristics from regions with the opposite trend.
[0096] In one or more embodiments, different alphabet-partitions can be used according to the value or a function of the value of the reconstructed samples from the corresponding positions in the lower LOD at the same quantization level. These samples from the corresponding positions can be obtained as a result of the nearest neighbor search in the LOD construction in GPCC. It can be understood that these samples can be obtained at the decoder and as a result of the LOD-by-LOD reconstruction in the transform technology in G-PCC.
[0097] Figure 6 is a flowchart showing a method 600 for point cloud coefficient coding according to an embodiment. In some implementations, Figure 6 one or more processing blocks of can be executed by the G-PCC decompressor 310. In some implementations, Figure 6One or more processing blocks may be performed by another device or a set of devices, such as G-PCC compressor 303, that is separate from or includes G-PCC decompressor 310.
[0098] Referring Figure 6 , in the first block 610, method 600 includes decomposing transform coefficients associated with point cloud data into set-index values and sign-index values, where the sign-index values specify the positions of the transform coefficients within the sets.
[0099] In the second block 620, method 600 includes partitioning the decomposed transform coefficients into one or more sets based on the set-index values and the sign-index values.
[0100] In the third block 630, method 600 includes entropy encoding the set-index values of the partitioned transform coefficients.
[0101] In the fourth block 640, method 600 includes bypass encoding the sign-index values of the partitioned transform coefficients.
[0102] In the fifth block 650, method 600 includes compressing the point cloud data based on the entropy-encoded set-index values and the bypass-encoded sign-index values.
[0103] Although Figure 6 illustrates example blocks of method 600, in some implementations, method 600 may include additional blocks, fewer blocks, different blocks, or differently arranged blocks compared to these blocks described in Figure 6 . Additionally or alternatively, two or more of the blocks of method 600 may be performed in parallel.
[0104] Furthermore, the proposed method may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In an example, one or more processors execute a program stored in a non-transitory computer-readable medium, the program being for performing one or more of the proposed methods.
[0105] Figure 7 is a block diagram of an apparatus 700 for point cloud coefficient encoding according to an embodiment.
[0106] Referring Figure 7 , apparatus 700 includes decomposition code 710, partitioning code 720, entropy encoding code 730, and bypass encoding code 740.
[0107] Decomposition code 710 is configured to cause at least one processor to decompose transform coefficients associated with point cloud data into set-index values and sign-index values, where the sign-index values specify the positions of the transform coefficients within the sets.
[0108] The partitioning code 720 is configured to cause at least one processor to partition the decomposed transform coefficients into one or more sets based on a set-index value and a symbol-index value.
[0109] The entropy coding code 730 is configured to cause at least one processor to entropy code the set-index values of the partitioned transform coefficients.
[0110] The bypass coding code 740 is configured to cause at least one processor to bypass code the symbol-index values of the partitioned transform coefficients.
[0111] The compression code 750 is configured to cause at least one processor to compress the point cloud data based on the entropy-coded set-index values and the bypass-coded symbol-index values.
[0112] Figure 8 It is a diagram of a computer system 800 suitable for implementing the embodiments.
[0113] Computer software can be encoded using any suitable machine code or computer language, and any suitable machine code or computer language can be assembled, compiled, linked, etc. through mechanisms to create code including instructions that can be directly executed by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or executed through interpretation, microcode execution, etc.
[0114] The instructions can be executed on various types of computers or their components - including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0115] Figure 8 The components shown for the computer system 800 are exemplary in nature and are not intended to impose any limitations on the scope of use or functionality of the computer software implementing the embodiments. The configuration of the components should also not be construed as having any dependencies or requirements related to any one component or combination of components shown in the embodiments of the computing system 800.
[0116] The computer system 800 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to inputs by one or more human users through, for example, tactile inputs (such as keystrokes, swipes, data glove movements), audio inputs (such as voice, taps), visual inputs (such as gestures), and olfactory inputs (not depicted). The human-machine interface devices may also be used to capture certain media that may not be directly related to conscious human input, such as audio (such as voice, music, ambient sounds), images (such as scanned images, photographic images obtained from a still image camera), and video (such as two-dimensional video, three-dimensional video including stereoscopic video).
[0117] The input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard 801, mouse 802, touchpad 803, touch screen 810, joystick 805, microphone 806, scanner 807, and camera 808.
[0118] The computer system 800 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: tactile output devices (e.g., tactile feedback through the touch screen 810 or joystick 805, but there may also be tactile feedback devices that do not function as input devices), audio output devices (such as speakers 809, headphones (not depicted)), visual output devices (such as screen 810, including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light-emitting diode (OLED) screens, each screen having or not having touch screen input capabilities, each screen having or not having tactile feedback capabilities - some of which screens may be capable of outputting two-dimensional visual output or more than three-dimensional output through means such as stereoscopic graphics output; virtual reality glasses (not depicted), holographic displays, and smoke cans (not depicted), as well as printers (not depicted). The graphics adapter 850 generates images and outputs the images to the touch screen 810.
[0119] The computer system 800 may also include human-accessible storage devices and their associated media, such as optical media including a CD / DVD read only memory (ROM) / read write (RW) drive 820 with media 821 such as a CD / DVD, a thumb drive 822, a removable hard disk drive or solid state drive 823, legacy magnetic media (such as tapes and floppy disks (not depicted)), devices based on a dedicated ROM / application specific integrated circuit (ASIC) / programmable logic device (PLD) (such as a security dongle (not depicted)), etc.
[0120] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass a transmission medium, a carrier wave, or other transient signals.
[0121] The computer system 800 may also include an interface to one or more communication networks 855. The communication network 855 may be, for example, wireless, wired, or optical. The network 855 may also be a local area network, a wide area network, a metropolitan area network, a vehicular and industrial network, a real-time network, a delay-tolerant network, and so on. Examples of the network 855 include: local area networks such as Ethernet, wireless local area network (LAN), cellular networks (including Global System for Mobile Communications (GSM), the third generation (3G), the fourth generation (4G), the fifth generation (5G), Long Term Evolution (LTE), etc.), television cable or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), vehicular and industrial (including Controller Area Network (CAN) bus), and so on. The network 855 generally requires an external network interface adapter attached to some general-purpose data port or peripheral bus 849 (such as, for example, the Universal Serial Bus (USB) port of the computer system 800); other networks are generally integrated into the core of the computer system 800 by attaching to a system bus as described below (for example, a network interface 854 including an Ethernet interface integrated into a PC computer system and / or a cellular network interface integrated into a smart phone computer system). Using any of these networks 855, the computer system 800 can communicate with other entities. Such communication can be unidirectional and only for receiving (such as, for example, broadcast television), unidirectional and only for sending (such as, for example, the CAN bus to certain CAN bus devices), or bidirectional (such as, for example, using a local digital network or a wide area digital network to other computer systems). Certain protocols and protocol stacks can be used on each of these networks 855 and network interfaces 854 as described above.
[0122] The above-mentioned human-machine interface device, human-accessible storage device, and network interface 854 can be attached to the core 840 of the computer system 800.
[0123] The core 840 may include one or more central processing units (CPUs) 841, graphics processing units (GPUs) 842, dedicated programmable processing units in the form of Field Programmable Gate Arrays (FPGAs) 843, hardware accelerators 844 for certain tasks, etc. These devices, together with read-only memory (ROM) 845, random-access memory (RAM) 846, and internal mass storage devices 847 such as internal non-user-accessible hard disk drives, solid-state drives (SSDs), etc., can be connected via a system bus 848. In some computer systems, the system bus 848 can be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices can be attached directly or via a peripheral bus 849 to the system bus 848 of the core. The architecture of the peripheral bus includes Peripheral Component Interconnect (PCI), USB, etc.
[0124] The CPU 841, GPU 842, FPGA 843, and hardware accelerator 844 can execute certain instructions, which, when combined, can constitute the aforementioned computer code. The computer code can be stored in the ROM 845 or RAM 846. Transitional data can also be stored in the RAM 846, while permanent data can be stored in, for example, the internal mass storage device 847. Fast storage and retrieval of any memory device in the memory devices can be achieved by using a cache memory, which can be closely associated with the CPU 841, GPU 842, internal mass storage device 847, ROM 845, RAM 846, etc.
[0125] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be media and computer code that are specifically designed and constructed for the purposes of the implementation, or the medium and the computer code can be of the types that are well-known and available to those skilled in the art of computer software.
[0126] By way of example and not limitation, a computer system 800 having an architecture, and in particular a core 840, can provide functionality due to software implemented in one or more tangible computer-readable media being executed by a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such computer-readable media can be media associated with a user-accessible mass storage device as introduced above, as well as certain storage devices of the core 840 having non-transitoriness, such as an on-core mass storage device 847 or a ROM 845. The software implementing the various embodiments can be stored in such devices and executed by the core 840. Depending on specific needs, the computer-readable media can include one or more storage devices or chips. The software can cause the core 840 and in particular the processors therein (including the CPU, GPU, FPGA, etc.) to perform specific processes or specific portions of specific processes described herein, including defining data structures stored in a RAM 846 and modifying such data structures in accordance with processes defined by the software. Additionally or alternatively, the computer system can provide functionality due to logic implemented hardwired or otherwise in circuitry (e.g., a hardware accelerator 844), which can operate in place of or in conjunction with the software to perform specific processes or specific portions of specific processes described herein. In suitable cases, the software involved can encompass the logic, and conversely, the logic involved can encompass the software. In suitable cases, the computer-readable media involved can encompass circuitry (e.g., an integrated circuit (IC)) storing software for execution, circuitry implementing logic for execution, or both of the above. Embodiments encompass any suitable combination of hardware and software.
[0127] Although the present disclosure has described several embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it will be understood that, although not explicitly shown or described herein, those skilled in the art can envision systems and methods that implement the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.
Claims
1. A method for encoding point cloud coefficients, the method comprising: Decomposing transform coefficients associated with point cloud data into set-index values and sign-index values, the sign-index values specifying the positions of the transform coefficients within the sets; Partitioning the decomposed transform coefficients into one or more sets based on the set-index values and the sign-index values; Entropy encoding the set-index values of the partitioned transform coefficients; Bypass encoding the sign-index values of the partitioned transform coefficients; And Compressing the point cloud data based on the entropy-encoded set-index values and the bypass-encoded sign-index values; Wherein frequency values associated with the transform coefficients are stored in descending order in a cache or a frequency-sorted lookup table, and the lowest set-index value is assigned to the transform coefficient having the largest frequency value.
2. The method according to claim 1, wherein, Signaling the sign-index values and the set-index values to indicate an alphabet-partition with associated boundary values based on one or more alphabet-partition types shared between an encoder and a decoder.
3. The method according to claim 2, wherein, The one or more alphabet-partition types are for one or more levels of detail layers corresponding to the transform coefficients.
4. The method according to claim 2, wherein, The one or more alphabet-partition types are for one or more quantization parameters based on quantization of the transform coefficients.
5. The method according to claim 2, wherein The one or more alphabet-partition types are for one or more layers of scalability of signal-to-noise ratio scalable coding based on correlations between the transform coefficients.
6. The method according to claim 1, wherein, Encoding the set-index values by multi-symbol arithmetic coding.
7. An apparatus for encoding point cloud coefficients, the apparatus comprising: At least one memory configured to store computer program code; And At least one processor configured to access the at least one memory and operate according to the computer program code, the computer program code comprising: Decomposition code configured to cause the at least one processor to decompose transform coefficients associated with point cloud data into set-index values and sign-index values, the sign-index values specifying the positions of the transform coefficients within the sets; Partitioning code configured to cause the at least one processor to partition the decomposed transform coefficients into one or more sets based on the set-index values and the sign-index values; Entropy encoding code configured to cause the at least one processor to entropy encode the set-index values of the partitioned transform coefficients; Bypass encoding code configured to cause the at least one processor to bypass encode the sign-index values of the partitioned transform coefficients; and Compression code configured to cause the at least one processor to compress the point cloud data based on the entropy-encoded set-index values and the bypass-encoded sign-index values; Among them, the frequency values associated with the transform coefficients are stored in descending order in a cache or a frequency-sorted lookup table, where the lowest set-index value is assigned to the transform coefficient with the largest frequency value.
8. The apparatus according to claim 7, wherein, The symbol-index value and the set-index value are signaled to indicate an alphabet partition with associated boundary values based on one or more alphabet-partition types shared between the encoder and the decoder.
9. The device according to claim 8, wherein The one or more alphabet-partition types are used for one or more detail-level layers corresponding to the transform coefficients.
10. The apparatus according to claim 8, wherein, The one or more alphabet-partition types are used for one or more quantization parameters based on the quantization of the transform coefficients.
11. The device according to claim 8, wherein, The one or more alphabet-partition types are used for one or more layers of the scalability of signal-to-noise ratio scalable coding based on the correlation between the transform coefficients.
12. The apparatus according to claim 7, wherein The set-index value is encoded by multi-symbol arithmetic coding.
13. A non-transitory computer-readable medium storing instructions configured to cause at least one processor to perform the method according to any one of claims 1-6, generate a video bitstream and store it.
14. An apparatus for encoding point cloud coefficients, the apparatus comprising: a decomposition module that decomposes transform coefficients associated with point cloud data into a set-index value and a symbol-index value, the symbol-index value specifying the position of the transform coefficient within the set; a partitioning module that partitions the decomposed transform coefficients into one or more sets based on the set-index value and the symbol-index value; an entropy coding module that performs entropy coding on the set-index value of the partitioned transform coefficients; a bypass coding module that performs bypass coding on the symbol-index value of the partitioned transform coefficients; and a compression module that compresses the point cloud data based on the entropy-coded set-index value and the bypass-coded symbol-index value; wherein the frequency values associated with the transform coefficients are stored in descending order in a cache or a frequency-sorted lookup table, where the lowest set-index value is assigned to the transform coefficient with the largest frequency value.
15. A method for storing a video stream, characterized in that, The video bitstream can be generated according to the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Adaptive coding and decoding of wide-range coefficients
WO2007021616A2